Abstract
<jats:p>Evaluating protein sequence similarity remains challenging in the protein-sequence twilightzone (20-35% sequence identity), where traditional methods often fail. In this study, we evaluate whether mean-pooled embeddings from four protein language models: ESM-1b, ESM-2, ProtT5, and ProstT5 can estimate pairwise structural similarity without performing sequence alignment. The benchmark dataset includes 20,445 PISCES protein pairs with sequence identity ≤ 30%, representing the protein-sequence twilight zone, with TM-align-derived TMmin used as the structural ground truth. Protein embeddings are compared using cosine similarity, Euclidean- and Manhattan-derived similarities, an RBF kernel, and dot product. Among these similarity metrics, cosine similarity performs best across all four models. Moreover, ProstT5 achieves the highest Spearman correlation with TMmin, followed by ESM-2, ProtT5, and ESM1b, while all four PLMs outperform BLASTP overall. Furthermore, the advantage of PLM embeddings is most pronounced for protein pairs with the lowest sequence identity. ProstT5 also provides the best discrimination between structurally similar and dissimilar protein pairs. Moreover, it offers a favorable balance between similarity performance and the computational requirements of residue-level embedding generation and storage. Overall, these findings support PLM embeddings as an effective alignment-free approach for detecting structural relationships among proteins in the twilight zone.</jats:p>