Back to Search View Original Cite This Article

Abstract

<jats:p>Machine learning models for protein properties are usually reported by a single accuracy figure, which says how a model behaves on average but not whether to act on any one prediction, especially for a protein unlike anything in the training set. That gap is both a black box problem and an out-of-distribution problem, and it is worst exactly where discovery work happens, on sequences the model has not seen. We present ProtTrust-XAI, a framework that scores each prediction by ensemble consensus and by the structural coherence of its own attribution, and separately tracks a third signal, distance from the training distribution, to catch cases the first two cannot see. We demonstrate it on per-protein thermostability, training a relational graph convolutional network on melting temperatures for over 20,000 proteins using AlphaFold-derived contact graphs and frozen protein language model embeddings. On family-level held-out proteins the model reaches a Spearman correlation of 0.65 and a mean absolute error of 4.1°C, and predictions the framework labels most trustworthy fall to 3.0°C, below the assay's own reproducibility floor, so a practitioner can act on the label with the same confidence as on the measurement itself. Applying the framework across the full dataset also exposes two representational blind spots, one around cofactor chemistry and one around membrane proteins, each with a distinct mechanistic explanation that points to a specific fix. Transferred to plastic-degrading enzymes at low sequence identity to the training data, absolute predictions collapse while the ranking survives, and a controlled ablation shows this is a general property of distribution shift rather than something particular to that external set. The same transfer identifies where the distance-based signal itself needs recalibrating before deployment, which is a diagnosis the framework produces about itself and not a hidden failure. The result is a practical rule. Inside a model's competence domain, trust its labels. Outside it, trust its ranking. A model that reports its own limits, rather than only its average accuracy, is one an experimentalist can actually build on.</jats:p>

Show More

Keywords

model training framework protein proteins

Related Articles

PORE

About

Connect