Back to Search View Original Cite This Article

Abstract

<title>Abstract</title> <p> <bold>Motivation.</bold> Sequence-based predictors of recombinant protein solubility are widely used to triage expression targets, and deep-learning models report strong benchmark AUROCs. Whether a model's headline benchmark predicts how it ranks on an independent distribution, and how much two state-of-the-art models can differ on the same proteins, is rarely tested directly. Here we show, across two <italic>Escherichia coli</italic> distributions with per-model leakage control, that a model's published benchmark AUROC is a poor guide to its rank on a new distribution, and that a simple interpretable baseline can match or outperform state-of-the-art deep models outside their training distribution. <bold>Results.</bold> We benchmarked two state-of-the-art deep models, RP3Net (ESM-2 650M, reported AUROC 0.83) and NetSolP (ESM1b, reported AUROC 0.76), against two dataset-naive composition baselines (an in-house heuristic and the published Solubility-Weighted Index, SWI) on two <italic>E. coli</italic> benchmarks, eSOL (native K-12 proteins) and the SoluProt held-out test set (heterologous constructs), with per-model MMseqs2 leakage control. On native cytoplasmic proteins the two deep models spanned the whole range and straddled the baselines: NetSolP was the best predictor (0.792) and RP3Net the worst (0.709), a gap of 0.083 (DeLong p &lt; 1e-8) larger than any model-to-baseline difference. Despite its 0.83 headline, RP3Net was beaten by the simple SWI score (0.745, p = 0.006) and only tied the naive heuristic (0.725), while NetSolP beat SWI (p &lt; 0.001). On the harder heterologous SoluProt set the two deep models instead agreed and led numerically (NetSolP 0.633, RP3Net 0.619, difference not significant), all predictors in a narrow 0.58–0.63 band; ProteinSol, calibrated on native eSOL, collapsed to near-chance there (0.542). A third deep model, PLM_Sol (ProtT5), sharpened the contrast further: our pipeline reproduced its ≈ 0.83 own-benchmark AUROC, yet it fell to 0.564 on SoluProt, below every baseline, a value unchanged by removing the 35.7% of the test set that overlapped its training data (identical 0.562 on the overlapping proteins), so a genuine distribution effect rather than memorisation. The native ordering (NetSolP &gt; SWI &gt; RP3Net) survived a homology-level screen (MMseqs2 clustering at ≤30% identity; 8.5% of eSOL removed), confirming it is not a residual-training-homology artefact. <bold>Conclusion.</bold> Solubility-predictor performance is model- and distribution-dependent: two SOTA deep models differed on native proteins by more than the gap to a simple interpretable baseline, and a model's published benchmark AUROC did not predict its generalisation rank. Evaluating RP3Net on native proteins probes generalisation outside its recombinant/heterologous design domain, so its low native ranking reflects distribution mismatch rather than a model deficiency; we therefore frame these results as a caution about distribution-matched evaluation, not a criticism of any model. A dataset-naive baseline is a useful sanity floor: a model falling below it on the target distribution (as RP3Net does on native proteins) is an informative signal, and predictors should be evaluated on the distribution of interest rather than trusted by headline number. </p>

Show More

Keywords

models native distribution proteins rp3net

Related Articles

PORE

About

Connect