Abstract
<jats:p>Virtual chemical libraries now exceed billions of compounds, placing joint demands on the scoring tools used to prioritize candidates for accuracy and scalability. Supervised affinity models meet scalability demands but remain vulnerable to dataset memorization and train-test leakage. Here we introduce AIMNet2(Score), a physics-informed scoring framework built on AIMNet2 machine-learned interatomic potentials (MLIPs) that uses no experimental binding-affinity labels. The score combines three energetic terms evaluated on a single bound structure: a protein–ligand interaction energy from AIMNet2(2025), a ligand desolvation penalty, and a local conformational strain penalty, the latter two computed with AIMNet2-CPCM. On KIN66, a new kinase benchmark of ~1,000-atom active sites with interaction energies at three DFT levels, AIMNet2(2025) reproduces its B97-3c training reference to 3.13 kcal mol-1 RMSE and transfers to systems larger than those used for training. On eight FEP+ congeneric series and a curated co-crystal set (HiQBind), the label-free score ranks affinities on par with or ahead of docking, semiempirical, and deep-learning baselines. On raw docked poses, all methods including ours perform comparably, isolating pose fidelity rather than scoring physics as the accuracy ceiling for single-structure scoring. AIMNet2(Score) thus provides a label-free, QM-informed route to candidate ranking and defines where such scoring succeeds and fails.</jats:p>