Back to Search View Original Cite This Article

Abstract

<jats:p>Structure-based scoring functions leveraging machine learning have recently demonstrated superior performance over classical scoring functions, particularly on virtual screening benchmarks. However, due to the fundamental differences between their underlying model principles and architectures, it remains unclear to what extent performance stems from an understanding of molecular binding or from exploitation of systemic biases. Thus, disentangling the factors underlying benchmark performance is essential for determining whether a scoring function will generalize to novel chemical space and succeed in prospective drug discovery. To address this need, we present a case study investigating the nature and impact of systemic biases on benchmark comparisons between different scoring function paradigms. By systematically analyzing the evaluation workflows of prominent models, we reveal pocket bias, a form of spatial coordinate frame leakage arising from static binding pocket extraction, which artificially inflates benchmark performance. To progressively eliminate these sources of bias, we benchmarked two selected graph neural network scoring functions against two minimalist machine learning models and a classical scoring function under four increasingly stringent evaluation levels, successively removing pocket bias, reducing structural data leakage, and finally evaluating on out-of-distribution (OOD) protein targets. Upon removal of pocket bias and structural data leakage, the performance of all machine learning models dropped substantially. When evaluated on out-of-distribution protein families, the classical baseline AutoDock Vina outperformed the machine learning models in five of seven virtual screening tasks and dominated the docking power evaluation. Our findings indicate that benchmark performance can be heavily shaped by evaluation design and dataset artifacts, potentially overshadowing algorithmic improvements. While the tested machine learning models remain heavily dependent on encountering familiar data distributions to achieve competitive results, AutoDock Vina demonstrated superior generalization capacity on OOD targets. This work underscores the critical need for rigorous, artifact-free benchmarking protocols to guide the development of truly prospective machine learning models for virtual screening.</jats:p>

Show More

Keywords

scoring machine learning performance models

Related Articles

PORE

About

Connect