Back to Search View Original Cite This Article

Abstract

<jats:p>In many molecular property prediction tasks, such as drug discovery and OLED material design, labeled data are often limited to only tens to thousands of samples due to high experimental costs. Despite the rapid development of machine learning models, it remains unclear which model classes are most effective under such low-data regimes and how their performance evolves with increasing data. In this work, we present a comprehensive learning curve benchmark of molecular property prediction across a broad model suite spanning traditional machine learning, graph neural networks (GNNs), transformer-based molecular language models, SE(3)-equivariant 3D networks, descriptor-based hybrids, multi-task learning variants, and multimodal graph–sequence fusion baselines. We systematically vary the training data size from 50 to 3,000 samples and evaluate on the QM9, ESOL, Lipophilicity, and BACE datasets using scaffold splits with repeated trials for statistical robustness. Our analysis reveals four principal patterns: (1) regime-dependent model optimality — deep learning models dominate continuous physicochemical property prediction (ESOL, Lipophilicity), while kernel-based methods such as Gaussian Process Regression remain superior for bioactivity prediction (BACE) across most of the tested range, with clear crossover points emerging between traditional and deep learning approaches; (2) target-dependent benefit of 3D inductive bias — on QM9 an equivariant 3D network trained on DFT-optimized geometries performs strongly on LUMO and the HOMO–LUMO gap but struggles on HOMO despite these exact geometries, isolating a representational mismatch rather than a geometry-quality limitation; on Lipophilicity and BACE, single-conformer ETKDG inputs yield essentially flat learning curves, where weak conformer geometry and the conformer-insensitivity of these endpoints cannot be fully disentangled — though both readings point to a mismatch between single-geometry 3D models and conformer-insensitive targets; (3) non-monotonic pretraining benefit — the advantage of modern transformer and 3D foundation models (ChemBERTa-2, MoLFormer-XL, SELFormer, Uni-Mol-PT) over GNNs and fingerprints is most pronounced at intermediate training sizes on most tasks (and at the smallest on bioactivity), and neither parameter count nor pretraining-corpus size predicts downstream accuracy (the compact ChemBERTa-2 outperforms the far larger MoLFormer and SELFormer on ESOL and Lipophilicity); and (4) model selection over model averaging — in a controlled post-hoc analysis, even oracle ensembles that select their members by test-set error barely improve on the best single model in the lowest-data regime (at most +1.2% at N=50, and negative on ESOL), indicating that the dominant lever under data scarcity is choosing the right single model rather than combining several. Our findings provide actionable guidelines for model selection under data scarcity and highlight learning curve analysis as a principled tool for understanding model behavior across the property–scale plane.</jats:p> <jats:p> <jats:bold>Scientific Contribution:</jats:bold> This study establishes a comprehensive learning curve benchmark for molecular property prediction in low-data regimes, systematically comparing a broad suite of model families—including traditional machine learning, graph neural networks, transformer-based molecular language models, SE(3)-equivariant 3D networks, descriptor-based hybrids, multi-task learning variants, and multimodal graph–sequence fusion baselines—across four datasets spanning physicochemical, quantum mechanical, and biological activity properties. We identify regime- and target-dependent model optimality and crossover points between traditional machine learning and deep learning methods, quantify the target-dependent benefit and conformer-fidelity limits of 3D equivariant networks, and provide actionable, data-scale-aware guidelines for model selection in low-data molecular property prediction. </jats:p>

Show More

Keywords

learning model molecular prediction models

Related Articles

PORE

About

Connect