Deprecated: Function curl_close() is deprecated since 8.5, as it has no effect since PHP 8.0 in /home/u483256323/domains/poorvam.com/public_html/subdomains/pore/includes/api.php on line 184
Abstract
<jats:title>Abstract</jats:title> <jats:p>Accurate benchmarking of intermolecular interaction energies is central to evaluating quantum chemical methods and guiding the development of reliable machine-learned interatomic potentials (MLIPs). We benchmark five MLIPs (AIMNet2(2023), AIMNet2(2025), MACE-OFF23(M), MACE-OMol, and UMA-S-OMol) across twenty-one datasets spanning hydrogen-bonded, dispersion- and pi-dominated, sigma-hole, ionic and charge transfer, and repulsive nonequilibrium interactions, with reference values at or near CCSD(T)/CBS accuracy. AIMNet2(2025) is a continually pretrained variant of AIMNet2(2023) that retains the original architecture but incorporates 3.8 million additional structures curated to improve noncovalent interactions (NCIs). AIMNet2(2025) improves on its predecessor across nearly all benchmark categories, with the largest gains in the hydrogen-bonded, sigma-hole, and repulsive regimes, while remaining competitive with the much larger MACE-OMol and UMA-S-OMol. The supramolecular S12L and L7 benchmarks show only marginal improvement: every evaluated MLIP exhibits large errors driven by a small number of pathological complexes. Two factors beyond intrinsic model quality significantly influence the reported performance. First, partial overlap between training and benchmark data, quantified here via systematic overlap detection, inflates apparent accuracy for all models, most strongly for those trained on OMol25. Second, differences in the DFT reference level used for MLIP training establish irreducible error floors, so superior benchmark performance may partly reflect closer proximity of the training functional to the CCSD(T)/CBS reference rather than stronger modeling capability. Sigma-hole interactions emerge as the category with the lowest training-benchmark overlap across all models and therefore provide the most discriminating test of true generalization. Meaningful MLIP evaluation must account for data provenance, reference theory consistency, and the distinction between interpolation and out-of-distribution generalization, particularly as standard NCI benchmark sets become&#xD; absorbed into large-scale training datasets.&#xD;</jats:p>