Abstract
<jats:p>There is a growing demand for tailormade molecules addressing specific applications, such as green solvents, pharmaceutical ingredients and other functional materials. High-quality chemical encodings are essential for effective molecular optimization to fulfill this demand. Current molecular encodings, predominantly based on intramolecular properties, inadequately represent crucial intermolecular interactions critical for predicting behavior in complex environments. Recent advances in atomistic foundation models, trained on large-scale chemical interaction datasets, offer potentially strong inductive bias for intermolecular phenomena. We propose Foundation Feature-enhanced Molecular Representation (FFEMR), a general framework for augmenting conventional molecular encodings with implicit representations derived from chemical foundation models within a Bayesian optimization loop. Within this framework, we systematically evaluate multiple fusion strategies for integrating foundation features with traditional descriptors, and characterize their trade-offs across data regimes and model complexity. We further introduce a lightweight adaptation strategy for targets that are not directly aligned with the foundation model’s training objective. Our experiments on both aligned and misaligned optimization tasks show that all FFEMR variants significantly outperform traditional encodings alone. Notably, even the simplest concatenation strategy delivers substantial gains, while more sophisticated fusion methods can further improve performance given sufficient data. We are deploying FFEMR-enhanced optimization on a high-throughput experimentation (HTE) platform to accelerate the discovery and development of sustainable chemicals.</jats:p>