Abstract
<jats:p>Accurate prediction of absorption, distribution, metabolism, excretion, and toxicity (ADMET) depends strongly on how a molecule is encoded, and no single representation dominates across endpoints. We ask whether heterogeneous, already-available encoders can be integrated effectively without training a new foundation model or performing additional cross-modal pre-training. Operating within this constraint, four encoders with complementary priors (descriptor-, SMILES-, 2D-, and 3Dinformed) are reused through parameter-efficient fine-tuning and combined at two controlled levels: feature-level joint fusion and prediction-level late fusion. These four encoders plus a joint-fusion model, each under single- and multi-task learning, yielded 10 base models. To avoid test-driven selection, all 1,023 ensemble subsets were scored exclusively on held-out validation predictions, with ensemble size and membership fixed before any test-set analysis. The selected five-member uniformmean ensemble, which we name Tetra-Fuse—D-MPNN·MTL, MoLFormer·MTL, CheMeleon·STL, GraphMVP·STL, and a Modality Attention Pooling joint-fusion model (MAP·STL)—uses no metalearner. Across the 21 evaluable TDC ADMET tasks, it reached a mean rank of 3.95 against an archived leaderboard snapshot in the primary matched-seed track—significantly outperforming three deployable single-model baselines (paired Wilcoxon, Holm-corrected p < 0.001)—and improved further to 3.10 when seed-averaged for practical deployment. Despite highly correlated member predictions (matched-seed mean |r| = 0.767), Leave-One-Out and Shapley analyses associate the gain primarily with task-wise specialization rather than error decorrelation. On an independent Polaris ADME-Fang benchmark, the identical configuration ranked first on two of six leaderboards and remained highly competitive with billion-parameter graph models despite using far lighter backbones—demonstrating that late fusion transfers robustly as a stable, reusable recipe.</jats:p>