Abstract
<title>Abstract</title> <p>QSAR models for Ames mutagenicity are increasingly used in regulatory risk assessment under EU REACH, K-REACH, and ICH M7, but their reported performance varies widely across evaluation protocols. 13 In this study, 225 model congurations seven algorithms, ve feature sets, and three preprocessing scenarioswere systematically evaluated across three genotoxicity datasets: Ames (n = 13,560), in vitro chromosome aberration (n = 1,619), and in vivo micronucleus (n = 2,153). Using a xed model conguration (XGBoost, compact features, raw SMILES), Ames MCC dropped from 0.670 under random-split CV to 0.624 under scaold-split CV. Cross-source leave-domain-out (LDO) evaluation produced markedly larger declines: 0.234 (drug→industrial) and 0.122 (industrial→drug). The source heterogeneity gap (∆ = 0.390) considerably exceeded the split-method gap (∆ = 0.046). A simple source-based classier reached MCC = 0.608 with no structural information, reecting a 19.5-fold dierence in positive rates between drug-sourced (58.3%) and industrial-sourced (3.0%) compounds. Threshold decomposition showed that prior shift accounted for only part of this gap, with oracle MCC plateauing at 0.270.36. Algorithm robustness 1 under source shift varied considerably: Random Forest maintained MCC = 0.325 in the hardest direction, whereas XGBoost collapsed to 0.122. In vivo micronucleus models showed strong ranking performance (ROC-AUC = 0.860, PR-AUC = 0.237) but unreliable classication (MCC 95% CI crossed zero). These results call for a source-aware evaluation cascade as a standard reporting requirement for mutagenicity QSAR.</p>