Abstract
<jats:p>Feature attributions are routinely used to justify molecular graph neural network (GNN) predictions to chemists, yet they are almost never audited for reliability: existing frameworks ask whether an explanation is faithful to the model, not whether it identifies the chemistry that determines the property, nor where it stops being trustworthy under scaffold shift. MolSanity is a reliability-audit framework that wraps canonical implementations (Captum, PyTorch Geometric, RDKit) rather than proposing a new attributor, scoring every (dataset × backbone × attributor × split) cell on six axes: motif-native coherence, occlusion–attribution faithfulness, ground-truth localisation where node labels exist, cross-checkpoint stability, calibration linkage, and confidence/correctness regime stratification. Our central finding is that faithfulness is not correctness, and that the two carry no dependable relationship in either regime. Across 30 selection tests — 5 molecular ground-truth arms × two splits × 3 ranking metrics — a faithfulness-only ranking picks an attributor other than the ground-truth-best one in 26, and this is no less true in distribution (14 of 15) than under scaffold shift (12 of 15); on MUTAG under shift it prefers an attributor anti-aligned with the nitro motif (GT AUROC 0.013) over one at 0.826. Pooled over 47 cells the faithfulness–correctness rank correlation is +0.222 in distribution (p = 0.134) and −0.124 under shift (p = 0.405), neither distinguishable from zero, while per arm under shift it runs from −0.564 to +0.786: reliability is a property of the individual (dataset, backbone, attributor, split) cell rather than of the attributor. Restricted to the 3 arms of an earlier analysis it gives −0.356 (p = 0.042); two externally authored rationale benchmarks remove the effect, and we report the 5-arm result. Faithfulness itself does not fall under shift — it rises, from 0.049 to 0.132, while ground-truth localisation does not move — so the metric gives no warning either way. Over 31,785 per-molecule records, localisation degrades on confidently-wrong predictions (0.769 to 0.681) while their measured faithfulness improves, and the calibration–reliability link attenuates from a per-cell median of 0.144 to 0.074 when cells are pooled. Every number, figure and table regenerates from the committed artifacts.</jats:p>