Abstract
<title>Abstract</title> <p>Motivation: Spatial transcriptomics (ST) benchmarks are routinely reported to be “inflated” by randomly interleaved cross-validation, with the inflation attributed to spatial leakage. That attribution has never been tested. The random-versus-block performance gap simultaneously changes the training-set size, the set of classes available in training and test, the amount of message passing that crosses the train/test boundary, and the scope over which preprocessing is fit. A single gap statistic cannot separate them. Results: We build a graph in which train, validation and test induced subgraphs are mutually disconnected—zero cross-role edges and infinite minimum train–test hop distance in every fold, verified by two independent implementations—and then walk a ladder of eleven protocols that change one design axis at a time on 10x Visium human breast cancer (3,798 spots, 11 Leiden domains; 6 methods × 5 folds × 5 seeds). Folds 1 and 3 are the same spatial partition, so all inference uses n = 4 independent spatial blocks with hierarchical bootstrap intervals. The closed- set random-to-block gap of +0.447 macro-F1 [95% CI +0.332, +0.562] decomposes additively into class-support mismatch +0.276 [+0.121, +0.429], residual spatial extrapolation +0.098 [+0.065, +0.135], preprocessing scope +0.051 [+0.011, +0.096] and training-set size +0.022 [+0.014, +0.029]. Class-support mismatch, not leakage, is the largest single component. After matching size, class support, graph masking rule and preprocessing scope, the genuine spatial- extrapolation residual is +0.214 [+0.160, +0.273] for graph neural networks but +0.042 [−0.039, +0.143] for non-graph classifiers, a family gap of +0.172 [+0.054, +0.274]. Cross-role message passing is worth +0.050 (GCN) and +0.016 (GAT) macro-F1 and is exactly 0.000 for every non-graph method by construction. Recomputing persistent-homology features within each fold removes the previously reported +0.0154 topology margin entirely: the leakage channel alone accounts for +0.0141 of it, and the clean estimate is −0.0144 (SD 0.0269, exact sign-permutation p = 0.50 against a floor of 0.125), i.e. no difference detected in either direction. On 12 expert- annotated DLPFC sections (47,280 spots, 3 donors) the mixed-model protocol effect is −0.5488 (SE 0.0089) and the method ranking reverses: XGBoost-PCA+XY beats SVM-PCA in 12/12 sections under random CV and loses under block CV (2/12), leave-one-section-out (1/24) and leave-one-donor-out (1/6). Two published methods behave the same way (STAGATE +100.7%, GraphST +62.4%). Conclusion: Protocol-driven optimism is not one quantity. In the deeply audited breast-cancer case study the largest component is a class-support artifact of spatial blocking; the residual spatial-extrapolation penalty that survives full matching is substantially larger for graph neural networks and is statistically resolvable only for them, while for non-graph classifiers it is small and unresolved. Benchmarks should report the decomposition, not a single inflation figure, and should match class support before attributing a gap to space. Availability: Code, fixed split definitions, pruned graphs, machine-readable result tables and scripts for reproducing all analyses and figures are available at https://github.com/ChimdiWalter/Topospatial,ChimdiWalter/Topospatial. The exact release corresponding to this manuscript is deposited in the Zenodo archive.</p>