Back to Search View Original Cite This Article

Abstract

<title>Abstract</title> <p>Biomedical benchmark leaderboards often combine several desirable properties into one score. The weights in that score express a scientific preference, but their effect on the reported ranking is rarely measured. We reanalyzed the public scIB single-cell data-integration benchmark to ask a narrow question: does the winning method change when the weight assigned to batch correction, relative to biological conservation, changes? The analysis covered 327 task-pipeline evaluations from five real atlas tasks, 17 method entries including an unintegrated baseline, and 14 metrics. We reproduced the benchmark’s within-task scaling and swept the batch-correction weight over 1,001 values from 0 to 1. At the published weight of 0.4, scANVI ranked first with a mean task rank of 3.2; its task ranks ranged from 1 to 6. The winner was unchanged in a centered sensitivity band from 0.3 to 0.5, although as many as 6.0% of strictly ordered method pairs reversed. Across the full sweep, scANVI won through weight 0.641, Scanorama won from 0.642 to 0.644, and Harmony won from 0.645 onward. At the most discordant weight, Kendall’s τb with the reference ranking was 0.494 and 25.2% of comparable method pairs were reversed. Resampling the five atlas tasks showed a separate source of uncertainty: at the reference weight, scANVI was the bootstrap winner in only 56.0% of 10,000 resamples. The result is therefore qualified. The published weight is locally stable in this reanalysis, but neither the winner nor the ordering is invariant to defensible changes in scientific priorities or task composition. Benchmark reports should include weight-sensitivity curves and task-resampling uncertainty rather than a single rank table.</p>

Show More

Keywords

weight from method benchmark scanvi

Related Articles

PORE

About

Connect