Abstract
<title>Abstract</title> <p>Transformers for single-cell RNA sequencing must convert a sparse, unordered count vector into tokens. Two common choices are to order genes by within-cell expression after corpus-level scaling, as in Geneformer, or to assign each gene a discrete expression value, as in scBERT. Rank encodings are often described as robust to differences in sequencing depth, but reduced depth is a stochastic loss of molecules rather than a uniform rescaling. We tested whether a corpus-median rank encoding preserves a full-depth cell partition better than fixed value bins when held-out transcriptomes are thinned. Raw counts from 2,638 peripheral blood mononuclear cells were aligned to eight Scanpy tutorial labels and restricted to 18,974 genes shared with the released Geneformer-104M vocabulary. Class-balanced linear probes were trained in five folds at full depth and evaluated after binomial thinning to 75%, 50%, 25%, 10%, and 5% of observed UMIs, with 30 paired realizations per depth. At 25% retention, value-bin macro-F1 was 0.857 while corpus-median rank macro-F1 was 0.776. Relative to each encoding’s full-depth score, value bins retained 5.9 macro-F1 points more than corpus-median ranks (95% bootstrap interval, 5.6–6.3 points; Holm-adjusted p < 10−8 ). An unscaled rank control improved macro-F1 to 0.799 but remained below value bins. Mechanistically, corpus-median scaling promoted low-count genes: half of its top 128 tokens were supported by a single UMI, compared with less than 1% under plain count ranks. These results do not compare pretrained transformers. They show, in one controlled dataset, that invariance to library-size multiplication does not imply robustness to molecule loss, and that corpus normalization can further alter token stability.</p>