Abstract
<title>Abstract</title> <p>In language models, the leading directions of the embedding spectrum encode how often a token appears rather than what it means. Single-cell foundation models (scFMs) tokenise non-zero expression counts, so they have an exact analogue of token frequency: how abundantly a gene is expressed and in how many cells it is detected. This paper asks two questions of the gene-embedding tables of eight released models — seven single-cell transformers and the frozen ESM2 protein language model that Universal Cell Embeddings uses as input — over a common universe of 17,846 genes. (A) A confound audit. Regressing every spectral axis on a spline basis in abundance and detection breadth shows that the leading axis of six of the seven count-trained models is dominated by abundance (R2 = 0.51–0.76 against a permutation null of 7 × 10−4 ), while the protein language model that never saw an expression value sits at R2 = 0.009. The loading is concentrated: beyond axis 64 it falls below 0.01. We then show that deleting the leading axes improves curated-relation retrieval — the “all-but-thetop” effect, replicated in gene embeddings — and that band-limited retrieval peaks in the spectral band [32, 64), not at the top. Under abundance-matched negatives, roughly a third of the apparent complex- and pathway-retrieval margin over chance disappears, a co-expression control is almost entirely abundance, and regulator–target retrieval is unaffected. A supervised 29-way compartment probe reaches 4.4× chance while an abundance-only classifier reaches exactly chance, so compartment structure is real — but no single axis carries it. (B) A compaction law. For each relation type we measure the bandwidth k∗ , the smallest number of leading axes recovering 95% of a model’s own ceiling. The hypothesis that k∗ is a property of the biology is rejected: a two-way analysis of variance attributes η^2 = 0.44 to the relation and η^2 = 0.34 to the model under the standard protocol, and under abundance matching the model term dominates (0.62 versus a non-significant 0.17). The model effect is not explained by embedding width, and a cluster bootstrap shows the across-model spread exceeds sampling noise for pathway and complex membership. What transfers better is the ordering: the ranking of relations by bandwidth agrees across architectures (Kendall’s W = 0.64, p = 0.004). Compaction curves follow a stretched exponential with exponent at or below one, and a rank-64 truncation retains 0.85–0.95 of the ceiling on average, with the worst individual cases near 0.74. Both confound controls raise k∗ , so bandwidths measured without them under-provision rank. The practical recommendations: report the abundance R2 of any axis claimed to encode a biological property, evaluate with abundance-matched negatives, and choose a truncation rank from the relation the surrogate must carry rather than from a single tuned number.</p>