Abstract
<title>Abstract</title> <p> <bold>Background</bold> Single-cell studies routinely draw donor-level conclusions — a signal that tracks age, sex or disease — from representations that mix two different things: which cell types a donor has, and what those cells express. Analyses commonly “control for cell-type composition”, but there is no reference scale for how much signal that removes, and no convention for reporting the resolution at which the control was applied. <bold>Methods</bold> We audited 5.85 million cells from 2,148 analysed donors across five public human blood datasets and six donor-level traits. For each dataset and trait we scored three donorlevel representations with the same estimator on the same donors: cell-type proportions alone (comp); pseudobulk built from within-cell-type profiles with cohort-fixed mixing weights, so that every donor has the same composition (state); and ordinary pseudobulk, which uses the donor’s own proportions (pb). state and pb share the same within-cell-type profiles and differ only in the weights, so the gap between them is composition and nothing else. We report composition sufficiency CS = Scomp/Spb and state retention SR = Sstate/Spb with donor-bootstrap intervals and permutation nulls. The statistic is calibrated on simulated traits: it is zero when a trait is orthogonal to composition and rises to 2.0 when a trait is purely compositional, so it is a monotone index rather than a fraction — a pseudobulk profile encodes composition only implicitly and a linear model cannot fully recover it. <bold>Results</bold> Traits differ widely and interpretably. Chromosomal sex is close to a pure state signal: the pseudobulk model reaches skill up to 1.000 while proportions reach at most 0.590, giving CS = 0.28–0.70. Chronological age sits at the compositional end, with median CS = 1.02, and cytomegalovirus serostatus with it — though a parity analysis shows that much of the composition channel’s strength is precision of measurement, since proportions are estimated from every cell while profiles come from a subsample. The second result is unaffected by that asymmetry and is the one we did not expect: holding composition fixed leaves the expression model essentially untouched. Across all 17 dataset–trait pairs with usable signal, median state retention is 1.00 (range 0.89–1.06). Replacing every donor’s own cell-type proportions with the cohort average costs a pseudobulk model nothing, for any trait we tested. Finally, the attribution is not a fixed property of the data. Refining the annotation each dataset ships from its coarsest to its finest level raises composition skill for age substantially in every dataset that offers the contrast (AIDA v1 0.001 at 9 classes to 0.499 at 34, AIDA v2 0.029 at 9 classes to 0.500 at 88, Allen atlas 0.243 at 9 classes to 0.675 at 71), turning a near-useless predictor into one that matches the full expression profile. An unsupervised sweep from 5 to 160 clusters reproduces the trend, and the one dataset where the effect is small is the one restricted to a single immune compartment. <bold>Conclusion</bold> Composition-matching an analysis built on pooled or averaged expression removes far less than the phrase suggests, so passing such a control is weak evidence that a finding concerns cell state. And because cell-type labels are derived from the expression data rather than measured independently, how much counts as “composition” depends on how finely those labels are drawn: a composition control is under-specified unless its resolution is stated. We recommend reporting proportions-only skill next to full-model skill at a named annotation resolution. </p>