Abstract
<title>Abstract</title> <p>Single-cell foundation models are often evaluated on broad benchmark suites, but the empirical question that matters for everyday use is narrower: when should a researcher expect a pretrained representation to beat a classical pipeline such as highly variable gene selection, PCA, and a simple readout? We study this question with a controlled, literature-calibrated simulation rather than a new large checkpoint benchmark. The simulation isolates two regimes where the current literature makes conflicting predictions: low-label cell type annotation, where foundation models may help by providing a reusable biological prior, and perturbation-response prediction, where recent benchmarks report that simple linear and PCA baselines remain very hard to beat. In 10 replicated annotation worlds and 28 replicated perturbation worlds, stylized foundation-model representations improved macro-F1 most strongly when only one to five labeled cells per class were available. At five labels per class, scGPT-like and scFoundation-like embeddings exceeded the HVG-PCA baseline by 0.11–0.19 macro-F1 across donor-holdout and rare-subtype regimes. With 50 labels per class, however, the advantage largely collapsed: HVG-PCA matched or nearly matched the best foundation-style representation in the matched-donor setting. In perturbation prediction, the PCA-ridge baseline tied the best foundation-style method on seen regulatory neighborhoods and outperformed it on novel neighborhoods (Pearson r = 0.700 versus 0.673 for the scFoundation-like readout). These results support a conservative interpretation of the field: pretraining is most useful as a label-efficiency prior for annotation, while perturbation generalization still requires stronger evidence than current foundation-model benchmarks usually provide.</p>