Abstract
<jats:p>Accurate predictions of transcriptomic responses to genetic perturbations could unlock our understanding of gene functions and regulatory networks. While a growing number of methods and benchmarks target this task, existing evaluations focus on mean expression accuracy alone. This overlooks differential expression (DE), which captures both mean and variance and forms the basis for biological interpretation and experimental follow-up. Here, we systematically evaluate a diverse set of deep learning and non-deep-learning methods for their ability to predict DE outcomes under two generalization regimes: unseen perturbations within the same cell line, and unseen cellular contexts across cell lines. We find that simple baselines, such as embedding-based nearest neighbors, are competitive and often outperform specialized deep learning models for DE classification across datasets and evaluation metrics. We further show that sparsity calibration, motivated by the structure of single-cell data, substantially improves DE classification for deep learning models that do not explicitly account for sparsity. Together, our findings establish practical baselines and evaluation principles for benchmarking perturbation models on DE prediction.</jats:p>