Abstract
<title>Abstract</title> <p>Virtual-cell research increasingly combines mechanistic simulators, single-cell foundation models, perturbation predictors, and multimodal integrators. Yet world-model terminology is often applied to systems that support substantially different capabilities. We introduce virtual cell world models (VCWMs) as a structured framework for evaluation rather than as settled nomenclature. The framework isolates three recurring gaps: representation is not dynamics, prediction is not intervention, and multimodality is not multiscale world modeling. Starting from a working definition of general world models, we formalize a VCWM as a maintained cellular state with biological and observational context, intervention-conditioned transition, and time-varying structure. We present three motivation experiments that provide empirical evidence for these gaps in current systems: foundation-model representations can improve present-state readouts without comparable future-fate signal; a perturbation predictor can forecast endpoints while failing state-space closure under iteration; and a multimodal model can learn cross-modal association without producing appreciable chromatin-to-RNA intervention effects. These results motivate three diagnostic axes—dynamics, intervention, and scale—each with its own L0–L3 capability ladder, and a staged roadmap from candidate VCWMs to multiscale, interactive systems.</p>