Abstract
<title>Abstract</title> <p>In large language models, factual and relational knowledge is often carried by a strikingly regular geometry: concepts correspond to linear directions, antonyms and analogies to parallel offsets, and taxonomies to nested, near-orthogonal subspaces. Single-cell foundation models (SCFMs)—“biological large language models” trained with the same masked-token recipe on transcriptomes—are now routinely mined for biological insight, yet whether their representations share this clean geometry of knowledge is largely untested. We conduct a systematic empirical analysis of four SCFMs spanning 8.9M–650M parameters (scBERT, Geneformer, scGPT, UCE) on a 0.42M-cell atlas panel with rich metadata and a hematopoietic differentiation trajectory, comparing every measurement against expression-space baselines (highly variable genes, PCA, and scVI). Across six families of analyses—global geometry, linear and nonlinear probing, concept-direction structure, categorical/hierarchical geometry, cross-model similarity, and trajectory-manifold geometry—we find that the geometry of knowledge in these models is real but modest. Representations are low-dimensional (intrinsic dimension 12–26 versus widths of 200–1280) and strongly anisotropic. Coarse categorical knowledge—compartment, broad cell type, sex—is linearly decodable and approximately hierarchically organized, but this structure largely coincides with what PCA already recovers from expression (median incremental gain of 3.5 balanced-accuracy points, and negative for batch). Finer subtypes, disease state, and continuous developmental time are present but nonlinearly entangled: the linear-to-nonlinear probe gap reaches 9 points, and a single linear direction recovers only 𝑅 2 ≈ 0.61 of pseudotime where a curved geodesic reaches 𝜌 ≈ 0.80—no better than a classical diffusion map. Concept directions are only weakly parallel (mean cosine 0.17–0.26, rising to 0.28–0.41 under a causal whitening) though unrelated concepts are near-orthogonal, and cross-model geometric convergence is moderate (CKA 0.55–0.74 between models but 0.36–0.45 to a cell-ontology kernel), increasing only slightly with scale. We conclude that current biological LLMs encode a partially linearized, low-dimensional geometry that is useful but far from the crisp structure documented in language models, and that most of its linearly accessible content is shared with classical expression embeddings.</p>