Abstract
<title>Abstract</title> <p>A growing program argues that biological foundation models internalize a rich complex geometry of knowledge—low-dimensional manifolds, multi-scale spectral structure, and linearly organized concepts— and much of the strongest evidence comes from single-cell transformers. We ask whether protein language models (pLMs), the other dominant family of biological foundation model, share this geometry. Applying a three-lens battery—manifold, spectral, and concept structure—to ESM-2 (8M–3B), ProtBERT, ProtT5, and Ankh, each against matched baselines and synthetic positive controls, we reach a largely negative but nuanced conclusion. Manifold: the nonlinear intrinsic dimension is small (TwoNN ≈26 for ESM-2 650M) yet the linear dimension is an order of magnitude larger (participation ratio ≈131; 64–128 linear directions needed for downstream saturation), and no nonlinear manifold coordinate ever beats a linear probe—the hallmark of a high-linear-dimensional, near-linear cloud, not a smooth low-dimensional manifold. Spectral: the covariance eigenspectrum is a single smooth power law 𝜆𝑖 ∝𝑖 −𝛼 (𝑅 2 >0.99) with 𝛼 near 1, which by the smoothness criterion of Stringer et al. implies an effectively high-dimensional code; there is no multi-band spectral structure, and no spectral gap aligns with a biological partition beyond the shuffled null, although the exponent 𝛼 is itself an informative, label-free quality summary (𝜌=−0.83 with downstream performance). Concept: many biological concepts are strongly linearly decodable (secondary structure, disorder, membrane, localization), so the linear representation hypothesis holds at the level of decodability—but the concept geometry is flat: taxonomy is barely reflected (cophenetic 𝜌=0.34), analogical directions are weakly parallel (mean cosine 0.14–0.24), and much apparent structure dissolves under composition and homology controls. The identical battery recovers manifolds, multi-band spectra, and hierarchies in synthetic controls, so the negatives reflect the data. We conclude that biological knowledge in pLMs is real and linearly accessible but geometrically simple, and argue that the discrete, homology-structured nature of sequence space explains why the complex geometry reported for single-cell models does not transfer.</p>