Back to Search View Original Cite This Article

Abstract

<title>Abstract</title> <p>Biological foundation models are often claimed to learn structured biological knowledge, but many interpretability studies still rely on qualitative examples or downstream task scores. This paper asks a narrow empirical question: does the learned gene-token embedding space of scGPT encode known physical interaction structure among human gene products? We analyzed the public scGPT human checkpoint, extracted its 512-dimensional gene-token embedding matrix, and compared cosine similarities between known physical interaction partners and endpointfrequency-matched non-interacting pairs. Physical interactions were taken from STRING v12.0 and BioGRID 5.0.259. Across 18,441 STRING-mapped genes and 11,250 BioGRID-mapped genes, scGPT embeddings showed a consistent but moderate interaction signal. AUROC increased with STRING confidence, from 0.622 for score ≥ 400 edges to 0.704 for score ≥ 900 edges, while gene-label permutation baselines remained at approximately 0.50. Mean embedding cosine also rose monotonically across STRING confidence bins, from 0.011 for matched nonedges to 0.111 for score 900–1000 edges. Nearest-neighbor analysis gave 52.9-fold Precision@10 enrichment over random expectation for STRING score ≥ 700 edges and 17.3-fold enrichment for BioGRID. These results support a limited mechanistic claim: before any cell-specific contextualization, scGPT’s static gene-token geometry already contains detectable physical interaction information. The signal is far from sufficient for interaction prediction on its own, and it should not be read as evidence that scGPT has learned causal regulatory mechanisms.</p>

Show More

Keywords

interaction scgpt physical from string

Related Articles

PORE

About

Connect