Back to Search View Original Cite This Article

Abstract

<title>Abstract</title> <p>Biological foundation models are often said to organize genes according to biological function. Such claims are difficult to interpret when the tested property is also present in the model’s training data. We examine a property that scGPT never receives explicitly: a gene’s coordinate in the human genome. We extracted the 512-dimensional input embeddings from the public wholehuman scGPT checkpoint and mapped 19,012 unique protein-coding gene symbols to GENCODE release 50. The primary benchmark distinguished same-chromosome gene pairs whose midpoints were at most 1 megabase (Mb) apart from an equal number of chromosome-matched pairs more than 10 Mb apart. Local pairs had mean cosine similarity 0.0618, compared with 0.0316 for far pairs (Cohen’s d = 0.323). The chromosome-macro-averaged AUROC was 0.589 (chromosomeblock 95% confidence interval, 0.575–0.603); all 24 chromosome-level AUROCs exceeded 0.50. One hundred within-chromosome gene-label permutations gave mean macro-AUROC 0.500. The association was concentrated at short distances: mean cosine was 0.1009 below 100 kilobases and 0.0563 from 100 kilobases to 1 Mb, then declined toward the permuted baseline. Genomic neighbors were also enriched 15.0-fold among each gene’s ten nearest embedding neighbors. These results establish a small but reproducible association between scGPT’s static gene geometry and linear genomic neighborhood. They do not establish that the model represents chromosome coordinates, uses the association causally, or captures three-dimensional chromatin contacts.</p>

Show More

Keywords

pairs genes from gene mean

Related Articles

PORE

About

Connect