Back to Search View Original Cite This Article

Abstract

<title>Abstract</title> <p>Single-cell foundation models are trained on millions of transcriptomes, but it remains unclear whether their internal representations contain biological relationships that were not supplied as labels. We test one narrow claim: genes whose proteins physically interact should be closer in the pretrained gene-token space of scGPT than carefully matched gene pairs without a recorded interaction. We extracted the 512-dimensional token embedding matrix from the official whole-human scGPT checkpoint, without fine-tuning. Positive pairs came from 89,187 multi-validated human physical interactions in BioGRID 5.0.259. For every positive pair, we sampled one unrecorded pair while matching endpoint degree in 20 strata. As an independent replication, we used STRING v12 physical links at confidence thresholds of 400, 700, and 900. Cosine similarity showed modest discrimination in BioGRID (ROC-AUC 0.581, 95% gene-block interval 0.573–0.589; average precision 0.604). The association strengthened monotonically with STRING confidence: ROC-AUC was 0.612, 0.657, and 0.686 at the three thresholds. Results were stable across ten negative draws, raw versus LayerNorm-transformed embeddings, and most degree strata; random 512-dimensional vectors produced chance performance. These findings provide direct evidence that a pretrained single-cell model’s static token geometry contains physical-interaction structure. The effect is not strong enough for stand-alone interaction prediction, and database absence is not proof of a negative interaction. The results therefore support biological containment, not causal understanding or downstream superiority.</p>

Show More

Keywords

interaction singlecell models biological pretrained

Related Articles

PORE

About

Connect