Back to Search View Original Cite This Article

Abstract

<jats:p>Single-cell RNA sequencing (scRNA-seq) simultaneously provides gene-expression profiles and genetic variants from individual cells, creating an opportunity to relate cellular phenotypes to their somatic evolutionary histories. However, delineation of genetic type (GTs) from scRNA-seq remains difficult because most variant positions are unobserved in individual cells and the observed base calls contain substantial false-positive and false-negative errors. We evaluated some existing phylogenetic and imputation methods using one simulated dataset and two empirical tumor datasets. We found that extreme sparsity prevented reliable recovery of known or independently inferred GTs when multiple GTs were present. This led us to adapt the STICI transformer architecture to train a separate model de novo on each sparse cell-variant (CV) matrix. These data-specific models predicted millions of missing bases, greatly reducing matrix sparsity. Phylogenetic analyses of the imputed CV matrices showed substantially improved recovery of GTs in both simulated and empirical datasets. In the empirical dataset, transformer-based analysis also suggested finer-scale genetic structure within some previously reported GTs that was not apparent with the existing methods. These results demonstrate that highly sparse scRNA-seq datasets contain substantially more recoverable lineage information than previously appreciated and that de novo transformer modeling provides an effective approach for recovering much of this hidden information. Nevertheless, sequencing errors persisted, limiting reconstruction of cellular lineage structure and leaving significant room for methodological improvement before expression phenotypes can be examined reliably in the context of their cellular evolutionary relationships.</jats:p>

Show More

Keywords

scrnaseq genetic cellular empirical datasets

Related Articles

PORE

About

Connect