Deprecated: Function curl_close() is deprecated since 8.5, as it has no effect since PHP 8.0 in /home/u483256323/domains/poorvam.com/public_html/subdomains/pore/includes/api.php on line 184
Abstract
<jats:p>RNA language models learn useful representations for structure and function, but the biological concepts encoded by their hidden states remain difficult to interpret. BiRNA-BERT is an RNA language model with adaptive byte-pair tokenization, making it an attractive target for mechanistic analysis. We present SPIRAL (Sparse autoencoders for Interpretable RNA Analysis), a layer-wise sparse-autoencoder (SAE) analysis of BiRNA-BERT that extracts interpretable sparse features while preserving the behaviour of the underlying language model. We train independent sparse autoencoders at layers 0, 5, and 11, each expanding the 768-dimensional hidden state into 6,144 dictionary features. Analysis of the learned features reveals biologically meaningful structure and family selectivity. Sparse features align with bpRNA nucleotide-level secondary-structure annotations and RNAcentral type labels: at layer 5, 44.3% of tested features show statistically significant structure association (mean enrichment of 1.61x), and all 1,237 eligible features show significant RNA-type association. RNA-type-selective sparse profiles modestly improve k-nearest-neighbour balanced accuracy over dense BiRNA-BERT embeddings at layer 5. Together, these results show that sparse autoencoders can recover faithful, nucleotide-aware, and biologically interpretable feature decompositions from BiRNA-BERT.</jats:p>