In the face of rapidly accumulating genomic data, our ability to predict key mature RNA properties that underlie transcript function and regulation remains limited. Pretrained genomic foundation models offer an avenue to adapt learned RNA representations to biological prediction tasks; however, existing models are trained using strategies borrowed from textual domains that do not leverage biological domain knowledge. Here we introduce Orthrus, a Mamba-based mature RNA foundation model pretrained using a self-supervised contrastive learning objective with biological augmentations. Orthrus is trained by maximizing embedding similarity between pairs of RNA transcripts that are formed from splice isoforms of ten model organisms and transcripts from orthologous genes in 400+ mammalian species. This training objective results in a latent representation that clusters RNA sequences with functional and evolutionary similarities. Orthrus' mature RNA isoform representations outperform genomic foundation models on mRNA property prediction tasks, requiring only a fraction of fine-tuning data. Finally, we show that Orthrus is capable of capturing divergent biological function of individual transcript isoforms.
Genomic loci associated with common traits and diseases are typically non-coding and likely impact gene expression, sometimes coinciding with rare loss-of-function variants in the target gene. However, our understanding of how gradual changes in gene dosage affect molecular, cellular, and organismal traits is currently limited. To address this gap, we induced gradual changes in gene expression of four genes using CRISPR activation and inactivation. Downstream transcriptional consequences of dosage modulation of three master trans-regulators associated with blood cell traits (GFI1B, NFE2, and MYB) were examined using targeted single-cell multimodal sequencing. We showed that guide tiling around the TSS is the most effective way to modulate cis gene expression across a wide range of fold-changes, with further effects from chromatin accessibility and histone marks that differ between the inhibition and activation systems. Our single-cell data allowed us to precisely detect subtle to large gene expression changes in dozens of trans genes, revealing that many responses to dosage changes of these three TFs are nonlinear, including non-monotonic behaviours, even when constraining the fold-changes of the master regulators to a copy number gain or loss. We found that the dosage properties are linked to gene constraint and that some of these nonlinear responses are enriched for disease and GWAS genes. Overall, our study provides a straightforward and scalable method to precisely modulate gene expression and gain insights into its downstream consequences at high resolution.
Characterizing cellular aging is essential for understanding age-related diseases. While tissue-level studies reveal broad age-associated changes, they often reflect compositional shifts rather than cell-level reprogramming. The cellular damage hypothesis posits that aging involves the accumulation of DNA, chromatin, and other damage across molecular layers, increasing transcriptional entropy. Existing supervised methods for detecting cellular senescence yield cell type-specific senescence scores but rely on labeled data and lack generalizability. Here, we introduce a first-principles framework for quantifying transcriptional entropy in single cells as each cell's deviation from a transcriptomic manifold, capturing breakdown of transcriptional coordination. This unsupervised approach identifies aging-affected cell types and distinguishes two cellular aging mechanisms: loss of expression precision and activation of stress-response pathways in high entropy cells. Applied to Tabula Muris Senis and SenNet Multiome datasets, transcriptional entropy correlates with chromatin-based mitotic age and highlights regenerative tissue compartments as most affected by aging.
Primary de novo high grade gliomas, such as glioblastoma and lower grade gliomas both converge on a common aggressive phenotype, and the basis for this progression is unknown. Glioma associated macrophages (GAM) have been strongly implicated in supporting tumor growth, however, robust isolation of functional subpopulations has been elusive. We hypothesize that functional populations of GAMs can be resolved through gene regulatory network (GRN) inference and show that a subpopulation of human GAMs, defined by a GRN centered around the activator protein-1 transcription factor FOSL2 is preferentially enriched in high grade gliomas. We nominate ANXA1 and HMOX1 as surrogate cell surface markers for a subpopulation we term malignancy associated GAMs (mGAMs) which possess distinct pro-tumorigenic properties, share partial ontogeny with peripheral blood monocytes, and are enriched in newly transformed regions of glioma. mGAMs potentially play a pivotal role in glioma progression and represent a plausible therapeutic target.
Transcript diversity including splicing and alternative 3′ end usage is crucial for cellular identity and adaptation, yet its spatial coordination remains poorly understood. Here we present SPLISOSM (spatial isoform statistical modeling), a method for detecting isoform-resolution patterns from spatial transcriptomics data. SPLISOSM uses multivariate testing with nonparametric kernels to account for spot-level and isoform-level dependencies, achieving high statistical power on sparse data. In the mouse brain, we identify over 1,000 spatially variable transcript diversity events, primarily in synaptic signaling pathways linked to neuropsychiatric disorders, and uncover both known and previously unknown regulatory relationships with region-specific RNA binding proteins. We further show that these patterns are evolutionarily conserved between mouse and human prefrontal cortex. Analysis of human glioblastoma highlights pervasive transcript diversity in antigen presentation and adhesion genes associated with specific microenvironmental conditions. Together, we present a comprehensive spatial splicing analysis in the brain under normal and neoplastic conditions. Differential isoform usage is identified with high statistical power from spatial transcriptomics data.