The piRNA biogenesis machinery localizes to phase separated nuage granules, but nuage function is not well understood. We therefore assayed nuage composition, piRNA expression and transposon silencing in Drosophila mutants that disrupt piRNA precursor production and nuclear export, ping-pong amplification and phased piRNA biogenesis. These mutations destabilize the genome and activate Chk2 signaling and chk2/mnk double mutants were therefore analyzed in parallel. Aub and Vasa are required for ping-pong amplification and Armi promotes phased piRNA processing. We show that Chk2 activation releases Aub and Vasa from nuage and that piRNA precursors are required for nuage localization of the ping-pong and phased biogenesis machinery. However, this analysis also indicates that Vasa, Aub, and Armi concentration in nuage is dispensable for piRNA production and transposon silencing, indicating that dispersed cytoplasmic proteins can drive these processes. We speculate that nuage sequesters silencing effectors, which are released by Chk2 in response to transposon mobilization.
Small nuclear RNAs (snRNAs) are essential components of the spliceosome and are encoded by large, multicopy gene families. However, their genome-wide identification and quantification have remained challenging due to high sequence similarity among family members. To address this, we utilized RAMPAGE (Rapid Amplification of cDNA Ends) data from the ENCODE project to comprehensively profile nascent transcription of spliceosomal snRNAs across 115 human biosamples. We identified 74 expressed snRNA variants, characterized by canonical promoter features including bidirectional transcription flanking a positioned nucleosome, active histone modifications, and evolutionary conservation- features largely absent from unexpressed variants. These transcriptional events were corroborated by total RNA-seq and Bru-seq data, yet the majority of these variants showed extremely low levels in small RNA-seq, indicating post-transcriptional bottlenecks for snRNA processing and maturation. Our findings reveal new layers of regulation in snRNA variant expression and suggest that selective post-transcriptional processing plays a critical role in shaping the functional snRNA repertoire and its contribution to splicing regulation.
Mammalian genomes contain millions of regulatory elements that control the complex patterns of gene expression1. Previously, the ENCODE consortium mapped biochemical signals across hundreds of cell types and tissues and integrated these data to develop a registry containing 0.9 million human and 300,000 mouse candidate cis-regulatory elements (cCREs) annotated with potential functions2. Here we have expanded the registry to include 2.37 million human and 967,000 mouse cCREs, leveraging new ENCODE datasets and enhanced computational methods. This expanded registry covers hundreds of unique cell and tissue types, providing a comprehensive understanding of gene regulation. Functional characterization data from assays such as STARR-seq3, massively parallel reporter assay4, CRISPR perturbation5,6 and transgenic mouse assays7 have profiled more than 90% of human cCREs, revealing complex regulatory functions. We identified thousands of novel silencer cCREs and demonstrated their dual enhancer and silencer roles in different cellular contexts. Integrating the registry with other ENCODE annotations facilitates genetic variation interpretation and trait-associated gene identification, exemplified by the identification of KLF1 as a novel causal gene for red blood cell traits. This expanded registry is a valuable resource for studying the regulatory genome and its impact on health and disease.
RNA-binding proteins (RBPs) are essential modulators in the regulation of mRNA processing. The binding patterns, interactions, and functions of most RBPs are not well-characterized. Previous studies have shown that motif context is an important contributor to RBP binding specificity, but its precise role remains unclear. Despite recent computational advances to predict RBP binding, existing methods are challenging to interpret and largely lack a categorical focus on RBP motif contexts and RBP-RBP interactions. There remains a need for interpretable predictive models to disambiguate the contextual determinants of RBP binding specificity in vivo. Here, we present a novel and comprehensive pipeline to address these knowledge gaps. We devise a natural language processing-based method to deconstruct sequences into entities comprising a target k-mer and its flanking regions, then use this representation to formulate RBP binding prediction as a weakly supervised multiple instance learning problem. To interpret our predictions, we introduce a deterministic motif discovery algorithm to leverage our data structure, recapitulating the established motifs of numerous RBPs as validation. Importantly, we characterize the binding motifs and binding contexts for 71 RBPs in HepG2 and 74 RBPs in K562, with many of them being novel. Finally, through feature integration, transitive inference, and a new cross-prediction approach, we propose novel cooperative and competitive RBP-RBP interaction partners and hypothesize their potential regulatory functions. In summary, we present a complete framework for investigating the contextual determinants of specific RBP binding, and we demonstrate the significance of our findings in delineating RBP binding patterns, interactions, and functions.
Transposable elements (TEs) in the human genome are the heritage of ancient parasitic infections. While most of human DNA comprises TEs and TE-derived elements, their repetitive nature poses technical challenges; thus, little is known about their positional identity and regulatory roles. Here, by integrating long-read and multidimensional transcriptional analyses, we investigate when, where and how TEs become part of a gene. We characterize how TE-derived isoforms change across mouse-human variation and how they are linked to gene regulatory networks controlling cell states during differentiation, organogenesis and health (aging and pathological states). Mechanistically, we identify an RNA degradation-dependent and splicing-dependent quality control mechanism that operates independently of conventional mechanisms of TE suppression, such as DNA methylation and heterochromatinization, and prevents TE-chimera expression and TE-induced cell differentiation. Overall, our findings unveil mechanisms by which viral-derived elements enhance transcriptome plasticity.
The Functional Annotation of Variants Online Resource (FAVOR), http://favor.genohub.org, is a whole genome variant annotation database and portal that provides comprehensive variant functional annotations of all possible variants across the genome. It can facilitate the analysis of whole-genome sequencing studies, support the interpretation of variant functional impacts, and help prioritize causal variants of diseases or traits. To support the growing popularity and expand the scope of FAVOR, we present here a substantial platform update. The new release features dramatically expanded annotations, a completely redesigned infrastructure powered by a newly implemented application programming interface (FAVOR-API), and a revamped web interface with advanced data-visualization capabilities and enhanced query performance. Key expansions include much more comprehensive variant annotations, including global, tissue- and cell-type-specific variant annotations; gene and protein annotations; support for both hg38 and hg19 reference genomes; and an interactive genome-browser for visualization of multi-faceted variant annotations. The updated platform also includes FAVOR-GPT, a large language model-powered interface for navigating the FAVOR database and interpreting results. FAVOR continues to evolve to keep pace with advances in research on interpreting the functional and phenotypic impact of genomic variation.
RNA-binding proteins (RBPs) regulate their RNA targets by binding to short sequence motifs, but the underlying mechanisms enabling sequence-specific recognition within the vast transcriptome remain unclear for the majority of human RBPs. Sequence contexts are believed to be a significant contributing factor to RBP binding specificity but are often overlooked. Further, existing motif discovery algorithms do not consider the structure and composition of the motif's flanking regions in their construction, which represents a consequential shortcoming. Herein, we present a novel linguistics-inspired RBP motif and context discovery algorithm that is consensus-based, deterministic, and flexible. Our algorithm draws multiple parallels between natural language and genomic language and relies on three important k-mer properties that impart lexical, syntactic, and semantic structures and rules to the process of motif and context discovery. Critically, our algorithm integrates information from sequence contexts when constructing RBP motifs. We demonstrate that our algorithm achieves strong discovery accuracy against a ground-truth set, and even outperforms existing methods in primary motif ranking.
Environmental stress activates transposons and is proposed to generate genetic diversity that facilitates adaptive evolution. piRNAs guide germline transposon silencing, but the impact of stress on the piRNA pathway is not well understood. In Drosophila, the Rhino-Deadlock-Cuff complex (RDC) drives transcription of clusters composed of nested transposon fragments, generating precursors that are processed into mature piRNAs in the cytoplasm. We show that acute heat shock triggers rapid, reversible loss of RDC localization and cluster transcript expression with coordinate changes in the cytoplasmic processing machinery. Maternal piRNAs bound to Piwi are proposed to guide Rhino localization to clusters during early embryogenesis. However, RDC relocalization after heat shock is accelerated in piwi mutants and delayed in thoc7 mutants, which disrupt piRNA precursor binding to THO complex, and we show that maternally deposited piRNAs are dispensable for RDC localization to the major 42AB cluster. Cluster specification is reconsidered in light of these findings.
Identifying transcriptional enhancers and their target genes is essential for understanding gene regulation and the effect of human genetic variation on disease1-6. Here we create and evaluate a resource of more than 92 million enhancer-gene regulatory interactions across 1,458 biosamples covering 369 cell types and tissues, by integrating predictive models, chromatin states, three-dimensional contacts and large-scale genetic perturbations generated by the ENCODE Consortium7. We first create a systematic benchmarking pipeline to compare predictive models, assembling a dataset of 10,356 element-gene pairs measured in CRISPR perturbation experiments, more than 30,000 fine-mapped expression quantitative trait loci and 569 fine-mapped genome-wide association study (GWAS) variants linked to a probable causal gene. Using this framework, we develop ENCODE-rE2G, a predictive model achieving state-of-the-art performance across several prediction tasks, demonstrating that iterative perturbations and supervised machine learning can build increasingly accurate predictive models of enhancer regulation. Using ENCODE-rE2G, we build an encyclopedia of enhancer-gene regulatory interactions in the human genome, revealing global properties of enhancer networks, identifying differences in regulatory complexity across genes and improving analyses linking noncoding variants to target genes and cell types for common complex diseases. By interpreting the model, we find that beyond enhancer activity and three-dimensional enhancer-promoter contacts, additional features that guide enhancer-promoter communication include promoter class and enhancer-enhancer synergy. These genome-wide maps of enhancer-gene regulatory interactions, benchmarking software, predictive models and insights about enhancer function provide a valuable resource for future studies of gene regulation and human genetics.
RNA-binding proteins (RBPs) are critical regulators of the human transcriptome, but the binding patterns of most RBPs are insufficiently characterized. While sequence context facilitates RBP binding specificity, its precise contribution remains unclear. Existing computational methods to decipher RBP binding patterns are limited by their architecture-dependence, challenging interpretability, and, importantly, lack of focus on context. We present a novel comprehensive approach to address the aforementioned knowledge gaps. We first introduce a natural language-based representation to model RNA sequences using lexical, syntactic, and semantic forms, then devise a sequence decomposition method based on these structures to deconstruct RNA sequences into regions, each containing a target k-mer and its flanking contexts. We leverage this linguistic conceptualization to predict RBP binding under a Multiple Instance Learning (MIL) framework, which we solve using a novel method of significant region extraction termed "iterative relabeling". We demonstrate that our bottom-up approach discovers key regions contributing to RBP binding in an architecture-dependent, accurate, and interpretable manner.
The oncogenic transcription factor MYB is a master regulator of self-renewal, differentiation and proliferation of hematopoietic cells. MYB is frequently aberrantly expressed in hematologic malignancies and has been revealed to be a central component of oncogenic complexes maintaining aberrant gene expression programs in AML. RGT-61159 is a potent, selective oral small molecule inhibitor of the oncogene MYB via the inclusions of a cryptic exon into MYB RNA transcripts, resulting in the activation of the nonsense-mediated decay pathway promoting MYB mRNA depletion and sequentially MYB protein degradation. The RGT-61159 Phase 1 study in adults with relapsed/refractory ACC or CRC has been initiated (NCT06462183). Here, we report genomic analyses of different AML cell lines treated with RGT-61159 that provided further insight into key oncogenes (e.g., BCl2, NOTCH, MYC) directly regulated by MYB. To further evaluate RGT-61159’s potential to treat leukemia, RGT-61159 single agent was profiled in a battery of murine models of AML harboring the most common genetic lesions (e.g., Flt3 ITD, MLL-fusion and NPM1 mutation). RGT-61159 showed significant anti-tumor activity across the different AML tumor models tested along with robust MYB RNA depletion in tumor cells at tolerated doses. In addition, the benefit of combining RGT-61159 with a panel of standard-of-care agents for AML treatment was explored in vitro across a panel of AML cell lines harboring different genetic alterations. The data revealed synergistic cell killing activity when RGT-61159 was combined with Flt3 inhibitors (e.g. gilteritinib and midostaurin) in MOLM-13 and MV4.11 cell lines harboring Flt3 IDT mutation, and with menin inhibitors (e.g. KO-539, revemenib) in AML cell lines with NPM1 mutant) or MLL-fusion. Finally, RGT-61159 combination with a Bcl2 inhibitor (venetoclax) resulted in synergistic cell killing activity in AML cell lines carrying NPM1 mutations or AML-1-ETO fusion protein. These studies provide compelling supportive evidence for combining RGT-61159 with standards-of-care for AML treatment in the relevant population of patients with AML. . Norman Lu, Patricia Soulard, Xiubin Gu, Chris Yates, Kai Li, Sam Hasson, Ibrahim Kay, Zhiping Weng, Travis Wager, Simon Xi. RGT-61159, best-in-class oral small molecule inhibitor of MYB via selective RNASplicing alteration, synergistic anti-tumor aActivity when combined with standards of care in leukemia disease models harboring AML common genetic lesions and with NOTCH inhibitors in ACC disease models [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2025; Part 1 (Regular Abstracts); 2025 Apr 25-30; Chicago, IL. Philadelphia (PA): AACR; Cancer Res 2025;85(8_Suppl_1):Abstract nr 7013.
Koala populations in Australia face a barrage of threats, chiefly, habitat degradation and the effects of climate change including drought and bushfire. Further, high rates of chlamydiosis, linked to koala retrovirus (KoRV) viral load, is a major contributing factor to northern population decline. However, recent work by Yu et al., (Cell, 2024) has provided a glimmer of hope: some koalas have evolved ‘adaptive genome immunity’, which is able to actively suppress endogenous KoRV transcription. A single KoRV-A provirus insertion within MAP4K4 gene's 3’ UTR is shown to be the trigger for production of sense and anti-sense piRNAs, and that MAP4K4 KoRV integration is linked to both a 20% reduction in proviral genome integrations and 10-fold reduction of KoRV transcription within male germline tissue. Here we discuss how this finding offers the potential to reduce koala disease burden and can be incorporated into conservation management to help save this iconic species.
Deciphering the regulatory syntax of the genome is essential to understand the genetic and molecular architecture of complex traits, as most trait-associated variants lie in non-coding regions. Yet, functional annotation of the bovine genome remains limited, hindering our ability to unravel the mechanisms underpinning complex traits of economic and ecological importance in cattle. Here, we present a comprehensive epigenetic atlas comprising 1,138 genome-wide epigenetic profiles, including chromatin accessibility, six histone modifications, CCCTC-binding factor (CTCF) transcription factor binding, DNA methylation, chromatin conformation, and transcriptomes across 53 adult tissues, five fetal tissues, and seven primary cell types. This atlas-level data enables us to annotate around 45% of the genome as putative regulatory elements exhibiting tissue- or cell-specific regulatory activity. Leveraging sequence-to-function deep learning models, we discovered 301 sequence motifs and predicted the functional impact of genetic variants through in silico mutagenesis, thereby facilitating the decoding of the regulatory syntax of the cattle genome and fine-mapping of GWAS loci for 22 complex traits. Cross-species analysis further revealed evolutionarily conserved features of regulatory architecture and provided evolutionary insights into complex traits and diseases in humans. Together, this atlas offers a foundational resource for advancing cattle functional genomics, sustainable breeding, and studies of regulatory evolution.
Koala retrovirus-A (KoRV-A) is spreading through wild koalas in a north-to-south wave while transducing the germ line, modifying the inherited genome as it transitions to an endogenous retrovirus. Previously, we found that KoRV-A is expressed in the germ line, but unspliced genomic transcripts are processed into sense-strand PIWI-interacting RNAs (piRNAs), which may provide an initial "innate" form of post-transcriptional silencing. Here, we show that this initial post-transcriptional response is prevalent south of the Brisbane River, whereas KoRV-A expression is suppressed, promoters are methylated, and sense and antisense piRNAs are equally abundant in a subpopulation of animals north of the river. These animals share a KoRV-A provirus in the MAP4K4 gene's 3' UTR that is spreading through northern koalas and produces hybrid transcripts that are processed into antisense piRNAs, which guide transcriptional silencing. We speculate that this provirus triggers adaptive transcriptional silencing of KoRV-A and is sweeping to fixation.
Transposons constitute ~45% of the human genome, driving gene evolution and contributing to disease, but their repetitive nature complicates the identification of new insertions. We present LOCATE (Long-read to Characterize All Transposable Elements), an algorithm using long-read sequencing to detect and assemble transposon insertions. LOCATE outperforms existing tools on simulated datasets and achieves the best performance in two previous benchmarks, as well as in a new benchmark we constructed using real biological datasets. Applying LOCATE to public datasets revealed that pre-existing Alu copies create two hotspots for Alu and LINE1 insertions: the A-rich linker and the poly(A) tail. We further observed a preference for self-insertions over non-self-insertions in Alu and LINE1, suggesting a "feedforward" transposition mechanism in which Alu and LINE1 RNA transcripts target the hotspots of their source copies to generate new insertions. LOCATE enhances our ability to study transposons and their role in genome dynamics. ### Competing Interest Statement Z. Weng is a co-founder of Rgenta Therapeutics and she serves on its scientific advisory board.
DNA motif discovery and, particularly, computational modeling of transcription factor binding motifs, has been a mecca of algorithmic bioinformatics for several decades. Here, we report the results of the largest open community challenge in Inferring BInding Specificities (IBIS), where participants all over the world were invited to construct binding specificity models from multi-assay experimental data for poorly studied human transcription factors. The submissions were rigorously tested against a rich held-out dataset. Benchmarking demonstrated a consistent advantage of properly designed deep learning models over traditional positional weight matrices and other machine learning methods. Yet, the positional weight matrices displayed a surprisingly strong performance out of the box, being only slightly behind the best deep learning models. A post-challenge assessment of a selection of other deep learning methods further solidified this finding. IBIS highlights the power of benchmarking in finding adequate DNA motif representations, emphasizes the pros and cons of various machine learning methods applied to DNA motif modeling, and establishes a rich dataset, benchmarking protocols, and computational framework for a fair cross-platform evaluation of future models of transcription factor binding motifs in DNA sequences.
Among the major classes of RNAs in the cell, tRNAs remain the most difficult to characterize via deep sequencing approaches, as tRNA structure and nucleotide modifications can each interfere with cDNA synthesis by commonly used reverse transcriptases (RTs). Here, we benchmark a recently developed RNA cloning protocol, termed Ordered Two-Template Relay (OTTR), to characterize intact tRNAs and tRNA fragments in budding yeast and in mouse tissues. We show that OTTR successfully captures both full-length tRNAs and tRNA fragments in budding yeast and in mouse reproductive tissues without any prior enzymatic treatment, and that tRNA cloning efficiency can be further enhanced via AlkB-mediated demethylation of modified nucleotides. As with other recent tRNA cloning protocols, we find that a subset of nucleotide modifications leave misincorporation signatures in OTTR datasets, enabling their detection without any additional protocol steps. Focusing on tRNA cleavage products, we compare OTTR with several standard small RNA-Seq protocols, finding that OTTR provides the most accurate picture of tRNA fragment levels by comparison to 'ground truth' Northern blots. Applying this protocol to mature mouse spermatozoa, our data dramatically alter our understanding of the small RNA cargo of mature mammalian sperm, revealing a far more complex population of tRNA fragments - including both 5' and 3' tRNA halves derived from the majority of tRNAs - than previously appreciated. Taken together, our data confirm the superior performance of OTTR to commercial protocols in analysis of tRNA fragments, and force a reappraisal of potential epigenetic functions of the sperm small RNA payload.
Aging brings dysregulation of various processes across organs and tissues, often stemming from stochastic damage to individual cells over time. Here, we used a combination of single-nucleus RNA-sequencing and single-cell whole-genome sequencing to identify transcriptomic and genomic changes in the prefrontal cortex of the human brain across life span, from infancy to centenarian. We identified infant-specific cell clusters enriched for the expression of neurodevelopmental genes, and a common down-regulation of cell-essential homeostatic genes that function in ribosomes, transport, and metabolism during aging across cell types. Conversely, expression of neuron-specific genes generally remains stable throughout life. We observed a decrease in specific DNA repair genes in aging, including genes implicated in generating brain somatic mutations as indicated by mutation signature analysis. Furthermore, we detected gene-length-specific somatic mutation rates that shape the transcriptomic landscape of the aged human brain. These findings elucidate critical aspects of human brain aging, shedding light on transcriptomic and genomics dynamics.