Previous TGIRT-seq analysis of RNAs in Inflammatory Breast Cancer (IBC) patient tumors, peripheral blood mononuclear cells (PBMCs) and plasma identified a short T-cell receptor mRNA fragment (TRBJ1-6) as a potential IBC biomarker that was detected in plasma samples from IBC patients but not patients with non-inflammatory breast cancer or healthy donors. Here, we traced the origin of this TRBJ1-6 RNA fragment to IBC patient PBMCs and used a high-throughput RT-PCR/Cas12a assay with larger numbers of samples to confirm its prevalence in IBC patient PBMCs. Detection of this RNA was enhanced by T4 polynucleotide kinase treatment, indicating the presence of a 2',3'-cyclic phosphate. Analysis of previous TGIRT-seq datasets revealed gene expression differences in IBC patient PBMCs that could contribute to TRBJ1-6 RNA prevalence in IBC patient PBMCs and plasma. Our results support the identification of the TRBJ1-6 RNA fragment as a novel, readily detectable blood-based RNA biomarker derived from IBC-patient immune cells, addressing a major unmet need for diagnosing IBC.
Gene disruption analysis revealed that an E. coli PPRT protein, which has an N-terminal Primase-Polymerase (PrimPol) domain fused to a group II intron-like reverse transcriptase (RT) domain followed by a long C-terminal domain (CTD), contributes to a cellular oxidative DNA damage response in addition to its previously described function in phage defense. Biochemical analysis showed that the PrimPol domain has an error-prone DNA polymerase activity that enables read through of oxidation-induced DNA damage. Surprisingly, we found that the RT-like domain, in addition to synthesizing protein-primed DNAs for phage defense, has a 3' to 5' DNA exonuclease activity that functions in proofreading DNAs synthesized by the PrimPol domain. Extending these findings, we identified structural features that contribute to this proofreading activity, enabling us to associate it with both a group II intron-encoded and retroviral RT and suggesting general methods for incorporating proofreading activity into RTs.
A previous study found that a bacterial group II intron-like reverse transcriptase (G2L4 RT) evolved to function in double-strand break repair (DSBR) via microhomology-mediated end-joining (MMEJ) and that a mobile group II intron-encoded RT has a basal DSBR activity that uses conserved structural features of non-long terminal repeat (non-LTR)-retroelement RTs. Here, we determined G2L4 RT apoenzyme and snap-back DNA synthesis structures revealing unique structural adaptations that optimized its cellular function in DSBR. These included an RT3a structure that stabilizes the apoenzyme in an inactive conformation until encountering a DNA substrate; a longer N-terminal extension/RT0-loop with conserved residues that together with a modified active site favors strand annealing; and a conserved dimer interface that localizes G2L4 RT homodimers to DSBR sites with both monomers positioned for MMEJ. Our findings reveal how an RT can function in DNA repair and suggest ways of optimizing related RTs for genome engineering applications.
Reverse transcriptase–Cas1 (RT-Cas1) fusion proteins found in some CRISPR systems enable spacer acquisition from both RNA and DNA, but the mechanism of RNA spacer acquisition has remained unclear. Here, we found that Marinomonas mediterranea RT-Cas1/Cas2 adds short 3′-DNA (dN) tails to RNA protospacers, enabling their direct integration into CRISPR arrays as 3′-dN-RNAs or 3′-dN-RNA/cDNA duplexes at rates comparable to similarly configured DNAs. Reverse transcription of RNA protospacers is initiated at 3′ proximal sites by multiple mechanisms, including recently described de novo initiation, protein priming with any dNTP, and use of short exogenous or synthesized DNA oligomer primers, enabling synthesis of near full-length cDNAs of diverse RNAs without fixed sequence requirements. The integration of 3′-dN-RNAs or single-stranded DNAs (ssDNAs) is favored over duplexes at higher protospacer concentrations, potentially relevant to spacer acquisition from abundant pathogen RNAs or ssDNA fragments generated by phage defense nucleases. Our findings reveal mechanisms for site-specifically integrating RNA into DNA genomes with potential biotechnological applications.
A previous study using Thermostable Group II Intron Reverse Transcriptase sequencing (TGIRT-seq) found human plasma contains short (≤300 nt) structured full-length excised linear intron (FLEXI) RNAs with potential to serve as blood-based biomarkers. Here, TGIRT-seq identified >9,000 different FLEXI RNAs in human cell lines, including relatively abundant FLEXIs with cell-type-specific expression patterns. Analysis of public CLIP-seq datasets identified 126 RNA-binding proteins (RBPs) that have binding sites within the region corresponding to the FLEXI or overlapping FLEXI splice sites in pre-mRNAs, including 53 RBPs with binding sites for ≥30 different FLEXIs. These included splicing factors, transcription factors, a chromatin remodeling protein, cellular growth regulators, and proteins with cytoplasmic functions. Analysis of ENCODE datasets identified subsets of these RBPs whose knockdown impacted FLEXI host gene mRNA levels or proximate alternative splicing, indicating functional interactions. Hierarchical clustering identified six subsets of RBPs whose FLEXI binding sites were co-enriched in six subsets of functionally related host genes: AGO1-4 and DICER, including but not limited to agotrons or mirtron pre-miRNAs; DKC1, NOLC1, SMNDC1, and AATF (Apoptosis Antagonizing Transcription Factor), including but not limited to snoRNA-encoding FLEXIs; two subsets of alternative splicing factors; and two subsets that included RBPs with cytoplasmic functions (e.g., LARP4, PABPC4, METAP2, and ZNF622) together with regulatory proteins. Cell fractionation experiments showed cytoplasmic enrichment of FLEXI RNAs with binding sites for RBPs with cytoplasmic functions. The subsets of host genes encoding FLEXIs with binding sites for different subsets of RBPs were co-enriched with non-FLEXI other short and long introns with binding sites for the same RBPs, suggesting overarching mechanisms for coordinately regulating expression of functionally related genes. Our findings identify FLEXIs as a previously unrecognized large class of cellular RNAs and provide a comprehensive roadmap for further analyzing their biological functions and the relationship of their RBPs to cellular regulatory mechanisms.
ATP-grasp superfamily enzymes contain a hand-like ATP-binding fold and catalyze a variety of reactions using a similar catalytic mechanism. More than 30 protein families are categorized in this superfamily, and they are involved in a plethora of cellular processes and human diseases. Here, we identify C12orf29 (RLIG1) as an atypical ATP-grasp enzyme that ligates RNA. Human RLIG1 and its homologs autoadenylate on an active site Lys residue as part of a reaction intermediate that specifically ligates RNA halves containing a 5’-phosphate and a 3’-hydroxyl. RLIG1 binds tRNA in cells and can ligate tRNA within the anticodon loop in vitro. Transcriptomic analyses of Rlig1 knockout mice revealed significant alterations in global tRNA levels in the brains of female mice, but not in those of male mice. Furthermore, crystal structures of a RLIG1 homolog from Yasminevirus bound to nucleotides revealed a minimal and atypical RNA ligase fold with a conserved active site architecture that participates in catalysis. Collectively, our results identify RLIG1 as an RNA ligase and suggest its involvement in tRNA biology.
Abstract Background Understanding the mechanisms of resistance to CDK4/6 inhibitors (CDK4/6i) and endocrine therapy (ET) is pivotal in exploring new therapeutic strategies for hormone receptor-positive (HR+), HER2-negative metastatic breast cancer. To decipher these resistance mechanisms, we analyzed the alterations in comprehensive RNA-seq (coding and non-coding RNAs) of HR+ HER2-negative BC cell lines via thermostable group II intron reverse transcriptase sequencing (TGIRT-seq). Methods We established Tamoxifen-resistant (TMR), Abemaciclib-resistant (ACR), Palbociclib-resistant (PCR), Tamoxifen/Abemaciclib double-resistant (TMR-ACR), and Tamoxifen/Palbociclib double-resistant (TMR-PCR) BC cell lines from MCF7 and T47D HR+/HER2- BC cell lines through stepwise dose-escalation continuous drug exposure. We performed a TGIRT-seq transcriptomic analysis using a protocol that allows the sequencing of both long and short non-coding and protein-coding RNAs in a single library. All libraries were sequenced using paired-end 150 bp on the Novaseq platform, resulting in an average of 50 million reads per library. For analysis, raw reads underwent adapter trimming, small RNA mapping, whole genome mapping, and the generation of annotated genes' read count. We used raw counts to detect differentially expressed genes (DEGs) with DESeq2 in R using the cut-off (log2 [fold change] > 1, FDR < 1e-3). We then used the selected DEGs for Gene Set Enrichment Analysis (GSEA) and Kaplan-Meier survival analysis from the TCGA database. Results The TGIRT-seq analysis identified 1171 to 3472 DEGs in different drug resistance and cell line combinations. Most DEGs in resistant cells were consistently downregulated compared to parent cells. Seventy-five percent of DEGs were protein-coding genes, with the rest being non-coding RNAs, such as small nucleolar RNAs, microRNAs, long non-coding RNAs, and transposable element RNAs, which conventional RNA-seq poorly detects. In principal component analysis (PCA) plots, replicates per drug resistance clustered clearly in both cell line backgrounds. PCA also indicated distinctive pathways for acquiring drug resistance in MCF7 and T47D cells. The estrogen receptor (ER) expression decreased, while HER2 increased in resistant cell lines. DEGs identified were enriched in ESR1-related pathways in all MCF7 and T47D cell lines resistant models. Several ER-regulating genes like IL-1R1 and RET were upregulated, while ADCY1 was consistently downregulated across different resistance types. A significant overlap with single-resistance DEGs was observed among the DEGs in double-resistance cell lines, but 20-50% of up/down-regulated DEGs in double-resistance cell lines were uniquely altered. Several previously identified DEGs (e.g., SALL4, TOP2A) and 21 novel candidate genes correlated with poor survival outcomes. Conclusion The analysis identified unique DEGs in double resistance cell lines, suggesting that double resistance mechanisms may not merely be a cumulative effect of single resistance mechanisms, necessitating further investigation and validation. CDK4/6i and/or ET-resistant BC cell lines displayed significant transcriptome reprogramming during the development of drug resistance. In previously published research, ESR1-related pathway alterations were proposed in tamoxifen resistance cell lines. Here we find that CDK4/6i-resistant cells also modify these ESR1-related pathways. Additionally, we identified several targetable genes, such as IL-1R1 and RET, involved in ESR1-related pathways that could pave the way for developing new treatment strategies. Citation Format: Toshiaki Iwase, Hengyi Xu, Nakyung Oh, Jangsoon Lee, Alan M Lambowitz, Naoto Ueno. Resistance Mechanisms to CDK4/6 Inhibitors and/or Tamoxifen Using Comprehensive Thermostable Group II Intron Reverse Transcriptase Sequencing [abstract]. In: Proceedings of the 2023 San Antonio Breast Cancer Symposium; 2023 Dec 5-9; San Antonio, TX. Philadelphia (PA): AACR; Cancer Res 2024;84(9 Suppl):Abstract nr PO4-23-11.
Senataxin is an evolutionarily conserved RNA-DNA helicase involved in DNA repair and transcription termination that is associated with human neurodegenerative disorders. Here, we investigated whether Senataxin loss affects protein homeostasis based on previous work showing R-loop-driven accumulation of DNA damage and protein aggregates in human cells. We find that Senataxin loss results in the accumulation of insoluble proteins, including many factors known to be prone to aggregation in neurodegenerative disorders. These aggregates are located primarily in the nucleolus and are promoted by upregulation of non-coding RNAs expressed from the intergenic spacer region of ribosomal DNA. We also map sites of R-loop accumulation in human cells lacking Senataxin and find higher RNA-DNA hybrids within the ribosomal DNA, peri-centromeric regions, and other intergenic sites but not at annotated protein-coding genes. These findings indicate that Senataxin loss affects the solubility of the proteome through the regulation of transcription-dependent lesions in the nucleus and the nucleolus.
A recent study found that a bacterial chromosomally encoded group II intron-like reverse transcriptase (G2L4 RT) functions in double-strand break repair (DSBR) via microhomology-mediated end joining (MMEJ) and that this function is dependent upon conserved structural features of non-LTR-retroelement RTs, a family of enzymes that includes bacterial RTs as well as human LINE-1 and other eukaryotic non-LTR-retrotransposon RTs. Here, we determined apoenzyme and co-crystal structures of G2L4 RT that revealed the structual basis of its MMEJ mechanism, including its regulation by transitioning between inactive and active conformations, the role of the conserved RT0 loop in annealing microhomologies, and a contribution of G2L4 RT homodimer formation. Our findings reveal molecular mechanisms by which a non-LTR-retroelement RT functions in DSBR, identify structural features of G2L4 RT that evolved to optimize this function, and provide a structural basis for engineering this activity of non-LTR-retroelement RTs for genome editing and other biotechnological applications. R35 GM131777/GM/NIGMS NIH HHS/United States R35 GM136216/GM/NIGMS NIH HHS/United States.
Inflammatory breast cancer (IBC) is the most aggressive and lethal breast cancer subtype, but lags in biomarker identification. Here, we used an improved Thermostable Group II Intron Reverse Transcriptase RNA sequencing (TGIRT-seq) method to simultaneously profile coding and non-coding RNAs from tumors, PBMCs, and plasma of IBC and non-IBC patients and healthy donors. Besides RNAs from known IBC-relevant genes, we identified hundreds of other overexpressed coding and non-coding RNAs (p≤0.001) in IBC tumors and PBMCs, including higher proportions with elevated intron-exon depth ratios (IDRs), likely reflecting enhanced transcription resulting in accumulation of intronic RNAs. As a consequence, differentially represented protein-coding gene RNAs in IBC plasma were largely intron RNA fragments, whereas those in healthy donor and non-IBC plasma were largely fragmented mRNAs. Potential IBC biomarkers in plasma included T-cell receptor pre-mRNA fragments traced to IBC tumors and PBMCs; intron RNA fragments correlated with high IDR genes; and LINE-1 and other retroelement RNAs that we found globally up-regulated in IBC and preferentially enriched in plasma. Our findings provide new insights into IBC and demonstrate advantages of broadly analyzing transcriptomes for biomarker identification. The RNA-seq and data analysis methods developed for this study may be broadly applicable to other diseases.
Cells respond to perturbations like inflammation by sensing changes in metabolite levels. Especially prominent is arginine, which has known connections to the inflammatory response. Here, we found that depletion of arginine during inflammation decreased levels of a nuclear form of arginyl-tRNA synthetase (ArgRS). Surprisingly, we found that nuclear ArgRS interacts with serine/arginine repetitive matrix protein 2 (SRRM2), a spliceosomal protein and nuclear speckle component and that arginine depletion impacted both condensate-like nuclear trafficking of SRRM2 and splice-site usage in certain genes. These splice-site usage changes cumulated in synthesis of different protein isoforms that altered cellular metabolism and peptide presentation to immune cells. Our findings uncover a novel mechanism whereby a tRNA synthetase cognate to a key amino acid that is metabolically controlled during inflammation modulates the splicing machinery.
Small nucleolar RNAs (snoRNAs) are an omnipresent class of non-coding RNAs involved in the modification and processing of ribosomal RNA (rRNA). As snoRNAs are required for ribosome production, the increase of which is a hallmark of cancer development, their expression would be expected to increase in proliferating cancer cells. However, the nature and extent of snoRNAs contribution to the biology of cancer cells remain largely unexplored. In this study, we examined the abundance patterns of snoRNA in high-grade serous ovarian carcinomas (HGSC) and serous borderline tumours (SBT) and identified a subset of snoRNA associated with increased invasiveness. This subgroup of snoRNA accurately discriminates between SBT and HGSC underlining their potential as biomarkers of tumour aggressiveness. Remarkably, knockdown of HGSC-associated H/ACA snoRNAs, but not their host genes, inhibits cell proliferation and induces apoptosis of model ovarian cancer cell lines. Wound healing and cell migration assays confirmed the requirement of these HGSC-associated snoRNA for cell invasion and increased tumour aggressiveness. Together our data indicate that H/ACA snoRNAs promote tumour aggressiveness through the induction of cell proliferation and resistance to apoptosis.
Background: Inflammatory breast cancer (IBC) is the most aggressive and lethal breast cancer subtype but lags in disease-specific RNA biomarkers due in part to its paucity of large discrete tumors. A strategy to overcome this challenge is to identify blood-based RNA biomarkers that are minimally invasive and reflect the state of both the diseased breast tissue and the patient's immune response. Here, we identified IBC-specific RNA biomarkers by thermostable group II intron reverse transcriptase sequencing (TGIRT-seq), a recently developed comprehensive RNA-seq technology that enables simultaneous profiling of all RNA biotypes from small amounts of starting material. We used these biomarkers to develop novel disease classification models for IBC based on coding and non-coding RNAs from FFPE tumor slices, PBMCs, and plasma. Methods: We obtained biological samples including FFPE, PBMC, and plasma from a cohort of ten patients with IBC and compared them to samples from six patients with non-IBC and sixteen healthy donors using TGIRT-seq technology. Results: TGIRT-seq of FFPE tumor slices identified differentially expressed mRNAs and miRNAs found previously to distinguish IBC from non-IBC tumors, as well as numerous additional differentially expressed mRNAs and small non-coding RNAs characteristic of IBC. Surprisingly, TGIRT-seq revealed that the differentially expressed protein-coding gene transcripts fall into two categories: mature mRNAs with reads confined to exons, and pre-mRNAs-derived transcripts with reads distributed across exons and introns, to our knowledge, a distinction not made previously for any cancer type. Differentially expressed miRNAs included both mature miRNAs and other transcripts of miRNA loci. IBC PBMCs showed a characteristic inflammatory response not seen in PBMCs from non-IBC patients, as well as differentially expressed tRNAs, snoRNAs, and other sncRNAs, while plasma samples, although of variable quality, included coding and non-coding RNAs distinctive of IBC. Classification models using panels consisting of sets of 50 selected biomarkers profiled by TGIRT-seq achieved a high degree of accuracy under cross-validation, with models based on PBMCs and plasma RNAs correlating with those based on tumor RNAs, and models using both coding and non-coding RNA biomarkers outperforming those based on either alone. Conclusions: Our findings are the first to define a distinct IBC profile across three different tissue types and advance TGIRT-seq as a promising method for high-resolution RNA biomarker profiling of both primary tumors and liquid biopsies with potentially broad utility for diagnosing and defining treatment response in IBC and other cancers. COI: Thermostable group II intron reverse transcriptase (TGIRT) enzymes and methods for their use are the subject of patents and patent applications that have been licensed by the University of Texas to InGex, LLC. A.M.L., some former and present members of the Lambowitz laboratory, and the University of Texas are minority equity holders in InGex, and receive royalty payments from the sale of TGIRT enzymes and kits and from sublicensing of intellectual property to other companies. Citation Format: Dennis C. Wylie, Xiaoping Wang, Jun Yao, Hengyi Xu, Toshiaki Iwase, Savitri Krishnamurthy, Naoto T. Ueno, Alan M. Lambowitz. Disease classification modeling of inflammatory breast cancer based on simultaneous profiling of coding and non-coding RNAs in tumor and blood samples by TGIRT-sequencing [abstract]. In: Proceedings of the 2021 San Antonio Breast Cancer Symposium; 2021 Dec 7-10; San Antonio, TX. Philadelphia (PA): AACR; Cancer Res 2022;82(4 Suppl):Abstract nr P5-07-03.
Bacteria encode reverse transcriptases (RTs) of unknown function that are closely related to group II intron-encoded RTs. We found that a Pseudomonas aeruginosa group II intron-like RT (G2L4 RT) with YIDD instead of YADD at its active site functions in DNA repair in its native host and when expressed in Escherichia coli. G2L4 RT has biochemical activities strikingly similar to those of human DNA repair polymerase θ and uses them for translesion DNA synthesis and double-strand break repair (DSBR) via microhomology-mediated end-joining (MMEJ). We also found that a group II intron RT can function similarly in DNA repair, with reciprocal active-site substitutions showing isoleucine favors MMEJ and alanine favors primer extension in both enzymes. These DNA repair functions utilize conserved structural features of non-LTR-retroelement RTs, including human LINE-1 and other eukaryotic non-LTR-retrotransposon RTs, suggesting such enzymes may have inherent ability to function in DSBR in a wide range of organisms.
Reverse transcriptases (RTs) can switch template strands during complementary DNA synthesis, enabling them to join discontinuous nucleic acid sequences. Template switching (TS) plays crucial roles in retroviral replication and recombination, is used for adapter addition in RNA-Seq, and may contribute to retroelement fitness by increasing evolutionary diversity and enabling continuous complementary DNA synthesis on damaged templates. Here, we determined an X-ray crystal structure of a TS complex of a group II intron RT bound simultaneously to an acceptor RNA and donor RNA template-DNA primer heteroduplex with a 1-nt 3'-DNA overhang. The structure showed that the 3' end of the acceptor RNA binds in a pocket formed by an N-terminal extension present in non-long terminal repeat-retroelement RTs and the RT fingertips loop, with the 3' nucleotide of the acceptor base paired to the 1-nt 3'-DNA overhang and its penultimate nucleotide base paired to the incoming dNTP at the RT active site. Analysis of structure-guided mutations identified amino acids that contribute to acceptor RNA binding and a phenylalanine residue near the RT active site that mediates nontemplated nucleotide addition. Mutation of the latter residue decreased multiple sequential template switches in RNA-Seq. Our results provide new in-sights into the mechanisms of TS and nontemplated nucleotide addition by RTs, suggest how these reactions could be improved for RNA-Seq, and reveal common structural features for TS by non-long terminal repeat-retroelement RTs and viral RNA-dependent RNA polymerases.
High-throughput RNA sequencing (RNA-seq) has extraordinarily advanced our understanding of gene expression and disease etiology, and is a powerful tool for the identification of biomarkers in a wide range of organisms. However, most RNA-seq methods rely on retroviral reverse transcriptases (RTs), enzymes that have inherently low fidelity and processivity, to convert RNAs into cDNAs for sequencing. Here, we describe an RNA-seq protocol using Thermostable Group II Intron Reverse Transcriptases (TGIRTs), which have high fidelity, processivity, and strand-displacement activity, as well as a proficient template-switching activity that enables efficient and seamless RNA-seq adapter addition. By combining these activities, TGIRT-seq enables the simultaneous profiling of all RNA biotypes from small amounts of starting material, with superior RNA-seq metrics, and unprecedented ability to sequence structured RNAs. The TGIRT-seq protocol for Illumina sequencing consists of three steps: (i) addition of a 3' RNA-seq adapter, coupled to the initiation of cDNA synthesis at the 3' end of a target RNA, via template switching from a synthetic adapter RNA/DNA starter duplex; (ii) addition of a 5' RNA-seq adapter, by using thermostable 5' App DNA/RNA ligase to ligate an adapter oligonucleotide to the 3' end of the completed cDNA; (iii) minimal PCR amplification, to add capture sites and indices for Illumina sequencing. TGIRT-seq for the Illumina sequencing platform has been used for comprehensive profiling of coding and non-coding RNAs in ribodepleted, chemically fragmented cellular RNAs, and for the analysis of intact (non-chemically fragmented) cellular, extracellular vesicle (EV), and plasma RNAs, where it yields continuous full-length end-to-end sequences of structured small non-coding RNAs (sncRNAs), including tRNAs, snoRNAs, snRNAs, pre-miRNAs, and full-length excised linear intron (FLEXI) RNAs. Graphic abstract: Figure 1.Overview of the TGIRT-seq protocol for Illumina sequencing.Major steps are: (1) Template switching from a synthetic R2 RNA/R2R DNA starter duplex with a 1-nt 3' DNA overhang (a mixture of A, C, G, and T residues, denoted N) that base pairs to the 3' nucleotide of a target RNA, and upon initiating reverse transcription by adding dNTPs, seamlessly links an R2R adapter to the 5' end of the resulting cDNA; (2) Ligation of an R1R adapter to the 3' end of the completed cDNA; and (3) Minimal PCR amplification with primers that add Illumina capture sites (P5 and P7) and barcode sequences (indices 5 and 7). The index 7 barcode is required, while the index 5 barcode is optional, to provide unique dual indices (UDIs).