Abstract Although the impact of single nucleotide variants (SNVs) and changes in transcription and RNA processing are often analyzed separately, a comprehensive analysis facilitates a complete understanding of how cancer gene alterations impact oncogenesis. In traditional short-read RNA sequencing, phasing of alternative exons and cancer variants is lost because the read lengths are much shorter than typical mRNA transcripts (average > 1kb). Here, we show that long-read RNA-seq (lrRNA-seq) can identify full-length transcript isoforms on which variants are expressed, which can be used to more accurately identify the functional impact of oncogenic variants. We developed FLAIR3, which performs an integrated analysis of SNVs, gene fusions, and alternative splicing using lrRNA-seq and predicts functional changes to the amino acid sequence. We performed ONT lrRNA-seq on three osteosarcoma cell lines and PacBio lrRNA-seq on paired normal and tumor tissue from two lung adenocarcinomas. We then used FLAIR3 to identify cancer driver variants and to determine how splicing modulates their expression and function. In the osteosarcoma samples, FLAIR3 revealed alternatively spliced gene fusions in cancer driver genes and TP53 gene fusions with intergenic regions, predicted to cause TP53 truncations. In the lung adenocarcinomas, FLAIR3 revealed isoform-biased expression of oncogenic BRAF V600E. Through an isoform-specific analysis of somatic SNVs in CDKN2A, we found that TP53 loss significantly co-occurs with CDKN2A missense or nonsense variants of the p16 isoform, but not with CDKN2A deep deletion. Damaging variants in the p16 isoform would not have the same damaging effects in p14 isoform, which functions through TP53; therefore, TP53 loss would be necessary to have complete loss of CDKN2A functions. A deep deletion of CDKN2A removes both p16 and p14 isoform function and would not need to have additional TP53 loss. These findings reveal how alternative splicing interacts with and modulates the function of oncogenic variants. Citation Format: Colette Felton, Andrea Galvez, Tanvi Damle, Kevin Levine, Mark Diekhans, Eunice Lopez Fuentes, Taylor Won, Christopher Vollmers, Alejandro Sweet-Cordero, Alice Berger, Angela N. Brooks. Cancer gene variant identification and functional interpretation using long-read RNA sequencing with FLAIR3 [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2026; Part 1 (Regular Abstracts); 2026 Apr 17-22; San Diego, CA. Philadelphia (PA): AACR; Cancer Res 2026;86(7 Suppl):Abstract nr 1501.
Plasmids are ubiquitous tools in molecular biology which are used for a large variety of experiments within academic and commercial labs. Both new and old plasmids have to undergo sequencing-based analysis to determine whether or not they are functional, i.e., contain the correct insert in the correct backbone. While traditional Sanger sequencing based analysis was most often limited to the inserts, new high-throughput sequencing based methods and services can now provide the complete sequence of a plasmid. Currently available methods and services vary in throughput and cost. Here, we adapted the Oxford Nanopore Technologies-based R2C2 sequencing method to - rapidly and at low cost - sequence complete plasmids, either individually or in a pool. We also developed an analysis pipeline, Chopper, that produces full-length plasmid sequences. We tested our workflow with commonly used plasmids we ordered from Addgene and produced highly accurate sequences for each plasmid from both their individual and pooled sequencing runs.
The mammalian circadian clock is an autoregulatory feedback process that is responsible for homeostasis in mouse livers. These circadian processes are well understood at the gene level; however, they are not well understood at the isoform level. To investigate circadian oscillations at the isoform level, we used the nanopore-based R2C2 method to create over 78 million highly accurate, full-length complementary DNA reads for 12 RNA samples extracted from mouse livers collected at 2 h intervals. To generate a circadian mouse liver isoform-level transcriptome, we processed these reads using the Mandalorion tool, which identified and quantified 58 612 isoforms, 1806 of which showed circadian oscillations. We performed detailed analysis on the circadian oscillation of these isoforms, their coding sequences, and transcription start sites and compiled easy-to-access resources for other researchers. This study and its results add a new layer of detail to the quantitative analysis of transcriptomes.
The spatial organization of adaptive immune cells within lymph nodes is critical for understanding immune responses during infection and disease. Here, we introduce AIR-SPACE, an integrative approach that combines high-resolution spatial transcriptomics with paired, high-fidelity long-read sequencing of T and B cell receptors. This method enables the simultaneous analysis of cellular transcriptomes and adaptive immune receptor (AIR) repertoires within their native spatial context. We applied AIR-SPACE to mouse popliteal lymph nodes at five distinct time points after Vaccinia virus footpad infection and constructed a comprehensive map of the developing adaptive immune response. Our analysis revealed heterogeneous activation niches, characterized by Interferon-gamma (IFN-γ) production, during the early stages of infection. At later stages, we delineated sub-anatomical structures within the germinal center (GC) and observed evidence that antibody-producing plasma cells differentiate and exit the GC through the dark zone. Furthermore, by combining clonotype data with spatial lineage tracing, we demonstrate that B cell clones are shared among multiple GCs within the same lymph node, reinforcing the concept of a dynamic, interconnected network of GCs. Overall, our study demonstrates how AIR-SPACE can be used to gain insight into the spatial dynamics of infection responses within lymphoid organs.
Although most cancer variant profiling is done with short-read-based methods, many cancers are driven by structural variants that are difficult to detect with these methods. Our current approach to understanding driver mutations is also limited to single variants and rarely considers the context in which they are expressed. Alterations in the expression and splicing of genes containing variants can both impact their tumorigenicity and allow them to develop resistance to therapies. We present FLAIR3, which generates a custom transcriptome from long-read RNA sequencing and identifies SNVs, insertions, deletions, and gene fusions from alignment to this transcriptome. We show that this improves accuracy above alignment to the genome or annotated transcripts. FLAIR3 then integrates these variants with the transcripts to predict functional changes to the amino acid sequence. We apply this approach to patient-derived osteosarcoma cell lines, a cancer whose pathology is driven by complex structural variation. In these samples, we identify more complex changes to previously identified amplified drivers such as alternative splicing of MYC and alternative splicing of a gene fusion in CCNE1. We also identify a novel set of actionable cancer gene alterations not previously detected by short-read methods. This includes a large deletion in KEAP1 and a number of novel gene fusions, including TP53 and other genes fused with intergenic regions and a complex 4-locus fusion in NOTCH1 not detected by any preexisting tools. We then validated the protein sequence of novel fusions with Quantum-Si protein sequencing. This analysis shows that long-read RNA sequencing can detect novel variants in actionable cancer genes and that integrating splicing and variant alterations provides the most complete picture of gene alterations in cancer. Colette Felton, Andrea Galvez, Leanne Sayles, Christopher Vollmers, Alejandro Sweet-Cordero, Angela Brooks. Gene fusion and variant-aware isoform detection with functional prediction from long-read RNA sequencing [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2025; Part 1 (Regular Abstracts); 2025 Apr 25-30; Chicago, IL. Philadelphia (PA): AACR; Cancer Res 2025;85(8_Suppl_1):Abstract nr 2396.
BackgroundIdentification of COPD disease-causing genes is an important tool for understanding why COPD develops, who is at highest COPD risk, and how new COPD treatments can be developed. Previous COPD genetic studies have identified a highly significant genetic association near nephronectin(NPNT), a gene involved in tissue repair, but the biological mechanisms underlying this association are unknown.MethodsSplicing quantitative trait locus analysis (sQTL) was performed to identify common genetic variants that alter RNA splicing in lung tissues. These lung sQTL signals were compared to COPD genetic association results near theNPNTgene using colocalization analysis to determine whether genetic risk for COPD in this region may act through altered splicing. Long read sequencing characterized COPD-associated splicing events at isoform-level resolution, and in silico protein structural analysis identified likely functional effects of this alternative splicing.ResultsAn established COPD genetic risk variant, rs34712979_A, creates a cryptic splice acceptor site that causes four separate splicing changes imNPNT. The only of these splicing changes that was associated with COPD phenotypes involved a cassette exon (exon 3). Long read RNA sequencing demonstrated that the COPD risk allele causes a shift in isoform usage away from the dominant NPNT Isoform B precursor, which excludes exon 3, to the Isoform A precursor which splices-in exon 3. Alpha-fold protein structural analysis reveals that inclusion of this exon disrupts an EGF-like functional domain in NPNT.ConclusionGenetic variants in the nephronectin (NPNT) gene increase COPD risk by changing RNA splicing ofNPNTin the lung.
Genome-wide association studies (GWASs) have identified multiple genetic loci associated with chronic obstructive pulmonary disease (COPD). Here, we identify SNPs that are associated with alternative splicing (splicing quantitative trait loci [sQTLs]) and gene expression (expression QTLs [eQTLs]) to identify functions for COPD-associated genetic variants. RNA sequencing on whole blood from 3,743 subjects in the COPDGene Study and from lung tissue of 1,241 subjects from the Lung Tissue Research Consortium (LTRC) was analyzed. Associations between all SNPs within 1,000 kb of a gene (cis-) and splice and gene expression quantifications were tested using tensorQTL. We assessed colocalization with COPD-associated SNPs from a published GWAS. After adjustment for multiple statistical testing, we identified 28,110 splice sites corresponding to 3,889 unique genes that were significantly associated with genotype in COPDGene whole blood and 58,258 splice sites corresponding to 10,307 unique genes associated with genotype in LTRC lung tissue. To determine what proportion of COPD-associated SNPs were associated with transcriptional splicing, we performed colocalization analysis between COPD GWAS and sQTL data and found that 38 genomic windows, corresponding to 33 COPD GWAS loci, had evidence of colocalization between QTLs and COPD. The top five colocalizations between COPD and lung sQTLs include Nephronectin (NPNT), F box protein 38 (FBXO38), Hedgehog interacting protein (HHIP), Netrin 4 (NTN4), and Betacellulin (BTC). Overall, a total of 38 COPD GWAS loci contain evidence of sQTLs, suggesting that analysis of sQTLs in whole blood and lung tissue can provide insights into disease mechanisms.
RationaleCigarette smoking (CS) impairs B-cell function and antibody production, increasing infection risk. The impact of e-cigarette use ('vaping') and combined CS and vaping ('dual-use') on B-cell activity is unclear.ObjectiveTo examine B-cell receptor sequencing (BCR-seq) profiles associated with CS, vaping, and dual-use.MethodsBCR-seq was performed on blood RNA samples from 234 participants in the COPDGene study. We assessed multivariable associations of B-cell function measures (immunoglobulin heavy chain (IGH) subclass expression and usage, class-switching, V allele usage, and clonal expansion) with CS, vaping, and dual-use. We adjusted for multiple comparisons using the Benjamini-Hochberg method, identifying significant associations at 5% FDR and suggestive associations at 10% FDR.ResultsAmong 234 non-Hispanic white (NHW) and African American (AA) participants, CS and dual-use were significantly positively associated with increased secretory IgA production, with dual-use showing the strongest associations. Dual-use was positively associated with class switching and B-cell clonal expansion, indicating increased B-cell activation, with similar trends in those only smoking or only vaping. The IGHV5-51*01 allele was increased in dual users.ConclusionsCS and vaping additively enhance B-cell activation, most notably in dual-users. CS and vaping are significantly associated to multiple alterations in B-cell function including increased class switching, clonal expansion, and a shift towards IgA-producing cell populations. These changes could be relevant to response to infection and vaccinations and merit further study.
The spatial organization of adaptive immune cells within lymph nodes is essential for understanding immune responses during infection and disease. Here, we sought to investigate the spatial and temporal changes to the adaptive immune receptor repertoire (AIRR) in the draining lymph node after footpad infection with Vaccinia virus in mice. We developed a novel method that combines high-resolution spatial transcriptomics (Slide-seq) with high-fidelity long-read adaptive immune receptor sequencing from tissue sections. This integration enables simultaneous analysis of whole transcriptomes and the AIRR in their spatial context. Applying our method to the model, we demonstrated its capability to map the spatial distribution and capture temporal dynamics of the AIRR at 3, 7, 10, 14, and 21 days post-infection. Our approach revealed heterogenous niches of activation from Interferon-gamma (IFN-γ) during early stages of infection. We also observed sub-anatomical structures within the germinal center (GC), providing evidence that antibody-producing plasma cells differentiate and exit the GC from the dark zone. Additionally, we traced the spatial lineage trajectory of B cell clones and found evidence to suggest that their maturation occurs across multiple GCs. Thus, our methodology offers valuable insights into the spatial dynamics of immune responses, presenting a powerful tool for studying the immune system and disease pathogenesis. Technological Innovations in Immunology (TECH)
BACKGROUND:Identification of COPD disease-causing genes is an important tool for understanding why COPD develops, who is at highest COPD risk and how new COPD treatments can be developed. Previous COPD genetic studies have identified a highly significant genetic association near NPNT (nephronectin), a gene involved in tissue repair, but the biological mechanisms underlying this association are unknown. METHODS:Splicing quantitative trait locus (sQTL) analysis was performed to identify common genetic variants that alter RNA splicing in lung tissues. These lung sQTL signals were compared to COPD genetic association results near the NPNT gene using colocalisation analysis to determine whether genetic risk for COPD in this region may act through altered splicing. Long-read sequencing characterised COPD-associated splicing events at isoform-level resolution and in silico protein structural analysis identified likely functional effects of this alternative splicing. RESULTS:An established COPD genetic risk variant, rs34712979-A, creates a cryptic splice acceptor site that causes four separate splicing changes in NPNT. The only one of these splicing changes that was associated with COPD phenotypes involved a cassette exon (exon 3). Long-read RNA sequencing demonstrated that the COPD risk allele causes a shift in isoform usage away from the dominant NPNT isoform B precursor, which excludes exon 3, to the isoform A precursor, which splices-in exon 3. AlphaFold protein structural analysis reveals that inclusion of this exon disrupts an epidermal growth factor-like functional domain in NPNT. CONCLUSION:Genetic variants in the NPNT gene increase COPD risk by changing RNA splicing of NPNT in the lung.
Plasmids are ubiquitous tools in molecular biology which are used for a large variety of experiments within academic and commercial labs. Both new and old plasmids have to undergo sequencing-based analysis to determine whether or not they are functional, i.e. contain the correct insert in the correct backbone. While traditional Sanger sequencing based analysis was most often limited to the inserts, new high-throughput sequencing based methods and services can now provide the complete sequence of a plasmid. Currently available methods and services vary in throughput and cost. Here, we adapted the Oxford Nanopore Technologies-based R2C2 sequencing method to - rapidly and at low cost - sequence complete plasmids, either individually or in a pool. We also developed an analysis pipeline, Chopper, that produces full-length plasmid sequences. We tested our workflow with commonly used plasmids we ordered from Addgene and produced highly accurate sequences for each plasmid from both their individual and pooled sequencing runs. ### Competing Interest Statement CV has filed a patent application on the R2C2 method
The Long-read RNA-Seq Genome Annotation Assessment Project (LRGASP) Consortium was formed to evaluate the effectiveness of long-read approaches for transcriptome analysis. The consortium generated over 427 million long-read sequences from cDNA and direct RNA datasets, encompassing human, mouse, and manatee species, using different protocols and sequencing platforms. These data were utilized by developers to address challenges in transcript isoform detection and quantification, as well as de novo transcript isoform identification. The study revealed that libraries with longer, more accurate sequences produce more accurate transcripts than those with increased read depth, whereas greater read depth improved quantification accuracy. In well-annotated genomes, tools based on reference sequences demonstrated the best performance. When aiming to detect rare and novel transcripts or when using reference-free approaches, incorporating additional orthogonal data and replicate samples are advised. This collaborative study offers a benchmark for current practices and provides direction for future method development in transcriptome analysis.
Generating an accurate and complete genome annotation for an organism is complex because the cells within each tissue can express a unique set of transcript isoforms from a unique set of genes. A comprehensive genome annotation should contain information on what tissues express what transcript isoforms at what level. This tissue-level isoform information can then inform a wide range of research questions as well as experiment designs. Long-read sequencing technology combined with advanced full-length cDNA library preparation methods has now achieved throughput and accuracy where generating these types of annotations is achievable. Here, we show this by generating a genome annotation of the mouse (Mus musculus). We used the nanopore-based R2C2 long-read sequencing method to generate 64 million highly accurate full-length cDNA consensus reads—averaging 5.4 million reads per tissue for a dozen tissues. Using the Mandalorion tool, we processed these reads to generate the Tissue-level Atlas of Mouse Isoforms which is available as a trackhub for the UCSC Genome Browser and contains at least one full-length isoform for the vast majority of expressed genes in each tissue.
BackgroundRNA-seq has brought forth significant discoveries regarding aberrations in RNA processing, implicating these RNA variants in a variety of diseases. Aberrant splicing and single nucleotide variants (SNVs) in RNA have been demonstrated to alter transcript stability, localization, and function. In particular, the upregulation of ADAR, an enzyme that mediates adenosine-to-inosine editing, has been previously linked to an increase in the invasiveness of lung adenocarcinoma cells and associated with splicing regulation. Despite the functional importance of studying splicing and SNVs, the use of short-read RNA-seq has limited the community's ability to interrogate both forms of RNA variation simultaneously.ResultsWe employ long-read sequencing technology to obtain full-length transcript sequences, elucidating cis-effects of variants on splicing changes at a single molecule level. We develop a computational workflow that augments FLAIR, a tool that calls isoform models expressed in long-read data, to integrate RNA variant calls with the associated isoforms that bear them. We generate nanopore data with high sequence accuracy from H1975 lung adenocarcinoma cells with and without knockdown of ADAR. We apply our workflow to identify key inosine isoform associations to help clarify the prominence of ADAR in tumorigenesis.ConclusionsUltimately, we find that a long-read approach provides valuable insight toward characterizing the relationship between RNA variants and splicing patterns.
The sequencing of PCR amplicons is a core application of high-throughput sequencing technology. Using unique molecular identifiers (UMIs), individual amplified molecules can be sequenced to very high accuracy on an Illumina sequencer. However, Illumina sequencers have limited read length and are therefore restricted to sequencing amplicons shorter than 600bp unless using inefficient synthetic long-read approaches. Native long-read sequencers from Pacific Biosciences and Oxford Nanopore Technologies can, using consensus read approaches, match or exceed Illumina quality while achieving much longer read lengths. Using a circularization-based concatemeric consensus sequencing approach (R2C2) paired with UMIs (R2C2+UMI) we show that we can sequence ∼550nt antibody heavy-chain (IGH) and ∼1500nt 16S amplicons at accuracies up to and exceeding Q50 (<1 error in 100,0000 sequenced bases), which exceeds accuracies of UMI-supported Illumina paired sequencing as well as synthetic long-read approaches.
Alternative splicing (AS) alters the cis-regulatory landscape of mRNA isoforms, leading to transcripts with distinct localization, stability, and translational efficiency. To rigorously investigate mRNA isoform-specific ribosome association, we generated subcellular fractionation and sequencing (Frac-seq) libraries using both conventional short reads and long reads from human embryonic stem cells (ESCs) and neural progenitor cells (NPCs) derived from the same ESCs. We performed de novo transcriptome assembly from high-confidence long reads from cytosolic, monosomal, light, and heavy polyribosomal fractions and quantified their abundance using short reads from their respective subcellular fractions. Thousands of transcripts in each cell type exhibited association with particular subcellular fractions relative to the cytosol. Of the multi-isoform genes, 27% and 19% exhibited significant differential isoform sedimentation in ESCs and NPCs, respectively. Alternative promoter usage and internal exon skipping accounted for the majority of differences between isoforms from the same gene. Random forest classifiers implicated coding sequence (CDS) and untranslated region (UTR) lengths as important determinants of isoform-specific sedimentation profiles, and motif analyses reveal potential cell type-specific and subcellular fraction-associated RNA-binding protein signatures. Taken together, our data demonstrate that alternative mRNA processing within the CDS and UTRs impacts the translational control of mRNA isoforms during stem cell differentiation, and highlight the utility of using a novel long-read sequencing-based method to study translational control.
The Javan gibbon, Hylobates moloch, is an endangered gibbon species restricted to the forest remnants of western and central Java, Indonesia, and one of the rarest of the Hylobatidae family. Hylobatids consist of 4 genera (Holoock, Hylobates, Symphalangus, and Nomascus) that are characterized by different numbers of chromosomes, ranging from 38 to 52. The underlying cause of this karyotype plasticity is not entirely understood, at least in part, due to the limited availability of genomic data. Here we present the first scaffold-level assembly for H. moloch using a combination of whole-genome Illumina short reads, 10X Chromium linked reads, PacBio, and Oxford Nanopore long reads and proximity-ligation data. This Hylobates genome represents a valuable new resource for comparative genomics studies in primates.
Promoters and the noncoding sequences that drive their function are fundamental aspects of genes that are critical to their regulation. The transcription preinitiation complex binds and assembles on promoters where it facilitates transcription. The transcription start site (TSS) is located downstream of the promoter sequence and is defined as the location in the genome where polymerase begins transcribing DNA into RNA. Knowing the location of TSSs is useful for annotation of genes, identification of non-coding sequences important to gene regulation, detection of alternative TSSs, and understanding of 5 ' UTR content. Several existing techniques make it possible to accurately identify TSSs, but are often difficult to perform experimentally, require large amounts of input RNA, or are unable to identify a large number of TSSs from a single sample. Many of these protocols take advantage of template switching reverse transcriptases (TSRTs), which reliably place an adaptor at the 5 ' end of a first strand synthesis of cDNA. Here, we introduce a protocol that exploits TSRT activity combined with rolling circle amplification to identify TSSs with several unique advantages over existing methods. Sequence adaptors are placed on the 5 ' and 3 ' end of the full-length cDNA copy of a transcript. A splint compatible with those adaptors is then used to circularize the full-length cDNA. Linear DNA containing concatemers of the cDNA are generated using rolling circle amplification, and a sequencing library is formed by fragmenting the concatemers. This protocol is straightforward to execute, requiring limited bench time with relatively stable reagents. Using extremely low amounts of RNA input, this protocol produces large numbers of accurate, deduplicated TSSs genome wide. (c) 2023 The Authors. Current Protocols published by Wiley Periodicals LLC.Basic Protocol 1: Splint generationBasic Protocol 2: RNA extractionBasic Protocol 3: cDNA synthesisBasic Protocol 4: cDNA circularization and amplificationBasic Protocol 5: Library generation
In this manuscript, we introduce and benchmark Mandalorion v4.1 for the identification and quantification of full-length transcriptome sequencing reads. It further improves upon the already strong performance of Mandalorion v3.6 used in the LRGASP consortium challenge. By processing real and simulated data, we show three main features of Mandalorion: first, Mandalorion-based isoform identification has very high precision and maintains high recall even in the absence of any genome annotation. Second, isoform read counts as quantified by Mandalorion show a high correlation with simulated read counts. Third, isoforms identified by Mandalorion closely reflect the full-length transcriptome sequencing data sets they are based on.
Polar bears are susceptible to climate warming because of their dependence on sea ice, which is declining rapidly. We present the first evidence for a genetically distinct and functionally isolated group of polar bears in Southeast Greenland. These bears occupy sea-ice conditions resembling those projected for the High Arctic in the late 21st century, with an annual ice-free period that is >100 days longer than the estimated fasting threshold for the species. Whereas polar bears in most of the Arctic depend on annual sea ice to catch seals, Southeast Greenland bears have a year-round hunting platform in the form of freshwater glacial mélange. This suggests that marine-terminating glaciers, although of limited availability, may serve as previously unrecognized climate refugia. Conservation of Southeast Greenland polar bears, which meet criteria for recognition as the world's 20th polar bear subpopulation, is necessary to preserve the genetic diversity and evolutionary potential of the species.