Comprehensive genomic analysis is essential for advancing our understanding of human genetics and disease. However, short-read sequencing technologies are inherently limited in their ability to resolve highly repetitive, structurally complex, and low-mappability genomic regions, previously coined as “dark” regions. Long-read sequencing technologies, such as PacBio and Oxford Nanopore Technologies (ONT), offer improved resolution of these regions, yet they are not perfect. With the advent of the new Telomere-to-Telomere (T2T) CHM13 reference genome, exploring its effect on dark regions is prudent. In this study, we systematically analyze dark regions across four human genome references—HG19, HG38 (with and without alternate contigs), and CHM13—using both short- and long-read sequencing data. We found that dark regions increase as the reference becomes more complete, especially dark-by-MAPQ regions, but that long-read sequencing significantly reduces the number of dark regions in the genome, particularly within gene bodies. However, we identify potential alignment challenges in long-read data, such as centromeric regions. These findings highlight the importance of both reference genome selection and sequencing technology choice in achieving a truly comprehensive genomic analysis.
The Apolipoprotein E (APOE) e4 and e2 alleles are respectively the most risk increasing and risk decreasing, common genetic risk factors for Alzheimer's disease (AD)1,2. They strongly affect Aβ burden in the brain parenchyma1, a core hallmark of AD, but also at the level of the brain vasculature, i.e. cerebral amyloid angiopathy (CAA)1,3, which in turn relates to increased risk for amyloid-related imaging abnormalities (ARIA) in APOE*4 carriers when receiving anti-Aβ antibody treatments4. This makes APOE a highly pursued AD drug target. A crucial question in the field is whether it would be beneficial to either increase or decrease APOE (particularly APOE*4) levels5. The answer from rodent work appears to converge on "decreasing APOE levels"5-7, with initial human studies supporting this5,8,9. Human genetic evidence however remains scarce and new insights are crucially needed to support clinical translation. Shade et al. 2024 conducted the largest to date genome-wide association study (GWAS) of various neuropathological traits, identifying a variant protective of CAA in the APOE locus independent of APOE*4 and APOE*2 genotypes10. Downstream analyses suggested this signal links to the nearby APOC2 gene through local effects on methylation. We applaud the authors on their timely, relevant, and well-conducted study. Here, we extend on these findings, highlighting there is compelling evidence that their genetic signal for reduced CAA relates to an effect on reduced microglial APOE expression, which would importantly support the evidence in favor of "decreasing APOE levels" and further herald this promising therapeutic avenue, not just for AD, but also for CAA. We additionally provide complimentary results regarding this locus' association with CAA and AD risk from analyses that we conducted parallel to Shade et al. 2024.
Recent advances in large language models have extended to genomic applications, yet model robustness relative to context is unclear. Here, we demonstrate two intrinsic biases (input sequence length and nucleotide position) affecting SegmentNT results, a model included with the Nucleotide Transformer that provides nucleotide-level predictions of biological features. We demonstrate that nucleotide position within the input sequence (beginning, middle, or end) alters the nature of SegmentNT's raw prediction probabilities, which can be standardized to improve prediction consistency. While longer input sequence length improves model performance, diminishing returns suggest a surprisingly small input length of ∼3072 nucleotides might be sufficient for many applications. We further identify a 24-nucleotide periodic oscillation in SegmentNT's prediction probabilities, revealing an intrinsic bias potentially linked to the model's training tokenization (6-mers) and architecture. We identify potential approaches to account for these biases and provide generalizable insights for utilizing nucleotide-resolution functional prediction models.
Long-read single-cell RNA sequencing provides an opportunity to understand human health and disease at a level difficult to resolve with bulk or short-read methods. This approach enables isoform-level investigation of cellular diversity and disease mechanisms and definition of cell-types, rather than using genes alone. Using a modified, microfluidic-free PIPseq workflow and computational pipeline adapted for Oxford Nanopore long-read sequencing, we generated the largest long-read single-cell dataset of human peripheral blood mononuclear cells (PBMCs) from a single donor to date, the first with sufficient cell numbers to detect megakaryocytes. This study profiled isoform usage across immune cells, integrating marker expression and isoform discovery. We identified 126 novel isoforms from known and new genes, several with distinct cell-type-specific patterns, and characterized marker gene isoform expression across cell-types. Non-canonical protein-coding variants of GZMB and CD3G were enriched in unexpected cell-types, including megakaryocytes and monocyte-derived populations. We also discovered novel transcripts from CMC1 and LYAR with cell-type-specific signatures that were also the predominantly expressed transcript within the gene. This study expands the versatility of long-read single-cell studies to not only relay changes in isoform signatures, but to position them within the functional context of the biology they impact. These results demonstrate the power of long-read single-cell sequencing for mapping the isoform landscape—the isonome—across tissues and disease contexts.
We systematically reviewed and meta-analyzed bulk RNA sequencing (RNAseq) studies comparing Alzheimer's disease (AD) patients to controls in human brain tissue. We searched PubMed, Web of Science, and Scopus for human brain bulk RNAseq studies, excluding re-analyses and studies limited to small RNAs or gene panels. We developed 10 criteria for quality assessment and performed a meta-analysis on three high-quality datasets. Of 3266 records, 24 qualified for the systematic review, and one study with three datasets qualified for the meta-analysis. The meta-analysis identified 571 differentially expressed genes (DEGs) in the temporal lobe and 189 in the frontal lobe, including CLU and GFAP. Pathway analysis suggested reactivation of developmental processes in the adult AD brain. Limited data availability constrained the meta-analysis. These findings underscore the need for rigorous methods in AD transcriptomic research to better identify transcriptomic changes and advance biomarker and therapeutic development. This review is registered in PROSPERO (CRD42023466522). HIGHLIGHTS:Comprehensive review: Conducted the first systematic review and meta-analysis of bulk RNA sequencing (RNAseq) studies comparing Alzheimer's disease (AD) patients with non-demented controls using primary human brain tissue. KEY FINDINGS:Identified 571 differentially expressed genes (DEGs) in the temporal lobe and 189 in the frontal lobe of patients with AD, revealing potential therapeutic targets. Pathway discovery: Highlighted key overlapping pathways such as "tube morphogenesis" and "neuroactive ligand-receptor interaction" that may play critical roles in AD. QUALITY ASSESSMENT:Emphasized the importance of methodological rigor in transcriptomic studies, including quality assessment tools to guide future research in AD. STUDY LIMITATION:Acknowledged limited access to complete data tables and lack of diversity in existing datasets, which constrained some of the analysis.
Alternative splicing generates multiple RNA isoforms from a single gene, enriching genetic diversity and impacting gene function. Effective visualization of these isoforms and their expression patterns is crucial but challenging due to limitations in existing tools. Traditional genome browsers lack programmability, while other tools offer limited customization, produce static plots, or cannot simultaneously display structures and expression levels. RNApysoforms was developed to address these gaps by providing a Python-based package that enables concurrent visualization of RNA isoform structures and expression data. Leveraging plotly and polars libraries, it offers an interactive, customizable, and faster-rendering framework suitable for web applications, enhancing the analysis and dissemination of RNA isoform research. RNApysoforms is a Python package available at (https://github.com/UK-SBCoA-EbbertLab/RNApysoforms) and (https://zenodo.org/records/14941190) via an open-source MIT license. It can be easily installed using the pip package installer for Python. Thorough documentation and usage vignettes are available at: https://rna-pysoforms.readthedocs.io/en/latest/.
Even though alternative RNA splicing was discovered nearly 50 years ago (1977), we still understand very little about most isoforms arising from a single gene, including in which tissues they are expressed and if their functions differ. Human gene annotations suggest remarkable transcriptional complexity, with approximately 252,798 distinct RNA isoform annotations from 62,710 gene bodies (Ensembl v109; 2023), emphasizing the need to understand their biological effects. For example, 256 gene bodies have ≥ 50 annotated isoforms, and 30 have ≥ 100, where one protein-coding gene (MAPK10) even has 192 distinct RNA isoform annotations. Whether such isoform diversity results from biological redundancy or spurious alternative splicing (i.e., noise), or whether individual isoforms have specialized functions (even if subtle) remains a mystery for most genes. Three recent studies demonstrated that long-read RNAseq enables improved RNA isoform quantification for essentially any tissue, cell type, or biological condition (e.g., disease, development, aging, etc.), making it possible to better assess individual isoform expression and function. While each study provided important discoveries related to RNA isoform diversity, deeper exploration is needed. We sought to quantify and characterize real isoform usage across tissues (compared to annotations). We used long-read RNAseq data from 58 GTEx samples across nine tissues (three brain, two heart, muscle, lung, liver, and cultured fibroblasts) generated by Glinos et al. and found considerable isoform diversity within and across tissues. Cerebellar hemisphere was the most transcriptionally complex tissue (22,522 distinct isoforms; 3,726 unique); liver was the least diverse (12,435 distinct isoforms; 1,039 unique). We highlight gene clusters exhibiting high tissue-specific isoform diversity per tissue (e.g., TPM1 expresses 19 in heart’s atrial appendage). We also validated 447 of the 700 new isoforms discovered by Aguzzoli-Heberle et al. and found that 88 were expressed in all nine tissues, while 58 were specific to a single tissue. This study represents a broad bioinformatic survey of the RNA isoform landscape, demonstrating isoform diversity across nine tissues and emphasizes the need for further verification, validation, and functional annotation research to better understand how individual isoforms from a single gene body contribute to human health and disease.
BACKGROUND:An accurate genome annotation is essential in many contexts, including RNA sequencing studies. Annotations include known genes and isoforms, detailing their location (chromosome, start, and end) and coding sequence, among other important metadata. RESULTS:We characterized changes in human Ensembl annotations from 2014 to 2023 and the important gains in our biological understanding in recent years. While generally gene and isoform annotations increased (2014: 58,812 genes ; 2023: 62,710), some years dropped (e.g., 2016). A similar pattern exists for the gene and isoform biotypes; both 2015 (19,825) and 2017 (19,828) have fewer genes annotated as protein-coding than 2014 (19,953) and 2016 (19,961)- 2023 has the most (20,048). PCBP1-AS1 had the most annotated isoforms (296). We quantified expression for isoforms that were new between 2019 and 2023 across nine GTEx tissues (58 samples) to demonstrate our significant gains in understanding recently. We saw 2,054 of these 'new' isoforms expressed in cerebellar hemisphere (594 in liver). For many genes, we saw that the relative expression of the 'new' isoforms was much greater than the previously known isoforms. CONCLUSIONS:This study demonstrates the importance of an accurate genome annotation to truly understand the underlying complexity of biology that is often oversimplified by ignoring transcriptional complexity.
Background: The synonymous variant NC_000007.14:g.100373690T>C (rs2405442:T>C) in the Paired Immunoglobulin-like Type 2 Receptor Alpha (PILRA) gene was previously associated with decreased risk for Alzheimer’s disease (AD) in genome-wide association studies, but its biological impact is largely unknown. Objective: We hypothesized that rs2405442:T>C decreases mRNA and protein levels by destroying a ramp of slowly translated codons at the 5′ end of PILRA. Methods: We assessed rs2405442:T>C predicted effects on PILRA through quantitative polymerase chain reactions (qPCRs) and enzyme-linked immunosorbent assays (ELISAs) using Chinese hamster ovary (CHO) cells. RESULTS: Both mRNA (p = 1.9184 × 10−13) and protein (p = 0.01296) levels significantly decreased in the mutant versus the wildtype in the direction that we predicted based on the destruction of a ramp sequence. Conclusions: We show that rs2405442:T>C alone directly impacts PILRA mRNA and protein expression, and ramp sequences may play a role in regulating AD-associated genes without modifying the protein product.
Comprehensive genomic analysis is essential for advancing our understanding of human genetics and disease. However, short-read sequencing technologies are inherently limited in their ability to resolve highly repetitive, structurally complex, and low-mappability genomic regions, previously coined as "dark" regions. Long-read sequencing technologies, such as PacBio and Oxford Nanopore Technologies (ONT), offer improved resolution of these regions, yet they are not perfect. With the advent of the new Telomere-to-Telomere (T2T) CHM13 reference genome, exploring its effect on dark regions is prudent. In this study, we systematically analyze dark regions across four human genome references-HG19, HG38 (with and without alternate contigs), and CHM13-using both short- and long-read sequencing data. We found that dark regions increase as the reference becomes more complete, especially dark-by-MAPQ regions, but that long-read sequencing significantly reduces the number of dark regions in the genome, particularly within gene bodies. However, we identify potential alignment challenges in long-read data, such as centromeric regions. These findings highlight the importance of both reference genome selection and sequencing technology choice in achieving a truly comprehensive genomic analysis.
Genome-wide association studies (GWAS) have identified >80 Alzheimer's disease and related dementias (ADRD)-associated genetic loci. However, the clinical outcomes used in most previous studies belie the complex nature of underlying neuropathologies. Here we performed GWAS on 11 ADRD-related neuropathology endophenotypes with participants drawn from the following three sources: the National Alzheimer's Coordinating Center, the Religious Orders Study and Rush Memory and Aging Project, and the Adult Changes in Thought study (n = 7,804 total autopsied participants). We identified eight independent significantly associated loci, of which four were new (COL4A1, PIK3R5, LZTS1 and APOC2). Separately testing known ADRD loci, 19 loci were significantly associated with at least one neuropathology after false-discovery rate adjustment. Genetic colocalization analyses identified pleiotropic effects and quantitative trait loci. Methylation in the cerebral cortex at two sites near APOC2 was associated with cerebral amyloid angiopathy. Studies that include neuropathology endophenotypes are an important step in understanding the mechanisms underlying genetic ADRD risk.
Abstract Background The gene C9orf72 harbors a non-coding hexanucleotide repeat expansion known to cause amyotrophic lateral sclerosis and frontotemporal dementia. While previous studies have estimated the length of this repeat expansion in multiple tissues, technological limitations have impeded researchers from exploring additional features, such as methylation levels. Methods We aimed to characterize C9orf72 repeat expansions using a targeted, amplification-free long-read sequencing method. Our primary goal was to determine the presence and subsequent quantification of observed methylation in the C9orf72 repeat expansion. In addition, we measured the repeat length and purity of the expansion. To do this, we sequenced DNA extracted from blood for 27 individuals with an expanded C9orf72 repeat. Results For these individuals, we obtained a total of 7,765 on-target reads, including 1,612 fully covering the expanded allele. Our in-depth analysis revealed that the expansion itself is methylated, with great variability in total methylation levels observed, as represented by the proportion of methylated CpGs (13 to 66%). Interestingly, we demonstrated that the expanded allele is more highly methylated than the wild-type allele (P-Value = 2.76E-05) and that increased methylation levels are observed in longer repeat expansions (P-Value = 1.18E-04). Furthermore, methylation levels correlate with age at collection (P-Value = 3.25E-04) as well as age at disease onset (P-Value = 0.020). Additionally, we detected repeat lengths up to 4,088 repeats (~ 25 kb) and found that the expansion contains few interruptions in the blood. Conclusions Taken together, our study demonstrates robust ability to quantify methylation of the expanded C9orf72 repeat, capturing differences between individuals harboring this expansion and revealing clinical associations.
Background Alzheimer’s disease is highly heritable and exhibits neuropathological hallmarks of neurofibrillary tau tangles and neuritic amyloid plaques. Previous genome-wide association studies (GWAS) have identified over 70 genomic risk loci of clinically diagnosed Alzheimer’s disease. However, upon autopsy, many Alzheimer’s disease patients have multiple comorbid neuropathologies that may have independent or pleiotropic genomic risk factors. Autopsy data combined with GWAS provides the opportunity to study the genetic risk factors of individual neuropathologies. Methods We studied the genome-wide risk factors of eleven Alzheimer’s disease-related neuropathology endophenotypes. We used four sources of neuropathological data: National Alzheimer’s Coordinating Center, Religious Orders Study and Rush Memory and Aging Project, Adult Changes in Thought study, and Alzheimer’s Disease Neuroimaging Initiative. We used generalized linear mixed models to identify risk loci, followed by Bayesian colocalization analyses to identify potential functional mechanisms by which genetic loci influence neuropathology risk. Results We identified two novel loci associated with neuropathology: one PIK3R5 locus (lead variant rs72807981) with neurofibrillary pathology, and one COL4A1 locus (lead variant rs2000660) with cerebral atherosclerosis. We also confirmed associations between known Alzheimer’s genes and multiple neuropathology endophenotypes, including APOE (neurofibrillary tangles, neuritic plaques, diffuse plaques, cerebral amyloid angiopathy, and TDP-43 pathology); BIN1 (neurofibrillary tangles and neuritic plaques); and TMEM106B (TDP-43 pathology and hippocampal sclerosis). After adjusting for APOE genotype, we identified a locus near APOC2 (lead variant rs4803778) associated with cerebral amyloid angiopathy that influences DNA methylation at nearby CpG sites in the cerebral cortex. Conclusions rs2000660 is in strong linkage disequilibrium with a synonymous coding variant (rs650724) of COL4A1 , providing a candidate functional variant. Two CpG sites affected by the cerebral amyloid angiopathy-associated APOC2 locus were previously associated with dementia in an independent cohort, suggesting that the effect of this locus on disease may be mediated by DNA methylation. BIN1 is associated with neurofibrillary tangles and neuritic plaques but not with amyloid pathology. TMEM106B is associated with hippocampal sclerosis and TDP-43 pathology but not the canonical Alzheimer’s disease pathologies. These findings provide insights into known Alzheimer’s disease risk loci by refining the pathways affected by these risk genes.
Objective:To systematically review and meta-analyze bulk RNA sequencing studies comparing Alzheimer's disease (AD) patients with controls in human brain tissue, assessing study quality and identifying key genes and pathways. Methods:We searched PubMed, Web of Science, and Scopus on September 23, 2023, for studies using bulk RNAseq on primary human brain tissue from AD patients and controls. Excluded were non-primary tissue, re-analyses without new data, limited RNA types and gene panels. Quality was assessed with a 10-category tool. Meta-analysis used high-quality datasets. Results:From 3,266 records, 24 studies met criteria. Meta-analysis found 571 differentially expressed genes (DEGs) in temporal lobe and 189 in frontal lobe; overlapping pathways included "Tube morphogenesis" and "Neuroactive ligand-receptor interaction." Limitations:Study heterogeneity and limited data tables constrained the review. Conclusions:Rigorous methods are vital in AD transcriptomic studies. Findings enhance understanding of transcriptomic changes, aiding biomarker and therapeutic development. Registration:PROSPERO (CRD42023466522).
Determining whether the RNA isoforms from medically relevant genes have distinct functions could facilitate direct targeting of RNA isoforms for disease treatment. Here, as a step toward this goal for neurological diseases, we sequenced 12 postmortem, aged human frontal cortices (6 Alzheimer disease cases and 6 controls; 50% female) using one Oxford Nanopore PromethION flow cell per sample. We identified 1,917 medically relevant genes expressing multiple isoforms in the frontal cortex where 1,018 had multiple isoforms with different protein-coding sequences. Of these 1,018 genes, 57 are implicated in brain-related diseases including major depression, schizophrenia, Parkinson's disease and Alzheimer disease. Our study also uncovered 53 new RNA isoforms in medically relevant genes, including several where the new isoform was one of the most highly expressed for that gene. We also reported on five mitochondrially encoded, spliced RNA isoforms. We found 99 differentially expressed RNA isoforms between cases with Alzheimer disease and controls. The landscape of RNA isoforms in human cortex is revealed by deep long-read RNA sequencing.
Even though alternative RNA splicing was discovered nearly 50 years ago (1977), we still understand very little about most isoforms arising from a single gene, including in which tissues they are expressed and if their functions differ. Human gene annotations suggest remarkable transcriptional complexity, with approximately 252,798 distinct RNA isoform annotations from 62,710 gene bodies (Ensembl v109; 2023), emphasizing the need to understand their biological effects. For example, 256 gene bodies have ≥50 annotated isoforms and 30 have ≥100, where one protein-coding gene (MAPK10) even has 192 distinct RNA isoform annotations. Whether such isoform diversity results from biological redundancy or spurious alternative splicing (i.e., noise), or whether individual isoforms have specialized functions (even if subtle) remains a mystery for most genes. Recent studies by Aguzzoli-Heberle et al., Leung et al., and Glinos et al. demonstrated long-read RNAseq enables improved RNA isoform quantification for essentially any tissue, cell type, or biological condition (e.g., disease, development, aging, etc.), making it possible to better assess individual isoform expression and function. While each study provided important discoveries related to RNA isoform diversity, deeper exploration is needed. We sought to quantify and characterize real isoform usage across tissues (compared to annotations). We used long-read RNAseq data from 58 GTEx samples across nine tissues (three brain, two heart, muscle, lung, liver, and cultured fibroblasts) generated by Glinos et al. and found considerable isoform diversity within and across tissues. Cerebellar hemisphere was the most transcriptionally complex tissue (22,522 distinct isoforms; 3,726 unique); liver was least diverse (12,435 distinct isoforms; 1,039 unique). We highlight gene clusters exhibiting high tissue-specific isoform diversity per tissue (e.g., TPM1 expresses 19 in heart's atrial appendage). We also validated 447 of the 700 new isoforms discovered by Aguzzoli-Heberle et al. and found that 88 were expressed in all nine tissues, while 58 were specific to a single tissue. This study represents a broad survey of the RNA isoform landscape, demonstrating isoform diversity across nine tissues and emphasizes the need to better understand how individual isoforms from a single gene body contribute to human health and disease.
Due to alternative splicing, human protein-coding genes average over eight RNA isoforms, resulting in nearly four distinct protein coding sequences per gene. Long-read RNAseq (IsoSeq) enables more accurate quantification of isoforms, shedding light on their specific roles. To assess the medical relevance of measuring RNA isoform expression, we sequenced 12 aged human frontal cortices (6 Alzheimer's disease cases and 6 controls; 50% female) using one Oxford Nanopore PromethION flow cell per sample. Our study uncovered 53 new high-confidence RNA isoforms in medically relevant genes, including several where the new isoform was one of the most highly expressed for that gene. Specific examples include WDR4 (61%; microcephaly), MYL3 (44%; hypertrophic cardiomyopathy), and MTHFS (25%; major depression, schizophrenia, bipolar disorder). Other notable genes with new high-confidence isoforms include CPLX2 (10%; schizophrenia, epilepsy) and MAOB (9%; targeted for Parkinson's disease treatment). We identified 1,917 medically relevant genes expressing multiple isoforms in human frontal cortex, where 1,018 had multiple isoforms with different protein coding sequences, demonstrating the need to better understand how individual isoforms from a single gene body are involved in human health and disease, if at all. Exactly 98 of the 1,917 genes are implicated in brain-related diseases, including Alzheimer's disease genes such as APP (Aβ precursor protein; five), MAPT (tau protein; four), and BIN1 (eight). As proof of concept, we also found 99 differentially expressed RNA isoforms between Alzheimer's cases and controls, despite the genes themselves not exhibiting differential expression. Our findings highlight the significant knowledge gaps in RNA isoform diversity and their medical relevance. Deep long-read RNA sequencing will be necessary going forward to fully comprehend the medical relevance of individual isoforms for a "single" gene.