Abstract Tandem repeats are highly mutable genomic elements linked to human traits and diseases. Profiling large catalogs of tandem repeats from population-scale long-read sequencing data requires accurate and efficient tools. We introduce inquiSTR, a command-line toolkit for fast genome-wide tandem repeat length genotyping. inquiSTR, with efficient parallel processing and low-memory streaming algorithms, genotypes a genome-wide repeat catalog of 1.78 million loci in less than two minutes. Benchmarking shows high accuracy and significantly faster performance compared to existing tools and truth sets. inquiSTR also provides methods for downstream analyses such as population structure inference, association testing, and outlier detection.
Neurodegenerative diseases are characterised by the assembly of a limited number of disease-specific proteins into amyloid filaments, which form intracellular inclusions or extracellular deposits in the central nervous system (CNS)1,2. We previously found that amyloid filaments of TATA-binding protein-associated factor 15 (TAF15) characterise a subtype of frontotemporal lobar degeneration with FET protein-immunoreactive inclusions (FTLD-FET)3, termed atypical FTLD with ubiquitin-positive inclusions (aFTLD-U)4, which causes early-onset, rapidly progressive behavioural variant frontotemporal dementia (FTD). However, it was not clear if TAF15 proteinopathy was more widespread in neurodegenerative diseases. Two additional FTLD-FET subtypes have been proposed, neuronal intermediate filament inclusion body disease (NIFID) and basophilic inclusion body disease (BIBD)5,6, which have more heterogenous clinical presentations including FTD, motor neuron diseases (MND) and movement disorders. Here, we used electron cryo-microscopy (cryo-EM) to determine a total of 32 amyloid filament structures from the brains of 17 individuals encompassing all three proposed subtypes of FTLD-FET and their diverse clinical presentations. All cases were characterised by TAF15 filaments, in the absence of filaments of the other FET proteins, fused in sarcoma (FUS) and Ewing's sarcoma (EWS). All three aFTLD-U cases had the previously-reported TAF15 fold3. Unexpectedly, we found four distinct TAF15 folds among 11 NIFID cases. Eight of these cases shared a common fold, while the remaining three were each distinct. Furthermore, we found distinct TAF15 folds for each of the three BIBD cases. Neuropathological reassessment of the neocortical TAF15 inclusion pathology of these cases distinguished the NIFID cases with the common fold from the others. Thus, TAF15 filament structures form the basis of a new, expanded classification of FTLD-FET subtypes. Moreover, we discovered a TAF15 Y38C variant in the filament fold of one of the individuals with BIBD. The structure is unable to incorporate wild-type TAF15, despite the individual being heterozygous, suggesting that this variant drives TAF15 filament assembly. This study provides structural and genetic evidence that TAF15 amyloid filaments underlie the diverse group of neurodegenerative diseases currently termed FTLD-FET, which we therefore rename FTLD-TAF15.
Tandem repeats play critical roles in human disease and phenotypic diversity but are among the most challenging classes of genomic variation to measure accurately. Long-read sequencing has the potential to accurately characterize long and complex tandem repeats. While an increasing number of genotyping methods are available, no systematic effort has been undertaken to evaluate their usability, accuracy, and performance across motifs and allele lengths. We reviewed 25 bioinformatic tools and selected seven actively maintained methods for benchmarking using publicly available Oxford Nanopore genome sequencing data from more than 100 individuals. We assessed performance across 43,009 genome-wide tandem repeat loci using four complementary strategies: concordance with haplotype-resolved Human Pangenome Reference Consortium assemblies, Mendelian consistency, cross-tool consistency, and sensitivity to molecularly confirmed pathogenic expansions. Most methods achieved high concordance with assemblies, with higher accuracy using R10 Oxford Nanopore pore chemistry than older R9 chemistry. Accuracy declined with increasing allele length, and most tools performed worse on homopolymers, heterozygous loci, and alleles differing from the reference genome. Assembly concordance and Mendelian consistency did not predict sensitivity to pathogenic expansions, suggesting that these metrics captured distinct aspects of performance. No single genotyper performs consistently best across all assessments, but strong contenders emerge in each. Our results demonstrate that length accuracy overestimates tandem repeat genotyping performance. Sequence-level benchmarking is essential for selecting tools best-suited for population studies and clinical diagnostics. This work provides practical guidance for tool selection and highlights key priorities for future long-read tandem repeat genotyping method development.
Atypical frontotemporal lobar degeneration with ubiquitin-positive inclusions (aFTLD-U) is neuropathologically characterized by aggregation of the FET family of proteins and clinically manifests as sporadic young-onset frontotemporal dementia. Here we describe a major risk locus on chr15q14 identified through a genome-wide association study in 59 pathologically confirmed aFTLD-U cases and 3,153 controls (lead single nucleotide polymorphism rs549846383, P = 5.85 × 10-21, odds ratio 26.7). When combined with data from 28 additional aFTLD-U cases, 3,712 controls and 3,215 individuals with other neurodegenerative diseases and by leveraging in-house and public long-read genome sequencing data from 1,715 individuals, we identified a tandem repeat expansion on the associated haplotypes in an intron of GOLGA8A. We found variation in repeat length, motif length, and motif sequence, with long CT-dimer expansions strongly associated with aFTLD-U. Although the functional consequence of this repeat remains unknown, its presence in nearly 60% of aFTLD-U cases points to a fundamental role in disease pathogenesis.
Aggregation of TAR-DNA-binding protein 43 (TDP-43) is strongly associated with frontotemporal lobar degeneration (FTLD-TDP), motor neuron disease (MND-TDP), and overlap disorders like FTLD-MND. Three major forms of motor neuron disease are recognized and include primary lateral sclerosis (PLS), amyotrophic lateral sclerosis (ALS), and progressive muscular atrophy (PMA). Annexin A11 (ANXA11) is understood to aggregate in amyotrophic lateral sclerosis (ALS-TDP) associated with pathogenic variants in ANXA11, as well as in FTLD-TDP type C. Given these observations and recent reports of ANXA11 variants in patients with semantic variant frontotemporal dementia (svFTD) and FTD-MND presentations, we sought to characterize ANXA11 proteinopathy in an autopsy cohort of 379 cases diagnosed with a primary TDP-43 proteinopathy, including FTLD-TDP, FTLD-MND, and MND-TDP. Cases with FTLD-MND and MND-TDP were classified further into PLS, ALS, and PMA based on the relative loss of upper and lower motor neurons. ANXA11 proteinopathy was present in over 40% of FTLD-MND cases. Further, ANXA11 colocalized with TDP-43 in the pathologic inclusions of all FTLD-TDP type C cases, as well as 38 out of 40 FTLD-PLS cases (95%), of which 84% had TDP type B or an unclassifiable TDP-43 proteinopathy and 16% had TDP type C. Genetic analysis excluded pathogenic ANXA11 variants in all ANXA11-positive cases. We thus demonstrated two novel ANXA11 proteinopathies strongly associated with FTLD-PLS, but not with TDP type C or pathogenic ANXA11 variants. Given the emerging relationship between TDP-43 and ANXA11 in neurodegenerative disease, we propose that TDP-43 and ANXA11 proteinopathy (TAP) comprises a distinct group of molecular pathologies and define three TAP types based on key clinical and neuropathologic characteristics.
BACKGROUND:Over the years, there has been growing interest in epigenetics, where nucleotide modifications are increasingly recognized for their roles in health and disease. Understanding methylation patterns at the nucleotide level has become pivotal for advancing this field. However, visualizing these modifications, particularly in cohorts of more than a few individuals, remains a challenge. RESULTS:Here, we present methylmap, a tool developed to visualize modified nucleotide frequencies for regions of interest, specifically optimized for cohort sizes with more than a few individuals. Furthermore, methylmap features the visualization of the haplotype-specific methylation status of 226 individuals of the 1000 Genomes Project ONT Sequencing Consortium, sequenced using the Oxford Nanopore Technologies PromethION. This resource provides the research community with a comprehensive and complete overview of genome-wide methylation patterns. CONCLUSIONS:Methylmap offers an easy-to-use platform to facilitate epigenetic research. It is available both as a web application at https://methylmap.bioinf.be and as a command-line tool through Bioconda and PyPI. As such, we provide a valuable resource for advancing the understanding of epigenetic modifications in health and disease.
Atypical frontotemporal lobar degeneration with ubiquitin-positive inclusions (aFTLD-U) is a rare cause of frontotemporal lobar degeneration (FTLD), characterized postmortem by neuronal inclusions of the FET family of proteins (FTLD-FET). The recent discovery of TAF15 amyloid filaments in aFTLD-U brains represents a significant step toward improved diagnostic and therapeutic strategies. However, our understanding of the etiology of this FTLD subtype remains limited, which severely hampers translational research efforts. To explore the transcriptomic changes in aFTLD-U, we performed bulk RNA sequencing on the frontal cortex tissue of 21 aFTLD-U patients and 20 control individuals. Cell-type deconvolution revealed loss of excitatory neurons and a higher proportion of astrocytes in aFTLD-U relative to controls. Differential gene expression and co-expression network analysis, adjusted for the shift in cell-type proportions, showed dysregulation of mitochondrial pathways, transcriptional regulators, and upregulation of the Sonic hedgehog (Shh) pathway, including the GLI1 transcription factor, in aFTLD-U. Overall, oligodendrocyte and astrocyte-enriched genes were significantly over-represented among the differentially expressed genes. Differential splicing analysis confirmed the dysregulation of non-neuronal cell types with significant splicing alterations, particularly in oligodendrocyte-enriched genes, including myelin basic protein (MBP), a crucial component of myelin. Immunohistochemistry in frontal cortex brain tissue also showed reduced myelin levels in aFTLD-U patients compared to controls. Together, these findings highlight a central role for glial cells, particularly astrocytes and oligodendrocytes, in the pathogenesis of aFTLD-U, with disruptions in mitochondrial activity, RNA metabolism, Shh signaling, and myelination as possible disease mechanisms. This study offers the first transcriptomic insight into aFTLD-U and presents new avenues for research into FTLD-FET.
The first edition of the joint international Cambridge-Antwerp Bioinformatics Hackathon welcomed an international group of participants from the Babraham Institute, University of Cambridge (UK), VIB-UAntwerp, VIB-KU-Leuven and University of Antwerp. The three-day event, running in parallel in Antwerp and Cambridge on 11th-13th September 2023, brought together programming enthusiasts in an interdisciplinary setting, to improve their coding skills, develop collaborative projects, and network. The Joint Hackathon model encourages participants to design and develop collaborative cross-site projects of interest spanning life sciences, statistics and bioinformatics, with projects often focusing on creating new software or improving existing tools.
Genetic variation in Transmembrane protein 106B (TMEM106B) is known to influence the risk and presentation in several neurodegenerative diseases and modifies healthy aging. While evidence from human studies suggests that the risk allele is associated with higher levels of TMEM106B, the contribution of elevated levels of TMEM106B to neurodegeneration and aging has not been assessed and it remains unclear how TMEM106B modulates disease risk. To study the effect of increased TMEM106B levels, we generated Cre-inducible transgenic mice expressing human wild-type TMEM106B. We evaluated lysosomal and neuronal health using in vitro and in vivo assays including transmission electron microscopy, immunostainings, behavioral testing, electrophysiology, and bulk RNA sequencing. We created the first transgenic mouse model that successfully overexpresses TMEM106B, with a 4- to 8-fold increase in TMEM106B protein levels in heterozygous (hTMEM106B(+)) and homozygous (hTMEM106B(++)) animals, respectively. We showed that the increase in TMEM106B protein levels induced lysosomal dysfunction and age-related downregulation of genes associated with neuronal plasticity, learning, and memory. Increased TMEM106B levels led to altered synaptic signaling in 12-month-old animals which further exhibited an anxiety-like phenotype. Finally, we observed mild neuronal loss in the hippocampus of 21-month-old animals. Characterization of the first transgenic mouse model that overexpresses TMEM106B suggests that higher levels of TMEM106B negatively impacts brain health by modifying brain aging and impairing the resilience of the brain to the pathomechanisms of neurodegenerative disorders. This novel model will be a valuable tool to study the involvement and contribution of increased TMEM106B levels to aging and will be essential to study the many age-related diseases in which TMEM106B was genetically shown to be a disease- and risk-modifier.
In the last decade, the importance of DNA methylation in the functioning of the central nervous system has been highlighted through associations between methylation changes and differential expression of key genes involved in aging and neurodegenerative diseases. In frontotemporal lobar degeneration (FTLD), aberrant methylation has been reported in causal disease genes including GRN and C9orf72; however, the genome-wide contribution of epigenetic changes to the development of FTLD remains largely unexplored. We performed reduced representation bisulfite sequencing of matched pairs of post-mortem tissue from frontal cortex (FCX) and cerebellum (CER) from pathologically confirmed FTLD patients with TDP-43 pathology (FTLD-TDP) further divided into five subtypes and including both sporadic and genetic forms (N = 25 pairs per group), and neuropathologically normal controls (N = 42 pairs). Case-control differential methylation analyses were performed, both at the individual CpG level, and in regions of grouped CpGs (differentially methylated regions; DMRs), either including all genomic locations or only gene promoters. Gene Ontology (GO) analyses were then performed using all differentially methylated genes in each group of sporadic patients. Finally, additional datasets were queried to prioritize candidate genes for follow-up. Using the largest FTLD-TDP DNA methylation dataset generated to date, we identified thousands of differentially methylated CpGs (FCX = 6,520; CER = 7,134) and several hundred DMRs in FTLD-TDP brains (FCX = 134; CER = 219). Of these, less than 10
Research and diagnostics for medically relevant tandem repeats and repeat expansions are hampered by the lack of population-scale databases. We attempt to fill this gap using our pathSTR web tool, which leverages long-read sequencing of large cohorts to determine repeat length and sequence composition in the general population. The current version includes 878 individuals of the 1000 Genomes Project cohort sequenced on the Oxford Nanopore Technologies PromethION. A comprehensive set of medically relevant tandem repeats were genotyped using STRdust to determine the tandem repeat length and sequence composition. PathSTR provides rich visualizations of this dataset, as well as the feature to upload one’s own data for comparison along the control cohort. We demonstrate the implementation of this application using data from targeted nanopore sequencing of a patient with Myotonic Dystrophy type 1. This resource will empower the genetics community to get a more complete overview of normal variation in tandem repeat length and sequence composition, and enable a better assessment of the pathogenic impact of tandem repeats observed in patients. PathSTR is available at <https://pathstr.bioinf.be> ### Competing Interest Statement WDC has received free consumables and travel reimbursement from Oxford Nanopore Technologies. ### Funding Statement WDC is a recipient of a postdoctoral fellowship from FWO [12ASR24N]. ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: The study concerning the DM1 patient was approved by the Swedish Ethical Review Authority (2019-04746), and written informed consent was obtained from the participating individual or their respective legal guardians. I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes The data generated in this project can be accessed at <https://pathstr.bioinf.be>, where the data can be queried, visualized, and downloaded in the form of a tab-separated file or individual VCF files as generated by STRdust. The original sequencing data is available at [https://ftp.1000genomes.ebi.ac.uk/vol1/ftp/data\_collections/1KG\_ONT_VIENNA/hg38/][1] [1]: https://ftp.1000genomes.ebi.ac.uk/vol1/ftp/data_collections/1KG_ONT_VIENNA/hg38/
Frontotemporal lobar degeneration with neuronal inclusions of the TAR DNA-binding protein 43 (FTLD-TDP) is a fatal neurodegenerative disorder with only a limited number of risk loci identified. We report our comprehensive genome-wide association study as part of the International FTLD-TDP Whole-Genome Sequencing Consortium, including 985 patients and 3,153 controls compiled from 26 institutions/brain banks in North America, Europe and Australia, and meta-analysis with the Dementia-seq cohort. We confirm UNC13A as the strongest overall FTLD-TDP risk factor and identify TNIP1 as a novel FTLD-TDP risk factor. In subgroup analyzes, we further identify genome-wide significant loci specific to each of the three main FTLD-TDP pathological subtypes (A, B and C), as well as enrichment of risk loci in distinct tissues, brain regions, and neuronal subtypes, suggesting distinct disease aetiologies in each of the subtypes. Rare variant analysis confirmed TBK1 and identified C3AR1, SMG8, VIPR1, RBPJL, L3MBTL1 and ANO9, as novel subtype-specific FTLD-TDP risk genes, further highlighting the role of innate and adaptive immunity and notch signaling pathway in FTLD-TDP, with potential diagnostic and novel therapeutic implications.
Less than half of individuals with a suspected Mendelian condition receive a precise molecular diagnosis after comprehensive clinical genetic testing. Improvements in data quality and costs have heightened interest in using long-read sequencing (LRS) to streamline clinical genomic testing, but the absence of control datasets for variant filtering and prioritization has made tertiary analysis of LRS data challenging. To address this, the 1000 Genomes Project ONT Sequencing Consortium aims to generate LRS data from at least 800 of the 1000 Genomes Project samples. Our goal is to use LRS to identify a broader spectrum of variation so we may improve our understanding of normal patterns of human variation. Here, we present data from analysis of the first 100 samples, representing all 5 superpopulations and 19 subpopulations. These samples, sequenced to an average depth of coverage of 37x and sequence read N50 of 54 kbp, have high concordance with previous studies for identifying single nucleotide and indel variants outside of homopolymer regions. Using multiple structural variant (SV) callers, we identify an average of 24,543 high-confidence SVs per genome, including shared and private SVs likely to disrupt gene function as well as pathogenic expansions within disease-associated repeats that were not detected using short reads. Evaluation of methylation signatures revealed expected patterns at known imprinted loci, samples with skewed X-inactivation patterns, and novel differentially methylated regions. All raw sequencing data, processed data, and summary statistics are publicly available, providing a valuable resource for the clinical genetics community to discover pathogenic SVs.
Tandem repeats (TRs) are highly polymorphic in the human genome, have thousands of associated molecular traits, and are linked to over 60 disease phenotypes. However, their complexity often excludes them from at-scale studies due to challenges with variant calling, representation, and lack of a genome-wide standard. To promote TR methods development, we create a comprehensive catalog of TR regions and explore its properties across 86 samples. We then curate variants from the GIAB HG002 individual to create a tandem repeat benchmark. We also present a variant comparison method that handles small and large alleles and varying allelic representation. The 8.1% of the genome covered by the TR catalog holds ∼24.9% of variants per individual, including 124,728 small and 17,988 large variants for the GIAB HG002 TR benchmark. We work with the GIAB community to demonstrate the utility of this benchmark across short and long read technologies.
MOTIVATION:Existing nanopore single-cell data analysis tools showed severe limitations in handling current data sizes. RESULTS:We introduce scywalker, an innovative and scalable package developed to comprehensively analyze long-read sequencing data of full-length single-cell or single-nuclei cDNA. We developed novel scalable methods for cell barcode demultiplexing and single-cell isoform calling and quantification and incorporated these in an easily deployable package. Scywalker streamlines the entire analysis process, from sequenced fragments in FASTQ format to demultiplexed pseudobulk isoform counts, into a single command suitable for execution on either server or cluster. Scywalker includes data quality control, cell type identification, and an interactive report. Assessment of datasets from the human brain, Arabidopsis leaves, and previously benchmarked data from mixed cell lines demonstrate excellent correlation with short-read analyses at both the cell-barcoding and gene quantification levels. At the isoform level, we show that scywalker facilitates the direct identification of cell-type-specific expression of novel isoforms. AVAILABILITY AND IMPLEMENTATION:Scywalker is available on github.com/derijkp/scywalker under the GNU General Public License (GPL) and at https://zenodo.org/records/13359438/files/scywalker-0.108.0-Linux-x86_64.tar.gz.
Structural variants (SVs) are important contributors to human disease. Their characterization remains however difficult due to their size and association with repetitive regions. Long-read sequencing (LRS) and optical genome mapping (OGM) can aid as their molecules span multiple kilobases and capture SVs in full. In this study, we selected six individuals who presented with unresolved SVs. We applied LRS onto all individuals and OGM to a subset of three complex cases. LRS detected and fully resolved the interrogated SV in all samples. This enabled a precise molecular diagnosis in two individuals. Overall, LRS identified 100% of the junctions at single-basepair level, providing valuable insights into their formation mechanisms without need for additional data sources. Application of OGM added straightforward variant phasing, aiding in the unravelment of complex rearrangements. These results highlight the potential of LRS and OGM as follow-up molecular tests for complete SV characterization. We show that they can assess clinically relevant structural variation at unprecedented resolution. Additionally, they detect (complex) cryptic rearrangements missed by conventional methods. This ultimately leads to an increased diagnostic yield, emphasizing their added benefit in a diagnostic setting. To aid their rapid adoption, we provide detailed laboratory and bioinformatics workflows in this manuscript.
The lack of population-scale databases hampers research and diagnostics for medically relevant tandem repeats and repeat expansions. We attempt to fill this gap using our pathSTR web tool, which leverages long-read sequencing of large cohorts to determine repeat length and sequence composition in a healthy population. The current version includes 1040 individuals of The 1000 Genomes Project cohort sequenced on the Oxford Nanopore Technologies PromethION. A comprehensive set of medically relevant tandem repeats has been genotyped using STRdust and LongTR to determine the tandem repeat length and sequence composition. PathSTR provides rich visualizations of this data set and the feature to upload one's data for comparison along the control cohort. We demonstrate the implementation of this application using data from targeted nanopore sequencing of a patient with myotonic dystrophy type 1. This resource will empower the genetics community to get a more complete overview of normal variation in tandem repeat length and sequence composition and, as such, enable a better assessment of rare tandem repeat alleles observed in patients.
Summary Increases in the cohort size in long-read sequencing projects necessitate more efficient software for quality assessment and processing of sequencing data from Oxford Nanopore Technologies and Pacific Biosciences. Here we describe novel tools for summarizing experiments, filtering datasets and visualizing phased alignments results, as well as updates to the NanoPack software suite. Availability and implementation Cramino, chopper, and phasius are written in Rust and available as executable binaries without requiring installation or managing dependencies. NanoPlot and NanoComp are written in Python3. Links to the separate tools and their documentation can be found at https://github.com/wdecoster/nanopack . All tools are compatible with Linux, Mac OS, and the MS Windows 10 Subsystem for Linux and are released under the MIT license. The repositories include test data, and the tools are continuously tested using GitHub Actions. Contact wouter.decoster@uantwerpen.vib.be