Loss-of-function mutations in the X chromosome gene PIGA lead to phosphatidylinositol glycan class A congenital disorder of glycosylation (PIGA-CDG), an ultra-rare CDG typically presenting with seizures, hypotonia, and neurodevelopmental delay. We identified two brothers (probands) with PIGA-CDG, presenting with epilepsy and mild developmental delay. Both probands carry PIGA c.395C>G (p.Ser132Cys), an ultra-rare variant predicted to be damaging. Strikingly, the maternal grandfather and a great uncle also carry the same PIGA variant, but neither presents with symptoms associated with PIGA-CDG. We hypothesized that genetic modifiers might contribute to this reduced penetrance. Using whole-genome sequencing and pedigree analysis, we identified possible susceptibility variants found in the probands and not in the carriers and possible protective variants found in the carriers and not in the probands. Candidate genetic modifier variants included heterozygous, damaging variants in three genes involved directly in glycosylphosphatidylinositol (GPI)-anchor biosynthesis and additional variants in other glycosylation pathways or encoding GPI-anchored proteins. Using a Drosophila eye-based model, we tested modifiers identified through genome sequencing. Loss of CNTN2, a predicted protective modifier that encodes a GPI-anchored protein responsible for neuron/glial interactions, rescues loss of PIGA in the eye-based model, as we predict in the family. Further testing found that the loss of CNTN2 also rescues PIGA-CDG-specific phenotypes, including seizures and climbing defects in Drosophila neurological models of PIGA-CDG. Using pedigree information, genome sequencing, and in vivo testing, we identified CNTN2 as a strong candidate modifier that could explain the incomplete penetrance in this family. Identifying and studying rare disease modifier genes in families may lead to therapeutic targets.
Understanding the human de novo mutation (DNM) rate requires complete sequence information1. Here using five complementary short-read and long-read sequencing technologies, we phased and assembled more than 95% of each diploid human genome in a four-generation, twenty-eight-member family (CEPH 1463). We estimate 98-206 DNMs per transmission, including 74.5 de novo single-nucleotide variants, 7.4 non-tandem repeat indels, 65.3 de novo indels or structural variants originating from tandem repeats, and 4.4 centromeric DNMs. Among male individuals, we find 12.4 de novo Y chromosome events per generation. Short tandem repeats and variable-number tandem repeats are the most mutable, with 32 loci exhibiting recurrent mutation through the generations. We accurately assemble 288 centromeres and six Y chromosomes across the generations and demonstrate that the DNM rate varies by an order of magnitude depending on repeat content, length and sequence identity. We show a strong paternal bias (75-81%) for all forms of germline DNM, yet we estimate that 16% of de novo single-nucleotide variants are postzygotic in origin with no paternal bias, including early germline mosaic mutations. We place all this variation in the context of a high-resolution recombination map (~3.4 kb breakpoint resolution) and find no correlation between meiotic crossover and de novo structural variants. These near-telomere-to-telomere familial genomes provide a truth set to understand the most fundamental processes underlying human genetic variation.
Rapid genomic diagnostics in the Neonatal Intensive Care Unit represents a paradigm shift in medicine with increasing evidence of the utility of early diagnosis, impacting management. The goal of the Utah NeoSeq Project was to implement and evaluate a multidisciplinary and longitudinal rapid sequencing program while transitioning to CLIA-certified sequencing. Enrollment of 65 infants resulted in 26 (40%) with a diagnostic variant(s) and 7 (11%) harboring a strong candidate. This includes re-analyses resulting in four additional diagnoses. Parental surveys indicated that 7% (4/59) of parents had a decisional conflict after consent, and 3% (2/59) experienced decisional regret after the results. Fifty-two provider surveys were conducted. Seventy-nine percent (41/52) of results and 86% (19/22) of diagnostic results were “very useful” or “useful” and associated with management changes. The NeoSeq Project demonstrates that a multidisciplinary collaborative approach to diagnosis is feasible. We have developed a generalizable, collaborative protocol that addresses the need for expedited genetic evaluation with emerging technologies.
MOTIVATION:Variant call format (VCF) files are the standard output format for various software tools that identify genetic variation from DNA sequencing experiments. Downstream analyses require the ability to query, filter, and modify them simply and efficiently. Several tools are available to perform these operations from the command line, including BCFTools, vembrane, slivar, and others. RESULTS:Here, we introduce vcfexpress, a new, high-performance toolset for the analysis of VCF files, written in the Rust programming language. It is nearly as fast as BCFTools, but adds functionality to execute user expressions in the lua programming language for precise filtering and reporting of variants from a VCF or BCF file. We demonstrate performance and flexibility by comparing vcfexpress to other tools using the vembrane benchmark. AVAILABILITY AND IMPLEMENTATION:vcfexpress is available under the MIT license at https://github.com/brentp/vcfexpress with code used for the manuscript deposited in https://doi.org/10.5281/zenodo.14756838.
Using five complementary short- and long-read sequencing technologies, we phased and assembled >95% of each diploid human genome in a four-generation, 28-member family (CEPH 1463) allowing us to systematically assess de novo mutations (DNMs) and recombination. From this family, we estimate an average of 192 DNMs per generation, including 75.5 de novo single-nucleotide variants (SNVs), 7.4 non-tandem repeat indels, 79.6 de novo indels or structural variants (SVs) originating from tandem repeats, 7.7 centromeric de novo SVs and SNVs, and 12.4 de novo Y chromosome events per generation. STRs and VNTRs are the most mutable with 32 loci exhibiting recurrent mutation through the generations. We accurately assemble 288 centromeres and six Y chromosomes across the generations, documenting de novo SVs, and demonstrate that the DNM rate varies by an order of magnitude depending on repeat content, length, and sequence identity. We show a strong paternal bias (75-81%) for all forms of germline DNM, yet we estimate that 17% of de novo SNVs are postzygotic in origin with no paternal bias. We place all this variation in the context of a high-resolution recombination map (~3.5 kbp breakpoint resolution). We observe a strong maternal recombination bias (1.36 maternal:paternal ratio) with a consistent reduction in the number of crossovers with increasing paternal (r=0.85) and maternal (r=0.65) age. However, we observe no correlation between meiotic crossover locations and de novo SVs, arguing against non-allelic homologous recombination as a predominant mechanism. The use of multiple orthogonal technologies, near-telomere-to-telomere phased genome assemblies, and a multi-generation family to assess transmission has created the most comprehensive, publicly available "truth set" of all classes of genomic variants. The resource can be used to test and benchmark new algorithms and technologies to understand the most fundamental processes underlying human genetic variation.
Introduction: Xia-Gibbs syndrome (XGS) is a rare syndromic disorder characterized by developmental delay with intellectual disability, muscular hypotonia, brain anomalies, and nonspecific dysmorphic features. Different heterozygous variants in AHDC1 have been reported as causal for XGS, comprising mainly de novo stop-gain and frameshift events, but also missense variants, deletions, and a duplication of the locus. Case Presentation: We hereby report 2 patients with clinical features of XGS. In the first patient, a de novo interstitial deletion in 1p36.11p35.3 encompassing the entire coding region of AHDC1 was initially suspected by trio exome sequencing and subsequently confirmed by shallow genome sequencing. In the second patient, a de novo deletion comprising most of the 5 ' untranslated region of AHDC1 was detected by genome sequencing. Conclusion: We identified the smallest deletion comprising AHDC1 reported so far by shallow genome sequencing as well as another small AHDC1 deletion by genome sequencing. These methods represent useful techniques for the identification and confirmation of small deletions and structural variants. Furthermore, our data provide additional evidence of AHDC1 haploinsufficiency as a disease mechanism in XGS. Clinically, foot deformity, skin and connective tissue abnormalities observed in one of the patients are consistent with other reported cases of XGS. These findings suggest that these manifestations could be considered as more prevalent characteristics, underscoring the importance of in-depth phenotyping.
AbstractMotivationIdentifyingde novotandem repeat (TR) mutations on a genome-wide scale is essential for understanding genetic variability and its implications in rare diseases. While PacBio HiFi sequencing data enhances the accessibility of the genome’s TR regions for genotyping, simplede novocalling strategies often generate an excess of likely false positives, which can obscure true positive findings, particularly as the number of surveyed genomic regions increases.ResultsWe developed TRGT-denovo, a computational method designed to accurately identify all types ofde novoTR mutations—including expansions, contractions, and compositional changes— within family trios. TRGT-denovo directly interrogates read evidence, allowing for the detection of subtle variations often overlooked in variant call format (VCF) files. TRGT-denovo improves the precision and specificity ofde novomutation (DNM) identification, reducing the number ofde novocandidates by an order of magnitude compared to genotype-based approaches. In our experiments involving eight rare disease trios previously studied TRGT-denovo correctly reclassified all false positive DNM candidates as true negatives. Using an expanded repeat catalog, it identified new candidates, of which 95% (19/20) were experimentally validated, demonstrating its effectiveness in minimizing likely false positives while maintaining high sensitivity for true discoveries.Availability and implementationBuilt in Rust, TRGT-denovo is available as source code and a pre-compiled Linux binary along with a user guide at:https://github.com/PacificBiosciences/trgt-denovo.
Loss of function mutations in the X-linked PIGA gene lead to PIGA-CDG, an ultra-rare congenital disorder of glycosylation (CDG), typically presenting with seizures, hypotonia, and neurodevelopmental delay. We identified two brothers (probands) with PIGA-CDG, presenting with epilepsy and mild developmental delay. Both probands carry PIGA S132C , an ultra-rare variant predicted to be damaging. Strikingly, the maternal grandfather and a great-uncle also carry PIGA S132C , but neither presents with symptoms associated with PIGA-CDG. We hypothesized genetic modifiers may contribute to this reduced penetrance. Using whole genome sequencing and pedigree analysis, we identified possible susceptibility variants found in the probands and not in carriers and possible protective variants found in the carriers and not in the probands. Candidate variants included heterozygous, damaging variants in three genes also involved directly in GPI-anchor biosynthesis and a few genes involved in other glycosylation pathways or encoding GPI-anchored proteins. We functionally tested the predicted modifiers using a Drosophila eye-based model of PIGA-CDG. We found that loss of CNTN2 , a predicted protective modifier, rescues loss of PIGA in Drosophila eye-based model, like what we predict in the family. Further testing found that loss of CNTN2 also rescues patient-relevant phenotypes, including seizures and climbing defects in Drosophila neurological models of PIGA-CDG. By using pedigree information, genome sequencing, and in vivo testing, we identified CNTN2 as a strong candidate modifier that could explain the incomplete penetrance in this family. Identifying and studying rare disease modifier genes in human pedigrees may lead to pathways and targets that may be developed into therapies.
Structural variant (SV) detection in human genomes using short-read sequencing data is hindered by false positives, arising from sequencing and mapping artifacts that mimic genuine SV signals. Despite advances, state-of-the-art SV callers like GRIDSS and Manta exhibit trade-offs between precision and recall, with GRIDSS offering the highest precision and Manta excelling in recall. To address these limitations, we introduce sv-channels, a novel deep learning model designed to improve the precision of SV detection by leveraging read information at call sites. Our method effectively reduces false positives in Manta's deletion callsets, achieving precision that surpasses GRIDSS while maintaining a recall rate comparable to Manta. This represents a significant improvement in SV detection, leveraging Manta's high recall through deep learning and paving the way for more accurate genomic analyses. The sv-channels codebase is openly accessible on GitHub at https://github.com/GooglingTheCancerGenome/sv-channels enabling further research and application in the field. ### Competing Interest Statement J.d.R and W.P.K are co-founders and directors of Cyclomics, a genomics company, they declare no competing interests. L.S is an employer of JSR Life Sciences, he declares no competing interests. S.G, A.K, B.S.P, C.S, S.M, L.R declare no competing interests.
Germline and somatic variants within an individual or cohort are interpreted with information from large cohorts. Annotation with this information becomes a computational bottleneck as population sets grow to terabytes of data. Here, we introduce echtvar, which efficiently encodes population variants and annotation fields into a compressed archive that can be used for rapid variant annotation and filtering. Most variants, represented by chromosome, position and alleles are encoded into 32-bits-half the size of previous encoding schemes and at least 4 times smaller than a naive encoding. The annotations, stored separately within the same archive, are also encoded and compressed. We show that echtvar is faster and uses less space than existing tools and that it can effectively reduce the number of candidate variants. We give examples on germ-line and somatic variants to document how echtvar can facilitate exploratory data analysis on genetic variants. Echtvar is available at https://github.com/brentp/echtvar under an MIT license.
Expansions of short tandem repeats (STRs) cause many rare diseases. Expansion detection is challenging with short-read DNA sequencing data since supporting reads are often mapped incorrectly. Detection is particularly difficult for “novel” STRs, which include new motifs at known loci or STRs absent from the reference genome. We developed STRling to efficiently count k-mers to recover informative reads and call expansions at known and novel STR loci. STRling is sensitive to known STR disease loci, has a low false discovery rate, and resolves novel STR expansions to base-pair position accuracy. It is fast, scalable, open-source, and available at: github.com/quinlan-lab/STRling .
Background Despite numerous molecular and computational advances, roughly half of patients with a rare disease remain undiagnosed after exome or genome sequencing. A particularly challenging barrier to diagnosis is identifying variants that cause deleterious alternative splicing at intronic or exonic loci outside of canonical donor or acceptor splice sites. Results Several existing tools predict the likelihood that a genetic variant causes alternative splicing. We sought to extend such methods by developing a new metric that aids in discerning whether a genetic variant leads to deleterious alternative splicing. Our metric combines genetic variation in the Genome Aggregate Database with alternative splicing predictions from SpliceAI to compare observed and expected levels of splice-altering genetic variation. We infer genic regions with significantly less splice-altering variation than expected to be constrained. The resulting model of regional splicing constraint captures differential splicing constraint across gene and exon categories, and the most constrained genic regions are enriched for pathogenic splice-altering variants. Building from this model, we developed ConSpliceML. This ensemble machine learning approach combines regional splicing constraint with multiple per-nucleotide alternative splicing scores to guide the prediction of deleterious splicing variants in protein-coding genes. ConSpliceML more accurately distinguishes deleterious and benign splicing variants than state-of-the-art splicing prediction methods, especially in “cryptic” splicing regions beyond canonical donor or acceptor splice sites. Conclusion Integrating a model of genetic constraint with annotations from existing alternative splicing tools allows ConSpliceML to prioritize potentially deleterious splice-altering variants in studies of rare human diseases.
Structural variants are associated with cancers and developmental disorders, but challenges with estimating population frequency remain a barrier to prioritizing mutations over inherited variants. In particular, variability in variant calling heuristics and filtering limits the use of current structural variant catalogs. We present STIX, a method that, instead of relying on variant calls, indexes and searches the raw alignments from thousands of samples to enable more comprehensive allele frequency estimation.
Modern DNA sequencing is used as a readout for diverse assays, with the count of aligned sequences (read depth) representing the quantitative signal for each underlying cellular phenomena. Existing data formats for quantitative genomics assays are, however, limited in either the analysis speeds they enable, the disk space they require or both. We have developed the dense depth data dump (D4) format and tool suite, with the goal of balancing improved analysis speeds with file size. The D4 format is adaptive in that it profiles a random sample of aligned sequence depth from the input sequence file to determine an optimal encoding that enables fast data access. We demonstrate that the D4 format offers substantial speed improvements over existing formats for random access, aggregation and summarization, while also achieving better or comparable file sizes. This performance enables scalable downstream analyses that would be otherwise difficult.
Since its introduction in 2011 the variant call format (VCF) has been widely adopted for processing DNA and RNA variants in practically all population studies-as well as in somatic and germline mutation studies. The VCF format can represent single nucleotide variants, multi-nucleotide variants, insertions and deletions, and simple structural variants called and anchored against a reference genome. Here we present a spectrum of over 125 useful, complimentary free and open source software tools and libraries, we wrote and made available through the multiple vcflib, bio-vcf, cyvcf2, hts-nim and slivar projects. These tools are applied for comparison, filtering, normalisation, smoothing and annotation of VCF, as well as output of statistics, visualisation, and transformations of files variants. These tools run everyday in critical biomedical pipelines and countless shell scripts. Our tools are part of the wider bioinformatics ecosystem and we highlight best practices. We shortly discuss the design of VCF, lessons learnt, and how we can address more complex variation through pangenome graph formats, variation that can not easily be represented by the VCF format.