Pediatric solid tumors are a leading cause of childhood disease mortality. In this work, we examined germline structural variants (SVs) as risk factors for pediatric extracranial solid tumors using germline genome sequencing of 1765 affected children, their 943 unaffected parents, and 6665 adult controls. We discovered a sex-biased association between very large (>1 megabase) germline chromosomal abnormalities and increased risk of solid tumors in male children. The overall impact of germline SVs was greatest in neuroblastoma, where we uncovered burdens of ultrarare SVs that cause loss of function of highly expressed, mutationally constrained genes, as well as noncoding SVs predicted to disrupt chromatin domain boundaries. Collectively, we estimate that rare germline SVs explain 1.1 to 5.6% of pediatric cancer liability, establishing them as an important component of disease predisposition.
Pediatric solid tumors are rare malignancies that represent a leading cause of death by disease among children in developed countries. The early age-of-onset of these tumors suggests that germline genetic factors are involved, yet conventional germline testing for short coding variants in established predisposition genes only identifies pathogenic events in 10-15% of patients. Here, we examined the role of germline structural variants (SVs)-an underexplored form of germline variation-in pediatric extracranial solid tumors using germline genome sequencing of 1,766 affected children, their 943 unaffected relatives, and 6,665 adult controls. We discovered a sex-biased association between very large (>1 megabase) germline chromosomal abnormalities and a four-fold increased risk of solid tumors in male children. The overall impact of germline SVs was greatest in neuroblastoma, where we revealed burdens of ultra-rare SVs that cause loss-of-function of highly expressed, mutationally intolerant, neurodevelopmental genes, as well as noncoding SVs predicted to disrupt three-dimensional chromatin domains in neural crest-derived tissues. Collectively, our results implicate rare germline SVs as a predisposing factor to pediatric solid tumors that may guide future studies and clinical practice.
Short-read genome sequencing (GS) holds the promise of becoming the primary diagnostic approach for the assessment of autism spectrum disorder (ASD) and fetal structural anomalies (FSAs). However, few studies have comprehensively evaluated its performance against current standard-of-care diagnostic tests: karyotype, chromosomal microarray (CMA), and exome sequencing (ES). To assess the clinical utility of GS, we compared its diagnostic yield against these three tests in 1,612 quartet families including an individual with ASD and in 295 prenatal families. Our GS analytic framework identified a diagnostic variant in 7.8% of ASD probands, almost 2-fold more than CMA (4.3%) and 3-fold more than ES (2.7%). However, when we systematically captured copy-number variants (CNVs) from the exome data, the diagnostic yield of ES (7.4%) was brought much closer to, but did not surpass, GS. Similarly, we estimated that GS could achieve an overall diagnostic yield of 46.1% in unselected FSAs, representing a 17.2% increased yield over karyotype, 14.1% over CMA, and 4.1% over ES with CNV calling or 36.1% increase without CNV discovery. Overall, GS provided an added diagnostic yield of 0.4% and 0.8% beyond the combination of all three standard-of-care tests in ASD and FSAs, respectively. This corresponded to nine GS unique diagnostic variants, including sequence variants in exons not captured by ES, structural variants (SVs) inaccessible to existing standard-of-care tests, and SVs where the resolution of GS changed variant classification. Overall, this large-scale evaluation demonstrated that GS significantly outperforms each individual standard-of-care test while also outperforming the combination of all three tests, thus warranting consideration as the first-tier diagnostic approach for the assessment of ASD and FSAs.
Objective Identification of genetic risk factors for Parkinson disease (PD) has to date been primarily limited to the study of single nucleotide variants, which only represent a small fraction of the genetic variation in the human genome. Consequently, causal variants for most PD risk are not known. Here we focused on structural variants (SVs), which represent a major source of genetic variation in the human genome. We aimed to discover SVs associated with PD risk by performing the first large‐scale characterization of SVs in PD. Methods We leveraged a recently developed computational pipeline to detect and genotype SVs from 7,772 Illumina short‐read whole genome sequencing samples. Using this set of SV variants, we performed a genome‐wide association study using 2,585 cases and 2,779 controls and identified SVs associated with PD risk. Furthermore, to validate the presence of these variants, we generated a subset of matched whole‐genome long‐read sequencing data. Results We genotyped and tested 3,154 common SVs, representing over 412 million nucleotides of previously uncatalogued genetic variation. Using long‐read sequencing data, we validated the presence of three novel deletion SVs that are associated with risk of PD from our initial association analysis, including a 2 kb intronic deletion within the gene LRRN4 . Interpretation We identified three SVs associated with genetic risk of PD. This study represents the most comprehensive assessment of the contribution of SVs to the genetic risk of PD to date. ANN NEUROL 2023;93:1012–1022
Parkinson’s disease is a complex neurodegenerative disorder, affecting approximately one million individuals in the USA alone. A significant proportion of risk for Parkinson’s disease is driven by genetics. Despite this, the majority of the common genetic variation that contributes to disease risk is unknown, in-part because previous genetic studies have focussed solely on the contribution of single nucleotide variants. Structural variants represent a significant source of genetic variation in the human genome. However, because assay of this variability is challenging, structural variants have not been cataloged on a genome-wide scale, and their contribution to the risk of Parkinson’s disease remains unknown. In this study, we 1) leveraged the GATK-SV pipeline to detect and genotype structural variants in 7,772 short-read sequencing data and 2) generated a subset of matched whole-genome Oxford Nanopore Technologies long-read sequencing data from the PPMI cohort to allow for comprehensive structural variant confirmation. We detected, genotyped, and tested 3,154 “high-confidence” common structural variant loci, representing over 412 million nucleotides of non-reference genetic variation. Using the long-read sequencing data, we validated three structural variants that may drive the association signals at known Parkinson’s disease risk loci, including a 2kb intronic deletion within the gene LRRN4 . Further, we confirm that the majority of structural variants in the human genome cannot be detected using short-read sequencing alone, encompassing on average around 4 million nucleotides of inaccessible sequence per genome. Therefore, although these data provide the most comprehensive survey of the contribution of structural variants to the genetic risk of Parkinson’s disease to date, this study highlights the need for large-scale long-read datasets to fully elucidate the role of structural variants in Parkinson’s disease.### Competing Interest StatementD.V. and M.A.N.'s participation in this project was part of a competitive contract awarded to Data Tecnica International LLC by the National Institutes of Health to support open science research. M.A.N. also currently serves on the scientific advisory board for Clover Therapeutics and is an advisor to Neuron23 Inc. BT currently serves on the Editorial board of EclinicalMedicine, JNNP, NBA, is an Associate Editor for Brain and has a collaborative research agreement with Ionis Pharmaceuticals, Roche and Optimeos. FJS receives research support from Illumina, ONT and PacBio. AM works for ONT.M.E.T. receives research funding and/or reagents from Levo Therapeutics, Microsoft Inc., and Illumina Inc.* PD : Parkinson’s disease SVs : Structural variants SNVs : Single nucleotide variants ONT : Oxford Nanopore Technologies SRS : Short-read sequencing LRS : Long-read sequencing GWAS : Genome-wide association study DEL : Deletions DUP : Duplications INS : Insertions INV : Inversions CTX : Translocations CPX : Complex structural variants CNV : Copy number variants MEI : Mobile element insertions SVA : SINE-VNTR- Alu HBS : Harvard Biomarkers Study NABEC : North American Brain Expression Consortium PDBP : Parkinson’s Disease Biomarker Program PPMI : Parkinson’s Progression Markers Initiative UKBEC : United Kingdom Brain Expression Consortium WGS : Whole genome sequencing
Focal segmental glomerulosclerosis (FSGS) is the main pathology underlying steroid-resistant nephrotic syndrome (SRNS) and a leading cause of chronic kidney disease. Monogenic forms of pediatric SRNS are predominantly caused by recessive mutations, while the contribution of de novo variants (DNVs) to this trait is poorly understood. Using exome sequencing (ES) in a proband with FSGS/SRNS, developmental delay, and epilepsy, we discovered a nonsense DNV in TRIM8, which encodes the E3 ubiquitin ligase tripartite motif containing 8. To establish whether TRIM8 variants represent a cause of FSGS, we aggregated exome/genome-sequencing data for 2,501 pediatric FSGS/SRNS-affected individuals and 48,556 control subjects, detecting eight heterozygous TRIM8 truncating variants in affected subjects but none in control subjects (p = 3.28 × 10-11). In all six cases with available parental DNA, we demonstrated de novo inheritance (p = 2.21 × 10-15). Reverse phenotyping revealed neurodevelopmental disease in all eight families. We next analyzed ES from 9,067 individuals with epilepsy, yielding three additional families with truncating TRIM8 variants. Clinical review revealed FSGS in all. All TRIM8 variants cause protein truncation clustering within the last exon between residues 390 and 487 of the 551 amino acid protein, indicating a correlation between this syndrome and loss of the TRIM8 C-terminal region. Wild-type TRIM8 overexpressed in immortalized human podocytes and neuronal cells localized to nuclear bodies, while constructs harboring patient-specific variants mislocalized diffusely to the nucleoplasm. Co-localization studies demonstrated that Gemini and Cajal bodies frequently abut a TRIM8 nuclear body. Truncating TRIM8 DNVs cause a neuro-renal syndrome via aberrant TRIM8 localization, implicating nuclear bodies in FSGS and developmental brain disease.
A Correction to this paper has been published: https://doi.org/10.1038/s41467-021-21077-8.
A Correction to this paper has been published: https://doi.org/10.1038/s41467-021-21077-8.
Genetic variants that inactivate protein-coding genes are a powerful source of information about the phenotypic consequences of gene disruption: genes that are crucial for the function of an organism will be depleted of such variants in natural populations, whereas non-essential genes will tolerate their accumulation. However, predicted loss-of-function variants are enriched for annotation errors, and tend to be found at extremely low frequencies, so their analysis requires careful variant annotation and very large sample sizes1. Here we describe the aggregation of 125,748 exomes and 15,708 genomes from human sequencing studies into the Genome Aggregation Database (gnomAD). We identify 443,769 high-confidence predicted loss-of-function variants in this cohort after filtering for artefacts caused by sequencing and annotation errors. Using an improved model of human mutation rates, we classify human protein-coding genes along a spectrum that represents tolerance to inactivation, validate this classification using data from model organisms and engineered human cells, and show that it can be used to improve the power of gene discovery for both common and rare diseases.
ABSTRACTCurrent clinical guidelines recommend three genetic tests for the assessment of fetal structural anomalies: karyotype to detect microscopically-visible balanced and unbalanced chromosomal rearrangements, chromosomal microarray (CMA) to detect sub-microscopic copy number variants (CNVs), and exome sequencing (ES) to identify individual nucleotide changes in coding sequence. Advances in genome sequencing (GS) analysis suggest that it is poised to displace the sequential application of all three conventional tests to become a single diagnostic approach for the assessment of fetal structural anomalies. However, systematic benchmarking is required to assure that GS can capture the full mutational spectrum associated with fetal structural anomalies and to accurately quantify the added diagnostic yield of GS. We applied a novel GS analytic framework that included the discovery, filtration, and interpretation of nine classes of genomic variation to 7,195 individuals. We assessed the sensitivity of GS to detect diagnostic variants (pathogenic or likely pathogenic) from three standard-of-care tests using 1,612 autism spectrum disorder quartet families (ASD; n=6,448) with matched GS, ES, and CMA data, and validated these findings in 46 fetuses with a clinically reportable variant originally identified by karyotype, CMA, or ES. We then assessed the added diagnostic yield of GS in 249 trios (n=747) comprising a fetus with a structural anomaly detected by ultrasound and two unaffected parents that were pre-screened with a combination of all three standard-of-care tests. Across both cohorts, our GS analytic framework identified 98.2% of all diagnostic variants detected by standard-of-care tests, including 100% of those originally detected by CMA (n=88) and ES (n=61), as well as 78.6% (n=11/14) of the chromosomal rearrangements identified by karyotype. The diagnostic yield from GS was 7.8% across all 1,612 ASD probands, almost two-fold more than CMA (4.4%) and three-fold more than ES (3.0%). We also demonstrated that the yield of ES can approach that of GS when CNVs are captured with high sensitivity from exome data (7.4% vs. 7.8%, respectively). In 249 pre-screened fetuses with structural anomalies, GS provided an additional diagnostic yield of 0.4% beyond the combination of all three tests (karyotype, CMA, and ES). Applying our benchmarking results to existing data indicates that GS can achieve an overall diagnostic yield of 46.1% in unselected fetuses with fetal structural anomalies, providing an estimated 17.2% increase in diagnostic yield over karyotype, 14.1% over CMA, and 36.1% over ES when sequence variants are assessed, and 4.1% when CNVs are also identified from exome data. In this study we demonstrate that GS is sensitive to the detection of almost all pathogenic variation captured by karyotype, CMA, and ES, provides a superior diagnostic yield than any individual test by a wide margin, and contributes a modest increase in diagnostic yield beyond the combination of all three tests. We also outline several strategies to aid the interpretation of GS variants that are cryptic to conventional technologies, which we anticipate will be increasingly encountered as comprehensive variant identification from GS is performed. Taken together, these data suggest GS warrants consideration as a first-tier diagnostic approach for fetal structural anomalies.
t-SNE and hierarchical clustering are popular methods of exploratory data analysis, particularly in biology. Building on recent advances in speeding up t-SNE and obtaining finer-grained structure, we combine the two to create tree-SNE, a hierarchical clustering and visualization algorithm based on stacked one-dimensional t-SNE embeddings. We also introduce alpha-clustering, which recommends the optimal cluster assignment, without foreknowledge of the number of clusters, based off of the cluster stability across multiple scales. We demonstrate the effectiveness of tree-SNE and alpha-clustering on images of handwritten digits, mass cytometry (CyTOF) data from blood cells, and single-cell RNA-sequencing (scRNA-seq) data from retinal cells. Furthermore, to demonstrate the validity of the visualization, we use alpha-clustering to obtain unsupervised clustering results competitive with the state of the art on several image data sets. Software is available at https://github.com/isaacrob/treesne.
ObjectiveRecently, the ASC‐1 complex has been identified as a mechanistic link between amyotrophic lateral sclerosis and spinal muscular atrophy (SMA), and 3 mutations of the ASC‐1 gene TRIP4 have been associated with SMA or congenital myopathy. Our goal was to define ASC‐1 neuromuscular function and the phenotypical spectrum associated with TRIP4 mutations.MethodsClinical, molecular, histological, and magnetic resonance imaging studies were made in 5 families with 7 novel TRIP4 mutations. Fluorescence activated cell sorting and Western blot were performed in patient‐derived fibroblasts and muscles and in Trip4 knocked‐down C2C12 cells.ResultsAll mutations caused ASC‐1 protein depletion. The clinical phenotype was purely myopathic, ranging from lethal neonatal to mild ambulatory adult patients. It included early onset axial and proximal weakness, scoliosis, rigid spine, dysmorphic facies, cutaneous involvement, respiratory failure, and in the older cases, dilated cardiomyopathy. Muscle biopsies showed multiminicores, nemaline rods, cytoplasmic bodies, caps, central nuclei, rimmed fibers, and/or mild endomysial fibrosis. ASC‐1 depletion in C2C12 and in patient‐derived fibroblasts and muscles caused accelerated proliferation, altered expression of cell cycle proteins, and/or shortening of the G0/G1 cell cycle phase leading to cell size reduction.InterpretationOur results expand the phenotypical and molecular spectrum of TRIP4‐associated disease to include mild adult forms with or without cardiomyopathy, associate ASC‐1 depletion with isolated primary muscle involvement, and establish TRIP4 as a causative gene for several congenital muscle diseases, including nemaline, core, centronuclear, and cytoplasmic‐body myopathies. They also identify ASC‐1 as a novel cell cycle regulator with a key role in cell proliferation, and underline transcriptional coregulation defects as a novel pathophysiological mechanism. ANN NEUROL 2020;87:217–232
Given increasing numbers of patients who are undergoing exome or genome sequencing, it is critical to establish tools and methods to interpret the impact of genetic variation. While the ability to predict deleteriousness for any given variant is limited, missense variants remain a particularly challenging class of variation to interpret, since they can have drastically different effects depending on both the precise location and specific amino acid substitution of the variant. In order to better evaluate missense variation, we leveraged the exome sequencing data of 60,706 individuals from the Exome Aggregation Consortium (ExAC) dataset to identify sub-genic regions that are depleted of missense variation. We further used this depletion as part of a novel missense deleteriousness metric named MPC. We applied MPC to de novo missense variants and identified a category of de novo missense variants with the same impact on neurodevelopmental disorders as truncating mutations in intolerant genes, supporting the value of incorporating regional missense constraint in variant interpretation.
Monkol Lek, Konrad J Karczewski, Eric V Minikel, Kaitlin E Samocha, Eric Banks, Timothy Fennell, Anne H O'Donnell-Luria, James S Ware, Andrew J Hill, Beryl B Cummings, Taru Tukiainen, Daniel P Birnbaum, Jack A Kosmicki, Laramie E Duncan, Karol Estrada, Fengmei Zhao, James Zou, Emma Pierce-Hoffman, Joanne Berghout, David N Cooper, Nicole Deflaux, Mark DePristo, Ron Do, Jason Flannick, Menachem Fromer, Laura Gauthier, Jackie Goldstein, Namrata Gupta, Daniel Howrigan, Adam Kiezun, Mitja I Kurki, Ami Levy Moonshine, Pradeep Natarajan, Lorena Orozco, Gina M Peloso, Ryan Poplin, Manuel A Rivas, Valentin Ruano-Rubio, Samuel A Rose, Douglas M Ruderfer, Khalid Shakir, Peter D Stenson, Christine Stevens, Brett P Thomas, Grace Tiao, Maria T Tusie-Luna, Ben Weisburd, HongHee Won, Dongmei Yu, David M Altshuler, Diego Ardissino, Michael Boehnke, John Danesh, Stacey Donnelly, Roberto Elosua, Jose C Florez, Stacey B Gabriel, Gad Getz, Stephen J Glatt, Christina M Hultman, Sekar Kathiresan, Markku Laakso, Steven McCarroll, Mark I McCarthy, Dermot McGovern, Ruth McPherson, Benjamin M Neale, Aarno Palotie, Shaun M Purcell, Danish Saleheen, Jeremiah M Scharf, Pamela Sklar, Patrick F Sullivan, Jaakko Tuomilehto, Ming T Tsuang, Hugh C Watkins, James G Wilson, Mark J Daly, Daniel G MacArthur, Hanna E Abboud, Goncalo Abecasis, Carlos A Aguilar-Salinas, Olimpia Arellano-Campos, Gil Atzmon, Ingvild Aukrust, Cathy L Barr, Graeme I Bell, Graeme I Bell, Sarah Bergen, Lise Bjørkhaug, John Blangero, Donald W Bowden, Cathy L Budman, Noël P Burtt, Federico Centeno-Cruz, John C Chambers, Kimberly Chambert, Robert Clarke, Rory Collins, Giovanni Coppola, Emilio J Córdova, Maria L Cortes, Nancy J Cox, Ravindranath Duggirala, Martin Farrall, Juan C Fernandez-Lopez, Pierre Fontanillas, Timothy M Frayling, Nelson B Freimer, Christian Fuchsberger, Humberto García-Ortiz, Anuj Goel, María J Gómez-Vázquez, María E González-Villalpando, Clicerio González-Villalpando, Marco A Grados, Leif Groop, Christopher A Haiman, Craig L Hanis, Craig L Hanis, Andrew T Hattersley, Brian E Henderson, Jemma C Hopewell, Alicia Huerta-Chagoya, Sergio Islas-Andrade, Suzanne BR Jacobs, Shapour Jalilzadeh, Christopher P Jenkinson, Jennifer Moran, Silvia JiménezMorale, Anna Kähler, Robert A King, George Kirov, Jaspal S Kooner, SUPPLEMENTARY INFORMATION doi:10.1038/nature19057
Large-scale reference data sets of human genetic variation are critical for the medical and functional interpretation of DNA sequence changes. Here we describe the aggregation and analysis of high-quality exome (protein-coding region) DNA sequence data for 60,706 individuals of diverse ancestries generated as part of the Exome Aggregation Consortium (ExAC). This catalogue of human genetic diversity contains an average of one variant every eight bases of the exome, and provides direct evidence for the presence of widespread mutational recurrence. We have used this catalogue to calculate objective metrics of pathogenicity for sequence variants, and to identify genes subject to strong selection against various classes of mutation; identifying 3,230 genes with near-complete depletion of predicted protein-truncating variants, with 72% of these genes having no currently established human disease phenotype. Finally, we demonstrate that these data can be used for the efficient filtering of candidate disease-causing variants, and for the discovery of human 'knockout' variants in protein-coding genes.