Cancer frequently clusters in families due to shared environment and genetics. However, many familial cancer cases lack a clinically recognized pathogenic germline variant (PGV). We analyzed germline genomes and family history from 2,726 individuals without a PGV in the All of Us Research Program, including 1,496 cases across 18 cancer types with extensive family history and 1,230 family history-negative, cancer-free controls. We identified allelic series of rare structural variants inactivating MSH2 in individuals with phenotypes consistent with Lynch syndrome and BRCA1 in breast cancer. Cancer polygenic risk scores were enriched in cases and correlated with patterns of cancer diagnoses within families. Exome-wide rare variant analyses nominated six candidate predisposition genes, including TSTD2 and BRAT1 in thyroid and breast cancer, respectively. Overall, polygenic risk and rare variants impacting known genes explained a median of 5% of unexplained familial cancers, increasing to 11% when including newly nominated risk factors.
Young-onset lung cancer is enriched for never-smoking and oncogene-driven tumors, yet its inherited genetic basis remains poorly defined. We performed germline whole-genome sequencing in 251 young-onset lung cancer cases (median age 37), which we jointly analyzed with never-smoking cases (n=196; median age 68) and cancer-free controls (n=1,883). We identified enrichments of rare deleterious coding variants across 55 cancer-related gene sets, including EGFR/ERBB2 signaling and genes implicated by prior lung cancer GWAS. Exome-wide analyses of rare coding variants affirmed TP53 as a penetrant lung cancer predisposition gene (odds ratio [OR]=36.1, p=1.02×10 -7 ) and discovered two novel exome-wide significant tumor subtype-dependent associations: IREB2 in cases with fusion-driven tumors (p=1.39×10 -6 ) and SMAD6 in fusion-negative tumors (p=2.05×10 -6 ). Structural variants contributed distinct risk, with enrichment in constrained, lung-expressed genes (OR=5.79, p=5.8×10 -5 ) and very large germline deletions being markedly enriched in cases with fusion-driven tumors. Polygenic risk scores for lung cancer were inversely correlated with rare variant burden, consistent with additive risk from rare and common variants. Collectively, these findings delineate a complex germline architecture underlying susceptibility and molecular subtype in young-onset lung cancer.
Recursive splice sites are rare motifs postulated to facilitate splicing across massive introns and shape isoform diversity, especially for long, brain-expressed genes. The necessity of this unique mechanism remains unsubstantiated, as does the role of recursive splicing (RS) in human disease. From analyses of rare copy number variants (CNVs) from almost one million individuals, we previously identified large, heterozygous deletions eliminating an RS site (RS1) in the first intron of CADM2 that conferred substantial risk for attention deficit hyperactivity disorder (ADHD) and other neurobehavioral traits. CADM2 encodes a neuronally expressed cell adhesion molecule that has repeatedly been associated with ADHD and numerous similar traits. To explore the molecular impact of RS ablation in CADM2 , we used CRISPR to model patient deletions and to target a smaller region (~500 base pairs) containing RS1 in both human induced neurons (iNs) and rats. Transcriptome analyses in unedited iNs provided a catalog of CADM2 transcripts, including novel transcripts that retained RS exons. Intriguingly, ablating RS1 altered the gradient of RNA abundance across the first intron of CADM2 , decreased the level of CADM2 expression, and impacted transcript usage. Decreased CADM2 expression was reflected in reduced exon usage downstream of the RS1 site and global alteration to genes involved in neuronal processes including synapse and axon development. Given the scale of our analyses and the widespread association of CADM2 with neurobehavioral traits, we sought to validate these findings using in vivo models and found that rodent models harboring Cadm2 RS1 deletions exhibited significant changes in relevant behaviors and functional brain connectivity. In summary, our analyses demonstrate a functional role for RS as a noncoding regulatory mechanism in a gene associated with a spectrum of neuropsychiatric and behavioral traits.
Pediatric solid tumors are a leading cause of childhood disease mortality. In this work, we examined germline structural variants (SVs) as risk factors for pediatric extracranial solid tumors using germline genome sequencing of 1765 affected children, their 943 unaffected parents, and 6665 adult controls. We discovered a sex-biased association between very large (>1 megabase) germline chromosomal abnormalities and increased risk of solid tumors in male children. The overall impact of germline SVs was greatest in neuroblastoma, where we uncovered burdens of ultrarare SVs that cause loss of function of highly expressed, mutationally constrained genes, as well as noncoding SVs predicted to disrupt chromatin domain boundaries. Collectively, we estimate that rare germline SVs explain 1.1 to 5.6% of pediatric cancer liability, establishing them as an important component of disease predisposition.
The biomedical community is increasingly invested in capturing all genetic variants across human genomes, interpreting their functional consequences and translating these findings to the clinic. A crucial component of this endeavour is the discovery and characterization of structural variants (SVs), which are ubiquitous in the human population, heterogeneous in their mutational processes, key substrates for evolution and adaptation, and profound drivers of human disease. The recent emergence of new technologies and the remarkable scale of sequence-based population studies have begun to crystalize our understanding of SVs as a mutational class and their widespread influence across phenotypes. In this Review, we summarize recent discoveries and new insights into SVs in the human genome in terms of their mutational patterns, population genetics, functional consequences, and impact on human traits and disease. We conclude by outlining three frontiers to be explored by the field over the next decade. Collins and Talkowski provide a broad overview of structural variation in the human genome that covers their mutational properties, the dynamics of population genetics and functional consequences in disease as well as promising directions for future research.
Defining genes that are somatically mutated in different cancers is a central goal of cancer genetics. Nevertheless, traditional definitions of “driver” genes based on pooled mutation frequency are biased toward common cancer types and tend to overlook genes that might be specific to rarer subtypes. As such, our understanding of the compendium of cancer genes is incomplete. In this study, we developed a statistical framework that defines genes enriched for functional somatic mutations in primary cancers to quantify incidence and tissue specificity for each gene. By applying this framework to the AACR GENIE v18.0 dataset, which comprised over 1.15 million mutations identified in 146,394 patients spanning 265 histologic subtypes, we identified a total of 95 genes significantly mutated in at least one subtype. We mined this dataset to derive tissue specificity scores for all 95 genes, demonstrating that tissue specificity in cancer is the norm, not the exception, for nearly all genes. We interrogated these new tissue specificity scores to reveal that oncogenes with restricted expression across normal tissues tend to exhibit higher tissue-specific mutation patterns in cancer. In contrast, tumor suppressors were often ubiquitously expressed irrespective of their underlying tissue specificity, which was partially correlated with differentially expressed compensatory genes and pathways in mutationally permissive tissues. Significance Statement We present a statistically robust, subtype-aware framework to identify tissue-specific cancer driver genes and uncover their underlying biology. Our findings reveal that tissue-selective oncogenicity arises from expression constraints and tissue-restricted compensatory pathways. These findings provide a systematic, well-powered resource on mutational significance, tissue tropism, and biological consequences for the cancer research community. ### Competing Interest Statement The authors have declared no competing interest. Cancer Research UK Grand Challenge Mark Foundation For Cancer Research (The Mark Foundation for Cancer Research) HHS | NIH | National Cancer Institute (NCI), K99CA286805
Pediatric solid tumors are rare malignancies that represent a leading cause of death by disease among children in developed countries. The early age-of-onset of these tumors suggests that germline genetic factors are involved, yet conventional germline testing for short coding variants in established predisposition genes only identifies pathogenic events in 10-15% of patients. Here, we examined the role of germline structural variants (SVs)-an underexplored form of germline variation-in pediatric extracranial solid tumors using germline genome sequencing of 1,766 affected children, their 943 unaffected relatives, and 6,665 adult controls. We discovered a sex-biased association between very large (>1 megabase) germline chromosomal abnormalities and a four-fold increased risk of solid tumors in male children. The overall impact of germline SVs was greatest in neuroblastoma, where we revealed burdens of ultra-rare SVs that cause loss-of-function of highly expressed, mutationally intolerant, neurodevelopmental genes, as well as noncoding SVs predicted to disrupt three-dimensional chromatin domains in neural crest-derived tissues. Collectively, our results implicate rare germline SVs as a predisposing factor to pediatric solid tumors that may guide future studies and clinical practice.
Background and objective: Previous germline studies on renal cell carcinoma (RCC) have usually pooled clear and non-clear cell RCCs and have not adequately accounted for population stratification, which might have led to an inaccurate estimation of genetic risk. Here, we aim to analyze the major germline drivers of RCC risk and clinically relevant but underexplored germline variant types. Methods: We first characterized germline pathogenic variants (PVs), cryptic splice variants, and copy number variants (CNVs) in 1436 unselected RCC patients. To evaluate the enrichment of PVs in RCC, we conducted a case-control study of 1356 RCC patients ancestry matched with 16 512 cancer-free controls using approaches accounting for population stratification and histological subtypes, followed by characterization of secondary somatic events. Key findings and limitations: Clear cell RCC patients (n = 976) exhibited a significant burden of PVs in VHL compared with controls (odds ratio [OR]: 39.1, p = 4.95e-05). Non-clear cell RCC patients (n = 380) carried enrichment of PVs in FH (OR: 77.9, p = 1.55e-08) and MET (OR: 1.98e11, p = 2.07e-05). In a CHEK2-focused analysis with European participants, clear cell RCC (n = 906) harbored nominal enrichment of low-penetrance CHEK2 variants-p.Ile157Thr (OR: 1.84, p = 0.049) and p. Ser428Phe (OR: 5.20, p = 0.045), while non-clear cell RCC (n = 295) exhibited nominal enrichment of CHEK2 loss of function PVs (OR: 3.51, p = 0.033). Patients with germline PVs in FH, MET, and VHL exhibited significantly earlier age of cancer onset than patients without germline PVs (mean: 46.0 vs 60.2 yr, p < 0.0001), and more than half had secondary somatic events affecting the same gene (n = 10/15, 66.7%). Conversely, CHEK2 PV carriers exhibited a similar age of onset to patients without germline PVs (mean: 60.1 vs 60.2 yr, p = 0.99), and only 30.4% carried somatic events in CHEK2 (n = 7/23). Finally, pathogenic germline cryptic splice variants were identified in SDHA and TSC1, and pathogenic germline CNVs were found in 18 patients, including CNVs in FH, SDHA, and VHL. Conclusions and clinical implications: This analysis supports the existing link between several RCC risk genes and RCC risk manifesting in earlier age of onset. It calls for caution when assessing the role of CHEK2 due to the burden of founder variants with varying population frequency. It also broadens the definition of the RCC germline landscape of pathogenicity to incorporate previously understudied types of germline variants. Patient summary: In this study, we carefully compared the frequency of rare inherited mutations with a focus on patients' genetic ancestry. We discovered that subtle variations in genetic background may confound a case -control analysis, especially in evaluating the cancer risk associated with specific genes, such as CHEK2. We also identified previously less explored forms of rare inherited mutations, which could potentially increase the risk of kidney cancer. (c) 2024 The Author(s). Published by Elsevier B.V. on behalf of European Association of Urology. This is an open access article under the CC BY -NC -ND license (http://creativecommons.org/licenses/by-nc-nd/4.0/).
Structural variants (SVs) are an important class of variation that contribute to the diversity of each human genome and a major component of the genetic architecture of human traits and disease. Unfortunately, the field of human genetics has lacked accessible maps of SVs at comparable scale and diversity to those routinely used for clinical screening of point mutations. In the Genome Aggregation Database (gnomAD), we describe an SV atlas comprised of ∼1.2 million SVs generated from 464,297 samples with exome sequencing (ES) and 63,046 individuals with genome sequencing (GS).
The depletion of disruptive variation caused by purifying natural selection (constraint) has been widely used to investigate protein-coding genes underlying human disorders1-4, but attempts to assess constraint for non-protein-coding regions have proved more difficult. Here we aggregate, process and release a dataset of 76,156 human genomes from the Genome Aggregation Database (gnomAD)-the largest public open-access human genome allele frequency reference dataset-and use it to build a genomic constraint map for the whole genome (genomic non-coding constraint of haploinsufficient variation (Gnocchi)). We present a refined mutational model that incorporates local sequence context and regional genomic features to detect depletions of variation. As expected, the average constraint for protein-coding sequences is stronger than that for non-coding regions. Within the non-coding genome, constrained regions are enriched for known regulatory elements and variants that are implicated in complex human diseases and traits, facilitating the triangulation of biological annotation, disease association and natural selection to non-coding DNA analysis. More constrained regulatory elements tend to regulate more constrained protein-coding genes, which in turn suggests that non-coding constraint can aid the identification of constrained genes that are as yet unrecognized by current gene constraint metrics. We demonstrate that this genome-wide constraint map improves the identification and interpretation of functional human genetic variation.
ABSTRACT Balanced chromosomal rearrangements (BCRs), including inversions, translocations, and insertions, reorganize large sections of the genome and contribute substantial risk for developmental disorders (DDs). However, the rarity and lack of systematic screening for BCRs in the population has precluded unbiased analyses of the genomic features and mechanisms associated with risk for DDs versus normal developmental outcomes. Here, we sequenced and analyzed 1,420 BCR breakpoints across 710 individuals, including 406 DD cases and the first large-scale collection of 304 control BCR carriers. We found that BCRs were not more likely to disrupt genes in DD cases than controls, but were seven-fold more likely to disrupt genes associated with dominant DDs (21.3% of cases vs. 3.4% of controls; P = 1.60×10 −12 ). Moreover, BCRs that did not disrupt a known DD gene were significantly enriched for breakpoints that altered topologically associated domains (TADs) containing dominant DD genes in cases compared to controls (odds ratio [OR] = 1.43, P = 0.036). We discovered six TADs enriched for noncoding BCRs (false discovery rate < 0.1) that contained known DD genes ( MEF2C, FOXG1, SOX9, BCL11A, BCL11B , and SATB2 ) and represent candidate pathogenic long-range positional effect (LRPE) loci. These six TADs were collectively disrupted in 7.4% of the DD cohort. Phased Hi-C analyses of five cases with noncoding BCR breakpoints localized to one of these putative LRPEs, the 5q14.3 TAD encompassing MEF2C , confirmed extensive disruption to local 3D chromatin structures and reduced frequency of contact between the MEF2C promoter and annotated enhancers. We further identified six genomic features enriched in TADs preferentially disrupted by noncoding BCRs in DD cases versus controls and used these features to build a model to predict TADs at risk for LRPEs across the genome. These results emphasize the potential impact of noncoding structural variants to cause LRPEs in unsolved DD cases, as well as the complex interaction of features associated with predicting three-dimensional chromatin structures intolerant to disruption.
We characterized the role of structural variants, a largely unexplored type of genetic variation, in two non-Alzheimer's dementias, namely Lewy body dementia (LBD) and frontotemporal dementia (FTD)/amyotrophic lateral sclerosis (ALS). To do this, we applied an advanced structural variant calling pipeline (GATK-SV) to short-read whole-genome sequence data from 5,213 European-ancestry cases and 4,132 controls. We discovered, replicated, and validated a deletion in TPCN1 as a novel risk locus for LBD and detected the known structural variants at the C9orf72 and MAPT loci as associated with FTD/ALS. We also identified rare pathogenic structural variants in both LBD and FTD/ALS. Finally, we assembled a catalog of structural variants that can be mined for new insights into the pathogenesis of these understudied forms of dementia.
ABSTRACTPurposeLarge copy number variants (CNVs) can cause a heterogeneous spectrum of rare and severe disorders. However, most CNVs are benign and are part of natural variation in human genomes. CNV pathogenicity classification, genotype-phenotype analyses, and therapeutic target identification are challenging and time-consuming tasks that require the integration and analysis of information from multiple scattered sources by experts.MethodsWe developed a web-application combining >250,000 patient and population CNVs together with a large set of biomedical annotations and provide tools for CNV classification based on ACMG/ClinGen guidelines and gene-set enrichment analyses.ResultsHere, we introduce the CNV-ClinViewer (https://cnv-ClinViewer.broadinstitute.org), an open-source web-application for clinical evaluation and visual exploration of CNVs. The application enables real-time interactive exploration of large CNV datasets in a user-friendly designed interface.ConclusionOverall, this resource facilitates semi-automated clinical CNV interpretation and genomic loci exploration and, in combination with clinical judgment, enables clinicians and researchers to formulate novel hypotheses and guide their decision-making process. Subsequently, the CNV-ClinViewer enhances for clinical investigators patient care and for basic scientists translational genomic research.
Copy number variants (CNVs) are major contributors to genetic diversity and disease. While standardized methods, such as the genome analysis toolkit (GATK), exist for detecting short variants, technical challenges have confounded uniform large-scale CNV analyses from whole-exome sequencing (WES) data. Given the profound impact of rare and de novo coding CNVs on genome organization and human disease, we developed GATK-gCNV, a flexible algorithm to discover rare CNVs from sequencing read-depth information, complete with open-source distribution via GATK. We benchmarked GATK-gCNV in 7,962 exomes from individuals in quartet families with matched genome sequencing and microarray data, finding up to 95% recall of rare coding CNVs at a resolution of more than two exons. We used GATK-gCNV to generate a reference catalog of rare coding CNVs in WES data from 197,306 individuals in the UK Biobank, and observed strong correlations between per-gene CNV rates and measures of mutational constraint, as well as rare CNV associations with multiple traits. In summary, GATK-gCNV is a tunable approach for sensitive and specific CNV discovery in WES data, with broad applications. GATK-gCNV uses a probabilistic model and inference framework to discover rare copy number variants (CNVs) from sequencing read-depth information. This algorithm is used to generate a reference catalog of rare coding CNVs in exome sequencing data from UK Biobank.