Abstract Determining the clinical relevance of BRCA2 variants of uncertain significance is critical for informed risk management. Recently, two saturation genome editing studies assessed the functional effects of all single nucleotide variants in the BRCA2 C-terminal DNA Binding Domain. To improve the accuracy of functional data used for ACMG/AMP variant classification, we combined results from these studies in four composite models and evaluated the performance of each model using variants with known classifications. Here, we show that an “Integrated VarCall Model”, which combined raw functional data for 6383 variants from the original studies, yielded 98.8% accuracy and out-performed the original studies and other combined data models. Incorporation of the “Integrated VarCall Model” functional data with other sources of evidence according to ClinGen BRCA1/2 variant curation expert panel specifications resulted in classification of 5926 (92.8%) BRCA2 variants as pathogenic (n = 735) or benign (n = 5191) and provides valuable insights for individuals with BRCA2 variants.
Although sequencing costs have steadily decreased with advances in technology, they remain high for large scale studies. The design of traditional individual-disease sequencing studies is either case only or cases with relatively few controls, resulting in potential loss of statistical power for discovery of disease associated genes. Here we show that for a given number of sequenced cases, a large control sample size is critical to maximize power for rare variant burden analysis. Furthermore, we have developed an end-to-end workflow based tool (CoCoRV-nf) to facilitate the use of external biobank sequence resources as controls. The modules include consistent variant QC, variant annotation, ancestry population prediction, and gene based burden analysis using summary genotype information, and combined analysis from multiple independent results. The tool supports exomes and genomes from gnomAD and All of Us as controls with preprocessed datasets. We apply the tool in two rare neurological diseases: amyotrophic lateral sclerosis and neuroblastoma. For each disease, two case cohorts are paired with gnomAD and All of Us data, respectively, followed by a combined analysis. Not only did we recapture known genes, but also, we identified new candidate genes for both diseases. By leveraging multiple large external biobank sequence data, we demonstrate the feasibility of using our tool to maximize statistical power to identify new disease predisposition genes.
By integrating short-read WGS and RNA-seq data with long-read RNA sequencing, we dissect the complex genomic architecture of PAX5 intragenic tandem multiplication (PAX5-ITM), revealing that these complex rearrangements result in in-frame transcripts that likely encode proteins with altered domains.
Germline BRCA2 loss-of function variants, which can be identified through clinical genetic testing, predispose to several cancers1–5. However, variants of uncertain significance limit the clinical utility of test results. Thus, there is a need for functional characterization and clinical classification of all BRCA2 variants to facilitate the clinical management of individuals with these variants. Here we analysed all possible single-nucleotide variants from exons 15 to 26 that encode the BRCA2 DNA-binding domain hotspot for pathogenic missense variants. To enable this, we used saturation genome editing CRISPR–Cas9-based knock-in endogenous targeting of human haploid HAP1 cells6. The assay was calibrated relative to nonsense and silent variants and was validated using pathogenic and benign standards from ClinVar and results from a homology-directed repair functional assay7. Variants (6,959 out of 6,960 evaluated) were assigned to seven categories of pathogenicity based on a VarCall Bayesian model8. Single-nucleotide variants that encode loss-of-function missense variants were associated with increased risks of breast cancer and ovarian cancer. The functional assay results were integrated into models from ClinGen, the American College of Medical Genetics and Genomics, and the Association for Molecular Pathology9 for clinical classification of BRCA2 variants. Using this approach, 91% were classified as pathogenic or likely pathogenic or as benign or likely benign. These classified variants can be used to improve clinical management of individuals with a BRCA2 variant. Results from a comprehensive evaluation of the function of BRCA2 variants, particularly variants of uncertain significance, provide a useful resource to improve the clinical management of individuals who carry such genetic variants.
Somatic mitochondrial DNA (mtDNA) mutations are frequently observed in tumors, yet their role in pediatric cancers remains poorly understood. The heteroplasmic nature of mtDNA-where mutant and wild-type mtDNA coexist-complicates efforts to define its contribution to disease progression. In this study, bulk whole-genome sequencing of 637 matched tumor-normal samples from the Pediatric Cancer Genome Project revealed an enrichment of functionally impactful mtDNA variants in specific pediatric leukemia subtypes. Collectively, the results from single-cell sequencing of five diagnostic leukemia samples demonstrated that somatic mtDNA mutations can arise early in leukemogenesis and undergo positive selection during disease progression, achieving intermediate heteroplasmy-a "sweet spot" that balances mitochondrial dysfunction with cellular fitness. Network-based systems biology analyses link specific heteroplasmic mtDNA mutations to metabolic reprogramming and therapy resistance. We reveal somatic mtDNA mutations as a potential source of functional heterogeneity and cellular diversity among leukemic cells, influencing their fitness and shaping disease progression.
PURPOSE:Recent studies reveal that 5%-18% of children with cancer harbor pathogenic variants in known cancer-predisposing genes. However, DNA damage repair (DDR) genes, which are frequently somatically altered in pediatric tumors, have not been systematically examined as a source of novel cancer-predisposing signals. METHODS:To address this gap, we interrogated 189 DDR genes for presence of germline predisposing variants (PV) among 5,993 childhood cancer cases and 14,477 adult noncancer controls (discovery cohort). PV were determined using a tiered approach incorporating ClinVar annotations, InterVar classification, and in silico tools (REVEL, CADD, and MetaSVM). Using logistic and firth regression, we identified genes with PV statistically enriched in the germline of children with tumors and replicated findings among 1,497 additional childhood cancer cases across three independent cohorts. RESULTS:Analysis across all cases with cancer revealed enrichment of TP53 PV. Cancer-specific analyses confirmed known associations including germline TP53 PV in adrenocortical carcinoma, high-grade glioma (HGG), and medulloblastoma (MB), PMS2 in HGG and non-Hodgkin lymphoma (NHL), MLH1 in HGG, BRCA2 in NHL, and BARD1 in neuroblastoma. In addition, four novel associations were uncovered, including BRCA1 in ependymoma, SPIDR in HGG, SMC5 in MB, and SMARCAL1 in osteosarcoma (OS). Importantly, the SMARCAL1:OS association was significant in the discovery (6/230, 2.6%, false discovery rate [FDR]logistic = 0.0189) as well as all three replication cohorts (Childhood Cancer Survivor Study: 8/275, 2.9%; PFisher < .0001; Cancer Predisposition Syndrome-German Childhood Cancer Registry: 4/135, 3%, PFisher = .002; Individualized Therapy for Relapsed Malignancies in Childhood: 4/217, 1.8%, PFisher = .012). The remaining wild-type SMARCAL1 allele was deleted in three of four OS tumors with available data. CONCLUSION:Our study confirms the relevance DDR genetic variation in pediatric cancer risk and establishes SMARCAL1 as a novel OS predisposing gene, providing insights into tumor biology and creating opportunities to optimize care for patients with this challenging tumor.
BACKGROUND:Diagnosing Mendelian and rare genetic conditions requires identifying phenotype-associated genetic findings and prioritizing likely disease-causing genes. This task is labor-intensive for molecular and clinical geneticists, who must review extensive literature and databases to link patient phenotypes with causal genotypes. The challenge is further complicated by the large number of genetic variants detected through next-generation sequencing, which impacts both diagnosis timelines and patient care strategies. To address this, in silico methods that prioritize causal genes based on patient-derived phenotypes offer an effective solution, reducing the time involved in diagnostic case reviews and enhancing the efficiency of clinical diagnosis. RESULTS:We developed the phenotype prioritization and analysis for rare diseases (PPAR) to rank genes based on human phenotype ontology (HPO) terms, with the specific goal of aiding the interpretation of genetic testing for Mendelian and rare diseases. PPAR leverages embeddings from a knowledge graph and incorporates knowledge from connections between genes, HPO terms, and gene ontology annotations. When applied on a clinical rare disease cohort and the publicly available deciphering developmental disorders (DDD) dataset. PPAR ranked the causal gene in the top 10 for 27% of cases in the clinical cohort and for 85% of cases in the DDD dataset, outperforming other established HPO-based methods. CONCLUSION:Our findings demonstrate that PPAR, a method developed from the clinical knowledge graph, effectively ranks causal genes based on patient-derived HPO terms in rare and Mendelian disease contexts. PPAR has shown superior performance compared to other well-established HPO-only methods and provides an efficient, accessible solution for clinical geneticists. The Python-based tool is publicly available at https://github.com/dimi-lab/PPAR , offering a user-friendly platform for gene prioritization.
Background:Recent large-scale genomic sequencing studies reveal that 5-18% of children with cancer harbor pathogenic variants (PV) in known cancer predisposing genes (CPG). However, DNA damage repair (DDR) genes, which are frequently somatically altered in pediatric tumors, have not been systematically examined as a source of novel cancer predisposing signals. Methods:To address this gap, we interrogated 189 genes across six DDR pathways for the presence of PV among 5,993 childhood cancer cases and 14,477 adult non-cancer controls. PV were defined as rare (allele frequency <0.05% in the gnomAD v2.1 non-cancer subset), nonsense, frameshift, affecting canonical splice sites, and missense with REVEL score >0.7. Using logistic and firth regression, we identified genes with statistically enriched PV and replicated findings among 1,494 additional childhood cancer cases across three independent cohorts. Findings:Analysis across all cancers revealed enrichment of TP53 PV (0.6%, false discovery rate [FDR]logistic=0.0066, FDRFirth=0.0064). Cancer-specific analyses confirmed previously identified associations for germline TP53 PV in adrenocortical carcinoma (37%, FDRlogistic<0.0001, FDRFirth=0) and high-grade glioma (2.4%, FDRlogistic=0.0022, FDRFirth=0.1082), as well as BARD1 PV in neuroblastoma (1.2%, FDRlogistic=0.0341, FDRFirth=0.2682). Three novel gene-tumor associations were identified, including POLL PV in Ewing sarcoma (1.7%, FDRlogistic=0.0319, FDRFirth=0.3101), SMC5 PV in medulloblastoma (1.6%, FDRlogistic=0.0005, FDRFirth=0.0499) and SMARCAL1 PV in osteosarcoma (2.6%, FDRlogistic=0.0250, FDRFirth=0.2180). Among these putative CPG, SMARCAL1 PV were enriched in osteosarcoma across each of the replication pediatric cancer cohorts (2.5%, PFisher <0.0001). All three osteosarcomas with available tumor data exhibited deletion of the wild-type SMARCAL1 allele. Interpretation:Our study identifies SMARCAL1 PV as a predisposing factor for osteosarcoma, providing insights into tumor biology and creating opportunities for development of novel therapeutic, surveillance, and preventive interventions for this aggressive childhood cancer.
Clinical genetic testing identifies variants causal for hereditary cancer, information that is used for risk assessment and clinical management. Unfortunately, some variants identified are of uncertain clinical significance (VUS), complicating patient management. Case-control data is one evidence type used to classify VUS. As an initiative of the Evidence-based Network for the Interpretation of Germline Mutant Alleles (ENIGMA) Analytical Working Group we analyze germline sequencing data of BRCA1 and BRCA2 from 96,691 female breast cancer cases and 302,116 controls from three studies: the BRIDGES study of the Breast Cancer Association Consortium, the Cancer Risk Estimates Related to Susceptibility consortium, and the UK Biobank. We observe 11,207 BRCA1 and BRCA2 variants, with 6909 being coding, covering 23.4% of BRCA1 and BRCA2 VUS in ClinVar and 19.2% of ClinVar curated (likely) benign or pathogenic variants. Case-control likelihood ratio (ccLR) evidence is highly consistent with ClinVar assertions for (likely) benign or pathogenic variants; exhibiting 99.1% sensitivity and 95.3% specificity for BRCA1 and 93.3% sensitivity and 86.6% specificity for BRCA2. This approach provides case-control evidence for 787 unclassified variants; these include 579 with strong or moderate benign evidence and 10 with strong pathogenic evidence for which ccLR evidence is sufficient to alter clinical classification.
PURPOSE To reduce costs in genomic studies of time-to-event phenotypes like survival, researchers often sequence a subset of samples from a larger cohort. This process usually involves two phases: first, collecting inexpensive variables from all samples, and second, selecting a subset for expensive measurements, for example, sequencing-based biomarkers. Common two-phase designs include nested case-control and case-cohort designs. Additional designs include sampling subjects based on follow-up time, like extreme case-control designs. Recently an optimal two-phase design using a maximum likelihood-based method was proposed, which could accommodate arbitrary sample selection in the second phase. However, direct comparisons of this optimal design with others in terms of power and computational cost is lacking. METHODS This study performs a direct evaluation of typical two-phase designs, including Tao's optimal design, on type I error, power, effect size estimation, and computational time, using both simulated and real data sets. RESULTS Results show that the optimal design had the highest power and accurate effect size estimation under the Cox regression model. Surprisingly, logistic regression achieved similar power with much lower computational cost than a more sophisticated method. The study further applied these methods to the MP2PRT study, reporting hazard ratios of cancer subtypes on relapse risk. CONCLUSION Recommendations for selecting two-phase designs and analysis methods are regarding power, bias of estimated effect size, and computational time.
Immunophenotyping of out -of -hospital cardiac arrest (OHCA) patients is of increasing interest but has challenges. Here, we describe steps for the design of the clinical cohort, planning patient enrollment and sample collection, and ethical review of the study protocol. We detail procedures for blood sample collection and cryopreservation of peripheral blood mononuclear cells (PBMCs). We detail steps to modulate immune checkpoints in OHCA PBMC ex vivo. This protocol also has relevance for immunophenotyping other types of critical illness. For complete details on the use and execution of this protocol, please refer to Tamura et al. (2023).1
10015 Background: While cure rates for childhood acute lymphoblastic leukemia (ALL) exceed 90%, half of relapses arise in those originally classified with standard risk (SR) disease. Methods: We performed genome/transcriptome sequencing of diagnostic and germline samples of children with SR (n=1381) B-ALL or high-risk (HR) B-ALL with favorable cytogenetics ( ETV6: RUNX1 or double trisomy (DT) of chromosomes (chr) 4+10; n=115) to identify predictors of relapse. We used a case-control study to analyze 439 patients who relapsed and 1057 who remained in complete remission for > 5 years. Results: Genomic subtype was associated with relapse. Unbalanced ETV6:RUNX1 translocations were more common than balanced in relapse patients (OR=2.01, CI=1.25-3.20, P=0.002). Conversely, balanced TCF3:PBX1 translocations were more often associated with relapse than unbalanced in TCF3:PBX1 ALL (OR=0.11, CI=0.01-0.50, P=0.003). A striking finding was the high relapse rate in PAX5 altered ALL (57 of 116 cases (49%); OR=3.29, CI=2.16-5.01, P=3.49x10 -8 ). The nature of the heterogeneous PAX5 driver alterations of this subtype influenced relapse risk, with internal PAX5 amplifications and biallelic PAX5 alterations associated with the highest risk. Specific chr gains influenced outcome in hyperdiploid ALL, with gain of chr 10 and disomy of chr 7 associated with favorable outcome (OR=0.27, CI=0.17-0.42, P=8.02x10 -10 , St Jude Children’s Research Hospital (SJCRH) validation cohort: OR=0.22, CI=0.05-0.80, P=0.009), while disomy of chr 10 and 17 and gain of chr 6 were enriched in patients that relapsed (OR=7.16, CI=2.63-21.51, P=2.19x10 -5 ; SJCRH cohort: OR=21.32, CI=3.62-119.30, P=0.0004). Genomic alterations were also associated with relapse in a subtype-dependent manner, including alterations of INO80 in ETV6:RUNX1, IKZF1 and CREBBP in hyperdiploid, and FHIT in Ph-like ALL. Conclusions: Genetic subtype, aneuploidy patterns, and secondary genomic alterations influence risk of relapse in children otherwise classified with SR ALL, or HR ALL with favorable genetics. Comprehensive genomic analysis is required for optimal risk stratification and treatment allocation, and particularly to study reduction of therapy in the lowest risk patients. [Table: see text]
Somatic mitochondrial DNA (mtDNA) mutations are prevalent in tumors, yet defining their biological significance remains challenging due to the intricate interplay between selective pressure, heteroplasmy, and cell state. Utilizing bulk whole-genome sequencing data from matched tumor and normal samples from two cohorts of pediatric cancer patients, we uncover differences in the accumulation of synonymous and nonsynonymous mtDNA mutations in pediatric leukemias, indicating distinct selective pressures. By integrating single-cell sequencing (SCS) with mathematical modeling and network-based systems biology approaches, we identify a correlation between the extent of cell-state changes associated with tumor-enriched mtDNA mutations and the selective pressures shaping their distribution among individual leukemic cells. Our findings also reveal an association between specific heteroplasmic mtDNA mutations and cellular responses that may contribute to functional heterogeneity among leukemic cells and influence their fitness. This study highlights the potential of SCS strategies for distinguishing between pathogenic and passenger somatic mtDNA mutations in cancer.
Biallelic mutation in the DNA-damage repair gene NBN is the genetic cause of Nijmegen Breakage Syndrome, which is associated with predisposition to lymphoid malignancies. Heterozygous carriers of germline NBN variants may also be at risk for leukemia development, although this is much less characterized. We systematically examined the frequency of germline NBN variants in pediatric B-ALL and identified 25 putatively damaging NBN coding variants in 50 of 4,183 B-ALL patients. Compared with the frequency of NBN variants in 118,479 gnomAD non-cancer controls we found significant overrepresentation in pediatric B-ALL (p=0.004, OR=1.77). Most B-ALL-risk variants were missense and cluster within the NBN N-terminal domains. Using two functional assays, we verified 14 of 25 variants with severe loss-of-function phenotypes and thus classified these as pathogenic or likely pathogenic. Finally, we found that heterozygous germline NBN variant carriers showed similar survival outcomes relative to those with WT status. Taken together, our findings provide novel insights into the genetic predisposition to B-ALL, the impact of NBN variants on protein function and suggest that heterozygous NBN variant carriers may safely receive B-ALL therapy.
<p>Strong nuclear staining of mutant TP53-R337H protein in normal tissue of carriers who developed ACC.</p>
<p>Senescence associated β-galactosidase expression and ROS levels in mutant p53-R334H primary MEFs.</p>