Pathogenic TP53 germline variants cause young-onset breast cancer and other cancers of the Li-Fraumeni syndrome (LFS) spectrum, but the clinical consequences of partial-loss-of function TP53 variants are incompletely understood. In the consecutive cohort of Palestinian breast cancer patients of the Middle East Breast Cancer Study (MEBCS), breast cancer risk among TP53 p.R181C heterozygotes was 50% by age 50 years and 81% by age 80 years. In contrast, prevalence of pediatric cancers in the MEBCS was similar among first-degree relatives of TP53 p.R181C carriers (3/519 = 0.0058) and first-degree relatives of MEBCS patients with no pathogenic germline variant in any known breast cancer gene (7/1082 = 0.0065; odds ratio [OR] = 0.90, 95% confidence interval [CI] [0.23 to 3.49], Fisher P = .90 [2-tailed]). This result suggests that in families harboring this TP53 allele, genetic testing in children is unwarranted, and screening children for LFS tumors is unnecessary. More generally, some TP53 missense alleles can predispose to very high risk of breast cancer without pleiotropic effects.
Alzheimer’s Disease (AD) incidence is almost double in female than male, suggesting sex-specific AD risk genes remain unknown. We designed a statistical physics approach that exploits freely available but massive evolutionary and phylogenetic coupling data on sequence variation and speciation. These couplings lead to quantifiable values for the selection pressure exerted on the genes within a population. We may then compare a gene’s influence in sequenced cases vs controls cohorts and test the hypothesis that significant deviations identify genes linked to disease risk. In 4768 AD cases and 4689 healthy controls (HC), we discovered 122 genes under greater selection pressure (q < 0.01). These genes overlapped (p = 3.10 -5 ) and interacted (z = 7.16) with AD GWAS genes. They also interacted mutually (n = 57, p = 0.0019) and with AD-related processes (p = 1.0 -16 ). More than 50% of the genes exhibited dysregulation in AD brains in snRNAseq analysis, suggesting participation in pathogenic or protective processes in AD. Furthermore, expression of these genes correlated with increased or decreased deposition of plaques and tangles in patient brains. Moreover, the candidates are enriched in modifiers of neurodegeneration in Drosophila : knockdown or overexpression of 64 genes ameliorated or worsened age-dependent neuronal dysfunction (p<0.05). Robustness to down-sampling allowed to analyze smaller, sex-separated cohorts. We identified 82 genes in males and 69 genes in females (15 genes overlapped, p < 10 -53 ), indicating shared as well as sex-specific AD mechanisms. The male and the female genes overlapped (p = 6.5 -6 and 1.5 -4 ) and interacted (z = 5.63 and 6.18) with the AD genes. Remarkably, using these gene sets as features for predicting AD risk in separate training and testing cohorts successfully differentiated between AD cases and controls with very high accuracy, even when blinded to APOE genotype. Notably, we predicted the risk with AUCs of 0.83 with APOE, 0.83, and 0.82 in combined, males, and females, respectively. A new statistical physics approach discovered male and female AD genes, predicting AD risk with very high accuracy. These results identify further genetic differences leading to AD in males and females, and show the power of quantitative phylogenetics to probe complex human diseases.
BACKGROUND:Molecular genetic diagnoses are critical to prevention and treatment of inherited polyposis and colorectal cancer. 19 genes responsible for these conditions are known, but many severely affected patients and families remain unsolved. Cryptic intronic variants that alter splicing of these genes and incomplete characterisation of recessive predisposition contribute to these diagnostic gaps. METHODS:Adaptive sampling long-read DNA sequencing targeted to 19 colon cancer genes, paired with direct long-read RNA whole-transcriptome sequencing, was undertaken for four patients referred for deficiency of mismatch repair proteins and/or familial polyposis, for whom multigene panel testing yielded negative or uncertain germline results. RESULTS:Genetic diagnoses were obtained for all four patients. Each patient carried a cryptic intronic germline variant in a different colon cancer gene. The variants abrogated splicing by various mechanisms, all leading to loss of gene function. Patient 1 was heterozygous for intronic insertion into MSH2 of an Alu element, leading to extension of transcription into the affected intron and a stop. Patient 2 was heterozygous for deep intronic insertion into APC of a Long Interspersed Nuclear Element (LINE), creating a pseudoexon and a stop. Patient 3 was compound heterozygous at MLH3, including a cryptic intronic substitution leading to exon skipping and a stop. Patient 4 was compound heterozygous at MUTYH, including a deep intronic deletion yielding an extremely short intron and transcriptional loss of an exon encoding a critical protein domain. CONCLUSION:Paired long-read DNA and RNA sequencing can enhance diagnostic yield through detection of cryptic intronic variants that impact cancer predisposition.
For many families severely affected with breast cancer, no inherited causal allele has been detected in any tumor suppressor gene. In an effort to understand the genetics underlying breast cancer in these families, we evaluated 136 such families for coinheritance of breast cancer with each of 79 common variants reported as high-confidence "risk alleles" for breast cancer by meta-analyses of genome-wide association studies. Simulations based on allele frequencies and family structures revealed one (and only one) of these 79 variants to cosegregate with breast cancer in the families significantly more frequently than expected by chance. This variant (rs2046210) is located 180 kb proximal to ESR1, encoding the estrogen receptor alpha. Reporter assays in MCF7 cells revealed enhancement by the genomic segment at this site of activity of ESR1 promoters, but no difference in effect among alternative haplotypes. In contrast, the 600 kb genomic region including ESR1 and rs2046210 harbored 11 rare variants, each of which cosegregated with breast cancer in one or a few families. For 9 of these 11 variants, reporter assays indicated significant allele-specific effects on ESR1 promoters, with the breast-cancer-linked allele of each variant yielding higher promoter activity. At the site with the most striking effect, the breast-cancer-linked allele was associated with increased binding by transcription factor AP2-gamma TFAP2C in both MCF7 and T47D cells. These results demonstrate coinheritance with breast cancer of rare alleles that increase activity of ESR1 promoters, and suggest that rare ESR1 regulatory alleles may contribute to inherited predisposition to breast cancer.
The vast majority of deeply intronic genomic variants are benign, but some extremely rare or private deep intronic variants lead to exonification of intronic sequence with abnormal transcriptional consequences. Damaging variants of this class are likely underreported as causes of disease for several reasons: Most clinical DNA and RNA testing does not include full intronic sequences; many of these variants lie in complex repetitive regions that cannot be aligned from short-read whole-genome sequence; and, until recently, consequences of deep intronic variants were not accurately predicted by in silico tools. We evaluated the frequency and consequences of rare deep intronic variants for families severely affected with breast, ovarian, pancreatic, and/or metastatic prostate cancer, but with no causal variant identified by any previous genomic or cDNA-based approach. For 10 tumor-suppressor genes, we used multiplexed adaptive sampling long-read DNA sequencing and cDNA sequencing, based on patient-derived DNA and RNA, to systematically evaluate deep intronic variation. We identified all variants across the full genomic loci of targeted genes, applied the in silico tools SpliceAI and Pangolin to predict variants of functional consequence, and then carried out long-read cDNA sequencing to identify aberrant transcripts. For eight of the 120 (6%) previously unsolved families, rare deep intronic variants in BRCA1 , PALB2 , and ATM create intronic pseudoexons that are spliced into transcripts, leading to premature truncations. These results suggest that long-read DNA and cDNA sequencing can be integrated into variant discovery, with strategies for accurately characterizing pathogenic variants.
Table of non-BRCA mutations with clinical characteristics
Background Coronary artery disease is a primary cause of death around the world, with both genetic and environmental risk factors. Although genome‐wide association studies have linked >100 unique loci to its genetic basis, these only explain a fraction of disease heritability. Methods and Results To find additional gene drivers of coronary artery disease, we applied machine learning to quantitative evolutionary information on the impact of coding variants in whole exomes from the Myocardial Infarction Genetics Consortium. Using ensemble‐based supervised learning, the Evolutionary Action–Machine Learning framework ranked each gene's ability to classify case and control samples and identified 79 significant associations. These were connected to known risk loci; enriched in cardiovascular processes like lipid metabolism, blood clotting, and inflammation; and enriched for cardiovascular phenotypes in knockout mouse models. Among them, INPP5F and MST1R are examples of potentially novel coronary artery disease risk genes that modulate immune signaling in response to cardiac stress. Conclusions We concluded that machine learning on the functional impact of coding variants, based on a massive amount of evolutionary information, has the power to suggest novel coronary artery disease risk genes for mechanistic and therapeutic discoveries in cardiovascular biology, and should also apply in other complex polygenic diseases.
Importance:In the US, most childhood-onset bilateral sensorineural hearing loss is genetic, with more than 120 genes and thousands of different alleles known. Primary treatments are hearing aids and cochlear implants. Genetic diagnosis can inform progression of hearing loss, indicate potential syndromic features, and suggest best timing for individualized treatment. Objective:To identify the genetic causes of childhood-onset hearing loss and characterize severity, progression, and cochlear implant success associated with genotype in a single large clinical cohort. Design, Setting, and Participants:This cross-sectional analysis (genomics) and retrospective cohort analysis (audiological measures) were conducted from 2019 to 2022 at the otolaryngology and audiology clinics of Seattle Children's Hospital and the University of Washington and included 449 children from 406 families with bilateral sensorineural hearing loss with an onset younger than 18 years. Data were analyzed between January and June 2022. Main Outcomes and Measures:Genetic diagnoses based on genomic sequencing and structural variant analysis of the DNA of participants; severity and progression of hearing loss as measured by audiologic testing; and cochlear implant success as measured by pediatric and adult speech perception tests. Hearing thresholds and speech perception scores were evaluated with respect to age at implant, months since implant, and genotype using a multivariate analysis of variance and covariance. Results:Of 406 participants, 208 (51%) were female, 17 (4%) were African/African American, 32 (8%) were East Asian, 219 (54%) were European, 53 (13%) were Latino/Admixed American, and 16 (4%) were South Asian. Genomic analysis yielded genetic diagnoses for 210 of 406 families (52%), including 55 of 82 multiplex families (67%) and 155 of 324 singleton families (48%). Rates of genetic diagnosis were similar for children of all ancestries. Causal variants occurred in 43 different genes, with each child (with 1 exception) having causative variant(s) in only 1 gene. Hearing loss severity, affected frequencies, and progression varied by gene and, for some genes, by genotype within gene. For children with causative mutations in MYO6, OTOA, SLC26A4, TMPRSS3, or severe loss-of-function variants in GJB2, hearing loss was progressive, with losses of more than 10 dB per decade. For all children with cochlear implants, outcomes of adult speech perception tests were greater than preimplanted levels. Yet the degree of success varied substantially by genotype. Adjusting for age at implant and interval since implant, speech perception was highest for children with hearing loss due to MITF or TMPRSS3. Conclusions and Relevance:The results of this cross-sectional study suggest that genetic diagnosis is now sufficiently advanced to enable its integration into precision medical care for childhood-onset hearing loss.
Journal Article Corrected proof A paradoxical genotype-phenotype relationship: Low level of GOSR2 translation from a non-AUG start codon in a family with profound hearing loss Get access Amal Aburayyan, Amal Aburayyan Department of Genome Sciences and Department of Medicine, University of Washington, Seattle, WA, USADepartment of Human Molecular Genetics and Biochemistry, Faculty of Medicine and Sagol School of Neuroscience, Tel Aviv University, Tel Aviv, IsraelHereditary Research Laboratory, Department of Biology, Bethlehem University, Bethlehem, Palestine Search for other works by this author on: Oxford Academic PubMed Google Scholar Ryan J Carlson, Ryan J Carlson Department of Genome Sciences and Department of Medicine, University of Washington, Seattle, WA, USA Search for other works by this author on: Oxford Academic PubMed Google Scholar Grace N Rabie, Grace N Rabie Hereditary Research Laboratory, Department of Biology, Bethlehem University, Bethlehem, Palestine Search for other works by this author on: Oxford Academic PubMed Google Scholar Ming K Lee, Ming K Lee Department of Genome Sciences and Department of Medicine, University of Washington, Seattle, WA, USA Search for other works by this author on: Oxford Academic PubMed Google Scholar Suleyman Gulsuner, Suleyman Gulsuner Department of Genome Sciences and Department of Medicine, University of Washington, Seattle, WA, USA Search for other works by this author on: Oxford Academic PubMed Google Scholar Tom Walsh, Tom Walsh Department of Genome Sciences and Department of Medicine, University of Washington, Seattle, WA, USA Search for other works by this author on: Oxford Academic PubMed Google Scholar Karen B Avraham, Karen B Avraham Department of Human Molecular Genetics and Biochemistry, Faculty of Medicine and Sagol School of Neuroscience, Tel Aviv University, Tel Aviv, Israel Search for other works by this author on: Oxford Academic PubMed Google Scholar Moien N Kanaan, Moien N Kanaan Hereditary Research Laboratory, Department of Biology, Bethlehem University, Bethlehem, Palestine Search for other works by this author on: Oxford Academic PubMed Google Scholar Mary-Claire King Mary-Claire King Department of Genome Sciences and Department of Medicine, University of Washington, Seattle, WA, USA To whom correspondence should be addressed at: Health Sciences Room K-160, University of Washington, 1959 NE Pacific Street, Seattle, WA 98195-7720, USA. Tel: 206-616-4294; Fax: 206-616-4295; Email: mcking@uw.edu Search for other works by this author on: Oxford Academic PubMed Google Scholar Human Molecular Genetics, ddad066, https://doi.org/10.1093/hmg/ddad066 Published: 19 April 2023 Article history Received: 25 March 2023 Revision received: 25 March 2023 Accepted: 10 April 2023 Published: 19 April 2023 Corrected and typeset: 04 May 2023
The incidence of Alzheimer’s Disease in females is almost double that of males. To search for sex-specific gene associations, we build a machine learning approach focused on functionally impactful coding variants. This method can detect differences between sequenced cases and controls in small cohorts. In the Alzheimer’s Disease Sequencing Project with mixed sexes, this approach identified genes enriched for immune response pathways. After sex-separation, genes become specifically enriched for stress-response pathways in male and cell-cycle pathways in female. These genes improve disease risk prediction in silico and modulate Drosophila neurodegeneration in vivo. Thus, a general approach for machine learning on functionally impactful variants can uncover sex-specific candidates towards diagnostic biomarkers and therapeutic targets.
Exome sequencing of genes associated with heritable thoracic aortic disease (HTAD) failed to identify a pathogenic variant in a large family with Marfan syndrome (MFS). A genome-wide linkage analysis for thoracic aortic disease identified a peak at 15q21.1, and genome sequencing identified a novel deep intronic FBN1 variant that segregated with thoracic aortic disease in the family (LOD score 2.7) and was predicted to alter splicing. RT-PCR and bulk RNA sequencing of RNA harvested from fibroblasts explanted from the affected proband revealed an insertion of a pseudoexon between exons 13 and 14 of the FBN1 transcript, predicted to lead to nonsense mediated decay (NMD). Treating the fibroblasts with an NMD inhibitor, cycloheximide, greatly improved the detection of the pseudoexon-containing transcript. Family members with the FBN1 variant had later onset aortic events and fewer MFS systemic features than typical for individuals with haploinsufficiency of FBN1. Variable penetrance of the phenotype and negative genetic testing in MFS families should raise the possibility of deep intronic FBN1 variants and the need for additional molecular studies.
Supplementary Tables 1-4, Figure 1 from Contribution of Inherited Mutations in the BRCA2-Interacting Protein PALB2 to Familial Breast Cancer
PDF file 80K, Cases with deleterious germline mutations, somatic HR mutations, and somatic PTEN mutations
To study how genetic backgrounds modulate neurodegenerative risk requires genotype-phenotype association methods effective in small patient cohorts. We developed a genome sequence analysis framework rooted in Statistical Physics, which yields a reliable measure of gene functional impact in any chosen population. When comparing two populations, such as cases and controls, shifts in gene importance identify those likely to drive phenotype differences. In repeated case-control genome sequence studies of Alzheimer’s Disease (AD) focused on a single sex, APOE allele, or ethnic ancestry, we identified gene sets that met criteria for success, including prior GWAS studies, post-mortem of AD brain tissues expression, live Drosophila experiments, and risk modeling. Statistical Physics of fitness landscape is a new tool to characterize genes and mutations that enhance or protect from AD, powerful enough to identify similarities, complementarities, and differences in risk gene in target subgroups of about 1000 subjects.
Moyamoya disease, a cerebrovascular disease leading to strokes in children and young adults, is characterized by progressive occlusion of the distal internal carotid arteries and the formation of collateral vessels. Altered genes play a prominent role in the aetiology of moyamoya disease, but a causative gene is not identified in the majority of cases. Exome sequencing data from 151 individuals from 84 unsolved families were analysed to identify further genes for moyamoya disease, then candidate genes assessed in additional cases (150 probands). Two families had the same rare variant in ANO1, which encodes a calcium-activated chloride channel, anoctamin-1. Haplotype analyses found the families were related, and ANO1 p.Met658Val segregated with moyamoya disease in the family with an LOD score of 3.3. Six additional ANO1 rare variants were identified in moyamoya disease families. The ANO1 rare variants were assessed using patch-clamp recordings, and the majority of variants, including ANO1 p.Met658Val, displayed increased sensitivity to intracellular Ca2+. Patients harbouring these gain-of-function ANO1 variants had classic features of moyamoya disease, but also had aneurysm, stenosis and/or occlusion in the posterior circulation. Our studies support that ANO1 gain-of-function pathogenic variants predispose to moyamoya disease and are associated with unique involvement of the posterior circulation.