Long-read sequencing (LRS) and diploid genome assembly have enabled nearly complete structural variant (SV) discovery. Using 293 nearly complete genomes, we characterize the full spectrum of genetic variation and show that while 99% of the variants between any two genomes are single base-pair substitutions, 88% of the euchromatic variant base pairs are SVs, including insertions, deletions, duplications, and inversions. We identify 24 gene-rich regions subject to megabase-scale variation, 2,293 potentially unstable tandem repeats, and 890 novel expression quantitative trait loci associated with SVs in humans. Expanding to 1,218 LRS samples from the 1000 Genomes Project and applying a newly developed cross-platform breakpoint evaluation tool, BoostSV, we construct a nonredundant callset comprising 614,522 SVs. We demonstrate the utility of this population-level SV reference callset by filtering >99% of the common variation from 44 unsolved LRS probands from the Undiagnosed Diseases Network to discover likely disease-causing SVs. Second, we genotype 1,053 high-impact biallelic SVs from the pangenome callset in 232,090 samples from All of Us and discover 105 SVs with significant associations, including 26% where the SV is the lead variant. This publicly available pangenome SV resource will drive new disease associations and further our understanding of the missing heritability of human genetic disease.
Identifying pathogenic non-coding variants that contribute to Mendelian conditions remains challenging, as the functional impact of these variants on gene function is often unknown. We present IsoRanker, a long-read transcriptome sequencing-based framework that prioritizes functionally relevant variants by detecting genes and isoforms with outlier expression, allelic imbalance, and/or nonsense-mediated decay (NMD). We generated paired cycloheximide-treated and untreated fibroblast transcriptomes from 31 individuals (3 individuals with known transcript-altering rare variants and 28 individuals with unsolved conditions) and linked transcripts to phased long-read genomes. IsoRanker successfully recovered known transcript alterations in this cohort, and exploratory subsampling analyses suggested that their prioritization was largely preserved down to cohorts of 11 individuals and ∼5 million full-length transcripts per individual. Performance was dependent upon de novo isoform caller choice, particularly for NMD-sensitive and previously unannotated isoforms. Among 28 previously unsolved cases, IsoRanker deprioritized 8 out of 10 fibroblast-expressed candidate splice-site variants while nominating 4 new leads. In one individual, IsoRanker prioritized HARS1, revealing bi-allelic non-coding variants that together produced a partial HARS1 loss of function and informed targeted therapy in this individual. These findings support long-read, NMD-aware transcriptomics with IsoRanker as an effective approach for generating isoform-level functional evidence, improving classification of non-coding variants and supporting the diagnosis of individuals with rare genetic conditions.
Interpreting protein-truncating variants predicted to escape nonsense-mediated decay (NMDe) based on the 50-bp rule is challenging due to their variable consequences. The Clinical Genome Resource, Cancer Genomics Consortium, and Variant Interpretation for Cancer Consortium recommendations focus on tumor suppressor genes where NMDe variants may result in loss-of-function. However, guidance for interpreting NMDe variants in oncogenes and dual-function genes remains limited. To address this gap, we screened the Catalogue of Somatic Mutations in Cancer, focusing on oncogenes and dual-function genes with ≥10 NMDe variants supported by published functional studies. This analysis prioritized 15 genes exhibiting two distinct NMDe patterns resulting in gene activation. The first pattern and gene examples involve NMDe variants causing loss of C-terminal regulatory regions that mediate protein inhibition and/or degradation, frequently manifesting as attenuated cell surface receptor internalization/degradation (e.g., CSF3R) or increased intracellular protein stability (e.g., CCND3). The second is driven exclusively by frameshift NMDe variants that generate novel peptide fragments that alter protein interactions (e.g., CALR). Interrogation of these genes in the St. Jude Children's Research Hospital clinical genomics pediatric cohort identified 119 NMDe variants across 8 genes in ~3% (113/3,492) of unique patient samples, with 118 of these classified as likely oncogenic or higher (99 relating to the first pattern and 19 to the second). One variant was classified as "uncertain significance". Our data emphasize the need to integrate gene function, variant type, and effects on the C-terminus to comprehensively evaluate somatic NMDe variants and predict their consequences across adult and pediatric cancer cohorts.
Journal Article Genomic Data and Privacy Get access Candace T Myers, Candace T Myers Laboratory Genetics and Genomics Fellow, Department of Laboratory Medicine and Pathology, University of Washington, Seattle, WA, United States Address correspondence to: C.T.M. at Department of Laboratory Medicine and Pathology, University of Washington, Box 357470, 1959 NE Pacific St., Seattle, WA 98195-7110, United States. E-mail [email protected]. R.D.K. at Department of Laboratory Medicine and Pathology, University of Washington, Box 357110, 1959 NE Pacific St., Seattle, WA 98195-7110, United States. E-mail [email protected]. Search for other works by this author on: Oxford Academic Google Scholar Runjun D Kumar, Runjun D Kumar Assistant Professor, Department of Laboratory Medicine and Pathology, University of Washington, Seattle, WA, United States Address correspondence to: C.T.M. at Department of Laboratory Medicine and Pathology, University of Washington, Box 357470, 1959 NE Pacific St., Seattle, WA 98195-7110, United States. E-mail [email protected]. R.D.K. at Department of Laboratory Medicine and Pathology, University of Washington, Box 357110, 1959 NE Pacific St., Seattle, WA 98195-7110, United States. E-mail [email protected]. Search for other works by this author on: Oxford Academic Google Scholar Lisa Pilgram, Lisa Pilgram Postdoctoral Fellow, Electronic Health Information Laboratory, University of Ottawa, Ottawa, ON, Canada Search for other works by this author on: Oxford Academic Google Scholar Luca Bonomi, Luca Bonomi Assistant Professor, Department of Biomedical Informatics, Vanderbilt University, Nashville, TN, United States Search for other works by this author on: Oxford Academic Google Scholar Mara Thomas, Mara Thomas Senior Data Scientist, Roche Informatics Solutions, F. Hoffmann-La-Roche AG, Basel, Switzerland Search for other works by this author on: Oxford Academic Google Scholar Obi L Griffith, Obi L Griffith Associate Professor, Departments of Medicine and Genetics, Washington University, St. Louis, MO, United States https://orcid.org/0000-0002-0843-4271 Search for other works by this author on: Oxford Academic Google Scholar Stephanie M Fullerton, Stephanie M Fullerton Professor, Department of Bioethics and Humanities, University of Washington School of Medicine, Seattle, WA, United States Search for other works by this author on: Oxford Academic Google Scholar Richard A Gibbs Richard A Gibbs Wofford Cain Chair and Professor, Department of Molecular and Human Genetics, Baylor College of Medicine, Houston, TX, United States Search for other works by this author on: Oxford Academic Google Scholar Clinical Chemistry, Volume 71, Issue 1, January 2025, Pages 10–17, https://doi.org/10.1093/clinchem/hvae184 Published: 03 January 2025 Article history Received: 02 August 2024 Accepted: 07 October 2024 Published: 03 January 2025
Variant-level functional data are a core component of clinical variant classification and can aid in reinterpreting variants of uncertain significance (VUSs). However, the usage of functional data by genetics professionals is currently unknown. An online survey was developed and distributed in the spring of 2024 to individuals actively engaged in variant interpretation. Quantitative and qualitative methods were used to assess responses. 190 eligible individuals responded, with 93% reporting interpreting 26 or more variants per year. The median respondent reported 11-20 years of experience. The most common professional roles were laboratory medical geneticists (23%) and variant review scientists (23%). 77% reported using functional data for variant interpretation in a clinical setting, and overall, respondents felt confident assessing functional data. However, 67% indicated that functional data for variants of interest were rarely or never available, and 91% considered insufficient quality metrics or confidence in the accuracy of data as barriers to their use. 94% of respondents noted that better access to primary functional data and standardized interpretation of functional data would improve usage. Respondents also indicated that handling conflicting functional data is a common challenge in variant interpretation that is not performed in a systematic manner across institutions. The results from this survey showed a demand for a comprehensive database with reliable quality metrics to support the use of functional evidence in clinical variant interpretation. The results also highlight a need for guidelines regarding how putatively conflicting functional data should be used for variant classification.
Clinical interpretation of genetic variants depends on robust and representative population genomic datasets as a source of evidence for benign or pathogenic effect. In 2023, a new version of the widely used Genome Aggregation Database (gnomAD 4.0) was released, encompassing 807,000 individual exomes and genomes, or almost three-fold more than the widely-used version 2.1.1. Furthermore, the AllofUs research program made available whole-genome sequencing for nearly 250,000 individuals. The effects of these new databases on variant interpretation are unknown.
Recognition of patients with multiple diagnoses, and the unique challenges they pose to clinicians and laboratorians, is increasing rapidly as genome-wide genetic testing grows in prevalence. We describe a unique patient with dual diagnoses of PDCD10-related cerebral cavernous malformations and ETV6-related thrombocytopenia with associated neutropenia. She presented with brain abscesses as an infant, which is highly atypical for these disorders in isolation. Confirming her diagnoses depended on thorough phenotyping both during and after her acute illness. Furthermore, the causative variant in ETV6 is a novel single-exon deletion that required multiple modalities with manual review to confirm, including unique use of polymorphic nucleotides in trio exome data. She illustrates the special challenges of patients with multiple diagnoses, and the multiple tools clinicians and laboratorians must use to treat them.
PURPOSE:Genome sequencing (GS) may shorten the diagnostic odyssey for patients, but clinical experience with this assay in nonresearch settings remains limited. Texas Children's Hospital began offering GS as a clinical test to admitted patients in 2020, providing an opportunity to study GS utilization, possibilities for test optimization, and testing outcomes.METHODS:We retrospectively reviewed GS orders for admitted patients for a nearly 3-year period from March 2020 through December 2022. We gathered anonymized clinical data from the electronic health record to answer the study questions.RESULTS:The diagnostic yield over 97 admitted patients was 35%. The majority of GS clinical indications were neurologic or metabolic (61%) and most patients were in intensive care (58%). Tests were often characterized as candidates for intervention/improvement (56%), frequently because of redundancy with prior testing. Patients receiving GS without prior exome sequencing (ES) had higher diagnostic rates (45%) than the cohort as a whole. In 2 cases, GS revealed a molecular diagnosis that is unlikely to be detected by ES.CONCLUSION:The performance of GS in clinical settings likely justifies its use as a first-line diagnostic test, but the incremental benefit for patients with prior ES may be limited.
Background Post-zygotic genetic variants in cancer genes cause a group of syndromic and non-syndromic conditions termed somatic mosaic disorders (SMD). Tissue-based deep sequencing is becoming the diagnostic method of choice given the diverse variant allele fractions (VAF) involved. We report the diagnostic yield of high-depth tissue-based next-generation sequencing (NGS) in 57 patients with SMD. Methods We retrospectively identified patients referred for testing with clinical diagnosis of SMD (e.g., CLOVES) or characteristic clinical manifestations. 35 patients were tested with a 124 gene solid tumor panel at 500-1000x coverage. Beginning in 2021, 22 patients were tested with a 29 gene SMD panel. Testing was predominantly performed on formalin-fixed paraffin-embedded (FFPE) tissue. Results An AMP/ASCO/CAP tier I/II variant was reported in 25/57 (43.8%) patients, including 17/35 (48.6%) patients tested with the solid tumor panel, and 8/22 (36.3%) tested with the SMD panel (p=0.37). Six patients had a variant of unknown significance at VAF consistent with mosaicism. Diagnosed disorders were related to PIK3CA (n=15), KRAS (3), GNAQ (2), AKT1, BRAF, GNA11, GNAS, MAP2K1, PTEN. The median VAF was 9% (4-80%), and 24/26 tier I/II variants are reported in COSMIC. 9/18 patients with prior germline genetic testing received a new diagnosis. 11/21 diagnosed patients who are following at our center are receiving targeted therapies. Conclusion Patients with SMDs can benefit from molecular diagnoses when clinical, radiologic and pathologic findings are inconclusive. In a significant proportion of patients, high-depth NGS of lesional tissue can reveal somatic mosaic alterations at low allelic fractions amenable to targeted therapy. Post-zygotic genetic variants in cancer genes cause a group of syndromic and non-syndromic conditions termed somatic mosaic disorders (SMD). Tissue-based deep sequencing is becoming the diagnostic method of choice given the diverse variant allele fractions (VAF) involved. We report the diagnostic yield of high-depth tissue-based next-generation sequencing (NGS) in 57 patients with SMD. We retrospectively identified patients referred for testing with clinical diagnosis of SMD (e.g., CLOVES) or characteristic clinical manifestations. 35 patients were tested with a 124 gene solid tumor panel at 500-1000x coverage. Beginning in 2021, 22 patients were tested with a 29 gene SMD panel. Testing was predominantly performed on formalin-fixed paraffin-embedded (FFPE) tissue. An AMP/ASCO/CAP tier I/II variant was reported in 25/57 (43.8%) patients, including 17/35 (48.6%) patients tested with the solid tumor panel, and 8/22 (36.3%) tested with the SMD panel (p=0.37). Six patients had a variant of unknown significance at VAF consistent with mosaicism. Diagnosed disorders were related to PIK3CA (n=15), KRAS (3), GNAQ (2), AKT1, BRAF, GNA11, GNAS, MAP2K1, PTEN. The median VAF was 9% (4-80%), and 24/26 tier I/II variants are reported in COSMIC. 9/18 patients with prior germline genetic testing received a new diagnosis. 11/21 diagnosed patients who are following at our center are receiving targeted therapies. Patients with SMDs can benefit from molecular diagnoses when clinical, radiologic and pathologic findings are inconclusive. In a significant proportion of patients, high-depth NGS of lesional tissue can reveal somatic mosaic alterations at low allelic fractions amenable to targeted therapy.
Stroke causes significant disability and is a common cause of death worldwide. Previous studies have estimated that 1%-5% of stroke is attributable to monogenic etiologies. We set out to assess the utility of clinical exome sequencing (ES) in the evaluation of stroke. We retrospectively analyzed 124 individuals who received ES at the Baylor Genetics reference lab between 2012 and 2021 who had stroke as a major part of their reported phenotype. Ages ranged from 10 days to 69 years. 8.9% of the cohort received a diagnosis, including 25% of infants less than 1 year old; an additional 10.5% of the cohort received a probable diagnosis. We identified several syndromes that predispose to stroke such as COL4A1-related brain small vessel disease, homocystinuria caused by CBS mutation, POLG-related disorders, TTC19-linked mitochondrial disease, and RNASEH2A associated Aicardi-Goutieres syndrome. We also observed pathogenic variants in NSD1, PKHD1, HRAS, and ATP13A2, which are genes rarely associated with stroke. Although stroke is a complex phenotype with varying pathologies and risk factors, these results show that use of exome sequencing can be highly relevant in stroke, especially for those presenting <1 year of age.
The tumor suppressor p53 has well known roles in cancer development and germline cancer predisposition disorders, but increasing evidence supports the role of activation of this transcription factor in the pathogenesis of inherited bone marrow failure and chromosomal instability disorders. Here we report a patient with red cell aplasia, which was steroid responsive, as well as intellectual disability, seizures, microcephaly, short stature, cellular radiosensitivity, and normal telomere lengths, who had a germline heterozygous C-terminal frameshift variant in TP53 similar to others that activate the transcription factor. This is the third reported individual with a germline p53 activation syndrome, with several unique features that refine the clinical disease associated with these variants.
Escape from nonsense mediated mRNA decay (NMD-) can produce activated or inactivated gene products, and bias in rates of escape can identify functionally important genes in germline disease. We hypothesized that the same would be true of cancer genes, and tested for NMD- bias within The Cancer Genome Atlas pan-cancer somatic mutation dataset. We identify 29 genes that show significantly elevated or suppressed rates of NMD-. This novel approach to cancer gene discovery reveals genes not previously cataloged as potentially tumorigenic, and identifies many potential driver mutations in known cancer genes for functional characterization.
Pathogenic variants in non-coding regions of genes encoding enzymes or transporters of the urea cycle can lead to urea cycle disorders (UCDs). However, not all commercially available testing platforms interrogate these regions. Here, we used a gene panel based on massively parallel sequencing (MPS) in 10 individuals with clinical or pedigree-based evidence of a proximal UCD but without a molecular confirmation of the diagnosis. We identified causal variant(s) in 5 of 10 individuals, including in 3 of 7 individuals in whom prior molecular testing was unrevealing. We show that a deep-intronic pathogenic variant in OTC, c.540+265G>A, is an important cause of ornithine transcarbamylase (OTC) deficiency.
A previously healthy and fully vaccinated 5-year-old girl presented in October 2018 with 19 days of fever, bitemporal headaches, fatigue, 6-pound weight loss, and recent diagnosis of streptococcal pharyngitis. Her fevers occurred daily, primarily in the afternoons and early evenings, and temperature measured orally as high as 104°F. The patient’s review of systems was negative for sore throat, oral ulcers, rashes, arthralgia, gastrointestinal symptoms, or lymphadenopathy. Her family history was significant for a grandmother with rheumatoid arthritis. Her exposure history was notable for exposure to cows, unvaccinated cats, and mosquitos. She had no history of recent travel. Patient’s vital signs on admission were within normal limits, and she was afebrile throughout admission. Physical examination was normal, with pertinent negatives including no lymphadenopathy, hepatosplenomegaly, or rash. On the second day of hospitalization, the patient developed left-sided facial droop consistent with Bell’s palsy (Figure 1). Diagnostic evaluation was notable for an elevated C-reactive protein of 8.1 mg/dL (reference range, <1.0 mg/dL) and erythrocyte sedimentation rate of 70 mm/h (reference range, 0-20 mm/h). Blood cultures remained negative. To evaluate the patient’s new facial droop, she underwent computed tomography head scan without contrast, which demonstrated no acute intracranial process. Infectious Diseases was consulted and recommended additional serologic and polymerase chain reaction testing for bacterial and viral pathogens (including for Rickettsial infection, Lyme disease, West Nile virus, varicella, mumps, Epstein-Barr virus, cytomegalovirus, and herpes simplex viruses 1 and 2), high-resolution abdominal ultrasound, and a dilated ophthalmologic examination. An abdominal ultrasound was notable for multiple subcentimeter hypoechoic lesions throughout her liver and spleen, consistent with disseminated cat scratch disease (CSD; Figure 2).
Dear Editor, The transmembrane protease serine 2-ETS-related gene (TMPRSS2-ERG) fusion occurs in >50% of prostate cancers, leading to upregulation of the transcription factor ERG and tumor cell sensitivity to androgen.1 The fusion is associated with more aggressive manifestations of prostate cancer.2 Here, we developed a gene signature that recapitulated the pathway activity downstream of the TMPRSS2-ERG fusion event and applied it to predict patient prognosis in prostate cancer. The gene signature was defined by performing a logistic regression on every gene in The Cancer Genome Atlas prostate adenocarcinoma (TCGA-PRAD) dataset. TMPRSS2-ERG fusion status was used as the response variable, while gene expression level, age, and Gleason score were used as predictor variables. Based on these results, the 700 most significant genes were selected and for each of them a weight within [−1, 1] was assigned with sign indicating up- and downregulation, respectively. Given a new prostate cancer gene expression dataset, the weighted gene signature was applied to calculate sample-specific scores for all samples by using a rank-based statistic method named BASE.3 The resultant signature scores recapitulate the deregulated pathways downstream of the TMPRSS2-ERG fusion event. GO enrichment analysis indicates that genes associated with hormone secretion and regulation were highly enriched in this signature (Table S3). A detailed description about the signature can be found in the Supporting Information Materials. First, we tested whether the signature can identify tumors with TMPRSS2-ERG fusion in the TCGA and two additional prostate cancer datasets, the Sboner data (GSE16560) and the Setlur data (GSE8402) (Table S1). We found that fusion-positive samples exhibited significantly higher signature scores than fusion-negative samples in all datasets (Figure 1A and Figure S1A and B). When the signature score was used to classify the two sample groups, a fairly high accuracy was achieved in all datasets as shown by the response operating characteristic curves and area under the curve scores (Figure 1B). A direct consequence of TMPRSS2-ERG fusion is the upregulated expression of ERG. Indeed, we found that the signature score is highly correlated with the ERG mRNA level in both fusion-positive and fusion-negative samples (Figure S1C and D). Interestingly, a small subset of TMPRSS2-ERG fusion-negative samples was associated with high signature scores (Figure S1A). Of these samples, nine can be explained by the fusion of ERG with other genes such as SLC45A3. Signature scores of these samples (ERG-Other) were lower than samples with TMPRSS2-ERG fusions, but were significantly higher than the samples with no ERG fusion (Figure 1C). These results indicate that genomic events alternative to TMPRSS2-ERG fusion might deregulate the same downstream pathways and thus result in similar gene expression patterns. Second, we examined the ability of signature score to predict patient prognosis using the Sboner data, for which disease-specific survival information was available. We calculated the signature scores for all samples and stratified patients into two groups using the median score as the threshold. Patients with high scores have significantly poorer prognosis (P = 7 × 10−05) than those with low scores (Figure S2A). When this analysis was restricted to samples without TMPRSS2-ERG fusion, the same result was observed: high score was associated with poor prognosis (P = .002, Figure 1D). Interestingly, in fusion-positive samples high score was associated with good prognosis (P = .02, Figure 1D), in contrast to the negative association observed in fusion-negative samples. Similar results but lower significance was obtained when ERG gene expression was used to stratify prostate cancer patients (Figure S2B-D). Gleason score has been defined to categorize morphological differences and found to have high prognostic value in prostate cancer.4 The most common Gleason score at diagnosis5 and within this dataset is 7 (Figure S3A, Supporting Information Table 1), therefore we investigated the prognostic value of our signature in Gleason 7 (G7) samples. The result indicated that high score was associated with poor prognosis in all G7 (P = .04) and fusion-negative G7 (P = .02) samples (Figure 1E, Figure S3B). In addition, signature score can accurately differentiate indolent from lethal tumor samples (Figure 1F), with lethal samples exhibiting significantly higher scores than indolent samples (Figure S3C-F). Finally, we examined the impact of TMPRSS2-ERG fusion on intratumoral immune infiltration in order to understand why fusion-positive patients have poor survival. Previous studies have reported the association of immune infiltration with cancer development and prognosis in prostate cancer.6 The leukocyte abundance in TCGA prostate cancer samples was obtained from the Thorsson et al's study.7 A comparison between TMPRSS2-ERG fusion-positive and fusion-negative samples indicated significantly lower leukocyte level in the former group (Figure 2A). Then, we applied a computation method8 to infer the infiltration levels of six immune cell types (naïve B cell, memory B cell, CD8+ T cell, CD4+ T cell, natural killer cell, and monocyte) in all TCGA prostate cancer samples based on their gene expression profiles. Our results indicate that naive B cells (P = .02), natural killer cells (P = 2 × 10−06), and monocytes (P = 3 × 10−07) have significantly lower infiltration in fusion-positive than fusion-negative samples, while CD4+ T cells (P = 1× 10−04) have significantly higher infiltration in fusion-positive samples than fusion-negative samples (Figure 2B). These findings were further supported by significant correlation between immune infiltration and signature score (Table S2). Taken together, our results suggested that TMPRSS2-ERG fusion is associated with reduced level of immune infiltration. Previous studies have shown that higher nonsynonymous mutation rate and lower copy number variation (CNV) in tumor samples are associated with higher immune infiltration.9, 10 We found that TMPRSS2-ERG fusion was associated with lower nonsynonymous mutation rate (Figure 2C). However, we also observed lower level of fraction altered (Figure 2D) and lower homologous recombination deficiency scores (HRD) in fusion-positive samples (Figure 2E). Both fraction-altered and HRD represent CNV levels. Thereby, the immune infiltration differences between fusion-positive and fusion-negative samples are not due to CNV but nonsynonymous mutation rate. In summary, we defined a novel gene signature for TMPRSS2-ERG fusion with great clinical value. The gene signature recapitulates the deregulated pathway downstream of the TMPRSS2-ERG fusion event and is predictive of patient prognosis in prostate cancer. This work is supported by the Cancer Prevention Research Institute of Texas (CPRIT) (RR180061 to CC) and the National Cancer Institute of the National Institutes of Health (1R21CA227996 to CC). CC is a CPRIT Scholar in Cancer Research. The authors declare that they have no conflict of interest. Cancer Prevention Research Institute of Texas (CPRIT); Grant Number: RR180061; National Cancer Institute of the National Institutes of Health; 1R21CA227996 Chao Cheng conceived of the idea and designed data analysis. Emily Zhou and Baoyi Zhang performed the data analysis and prepared figures and tables. Emily Zhou wrote the manuscript with the help from Chao Cheng, Runjun D. Kumar, Evelien Schaafsma, Baoyi Zhang, and Kenneth Zhu. All authors discussed the results and contributed to the final manuscript. The datasets supporting the conclusions of this article are available in the Gene Expression Omnibus repository (GSE16560 and GSE8402). Please note: The publisher is not responsible for the content or functionality of any supporting information supplied by the authors. Any queries (other than missing content) should be directed to the corresponding author for the article.
The original version of this Article contained errors in the depiction of confidence intervals in the NF1 BCSS data illustrated in Figure 3b. These have now been corrected in both the PDF and HTML versions of the Article. The incorrect version of Figure 3b is presented in the associated Author Correction.
In this study we use somatic cancer mutations to identify important functional residues within sets of related genes. We focus on protein kinases, a superfamily of phosphotransferases that share homologous sequences and structural motifs and have many connections to cancer. We develop several statistical tests for identifying Significantly Mutated Positions (SMPs), which are positions in an alignment with mutations that show signs of selection. We apply our methods to 21,917 mutations that map to the alignment of human kinases and identify 23 SMPs. SMPs occur throughout the alignment, with many in the important A-loop region, and others spread between the N and C lobes of the kinase domain. Since mutations are pooled across the superfamily, these positions may be important to many protein kinases. We select eleven mutations from these positions for functional validation. All eleven mutations cause a reduction or loss of function in the affected kinase. The tested mutations are from four genes, including two tumor suppressors (TGFBR1 and CHEK2) and two oncogenes (KDR and ERBB2). They also represent multiple cancer types, and include both recurrent and non-recurrent events. Many of these mutations warrant further investigation as potential cancer drivers.