Expanded Materials and Methods, Supplemental Figures, and Supplemental Tables Figure S1. Flow cytometric analysis of peripheral blood from SS, SS.BN3, SSIL2Rγ, and SS.BN3IL2Rγ rats. Figure S2.SH2B3 specifically marks CD31+ blood vessels, but not LYVE1+ lymphatic vessels. Table S1. Peripheral blood analysis of SS, SS.BN3, SSIL2Rγ and SS.BN3IL2Rγ rats. Table S2. Genes located in the overlapping portions of the rat chromosome 3 cancer QTLs. Table S3. Variants that cosegregate between tumor resistant (BN and Cop) and tumor susceptible (ACI, F344, and SS) rat strains.
Population isolates such as those in Finland benefit genetic research because deleterious alleles are often concentrated on a small number of low-frequency variants (0.1% ≤ minor allele frequency < 5%). These variants survived the founding bottleneck rather than being distributed over a large number of ultrarare variants. Although this effect is well established in Mendelian genetics, its value in common disease genetics is less explored 1 , 2 . FinnGen aims to study the genome and national health register data of 500,000 Finnish individuals. Given the relatively high median age of participants (63 years) and the substantial fraction of hospital-based recruitment, FinnGen is enriched for disease end points. Here we analyse data from 224,737 participants from FinnGen and study 15 diseases that have previously been investigated in large genome-wide association studies (GWASs). We also include meta-analyses of biobank data from Estonia and the United Kingdom. We identified 30 new associations, primarily low-frequency variants, enriched in the Finnish population. A GWAS of 1,932 diseases also identified 2,733 genome-wide significant associations (893 phenome-wide significant (PWS), P < 2.6 × 10 –11 ) at 2,496 (771 PWS) independent loci with 807 (247 PWS) end points. Among these, fine-mapping implicated 148 (73 PWS) coding variants associated with 83 (42 PWS) end points. Moreover, 91 (47 PWS) had an allele frequency of <5% in non-Finnish European individuals, of which 62 (32 PWS) were enriched by more than twofold in Finland. These findings demonstrate the power of bottlenecked populations to find entry points into the biology of common diseases through low-frequency, high impact variants.
Accurately identifying somatic mutations is essential for precision oncology and crucial for calculating tumor-mutational burden (TMB), an important predictor of response to immunotherapy. For tumor-only variant calling (i.e., when the cancer biopsy but not the patient’s normal tissue sample is sequenced), accurately distinguishing somatic mutations from germline variants is a challenging problem that, when unaddressed, results in unreliable, biased, and inflated TMB estimates. Here, we apply machine learning to the task of somatic vs germline classification in tumor-only solid tumor samples using TabNet, XGBoost, and LightGBM, three machine-learning models for tabular data. We constructed a training set for supervised classification using features derived exclusively from tumor-only variant calling and drawing somatic and germline truth labels from an independent pipeline using the patient-matched normal samples. All three trained models achieved state-of-the-art performance on two holdout test datasets: a TCGA dataset including sarcoma, breast adenocarcinoma, and endometrial carcinoma samples (AUC > 94%), and a metastatic melanoma dataset (AUC > 85%). Concordance between matched-normal and tumor-only TMB improves from R2 = 0.006 to 0.71–0.76 with the addition of a machine-learning classifier, with LightGBM performing best. Notably, these machine-learning models generalize across cancer subtypes and capture kits with a call rate of 100%. We reproduce the recent finding that tumor-only TMB estimates for Black patients are extremely inflated relative to that of white patients due to the racial biases of germline databases. We show that our approach with XGBoost and LightGBM eliminates this significant racial bias in tumor-only variant calling.
Turan et. al Supplementary Methods File S1
Genome-wide association studies have successfully discovered thousands of common variants associated with human diseases and traits, but the landscape of rare variation in human disease has not been explored at scale. Exome sequencing studies of population biobanks provide an opportunity to systematically evaluate the impact of rare coding variation across a wide range of phenotypes to discover genes and allelic series relevant to human health and disease. Here, we present results from systematic association analyses of 4,529 phenotypes using single-variant and gene tests of 426,370 individuals in the UK Biobank with exome sequence data. We find that the discovery of genetic associations is tightly linked to frequency as well as correlated with metrics of deleteriousness and natural selection. We highlight biological findings elucidated by these data and release the dataset as a public resource alongside the Genebass browser for rapidly exploring rare variant association results. ### Competing Interest Statement KJK is a consultant for Vor Biopharma. BRG, XZ, FR, SE, AJG, MR, JW, HJ, JWD are employees of AbbVie Inc. MR is an employee of and owns stock in AbbVie Inc. EAT, DS, PGB, HR are employees of Biogen and hold stocks/stock options in Biogen. HK, XC, XH and MRM are employees of Pfizer. DK holds stock in the private company TriNetX, LLC. DSP was an employee of Genomics plc. All the analyses reported in this paper were performed as part of DSP's previous employment at the Massachusetts General Hospital and Broad Institute. NAW owns stock in Pfizer. LDG receives funding from Intel and Illumina. DGM is a founder with equity of Goldfinch Bio, and serves as a paid advisor to GSK, Variant Bio, Insitro, and Foresite Labs. HLR is a member of the scientific advisory board at Genome Medical. AAP is a Venture Partner at GV. He has received consulting fees from Novartis, and receives funding from Bayer, IBM, Microsoft, Alphabet, Intel, GSK, Pfizer, Illumina. MJD is a founder of Maze Therapeutics. BMN is a member of the scientific advisory board at Deep Genomics and RBNC Therapeutics, member of the scientific advisory committee at Milken and a consultant for Camp4 Therapeutics, Merck and Biogen. ### Funding Statement This work was funded by Abbvie, Biogen, and Pfizer, who are represented by co-authors on this manuscript ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: UK Biobank applications 26041 and 48511 I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines and uploaded the relevant EQUATOR Network research reporting checklist(s) and other pertinent material as supplementary files, if applicable. Yes All data is publicly available at <https://genebass.org> <https://genebass.org>
Drug development is a resource and time-intensive process resulting in attrition rates of up to 90%. As a result, repurposing existing drugs with established safety and pharmacokinetic profiles is gaining traction as a way of accelerating therapeutics development. Here we have developed unique machine learning-driven Natural Language Processing and biomedical semantic technologies that mine over 53 million biomedical documents to automate the generation of a 911M edge knowledge graph. We then applied subgraph queries that relate drugs to diseases using genetic evidence to identify potential drug repurposing candidates for a broad range of diseases. We use Carney Complex, a disease with no known treatment, to illustrate our approach. This analysis revealed Ruxolitinib (Incyte, trade name Jakafi), a JAK1/2 inhibitor with an established safety and efficacy profile approved to treat myelofibrosis, as a potential candidate for the treatment of Carney Complex through off-target drug activity.
ABSTRACT Population isolates such as Finland provide benefits in genetic studies because the allelic spectrum of damaging alleles in any gene is often concentrated on a small number of low-frequency variants (0.1% ≤ minor allele frequency < 5%), which survived the founding bottleneck, as opposed to being distributed over a much larger number of ultra--rare variants. While this advantage is well-- established in Mendelian genetics, its value in common disease genetics has been less explored. FinnGen aims to study the genome and national health register data of 500,000 Finns, already reaching 224,737 genotyped and phenotyped participants. Given the relatively high median age of participants (63 years) and dominance of hospital-based recruitment, FinnGen is enriched for many disease endpoints often underrepresented in population-based studies (e.g., rarer immune-mediated diseases and late onset degenerative and ophthalmologic endpoints). We report here a genome-wide association study (GWAS) of 1,932 clinical endpoints defined from nationwide health registries. We identify genome--wide significant associations at 2,491 independent loci. Among these, finemapping implicates 148 putatively causal coding variants associated with 202 endpoints, 104 with low allele frequency (AF<10%) of which 62 were over two-fold enriched in Finland. We studied a benchmark set of 15 diseases that had previously been investigated in large genome-wide association studies. FinnGen discovery analyses were meta-analysed in Estonian and UK biobanks. We identify 30 novel associations, primarily low-frequency variants strongly enriched, in or specific to, the Finnish population and Uralic language family neighbors in Estonia and Russia. These findings demonstrate the power of bottlenecked populations to find unique entry points into the biology of common diseases through low-frequency, high impact variants. Such high impact variants have a potential to contribute to medical translation including drug discovery.
Background The remarkable growth of genome-wide association studies (GWAS) has created a critical need to experimentally validate the disease-associated variants, 90% of which involve non-coding variants. Methods To determine how the field is addressing this urgent need, we performed a comprehensive literature review identifying 36,676 articles. These were reduced to 1454 articles through a set of filters using natural language processing and ontology-based text-mining. This was followed by manual curation and cross-referencing against the GWAS catalog, yielding a final set of 286 articles. Results We identified 309 experimentally validated non-coding GWAS variants, regulating 252 genes across 130 human disease traits. These variants covered a variety of regulatory mechanisms. Interestingly, 70% (215/309) acted through cis-regulatory elements, with the remaining through promoters (22%, 70/309) or non-coding RNAs (8%, 24/309). Several validation approaches were utilized in these studies, including gene expression (n = 272), transcription factor binding (n = 175), reporter assays (n = 171), in vivo models (n = 104), genome editing (n = 96) and chromatin interaction (n = 33). Conclusions This review of the literature is the first to systematically evaluate the status and the landscape of experimentation being used to validate non-coding GWAS-identified variants. Our results clearly underscore the multifaceted approach needed for experimental validation, have practical implications on variant prioritization and considerations of target gene nomination. While the field has a long way to go to validate the thousands of GWAS associations, we show that progress is being made and provide exemplars of validation studies covering a wide variety of mechanisms, target genes, and disease areas.
Genome-wide association studies have successfully discovered thousands of common variants associated with human diseases and traits, but the landscape of rare variations in human disease has not been explored at scale. Exome-sequencing studies of population biobanks provide an opportunity to systematically evaluate the impact of rare coding variations across a wide range of phenotypes to discover genes and allelic series relevant to human health and disease. Here, we present results from systematic association analyses of 4,529 phenotypes using single-variant and gene tests of 394,841 individuals in the UK Biobank with exome-sequence data. We find that the discovery of genetic associations is tightly linked to frequency and is correlated with metrics of deleteriousness and natural selection. We highlight biological findings elucidated by these data and release the dataset as a public resource alongside the Genebass browser for rapidly exploring rare-variant association results.
Dupuytren’s contracture or disease (DD) is a disabling, fibroproliferative disease of the hand that affects up to 25% of people of northwestern European descent. It typically manifests in adulthood and many affected individuals have a positive family history, yet the genetic architecture of DD is not completely understood. We conducted genome-wide association studies (GWAS) of DD in 8,846 cases and 347,659 population controls from the UK Biobank resource, and 4,616 cases and 253,122 population controls from the FinnGen study. We combined these datasets with a meta-analysis, which represents the largest GWAS conducted in DD to date including 13,462 cases. We identified 83 loci with genome-wide significance of p < 5 × 10 -8 . We replicated association at the 24 previously reported loci and discovered 59 novel loci, substantially increasing the number of risk loci for DD. Colocalization with expression quantitative trait loci and overlap with genes linked to phenotypically matched human Mendelian disorders or animal models support causal roles for at least 30 genes with high confidence. Among these, fifteen genes cause rare limb abnormalities when mutated, an observation that may potentially shed light on hand specificity of DD phenotype. Gene enrichment analysis revealed predominant role of connective tissue development and maintenance and extracellular matrix homeostasis but limited or no role of inflammatory processes in disease causality. These findings provide key insights into the biological mechanisms underlying DD and identify genetically informed therapeutic targets for DD and possibly for other fibrotic diseases.
J.S. Lee: None. K.C. Lee: Royalty; Regents of the University of Michigan. Patent/License Fees/Copyright; Regents of the University of Michigan. D. Bojrab: None. P.Y. Chen: Stock; Greater Michigan Gamma Knife (GMGK). J. Jacob: None. I.S. Grills: Greater Michigan Gamma Knife.
The UK Biobank Exome Sequencing Consortium (UKB-ESC) is a private–public partnership between the UK Biobank (UKB) and eight biopharmaceutical companies that will complete the sequencing of exomes for all ~500,000 UKB participants. Here, we describe the early results from ~200,000 UKB participants and the features of this project that enabled its success. The biopharmaceutical industry has increasingly used human genetics to improve success in drug discovery. Recognizing the need for large-scale human genetics data, as well as the unique value of the data access and contribution terms of the UKB, the UKB-ESC was formed. As a result, exome data from 200,643 UKB enrollees are now available. These data include ~10 million exonic variants—a rich resource of rare coding variation that is particularly valuable for drug discovery. The UKB-ESC precompetitive collaboration has further strengthened academic and industry ties and has provided teams with an opportunity to interact with and learn from the wider research community. The UK Biobank Exome Sequencing Consortium aims to sequence all the exomes of approximately 500,000 UK Biobank participants. This Perspective describes the results from approximately 200,000 exomes and discusses the lessons learned from this UK Biobank–biopharmaceutical company collaboration.
Clinical applications of precision oncology require accurate tests that can distinguish true cancer-specific mutations from errors introduced at each step of next-generation sequencing (NGS). To date, no bulk sequencing study has addressed the effects of cross-site reproducibility, nor the biological, technical and computational factors that influence variant identification. Here we report a systematic interrogation of somatic mutations in paired tumor–normal cell lines to identify factors affecting detection reproducibility and accuracy at six different centers. Using whole-genome sequencing (WGS) and whole-exome sequencing (WES), we evaluated the reproducibility of different sample types with varying input amount and tumor purity, and multiple library construction protocols, followed by processing with nine bioinformatics pipelines. We found that read coverage and callers affected both WGS and WES reproducibility, but WES performance was influenced by insert fragment size, genomic copy content and the global imbalance score (GIV; G > T/C > A). Finally, taking into account library preparation protocol, tumor content, read coverage and bioinformatics processes concomitantly, we recommend actionable practices to improve the reproducibility and accuracy of NGS experiments for cancer mutation detection. Recommendations are given on optimal read coverage and selection of calling algorithm to maximize the reproducibility of cancer mutation detection in whole-genome or whole-exome sequencing.
In precision oncology, reliable identification of tumor-specific DNA mutations requires sequencing tumor DNA and non-tumor DNA (so-called “matched normal”) from the same patient. The normal sample allows researchers to distinguish acquired (somatic) and hereditary (germline) variants. The ability to distinguish somatic and germline variants facilitates estimation of tumor mutation burden (TMB), which is a recently FDA-approved pan-cancer marker for highly successful cancer immunotherapies; in tumor-only variant calling (i.e., without a matched normal), the difficulty in discriminating germline and somatic variants results in inflated and unreliable TMB estimates. We apply machine learning to the task of somatic vs germline classification in tumor-only samples using TabNet, a recently developed attentive deep learning model for tabular data that has achieved state of the art performance in multiple classification tasks (Arik and Pfister 2019). We constructed a training set for supervised classification using features derived from tumor-only variant calling and drawing somatic and germline truth-labels from an independent pipeline incorporating the patient-matched normal samples. Our trained model achieved state-of-the-art performance on two hold-out test datasets: a TCGA dataset including sarcoma, breast adenocarcinoma, and endometrial carcinoma samples (F1-score: 88.3), and a metastatic melanoma dataset, (F1-score 79.8). Concordance between matched-normal and tumor-only TMB improves from R 2 = 0.006 to 0.705 with the addition of our classifier. And importantly, this approach generalizes across tumor tissue types and capture kits and has a call rate of 100%. The interpretable feature masks of the attentive deep learning model explain the reasons for misclassified variants. We reproduce the recent finding that tumor-only TMB estimates for Black patients are extremely inflated relative to that of White patients due to the racial biases of germline databases. We show that our machine learning approach appreciably reduces this racial bias in tumor-only variant-calling.
The lack of samples for generating standardized DNA datasets for setting up a sequencing pipeline or benchmarking the performance of different algorithms limits the implementation and uptake of cancer genomics. Here, we describe reference call sets obtained from paired tumor–normal genomic DNA (gDNA) samples derived from a breast cancer cell line—which is highly heterogeneous, with an aneuploid genome, and enriched in somatic alterations—and a matched lymphoblastoid cell line. We partially validated both somatic mutations and germline variants in these call sets via whole-exome sequencing (WES) with different sequencing platforms and targeted sequencing with >2,000-fold coverage, spanning 82% of genomic regions with high confidence. Although the gDNA reference samples are not representative of primary cancer cells from a clinical sample, when setting up a sequencing pipeline, they not only minimize potential biases from technologies, assays and informatics but also provide a unique resource for benchmarking ‘tumor-only’ or ‘matched tumor–normal’ analyses. Tumor–normal paired DNA samples from a breast cancer cell line and a matched lymphoblastoid cell line enable calibration of clinical sequencing pipelines and benchmarking ‘tumor-only’ or ‘matched tumor–normal’ analyses.
A Correction to this paper has been published: https://doi.org/10.1038/s41416-020-01193-w
Significance Statement The genes and mechanisms underlying the association between diabetes or hypertension and CKD risk are unclear. The authors identified a recessive K572Q mutation in γ-adducin (Add3), which encodes a cytoskeletal protein (ADD3), in fawn-hooded hypertensive (FHH) rats—a mutation also reported in Milan normotensive (MNS) rats that develop renal disease. They demonstrated that FHH and Add3 knockout rats had impairments in the myogenic response of afferent arterioles and in renal blood flow autoregulation, which were rescued in Add3 transgenic rats. They confirmed the K572Q mutation’s role in altering the myogenic response in a genetic complementation study that involved crossing FHH and MNS rats. The work is the first to demonstrate that a mutation in ADD3 that causes renal vascular dysfunction also promotes susceptibility to kidney disease. Background The genes and mechanisms involved in the association between diabetes or hypertension and CKD risk are unclear. Previous studies have implicated a role for γ-adducin (ADD3), a cytoskeletal protein encoded by Add3. Methods We investigated renal vascular function in vitro and in vivo and the susceptibility to CKD in rats with wild-type or mutated Add3 and in genetically modified rats with overexpression or knockout of ADD3. We also studied glomeruli and primary renal vascular smooth muscle cells isolated from these rats. Results This study identified a K572Q mutation in ADD3 in fawn-hooded hypertensive (FHH) rats—a mutation previously reported in Milan normotensive (MNS) rats that also develop kidney disease. Using molecular dynamic simulations, we found that this mutation destabilizes a critical ADD3-ACTIN binding site. A reduction of ADD3 expression in membrane fractions prepared from the kidney and renal vascular smooth muscle cells of FHH rats was associated with the disruption of the F-actin cytoskeleton. Compared with renal vascular smooth muscle cells from Add3 transgenic rats, those from FHH rats had elevated membrane expression of BKα and BK channel current. FHH and Add3 knockout rats exhibited impairments in the myogenic response of afferent arterioles and in renal blood flow autoregulation, which were rescued in Add3 transgenic rats. We confirmed these findings in a genetic complementation study that involved crossing FHH and MNS rats that share the ADD3 mutation. Add3 transgenic rats showed attenuation of proteinuria, glomerular injury, and kidney fibrosis with aging and mineralocorticoid-induced hypertension. Conclusions This is the first report that a mutation in ADD3 that alters ACTIN binding causes renal vascular dysfunction and promotes the susceptibility to kidney disease.
We characterized two reference samples for NGS technologies: a human triple-negative breast cancer cell line and a matched normal cell line. Leveraging several whole-genome sequencing (WGS) platforms, multiple sequencing replicates, and orthogonal mutation detection bioinformatics pipelines, we minimized the potential biases from sequencing technologies, assays, and informatics. Thus, our “truth sets” were defined using evidence from 21 repeats of WGS runs with coverages ranging from 50X to 100X (a total of 140 billion reads). These “truth sets” present many relevant variants/mutations including 193 COSMIC mutations and 9,016 germline variants from the ClinVar database, nonsense mutations in BRCA1/2 and missense mutations in TP53 and FGFR1. Independent validation in three orthogonal experiments demonstrated a successful stress test of the truth set. We expect these reference materials and “truth sets” to facilitate assay development, qualification, validation, and proficiency testing. In addition, our methods can be extended to establish new fully characterized reference samples for the community.
The authors reviewed data on 1519 patients referred to the Undiagnosed Diseases Network (UDN), a US NIH funded network linking seven clinical sites. 53% of patients were female and their symptoms were neurologic (40%), musculoskeletal (10%), immunological (7%), gastrointestinal (7%), or rheumatological (6%). Of the 382 patients who had a complete evaluation, the UDN was able to establish the diagnosis in 132 (35%), many of which resulted in changes in management. 31 new syndromes were defined.
BackgroundRare variants (RV) in immunoglobulin mu-binding protein 2 (IGHMBP2) [OMIM 600502] can cause an autosomal recessive type of Charcot-Marie-Tooth (CMT) disease [OMIM 616155], an inherited peripheral neuropathy. Over 40 different genes are associated with CMT, with different possible inheritance patterns. Methods and ResultsAn 11-year-old female with motor delays was found to have distal atrophy, weakness, and areflexia without bulbar or sensory findings. Her clinical evaluation was unrevealing. Whole exome sequencing (WES) revealed a maternally inherited IGHMBP2 RV (c.1730T>C) predicted to be pathogenic, but no variant on the other allele was identified. Deletion and duplication analysis was negative. She was referred to the Undiagnosed Disease Network (UDN) for further evaluation. Whole genome sequencing (WGS) confirmed the previously identified IGHMBP2 RV and identified a paternally inherited non-coding IGHMBP2 RV. This was predicted to activate a cryptic splice site perturbing IGHMBP2 splicing. Reverse transcriptase polymerase chain reaction (RT-PCR) analysis was consistent with activation of the cryptic splice site. The abnormal transcript was shown to undergo nonsense-mediated decay (NMD), resulting in halpoinsufficiency. ConclusionThis case demonstrates the deficiencies of WES and traditional molecular analyses and highlights the advantages of utilization of WGS and functional studies.