Despite the introduction of genome sequencing (GS) for rare disease diagnostics, a genetic cause is not identified in most patients. Here, we explored the potential of proteomics to improve the diagnostic yield in 424 patients with rare diseases from the 100,000 Genomes Project (100kGP) without a genetic diagnosis. Serum proteomic profiling was performed using the Olink Explore 1536 assay ( N = 1463 proteins). For 13 patients without genetic diagnoses, detection of lower serum protein “outliers” ( z -score < −2) led to confirmed genetic diagnoses by resolving variants of uncertain significance or prioritizing genes for targeted GS reanalysis. For 23 additional patients without genetic diagnoses (64% of findings), we identified candidate gene-disease links and variants through convergent evidence from lower protein outliers and variants ranked through the variant prioritization tool Exomiser. For example, we identified a candidate heterozygous missense variant [Genome Aggregation Database (gnomAD) minor allele frequency = 0.006%] in tyrosine kinase with immunoglobulin-like and epidermal growth factor homology domains 1 ( TIE1 ) that was only present in a patient with lower TIE1 serum abundance ( z -score = −5.12) and their father, both of whom were affected by the same monogenic cardiac disorder, but in no other individuals from the 100kGP. Missense (52.5%) and splice region (27.5%) variants accounted for most diagnostic or candidate variants prioritized. This proof-of-principle study demonstrated that serum proteomics can support rare disease diagnosis and identify disease-causing genes in patients undiagnosed after GS, although successful implementation will likely depend on tissue specificity of protein expression, detectability in blood, proteomic platform coverage, and sensitivity.
Rare coding variants can have large effects on disease risk and provide direct routes from human genetics to disease mechanisms and therapeutic targets, but their discovery is constrained by sample size, particularly for low-prevalence diseases. Here we establish the Biobank Rare Variant Analysis (BRaVa) consortium, a global rare variant association resource that integrates sequencing and linked health-record data from ten biobanks and cohorts comprising over 1.2 million individuals across diverse ancestries. We performed gene-based meta-analyses of rare coding variation across 33 clinical endpoints and 11 quantitative traits. Aggregating evidence across biobanks and ancestries identified 514 gene-trait associations, including 31 not previously reported in prior studies or curated association resources following systematic literature review. Notably, 36.1% of gene-level associations were undetectable in any individual biobank, and 91 emerged only through cross-ancestry meta-analysis, demonstrating that federated integration enables discovery beyond the reach of single cohorts. Similar gains were observed at the variant level, where 25.0% of phenotype-locus associations were detectable only through meta-analysis. Effect size estimates were correlated across ancestries with concordant directions of effect, supporting the generalizability of rare variant associations. The identified signals implicate pathways involved in transcriptional and epigenetic regulation, metabolism, vascular and epithelial biology, and immune function, highlighting rare coding variation as an engine for biological discovery across medical record phenotypes. For example, damaging variation in ANKRD12 implicates inflammatory transcriptional dysregulation in asthma and chronic obstructive pulmonary disease, and ultra-rare predicted loss-of-function variants in NAA15 link protein acetylation processes to type 2 diabetes risk. BRaVa establishes a scalable framework and freely available community resource for rare variant meta-analysis across global biobanks. Public release of gene- and variant-level association summary statistics provides a reference map of rare coding variant associations to support disease gene discovery, biological interpretation, and therapeutic target prioritization as sequencing-linked health-record resources continue to expand.
BACKGROUND:Idiopathic pulmonary fibrosis (IPF) is a progressive and debilitating respiratory disease with limited therapeutic options. Genetic association studies for IPF have identified several associations and probable effector genes that could not only help understanding IPF pathogenesis but also develop effective treatments. Assessing genetic overlap between IPF and severe COVID-19, an acute respiratory disease that can trigger pulmonary fibrosis, may reveal shared aetiology and mechanisms, thereby supporting the development of common treatments. METHODS:We carried out genome-wide association studies (GWAS), post-GWAS, and rare variant analyses using whole genome sequencing data from the 100,000 Genomes Project IPF cohort (n = 586). We performed a meta-analysis combining 100 kGP with published IPF GWASs (total 11,746 cases and 1,416,493 controls). We tested inhibition in vitro for a probable effector gene of an identified association. We also investigated genetic colocalisation between IPF and severe COVID-19 and leveraged their genetic correlation through multi-trait meta-analysis for discovery. FINDINGS:IPF meta-analysis identified an additional association at 1q21.2 (rs16837903, OR [95% CI] = 0.88 [0.85, 0.92], P = 9.5 × 10-9), which was replicated in independent data. MCL1, one of the probable effector genes of the 1q21.2 signal has a known antiapoptotic role, but MCL1 inhibition in vitro did not selectively deplete senescent alveolar epithelial cells. Rare variant burden analysis identified ANGPTL7, a secreted glycoprotein involved in the regulation of angiogenesis, as an IPF candidate gene (OR [95% CI] = 28.8 [8.51, 97.4], P = 6.7 × 10-8). We discovered additional shared genetic loci between IPF and severe COVID-19 at 1q21.2, 6p24.3, and 16p13.3, with probable effector genes MCL1, DSP, and RHBDF1, implicating regulation of apoptosis, cell adhesion, and epidermal growth factor signalling, respectively. The genetic correlation between IPF and severe COVID-19 was rg [95% CI] = 0.39 [0.25, 0.53]. Using multi-trait meta-analysis, we identified and replicated an additional candidate IPF signal at 2p16.1 with probable effector gene BCL11A, a regulator of haematopoiesis and lymphocyte development. INTERPRETATION:These findings prioritise probable effector genes mediating IPF risk and identify potential therapeutic targets that require validation, with genes colocalising with severe COVID-19 suggesting potential for developing common treatments. FUNDING:None.
In susceptible patients, COVID-19 causes life-threatening disease driven by immune-mediated inflammatory lung injury. We have previously shown that multiple common host genetic variants are significantly associated with susceptibility to critical Covid-19, and in one case, we demonstrated that such variants can inform development of new, effective drug treatment. Here we report an association analysis of whole-genome sequences (WGS) from 11,423 cases from the GenOMICC study and 60,628 controls, together with meta-analyses with available genome-wide data. We identify a rare association signal at SLC50A1, primarily driven by a missense variant rs147850817 (1:155138217:G:T, Arg201Leu) that may interfere with transport function, and we identify four common association signals near ARF1, ZNF462, KLF13 and MVP genes. Finally, we build a WGS-derived polygenic risk score (PRS) for critical Covid-19, which offers only marginal improvement in risk estimation for the general population but may provide clinically-valuable discrimination for extreme susceptibility. ### Competing Interest Statement The authors have declared no competing interest. ### Clinical Protocols ### Funding Statement GenOMICC was funded by Sepsis Research (the Fiona Elizabeth Agnew Trust), the Intensive Care Society, a Wellcome Trust Senior Research Fellowship (J.K.Baillie, 223164/Z/21/Z), the Department of Health and Social Care (DHSC), Illumina, LifeArc, the Medical Research Council, UKRI, a BBSRC Institute Strategic Program Support Grant to the Roslin Institute (BBS/E/D/20002172, BBS/E/D/10002070 and BBS/E/D/30002275) and UKRI grants MC PC 20004, MC PC 19025, MC PC 1905, and MRNO2995X/1. ADB acknowledges funding from the Wellcome PhD training fellowship for clinicians (204979/Z/16/Z), the Edinburgh Clinical Academic Track (ECAT) programme. This research is supported in part by the Data and Connectivity National Core Study, led by Health Data Research UK in partnership with the Office for National Statistics and funded by UK Research and Innovation (grant ref MC PC 20029). This study owes a great deal to the National Institute for Healthcare Research Clinical Research Network (NIHR CRN) and the Chief Scientist's Office (Scotland), who facilitate recruitment into research studies in NHS hospitals, and to the global ISARIC and InFACT consortia. This work forms part of the translational research portfolio of the National Institute for Health and Care Research Barts Biomedical Research Centre. T.M. is supported by Cancer Research UK grant DRCRPG-May23/100002 to C. Siebold. Genomics England: This research was made possible through access to data in the National Genomic Research Library, which is managed by Genomics England Limited (a wholly owned company of the Department of Health and Social Care). The National Genomic Research Library (\url{https://www.genomicsengland.co.uk/research}) holds data provided by patients and collected by the NHS as part of their care and data collected as part of their participation in research. The National Genomic Research Library is funded by the National Institute for Health Research and NHS England. The Wellcome Trust, Cancer Research UK and the Medical Research Council have also funded research infrastructure. REACT: National Institute for Health and Care Research (NIHR) and UK Research and Innovation (UKRI) - REACT-Genomics England (REACT-GE) (MR/V030841/1) and REACT-Long COVID (REACT-LC) (COV-LT-0040). The REACT study was funded by the UK Department of Health and Social Care with supplemental funding from the Huo Family Foundation. ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: GenOMICC was approved by the following research ethics committees: Scotland A Research Ethics Committee (15/SS/0110) and Coventry and Warwickshire Research Ethics Committee (England, Wales and Northern Ireland) (19/WM/0247). I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes All other data produced in the present study are available upon reasonable request to the authors
Introduction Cystic kidney disease (CyKD) is frequently a familial disease, with ~85% of probands receiving a monogenic diagnosis. However, gene discovery has been led by family-based and candidate gene studies, limiting the ascertainment of non-Mendelian genetic contributors to the disease. Using whole genome sequencing data provided by the 100,000 Genomes Project (100KGP), we used hypothesis-free approaches to systematically characterize and quantify the genetic contributors to CyKD across variant types and the allele frequency spectrum. Methods We performed a sequencing-based genome-wide association study in 1,209 unrelated patients recruited to the 100,000 Genomes Project with CyKD and 26,096 ancestry-matched unaffected controls. The analysis was inclusive of individuals with diverse genetic ancestries. Enrichment of common, low-frequency (minor allele frequency [MAF] > 0.1%) and rare (MAF < 0.1%) single-nucleotide variants (SNV), indels and rare structural variants (SV) on a genome-wide and per-gene basis was sought using a generalised linear mixed model approach to account for population structure. Meta-analysis of CyKD cohorts from Finngen, the UK Biobank and BioBank Japan was performed. Results In 995 of the 1209 (82.30%) CyKD cases a likely disease-causing monogenic variant was identified. Gene-based analysis of rare SNVs/indels predicted to be damaging revealed PKD1 (P=1.13x10-309), PKD2 (P=1.96x10-150), DNAJB11 (P=3.52x10-7), COL4A3 (P=1.26x10-6) and truncating monoallelic PKHD1 (P=2.98x10-8) variants to be significantly associated with disease. Depleting for solved cases led to the emergence of a significant association at IFT140 (P=3.46x10-17) and strengthening of the COL4A3 (P=9.27x10-7) association, driven exclusively by heterozygous variants for both genes. After depleting for those harbouring IFT140 and COL4A3 variants , no other genes were identified. Risk of disease attributable to monoallelic defects of multiple genes linked with CyKD was quantified, with lower risk seen in rarer and more recently described genetic diagnoses. Genome-wide structural variant associations highlighted deletions in PKD1 (P=2.17x10-22), PKD2 (P=7.48x10-12) and the 17q12 locus containing HNF1B (P=4.12x10-8) as statistically significant contributors to disease. Genome-wide analysis of over 18 million common and low-frequency variants in the Finnish population revealed evidence of association (P=1.4x10-149) of a heterozygous stop-gain variant in PKHD1 that is endemic (MAF=4.7x10-03) in this population. Meta-analysis of 2,923 cases and 900,824 controls across 6,641,351 common and low frequency variants including UK, Japanese and Finnish biobanks did not reveal any novel significant associations. SNVs with a MAF>0.1% accounted for between 3 and 9% of the heritability of CyKD across three different European ancestry cohorts. Conclusions These findings represent an unbiased examination of the genetic architecture of a national CyKD cohort using robust statistical methodology. Causative monoallelic mutations in IFT140 have recently been reported in other cohorts associated with a milder phenotype than PKD1/2-associated disease. The association with COL4A3 suggests that in some circumstances CyKD may be the presenting feature of collagen IV-related kidney disease and the significant association observed with monoallelic predicted loss-of-function PKHD1 variants extends the spectrum of phenotypic abnormalities associated with this gene. In addition to quantification of the contribution of non-coding and structural variants to CyKD, the per gene quantification of CyKD risk presented could be used to inform genetic testing and counselling strategies clinically and we also show that common variants make a small contribution to CyKD heritability. Keywords: genomics, cystic kidney disease, renal, ADPKD ### Competing Interest Statement The authors have declared no competing interest. ### Funding Statement OSA is funded by an MRC Clinical Research Training Fellowship (MR/S021329/1). MYC is funded by a Kidney Research UK Clinical Research Fellowship (TF\_004\_20161125). DPG is supported by the St Peters Trust for Kidney, Bladder and Prostate Research. ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: Ethical approval for the 100,000 genome project was granted by the Research Ethics Committee for East of England Cambridge South (REC Ref: 14/EE/1112). I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes All data produced in the present work are contained in the manuscript and online at https://zenodo.org/records/10613736
BACKGROUNDCystic kidney disease (CyKD) is a predominantly familial disease in which gene discovery has been led by family-based and candidate gene studies, an approach that is susceptible to ascertainment and other biases.METHODSUsing whole-genome sequencing data from 1,209 cases and 26,096 ancestry-matched controls participating in the 100,000 Genomes Project, we adopted hypothesis-free approaches to generate quantitative estimates of disease risk for each genetic contributor to CyKD, across genes, variant types and allelic frequencies.RESULTSIn 82.3% of cases, a qualifying potentially disease-causing rare variant in an established gene was found. There was an enrichment of rare coding, splicing, and structural variants in known CyKD genes, with statistically significant gene-based signals in COL4A3 and (monoallelic) PKHD1. Quantification of disease risk for each gene (with replication in the separate UK Biobank study) revealed substantially lower risk associated with genes more recently associated with autosomal dominant polycystic kidney disease, with odds ratios for some below what might usually be regarded as necessary for classical Mendelian inheritance. Meta-analysis of common variants did not reveal significant associations, but suggested this category of variation contributes 3%-9% to the heritability of CyKD across European ancestries.CONCLUSIONBy providing unbiased quantification of risk effects per gene, this research suggests that not all rare variant genetic contributors to CyKD are equally likely to manifest as a Mendelian trait in families. This information may inform genetic testing and counseling in the clinic.
Critical illness in COVID-19 is an extreme and clinically homogeneous disease phenotype that we have previously shown 1 to be highly efficient for discovery of genetic associations 2 . Despite the advanced stage of illness at presentation, we have shown that host genetics in patients who are critically ill with COVID-19 can identify immunomodulatory therapies with strong beneficial effects in this group 3 . Here we analyse 24,202 cases of COVID-19 with critical illness comprising a combination of microarray genotype and whole-genome sequencing data from cases of critical illness in the international GenOMICC (11,440 cases) study, combined with other studies recruiting hospitalized patients with a strong focus on severe and critical disease: ISARIC4C (676 cases) and the SCOURGE consortium (5,934 cases). To put these results in the context of existing work, we conduct a meta-analysis of the new GenOMICC genome-wide association study (GWAS) results with previously published data. We find 49 genome-wide significant associations, of which 16 have not been reported previously. To investigate the therapeutic implications of these findings, we infer the structural consequences of protein-coding variants, and combine our GWAS results with gene expression data using a monocyte transcriptome-wide association study (TWAS) model, as well as gene and protein expression using Mendelian randomization. We identify potentially druggable targets in multiple systems, including inflammatory signalling ( JAK1 ), monocyte–macrophage activation and endothelial permeability ( PDE4A ), immunometabolism ( SLC2A5 and AK5 ), and host factors required for viral entry and replication ( TMPRSS2 and RAB2A ).
Host-pathogen interactions impose recurrent selective pressures that lead to constant adaptation and counter-adaptation in both competing species. Here, we sought to study this evolutionary arms-race and assessed the impact of the innate immune system on viral population diversity and evolution, using Drosophila melanogaster as model host and its natural pathogen Drosophila C virus (DCV). We isogenized eight fly genotypes generating animals defective for RNAi, Imd and Toll innate immune pathways as well as pathogen-sensing and gut renewal pathways. Wild-type or mutant flies were then orally infected with DCV and the virus was serially passaged ten times via reinfection in naive flies. Viral population diversity was studied after each viral passage by high-throughput sequencing and infection phenotypes were assessed at the beginning and at the end of the evolution experiment. We found that the absence of any of the various immune pathways studied increased viral genetic diversity while attenuating virulence. Strikingly, these effects were observed in a range of host factors described as having mainly antiviral or antibacterial functions. Together, our results indicate that the innate immune system as a whole and not specific antiviral defence pathways in isolation, generally constrains viral diversity and evolution.
Pulmonary inflammation drives critical illness in Covid-19, [1][1];[2][2] creating a clinically homogeneous extreme phenotype, which we have previously shown to be highly efficient for discovery of genetic associations. [3][3];[4][4] Despite the advanced stage of illness, we have found that immunomodulatory therapies have strong beneficial effects in this group. [1][1];[5][5] Further genetic discoveries may identify additional therapeutic targets to modulate severe disease. [6][6] In this new data release from the GenOMICC (Genetics Of Mortality in Critical Care) study we include new microarray genotyping data from additional critically-ill cases in the UK and Brazil, together with cohorts of severe Covid-19 from the ISARIC4C [7][7] and SCOURGE [8][8] studies, and meta-analysis with previously-reported data. We find an additional 14 new genetic associations. Many are in potentially druggable targets, in inflammatory signalling (JAK1, PDE4A), monocyte-macrophage differentiation (CSF2), immunometabolism (SLC2A5, AK5), and host factors required for viral entry and replication (TMPRSS2, RAB2A). As with our previous work, these results provide tractable therapeutic targets for modulation of harmful host-mediated inflammation in Covid-19. ### Competing Interest Statement The authors have declared no competing interest. ### Funding Statement GenOMICC was funded by Sepsis Research (the Fiona Elizabeth Agnew Trust), the Intensive Care Society, a Wellcome Trust Senior Research Fellowship (J.K.Baillie, 223164/Z/21/Z), the Department of Health and Social Care (DHSC), Illumina, LifeArc, the Medical Research Council, UKRI, a BBSRC Institute Program Support Grant to the Roslin Institute (BBS/E/D/20002172, BBS/E/D/10002070 and BBS/E/D/30002275) and UKRI grants MC\_PC\_20004, MC\_PC\_19025, MC\_PC\_1905, and MRNO2995X/1. This research is supported in part by the Data and Connectivity National Core Study, led by Health Data Research UK in partnership with the Office for National Statistics and funded by UK Research and Innovation (grant ref MC\_PC\_20029). We acknowledge NHS Digital, Public Health England and the Intensive Care National Audit and Research Centre who provided clinical data on the participants. This study owes a great deal to the National Institute for Healthcare Research Clinical Research Network (NIHR CRN) and the Chief Scientist Office (Scotland), who facilitate recruitment into research studies in NHS hospitals, and to the global ISARIC and InFACT consortia. GenOMICC genotype controls were obtained using UK Biobank Resource under project 788 funded by Roslin Institute Strategic Programme Grants from the BBSRC (BBS/E/D/10002070 and BBS/E/D/30002275) and Health Data Research UK (references HDR-9004 and HDR-9003) ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: We performed TWAS in the MetaXcan framework and the GTExv8 eQTL and sQTL MASHR-M models available for download in http://predictdb.org/. Publicly-available HGI data was downloaded from https://www.covid19hg.org/results/r6/. I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines and uploaded the relevant EQUATOR Network research reporting checklist(s) and other pertinent material as supplementary files, if applicable. Yes All data, including downloadable summary data and access applications for individual-level data, can be obtained through the GenOMICC gateway site https://genomicc.org/data. [1]: #ref-1 [2]: #ref-2 [3]: #ref-3 [4]: #ref-4 [5]: #ref-5 [6]: #ref-6 [7]: #ref-7 [8]: #ref-8
Posterior urethral valves (PUV) are the commonest cause of end-stage renal disease in children, but the genetic architecture of this rare disorder remains unknown. We performed a sequencing-based genome-wide association study (seqGWAS) in 132 unrelated male PUV cases and 23,727 controls of diverse ancestry, identifying statistically significant associations with common variants at 12q24.21 (p=7.8 × 10−12; OR 0.4) and rare variants at 6p21.1 (p=2.0 × 10-8; OR 7.2), that were replicated in an independent European cohort of 395 cases and 4151 controls. Fine mapping and functional genomic data mapped these loci to the transcription factor TBX5 and planar cell polarity gene PTK7, respectively, the encoded proteins of which were detected in the developing urinary tract of human embryos. We also observed enrichment of rare structural variation intersecting with candidate cis-regulatory elements, particularly inversions predicted to affect chromatin looping (p=3.1 × 10-5). These findings represent the first robust genetic associations of PUV, providing novel insights into the underlying biology of this poorly understood disorder and demonstrate how a diverse ancestry seqGWAS can be used for disease locus discovery in a rare disease.
Critical COVID-19 is caused by immune-mediated inflammatory lung injury. Host genetic variation influences the development of illness requiring critical care1 or hospitalization2-4 after infection with SARS-CoV-2. The GenOMICC (Genetics of Mortality in Critical Care) study enables the comparison of genomes from individuals who are critically ill with those of population controls to find underlying disease mechanisms. Here we use whole-genome sequencing in 7,491 critically ill individuals compared with 48,400 controls to discover and replicate 23 independent variants that significantly predispose to critical COVID-19. We identify 16 new independent associations, including variants within genes that are involved in interferon signalling (IL10RB and PLSCR1), leucocyte differentiation (BCL11A) and blood-type antigen secretor status (FUT2). Using transcriptome-wide association and colocalization to infer the effect of gene expression on disease severity, we find evidence that implicates multiple genes-including reduced expression of a membrane flippase (ATP11A), and increased expression of a mucin (MUC1)-in critical disease. Mendelian randomization provides evidence in support of causal roles for myeloid cell adhesion molecules (SELE, ICAM5 and CD209) and the coagulation factor F8, all of which are potentially druggable targets. Our results are broadly consistent with a multi-component model of COVID-19 pathophysiology, in which at least two distinct mechanisms can predispose to life-threatening disease: failure to control viral replication; or an enhanced tendency towards pulmonary inflammation and intravascular coagulation. We show that comparison between cases of critical illness and population controls is highly efficient for the detection of therapeutically relevant mechanisms of disease.
While recent advancements in computation and modelling have improved the analysis of complex traits, our understanding of the genetic basis of the time at symptom onset remains limited. Here, we develop a Bayesian approach (BayesW) that provides probabilistic inference of the genetic architecture of age-at-onset phenotypes in a sampling scheme that facilitates biobank-scale time-to-event analyses. We show in extensive simulation work the benefits BayesW provides in terms of number of discoveries, model performance and genomic prediction. In the UK Biobank, we find many thousands of common genomic regions underlying the age-at-onset of high blood pressure (HBP), cardiac disease (CAD), and type-2 diabetes (T2D), and for the genetic basis of onset reflecting the underlying genetic liability to disease. Age-at-menopause and age-at-menarche are also highly polygenic, but with higher variance contributed by low frequency variants. Genomic prediction into the Estonian Biobank data shows that BayesW gives higher prediction accuracy than other approaches.
Due to the complexity of linkage disequilibrium (LD) and gene regulation, understanding the genetic basis of common complex traits remains a major challenge. We develop a Bayesian model (BayesRR-RC) implemented in a hybrid-parallel algorithm that scales to whole-genome sequence data on many hundreds of thousands of individuals, taking 22 seconds per iteration to estimate the inclusion probabilities and effect sizes of 8.4 million markers and 78 SNP-heritability parameters in the UK Biobank. We show in theory and simulation that BayesRR-RC provides robust variance component and enrichment estimates, improved marker discovery and effect estimates over mixed-linear model association approaches, and accurate genomic prediction. Of the genetic variation captured for height, body mass index, cardiovascular disease, and type-2 diabetes in the UK Biobank, only ≤ 10% is attributable to proximal regulatory regions within 10kb upstream of genes, while 12-25% is attributed to coding regions, 32-44% to intronic regions, and 22-28% to distal 10-500kb upstream regions. ≥ 60% of the variance contributed by these exonic, intronic and distal 10-500kb regions is underlain by many thousands of common variants, which on average have larger effect sizes than for other annotation groups. Up to 24% of all cis and coding regions of each chromosome are associated with each trait, with over 3,100 independent exonic and intronic regions and over 5,400 independent regulatory regions having ≥ 95% probability of contributing ≥ 0.001% to the genetic variance of these four traits. Thus, these quantitative and disease traits are truly complex. The BayesRR-RC prior gives robust model performance across the data analysed, providing an alternative to current approaches.
Severe acute respiratory syndrome coronavirus-2 (SARS-CoV-2) causes coronavirus disease-19 (COVID-19), a respiratory illness that can result in hospitalization or death. We investigated associations between rare genetic variants and seven COVID-19 outcomes in 543,213 individuals, including 8,248 with COVID-19. After accounting for multiple testing, we did not identify any clear associations with rare variants either exome-wide or when specifically focusing on (i) 14 interferon pathway genes in which rare deleterious variants have been reported in severe COVID-19 patients; (ii) 167 genes located in COVID-19 GWAS risk loci; or (iii) 32 additional genes of immunologic relevance and/or therapeutic potential. Our analyses indicate there are no significant associations with rare protein-coding variants with detectable effect sizes at our current sample sizes. Analyses will be updated as additional data become available, with results publicly browsable at https://rgc-covid19.regeneron.com.
Severe acute respiratory syndrome coronavirus-2 (SARS-CoV-2) causes coronavirus disease 2019 (COVID-19), a respiratory illness that can result in hospitalization or death. We used exome sequence data to investigate associations between rare genetic variants and seven COVID-19 outcomes in 586,157 individuals, including 20,952 with COVID-19. After accounting for multiple testing, we did not identify any clear associations with rare variants either exome wide or when specifically focusing on (1) 13 interferon pathway genes in which rare deleterious variants have been reported in individuals with severe COVID-19, (2) 281 genes located in susceptibility loci identified by the COVID-19 Host Genetics Initiative, or (3) 32 additional genes of immunologic relevance and/or therapeutic potential. Our analyses indicate there are no significant associations with rare protein-coding variants with detectable effect sizes at our current sample sizes. Analyses will be updated as additional data become available, and results are publicly available through the Regeneron Genetics Center COVID-19 Results Browser.
AbstractCritical illness in COVID-19 is caused by inflammatory lung injury, mediated by the host immune system. We and others have shown that host genetic variation influences the development of illness requiring critical care1or hospitalisation2;3;4following SARS-Co-V2 infection. The GenOMICC (Genetics of Mortality in Critical Care) study recruits critically-ill cases and compares their genomes with population controls in order to find underlying disease mechanisms.Here, we use whole genome sequencing and statistical fine mapping in 7,491 critically-ill cases compared with 48,400 population controls to discover and replicate 22 independent variants that significantly predispose to life-threatening COVID-19. We identify 15 new independent associations with critical COVID-19, including variants within genes involved in interferon signalling (IL10RB, PLSCR1), leucocyte differentiation (BCL11A), and blood type antigen secretor status (FUT2). Using transcriptome-wide association and colocalisation to infer the effect of gene expression on disease severity, we find evidence implicating expression of multiple genes, including reduced expression of a membrane flippase (ATP11A), and increased mucin expression (MUC1), in critical disease.We show that comparison between critically-ill cases and population controls is highly efficient for genetic association analysis and enables detection of therapeutically-relevant mechanisms of disease. Therapeutic predictions arising from these findings require testing in clinical trials.
Abstract Background The molecular factors which control circulating levels of inflammatory proteins are not well understood. Furthermore, association studies between molecular probes and human traits are often performed by linear model-based methods which may fail to account for complex structure and interrelationships within molecular datasets. Methods In this study, we perform genome- and epigenome-wide association studies (GWAS/EWAS) on the levels of 70 plasma-derived inflammatory protein biomarkers in healthy older adults (Lothian Birth Cohort 1936; n = 876; Olink® inflammation panel). We employ a Bayesian framework (BayesR+) which can account for issues pertaining to data structure and unknown confounding variables (with sensitivity analyses using ordinary least squares- (OLS) and mixed model-based approaches). Results We identified 13 SNPs associated with 13 proteins (n = 1 SNP each) concordant across OLS and Bayesian methods. We identified 3 CpG sites spread across 3 proteins (n = 1 CpG each) that were concordant across OLS, mixed-model and Bayesian analyses. Tagged genetic variants accounted for up to 45% of variance in protein levels (for MCP2, 36% of variance alone attributable to 1 polymorphism). Methylation data accounted for up to 46% of variation in protein levels (for CXCL10). Up to 66% of variation in protein levels (for VEGFA) was explained using genetic and epigenetic data combined. We demonstrated putative causal relationships between CD6 and IL18R1 with inflammatory bowel disease and between IL12B and Crohn’s disease. Conclusions Our data may aid understanding of the molecular regulation of the circulating inflammatory proteome as well as causal relationships between inflammatory mediators and disease.