Genes & Health (G&H) is a biomedical study of adult British Pakistani and Bangladeshi research volunteers enriched for autozygosity. Here we performed whole-exome sequencing in 44,028 G&H participants, establishing a large publicly available South Asian exome resource linked to longitudinal electronic health records. We performed exome-wide association analyses for 645 electronic health record-derived traits under additive and recessive models, and meta-analyses of 33 cardiometabolic traits with UK Biobank, finding more than 100 novel gene-phenotype associations. We identified 2,991 genes with rare biallelic predicted loss-of-function ('knockout') genotypes, 546 of which had not been previously reported. We show that drugs targeting genes with knockouts in adults are associated with a 2.2-fold higher likelihood of progressing beyond phase 1 clinical trials. We further illustrate how phenotypic profiles associated with knockout genotypes can enhance efficacy and safety assessment of drug targets and aid in the interpretation of variants with ambiguous clinical significance in autosomal recessive disease genes.
Understanding the function of genetic variants associated with human traits and diseases remains a significant challenge. Here, we combined analyses based on natural genetic variation and genetic engineering to dissect the function of 94 non-coding variants associated with hematological traits. We describe 22 genetic variants impacting hematological variation through gene expression. Further, through in-depth functional analysis, we illustrate how a rare, non-coding variant near the CUX1 transcription factor impacts megakaryopoiesis through the modulation of the CUX1 transcriptional cascade. Collectively, our findings enhance the functional interpretation of genetic association studies and advance understanding of how non-coding variants contribute to blood and immune system variation.
Type 2 diabetes (T2D) is a common and complex metabolic condition with significant heterogeneity within and across ancestries 1-4 . Compared with individuals of European ancestry (EUR), people of south Asian ancestry (SAS) have two to four-fold higher risk of T2D, develop the disease at younger ages and lower body mass index (BMI), and experience more rapid progression to complications 5-10 . Understanding the genetic basis of this is hindered by low representation of south Asians in genetic studies. Here, we perform an exome-wide association study of T2D in 13,674 cases and 41,024 controls from the Genes & Health study of British Pakistani and Bangladeshi individuals. We identify a novel rare variant in HNF4A - a canonical monogenic diabetes / MODY gene, in which missense variants would be expected to increase T2D risk. Surprisingly, HNF4A Pro437Ser is associated with a halved risk of T2D and reduced risk of diabetes-related complications but increased non-HDL cholesterol. We additionally characterise a T2D risk-increasing variant which is common only in South and East Asian ancestral groups ( GP2 Val429Met), which is associated with lower BMI and phenotypic and genetic markers of insulin deficiency. We validate our findings through replication in independent multi-ancestry cohorts, in vitro functional assays, and integration of proteogenomic analysis. These findings highlight how the study of under-represented populations can identify biological mechanisms associated with disease phenotypes enriched in those populations.
IntroductionIt is becoming increasingly evident that SARS-CoV-2 infection is here to stay. Therefore, understanding whether genetic variants may impact the response to the virus or vaccination is crucial. Studies on the genetic determinants of immune responses to SARS-CoV-2 have been limited by the scarcity of genetically homogenous populations and longitudinal designs that assess responses to both infection and vaccination in relation to individual genetic variation. MethodsHere we performed genotyping and whole-genome sequencing in a well-annotated and intensively followed population from the municipality of Vo’, which has previously provided critical insights into SARS-CoV-2 transmission, infection dynamics and COVID-19 clinical manifestations. ResultsWe identified 99 variants within the major histocompatibility complex (MHC) associated with altered T cell response dynamics following infection. These variants clustered into two semi-independent linkage disequilibrium (LD) blocks, respectively tagged by the HLA-A*01:01 allele and by SNP rs1611581. Additionally, when examining the response to vaccination, we identified 617 MHC genetic variants clustering into 27 semi-independent LD blocks that correlated with either increased or decreased TCR responses. We constructed a polygenic risk score (PRS) that comprehensively captures this genetic variation. Finally, structural modelling of selected variants affecting HLA proteins identified specific amino acid residuals most likely to influence interactions with SARS-CoV-2 epitopes, including arginine at position 114, isoleucine at position 97, and alanine at position 152 of the HLA-A molecule. ConclusionTogether, these findings provide robust evidence that genetic profiles modulate the immune response to SARS-CoV-2 in a longitudinal setting, offering insights that may inform further public health interventions.
Two decades of Genome Wide Association Studies (GWAS) have yielded hundreds of thousands of robust genetic associations to human complex traits and diseases. Nevertheless, the dissection of the functional consequences of variants lags behind, especially for non-coding variants (RNVs). Here we have characterised a set of rare, non-coding variants with large effects on haematological traits by integrating (i) a massively parallel reporter assay with (ii) a CRISPR/Cas9 screen and (iii) in vivo gene expression and transcript relative abundance analysis of whole blood and immune cells. After extensive manual curation we identify 22 RNVs with robust mechanistic hypotheses and perform an in-depth characterization of one of them, demonstrating its impact on megakaryopoiesis through regulation of the CUX1 transcriptional cascade. With this work we advance the understanding of the translational value of GWAS findings for variants implicated in blood and immunity. ### Competing Interest Statement M.I. is a trustee of the Public Health Genomics (PHG) Foundation, a member of the Scientific Advisory Board of Open Targets, and has research collaborations with AstraZeneca, Nightingale Health and Pfizer which are unrelated to this study. T.V. has received PhD studentship funding from AstraZeneca. K.K and D.P. are current employees and stockholders of AstraZeneca. P.A. is a current employee of Glaxosmithkline.
Genetic association studies have focused on testing additive models in cohorts with European ancestry. Little is known about recessive effects on common diseases, specifically for non-European ancestry. Genes & Health is a cohort of British Pakistani and Bangladeshi individuals with elevated rates of consanguinity and endogamy, making it suitable to study recessive effects. We imputed variants into a genotyped dataset (n = 44,190) by using two reference panels: a set of 4,982 whole-exome sequences from within the cohort and the Trans-Omics for Precision Medicine (TOPMed-r2) panel. We performed association testing with 898 diseases from electronic health records. 185 independent loci reached genome-wide significance (p < 5 × 10-8) under the recessive model, with p values lower than under the additive model, and >40% of these were novel. 140 loci demonstrated nominally significant (p < 0.05) dominance deviation p values, confirming a recessive association pattern. Sixteen loci in three clusters were significant at a Bonferroni threshold, accounting for multiple phenotypes tested (p < 5.4 × 10-12). In FinnGen, we replicated 44% of the expected number of Bonferroni-significant loci we were powered to replicate, at least one from each cluster, including an intronic variant in patatin-like phospholipase domain-containing protein 3 (PNPLA3; rs66812091) and non-alcoholic fatty liver disease, a previously reported additive association. We present evidence suggesting that the association is recessive instead (odds ratio [OR] = 1.3, recessive p = 2 × 10-12, additive p = 2 × 10-11, dominance deviation p = 3 × 10-2, and FinnGen recessive OR = 1.3 and p = 6 × 10-12). We identified a novel protective recessive association between a missense variant in SGLT4 (rs61746559), a sodium-glucose transporter with a possible role in the renin-angiotensin-aldosterone system, and hypertension (OR = 0.2, p = 3 × 10-8, dominance deviation p = 7 × 10-6). These results motivate interrogating recessive effects on common diseases more widely.
Human loss-of-function (LoF) variants affecting both copies of a gene ('human knockouts') provide a unique opportunity to directly study function and clinical impact of genes but are very rare in most populations sequenced to date. Here we study 1,569 British Bangladeshi and Pakistani adults who were recalled for plasma sampling for proteomic profiling using three distinct technologies (covering >12,000 proteins) from 55k whole exome sequenced Genes & Health adults - a cohort enriched for rare, biallelic (homozygous) variants due to high autozygosity. We identified 199 individuals with rare homozygous predicted LoF genotypes (pLoF) for which the respective cis-protein was measured by at least one technology, and observed extreme (> 3SDs) cis-protein underexpression in 41 individuals (median z-score = -9.72 (range: -19.61 to -4.78) and overexpression in 2 individuals (median z-score = 8.1 (range 4.80 - 11.40)), representing 19% of these variants. For missense homozygotes, we observed 158 individuals with significantly under-expressed cis-protein (median z-score = -6.95 (range: -28.16 to -3.95)) and 62 individuals over-expressed, median z-score = 5.65 (range: 4.57 to 25.08)). The majority (62%) of LoF knockout genes with an identified cis-protein effect had evidence from 2 or more platforms, highlighting the high confidence nature of these discoveries. Systematic clinical assessment of human knockouts with strong evidence of an impact on cis-protein abundance through multi-source electronic health record linkage enabled identification of 1) knockout carriers with rare disease features based on phenotypic similarity, 2) novel rare disease-causing variants, 3) evidence for reclassification of genes and variants of uncertain significance from ClinVar and rare disease panels, and 4) novel gene-phenotype associations in humans. Based on high confidence examples, we developed a machine learning model that predicted 1 in 4 pLOF and 9 in 10 missense variants are likely benign. In summary, our study provides strong human derived insights into the fundamental biology and clinical relevance of many genes and shows the value of proteogenomic studies of human knockout carriers. ### Competing Interest Statement The authors have declared no competing interest. ### Funding Statement Genes & Health is/has recently been core-funded by Wellcome (WT102627, WT210561), the Medical Research Council (UK) (M009017, MR/X009777/1, MR/X009920/1), Higher Education Funding Council for England Catalyst, Barts Charity (845/1796), Health Data Research UK (for London substantive site), and research delivery support from the NHS National Institute for Health Research Clinical Research Network (North Thames). We acknowledge the support of the National Institute for Health and Care Research Barts Biomedical Research Centre (NIHR203330); a delivery partnership of Barts Health NHS Trust, Queen Mary University of London, St George's University Hospitals NHS Foundation Trust and St George's University of London Genes & Health is/has recently been funded by Alnylam Pharmaceuticals, Genomics PLC; and a Life Sciences Industry Consortium of AstraZeneca PLC, Bristol-Myers Squibb Company, GlaxoSmithKline Research and Development Limited, Maze Therapeutics Inc, Merck Sharp & Dohme LLC, Novo Nordisk A/S, Pfizer Inc, Takeda Development Centre Americas Inc. We thank Social Action for Health, Centre of The Cell, members of our Community Advisory Group, and staff who have recruited and collected data from volunteers. We thank the NIHR National Biosample Centre (UK Biocentre), the Social Genetic & Developmental Psychiatry Centre (King's College London), Wellcome Sanger Institute, and Broad Institute for sample processing, genotyping, sequencing and variant annotation. This work uses data provided by patients and collected by the NHS as part of their care and support. This research utilised Queen Mary University of London's Apocrita HPC facility, supported by QMUL Research-IT. ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: G&H was approved by the London Southeast NRES Committee of the Health Research Authority (reference 14/LO/1240) on 16 Sept 2014. I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes Individual-level data from Genes & Health are available for bona fide researchers on application (https://www.genesandhealth.org/).
Type 2 diabetes (T2D) disproportionately affects individuals of South Asian ancestry (SAS), yet they remain underrepresented in genetic studies. We performed an exome-wide association study in 54,698 SAS T2D case: controls and follow-up metabolic trait evaluation. We identified ancestry-specific genes and protein-coding variants, including a SAS-specific variant in the known monogenic diabetes gene HNF4A (rs150776703, Pro437Ser), which was associated with protection from T2D (OR = 0.48, p = 2.8×10⁻¹□), diabetic eye disease, and gestational diabetes. Experimental interrogation of HNF4A Pro437Ser revealed context-dependent enhancement of HNF4A transcriptional activity, suggesting gain-of-function. We additionally describe a T2D risk increasing variant in GP2 (rs78193826, Val429Met, OR = 1.21, p = 5.14×10 -6 ), which was also associated with beta-cell dysfunction. These findings underscore how studying ancestrally distinct populations disproportionately affected by T2D can reveal novel disease genes, therapeutic hypotheses, and biological insight.
Autosomal recessive coding variants are well-known causes of rare disorders. We quantified the contribution of these variants to developmental disorders in a large, ancestrally diverse cohort comprising 29,745 trios, of whom 20.4% had genetically inferred non-European ancestries. The estimated fraction of patients attributable to exome-wide autosomal recessive coding variants ranged from ~2-19% across genetically inferred ancestry groups and was significantly correlated with average autozygosity. Established autosomal recessive developmental disorder-associated (ARDD) genes explained 84.0% of the total autosomal recessive coding burden, and 34.4% of the burden in these established genes was explained by variants not already reported as pathogenic in ClinVar. Statistical analyses identified two novel ARDD genes: KBTBD2 and ZDHHC16. This study expands our understanding of the genetic architecture of developmental disorders across diverse genetically inferred ancestry groups and suggests that improving strategies for interpreting missense variants in known ARDD genes may help diagnose more patients than discovering the remaining genes.
Myeloproliferative neoplasms (MPNs) are chronic cancers characterized by overproduction of mature blood cells. Their causative somatic mutations, for example, JAK2 V617F , are common in the population, yet only a minority of carriers develop MPN. Here we show that the inherited polygenic loci that underlie common hematological traits influence JAK2 V617F clonal expansion. We identify polygenic risk scores (PGSs) for monocyte count and plateletcrit as new risk factors for JAK2 V617F positivity. PGSs for several hematological traits influenced the risk of different MPN subtypes, with low PGSs for two platelet traits also showing protective effects in JAK2 V617F carriers, making them two to three times less likely to have essential thrombocythemia than carriers with high PGSs. We observed that extreme hematological PGSs may contribute to an MPN diagnosis in the absence of somatic driver mutations. Our study showcases how polygenic backgrounds underlying common hematological traits influence both clonal selection on somatic mutations and the subsequent phenotype of cancer.
AbstractUnderstanding the genetic basis of routinely-acquired blood tests can provide insights into several aspects of human physiology. We report a genome-wide association study of 42 quantitative blood test traits defined using Electronic Healthcare Records (EHRs) of ~50,000 British Bangladeshi and British Pakistani adults. We demonstrate a causal variant within the PIEZO1 locus which was associated with alterations in red cell traits and glycated haemoglobin. Conditional analysis and within-ancestry fine mapping confirmed that this signal is driven by a missense variant - chr16-88716656-G-TT - which is common in South Asian ancestries (MAF 3.9%) but ultra-rare in other ancestries. Carriers of the T allele had lower mean HbA1c values, lower HbA1c values for a given level of random or fasting glucose, and delayed diagnosis of Type 2 Diabetes Mellitus. Our results shed light on the genetic basis of clinically-relevant traits in an under-represented population, and emphasise the importance of ancestral diversity in genetic studies.
Gene misexpression is the aberrant transcription of a gene in a context where it is usually inactive. Despite its known pathological consequences in specific rare diseases, we have a limited understanding of its wider prevalence and mechanisms in humans. To address this, we analyzed gene misexpression in 4,568 whole-blood bulk RNA sequencing samples from INTERVAL study blood donors. We found that while individual misexpression events occur rarely, in aggregate they were found in almost all samples and a third of inactive protein-coding genes. Using 2,821 paired whole-genome and RNA sequencing samples, we identified that misexpression events are enriched in cis for rare structural variants. We established putative mechanisms through which a subset of SVs lead to gene misexpression, including transcriptional readthrough, transcript fusions, and gene inversion. Overall, we develop misexpression as a type of transcriptomic outlier analysis and extend our understanding of the variety of mechanisms by which genetic variants can influence gene expression.
Few genome-wide association studies (GWAS) analyzing genetic regulation of morphological traits of white blood cells have been reported. We carried out a GWAS of 12 morphological traits in 869 individuals from the general population of Sardinia, Italy. These traits, included measures of cell volume, conductivity and light scatter in four white-cell populations (eosinophils, lymphocytes, monocytes, neutrophils). This analysis yielded seven statistically significant signals, four of which were novel (four novel, PRG2, P2RX3, two of CDK6). Five signals were replicated in the independent INTERVAL cohort of 11 822 individuals. The most interesting signal with large effect size on eosinophil scatter (P-value = 8.33 x 10(-32), beta = -1.651, se = 0.1351) falls within the innate immunity cluster on chromosome 11, and is located in the PRG2 gene. Computational analyses revealed that a rare, Sardinian-specific PRG2:p.Ser148Pro mutation modifies PRG2 amino acid contacts and protein dynamics in a manner that could possibly explain the changes observed in eosinophil morphology. Our discoveries shed light on genetics of morphological traits. For the first time, we describe such large effect size on eosinophils morphology that is relatively frequent in Sardinian population.
Blood cells contain functionally important intracellular structures, such as granules, critical to immunity and thrombosis. Quantitative variation in these structures has not been subjected previously to large-scale genetic analysis. We perform genome-wide association studies of 63 flow-cytometry derived cellular phenotypes—including cell-type specific measures of granularity, nucleic acid content and reactivity—in 41,515 participants in the INTERVAL study. We identify 2172 distinct variant-trait associations, including associations near genes coding for proteins in organelles implicated in inflammatory and thrombotic diseases. By integrating with epigenetic data we show that many intracellular structures are likely to be determined in immature precursor cells. By integrating with proteomic data we identify the transcription factor FOG2 as an early regulator of platelet formation and α-granularity. Finally, we show that colocalisation of our associations with disease risk signals can suggest aetiological cell-types—variants in IL2RA and ITGA4 respectively mirror the known effects of daclizumab in multiple sclerosis and vedolizumab in inflammatory bowel disease.
Autosomal recessive (AR) coding variants are a well-known cause of rare disorders. We quantified the contribution of these variants to developmental disorders (DDs) in the largest and most ancestrally diverse sample to date, comprising 29,745 trios from the Deciphering Developmental Disorders (DDD) study and the genetic diagnostics company GeneDx, of whom 20.4% have genetically-inferred non-European ancestries. The estimated fraction of patients attributable to exome-wide AR coding variants ranged from ∼2% to ∼18% across genetically-inferred ancestry groups, and was significantly correlated with the average autozygosity (r=0.99, p=5x10-6). Established AR DD-associated (ARDD) genes explained 90% of the total AR coding burden, and this was not significantly different between probands with genetically-inferred European versus non-European ancestries. Approximately half the burden in these established genes was explained by variants not already reported as pathogenic in ClinVar. We estimated that ∼1% of undiagnosed patients in both cohorts were attributable to damaging biallelic genotypes involving missense variants in established ARDD genes, highlighting the challenge in interpreting these. By testing for gene-specific enrichment of damaging biallelic genotypes, we identified two novel ARDD genes passing Bonferroni correction, KBTBD2 (p=1x10-7) and CRELD1 (p=9x10-8). Several other novel or recently-reported candidate genes were identified at a more lenient 5% false-discovery rate, including ZDHHC16 and HECTD4 . This study expands our understanding of the genetic architecture of DDs across diverse genetically-inferred ancestry groups and suggests that improving strategies for interpreting missense variants in known ARDD genes may allow us to diagnose more patients than discovering the remaining genes.### Competing Interest StatementKM and VDU are employees of GeneDx. ZZ, KR and RT were formerly employees of GeneDx, and KR and RT are now employees of Geisinger Health System. EJG is an employee of and holds shares in Adrestia Therapeutics. MEH is a co-founder of, consultant to and holds shares in Congenica, a genetics diagnostic company.### Funding StatementThe DDD study presents independent research commissioned by the Health Innovation Challenge Fund (grant number HICF-1009-003). This study makes use of DECIPHER, which is funded by the Wellcome Trust. The full acknowledgements can be found at www.ddduk.org/access.html. We additionally thank Rachel Hobson, Erwan Delage and the Human Genetics Informatics team at Sanger for their input on the DDD study. This research was funded in whole, or in part, by the Wellcome Trust Grant 220540/Z/20/A, 'Wellcome Sanger Institute Quinquennial Review 2021-2026'. For the purpose of Open Access, the author has applied a CC BY public copyright licence to any Author Accepted Manuscript version arising from this submission. DSM is supported by a Gates Cambridge Scholarship (OPP1144).### Author DeclarationsI confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained.YesThe details of the IRB/oversight body that provided approval or exemption for the research described are given below:The Cambridge South Research Ethics Committee (10/H0305/83) and the Republic of Ireland Research Ethics Committee (GEN/284/12) gave ethical approval for DDD The Western Institutional Review Board, Puyallup, gave ethical approval for research on the GeneDx data (WIRB 20162523)I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals.YesI understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance).YesI have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable.Yes