Safety-related issues account for approximately 25% of failures in new drug discovery programs. On top of that, many are discovered during post-marketing surveillance, significantly limiting drug utility and application. To proactively address these concerns, we developed a genetics-led strategy leveraging Mendelian Randomization (MR) across large-scale genetic datasets from the Million Veteran Program, FinnGen, and UK Biobank. By mapping genetic variants associated with gene expression and protein abundance to 1,449 harmonized human phenotypes, we systematically identified potential adverse drug reactions (ADR). Our extensive MR analysis, encompassing 16,915 protein-coding genes, demonstrated the capacity to predict hundreds of known ADR for approved medications, with approximately 40% corroborated by FDA Adverse Event Reporting System (FAERS) data. Additionally, we found significant enrichment of identified gene-mechanism pairs in clinical trials terminated early due to safety concerns, highlighting the clinical utility of genetics-informed safety prediction. Notably, immune-related pathways were prominently associated with ADR, indicating particular sensitivity within immune modulation targets. Our comprehensive atlas, integrating genetic evidence with pharmacological mechanisms, provides a robust predictive framework for anticipating drug safety, potentially enhancing decision-making in drug development and pharmacovigilance. An interactive web interface allowing filtering by gene, phenotype, drug phase, and mechanism of action is available at https://shiny.parse-health.org/safety/.
New obesity medications have demonstrated efficacy in trials, but their real-world deployment is partly limited by the absence of approaches that identify individuals for treatment based on risks for obesity-related complications. Here we present a risk prediction model to guide prioritization of high-risk individuals. In a population-based sample of ~200,000 individuals with a body mass index (BMI) exceeding 27 kg m-2, our machine learning framework identified the 20 most informative features, from among thousands tested, that predict future onset of 18 complications of obesity, providing information beyond BMI. An integrated model (OBSCORE) successfully stratified individuals into risk groups based on incidence over 10 years: for example, 5.7%, 1.8%, 0.9%, 0.4% and 0.1% for cardiovascular mortality. We demonstrate generalizability of the model in independent populations of European and non-European ancestry and, in SURMOUNT-1 trial participants, show that weight loss was similar across baseline OBSCORE risk groups and that predicted risks decreased following treatment with tirzepatide. In summary, OBSCORE provides a framework for prioritizing high-risk individuals with overweight or obesity based on their risk of obesity-related complications, complementing BMI-based frameworks.
The glucose-dependent insulinotropic polypeptide receptor (GIPR) is a major therapeutic target in type 2 diabetes and obesity. Missense variation in GIPR could confer phenotypic effects through alterations to constitutive activity or functional responses to GIP or pharmacological agonists. In this study, we aimed to provide a deep understanding of the molecular mechanisms that underpin the cellular and physiological impacts of GIPR coding variation by studying 30 GIPR coding variants in cellular models and pancreatic islets. Many variants showed impaired GIP-induced cyclic adenosine monophosphate responses, and population-based association analysis highlighted that these loss-of-function variants decrease body mass index but increase glycemia. In many cases, reduced function was partly driven by reduced expression at the cell surface due to impaired stability and redirection toward proteasomal degradation. Molecular dynamics simulations suggest distinct variant-induced perturbations in inter- and intrahelical interactions, which interfere with receptor stability. This study highlights the mechanisms and consequences of GIPR coding variation, which may have implications for the therapeutic targeting of this receptor.
Genes & Health (G&H) is a biomedical study of adult British Pakistani and Bangladeshi research volunteers enriched for autozygosity. Here we performed whole-exome sequencing in 44,028 G&H participants, establishing a large publicly available South Asian exome resource linked to longitudinal electronic health records. We performed exome-wide association analyses for 645 electronic health record-derived traits under additive and recessive models, and meta-analyses of 33 cardiometabolic traits with UK Biobank, finding more than 100 novel gene-phenotype associations. We identified 2,991 genes with rare biallelic predicted loss-of-function ('knockout') genotypes, 546 of which had not been previously reported. We show that drugs targeting genes with knockouts in adults are associated with a 2.2-fold higher likelihood of progressing beyond phase 1 clinical trials. We further illustrate how phenotypic profiles associated with knockout genotypes can enhance efficacy and safety assessment of drug targets and aid in the interpretation of variants with ambiguous clinical significance in autosomal recessive disease genes.
Understanding the genetic regulation of circulating protein levels can provide new insights into disease mechanisms. Here, we present the largest proteogenomic study to date (n = 78,664 participants across 38 studies), identifying >24,000 protein quantitative trait loci (QTLs) associated with 1,116 proteins, acting near to (n = 5,040) or distant (n = 19,698) from the cognate gene. Using machine learning-guided effector gene assignment, we provide genetic evidence for pathways, cell types, and tissues that modulate circulating protein levels, highlighting N-linked glycosylation as an important regulatory pathway. We demonstrate that genetic instruments of protein production/function (“cis”) versus modulation (“trans”) reveal distinct phenotypic insights. We identify proteins as candidates for drug targets and engagement (e.g., plasma furin and cardiovascular diseases) by comparing cis-based genetic evidence with protein-disease associations. Systematic triangulation of trans-protein QTLs (pQTLs) with genetic and protein associations across many diseases highlights potential drug repurposing opportunities, e.g., tyrosine kinase 2 (TYK2) inhibitors for rheumatoid arthritis. Our multi-cohort meta-analyses generate proteogenomic insights into disease mechanisms and new treatment opportunities.
Heart failure with preserved ejection fraction (HFpEF) and metabolic dysfunction-associated steatotic liver disease (MASLD) are increasingly prevalent, interrelated conditions driven by the global rise in obesity and metabolic syndrome. Once viewed in isolation, HFpEF and MASLD are now recognized as organ-specific manifestations of shared systemic metabolic dysfunction. Evidence from the past decade highlights not only overlapping risk factors but also a dynamic, bidirectional inter-organ crosstalk between the liver and the heart that shapes their natural history. In this Review, we explore the epidemiological and mechanistic basis of the MASLD-HFpEF connection, focusing on shared metabolic drivers such as lipotoxicity, meta-inflammation and oxidative stress. We also discuss emerging liver-derived mediators, including hepatokines, metabolites and extracellular vesicles, that influence cardiac structure and function. Finally, we highlight diagnostic and therapeutic strategies relevant to both conditions and propose a multiorgan framework to improve their clinical recognition and management. Understanding the liver-heart axis is key to rethinking cardiometabolic disease beyond organ silos and towards more integrated, mechanism-based approaches.
Circulating proteins act as important hormonal signals of nutrient intake. We aimed to systematically characterise the time-resolved proteomic response to glucose ingestion in humans, and to assess its robustness following prolonged complete caloric restriction. We conducted oral glucose tolerance tests (OGTTs) in 11 healthy volunteers before and after 7 days of complete caloric restriction and measured the response of >2900 targets through high-resolution plasma protein profiling. We identified a signature of 44 proteins that changed significantly following glucose ingestion, which was reproducible after 7 days without food, and was strongly (20-fold) enriched for ‘stomach-specific’ proteins. We report that annexin A10 (ANXA10) shows the most significant post-glucose change observed, similar to the trajectories of secreted hormones. We present observational human evidence from multiple sources suggesting that ANXA10 is secreted upon sensing an increase in gastric pH, with the stomach as the major contributing tissue. Despite a profound metabolic shift after 7 days of complete caloric restriction, characterised by delayed insulin secretion and postprandial hyperglycaemia, only four proteins showed robust evidence for a differential trajectory during both OGTTs. This included plasma levels of tryptophanyl-tRNA synthetase 1 (WARS), for which we found a genetic association with glucose homeostasis and coronary artery disease. Our exploratory study identifies the proteomic response to glucose ingestion and demonstrates its reproducibility despite major shifts in glucose homeostasis. We characterise the gastrointestinal origin of these changes, and hypothesise a hitherto under-recognised role for sensing of changes in gastric pH on the plasma proteome.
Type 2 diabetes (T2D) is a common and complex metabolic condition with significant heterogeneity within and across ancestries 1-4 . Compared with individuals of European ancestry (EUR), people of south Asian ancestry (SAS) have two to four-fold higher risk of T2D, develop the disease at younger ages and lower body mass index (BMI), and experience more rapid progression to complications 5-10 . Understanding the genetic basis of this is hindered by low representation of south Asians in genetic studies. Here, we perform an exome-wide association study of T2D in 13,674 cases and 41,024 controls from the Genes & Health study of British Pakistani and Bangladeshi individuals. We identify a novel rare variant in HNF4A - a canonical monogenic diabetes / MODY gene, in which missense variants would be expected to increase T2D risk. Surprisingly, HNF4A Pro437Ser is associated with a halved risk of T2D and reduced risk of diabetes-related complications but increased non-HDL cholesterol. We additionally characterise a T2D risk-increasing variant which is common only in South and East Asian ancestral groups ( GP2 Val429Met), which is associated with lower BMI and phenotypic and genetic markers of insulin deficiency. We validate our findings through replication in independent multi-ancestry cohorts, in vitro functional assays, and integration of proteogenomic analysis. These findings highlight how the study of under-represented populations can identify biological mechanisms associated with disease phenotypes enriched in those populations.
Precision medicine tailors prevention, diagnosis and treatment of cardiometabolic diseases to individual genetic, environmental and lifestyle determinants, with the potential to fundamentally change healthcare. However, low-income and middle-income countries (LMICs) and small island developing states (SIDS) experience severe implementation barriers: inadequate healthcare infrastructure, prohibitive costs, under-representation in genomic datasets and additional SIDS-specific constraints. This Perspective advances three specific contributions beyond generic equity calls. First, it delineates distinct precision medicine pathways for larger LMICs versus SIDS, highlighting SIDS opportunities for regional consortia, shared sequencing and/or biobanking hubs and technological leapfrogging via mobile health platforms and digital phenotyping. Second, it emphasizes practical and high-impact entry points that are financially sustainable. Additionally, it advocates for integrating polygenic risk-based stratification into existing non-communicable disease care pathways rather than establishing separate specialist services. Third, it delineates a staged implementation framework that prioritizes ethical oversight and robust data governance, underscoring the importance of privacy safeguards, data sovereignty, equitable benefit sharing, community consent mechanisms and alignment with the Sustainable Development Goals to minimize associated risks of exploitation. Equitable partnerships between LMICs and high-income countries, expansion of diverse genomic data and community-driven innovation will ensure that precision tools effectively target metabolic phenotypes in LMICs and SIDS while advancing global health equity.
Disease-associated variants reside frequently in noncoding cis-regulatory elements (CREs), yet their functional consequences remain poorly understood. We performed a large-scale lentiMPRA in human excitatory neurons, quantifying the impact of >46,000 naturally occurring variants across >27,000 candidate CREs near 524 disease-associated genes. These data improved regulatory variant effect predictions beyond state-of-the-art models. Significant allelic effects occurred at comparable rates across common, rare, and singleton variants, demonstrating that, within MPRA-measurable effects, population frequency carries limited information about per-variant regulatory impact. Variant effect detectability and magnitude were governed primarily by baseline activity of the enclosing regulatory element and local sequence context. Regulatory effects were distributed across numerous transcription factors rather than concentrated in master regulators, consistent with a combinatorial enhancer architecture. We establish a large-scale functional variant catalog and provide a complementary benchmark and resource for developing and evaluating models of noncoding regulatory variation.
Kidney disease disproportionately affects populations of African ancestry, yet most genetic studies have focused on Europeans. Here, we present a three-stage genome-wide association study meta-analysis of estimated glomerular filtration rate in ~26,000 individuals across Eastern, Western, and Southern Africa and ~81,000 African-ancestry individuals in the diaspora. Continental African meta-analysis identifies four independent genome-wide significant loci, including two previously unreported loci. Pan-African meta-analysis identifies 19 independent loci, including three previously unreported loci. Fine-mapping reveals four loci with high causality probability, and phenome-wide analyses demonstrate pleiotropic effects on cardiometabolic and immunological traits. Notably, APOL1 high-risk variants strongly associated with kidney disease in African Americans show markedly lower frequency and attenuated effects in continental Africa, indicating potential distinct genetic architectures. Polygenic scores from genetically similar populations significantly outperformed those from distant cohorts. These findings demonstrate the necessity of conducting genomic research across diverse African populations to enable equitable health outcomes.
Foundation models trained on electronic healthcare records (EHRs) have gained traction with the aim to transform personalised medicine. However, their interpretability is bound to redescribing the records the models were trained on, missing implicitly learned concepts and biases. Here, we show that human genetics provides an orthogonal layer to surface implicitly learned biological concepts and otherwise hidden risk factors. Re‑implementing the generative transformer Delphi‑2M in >500,000 UK Biobank participants, we performed genome‑wide association testing on its 120 learned embeddings and identified 434 genome‑wide‑significant signals across 151 independent loci and 98 embeddings, revealing a heritable structure that feature‑attribution methods cannot recover. Effector-gene mapping implicated cholesterol metabolism and an IL-1-family epithelial-alarmin pathway, supported by strong (>50-fold) enrichment for variants previously associated with blood lipids, body-mass index, and asthma. Loci recovered the targets of essentially all approved lipid-lowering and severe-asthma therapies, and another twelve drugs not obvious from genetic results based on single ICD-10 GWAS. Yet, embeddings poorly explained variation in pleiotropic risk factors, while still retaining most of their predictive value. Substantial improvements in predictive performance were hence confined to a minority of common diseases by adding specific diagnostic or organ‑derived markers. Our findings suggest that human genetics might be most powerful as an orthogonal explanatory or regularising layer to train the next generation of EHR-based foundation models that likely benefit most from the addition of targeted biomarkers to advance personalised medicine.
Understanding genetic variation associated with differences in plasma protein levels can elucidate human disease mechanisms. Here we demonstrate how untargeted nanoparticle-enriched mass spectrometry (MS)-based plasma proteomics delivers quantitatively and qualitatively different insights compared to two affinity-based assays in a sample of ~1,400 British South Asian individuals. We identify >1,200 significant locus-protein associations (P < 8.7 × 10-12; n = 895 cis-protein quantitative trait loci (pQTLs)), more than half of which have not been reported previously. Cross-platform comparison demonstrated that multiple platforms are required to capture the full spectrum of pQTLs of blood proteins. We combine proteogenomic results with evidence from multiple biological domains to suggest a potential role of 21 proteins in the pathology of 44 diseases, including a previously uncharacterized role of immunoglobulin λ variable 3-21 in the development of Graves' disease. Our results demonstrate the potential of MS-based blood proteomics in non-European ancestries for pQTL discovery and the need to consolidate proteogenomic evidence to confidently assign proteins to disease pathology.
Abstract BACKGROUND Individuals who develop coronary artery disease (CAD) are clinically and mechanistically heterogeneous, and understanding this variation is crucial for precise risk stratification and tailored interventions. However, the molecular mechanisms that connect these two kinds of heterogeneity remain unclear, limiting progress toward biologically grounded risk stratification and targeted interventions. Here, we investigated the heterogeneity of individuals who develop CAD by leveraging plasma proteomic signatures, placed individuals along continuous metabolic gradients and revealed the molecular programs underlying these patterns, thereby linking mechanistic variation to clinical heterogeneity. METHODS AND RESULTS From 42,803 UK Biobank participants, including 3,713 individuals who developed CAD within 10 years (incident CAD), we first identified a 320-protein panel from 2,923 baseline proteins that improved prediction of incident CAD beyond clinical risk scores. Using reverse graph embedding, we reduced the proteomic data to two dimensions and mapped each incident case onto the resulting two-dimensional latent proteomic space. These proteomic dimensions show significant associations with cardiometabolic and kidney-related clinical markers. The patterns were replicated in the EPIC-Norfolk study. Phenome-wide Cox regression analyses further linked these proteomic dimensions to 10-year incidence rates for various diseases, including type 2 diabetes, obesity, and chronic kidney disease (CKD). Furthermore, adding the proteomic dimensions to clinical variable-based Cox regression model improved prediction of 10-year incidence of CKD and other diseases, demonstrating the value of proteomic dimensions beyond conventional clinical risk factors. Moreover, individuals with prevalent CAD (diagnosed before proteomic sampling) exhibited high, metabolically adverse dimension values, indicating that these axes capture cumulative metabolic burden. Pathway enrichment analyses implicated altered extracellular matrix organization and immune programs among the proteins contributing to the proteomic dimensions. CONCLUSIONS Our findings demonstrate that plasma proteomic signatures can dissect the heterogeneity of individuals who develop CAD in continuous phenotypic gradients, improve prediction of CAD and comorbidities, and map underlying biological mechanisms. Clinical Perspective What is New? In 42,803 UK Biobank participants, baseline plasma proteomics identified a protein panel that improved prediction of incident coronary artery disease (CAD) beyond conventional clinical risk scores, including AHA PREVENT, and defined two continuous proteomic dimensions that captured clinically relevant heterogeneity among individuals who later developed coronary artery disease; these patterns were replicated in EPIC-Norfolk cohort. These proteomic dimensions captured distinct cardiometabolic and kidney-related patterns, improved prediction of multiple comorbidities, including chronic kidney disease, beyond clinical risk factors, and reflected biological programs involving metabolic, immune, and extracellular matrix pathways. What Are the Clinical Implications? Plasma proteomics can serve as a biologically grounded tool to improve prediction of CAD and related comorbidities, characterize clinically relevant heterogeneity, and provide disease-relevant information beyond standard clinical risk-factor measurements.
Announced in this Comment and in collaboration with Nature Medicine is the convening of the Data-Driven Decision Support in Obesity Management Commission, to promote adequate scientific evidence to support obesity management across global populations.
Assessment of biological aging using proteomic clocks may enhance risk prediction and elucidate the molecular links between aging and chronic diseases. Here, among 17,473 participants of the European Prospective Investigation into Cancer and Nutrition, we examined associations of plasma SomaScan-based proteomic clocks, including organ-specific clocks, with risk factors, 24 incident chronic diseases and all-cause mortality, over up to 28 years of follow-up. Replication was conducted in the Whitehall II study. We show that the global age gap, an age acceleration score combining proteomic clocks, was associated with smoking, alcohol consumption, physical inactivity and higher risk of mortality, cardiovascular diseases, dementia and cancers of the liver, upper aero-digestive tract, lung and kidney. Lung, kidney and stomach cancers were more strongly associated with related organ-specific age gaps. Predictive performance of proteomic clocks for mortality was comparable to that of classical lifestyle risk factors. In summary, proteomic clocks appear promising biomarkers of generalized age-related disease risk.
BACKGROUND:Limited evidence exists for effect modification of genetic characteristics on the associations of food consumption and incident type 2 diabetes (T2D). OBJECTIVES:We aimed to investigate whether the food-T2D association would vary by genetic susceptibility to metabolic traits. METHODS:We analyzed data from 9542 incident T2D cases and a subcohort of 12,477 participants nested within the 340,234-participant cohort recruited in 1991-1998 and followed up for 10.9 y on average in 8 European countries. Polygenic risk scores (PRSs) for higher body mass index, insulin resistance, and T2D were constructed. Fifteen dietary variables potentially associated with T2D, obtained with cohort-specific self-reported dietary assessment, were examined: fruits, green leafy vegetables, root vegetables, wholegrains, rice, legumes, nuts and seeds, fermented dairy, red meat, processed meat, fish, eggs and egg products, sugar-sweetened beverages, coffee, and tea. A cross-product term between each PRS and each food/beverage was evaluated by genotyping chip and country with Prentice-weighted Cox regression for incident T2D, and stratum-specific estimates were meta analyzed, followed by Benjamini-Yekutieli multiple-testing correction. RESULTS:Accounting for multiple tests of 3 PRSs × 15 dietary items, no evidence of statistical interaction was evident on either a multiplicative or additive scale, with exp(β for a multiplicative interaction) (95% confidence interval) ranging from 0.84 (0.64, 1.10) (root vegetables and PRS for T2D) to 1.45 (0.78-2.76) (fish and PRS for T2D). CONCLUSIONS:Genetic susceptibility to high-risk metabolic traits did not modify the diet-T2D associations in European populations. Acknowledging the limitations of current PRS-based methods to detect gene-diet interactions, research should continue into the potential for precision nutrition and tailored food-based dietary guidance for T2D prevention.
Understanding genetic variation underlying differences in plasma protein levels can elucidate human disease mechanisms, but prior evidence was almost entirely derived from white Europeans using protein-preselected affinity reagents. Here, we integrate exome sequencing and common non-coding variation with untargeted nanoparticle enriched mass spectrometry (MS)-based plasma proteomics (Seer Proteograph XT: n=8067 protein groups) in >1,400 British-Bangladeshi and British-Pakistani individuals. We gain quantitatively and qualitatively different insights compared to two affinity-based assays in the same samples (SomaLogic 11k: n=9685 and Olink HT: n=5416 proteins), both in terms of proteins covered, and new proteogenomic insights into disease biology. Considering both additive and non-additive genetic effects, we identify >1,200 significant variant-protein associations (n=895 cis-protein quantitative trait loci (pQTL)), half of which are novel. Cross-platform comparison demonstrated that inconsistencies in pQTL discovery are mostly explained by technical variation, and that multiple platforms are required to capture the full spectrum of pQTLs of blood proteins. We integrate proteogenomic evidence with orthogonal human genetic, experimental, and single cell expression data to consolidate a potential role of 21 proteins in the pathology of 44 diseases: e.g., a novel role of high IGLV3-21 in the development of Grave's disease elucidating B-cell mediated autoimmunity. Our results demonstrate the potential of MS-based blood proteomics in diverse ancestries for pQTL discovery, including functional characterization of missense variants, and the need to consolidate evidence from multiple biological domains to confidently assign proteins to disease pathology to guide drug target identification and drug repurposing. ### Competing Interest Statement The authors have declared no competing interest. ### Funding Statement Genes & Health is/has recently been core-funded by Wellcome (WT102627, WT210561), the Medical Research Council (UK) (M009017, MR/X009777/1, MR/X009920/1), Higher Education Funding Council for England Catalyst, Barts Charity (845/1796), Health Data Research UK (for London substantive site), and research delivery support from the NHS National Institute for Health Research Clinical Research Network (North Thames). We acknowledge the support of the National Institute for Health and Care Research Barts Biomedical Research Centre (NIHR203330); a delivery partnership of Barts Health NHS Trust, Queen Mary University of London, St George's University Hospitals NHS Foundation Trust and St George's University of London Genes & Health is/has recently been funded by Alnylam Pharmaceuticals, Genomics PLC; and a Life Sciences Industry Consortium of AstraZeneca PLC, Bristol-Myers Squibb Company, GlaxoSmithKline Research and Development Limited, Maze Therapeutics Inc, Merck Sharp & Dohme LLC, Novo Nordisk A/S, Pfizer Inc, Takeda Development Centre Americas Inc. We thank Social Action for Health, Centre of The Cell, members of our Community Advisory Group, and staff who have recruited and collected data from volunteers. We thank the NIHR National Biosample Centre (UK Biocentre), the Social Genetic & Developmental Psychiatry Centre (King's College London), Wellcome Sanger Institute, and Broad Institute for sample processing, genotyping, sequencing and variant annotation. This work uses data provided by patients and collected by the NHS as part of their care and support. This research utilised Queen Mary University of London's Apocrita HPC facility, supported by QMUL Research-IT, http://doi.org/10.5281/zenodo.438045 The work was co-funded by the European Union (ERC, GenDrug, 101116072) to M.P.. Views and opinions expressed are however those of the author(s) only and do not necessarily reflect those of the European Union or the European Research Council. Neither the European Union nor the granting authority can be held responsible for them. The funders had no role in study design, data collection and analysis, decision to publish or preparation of the manuscript. This work was supported by the DFG (German Research Foundation) to K.D. (Walter Benjamin Fellowship, Grant Number: 547107463), and the Friede Springer Cardiovascular Prevention Center at Charite - Universitaetsmedizin Berlin, Germany to A.W. ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: G&H was approved by the London Southeast NRES Committee of the Health Research Authority (14/LO/1240). I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes Individual-level data from Genes & Health are available for bona fide researchers on application (https://www.genesandhealth.org/). Genome-wide summary statistics well be released upon publication.
Human loss-of-function (LoF) variants affecting both copies of a gene ('human knockouts') provide a unique opportunity to directly study function and clinical impact of genes but are very rare in most populations sequenced to date. Here we study 1,569 British Bangladeshi and Pakistani adults who were recalled for plasma sampling for proteomic profiling using three distinct technologies (covering >12,000 proteins) from 55k whole exome sequenced Genes & Health adults - a cohort enriched for rare, biallelic (homozygous) variants due to high autozygosity. We identified 199 individuals with rare homozygous predicted LoF genotypes (pLoF) for which the respective cis-protein was measured by at least one technology, and observed extreme (> 3SDs) cis-protein underexpression in 41 individuals (median z-score = -9.72 (range: -19.61 to -4.78) and overexpression in 2 individuals (median z-score = 8.1 (range 4.80 - 11.40)), representing 19% of these variants. For missense homozygotes, we observed 158 individuals with significantly under-expressed cis-protein (median z-score = -6.95 (range: -28.16 to -3.95)) and 62 individuals over-expressed, median z-score = 5.65 (range: 4.57 to 25.08)). The majority (62%) of LoF knockout genes with an identified cis-protein effect had evidence from 2 or more platforms, highlighting the high confidence nature of these discoveries. Systematic clinical assessment of human knockouts with strong evidence of an impact on cis-protein abundance through multi-source electronic health record linkage enabled identification of 1) knockout carriers with rare disease features based on phenotypic similarity, 2) novel rare disease-causing variants, 3) evidence for reclassification of genes and variants of uncertain significance from ClinVar and rare disease panels, and 4) novel gene-phenotype associations in humans. Based on high confidence examples, we developed a machine learning model that predicted 1 in 4 pLOF and 9 in 10 missense variants are likely benign. In summary, our study provides strong human derived insights into the fundamental biology and clinical relevance of many genes and shows the value of proteogenomic studies of human knockout carriers. ### Competing Interest Statement The authors have declared no competing interest. ### Funding Statement Genes & Health is/has recently been core-funded by Wellcome (WT102627, WT210561), the Medical Research Council (UK) (M009017, MR/X009777/1, MR/X009920/1), Higher Education Funding Council for England Catalyst, Barts Charity (845/1796), Health Data Research UK (for London substantive site), and research delivery support from the NHS National Institute for Health Research Clinical Research Network (North Thames). We acknowledge the support of the National Institute for Health and Care Research Barts Biomedical Research Centre (NIHR203330); a delivery partnership of Barts Health NHS Trust, Queen Mary University of London, St George's University Hospitals NHS Foundation Trust and St George's University of London Genes & Health is/has recently been funded by Alnylam Pharmaceuticals, Genomics PLC; and a Life Sciences Industry Consortium of AstraZeneca PLC, Bristol-Myers Squibb Company, GlaxoSmithKline Research and Development Limited, Maze Therapeutics Inc, Merck Sharp & Dohme LLC, Novo Nordisk A/S, Pfizer Inc, Takeda Development Centre Americas Inc. We thank Social Action for Health, Centre of The Cell, members of our Community Advisory Group, and staff who have recruited and collected data from volunteers. We thank the NIHR National Biosample Centre (UK Biocentre), the Social Genetic & Developmental Psychiatry Centre (King's College London), Wellcome Sanger Institute, and Broad Institute for sample processing, genotyping, sequencing and variant annotation. This work uses data provided by patients and collected by the NHS as part of their care and support. This research utilised Queen Mary University of London's Apocrita HPC facility, supported by QMUL Research-IT. ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: G&H was approved by the London Southeast NRES Committee of the Health Research Authority (reference 14/LO/1240) on 16 Sept 2014. I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes Individual-level data from Genes & Health are available for bona fide researchers on application (https://www.genesandhealth.org/).