Proteomics holds great promise for identifying potentially druggable effectors of common diseases, yet its application at population-scale across diverse ancestries, remains challenging. Here, we developed genetic imputation models for 2,594 plasma proteins using proteomic and genetic data from 54,219 UK Biobank participants, validating their performance across multiple ancestry groups and in an independent cohort. Plasma proteomes were then imputed for over 640,000 participants in the UK Biobank and the All of Us Research Program. To assess its aetiological value at population-scale, a further proteome-wide association study of cardiovascular diseases was performed across six genetic ancestries. We identified ∼9000 protein-disease associations across 89 cardiovascular conditions (PheCodes), the majority of which show consistent effects across ancestries and biobanks, with many comprising known targets of drugs either approved or under development. The associations reveal both shared and distinct proteomic signatures across cardiovascular conditions and defined clusters of distinct pathophysiology with shared underlying molecular pathways. Integration of data on tissue specificity and single-cell transcriptomics prioritised liver-derived proteins in circulation as candidate effectors of coronary artery disease, highlighting inter-alpha-trypsin inhibitor heavy chain H4 (ITIH4) as a putative effector. Using a liver-targeted CRISPR gene-editing platform, we show that in vivo disruption of ITIH4 reduces plasma cholesterol and pro-atherogenic lipid species in a preclinical model, consistent with a causal role in cardiovascular disease. Our study enables study of large-scale proteomics in diverse populations, provides a systematic map of protein associations of cardiovascular diseases, and demonstrates the utility of genetically imputed proteomes for target discovery and experimental validation. To facilitate proteomic analyses for the research community, the resultant models and association results have been made freely available through the OmicsPred platform.
Abstract Clonal haematopoiesis (CH) becomes ubiquitous as humans age. The role of somatic driver mutations in its development has been studied widely, but little is known about CH without identified genetic drivers, also known as “CH with unknown drivers” (CH-UD). A fundamental unresolved question is whether CH-UD is driven by undiscovered somatic genetic drivers or by other cell-heritable traits. Here, to investigate this, we develop a new machine learning classifier to improve CH-UD detection from whole-genome sequencing data. After excluding 77,885 individuals with previously documented driver CH or mosaic chromosomal alterations (mCA), we applied our classifier to 407,512 UK Biobank participants and identified 26,963 (6.6%) with CH-UD. A genome-wide association study (GWAS) of common germline variants identified 31 polymorphic loci associated with predisposition to CH-UD. Of these, 25 were associated with other forms of CH at genome-wide significance. Linkage Disequilibrium Score Regression analyses revealed an unexpectedly high genetic correlation (r g =0.794) between CH-UD and non- DNMT3A driver CH, indicative of a remarkable overlap between the genetic aetiologies of the two phenomena. Analysis of 2,941 plasma protein measurements in 47,757 individuals revealed that TCL1A was the most significantly elevated plasma protein in CH-UD, mirroring the finding that the TCL1A locus was in the top two most significant associations of CH-UD GWAS and TET2 -CH and ASXL1 -CH GWAS, the two most common forms of non-DNMT3A-CH. Furthermore, TCL1A plasma levels rose steadily with age even in those without detectable CH, particularly among carriers of the common TCL1A risk variant (rs2887399-G), potentially via stochastic promoter demethylation as described in TET2 -CH and ASXL1 -CH. Phenome-wide association analysis of 13,225 binary and 1,682 quantitative traits revealed that, similarly to non- DNMT3A -CH, CH-UD was significantly associated with several malignant (haematological and solid organ) and non-malignant (including cardiovascular and renal) diseases. Our findings reveal striking genetic and phenotypic similarities between CH-UD and non- DNMT3A driver CH, including a strong dependence on TCL1A, a protein recently found to inhibit DNA methylation. Collectively, these observations propose that CH-UD develops through selection acting on ageing-associated epigenetic changes that mirror those of non- DNMT3A -CH, but without the need for somatic genetic drivers.
To assess the contribution of rare coding germline genetic variants to prostate cancer risk and severity, we perform here a meta-analysis of 37,184 prostate cancer cases and 331,329 male controls from five cohorts with germline whole exome or genome sequencing data, and one cohort with imputed array data. At the gene level, our case-control collapsing analysis confirms associations between rare damaging variants in four genes and increased prostate cancer risk: SAMHD1, BRCA2 and ATM at the study-wide significance level (P < 1x10(-8)), and CHEK2 at the suggestive threshold (P < 2.6x10(-6)). Our case-only analysis, reveals that rare damaging variants in AOX1 are associated with more aggressive disease (OR = 2.60 [1.75-3.83], P = 1.35x10(-6)), as well as confirming the role of BRCA2 in determining disease severity. At the single-variant level, our study reveals that a rare missense variant in TERT is associated with substantially reduced prostate cancer risk (OR = 0.13 [0.07-0.25], P = 4.67x10(-10)), and confirms rare non-synonymous variants in a further three genes associated with reduced risk (ANO7, SPDL1, AR) and in three with increased risk (HOXB13, CHEK2, BIK). Altogether, this work provides deeper insights into the genetic architecture and biological basis of prostate cancer risk and severity, with potential implications for clinical risk prediction and therapeutic strategies.
The impact of genetic ancestry on the development of clonal hematopoiesis (CH) remains largely unexplored. Here, we compared CH in 136,401 participants from the Mexico City Prospective Study (MCPS) to 416,118 individuals from the UK Biobank (UKB) and observed CH to be significantly less common in MCPS compared to UKB (adjusted odds ratio = 0.59, 95% confidence interval (CI) = [0.57, 0.61], P = 7.31 × 10-185). Among MCPS participants, CH frequency was positively correlated with the percentage of European ancestry (adjusted beta = 0.84, 95% CI = [0.66, 1.03], P = 7.35 × 10-19). Genome-wide and exome-wide association analyses in MCPS identified ancestry-specific variants in the TCL1B locus with opposing effects on DNMT3A-CH versus non-DNMT3A-CH. Meta-analysis of MCPS and UKB identified five novel loci associated with CH, including polymorphisms at PARP11/CCND2, MEIS1 and MYCN. Our CH study, the largest in a non-European population to date, demonstrates the power of cross-ancestry comparisons to derive novel insights into CH pathogenesis.
Genomics can provide insight into the etiology of type 2 diabetes and its comorbidities, but assigning functionality to non-coding variants remains challenging. Polygenic scores, which aggregate variant effects, can uncover mechanisms when paired with molecular data. Here, we test polygenic scores for type 2 diabetes and cardiometabolic comorbidities for associations with 2,922 circulating proteins in the UK Biobank. The genome-wide type 2 diabetes polygenic score associates with 617 proteins, of which 75% also associate with another cardiometabolic score. Partitioned type 2 diabetes scores, which capture distinct disease biology, associate with 342 proteins (20% unique). In this work, we identify key pathways (e.g., complement cascade), potential therapeutic targets (e.g., FAM3D in type 2 diabetes), and biomarkers of diabetic comorbidities (e.g., EFEMP1 and IGFBP2) through causal inference, pathway enrichment, and Cox regression of clinical trial outcomes. Our results are available via an interactive portal ( https://public.cgr.astrazeneca.com/t2d-pgs/v1/ ).
Telomeres protect chromosome ends from damage and their length is linked with human disease and aging. We developed a joint telomere length metric, combining quantitative PCR and whole-genome sequencing measurements from 462,666 UK Biobank participants. This metric increased SNP heritability, suggesting that it better captures genetic regulation of telomere length. Exome-wide rare-variant and gene-level collapsing association studies identified 64 variants and 30 genes significantly associated with telomere length, including allelic series in ACD and RTEL1. Notably, 16% of these genes are known drivers of clonal hematopoiesis-an age-related somatic mosaicism associated with myeloid cancers and several nonmalignant diseases. Somatic variant analyses revealed gene-specific associations with telomere length, including lengthened telomeres in individuals with large SRSF2-mutant clones, compared with shortened telomeres in individuals with clonal expansions driven by other genes. Collectively, our findings demonstrate the impact of rare variants on telomere length, with larger effects observed among genes also associated with clonal hematopoiesis. Genome-wide association analysis of an improved telomere length score, calculated from quantitative PCR and whole-genome sequencing measurements in 462,666 individuals in the UK Biobank, identifies novel genes and variants underlying this trait.
The etiology of prostate cancer, the second most common cancer in men globally, has a strong heritable component. While rare coding germline variants in several genes have been identified as risk factors from candidate gene and linkage studies, the exome-wide spectrum of causal rare variants remains to be fully explored. To more comprehensively address their contribution, we analysed data from 37,184 prostate cancer cases and 331,329 male controls from five cohorts with germline exome/genome sequencing and one cohort with imputed array data from a population enriched in low-frequency deleterious variants. Our gene-level collapsing analysis revealed that rare damaging variants in SAMHD1 as well as genes in the DNA damage response pathway (BRCA2, ATM and CHEK2) are associated with the risk of overall prostate cancer. We also found that rare damaging variants in AOX1 and BRCA2 were associated with increased severity of prostate cancer in a case-only analysis of aggressive versus non-aggressive prostate cancer. At the single-variant level, we found rare non-synonymous variants in three genes (HOXB13, CHEK2, BIK) significantly associated with increased risk of overall prostate cancer and in four genes (ANO7, SPDL1, AR, TERT) with decreased risk. Altogether, this study provides deeper insights into the genetic architecture and biological basis of prostate cancer risk and severity.
Introduction Type 2 diabetes (T2D) is a heterogeneous disorder for which disease-causing pathways are incompletely understood. Here, we mapped genetic risk for T2D and its comorbidities to proteins, mechanistic pathways and clinical outcomes using proteogenomic data from a population-scale biobank and two randomized controlled trials. Methods We tested polygenic scores (PGS) for T2D and its cardiometabolic comorbidities, plus five partitioned T2D PGS (beta cell, lipodystrophy, liver lipid, obesity, and liver lipid), for association with 2,922 circulating proteins in 54,306 multi-ancestry participants (of which 42,452 were unrelated and without prevalent cardiometabolic disease) from the UK Biobank (UKB). Then, we tested the PGS-associated proteins for association with incident cardiometabolic complications in two cardiovascular outcome trials among T2D patients with proteogenomic data: EXSCEL (N=2,823) and DECLARE-TIMI 58 (N=915). We assessed causality using two-sample Mendelian randomization and mediation. Results We identified 839 unique proteins significantly associated with any T2D PGS and 1,005 proteins that were associated with at least one cardiometabolic PGS. Some PGS-associated proteins such as TFF3, EFEMP1, and MMP12 were in turn associated with renal and cardiovascular trial outcomes. PGS association patterns revealed shared pathways, e.g., complement cascade, cholesterol metabolism, IGF signaling. The proteins underlying these pathways, such as LPA, C1S, and IGFBP2, were consistently associated with clinical trial outcomes or identified via causal inference. Conclusions This proteogenomic study revealed proteins and mechanistic pathways underlying T2D and related comorbidities, advancing our understanding of T2D pathobiology and identifying putative biomarkers. All our results are available in an online data portal (<https://public.cgr.astrazeneca.com/t2d-pgs/v1/>). ### Competing Interest Statement D.P.L., M.G., D.M., D.V., X.J., I.A.G., S.P., J.O., A.N., and D.S.P. are employees of AstraZeneca and may hold AstraZeneca stock options. B.B.S. and H.R. are employees of Biogen and may hold stock options. C.D.W. is an employee of Janssen Pharmaceuticals, a Johnson & Johnson company, and may hold stock options. R.R.H. reports personal fees from Anji Pharmaceuticals, AstraZeneca and Novartis. R.J.M. received research support and honoraria from Abbott, American Regent, Amgen, AstraZeneca, Bayer, Boehringer Ingelheim, Boston Scientific, Cytokinetics, Fast BioMedical, Gilead, Innolife, Eli Lilly, Medtronic, Medable, Merck, Novartis, Novo Nordisk, Pfizer, Pharmacosmos, Relypsa, Respicardia, Roche, Rocket Pharmaceuticals, Sanofi, Verily, Vifor, Windtree Therapeutics, and Zoll. M.I. is a trustee of the Public Health Genomics (PHG) Foundation, a member of the Scientific Advisory Board of Open Targets and has research collaborations with Nightingale Health and Pfizer which are unrelated to this study. ### Funding Statement UK Biobank proteomics data was funded by a consortia of 13 participating pharmaceutical companies (Alnylam Pharmaceuticals, Amgen, AstraZeneca, Biogen, Bristol Myers Squibb, Calico, Genentech, GlaxoSmithKline, The Janssen Pharmaceutical Companies of Johnson & Johnson, Novo Nordisk, Pfizer, Regeneron and Takeda). The DECLARE-TIMI 58 and EXSCEL clinical trial were funded by AstraZeneca. ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: Information on EXSCEL and DECLARE-TMI 58 trials can be found on [clinicaltrials.gov][1] ([NCT01144338][2] for EXSCEL and [NCT01144338][2] for DECLARE). The trial protocols were approved by institutional review board at each participating site. UK Biobank has approval from the North West Multi-centre Research Ethics Committee (MREC) as a Research Tissue Bank (RTB). I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as [ClinicalTrials.gov][3]. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes All data produced in the present study are available upon reasonable request to the authors. Polygenic scores will be uploaded to the PGS Catalog. [1]: http://clinicaltrials.gov [2]: /lookup/external-ref?link_type=CLINTRIALGOV&access_num=NCT01144338&atom=%2Fmedrxiv%2Fearly%2F2024%2F03%2F19%2F2024.03.15.24304200.atom [3]: http://ClinicalTrials.gov
Introduction: Somatic mutations in blood stem cells can drive clonal expansion and result in clonal haematopoiesis (CH). CH is a precursor to myeloid malignancies, and is increasingly also recognised as a risk factor for non-malignant diseases. CH has been investigated in Europeans, but remains understudied in non-European populations. Here, we investigate the causes and consequences of CH in two large prospective cohort studies, specifically the Mexico City Prospective Study (MCPS) and UK Biobank (UKB). MCPS represents the largest CH study in a non-European population to date. Methods: Admixed Americans were identified from MCPS (n=136,401) while Europeans were identified from the UKB (n=419,228) participants. Somatic variant calling was performed using Mutect2 on whole exome sequencing (WES) data across a selected panel of genes to identify putative CH driver mutations. WES was performed at an average depth of 55x, and CH was defined using a variant allele frequency of ≥3%. Genome-wide association studies (GWAS) for CH were performed using imputed array data, and exome-wide association studies (ExWAS) and gene-level association analyses using WES data. Global/Continental-level ancestry was assigned based on peddy score ≥95% while local ancestry was inferred using RFMix (Ziyatdinov et al., 2022). Results: The most recurrently mutated genes in MCPS and UKB were DNMT3A, TET2, ASXL1, PPM1D, TP53, SF3B1 and SRSF2. The prevalence of CH increased progressively with age to approximately 8/22 (36%) detected as carriers at age 100 and above. CH was 40% more prevalent in age-matched Europeans (4.96%) compared to Admixed Americans (3.10%). Inter-population comparison revealed overall CH, DNMT3A-, TET2-, ASXL1-, PPM1D-, TP53-, JAK2-, and SRSF2-mutant CH to be more prevalent in UKB relative to MCPS (Figure A). Intra-population analysis of the Admixed American cohort further revealed that individuals with a higher fraction of European ancestry were at higher risk of overall CH, DNMT3A-, ASXL1-, and SRSF2-mutant CH, but not TET2-, PPM1D-, TP53-, and JAK2-mutant CH. These suggest differences in relative contribution of genetic, lifestyle or environment factors to specific CH genes. CH GWAS performed in Admixed Americans recapitulated previously reported variants in Europeans, and also identified novel, ancestry-specific variants associated with CH risk, including SNPs upstream of TCL1B (rs968294563: OR=1.79, P=2.01x10 -9; rs187319135: OR=1.85, P=2.69x10 -9). Notably, TCL1B variants were associated with an increased risk of TET2- and ASXL1-mutant CH (rs968294563: OR=3.17, P=3.82x10 -16 for TET2; OR=2.36, P=2.40x10 -3 for ASXL1), but a decreased risk of DNMT3A-mutantCH (rs968294563: OR=0.51, P=6.32x10 -4) (Figure B). The minor allele frequency (MAF) of the most common TCL1B variant in Admixed Americans was 1.23% but was virtually absent in Europeans. CH ExWAS can identify rare causal variants not captured by genotyping or imputation, and these variants may be in linkage disequilibrium with common variants detectable via GWAS. Indeed, ExWAS in Admixed Americans identified one rare SNP on the TCL1B promoter (rs774615666: OR=2.24, P=1.94x10 -8) that was associated with an increased risk of TET2-mutantCH (OR=4.19, P=2.77x10 -10) and a decreased risk of DNMT3A-mutantCH (OR=0.13, P=1.65x10 -4), mirroring our GWAS findings. The rs774615666 risk allele was >200-fold more common in Admixed Americans (MAF: 0.33%) compared to Europeans (MAF: 0.0011%) and was not in linkage disequilibrium with the previously reported TCL1A promoter CH risk SNP in Europeans (Weinstock et al., 2023). Meta-analyses were performed using 555,629 individuals, from both MCPS Admixed American CH GWAS and UKB European CH GWAS. Of the 11 loci reported for overall CH, one was novel, namely GACAT3/ CYRIA. Lastly, we investigated the phenotypic associations of CH in Admixed Americans. Gene-specific CH was associated with increased risk of death from haematological malignancies ( TET2, TP53, SF3B1, and SRSF2), other cancer types ( ASXL1, DNMT3A, and SRSF2) and cardiovascular diseases( DNMT3A, TP53). Conclusions: The substantial difference in CH prevalence between populations and the ancestry-specific genetic associations, demonstrate how the analysis of non-European cohorts can generate novel insights and highlights the importance of such analyses in advancing health equality amongst different human populations.
The Mexico City Prospective Study is a prospective cohort of more than 150,000 adults recruited two decades ago from the urban districts of Coyoacán and Iztapalapa in Mexico City 1 . Here we generated genotype and exome-sequencing data for all individuals and whole-genome sequencing data for 9,950 selected individuals. We describe high levels of relatedness and substantial heterogeneity in ancestry composition across individuals. Most sequenced individuals had admixed Indigenous American, European and African ancestry, with extensive admixture from Indigenous populations in central, southern and southeastern Mexico. Indigenous Mexican segments of the genome had lower levels of coding variation but an excess of homozygous loss-of-function variants compared with segments of African and European origin. We estimated ancestry-specific allele frequencies at 142 million genomic variants, with an effective sample size of 91,856 for Indigenous Mexican ancestry at exome variants, all available through a public browser. Using whole-genome sequencing, we developed an imputation reference panel that outperforms existing panels at common variants in individuals with high proportions of central, southern and southeastern Indigenous Mexican ancestry. Our work illustrates the value of genetic studies in diverse populations and provides foundational imputation and allele frequency resources for future genetic studies in Mexico and in the United States, where the Hispanic/Latino population is predominantly of Mexican descent.
Genome-wide association studies (GWASs) have established the contribution of common and low-frequency variants to metabolic blood measurements in the UK Biobank (UKB). To complement existing GWAS findings, we assessed the contribution of rare pro-tein-coding variants in relation to 355 metabolic blood measurements-including 325 predominantly lipid-related nuclear magnetic resonance (NMR)-derived blood metabolite measurements (Nightingale Health Plc) and 30 clinical blood biomarkers-using 412,393 exome sequences from four genetically diverse ancestries in the UKB. Gene-level collapsing analyses were conducted to evaluate a diverse range of rare-variant architectures for the metabolic blood measurements. Altogether, we identified significant associations (p < 1 3 10-8) for 205 distinct genes that involved 1,968 significant relationships for the Nightingale blood metabolite measure-ments and 331 for the clinical blood biomarkers. These include associations for rare non-synonymous variants in PLIN1 and CREB3L3 with lipid metabolite measurements and SYT7 with creatinine, among others, which may not only provide insights into novel biology but also deepen our understanding of established disease mechanisms. Of the study-wide significant clinical biomarker associations, 40% were not previously detected on analyzing coding variants in a GWAS in the same cohort, reinforcing the importance of studying rare variation to fully understand the genetic architecture of metabolic blood measurements.
Telomeres protect the ends of chromosomes from damage, and genetic regulation of their length is associated with human disease and ageing. We developed a joint telomere length (TL) metric, combining both qPCR and whole genome sequencing (WGS) measurements across 462,675 UK Biobank participants that increased our ability to capture TL heritability by 36% (h 2 mean= 0.058 to h 2 combined= 0.079) and improved predictions of age. Exome-wide rare variant (minor allele frequency<0.001) and gene-level collapsing association studies identified 53 variants and 22 genes significantly associated with TL that included allelic series in ACD and RTEL1 . Five of the 31 rare-variant TL associated genes (16%) were also known drivers of clonal haematopoiesis (CH), prompting somatic variant analyses. Stratifying by CH clone size, we uncovered novel gene-specific associations with TL, including lengthened telomeres in individuals with large SRSF2 -mutant clones, in contrast to the progressive telomere shortening observed with increasing clonal expansions driven by other CH genes. Our findings demonstrate the impact of rare variants on TL with larger effects in genes associated with CH, a precursor of myeloid cancers and several other non-malignant human diseases. Telomere biology is likely to be an important focus for the prevention and treatment of these conditions.
The Mexico City Prospective Study (MCPS) is a prospective cohort of over 150,000 adults recruited two decades ago from the urban districts of Coyoacán and Iztapalapa in Mexico City. We generated genotype and exome sequencing data for all individuals, and whole genome sequencing for 10,000 selected individuals. We uncovered high levels of relatedness and substantial heterogeneity in ancestry composition across individuals. Most sequenced individuals had admixed Native American, European and African ancestry, with extensive admixture from indigenous groups in Central, Southern and South Eastern Mexico. Native Mexican segments of the genome had lower levels of coding variation, but an excess of homozygous loss of function variants compared with segments of African and European origin. We estimated population specific allele frequencies at 142 million genomic variants, with an effective sample size of 91,856 for Native Mexico at exome variants, all available via a public browser. Using whole genome sequencing, we developed an imputation reference panel which outperforms existing panels at common variants in individuals with high proportions of Central, South and South Eastern Native Mexican ancestry. Our work illustrates the value of genetic studies in populations with diverse ancestry and provides foundational imputation and allele frequency resources for future genetic studies in Mexico and in the United States where the Hispanic/Latino population is predominantly of Mexican descent.### Competing Interest StatementAndrey Ziyatdinov, Joshua Backman, Joelle Mbatchou, Sheila M. Gaynor, Tyler Joseph, Yuxin Zou, Daren Liu, Jeffrey Staples, Razvan Panea, Xiaodong Bai, Suganthi Balasubramanian, Lukas Habegger, Rouel Lanche, Alex Lopez, Evan Maxwell, Marcus Jones, Eric Jorgenson, William Salerno, John Overton, Jeffrey Reid, Timothy Thornton, Goncalo Abecasis, Aris Baras, Jonathan Marchini are current employees and/or stockholders of Regeneron Genetics Center or Regeneron Pharmaceuticals. Abhishek Nag, Katherine Smith and Slave Petrovski are current employees and/or stockholders of AstraZeneca Mark Reppell is current employee and/or stockholder of AbbVie.
Large-scale phenome-wide association studies performed using densely-phenotyped cohorts such as the UK Biobank (UKB), reveal many statistically robust gene-phenotype relationships for both clinical and continuous traits. Here, we present Gene-SCOUT, a tool used to identify genes with similar continuous trait fingerprints to a gene of interest. A fingerprint reflects the continuous traits identified to be statistically associated with a gene of interest based on multiple underlying rare variant genetic architectures. Similarities between genes are evaluated by the cosine similarity measure, to capture concordant effect directionality, elucidating clusters of genes in a high dimensional space. The underlying gene-biomarker population-scale association statistics were obtained from a gene-level rare variant collapsing analysis performed on over 1500 continuous traits using 394 692 UKB participant exomes, with additional metabolomic trait associations provided through Nightingale Health's recent study of 121 394 of these participants. We demonstrate that gene similarity estimates from Gene-SCOUT provide stronger enrichments for clinical traits compared to existing methods. Furthermore, we provide a fully interactive web-resource (http: //genescout.public.cgr.astrazeneca.com) to explore the pre-calculated exome-wide similarities. This resource enables a user to examine the biological relevance of the most similar genes for Gene Ontology (GO) enrichment and UKB clinical trait enrichment statistics, as well as a detailed breakdown of the traits underpinning a given fingerprint.