Genetic support for drug targets substantially increases clinical success rates, establishing genomewide association studies (GWAS) as central to therapeutic hypothesis generation. However, the same genetic evidence that reveals causal gene-disease relationships simultaneously exposes organismlevel safety liabilities-a dimension requiring principled, genome-wide quantification. Here we systematically analyse 100,526 GWAS to yield 789,453 credible sets and gene prioritisations for 15,641 genes, with discovery showing no saturation as GWAS expand and increase diversity. We find that 64% of GWAS-implicated genes are pleiotropic, associated with traits across multiple diseases and showing a non-linear relationship between the degree of pleiotropy and clinical success. Highly pleiotropic genes-concentrated in immune, inflammatory, and oncogenic signalling programmes-are enriched in safety-terminated clinical programmes, mouse lethal knockouts, and cancer driver genes, establishing gene-level pleiotropy as a potential measure of genetically-informed organism-level safety liability. Protein-altering variant (PAV) support amplifies therapeutic signal (OR = 6.0), yet PAV targets show higher average pleiotropy, introducing a competing safety liability. Combining PAV support with intermediate pleiotropy (2-5 therapeutic areas) resolves this tension, yielding OR = 10.3 and relative success = 4.8-a profile already satisfied by 52 approved therapies. As GWAS continue to expand in scale and resolution, these findings lay the groundwork for increasingly sophisticated target discovery strategies that yield safer and more effective therapeutic hypotheses.
SUMMARY:The lit-OTAR framework, developed through a collaboration between Europe PMC and Open Targets, leverages deep learning to revolutionize drug discovery by extracting evidence from scientific literature for drug target identification and validation. This novel framework combines named entity recognition for identifying gene/protein (target), disease, organism, and chemical/drug within scientific texts, and entity normalization to map these entities to databases like Ensembl, Experimental Factor Ontology, and ChEMBL. Continuously operational, it has processed over 39 million abstracts and 4.5 million full-text articles and preprints to date, identifying more than 48.5 million unique associations that significantly help accelerate the drug discovery process and scientific research >29.9 m distinct target-disease, 11.8 m distinct target-drug, and 8.3 m distinct disease-drug relationships. AVAILABILITY AND IMPLEMENTATION:The results are accessible through Europe PMC's SciLite web app (https://europepmc.org/) and its annotations API (https://europepmc.org/annotationsapi), as well as via the Open Targets Platform (https://platform.opentargets.org/). The daily pipeline is available at https://github.com/ML4LitS/otar-maintenance, and the Open Targets ETL processes are available at https://github.com/opentargets.
The Open Targets Platform (https://platform.opentargets.org) is a unique, open-source, publicly-available knowledge base providing data and tooling for systematic drug target identification, annotation, and prioritisation. Since our last report, we have expanded the scope of the Platform through a number of significant enhancements and data updates, with the aim to enable our users to formulate more flexible and impactful therapeutic hypotheses. In this context, we have completely revamped our target-disease associations page with more interactive facets and built-in functionalities to empower users with additional control over their experience using the Platform, and added a new Target Prioritisation view. This enables users to prioritise targets based upon clinical precedence, tractability, doability and safety attributes. We have also implemented a direction of effect assessment for eight sources of target-disease association evidence, showing the effect of genetic variation on the function of a target is associated with risk or protection for a trait to inform on potential mechanisms of modulation suitable for disease treatment. These enhancements and the introduction of new back and front-end technologies to support them have increased the impact and usability of our resource within the drug discovery community.
Interacting proteins tend to have similar functions, influencing the same organismal traits. Interaction networks can be used to expand the list of candidate trait-associated genes from genome-wide association studies. Here, we performed network-based expansion of trait-associated genes for 1,002 human traits showing that this recovers known disease genes or drug targets. The similarity of network expansion scores identifies groups of traits likely to share an underlying genetic and biological process. We identified 73 pleiotropic gene modules linked to multiple traits, enriched in genes involved in processes such as protein ubiquitination and RNA processing. In contrast to gene deletion studies, pleiotropy as defined here captures specifically multicellular-related processes. We show examples of modules linked to human diseases enriched in genes with known pathogenic variants that can be used to map targets of approved drugs for repurposing. Finally, we illustrate the use of network expansion scores to study genes at inflammatory bowel disease genome-wide association study loci, and implicate inflammatory bowel disease-relevant genes with strong functional and genetic support.
The Open Targets Platform (https://platform.opentargets.org/) is an open source resource to systematically assist drug target identification and prioritisation using publicly available data. Since our last update, we have reimagined, redesigned, and rebuilt the Platform in order to streamline data integration and harmonisation, expand the ways in which users can explore the data, and improve the user experience. The gene-disease causal evidence has been enhanced and expanded to better capture disease causality across rare, common, and somatic diseases. For target and drug annotations, we have incorporated new features that help assess target safety and tractability, including genetic constraint, PROTACtability assessments, and AlphaFold structure predictions. We have also introduced new machine learning applications for knowledge extraction from the published literature, clinical trial information, and drug labels. The new technologies and frameworks introduced since the last update will ease the introduction of new features and the creation of separate instances of the Platform adapted to user requirements. Our new Community forum, expanded training materials, and outreach programme support our users in a range of use cases.
Haematological traits are linked to cardiovascular, metabolic, infectious and immune disorders, as well as cancer. Here, we examine the role of genetic variation in shaping haematological traits in two isolated Mediterranean populations. Using whole-genome sequencing data at 22× depth for 1457 individuals from Crete (MANOLIS) and 1617 from the Pomak villages in Greece, we carry out a genome-wide association scan for haematological traits using linear mixed models. We discover novel associations ( p < 5 × 10 –9 ) of five rare non-coding variants with alleles conferring effects of 1.44–2.63 units of standard deviation on red and white blood cell count, platelet and red cell distribution width. Moreover, 10.0% of individuals in the Pomak population and 6.8% in MANOLIS carry a pathogenic mutation in the Haemoglobin Subunit Beta (HBB) gene. The mutational spectrum is highly diverse (10 different mutations). The most frequent mutation in MANOLIS is the common Mediterranean variant IVS-I-110 (G>A) (rs35004220). In the Pomak population, c.364C>A (“HbO-Arab”, rs33946267) is most frequent (4.4% allele frequency). We demonstrate effects on haematological and other traits, including bilirubin, cholesterol, and, in MANOLIS, height and gestation age. We find less severe effects on red blood cell traits for HbS, HbO, and IVS-I-6 (T>C) compared to other b+ mutations. Overall, we uncover allelic diversity of HBB in Greek isolated populations and find an important role for additional rare variants outside of HBB .
Open Targets Genetics (https://genetics.opentargets.org) is an open-access integrative resource that aggregates human GWAS and functional genomics data including gene expression, protein abundance, chromatin interaction and conformation data from a wide range of cell types and tissues to make robust connections between GWAS-associated loci, variants and likely causal genes. This enables systematic identification and prioritisation of likely causal variants and genes across all published trait-associated loci. In this paper, we describe the public resources we aggregate, the technology and analyses we use, and the functionality that the portal offers. Open Targets Genetics can be searched by variant, gene or study/phenotype. It offers tools that enable users to prioritise causal variants and genes at disease-associated loci and access systematic cross-disease and disease-molecular trait colocalization analysis across 92 cell types and tissues including the eQTL Catalogue. Data visualizations such as Manhattan-like plots, regional plots, credible sets overlap between studies and PheWAS plots enable users to explore GWAS signals in depth. The integrated data is made available through the web portal, for bulk download and via a GraphQL API, and the software is open source. Applications of this integrated data include identification of novel targets for drug discovery and drug repurposing.
Podoconiosis, a debilitating lymphoedema of the leg, results from barefoot exposure to volcanic clay soil in genetically susceptible individuals. A previous genome-wide association study (GWAS) conducted in the Wolaita ethnic group from Ethiopia showed association between single nucleotide polymorphisms (SNPs) in the HLA class II region and podoconiosis. We aimed to conduct a second GWAS in a new sample (N = 1892) collected from the Wolaita and two other Ethiopian populations, the Amhara and the Oromo, also affected by podoconiosis. Fourteen SNPs in the HLA class II region showed significant genome-wide association ( P < 5.0 × 10 −8 ) with podoconiosis. The lead SNP was rs9270911 ( P = 5.51 × 10 −10 ; OR 1.53; 95% CI 1.34–1.74), located near HLA-DRB1 . Inclusion of data from the first GWAS (combined N = 2289) identified 47 SNPs in the class II HLA region that were significantly associated with podoconiosis (lead SNP also rs9270911 ( P = 2.25 × 10 −12 ). No new loci outside of the HLA class II region were identified in this more highly-powered second GWAS. Our findings confirm the HLA class II association with podoconiosis suggesting HLA-mediated abnormal induction and regulation of immune responses may have a direct role in its pathogenesis.
The human proteome is a crucial intermediate between complex diseases and their genetic and environmental components, and an important source of drug development targets and biomarkers. Here, we comprehensively assess the genetic architecture of 257 circulating protein biomarkers of cardiometabolic relevance through high-depth (22.5x) whole-genome sequencing (WGS) in 1,328 individuals. We discover 131 independent sequence variant associations ( P <7.45×10−11) across the allele frequency spectrum, all of which replicate in an independent cohort (n=1,605, 18.4x WGS). We identify for the first time replicating evidence for rare-variant cis -acting protein quantitative trait loci for five genes, involving both coding and non-coding variation. We construct and validate polygenic scores that explain up to 45% of protein level variation. We find causal links between protein levels and disease risk, identifying high-value biomarkers and drug development targets.
Abstract The Open Targets Platform (https://www.targetvalidation.org/) provides users with a queryable knowledgebase and user interface to aid systematic target identification and prioritisation for drug discovery based upon underlying evidence. It is publicly available and the underlying code is open source. Since our last update two years ago, we have had 10 releases to maintain and continuously improve evidence for target–disease relationships from 20 different data sources. In addition, we have integrated new evidence from key datasets, including prioritised targets identified from genome-wide CRISPR knockout screens in 300 cancer models (Project Score), and GWAS/UK BioBank statistical genetic analysis evidence from the Open Targets Genetics Portal. We have evolved our evidence scoring framework to improve target identification. To aid the prioritisation of targets and inform on the potential impact of modulating a given target, we have added evaluation of post-marketing adverse drug reactions and new curated information on target tractability and safety. We have also developed the user interface and backend technologies to improve performance and usability. In this article, we describe the latest enhancements to the Platform, to address the fundamental challenge that developing effective and safe drugs is difficult and expensive.
The NHGRI-EBI GWAS Catalog is a comprehensive and widely-used database of published genome-wide association studies (GWAS), providing links between genetic loci and complex traits. As of December 2018, the Catalog contains over 3,600 publications and 81,000 unique SNP-trait associations. The increasing complexity of GWAS publications in recent years presents challenges for the accurate and user-friendly display of Catalog data. To meet this challenge, we engaged with the GWAS community to develop a new user interface, released in September 2018. We analysed the most common search terms and conducted user interviews to define and prioritise improvements. We also published a “labs” site for users to test proposed features and conducted face-to-face user testing throughout the development process. Following release, we sought feedback and continued to make incremental changes in response to users’ needs. The new interface enables more specific and intuitive access to Catalog data. Individual pages for each publication, study, trait, variant and gene, allow us to make available information tailored to each of these key concepts. Each page includes supporting information to provide context, a table of associations and links to external resources. New visualisations support further analyses of Catalog data. A plot of the linkage disequilibrium landscape allows users to investigate a region of association tagged by a GWAS variant and prioritise causal variants. Users can also visualise associations across the genome by trait in a LocusZoom plot. The GWAS Catalog now hosts full p-value summary statistics in addition to curated associations. Access to summary statistics is now integrated within the user interface, providing users with direct links to available files, in addition to API access. These changes respond to developments in the GWAS field as well as user feedback and will ensure that the GWAS Catalog continues to be a valuable resource for the scientific community.
Motivation Very low depth sequencing has been proposed as a cost-effective approach to capture low-frequency and rare variation in complex trait association studies. However, a full characterisation of the genotype quality and association power for very low depth sequencing designs is still lacking. Results We perform cohort-wide whole genome sequencing (WGS) at low depth in 1,239 individuals (990 at 1x depth and 249 at 4x depth) from an isolated population, and establish a robust pipeline for calling and imputing very low depth WGS genotypes from standard bioinformatics tools. Using genotyping chip, whole-exome sequencing (WES, 75x depth) and high-depth (22x) WGS data in the same samples, we examine in detail the sensitivity of this approach, and show that imputed 1x WGS recapitulates 95.2% of variants found by imputed GWAS with an average minor allele concordance of 97% for common and low-frequency variants. In our study, 1x further allowed the discovery of 140,844 true low-frequency variants with 73% genotype concordance when compared to high-depth WGS data. Finally, using association results for 57 quantitative traits, we show that very low depth WGS is an efficient alternative to imputed GWAS chip designs, allowing the discovery of up to twice as many true association signals than the classical imputed GWAS design. Supplementary Data Supplementary Data are appended to this manuscript.
MotivationCopy number variants (CNVs) are large deletions or duplications at least 50 to 200 base pairs long. They play an important role in multiple disorders, but accurate calling of CNVs remains challenging. Most current approaches to CNV detection use raw read alignments, which are computationally intensive to process.ResultsWe use a regression tree-based approach to call CNVs from whole-genome sequencing (WGS, > 18x) variant call-sets in 6,898 samples across four European cohorts, and describe a rich large variation landscape comprising 1,320 CNVs. 61.8% of detected events have been previously reported in the Database of Genomic Variants. 23% of high-quality deletions affect entire genes, and we recapitulate known events such as theGSTM1andRHDgene deletions. We test for association between the detected deletions and 275 protein levels in 1,457 individuals to assess the potential clinical impact of the detected CNVs. We describe the LD structure and copy number variation underlying the association between levels of the CCL3 protein and a complex structural variant (MAF = 0.15, p = 3.6×10-12) affectingCCL3L3, a paralog of theCCL3gene. We also identify acis-association between a low-frequencyNOMO1deletion and the protein product of this gene (MAF = 0.02, p = 2.2×10-7), for which nocis-ortrans-single nucleotide variant-driven protein quantitative trait locus (pQTL) has been documented to date. This work demonstrates that existing population-wide WGS call-sets can be mined for CNVs with minimal computational overhead, delivering insight into a less well-studied, yet potentially impactful class of genetic variant.AvailabilityThe regression tree based approach, UN-CNVc, is available as an R and bash executable on GitHub athttps://github.com/agilly/un-cnvc.Contacteleftheria.zeggini@helmholtz-muenchen.de;arthur.gilly@helmholtz-muenchen.deSupplementary InformationSupplementary information is appended.
Osteoarthritis is the most common musculoskeletal disease and the leading cause of disability globally. Here, we performed a genome-wide association study for osteoarthritis (77,052 cases and 378,169 controls), analyzing four phenotypes: knee osteoarthritis, hip osteoarthritis, knee and/or hip osteoarthritis, and any osteoarthritis. We discovered 64 signals, 52 of them novel, more than doubling the number of established disease loci. Six signals fine-mapped to a single variant. We identified putative effector genes by integrating expression quantitative trait loci (eQTL) colocalization, fine-mapping, and human rare-disease, animal-model, and osteoarthritis tissue expression data. We found enrichment for genes underlying monogenic forms of bone development diseases, and for the collagen formation and extracellular matrix organization biological pathways. Ten of the likely effector genes, including TGFB1 (transforming growth factor beta 1), FGF18 (fibroblast growth factor 18), CTSK (cathepsin K), and IL11 (interleukin 11), have therapeutics approved or in clinical trials, with mechanisms of action supportive of evaluation for efficacy in osteoarthritis.