Knowledge graphs represent information in the form of entities and relationships between those entities. Such a representation has multiple potential applications in drug discovery, including democratizing access to biomedical data, contextualizing or visualizing that data, and generating novel insights through the application of machine learning approaches. Knowledge graphs put data into context and therefore offer the opportunity to generate explainable predictions, which is a key topic in contemporary artificial intelligence. In this chapter, we outline some of the factors that need to be considered when constructing biomedical knowledge graphs, examine recent advances in mining such systems to gain insights for drug discovery, and identify potential future areas for further development.
Building and analyzing knowledge graphs (KGs) to aid drug discovery is a topical area of research. A salient feature of KGs is their ability to combine many heterogeneous data sources in a format that facilitates discovering connections. The utility of KGs has been exemplified in areas such as drug repurposing, with insights made through manual exploration and modeling of the data. In this chapter, we discuss promises and pitfalls of using natural language processing (NLP) to mine "unstructured text"- typically from scientific literature- as a data source for KGs. This draws on our experience of initially parsing "structured" data sources-such as ChEMBL-as the basis for data within a KG, and then enriching or expanding upon them using NLP. The fundamental promise of NLP for KGs is the automated extraction of data from millions of documents-a task practically impossible to do via human curation alone. However, there are many potential pitfalls in NLP-KG pipelines, such as incorrect named entity recognition and ontology linking, all of which could ultimately lead to erroneous inferences and conclusions.
One area of active research is the use of natural language processing (NLP) to mine biomedical texts for sets of triples (subject-predicate-object) for knowledge graph (KG) construction. While statistical methods to mine co-occurrences of entities within sentences are relatively robust, accurate relationship extraction is more challenging. Herein, we evaluate the Global Network of Biomedical Relationships (GNBR), a dataset that uses distributional semantics to model relationships between biomedical entities. The focus of our paper is an evaluation of a subset of the GNBR data; the relationships between chemicals and genes/proteins. We use Evotec's structured 'Nexus' database of >2.76M chemical-protein interactions as a ground truth to compare with GNBRs relationships and find a micro-averaged precision-recall area under the curve (AUC) of 0.50 and a micro-averaged receiver operating characteristic (ROC) curve AUC of 0.71 across the relationship classes 'inhibits', 'binding', 'agonism' and 'antagonism', when a comparison is made on a sentence-by-sentence basis. We conclude that, even though these micro-average scores are modest, using a high threshold on certain relationship classes like 'inhibits' could yield high fidelity triples that are not reported in structured datasets. We discuss how different methods of processing GNBR data, and the factuality of triples could affect the accuracy of NLP data incorporated into knowledge graphs. We provide a GNBR-Nexus(ChEMBL-subset) merged datafile that contains over 20,000 sentences where a protein/gene-chemical co-occur and includes both the GNBR relationship scores as well as the ChEMBL (manually curated) relationships (e.g., 'agonist', 'inhibitor') -this can be accessed at https://doi.org/10.5281/zenodo.8136752. We envisage this being used to aid curation efforts by the drug discovery community.
Within the context of the latest resurgence in the application of artificial intelligence approaches, deep learning has undergone a renaissance over recent years. These methods have been applied to a number of problems in computational chemistry. Compared to other machine learning approaches, the practical performance advantages of deep neural networks are often unclear. However, deep learning does appear to offer a number of other advantages such as the facile incorporation of multitask learning and the enhancement of generative modeling. The high complexity of contemporary network architectures represents a potentially significant barrier to their future adoption due to the costs of training such models and challenges in interpreting their predictions. When combined with the relative paucity of very large datasets, it is interesting to reflect on whether deep learning is likely to have the kind of transformational impact on computational chemistry that it is commonly held to have had in other domains such as image recognition.
Artificial Intelligence (AI) and Machine Learning (ML) are becoming some of the most dominant tools in scientific research. Despite this, little is often understood about the complex decisions taken by the models in predicting their results. This disproportionately affects biomedical and healthcare research where explainability of AI is one of the requirements for its wide adoption. To help answer the question of what the network is looking at when the labels do not correspond to the presence of objects in the image but the context in which they are found, we propose a novel framework for Explainable AI that combines and simultaneously analyses Class Activation and Segmentation Maps for thousands of images. We apply our approach to two distinct, complex examples of real-world biomedical research, and demonstrate how it can be used to provide a global and concise numerical measurement of how distinct classes of objects affect the final classification. We also show how this can be used to inform model selection, architecture design and aid traditional domain researchers in interpreting the model results.
Estimating the range of three-dimensional structures (conformations) that are available to a molecule is a key component of computer-aided drug design. Quantum mechanical simulation offers improved accuracy over forcefield methods, but at a high computational cost. The question is whether this increased cost can be justified in a context in which high-throughput analysis of large numbers of molecules is often key. This chapter discusses the application of quantum mechanics to conformational searching, with a focus on three key challenges: (1) the generation of ensembles that include a good approximation to a molecule's bioactive conformation at as prominent a ranking as possible; (2) rational analysis and modification of a pre-established bioactive conformation in terms of its energetics; and (3) approximation of real solution-phase conformational ensembles in tandem with NMR data. The impact of QM on the high-throughput application (1) is debatable, meaning that for the moment its primary application is still lower-throughput applications such as (2) and (3). The optimal choice of QM method is also discussed. Rigorous benchmarking suggests that DFT methods are only acceptable when used with large basis sets, but a trickle of papers continue to obtain useful results with relatively low-cost methods, leading to a dilemma that the literature has yet to fully resolve.
G-protein coupled receptors (GPCRs) are the largest superfamily of membrane proteins, regulating almost every aspect of cellular activity and serving as key targets for drug discovery. We have identified an accurate and reliable computational method to characterize the strength and chemical nature of the interhelical interactions between the residues of transmembrane (TM) domains during different receptor activation states, something that cannot be characterized solely by visual inspection of structural information. Using the fragment molecular orbital (FMO) quantum mechanics method to analyze 35 crystal structures representing different branches of the class A GPCR family, we have identified 69 topologically equivalent TM residues that form a consensus network of 51 inter-TM interactions, providing novel results that are consistent with and help to rationalize experimental data. This discovery establishes a comprehensive picture of how defined molecular forces govern specific interhelical interactions which, in turn, support the structural stability, ligand binding, and activation of GPCRs.
G-protein-coupled receptors (GPCRs) have enormous physiological and biomedical importance, and therefore it is not surprising that they are the targets of many prescribed drugs. Further progress in GPCR drug discovery is highly dependent on the availability of protein structural information. However, the ability of X-ray crystallography to guide the drug discovery process for GPCR targets is limited by the availability of accurate tools to explore receptor-ligand interactions. Visual inspection and molecular mechanics approaches cannot explain the full complexity of molecular interactions. Quantum mechanics (QM) approaches are often too computationally expensive to be of practical use in time-sensitive situations, but the fragment molecular orbital (FMO) method offers an excellent solution that combines accuracy, speed, and the ability to reveal key interactions that would otherwise be hard to detect. Integration of GPCR crystallography or homology modelling with FMO reveals atomistic details of the individual contributions of each residue and water molecule toward ligand binding, including an analysis of their chemical nature. Such information is essential for an efficient structure-based drug design (SBDD) process. In this chapter, we describe how to use FMO in the characterization of GPCR-ligand interactions.
The understanding of binding interactions between a protein and a small molecule plays a key role in the rationalization of potency and selectivity and in design of new ideas. However, even when a target of interest is structurally enabled, visual inspection and force field-based molecular mechanics calculations cannot always explain the full complexity of the molecular interactions that are critical in drug design. Quantum mechanical methods have the potential to address this shortcoming, but traditionally, computational expense has made the application of these calculations impractical. The fragment molecular orbital (FMO) method offers a solution that combines accuracy, speed, and the ability to characterize important interactions (i.e. its strength in kcal/mol and chemical nature: hydrophobic, electrostatic, etc) that would otherwise be hard to detect. In this chapter, we describe the FMO method and illustrate its application in the discovery of the benzothiazole (BZT) series as novel tyrosine kinase ITK inhibitors for treatment of allergic asthma.
One of the grand challenges in contemporary chemical biology is the generation of a probe for every member of the human proteome. Probe selection and optimization strategies typically rely on experimental bioactivity data to determine the potency and selectivity of candidate molecules. However, this approach is profoundly limited by the sparsity of the known data, the annotation bias often found in the literature, and the cost of physical screening. Recent advancements in predictive pharmacology, such as the application of multitask and transfer learning, as well as the use of biologically motivated, structure-agnostic features to characterize molecules, should serve to mitigate these issues. Computational modeling likely offers the only cost-effective approach to substantially increasing the bioactivity annotation density both on the local and global scale and thus, we argue, will need to make a substantial contribution if the ambitious goals of probing the human proteome are to be realized in the foreseeable future.
There has been fantastic progress in solving GPCR crystal structures. However, the ability of X-ray crystallography to guide the drug discovery process for GPCR targets is limited by the availability of accurate tools to explore receptor-ligand interactions. Visual inspection and molecular mechanics approaches cannot explain the full complexity of molecular interactions. Quantum mechanical approaches (QM) are often too computationally expensive, but the fragment molecular orbital (FMO) method offers an excellent solution that combines accuracy, speed and the ability to reveal key interactions that would otherwise be hard to detect. Integration of GPCR crystallography or homology modelling with FMO reveals atomistic details of the individual contributions of each residue and water molecule towards ligand binding, including an analysis of their chemical nature.
The understanding of binding interactions between any protein and a small molecule plays a key role in the rationalization of affinity and selectivity. It is essential for an efficient structure-based drug design (SBDD) process. FMO enables ab initio approaches to be applied to systems that conventional quantum-mechanical (QM) methods would find challenging. The key advantage of the Fragment Molecular Orbital Method (FMO) is that it can reveal atomistic details about the individual contributions and chemical nature of each residue and water molecule toward ligand binding which would otherwise be difficult to detect without using QM methods. In this chapter, we demonstrate the typical use of FMO to analyze 19 crystal structures of β1 and β2 adrenergic receptors with their corresponding agonists and antagonists.
Cheminformatics is a broad discipline covering a wide range of computational approaches, including the characterization of molecular similarity, pattern recognition, and predictive modeling. The unifying theme that these apparently disparate methods have in common is the aim of extracting useable information from the increasing amounts of data that are associated with contemporary drug discovery projects. Both proprietary and publically available data can be exploited to help inform and improve the process of developing novel therapeutic molecules targeting the GPCR family of proteins.
G-protein coupled receptor (GPCR) modeling approaches are widely used in the hit-to-lead and lead optimization stages of drug discovery. Modern protocols that involve molecular dynamics simulation can address key issues such as the free energy of binding (affinity), ligand-induced GPCR flexibility, ligand binding kinetics, conserved water positions and their role in ligand binding and the effects of mutations. The goals of these calculations are to predict the structures of the complexes between existing ligands and their receptors, to understand the key interactions and to utilize these insights in the design of new molecules with improved binding, selectivity or other pharmacological properties. In this review we present a brief survey of various computational approaches illustrated through a hierarchical GPCR modeling protocol and its prospective application in three industrial drug discovery projects.
Agonism of the 5-HT2C serotonin receptor has been associated with the treatment of a number of diseases including obesity, psychiatric disorders, sexual health, and urology. However, the development of effective 5-HT2C agonists has been hampered by the difficulty in obtaining selectivity over the closely related 5-HT2B receptor, agonism of which is associated with irreversible cardiac valvulopathy. Understanding how to design selective agonists requires exploration of the structural features governing the functional uniqueness of the target receptor relative to related off targets. X-ray crystallography, the major experimental source of structural information, is a slow and challenging process for integral membrane proteins, and so is currently not feasible for every GPCR or GPCR-ligand complex. Therefore, the integration of existing ligand SAR data with GPCR modeling can be a practical alternative to provide this essential structural insight. To demonstrate this, we integrated SAR data from 39 azepine series 5-HT2C agonists, comprising both selective and unselective examples, with our hierarchical GPCR modeling protocol (HGMP). Through this work we have been able to demonstrate how relatively small differences in the amino acid sequences of GPCRs can lead to significant differences in secondary structure and function, as supported by experimental data. In particular, this study suggests that conformational differences in the tilt of TM7 between 5-HT2B and 5-HT2C, which result from differences in interhelical interactions, may be the major source of selectivity in G-protein activation between these two receptors. Our approach also demonstrates how the use of GPCR models in conjunction with SAR data can be used to explain activity cliffs.
The Jumonji (Jmj) family of demethylases has a crucial role in regulating epigenetic processes through the removal of methyl groups from histone tails. The ability of Jmj demethylases to recognise their targets selectively has been elegantly addressed by structural studies. Reviewing recent structural literature, we provide an overview of selectivity mechanisms that demethylases use, including specific residues, methylation states and contextual requirements. We also report the presence of a common JmjN support domain across the family. The ability to use structural information for this enzyme class will be a crucial component of future drug discovery.
We describe the first targeted validation of fFLASH, a molecular similarity program from IBM that has been previously proposed as suitable for the virtual screening (VS) of compound libraries based on explicit 3D flexible superimpositions, as part of its deployment within a novel consensus ligand-based virtual screening cascade. A virtual screening protocol using fFLASH for the human estrogen receptor alpha (ERα) was advanced and benchmarked against screens completed using established commercial screening softwares - Catalyst and ROCS. The optimised protocol was applied to a ∼6000 member physical screening collection and virtual 'hits' sourced and biologically assayed. The approach identified a novel, potent and highly selective partial antagonist of the ERα. This study firstly validates the clique detection algorithm utilised by fFLASH and secondly, emphasises the benefits of the consensus approach of employing more than one program in a VS protocol.
The cytochrome P450 (CYP) gene family catalyzes drug metabolism and bioactivation and is therefore relevant to drug development. We determined potency values for 17,143 compounds against five recombinant CYP isozymes (1A2, 2C9, 2C19, 2D6 and 3A4) using an in vitro bioluminescent assay. The compounds included libraries of US Food and Drug Administration (FDA)-approved drugs and screening libraries. We observed cross-library isozyme inhibition (30-78%) with important differences between libraries. Whereas only 7% of the typical screening library was inactive against all five isozymes, 33% of FDA-approved drugs were inactive, reflecting the optimized pharmacological properties of the latter. Our results suggest that low CYP 2C isozyme activity is a common property of drugs, whereas other isozymes, such as CYP 2D6, show little discrimination between drugs and unoptimized compounds found in screening libraries. We also identified chemical substructures that differentiated between the five isozymes. The pharmacological compendium described here should further the understanding of CYP isozymes.