Chemical mixtures have recently come to the attention of open standards and data structures for capturing machine-readable descriptions for informatics uses. At the present time, essentially all transmission of information about mixtures is done using short text descriptions that are readable only by trained scientists, and there are no accessible repositories of marked-up mixture data. We have designed a machine learning tool that can interpret mixture descriptions and upgrade them to the high-level Mixfile format, which can in turn be used to generate Mixtures InChI notation. The interpretation achieves a high success rate and can be used at scale to markup large catalogs and inventories, with some expert checking to catch edge cases. The training data that was accumulated during the project is made openly available, along with previously released mixture editing tools and utilities.
We describe a file format that is designed to represent mixtures of compounds in a way that is fully machine readable. This Mixfile format is intended to fill the same role for substances that are composed of multiple components as the venerable Molfile does for specifying individual structures. This much needed datastructure is intended to replace current practices for communicating information about mixtures, which usually relies on human-readable text descriptions, drawing several species within a single molecular diagram, or mutually incompatible ad hoc solutions. We describe an open source software application for editing mixture files, which can also be used as web-ready tools for manipulating the file format. We also present a corpus of mixture examples, which we have extracted from collections of text-based descriptions. Furthermore, we present an early look at the proposed IUPAC Mixtures InChI specification, instances of which can be automatically generated using the Mixfile format as a precursor.
We are now seeing the benefit of investments made over the last decade in high-throughput screening (HTS) that is resulting in large structure activity datasets entering public and open databases such as ChEMBL and PubChem. The growth of academic HTS screening centers and the increasing move to academia for early stage drug discovery suggests a great need for the informatics tools and methods to mine such data and learn from it. Collaborative Drug Discovery, Inc. (CDD) has developed a number of tools for storing, mining, securely and selectively sharing, as well as learning from such HTS data. We present a new web based data mining and visualization module directly within the CDD Vault platform for high-throughput drug discovery data that makes use of a novel technology stack following modern reactive design principles. We also describe CDD Models within the CDD Vault platform that enables researchers to share models, share predictions from models, and create models from distributed, heterogeneous data. Our system is built on top of the Collaborative Drug Discovery Vault Activity and Registration data repository ecosystem which allows users to manipulate and visualize thousands of molecules in real time. This can be performed in any browser on any platform. In this chapter we present examples of its use with public datasets in CDD Vault. Such approaches can complement other cheminformatics tools, whether open source or commercial, in providing approaches for data mining and modeling of HTS data.
Collaborative Drug Discovery (CDD) provides a whole solution for today?s biological and chemical data needs, differentiated by ease-of-use and superior collaborative capabilities. CDD Vault
Neglected disease drug discovery is generally poorly funded compared with major diseases and hence there is an increasing focus on collaboration and precompetitive efforts such as public-private partnerships (PPPs). The More Medicines for Tuberculosis (MM4TB) project is one such collaboration funded by the EU with the goal of discovering new drugs for tuberculosis. Collaborative Drug Discovery has provided a commercial web-based platform called CDD Vault which is a hosted collaborative solution for securely sharing diverse chemistry and biology data. Using CDD Vault alongside other commercial and free cheminformatics tools has enabled support of this and other large collaborative projects, aiding drug discovery efforts and fostering collaboration. We will describe CDD's efforts in assisting with the MM4TB project.
Annotation of bioassay protocols using semantic web vocabulary is a way to make experiment descriptions machine-readable. Protocols are communicated using concise scientific English, which precludes most kinds of analysis by software algorithms. Given the availability of a sufficiently expressive ontology, some or all of the pertinent information can be captured by asserting a series of facts, expressed as semantic web triples (subject, predicate, object). With appropriate annotation, assays can be searched, clustered, tagged and evaluated in a multitude of ways, analogous to other segments of drug discovery informatics. The BioAssay Ontology (BAO) has been previously designed for this express purpose, and provides a layered hierarchy of meaningful terms which can be linked to. Currently the biggest challenge is the issue of content creation: scientists cannot be expected to use the BAO effectively without having access to software tools that make it straightforward to use the vocabulary in a canonical way. We have sought to remove this barrier by: (1) defining a BioAssay Template (BAT) data model; (2) creating a software tool for experts to create or modify templates to suit their needs; and (3) designing a common assay template (CAT) to leverage the most value from the BAO terms. The CAT was carefully assembled by biologists in order to find a balance between the maximum amount of information captured vs. low degrees of freedom in order to keep the user experience as simple as possible. The data format that we use for describing templates and corresponding annotations is the native format of the semantic web (RDF triples), and we demonstrate some of the ways that generated content can be meaningfully queried using the SPARQL language. We have made all of these materials available as open source (http://github.com/cdd/bioassay-template), in order to encourage community input and use within diverse projects, including but not limited to our own commercial electronic lab notebook products.
PURPOSE:We propose a framework with simple proxies to dissect the relative energy contributions responsible for standard drug discovery binding activity.METHODS:We explore a rule of thumb using hydrogen-bond donors, hydrogen-bond acceptors and rotatable bonds as relative proxies for the thermodynamic terms. We apply this methodology to several datasets (e.g., multiple small molecules profiled against kinases, Mycobacterium tuberculosis (Mtb) high throughput screening (HTS) and structure based drug design (SBDD) derived compounds, and FDA approved drugs).RESULTS:We found that Mtb active compounds developed through SBDD methods had statistically significantly larger PEnthalpy values than HTS derived compounds, suggesting these compounds had relatively more hydrogen bond donor and hydrogen bond acceptors compared to rotatable bonds. In recent FDA approved medicines we found that compounds identified via target-based approaches had a more balanced enthalpic relationship between these descriptors compared to compounds identified via phenotypic screensCONCLUSIONS:As it is common to experimentally optimize directly for total binding energy, these computational methods provide alternative calculations and approaches useful for compound optimization alongside other common metrics in available software and databases.
The past decade has seen an increased availability of databases with both molecular structures and bioactivity (e.g. binding affinity data) curated from publications or submitted by high throughput screening centers. Early collections of curated data, such as BindingDB or ChemBank, were initially eclipsed by PubChem, ChEMBL and commercial literature databases. In 2016, there are now many smaller focused databases that are freely available, such as the GtoPdb and the public data in Collaborative Drug Discovery (CDD) Vault. In this chapter, we will describe why these types of databases are important, what they can be used for and some of the issues we need to be aware of. The reader will learn more about the evolution of these databases, how they benefit the scientific community and potential improvements that could be made to make them more accessible.