In silico approaches are viewed as ways to complement and enhance traditional risk assessment and safety evaluation. Modern regulatory science therefore presents new opportunities and challenges for the field of chemoinformatics. Risk assessment is an important part of regulatory science that is conducted when the safety of a test substance needs to be addressed in a specific exposure condition. Once identified, risks arisen from chemical exposure are characterized and managed. Quite often researchers are faced with data gaps. Filling these data gaps is therefore an important activity of risk assessment. To design databases suitable for risk assessment of chemicals. This chapter introduces new types of descriptors applicable to address the new approach methodologies (NAMs), including metabolomics. As the NAM framework opens the door to in silico approaches, quantitative structure-activity relationship (QSAR) and other computational methods can sometimes be used to fill data gaps.
Chemotypes are a new approach for representing molecules, chemical substructures and patterns, reaction rules, and reactions. Chemotypes are capable of integrating types of information beyond what is possible using current representation methods (e.g., SMARTS patterns) or reaction transformations (e.g., SMIRKS, reaction SMILES). Chemotypes are expressed in the XML-based Chemical Subgraphs and Reactions Markup Language (CSRML), and can be encoded not only with connectivity and topology but also with properties of atoms, bonds, electronic systems, or molecules. CSRML has been developed in parallel with a public set of chemotypes, i.e., the ToxPrint chemotypes, which are designed to provide excellent coverage of environmental, regulatory, and commercial-use chemical space, as well as to represent chemical patterns and properties especially relevant to various toxicity concerns. A software application, ChemoTyper has also been developed and made publicly available in order to enable chemotype searching and fingerprinting against a target structure set. The public ChemoTyper houses the ToxPrint chemotype CSRML dictionary, as well as reference implementation so that the query specifications may be adopted by other chemical structure knowledge systems. The full specifications of the XML-based CSRML standard used to express chemotypes are publicly available to facilitate and encourage the exchange of structural knowledge.
Early prediction of safety issues in drug development is at the same time highly desirable and highly challenging. Recent advances emphasize the importance of understanding the whole chain of causal events leading to observable toxic outcomes. Here we describe an integrative modeling strategy based on these ideas that guided the design of eTOXsys, the prediction system used by the eTOX project. Essentially, eTOXsys consists of a central server that marshals requests to a collection of independent prediction models and offers a single user interface to the whole system. Every of such model lives in a self-contained virtual machine easy to maintain and install. All models produce toxicity-relevant predictions on their own but the results of some can be further integrated and upgrade its scale, yielding in vivo toxicity predictions. Technical aspects related with model implementation, maintenance and documentation are also discussed here. Finally, the kind of models currently implemented in eTOXsys is illustrated presenting three example models making use of diverse methodology (3D-QSAR and decision trees, Molecular Dynamics simulations and Linear Interaction Energy theory, and fingerprint-based QSAR).
The high-quality in vivo preclinical safety data produced by the pharmaceutical industry during drug development, which follows numerous strict guidelines, are mostly not available in the public domain. These safety data are sometimes published as a condensed summary for the few compounds that reach the market, but the majority of studies are never made public and are often difficult to access in an automated way, even sometimes within the owning company itself. It is evident from many academic and industrial examples, that useful data mining and model development requires large and representative data sets and careful curation of the collected data. In 2010, under the auspices of the Innovative Medicines Initiative, the eTOX project started with the objective of extracting and sharing preclinical study data from paper or pdf archives of toxicology departments of the 13 participating pharmaceutical companies and using such data for establishing a detailed, well-curated database, which could then serve as source for read-across approaches (early assessment of the potential toxicity of a drug candidate by comparison of similar structure and/or effects) and training of predictive models. The paper describes the efforts undertaken to allow effective data sharing intellectual property (IP) protection and set up of adequate controlled vocabularies) and to establish the database (currently with over 4000 studies contributed by the pharma companies corresponding to more than 1400 compounds). In addition, the status of predictive models building and some specific features of the eTOX predictive system (eTOXsys) are presented as decision support knowledge-based tools for drug development process at an early stage.
Two quantitative pKa prediction models for aliphatic carboxylic acids and for alcohols were developed by multiple linear-regression (MLR) analysis with empirical atomic descriptors. The acid and alcohol molecules were described by a set of five and four atomic descriptors, respectively. For the pKa model of 1122 aliphatic carboxylic acids, the squared correlation coefficient is 0.813 with a standard error of prediction of 0.423; for the pKa model of 288 alcohols, the squared correlation coefficient is 0.817 with a standard error of prediction of 0.755, respectively. The good predictive abilities of the models obtained were indicated by both cross-validation and by external validation. An atomic descriptor was developed to model the inductive effect of the neighboring atoms for a central atom in a molecule. The ability of the descriptor to measure the inductive effect of substituent groups was demonstrated by a good correlation of this descriptor with Taft sigma* constants in aliphatic carboxylic acids. It provides a new approach to estimate Taft sigma* constants directly from molecular structures. An algorithm using Kohonen neural networks for splitting a data set into a training set and a test set is also presented.