A video describing PDX mice, minimal information, and why minimal information is needed in the PDX research field.
Drug discovery and development pipelines are long, complex and depend on numerous factors. Machine learning (ML) approaches provide a set of tools that can improve discovery and decision making for well-specified questions with abundant, high-quality data. Opportunities to apply ML occur in all stages of drug discovery. Examples include target validation, identification of prognostic biomarkers and analysis of digital pathology data in clinical trials. Applications have ranged in context and methodology, with some approaches yielding accurate predictions and insights. The challenges of applying ML lie primarily with the lack of interpretability and repeatability of ML-generated results, which may limit their application. In all areas, systematic and comprehensive high-dimensional data still need to be generated. With ongoing efforts to tackle these issues, as well as increasing awareness of the factors needed to validate ML approaches, the application of ML can promote data-driven decision making and has the potential to speed up the process and reduce failure rates in drug discovery and development.
Abstract Patient-derived tumor xenograft (PDX) mouse models have emerged as an important oncology research platform to study tumor evolution, mechanisms of drug response and resistance, and tailoring chemotherapeutic approaches for individual patients. The lack of robust standards for reporting on PDX models has hampered the ability of researchers to find relevant PDX models and associated data. Here we present the PDX models minimal information standard (PDX-MI) for reporting on the generation, quality assurance, and use of PDX models. PDX-MI defines the minimal information for describing the clinical attributes of a patient's tumor, the processes of implantation and passaging of tumors in a host mouse strain, quality assurance methods, and the use of PDX models in cancer research. Adherence to PDX-MI standards will facilitate accurate search results for oncology models and their associated data across distributed repository databases and promote reproducibility in research studies using these models. Cancer Res; 77(21); e62–66. ©2017 AACR.
Motivation: The field of toxicogenomics (the application of ‘-omics’ technologies to risk assessment of compound toxicities) has expanded in the last decade, partly driven by new legislation, aimed at reducing animal testing in chemical risk assessment but mainly as a result of a paradigm change in toxicology towards the use and integration of genome wide data. Many research groups worldwide have generated large amounts of such toxicogenomics data. However, there is no centralized repository for archiving and making these data and associated tools for their analysis easily available. Results: The Data Infrastructure for Chemical Safety Assessment (diXa) is a robust and sustainable infrastructure storing toxicogenomics data. A central data warehouse is connected to a portal with links to chemical information and molecular and phenotype data. diXa is publicly available through a user-friendly web interface. New data can be readily deposited into diXa using guidelines and templates available online. Analysis descriptions and tools for interrogating the data are available via the diXa portal. Availability and implementation: http://www.dixa-fp7.eu Contact: d.hendrickx@maastrichtuniversity.nl; info@dixa-fp7.eu Supplementary information: Supplementary data are available at Bioinformatics online.
More and more antibody therapeutics are being approved every year, mainly due to their high efficacy and antigen selectivity. However, it is still difficult to identify the antigen, and thereby the function, of an antibody if no other information is available. There are obstacles inherent to the antibody science in every project in antibody drug discovery. Recent experimental technologies allow for the rapid generation of large-scale data on antibody sequences, affinity, potency, structures, and biological functions; this should accelerate drug discovery research. Therefore, a robust bioinformatic infrastructure for these large data sets has become necessary. In this article, we first identify and discuss the typical obstacles faced during the antibody drug discovery process. We then summarize the current status of three sub-fields of antibody informatics as follows: (i) recent progress in technologies for antibody rational design using computational approaches to affinity and stability improvement, as well as ab-initio and homology-based antibody modeling; (ii) resources for antibody sequences, structures, and immune epitopes and open drug discovery resources for development of antibody drugs; and (iii) antibody numbering and IMGT. Here, we review “antibody informatics,” which may integrate the above three fields so that bridging the gaps between industrial needs and academic solutions can be accelerated. This article is part of a Special Issue entitled: Recent advances in molecular engineering of antibody.
In the Semantic Enrichment of the Scientific Literature (SESL) project, researchers from academia and from life science and publishing companies collaborated in a pre-competitive way to integrate and share information for type 2 diabetes mellitus (T2DM) in adults. This case study exposes benefits from semantic interoperability after integrating the scientific literature with biomedical data resources, such as UniProt Knowledgebase (UniProtKB) and the Gene Expression Atlas (GXA). We annotated scientific documents in a standardized way, by applying public terminological resources for diseases and proteins, and other text-mining approaches. Eventually, we compared the genetic causes of T2DM across the data resources to demonstrate the benefits from the SESL triple store. Our solution enables publishers to distribute their content with little overhead into remote data infrastructures, such as into any Virtual Knowledge Broker.
The field of predictive toxicology requires the development of open, public, computable, standardized toxicology vocabularies and ontologies to support the applications required by in silico, in vitro, and in vivo toxicology methods and related analysis and reporting activities. In this article we review ontology developments based on a set of perspectives showing how ontologies are being used in predictive toxicology initiatives and applications. Perspectives on resources and initiatives reviewed include OpenTox, eTOX, Pistoia Alliance, ToxWiz, Virtual Liver, EU-ADR, BEL, ToxML, and Bioclipse. We also review existing ontology developments in neighboring fields that can contribute to establishing an ontological framework for predictive toxicology. A significant set of resources is already available to provide a foundation for an ontological framework for 21st century mechanistic-based toxicology research. Ontologies such as ToxWiz provide a basis for application to toxicology investigations, whereas other ontologies under development in the biological, chemical, and biomedical communities could be incorporated in an extended future framework. OpenTox has provided a semantic web framework for the implementation of such ontologies into software applications and linked data resources. Bioclipse developers have shown the benefit of interoperability obtained through ontology by being able to link their workbench application with remote OpenTox web services. Although these developments are promising, an increased international coordination of efforts is greatly needed to develop a more unified, standardized, and open toxicology ontology framework.
Integrating knowledge from a variety of published reports is a crucial aspect in the development of bioactive entities. Unfortunately, the analysis of the enormous body of items of information reported in the literature for biologically active compounds is hampered by the lack of both uniformity and completeness of data. In order to over come the difficulties arising from the hetero geneity and defectiveness of data formats, 26 scholars from data resource providers, pharmaceutical companies and academic groups have recently proposed a formal list of the items of information that should be provided when describing the preparation and biological evaluation of compounds (Minimum information about a bioactive entity (MIABE). Nature Rev. Drug Discov. 10, 661–669 (2011)) 1 . By following this checklist, researchers would be able to offer the ‘minimum information about a bioactive entity’ (MIABE) to the international scien tific community for each compound to be published. Despite their commendable efforts towards exhaustiveness, in the compilation of the MIABE guidelines the authors over looked an aspect that is crucial, in my opinion, for correctly interpreting biological data obtained from studies on chiral compounds: enantiomeric purity. Since Barlow’s 2 studies were published, the stereochemical purity of chiral compounds has been assumed to be a mandatory aspect to consider when evaluating the activity of single enantiomeric forms of drugs if one enantiomer has an appreciably higher biological activity than the other. In these cases, even low percent ages of the contaminant enantiomer may heavily influence the activity of the principal enantiomer 3 . Furthermore, and in particular, when in vivo activities are considered, the degree of enantiomeric purity may influence both qualitatively and quantitatively the bio logical outcome, often in an unpredictable way 4 . Thus, without specification of enantio meric purity data, the biological activity data reported in the MIABE guidelines for many homochiral compounds could be biased and could expose researchers to potential pitfalls associated with compound chirality 5,6
Foreign substances can have a dramatic and unpredictable adverse effect on human health. In the development of new therapeutic agents, it is essential that the potential adverse effects of all candidates be identified as early as possible. The field of predictive toxicology strives to profile the potential for adverse effects of novel chemical substances before they occur, both with traditional in vivo experimental approaches and increasingly through the development of in vitro and computational methods which can supplement and reduce the need for animal testing. To be maximally effective, the field needs access to the largest possible knowledge base of previous toxicology findings, and such results need to be made available in such a fashion so as to be interoperable, comparable, and compatible with standard toolkits. This necessitates the development of open, public, computable, and standardized toxicology vocabularies and ontologies so as to support the applications required by in silico, in vitro, and in vivo toxicology methods and related analysis and reporting activities. Such ontology development will support data management, model building, integrated analysis, validation and reporting, including regulatory reporting and alternative testing submission requirements as required by guidelines such as the REACH legislation, leading to new scientific advances in a mechanistically-based predictive toxicology. Numerous existing ontology and standards initiatives can contribute to the creation of a toxicology ontology supporting the needs of predictive toxicology and risk assessment. Additionally, new ontologies are needed to satisfy practical use cases and scenarios where gaps currently exist. Developing and integrating these resources will require a well-coordinated and sustained effort across numerous stakeholders engaged in a public-private partnership. In this communication, we set out a roadmap for the development of an integrated toxicology ontology, harnessing existing resources where applicable. We describe the stakeholders’ requirements analysis from the academic and industry perspectives, timelines, and expected benefits of this initiative, with a view to engagement with the wider community.
Bioactive molecules such as drugs, pesticides and food additives are produced in large numbers by many commercial and academic groups around the world. Enormous quantities of data are generated on the biological properties and quality of these molecules. Access to such data - both on licensed and commercially available compounds, and also on those that fail during development - is crucial for understanding how improved molecules could be developed. For example, computational analysis of aggregated data on molecules that are investigated in drug discovery programmes has led to a greater understanding of the properties of successful drugs. However, the information required to perform these analyses is rarely published, and when it is made available it is often missing crucial data or is in a format that is inappropriate for efficient data-mining. Here, we propose a solution: the definition of reporting guidelines for bioactive entities - the Minimum Information About a Bioactive Entity (MIABE) - which has been developed by representatives of pharmaceutical companies, data resource providers and academic groups.
BACKGROUND:The SYMBIOmatics Specific Support Action (SSA) is "an information gathering and dissemination activity" that seeks "to identify synergies between the bioinformatics and the medical informatics" domain to improve collaborative progress between both domains (ref. to http://www.symbiomatics.org). As part of the project experts in both research fields will be identified and approached through a survey. To provide input to the survey, the scientific literature was analysed to extract topics relevant to both medical informatics and bioinformatics.RESULTS:This paper presents results of a systematic analysis of the scientific literature from medical informatics research and bioinformatics research. In the analysis pairs of words (bigrams) from the leading bioinformatics and medical informatics journals have been used as indication of existing and emerging technologies and topics over the period 2000-2005 ("recent") and 1990-1990 ("past"). We identified emerging topics that were equally important to bioinformatics and medical informatics in recent years such as microarray experiments, ontologies, open source, text mining and support vector machines. Emerging topics that evolved only in bioinformatics were system biology, protein interaction networks and statistical methods for microarray analyses, whereas emerging topics in medical informatics were grid technology and tissue microarrays.CONCLUSION:We conclude that although both fields have their own specific domains of interest, they share common technological developments that tend to be initiated by new developments in biotechnology and computer science.
This paper reports on an analysis of the bioinformatics and medical informatics literature with the objective to identify upcoming trends that are shared among both research fields to derive benefits from potential collaborative initiatives for their future. Our results present the main characteristics of the two fields and show that these domains are still relatively separated
The APPLAUSE ESPRIT Project is building major applications using the ElipSys parallel constraint logic programming system developed at ECRC. Two major aims of the project are to advance the state of the art in four commercially significant application areas and to promote the use of ElipSys-like languages among applications developers. This brief paper gives an outline of ElipSys and an overview of the applications being developed within the APPLAUSE Project.
The goal of this research is !o produce more effective methods for protein sequence analysis and strucure prediction through the rse of krnwledge-based techniques for orchestrating protein sequerce and other analyses. We are developing a system (PAPAIN) that will provide intelligent assistance in manipulating and integrating diverse sources of information in a manner ttrat will permit experimentation with hypothesis formation and reasoning styles. This paper describes foundational knowledge engineering studies, ttre resulting logical simuliuion and outlines curent research.
Expression arrays facilitate the monitoring of changes in the expression patterns of large collections of genes. The analysis of expression array data has become a computationally-intensive task that requires the development of bioinformatics technology for a number of key stages in the process, such as image analysis, database storage, gene clustering and information extraction. Here, we review the current trends in each of these areas, with particular emphasis on the development of the related technology being carried out within our groups.
Abstract This paper describes a framework for the use of machine learning as a tool to aid scientists in the discovery of patterns in data. The framework is tested by the application of the inductive logic programming (ILP) program GOLEM to the discovery of constraints in the packing of beta-sheets in alpha/beta proteins. These constraints (rules) play a part in the protein folding problem, an important unsolved problem in molecular biology. Constraints were learnt for four features of beta-sheet packing: the winding direction of two sequential sheets, whether two sequential sheets pack parallel or anti-parallel, whether two sheets pack adjacently, and whether a beta-sheet is at an edge. Investigation of the constraints found revealed interesting patterns, some of which were previously known, others that were novel. Novel features include the discovery that the relationship between pairs of sequential strands is in general one of decreasing size, and that more sequential pairs of strands wind in the direction out than the direction in. We conclude that machine learning has a role for scientists as a pattern discovery tool.
This review provides a description of the background to the first international conference on Intelligent Systems for Molecular Biology and a problem-oriented overview of the papers presented. It focuses on genome analysis, gene identification, protein and RNA structure prediction and function, the modelling of biochemical pathways and data and knowledge bases. The range and quality of papers indicates that intelligent systems is emerging as an important sub-field of computational applications in the biological sciences.