
Information visualization techniques, which take advantage of the bandwidth of human vision, are powerful tools for organizing and analyzing a large amount of data. In the postgenomic era, information visualization tools are indispensable for biomedical research. This paper aims to present an overview of current applications of information visualization techniques in bioinformatics for visualizing different types of biological data, such as from genomics, proteomics, expression profiling and structural studies. Finally, we discuss the challenges of information visualization in bioinformatics related to dealing with more complex biological information in the emerging fields of systems biology and systems medicine.
Bayesian regularized artificial neural networks (BRANNs) are used in the development of quantitative SAR models. These networks have the potential to solve several problems that arise in QSAR modeling such as choice of model, robustness of model, choice of validation set, size of validation effort, and optimization of network architecture. The application of the methods to a wide range of problems, including target-based QSAR, ADMET modeling and eukaryotic promoter finding, is illustrated.
…Can we predict a global genetic network on the basis of genomic/microarray data and is this the ‘right’ question to ask?
A quick search through high-throughput proteomics and genomics data can reveal information on many aspects of protein function, such as mutant phenotypes, protein interactions, mRNA expression patterns, transcriptional regulation, and even protein structure. The computational integration of such data is proving to be the most effective route to protein function.
In this review, we present a survey of the scientific literature on biological-function attribution by computer. The focus is on methods of predicting protein-based biological function in terms of intracellular protein–protein interactions. For each methodology, a critical evaluation of strengths and weaknesses is presented from the literature in the field. A conceptual classification scheme is proposed that separates computational methodologies into those based on biological hypotheses and those based on machine learning hypotheses. This represents a different perspective on in silico function attribution. The scope of the discussion here centers on various machine-learning approaches reported in the literature. Machine-generated hypotheses implicitly model biological function by learning from patterns inherent in data.
Rapid advances in mass spectrometry have positioned it as a prime tool for diagnosis and biomarker discovery. Distilling a handful of accurate disease markers from the thousands of mass-to-charge ratios that comprise a mass spectrum raises non-trivial data analytic challenges. Thus, mass-spectral analysis is turning more and more to data mining, a technology that lies at the crossroads of artificial intelligence and statistical data analysis. This review describes recent attempts at applying data mining techniques to extract diagnostic biomarkers for cancer from SELDI-TOF and MALDI-TOF mass spectra.
Computer models usefully complement experimentation in the efficient discovery of MHC-binding peptides and T-cell epitopes, and have been applied successfully to predict T-cell epitopes in infectious disease, cancer, autoimmunity and allergy. Prediction methods include binding motifs, quantitative matrices, various artificial intelligence techniques and molecular modelling. Computational modelling should be performed according to strict standards, requiring careful data selection for model building, followed by adequate testing and validation. Many web-based databases and binding prediction programs are now available. Although certain prediction programs are reasonably accurate, at least for some MHC alleles, one cannot guarantee that all models produce results of adequate predictivity and therefore these prediction results should be used with care.
The pharmaceutical industry is rapidly adopting virtual screening techniques aimed at identifying chemical compounds that have the required ingredients to become successful drugs. The need for a high-throughput yet inexpensive evaluation of the molecules in silico, before they are tested or even made, is necessitated by the increasing costs of drug discovery and the current ‘drought’ in the new drug approvals. The computational filtering step is especially important for combinatorial chemistry, where billions of compounds can be synthesized from the commodity reagents. Neural networks have a proven ability to model complex relationships between pharmaceutically relevant properties and chemical structures of compounds, and have the potential to improve diversity and quality of virtual screening. This review describes how neural networks are currently used and might be further used in virtual screening of combinatorial libraries.
The automatic integration of information resources in the life sciences is one of the most challenging goals facing biomedical informatics today. Controlled vocabularies have played an important role in realizing this goal, by making it possible to draw together information from heterogeneous sources secure in the knowledge that the same terms will also represent the same entities on all occasions of use. One of the most impressive achievements in this regard is the Gene Ontology (GO), which is rapidly acquiring the status of a de facto standard in the field of gene and gene product annotations, and whose methodology has been much intimated in attempts to develop controlled vocabularies for shared use in different domains of biology. The GO Consortium has recognized, however, that its controlled vocabulary as currently constituted is marked by several problematic features - features which are characteristic of much recent work in bioinformatics and which are destined to raise increasingly serious obstacles to the automatic integration of biomedical information in the future. Here, we survey some of these problematic features, focusing especially on issues of compositionality and syntactic regimentation.
To cope with increasing numbers of compounds and much higher demand on early ADME data, higher-throughput absorption, distribution, metabolism, and elimination (ADME) in vitro screens have been developed. However, the realisation that screening is a costly business urged pharma and biotech companies to investigate in silico prediction and simulation of pharmacokinetic data and processes. Good progress has been made in simulating the extent and rate of absorption, plasma concentration–time profiles, the site of metabolism within the compound and formed metabolites, and the effect of drug–drug interactions. Further development of robust in silico tools, in combination with high-throughput in vitro screening, will lead to an in combo approach towards drug discovery.
Bioinformatics is bimodal, with results gathered either from rigid pipelines, their functions encoded in programs accessible only to specialists, or from ad hoc collections of scripts or web pages, used once, then lost track of or thrown away. We discuss the possibility of developing tools to straddle this divide, in particular, tools using graphical dataflow programming, which have the potential to make bioinformatics flexible, approachable and repeatable.
Various knowledge-based methods can be applied to the analysis and modeling of gene expression and clinical data. I focus on knowledge-based neural networks (KBNN), illustrating how they have been used in cancer profiling and gene regulatory network discovery. I demonstrate that the KBNN approach facilitates adaptive learning and knowledge discovery, along with the creation of more accurate classification and prognostic systems, both global and personalized.
Mark is responsible for the strategic evaluation, identification, and integration of new technologies into worldwide R&D. He joined Vertex as a Founding Scientist in 1990 and started the company's molecular modeling, bioinformatics, IS and chemoinformatics groups. He is a co-inventor of a half-dozen compounds in Vertex's clinical pipeline. Notable among these are the HIV-protease inhibitors Agenerase® and Lexiva®, Vertex's first two marketed drugs, as well as the anti-inflammatory compounds Pralnacasan® and VX-765, both oral blockers of the IL-1 beta converting enzyme. Mark also played a key role in the creation and launch of Vertex's gene family approach, dubbed chemogenomics. Mark was named Chief Technology Officer in 2001.
The ligand-target matrix is the core element of the chemogenomics knowledge space. Its content is derived from different chemical and biological data sources, which need to be integrated, and a large variety of molecular informatics tools available to mine it. Initial, although incomplete, approaches to assemble this knowledge space are reported together with their use in virtual screening. The existing relationship between the structural properties of ligands, especially when described with pharmacophore descriptors, and their biological activity has been demonstrated to be useful for drug discovery applications.