
Personalized medicine is strongly tied with human variome research: understanding the impact of specific genetic sequence mutations on observable human traits will play a key role in the quest for custom drugs therapies and improved patient care. Recent growth in this particular field leveraged the appearance of locus-specific databases (LSDBs). Although these systems are praised in the scientific community, they lack some features that can promote a more widespread usage. Existing systems are closed, independent and designed solely for gene curators. In this paper we present a new approach based on a holistic perspective of the genomic variation field, envisaging the integration of LSDBs, genes and variants, as well as a broad set of related resources in an innovative workspace. A prototype implementation for this approach is deployed online at .
The Spanish National Bioinformatics Institute (Instituto Nacional de Bioinformática in Spanish, or short INB) is an academic service institution founded in 2003 by the mayor research groups in Spain at that time. The INB serves in the coordination, integration and development of Spanish Bioinformatics Resources in projects in the areas of genomics, proteomics and translational medicine. Its mission is to consolidate Bioinformatics as a scientific discipline, providing technical support in Bioinformatics to laboratories, institutions and companies throughout the territory. The JBI2010 conference featured two sessions, "INB Technicians internal session" and "Bioinformatic Software Developments in Spain and beyond", that introduced the state of the art of bioinformatic software developments at the INB and its role at the national and international level. This paper gives a summary of those sessions and presents an overview of the activities and contributions of the INB to the field of bioinformatics.
It becomes increasingly important to support automated service discovering and composition due to the growing number of Web Services and data types in bioinformatics and biomedicine. jORCA is a user-friendly desktop client which is able to discover and invoke Web Services from different metadata repositories for services. This paper demonstrates the usefulness of jORCA for service composition by recreating a previously published workflow, starting with the discovery of data types, service composition (workflow generation) and refinement; to enactment, monitoring and visualization of results. The system has been exhaustively tested and documented and is freely available at .
In the Microbial typing field, the need to have a common understanding of the concepts described and the ability to share results within the community is an increasingly important requisite for the continued development of portable and accurate sequence-based typing methods. These methods are used for bacterial strain identification and are fundamental tools in Clinical Microbiology and Bacterial Population Genetics studies. In this article we propose an ontology designed for the microbial typing field, focusing on the widely used Multi Locus Sequence Typing methodology, and a RESTful API for accessing information systems based on the proposed ontology. This constitutes an important first step to accurately describe, analyze, curate, and manage information for microbial typing methodologies based on sequence based typing methodologies, and allows for the future integration with data analysis Web services.
Deep DNA or RNA sequencing and posterior mapping to a reference sequence is becoming a standard procedure in molecular biology research. Analyzing millions of mapped reads is a challenging task that doesn’t have a unique solution, because experiments using deep sequencing technology vary a great deal among each other. This is why we have developed a flexible tool library called Pyicos, which aims to help biologists in their research when performing their analysis on mapped reads.
The recently published 3D-footprint database contains an up-to-date repository of protein-DNA complexes of known structure that belong to different superfamilies and bind to DNA with distinct specificities. This repository can be scanned by means of sequence alignments in order to look for similar DNA-binding proteins, which might in turn recognize similar DNA motifs. Here we take the complete set of Homeobox proteins from Drosophila melanogaster and their preferred DNA motifs, which would fall in the largest 3D-footprint superfamily and were recently characterized by Noyes and collaborators, and annotate their interface residues. We then analyze the observed amino acid substitutions at equivalent interface positions and their effect on recognition. Finally we estimate to what extent interface similarity, computed over the set of residues which mediate DNA recognition, outperforms BLAST expectation values when deciding whether two aligned Homeobox proteins might bind to the same DNA motif.
In this work we present a significance curve to segregate random alignments from true matches in by identity sequence comparison, especially suitable for sequencing data produced by NGS-technologies. The experimental approach reproduces the random local ungapped similarities distribution by score and length from which it is possible to asses the statistical significance of any particular ungapped similarity. This work includes the study of the distribution behaviour as a function of the experimental technology used to produce the raw sequences, as well as the scoring system used in the comparison. Our approach reproduces the expected behaviour and completes the proposal of Rost and Sander for homology based sequence comparisons. Results can be exploited by computational applications to reduce the computational cost and memory usage.
In order to model protein networks we must extend our knowledge of the protein associations occurring in molecular systems and their functional relationships. We have significantly increased the accuracy of protein association predictions by the meta-statistical integration of three computational methods specifically designed for eukaryotic proteomes. From this former work it was discovered that high-throughput experimental assays seem to perform biased screenings of the real protein networks and leave important areas poorly characterized. This finding supports the convenience to combine computational prediction approaches to model protein interaction networks. We address in this work the challenge of integrating context information, present in predicted and known protein network models, to functionally characterize novel proteins. We applied a random walk-with-restart kernel to our models aiming at fixing some poorly described or unknown proteins involve in angiogenesis. This approach reveals some novel key angiogenic components within the human interactome.
Over the last three decades, the power, resolution and sophistication of scientific experiments has vastly increased, allowing the generation of vast volumes of biological data that need to be stored and processed. Array-oriented Scientific Data Formats are part of an effort by diverse scientific communities to solve the increasing problems of data storage and manipulations. Genome-wide Association Studies (GWAS) based on Single Nucleotide Polymorphism (SNP) arrays are one of the technologies that produce large volumes of data, particularly information on genomic variability. Due to the complexity of the methods and software packages available, each with its particular and intricate formats and work-flows, the analysis of GWAS confronts scientists with a complex hardware and software problematic. To help easing these issues, we have introduced the use of Array-oriented Scientific Data Format databases (NetCDF) in the GWASpi application, a user-friendly, multi-platform, desktop-able software for the management and analysis of GWAS data. The achieved leap of performance has permitted to leverage the most out of commonly available desktop hardware, on which GWASpi now enables ”start- to-end” GWAS management, from raw data to end results and charts. Not only NetCDF allows storing the data efficiently, but it reduces the time needed to achieve the basic results of a GWAS in up to two orders of magnitude. Additionally, the same principles can be used to store and analyze variability data generated by means of ultrasequencing technologies. Available at http://www.gwaspi.org .
De novo identification of genes in newly-sequenced eukaryotic genomes is based on sensors, which are not available in non-model organisms. Many annotation tools have been developed and most of them require sequence training, computer skills and accessibility to sufficient computational power. The main need of non-model organisms is finding genes, transposable elements, repetitions, etc., in reliable assemblies. GENote v.β is intended to cope with these aspects as a web tool for researchers without bioinformatics skills. It facilitates the annotation of new, unfinished sequences with descriptions, GO terms, EC numbers and KEEG pathways. It currently localises genes and transposons, which enable the sorting of contigs or scaffolds from a BAC clone, and reveals some putative assembly inconsistencies. Results are provided in GFF3 format and in tab-delimited text readable in viewers; a summary of findings is provided also as a PNG file.
As the developments in high throughput technologies have become more common and accessible it is becoming usual to take several distinct simultaneous approaches to study the same problem. In practice, this means that data of different types (expression, proteins, metabolites...) may be available for the same study, highlighting the need for methods and tools to analyze them in a combined way. In recent years there have been developed many methods that allow for the integrated analysis of different types of data. Corresponding to a certain tradition in bioinformatics many methodologies are rooted in machine learning such as bayesian networks, support vector machines or graph-based methods. In contrast with the high number of applications from these fields, another that seems to have contributed less to “omic” data integration is multivariate statistics, which has however a long tradition in being used to combine and visualize multidimensional data. In this work, we discuss the application of multivariate statistical approaches to integrate bio-molecular information by using multiple factorial analysis. The techniques are applied to a real unpublished data set consisting of three different data types: clinical variables, expression microarrays and DNA Gel Electrophoretic bands. We show how these statistical techniques can be used to perform reduction dimension and then visualize data of one type useful to explain those from other types. Whereas this is more or less straightforward when we deal with two types of data it turns to be more complicated when the goal is to visualize simultaneously more than two types. Comparison between the approaches shows that the information they provide is complementary suggesting their combined use yields more information than simply using one of them.
BioPax Level 3 is a novel approach to describe pathways at a semantic level by means of an owl ontology. Data provided as BioPax instances is distributed in several databases, and so it is difficult to find integrated information as instances of this ontology. Biopax is a biology ontology that aims to facilitate the integration and exchanged data maintained in biological pathways data. In this paper we present an approach to integrate pathway information by means of an ontology-based mediator (SB-KOM). This mediator has been enabled to produce instances of BioPax Level 3 from integrated data. Thus, it is possible to obtain information about a specific pathway extracting data from distributed databases.
We implement a solution for the automatic ordering of BAC clones from fingerprint data obtained through full digestion of the clones with four enzymes that leave different 3’ recessed ends and fluorescent labeling. The algorithm is inspired to a previous existing approach for building restriction maps.
Since the first papers published in the late nineties, including, for the first time, a comprehensive analysis of microarray data, the number of questions that have been addressed through this technique have both increased and diversified. Initially, interest focussed on genes coexpressing across sets of experimental conditions, implying, essentially, the use of clustering techniques. Recently, however, interest has focussed more on finding genes differentially expressed among distinct classes of experiments, or correlated to diverse clinical outcomes, as well as in building predictors. In addition to this, the availability of accurate genomic data and the recent implementation of CGH arrays has made mapping expression and genomic data on the chromosomes possible. There is also a clear demand for methods that allow the automatic transfer of biological information to the results of microarray experiments. Different initiatives, such as the Gene Ontology (GO) consortium, pathways databases, protein functional motifs, etc., provide curated annotations for genes. Whereas many resources on the web focus mainly on clustering methods, GEPAS has evolved to cope with the aforementioned new challenges that have recently arisen in the field of microarray data analysis. The web-based pipeline for microarray gene expression data, GEPAS, is available athttp://gepas.bioinfo.cnio.es.