The HmtDB resource hosts a database of human mitochondrial genome sequences from individuals with healthy and disease phenotypes. The database is intended to support both population geneticists as well as clinicians undertaking the task to assess the pathogenicity of specific mtDNA mutations. The wide application of next-generation sequencing (NGS) has provided an enormous volume of high-resolution data at a low price, increasing the availability of human mitochondrial sequencing data, which called for a cogent and significant expansion of HmtDB data content that has more than tripled in the current release. We here describe additional novel features, including: (i) a complete, user-friendly restyling of the web interface, (ii) links to the command-line stand-alone and web versions of the MToolBox package, an up-to-date tool to reconstruct and analyze human mitochondrial DNA from NGS data and (iii) the implementation of the Reconstructed Sapiens Reference Sequence (RSRS) as mitochondrial reference sequence. The overall update renders HmtDB an even more handy and useful resource as it enables a more rapid data access, processing and analysis. HmtDB is accessible at http://www.hmtdb.uniba.it/.
Currently, there is very little information available regarding the microbiome associated with the wine production chain. Here, we used an amplicon sequencing approach based on high-throughput sequencing (HTS) to obtain a comprehensive assessment of the bacterial community associated with the production of three Apulian red wines, from grape to final product. The relationships among grape variety, the microbial community, and fermentation was investigated. Moreover, the winery microbiota was evaluated compared to the autochthonous species in vineyards that persist until the end of the winemaking process. The analysis highlighted the remarkable dynamics within the microbial communities during fermentation. A common microbial core shared among the examined wine varieties was observed, and the unique taxonomic signature of each wine appellation was revealed. New species belonging to the genus Halomonas were also reported. This study demonstrates the potential of this metagenomic approach, supported by optimized protocols, for identifying the biodiversity of the wine supply chain. The developed experimental pipeline offers new prospects for other research fields in which a comprehensive view of microbial community complexity and dynamics is desirable.
The use of "mashups" is expanding considerably in the business environment. Business mashups are usually adopted within integrating business and data-service frameworks to provide the ability to develop new integrated services quickly. Typically, mashups provide organisations with a pronounced and flexible commodity to combine internal with external services in order to create new services, usually accessed through user-friendly Web-browser interfaces. In this study, a Web 2.0 technology was adopted to promote a key field of bioinformatics research through the management and automation of bioinformatics workflows. Consumables (widgets and services) have been developed using the Lotus Widget Factory, an Eclipse plug-in providing an easy-to-use development environment enabling developers of all skill levels to create dynamic widgets rapidly. A workflow built from widgets works as follows: the core widget receives data from one or more widgets, invokes a generic Web service, performing iteration and/or recursion, and sends the results to all other connected widgets. The number of iterations and recursions depends on the input data-set dimension and user-defined parameter values related to each specific application. Some prototype workflows have been assembled and tested with a number of widgets created with algorithms from the European Molecular Biology Open Software Suite (EMBOSS), exposed as Web services. The adoption of recent Web 2.0 technologies, such as mashup platforms, has enabled rapid generation, sharing and discovery of reusable application building-blocks (widgets, feeds, mashups), and has shown to be a plausible alternative environment for supporting bioinformatics workflow design, management and execution.
Motivation and Objectives The biodiversity is nowadays one of the main scientific area of interest because of its importance for a sustainable development in many technological domains such as biotechnologies as well as for agriculture and human health. For instance, plant genetic resources are the basis of food security and consist of diversity of seeds and planting material of traditional varieties or modern cultivars and crop wild relatives. These resources are used as food, feed for domesticated animals and in recent years for the identification of new chemical compounds to be used in clinical therapeutic protocols. Biodiversity research communities have to deal with data coming from many different domains (e.g., biology, geography, evolutionary studies, genomics, taxonomy, environmental sciences, etc.). Collecting and integrating data from so many disparate resources is not a trivial task, data are extremely scattered, heterogeneous in format and purpose, often protected in repositories of diverse research institutes. With the advent of next generation technologies, molecular biodiversity research is producing large amounts of data that researchers use for complex comparative analyses exploiting information present both in public databases (like GenBank) and in their personal repositories. Improving the management of molecular data and their integration with related information present in the genetic resources databases such as morphologic, geographic and ecologic data will lead to new valuable biodiversity knowledge. Driven by the widely diffused trend of the web of sharing information through aggregation of people with the same interests (social networks), and by the new type of database architecture defined as dynamic distributed federated database, here we present MBlabDB, a tool representing a new paradigm of data integration in the biodiversity domain. Methods MBLabDB uses a hybrid approach of data federation and data warehousing. The system architecture (Figure 1) is based on the integrated cooperation of several components: a robust Database Management System, managing the large volume of molecular data and information available in public resources such as GenBank; a set of federated databases implemented with GaianDB (Bent G. et al., 2008) tool, managing remote specialized biodiversity databases; the IBM Information Integrator, implementing the database conceptual schema and integrating all federated databases with public molecular data using a data warehouse approach. The conceptual schema of MBLabDB named MolecularBiodiversity Database Schema (Pannarale et al.,2012), is tailored to biodiversity data collection, integrationand analysis. It is modeled on six main sections: Individual, MolecularData, Experiment, Collection, Supply chain and Taxonomy. The MolecularData section is structured following a Chado-like model (Mungall CJ et al., 2007), using Sequence Ontology (Eilbeck K et al., 2005) entities and relations. Similarly the Taxonomy section has been designed in order to incorporate and integrate more than one taxonomy, because of different reference taxonomies that could be related to a taxonomic kingdom. The federated databases have been implemented by GaianDB (Bent G. et al., 2008), a Dynamic Distributed Federated Database of sources whose growth is regulated by biologically inspired principles and graph theoretic methods. The idea is to create a network of database nodes, each containing specialised collections of biodiversity data, and to expose their content by means of a GaianDB data server. Information coming from the network nodes are collected by a GaianDB hub and are integrated with public data by means of the Information Integrator server. Two steps are needed to add a new GaianDB node: the installation of a GaianDB server instance and the writing of a wrapper for the mapping of the local schema with the general MBLabDB schema. An efficient and reliable ETL (Extraction, Transformation and Load) module, implemented with CLIPS Rule Based Programming Language (Pannarale et al., 2012), has been used to integrate GenBank data in MBLabDB. The ETL procedure extracts information from the GenBank entries and fits them into the MBLabDB schema. The MBLabDB graphical user interface (GUI) has been developed as a Java platform web application. In the GUI the public-private data integration is highlighted through the implementation of taxonomic and ontology based queries. Results and Discussion Currently, MBLabDB integrates 4,360,218 entries from the GenBank database and two biodiversity data collections: the ITEM Collection (http://www.ispa.cnr.it/Collection), located at the ISPA-CNR server (containing 9,181 specimen and 3,584 sequences), and the IGV Germoplasm Database (http://www.igv.cnr.it), located at the IGV-CNR server (containing 11,113 accessions). Furthermore the NCBI Taxonomy (www.ncbi.nlm.nih.gov/Taxonomy) and the Catalogue of Life (http://www.catalogueoflife.org/) taxonomic classifications have been included in the Taxonomy section. Two search and retrieval modalities are available in MBLabDB, an advanced query mode, where search criteria and results can be combined using an incremental composition of “querying & filtering”, and an ontology based retrieval that queries data using the biological concepts expressed by the Sequence Ontology. Therefore, MBLabDB combines public molecular data with biodiversity data contained in genetic resource collections, that are typical of the biodiversity domain. By way of example, using MBLabDB a researcher can extract datasets of sequences related to specimen of his own interest using biodiversity criteria such as species/varieties, geolocation, morphology and passport data. Using the MBLabDB paradigm of data integration, database hosting, management and information sharing strategy of specialised resources are left to the research group owner of the data collection. So the biodiversity research groups can contribute to the information network by sharing their data sources with a reasonable effort. In this network, named Social Database for Molecular Biodiversity Data, information remains scattered, but knowledge are shared. Acknowledgements This work was supported by DM19410 - Bioinformatics Molecular Biodiversity LABoratory - MBLab (www.mblabproject.it). References Bent G. et al. (2008) A dynamic distributed federated database. Second Annual Conference of ITA, Imperial College, London Eilbeck K et al. (2005) The Sequence Ontology: A tool for the unification of genome annotations. Genome Biology 6:R44 Mungall CJ et al. (2007) A Chado case study: an ontology-based modular schema for representing genome-associated biological information. Bioinformatics 23: i337-i346 Pannarale P et al. (2012) GIDL: a rule based expert system for GenBank Intelligent Data Loading into the Molecular Biodiversity database. BMC Bioinformatics 13 Suppl 4:S4 Note: Figures and tables are available in PDF version only.
BACKGROUND:In the scientific biodiversity community, it is increasingly perceived the need to build a bridge between molecular and traditional biodiversity studies. We believe that the information technology could have a preeminent role in integrating the information generated by these studies with the large amount of molecular data we can find in bioinformatics public databases. This work is primarily aimed at building a bioinformatic infrastructure for the integration of public and private biodiversity data through the development of GIDL, an Intelligent Data Loader coupled with the Molecular Biodiversity Database. The system presented here organizes in an ontological way and locally stores the sequence and annotation data contained in the GenBank primary database.METHODS:The GIDL architecture consists of a relational database and of an intelligent data loader software. The relational database schema is designed to manage biodiversity information (Molecular Biodiversity Database) and it is organized in four areas: MolecularData, Experiment, Collection and Taxonomy. The MolecularData area is inspired to an established standard in Generic Model Organism Databases, the Chado relational schema. The peculiarity of Chado, and also its strength, is the adoption of an ontological schema which makes use of the Sequence Ontology. The Intelligent Data Loader (IDL) component of GIDL is an Extract, Transform and Load software able to parse data, to discover hidden information in the GenBank entries and to populate the Molecular Biodiversity Database. The IDL is composed by three main modules: the Parser, able to parse GenBank flat files; the Reasoner, which automatically builds CLIPS facts mapping the biological knowledge expressed by the Sequence Ontology; the DBFiller, which translates the CLIPS facts into ordered SQL statements used to populate the database. In GIDL Semantic Web technologies have been adopted due to their advantages in data representation, integration and processing.RESULTS AND CONCLUSIONS:Entries coming from Virus (814,122), Plant (1,365,360) and Invertebrate (959,065) divisions of GenBank rel.180 have been loaded in the Molecular Biodiversity Database by GIDL. Our system, combining the Sequence Ontology and the Chado schema, allows a more powerful query expressiveness compared with the most commonly used sequence retrieval systems like Entrez or SRS.
The increasing use of phylogeny in biological studies is limited by the need to make available more efficient tools for computing distances between trees. The geodesic tree distance-introduced by Billera, Holmes, and Vogtmann-combines both the tree topology and edge lengths into a single metric. Despite the conceptual simplicity of the geodesic tree distance, algorithms to compute it don't scale well to large, real-world phylogenetic trees composed of hundred or even thousand leaves. In this paper, we propose the geodesic distance as an effective tool for exploring the likelihood profile in the space of phylogenetic trees, and we give a cubic time algorithm, GeoHeuristic, in order to compute an approximation of the distance. We compare it with the GTP algorithm, which calculates the exact distance, and the cone path length, which is another approximation, showing that GeoHeuristic achieves a quite good trade-off between accuracy (relative error always lower than 0.0001) and efficiency. We also prove the equivalence among GeoHeuristic, cone path, and Robinson-Foulds distances when assuming branch lengths equal to unity and we show empirically that, under this restriction, these distances are almost always equal to the actual geodesic.
The LIBI project (International Laboratory of BioInformatics), which started in 2005 and will end in 2009, was initiated with the aim of setting up an advanced bioinformatics and computational biology laboratory, focusing on basic and applied research in modern biology and biotechnologies. One of the goals of this project has been the development of a Grid Problem Solving Environment, built on top of EGEE, DEISA and SPACI infrastructures, to allow the submission and monitoring of jobs mapped to complex experiments in bioinformatics. In this work we describe the architecture of this environment and describe several case studies and related results which have been obtained using it.
At present, 51 genes are already known to be responsible for Non-Syndromic hereditary Hearing Loss (NSHL), but the knowledge of 121 NSHL-linked chromosomal regions brings to the hypothesis that a number of disease genes have still to be uncovered. To help scientists to find new NSHL genes, we built a gene-scoring system, integrating Gene Ontology, NCBI Gene and Map Viewer databases, which prioritizes the candidate genes according to their probability to cause NSHL. We defined a set of candidates and measured their functional similarity with respect to the disease gene set, computing a score ( S S M avg) that relies on the assumption that functionally related genes might contribute to the same (disease) phenotype. A Kolmogorov-Smirnov test, comparing the pair-wise distribution on the disease gene set with the distribution on the remaining human genes, provided a statistical assessment of this assumption. We found at a p-value < 2.2.10 (-16) that the former pair-wise is greater than the latter, justifying a prioritization strategy based on the functional similarity of candidate genes respect to the disease gene set. A cross-validation test measured to what extent the S S M avg ranking for NSHL is different from a random ordering: adding 15% of the disease genes to the candidate gene set, the ranking of the disease genes in the first eight positions resulted statistically different from a hypergeometric distribution with a p-value = 2.04.10(-5) and a power > 0.99. The twenty top-scored genes were finally examined to evaluate their possible involvement in NSHL. We found that half of them are known to be expressed in human inner ear or cochlea and are mainly involved in remodeling and organization of actin formation and maintenance of the cilia and the endocochlear potential. These findings strongly indicate that our metric was able to suggest excellent NSHL candidates to be screened in patients and controls for causative mutations.
BACKGROUND:A standardized and cost-effective molecular identification system is now an urgent need for Fungi owing to their wide involvement in human life quality. In particular the potential use of mitochondrial DNA species markers has been taken in account. Unfortunately, a serious difficulty in the PCR and bioinformatic surveys is due to the presence of mobile introns in almost all the fungal mitochondrial genes. The aim of this work is to verify the incidence of this phenomenon in Ascomycota, testing, at the same time, a new bioinformatic tool for extracting and managing sequence databases annotations, in order to identify the mitochondrial gene regions where introns are missing so as to propose them as species markers.METHODS:The general trend towards a large occurrence of introns in the mitochondrial genome of Fungi has been confirmed in Ascomycota by an extensive bioinformatic analysis, performed on all the entries concerning 11 mitochondrial protein coding genes and 2 mitochondrial rRNA (ribosomal RNA) specifying genes, belonging to this phylum, available in public nucleotide sequence databases. A new query approach has been developed to retrieve effectively introns information included in these entries.RESULTS:After comparing the new query-based approach with a blast-based procedure, with the aim of designing a faithful Ascomycota mitochondrial intron map, the first method appeared clearly the most accurate. Within this map, despite the large pervasiveness of introns, it is possible to distinguish specific regions comprised in several genes, including the full NADH dehydrogenase subunit 6 (ND6) gene, which could be considered as barcode candidates for Ascomycota due to their paucity of introns and to their length, above 400 bp, comparable to the lower end size of the length range of barcodes successfully used in animals.CONCLUSION:The development of the new query system described here would answer the pressing requirement to improve drastically the bioinformatics support to the DNA Barcode Initiative. The large scale investigation of Ascomycota mitochondrial introns performed through this tool, allowing to exclude the introns-rich sequences from the barcode candidates exploration, could be the first step towards a mitochondrial barcoding strategy for these organisms, similar to the standard approach employed in metazoans.
Margherita Berardi 1,5, Donato Malerba1 , Roberta Piredda2, Marcella Attimonelli2, Gaetano Scioscia3,4 and Pietro Leo3,4 1Dipartimento di Informatica – Università degli Studi di Bari 2Dipartimento di Biochimica e Biologia Molecolare “E.Quagliariello" – Università degli Studi di Bari 3IBM Italia S.p.A. Molecular Biodiversity Laboratory 4IBM Italia S.p.A. GBS Innovation Centre 5Exhicon S.r.l., Bari Italy
Background Population genetics studies based on the analysis of mtDNA and mitochondrial disease studies have produced a huge quantity of sequence data and related information. These data are at present worldwide distributed in differently organised databases and web sites not well integrated among them. Moreover it is not generally possible for the user to submit and contemporarily analyse its own data comparing them with the content of a given database, both for population genetics and mitochondrial disease data. Results HmtDB is a well-integrated web-based human mitochondrial bioinformatic resource aimed at supporting population genetics and mitochondrial disease studies, thanks to a new approach based on site-specific nucleotide and aminoacid variability estimation. HmtDB consists of a database of Human Mitochondrial Genomes, annotated with population data, and a set of bioinformatic tools, able to produce site-specific variability data and to automatically characterize newly sequenced human mitochondrial genomes. A query system for the retrieval of genomes and a web submission tool for the annotation of new genomes have been designed and will soon be implemented. The first release contains 1255 fully annotated human mitochondrial genomes. Nucleotide site-specific variability data and multialigned genomes can be downloaded. Intra-human and inter-species aminoacid variability data estimated on the 13 coding for proteins genes of the 1255 human genomes and 60 mammalian species are also available. HmtDB is freely available, upon registration, at http://www.hmdb.uniba.it . Conclusion The HmtDB project will contribute towards completing and/or refining haplogroup classification and revealing the real pathogenic potential of mitochondrial mutations, on the basis of variability estimation.
Finding disease relationships requires laborious examination of hundreds of possible candidate heterogeneous factors. Much of the related information is currently contained in biological and medical journals, making biomedical text mining a central bioinformatic problem. More than 14 million abstracts of such papers are contained in the Medline collection and are available online. In this paper we present a data mining engine, namely MeSH Terms Associator (MTA), that has been employed in a distributed architecture to refine a generic PubMed query by means of discovery of concept relations in the form of association rules. However, the number of discovered association rules is usually high and the interest of most of them does not fulfil user expectations. In addition, the presentation of thousands of rules can discourage users from interpreting them. To overcome this problem we investigate the application of some filtering techniques. Experimental results on datasets corresponding to real-world biomedical queries are discussed and future directions are drawn.
Building integrated bioinformatic platforms is one of the most challenging tasks which Bioinformatics community is dealing with in recent years [1-2]. Facing this task, a number of specific problems arises connected to data integration, integration of specialized tools and algorithms. The solution described in this paper goes in the direction to solve this challenge. It is characterized by two original assumptions: 1) a quite sharp division between the data realm of a bioinformatics analysis and its components in terms of algorithms and processes, 2) the conception of a rigorous algebra that allows researchers to formalize their analyses in terms of atomic process workflows. As a result of this approach two bioinformatics web tools, BioWBI and WEE, have been designed and prototyped by our group to provide researchers with a virtual collaborative workspace in which defining their data-sources, drawing graphically as well as executing analysis workflows. These tools constitute the basic components of a much more general bioinformatic e-workplace.
Margherita Berardi合作论文数Dipartimento di Informatica
Universit?? degli Studi di Bari2