Modern academic and industrial research in life sciences generates huge amounts of data and information. To extract knowledge from this information space, optimized integration and retrieval software tools are essential. In the last years, a number of academic as well as commercial systems have been developed to solve this problem. However, as scientific projects are distributed at different locations (e.g., subsidiaries of companies, academic partnerships), data exchange and availability must be realized in a way that avoids data replication.In this article, we describe a global solution for integrating distributed information by applying the BioRSTM Integration and Retrieval System and its inter-BioRS communication capability that goes beyond the standard issue of local data integration. Each site integrates and maintains locally generated data using a local copy of the BioRS software. Applying the inter-BioRS communication, all available BioRS instances can communicate with each other realizing a global network of integrated databanks. All databanks integrated in this network can be accessed from any site without any data replication. This open system allows the addition of new information and sites dynamically. However, access privileges for certain databanks can be maintained on a per user and databank level ensuring data security when required.
In April 1996, the complete sequence of the yeast genome has been released. Now, the systematic analysis of an entire eukaryotic genome is possible. To allow for a global view on the genomic data and for the visualization of sequence homologies within a whole genome, the Genomebrowser has been developed. Furthermore, a structural characterization of the yeast proteins using sequence comparison and structure prediction techniques has been performed. The complete sequence of the yeast genome associated with biological information is available through the W W W. In addition, an intranet solution has been developed to allow the detailed inspection of the Yeast Sequence Database. The CD-ROM includes the Genomebrowser, the functional classification of the open reading frames, aligned protein families and other services based on Java applications and the Netscape browser.
The Comprehensive Yeast Genome Database (CYGD) compiles a comprehensive data resource for information on the cellular functions of the yeast Saccharomyces cerevisiae and related species, chosen as the best understood model organism for eukaryotes. The database serves as a common resource generated by a European consortium, going beyond the provision of sequence information and functional annotations on individual genes and proteins. In addition, it provides information on the physical and functional interactions among proteins as well as other genetic elements. These cellular networks include metabolic and regulatory pathways, signal transduction and transport processes as well as co-regulated gene clusters. As more yeast genomes are published, their annotation becomes greatly facilitated using S.cerevisiae as a reference. CYGD provides a way of exploring related genomes with the aid of the S.cerevisiae genome as a backbone and SIMAP, the Similarity Matrix of Proteins. The comprehensive resource is available under http://mips.gsf.de/genre/proj/yeast/.
The PEDANT genome database (http://pedant.gsf.de) provides exhaustive automatic analysis of genomic sequences by a large variety of established bioinformatics tools through a comprehensive Web-based user interface. One hundred and seventy seven completely sequenced and unfinished genomes have been processed so far, including large eukaryotic genomes (mouse, human) published recently. In this contribution, we describe the current status of the PEDANT database and novel analytical features added to the PEDANT server in 2002. Those include: (i) integration with the BioRS data retrieval system which allows fast text queries, (ii) pre-computed sequence clusters in each complete genome, (iii) a comprehensive set of tools for genome comparison, including genome comparison tables and protein function prediction based on genomic context, and (iv) computation and visualization of protein-protein interaction (PPI) networks based on experimental data. The availability of functional and structural predictions for 650 000 genomic proteins in well organized form makes PEDANT a useful resource for both functional and structural genomics.
The review begins by providing a brief typology of biological databases on the Internet, illustrated by examples of the most influential resources of each kind. We then take an insider look at one typical on-line genomic resource -- the yeast genome database hosted at the Munich Information Center for Protein Sequences (MIPS) -- and explain how and why it has evolved from a basic sequence repository to a multidomain knowledge base. The role of community efforts in curating and annotating genome data is discussed. The crucial role of data integration and interoperability in developing next-generation genomic facilities is underscored.
The Munich Information Center for Protein Sequences (MIPS-GSF), Martinsried near Munich, Germany, develops and maintains genome oriented databases. It is commonplace that the amount of sequence data available increases rapidly, but not the capacity of qualified manual annotation at the sequence databases. Therefore, our strategy aims to cope with the data stream by the comprehensive application of analysis tools to sequences of complete genomes, the systematic classification of protein sequences and the active support of sequence analysis and functional genomics projects. This report describes the systematic and up-to-date analysis of genomes (PEDANT), a comprehensive database of the yeast genome (MYGD), a database reflecting the progress in sequencing the Arabidopsis thaliana genome (MATD), the database of assembled, annotated human EST clusters (MEST), and the collection of protein sequence data within the framework of the PIR-International Protein Sequence Database (described elsewhere in this volume). MIPS provides access through its WWW server (http://www.mips.biochem.mpg.de) to a spectrum of generic databases, including the above mentioned as well as a database of protein families (PROTFAM), the MITOP database, and the all-against-all FASTA database.
Von links nach rechts: H.-W. Mewes, A. Kaps, K. Heumann und A. Maierl. Dr. Hans-Werner Mewes, Arbeitsgruppenleiter von MIPS, studierte Chemie an der Philipps-Universität Marburg und promovierte über die Identifizierung von Proteinen durch computergestützte Aminosäurenanalyse. Seit Beginn der Achtziger Jahre beschäftigt er sich mit Konzepten der Informatik und ihrer Anwendung auf biologische Problemstellungen. Dipl.-Inform. Andreas Kaps, Dipl.-Inform. Klaus Heumann und Dipl.-Inform. Andreas Maierl studierten Informatik an der Technischen Universität München. Sie sind als wissenschaftliche Mitarbeiter bei MIPS beschäftigt. K. Heumann beschäftigt sich mit Algorithmen und Datenstrukturen zur biologischen Sequenzdatenanalyse, A. Kaps mit Geschäftsprozeßmodellierung und Workflow-Management im Bereich biologischer Sequenzdatenbanken sowie verteilten Anwendungen, und A. Maierl mit objektorientierter Analyse und Design sowie objektorientierten Datenbanken.
This chapter contains sections titled: Introduction Concept Components of the service layer Integration variants of the gateway layer Types of linking Synchronization of Databases Discussion and outlook