New molecular resources regarding the so-called “non-standard models” in biology extend the present knowledge and are essential for molecular evolution and diversity studies (especially during the development) and evolutionary inferences about these zoological groups, or more practically for their fruitful management. Sepia officinalis, an economically important cephalopod species, is emerging as a new lophotrochozoan developmental model. We developed a large set of expressed sequence tags (ESTs) from embryonic stages of S. officinalis, yielding 19,780 non-redundant sequences (NRS). Around 75% of these sequences have no homologs in existing available databases. This set is the first developmental ESTs library in cephalopods. By exploring these NRS for tubulin, a generic protein family, and reflectin, a cephalopod specific protein family,we point out for both families a striking molecular diversity in S. officinalis.
Metagenomics aims at exploring microbial communities concerning their composition and functioning. Application of high-throughput sequencing technologies for the analysis of environmental DNA-preparations can generate large sets of metagenome sequence data which have to be analyzed by means of bioinformatics tools to unveil the taxonomic composition of the analyzed community as well as the repertoire of genes and gene functions. A bioinformatics software platform is required that allows the automated taxonomic and functional analysis and interpretation of metagenome datasets without manual effort. To address current demands in metagenome data analyses, the novel platform MetaSAMS was developed. MetaSAMS automatically accomplishes the tasks necessary for analyzing the composition and functional repertoire of a given microbial community from metagenome sequence data by implementing two software pipelines: (i) the first pipeline consists of three different classifiers performing the taxonomic profiling of metagenome sequences and (ii) the second functional pipeline accomplishes region predictions on assembled contigs and assigns functional information to predicted coding sequences. Moreover, MetaSAMS provides tools for statistical and comparative analyses based on the taxonomic and functional annotations. The capabilities of MetaSAMS are demonstrated for two metagenome datasets obtained from a biogas-producing microbial community of a production-scale biogas plant. The MetaSAMS web interface is available at https://metasams.cebitec.uni-bielefeld.de.
In order to improve the genetic characterisation of the barnacle Balanus amphitrite, normalised EST libraries for the developmental stages, viz. nauplius (a mix of instars I and II), cyprid and adult, were generated. The libraries were sequenced independently using 454 technologies and 575,666 reads were generated. For adults, 4843 unique isotigs were estimated and 6754 and 7506 in the cyprid and naupliar stage, respectively. It was found that some of the previously proposed cyprid-specific bcs genes were also expressed during the naupliar and adult stage. Furthermore, as lectins have been hypothesised to influence settlement cue recognition in barnacles, the database was searched for lectin-like isotigs. Two proteins, uniquely expressed in either the cyprid or the adult stage, matched a mannose receptor, and their nucleotide sequences were 33% and 31% identical to a lectin (BRA-3) isolated from Megabalanus rosa. Further characterisation of these genes may suggest their involvement in settlement.
The pyrosequencing technology from 454 Life Sciences and a novel assembly approach for cDNA sequences with the Newbler Assembler were used to achieve a major step forward to unravel the transcriptome of Chinese hamster ovary (CHO) cells. Normalized cDNA libraries originating from several cell lines and diverse culture conditions were sequenced and the resulting 1.84 million reads were assembled into 32,801 contiguous sequences, 29,184 isotigs, and 24,576 isogroups. A taxonomic classification of the isotigs showed that more than 70% of the assembled data is most similar to the transcriptome of Mus musculus, with most of the remaining isotigs being homologous to DNA sequences from Rattus norvegicus. Mapping of the CHO cell line contigs to the mouse transcriptome demonstrated that 9124 mouse transcripts, representing 6701 genes, are covered by more than 95% of their sequence length. Metabolic pathways of the central carbohydrate metabolism and biosynthesis routes of sugars used for protein N-glycosylation were reconstructed from the transcriptome data. All relevant genes representing major steps in the N-glycosylation pathway of CHO cells were detected. The present manuscript represents a data set of assembled and annotated genes for CHO cells that can now be used for a detailed analysis of the molecular functioning of CHO cell lines.
Biogas production from renewable resources is attracting increased attention as an alternative energy source due to the limited availability of traditional fossil fuels. Many countries are promoting the use of alternative energy sources for sustainable energy production. In this study, a metagenome from a production-scale biogas fermenter was analysed employing Roche's GS FLX Titanium technology and compared to a previous dataset obtained from the same community DNA sample that was sequenced on the GS FLX platform. Taxonomic profiling based on 16S rRNA-specific sequences and an Environmental Gene Tag (EGT) analysis employing CARMA demonstrated that both approaches benefit from the longer read lengths obtained on the Titanium platform. Results confirmed Clostridia as the most prevalent taxonomic class, whereas species of the order Methanomicrobiales are dominant among methanogenic Archaea. However, the analyses also identified additional taxa that were missed by the previous study, including members of the genera Streptococcus, Acetivibrio, Garciella, Tissierella, and Gelria, which might also play a role in the fermentation process leading to the formation of methane. Taking advantage of the CARMA feature to correlate taxonomic information of sequences with their assigned functions, it appeared that Firmicutes, followed by Bacteroidetes and Proteobacteria, dominate within the functional context of polysaccharide degradation whereas Methanomicrobiales represent the most abundant taxonomic group responsible for methane production. Clostridia is the most important class involved in the reductive CoA pathway (Wood-Ljungdahl pathway) that is characteristic for acetogenesis. Based on binning of 16S rRNA-specific sequences allocated to the dominant genus Methanoculleus, it could be shown that this genus is represented by several different species. Phylogenetic analysis of these sequences placed them in close proximity to the hydrogenotrophic methanogen Methanoculleus bourgensis. While rarefaction analyses still indicate incomplete coverage, examination of the GS FLX Titanium dataset resulted in the identification of additional genera and functional elements, providing a far more complete coverage of the community involved in anaerobic fermentative pathways leading to methane formation.
Isolates of the symbiotic nitrogen-fixing species Sinorhizobium meliloti usually contain a chromosome and two large megaplasmids encoding functions that are absolutely required for the specific interaction of the microsymbiont with corresponding host plants leading to an effective symbiosis. The complete genome sequence, including the megaplasmids pSmeSM11c (related to pSymA) and pSmeSM11d (related to pSymB), was established for the dominant, indigenous S. meliloti strain SM11 that had been isolated during a long-term field release experiment with genetically modified S. meliloti strains. The chromosome, the largest replicon of S. meliloti SM11, is 3,908,022bp in size and codes for 3785 predicted protein coding sequences. The size of megaplasmid pSmeSM11c is 1,633,319bp and it contains 1760 predicted protein coding sequences whereas megaplasmid pSmeSM11d is 1,632,395bp in size and comprises 1548 predicted coding sequences. The gene content of the SM11 chromosome is quite similar to that of the reference strain S. meliloti Rm1021. Comparison of pSmeSM11c to pSymA of the reference strain revealed that many gene regions of these replicons are variable, supporting the assessment that pSymA is a major hot-spot for intra-specific differentiation. Plasmids pSymA and pSmeSM11c both encode unique genes. Large gene regions of pSmeSM11c are closely related to corresponding parts of Sinorhizobium medicae WSM419 plasmids. Moreover, pSmeSM11c encodes further novel gene regions, e.g. additional plasmid survival genes (partition, mobilisation and conjugative transfer genes), acdS encoding 1-aminocyclopropane-1-carboxylate deaminase involved in modulation of the phytohormone ethylene level and genes having predicted functions in degradative capabilities, stress response, amino acid metabolism and associated pathways. In contrast to Rm1021 pSymA and pSmeSM11c, megaplasmid pSymB of strain Rm1021 and pSmeSM11d are highly conserved showing extensive synteny with only few rearrangements. Most remarkably, pSmeSM11b contains a new gene cluster predicted to be involved in polysaccharide biosynthesis. Compilation of the S. meliloti SM11 genome sequence contributes to an extension of the S. meliloti pan-genome.
In recent years, modern high-throughput techniques in genome and post-genome research have made a marked impact on the marine sciences. Today, massively parallel DNA sequencing and hybridization approaches allow the identification of not only the gene repertoire but also the gene regulatory networks that function within an organism. The huge amounts of data acquired from such experiments can only be handled with intensive bioinformatics support that has to provide an adequate infrastructure for storing and analysing these data. Bioinformatics has to deliver efficient data analysis algorithms, user-friendly tools and software applications, as well as extensive hardware infrastructure to deal with these genome-scale analyses.The following chapter briefly introduces not only the most relevant topics of bioinformatics for functional and structural genomics but also addresses the practical aspects of other steps of a genome project such as sequencing or data management issues. The chapter will take the reader through the different technical approaches that can be applied in marine genomics projects.In the first part, we will mainly focus on data generation, introducing classical genome sequencing approaches such as the Sanger method and the shotgun technique. Moreover, a short overview of the current status of the next generation of sequencing techniques will be given. In the second part, we briefly introduce the concept of data management for bioinformatics applications. In the third part, we describe the basic principles of genome sequence analysis and address topics like EST clustering and genome assembly, gene prediction, gene function assignment and classification as well as whole genome annotation. In the fourth part of this chapter, we present an overview of transcriptome data analysis using microarray hybridization technology. After a brief introduction to microarray technology we describe state-of-the-art methods for image processing, data normalization, significance testing and cluster analysis.
The plant pathogenic basidiomycete Sclerotium rolfsii produces the industrially exploited exopolysaccharide scleroglucan, a polymer that consists of (1 → 3)-β-linked glucose with a (1 → 6)-β-glycosyl branch on every third unit. Although the physicochemical properties of scleroglucan are well understood, almost nothing is known about the genetics of scleroglucan biosynthesis. Similarly, the biosynthetic pathway of oxalate, the main by-product during scleroglucan production, has not been elucidated yet. In order to provide a basis for genetic and metabolic engineering approaches, we studied scleroglucan and oxalate biosynthesis in S. rolfsii using different transcriptomic approaches.
Background: Corynebacterium aurimucosum is a slightly yellowish, non-lipophilic, facultative anaerobic member of the genus Corynebacterium and predominantly isolated from human clinical specimens. Unusual black-pigmented variants of C. aurimucosum (originally named as C. nigricans) continue to be recovered from the female urogenital tract and they are associated with complications during pregnancy. C. aurimucosum ATCC 700975 (C. nigricans CN-1) was originally isolated from a vaginal swab of a 34-year-old woman who experienced a spontaneous abortion during month six of pregnancy. For a better understanding of the physiology and lifestyle of this potential urogenital pathogen, the complete genome sequence of C. aurimucosum ATCC 700975 was determined.Results: Sequencing and assembly of the C. aurimucosum ATCC 700975 genome yielded a circular chromosome of 2,790,189 bp in size and the 29,037-bp plasmid pET44827. Specific gene sets associated with the central metabolism of C. aurimucosum apparently provide enhanced metabolic flexibility and adaptability in aerobic, anaerobic and low-pH environments, including gene clusters for the uptake and degradation of aromatic amines, L-histidine and L-tartrate as well as a gene region for the formation of selenocysteine and its incorporation into formate dehydrogenase. Plasmid pET44827 codes for a non-ribosomal peptide synthetase that plays the pivotal role in the synthesis of the characteristic black pigment of C. aurimucosum ATCC 700975.Conclusions: The data obtained by the genome project suggest that C. aurimucosum could be both a resident of the human gut and possibly a pathogen in the female genital tract causing complications during pregnancy. Since hitherto all black-pigmented C. aurimucosum strains have been recovered from female genital source, biosynthesis of the pigment is apparently required for colonization by protecting the bacterial cells against the high hydrogen peroxide concentration in the vaginal environment. The location of the corresponding genes on plasmid pET44827 explains why black-pigmented (formerly C. nigricans) and non-pigmented C. aurimucosum strains were isolated from clinical specimens.
BACKGROUND:Databases for either sequence, annotation, or microarray experiments data are extremely beneficial to the research community, as they centrally gather information from experiments performed by different scientists. However, data from different sources develop their full capacities only when combined. The idea of a data warehouse directly adresses this problem and solves it by integrating all required data into one single database - hence there are already many data warehouses available to genetics. For the model legume Medicago truncatula, there is currently no such single data warehouse that integrates all freely available gene sequences, the corresponding gene expression data, and annotation information. Thus, we created the data warehouse TRUNCATULIX, an integrative database of Medicago truncatula sequence and expression data.RESULTS:The TRUNCATULIX data warehouse integrates five public databases for gene sequences, and gene annotations, as well as a database for microarray expression data covering raw data, normalized datasets, and complete expression profiling experiments. It can be accessed via an AJAX-based web interface using a standard web browser. For the first time, users can now quickly search for specific genes and gene expression data in a huge database based on high-quality annotations. The results can be exported as Excel, HTML, or as csv files for further usage.CONCLUSION:The integration of sequence, annotation, and gene expression data from several Medicago truncatula databases in TRUNCATULIX provides the legume community with access to data and data mining capability not previously available. TRUNCATULIX is freely available at http://www.cebitec.uni-bielefeld.de/truncatulix/.
BACKGROUND:The rapid progress of post-genomic analyses, such as transcriptomics, proteomics, and metabolomics has resulted in the generation of large amounts of quantitative data covering and connecting the complete cascade from genotype to phenotype for individual organisms. Various benefits can be achieved when these "Omics" data are integrated, such as the identification of unknown gene functions or the elucidation of regulatory networks of whole organisms. In order to be able to obtain deeper insights in the generated datasets, it is of utmost importance to present the data to the researcher in an intuitive, integrated, and knowledge-based environment. Therefore, various visualization paradigms have been established during the last years. The visualization of "Omics" data using metabolic pathway maps is intuitive and has been applied in various software tools. It has become obvious that the application of web-based and user driven software tools has great potential and benefits from the use of open and standardized formats for the description of pathways.RESULTS:In order to combine datasets from heterogeneous "Omics" sources, we present the web-based ProMeTra system that visualizes and combines datasets from transcriptomics, proteomics, and metabolomics on user defined metabolic pathway maps. Therefore, structured exchange of data with our "Omics" applications Emma 2, Qupe and MeltDB is employed. Enriched SVG images or animations are generated and can be obtained via the user friendly web interface. To demonstrate the functionality of ProMeTra, we use quantitative data obtained during a fermentation experiment of the L-lysine producing strain Corynebacterium glutamicum DM1730. During fermentation, oxygen supply was switched off in order to perturb the system and observe its reaction. At six different time points, transcript abundances, intracellular metabolite pools, as well as extracellular glucose, lactate, and L-lysine levels were determined.CONCLUSION:The interpretation and visualization of the results of this complex experiment was facilitated by the ProMeTra software. Both transcriptome and metabolome data were visualized on a metabolic pathway map. Visual inspection of the combined data confirmed existing knowledge but also delivered novel correlations that are of potential biotechnological importance.