Journal of Data Mining in Genomics & Proteomics(2013)
被引用23|浏览54
摘要
The major task of the informatics work was to design and implement a complete data management solution for the outcomes of molecular phenotyping experiments (transcriptomics, proteomics, metabonomics, and others).A 2-tier system was built, reflecting the two main stages of the data management lifecycle (Figure 1):1. Information submission, storage and unification, facilitated by data repository components and the reannotation system; 2. Data integration, search and retrieval, facilitated by the MoDa warehouse interface.The information system for data submissions and storage, SIMBioMS [1], includes two components interlinked through identifiers: Sample Information Management System (SIMS) and Assay Information Management System (AIMS).These are two databases with associated web-based submission tools, enabling the collection of sample information (SIMS) and experimental results and metadata (AIMS).Apart from data submission these systems also support sample and assay metadata search and export, as well as data file exchange.The system for sample information management, SIMS, has been described separately in Viksna et al. [2] (referred to as "Sample Management Database" in that article).A critical and often costly aspect of any data warehousing effort is data cleansing and transformation to ensure that the data that can be queried through the warehouse is consistent.After experiment results of a study have been collected in AIMS, all supporting metaand sample data can be exported from SIMS and AIMS, the data files parsed, and the data points linked to the metadata and transferred to the data warehouse.As a part of data processing, biomolecular entities (genes, proteins, metabolites) are passed through the reannotation system, mapping the measured molecular entities (transcripts, proteins, metabolites etc) against a uniform reference system.