Genome-scale metabolic models (GEMs) are used extensively for analysis of mechanisms underlying human diseases and metabolic malfunctions. However, the lack of comprehensive and high-quality GEMs for model organisms restricts translational utilization of omics data accumulating from the use of various disease models. Here we present a unified platform of GEMs that covers five major model animals, including Mouse1 (Mus musculus), Rat1 (Rattus norvegicus), Zebrafish1 (Danio rerio), Fruitfly1 (Drosophila melanogaster), and Worm1 (Caenorhabditis elegans). These GEMs represent the most comprehensive coverage of the metabolic network by considering both orthology-based pathways and species-specific reactions. All GEMs can be interactively queried via the accompanying web portal Metabolic Atlas. Specifically, through integrative analysis of Mouse1 with RNA-sequencing data from brain tissues of transgenic mice we identified a coordinated up-regulation of lysosomal GM2 ganglioside and peptide degradation pathways which appears to be a signature metabolic alteration in Alzheimer's disease (AD) mouse models with a phenotype of amyloid precursor protein overexpression. This metabolic shift was further validated with proteomics data from transgenic mice and cerebrospinal fluid samples from human patients. The elevated lysosomal enzymes thus hold potential to be used as a biomarker for early diagnosis of AD. Taken together, we foresee that this evolving open-source platform will serve as an important resource to facilitate the development of systems medicines and translational biomedical applications.
Genome-scale metabolic models (GEMs) are valuable tools to study metabolism and provide a scaffold for the integrative analysis of omics data. Researchers have developed increasingly comprehensive human GEMs, but the disconnect among different model sources and versions impedes further progress. We therefore integrated and extensively curated the most recent human metabolic models to construct a consensus GEM, Human1. We demonstrated the versatility of Human1 through the generation and analysis of cell- and tissue-specific models using transcriptomic, proteomic, and kinetic data. We also present an accompanying web portal, Metabolic Atlas (https://www.metabolicatlas.org/), which facilitates further exploration and visualization of Human1 content. Human1 was created using a version-controlled, open-source model development framework to enable community-driven curation and refinement. This framework allows Human1 to be an evolving shared resource for future studies of human health and disease.
Genome-scale metabolic models (GEMs) are valuable tools to study metabolism and provide a scaffold for the integrative analysis of omics data. GEMs often encompass thousands of reactions, metabolites and genes, whose manipulation require complex data structures. Metabolic Atlas, through the web platform available at https://metabolicatlas.org , aims to make the entire GEM content available for easy navigation. This is achieved through both tabular and map views (2D and 3D), each suited for different usage scenarios. In addition, Metabolic Atlas aims to meet the needs of the community through the development of specific tools and features though iterative releases. Currently, Metabolic Atlas facilitates exploration and visualization of two open-source GEMs: Human1, an integration and extensive curation of the most recent human metabolic models ( Robinson et al., 2020 ), and Yeast8, a consensus metabolic model for S. cerevisiae ( Lu et al., 2019 ). The history of Metabolic Atlas starts before 2015 ( Pornputtapong et al., 2015 ). The present website has been re-developed from the ground up, following open-source standards. It was made publicly available mid 2019, and is now at version 1.7 ( Robinson et al., 2020 ). We plan to continue the development in a tic-toc method: a major release altering the foundation of the website, followed by several smaller releases. We are actively engaging the community to address the challenges in accessing GEMs for curation, analysis and biological understanding, under the guidance of FAIR principles.
The enormous amount of freely accessible functional genomics data is an invaluable resource for interrogating the biological function of multiple DNA-interacting players and chromatin modifications by large-scale comparative analyses. However, in practice, interrogating large collections of public data requires major efforts for (i) reprocessing available raw reads, (ii) incorporating quality assessments to exclude artefactual and low-quality data, and (iii) processing data by using high-performance computation. Here, we present qcGenomics , a user-friendly online resource for ultrafast retrieval, visualization, and comparative analysis of tens of thousands of genomics datasets to gain new functional insight from global or focused multidimensional data integration.
Complex organisms originate from and are maintained by the information encoded in the genome. A major challenge of systems biology is to develop algorithms that describe the dynamic regulation of genome functions from large omics datasets. Here, we describe TETRAMER, which reconstructs gene-regulatory networks from temporal transcriptome data during cell fate transitions to predict "master" regulators by simulating cascades of temporal transcription-regulatory events.
BACKGROUND:Exponentially increasing numbers of NGS-based epigenomic datasets in public repositories like GEO constitute an enormous source of information that is invaluable for integrative and comparative studies of gene regulatory mechanisms. One of today's challenges for such studies is to identify functionally informative local and global patterns of chromatin states in order to describe the regulatory impact of the epigenome in normal cell physiology and in case of pathological aberrations. Critically, the most preferred Chromatin ImmunoPrecipitation-Sequencing (ChIP-Seq) is inherently prone to significant variability between assays, which poses significant challenge on comparative studies. One challenge concerns data normalization to adjust sequencing depth variation.RESULTS:Currently existing tools either apply linear scaling corrections and/or are restricted to specific genomic regions, which can be prone to biases. To overcome these restrictions without any external biases, we developed Epimetheus, a genome-wide quantile-based multi-profile normalization tool for histone modification data and related datasets.CONCLUSIONS:Epimetheus has been successfully used to normalize epigenomics data in previous studies on X inactivation in breast cancer and in integrative studies of neuronal cell fate acquisition and tumorigenic transformation; Epimetheus is freely available to the scientific community.
Studying living organisms as an ensemble of components in which the whole is the consequence of the complexity of their interactions represents the biggest challenge of the current “big-data omics” era. Specifically since the release of the first draft of the human genome in 2001, followed by the rapid development of the massive parallel sequencing technologies, the avenue towards the analysis of genome functions from a holistic point of view has been opened. Importantly, the combination of a multiplicity of genomic readouts will provide means to describe living systems through the reconstitution of their genomic-regulatory functions which are at the basis of their defined state. Moreover, understanding the reorganization of their regulatory wires – as a consequence of external/internal cues – represents a new approach to interpret the acquisition of novel physiological or aberrant system states. In a cellular context, the detailed comprehension of these reorganizations, known as cell fate transitions , is a major component of the novel therapeutic developments in regenerative medicine. In our laboratory, we take advantage of retinoic acid (RA)-driven cell differentiation model systems as a systematic approach for enhancing our understanding of cell fate transitions. Specifically it is based on (i) the integration of temporal functional genomic readouts for the reconstruction of gene regulatory networks (GRNs) describing the regulatory principles taken place during the cell fate transition process; (ii) the use of computational approaches for modelling signal transduction propagation over the reconstructed GRN; and finally (iii) the validation of the prediction readouts by taken advantage of the current genome editing approaches (CRISPR). Here we present our implemented signal transduction model able to verify the coherence of the reconstructed GRN with the temporal transcriptional information describing the cell fate transition; but also predict the capacity of any nodes composing the GRN to drive the expected cell-fate transition behavior. This methodology mimics signalling propagation over multiple temporal transcriptional response layers by (i) verifying a change in the temporal transcriptional state between interconnected nodes; (ii) a coherent temporal directionality in the context of TF-TG interactions; and (iii) a signal propagation interconnections derived only from the initial signalling cue. The presented approach has been challenged over large number of nodes, and complex interconnected GRNs describing processes like cell differentiation, tumorigenesis and cell fate reprogramming. Following the strong prediction power – verified by the use of CRISPR-dCas9 activation assays – observed in our RA-driven cell differentiation studies, we are currently working in a second version able to incorporate in-silico knock-out modelling, but also combinatorial nodes activation assays in order to predict cooperative situations able to enhance signal transduction performance. Taken in consideration the interest of the scientific community to access to this type of signal transduction modelling instruments, we are currently preparing a Cytoscape app, such that users could benefit of the multiple options available in such environment.
The combination of massive parallel sequencing with a variety of modern DNA/RNA enrichment technologies provides means for interrogating functional protein-genome interactions (ChIP-seq), genome-wide transcriptional activity (RNA-seq; GRO-seq), chromatin accessibility (DNase-seq, FAIRE-seq, MNase-seq), and more recently the three-dimensional organization of chromatin (Hi-C, ChIA-PET). In systems biology-based approaches several of these readouts are generally cumulated with the aim of describing living systems through a reconstitution of the genome-regulatory functions. However, an issue that is often underestimated is that conclusions drawn from such multidimensional analyses of NGS-derived datasets critically depend on the quality of the compared datasets. To address this problem, we have developed the NGS-QC Generator, a quality control system that infers quality descriptors for any kind of ChIP-sequencing and related datasets. In this chapter we provide a detailed protocol for (1) assessing quality descriptors with the NGS-QC Generator; (2) to interpret the generated reports; and (3) to explore the database of QC indicators (www.ngs-qc.org) for >21,000 publicly available datasets.
Background: Proximity ligation-mediated methods are essential to study the impact of three-dimensional chromatin organization on gene programming. Albeit significant progress has been made in the development of computational tools that assess long-range chromatin interactions, next to nothing is known about the quality of the generated datasets.Method: We have developed LOGIQA (www.ngs-qc.org/logiqa), a database hosting quality scores for long-range genome interaction assays, accessible through a user-friendly web-based environment.Results: Currently, LOGIQA harbors QC scores for >900 datasets, which provides a global view of their relative quality and reveals the impact of genome size, coverage and other technical aspects. LOGIQA provides a user-friendly dataset query panel and a genome viewer to assess local genome-interaction maps at different resolution and quality-assessment conditions.Conclusions: LOGIQA is the first database hosting quality scores dedicated to long-range chromatin interaction assays, which in addition provides a platform for visualizing genome interactions made available by the scientific community.
We have established a certification system for antibodies to be used in chromatin immunoprecipitation assays coupled to massive parallel sequencing (ChIP-seq). This certification comprises a standardized ChIP procedure and the attribution of a numerical quality control indicator (QCi) to biological replicate experiments. The QCi computation is based on a universally applicable quality assessment that quantitates the global deviation of randomly sampled subsets of ChIP-seq dataset with the original genome-aligned sequence reads. Comparison with a QCi database for >28,000 ChIP-seq assays were used to attribute quality grades (ranging from ‘AAA’ to ‘DDD’) to a given dataset. In the present report we used the numerical QC system to assess the factors influencing the quality of ChIP-seq assays, including the nature of the target, the sequencing depth and the commercial source of the antibody. We have used this approach specifically to certify mono and polyclonal antibodies obtained from Active Motif directed against the histone modification marks H3K4me3, H3K27ac and H3K9ac for ChIP-seq. The antibodies received the grades AAA to BBC (www.ngs-qc.org). We propose to attribute such quantitative grading of all antibodies attributed with the label “ChIP-seq grade”.
Acid mine drainage (AMD) is a highly toxic environment for most living organisms due to the presence of many lethal elements including arsenic (As). Thiomonas (Tm.) bacteria are found ubiquitously in AMD and can withstand these extreme conditions, in part because they are able to oxidize arsenite. In order to further improve our knowledge concerning the adaptive capacities of these bacteria, we sequenced and assembled the genome of six isolates derived from the Carnoulès AMD, and compared them to the genomes of Tm. arsenitoxydans 3As (isolated from the same site) and Tm. intermedia K12 (isolated from a sewage pipe). A detailed analysis of the Tm. sp. CB2 genome revealed various rearrangements had occurred in comparison to what was observed in 3As and K12 and over 20 genomic islands (GEIs) were found in each of these three genomes. We performed a detailed comparison of the two arsenic-related islands found in CB2, carrying the genes required for arsenite oxidation and As resistance, with those found in K12, 3As, and five other Thiomonas strains also isolated from Carnoulès (CB1, CB3, CB6, ACO3 and ACO7). Our results suggest that these arsenic-related islands have evolved differentially in these closely related Thiomonas strains, leading to divergent capacities to survive in As rich environments.