
Background: Systematic fusion of multiple data sources for Gene Regulatory Networks (GRN) inference remains a key challenge in systems biology. We incorporate information from protein-protein interaction networks (PPIN) into the process of GRN inference from gene expression (GE) data. However, existing PPIN remain sparse and transitive protein interactions can help predict missing protein interactions. We therefore propose a systematic probabilistic framework on fusing GE data and transitive protein interaction data to coherently build GRN. Results: We use a Gaussian Mixture Model (GMM) to soft-cluster GE data, allowing overlapping cluster memberships. Next, a heuristic method is proposed to extend sparse PPIN by incorporating transitive linkages. We then propose a novel way to score extended protein interactions by combining topological properties of PPIN and correlations of GE. Following this, GE data and extended PPIN are fused using a Gaussian Hidden Markov Model (GHMM) in order to identify gene regulatory pathways and refine interaction scores that are then used to constrain the GRN structure. We employ a Bayesian Gaussian Mixture (BGM) model to refine the GRN derived from GE data by using the structural priors derived from GHMM. Experiments on real yeast regulatory networks demonstrate both the feasibility of the extended PPIN in predicting transitive protein interactions and its effectiveness on improving the coverage and accuracy the proposed method of fusing PPIN and GE to build GRN. Conclusion: The GE and PPIN fusion model outperforms both the state-of-the-art single data source models (CLR, GENIE3, TIGRESS) as well as existing fusion models under various constraints.
BACKGROUND:One of the questions in the design of cancer clinical trials with combination of two drugs is in which order to administer the drugs. This is an important question, especially in the case where one agent may interfere with the effectiveness of the other agent. RESULTS:In the present paper we develop a mathematical model to address this scheduling question in a specific case where one of the drugs is anti-VEGF, which is known to affect the perfusion of other drugs. As a second drug we take anti-PD-1. Both drugs are known to increase the activation of anticancer T cells. Our simulations show that in the case where anti-VEGF reduces the perfusion, a non-overlapping schedule is significantly more effective than a simultaneous injection of the two drugs, and it is somewhat more beneficial to inject anti-PD-1 first. CONCLUSION:The method and results of the paper can be extended to other combinations, and they could play an important role in the design of clinical trials with combination therapy, where scheduling strategies may significantly affect the outcome.
BACKGROUND:Genome-scale models of metabolism and macromolecular expression (ME models) enable systems-level computation of proteome allocation coupled to metabolic phenotype.RESULTS:We develop DynamicME, an algorithm enabling time-course simulation of cell metabolism and protein expression. DynamicME correctly predicted the substrate utilization hierarchy on a mixed carbon substrate medium. We also found good agreement between predicted and measured time-course expression profiles. ME models involve considerably more parameters than metabolic models (M models). We thus generate an ensemble of models (each model having its rate constants perturbed), and then analyze the models by identifying archetypal time-course metabolite concentration profiles. Furthermore, we use a metaheuristic optimization method to calibrate ME model parameters using time-course measurements such as from a (fed-) batch culture. Finally, we show that constraints on protein concentration dynamics ("inertia") alter the metabolic response to environmental fluctuations, including increased substrate-level phosphorylation and lowered oxidative phosphorylation.CONCLUSIONS:Overall, DynamicME provides a novel method for understanding proteome allocation and metabolism under complex and transient environments, and to utilize time-course cell culture data for model-based interpretation or model refinement.
BACKGROUND:The topological landscape of gene interaction networks provides a rich source of information for inferring functional patterns of genes or proteins. However, it is still a challenging task to aggregate heterogeneous biological information such as gene expression and gene interactions to achieve more accurate inference for prediction and discovery of new gene interactions. In particular, how to generate a unified vector representation to integrate diverse input data is a key challenge addressed here.RESULTS:We propose a scalable and robust deep learning framework to learn embedded representations to unify known gene interactions and gene expression for gene interaction predictions. These low- dimensional embeddings derive deeper insights into the structure of rapidly accumulating and diverse gene interaction networks and greatly simplify downstream modeling. We compare the predictive power of our deep embeddings to the strong baselines. The results suggest that our deep embeddings achieve significantly more accurate predictions. Moreover, a set of novel gene interaction predictions are validated by up-to-date literature-based database entries.CONCLUSION:The proposed model demonstrates the importance of integrating heterogeneous information about genes for gene network inference. GNE is freely available under the GNU General Public License and can be downloaded from GitHub ( https://github.com/kckishan/GNE ).
BACKGROUND:A recurrent problem in genome-scale metabolic models (GEMs) is to correctly represent lipids as biomass requirements, due to the numerous of possible combinations of individual lipid species and the corresponding lack of fully detailed data. In this study we present SLIMEr, a formalism for correctly representing lipid requirements in GEMs using commonly available experimental data.RESULTS:SLIMEr enhances a GEM with mathematical constructs where we Split Lipids Into Measurable Entities (SLIME reactions), in addition to constraints on both the lipid classes and the acyl chain distribution. By implementing SLIMEr on the consensus GEM of Saccharomyces cerevisiae, we can represent accurate amounts of lipid species, analyze the flexibility of the resulting distribution, and compute the energy costs of moving from one metabolic state to another.CONCLUSIONS:The approach shows potential for better understanding lipid metabolism in yeast under different conditions. SLIMEr is freely available at https://github.com/SysBioChalmers/SLIMEr .
BACKGROUND:There is little published regarding metabolism of Salinispora species. In continuation with efforts performed towards this goal, this study is focused on new insights into the metabolism of the three-identified species of Salinispora using constraints-based modeling. At present, only one manually curated genome-scale metabolic model (GSM) for Salinispora tropica strain CNB-440T has been built despite the role of Salinispora strains in drug discovery.RESULTS:Here, we updated, and expanded the scope of the model of Salinispora tropica CNB-440T, and GSMs were constructed for two sequenced type strains covering the three-identified species. We also constructed a Salinispora core model that contains the genes shared by 93 sequenced strains and a few non-conserved genes associated with essential reactions. The models predicted no auxotrophies for essential amino acids, which was corroborated experimentally using a defined minimal medium (DMM). Experimental observations suggest possible sulfur accumulation. The Core metabolic content shows that the biosynthesis of specialised metabolites is the less conserved subsystem. Sets of reactions were analyzed to explore the differences between the reconstructions. Unique reactions associated to each GSM were mainly due to genome sequence data except for the ST-CNB440 reconstruction. In this case, additional reactions were added from experimental evidence. This reveals that by reaction content the ST-CNB440 model is different from the other species models. The differences identified in reaction content between models gave rise to different functional predictions of essential nutrient usage by each species in DMM. Furthermore, models were used to evaluate in silico single gene knockouts under DMM and complex medium. Cluster analysis of these results shows that ST-CNB440, and SP-CNR114 models are more similar when considering predicted essential genes.CONCLUSIONS:Models were built for each of the three currently identified Salinispora species, and a core model representing the conserved metabolic capabilities of Salinispora was constructed. Models will allow in silico metabolism studies of Salinispora strains, and help researchers to guide and increase the production of specialised metabolites. Also, models can be used as templates to build GSMs models of closely related organisms with high biotechnology potential.
BACKGROUND:The JAK/STAT signaling pathway is involved in many aging-related cellular functions. However, effects of overexpression of genes controlling JAK/STAT signal transduction on longevity of model organisms have not been studied. Here we evaluate the effect of overexpression of the unpaired 1 (upd1) gene, which encodes an activating ligand for JAK/STAT pathway, on the lifespan of Drosophila melanogaster.RESULTS:Overexpression of upd1 in the intestine caused a pronounced shortening of the median lifespan by 54.1-18.9%, and the age of 90% mortality by 40.9-19.1% in males and females, respectively. In fat body and in nervous system of male flies, an induction of upd1 overexpression increased the age of 90% mortality and median lifespan, respectively. An increase in upd1 expression enhanced mRNA levels of the JAK/STAT target genes domeless and Socs36E.CONCLUSIONS:Conditional overexpression of upd1 in different tissues of Drosophila imago induces pro-aging or pro-longevity effects in tissue-dependent manner. The effects of upd1 overexpression on lifespan are accompanied by the transcription activation of genes for the components of JAK/STAT pathway.
BACKGROUND:Improving efficiency of disease diagnosis based on phenotype ontology is a critical yet challenging research area. Recently, Human Phenotype Ontology (HPO)-based semantic similarity has been affectively and widely used to identify causative genes and diseases. However, current phenotype similarity measurements just consider the annotations and hierarchy structure of HPO, neglecting the definition description of phenotype terms.RESULTS:In this paper, we propose a novel phenotype similarity measurement, termed as DisPheno, which adequately incorporates the definition of phenotype terms in addition to HPO structure and annotations to measure the similarity between phenotype terms. DisPheno also integrates phenotype term associations into phenotype-set similarity measurement using gene and disease annotations of phenotype terms.CONCLUSIONS:Compared with five existing state-of-the-art methods, DisPheno shows great performance in HPO-based phenotype semantic similarity measurement and improves the efficiency of disease identification, especially on noisy patients dataset.
BACKGROUND:Liver has the unique ability to regenerate following injury, with a wide range of variability of the regenerative response across individuals. Existing computational models of the liver regeneration are largely tuned based on rodent data and hence it is not clear how well these models capture the dynamics of human liver regeneration. Recent availability of human liver volumetry time series data has enabled new opportunities to tune the computational models for human-relevant time scales, and to predict factors that can significantly alter the dynamics of liver regeneration following a resection.METHODS:We utilized a mathematical model that integrates signaling mechanisms and cellular functional state transitions. We tuned the model parameters to match the time scale of human liver regeneration using an elastic net based regularization approach for identifying optimal parameter values. We initially examined the effect of each parameter individually on the response mode (normal, suppressed, failure) and extent of recovery to identify critical parameters. We employed phase plane analysis to compute the threshold of resection. We mapped the distribution of the response modes and threshold of resection in a virtual patient cohort generated in silico via simultaneous variations in two most critical parameters.RESULTS:Analysis of the responses to resection with individual parameter variations showed that the response mode and extent of recovery following resection were most sensitive to variations in two perioperative factors, metabolic load and cell death post partial hepatectomy. Phase plane analysis identified two steady states corresponding to recovery and failure, with a threshold of resection separating the two basins of attraction. The size of the basin of attraction for the recovery mode varied as a function of metabolic load and cell death sensitivity, leading to a change in the multiplicity of the system in response to changes in these two parameters.CONCLUSIONS:Our results suggest that the response mode and threshold of failure are critically dependent on the metabolic load and cell death sensitivity parameters that are likely to be patient-specific. Interventions that modulate these critical perioperative factors may be helpful to drive the liver regenerative response process towards a complete recovery mode.
It was highlighted that the original article [1] contained a typesetting error in the last name of Allon Canaan. This was incorrectly captured as Allon Canaann in the original article which has since been updated.
Identification of Hürthle cell cancers by non-operative fine-needle aspiration biopsy (FNAB) of thyroid nodules is challenging. Resultingly, non-cancerous Hürthle lesions were conventionally distinguished from Hürthle cell cancers by histopathological examination of tissue following surgical resection. Reliance on histopathological evaluation requires patients to undergo surgery to obtain a diagnosis despite most being non-cancerous. It is highly desirable to avoid surgery and to provide accurate classification of benignity versus malignancy from FNAB preoperatively. In our first-generation algorithm, Gene Expression Classifier (GEC), we achieved this goal by using machine learning (ML) on gene expression features. The classifier is sensitive, but not specific due in part to the presence of non-neoplastic benign Hürthle cells in many FNAB. We sought to overcome this low-specificity limitation by expanding the feature set for ML using next-generation whole transcriptome RNA sequencing and called the improved algorithm the Genomic Sequencing Classifier (GSC). The Hürthle identification leverages mitochondrial expression and we developed novel feature extraction mechanisms to measure chromosomal and genomic level loss-of-heterozygosity (LOH) for the algorithm. Additionally, we developed a multi-layered system of cascading classifiers to sequentially triage Hürthle cell-containing FNAB, including: 1. presence of Hürthle cells, 2. presence of neoplastic Hürthle cells, and 3. presence of benign Hürthle cells. The final Hürthle cell Index utilizes 1048 nuclear and mitochondrial genes; and Hürthle cell Neoplasm Index leverages LOH features as well as 2041 genes. Both indices are Support Vector Machine (SVM) based. The third classifier, the GSC Benign/Suspicious classifier, utilizes 1115 core genes and is an ensemble classifier incorporating 12 individual models. The accurate algorithmic depiction of this complex biological system among Hürthle subtypes results in a dramatic improvement of classification performance; specificity among Hürthle cell neoplasms increases from 11.8% with the GEC to 58.8% with the GSC, while maintaining the same sensitivity of 89%.
Anti-tumor necrosis factor alpha (TNF- α) therapy has made a significant impact on treating psoriasis. Despite these agents being designed to block TNF- α activity, their mechanism of action in the remission of psoriasis is still not fully understood at the molecular level. To better understand the molecular mechanisms of Anti-TNF- α therapy, we analysed the global gene expression profile (using mRNA microarray) in peripheral blood mononuclear cells (PBMCs) that were collected from 6 psoriasis patients before and 12 weeks after the treatment of etanercept. First, we identified 176 differentially expressed genes (DEGs) before and after treatment by using paired t-test. Then, we constructed the gene co-expression modules by weighted correlation network analysis (WGCNA), and 22 co-expression modules were found to be significantly correlated with treatment response. Of these 176 DEGs, 79 DEGs (M_DEGs) were the members of these 22 co-expression modules. Of the 287 GO functional processes and pathways that were enriched for these 79 M_DEGs, we identified 30 pathways whose overall gene expression activities were significantly correlated with treatment response. Of the original 176 DEGs, 19 (GO_DEGs) were found to be the members of these 30 pathways, whose expression profiles showed clear discrimination before and after treatment. As expected, of the biological processes and functionalities implicated by these 30 treatment response-related pathways, the inflammation and immune response was the top pathway in response to etanercept treatment, and some known TNF- α related pathways, such as molting cycle process, hair cycle process, skin epidermis development, regulation of hair follicle development, were implicated. Furthermore, additional novel pathways were also suggested, such as heparan sulfate proteoglycan metabolic process, vascular endothelial growth factor production, whose transcriptional regulation may mediate the response to etanercept treatment. Through global gene expression analysis in PBMC of psoriasis patient and subsequent co-expression module based pathway analyses, we have identified a group of functionally coherent and differentially expressed genes (DEGs) and related pathways, which has not only provided new biological insight about the molecular mechanism of anti-TNF- α treatment, but also identified several genes whose expression profiles can be used as potential biomarkers for anti-TNF- α treatment response in psoriasis.
Flow cytometry is a popular technology for quantitative single-cell profiling of cell surface markers. It enables expression measurement of tens of cell surface protein markers in millions of single cells. It is a powerful tool for discovering cell sub-populations and quantifying cell population heterogeneity. Traditionally, scientists use manual gating to identify cell types, but the process is subjective and is not effective for large multidimensional data. Many clustering algorithms have been developed to analyse these data but most of them are not scalable to very large data sets with more than ten million cells. Here, we present a new clustering algorithm that combines the advantages of density-based clustering algorithm DBSCAN with the scalability of grid-based clustering. This new clustering algorithm is implemented in python as an open source package, FlowGrid. FlowGrid is memory efficient and scales linearly with respect to the number of cells. We have evaluated the performance of FlowGrid against other state-of-the-art clustering programs and found that FlowGrid produces similar clustering results but with substantially less time. For example, FlowGrid is able to complete a clustering task on a data set of 23.6 million cells in less than 12 seconds, while other algorithms take more than 500 seconds or get into error. FlowGrid is an ultrafast clustering algorithm for large single-cell flow cytometry data. The source code is available at https://github.com/VCCRI/FlowGrid .
BACKGROUND:Single-cell RNA sequencing (scRNA-seq) technology provides an effective way to study cell heterogeneity. However, due to the low capture efficiency and stochastic gene expression, scRNA-seq data often contains a high percentage of missing values. It has been showed that the missing rate can reach approximately 30% even after noise reduction. To accurately recover missing values in scRNA-seq data, we need to know where the missing data is; how much data is missing; and what are the values of these data.METHODS:To solve these three problems, we propose a novel model with a hybrid machine learning method, namely, missing imputation for single-cell RNA-seq (MISC). To solve the first problem, we transformed it to a binary classification problem on the RNA-seq expression matrix. Then, for the second problem, we searched for the intersection of the classification results, zero-inflated model and false negative model results. Finally, we used the regression model to recover the data in the missing elements.RESULTS:We compared the raw data without imputation, the mean-smooth neighbor cell trajectory, MISC on chronic myeloid leukemia data (CML), the primary somatosensory cortex and the hippocampal CA1 region of mouse brain cells. On the CML data, MISC discovered a trajectory branch from the CP-CML to the BC-CML, which provides direct evidence of evolution from CP to BC stem cells. On the mouse brain data, MISC clearly divides the pyramidal CA1 into different branches, and it is direct evidence of pyramidal CA1 in the subpopulations. In the meantime, with MISC, the oligodendrocyte cells became an independent group with an apparent boundary.CONCLUSIONS:Our results showed that the MISC model improved the cell type classification and could be instrumental to study cellular heterogeneity. Overall, MISC is a robust missing data imputation model for single-cell RNA-seq data.
systems-computational-biology/) and at 10th International Young Scientists School (SBB-2018) (
BACKGROUND:Microscopic images are widely used in plant biology as an essential source of information on morphometric characteristics of the cells and the topological characteristics of cellular tissue pattern due to modern computer vision algorithms. High-resolution 3D confocal images allow extracting quantitative characteristics describing the cell structure of leaf epidermis. For some issues in the study of cereal leaves development, it is required to apply the staining techniques with fluorescent dyes and to scan rather large fragments consisting of several frames. We aimed to develop a tool for processing multi-frame multi-channel 3D images obtained from confocal laser scanning microscopy and taking into account the peculiarities of the cereal leaves staining.RESULTS:We elaborated an ImageJ-plugin LSM-W2 that allows extracting data on Leaf Surface Morphology from Laser Scanning Microscopy images. The plugin is a crucial link in a workflow for obtaining data on structural properties of leaf epidermis and morphological properties of epidermal cells. It allows converting large lsm-files (laser scanning microscopy) into segmented 2D/3D images or tables with data on cells and/or nuclei sizes. In the article, we also represent some case studies showing the plugin application for solving biological tasks. Namely the plugin is applied in the following cases: defining parameters of jigsaw-puzzle pattern for maize leaf epidermal cells, analysis of the pavement cells morphological parameters for the mature wheat leaf grown under control and water deficit conditions, initiation of cell longitudinal rows, and detection of guard mother cells emergence at the initial stages of the stomatal morphogenesis in the growth zone of a wheat leaf.CONCLUSION:The proposed plugin is efficient for high-throughput analysis of cellular architecture for cereal leaf epidermis. The workflow implies using inexpensive and rapid sample preparation and does not require the applying of transgenesis and reporter genetic structures expanding the range of species and varieties to study. Obtained characteristics of the cell structure and patterns further could act as a basis for the development and verification for spatial models of plant tissues formation mechanisms accounting for structural features of cereal leaves.AVAILABILITY:The implementation of this workflow is available as an ImageJ plugin distributed as a part of the Fiji project (FijiisjustImageJ: https://fiji.sc/ ). The plugin is freely available at https://imagej.net/LSM_Worker , https://github.com/JmanJ/LSM_Worker and http://pixie.bionet.nsc.ru/LSM_WORKER/ .
BACKGROUND:The Notch signaling pathway is involved in cell fate decision and developmental patterning in diverse organisms. A receptor molecule, Notch (N), and a ligand molecule (in this case Delta or Dl) are the central molecules in this pathway. In early Drosophila embryos, these molecules determine neural vs. skin fates in a reproducible rosette pattern.RESULTS:We have created an agent-based model (ABM) that simulates the molecular components for this signaling pathway as agents acting within a spatial representation of a cell. The model captures the changing levels of these components, their transition from one state to another, and their movement from the nucleus to the cell membrane and back to the nucleus again. The model introduces stochastic variation into the system using a random generator within the Netlogo programming environment. The model uses these representations to understand the biological systems at three levels: individual cell fate, the interactions between cells, and the formation of pattern across the system. Using a set of assessment tools, we show that the current model accurately reproduces the rosette pattern of neurons and skin cells in the system over a wide set of parameters. Oscillations in the level of the N agent eventually stabilize cell fate into this pattern. We found that the dynamic timing and the availability of the N and Dl agents in neighboring cells are central to the formation of a correct and stable pattern. A feedback loop to the production of both components is necessary for a correct and stable pattern.CONCLUSIONS:The signaling pathways within and between cells in our model interact in real time to create a spatially correct field of neurons and skin cells. This model predicts that cells with high N and low Dl drive the formation of the pattern. This model also be used to elucidate general rules of biological self-patterning and decision-making.
BACKGROUND:Psoriasis is a complex multi-factorial disease, involving both genetic susceptibilities and environmental triggers. Genome-wide association studies (GWAS) and epigenome-wide association studies (EWAS) have been carried out to identify genetic and epigenetic variants that are associated with psoriasis. However, these loci cannot fully explain the disease pathogenesis.METHODS:To achieve a comprehensive mechanistic understanding of psoriasis, we conducted a systems biology study, integrating multi-omics datasets including GWAS, EWAS, tissue-specific transcriptome, expression quantitative trait loci (eQTLs), gene networks, and biological pathways to identify the key genes, processes, and networks that are genetically and epigenetically associated with psoriasis risk.RESULTS:This integrative genomics study identified both well-characterized (e.g., the IL17 pathway in both GWAS and EWAS) and novel biological processes (e.g., the branched chain amino acid catabolism process in GWAS and the platelet and coagulation pathway in EWAS) involved in psoriasis. Finally, by utilizing tissue-specific gene regulatory networks, we unraveled the interactions among the psoriasis-associated genes and pathways in a tissue-specific manner and detected potential key regulatory genes in the psoriasis networks.CONCLUSIONS:The integration and convergence of multi-omics signals provide deeper and comprehensive insights into the biological mechanisms associated with psoriasis susceptibility.
It was highlighted that the original article [1] contained errors in the figures and their legends and by extension the in-text figure citations. This Corrections article shows the correct figures and correct figure legends.