The transcriptome and proteome change dynamically as cells respond to environmental stress; however, prior proteomic studies reported poor correlation between mRNA and protein, rendering their relationships unclear. To address this, we combined high mass accuracy mass spectrometry with isobaric tagging to quantify dynamic changes in ~2500 Saccharomyces cerevisiae proteins, in biological triplicate and with paired mRNA samples, as cells acclimated to high osmolarity. Surprisingly, while transcript induction correlated extremely well with protein increase, transcript reduction produced little to no change in the corresponding proteins. We constructed a mathematical model of dynamic protein changes and propose that the lack of protein reduction is explained by cell-division arrest, while transcript reduction supports redistribution of translational machinery. Furthermore, the transient 'burst' of mRNA induction after stress serves to accelerate change in the corresponding protein levels. We identified several classes of post-transcriptional regulation, but show that most of the variance in protein changes is explained by mRNA. Our results present a picture of the coordinated physiological responses at the levels of mRNA, protein, protein-synthetic capacity, and cellular growth.
Determining the underlying regulatory mechanism of genetic networks is one of the central challenges of computational biology. Numerous methods have been developed and applied to the important but complex task of reverse engineering regulatory networks from high-throughput gene expression data. However, many challenges remain. In this paper, we are interested in learning rules that will reveal the causal genes for the expression variation from various relational data sources in addition to gene expression data. Following our previous work where we showed that time series gene expression data could potentially uncover causal effects, we describe an application of an inductive logic programming (ILP) system, to the task of identifying important regulatory relationships from discretized time series gene expression data, protein-protein interaction, protein phosphorylation and transcription factor data about the organism. Specifically, we learn rules for predicting gene expression levels at the next time step based on the available relational data and then generalize the learned theory to visualize a pruned network of important interactions. We evaluate and present experimental results on microarray experiments from Gasch et alon Saccharomyces cerevisiae.
In this era where healthcare is one of the world's largest and fastest growing industries, there is great interest in understanding cells at the molecular level. Fortunately, innovations in experimental technology continue to spur the quantity and types of high-throughput biological data that can be measured. Analyzing the combination of these data sets could lead to novel discoveries in biology and medicine. The main focus of this thesis is the investigation of machine learning techniques for inferring gene regulatory networks from the combination of high-throughput time series gene expression array data and other data sources. The main contribution of this thesis is to computational biology. We exploit background knowledge and temporal information to determine causality in gene regulatory networks using a dynamic Bayesian network. We further combine other sources of relational data with temporal data and use an inductive learning method known as inductive logic programming (ILP) to learn rules for predicting gene expression. We also constructed a simplified theoretical model for the modeling of time series gene expression data to guide experimental decisions. A minor contribution of this thesis is to data mining. We show how we integrate an ILP system, FOIL, with a database and utilize statistics to avoid expensive database operations in certain cases. Another contribution involves a different ILP system, Aleph, and the use of pathfinding to search for rules that link the entities of interest. These extensions in ILP can also be useful for computational biology as ILP is often applied to diverse biological data because of its ability to accommodate multiple tables, take into account background knowledge, and produce rules comprehensible to biologists.
Jan Struyf合作论文数Declarative Languages and Artificial Intelligence research group1