
In this paper, a model that combining cell population and genetic regulation within a single cell by using stochastic hybrid systems is proposed. The objective is to study the response of a population of cancer cells to various drugs that targeting the proliferation and survival pathways. The proposed model captures both the dynamics of the cell population and the dynamics of gene regulations within each individual cell. We use drug Lapatinib applied to colon cancer cell line HCT-116 as an example to validate the proposed model. Simulation results demonstrate the phenomena that observed in TGen experiments.
Understanding the relationships between transcription factors (TFs) and genes in plants under abiotic stress responses, tolerance and adaptation to adverse environments is very important in developing resilient crop varieties. While experimental methods to characterize stress responsive TFs and their targets are highly accurate, identification and characterization of the role of a given gene in a given stress response event are often laborious and time consuming. Computational approaches, on the other hand, offer a platform to identify new knowledge by integrating high throughput omics data and mathematical methods/models. In this research, we have developed a generic linear model of transcriptional regulatory networks (TRNs) and a companion algorithm to identify and to characterize stress responsive genes and their roles in a given stress response event. The proposed methodology was applied to plants, by using Arabidopsis thaliana as an example, under abiotic stress. Well known interactions were inferred as well as putative novel ones that may play important roles in plants under abiotic stress conditions as confirmed by statistical and literature evidences.
The major challenge in reverse-engineering genetic regulatory networks is the small number of (time) measurements or experiments compared to the number of genes, which makes the system under-determined and hence unidentifiable. The only way to overcome the identifiability problem is to incorporate prior knowledge about the system. It is often assumed that genetic networks are sparse. In addition, if the measurements, in each experiment, present an unknown correlation structure, then the estimation problem becomes even more challenging. Estimating the covariance structure will improve the estimation of the network connectivity but will also make the estimation of the already under-determined problem even more challenging. In this paper, we formulate reverse-engineering genetic networks as a multiple linear regression problem. We show that, if the number of experiments is smaller than the number of genes and if the measurements present an unknown covariance structure, then the likelihood function diverges, making the maximum likelihood estimator senseless. We subsequently propose a normalized likelihood function that guarantees convergence while keeping the form of the Gaussian distribution. The optimal connectivity matrix is approximated as the solution of a convex optimization problem. Our simulation results show that the proposed maximum normalized-likelihood estimator outperforms the classical regularized maximum likelihood estimator, which assumes a known covariance structure.
Summary form only given. MicroRNAs (miRNAs) are short non-coding RNAs with the average length of 22 nucleotides. They are known to induce mRNA degradation or suppression of translation by complementarily binding to 3' untranslated regions (3' UTRs) of target mRNA transcripts. Recently, an alternative mechanism through which miRNAs participate in gene regulation was postulated and experimentally validated, namely the competing endogenous RNAs (ceRNAs). By competing for a limited pool of common targeting miRNAs (miRNA programs; miRP), pairs of genes (ceRNAs) sharing, fully or partially, identical miRNAs binding sites can “talk” to each other: when one ceRNA is up-regulated (or down-regulated) in cells, it attracts (or releases) the targeting miRNAs away from (or toward) the other ceRNA, and in turn have protective (or harmful) effects on expression of the other ceRNA. Based on in silico and in vitro analysis, recent reports suggested the dynamic and condition-specific properties of ceRNA regulation. The essential factors involved in ceRNA regulation include size of miRP, number of miRP binding sites, expression level of miRP, and expression level of ceRNAs. For better characterizing the optimal conditions for ceRNA regulation, in the present study we aim to confer how essential factors determine strength of ceRNA regulation in vivo, by analyzing TCGA datasets of glioblastoma multiforme (GBM) patients with 491 tumor samples profiled with paired miRNA and gene expression. Based on the definition that two genes sharing any number of common targeting miRNAs as a putative ceRNA pair, and by utilizing TargetScan algorithm, we identified 47,451,423 putative ceRNA pairs, involving 10,872 ceRNAs (genes). Pairwise correlation coefficients of gene expression profiles were then computed for each of the putative ceRNA pairs, and then the CDF. Varying size of miRP, for example, generated multiple CDFs, and then the goodness-of-fit was performed for pinpointing the essential factors and optimal conditions for intensified ceRNA activity. Our analysis results demonstrated that increased size of miRPs as well as the abundance of miRP binding sites stabilize ceRNA activity and strengthen coexpression of ceRNA pairs. Furthermore, the expression levels of both miRPs and ceRNAs affect ceRNA activity and lead to statistically significant differences in distributions of correlation coefficients. Taken together, the results indicated that ceRNA regulation depends on states of the essential factors and thus may involve complex and dynamic processes in vivo. Our findings bring biological insights into complex ceRNA crosstalk in glioblastoma multiforme and contribute to further unveiling complex mechanism governing ceRNA regulation.
The U-curve branch-and-bound algorithm for optimization was introduced recently by Ris and collaborators. In this paper we introduce an improved algorithm for finding the optimal set of features based on the U-curve assumption. Synthetic experiments are used to asses the performance of the proposed algorithm, and compare it to exhaustive search and the original algorithm. The results show that the modified U-curve BB algorithm makes fewer evaluations and is more robust than the original algorithm.
Cross-validation is commonly used to estimate the overall error rate of a designed classifier in a small-sample expression study. The true error of the classifier is a function of the prior probabilities of the classes. With random sampling these can be estimated consistently in terms of the class sample sizes, but when sampling is separate, meaning these sample sizes are determined prior to sampling, there are no reasonable estimates from the data and the prior probabilities must be “estimated” outside the experiment. We have conducted a set of simulations to study the bias of cross-validation as a function of these “estimates”. The results show that a poor choice for estimating these probabilities can significantly increase the bias of cross-validation as an estimator of the true error.
One of the main issues in systems biology is limited resources for conducting biological experiments. Therefore, a strategy for prioritizing the experiments seems to be inevitable. Experimental design is the process of planning experiments in such a way to make experiments as informative as possible. In this work, we propose a novel strategy for designing effective experiments that can optimally reduce the uncertainty in gene regulatory networks, based on the concept of mean objective cost of uncertainty (MOCU).
DBComposer is an R package with a graphical user interface (GUI) to analyze and integrate human gene expression microarray data. With DBComposer, the data can be easily annotated, preprocessed and analyzed in several ways. DBComposer can also serve as a personal expression microarray database allowing users to store multiple datasets together for later retrieval or data analysis. It takes advantage of many R packages for statistics and visualizations, and provides a flexible framework to implement custom workflows to extend the data analysis capabilities.
High dimensional data and small samples make genomic/proteomic classifier design and error estimation virtually impossible without the use of prior information [1]. Dalton and Dougherty utilize prior biological knowledge via a Bayesian approach that considers a prior distribution on an uncertainty class of feature-label distributions [2], [3]. While their general framework is very broad, the focus their attention on multinomial and Gaussian models, for which they derive closed-form solutions of the minimum mean squared error (MMSE) error estimate, the MSE of the error estimate, and an optimal Bayesian classifier (OBC) classifier relative to the prior distribution. Sequencing datasets consist of the number of reads found to map to specific regions of a reference genome. As such, they are often modeled with a discrete distribution, such as the Poisson. For this reason, Gaussian and multinomial distributions are not ideal for sequence-based datasets. Thus, we introduce a multivariate Poisson model (MP) and the associated MP OBC for classifying samples using sequencing data. Lacking closed-form solutions, we employ a Monte Carlo Markov Chain (MCMC) approach to perform classification. We demonstrate superior classification performance for more complex synthetic datasets and comparable performance to the top classifiers in other simpler synthetic datasets.
In this paper, we investigate a phenotypically constrained inference algorithm to reconstruct genetic regulatory networks modeled as Boolean networks (BNs). Based on a previous universal Minimum Description Length (uMDL) network inference algorithm, we study whether adding the prior information based on prescribed attractors or steady states can help better reconstruct the underlying gene regulatory relationships. Comparing the network inference performance with and without prescribed steady states, the experiments based on randomly generated networks as well as a metastatic melanoma network have shown that the phenotypically constrained inference obtains improved performance when we have small numbers of state transition observations.
Long non-coding RNAs (lncRNAs) are suspected to have a wide range of roles in cellular functions. The precise transcriptional mechanisms and the interactions with coding RNAs (genes) are yet to be elucidated. In this paper we present a novel methodology that explores interactions between coding genes and lncRNAs and constructs gene-lncRNA co-expression networks, taking into account their unique expression characteristics. We evaluated several similarity measures to associate a gene and a lncRNA from RNA sequencing data of breast cancer patients and determined correlation to be the metric appropriately suited to this kind of data. Based on an empirically determined threshold, we selected a number of pairs to construct co-expression networks and identified sub-networks that capture previously-unknown lncRNA partners of key players in breast cancer like estrogen receptor. In essence, we have developed a data-driven approach to identify important, functional, coding-lncRNA interactions that sets the stage for more in-depth analyses capturing how non-coding interactions influence expression of protein coding genes and modulate pathways contributing to cancer.
The coefficient of determination (CoD) has significant applications in genomics, for example, in the inference of gene regulatory networks. In previous publications, we have studied several nonparametric CoD estimators, based upon the resubstitution, leave-one-out, cross-validation, and bootstrap error estimators, and one parametric maximum-likelihood (ML) CoD estimator that allows the incorporation of available prior knowledge, from a frequentist perspective. However, none of these CoD estimators are rigorously optimized based on statistical inference across a family of possible distributions. Therefore, by following the idea of Bayesian error estimation for classification, we define a Bayesian CoD estimator that minimizes the mean-square error (MSE), based on a parametrized family of joint distributions between predictors and target as a function of random parameters characterized by assumed prior distributions. We derive an exact formulation of the sample-based Bayesian MMSE CoD estimator. Numerical experiments are carried out to estimate performance metrics of the Bayesian CoD estimator and compare them against those of resubstitution, leave-one-out, bootstrap and cross-validation CoD estimators over all the distributions, by employing the Monte Carlo sample method. Results show that the Bayesian CoD estimator has the best performance, displaying zero bias, small variance, and least root mean-square error (RMS).
Following the breakthrough of the microarray technology, the next generation sequencing (NGS) technology further advanced approaches in modern biomedical research. The high-throughput NGS technology is now frequently used in profiling tumor and control samples for the study of DNA copy number variants (CNVs). In particular, the ratio of read count of the tumor sample to that of the control sample is popularly used for identifying CNV regions. We illustrate that a change-point (or a breakpoint) detection method, along with a Bayesian approach, is particularly suitable for identifying CNVs in the reads ratio data. We have written our algorithm into a user friendly R-package, SeqBBS (stands for Bayesian breakpoints search for sequencing data) and applied our method to the sequencing data of reads ratio between the breast tumor cell lines HCC1954 and its matched normal cell line BL1954. Breakpoints that separate different CNV regions are successfully identified.
With development of new technologies applied to biological experiments, more and more data are generated every day. To make predictions in biological systems, mathematical modeling plays a critical role. Ordinary differential equations (ODEs) contribute to a large portion in mathematical modeling. In which parameters are inevitable. Noise is intrinsic in all experiments. Therefore, to think of parameters as statistical distributions is a realistic treatment. In this paper, we discuss in a 1st order ODE common in biological systems, how to calculate parameter distribution analytically according to the experimentally observed output assumed to be normal distribution. Conditions on when parameter can be correctly estimated are elucidated.
Summary form only given. Understanding the mechanism of transcriptional regulation remains to be an inspiring stage of molecular biology. Within the popular methods for modeling TFBS, position-specific weight matrix and k-mer based approaches have gained great success. However, both approaches fail to consider the structural properties of a binding site. Recently, a novel TFBS modeling and predicting approach is presented by Bauer et al. (2010), where the sequence-specific chemical and structural features of DNA are applied. However, the in vivo protein-DNA interactions observed in ChIP-chip assays, which were used in this study, are not necessarily direct, as some TFs tend to interact with DNAs extensively through other partners. Therefore, an evaluation on a proper in vitro dataset would be more appropriate to reveal the benefit of such physicochemical features in modeling TF-DNA interactions. Recently, in vitro protein-binding microarray experiment has greatly improved the understanding of transcription factor-DNA interaction. It is a high-throughput experiment used to measure the in vitro binding affinity of a given TF to the sequences on the probe array. Because typical confounding factors such as transcription co-factors present in ChIP-based experiments are eliminated, PBM data provide an excellent information source to develop structural models for TF-DNA interactions. On the other hand, directly mapping of the 3-mer or 4-mer based meta-features to the candidate DNA binding sequences as in their work may not reflect the TF-DNA binding nature, since a TFBS is usually an 8 to 12 base-pair. As a result, conventionally machine-learning algorithms, which rely on well-structured feature vector and label pairs, may not work well in modeling PBM data. In this paper we propose a novel approach to predicts in vitro transcription factor binding based on the structural properties of DNA using a so-called multiple-instance learning algorithm. Compared to conventional (single-instance based) learning algorithms, our multi-instance learning-based algorithm does not require the knowledge of the actual binding site within a candidate probe sequence, yet can still take full advantage of the physicochemical properties in modeling and predicting TF-DNA interactions. Evaluation on an in vitro protein binding microarray data of twenty mouse TFs shows that our new model performs significantly better than several k-mer or structure-based single-instance learning algorithms. It indicates that combining multi-instance learning and structural properties of DNA has promising potential for studying biological regulatory networks.
By using the competing endogenous RNA (ceRNA) concept, we implemented a web-based application TraceRNA. TraceRNA allows us to interactively construct a regulation network for a specific phenotype by using a disease-specific transcriptome data. In this work, we further extend the TraceRNA with a novel algorithm implementation where we examined the microRNA expression derived from same disease type. The proposed algorithm, NetceRNA, finds an optimized network representation under a certain phenotype context by iteratively perturbing the network and measuring the network configuration change with respect to the original ceRNA network. The resulting algorithm outputs an improved network together with a ranked list of genes and miRNAs which are characteristic of the specific phenotype. To illustrate the utility of NetceRNA, gene expression and microRNA expression data of breast cancer study from The Cancer Genome Atlas (TCGA) were used.
A central problem in cancer genomics is to identify interpretable biomarkers for better disease prognosis. Many of the biomarkers identified through Cox Proportional Hazard (PH) models are biologically uninterpretable. We propose the use of graph Laplacian regularized Cox PH model to integrate biological networks into the feature selection problem in survival analysis. Simulation studies demonstrate that the performance of the proposed algorithm is superior to L1 and L1+L2 regularized Cox PH models. Utility of this algorithm is also validated by its ability to identify key known biomarkers such as p53 and myc in estrogen receptor positive breast cancer patients using genomic abberration data generated by the Cancer Genome Altas consortium. With the rapid expansion of our knowledge of biological networks, this approach will become increasingly useful for mining high-throughput genomic datasets.
A novel assembly pipeline, MiB, employs Minimum Description Length (MDL), de-Bruijn graphs and Bayesian estimation for reference assisted assembly of the novel genome. In a previous study MiB assembly was compared with nine other assembly algorithms showing significant improvement in results coupled with very large execution times. This correspondence introduces ‘Supersonic MiB’, an extension to our previous study MiB. Supersonic MiB aims to stimulate the assembly pipeline of MiB showing significant improvement in execution time compared to its predecessor.
The etiology of chemically-induced neurotoxicity like seizures is poorly understood. Using reversible neurotoxicity induced by two neurotoxicants as example, we demonstrate that a bioinformatics-guided reverse engineering approach can be applied to analyze time series microarray gene expression data and uncover the underlying molecular mechanism. Our results reinforce previous findings that cholinergic and GABAergic synapse pathways are the target of carbaryl and RDX, respectively. We also conclude that perturbations to these pathways by sublethal concentrations of RDX and carbaryl were temporary, and earthworms were capable of fully recovering at the end of the 7-day recovery phase. In addition, our study indicates that many pathways other than those related to synaptic and neuronal activities were altered during the 6-day exposure phase.
It is commonplace in bioinformatics (and elsewhere) to build a classifier from sample data in which the sample sizes of the classes are not random; that is, they are selected prior to sampling. The result is that there is no estimate of the prior class probabilities available from the data. In this paper, we find an analytic result for the minimax solution for the class prior probabilities for a general Neyman-Pearson induced classifier. From that we derive Anderson's classical minimax prior probability “estimate.” Using synthetic and real data, we demonstrate the degradation in classifier performance from using inaccurate values for the prior probabilities.