
We reported previously that the persistence of complex immune, endocrine and neurological symptoms that afflict up to one third of veterans from the 1990-91 Gulf War might be supported by a misdirected regulatory drive. Here we use a detailed model of immune signaling in concert with an overarching circuit model of known sex and stress hormone co-regulation to explore how the failure of regulatory elements may further establish a self-perpetuating imbalance that closely resembles Gulf War Illness (GWI). Defects to the model were imparted iteratively and the stable regulatory modes supported by these altered immune-endocrine circuits were identified using repeated simulation experiments. In each case the predicted homeostatic regimes were compared to experimental data collected in male GWI (n=20 ) and matched healthy veterans (n=22 ). We found that alignment of GWI with a new homeostatic regime improved significantly when cortisol's normal anti-inflammatory activity was interrupted. Alignment improved further when this cortisol insensitivity was compounded by the loss of the normal antagonistic effects of Th1 cytokines on Th2 lymphocyte activation. Together these simulation results suggest altered glucocorticoid gene regulation compounded by possible changes in IGF-1 regulation of Th1:Th2 immune balance may be key underlying features of GWI.
In this study, we developed a transcriptomics based human in vitro model for predicting DILI in humans. The transcriptomics data (Affymetrix GeneChip Human Genome U133 Plus 2.0) from primary human hepatocytes were provided by the Japanese Toxicogenomics Project (TGP). The selected compounds were divided into two groups, i.e., most-DILI and no-DILI, based on FDA-approved drug labels. The compounds were further grouped in a training and validation set. The training set, containing the most extreme most-DILI and no-DILI compounds based on the in vivo rat clinical chemistry measurements from TGP, was used to develop the prediction model. The validation set showed high accuracy (> 90%) and performed better than splitting the compounds into training and validation set randomly.
We investigate the problem of detecting toxicogenomic associations that generalize across organisms, that is, statistical dependencies between transcriptional responses of multiple organisms and toxicological outcomes. We apply an interpretable probabilistic model to detect cross-organism toxicogenomic associations and propose an approach for drug toxicity analysis based on the interactive retrieval of drugs with similar toxicogenomic properties. We show that our approach can give relevant information about the properties of a drug even when direct prediction of toxicity is not feasible. Moreover, we show that a search from a cross-organism database can improve accuracy in the analysis.
The Japanese Toxicogenomics Project (TGP) provides large amount of data for the toxicology and safety framework. We focus on gene expression data of rat in vivo and human in vitro. We consider two different analyses for the TGP data. The first analysis is based on two-way analysis of variance model and the goal is to detect genes with significant dose-response relationship in both humans and rats. The second analysis consists of a trend analysis at each time point and the goal is to detect genes in the rat in order to predict gene expression in humans. The first analysis leads us to conclusions about the heterogeneity of the compound set and will suggest how to address this issue to improve future analyses. In the second part, we identify, for particular compounds, groups of genes that are translatable from rats to humans, so they can be used for prediction of human in vitro data based on rat in vivo data.
Our aim was to predict drug-induced liver injury potential from differential expression profiles in the Japanese Toxicogenomic Project's data. Additionally, we wondered which drug concentrations and treatment periods proved most informative for classification. The study confirmed that the effect of drugs on differential gene expression is dose and time-dependent suggesting that less information is contained in responses to low doses and short treatment periods. Cross-validation performance of predictions was low due to high false positive rate. Present results indicate that distinguishing between differential expression responses to nontoxic and toxic agents is non-trivial. The value of the Japanese Toxicogenomic Project's data for classification purposes would be increased if gene expression responses to additional nontoxic drugs were available.
Since experiments involving animal models are labor and time intensive, there is an attempt to replace these measurements on animal models with in vitro assays which has higher acceptance in the population concerning ethical issues. In this work, we explore to what extend animal models can be replaced by in vitro assays in the context of a toxicogenomics study. The data from the Japanese Toxicogenomics Project are gene expression profiles measured by microarrays from both in vitro and animal samples. We apply a comprehensive genomic association network analysis in order to study the comparative behavior of the genomic networks for the in vivo vs. in vitro data. The genomic networks are computed based on association scores of gene-gene pairs using a partial least squares modeling of gene expression values adjusted for sacrifice time and dosage. We apply permutation based statistical tests to compare the connectivity of a given gene, as well as a class of genes in the two networks which may be affected by a given drug. The goal is to identify parts of these networks including key genes that are not significantly altered for in vivo vs. in vitro samples for the majority of the drugs.
It is known that all agents that cause cancer (carcinogens) also cause a change in the DNA sequence. In order to identify such often subtle changes, we attempt to integrate multiple molecular profile data sets released by the International Cancer Genome Consortium (ICGC). The list of data sets includes matched gene and microRNA expression profiles, somatic copy number variation, DNA methylation, and protein expression profiles for lung adenocarcinoma patients receiving treatments. We consider both unsupervised and supervised learning techniques (clustering and penalized regression) to identify interesting molecular markers corresponding to each type of –omics profiles that can differentiate patients. Associations between important markers of 2 types have been studied. An adaptive ensemble binary regression model has been presented that uses the entirety of available –omics profiles leading to a more accurate clinical prognosis for the patients in the given sample. This integrated study provides a more comprehensive picture of lung adenocarcinoma.
The 12th Annual International Conference on the Critical Assessment of Massive Data Analysis (CAMDA) used data from the massive Japanese Toxicogenomics Project (TGP) to predict drug-induced liver injury (DILI) concern provided by the U.S. Food and Drug Administration (FDA). The challenge was to predict DILI concern by means of gene expression data. Analysis of this high-dimensional toxicogenomic data requires statistical methodologies that can detect the transcriptomic associations with toxicity. We propose an analysis technique that involves sparse principal component analysis to efficiently reduce the dimension of the analysis problem. Sparse principal component variables are composed of groups of expressed genes. Associations between DILI concern and sparse principal component variables were tested and further scrutinized with sparse regression methodology to identify concise transcriptomic structures potentially responsible for and predictive of drug toxicity. Working with a subset of the TGP data with FDA DILI concern classification, we identified 5 transcriptomic structures (sparse principal component variables) statistically associated with DILI concern. The most statistically significant structure consists of the genes ZBTB16, FLVCR2, TNS3, and ASB13. Sparse statistical methods offer a new way to handle analysis issues with massive omic data. Sparse PCA can efficiently extract groups of transcriptomic markers that may indicate drug toxicity.
Numerous methods of RNA-Seq data analysis have been developed, and there are more under active development. In this paper, our focus is on evaluating the impact of each processing stage; from pre-processing of sequencing reads to alignment/counting to count normalization to differential expression testing to downstream functional analysis, on the inferred functional pattern of biological response. We assess the impact of 6,912 combinations of technical and biological factors on the resulting signature of transcriptomic functional response. Given the absence of the ground truth, we use 2 complementary evaluation criteria: a) consistency of the functional patterns identified in 2 similar comparisons, namely effects of a naturally-toxic medium and a medium with artificially reconstituted toxicity, and b) consistency of the results in RNA-Seq and microarray versions of the same study. Our results show that despite high variability at the low-level processing stage (read pre-processing, alignment and counting) and the differential expression calling stage, their impact on the inferred pattern of biological response was surprisingly low; they were instead overshadowed by the choice of the functional enrichment method. The latter have an impact comparable in magnitude to the impact of biological factors per se.
Traditional studies of liver toxicity involve screening compounds through in vivo and in vitro tests. They need to distinguish between compounds that represent little or no health concern and those with the greatest likelihood to cause adverse effects in humans. High-throughput and toxicogenomic screening methods coupled with a plethora of circumstantial evidence provide a challenge for improved toxicity prediction and require appropriate computational methods that integrate various biological, chemical and toxicological data. We report on a data fusion approach for prediction of drug-induced liver injury potential in humans using microarray data from the Japanese Toxicogenomics Project (TGP) as provided for the contest by CAMDA 2013 Conference. Our aim was to investigate if the data from different TGP studies could be fused together to boost prediction accuracy. We were also interested if in vitro studies provided sufficient information to refrain from studies in animals. We show that our recently proposed matrix factorization-based data fusion provides an elegant computational framework for integration of the TGP and related data sets, 29 data sets in total. Fusion yields a high cross-validated accuracy (AUC of 0.819 for in vivo assays), which is above the accuracy of the established machine learning procedure of stacked classification with feature selection. Our data analysis shows that animal studies may be replaced with in vitro assays (AUC = 0.799) and that liver injury in humans can be predicted from animal data (AUC = 0.811). Our principal contribution is a demonstration that analysis of toxicogenomic data can substantially benefit from data fusion with directly and circumstantially related data sets.
Any knowledge discovery could in principal benefit from the fusion of directly or even indirectly related data sources. In this paper we explore whether data fusion by simultaneous matrix factorization could be adapted for survival regression. We propose a new method that jointly infers latent data factors from a number of heterogeneous data sets and estimates regression coefficients of a survival model. We have applied the method to CAMDA 2014 large-scale Cancer Genomes Challenge and modeled survival time as a function of gene, protein and miRNA expression data, and data on methylated and mutated regions. We find that both joint inference of data factors and regression coefficients and data fusion procedure are crucial for performance. Our approach is substantially more accurate than the baseline Aalen's additive model. Latent factors inferred by our approach could be mined further; for CAMDA challenge, we found that the most informative factors are related to known cancer processes.
The central nervous system (CNS) is composed of hundreds of distinct cell types, each expressing different subsets of genes from the genome. High throughput gene expression analysis of the CNS from patients and controls is a common method to screen for potentially pathological molecular mechanisms of psychiatric disease. One mechanism by which gene expression might be seen to vary across samples would be alterations in the cellular composition of the tissue. While the expressions of gene “markers” for each cell type can provide certain information of cellularity, for many rare cell types markers are not well characterized. Moreover, if only small sets of markers are known, any substantial variation of a marker’s expression pattern due to experiment conditions would result in poor sensitivity and specificity. Here, our proposed method combines prior information from mice cell-specific transcriptome profiling experiments with co-expression network analysis, to select large sets of potential cell type-specific gene markers in a systematic and unbiased manner. The method is efficient and robust, and identifies sufficient markers for further cellularity analysis. We then employ the markers to analytically detect changing cellular composition in human brain. Application of our method to temporal human brain microarray data successfully detects changes in cellularity over time that roughly correspond to known epochs of human brain development. Furthermore, application of our method to human brain samples with the neurodevelopmental disorder of autism supports the interpretation that the changes in astrocytes and neurons might contribute to the disorder.
Feedback mechanisms throughout the immune and endocrine systems play a significant role in maintaining physiological homeostasis. Specifically, the hypothalamic-pituitary-adrenal (HPA) and hypothal...
Sample classification, especially disease status prediction, is an important area of investigation for gene expression studies. Many machine learning methods have been developed to tackle this problem. To evaluate different prediction methods, the IMPROVER Challenge made several data sets available. Here we focus on one sub-challenge: chronic obstructive pulmonary disease (COPD). We outlined critical preprocessing steps to make training and test data comparable. We compared our recently introduced random generalized linear model (RGLM) predictor with Leo Breiman’s random forest (RF) predictor on the COPD data set. We discussed potential reasons for the superior performance of the RGLM predictor in this sub-challenge. Interestingly, we found that although several genes were highly predictive of COPD status, none were necessary to achieve accurate prediction when demographic features smoking status and age were used. In conclusion, RGLM achieved superior predictive accuracy for predicting COPD status with smoking status and age as mandatory features. Future cohort studies could evaluate whether the resulting predictor has clinical utility.
The sbv IMPROVER Diagnostic Signature Challenge used crowdsourcing to identify the best methods to classify clinical samples using transcriptomics data. Participating teams used public microarray data sets to develop prediction models in four disease areas, and then made predictions on blinded test data generated by the organizers. Here we describe the approach of the team for the Perinatology Research Branch (Team PRB; AL Tarca, R Romero), that was awarded the best performing entrant prize out of 54 entrants. The key elements of our approach included: (1) selection of training data sets by trial and error; (2) removal of batch effects by pre-processing the test and training data together; (3) the use of statistical significance and magnitude of change to select biomarkers; and (4) optimization of the number of biomarkers via the cross-validated performance of a simple linear discriminant analysis (LDA) model. Not only were our resulting models ranked consistently high, but they also generated parsimonious signatures of as low as two genes, unlike most of the other top-ranked teams that used hundreds of genes for prediction.
Motivation: Current clinical and biological studies apply different biotechnologies and subsequently combine the resulting -omics data to test biological hypotheses. The plethora of -omics data and their combination generates a large number of hypotheses and apparently increases the study power. Contrary to these expectations, the wealth of -omics data may even reduce the statistical power of a study because of a large correction factor for multiple testing. Typically, this loss of power in analyzing -omics data are caused by an increased false detection rate (FDR) in measurements, like falsely detected DNA copy number changes, or falsely identified differentially expressed genes. The false detections are random and, therefore, not related to the tested conditions. Thus, a high FDR considerably decreases the discovery power of studies, especially if different -omics data are involved. Results: On a HapMap data set, where known CNVs have to be re-detected, I/NI call filtering was much more efficient than variance-based filtering. In particular, the I/NI call filter outperforms variance-based filters on data with rare events like the CNVs in the HapMap data set. We assessed the efficiency of the I/NI call filter in reducing the FDR on two different cancer cell lines where it reduced the FDR 18- to 22-fold. Materials and Methods: A mitigation strategy for too high FDRs is to filter out putative false detections. We suggest using probabilistic latent variable models to identify putative false detections which may be found via such models by high estimated noise or by model-based measurement inconsistencies across samples. To select such a model, a Bayesian approach starts with the maximum a priori model that assumes no detection and selects the maximum a posteriori model. Hence detection results in a deviation of the maximal posterior from the maximal prior model measured by the information gain obtained by the data. If this information gain exceeds a threshold then the selected model obtains an Informative/Non-Informative (I/NI) call that indicates a detection. I/NI call filtering has been successfully applied in different projects, but it has so far not been shown that correction for multiple testing after I/NI call filtering still controls the type-I error rate. We prove this important property of the I/NI call and show that it is independent of commonly used test statistics for null hypotheses. We apply the I/NI call to transcriptomics (gene expression), where the prior model corresponds to a constant gene expression level across compared samples, and to genomics, analyzing copy number variation (CNV) data, where the prior model corresponds to a constant DNA copy number of 2 across compared samples.
For the first time, we report here that Illumina high-density methylation arrays can also be used to estimate DNA copy number variations. We used the Illumina HM450K methylation array data to characterize the DNA copy number aberrations in the HT-29 colon cancer cell line to test our statistical model. Results were validated using an Affymetrix SNP array. Utilizing the CAMDA 2011 glioblastoma data set, we have demonstrated that our novel statistical method can potentially lower the cost and reduce the processing time of large-scale profiling studies where both DNA copy number and methylation status are of interest. Our new method, named methylCNV, is implemented in the Lumi package of Bioconductor.
Western medicine, according to definitions from the NCI Dictionary, Macmillan Dictionary and Oxford Dictionary, is based on treating symptoms and diseases using drugs, surgery, and such. An alterna...
Identifying robust biomarkers for cancer phenotypes has challenged the biological and pharmacological communities for many years, more so since the availability of screening methods that reveal the expression levels of all the genes in the genome. A host of different approaches have been used to address this lack of robustness. These methods have included a spectrum of approaches from gene enrichment analysis to network inference analysis. More recently, some methods that use the network properties of genes have demonstrated an ability to provide a more robust signature. In this review, we survey different network-as-biomarker methods used to identify various biomarkers and we discuss the critical role of networks in the progress toward personalized medicine. We also discuss the ability of the network to identify misguided processes, rather than the gene itself, as the core of distinctions among phenotypes. Discussions about the importance of the molecular pathway view and about processes (rather than the gene per se) at the core of understanding cancer are not new. However, this review focuses on the set of tools available for actually measuring the pathway, or the process, when the expression levels of their components are available.
The central nervous system (CNS) is composed of hundreds of distinct cell types, each expressing different subsets of genes from the genome. High throughput gene expression analysis of complex tissues like the CNS from patients and controls is a common method to screen for potentially pathological molecular mechanisms of psychiatric disease. One mechanism by which gene expression might be seen to vary across samples would be alterations in the cellular composition of the tissue. While there are a few gene `markers' from literature for each cell type, their expression patterns vary significantly resulting in poor sensitivity and specificity. Here, we propose a method utilizing prior information from cell specific transcriptome profiling experiments in mice and co-expression network analysis to select cell type specific gene markers, and further to analytically detect changing cellular composition in human tissues. Our method successfully detects changes in cellularity over time that roughly correspond to known epochs of human brain development.