Protein quality control (PQC) systems are essential for cellular resilience to proteotoxic stress. Despite intensive study, functional redundancies in the system obscure the contributions of important individual genes. Here, we leverage transposon sequencing (Tn-seq) across bacterial strains lacking key chaperones and proteases to reveal hidden determinants of stress response in protein homeostasis. By profiling fitness under multiple proteotoxic stresses, we uncover stress-specific vulnerabilities and reveal how major players of PQC mask correlations between transcriptomic responses and gene fitness. We identify a heat-specific synthetic lethality between the ClpB disaggregase and DNA polymerase 1 mediated by prolific aggregation of the RecA recombinase and persistent induction of the heat shock regulon, supporting a conclusion that vulnerabilities in PQC are genetic- and environmental-context specific. Overall, our work presents a framework to reveal critical addressable fragilities in stress responses using gene fitness scores adaptable to a variety of systems.
Many Bayesian statistical inference problems come down to computing a maximum a-posteriori (MAP) assignment of latent variables. Yet, standard methods for estimating the MAP assignment do not have a finite time guarantee that the algorithm has converged to a fixed point. Previous research has found that MAP inference can be represented in dual form as a linear programming problem with a non-polynomial number of constraints. A Lagrangian relaxation of the dual yields a statistical inference algorithm as a linear programming problem. However, the decision as to which constraints to remove in the relaxation is often heuristic. We present a method for maximum a-posteriori inference in general Bayesian factor models that sequentially adds constraints to the fully relaxed dual problem using Benders' decomposition. Our method enables the incorporation of expressive integer and logical constraints in clustering problems such as must-link, cannot-link, and a minimum number of whole samples allocated to each cluster. Using this approach, we derive MAP estimation algorithms for the Bayesian Gaussian mixture model and latent Dirichlet allocation. Empirical results show that our method produces a higher optimal posterior value compared to Gibbs sampling and variational Bayes methods for standard data sets and provides certificate of convergence.
Every protein progresses through a natural lifecycle from birth to maturation to death; this process is coordinated by the protein homeostasis system. Environmental or physiological conditions trigger pathways that maintain the homeostasis of the proteome. An open question is how these pathways are modulated to respond to the many stresses that an organism encounters during its lifetime. To address this question, we tested how the fitness landscape changes in response to environmental and genetic perturbations using directed and massively parallel transposon mutagenesis in Caulobacter crescentus. We developed a general computational pipeline for the analysis of gene-by-environment interactions in transposon mutagenesis experiments. This pipeline uses a combination of general linear models, statistical knockoffs, and a nonparametric Bayesian statistical model to identify essential genetic network components that are shared across environmental perturbations. This analysis allows us to quantify the similarity of proteotoxic environmental perturbations from the perspective of the fitness landscape. We find that essential genes vary more by genetic background than by environmental conditions, with limited overlap among mutant strains targeting different facets of the protein homeostasis system. We also identified 146 unique fitness determinants across different strains, with 19 genes common to at least two strains, showing varying resilience to proteotoxic stresses. Experiments exposing cells to a combination of genetic perturbations and dual environmental stressors show that perturbations that are quantitatively dissimilar from the perspective of the fitness landscape are likely to have a synergistic effect on the growth defect.
Abstract Background Multiple organ failure/dysfunction syndrome (MOF/MODS) is a major cause of mortality and morbidity among severe trauma patients. Current clinical practices entail monitoring physiological measurements and applying clinical score systems to diagnose its onset. Instead, we aimed to develop an early prediction model for MOF outcome evaluated soon after traumatic injury by performing machine learning analysis of genome-wide transcriptome data from blood samples drawn within 24 h of traumatic injury. We then compared its performance to baseline injury severity scores and detection of infections. Methods Buffy coat transcriptome and linked clinical datasets from blunt trauma patients from the Inflammation and the Host Response to Injury Study (“Glue Grant”) multi-center cohort were used. According to the inclusion/exclusion criteria, 141 adult (age ≥ 16 years old) blunt trauma patients (excluding penetrating) with early buffy coat (≤ 24 h since trauma injury) samples were analyzed, with 58 MOF-cases and 83 non-cases. We applied the Least Absolute Shrinkage and Selection Operator (LASSO) and eXtreme Gradient Boosting (XGBoost) algorithms to select features and develop models for MOF early outcome prediction. Results The LASSO model included 18 transcripts (AUROC [95% CI]: 0.938 [0.890–0.987] (training) and 0.833 [0.699–0.967] (test)), and the XGBoost model included 41 transcripts (0.999 [0.997–1.000] (training) and 0.907 [0.816–0.998] (test)). There were 16 overlapping transcripts comparing the two panels (0.935 [0.884–0.985] (training) and 0.836 [0.703–0.968] (test)). The biomarker models notably outperformed models based on injury severity scores and sex, which we found to be significantly associated with MOF (APACHEII + sex—0.649 [0.537–0.762] (training) and 0.493 [0.301–0.685] (test); ISS + sex—0.630 [0.516–0.744] (training) and 0.482 [0.293–0.670] (test); NISS + sex—0.651 [0.540–0.763] (training) and 0.525 [0.335–0.714] (test)). Conclusions The accurate assessment of MOF from blood samples immediately after trauma is expected to aid in improving clinical decision-making and may contribute to reduced morbidity, mortality and healthcare costs. Moreover, understanding the molecular mechanisms involving the transcripts identified as important for MOF prediction may eventually aid in developing novel interventions.
We consider the problem of developing interpretable and computationally efficient matrix decomposition methods for matrices whose entries have bounded support. Such matrices are found in large-scale DNA methylation studies and many other settings. Our approach decomposes the data matrix into a Tucker representation wherein the number of columns in the constituent factor matrices is not constrained. We derive a computationally efficient sampling algorithm to solve for the Tucker decomposition. We evaluate the performance of our method using three criteria: predictability, computability, and stability. Empirical results show that our method has similar performance as other state-of-the-art approaches in terms of held-out prediction and computational complexity, but has significantly better performance in terms of stability to changes in hyper-parameters. The improved stability results in higher confidence in the results in applications where the constituent factors are used to generate and test scientific hypotheses such as DNA methylation analysis of cancer samples.
Experimentation involves risk. The investigator expends time and money in the pursuit of data that supports a hypothesis. In the end, the investigator may find that all of these costs were for naught and the data fail to reject the null. Furthermore, the investigator may not be able to test other hypotheses with the same data set in order to avoid false positives due to p-hacking. Therefore, there is a need for a mechanism for investigators to hedge the risk of financial and statistical bankruptcy in the business of experimentation. In this work, we build on the game-theoretic statistics framework to enable an investigator to hedge their bets against the null hypothesis and thus avoid ruin. First, we describe a method by which the investigator's test martingale wealth process can be capitalized by solving for the risk-neutral price. Then, we show that a portfolio that comprises the risky test martingale and a risk-free process is still a test martingale which enables the investigator to select a particular risk-return position using Markowitz portfolio theory. Finally, we show that a function that is derivative of the test martingale process can be constructed and used as a hedging instrument by the investigator or as a speculative instrument by a risk-seeking investor who wants to participate in the potential returns of the uncertain experiment wealth process. Together, these instruments enable an investigator to hedge the risk of ruin and they enable a investigator to efficiently hedge experimental risk.
Supplementary Table 2 from Oncogenic <i>BRAF</i> Mutation with <i>CDKN2A</i> Inactivation Is Characteristic of a Subset of Pediatric Malignant Astrocytomas
ABSTRACT Introduction: Despite significant advances in pediatric burn care, bloodstream infections (BSIs) remain a compelling challenge during recovery. A personalized medicine approach for accurate prediction of BSIs before they occur would contribute to prevention efforts and improve patient outcomes. Methods: We analyzed the blood transcriptome of severely burned (total burn surface area [TBSA] ≥20%) patients in the multicenter Inflammation and Host Response to Injury (“Glue Grant”) cohort. Our study included 82 pediatric (aged <16 years) patients, with blood samples at least 3 days before the observed BSI episode. We applied the least absolute shrinkage and selection operator (LASSO) machine-learning algorithm to select a panel of biomarkers predictive of BSI outcome. Results: We developed a panel of 10 probe sets corresponding to six annotated genes (ARG2 [arginase 2], CPT1A [carnitine palmitoyltransferase 1A], FYB [FYN binding protein], ITCH [itchy E3 ubiquitin protein ligase], MACF1 [microtubule actin crosslinking factor 1], and SSH2 [slingshot protein phosphatase 2]), two uncharacterized (LOC101928635, LOC101929599), and two unannotated regions. Our multibiomarker panel model yielded highly accurate prediction (area under the receiver operating characteristic curve, 0.938; 95% confidence interval [CI], 0.881–0.981) compared with models with TBSA (0.708; 95% CI, 0.588–0.824) or TBSA and inhalation injury status (0.792; 95% CI, 0.676–0.892). A model combining the multibiomarker panel with TBSA and inhalation injury status further improved prediction (0.978; 95% CI, 0.941–1.000). Conclusions: The multibiomarker panel model yielded a highly accurate prediction of BSIs before their onset. Knowing patients' risk profile early will guide clinicians to take rapid preventive measures for limiting infections, promote antibiotic stewardship that may aid in alleviating the current antibiotic resistance crisis, shorten hospital length of stay and burden on health care resources, reduce health care costs, and significantly improve patients' outcomes. In addition, the biomarkers' identity and molecular functions may contribute to developing novel preventive interventions.
Supplementary Table 1 from Oncogenic <i>BRAF</i> Mutation with <i>CDKN2A</i> Inactivation Is Characteristic of a Subset of Pediatric Malignant Astrocytomas
Large-scale multiple perturbation experiments have the potential to reveal a more detailed understanding of the molecular pathways that respond to genetic and environmental changes. A key question in these studies is which gene expression changes are important for the response to the perturbation. This problem is challenging because (i) the functional form of the nonlinear relationship between gene expression and the perturbation is unknown and (ii) identification of the most important genes is a high-dimensional variable selection problem. To deal with these challenges, we present here a method based on the model-X knockoffs framework and Deep Neural Networks to identify significant gene expression changes in multiple perturbation experiments. This approach makes no assumptions on the functional form of the dependence between the responses and the perturbations and it enjoys finite sample false discovery rate control for the selected set of important gene expression responses. We apply this approach to the Library of Integrated Network-Based Cellular Signature data sets which is a National Institutes of Health Common Fund program that catalogs how human cells globally respond to chemical, genetic and disease perturbations. We identified important genes whose expression is directly modulated in response to perturbation with anthracycline, vorinostat, trichostatin-a, geldanamycin and sirolimus. We compare the set of important genes that respond to these small molecules to identify co-responsive pathways. Identification of which genes respond to specific perturbation stressors can provide better understanding of the underlying mechanisms of disease and advance the identification of new drug targets.
We consider the problem of sequential multiple hypothesis testing with nontrivial data collection costs. This problem appears, for example, when conducting biological experiments to identify differentially expressed genes of a disease process. This work builds on the generalized α-investing framework which enables control of the marginal false discovery rate in a sequential testing setting. We make a theoretical analysis of the long term asymptotic behavior of α-wealth which motivates a consideration of sample size in the α-investing decision rule. Posing the testing process as a game with nature, we construct a decision rule that optimizes the expected α-wealth reward (ERO) and provides an optimal sample size for each test. Empirical results show that a cost-aware ERO decision rule correctly rejects more false null hypotheses than other methods for $n=1$ where n is the sample size. When the sample size is not fixed cost-aware ERO uses a prior on the null hypothesis to adaptively allocate of the sample budget to each test. We extend cost-aware ERO investing to finite-horizon testing which enables the decision rule to allocate samples in a non-myopic manner. Finally, empirical tests on real data sets from biological experiments show that cost-aware ERO balances the allocation of samples to an individual test against the allocation of samples across multiple tests.
The understanding of bacterial gene function has been greatly enhanced by recent advancements in the deep sequencing of microbial genomes. Transposon insertion sequencing methods combines next-generation sequencing techniques with transposon mutagenesis for the exploration of the essentiality of genes under different environmental conditions. We propose a model-based method that uses regularized negative binomial regression to estimate the change in transposon insertions attributable to gene-environment changes in this genetic interaction study without transformations or uniform normalization. An empirical Bayes model for estimating the local false discovery rate combines unique and total count information to test for genes that show a statistically significant change in transposon counts. When applied to RB-TnSeq (randomized barcode transposon sequencing) and Tn-seq (transposon sequencing) libraries made in strains of Caulobacter crescentus using both total and unique count data the model was able to identify a set of conditionally beneficial or conditionally detrimental genes for each target condition that shed light on their functions and roles during various stress conditions.
There are distinguishing features or "hallmarks" of cancer that are found across tumors, individuals and types of cancer, and these hallmarks can be driven by specific genetic mutations. Yet within a single tumor there is often extensive genetic heterogeneity as evidenced by single-cell and bulk DNA sequencing data. The goal of this work is to jointly infer the underlying genotypes of tumor subpopulations and the distribution of those subpopulations in individual tumors by integrating single-cell and bulk sequencing data. Understanding the genetic composition of the tumor at the time of treatment is important in the personalized design of targeted therapeutic combinations and monitoring for possible recurrence after treatment. We propose a hierarchical Dirichlet process mixture model that incorporates the correlation structure induced by a structured sampling arrangement, and we show that this model improves the quality of inference. We develop a representation of the hierarchical Dirichlet process prior as a Gamma-Poisson hierarchy, and we use this representation to derive a fast Gibbs sampling inference algorithm using the augment-and-marginalize method. Experiments with simulation data show that our model outperforms standard numerical and statistical methods for decomposing admixed count data. Analyses of real acute lymphoblastic leukemia cancer sequencing dataset shows that our model improves upon state-of-the-art bioinformatic methods. An interpretation of the results of our model on this real dataset reveals comutated loci across samples.
The epithelial to mesenchymal transition (EMT) is characterized by a loss of cell polarity, a decrease in the epithelial cell marker E-cadherin, and an increase in mesenchymal markers including the zinc-finger E-box bind-ing homeobox (ZEB1). The EMT is also associated with an increase in cell migration and anchorage-independent growth. Induction of a reversal of the EMT, a mesenchymal to epithelial transition (MET), is an emerging strategy being explored to attenuate the metastatic potential of aggressive cancer types, such as triple-negative breast can-cers (TNBCs) and tamoxifen-resistant (TAMR) ER-positive breast cancers, which have a mesenchymal phenotype. Patients with these aggressive cancers have poor prognoses, quick relapse, and resistance to most chemother-apeutic drugs. Overexpression of extracellular signal-regulated kinase (ERK) 1/2 and ERK5 is associated with poor patient survival in breast cancer. Moreover, TNBC and tamoxifen resistant cancers are unresponsive to most targeted clinical therapies and there is a dire need for alternative therapies. In the current study, we found that MAPK3, MAPK1, and MAPK7 gene expression correlated with EMT mark-ers and poor overall survival in breast cancer patients using publicly available datasets. The effect of ERK1/2 and ERK5 pathway inhibition on MET was evaluated in MDA-MB-231, BT-549 TNBC cells, and tamoxifen-resistant MCF-7 breast cancer cells. Moreover, TU-BcX-4IC patient-derived primary TNBC cells were included to enhance the translational relevance of our study. We evaluated the effect of pharmacological inhibitors and lentivirus-induced activation or inhibition of the MEK1/2-ERK1/2 and MEK5-ERK5 pathways on cell morphology, E-cadherin, vimentin and ZEB1 expression. Additionally, the effects of pharmacological inhibition of trametinib and XMD8-92 on nuclear localization of ERK1/2 and ERK5, cell migration, proliferation, and spheroid formation were evaluated. Novel compounds that target the MEK1/2 and MEK5 pathways were used in combination with the AKT inhibitor ipatasertib to understand cell-specific responses to kinase inhibition. The results from this study will aid in the design of innovative therapeutic strategies that target cancer metastases.
Hierarchical clustering is a fundamental task often used to discover meaningful structures in data, such as phylogenetic trees, taxonomies of concepts, subtypes of cancer, and cascades of particle decays in particle physics. Typically approximate algorithms are used for inference due to the combinatorial number of possible hierarchical clusterings. In contrast to existing methods, we present novel dynamic-programming algorithms for \emph{exact} inference in hierarchical clustering based on a novel trellis data structure, and we prove that we can exactly compute the partition function, maximum likelihood hierarchy, and marginal probabilities of sub-hierarchies and clusters. Our algorithms scale in time and space proportional to the powerset of $N$ elements which is super-exponentially more efficient than explicitly considering each of the (2N-3)!! possible hierarchies. Also, for larger datasets where our exact algorithms become infeasible, we introduce an approximate algorithm based on a sparse trellis that compares well to other benchmarks. Exact methods are relevant to data analyses in particle physics and for finding correlations among gene expression in cancer genomics, and we give examples in both areas, where our algorithms outperform greedy and beam search baselines. In addition, we consider Dasgupta's cost with synthetic data.
Hierarchical clustering is a critical task in numerous domains. Many approaches are based on heuristics and the properties of the resulting clusterings are studied post hoc. However, in several applications, there is a natural cost function that can be used to characterize the quality of the clustering. In those cases, hierarchical clustering can be seen as a combinatorial optimization problem. To that end, we introduce a new approach based on A* search. We overcome the prohibitively large search space by combining A* with a novel \emph{trellis} data structure. This combination results in an exact algorithm that scales beyond previous state of the art, from a search space with $10^{12}$ trees to $10^{15}$ trees, and an approximate algorithm that improves over baselines, even in enormous search spaces that contain more than $10^{1000}$ trees. We empirically demonstrate that our method achieves substantially higher quality results than baselines for a particle physics use case and other clustering benchmarks. We describe how our method provides significantly improved theoretical bounds on the time and space complexity of A* for clustering.
Triple negative breast cancer (TNBC) is an aggressive subtype of breast cancer with limited targeted therapeutic options. A defining feature of TNBC is the propensity to metastasize and acquire resistance to cytotoxic agents. Mitogen activated protein kinase (MAPK) and extracellular regulated kinase (ERK) signaling pathways have integral roles in cancer development and progression. While MEK5/ERK5 signaling drives mesenchymal and migratory cell phenotypes in breast cancer, the specific mechanisms underlying these actions remain under-characterized. To elucidate the mechanisms through which MEK5 regulates the mesenchymal and migratory phenotype, we generated stably transfected constitutively active MEK5 (MEK5-ca) TNBC cells. Downstream signaling pathways and candidate targets of MEK5-ca cells were based on RNA sequencing and confirmed using qPCR and Western blot analyses. MEK5 activation drove a mesenchymal cell phenotype independent of cell proliferation effects. Transwell migration assays demonstrated MEK5 activation significantly increased breast cancer cell migration. In this study, we provide supporting evidence that MEK5 functions through FRA-1 to regulate the mesenchymal and migratory phenotype in TNBC.
Extracellular signal-regulated kinase (ERK5) is an essential regulator of cancer progression, tumor relapse, and poor patient survival. Epithelial to mesenchymal transition (EMT) is a complex oncogenic process, which drives cell invasion, stemness, and metastases. Activators of ERK5, including mitogen-activated protein kinase 5 (MEK5), tumor necrosis factor α (TNF-α), and transforming growth factor-β (TGF-β), are known to induce EMT and metastases in breast, lung, colorectal, and other cancers. Several downstream targets of the ERK5 pathway, such as myocyte-specific enhancer factor 2c (MEF2C), activator protein-1 (AP-1), focal adhesion kinase (FAK), and c-Myc, play a critical role in the regulation of EMT transcription factors SNAIL, SLUG, and β-catenin. Moreover, ERK5 activation increases the release of extracellular matrix metalloproteinases (MMPs), facilitating breakdown of the extracellular matrix (ECM) and local tumor invasion. Targeting the ERK5 signaling pathway using small molecule inhibitors, microRNAs, and knockdown approaches decreases EMT, cell invasion, and metastases via several mechanisms. The focus of the current review is to highlight the mechanisms which are known to mediate cancer EMT via ERK5 signaling. Several therapeutic approaches that can be undertaken to target the ERK5 pathway and inhibit or reverse EMT and metastases are discussed.
We present a new non-negative matrix factorization model for (0, 1) bounded-support data based on the doubly non-central beta (DNCB) distribution, a generalization of the beta distribution. The expressiveness of the DNCB distribution is particularly useful for modeling DNA methylation datasets, which are typically highly dispersed and multi-modal; however, the model structure is sufficiently general that it can be adapted to many other domains where latent representations of (0, 1) bounded-support data are of interest. Although the DNCB distribution lacks a closed-form conjugate prior, several augmentations let us derive an efficient posterior inference algorithm composed entirely of analytic updates. Our model improves out-of-sample predictive performance on both real and synthetic DNA methylation datasets over stateof-the-art methods in bioinformatics. In addition, our model yields meaningful latent representations that accord with existing biological knowledge.
Triple-negative breast cancer (TNBC) presents a clinical challenge due to the aggressive nature of the disease and a lack of targeted therapies. Constitutive activation of the mitogen-activated protein kinase (MAPK)/extracellular signal-regulated kinase (ERK) pathway has been linked to chemoresistance and metastatic progression through distinct mechanisms, including activation of epithelial-to-mesenchymal transition (EMT) when cells adopt a motile and invasive phenotype through loss of epithelial markers (CDH1), and acquisition of mesenchymal markers (VIM, CDH2). Although MAPK/ERK1/2 kinase inhibitors (MEKi) are useful antitumor agents in a clinical setting, including the Food and Drug Administration (FDA)-approved MEK1,2 dual inhibitors cobimetinib and trametinib, there are limitations to their clinical utility, primarily adaptation of the BRAF pathway and ocular toxicities. The MEK5 (HGNC: MAP2K5) pathway has important roles in metastatic progression of various cancer types, including those of the prostate, colon, bone and breast, and elevated levels of ERK5 expression in breast carcinomas are linked to a worse prognoses in TNBC patients. The purpose of this study is to explore MEK5 regulation of the EMT axis and to evaluate a novel pan-MEK inhibitor on clinically aggressive TNBC cells. Our results show a distinction between the MEK1/2 and MEK5 cascades in maintenance of the mesenchymal phenotype, suggesting that the MEK5 pathway may be necessary and sufficient in EMT regulation while MEK1/2 signaling further sustains the mesenchymal state of TNBC cells. Furthermore, additive effects on MET induction are evident through the inhibition of both MEK1/2 and MEK5. Taken together, these data demonstrate the need for a better understanding of the individual roles of MEK1/2 and MEK5 signaling in breast cancer and provide a rationale for the combined targeting of these pathways to circumvent compensatory signaling and subsequent therapeutic resistance.