Escherichia coli encounters chemically diverse carbon sources, and the observed outputs of its transcriptional regulatory network (TRN) vary with substrate chemistry, metabolic entry route, and growth physiology. Here, we compiled PRECISE-NP881, an 881-condition transcriptome compendium comprising 346 RNA-seq profiles generated for this study during growth on 43 individual carbon sources, and used independent component analysis to quantify condition-specific activities of 137 iModulons, defined here as statistically independent gene-expression modules. We identified 25 carbon-catabolism iModulons and summarized their activity patterns across the 43 substrates into four activity-defined substrate groups. These activity patterns were associated with measured growth rates, substrate chemical classes, central-metabolic entry routes, carbon-normalized stoichiometric yield, and model-estimated proteome allocation. Faster-growing sugar conditions showed low CRP-linked iModulon activity, whereas slower-growing conditions showed elevated, condition-specific activity of CRP-linked and substrate-specific catabolic iModulons. TCA-entry and amino acid-associated conditions were linked with NtrC-1 and Propionate iModulon activities, with targeted knock-out assays supporting the conditional physiological relevance of selected propionyl-CoA-associated genes. A subset of nitrogen-containing, slower-growth conditions with predicted ammonium release induced the cryptic prophage-associated SgcABCEQX iModulon. Projection of an independent glucose starvation/refeeding time-course dataset revealed overlapping dynamics among selected carbon-catabolism iModulons and coordinated changes in growth- and stress-associated TRN outputs. Together, these results provide a systems-level atlas of observed carbon-responsive transcriptional states and systematize carbon physiology at scale.
Pseudomonas putida is a gram-negative bacterial species increasingly utilized in biotechnology due to its robust growth, ability to degrade aromatic compounds, solvent tolerance, and genetic tractability. In this study, we report a comprehensive multi-strain analysis of 164 P. putida strains based on the reconstruction of a pan-putida metabolic network and the formulation of strain-specific genome-scale metabolic models (GEMs). We performed whole-genome sequencing and hybrid assembly for 40 strains, contributing a ~8% increase to the available genomic data for P. putida. Furthermore, high-throughput phenotypic profiling using the Biolog phenotype microarray system for 24 strains on 190 unique carbon sources, along with 15 aromatic compounds not present on Biolog plates, yielded 4,920 unique strain-phenotype measurements. These data were leveraged to curate GEMs for 24 representative strains, including a refined model for strain KT2440, which comprised 1,480 genes and 2,191 metabolites, achieving a prediction accuracy of 91.2% in carbon utilization. Systematic comparison of genomes and GEMs revealed both conserved core pathways and significant allelic and functional divergence across strains, highlighting strain-specific variation in aromatic degradation. While pathways for protocatechuate and phenylacetate degradation were widely conserved, metabolic capabilities for compounds such as ferulate, phenol, and cresols varied markedly, suggesting adaptation to distinct ecological niches. Alleleome analysis of enzymes, such as PcaI and PcaJ, revealed distinct, functionally similar clades, indicating possible convergent evolution or horizontal gene transfer. These results provide computable resources and informative models for selecting P. putida strains with desired traits for biomanufacturing and bioremediation and offer insights into the evolution and phylogeny of the P. putida species.IMPORTANCEPseudomonas putida has become an organism of interest for biotechnological applications, but a species-level understanding of its metabolic diversity remains incomplete. In this study, we analyzed 164 P. putida strains using a combination of genome sequencing, phenotypic profiling, and metabolic modeling. Our results indicate that while many metabolic pathways are conserved, notable differences exist across strains, particularly in aromatic compound degradation. These observations may inform future strain selection and engineering strategies tailored to specific industrial or environmental goals. In addition, the genome-scale models and phenotypic data generated here can serve as a foundation for broader studies of metabolism and functional variation within this species.
Escherichia coli strains are widely used across numerous industrial and biotechnological applications. Yet their performance varies substantially in ways that can not be anticipated from genome annotation. Because transcriptional regulatory networks (TRNs) govern cellular functions such as motility, stress responses, metabolic flexibility, and production efficiency, differences in TRN organization and use may underlie many observed phenotypic differences. To investigate TRN differences between strains, we generated a compendium of 433 matched RNA-Seq profiles for six commonly used industrial E. coli strains (BL21, C, Crooks, MG1655, W, and W3110) and applied iModulon analysis to compare the state of their TRNs under similar growth conditions. This analysis revealed that core regulatory programs with similar functions are wired differently across the strains, and that the strains engage these programs in distinct ways when exposed to the same environmental challenges. Together, these findings highlight transcriptional regulation diversity underlying phenotypic expression among industrial E. coli strains. By providing an integrated view of TRN differences across widely used hosts, this work offers a fundamental basis for interpreting strain-specific behaviors and supports more informed approaches to strain selection and optimization.
The earliest responses of pathogenic bacteria to antibiotics can affect the outcome of an infection. While long-term adaptations have been extensively studied, the immediate transcriptional changes that unfold immediately following antibiotic exposure remain poorly understood. Here, we applied iModulon analysis to time-resolved transcriptomic data from Escherichia coli exposed to subinhibitory concentrations of two antibiotics (ampicillin and ciprofloxacin), capturing transcriptional regulatory changes occurring within the first 30 min of exposure. This analysis proposes an integrated, three-phase response model: an immediate and sustained primary response that broadly activates stress programs, a transient secondary response that restores redox balance, and a tertiary response that supports long-term survival through metabolic remodeling and antibiotic-specific defenses. These results highlight a coordinated and dynamic regulatory strategy describing how metabolic, redox, and stress responses are integrated to manage the physiological challenges of antibiotic stress. By disentangling these overlapping transcriptional regulatory programs, this work offers a genome-scale understanding of how early regulatory programs are engaged immediately after antibiotic exposure. Together, these findings provide a structured framework for characterizing complex transcriptomic responses and generating testable hypotheses about the regulatory logic that shapes the understudied early phase of antibiotic exposure.IMPORTANCEInitial bacterial responses to antibiotics are important for survival and can influence the development of tolerance and resistance. However, this period remains poorly understood, in part, because the transcriptional responses that unfold within minutes of antibiotic exposure are complex and difficult to interpret. In this study, we applied novel data generation and data analytics approaches to resolve the regulatory structure of the initial response of Escherichia coli to two antibiotics. We identify a three-phase process that explains how E. coli coordinates stress responses, maintains redox homeostasis, and initiates downstream protective programs. The novel transcriptomic analytics elucidate independently regulated sets of genes that constitute cellular processes. By identifying the regulatory modules that change over this initial timescale, we can deconvolute the response based on first principles of cellular physiology.
The growth of RNA sequencing (RNA-seq) data accompanied by the development of novel scalable data analytic methods has revealed a deep understanding of the composition of bacterial transcriptomes. This new, first-biological-principles understanding has enabled a novel characterization of the function of the transcriptional regulatory network. Here, we present a single-strain wild-type transcriptomic knowledgebase for the model strain Escherichia coli MG1655. The associated transcriptomic compendium consists of 584 high-quality RNA-seq samples from wild-type E. coli MG1655 generated using a single protocol. These samples range over a wide condition space, including 45 carbon sources and 10 base media. Using independent component analysis, we decomposed the transcriptomic compendium to extract 115 independently modulated sets of genes (iModulons). We find that (i) iModulons explain 75% of variance in the dataset through knowledge enrichment; (ii) 67% of iModulons are associated with single/combined dominant regulators; (iii) iModulon activity profiles of samples can be utilized to elucidate patterns within the transcriptional regulatory network, such as differences in aerobicity; and (iv) the use of transcriptomic data derived from non-wild-type strains results in changes in iModulon gene membership, highlighting the malleability of the transcriptional regulatory network. Altogether, this knowledgebase serves as a resource for multi-scale knowledge mining for transcriptional regulation in E. coli MG1655.
Adaptive laboratory evolution is able to generate microbial strains, which exhibit extreme phenotypes, revealing fundamental biological adaptation mechanisms. Here, we use adaptive laboratory evolution to evolve Escherichia coli strains that grow at temperatures as high as 45.3 °C, a temperature lethal to wild-type cells. The strains adopted a hypermutator phenotype and employed multiple systems-level adaptations that made global analysis of the DNA mutations difficult. Given the challenge at the genomic level, we were motivated to uncover high-temperature tolerance adaptation mechanisms at the transcriptomic level. We employed independently modulated gene set (iModulon) analysis to reveal five transcriptional mechanisms underlying growth at high temperatures. These mechanisms were connected to acquired mutations, changes in transcriptome composition, sensory inputs, phenotypes, and protein structures. They are as follows: (i) downregulation of general stress responses while upregulating the specific heat stress responses, (ii) upregulation of flagellar basal bodies without upregulating motility and upregulation fimbriae, (iii) shift toward anaerobic metabolism, (iv) shift in regulation of iron uptake away from siderophore production, and (v) upregulation of yjfIJKL, a novel heat tolerance operon whose structures we predicted with AlphaFold. iModulons associated with these five mechanisms explain nearly half of all variance in the gene expression in the adapted strains. These thermotolerance strategies reveal that optimal coordination of known stress responses and metabolism can be achieved with a small number of regulatory mutations and may suggest a new role for large protein export systems. Adaptive laboratory evolution with transcriptomic characterization is a productive approach for elucidating and interpreting adaptation to otherwise lethal stresses.
The Staphylococcus aureus clonal complex 8 (CC8) is made up of several subtypes with varying levels of clinical burden; from community-associated methicillin-resistant S. aureus USA300 strains to hospital-associated (HA-MRSA) USA500 strains and ancestral methicillin-susceptible (MSSA) strains. This phenotypic distribution within a single clonal complex makes CC8 an ideal clade to study the emergence of mutations important for antibiotic resistance and community spread. Gene-level analysis comparing USA300 against MSSA and HA-MRSA strains have revealed key horizontally acquired genes important for its rapid spread in the community. However, efforts to define the contributions of point mutations and indels have been confounded by strong linkage disequilibrium resulting from clonal propagation. To break down this confounding effect, we combined genetic association testing with a model of the transcriptional regulatory network (TRN) to find candidate mutations that may have led to changes in gene regulation. First, we used a De Bruijn graph genome-wide association study to enrich mutations unique to the USA300 lineages within CC8. Next, we reconstructed the TRN by using independent component analysis on 670 RNA-sequencing samples from USA300 and non-USA300 CC8 strains which predicted several genes with strain-specific altered expression patterns. Examination of the regulatory region of one of the genes enriched by both approaches, isdH , revealed a 38-bp deletion containing a Fur-binding site and a conserved single-nucleotide polymorphism which likely led to the altered expression levels in USA300 strains. Taken together, our results demonstrate the utility of reconstructed TRNs to address the limits of genetic approaches when studying emerging pathogenic strains.
Bacteria showcase remarkable metabolic diversity and traits, even among strains of the same species. In recent years, a large number of bacterial genomes have been sequenced, leading to the elucidation and documentation of genomic differences and commonalities across and within species. Genome-scale metabolic reconstructions, which are often defined and curated using data from phenotype microarrays, elucidate the differences in metabolic traits resulting from genomic diversity. These microarrays measure cellular respiration on a variety of carbon, nitrogen, phosphorus, and sulfur sources and various stressors and inhibitors over a period of time to determine the metabolic activity of a given strain. Despite their popularity in measuring bacterial metabolic activity and traits, no public databases that allow researchers to warehouse, access, and analyze this information currently exist. Additionally, there are no publicly available tools that allow researchers to view the variance of these metabolic traits across bacterial strains. To address this need, we present Phenotype Microarray Knowledgebase (PMkbase [version 1.0], https://pmkbase.com/), an interactive database that acts as a repository of phenotype microarray (PM) data with integrated sequence information. Binarized activity calls, along with associated kinetic parameters, are made for all metabolic substrates and inhibitors. Users can upload their own data for analysis and visualization and to perform quality checks on their experiments. PMkbase will address an unmet need to track and view bacterial metabolic traits and provide researchers with valuable information to develop metabolic models, enrich pangenomic analyses, and design new experiments.IMPORTANCEBacterial species can be differentiated by their metabolic profiles or the type of nutrients they consume. Interestingly, strains within the same species also display differences in nutrient consumption. Phenotype microarrays are a high-throughput, widely used technology to measure which substrates can be metabolized by various microbial strains and the extent to which inhibitors can affect it. Despite their widespread use, public databases to parse and access this data type at scale do not exist. PMkbase, which contains 9,024 data points for nitrogen substrate utilization, 41,664 data points for carbon substrate utilization, 8,448 data points for phosphorus/sulfur substrate utilization, and 27,264 data points on various antibiotics across three species (Escherichia coli, Pseudomonas putida, and Staphylococcus aureus), has been developed to allow researchers to freely access PM data, along with enriching the data with sequence information.
It has proved challenging to quantitatively relate the proteome to the transcriptome on a per-gene basis. Recent advances in data analytics have enabled a biologically meaningful modularization of the bacterial transcriptome. We thus investigated whether matched datasets of transcriptomes and proteomes from bacteria under diverse conditions could be modularized in the same way to reveal novel relationships between their compositions. We found that; 1) the modules of the proteome and the transcriptome are comprised of a similar list of gene products, 2) the modules in the proteome often represent combinations of modules from the transcriptome, 3) known transcriptional and post-translational regulation is reflected in differences between two sets of modules, allowing for knowledge-mapping when interpreting module functions, and 4) through statistical modeling, absolute proteome allocation can be inferred from the transcriptome alone. Quantitative and knowledge-based relationships can thus be found at the genome-scale between the proteome and transcriptome in bacteria.
Limosilactobacillus reuteri, a probiotic microbe instrumental to human health and sustainable food production, adapts to diverse environmental shifts via dynamic gene expression. We applied the independent component analysis (ICA) to 117 RNA-seq data sets to decode its transcriptional regulatory network (TRN), identifying 35 distinct signals that modulate specific gene sets. Our findings indicate that the ICA provides a qualitative advancement and captures nuanced relationships within gene clusters that other methods may miss. This study uncovers the fundamental properties of L. reuteri’s TRN and deepens our understanding of its arginine metabolism and the co-regulation of riboflavin metabolism and fatty acid conversion. It also sheds light on conditions that regulate genes within a specific biosynthetic gene cluster and allows for the speculation of the potential role of isoprenoid biosynthesis in L. reuteri’s adaptive response to environmental changes. By integrating transcriptomics and machine learning, we provide a system-level understanding of L. reuteri’s response mechanism to environmental fluctuations, thus setting the stage for modeling the probiotic transcriptome for applications in microbial food production. IMPORTANCE We have studied Limosilactobacillus reuteri, a beneficial probiotic microbe that plays a significant role in our health and production of sustainable foods, a type of foods that are nutritionally dense and healthier and have low-carbon emissions compared to traditional foods. Similar to how humans adapt their lifestyles to different environments, this microbe adjusts its behavior by modulating the expression of genes. We applied machine learning to analyze large-scale data sets on how these genes behave across diverse conditions. From this, we identified 35 unique patterns demonstrating how L. reuteri adjusts its genes based on 50 unique environmental conditions (such as various sugars, salts, microbial cocultures, human milk, and fruit juice). This research helps us understand better how L. reuteri functions, especially in processes like breaking down certain nutrients and adapting to stressful changes. More importantly, with our findings, we become closer to using this knowledge to improve how we produce more sustainable and healthier foods with the help of microbes.
The transcriptional regulatory network (TRN) in bacteria is thought to rapidly evolve in response to selection pressures, modulating transcription factor (TF) activities and interactions. In order to probe the limits and mechanisms surrounding the short-term adaptability of the TRN, we generated, evolved, and characterized knockout (KO) strains in Escherichia coli for 11 regulators selected based on measured growth impact on glucose minimal media. All but one knockout strain (Δlrp) were able to recover growth and did so requiring few convergent mutations. We found that the TF knockout adaptations could be divided into four categories: (i) Strains (ΔargR, ΔbasR, Δlon, ΔzntR, and Δzur) that recovered growth without any regulator-specific adaptations, likely due to minimal activity of the regulator on the growth condition, (ii) Strains (ΔcytR, ΔmlrA, and ΔybaO) that recovered growth without TF-specific mutations but with differential expression of regulators with overlapping regulons to the KO'ed TF, (iii) Strains (Δcrp and Δfur) that recovered growth using convergent mutations within their regulatory networks, including regulated promoters and connected regulators, and (iv) Strains (Δlrp) that were unable to fully recover growth, seemingly due to the broad connectivity of the TF within the TRN. Analyzing growth capabilities in evolved and unevolved strains indicated that growth adaptation can restore fitness to diverse substrates often despite a lack of TF-specific mutations. This work reveals the breadth of TRN adaptive mechanisms and suggests these mechanisms can be anticipated based on the network and functional context of the perturbed TFs.
Surveillance programs for managing antimicrobial resistance (AMR) have yielded thousands of genomes suited for data-driven mechanism discovery. We present a workflow integrating pangenomics, gene annotation, and machine learning to identify AMR genes at scale. When applied to 12 species, 27,155 genomes, and 69 drugs, we 1) find AMR gene transfer mostly confined within related species, with 925 genes in multiple species but just eight in multiple phylogenetic classes, 2) demonstrate that discovery-oriented support vector machines outperform contemporary methods at recovering known AMR genes, recovering 263 genes compared to 145 by Pyseer, and 3) identify 142 AMR gene candidates. Validation of two candidates in E. coli BW25113 reveals cases of conditional resistance: ΔcycA confers ciprofloxacin resistance in minimal media with D-serine, and frdD V111D confers ampicillin resistance in the presence of ampC by modifying the overlapping promoter. We expect this approach to be adaptable to other species and phenotypes.
I Abstract Limosilactobacillus reuteri , a probiotic microbe instrumental to human health and sustainable food production, adapts to diverse environmental shifts via dynamic gene expression. We applied independent component analysis to 117 high-quality RNA-seq datasets to decode its transcriptional regulatory network (TRN), identifying 35 distinct signals that modulate specific gene sets. This study uncovers the fundamental properties of L. reuteri’s TRN, deepens our understanding of its arginine metabolism, and the co-regulation of riboflavin metabolism and fatty acid biosynthesis. It also sheds light on conditions that regulate genes within a specific biosynthetic gene cluster and the role of isoprenoid biosynthesis in L. reuteri’s adaptive response to environmental changes. Through the integration of transcriptomics and machine learning, we provide a systems-level understanding of L. reuteri’s response mechanism to environmental fluctuations, thus setting the stage for modeling the probiotic transcriptome for applications in microbial food production. Graphical Abstract Comprehensive iModulon Workflow Overview. Our innovative workflow is grounded in the analysis of the LactoPRECISE compendium, a curated dataset containing 117 internally sequenced RNA-seq samples derived from a diversity of 50 unique conditions, encompassing an extensive range of 13 distinct condition types. We employ the power of Independent Component Analysis (ICA), a cutting-edge machine learning algorithm, to discern the underlying structure of iModulons within this wealth of data. In the subsequent stage of our workflow, the discovered iModulons undergo detailed scrutiny to uncover media-specific regulatory mechanisms governing metabolism, illuminate the context-dependent intricacies of gene expression, and predict pathways leading to the biosynthesis of probiotic secondary metabolites. Our workflow offers an invaluable and innovative lens through which to view probiotic strain design while simultaneously highlighting transformative approaches to data analytics in the field.
ABSTRACT Fast growth phenotypes are achieved through optimal transcriptomic allocation, in which cells must balance tradeoffs in resource allocation between diverse functions. One such balance between stress readiness and unbridled growth in E. coli has been termed the fear versus greed (f/g) tradeoff. Two specific RNA polymerase (RNAP) mutations observed in adaptation to fast growth have been previously shown to affect the f/g tradeoff, suggesting that genetic adaptations may be primed to control f/g resource allocation. Here, we conduct a greatly expanded study of the genetic control of the f/g tradeoff across diverse conditions. We introduced 12 RNA polymerase (RNAP) mutations commonly acquired during adaptive laboratory evolution (ALE) and obtained expression profiles of each. We found that these single RNAP mutation strains resulted in large shifts in the f/g tradeoff primarily in the RpoS regulon and ribosomal genes, likely through modifying RNAP-DNA interactions. Two of these mutations additionally caused condition-specific transcriptional adaptations. While this tradeoff was previously characterized by the RpoS regulon and ribosomal expression, we find that the GAD regulon plays an important role in stress readiness and ppGpp in translation activity, expanding the scope of the tradeoff. A phylogenetic analysis found the greed-related genes of the tradeoff present in numerous bacterial species. The results suggest that the f/g tradeoff represents a general principle of transcriptome allocation in bacteria where small genetic changes can result in large phenotypic adaptations to growth conditions. IMPORTANCE To increase growth, E. coli must raise ribosomal content at the expense of non-growth functions. Previous studies have linked RNAP mutations to this transcriptional shift and increased growth but were focused on only two mutations found in the protein’s central region. RNAP mutations, however, commonly occur over a large structural range. To explore RNAP mutations’ impact, we have introduced 12 RNAP mutations found in laboratory evolution experiments and obtained expression profiles of each. The mutations nearly universally increased growth rates by adjusting said tradeoff away from non-growth functions. In addition to this shift, a few caused condition-specific adaptations. We explored the prevalence of this tradeoff across phylogeny and found it to be a widespread and conserved trend among bacteria.
SummaryRelationships between the genome, transcriptome, and metabolome underlie all evolved phenotypes. However, it has proved difficult to elucidate these relationships because of the high number of variables measured. A recently developed data analytic method for characterizing the transcriptome can simplify interpretation by grouping genes into independently modulated sets (iModulons). Here, we demonstrate how iModulons reveal deep understanding of the effects of causal mutations and metabolic rewiring. We use adaptive laboratory evolution to generateE. colistrains that tolerate high levels of the redox cycling compound paraquat, which produces reactive oxygen species (ROS). We combine resequencing, iModulons, and metabolic models to elucidate six interacting stress tolerance mechanisms: 1) modification of transport, 2) activation of ROS stress responses, 3) use of ROS-sensitive iron regulation, 4) motility, 5) broad transcriptional reallocation toward growth, and 6) metabolic rewiring to decrease NADH production. This work thus reveals the genome-scale systems biology of ROS tolerance.Graphical Abstract
The bacterial respiratory electron transport system (ETS) is branched to allow condition-specific modulation of energy metabolism. There is a detailed understanding of the structural and biochemical features of respiratory enzymes; however, a holistic examination of the system and its plasticity is lacking. Here we generate four strains of Escherichia coli harboring unbranched ETS that pump 1, 2, 3, or 4 proton(s) per electron and characterized them using a combination of synergistic methods (adaptive laboratory evolution, multi-omic analyses, and computation of proteome allocation). We report that: (a) all four ETS variants evolve to a similar optimized growth rate, and (b) the laboratory evolutions generate specific rewiring of major energy-generating pathways, coupled to the ETS, to optimize ATP production capability. We thus define an Aero-Type System (ATS), which is a generalization of the aerobic bioenergetics and is a metabolic systems biology description of respiration and its inherent plasticity.
Staphylococcus aureus is a versatile pathogen with an expanding antibiotic resistance profile. The biology underlying its clinical success emerges from an interplay of many systems such as metabolism and gene regulatory networks.
Respiration requires organisms to have an electron transport system (ETS) for the generation of proton motive force across the membrane that drives ATP synthase. Although the molecular details of the ETS are well studied and constitute textbook material, few studies have appeared to elucidate its systems biology. The most thermodynamically efficient ETS consists of two enzymes, an NADH: quinone oxidoreductase (NqRED) and a dioxygen reductase (O 2 RED), which facilitate the shuttling of electrons from NADH to oxygen. However, evolution has produced variations within ETS which modulate the overall energy efficiency of the system even within the same organism 1–3 . The system-level impact of these variations and their individual physiological optimality remain poorly determined. To mimic varying ETS efficiency we generated four Escherichia coli deletion strains (named ETS-1H, 2H, 3H, and 4H) harboring unbranched ETS variants that pump 1, 2, 3, or 4 proton(s) per electron respectively. We then used a combination of synergistic methods (laboratory evolution, multi-omic analyses, and computation of proteome allocation) to characterize these ETS variants. We found that: (a) all four ETS variants evolved to a similar optimized growth rate, (b) the evolution of ETS variants was enabled by specific rewiring of major energy-generating pathways that couple to the ETS to optimize their ATP production capability, (c) proteome allocation per ATP generated was the same for all the variants, (d) the aero-type, that designates the overall ATP generation strategy 4 of a variant, remained conserved during its laboratory evolution, with the exception of the ETS-4H variant, and (e) integrated computational analysis of then data supported a proton-to-ATP ratio of 10 protons per 3 ATP for ATP synthase for all four ETS variants. We thus have defined the Aero-Type System (ATS) as a generalization of the aerobic bioenergetics, which is descriptive of the metabolic systems biology of respiration and demonstrates its plasticity.
In vitro antibiotic susceptibility testing often fails to accurately predict in vivo drug efficacies, in part due to differences in the molecular composition between standardized bacteriologic media and physiological environments within the body. Here, we investigate the interrelationship between antibiotic susceptibility and medium composition in Escherichia coli K-12 MG1655 as contextualized through machine learning of transcriptomics data. Application of independent component analysis, a signal separation algorithm, shows that complex phenotypic changes induced by environmental conditions or antibiotic treatment are directly traced to the action of a few key transcriptional regulators, including RpoS, Fur, and Fnr. Integrating machine learning results with biochemical knowledge of transcription factor activation reveals medium-dependent shifts in respiration and iron availability that drive differential antibiotic susceptibility. By extension, the data generation and data analytics workflow used here can interrogate the regulatory state of a pathogen under any measured condition and can be applied to any strain or organism for which sufficient transcriptomics data are available. IMPORTANCE Antibiotic resistance is an imminent threat to global health. Patient treatment regimens are often selected based on results from standardized antibiotic susceptibility testing (AST) in the clinical microbiology lab, but these in vitro tests frequently misclassify drug effectiveness due to their poor resemblance to actual host conditions. Prior attempts to understand the combined effects of drugs and media on antibiotic efficacy have focused on physiological measurements but have not linked treatment outcomes to transcriptional responses on a systems level. Here, application of machine learning to transcriptomics data identified medium -dependent responses in key regulators of bacterial iron uptake and respiratory activity. The analytical workflow presented here is scalable to additional organisms and conditions and could be used to improve clinical AST by identifying the key regulatory factors dictating antibiotic susceptibility.
Abstract Many genes in bacterial genomes are of unknown function, often referred to as y-genes. Recently, novel analytic methods have divided bacterial transcriptomes into independently modulated sets of genes (iModulons). Functionally annotated iModulons that contain y-genes lead to testable hypotheses to elucidate y-gene function. Inversely correlated expression of a putative transporter gene, ydhC, relative to purine biosynthetic genes, has led to the hypothesis that it encodes a purine-related transporter and revealed a LysR-family regulator, YdhB, with a predicted 23-bp palindromic binding motif. RNA-Seq analysis of a ydhB knockout mutant confirmed the YdhB-dependent activation of ydhC in the presence of adenosine. The deletion of either the ydhC or the ydhB gene led to a substantially decreased growth rate for E. coli in minimal medium with adenosine as the nitrogen source, as well as with inosine or guanosine. Taken together, we provide clear evidence that YdhB activates the expression of the ydhC gene that encodes a novel purine transporter in E. coli. We propose that the genes ydhB and ydhC be re-named as punR and punC, respectively.