Metagenome sequencing enables the genetic characterization of complex microbial communities. However, determining the activity of isolates within a community presents several challenges, including the wide range of organismal and gene expression abundances, the presence of host RNA, and low microbial biomass at many sites. To address these limitations, we developed "targeted expression analysis sequencing" or TEAL-seq, enabling sensitive species-specific analyses of gene expression using highly multiplexed custom probe pools. For proof of concept, we targeted about 1,700 core and accessory genes of Staphylococcus aureus and S. epidermidis, two key species of the skin microbiome. Two targeting methods were applied to laboratory cultures and human nasal swab specimens. Both methods showed a high degree of specificity, with >90% reads on target, even in the presence of complex microbial or human background DNA/RNA. Targeting using molecular inversion probes demonstrated excellent correlation in inferred expression levels with bulk RNA-seq. Furthermore, we show that a linear pre-amplification step to increase the number of nucleic acids for analysis yielded consistent and predictable results when applied to complex samples and enabled profiling of expression from as little as 1 ng of total RNA. TEAL-seq is much less expensive than bulk metatranscriptomic profiling, enables detection across a greater dynamic range, and uses a strategy that is readily configurable for determining the transcriptional status of organisms in any microbial community.IMPORTANCEThe gene expression patterns of bacteria in microbial communities reflect their activity and interactions with other community members. Measuring gene expression in complex microbiome contexts is challenging, however, due to the large dynamic range of microbial abundances and transcript levels. Here we describe an approach to assessing gene expression for specific species of interest using highly multiplexed pools of targeting probes. We show that an isothermal amplification step enables the profiling of low biomass samples. TEAL-seq should be widely adaptable to the study of microbial activity in natural environments.
E.PathDash facilitates re-analysis of gene expression data from pathogens clinically relevant to chronic respiratory diseases, including a total of 48 studies, 548 samples, and 404 unique treatment comparisons. The application enables users to assess broad biological stress responses at the KEGG pathway or gene ontology level and also provides data for individual genes. E.PathDash reduces the time required to gain access to data from multiple hours per data set to seconds. Users can download high-quality images such as volcano plots and boxplots, differential gene expression results, and raw count data, making it fully interoperable with other tools. Importantly, users can rapidly toggle between experimental comparisons and different studies of the same phenomenon, enabling them to judge the extent to which observed responses are reproducible. As a proof of principle, we invited two cystic fibrosis scientists to use the application to explore scientific questions relevant to their specific research areas. Reassuringly, pathway activation analysis recapitulated results reported in original publications, but it also yielded new insights into pathogen responses to changes in their environments, validating the utility of the application. All software and data are freely accessible, and the application is available at scangeo.dartmouth.edu/EPathDash. IMPORTANCE:Chronic respiratory illnesses impose a high disease burden on our communities and people with respiratory diseases are susceptible to robust bacterial infections from pathogens, including Pseudomonas aeruginosa and Staphylococcus aureus, that contribute to morbidity and mortality. Public gene expression datasets generated from these and other pathogens are abundantly available and an important resource for synthesizing existing pathogenic research, leading to interventions that improve patient outcomes. However, it can take many hours or weeks to render publicly available datasets usable; significant time and skills are needed to clean, standardize, and apply reproducible and robust bioinformatic pipelines to the data. Through collaboration with two microbiologists, we have shown that E.PathDash addresses this problem, enabling them to elucidate pathogen responses to a variety of over 400 experimental conditions and generate mechanistic hypotheses for cell-level behavior in response to disease-relevant exposures, all in a fraction of the time.
Chronic Pseudomonas aeruginosa lung infections are a feature of cystic fibrosis (CF) that many patients experience even with the advent of highly effective modulator therapies. Identifying factors that impact P. aeruginosa in the CF lung could yield novel strategies to eradicate infection or otherwise improve outcomes. To complement published P. aeruginosa studies using laboratory models or RNA isolated from sputum, we analyzed transcripts of strain PAO1 after incubation in sputum from different CF donors prior to RNA extraction. We compared PAO1 gene expression in this "spike-in" sputum model to that for P. aeruginosa grown in synthetic cystic fibrosis sputum medium to determine key genes, which are among the most differentially expressed or most highly expressed. Using the key genes, gene sets with correlated expression were determined using the gene expression analysis tool eADAGE. Gene sets were used to analyze the activity of specific pathways in P. aeruginosa grown in sputum from different individuals. Gene sets that we found to be more active in sputum showed similar activation in published data that included P. aeruginosa RNA isolated from sputum relative to corresponding in vitro reference cultures. In the ex vivo samples, P. aeruginosa had increased levels of genes related to zinc and iron acquisition which were suppressed by metal amendment of sputum. We also found a significant correlation between expression of the H1-type VI secretion system and CFTR corrector use by the sputum donor. An ex vivo sputum model or synthetic sputum medium formulation that imposes metal restriction may enhance future CF-related studies.IMPORTANCEIdentifying the gene expression programs used by Pseudomonas aeruginosa to colonize the lungs of people with cystic fibrosis (CF) will illuminate new therapeutic strategies. To capture these transcriptional programs, we cultured the common P. aeruginosa laboratory strain PAO1 in expectorated sputum from CF patient donors. Through bioinformatic analysis, we defined sets of genes that are more transcriptionally active in real CF sputum compared to a synthetic cystic fibrosis sputum medium. Many of the most differentially active gene sets contained genes related to metal acquisition, suggesting that these gene sets play an active role in scavenging for metals in the CF lung environment which may be inadequately represented in some models. Future studies of P. aeruginosa transcript abundance in CF may benefit from the use of an expectorated sputum model or media supplemented with factors that induce metal restriction.
Pseudomonas aeruginosa is a causative agent of a wide range of infections, including chronic infections associated with cystic fibrosis. These P. aeruginosa infections are difficult to treat and often have negative outcomes.
Overexpression of interleukin (IL-)17 has recently been shown to be associated with a number of pathological conditions. Because IL-17 is found at high levels in the synovial fluid surrounding cartilage in patients with inflammatory arthritis, the present study determined the direct effect of IL-17 on articular cartilage. As shown herein, IL-17 was a direct and potent inducer of matrix breakdown and an inhibitor of matrix synthesis in articular cartilage explants. These effects were mediated in part by leukemia inhibitory factor (LIF), but did not depend on interleukin-1 activity. The mechanism whereby IL-17 induced matrix breakdown in cartilage tissue appeared to be due to stimulation of activity of aggrecanase(s), not matrix metalloproteinase(s). However, IL-17 upregulated expression of matrix metalloproteinase(s) in chondrocytes cultured in monolayer. In vivo, IL-17 induced a phenotype similar to inflammatory arthritis when injected into the intra-articular space of mouse knee joints. Furthermore, a related protein, IL-17E, was found to have catabolic activity on human articular cartilage. This study characterizes the mechanism whereby IL-17 acts directly on cartilage matrix turnover. Such findings have important implications for the treatment of degenerative joint diseases such as arthritis.
AbstractChronicPseudomonas aeruginosalung infections are a distinctive feature of cystic fibrosis (CF) pathology, that challenge adults with CF even with the advent of highly effective modulator therapies. CharacterizingP. aeruginosatranscription in the CF lung and identifying factors that drive gene expression could yield novel strategies to eradicate infection or otherwise improve outcomes. To complement publishedP. aeruginosagene expression studies in laboratory culture models designed to model the CF lung environment, we employed an ex vivo sputum model in which laboratory strain PAO1 was incubated in sputum from different CF donors. As part of the analysis, we compared PAO1 gene expression in this “spike-in” sputum model to that forP. aeruginosagrown in artificial sputum medium (ASM). Analyses focused on genes that were differentially expressed between sputum and ASM and genes that were most highly expressed in sputum. We present a new approach that used sets of genes with correlated expression, identified by the gene expression analysis tool eADAGE, to analyze the differential activity of pathways inP. aeruginosagrown in CF sputum from different individuals. A key characteristic ofP. aeruginosagrown in expectorated CF sputum was related to zinc and iron acquisition, but this signal varied by donor sputum. In addition, a significant correlation betweenP. aeruginosaexpression of the H1-type VI secretion system and corrector use by the sputum donor was observed. These methods may be broadly useful in looking for variable signals across clinical samples.ImportanceIdentifying the gene expression programs used byPseudomonas aeruginosato colonize the lungs of people with cystic fibrosis (CF) will illuminate new therapeutic strategies. To capture these transcriptional programs, we cultured the commonP. aeruginosalaboratory strain PAO1 in expectorated sputum from CF patient donors. Through bioinformatics analysis, we defined sets of genes that are more transcriptionally active in real CF sputum compared to artificial sputum media (ASM). Many of the most differentially active gene sets contained genes related to metal acquisition, suggesting that these gene sets play an active role in scavenging for metals in the CF lung environment which is inadequately represented in ASM. Future studies ofP. aeruginosatranscription in CF may benefit from the use of an expectorated sputum model or modified forms of ASM supplemented with metals.
Strains of Pseudomonas aeruginosa , an opportunistic pathogen that causes difficult to treat infections, have significant genomic heterogeneity including the presence of diverse accessory genes that are only present in some strains or clades. Both core genes, which are conserved across strains, and accessory genes have been associated with traits such as biofilm formation and virulence. Much of what we know about core and accessory gene content comes from genome analyses. Here, we use a newly assembled transcriptome compendium to analyze the transcriptional patterns of core and accessory gene expression in PAO1 and PA14 strains across thousands of samples from hundreds of distinct experiments. We found that a subset of core genes were stable, having consistent correlated expression patterns across samples regardless of strain background, with a focus on strains PAO1 and PA14. These stable core genes had fewer co-expressed neighbors that were accessory genes. Significance Pseudomonas aeruginosa is a ubiquitous pathogen. There is a lot of diversity amongst P. aeruginosa strains, some which are clinically relevant. Understanding how these different strain-level traits manifest is important for identifying targets that regulate different traits of interest. With the availability of a PAO1-mapped and PA14-mapped RNA-seq compendium, which contain hundreds of strains, it is now possible to examine the effect of different strains on expression, which can mediate different traits. In this study we developed an approach to compare expression profiles across different P. aeruginosa gene groups – core and accessory genes. This approach revealed a subset of core genes with different transcriptional patterns across strains, which could contribute to trait differences.
Genome-wide transcriptome profiling identifies genes that are prone to differential expression (DE) across contexts, as well as genes with changes specific to the experimental manipulation. Distinguishing genes that are specifically changed in a context of interest from common differentially expressed genes (DEGs) allows more efficient prediction of which genes are specific to a given biological process under scrutiny. Currently, common DEGs or pathways can only be identified through the laborious manual curation of experiments, an inordinately time-consuming endeavor. Here we pioneer an approach, Specific cOntext Pattern Highlighting In Expression data (SOPHIE), for distinguishing between common and specific transcriptional patterns using a generative neural network to create a background set of experiments from which a null distribution of gene and pathway changes can be generated. We apply SOPHIE to diverse datasets including those from human, human cancer, and bacterial pathogen Pseudomonas aeruginosa. SOPHIE identifies common DEGs in concordance with previously described, manually and systematically determined common DEGs. Further molecular validation indicates that SOPHIE detects highly specific but low-magnitude biologically relevant transcriptional changes. SOPHIE's measure of specificity can complement log2 fold change values generated from traditional DE analyses. Forexample, by filtering the set of DEGs, one can identify genes that are specifically relevant to the experimental condition of interest. Consequently, these results can inform future research directions. All scripts used in these analyses are available at https://github.com/greenelab/generic-expression-patterns. Users can access https://github.com/greenelab/sophie to run SOPHIE on their own data.
Researchers studying cystic fibrosis (CF) pathogens have produced numerous RNA-seq datasets which are available in the gene expression omnibus (GEO). Although these studies are publicly available, substantial computational expertise and manual effort are required to compare similar studies, visualize gene expression patterns within studies, and use published data to generate new experimental hypotheses. Furthermore, it is difficult to filter available studies by domain-relevant attributes such as strain, treatment, or media, or for a researcher to assess how a specific gene responds to various experimental conditions across studies. To reduce these barriers to data re-analysis, we have developed an R Shiny application called CF-Seq, which works with a compendium of 128 studies and 1,322 individual samples from 13 clinically relevant CF pathogens. The application allows users to filter studies by experimental factors and to view complex differential gene expression analyses at the click of a button. Here we present a series of use cases that demonstrate the application is a useful and efficient tool for new hypothesis generation. (CF-Seq: http://scangeo.dartmouth.edu/CFSeq/ )
A gene expression compendium is a heterogeneous collection of gene expression experiments assembled from data collected for diverse purposes. The widely varied experimental conditions and genetic backgrounds across samples creates a tremendous opportunity for gaining a systems level understanding of the transcriptional responses that influence phenotypes. Variety in experimental design is particularly important for studying microbes, where the transcriptional responses integrate many signals and demonstrate plasticity across strains including response to what nutrients are available and what microbes are present. Advances in high-throughput measurement technology have made it feasible to construct compendia for many microbes. In this review we discuss how these compendia are constructed and analyzed to reveal transcriptional patterns.
Motivation: In the past two decades, scientists in different laboratories have assayed gene expression from millions of samples. These experiments can be combined into compendia and analyzed collectively to extract novel biological patterns. Technical variability, or "batch effects," may result from combining samples collected and processed at different times and in different settings. Such variability may distort our ability to extract true underlying biological patterns. As more integrative analysis methods arise and data collections get bigger, we must determine how technical variability affects our ability to detect desired patterns when many experiments are combined. Objective: We sought to determine the extent to which an underlying signal was masked by technical variability by simulating compendia comprising data aggregated across multiple experiments. Method: We developed a generative multi-layer neural network to simulate compendia of gene expression experiments from large-scale microbial and human datasets. We compared simulated compendia before and after introducing varying numbers of sources of undesired variability. Results: The signal from a baseline compendium was obscured when the number of added sources of variability was small. Applying statistical correction methods rescued the underlying signal in these cases. However, as the number of sources of variability increased, it became easier to detect the original signal even without correction. In fact, statistical correction reduced our power to detect the underlying signal. Conclusion: When combining a modest number of experiments, it is best to correct for experiment-specific noise. However, when many experiments are combined, statistical correction reduces our ability to extract underlying patterns.
Pseudomonas aeruginosa and Candida albicans are opportunistic pathogens whose interactions involve the secreted products ethanol and phenazines. Here, we describe the role of ethanol in mixed-species co-cultures by dual-seq analyses. P. aeruginosa and C. albicans transcriptomes were assessed after growth in mono-culture or co-culture with either ethanol-producing C. albicans or a C. albicans mutant lacking the primary ethanol dehydrogenase, Adh1. Analysis of the RNA-Seq data using KEGG pathway enrichment and eADAGE methods revealed several P. aeruginosa responses to C. albicans-produced ethanol including the induction of a non-canonical low-phosphate response regulated by PhoB. C. albicans wild type, but not C. albicans adh1Δ/Δ, induces P. aeruginosa production of 5-methyl-phenazine-1-carboxylic acid (5-MPCA), which forms a red derivative within fungal cells and exhibits antifungal activity. Here, we show that C. albicans adh1Δ/Δ no longer activates P. aeruginosa PhoB and PhoB-regulated phosphatase activity, that exogenous ethanol complements this defect, and that ethanol is sufficient to activate PhoB in single-species P. aeruginosa cultures at permissive phosphate levels. The intersection of ethanol and phosphate in co-culture is inversely reflected in C. albicans; C. albicans adh1Δ/Δ had increased expression of genes regulated by Pho4, the C. albicans transcription factor that responds to low phosphate, and Pho4-dependent phosphatase activity. Together, these results show that C. albicans-produced ethanol stimulates P. aeruginosa PhoB activity and 5-MPCA-mediated antagonism, and that both responses are dependent on local phosphate concentrations. Further, our data suggest that phosphate scavenging by one species improves phosphate access for the other, thus highlighting the complex dynamics at play in microbial communities.
Pseudomonas aeruginosa is often found with bacteria and fungi that produce fermentation products, including ethanol. At concentrations similar to those produced by environmental microbes, we found that ethanol stimulated expression of trehalose-biosynthetic genes and cellular levels of trehalose, a disaccharide that protects against environmental stresses. The induction of trehalose by ethanol required the alternative sigma factor AlgU through DksA- and SpoT-dependent (p)ppGpp. Trehalose accumulation also required AHL quorum sensing and occurred only in post-exponential-phase cultures. This work highlights how cells integrate cell density and growth cues in their responses to products made by other microbes and reveals a new role for (p)ppGpp in the regulation of AlgU activity.
ABSTRACT The Pseudomonas fluorescens genome encodes more than 50 proteins predicted to be involved in c-di-GMP signaling. Here, we demonstrated that, tested across 188 nutrients, these enzymes and effectors appeared capable of impacting biofilm formation. Transcriptional analysis of network members across ∼50 nutrient conditions indicates that altered gene expression can explain a subset of but not all biofilm formation responses to the nutrients. Additional organization of the network is likely achieved through physical interaction, as determined via probing ∼2,000 interactions by bacterial two-hybrid assays. Our analysis revealed a multimodal regulatory strategy using combinations of ligand-mediated signals, protein-protein interaction, and/or transcriptional regulation to fine-tune c-di-GMP-mediated responses. These results create a profile of a large c-di-GMP network that is used to make important cellular decisions, opening the door to future model building and the ability to engineer this complex circuitry in other bacteria. IMPORTANCE Cyclic diguanylate (c-di-GMP) is a key signaling molecule regulating bacterial biofilm formation, and many microbes have up to dozens of proteins that make, break, or bind this dinucleotide. A major open issue in the field is how signaling specificity is conferred in the unpartitioned space of a bacterial cell. Here, we took a systems approach, using mutational analysis, transcriptional studies, and bacterial two-hybrid analysis to interrogate this network. We found that a majority of enzymes are capable of impacting biofilm formation in a context-dependent manner, and we revealed examples of two or more modes of regulation (i.e., transcriptional control with protein-protein interaction) being utilized to generate an observable impact on biofilm formation.
ABSTRACT Iodobacter species are among a number of freshwater Gram-negative violacein-producing bacteria. Janthinobacterium lividum and Chromobacterium violaceum have had their whole genomes sequenced and annotated. This is the first report of a draft whole-genome sequence of a violacein-producing Iodobacter strain that was isolated from the Hudson Valley watershed.
Background Investigators often interpret genome-wide data by analyzing the expression levels of genes within pathways. While this within-pathway analysis is routine, the products of any one pathway can affect the activity of other pathways. Past efforts to identify relationships between biological processes have evaluated overlap in knowledge bases or evaluated changes that occur after specific treatments. Individual experiments can highlight condition-specific pathway-pathway relationships; however, constructing a complete network of such relationships across many conditions requires analyzing results from many studies. Results We developed PathCORE-T framework by implementing existing methods to identify pathway-pathway transcriptional relationships evident across a broad data compendium. PathCORE-T is applied to the output of feature construction algorithms; it identifies pairs of pathways observed in features more than expected by chance as functionally co-occurring . We demonstrate PathCORE-T by analyzing an existing eADAGE model of a microbial compendium and building and analyzing NMF features from the TCGA dataset of 33 cancer types. The PathCORE-T framework includes a demonstration web interface, with source code, that users can launch to (1) visualize the network and (2) review the expression levels of associated genes in the original data. PathCORE-T creates and displays the network of globally co-occurring pathways based on features observed in a machine learning analysis of gene expression data. Conclusions The PathCORE-T framework identifies transcriptionally co-occurring pathways from the results of unsupervised analysis of gene expression data and visualizes the relationships between pathways as a network. PathCORE-T recapitulated previously described pathway-pathway relationships and suggested experimentally testable additional hypotheses that remain to be explored.
Investigation of the Hudson Valley watershed reveals many violacein-producing bacteria. These are of interest for their biotherapeutic potential in treating chytrid infections of amphibians. The draft whole-genome sequences for seven Janthinobacterium isolates with a variety of phenotypes are provided in this study.
Cross-experiment comparisons in public data compendia are challenged by unmatched conditions and technical noise. The ADAGE method, which performs unsupervised integration with denoising autoencoder neural networks, can identify biological patterns, but because ADAGE models, like many neural networks, are over-parameterized, different ADAGE models perform equally well. To enhance model robustness and better build signatures consistent with biological pathways, we developed an ensemble ADAGE (eADAGE) that integrated stable signatures across models. We applied eADAGE to a compendium of Pseudomonas aeruginosa gene expression profiling experiments performed in 78 media. eADAGE revealed a phosphate starvation response controlled by PhoB in media with moderate phosphate and predicted that a second stimulus provided by the sensor kinase, KinB, is required for this PhoB activation. We validated this relationship using both targeted and unbiased genetic approaches. eADAGE, which captures stable biological patterns, enables cross-experiment comparisons that can highlight measured but undiscovered relationships.
While the large sets of publicly available gene expression data contain substantial information about relationships between mRNA expression profiles and genetic background, environment, and cellular state, cross experiment comparisons of public data are challenged by technical noise that masks biological signals. We previously showed that one could reveal biological signatures within compendia of expression data using an unsupervised neural network algorithm, called ADAGE, which excels at detecting patterns in noisy datasets. Here, we show that the generation and integration of multiple ADAGE models, resulting in an ensemble ADAGE (eADAGE), better identified biological pathways. For the bacterium Pseudomonas aeruginosa, our analysis found that on the order of 1000 samples were needed to build pathway-level gene expression signatures. The P. aeruginosa gene expression compendium contains experiments performed in 78 different media, and we used eADAGE to identify expression signatures associated with medium-type across multiple experiments. We identified a subset of media, including several complex media that were not designed to limit phosphate, in which P. aeruginosa exhibited a phosphate starvation response controlled by PhoB. Furthermore, while it was expected that PhoB activates the phosphate starvation response in low phosphate, our analyses found that PhoB was also active in moderate phosphate concentrations and predicted that activity required a second stimulus provided by a sensor kinase, KinB, which was validated in subsequent experiments including a screen of a histidine kinase knock out collection confirmed the specificity of its role in the activation of the Pho regulon. Algorithms that extract biological signal from large collections of public gene expression data, such as eADAGE, can highlight opportunities to discover mechanisms that are currently unrecognized from public data.
Abundant public expression data capture gene expression across diverse conditions. These steady state mRNA measurements could reveal the transcriptional consequences of cells' genetic backgrounds or their responses to the environment. However, public data remain relatively untapped, in part because extracting biological signal as opposed to technical noise remains challenging. Here we introduce a procedure, termed eADAGE, that performs unsupervised integration of public expression data using an ensemble of neural networks as well as heuristics that, given a dataset, help users identify an appropriate level of model complexity. This ensemble modeling approach captures biological pathways more clearly than existing methods, enabling analyses that span entire public gene expression compendia such as that for the bacterium Pseudomonas aeruginosa. These analyses reveal a previously undiscovered feature of the phosphate starvation response apparent in public data: a sensor kinase, KinB, that is required for full activation of the response to phosphate at intermediate concentrations. Our molecular validation experiments confirm this role of KinB and our screen of a histidine kinase knock out collection confirmed the prediction's specificity. Public data are captured from a broad range of conditions in diverse organism backgrounds and may provide a unique opportunity to identify these subtle and context-specific regulatory interactions. Algorithms that extract biological signal from these data, such as eADAGE, can highlight opportunities to discover mechanisms that are apparent from but unrealized in public data.