Because large brains are energetically expensive, they are associated with metabolic traits that facilitate energy availability across vertebrates. However, the biological underpinnings driving these traits are not known. Given its role in regulating host metabolism in disease studies, we hypothesized that the gut microbiome contributes to variation in normal cross-vertebrate species differences in metabolism, including those associated with the brain's energetic requirements. By inoculating germ-free mice with the gut microbiota (GM) of three primate species - two with relatively larger brains and one with a smaller brain - we demonstrated that the GM of larger-brained primates shifts host metabolism towards energy use and production, while that of smaller-brained primates stimulates energy storage in adipose tissues. Our findings establish a causal role of the GM in normal cross-host species differences in metabolism associated with relative brain size and suggest that the GM may have been an important facilitator of metabolic changes during human evolution that supported encephalization.
Autoimmune diseases such as systemic lupus erythematosus (SLE) display a strong female bias. Although sex hormones have been associated with protecting males from autoimmunity, the molecular mechanisms are incompletely understood. Here we report that androgen receptor (AR) expressed in T cells regulates genes involved in T cell activation directly, or indirectly via controlling other transcription factors. T cell-specific deletion of AR in mice leads to T cell activation and enhanced autoimmunity in male mice. Mechanistically, Ptpn22, a phosphatase and negative regulator of T cell receptor signaling, is downregulated in AR-deficient T cells. Moreover, a conserved androgen-response element is found in the regulatory region of Ptpn22 gene, and the mutation of this transcription element in non-obese diabetic mice increases the incidence of spontaneous and inducible diabetes in male mice. Lastly, Ptpn22 deficiency increases the disease severity of male mice in a mouse model of SLE. Our results thus implicate AR-regulated genes such as PTPN22 as potential therapeutic targets for autoimmune diseases.
B cells are a critical component of the adaptive immune system. Single cell RNA-sequencing (scRNA-seq) has allowed for both profiling of B cell receptor (BCR) sequences and gene expression. However, understanding the adaptive and evolutionary mechanisms of B cells in response to specific stimuli remains a significant challenge in the field of immunology. We introduce a new method, TRIBAL, which aims to infer the evolutionary history of clonally related B cells from scRNA-seq data. The key insight of TRIBAL is that inclusion of isotype data into the B cell lineage inference problem is valuable for reducing phylogenetic uncertainty that arises when only considering the receptor sequences. Consequently, the TRIBAL inferred B cell lineage trees jointly capture the somatic mutations introduced to the B cell receptor during affinity maturation and isotype transitions during class switch recombination. In addition, TRIBAL infers isotype transition probabilities that are valuable for gaining insight into the dynamics of class switching. Via in silico experiments, we demonstrate that TRIBAL infers isotype transition probabilities with the ability to distinguish between direct versus sequential switching in a B cell population. This results in more accurate B cell lineage trees and corresponding ancestral sequence and class switch reconstruction compared to competing methods. Using real-world scRNA-seq datasets, we show that TRIBAL recapitulates expected biological trends in a model affinity maturation system. Furthermore, the B cell lineage trees inferred by TRIBAL were equally plausible for the BCR sequences as those inferred by competing methods but yielded lower entropic partitions for the isotypes of the sequenced B cell. Thus, our method holds the potential to further advance our understanding of vaccine responses, disease progression, and the identification of therapeutic antibodies.
Commensal microbes have the capacity to affect development and severity of autoimmune diseases. Germ-free (GF) animals have proven to be a fine tool to obtain definitive answers to the queries about the microbial role in these diseases. Moreover, GF and gnotobiotic animals can be used to dissect the complex symptoms and determine which are regulated (enhanced or attenuated) by microbes. These include disease manifestations that are sex biased. Here, we review comparative analyses conducted between GF and Specific-Pathogen Free (SPF) mouse models of autoimmunity. We present data from the B6;NZM-Sle1NZM2410/AegSle2NZM2410/AegSle3NZM2410/Aeg-/LmoJ (B6.NZM) mouse model of systemic lupus erythematosus (SLE) characterized by multiple measurable features. We compared the severity and sex bias of SPF, GF, and ex-GF mice and found variability in the severity and sex bias of some manifestations. Colonization of GF mice with the microbiotas taken from B6.NZM mice housed in two independent institutions variably affected severity and sexual dimorphism of different parameters. Thus, microbes regulate both the severity and sexual dimorphism of select SLE traits. The sensitivity of particular trait to microbial influence can be used to further dissect the mechanisms driving the disease. Our results demonstrate the complexity of the problem and open avenues for further investigations.
Single-cell RNA sequencing (scRNA-seq) enables comprehensive characterization of the micro-evolutionary processes of B cells during an adaptive immune response, capturing features of somatic hypermutation (SHM) and class switch recombination (CSR). Existing phylogenetic approaches for reconstructing B cell evolution have primarily focused on the SHM process alone. Here, we present tree inference of B cell clonal lineages (TRIBAL), an algorithm designed to optimally reconstruct the evolutionary history of B cell clonal lineages undergoing both SHM and CSR from scRNA-seq data. Through simulations, we demonstrate that TRIBAL produces more comprehensive and accurate B cell lineage trees compared to existing methods. Using real-world datasets, TRIBAL successfully recapitulates expected biological trends in a model affinity maturation system while reconstructing evolutionary histories with more parsimonious class switching than state-of-the-art methods. Thus, TRIBAL significantly improves B cell lineage tracing, useful for modeling vaccine responses, disease progression, and the identification of therapeutic antibodies.
Diet and commensals can affect the development of autoimmune diseases like type 1 diabetes (T1D). However, whether dietary interventions are microbe-mediated was unclear. We found that a diet based on hydrolyzed casein (HC) as a protein source protects non-obese diabetic (NOD) mice in conventional and germ-free (GF) conditions via improvement in the physiology of insulin-producing cells to reduce autoimmune activation. The addition of gluten (a cereal protein complex associated with celiac disease) facilitates autoim-munity dependent on microbial proteolysis of gluten: T1D develops in GF animals monocolonized with Entero-coccus faecalis harboring secreted gluten-digesting proteases but not in mice colonized with protease defi-cient bacteria. Gluten digestion by E. faecalis generates T cell-activating peptides and promotes innate immunity by enhancing macrophage reactivity to lipopolysaccharide (LPS). Gnotobiotic NOD Toll4-negative mice monocolonized with E. faecalis on an HC + gluten diet are resistant to T1D. These findings provide insights into strategies to develop dietary interventions to help protect humans against autoimmunity.
Background and aims Normal gestation involves reprogramming of maternal gut microbiome (GM) that may contribute to maternal metabolic changes by unclear mechanisms. This study aimed to understand the mechanistic underpinnings of GM – maternal metabolism interaction. Methods The GM and plasma metabolome of CD1, NIH-Swiss and C57BL/6J mice were analyzed using 16S rRNA sequencing and untargeted LC-MS throughout gestation and postpartum. Pharmacologic and genetic knockout mouse models were used to identify the role of indoleamine 2,3-dioxygenase (IDO1) in pregnancy-associated insulin resistance (IR). Involvement of gestational GM in the process was studied using fecal microbial transplants (FMT). Results Significant variation in gut microbial alpha diversity occurred throughout pregnancy. Enrichment in gut bacterial taxa was mouse strain and pregnancy time-point specific, with species enriched at gestation day 15/19 (G15/19), a point of heightened IR, distinct from those enriched pre- or post- pregnancy. Untargeted and targeted metabolomics revealed elevated plasma kynurenine at G15/19 in all three mouse strains. IDO1, the rate limiting enzyme for kynurenine production, had increased intestinal expression at G15, which was associated with mild systemic and gut inflammation. Pharmacologic and genetic inhibition of IDO1 inhibited kynurenine levels and reversed pregnancy-associated IR. FMT revealed that IDO1 induction and local kynurenine levels effects on IR derive from the GM in both mouse and human pregnancy. Conclusions GM changes accompanying pregnancy shift IDO1-dependent tryptophan metabolism toward kynurenine production, intestinal inflammation and gestational IR, a phenotype reversed by genetic deletion or inhibition of IDO1.
The accuracy of methods for assembling transcripts from short-read RNA sequencing data is limited by the lack of long-range information. Here we introduce Ladder-seq, an approach that separates transcripts according to their lengths before sequencing and uses the additional information to improve the quantification and assembly of transcripts. Using simulated data, we show that a kallisto algorithm extended to process Ladder-seq data quantifies transcripts of complex genes with substantially higher accuracy than conventional kallisto. For reference-based assembly, a tailored scheme based on the StringTie2 algorithm reconstructs a single transcript with 30.8% higher precision than its conventional counterpart and is more than 30% more sensitive for complex genes. For de novo assembly, a similar scheme based on the Trinity algorithm correctly assembles 78% more transcripts than conventional Trinity while improving precision by 78%. In experimental data, Ladder-seq reveals 40% more genes harboring isoform switches compared to conventional RNA sequencing and unveils widespread changes in isoform usage upon m6A depletion by Mettl14 knockout.
Pregnancy is a dynamic state with multiple metabolic changes occurring including insulin resistance. Gestational diabetes mellitus (GDM), a form of diabetes that appears during pregnancy, develops if metabolic aberrations occur, in particular, in normal pregnancy-induced insulin resistance. Multi-omics is a powerful approach for uncovering the mechanisms driving metabolic change in different physiologic and pathologic states. A recent study demonstrated that the gestational gut microbiome mediates pregnancy metabolic adaptations through effects on gut indoleamine-2,3 dioxygenase 1 activity and the production of kynurenine. Using the dataset generated from this highly controlled study, we performed a comprehensive analysis of the pregnancy-specific physiological and metabolic profiles, 16S rRNA microbiome, and plasma untargeted LC-MS metabolome data. To facilitate the utilization of these analysis results by other researchers, we developed MOMMI-MP, a database that provides an easy-to-use platform to browse and search differential abundant microbial taxa and metabolites, and to examine metabolic pathways. The datasets consist of data collected from 3 genetically diverse strains of mice (C57BL/6J, CD1, and NIH-Swiss) over 6 time points during the gestational (days 0, 10, 15, and 19 during gestation) and postpartum (days 3 and 20 after delivery) states, totaling 180 samples for each strain. The computational results are presented in various tables and plots, and organized in MOMMI-MP to empower exploratory analyses by other researchers. In conclusion, MOMMI-MP is a resource to facilitate the investigation of novel mechanisms governing metabolic changes during pregnancy.
The alterations in myometrial biology during labor are not well understood. The myometrium is the contractile portion of the uterus and contributes to labor, a process that may be regulated by the steroid hormone progesterone. Thus, human myometrial tissues from term pregnant in-active-labor (TIL) and term pregnant not-in-labor (TNIL) subjects were used for genome-wide analyses to elucidate potential future preventive or therapeutic targets involved in the regulation of labor. Using myometrial tissues directly subjected to RNA sequencing (RNA-seq), progesterone receptor (PGR) chromatin immunoprecipitation sequencing (ChIP-seq), and histone modification ChIP-seq, we profiled genome-wide changes associated with gene expression in myometrial smooth muscle tissue in vivo. In TIL myometrium, PGR predominantly occupied promoter regions, including the classical progesterone response element, whereas it bound mainly to intergenic regions in TNIL myometrial tissue. Differential binding analysis uncovered over 1700 differential PGR-bound sites between TIL and TNIL, with 1361 sites gained and 428 lost in labor. Functional analysis identified multiple pathways involved in cAMP-mediated signaling enriched in labor. A three-way integration of the data for ChIP-seq, RNA-seq, and active histone marks uncovered the following genes associated with PGR binding, transcriptional activation, and altered mRNA levels: ATP11A, CBX7, and TNS1. In vitro studies showed that ATP11A, CBX7, and TNS1 are progesterone responsive. We speculate that these genes may contribute to the contractile phenotype of the myometrium during various stages of labor. In conclusion, we provide novel labor-associated genome-wide events and PGR-target genes that can serve as targets for future mechanistic studies.
Metagenomic studies of the microbiome community have revealed associations of the microbiome community to host disease state. The detection of these associations can rely on statistical analyses identifying differentially abundant taxa between diseased and healthy populations. Accurate prediction of the host phenotype from a metagenomic sample and identification of the associated microbial markers are important in understanding potential host-microbiome interactions related to disease initiation and progression. However, associations of individual microbes to a particular disease have shown contradictory results in past studies, possibly due to dynamic and complex natures of different microbes. To handle the complex nature of the microbiome, machine learning methods have begun being employed. Machine learning algorithms are a set of methods in which a model learns intrinsic patterns in data and use them to predict labels of data. In this chapter, we introduce the commonly used machine learning methods in metagenomic studies. We show readers how to use the currently available tools found in Python libraries. Our purpose is to demonstrate the proper training and analysis of machine learning models for microbiome researchers, who may not have experience in machine learning or Python programming.
The advance in microbiome and metabolome studies has generated rich omics data revealing the involvement of the microbial community in host disease pathogenesis through interactions with their host at a metabolic level. However, the computational tools to uncover these relationships are just emerging. Here, we present MiMeNet, a neural network framework for modeling microbe-metabolite relationships. Using ten iterations of 10-fold cross-validation on three paired microbiome-metabolome datasets, we show that MiMeNet more accurately predicts metabolite abundances (mean Spearman correlation coefficients increase from 0.108 to 0.309, 0.276 to 0.457, and -0.272 to 0.264) and identifies more well-predicted metabolites (increase in the number of well-predicted metabolites from 198 to 366, 104 to 143, and 4 to 29) compared to state-of-art linear models for individual metabolite predictions. Additionally, we demonstrate that MiMeNet can group microbes and metabolites with similar interaction patterns and functions to illuminate the underlying structure of the microbe-metabolite interaction network, which could potentially shed light on uncharacterized metabolites through "Guilt by Association". Our results demonstrated that MiMeNet is a powerful tool to provide insights into the causes of metabolic dysregulation in disease, facilitating future hypothesis generation at the interface of the microbiome and metabolomics.
Immune checkpoint blockade (ICB) is widely used to treat non-small cell lung cancer (NSCLC) patients and works by inhibiting the PD-1/PDL1 axis to reinvigorate exhausted T cells. In the prevailing model, the primary mechanism for direct tumor cell killing occurs through cytotoxic CD8+ T cells. Here, we use single-cell multi-omic profiling (a combination of single-cell RNA sequencing, TCR sequencing, and surface protein profiling) in 10 NSCLC patients to identify a population of CD4+ T cells that are tumor-infiltrating, clonally expanded, and express a cytotoxic gene program. Concordantly, we found in the same patients a subpopulation of tumor cells with elevated HLA class II expression, suggesting a mechanism for tumor-mediated antigen presentation to CD4+ T cells. Finally, we show that a cytotoxic CD4 gene signature is associated with improved progression-free survival in a cohort of 180 NSCLC patients treated with ICB regimens, including those with loss of heterozygosity at the HLA class I locus. Overall, these results suggest a model where cytotoxic CD4+ T cells can perform direct tumor cell killing in a class II restricted manner and that their presence is associated with favorable ICB outcomes in NSCLC. Citation Format: Denise Lau, Sonal Khare, Derek Reiman, Tim Rand, Ameen A. Salahudeen, Aly Khan. Cytotoxic CD4+ T cells contribute to anti-tumor immune responses in NSCLC [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2021; 2021 Apr 10-15 and May 17-21. Philadelphia (PA): AACR; Cancer Res 2021;81(13_Suppl):Abstract nr 1487.
Accurate prediction of the host phenotypes from a microbial sample and identification of the associated microbial markers are important in understanding the impact of the microbiome on the pathogenesis and progression of various diseases within the host. A deep learning tool, PopPhy-CNN, has been developed for the task of predicting host phenotypes using a convolutional neural network (CNN). By representing samples as annotated taxonomic trees and further representing these trees as matrices, PopPhy-CNN utilizes the CNN's innate ability to explore locally similar microbes on the taxonomic tree. Furthermore, PopPhy-CNN can be used to evaluate the importance of each taxon in the prediction of host status. Here, we describe the underlying methodology, architecture, and core utility of PopPhy-CNN. We also demonstrate the use of PopPhy-CNN on a microbial dataset.
Single cell RNA sequencing (scRNAseq) can be used to infer a temporal ordering of cellular states. Current methods for the inference of cellular trajectories rely on unbiased dimensionality reduction techniques. However, such biologically agnostic ordering can prove difficult for modeling complex developmental or differentiation processes. The cellular heterogeneity of dynamic biological compartments can result in sparse sampling of key intermediate cell states. To overcome these limitations, we develop a supervised machine learning framework, called Pseudocell Tracer, which infers trajectories in pseudospace rather than in pseudotime. The method uses a supervised encoder, trained with adjacent biological information, to project scRNAseq data into a low-dimensional manifold that maps the transcriptional states a cell can occupy. Then a generative adversarial network (GAN) is used to simulate pesudocells at regular intervals along a virtual cell-state axis. We demonstrate the utility of Pseudocell Tracer by modeling B cells undergoing immunoglobulin class switch recombination (CSR) during a prototypic antigen-induced antibody response. Our results revealed an ordering of key transcription factors regulating CSR to the IgG1 isotype, including the concomitant expression of Nfkb1 and Stat6 prior to the upregulation of Bach2 expression. Furthermore, the expression dynamics of genes encoding cytokine receptors suggest a poised IL-4 signaling state that preceeds CSR to the IgG1 isotype.
The advance of metagenomic studies provides the opportunity to identify microbial taxa that are associated with human diseases. Multiple methods exist for the association analysis. However, the results could be inconsistent, presenting challenges in interpreting the host-microbiome interactions. To address this issue, we develop Meta-Signer, a novel Metagenomic Signature Identifier tool based on rank aggregation of features identified from multiple machine learning models including Random Forest, Support Vector Machines, Logistic Regression, and Multi-Layer Perceptron Neural Networks. Meta-Signer generates ranked taxa lists by training individual machine learning models over multiple training partitions and aggregates the ranked lists into a single list by an optimization procedure to represent the most informative and robust microbial features. A User will receive speedy assessment on the predictive performance of each ma-chine learning model using different numbers of the ranked features and determine the final models to be used for evaluation on external datasets. Meta-Signer is user-friendly and customizable, allowing users to explore their datasets quickly and efficiently.
The concurrent profiles of the gut microbiome and metabolome can be used in the diagnosis of complex diseases. However, the establishment of robust predictive models is challenging due to the high dimensionality of data and complex interactions among microbiome, metabolites, and host. Using deep neural networks consisting of an autoencoder for extracting latent representations and a multilayer neural network for disease prediction, we show that gut metabolome is more predictive of inflammatory bowel disease (IBD) than gut microbiome. In addition, we design a new multi-task autoencoder to extract the latent profiles from the combined microbiome and metabolome data. We further demonstrate that the combined latent profiles can further improve the performance of prediction. In summary, our work shows that autoencoders are useful apparatuses in generating low dimensional profiles that contribute to the improved performance and robustness for IBD prediction.
The microbiome of the human body has been shown to have profound effects on physiological regulation and disease pathogenesis. However, association analysis based on statistical modeling of microbiome data has continued to be a challenge due to inherent noise, complexity of the data, and high cost of collecting large number of samples. To address this challenge, we employed a deep learning framework to construct a data-driven simulation of microbiome data using a conditional generative adversarial network. Conditional generative adversarial networks train two models against each other while leveraging side information learn from a given dataset to compute larger simulated datasets that are representative of the original dataset. In our study, we used a cohorts of patients with inflammatory bowel disease to show that not only can the generative adversarial network generate samples representative of the original data based on multiple diversity metrics, but also that training machine learning models on the synthetic samples can improve disease prediction through data augmentation. In addition, we also show that the synthetic samples generated by this cohort can boost disease prediction of a different external cohort.
ABSTRACT Single cell RNA sequencing (scRNA-seq) can be used to infer a temporal ordering of dynamic cellular states. Current methods for the inference of cellular trajectories rely on unbiased dimensionality reduction techniques. However, such biologically agnostic ordering can prove difficult for modeling complex developmental or differentiation processes. The cellular heterogeneity of dynamic biological compartments can result in sparse sampling of key intermediate cell states. This scenario is especially pronounced in dynamic immune responses of innate and adaptive immune cells. To overcome these limitations, we develop a supervised machine learning framework, called Pseudocell Tracer, which infers trajectories in pseudospace rather than in pseudotime. The method uses a supervised encoder, trained with adjacent biological information, to project scRNA-seq data into a low-dimensional cellular state space. Then a generative adversarial network (GAN) is used to simulate pesudocells at regular intervals along a virtual cell-state axis. We demonstrate the utility of Pseudocell Tracer by modeling B cells undergoing immunoglobulin class switch recombination (CSR) during a prototypic antigen-induced antibody response. Our results reveal an ordering of key transcription factors regulating CSR, including the concomitant induction of Nfkb1 and Stat6 prior to the upregulation of Bach2 expression. Furthermore, the expression dynamics of genes encoding cytokine receptors point to the existence of a regulatory mechanism that reinforces IL-4 signaling to direct CSR to the IgG1 isotype.