A wide variety of human diseases are associated with loss of microbial diversity in the human gut, inspiring a great interest in the diagnostic or therapeutic potential of the microbiota. However, the ecological forces that drive diversity reduction in disease states remain unclear, rendering it difficult to ascertain the role of the microbiota in disease emergence or severity. One hypothesis to explain this phenomenon is that microbial diversity is diminished as disease states select for microbial populations that are more fit to survive environmental stress caused by inflammation or other host factors. Here, we tested this hypothesis on a large scale, by developing a software framework to quantify the enrichment of microbial metabolisms in complex metagenomes as a function of microbial diversity. We applied this framework to over 400 gut metagenomes from individuals who are healthy or diagnosed with inflammatory bowel disease (IBD). We found that high metabolic independence (HMI) is a distinguishing characteristic of microbial communities associated with individuals diagnosed with IBD. A classifier we trained using the normalized copy numbers of 33 HMI-associated metabolic modules not only distinguished states of health vs IBD, but also tracked the recovery of the gut microbiome following antibiotic treatment, suggesting that HMI is a hallmark of microbial communities in stressed gut environments.
Plasmids are extrachromosomal genetic elements that often encode fitness-enhancing features. However, many bacteria carry "cryptic" plasmids that do not confer clear beneficial functions. We identified one such cryptic plasmid, pBI143, which is ubiquitous across industrialized gut microbiomes and is 14 times as numerous as crAssphage, currently established as the most abundant extrachromosomal genetic element in the human gut. The majority of mutations in pBI143 accumulate in specific positions across thousands of metagenomes, indicating strong purifying selection. pBI143 is monoclonal in most individuals, likely due to the priority effect of the version first acquired, often from one's mother. pBI143 can transfer between Bacteroidales, and although it does not appear to impact bacterial host fitness in vivo, it can transiently acquire additional genetic content. We identified important practical applications of pBI143, including its use in identifying human fecal contamination and its potential as an alternative approach to track human colonic inflammatory states.
Plasmids alter microbial evolution and lifestyles by mobilizing genes that often confer fitness in changing environments across clades. Yet our ecological and evolutionary understanding of naturally occurring plasmids is far from complete. Here we developed a machine-learning model, PlasX, which identified 68,350 non-redundant plasmids across human gut metagenomes and organized them into 1,169 evolutionarily cohesive ‘plasmid systems’ using our sequence containment-aware network-partitioning algorithm, MobMess. Individual plasmids were often country specific, yet most plasmid systems spanned across geographically distinct human populations. Cargo genes in plasmid systems included well-known determinants of fitness, such as antibiotic resistance, but also many others including enzymes involved in the biosynthesis of essential nutrients and modification of transfer RNAs, revealing a wide repertoire of likely fitness determinants in complex environments. Our study introduces computational tools to recognize and organize plasmids, and uncovers the ecological and evolutionary patterns of diverse plasmids in naturally occurring habitats through plasmid systems.
Plasmids are extrachromosomal genetic elements that often encode fitness enhancing features. However, many bacteria carry ‘cryptic’ plasmids that do not confer clear beneficial functions. We identified one such cryptic plasmid, pBI143, which is ubiquitous across industrialized gut microbiomes, and is 14 times as numerous as crAssphage, currently established as the most abundant genetic element in the human gut. The majority of mutations in pBI143 accumulate in specific positions across thousands of metagenomes, indicating strong purifying selection. pBI143 is monoclonal in most individuals, likely due to the priority effect of the version first acquired, often from one’s mother. pBI143 can transfer between Bacteroidales and although it does not appear to impact bacterial host fitness in vivo, can transiently acquire additional genetic content. We identified important practical applications of pBI143, including its use in identifying human fecal contamination and its potential as an inexpensive alternative for detecting human colonic inflammatory states.
Background Changes in microbial community composition as a function of human health and disease states have sparked remarkable interest in the human gut microbiome. However, establishing reproducible insights into the determinants of microbial succession in disease has been a formidable challenge. Results Here we use fecal microbiota transplantation (FMT) as an in natura experimental model to investigate the association between metabolic independence and resilience in stressed gut environments. Our genome-resolved metagenomics survey suggests that FMT serves as an environmental filter that favors populations with higher metabolic independence, the genomes of which encode complete metabolic modules to synthesize critical metabolites, including amino acids, nucleotides, and vitamins. Interestingly, we observe higher completion of the same biosynthetic pathways in microbes enriched in IBD patients. Conclusions These observations suggest a general mechanism that underlies changes in diversity in perturbed gut environments and reveal taxon-independent markers of “dysbiosis” that may explain why widespread yet typically low-abundance members of healthy gut microbiomes can dominate under inflammatory conditions without any causal association with disease.
Despite their prevalence and impact on microbial lifestyles, ecological and evolutionary insights into naturally occurring plasmids are far from complete. Here we developed a machine learning model, PlasX, which identified 68,350 non-redundant plasmids across human gut metagenomes, and we organized them into 1,169 evolutionarily cohesive ‘plasmid systems’ using our sequence containment-aware network partitioning algorithm, MobMess. Similar to microbial taxa, individuals from the same country tend to cluster together based on their plasmid diversity. However, we found no correlation between plasmid diversity and bacterial taxonomy. Individual plasmids were often country-specific, yet most plasmid systems spanned across geographically distinct human populations, revealing cargo genes that likely respond to environmental selection. Our study introduces powerful tools to recognize and organize plasmids, uncovers their tremendous diversity and intricate ecological and evolutionary patterns in naturally occurring habitats, and demonstrates that plasmids represent a dimension of ecosystems that is not explained by microbial taxonomy alone.
25 A detailed understanding of human gut microbial ecology is essential to engineer effective microbial therapeutics and to model microbial community assembly in health and disease. However, establishing generalizable insights into the functional determinants of microbial fitness in the gut has been a formidable challenge. Here we employ fecal microbiota transplantation (FMT) as an in natura experimental model to identify determinants of 30 microbial colonization and resilience. Our findings reveal adaptive ecological processes that favor high-fitness populations with higher metabolic competence as the main driver of microbial colonization outcomes after FMT. We further show that while healthy individuals harbor both low-fitness and high-fitness populations, individuals with inflammatory bowel disease are typically depleted of low-fitness populations. These 35 results offer a model to explain why common yet typically rare members of healthy guts can dominate under inflammatory conditions without any need for them to be causally associated with, or contribute to, such disease states.
A major goal of cancer research is to understand how mutations distributed across diverse genes affect common cellular systems, including multiprotein complexes and assemblies. Two challenges-how to comprehensively map such systems and how to identify which are under mutational selection-have hindered this understanding. Accordingly, we created a comprehensive map of cancer protein systems integrating both new and published multi-omic interaction data at multiple scales of analysis. We then developed a unified statistical model that pinpoints 395 specific systems under mutational selection across 13 cancer types. This map, called NeST (Nested Systems in Tumors), incorporates canonical processes and notable discoveries, including a PIK3CA-actomyosin complex that inhibits phosphatidylinositol 3-kinase signaling and recurrent mutations in collagen complexes that promote tumor proliferation. These systems can be used as clinical biomarkers and implicate a total of 548 genes in cancer evolution and progression. This work shows how disparate tumor mutations converge on protein assemblies at different scales.
Recent studies of the tumor genome seek to identify cancer pathways as groups of genes in which mutations are epistatic with one another or, specifically, “mutually exclusive.” Here, we show that most mutations are mutually exclusive not due to pathway structure but to interactions with disease subtype and tumor mutation load. In particular, many cancer driver genes are mutated preferentially in tumors with few mutations overall, causing mutations in these cancer genes to appear mutually exclusive with numerous others. Researchers should view current epistasis maps with caution until we better understand the multiple cause-and-effect relationships among factors such as tumor subtype, positive selection for mutations, and gross tumor characteristics including mutational signatures and load.
Gene networks are rapidly growing in size and number, raising the question of which networks are most appropriate for a particular application. Here, we evaluate 21 human genome-wide interaction networks for their ability to recover gene sets associated with 446 different diseases and 9 cancer hallmarks. While all networks have some ability in these recovery tasks, we observe a wide range of performance with STRING, GeneMANIA and GIANT networks having the best performance overall. A general tendency is that performance scales with network size, suggesting that new interaction discovery currently outweighs the detrimental effects of false positives. Correcting for size, we find that the DIP network provides the highest efficiency (value per interaction). Based on these results we create a parsimonious composite network with both high efficiency and absolute performance, which outperforms any single resource. This work provides a benchmark for selection of molecular networks in human disease research. Citation Format: Justin K. Huang, Daniel E. Carlin, Michael K. Yu, Wei Zhang, Jason F. Kreisberg, Pablo Tamayo, Trey Ideker. Systematic evaluation of gene networks for discovery of disease genes [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2018; 2018 Apr 14-18; Chicago, IL. Philadelphia (PA): AACR; Cancer Res 2018;78(13 Suppl):Abstract nr 1310.
Gene networks are rapidly growing in size and number, raising the question of which networks are most appropriate for particular applications. Here, we evaluate 21 human genome-wide interaction networks for their ability to recover 446 disease gene sets identified through literature curation, gene expression profiling, or genome-wide association studies. While all networks have some ability to recover disease genes, we observe a wide range of performance with STRING, ConsensusPathDB, and GIANT networks having the best performance overall. A general tendency is that performance scales with network size, suggesting that new interaction discovery currently outweighs the detrimental effects of false positives. Correcting for size, we find that the DIP network provides the highest efficiency (value per interaction). Based on these results, we create a parsimonious composite network with both high efficiency and performance. This work provides a benchmark for selection of molecular networks in human disease research.
Abstract Cancer is governed by modular systems of genes, the composition and organization of which remains poorly understood. Here, we integrate physical and functional networks from a wide range of molecular studies to assemble a comprehensive multi-scale map of human cancer cell biology. This map consists of a hierarchical catalog of protein complexes, signaling pathways and inter-pathway crosstalk implicated in cancer, and it suggests many uncharacterized functional modules as intriguing hypotheses for further validation. Analysis of the pattern of somatic mutations in The Cancer Genome Atlas (TCGA) reveals that these mutations target systems of varying scales above the level of individual genes. The map also provides a platform to integrate and interpret new 'omics data; we integrate new protein-protein interactions identified using AP-MS in multiple breast cancer cell lines, revealing how different functional modules are rewired in cancer cells. A general model browsing tool has been created to visualize and navigate these hierarchical cancer maps. This multi-scale mapping approach elucidates the molecular heterogeneity of cancer, connects tumor genotypes to phenotypes and, ultimately, enables a platform for cancer precision medicine. Citation Format: Fan Zheng, Michael K. Yu, Minkyu Kim, Keiichiro Ono, Mitchell Flagg, Jason F. Kreisberg, Nevan Krogan, Trey Ideker. Multi-scale mapping of the physical and functional architecture of the cancer cell [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2018; 2018 Apr 14-18; Chicago, IL. Philadelphia (PA): AACR; Cancer Res 2018;78(13 Suppl):Abstract nr 1317.
A major ambition of artificial intelligence lies in translating patient data to successful therapies. Machine learning models face particular challenges in biomedicine, however, including handling of extreme data heterogeneity and lack of mechanistic insight into predictions. Here, we argue for "visible'' approaches that guide model structure with experimental biology.
Pablo Tamayo合作论文数Theoretical Division and Advanced Computing Laboratory, Los Alamos National Laboratory, Los Alamos, NM2