The built environment houses diverse microbial communities whose diversity and composition differ among building materials and environmental conditions. Ecological theory makes predictions about how productivity and diversity shape communities, and experiments in the built environment provide an opportunity to test these. We manipulated moisture (constant or repeated wet–dry cycling) on three common building materials to test predictions about alpha and beta diversity. The most productive material (oriented strand board) supported the highest bacterial alpha and beta diversity, and these diversity levels were reduced by repeated drying disturbances. Diversity patterns for fungi were more variable, with the highest alpha diversity on low–moderate productivity material (gypsum wallboard). Fungal beta diversity was reduced by disturbance on high-productivity material, but increased on the other materials. These patterns were driven largely by members of Bacillaceae, Sphingomonadaceae, and Aspergillaceae that reached high abundances in some treatments. Differences between bacteria and fungi may be due to the scale-dependence of productivity–diversity relationships. Together, these results indicate that disturbances can interact with building materials, in some cases leading to variation in community composition that makes it difficult to predict the conditions under which microorganisms with potential importance to health and safety will occur. Importance The built environment—the homes, workplaces, vehicles, and other spaces where people spend most of their time—contains an enormous diversity of microorganisms with significance for human wellbeing, yet we know little about the factors shaping these microbial communities. We found that patterns of wetting and drying that mimic indoor leaks on common building materials affects the diversity of bacteria and fungi growing on the materials. But the identity of these microorganisms differed from one building material to another, and this was especially variable with wetting and drying. This means that common disturbances that lead to microbial growth in homes and offices can make it difficult to predict which microbes, including those that represent health threats to people, will occur in the built environment. ### Competing Interest Statement The authors have declared no competing interest.
Sparse feature tables, in which many features are present in very few samples, are common in big biological data (e.g. metagenomics). Ignoring issues of zero-laden datasets can result in biased statistical estimates and decreased power in downstream analyses. Zeros are also a particular issue for compositional data analysis using log-ratios since the log of zero is undefined. Researchers typically deal with this issue by removing low frequency features, but the thresholds for removal differ markedly between studies with little or no justification. Here, we present CurvCut, an unsupervised data-driven approach with human confirmation for rare-feature removal. CurvCut implements two distinct approaches for determining natural breaks in the feature distributions: a method based on curvature analysis borrowed from thermodynamics and the Fisher-Jenks statistical method. Our results show that CurvCut rapidly identifies data-specific breaks in these distributions that can be used as cutoff points for low-frequency feature removal that maximizes feature retention. We show that CurvCut works across different biological data types and rapidly generates clear visual results that allow researchers to confirm and apply feature removal cutoffs to individual datasets.
Phylogenetic analysis of protein sequences provides a powerful means of identifying novel protein functions and subfamilies, and for identifying and resolving annotation errors. However, automation of functional clustering based on phylogenetic trees has been challenging and most of it is done manually. Clustering phylogenetic trees usually requires the delineation of tree-based thresholds (e.g., distances), leading to an ad hoc problem. We propose a new phylogenetic clustering approach that identifies clusters without using ad hoc distances or other pre-defined values. Our workflow combines uniform manifold approximation and projection (UMAP) with Gaussian mixture models as a k-means like procedure to automatically group sequences into clusters. We then apply a "second pass" clade identification algorithm to resolve non-monophyletic groups. We tested our approach with several well-curated protein families (outer membrane porins, acyltransferase, and nuclear receptors) and showed our automated methods recapitulated known subfamilies. We also applied our methods to a broad range of different protein families from multiple databases, including Pfam, PANTHER, and UniProt, and to alignments of RNA viral genomes. Our results showed that AutoPhy rapidly generated monophyletic clusters (subfamilies) within phylogenetic trees evolving at very different rates both within and among phylogenies. The phylogenetic clusters generated by AutoPhy resolved misannotations and identified new protein functional groups and novel viral strains.
Background We introduce nyemtaay, a Python package for the calculation of classical population genetic statistics and inference of gene flow network connections and directionality in metapopulation networks using information theory. This genetic information flow network inference approach provided here is the only existing implementation of [[1][1]], and is applicable not only to ecological and evolutionary organism and landscape scale studies, but also has potential applications in, for example, cancer biology for analyzing clonal cell origins in metastasizing tumors. Results We demonstrate this potential through simulations and an analysis of metastasizing cancer cell lineages, showcasing its ability to identify the tissue site of origin in cancer networks. This work highlights the importance of considering demographic history and founder effects in interpreting gene flow directionality, and the benefits of this understanding in allowing application of this approach to gene flow network modeling to reach a broader range of domains. Conclusions nyemtaay is available under the MIT license from its public repository (), and can be installed locally using the Python package manager ‘pip’. ### Competing Interest Statement The authors have declared no competing interest. [1]: #ref-1
Sparse feature tables, in which many features are present in very few samples, are common in big biological data (e.g., metagenomics, transcriptomics). Ignoring the problem of zero-inflation can result in biased statistical estimates and decrease power in downstream analyses. Zeros are also a particular issue for compositional data analysis using log-ratios since the log of zero is undefined. Researchers typically deal with zero-inflated data by removing low frequency features, but the thresholds for removal differ markedly between studies with little or no justification. Here, we present CurvCut, a data-driven mathematical approach to zero-inflated feature removal based on curvature analysis of a “ball rolling down a hill”, where the hill is a histogram of feature distribution. These histograms typically contain a point of regime change, a discontinuity with a sharp change in the characteristics of the distribution, that can be used as a cutoff point for low frequency feature removal that considers the data-specific nature of the feature distribution. Our results show that CurvCut works well across a variety of biological data types, including ones with both right- and left-skewed feature distributions, and rapidly generates clear visual results allowing researchers to select data-appropriate cutoffs for feature removal.
Anaerobic fungi are emerging biotechnology platforms with genomes rich in biosynthetic potential. Yet, the heterologous expression of their biosynthetic pathways has had limited success in model hosts like E. coli. We find one reason for this is that the genome composition of anaerobic fungi like P. indianae are extremely AT-biased with a particular preference for rare and semi-rare AT-rich tRNAs in E coli, which are not explicitly predicted by standard codon adaptation indices (CAI). Native P. indianae genes with these extreme biases create drastic growth defects in E. coli (up to 69% reduction in growth), which is not seen in genes from other organisms with similar CAIs. However, codon optimization rescues growth, allowing for gene evaluation. In this manner, we demonstrate that anaerobic fungal homologs such as PI.atoB are more active than S. cerevisiae homologs in a hybrid pathway, increasing the production of mevalonate up to 2.5 g/L (more than two-fold) and reducing waste carbon to acetate by ~90% under the conditions tested. This work demonstrates the bioproduction potential of anaerobic fungal enzyme homologs and how the analysis of codon utilization enables the study of otherwise difficult to express genes that have applications in biocatalysis and natural product discovery.
Periodontal disease (PD) is a chronic, progressive polymicrobial disease that induces a strong host immune response. Culture-independent methods, such as next-generation sequencing (NGS) of bacteria 16S amplicon and shotgun metagenomic libraries, have greatly expanded our understanding of PD biodiversity, identified novel PD microbial associations, and shown that PD biodiversity increases with pocket depth. NGS studies have also found PD communities to be highly host-specific in terms of both biodiversity and the response of microbial communities to periodontal treatment. As with most microbiome work, the majority of PD microbiome studies use standard data normalization procedures that do not account for the compositional nature of NGS microbiome data. Here, we apply recently developed compositional data analysis (CoDA) approaches and software tools to reanalyze multiomics (16S, metagenomics, and metabolomics) data generated from previously published periodontal disease studies. CoDA methods, such as centered log-ratio (clr) transformation, compensate for the compositional nature of these data, which can not only remove spurious correlations but also allows for the identification of novel associations between microbial features and disease conditions. We validated many of the studies' original findings, but also identified new features associated with periodontal disease, including the genera Schwartzia and Aerococcus and the cytokine C-reactive protein (CRP). Furthermore, our network analysis revealed a lower connectivity among taxa in deeper periodontal pockets, potentially indicative of a more "random" microbiome. Our findings illustrate the utility of CoDA techniques in multiomics compositional data analysis of the oral microbiome.
BackgroundPlant biomass is an abundant but underused feedstock for bioenergy production due to its complex and variable composition, which resists breakdown into fermentable sugars. These feedstocks, however, are routinely degraded by many uncommercialized microbes such as anaerobic gut fungi. These gut fungi express a broad range of carbohydrate active enzymes and are native to the digestive tracts of ruminants and hindgut fermenters. In this study, we examine gut fungal performance on these substrates as a function of composition, and the ability of this isolate to degrade inhibitory high syringyl lignin-containing forestry residues.ResultsWe isolated a novel fungal specimen from a donkey in Independence, Indiana, United States. Phylogenetic analysis of the Internal Transcribed Spacer 1 sequence classified the isolate as a member of the genus Piromyces within the phylum Neocallimastigomycota (Piromyces sp. UH3-1, strain UH3-1). The isolate penetrates the substrate with an extensive rhizomycelial network and secretes many cellulose-binding enzymes, which are active on various components of lignocellulose. These activities enable the fungus to hydrolyze at least 58% of the glucan and 28% of the available xylan in untreated corn stover within 168h and support growth on crude agricultural residues, food waste, and energy crops. Importantly, UH3-1 hydrolyzes high syringyl lignin-containing poplar that is inhibitory to many fungi with efficiencies equal to that of low syringyl lignin-containing poplar with no reduction in fungal growth. This behavior is correlated with slight remodeling of the fungal secretome whose composition adapts with substrate to express an enzyme cocktail optimized to degrade the available biomass.ConclusionsPiromyces sp. UH3-1, a newly isolated anaerobic gut fungus, grows on diverse untreated substrates through production of a broad range of carbohydrate active enzymes that are robust to variations in substrate composition. Additionally, UH3-1 and potentially other anaerobic fungi are resistant to inhibitory lignin composition possibly due to changes in enzyme secretion with substrate. Thus, anaerobic fungi are an attractive platform for the production of enzymes that efficiently use mixed feedstocks of variable composition for second generation biofuels. More importantly, our work suggests that the study of anaerobic fungi may reveal naturally evolved strategies to circumvent common hydrolytic inhibitors that hinder biomass usage.