Development of effective pollution mitigation strategies require an understanding of the pollution sources and factors influencing fecal pollution loading. Fecal contamination of Turkey Creek in Gulfport, Mississippi, one of the nation's most endangered creeks, was studied through a multi-tiered approach. Over a period of approximately two years, four stations across the watershed were analyzed for nutrients, enumeration of E. coli, male-specific coliphages and bioinformatic analysis of sediment microbial communities. The results demonstrated that two stations, one adjacent to a lift station and one just upstream from the wastewater-treatment plant, were the most impacted. The station adjacent to land containing a few livestock was the least impaired. While genotyping of male-specific coliphage viruses generally revealed a mixed viral signature (human and other animals), fecal contamination at the station near the wastewater treatment plant exhibited predominant impact by municipal sewage. Fecal indicator loadings were positively associated with antecedent rainfall for three of four stations. No associations were noted between fecal indicator loadings and any of the nutrients. Taxonomic signatures of creek sediment were unique to each sample station, but the sediment microbial community did overlap somewhat following major rain events. No presence of Escherichia coli (E. coli) or enterococci were found in the sediment. At some of the stations it was evident that rainfall was not always the primary driver of fecal transport. Repeated monitoring and analysis of a variety of parameters presented in this study determined that point and non-point sources of fecal pollution varied spatially in association with treated and/or untreated sewage.
Next-generation sequencing (NGS) has been instrumental to the advancement of metagenomic analysis for complex microbial communities. As techniques and metagenome databases improve, shotgun metagenomic analysis is becoming more common. However, several factors including genomic sequence repeats, translocations, and duplications can be difficult for NGS short reads to resolve. Long-read sequencing, on the other hand, has the potential to overcome the limitations of short-read sequencing, enabling better resolution of structural variants, sequencing repetitive regions, phasing of alleles, and distinguishing highly homologous genomic regions. With NGS revolutionizing our ability to characterize diverse and complex microbial communities such as stool, there is an increased need for high-quality, long high-molecular weight (HMW) genomic DNA (gDNA) and therefore efficient and effective long HMW gDNA extraction methods, especially from metagenomic samples. Many factors affect the quality and length of microbial DNA when extracting from stool, including microbial blooms and evanescence during collection and storage, nuclease activity during processing, and inefficient lysis of recalcitrant bacteria, yeast, and archaea during DNA purification. Furthermore, the choice of bioinformatics tools can alter read alignment efficiency and relative abundance analysis. Here we present evaluations of stool sample preservation, pre-treatment, and long HMW gDNA extraction. Additionally, we utilized long-read sequencing and a defined mock microbial community standard comprised of 8 bacteria and 2 yeasts to compare metagenomic long-read data using individual alignment tools. These data and observations contribute to the improvement of our ability to utilize long-read sequencing for routine metagenomic analysis.
The human gut microbiome is receiving increasingly more attention in recent years due to growing awareness of its relationship to human health. Taxonomic identification and phylogenetic profiling are the fundamental steps in the characterization of the gut microbiome. However, many recent publications have highlighted data variations of gut microbiome profiling when utilizing different methods for sample preservation, DNA purification, library preparation, and bioinformatics analysis. In this study, we illustrate key biases and variations caused by improper sample processing and present optimized solutions to prevent such discrepancies. We found liquid preservatives protect and preserve the composition of a microbiome sample prior to DNA extraction. Also, exaggerated variations of microbial composition can stem from the use of different DNA extraction kits. Additionally, differences in NGS library preparations can have a dramatic effect on the final composition profile. And finally, bioinformatic analysis can have a large impact on the observed microbial profiles. The microbiome profiling workflow has many sequential steps where biases can compound, amplifying deviations from the true profile, and the complexity of fecal samples further convolutes this process. But, if these sources of bias are considered and validated methods are used, researchers can have confidence in the accuracy of microbiome data, from which meaningful conclusions can be drawn.
ABSTRACTSummaryMicrobiome studies continue to provide tremendous insight into the importance of microorganism populations to the macroscopic world. High-throughput DNA sequencing technology (i.e., Next-generation Sequencing) has enabled the cost-effective, rapid assessment of microbial populations when combined with bioinformatic tools capable of identifying microbial taxa and calculating the diversity and composition of biological and environmental samples. Ribosomal RNA gene sequencing, where 16S and 18S rRNA gene sequences are used to identify prokaryotic and eukaryotic species, respectively, is one of the most widely-used techniques currently employed in microbiome analysis. Prior to bioinformatic analysis of these sequences, trimming parameters must be set so that post-trimming sequence information is maximized while expected errors in the sequences themselves are minimized. In this application note, we present FIGARO: a Python–based application designed to maximize read retention after trimming and filtering for quality. FIGARO was designed specifically to increase reproducibility and minimize trial-and-error in trimming parameter selection for a DADA2–based pipeline and will likely be useful for optimizing trimming parameters and minimizing sequence errors in other pipelines as well where paired-end overlap is required.Availability and implementationThe FIGARO application is freely available as source code at https://github.com/Zymo-Research/figaro.
Metagenomics research has grown rapidly since 2010. This growth has contributed to a lack of reproducibility between different methods and laboratories, which risks limiting our ability to compare between studies and decreases confidence in previous conclusions. To address the diversity of methods available, we compared the performance of many commercially- and academically sourced lysis protocols. The methods were evaluated using mock microbial community standards with defined composition and known manufacturing tolerances to serve as a ground truth for measurement. In order to facilitate comparisons, we developed the Measurement Integrity Quotient (MIQ), providing a single, easy to understand numerical score that describes the accuracy of an observed composition relative to a known standard. Utilizing this method, we compared the effects of many different variables on lysis efficiency, including differences between thermal, enzymatic, and mechanical (bead) lysis as well as minor changes within a method, including over 40 different bead material/size combinations and cell disruptor type/intensity/run time combinations. Additionally, several replicates of typical sample types (feces, soil, skin, saliva, urine) were tested with different lysis methods with replicates at different laboratories for over 1500 samples tested. Hard, dense ceramic beads of mixed size on an appropriate cell disruptor with a validated protocol created the least biased lysis of all examined methods. The use of the MIQ score provided rapid, easy to understand analysis of accuracy and has been released in a rapidly deployable container image for public use. The use of these data, and methods can allow for the creation of more reproducible metagenomics pipelines with results that better represent the true composition of the sample. These methods were able to achieve minimal amounts of deviation from expected composition and a high degree of run-to-run and interlab reproducibility.