Maximizing the recovery factor achieved through water flooding depends on acquiring a detailed understanding of the vertical and areal sweep efficiency. DNA diagnostics can monitor changes in oil contributions from multiple zones and from injectors, becoming a leading indicator for the potential of water breakthrough, loss of injectivity, and the overall advancement of the water front when combined with subsurface information. This allows for proactive management of injection rates and timing to maximize recovery rates for green fields and brownfields alike. DNA diagnostics use DNA markers acquired from microbes. DNA markers of produced fluids are compared to the DNA markers of injected fluids to establish relationships and shared fluid flow. This paper will cover the end to end workflow for long term waterflood monitoring:Establishing end members, even for a mature field, with the use of new samples from offset wells, properly stored samples from existing wells, and the analysis of commingled samples in combination with the subsurface model.Establishing the level of similarity between injectors and producers as an indication for the progression of the waterflood front using methods including Principal Coordinate Analysis (PCoA) of DNA marker profiles.Performing time series analysis and establishing sampling periodicity for effective waterflood monitoring. A pilot project, consisting of 12 producers and 3 injectors in a conventional California reservoir, was conducted to prove the concepts and further develop the required analysis for waterflood monitoring. Fluid samples were taken weekly on each well over 3 weeks to establish the difference in DNA markers between the fluids. The DNA markers were used to determine the probability that injection fluid was being produced from the surrounding wells. These results were overlaid to temporal changes in the Total Fluid Logs. Taken together, the results correlated and confirmed previous water breakthrough information and provided insights into arial and vertical conformance changes. Additionally, the project provided new insights into strength of producer and injector connection based on geological features and with that informing future infill drilling decisions. Waterflood monitoring is a powerful application for DNA diagnostics that is deployable on new and existing waterfloods. The spatial and temporal monitoring limitations of modeling or tracer studies can be improved upon through this non-invasive diagnostic. Initial results demonstrate the insights that can be provided not just for monitoring the waterflood but also for further field development decisions.
Permian operators have dramatically increased the number of multi-stage fractured horizontal wells over the past 5 years and face challenges associated with maximizing production of existing wells while developing new acreage and benches, all the while meeting capital return requirements. Over that time, DNA diagnostics have been applied successfully to more than 1000 wells throughout the Permian Basin to help operators reduce uncertainties ranging from drained rock volume, well-well communication, and sources of water production. When subsurface conditions change, microbes change, and the DNA from microbes can be used to profile total fluid flow (water + oil phases) from benches and between wells. It therefore serves as a powerful tool to provide a range of answers, using advanced analytics and integration with various data sets. In this study, we will provide the background of DNA diagnostics and related analytics, along with the latest insights into viable operating environments. We also highlight recent Permian basin projects that have used DNA in conjunction with operator data to reduce uncertainty about subsurface conditions. We will show Total Fluid Logs, which are based on comparing DNA signatures from produced fluids with a DNA stratigraphy log. Total Fluid Logs are utilized to 1) constrain interpreted fracture heights, and 2) work in combination with pressure and production data for Rate Transient Analysis (RTA) for significantly improved estimation of the half-length. The case histories will illustrate the differences between production rates and confirmed fracture height and half-length, and a discussion of microseismic is included. We show how produced fluid collection during pad completions can elucidate well-well communication and demonstrate the impact of completion size and completion order on effective drainage heights. DNA changes in produced fluids can be compared to production data to reveal the timing and impact of frac hits between wells during zipper completions. Finally, we provide a suggested workflow for analyzing water contributions out of target in the diagnosis of problem wells. Petrophysical logs can be compared to drainage height assessments to help reveal from which depths water may be producing and can be integrated with production data for a more complete subsurface understanding. DNA diagnostics represent a complementary, cost effective, minimum environmental footprint and low risk tool for operators to easily integrate into existing production and engineering workflows for monitoring well health and subsurface conditions across time.
Abstract Subsurface DNA Diagnostics™ is a low-risk, high-resolution evaluation tool offering oil and gas operators a measurement of fluid movement in the subsurface. DNA sequencing methodologies that use subsurface DNA markers acquired from well cuttings and produced fluids are currently used in all major US unconventional basins to elucidate drainage heights for new and existing wells. Dynamic drainage height estimations are especially important in field-wide development, when actionable turnaround time of drainage height estimates is a priority to improve subsurface understanding and maximize reservoir economics. For this work, well cuttings were collected during drilling every 10’ MD from a 1000’ vertical section of interest in the Permian's highly stratified Wolfcamp formation. Subsurface DNA was extracted and sequenced from cuttings at the sub-formation level to create a robust DNA marker profile. Produced fluids were collected every 2 weeks through 140 days. DNA markers from each sub-formation were compared to every fluid time point using a Bayesian mixture modeling algorithm. This method produces an estimate of the relative contribution of DNA markers from each sub-formation for each fluid sample over time. The result of Subsurface DNA Diagnostics is a novel subsurface log that characterizes DNA markers as a function of depth, and identifies the intervals with relative higher productivity compared to the other intervals over time. These drainage metrics are then overlaid with existing petrophysical properties to refine knowledge of oil and water behavior within different sub formations during a well's lifetime. Subsurface DNA Diagnostic's use of well cuttings and produced fluids enables rapid scalability due to low-risk sampling (no downhole tools or lost production time) and minimal field personnel time. As such, an operator can relatively quickly characterize the drainage height of their field over time using several vertical baselines and a routine sampling protocol of producing wells in both exploratory, development, and producing phases of the asset.
Abstract The application of DNA sequencing in the oil industry provides a non-invasive, economical, and high-resolution data source for elucidating hydrocarbon fluid movement. DNA Diagnostics could be used to understand drainage height and inform the vertical spacing of wells, as well as provide the ability to derive/predict depositional environments based on biomarker signatures. To date, DNA sequencing has been applied in over 200 wells across 8 basins in North America, including the Permian, Eagle Ford, and Bakken. Subsurface DNA Diagnostics are used to guide well spacing, evaluate well interference (e.g. frac hits), determine oil potential, and understand production profiling. The effective application of DNA sequencing in the oilfield requires the use of advanced data science techniques and machine learning. To date, the industry's application of machine learning has been focused on production and completion parameters that provide a statistical view of the subsurface. In oilfield DNA sequencing, millions of data points per well are generated; and to apply these direct measurements to the subsurface requires novel machine learning techniques which are presented here. In this analysis, we present the results of DNA sequencing from a 33-well study in the Delaware Basin. Well cuttings and produced fluid samples were obtained in a non-invasive manner to generate stratigraphically unique signals for hydrocarbon fluid movement. Various data science techniques are presented to interpret the data with initial observations providing a novel view into production across formations. Additionally, we will provide an overview of future work and how we plan to augment the dataset with other more conventional techniques in order to fine-tune the uncertainties associated with a new data source and methodology.
Abstract DNA diagnostics is a new reservoir characterization tool with potential to maximize reservoir production in tight rock formations. DNA extracted from rock layers provides high resolution fingerprints that define a "DNA stratigraphy" for organic intervals like the Wolfcamp. DNA sequences originate from microbes feeding on organic matter or minerals within the formation. A DNA stratigraphic profile, or type section, was assembled from a vertical pilot well's cuttings and core. The DNA signature from produced oil from offset laterals was subsequently compared against the DNA type section to provide estimated effective drainage height. Cuttings from a lateral well were compared with DNA from its produced oil to construct a production profile comparable to a traditional production log. In addition, when oil samples are collected over time, the method provides insight on interference, completion effectiveness, and SRV (Stimulated Reservoir Volume) changes with time. An optimized development plan in unconventional reservoirs requires operators to understand parameters such as effective drainage height, hydraulic fracture half-length and individual stage contributions resulting from their completions. Wolfcamp reservoirs consist of highly laminated mudrocks interbedded with limestones that have quite different mechanical properties. These contrasting lithologies make it difficult to estimate resultant completion geometries, SRV, and well-to- well interactions. Also, using costly production logs, individual stage contributions are difficult to obtain in lower pressure reservoirs like the Wolfcamp. However, these reservoir performance parameters are required to set benchmarks and continuously uplift the EUR by taking advantage of insightful diagnostics. Production logs, micro-seismic, chemical or radioactive tracers are all useful in understanding the subsurface, but can be expensive and can pose operational challenges. Subsurface DNA sequencing is a relatively low cost new data source that can be used to gain subsurface insights in complicated reservoirs. DNA stratigraphy can help assess critical geometric parameters resulting from stimulation by employing non-invasive sampling that enables lifetime well monitoring to track the flow of oil and provide engineers the basis to optimize completions and development plans. An 8 well "subsurface" lab was selected for the experiment. The project included one vertical pilot hole with cuttings, and 8 horizontal wells landed in two Wolfcamp pay zones (one of the laterals was extended from the same vertical pilot). Three horizontals had been on production for 11 months before the pilot well and 6 additional laterals were drilled. The pilot well and its sidetracked lateral had cuttings extracted for DNA sequencing. DNA signatures from the pilot well and lateral well were compiled to produce vertical and lateral DNA stratigraphic profiles. The DNA stratigraphic profiles were then compared to DNA from oil produced in the 7 offset laterals. DNA profiles were also compared to standard geologic parameters using pilot well e-logs, particularly mechanical stratigraphy. Lateral wells were sampled at various times after initial production to assess changes with time. Blind tests were designed to check the method as a reasonable estimator for effective drainage height and communication. DNA stratigraphy provides a more informed view of well spacing, completion design and well performance to help increase efficiency and asset value.
Disruption of healthy microbial communities has been linked to numerous diseases, yet microbial interactions are little understood. This is due in part to the large number of bacteria, and the much larger number of interactions (easily in the millions), making experimental investigation very difficult at best and necessitating the nascent field of computational exploration through microbial correlation networks. We benchmark the performance of eight correlation techniques on simulated and real data in response to challenges specific to microbiome studies: fractional sampling of ribosomal RNA sequences, uneven sampling depths, rare microbes and a high proportion of zero counts. Also tested is the ability to distinguish signals from noise, and detect a range of ecological and time-series relationships. Finally, we provide specific recommendations for correlation technique usage. Although some methods perform better than others, there is still considerable need for improvement in current techniques.
Recent studies suggest that gut microbiomes of urban-industrialized societies are different from those of traditional peoples. Here we examine the relationship between lifeways and gut microbiota through taxonomic and functional potential characterization of faecal samples from hunter-gatherer and traditional agriculturalist communities in Peru and an urban-industrialized community from the US. We find that in addition to taxonomic and metabolic differences between urban and traditional lifestyles, hunter-gatherers form a distinct sub-group among traditional peoples. As observed in previous studies, we find that Treponema are characteristic of traditional gut microbiomes. Moreover, through genome reconstruction (2.2–2.5 MB, coverage depth × 26–513) and functional potential characterization, we discover these Treponema are diverse, fall outside of pathogenic clades and are similar to Treponema succinifaciens , a known carbohydrate metabolizer in swine. Gut Treponema are found in non-human primates and all traditional peoples studied to date, suggesting they are symbionts lost in urban-industrialized societies.
ABSTRACT Glycans form the primary nutritional source for microbes in the human gut, and understanding their metabolism is a critical yet understudied aspect of microbiome research. Here, we present a novel computational pipeline for modeling glycan degradation (GlyDeR) which predicts the glycan degradation potency of 10,000 reference glycans based on either genomic or metagenomic data. We first validated GlyDeR by comparing degradation profiles for genomes in the Human Microbiome Project against KEGG reaction annotations. Next, we applied GlyDeR to the analysis of human and mammalian gut microbial communities, which revealed that the glycan degradation potential of a community is strongly linked to host diet and can be used to predict diet with higher accuracy than sequence data alone. Finally, we show that a microbe’s glycan degradation potential is significantly correlated (R = 0.46) with its abundance, with even higher correlations for potential pathogens such as the class Clostridia (R = 0.76). GlyDeR therefore represents an important tool for advancing our understanding of bacterial metabolism in the gut and for the future development of more effective prebiotics for microbial community manipulation. IMPORTANCE The increased availability of high-throughput sequencing data has positioned the gut microbiota as a major new focal point for biomedical research. However, despite the expenditure of huge efforts and resources, sequencing-based analysis of the microbiome has uncovered mostly associative relationships between human health and diet, rather than a causal, mechanistic one. In order to utilize the full potential of systems biology approaches, one must first characterize the metabolic requirements of gut bacteria, specifically, the degradation of glycans, which are their primary nutritional source. We developed a computational framework called GlyDeR for integrating expert knowledge along with high-throughput data to uncover important new relationships within glycan metabolism. GlyDeR analyzes particular bacterial (meta)genomes and predicts the potency by which they degrade a variety of different glycans. Based on GlyDeR, we found a clear connection between microbial glycan degradation and human diet, and we suggest a method for the rational design of novel prebiotics.
Recent advances that allow us to collect more data on DNA sequences and metabolites have increased our understanding of connections between the intestinal microbiota and metabolites at a whole-systems level. We can also now better study the effects of specific microbes on specific metabolites. Here, we review how the microbiota determines levels of specific metabolites, how the metabolite profile develops in infants, and prospects for assessing a person's physiological state based on their microbes and/or metabolites. Although data acquisition technologies have improved, the computational challenges in integrating data from multiple levels remain formidable; developments in this area will significantly improve our ability to interpret current and future data sets.
The bacteria that colonize humans and our built environments have the potential to influence our health. Microbial communities associated with seven families and their homes over 6 weeks were assessed, including three families that moved their home. Microbial communities differed substantially among homes, and the home microbiome was largely sourced from humans. The microbiota in each home were identifiable by family. Network analysis identified humans as the primary bacterial vector, and a Bayesian method significantly matched individuals to their dwellings. Draft genomes of potential human pathogens observed on a kitchen counter could be matched to the hands of occupants. After a house move, the microbial community in the new house rapidly converged on the microbial community of the occupants' former house, suggesting rapid colonization by the family's microbiota.
We present a performance-optimized algorithm, subsampled open-reference OTU picking, for assigning marker gene (e.g., 16S rRNA) sequences generated on next-generation sequencing platforms to operational taxonomic units (OTUs) for microbial community analysis. This algorithm provides benefits over de novo OTU picking (clustering can be performed largely in parallel, reducing runtime) and closed-reference OTU picking (all reads are clustered, not only those that match a reference database sequence with high similarity). Because more of our algorithm can be run in parallel relative to “classic” open-reference OTU picking, it makes open-reference OTU picking tractable on massive amplicon sequence data sets (though on smaller data sets, “classic” open-reference OTU clustering is often faster). We illustrate that here by applying it to the first 15,000 samples sequenced for the Earth Microbiome Project (1.3 billion V4 16S rRNA amplicons). To the best of our knowledge, this is the largest OTU picking run ever performed, and we estimate that our new algorithm runs in less than 1/5 the time than would be required of “classic” open reference OTU picking. We show that subsampled open-reference OTU picking yields results that are highly correlated with those generated by “classic” open-reference OTU picking through comparisons on three well-studied datasets. An implementation of this algorithm is provided in the popular QIIME software package, which uses uclust for read clustering. All analyses were performed using QIIME’s uclust wrappers, though we provide details (aided by the open-source code in our GitHub repository) that will allow implementation of subsampled open-reference OTU picking independently of QIIME (e.g., in a compiled programming language, where runtimes should be further reduced). Our analyses should generalize to other implementations of these OTU picking algorithms. Finally, we present a comparison of parameter settings in QIIME’s OTU picking workflows and make recommendations on settings for these free parameters to optimize runtime without reducing the quality of the results. These optimized parameters can vastly decrease the runtime of uclust-based OTU picking in QIIME.
We present a performance-optimized algorithm, subsampled open-reference OTU picking, for assigning marker gene (e.g., 16S rRNA) sequences generated on next-generation sequencing platforms to operational taxonomic units (OTUs) for microbial community analysis. This algorithm provides benefits over de novo OTU picking (clustering can be performed largely in parallel, reducing runtime) and closed-reference OTU picking (all reads are clustered, not only those that match a reference database sequence with high similarity). Because more of our algorithm can be run in parallel relative to “classic” open-reference OTU picking, it makes open-reference OTU picking tractable on massive amplicon sequence data sets (though on smaller data sets, “classic” open-reference OTU clustering is often faster). We illustrate that here by applying it to the first 15,000 samples sequenced for the Earth Microbiome Project (1.3 billion V4 16S rRNA amplicons). To the best of our knowledge, this is the largest OTU picking run ever performed, and we estimate that our new algorithm runs in less than 1/5 the time than would be required of “classic” open reference OTU picking. We show that subsampled open-reference OTU picking yields results that are highly correlated with those generated by “classic” open-reference OTU picking through comparisons on three well-studied datasets. An implementation of this algorithm is provided in the popular QIIME software package, which uses uclust for read clustering. All analyses were performed using QIIME’s uclust wrappers, though we provide details (aided by the open-source code in our GitHub repository) that will allow implementation of subsampled open-reference OTU picking independently of QIIME (e.g., in a compiled programming language, where runtimes should be further reduced). Our analyses should generalize to other implementations of these OTU picking algorithms. Finally, we present a comparison of parameter settings in QIIME’s OTU picking workflows and make recommendations on settings for these free parameters to optimize runtime without reducing the quality of the results. These optimized parameters can vastly decrease the runtime of uclust-based OTU picking in QIIME.
To study how microbes establish themselves in a mammalian gut environment, we colonized germfree mice with microbial communities from human, zebrafish, and termite guts, human skin and tongue, soil, and estuarine microbial mats. Bacteria from these foreign environments colonized and persisted in the mouse gut; their capacity to metabolize dietary and host carbohydrates and bile acids correlated with colonization success. Cohousing mice harboring these xenomicrobiota or a mouse cecal microbiota, along with germ-free "bystanders," revealed the success of particular bacterial taxa in invading guts with established communities and empty gut habitats. Unanticipated patterns of ecological succession were observed; for example, a soil-derived bacterium dominated even in the presence of bacteria from other gut communities (zebrafish and termite), and human-derived bacteria colonized germ-free bystander mice before mouse-derived organisms. This approach can be generalized to address a variety of mechanistic questions about succession, including succession in the context of microbiota-directed therapeutics.
A new study explores the ancient oral microbiome from the well-preserved dental calculus samples of four human individuals who lived during medieval times, using a suite of genomic, proteomic and microscopic approaches. The authors investigate the evolution of dental pathogens by reconstructing the genome of the periodontal pathogen Tannerella forsythia and also identify antibiotic resistance genes, bacterial virulence factors and host immune defense proteins.
RATIONALE:Lung infections caused by opportunistic or virulent pathogens are a principal cause of morbidity and mortality in HIV infection. It is unknown whether HIV infection leads to changes in basal lung microflora, which may contribute to chronic pulmonary complications that increasingly are being recognized in individuals infected with HIV.OBJECTIVES:To determine whether the immunodeficiency associated with HIV infection resulted in alteration of the lung microbiota.METHODS:We used 16S ribosomal RNA targeted pyrosequencing and shotgun metagenomic sequencing to analyze bacterial gene sequences in bronchoalveolar lavage (BAL) and mouths of 82 HIV-positive and 77 HIV-negative subjects.MEASUREMENTS AND MAIN RESULTS:Sequences representing Tropheryma whipplei, the etiologic agent of Whipple's disease, were significantly more frequent in BAL of HIV-positive compared with HIV-negative individuals. T. whipplei dominated the community (>50% of sequence reads) in 11 HIV-positive subjects, but only 1 HIV-negative individual (13.4 versus 1.3%; P = 0.0018). In 30 HIV-positive individuals sampled longitudinally, antiretroviral therapy resulted in a significantly reduced relative abundance of T. whipplei in the lung. Shotgun metagenomic sequencing was performed on eight BAL samples dominated by T. whipplei 16S ribosomal RNA. Whole genome assembly of pooled reads showed that uncultured lung-derived T. whipplei had similar gene content to two isolates obtained from subjects with Whipple's disease.CONCLUSIONS:Asymptomatic subjects with HIV infection have unexpected colonization of the lung by T. whipplei, which is reduced by effective antiretroviral therapy and merits further study for a potential pathogenic role in chronic pulmonary complications of HIV infection.
High-throughput DNA sequencing technologies, coupled with advanced bioinformatics tools, have enabled rapid advances in microbial ecology and our understanding of the human microbiome. QIIME (Quantitative Insights Into Microbial Ecology) is an open-source bioinformatics software package designed for microbial community analysis based on DNA sequence data, which provides a single analysis framework for analysis of raw sequence data through publication-quality statistical analyses and interactive visualizations. In this chapter, we demonstrate the use of the QIIME pipeline to analyze microbial communities obtained from several sites on the bodies of transgenic and wild-type mice, as assessed using 16S rRNA gene sequences generated on the Illumina MiSeq platform. We present our recommended pipeline for performing microbial community analysis and provide guidelines for making critical choices in the process. We present examples of some of the types of analyses that are enabled by QIIME and discuss how other tools, such as phyloseq and R, can be applied to expand upon these analyses.