In mine planning, geospatial estimates of variables such as comminution indexes and metallurgical recovery are extremely important to locate blocks for which the energy consumption at the plant is minimized and for which the recovery of minerals is maximized. Unlike ore grades, these variables cannot be modeled with traditional geostatistical methods, which rely on the availability of a large number of samples for variogram estimation and on the additivity of variables for change of support, among other issues. Past attempts to build geospatial models of geometallurgical variables have failed to address some of these issues, and most importantly, did not consider adequate mathematical models for uncertainty quantification. In this work, we propose a new methodology that combines Bayesian predictive models with Kriging in Hilbert spaces to quantify the geospatial uncertainty of such variables in realistic industrial settings. The results we obtained with data from a real deposit indicate that the proposed approach may become an interesting alternative to geostatistical simulation.
Curated dataset with measurements of various geometallurgical variables as described in the paper "Modeling geospatial uncertainty of geometallurgical variables with Bayesian models and Hilbert-Kriging" by the same authors. The comminution.csv table contains measurements associated with the DWT and BWI indexes. The flotation.csv table contains measurements related to the metallurgical recovery of copper in a locked cycled test LCT and the drillholes.csv table contains chemical analysis data along with the geospatial coordinates of the samples.
The construction of conceptual geological models is an essential task in petroleum exploration, especially during the early stages of investment, when evidence about the subsurface is limited. In this task, geoscientists recreate the most likely geological scenarios that led to potential accumulation of reserves in a target block, based on past experience, historical analogues, and interpreted "signatures" that were left in the data by physical processes. Due to cognitive constraints, this task has traditionally focused on the single most likely conceptual scenario, or at most, a reduced set of scenarios chosen a priori via ad-hoc methods, which often lead to improper block valuation and severe money losses. In this work, we propose a probabilistic framework for reasoning about conceptual geological scenarios that helps domain experts maintain multiple hypotheses throughout the exploration program. The framework is extensible and can be instantiated automatically from simple knowledge templates, a form of "knowledge standard" in the company. We show how the acquired knowledge can be leveraged for uncertainty mitigation using concepts from information theory, and assess the framework qualitatively in a real case study.
Statistical learning theory provides the foundation to applied machine learning, and its various successful applications in computer vision, natural language processing and other scientific domains. The theory, however, does not take into account the unique challenges of performing statistical learning in geospatial settings. For instance, it is well known that model errors cannot be assumed to be independent and identically distributed in geospatial (a.k.a. regionalized) variables due to spatial correlation; and trends caused by geophysical processes lead to covariate shifts between the domain where the model was trained and the domain where it will be applied, which in turn harm the use of classical learning methodologies that rely on random samples of the data. In this work, we introduce thegeostatistical (transfer) learningproblem, and illustrate the challenges of learning from geospatial data by assessing widely-used methods for estimating generalization error of learning models, under covariate shift and spatial correlation. Experiments with synthetic Gaussian process data as well as with real data from geophysical surveys in New Zealand indicate that none of the methods are adequate for model selection in a geospatial context. We provide general guidelines regarding the choice of these methods in practice while new methods are being actively researched.
Machine learning (ML) models are being widely used in the geosciences for various tasks involving well log data, including prediction of missing well log curves, picking of stratigraphic surfaces, facies classification, and segmentation of different rock types. Even though various ML applications have been proposed in the literature, it is difficult to reproduce and advance the prior art without having access to the data and preprocessing steps used. In fact, there is an increasing need for benchmark cases to assess past and future solutions. The present dataset integrates well log data curated from the 2016 New Zealand Petroleum Exploration Public Data Pack.
Directional variograms were introduced in geostatistics as a tool for revealing major directions of correlation in spatial data. However, their estimation presents some practical challenges, particularly in the case of large irregularly-sampled data sets where efficient spectral-based estimation methods are not applicable. In this work, we propose a generalization of directional variograms to general partitions of spatial data, and introduce a parallel estimation algorithm that can efficiently handle large data sets with more than 105 points. This partition variogram generalization is motivated by a five-spot point pattern in the petroleum industry, which we named more generally as the isolated-lines arrangement. In such an arrangement, traditional estimators of directional variograms such as r-tube estimators very often fail to incorporate measurements from adjacent lines (e.g. vertical wells) without also incorporating measurements from other planes (e.g. horizontal layers). We provide illustrations of this new concept, and assess the approximation error of the proposed estimators with bootstrap methods and synthetic Gaussian process data.
Many Earth surface processes are studied using field, experimental, or numerical modeling data sets that represent a small subset of possible outcomes observed in nature. Based on these data, deterministic models can be built that describe the "average" evolution of a system. However, these models commonly cannot account for the complex variability of many processes or present a quantitative statement of uncertainty. To assess such uncertainty, stochastic models are needed that can mimic spatial as well as temporal variability. A common limitation for applying stochastic models to Earth surface processes is a lack of data and methods that allow constraining the full spatiotemporal variability of these models. In this paper, we propose a Bayesian framework for calibrating input parameters to stochastic models of morphodynamic systems using time series of image data from the field, or from numerical and laboratory experiments. The framework consists of generating synthetic time series of images using the stochastic model and rejecting those time series that do not reproduce key morphodynamic statistics of the available data sets. The calibrated stochastic model allows us to quantify both the spatial and temporal uncertainty about the evolution of the morphodynamic systems of interest. For demonstration purposes, we apply the framework to a single flume experiment of braided river channels evolving under steady water and sediment discharges, but it can be used more generally to quantify spatiotemporal uncertainty for any time series of morphodynamic data for which key statistics can be defined.
Summary Increasing the productivity of seismic imaging workflow through efficient simulation pipelines is a mandatory task for any Oil & Gas company nowadays. In this work, we propose a GPGPU pipeline for fast synthesis of seismic data that encompasses high-performance geostatistical simulation of rock properties and efficient numerical propagation of acoustic waves to deliver a large data set of 3D seismic cubes with spatially-varying properties, enabling the training and assessment of recently-proposed neural network architectures for seismic inversion.
Creating increasingly realistic groundwater models involves the inclusion of additional geological and geophysical data in the hydrostratigraphic modeling procedure. Using multiple-point statistics (MPS) for stochastic hydrostratigraphic modeling provides a degree of flexibility that allows the incorporation of elaborate datasets and provides a framework for stochastic hydrostratigraphic modeling. This paper focuses on comparing three MPS methods: snesim, DS and iqsim. The MPS methods are tested and compared on a real-world hydrogeophysical survey from Kasted in Denmark, which covers an area of 45 km2. A controlled test environment, similar to a synthetic test case, is constructed from the Kasted survey and is used to compare the modeling results of the three aforementioned MPS methods. The comparison of the stochastic hydrostratigraphic MPS models is carried out in an elaborate scheme of visual inspection, mathematical similarity and consistency with boreholes. Using the Kasted survey data, an example for modeling new survey areas is presented. A cognitive hydrostratigraphic model of one area is used as a training image (TI) to create a suite of stochastic hydrostratigraphic models in a new survey area. The advantage of stochastic modeling is that detailed multiple point information from one area can be easily transferred to another area considering uncertainty. The presented MPS methods each have their own set of advantages and disadvantages. The DS method had average computation times of 6–7 h, which is large, compared to iqsim with average computation times of 10–12 min. However, iqsim generally did not properly constrain the near-surface part of the spatially dense soft data variable. The computation time of 2–3 h for snesim was in between DS and iqsim. The snesim implementation used here is part of the Stanford Geostatistical Modeling Software, or SGeMS. The snesim setup was not trivial, with numerous parameter settings, usage of multiple grids and a search-tree database. However, once the parameters had been set it yielded comparable results to the other methods. Both iqsim and DS are easy to script and run in parallel on a server, which is not the case for the snesim implementation in SGeMS.
GeoStats.jl is an extensible framework for high-performance geostatistics in Julia, as well as a formal specification of statistical problems in the spatial setting.It provides highly optimized solvers for estimation and (conditional) simulation of variables defined over general spatial domains (e.g.regular grid, point collection), and can utilize high-performance hardware for parallel execution such as GPUs and computer clusters.
Process-based modeling offers a way to represent realistic geological heterogeneity in subsurface models. The main limitation lies in conditioning such models to data. Multiple-point geostatistics can use these process-based models as training images and address the data conditioning problem. In this work, we further develop image quilting as a method for 3D stochastic simulation capable of mimicking the realism of process-based geological models with minimal modeling effort (i.e. parameter tuning) and at the same time condition them to a variety of data. In particular, we develop a new probabilistic data aggregation method for image quilting that bypasses traditional ad-hoc weighting of auxiliary variables. In addition, we propose a novel criterion for template design in image quilting that generalizes the entropy plot for continuous training images. The criterion is based on the new concept of voxel reuse—a stochastic and quilting-aware function of the training image. We compare our proposed method with other established simulation methods on a set of process-based training images of varying complexity, including a real-case example of stochastic simulation of the buried-valley groundwater system in Denmark.
Fixes deprecation warnings from using old @functorize macro and update jupyter notbooks and plotting examples in doc directory.
We are exploring algorithms to predict the aggregate power output of many photovoltaic systems in a single geographic region on a 3-hour time horizon at 5-minute steps (a 36-step forecast) based only on observed system power. The goal is to correctly identify upcoming “ramp events,” or large positive or negative deviations from a long-term trend over a short time period (Sevlian and Rajagopal, 2013). Identifying these ramp events before they occur will allow grid operators to plan for large changes in net load, thereby allowing for deeper penetration of solar power generation on the grid.
Drilling activities in the oil and gas industry have been reported over decades for thousands of wells on a daily basis, yet the analysis of this text at large-scale for information retrieval, sequence mining, and pattern analysis is very challenging. Drilling reports contain interpretations written by drillers from noting measurements in downhole sensors and surface equipment, and can be used for operation optimization and accident mitigation. In this initial work, a methodology is proposed for automatic classification of sentences written in drilling reports into three relevant labels (EVENT, SYMPTOM and ACTION) for hundreds of wells in an actual field. Some of the main challenges in the text corpus were overcome, which include the high frequency of technical symbols, mistyping/abbreviation of technical terms, and the presence of incomplete sentences in the drilling reports. We obtain state-of-the-art classification accuracy within this technical language and illustrate advanced queries enabled by the tool.