Multi-level Monte Carlo methods have become an established technique in uncertainty quantification as they provide the same statistical accuracy as traditional Monte Carlo methods but with increased computational performance. Recently, similar techniques using multi-level ensembles have been applied to data assimilation problems. In this work we study the practical challenges and opportunities of applying multi-level methods to complex data assimilation problems, in the context of simplified ocean models. We simulate the shallow-water equations at different resolutions and employ a multi-level Kalman filter to assimilate sparse in-situ observations. In this context, where the shallow-water equations represent a simplified ocean model, we present numerical results from a synthetic test case, where small-scale perturbations lead to turbulent behaviour, conduct state estimation and forecast drift trajectories using multi-level ensembles. This represents a new advance towards making multi-level data assimilation feasible for real-world oceanographic applications.
Using distributed acoustic sensing data from a day of field testing on a fiber-optic cable along a railroad track in Norway, we detect and track cars and trains moving along a segment of the cable where the road runs parallel to the railroad tracks. We develop a method for automatic detection of events using signal processing, thresholding and density-based clustering, and then put data picks into a Kalman filter variant known as joint probabilistic data association filter for multiple object tracking and classification. Statistical model parameters are specified using in-situ labeling data along with the fiber-optic signals. Running the algorithm over time, we automatically track about 100 cars and 20 trains per hour. The velocities of cars coming from a zone with higher speed limit tend to be larger (35 km/h) than that of cars going in the opposite direction (30 km/h).
Potential CO2 storage sites need to perform risk assessments on the likelihood of anomalous events such as leakage. The intrinsic heterogeneity of the rock system with uncertain values for the capillary threshold pressures of the various rock elements is the most likely reason for unexpected vertical migration of CO2 within a storage complex. This study shows how the Invasion Percolation Markov Chain approach can be used to address this concern. We tested the approach using detailed 3D models of the multi-layer plume at Sleipner showing that even small variations in the threshold pressures of the shales can impact the flow of CO2 into multiple accumulations. Models with and without shale breaks reveal the importance of vertical feeders and/or faults, and the geometry of the shale layers is also crucial as the CO2 strongly conforms to topography. We demonstrate that the vertical migration of CO2 at Sleipner follows a Markovian model in which the probability of later migration events is highly dependent of the probability of preceding events. This case study illustrates how the initial migration events, which have the highest probability of occurring, should be the focus of CO2 storage risk assessments.
Local approaches have gained interest because they can provide fast approximate solutions for inverse problems. Following the idea of split-and-conquer, one aims to effectively condition variables to data using only small parts of the big model. We study and compare local approaches for conditioning in the context of seismic amplitude data and tomography, with two datasets relevant to improved oil and gas recovery in the North Sea and groundwater characterization in Denmark. In our comparison we study a local variant of an extended rejection sampler, termed localized extended rejection sampler (LERS), and a local ensemble transform Kalman filter (LETKF). Using various output statistics, we investigate the performance of the methods at marginal (e.g. mean and variance) level and joint properties (e.g. volume uncertainty and connectivity) of the subsurface variables of interest. Computed posterior statistics are compared with a reference Markov chain Monte Carlo solution. The results highlight benefits of the methods, such as fast reliable performance on the marginal properties, while joint properties in the more difficult cases show potential challenges of applying these local methodologies. Based on the results in our two cases, we discuss the applicability of the methods. We conclude that the localization methods are efficient and useful for estimating marginal properties and associated uncertainty, and can be an inexpensive tool for evaluating the need for further data processing. Local assimilation as outlined here is not suitable for generating posterior realizations of the spatial process variables.
Creating accurate and geologically realistic reservoir facies based on limited measurements is crucial for field development and reservoir management, especially in the oil and gas sector. Traditional two-point geostatistics, while foundational, often struggle to capture complex geological patterns. Multi-point statistics offers more flexibility, but comes with its own challenges. With the rise of Generative Adversarial Networks (GANs) and their success in various fields, there has been a shift towards using them for facies generation. However, recent advances in the computer vision domain have shown the superiority of diffusion models over GANs. Motivated by this, a novel Latent Diffusion Model is proposed, which is specifically designed for conditional generation of reservoir facies. The proposed model produces high-fidelity facies realizations that rigorously preserve conditioning data. It significantly outperforms a GAN-based alternative.
Stochastic reservoir characterization, a critical aspect of subsurface exploration for oil and gas reservoirs, relies on stochastic methods to model and understand subsurface properties using seismic data. This paper addresses the computational challenges associated with Bayesian reservoir inversion methods, focusing on two key obstacles: the demanding forward model and the high dimensionality of Gaussian random fields. Leveraging the generalized Bayesian approach, we replace the intricate forward function with a computationally efficient multivariate adaptive regression splines method, resulting in a 34 acceleration in computational efficiency. For handling high-dimensional Gaussian random fields, we employ a fast Fourier transform (FFT) technique. Additionally, we explore the preconditioned Crank-Nicolson method for sampling, providing a more efficient exploration of high-dimensional parameter spaces. The practicality and efficacy of our approach are tested extensively in simulations and its validity is demonstrated in application to the Alvheim field data.
The ecological status of lakes is important for understanding an ecosystem's biodiversity as well as for service water quality and policies related to land use and agricultural run-off. If the status is weak, then decisions about management alternatives need to be made. We assess the value of information of lake monitoring in Finland, where lakes are abundant. With reasonable ecological values and restoration costs, the value of information analysis can be compared with the survey's costs. Data are worth gathering if the expected value from the data exceeds the costs. From existing data, we specify a hierarchical Bayesian spatial logistic regression model for the ecological status of lakes. We then rely on functional approximations and Laplace approximations to get closed-form expressions for the value of information of a sampling design. The case study contains thousands of lakes. The combinatorially difficult design problem is to wisely pick the right subset of lakes for data gathering. To solve this optimization problem, we study the performance of various heuristics: greedy forward algorithms, exchange algorithms and Bayesian optimization approaches. The value of information increases quickly when adding lakes to a small design but then flattens out. Good designs are usually composed of lakes that are difficult to manage, while also balancing a variety of covariates and geographic coverage. The designs achieved by forward selection are reasonably good, but we can outperform them with the more nuanced search algorithms. Statistical designs clearly outperform other designs selected according to simpler criteria.
We describe, implement, and show the results of a localized ensemble-based approach for seismic amplitude-variation-with-offset (AVO) inversion with uncertainty quantification. Ensembles are simulated from prior probability distributions for fluid saturations and clay content. Starting with continuous saturations and clay content variables, we use depth-varying models for cementation and grain contact theory, Gassmann fluid substitution with mixed saturations, and approximations to the Zoeppritz equations for the AVO attributes at the top-reservoir. The local conditioning to seismic AVO observations relies on (1) the misfit between ensemble simulated seismic AVO data and the field observations in a local partition of the grid/local patch, of inlines/crosslines around the locations where we aim to predict, (2) correlations between the simulated reservoir properties and the data in local patches, and (3) local assessment to avoid unrealistic updates based on spurious correlations in the ensembles. Data from the Alvheim field in the North Sea are used to demonstrate the approach. The influence of the prior information from the well logs in combination with the seismic reflection data indicates the presence of higher oil and gas saturation in the lobe structures of the field and increased clay content at their edges.
Geophysical monitoring of CO2 storage projects enables informed decision making of injection strategies. When monitoring projects are designed, decisions should be made related to which geophysical data should be collected, for example seismic or electromagnetic data, and when surveys should be collected. In this work, we conduct value of information analysis to assess when to perform monitoring and which data to collect. The assessment is based on multiple reservoir simulations and uncertainty quantification of the CO2 plume characteristics as well as for the most relevant data variables. We consider multiple time decision-making processes, which require probability updating after collecting data at previous time steps. To demonstrate the suggested workflow, we perform dynamic simulations of the CO2 saturation fluid flow model based on a simplified version of the Smeaheia reservoir characterization model. The results show that the value of information depends on the acquisition time and on the previously measured geophysical data. We compare results of the suggested monitoring workflow with traditional fixed-interval strategies to demonstrate that the proposed method achieves accurate decision making with more cost-effective monitoring.
The coastal environment faces multiple challenges due to climate change and human activities. Sustainable marine resource management necessitates knowledge, and development of efficient ocean sampling approaches is increasingly important for understanding the ocean processes. Currents, winds, and freshwater runoff make ocean variables such as salinity very heterogeneous, and standard statistical models can be unreasonable for describing such complex environments. We employ a class of Gaussian Markov random fields that learns complex spatial dependencies and variability from numerical ocean model data. The suggested model further benefits from fast computations using sparse matrices, and this facilitates real-time model updating and adaptive sampling routines on an autonomous underwater vehicle. To justify our approach, we compare its performance in a simulation experiment with a similar approach using a more standard statistical model. We show that our suggested modeling framework outperforms the current state of the art for modeling such spatial fields. Then, the approach is tested in a field experiment using two autonomous underwater vehicles for characterizing the three-dimensional fresh-/saltwater front in the sea outside Trondheim, Norway. One vehicle is running an adaptive path planning algorithm while the other runs a preprogrammed path. The objective of adaptive sampling is to reduce the variance of the excursion set to classify freshwater and more saline fjord water masses. Results show that the adaptive strategy conducts effective sampling of the frontal region of the river plume.
Abstract. Multi-level Monte Carlo methods have established as a tool in uncertainty quantification for decreasing the computational costs while maintaining the same statistical accuracy as in single-level Monte Carlo. Lately, there have also been theoretical efforts to use similar ideas to facilitate multi-level data assimilation. By applying a multi-level ensemble Kalman filter for assimilating sparse observations of ocean currents into a simplified ocean model based on the shallow-water equations, we study the practical challenges of applying these method to more complex problems. We present numerical results from a realistic test case where small-scale perturbations lead to chaotic behaviour, and in this context we conduct state estimation and drift trajectories forecasting using multi-level ensembles. This represents a new step on the path of making multi-level data assimilation feasible for real-world oceanographic applications.
Models for areal data are traditionally defined using the neighborhood structure of the regions on which data are observed. The unweighted adjacency matrix of a graph is commonly used to characterize the relationships between locations, resulting in the implicit assumption that all pairs of neighboring regions interact similarly, an assumption which may not be true in practice. It has been shown that more complex spatial relationships between graph nodes may be represented when edge weights are allowed to vary. Christensen and Hoff (2023) introduced a covariance model for data observed on graphs which is more flexible than traditional alternatives, parameterizing covariance as a function of an unknown edge weights matrix. A potential issue with their approach is that each edge weight is treated as a unique parameter, resulting in increasingly challenging parameter estimation as graph size increases. Within this article we propose a framework for estimating edge weight matrices that reduces their effective dimension via a basis function representation of of the edge weights. We show that this method may be used to enhance the performance and flexibility of covariance models parameterized by such matrices in a series of illustrations, simulations and data examples.
For oceanographic applications, probabilistic forecasts typically have to deal with i) high-dimensional complex models, and ii) very sparse spatial observations. In search-and-rescue operations at sea, for instance, the short-term predictions of drift trajectories are essential to efficiently define search areas, but in-situ buoy observations provide only very sparse point measurements, while the mission is ongoing. Statistically optimal forecasts, including consistent uncertainty statements, rely on Bayesian methods for data assimilation to make the best out of both the complex mathematical modeling and the sparse spatial data. To identify suitable approaches for data assimilation in this context, we discuss localisation strategies and compare two state-of-the-art ensemble-based methods for applications with spatially sparse observations. The first method is a version of the ensemble-transform Kalman filter, where we tailor a localisation scheme for sparse point data. The second method is the implicit equal-weights particle filter which has recently been tested for related oceanographic applications. First, we study a linear spatio-temporal model for contaminant advection and diffusion, where the analytical Kalman filter provides a reference. Next, we consider a simplified ocean model for sea currents, where we conduct state estimation and predict drift. Insight is gained by comparing ensemble-based methods on a number of skill scores including prediction bias and accuracy, distribution coverage, rank histograms, spatial connectivity and drift trajectory forecasts.
Rule-based reservoir models incorporate rules that mimic actual sediment deposition processes for accurate representation of geological patterns of sediment accumulation. Bayesian methods combine rule-based reservoir modelling and well data, with geometry and placement rules as part of the prior and well data accounted for by the likelihood. The focus here is on a shallow marine shoreface geometry of ordered sedimentary packages called bedsets. Shoreline advance and sediment build-up are described through progradation and aggradation parameters linked to individual bedset objects. Conditioning on data from non-vertical wells is studied. The emphasis is on the role of ‘configurations’—the order and arrangement of bedsets as observed within well intersections in establishing the coupling between well observations and modelled objects. A conditioning algorithm is presented that explicitly integrates uncertainty about configurations for observed intersections between the well and the bedset surfaces. As data volumes increase and model complexity grows, the proposed conditioning method eventually becomes computationally infeasible. It has significant potential, however, to support the development of more complex models and conditioning methods by serving as a reference for consistency in conditioning.
Limnology and Oceanography BulletinVolume 32, Issue 1 p. 39-40 Meeting Highlights Symposium on Advances in Ocean Observation Kanna Rajan, Kanna Rajan University of Porto, Porto, PortugalSearch for more papers by this authorAida Alvera-Azcárate, Aida Alvera-Azcárate University of Liege, Liege, BelgiumSearch for more papers by this authorJo Eidsvik, Jo Eidsvik Norwegian University of Science and Technology, Trondheim, NorwaySearch for more papers by this authorAjit Subramaniam, Corresponding Author Ajit Subramaniam [email protected] orcid.org/0000-0003-1316-5827 Columbia University, New York, New York, USASearch for more papers by this author Kanna Rajan, Kanna Rajan University of Porto, Porto, PortugalSearch for more papers by this authorAida Alvera-Azcárate, Aida Alvera-Azcárate University of Liege, Liege, BelgiumSearch for more papers by this authorJo Eidsvik, Jo Eidsvik Norwegian University of Science and Technology, Trondheim, NorwaySearch for more papers by this authorAjit Subramaniam, Corresponding Author Ajit Subramaniam [email protected] orcid.org/0000-0003-1316-5827 Columbia University, New York, New York, USASearch for more papers by this author First published: 03 January 2023 https://doi.org/10.1002/lob.10536Read the full textAboutPDF ToolsRequest permissionExport citationAdd to favoritesTrack citation ShareShare Give accessShare full text accessShare full-text accessPlease review our Terms and Conditions of Use and check box below to share full-text version of article.I have read and accept the Wiley Online Library Terms and Conditions of UseShareable LinkUse the link below to share a full-text version of this article with your friends and colleagues. Learn more.Copy URL Share a linkShare onEmailFacebookTwitterLinkedInRedditWechat No abstract is available for this article. Volume32, Issue1February 2023Pages 39-40 RelatedInformation
Summary Distributed acoustic sensing (DAS) is a promising technology for monitoring for instance geo-energy and geohazards applications. However, the large amount of data generated by DAS can be challenging to analyze, and there is a need for developing tools that can process data effectively for decision support systems. We propose a two-step approach for detecting events in DAS data. First, we use DBSCAN to cluster the data into groups of points that are likely to be caused by the same event. Then, we use multiple convolutional neural networks (CNNs) to classify the clusters into different event types. The proposed approach has several advantages: it is able to detect events of arbitrary shape and size, it is robust to noise, and it is able to classify events into a variety of different types. We evaluate the methods on a DAS dataset collected at the Gløshaugen campus of the Norwegian University of Science and Technology. Our results show that the approach can achieve an accuracy of 87% in classifying events.
Summary Diffusion models have recently gained significant attention, especially the text-to-image models that produce realistic artworks. The challenge of conditional sampling is somewhat similar to our current focus: analyzing facies well logs and aiming to reconstruct complete 2 or 3-dimensional subsurface facies models based on well observations. We introduce a diffusion model that uses latent Gaussian variables and is paired with an encoder-decoder for the discrete representation of facies. Trained on a comprehensive collection of simulated samples, our model can generate samples. We have analyzed the statistical characteristics of the generated samples in comparison to others and believe that our method offers a promising approach for conditional sampling of subsurface geological models.
Autonomous underwater vehicles with onboard computing units foster innovative approaches for sampling oceanographic phenomena. Feedback of observations via the onboard model for planning algorithms enable adaptive sampling for such robotic units. In this work, we develop, implement, and test an adaptive sampling algorithm for efficient sampling of water masses in a 3-D frontal system. Focusing on a river plume, salinity variations are used to characterize the water masses. A threshold in salinity is assumed to distinguish the ocean and river waters, so that excursions below the threshold define river waters. The onboard model builds on a Gaussian random field representation of the salinity variations in (north, east, depth) coordinates. This model is initially trained from numerical ocean model data, and then updated with data gathered by the vehicle sensor. The Gaussian random field model further allows closed-form expressions of the expected spatially integrated Bernoulli variance of the salinity excursion set, which is used to reward sampling efforts. Combining these results with forward-looking planning algorithms, we suggest a workflow for 3-D adaptive sampling to map river plume systems. Simulation studies are used to compare the suggested approach with others. Results of field trials in the Nidelva river plume in Norway are presented and discussed.
Constructing efficient spatial sampling designs is important for monitoring, understanding and characterizing environmental processes acting over large-scale spatial domain. Multiple types of devices can accomplish informative sampling in various ways, and time and computational limitations should be considered too. In-situ sampling is often accurate, but it tends to be very sparse in the vast spatial areas considered for mapping purposes, and careful constructions of experimental sampling designs are needed. A criterion for mapping spatial presence-absence variables is presented. The expected integrated Bernoulli variance criterion is used in order to explore regions with more uncertainty about the binary outcomes. This approach uses a hierarchical Bayesian logistic regression model for the binary variable and develops approximate closed form expressions for the expected integrated Bernoulli variance. The approximations are fast to compute for many designs. The expressions are extended to find adaptive designs in the setting of sequential spatial exploration, where there are limited computation resources and operational time restrictions. In simulation studies the approximations are compared with tedious Monte Carlo sampling methods. Approximations are shown to be accurate and to work for a large spatial grid. An application in benthic habitat mapping is presented. It considers a coral dataset from Australia to demonstrate the suggested sampling design approach.
Discharge of mine tailings significantly impacts the ecological status of the sea. Methods to efficiently monitor the extent of dispersion is essential to protect sensitive areas. By combining underwater robotic sampling with ocean models, we can choose informative sampling sites and adaptively change the robot’s path based on in situ measurements to optimally map the tailings distribution near a seafill. This paper creates a stochastic spatio-temporal proxy model of dispersal dynamics using training data from complex numerical models. The proxy model consists of a spatio-temporal Gaussian process model based on an advection–diffusion stochastic partial differential equation. Informative sampling sites are chosen based on predictions from the proxy model using an objective function favoring areas with high uncertainty and high expected tailings concentrations. A simulation study and data from real-life experiments are presented.