I consider Bayesian model weighting (BMW), which is an essential component in realistic uncertainty quantification. The main motivation is field-scale porous-media-flow problems, where existing methods such as Bayesian model averaging (BMA), Bayesian stacking (BS), and Modified Bayesian stacking (MBS), typically are computationally infeasible. The main computational obstacle with BMA and MBS is use of Monte-Carlo (MC) integration where the integrand involves a data misfit in high dimensions, suffering from the curse of dimensionality. I propose a novel class of methods for BMW that in essence replaces a single high-dimensional MC integration with a modest number of MC integrations not suffering from the curse of dimensionality. The novel class of methods can be utilized both within the BMA and MBS frameworks. The computational gain with respect to plain BMA and MBS is quantified, and three methods from the novel class are applied to three examples. The first example is a modified version of Model 1 in the Society of Petroleum Engineers Comparative Solution Project, while the two other examples utilize data, simulation results, and additional information from the G-segment of the Norne field in the North Sea.
Summary The overall objective of the new RamonCO project is "to develop multi-dimensional robust conformance assessment methodology and advance risk-based monitoring and societal implementation strategies", and thereby contribute to acceleration of the CO2 storage project permitting phase and secure containment and conformance during the project execution phase. Specific objectives include: Develop a framework for determining site-specific detection thresholds for field scale, multi-physics, multi-modal and multi-scale data types. Develop and test inversion framework for field scale inversion of multi-physics monitoring data. Develop and implement a methodology for including societal and environmental risk in common risk evaluation tools and assessment. Develop and implement a framework for risk-based value of information (VoI) analysis of CO2 storage monitoring
We consider calculation of the Bayesian model evidence, which is an essential component in realistic uncertainty quantification. The main motivation is large-scale porous-media-flow problems, where plain Monte-Carlo (MC) integration is computationally infeasible. We propose a simplistic multilevel (ML) estimator and argue why this estimator is expected to perform well within a certain problem class, and point out that many important problems, including many porous-media-flow problems, belong to this class. Four test cases are utilized to assess the performance of the proposed ML estimator. Test cases I and II contain two strongly linked toy models, but only Test case II belongs to the problem class where the proposed ML estimator is expected to perform well. Test cases III and IV are concerned with single-phase and two-phase porous-media flows, respectively. The results show that the ML estimator clearly and consistently outperforms MC integration in Test cases II, III, and IV, while this is not the case for Test case I.
Summary This paper introduces a novel approach for Bayesian reservoir history matching using Multi-fidelity Data Assimilation (MFDA). Traditional ensemble-based DA methods face computational challenges in field-scale applications, leading to sampling errors due to limited ensemble sizes. Typically, these errors are mitigated using complex distance-based localization strategies. The MFDA method presented here combines simulations with lower and varying fidelity, allowing a larger total ensemble size without increasing computational costs. This technique balances increased bias against reduced variance errors. This technique has previously been proven to be superior when compared to traditional DA for simpler synthetic reservoir cases. Applied to a complex, realistic field case with multiple wells and faults, we employed efficient grid-generation techniques to create grids with four fidelity levels. This enabled a large ensemble for history-matching, resulting in a better data match and accurate estimation of critical parameters like pore volume and field oil in place, surpassing standard Ensemble Smoother with Multiple Data Assimilation (ES -MDA) methods. This successful application demonstrates MFDA's potential in practical, field-scale reservoir simulations.
I consider the problem of model diagnostics, that is, the problem of criticizing a model prior to history matching by comparing data to an ensemble of simulated data based on the prior model (prior predictions). If the data are not deemed as a credible prior prediction by the model diagnostics, some settings of the model should be changed before history matching is attempted. I particularly target methodologies that are computationally feasible for large models with large amounts of data. A multiscale methodology, that can be applied to analyze differences between data and prior predictions in a scale-by-scale fashion, is proposed for this purpose. The methodology is computationally inexpensive, straightforward to apply, and can handle correlated observation errors without making approximations. The multiscale methodology is tested on a set of toy models, on two simplistic reservoir models with synthetic data, and on real data and prior predictions from the Norne field. The tests include comparisons with a previously published method (termed the Mahalanobis methodology in this paper). For the Norne case, both methodologies led to the same decisions regarding whether to accept or discard the data as a credible prior prediction. The multiscale methodology led to correct decisions for the toy models and the simplistic reservoir models. For these models, the Mahalanobis methodology either led to incorrect decisions, and/or was unstable with respect to selection of the ensemble of prior predictions.
Traditional uncertainty analysis for subsurface models is typically based on a single dynamic model with a number of uncertain parameters. Improved and more robust forecasting can be obtained by combining several models in a Bayesian setting using model averaging. The traditional Bayesian Model Averaging (BMA), however, suffers from several drawbacks, such as too large sensitivity to prior model assumptions and instability with respect to measurement perturbations, especially when the number of measurements is large. We suggest a modified version of BMA (MBMA) where the calculations are stabilized using an ensemble of measurements. Bayesian stacking (BS) is a method that is directly focused on the performance of the combined predictive distribution of several models. The original version of BS (BSLOO) is based on leave-one-out cross-validation and requires a Bayesian inversion for each data point which may be very time consuming. We suggest a modified version of stacking (MBS) that requires only a single history match and uses an ensemble of measurements. MBS may be used with either prior (MBS-pri) or posterior (MBS-post) predictive distributions. The behavior of the methods is illustrated using three synthetic, linear examples. One is a simple mixture model. The other two are inspired by 4D seismic data. The results with MBS-pri are very similar to the results with MBMA. The results with MBS-post are similar to those of BSLOO when the data are uncorrelated. MBS can take into account correlated data or measurement errors, while correlations are neglected in the BSLOO weight calculations.
We consider estimation of absolute permeability from inverted seismic data. Large amounts of simultaneous data, such as inverted seismic data, enhance the negative effects of Monte Carlo errors in ensemble-based Data Assimilation (DA). Multilevel (ML) models consist of a selection of models with different fidelities. Multilevel Data Assimilation (MLDA) attempts to obtain a better statistical accuracy with a small sacrifice of the numerical accuracy. Spatial grid coarsening is one way of generating an ML model. It has been shown that coarsening the spatial grid results in a problem with weaker nonlinearity, and hence, in a less challenging problem than the problem on the original fine grid. Accordingly, formulating a sequential MLDA algorithm which uses the coarser models in the first steps of the DA, followed by the finer models, helps to find an approximation to the solution of the inverse problem at the first steps and gradually converge to the solution. We present two variants of a sequential MLDA algorithm and compare their performance with both conventional DA algorithms and a simultaneous (i.e., using all the models on the different grids simultaneously) MLDA algorithm using numerical experiments. Both posterior parameters and posterior model forecasts are compared qualitatively and quantitatively. The results from numerical experiments suggest that all MLDA algorithms generally perform better than the conventional DA algorithms. In estimation of the posterior parameter fields, the simultaneous MLDA algorithm and one of the variants of sequential MLDA (SMLES-H) perform similarly and slightly better than the other variant (SMLES-S). While in estimation of the posterior model forecasts, SMLES-S clearly performs better than both the simultaneous MLDA algorithm and SMLES-H.
A sequential inversion methodology for combining geophysical data types of different resolutions is developed and applied to monitoring of large-scale CO 2 injection. The methodology is a two-step approach within the Bayesian framework where lower resolution data are inverted first and subsequently used in the generation of the prior model for inversion of the higher resolution data. For the application of CO 2 monitoring, the first step is done with either controlled-source electromagnetic (CSEM), or gravimetric, data, while the second step is done with seismic amplitude-versus-offset (AVO) data. The Bayesian inverse problems are solved by sampling the posterior probability distributions using either the ensemble Kalman filter or ensemble smoother with multiple data assimilation. A carefully designed parameterization is used to represent the unknown geophysical parameters: electric conductivity, density, and seismic velocity. The parameterization is well suited for identification of CO 2 plume location and variation of geophysical parameters within the regions corresponding to inside and outside of the plume. The inversion methodology is applied to a synthetic monitoring test case where geophysical data are made from fluid-flow simulation of large-scale CO 2 sequestration in the Skade formation in the North Sea. The numerical experiments show that seismic AVO inversion results are improved with the sequential inversion methodology using prior information from either CSEM or gravimetric inversion.
Summary Harnessing sufficient computational resources is one of the main constraints in the domain of computational statistics. In ensemble-based data assimilation (DA), this constraint results in a limited ensemble size which in turn results in high sampling errors. In case of large amounts of simultaneous data, e.g. seismic data, these sampling errors manifest themselves in severe underestimation of uncertainties in posterior distributions. This problem is known as ensemble collapse. The traditional tool to mitigate this problem is localization. As the term implies, it annihilates spurious non-local correlations but does not allow for true non-local correlations. An alternative approach is use of lower-fidelity modeling which reduces the computational cost per realization and thus renders a larger ensemble size. However, it also entails larger numerical errors. A multilevel model is a set of models consisting hierarchies of both computational accuracy and cost. Accordingly, Multilevel Data Assimilation (MLDA) attempts to generate a better balance between the statistical and numerical errors by utilizing a multilevel model for the forecast step of the DA. There have been several MLDA algorithms developed recently which have demonstrated promising results in assimilation of inverted seismic data in simplistic reservoir models. The multilevel attribute of these algorithms has been the coarseness of the spatial grid in the forward model. In this research, we examine one of these algorithms for assimilation of inverted seismic data in a more realistic petroleum reservoir problem. In doing so, a new method for coarsening the spatial grid is performed. This method allows for better inclusion of important geological features, such as faults. Additionally, certain adjustments are implemented so that the algorithm can handle storing and operating on large amounts of data efficiently. To assess the performance of the devised method, a numerical experiment is conducted, and the results obtained from this algorithm are compared with those of traditional DA methods.
To enhance the safety and reduce risk of leakage during carbon capture and storage (CCS) operations, there is a need for cost-effective and reliable geophysical monitoring tools for assessing of the spatial evolution of the CO2 front. Here, we investigate the applicability of distributed acoustic sensing (DAS) data to estimate changes in CO2 saturation (DSg). The DAS technology has recently emerged as a promising geophysical motoring tool due to its low deployment and maintenance costs and densely spaced information. In this work, we propose a stochastic inversion framework, by combining the iterative Ensemble Smoother (iES) with sparse representation of data and correlation based localization, to estimate DSg using time-lapse DAS data coming from a vertical borehole and a horizontal seabottom receiver lines. The proposed framework is implemented on a realistic reservoir model which is a potential candidate for large-scale offshore CO2 storage in the North Sea. The numerical results demonstrate that DAS data is capable of monitoring CO2 plume movement and thus can be considered as an efficient geophysical tool in CCS.
In ensemble-based data assimilation (DA), the ensemble size is usually limited to around one hundred. Straightforward application of ensemble-based DA can therefore result in significant Monte Carlo errors, often manifesting themselves as severe underestimation of parameter uncertainties. Localization is the conventional remedy for this problem. Assimilation of large amounts of simultaneous data enhances the negative effects of Monte Carlo errors. Use of lower-fidelity models reduces the computational cost per ensemble member and therefore renders the possibility to reduce Monte Carlo errors by increasing the ensemble size, but it also adds to the modeling error. Multilevel data assimilation (MLDA) uses a selection of models forming hierarchies of both computational cost and computational accuracy, and tries to balance between Monte Carlo errors and modeling errors. In this work, we assess a recently developed MLDA algorithm, the Multilevel Hybrid Ensemble Smoother (MLHES), and introduce and assess an iterative version of this algorithm, the Iterative Multilevel Hybrid Ensemble Smoother (IMLHES). In our assessments, we compare these algorithms with conventional single-level DA algorithms with localization. To this end, a typical example of large amount of spatially distributed data, i.e. inverted seismic data, is considered and three data sets of this kind are assimilated in three different petroleum reservoir models. Qualitatively evaluating the DA outcomes, it is found that multilevel algorithms outperform their conventional single-level counterparts in obtaining the posterior statistics of both uncertain parameters and model forecasts. Additionally, it is observed that IMLHES performs better than MLHES in the same regard, and also successfully converges to the proximity of solution in a case where the considered iterative single-level algorithm did not converge to the global optimum.
Summary Carbon capture and storage (CCS) is a crucial component in reducing greenhouse gases in the atmosphere. To make CCS attractive and ensure control and safety, monitoring solutions are needed capable of measuring subsurface effects of storage operations. Distributed acoustic sensing (DAS) can play a pivotal role as an alternative to conventional seismic monitoring. Here, we present results of a forward modeling approach for calculating seismic time lapse effects of CO2 injection in reservoirs on particle velocity and strain rate data, respectively equivalent to geophone and DAS data. The developed modelling approach is capable of efficiently simulating the time lapse effect of CO2 injection, while accounting for realistic field noise levels. To overcome the angle sensitivity of the DAS cable we recommend an orthogonal cable layout combined with analysis of multiple seismic phases (P and S) is recommended. The approach can help to support the design of the monitoring layouts and tune this to the specific target to be monitored. In a next step, quantitative information derived from seismic phases can be used to invert for changes in reservoir properties.
With large amounts of simultaneous data, like inverted seismic data in reservoir modeling, negative effects of Monte Carlo errors in straightforward ensemble-based data assimilation (DA) are enhanced, typically resulting in underestimation of parameter uncertainties. Utilization of lower fidelity reservoir simulations reduces the computational cost per ensemble member, thereby rendering the possibility of increasing the ensemble size without increasing the total computational cost. Increasing the ensemble size will reduce Monte Carlo errors and therefore benefit DA results. The use of lower fidelity reservoir models will however introduce modeling errors in addition to those already present in conventional fidelity simulation results. Multilevel simulations utilize a selection of models for the same entity that constitute hierarchies both in fidelities and computational costs. In this work, we estimate and approximately account for the multilevel modeling error (MLME), that is, the part of the total modeling error that is caused by using a multilevel model hierarchy, instead of a single conventional model to calculate model forecasts. To this end, four computationally inexpensive approximate MLME correction schemes are considered, and their abilities to correct the multilevel model forecasts for reservoir models with different types of MLME are assessed. The numerical results show a consistent ranking of the MLME correction schemes. Additionally, we assess the performances of the different MLME-corrected model forecasts in assimilation of inverted seismic data. The posterior parameter estimates from multilevel DA with and without MLME correction are compared to results obtained from conventional single-level DA with localization. It is found that multilevel DA (MLDA) with and without MLME correction outperforms conventional DA with localization. The use of all four MLME correction schemes results in posterior parameter estimates with similar quality. Results obtained with MLDA without any MLME correction were also of similar quality, indicating some robustness of MLDA toward MLME.
Summary Carbon capture and storage in subsurface requires frequent monitoring of CO2 plume movement. The correct assessment of the spatial distribution of CO2 saturation front lowers the risk of leakage and thus environmental hazards. In this work, we propose a stochastic inversion framework, the iterative Ensemble Smoother (iES), to predict changes in CO2 saturation plume (ΔSg) using time-lapse seismic and gravity data simultaneously. Here, change in inverted acoustic impedance (ΔIp) is considered as seismic data. The methodology is based on a Bayesian inversion problem, where the prior is provided as an ensemble of changes in CO2 saturation. The realizations of ΔSg in the prior model are computed using geostatistical method and reservoir flow simulator. The updated ΔSg (posterior) are then obtained based on the misfit between simulated and measured geophysical data. The proposed framework is applied and validated on a 3D reservoir model based on the Johansen formation which is a potential large scale offshore CO2 storage site in the North Sea. The numerical results demonstrate that the proposed framework is capable to monitor CO2 plume movement efficiently.
A sequential inversion methodology for combining geophysical data types of different resolutions is developed and applied to monitoring of large-scale CO 2 injection. The methodology is a two-step approach within the Bayesian framework where lower resolution data are inverted first, and subsequently used in the generation of the prior model for inversion of the higher resolution data. For the application of CO 2 monitoring, the first step is done with either controlled source electromagnetic (CSEM) or gravimetric data, while the second step is done with seismic amplitude-versus-offset (AVO) data. The Bayesian inverse problems are solved by sampling the posterior probability distributions using either the ensemble Kalman filter or ensemble smoother with multiple data assimilation. A model-based parameterization is used to represent the unknown geophysical parameters: electric conductivity, density, and seismic velocity. The parameterization is well suited for identification of CO 2 plume location and variation of geophysical parameters within the regions corresponding to inside and outside of the plume. The inversion methodology is applied to a synthetic monitoring test case where geophysical data are made from fluid-flow simulation of large-scale CO 2 sequestration in the Skade formation. The numerical experiments show that seismic AVO inversion results are improved with the sequential inversion methodology using prior information from either CSEM or gravimetric inversion.
Summary There is an increasing interest in Multi-fidelity Modeling within computational statistics research in recent years. Multilevel ensemble-based data assimilation (MLDA), taking advantage of Multi-fidelity modeling, is a novel approach for reservoir history-matching. This method has been proposed to overcome the potential sampling errors that are encountered in conventional ensemble-based data assimilation techniques. Ensemble-based methods have been successful in history-matching of large cases but the limit in computational resources normally results in the ensemble size to be confined to about 100, which can yield to sampling error. In order to address the problem of sampling error, localization has been proposed which handles the problem of non-local spurious correlations but does not allow for true non-local correlations. The basic concept of MLDA revolves about allocation of resources for computation of models on a hierarchy of accuracy and computational cost. Utilization of models with a lower computational cost enables a significant increase in the ensemble size. Doing so, it brings about the opportunity to trade an appropriate amount of computational accuracy for a better statistical accuracy. In this research, the hierarchy of computational cost is established using a variation of spatial resolutions in the simulation models, and a new scheme called Simultaneous Spatial Multilevel Data Assimilation for multilevel data is investigated on a reservoir model. This method is designed to assimilate the inverted seismic data in a multilevel manner. Accordingly, a set of different spatial resolutions of the model is created and an ensemble of models and their corresponding inverted seismic data are considered for any of the resolutions. The simulations are run for all of the levels and an independent update is performed on any of the levels using the Ensemble Smoother (ES). The reduction of computational cost in coarser resolutions entails a multilevel error which can be quantified and accounted for, by a comparison with the simulations in the finest level. Finally, a cumulative statistical analysis over all ensembles is done to assess the performance in data assimilation. Results obtained from two variants of the new scheme are evaluated and compared to ES with localization and standard ensemble size.
Assimilation of a sequence of linearly dependent data vectors, $\{d_{l}\}^{L}_{l=1}$ such that ${d_{l} = B_{l}d_{L}}^{L-1}_{ l=1}$, is considered for a parameter estimation problem. Such a data sequence can occur, for example, in the context of multilevel data assimilation. Since some information is used several times when linearly dependent data vectors are assimilated, the associated data-error covariances must be modified. I develop a condition that the modified covariances must satisfy in order to sample correctly from the posterior probability density function of the uncertain parameter in the linear-Gaussian case. It is shown that this condition is a generalization of the well-known condition that must be satisfied when assimilating the same data vector multiple times. I also briefly discuss some qualitative and computational issues related to practical use of the developed condition.
Multilevel ensemble-based data assimilation (DA) as an alternative to standard (single-level) ensemble-based DA for reservoir history matching problems is considered. Restricted computational resources currently limit the ensemble size to about 100 for field-scale cases, resulting in large sampling errors if no measures are taken to prevent it. With multilevel methods, the computational resources are spread over models with different accuracy and computational cost, enabling a substantially increased total ensemble size. Hence, reduced numerical accuracy is partially traded for increased statistical accuracy. A novel multilevel DA method, the multilevel hybrid ensemble Kalman filter (MLHEnKF) is proposed. Both the expected and the true efficiency of a previously published multilevel method, the multilevel ensemble Kalman filter (MLEnKF), and the MLHEnKF are assessed for a toy model and two reservoir models. A multilevel sequence of approximations is introduced for all models. This is achieved via spatial grid coarsening and simple upscaling for the reservoir models, and via a designed synthetic sequence for the toy model. For all models, the finest discretization level is assumed to correspond to the exact model. The results obtained show that, despite its good theoretical properties, MLEnKF does not perform well for the reservoir history matching problems considered. We also show that this is probably caused by the assumptions underlying its theoretical properties not being fulfilled for the multilevel reservoir models considered. The performance of MLHEnKF, which is designed to handle restricted computational resources well, is quite good. Furthermore, the toy model is utilized to set up a case where the assumptions underlying the theoretical properties of MLEnKF are fulfilled. On that case, MLEnKF performs very well and clearly better than MLHEnKF.
Large-scale CO2 storage requires advanced understanding of the geomechanical response of the caprock subject to pressure buildup in the reservoir. CO2 injection into basin systems will need to approach rates of 100 Mt/y in order to achieve significant reduction in worldwide emissions and realize the large capacity of saline aquifers. High pressure build-up is expected, which may activate existing fractures and faults in the seal and create leakage pathways for the CO2 plume. New fractures may also evolve at higher pressures. To avoid creating leakage pathways for the CO2 plume, the pressure must be kept under an upper bound that is determined by the caprock weakest point or initial stress. To date, there has been great uncertainty in critical knowledge needed for planning and execution of safe large-scale CO2 storage projects. Within the PROTECT project, we advance methods and reduce uncertainty regarding the mechanical response. The research study focuses on three subareas: data collection, experimental studies and computational tasks. The studies are performed at small scale (scale of individual cracks, cm to m) and at large scale (reservoir or basin scale, 10 m to km). Data are synthesized, and a benchmark study of large-scale geomechanical simulation is performed. The Utsira Formation on the Norwegian continental shelf has been used as a common case study for integration and benchmarking activities. Finally, recommendations are made for future research based on knowledge gained in this study.