The selection of wells for well‑treatment (WT) in oil and gas fields has traditionally relied on expert judgment and multi‑criteria decision analysis, which are subjective, labor‑intensive, and often yield suboptimal rankings. This paper proposes a data‑driven framework that ranks candidate wells for WT, specifically acid‑stimulation. To identify the best approach, we systematically compare three learning paradigms: (1) regression for point‑wise prediction of a continuous uplift value for each well independently; (2) classification for point‑wise prediction of a discrete effectiveness class; and (3) learn‑to‑rank for pairwise comparison of wells to directly optimize their relative ordering. The framework integrates physically grounded graphical models, which incorporate geological, technical, and inter‑well pressure‑maintenance factors with data-driven models, including tabular, time-series foundation models, and gradient boosting models. We evaluate the system on 3072 historical WT events and demonstrate a marked improvement in regression, classification and ranking over conventional baselines metrics. Specifically, the fine‑tuned time‑series foundation model performed best achieving a $R^2$ score of 0.87 for predicting the 6‑month mean uplift oil production rate, while the tabular foundation model attains a $ROC‑AUC$ of 0.84 for success determination, and the gradient boosting model reaches an $NDCG@10$ of 0.77 for ranking candidate well for WT. Using $NDCG@10$ as the unified evaluation metric, learn‑to‑rank clearly outperforms both regression and classification methods. Furthermore, an ablation study isolating the contribution of physics-based features revealed a significant performance gains across all evaluated metrics of at least 21\%. The approach delivers an objective decision-support tool that shortens planning, reduces bias, and improves treatment selection.
Characterizing multiscale pore networks in natural and engineered porous media is critical for understanding fluid transport, storage capacity, and material performance, yet conventional imaging techniques often fail to resolve submicron porosity. While advanced methods like focused ion beam scanning electron microscopy (FIB-SEM) provide high-resolution three-dimensional (3D) data, they are limited by small sample sizes, high costs, and destructive sample preparation. This study presents an innovative algorithm to determine the permeability-porosity correlation of submicron pore space using high-resolution scanning electron microscopy (SEM) images, enabling non-destructive, large-scale analysis. The approach addresses key challenges in converting two-dimensional SEM data into reliable 3D models by (1) identifying submicron porosity regions, (2) processing and binarizing SEM fragments, and (3) generating statistically equivalent 3D reconstructions for fluid flow simulations. By integrating texture-based segmentation of SEM images with mineralogical data from energy-dispersive X-ray spectroscopy (EDS) mapping, new method allows for the characterization of heterogeneous submicron pore systems and their incorporation into multiclass digital rock models. This advancement enhances the prediction of petrophysical properties in unresolved pore networks, offering a practical alternative to labor-intensive 3D imaging techniques. The algorithm’s application spans hydrocarbon recovery, geothermal systems, and advanced material design, providing a scalable solution for multiscale porosity analysis in energy and environmental research.
Electric submersible pumps are mission-critical assets in oil production, where intelligent decision support and early-warning systems are required to reduce unplanned downtime and support proactive maintenance planning under resource constraints. This study proposes a machine-learning-based decision-support framework for multi-horizon early warning of electric submersible pump failures using sparse daily surface telemetry. The dataset comprises 1,273,208 telemetry records from 843 wells and 1,162 documented failure or intervention events. Eight prediction horizons of 1, 3, 7, 14, 21, 30, 45, and 60 days are formulated using strictly history-only feature construction and leakage-free validation with temporal holdout and well-grouped cross-validation. To support actionable decision making, daily risk scores are transformed into confirmed alert episodes using hysteresis and minimum confirmation rules. Performance is assessed using both conventional window-based matching and a strict horizon-aware evaluation protocol that enforces actionable lead time. The results show a gradual degradation of discrimination with increasing horizon, while episode-level performance remains stable under realistic reaction-time constraints. An ensemble combining horizon-specific classification with a time-to-failure proxy improves the robustness of risk trajectories and reduces spurious alerts. The framework is intended to support predictive maintenance and industrial monitoring workflows for electric submersible pumps.
Quantitative determination of mineralogy, through laboratory core studies and high-definition spectroscopic logging, is effective but underutilized due to cost and complexity. Unconventional formations present additional challenges, such as kerogen presence, heterogeneity, and anisotropy. This problem can be addressed by utilizing well logs and thermal profiling with specialized wrappers, such as multioutput regressor and regressor chain. Several machine learning models and strategies for combining well logs on multiscale data from an unconventional formation in West Siberia were tested to predict the mass and volumetric fractions of minerals obtained from the Litho Scanner. The gradient boosting regressor, wrapped in a regressor chain and combined with conventional well logs, demonstrated superior performance in predicting both mineral weight and volume fractions, effectively capturing the heterogeneity of the rock structure. A comparison between the machine learning-based model and the Litho Scanner showed an average discrepancy, measured by the root mean squared error for weight fraction, of 0.026 in the Bazhenov Formation. The relationship between certain minerals and the thermal properties of the rock was validated by assessing the importance of thermal core logging data for quartz and pyrite. Moreover, the volume fraction of the rock matrix, composed of total organic carbon and other minerals, was predicted more accurately by incorporating thermal core logging data. The mineral densities, required for obtaining mineral volumes, were determined by solving an optimization problem. Subsequently, a theoretical model was used to calculate thermal conductivity from the mineral volume fractions, revealing a significant similarity between the predicted and experimental values.
During the development of gas-condensate fields, companies often face flow assurance issues. They may lead to unplanned standstills of wells, damaged assets, and substantial financial losses. To help reservoir engineers to identify all the patterns of upcoming flow assurance issues and perform preventing actions timely, we propose to use a machine learning detector. In this paper, we consider the performance of a gradient boosting machine learning model while predicting two types of flow assurance issues appearing at the bottomhole zone: hydrate plugging and condensate banking. After expanding the limited set of production parameters with retrospective features and features of surrounding wells we trained the model as a binary classifier. At the blind test, such values of quality metrics were achieved as Recall ≈0.78 and Precision ≈0.64. The investigation of feature importance revealed that the changes in gas flow rate together with changes in bottomhole and wellhead pressure are the most important parameters for classification. That is fully consistent with the underlying physics.
The rapid development of Digital Rock Physics (DRP) requires the elaboration of robust techniques for closing the gaps between different scales of rock studies (upscaling). The upscaling workflows are especially needed to support the applicability of DRP for heterogeneous rocks. Basically, DRP involves two primary stages: model construction and simulation of physical processes on the models created. For heterogeneous rocks, there is an inherent trade-off between the spatial resolution of the data and the representativeness of the model size. The primary objective of this study was to implement and test a technique for upscaling digital core models from microscale to macroscale, enabling the computation of rock properties while accounting for heterogeneity of various scales. The upscaling is based on establishing correlations between tomography data of different resolutions and transforming low-resolution tomography into a multi-class model according to the defined correlation. The convolutional neural network for high-resolution tomography data was considered as the optimal algorithm for transforming low-resolution tomography into a multi-class model. The output of the neural network was an upscaled model of lower resolution than the original tomography image. Each cell in the upscaled model belonged to one of several types of formation, whose generalized characteristics were determined on the basis of the analysis of high-resolution tomography data. To validate the upscaling technique, we constructed a digital model of a complex carbonate reservoir based on data from multi-scale microtomography (mu CT). A Darcy-scale model has been used and validated as a multi-class model, enabling the computation of flows in pore samples of various scales. By incorporating diverse pore space structures as different classes in the Darcy-scale model, it is possible to preserve the substantial physical size of the model while enhancing its level of complexity.
Abstract Quantitative determination of mineralogy can be done using high-definition spectroscopic logging methods, however these methods are rarely used due to complexity and cost. Also, it is difficult to obtain mineralogical composition in unconventional formations due to presence of kerogen and high heterogeneity and anisotropy of such formations. This problem can be resolved by utilizing Machine Learning algorithms based on well logging and thermal profiling data which can improve and speed up reservoir characterisation. Special wrappers such as Multioutput Regressor and Regressor Chain were applied to test several machine learning models and strategies of well logs combinations on multiscale data from an unconventional formation in West Siberia to predict mass and volumetric fractions of minerals obtained from Litho Scanner. RMSE and MAE were used as regression metrics. To validate the results, theoretical model was used to calculate thermal conductivity based on mineral volume fractions and compared with experimental data. Regressor Chain showed better performance for weight fractions prediction when data was scarce. The Gradient Boosting Regressor encapsulated within a Regressor Chain exhibited the most favorable outcomes in relation to the precision of mapping. The evaluation contrasting the ML-based model with the LithoScanner exhibited an average discrepancy of 0.026, as measured by the RMSE metric.
Well workovers are an inevitable part of any oil- or gas-producing well’s lifecycle. Today, the selection of wells for performing workovers strongly relies on a set of rules evolved from the experts’ experience and the company’s standards. The rapid assessment of big data from oil and gas wells using modern machine learning (ML) models can make this process more effective, timely highlighting the decline in the efficiency of each well. One of the challenges of this task is preparing data for training. Together with such basic steps as data cleansing, extensive feature engineering and data balancing have to be performed. After that, ML models can be applied to solve a classification problem. Among all the analyzed models, gradient boosting appeared to be the most promising one. The quality metrics demonstrated that the model accurately predicts the points at which the expert planned a well workover or the possible issue that will eventually require the workover. However, together with true values, some “false alarms” appear, whose amount can be adjusted. Thus, the algorithm can be used to automate the process of analyzing vast amounts of data, prolong the lifetime of wells, and reduce operational expenses significantly.
The Capacitance-Resistance Model (CRM) has been a useful physics-based tool for obtaining production forecasts for decades. However, the model's limitations make it difficult to work with real field cases, where a lot of various events happen. Such events often include new well commissioning (NWC). We introduce a workflow that combines CRM concepts and kriging into a single tool to handle these types of events during history matching. Moreover, it can be used for selecting a new well placement during infill drilling. To make the workflow even more versatile, an improved version of CRM was used. It takes into account wells shut-ins and performed workovers by additional adjustment of the model coefficients. By preliminary re-weighing and interpolating these coefficients using kriging, the coefficients for potential wells can be determined. The approach was validated using both synthetic and real datasets, from which the cases of putting new wells into operation were selected. The workflow allows a fast assessment of future well performance with a minimal set of reservoir data. This way, a lot of well placement scenarios can be considered, and the best ones could be chosen for more detailed studies.
Global warming is considered as a severe growing crisis over the next years. Carbon Capture and Storage (CCS) has been proposed as a feasible solution to stop the uprising trend of temperature. The satisfactory performance of the CCS implementation is obtained only if the relevant mechanisms are understood well. Undoubtedly, multiphase flow is the most significant common aspect of all effective mechanisms. Relative permeability is recognized as the parameter to describe multiphase flow physics qualitatively and quantitatively. Moreover, previous studies have shown that pore-scale phenomena strongly influence relative permeability data. However, the classic experimental procedures of relative permeability measurements are inherently unable to describe how microscopic phenomena like pore geometries, fluid-fluid interactions, and flow regimes affect the relative permeability data. Based on micro X-ray Computed Tomography (mu xCT) images and Computational Fluid Dynamics (CFD), the current research has put forward a systematic workflow to figure out how pore geometry impacts the relative permeability data. Furthermore, the effect of computational domain size on calculated relative permeability data has thoroughly been investigated as well. The results show that realistic relative permeability data are acquired if the effects of pore-scale phenomena are considered.
Capacitance-Resistance model (CRM) has been a useful tool for fast production forecasts for decades. The unique combination of simplicity and physics-based nature in this data-driven approach allowed it to stay as an object of scientific interest and get its own place among other types of models capable of giving predictions on flow rates, such as full-scale 3D reservoir models. However, the model simplicity, assumptions, and limitations does not allow wide application of a conventional CRM in complex field cases. A vast majority of studies on CRM are about overcoming its limitations by introducing new coefficients, modifying the analytical form of the equation, etc. Integrating CRM with rapidly developing artificial intelligence (AI) methods seems to be a logical continuation of model evolution. Recently introduced Physics-Informed Neural Networks (PINN) can preserve CRM's governing equations and coefficients that gives some insights about wells and formations standing out from other popular machine learning and deep learning methods. Moreover, PINN type models give certain flexibility in the choice of architectures – it means that the model architecture can be changed in a way that may assist in solving different problems. Thus, we introduce end-to-end learning of neural networks (NN) while implying some physical constraints. It is intended to overcome one of the major limitations of CRM, which is obtaining predictions for oil and water production rates from total liquid. This way, the additional training of rough approximation fractional flow models that are either not suitable for the case or may require the knowledge of reservoir properties is not needed. In this work, the well-known concept of Capacitance-Resistance models appears in a new form, which allows performing history matching rather rapidly, achieving robustness and forecasting liquid, oil and water production rates simultaneously. To test this new approach, several datasets (both synthetic and real) were used. The results obtained by PINN are compared to those obtained by a conventional CRM. By conventional we mean the analytical solution, which was modified by our research group to take into consideration common real field cases such as shut-in wells, workover operations, etc. by introducing dynamic characteristic coefficients [1].
To ensure the required level of production of brown fields, it is necessary to plan and implement an efficient program of well stimulation in a timely manner. This program has a significant impact on the further development of the oil field in terms of its productivity and economics. Nowadays, the decisions about well stimulations are made by experts mostly based on their experience and the company’s heuristics. In the work, a novel approach to fast, accurate, and computationally efficient selection of candidates for well stimulations was presented. We explored the ability of Machine Learning algorithms to solve the problem. The predictive pipeline was built on the basis of production and pressure time series of historical data. We developed and applied a comprehensive prepossessing workflow to a real field data to prepare a training data set. Finally, two Gradient-Boosting models were developed, tuned, trained and validated. The first model was used for prediction of the necessity for the well stimulation and the second model—for identification of the required treatment type. Blind test of the models was provided. It resulted in a 0.80 recall score and 0.79 balanced accuracy score.
A basic mathematical model of the deformation of a large elastic element of a small spacecraft in its plane is constructed. Deformations are caused by a temperature shock after a small spacecraft leaves the Earth’s shadow on the solar portion of the orbit. The model is used to conduct a computational experiment with the aim of assessing perturbations acting on a small spacecraft due to temperature shock. The temperature distribution during thermal shock is described by a one-dimensional model of thermal conductivity. The classical theory of thin plates is used to determine the deformations. The results of the estimation of disturbing factors are obtained as a result of a computational experiment for a model small spacecraft. These results indicate the need to compensate for the impact of temperature shock for small technological spacecraft. The data obtained can be used in the design of small space-craft for technological purposes.
Well known oil recovery factor estimation techniques such as analogy, volumetric calculations, material balance, decline curve analysis, hydrodynamic simulations have certain limitations. Those techniques are time-consuming, require specific data and expert knowledge. Besides, though uncertainty estimation is highly desirable for this problem, the methods above do not include this by default. In this work, we present a data-driven technique for oil recovery factor estimation using reservoir parameters and representative statistics. We apply advanced machine learning methods to historical worldwide oilfields datasets (more than 2000 oil reservoirs). The data-driven model might be used as a general tool for rapid and completely objective estimation of the oil recovery factor. In addition, it includes the ability to work with partial input data and to estimate the prediction interval of the oil recovery factor. We perform the evaluation in terms of accuracy and prediction intervals coverage for several tree-based machine learning techniques in application to the following two cases: (1) using parameters only related to geometry, geology, transport, storage and fluid properties, (2) using an extended set of parameters including development and production data. For both cases model proved itself to be robust and reliable. We conclude that the proposed data-driven approach overcomes several limitations of the traditional methods and is suitable for rapid, reliable and objective estimation of oil recovery factor for hydrocarbon reservoir.
Waterflooding is a widely used secondary oil recovery technique. The oil and gas industry uses a complex reservoir numerical simulation and reservoir engineering analysis to forecast production curves from waterflooding projects. The application of such standard methods at the stage of assessing the potential of a huge number of projects could be computationally inefficient and requires a lot of effort. This paper demonstrates the applicability of machine learning to rate the outcome of waterflooding applied to an oil reservoir. We also explore the relationship of project evaluations by operators at the final stages with several performance metrics for forecasting. Real data about several thousand waterflooding projects in Texas are used in the current study. We compare the ML models rankings of the waterflooding efficiency and the expert rankings. Linear regression models along with neural networks and gradient boosting on decision threes are considered. We show that machine learning models allow reducing computational complexity and can be useful for rating the reservoirs, with respect to the effectiveness of waterflooding.
In many branches of earth sciences, the problem of rock study on the microlevel arises. However, a significant number of representative samples is not always feasible. Thus the problem of the generation of samples with similar properties becomes actual. In this paper we propose a deep learning architecture for three-dimensional porous medium reconstruction from two-dimensional slices. We fit a distribution on all possible three-dimensional structures of a specific type based on the given data set of samples. Then, given partial information (central slices), we recover the three-dimensional structure around such slices as the most probable one according to that constructed distribution. Technically, we implement this in the form of a deep neural network with encoder, generator, and discriminator modules. Numerical experiments show that this method provides a good reconstruction in terms of Minkowski functionals.
The efficient development of tight hydrocarbon resources based on the proper understanding of rock properties can be underlined as a game-changing scenario of the energy market during the coming years. Digital Rock Physics (DRP) is an advanced technique for the accurate assessment of formation properties with respect to the effects of pore structures. However, there are doubts about the usefulness of DRP for tight systems where a large portion of pores are sub-resolved, and subsequently, their effects on numerical computations cannot be taken into account. By taking advantage of Xenon-enhanced computed tomography and nano-scale imaging, the current research puts forward a hybrid workflow that reveals sub-resolved pores and assigns them a proper set of petrophysical properties that can be used for further numerical simulations. The introduced method was applied to a tight sandstone sample. Conventional micron-scale x-ray computed tomography images show no connectivity of the porous space. That is why computation of permeability is impossible. The developed workflow created a three-substance media including pores, solids and submicron porous segments. The triple reconstructed porous media allows introducing the Stokes-Brinkman equation to simulate flows. The results indicated that the computed permeability is reasonably close to the experimental value. The main advantage of the launched method is that the computations achieved satisfying results without increasing the spatial resolution of images or limiting the field of view.
Most methods for automated full-bore rock core image analysis (description, colour, properties distribution, etc.) are based on separate core column analyses. The core is usually imaged in a box because of the significant amount of time taken to get an image for each core column. The work presents an innovative method and algorithm for core columns extraction from core boxes. The conditions for core boxes imaging may differ tremendously. Such differences are disastrous for machine learning algorithms which need a large dataset describing all possible data variations. Still, such images have some standard features - a box and core. Thus, we can emulate different environments with a unique augmentation described in this work. It is called template-like augmentation (TLA). The method is described and tested on various environments, and results are compared on an algorithm trained on both 'traditional' data and a mix of traditional and TLA data. The algorithm trained with TLA data provides better metrics and can detect core on most new images, unlike the algorithm trained on data without TLA. The algorithm for core column extraction implemented in an automated core description system speeds up the core box processing by a factor of 20.
We disclose a new -age field -scale production forecast model that handles complex treatment of wellbores during their life cycle. Predictive production models have been an object of increased interest and research for a long time due to the need for a fast tool for forecasting production rates or choosing an optimal field development scheme. The existing approaches based on the material balance equation have several limitations and are not very applicable for real objects. Full -scale reservoir modeling is relatively slow and requires large computing resources. In this paper, we propose a proxy model based on advanced capacitance-resistance approach. The model predicts multiphase flow rates based on the available historical data of field production and information about well treatments. In addition, it pro-vides preferable transmissibility trends, the presence of sealed or leaking faults, and the degree of dissipation between injector-producer well pairs. The advanced feature of the model is time-dependent weight coefficients, which have not been studied previously. They help in accounting the shut -in and workover periods and can be found during the optimization procedure simultaneously. Another feature is fast calculations due to a vectorized form of the model and application of modern optimization techniques. All these options allow modeling real oil fields with a large number of wells and a complex system of production control.