
Subsurface media exhibit significant heterogeneity and uncertainty, making the construction of accurate and geologically consistent three-dimensional geological models under limited well constraints a central challenge in hydrogeological modeling. Traditional multiple-point simulation methods rely on the assumption of regional stationarity of spatial patterns, whereas existing generative adversarial network (GAN)-based approaches require massive training images to encode prior geological knowledge and remain limited in capturing multi-scale spatial features. However, spatial structures vary significantly across regions. Realistic and reliable training images are scarce and costly to obtain. These limitations severely restrict the generalization and applicability of these methods. In this study, we propose an attention-enhanced conditional SinGAN (AE-CSinGAN) method for generating three-dimensional subsurface structural models. Under limited well data and a single training image, the method uses a concurrent multi-stage network architecture to achieve multi-scale information fusion. A mixed attention mechanism is incorporated into the generator to efficiently model and dynamically focus on both the conditional constraints and training image. Results from five types of experiments show that AE-CSinGAN achieves high reconstruction accuracy at conditioned well locations. It also preserves reasonable geological continuity and realism. Comparative analyses with the direct sampling (DS) multiple-point simulation method and the baseline CSinGAN model further demonstrate the effectiveness and superiority of the proposed method.
A novel machine learning approach to seismic full waveform inversion is developed by adopting agentic artificial intelligence (AI) concepts. The method involves the concatenation of multiple learning algorithms to generate an autonomous system that efficiently explores and learns the parameter space to perform high-resolution full-waveform inversion. Physics-based exploration agents are used to generate high-quality supervisory signals in the form of pseudo labels to reinforce the learning of the local data domain. Reinforcement learning agents, rooted in either L2 local or stochastic global optimizers, are coupled to active learning sampling techniques to select the most informative data samples and enhance the learning process. Physics-based agents are subject to a long-term cumulative reward policy of data misfit reduction. The agentic system autonomously learns from the unlabeled pool of data and provides domain adaptation for effective self-supervised learning. The system performs full-waveform inversion in an autonomous and adaptive manner, using local or global optimizers to guide the inversion process. The surrogate machine learning inversion matches or outperforms the resolution of conventional optimization techniques and significantly reduces computing resource utilization. On the realistic test, the agentic framework achieves speed-ups of up to 10× for deterministic schemes and 37× for stochastic Markov chain Monte Carlo (MCMC) runs compared with standard inversion workflows. The developed self-supervised agentic approach effectively performs domain adaptation on any dataset, providing a highly efficient, autonomous, and adaptive framework for general physics-based inversion applications.
An extension of the Bayesian linearized AVO inversion framework is presented to jointly incorporate PP- and PS-wave seismic reflection data for improved prediction of elastic subsurface properties. The inversion is based on a convolutional model and a weak contrast linearization of the Zoeppritz equations. The inclusion of PS-wave data increases sensitivity to shear properties, enabling enhanced resolution of S-wave velocity and elastic parameter ratios such as V_P/V_S . The inverse problem is formulated in a Bayesian Gaussian framework, with explicit expressions for the posterior mean and covariance of the model variables, including P-wave velocity, S-wave velocity, and density. The joint inversion of PP- and PS-wave data leads to improved constraints on the posterior distributions, compared to the PP-only inversion. Synthetic tests, based on a field dataset from the Norwegian Sea, demonstrate that the joint inversion reduces prediction uncertainty and improves accuracy under realistic noise conditions. Application to a 3-D multicomponent field dataset from a Brazilian oil field shows better match of the inversion results with well log data compared to PP-only results, particularly for V_P/V_S , confirming the added value of incorporating PS information into the Bayesian AVO inversion framework.
Borehole thermal energy storage (BTES) systems represent a promising solution for long-term thermal energy management and the integration of renewable heating and cooling technologies. However, their performance remains strongly dependent on subsurface geological characteristics, particularly thermal conductivity heterogeneity, which governs conductive heat transport and storage efficiency. This study investigates the impact of geological variability on BTES performance and evaluates the extent to which borehole-level operational optimization can mitigate these effects. A three-dimensional, physics-based BTES model was developed in FEFLOW to simulate three consecutive years of operation under alternating 6-month charging and discharging cycles. Heterogeneous thermal conductivity fields were generated using sequential Gaussian simulation with variance levels of 1 and 4. A single realization from each variance level was selected to represent moderate and strong heterogeneity, respectively, while a homogeneous case employed a constant geometric mean thermal conductivity of 1.39 W m ^-1 K ^-1 for baseline comparison. A genetic algorithm was integrated with the numerical model to optimize flow rates under two operational strategies: global (uniform across all boreholes) and individual (borehole-specific) control. Results indicate that increasing heterogeneity reduces thermal recovery efficiency due to distorted heat propagation, localized heat trapping, and nonuniform plume development. Individualized flow-rate optimization improved recovery efficiency by approximately 3–4
Basalt formations have emerged as promising geological reservoirs for long-term CO2 storage due to their widespread distribution and ability to mineralize injected CO2 into stable carbonate minerals. However, their highly variable textures and porosities, resulting from different cooling histories, produce pronounced heterogeneity and anisotropy that influence reservoir properties and CO2–rock interactions. This study presents a micro X-ray computed tomography (microCT)-based framework to quantify heterogeneity and anisotropy in Deccan basalts and their evolution following CO2 exposure. Three metrics were employed: (i) grayscale heterogeneity using the coefficient of variation (CV), (ii) structural heterogeneity using the power-law spectral slope (β), and (iii) anisotropy using the angular autocorrelation-derived anisotropy ratio (AR). Textural analysis showed that vesicles and amygdales introduced strong heterogeneity and directional structural bias, resulting in high CV, more negative β, and larger AR values. In contrast, fine-grained dyke and massive basalts displayed lower variability and remained largely isotropic. Following CO2 exposure, the basalt samples showed an increase in porosity, accompanied by a 5.71–7.40
Faults control hydrocarbon migration, trap sealing, and reservoir compartmentalization, so their accurate identification is critical for petroleum exploration and development. As seismic data volumes grow rapidly, traditional fault interpretation methods based on manual analysis and simple image processing struggle with efficiency, consistency, and noise robustness. To overcome these limitations, this paper proposes an unsupervised, geologically constrained segmentation framework for extracting faults from complex coherence images. The method exploits the fractal and geometric characteristics of faults to build a training-free, geologically driven workflow. First, adaptive thresholding combining the Whale Optimization Algorithm (WOA) and a maximum entropy criterion extracts preliminary fault skeletons. Second, a dual-constraint denoising strategy based on fractal dimension and geometric morphology, especially aspect ratio, removes nonlinear geological and acquisition noise while preserving linear faults. Third, a topological reconstruction module guided by local centroid orientation and lateral offset reconnects fault segments fragmented by noise or tectonic shearing, thereby restoring geometric continuity. Finally, a multiscale structure tensor coherence framework refines fault edges and optimizes morphology through the selection of seed points exhibiting high confidence and region growing guided by the principal orientation, improving boundary localization and structural completeness. Applied to a 2D seismic dataset from the Xinglongtai structural zone in the Liaohe Depression, Bohai Bay Basin, the method outperforms traditional edge detectors and global thresholding by extracting more continuous and complete fault networks with clearer, better-defined edges. Overall, this workflow provides an efficient, robust, and geologically interpretable solution for fault interpretation in complex survey areas where extensive annotated data are unavailable and conventional supervised methods are difficult to apply.
Ground-penetrating radar (GPR) is a crucial tool for shallow subsurface investigation, enabling the identification of soil composition, utility lines, and other features. However, environmental factors and hardware limitations often introduce noise during data acquisition, leading to inaccurate hyperbola identification and interpretation. This research proposes a novel noise attenuation method, the GMD-BF algorithm, which combines geometric mode decomposition (GMD) with a bilateral filter. The process begins by using GMD to decompose noisy GPR data into multiple band-limited modes ranging from low to high frequencies. These modes are then classified into noise-dominant and signal-dominant sets using a similarity measurement test. Subsequently, a bilateral filter is applied exclusively to the noise-dominant modes, achieving selective noise attenuation while preserving critical edge and structural features. The GMD component is particularly effective at isolating directional signatures, while the bilateral filter excels at preserving geometric integrity during denoising. Experimental results on both synthetic and real GPR datasets demonstrate the superiority of the proposed GMD-BF method in improving the signal-to-noise ratio (SNR) and enhancing hyperbola clarity with minimal signal degradation, outperforming existing techniques.
Geochemical anomaly identification is a fundamental task in mineral exploration, aiming to extract mineralization-related geochemical signals from complex geological background. Recent advances in artificial intelligence (AI) have developed novel methods for geochemical anomaly identification. However, the challenge lies in developing models that improve predictive performance while remaining geologically interpretable. This study proposes an explainable correlation-augmentation graph reinforcement learning model that integrates a correlation-augmentation module with inverse reinforcement learning to identify geochemical anomalies associated with gold mineralization in northwestern Hubei Province of China. The study area, located along the southern margin of the East Qinling orogen and dominated by orogenic gold deposits, contains 5,099 1:200,000-scale stream sediment geochemical samples, with analyses of 39 elements. Comparison results showed that the correlation-augmentation module reduced redundant information and highlighted elements associated with key ore-controlling factors, thereby improving model performance and interpretability for geochemical anomaly identification. Inverse reinforcement learning showed that regions important to model decision-making were spatially correlated with ore-controlling factors, providing geological support for the predicted anomalies. Overall, this framework provides a promising way to bridge AI-driven geochemical anomaly detection and geological knowledge, offering a transparent tool for geochemical exploration.
An essential preprocessing step in geochemical data analysis is the identification and handling of censored data that occur when elemental concentrations fall below analytical detection limits. These censored observations must be estimated in a way that preserves the spatial coherence and multivariate structure of the dataset. Conventional imputation approaches, such as constant-value substitution or distribution-specific parametric models like the multiplicative lognormal method, often introduce biased estimates, distort variance structure, and weaken inter-element correlations. Modern machine-learning (ML) models, on the other hand, can capture the nonlinear dependencies typical of geochemical systems; however, they are not inherently designed to accommodate the compositional constraints associated with censored observations. This paper therefore introduces ReX-CoDA (Regression-based Flexible-Limit Compositional Data Imputation) as a model-agnostic iterative framework that integrates predictive modeling with Mills-ratio-based correction to enable the estimation of censored values in a compositionally coherent manner. The framework is demonstrated using different ML techniques, including random forest (RF) and extreme gradient boosting (XGB), on geochemical datasets from the Geological Survey of Norway, comprising both artificially censored copper data with known ground truth and naturally censored silver, mercury, and tantalum data. Results show that ReX-CoDA performs competitively against established imputation methods and provides a robust, scalable, and compositionally consistent alternative for handling censored geochemical data.
The global increase in flood occurrences has necessitated the development of flood susceptibility models (FSMs) and their mapping as essential tools for mitigating flood impact. While FSM studies have surged, they often assume that all conditioning factors (CFs) are static, an assumption that does not hold for dynamic factors such as land cover and rainfall. Also, current research has overlooked the distinction between mean and total rainfall inputs and the importance of seasonality and timing, creating a clear methodological gap. To address this, this study develops a workflow that integrates forecasted temporal seasonal climatic factors (TSCFs) with static geospatial features (SGFs) for flood susceptibility across distinct seasonal phases in Anambra and Delta States, Nigeria. A two-dimensional convolutional neural network–long short-term memory model forecasts seasonal rainfall and land cover for 2025 from historical data. These forecasts are combined with nine SGFs using a deep hybrid attention model (DHAM) to produce seasonal flood susceptibility maps, benchmarked against a random forest (RF) baseline. The DHAM shows strong and stable predictive performance and a smaller train-to-test generalisation drop than the RF model across both rainfall scenarios, indicating more robust spatial generalisation. Shapley Additive Explanations (SHAP) analysis shows that the DHAM dynamically reweights TSCFs and SGFs based on the seasonal phase and rainfall input type. These findings confirm that treating CFs as static obscures the seasonally dynamic nature of flood conditioning. Results identify the low-lying plains of the study area as the most susceptible region and show that using mean rainfall alone systematically underestimates the spatial extent of flood-susceptible areas. This study highlights the critical role of TSCFs in flood assessment and provides actionable, season-aware guidance for flood management and disaster preparedness.
The association of incomplete fault observations is a complex task for which several solutions exist. This problem has already been addressed using a probabilistic approach, where pairwise expert rules are considered. An approach considering multiple-point interactions could better represent geological knowledge, but defining such rules from expert knowledge is challenging. To address this limitation, a Random Forest is used to estimate, from representative three-dimensional models, the potential that a triplet or pair of sparse observations belongs to the same fault. In addition, a new analytical model expresses the probability density function on the graph that combines the triplet and pair potentials previously described. Two random forest learners are trained, one from pairs of fault features and the second from triplets of fault features. The features are computed from fault sticks extracted from a known three-dimensional geological model (e.g. the length of the fault stick, the throw value, etc.). The learners are tested on fault sticks extracted on a contiguous sector of the same model. The resulting three-point potentials are different from combinations of pair potentials. Also, the combination of triplet and pair potentials shows better results than considering pair potentials alone.
Geological and grade uncertainty influence the economic performance and feasibility of underground mining projects, particularly in sublevel stoping operations. Traditional deterministic approaches to stope layout optimization fail to account for this uncertainty, often leading to suboptimal designs and increased financial risks. This research addresses these challenges by proposing a novel reinforcement learning (RL) framework to optimize underground stope layouts under uncertainty. The framework employs proximal policy optimization (PPO), a deep reinforcement learning algorithm, to balance profitability and financial risk using the conditional-value-at-risk metric. An ensemble of geostatistical realizations captures the uncertainty in mineral grade and geological variability, enabling the evaluation of risk-adjusted profit distributions for stope designs. The proposed methodology generates a range of design options tailored to varying risk tolerances, enabling decision-makers to align designs with project-specific objectives and risk attitudes. A case study of an underground gold deposit, comprising 840,000 blocks and 100 realizations, demonstrates the effectiveness of the framework. The RL-based designs outperform conventional deterministic methods by forming an efficient frontier of stope layout options that balance expected profit and uncertainty. At the opportunity-seeking behavior, a 5
Whole-rock geochemical data have been utilized to construct intelligent discriminant models for predicting the magmatic fertility of porphyry systems. However, most studies focus on subduction-related systems, with little research on collision-related ones. Moreover, the key geochemical controls on fertility under distinct tectonic regimes remain unclear. The Tibetan Plateau, which records the full geodynamic evolution from Tethyan subduction to continental collision, hosts numerous porphyry deposits formed in diverse settings, providing an ideal natural laboratory. Here, we constructed predictive models of magmatic fertility using random forest, eXtreme gradient boosting, and support vector machine algorithms, applicable across the entire plateau. These models simultaneously encompass both subduction-related and collision-related porphyry systems, and successfully predict magmatic fertility records of the Tibetan Plateau since the Jurassic. Furthermore, by integrating machine learning feature importance analysis and Monte Carlo simulations, we identified that the features most strongly associated with fertile magmas across the plateau’s porphyry systems are high SiO2, K2O, and Na2O, coupled with low TiO2, MgO, and MnO. Further analysis focusing exclusively on subduction-related porphyry systems revealed that Al2O3, high SiO2, and K2O, along with low total Fe, MgO, and Na2O, are the most critical indicators of fertile magmas. In contrast, for collision-related systems, Al2O3, K2O, and high SiO2, combined with low MnO, Tb, and Rb, emerge as the most influential features. These results suggest that, regardless of tectonic setting, magmas that have undergone prolonged intracrustal differentiation processes exhibit higher metallogenic potential. Additionally, compared to barren magmas, fertile magmas generally display stronger signatures of crust–mantle magma mixing.
The extraction of mineralization-related geochemical anomalies through data mining has become a critical component of mineral exploration. With the accumulation of multi-element and high-dimensional geochemical survey datasets, traditional analytical methods have become increasingly inadequate for capturing the complex nonlinear patterns associated with mineralization. Consequently, deep learning algorithms have been applied to geochemical anomaly identification due to their strong capability for nonlinear feature learning and modeling of complex relationships among multiple geochemical variables. The geochemical spatial patterns are strongly controlled by geological processes and commonly exhibit spatial heterogeneity, anisotropy, and directional continuity. Although graph-based representations are well suited for modeling non-Euclidean spatial relationships, conventional static graph construction strategies often struggle to capture the spatial heterogeneity and directional characteristics inherent in geochemical survey data. In addition, static graph neural networks rely on a fixed graph topology throughout the training process, limiting their ability to modify node connectivity based on learned high-level features and to continuously focus on the most important nodes. In order to overcome these limitations, we propose a dynamic graph representation method. By introducing a dynamic edge construction mechanism while retaining the advantages of graph structure, the proposed approach enables edge weights to be adaptively updated according to both spatial distance and node feature similarity during initial graph construction and successive graph convolution operations. The case study conducted in Southwest Fujian Province, China, demonstrates that the proposed method effectively characterizes the spatial heterogeneity of geochemical patterns associated with mineralization, improving the performance of geochemical anomaly identification and interpretability of the obtained results. The identified geochemical anomalies provide valuable targets and geological insights for the next round of mineral exploration in the study area.
Mineral resource classification, based on confidence levels associated with tonnages and grades, is essential for public disclosure and for assessing deposit maturity and related risks. Although geometric classification methods are transparent and straightforward, they do not explicitly account for local variations in uncertainty. In contrast, geostatistical approaches focus primarily on the quality of grade estimation. The scorecard methodology attempts to integrate multiple parameters relevant to mineral resource classification; however, elements such as the definition of uncertainty classes and the weighting of individual criteria often rely on subjective judgment. This study proposes a hybrid methodology for mineral resource classification that integrates traditional geometric approaches with a quantitative risk map. The risk map incorporates three main sources of uncertainty: (1) geological uncertainty, (2) database quality, and (3) grade estimation quality. To minimize subjectivity in the integration of these factors, the risk map is constructed using a k-means clustering approach, resulting in three risk categories: low, medium, and high. The final classification is obtained through a hierarchical combination of geometric classification and the risk map. A machine learning post-processing step is applied to mitigate local classification artifacts, commonly referred to as “salt-and-pepper” or “spotted dog” effects. The proposed methodology was applied to a stratigraphic deposit case study and yielded robust results. It effectively integrates conventional geometric classification with key sources of uncertainty, reduces subjectivity compared with traditional scorecard methodology, and ensures reproducibility in mineral resource classification.
Clustering techniques for grouping georeferenced data points (Q-mode clustering) have advanced considerably, whereas the clustering of georeferenced variables themselves (R-mode clustering) remains largely unexplored. Traditional non-spatial R-mode clustering methods are generally inadequate in spatial contexts, as they ignore the spatial dependence structure of regionalized variables, which is a fundamental aspect of multivariate spatial data analysis. This paper proposes a geostatistical R-mode clustering framework that explicitly incorporates spatial correlation through a (dis)similarity measure derived from Matheron’s codispersion coefficient. This framework generalizes conventional non-spatial R-mode clustering approaches based on pairwise associations, such as the Pearson correlation coefficient. Applied to soil geochemical datasets, the method identifies two types of variable groupings: spatially coherent groups that remain stable across multiple spatial scales, and scale-dependent groups whose internal composition varies with the spatial scale considered. The proposed framework addresses a critical gap in multivariate spatial data analysis and provides a robust alternative to classical non-spatial R-mode clustering techniques for detecting spatially meaningful patterns.
We derive closed-form expressions for the Newtonian potential, gravitational acceleration, and gravitational gradient tensor generated by a homogeneous, planar, zero-thickness polygon embedded in three-dimensional space. The analysis is formulated for a general simple polygon with constant areal mass density σ and is implemented by decomposing the polygon into oriented triangular elements. For each triangle, the formulation expresses the potential and acceleration using edge log-factors and the signed solid angle, and provides an explicit dyadic representation for the tensor in terms of gradients of these quantities. For the important special case of a rectangle, compact corner-sum formulas provide efficient expressions for the potential, acceleration, and tensor. These formulas are numerically stable except at analytically singular boundary configurations, such as edges and corners of the ideal zero-thickness surface-mass model. Four independent benchmarks validate the theory: (i) a disk benchmark that demonstrates rapid convergence of the triangulated solution to the analytical axis field; (ii) the thin-cuboid limit, confirming that the rectangle model is the correct zero-thickness limit of a finite cuboid (or plate) with matching masses; (iii) reconstruction of the cuboid potential by stacking rectangles, confirming correct volume integration behavior; and (iv) agreement of rectangle and two-triangle decompositions at both acceleration and tensor levels. These results establish the correctness and consistency of the proposed polygonal gravitational primitives for geodetic and geophysical forward modeling.
Effective monitoring of industrial emissions is essential for environmental sustainability. This study presents the Graph Attention Sensor (GAS), a forecasting framework that integrates graph attention networks with Time2Vec to model complex multivariate dependencies among co-located emission sensors. Unlike conventional time-series models, GAS represents pollutant sensors as nodes in a relational graph, enabling selective attention to inter-pollutant interactions while preserving causal temporal dynamics. GAS is benchmarked against random forest, NBeats (neural basis expansion analysis for interpretable time series), temporal convolutional network (TCN), long short-term memory (LSTM), and temporal fusion transformer (TFT) under nested cross-validation on real-world emission data from a high-purity ferrosilicon industrial plant. Results reveal that GAS achieves the highest performance for TSP forecasting under constrained context windows ( P ≤ 22 ) across all forecast horizons, while NBeats dominates for gaseous pollutants ( SO_x , NO_x ) with extended historical context. Across all architectures, GAS exhibits the most stable cross-horizon degradation, maintaining positive R2 for TSP at t_6 , where random forest collapses to sub-baseline performance. Analysis of aggregated attention maps reveals that GAS identifies that NO_x is associated with the furnace thermal state, a physically grounded relationship that emerges from the data without explicit encoding. These findings support a pollutant-stratified deployment strategy and provide practical guidance for model selection in industrial emission monitoring.
The scarcity of labeled samples and pronounced spatial heterogeneity present significant challenges for mineral prospectivity mapping using geoscientific data. This study introduces GeoKIM, a label-informed masked representation learning framework designed for tabular geoscientific datasets under few-shot conditions. During pretraining, the model learns from unlabeled observations through masked feature reconstruction, while limited labeled information is used to construct a correlation-guided masking distribution and to supervise an auxiliary regression branch targeting key metallogenic indicators such as Au, As, and Sb. Evaluation on real-world geochemical data from the Xiahe–Hezuo region is conducted using a spatially aware few-shot protocol with strict spatial partitioning to prevent information leakage. In addition to classical machine learning baselines, GeoKIM is compared with four representative modern tabular learning frameworks, namely FT-Transformer, TabNet, SubTab, and VIME. The results show that, although no method is universally dominant in the extreme 1-shot regime, GeoKIM achieves the best overall accuracy, area under the receiver operating characteristic (ROC) curve (AUC), and F1-score from 5-shot onward and yields the clearest AUC advantage under limited supervision. Ablation results indicate that correlation-guided masking and the Transformer backbone both contribute to downstream performance, while representation analyses suggest improved latent-space organization and reconstruction of mineralization-related indicators. These findings position GeoKIM as a promising label-informed pretraining framework for few-shot mineral prospectivity mapping on tabular geoscientific data.