Lithologic identification plays a crucial role in petroleum geologic exploration, and machine learning (ML) has become increasingly prevalent in intelligent lithology identification in recent years. However, identifying lithologies presents challenges due to a lack of lithologic labels and an imbalanced distribution of lithologies. To address this issue and obtain satisfactory lithologic identification results, this study investigates a class-rebalancing self-training (CReST) lithology identification framework. This framework uses logging data and limited lithologic labels as input and achieves promising lithology classification through the CReST approach. Four ML algorithms with high overall performance are selected from 25 common algorithms to establish CReST models, such as bagging classifier, extra trees classifier, random forest classifier, and support vector classifier. The classification results of the models are compared and analyzed under three conditions. The experimental findings indicate that (1) under label scarcity, the effect of category recognition varies greatly with different sample numbers; (2) under self-training (ST), overall performance is improved, but the difference in performance caused by category imbalance also increases; and (3) under CReST framework, the model effectively resolves the identification problems caused by a lack of labels and an imbalanced category distribution. Specifically, the precision of identifying categories with fewer samples is improved by more than 20%.
ABSTRACT Numerical forward modelling and laboratory experiments suggest that autogenic factors in the sediment routing system serve as long‐pass filters, preserving only orbital cycles with a period exceeding the compensation timescale, T c , or thickness in the depth domain exceeding the compensation depth scale, H c . For a specific orbital cycle with a certain period, this preservation in alluvial strata occurs unless it exhibits a sufficiently large amplitude. This study stratigraphically confirms, for the first time, the long‐pass filtering of autogenic dynamics using elemental data from the alluvial–lacustrine Sifangtai and Mingshui formations in the Songliao Basin. Spectral analysis of the Si and Zr series in coarse‐grained sediments reveals no cyclic signal with thicknesses below the estimated lower limits of H c . This implies that the spatial storage threshold for orbital cycles in proxies of the coarse‐grained sediment component is equal to or less than H c . However, cyclic signals of obliquity and precession with smaller thicknesses are identified in Ti, Fe and Al enriched in the fine‐grained sediment components of the stratigraphy. Notably, previously reported proxies preserving high‐frequency orbital cycles are derived from fine‐grained sediment components, differing from the sedimentation rate series used in the reported experimental studies. Therefore, the authors hypothesize a grain‐size component‐dependent storage threshold, suggesting that the storage threshold of orbital cycles in proxies associated with fine‐grained components is lower. This hypothesis arises from the weaker effect of autogenic dynamics on the content of fine‐grained sediment components transported to the sampling site by a suspended load compared to coarser components that are subjected to stronger autogenic dynamics within or near channels. The hypothesis and model presented propose a dynamic process elucidating the nuanced roles of autogenic dynamics in preserving orbital cycles. This perspective, considering sediment composition, inspires prioritizing proxies enriched in the fine‐grained fraction for identifying allogenic cycles in alluvial strata.
Three-dimensional (3D) geological modeling is a process of interpretation that integrates multiple source inputs and knowledge into geometry to represent the understanding of geologists. When geologists build a high-quality 3D geological model, this process still involves some issues such as sparse drillhole data, imperfect prior knowledge, and sensitive modeling algorithms. Therefore, taking uncertainty as the measurement criterion for the variation extent of the posterior likelihood of the 3D geological model and assisting in increasing the quality of the model are crucial issues in this domain. This paper proposes a novel method based on a (1 + ε)-approximation global optimum strategy, which is a type of big data and machine learning technique, to determine and present the uncertainty hidden in geometry. Compared with previous approaches, our strategy made the following new contributions: (1) the global optimum solution calculated by potential models is utilized to represent the uncertainty at each location; (2) the strategy offers a quantifiable reliability to each model that is involved in the evaluation process, and values of reliability are unknown before the commencement, meaning that they do not depend on expert experience; moreover, they can also be verified by comparing prior knowledge with information that such 3D models possess; (3) compared with previous studies, the number of perturbing models is no longer a key prerequisite for this kind of study to evaluate the quality of one geological model, thereby greatly reducing the computational complexity and improving the practicability. Finally, a case study was conducted to assess the uncertainty of a real 3D geological model in northwest Hunan Province, China.
GIS-based Mineral Prospectivity Mapping (MPM) has been widely employed, however, the absence of correlative interpretation between the final prospectivity maps and metallogeneny has resulted in low credibility in the prediction results. Therefore, in this study, machine learning technology combined with knowledge embedding and explainable ensemble learning is utilized to improve interpretability of the result. The Best-worst method (BWM) was employed for the embedding of prior knowledge. Stacking ensemble learning integrated these weights into the predictive model and generated targets. Permutation Importance (PI), Partial Dependence Plots (PDP), and Local Interpretable Model-agnostic Explanations (LIME), were applied to calculate the global and local output weights of features, thereby enhancing the interpretability of the targets. An experiment was carried out at the Keeryin ore concentration in Sichuan, China. The experiment demonstrates effectiveness of this method, wherein 84 % of samples falling within high and extremely high mineralization probability zones, covering 6.58 % of the total area. Na-feldspar spectrum, Na2O + K2O, ring structures, Li/La, and two-mica granite have emerged as key predictive features in sequence, which displays orderliness and verifies their tight correlations with the metallogenic environment and prospecting indicators of pegmatite-type lithium deposit in Keeryin.
An interpretability model for intelligent lithology identification is proposed, which utilizes Ensemble Learning Stacking, Permutation Importance (PI), and Local Interpretable Model-agnostic Explanations (LIME) techniques. The aim of this method is to provide more accurate geological information and scientific support for oil and gas resource exploration. Two logging datasets from the public domain were used as experiments, and support vector machine (SVM), random forest (RF), and naive Bayes (NB) were employed as base learners, while SVM was utilized as the meta learner for lithology classification via stacking algorithm. The accuracy of the model was verified using evaluation metrics such as Area Under Curve (AUC), precision, recall, and F1-score. The PI and LIME techniques were employed to explain the lithology identification model. The results indicate that the stacking algorithm produced the best indexes and highest prediction accuracy. With respect to overall interpretation, PHIND, GR, and RT were found to have the most significant influence on lithology identification in a natural gas protection area in the United States, while DEN, CAL, and PEF were observed to be the most influential variables for lithology identification in the Daqing Oilfield in China. From the perspective of a single sample, the LIME algorithm can provide a quantitative prediction probability and degree of influence of the characteristic variables.
Geographic information system-based mineral prospectivity mapping (MPM) aims to generate targets by combining multiple proxy layers containing geology, geochemistry, and geophysics information based on an available understanding of geological processes and translating it into critical targeting criteria. However, factors such as an imperfect geological understanding with numerous heuristic natures and intrinsic biases, the inaccuracy and sparsity of datasets, and multiple selections of predictive methods adversely affect the results of MPM and jeopardize the reliability of decision-making in exploration. Thus, a series of knowledge-driven and data-driven MPM approaches to counterbalance these disadvantages have been proposed. Uncertainty is defined as a metric of the various scales of the likelihood and consequences, which is helpful to quantitatively represent the above risks. In this paper, the uncertainty in the final three-dimensional prospectivity map is analyzed and interpreted in terms of quantification, visualization, and comparison of different predictive approaches, and a novel technology is proposed based on a (1 + ε ) approximate global optimum strategy derived from the truth discovery society. It outputs the globally optimal truth ( p* ) as the overall mathematical expectation and a set of weights {w_i } as the representation of reliability. Here, a previous study of the Haoyaoerhudong gold deposit, which is one of the largest black-rock-series-type gold mines in China, was reevaluated. The method demonstrated the following advantages: (i) sorted the reliability of potential models built by multiple predictive variables and different mathematical methods, (ii) provided a “best-guess-decision” prospectivity result by combining the best reliability model from {w_i } and the signal-to-noise ratio (SNR) of the risk–return model, and (iii) provided a statistically final uncertainty model by combining p* and risk information.
Lithology identification is an important task in oil and gas exploration. In recent years, machine learning methods have become a powerful tool for intelligent lithology identification. To address the redundancy of conventional logging data and unbalanced distribution among formation lithology classes due to the complexity of depositional environment and inhomogeneity of subsurface space, this paper investigates the affiliation-weighted one-to-one support vector machine (WOVOSVM) lithology identification method based on geochemical logging data. This method uses geochemical logging data, which can directly reflect the formation lithology information, as input, and achieves intelligent and accurate lithology classification under the calculation of WOVOSVM. In this study, Shahezi Formation of Songke 2 Well in Songliao Basin, China is taken as the experimental object, and two data sets with different distribution characteristics are selected as the input. Use WOVOSVM, Adaboost, random forest (RF) and traditional support vector machine (SVM) to identify lithology, and compare and analyze the results. The results are as follows: (1) Accuracy metrics of most of the four classification models were above 60%, indicating the geochemical logging data can effectively reflect the formation lithology information, which is a reliable indicator for the intelligent identification of logging lithology. (2) When the data set has a strong imbalance, the lithology recognition performance of WOVOSVM is better than other methods, the average value of accuracy metrics is more than 72%, F1 value is 8.77% to 14.56% higher than other models, especially in the small sample lithology category recognition, 70% of the samples are correctly classified.
The mineral system modeling approach for prospectivity mapping is an efficient and economic method to assess undiscovered mineral potential quantitatively. It is a procedure of modeling, acquiring, and coupling the proxies of footprints of mineral systems at multiple scales (e.g., regional, district, and deposit scales). In this approach, the critical issue from multiple scales is that the data collected are asymmetrical from the superficial to the deep or from mine to its brown fields, so that it is hard to employ and integrate them. To complete this study, firstly, multi-tactic 3D geological modeling methods, including the explicit, the implicit, and inversion, were used to build geological models in the condition of asymmetrical datasets at the deposit and district scales. Secondly, indicators acquired in drill-intensive fields among multisource datasets composed of geology, geochemistry, geophysics and alteration data were transferred to studies in deep and brown fields. Finally, deep (~ 1,100 m) and circumjacent potentials of mine were targeted in the Haoyaoerhudong gold deposit situated in the Urad Middle Banner area, Inner Mongolia, which is one of the largest black-rock-series-type gold mines in China. This proposed procedure is more visual, clear, intuitive, and transferable to drive mineral system approach to exploration discovery than previous GIS-based studies.
The demand for shale gas has propelled researchers to focus on precise and high-resolution stratigraphic divisions for homogeneous shales, of which the late-Ordovician Wufeng (O(3)w) and the early-Silurian Longmaxi (S(1)l) formations in southwest China are two of the best candidates for shale gas exploration in China. However, systematic chemostratigraphic work for these strata is still sparse, and the existing chemostratigraphic work either lack representativeness in terms of the proxies used or are subjective during their division procedures. Thus, automatic division process based on multi proxies and an objective statistical technique was applied to establish a quantitative, high-resolution, and robust chemostratigraphic scheme for the Wufeng and lower Longmaxi shales. The geochemical analysis unveils that the Wufeng and Lower Longmaxi shales show prominent heterogeneities in terrigenous inputs, redox conditions, and paleoproductivity, enabling the potential application of chemostratigraphy to these strata. Based on these heterogeneities, the chemostratigraphic scheme for the Wufeng and Lower Longmaxi shales has been established, and the whole strata could be divided into 13 chemozones using constrained clustering analysis. The chemostratigraphic scheme could not only be comparable to the regional sequence stratigraphic scheme but also more objective and higher-resolution. The high TOC content and brittle minerals within chemozone C1 makes it the most preferable layer for shale gas exploration and development. This research gives a systematic chemostratigraphic analysis on Wufeng and Lower Longmaxi shales, which testifies the feasibility and potential of usage of chemostratigraphy for Chinese shale gas exploration and development.
已有勘探资料表明,西藏尼玛盆地古新统—始新统牛堡组地层具有良好的油气资源显示,然而目前有关于该套地层的地层格架划分仍然薄弱.化学地层学方法在北美页岩气勘探开发中取得了巨大成功,鉴于此,本文以尼玛盆地东部的协德乡南牛堡组剖面作为研究对象,通过对露头样品的主微量元素测试结果进行沉积地球化学、主成分分析、完备总体经验模态分解、以及自相关函数分析,从化学自—异旋回角度以及元素耦合行为出发,探讨地球化学基准面对化学地层格架的控制作用,从而为牛堡组地层提供化学地层划分方案.主成分分析结果表明,牛堡组地层沉积主要受控于细粒碎屑输入、碳酸盐岩、粗粒碎屑输入、氧化还原—生产力、以及盐度这五个因素;经验模态分解和自相关函数分析结果表明,牛堡组地层受到了明显的异旋回驱动,显示出多尺度基准面震荡特点.通过对异旋回信号分量(本征模函数,IMFs)进行重构,并且结合元素相互耦合特性,建立了牛堡组化学地层格架,该结果与岩石地层单元以及沉积相单元一致,证明了本文所提出的化学地层划分方案的可靠性和实用性.