Multi-element geochemical data derived from systematic surveys from a variety of media within a geospatial continuum contain information that provides insight into processes that reflect mineralogy, paragenesis, alteration and economic mineralization. The relationships of the elements and corresponding attributes using a range of data analytical methods, including univariate and multivariate statistical methods along with machine learning, provide the framework for building knowledge from which extended methods of artificial intelligence can enhance model building and mineral systems discovery. This contribution provides the elements for the design of workflows that meet the requirements of enhancing knowledge from geochemical and mineralogical surveys for the purposes of geological mapping, mineral resource prediction or environmental management. Further, it outlines a framework for a systematic evaluation of geochemical data plus attributes that enable the discovery of processes from which models can be constructed and tested using machine learning methods. Methods are described for ‘process discovery’ followed by additional methods for ‘process validation/prediction’. Six case studies are presented that highlight different approaches to the discovery of processes that assist in mineral exploration and geological mapping at various scales. The workflow includes caveats and flags potential issues or limitations on what can be discovered and validated.
Geochemical data are compositional in nature and are subject to the problems typically associated with data that are restricted to the real non-negative number space with constant-sum constraint, that is, the simplex. Geochemistry can be considered a proxy for mineralogy, comprised of atomically ordered structures that define the placement and abundance of elements in the mineral lattice structure. Based on the innovative contributions of John Aitchison, who introduced the logratio transformation into compositional data analysis, this contribution provides a systematic workflow for assessing geochemical data in a simple and efficient way, such that significant geochemical (mineralogical) processes can be recognized and validated. This workflow, called GeoCoDA and presented here in the form of a tutorial, enables the recognition of processes from which models can be constructed based on the associations of elements that reflect mineralogy. Both the original compositional values and their transformation to logratios are considered. These models can reflect rock-forming processes, metamorphism, alteration and ore mineralization. Moreover, machine learning methods, both unsupervised and supervised, applied to an optimized set of subcompositions of the data, provide a systematic, accurate, efficient and defensible approach to geochemical data analysis. The workflow is illustrated on lithogeochemical data from exploration of the Star kimberlite, consisting of a series of eruptions with five recognized phases.
Soil spectroscopy has the potential to replace wet chemical methods. However, high prediction accuracy of unknown samples is necessary to confidently map predicted values across large spatial scales. This study developed a systematic approach using MIR spectroscopy to predict soil particle size of 9,009 unknown samples, representing an area of approximately 35,716 km(2) in Ireland. A systematic approach for MIR spectroscopy combined with chemometrics was augmented by including spectral control charts to identify unrepresentative spectra in predicting of unknown samples. Moreover, similar to 2% of the predicted values (external validation) were selected, covering a wide range of clay, sand and silt, and analysing them using the classical reference method. This step assesses the accuracy of predicted values and provided confidence when projecting predicted values onto soil particle size maps at multiple regional scales. The MIR models were based on the support vector regression algorithm, and were built using a calibration dataset of similar to 1,000 mineral and organo-mineral (MOM) soils, i.e., non-peat soils, analysed using the pipette method. The model was used to predict soil particle size from 9009 unknown samples described as MOM soils. Spectral control charts correctly identified 97.59% of the peat soils (3078 out of 3154 samples) as 'out of control' samples, which the spectral model should not predict, and about 90% of the MOM soils (5254 out of 5855 samples) were classified as 'under control' samples, which the spectra models can predict. The accuracy calculated for the similar to 2% of samples selected from the MOM unknown samples (n = 5254) was similar to the accuracy in the internal validation set (n = 280), with R-2 val values of 0.90, 0.84 and 0.71, RMSEP values of 2.21, 4.31 and 3.91%; RPIQval values of 4.97, 3.71 and 2.30, for clay, sand and silt prediction, respectively. This systematic approach can predict soil attributes using large spectral libraries, providing confidence for building regional and national scale soil maps.
A mineral prospectivity map (MPM) focusing on gold mineralization in the Larder Lake region of Northern Ontario, Canada, has been produced in this study. We have used the Random Forest (RF) algorithm to use 32 predictor maps integrating geophysical, geochemical, and geological datasets from various sources that represent vectors to gold mineralization. It is evident from the efficiency of classification curves that MPMs generated are robust. The unsupervised algorithms, K -means and principal component analysis (PCA) were used to investigate and visualize the clustering nature of large geochemical and geophysical datasets. We used RQ-mode PCA to compute variable and object loadings simultaneously, which allows the displays of observations and the variables at the same scale. PCA biplots of the Larder Lake geochemical data show that Au is strongly correlated with W, S, Pb and K, but inversely correlated with Fe, Mn, Co, Mg, Ca, and Ni. The known gold mineralization locations were well classified by RF with the accuracy of 95.63 %. Furthermore, partial least squares -discriminant analysis (PLS-DA) model combines 3D geophysical clusters and geochemical compositions, which indicates the Au -rich areas are characterized with low to mid resistivity - low susceptibility properties. We conclude that the Larder Lake -Cadillac deformation zone (LLCDZ) is relatively more fertile than the Lincoln-Nipissing shear zone (LNSZ) with respect to gold mineralization due to deeper penetrating faults. The intersection of the LLCDZ and network of high -angle NE -trending cross faults acts as key conduits for gold endowments in the Larder Lake area. This study innovatively combined multivariate geological, geochemical, and geophysical datasets via machine learning algorithms, which improves identification of geochemical anomalies and interpretation of spatial features associated with gold mineralization.
The development of John Aitchison's approach to compositional data analysis is followed since his paper read to the Royal Statistical Society in 1982. Aitchison's logratio approach, which was proposed to solve the problematic aspects of working with data with a fixed sum constraint, is summarized and reappraised. It is maintained that the properties on which this approach was originally built, the main one being subcompositional coherence, are not required to be satisfied exactly -- quasi-coherence is sufficient, that is near enough to being coherent for all practical purposes. This opens up the field to using simpler data transformations, such as power transformations, that permit zero values in the data. The additional property of exact isometry, which was subsequently introduced and not in Aitchison's original conception, imposed the use of isometric logratio transformations, but these are complicated and problematic to interpret, involving ratios of geometric means. If this property is regarded as important in certain analytical contexts, for example unsupervised learning, it can be relaxed by showing that regular pairwise logratios, as well as the alternative quasi-coherent transformations, can also be quasi-isometric, meaning they are close enough to exact isometry for all practical purposes. It is concluded that the isometric and related logratio transformations such as pivot logratios are not a prerequisite for good practice, although many authors insist on their obligatory use. This conclusion is fully supported here by case studies in geochemistry and in genomics, where the good performance is demonstrated of pairwise logratios, as originally proposed by Aitchison, or Box-Cox power transforms of the original compositions where no zero replacements are necessary.
A novel method of estimating the silica (SiO 2 ) and loss-on-ignition (LOI) concentrations for the North American Soil Geochemical Landscapes (NASGL) project datasets is proposed. Combining the precision of the geochemical determinations with the completeness of the mineralogical NASGL data, we suggest a ‘reverse normative’ or inversion approach to first calculate the minimum SiO 2 , water (H 2 O) and carbon dioxide (CO 2 ) concentrations in weight percent (wt%) in these samples. These can be used in a first step to compute minimum and maximum estimates for SiO 2 . In a recursive step, a ‘consensus’ SiO 2 is then established as the average between the two aforementioned SiO 2 estimates, trimmed as necessary to yield a total composition (major oxides converted from reported Al, Ca, Fe, K, Mg, Mn, Na, P, S and Ti elemental concentrations + ‘consensus’ SiO 2 + reported trace element concentrations converted to wt% + ‘normative’ H 2 O + ‘normative’ CO 2 ) of no more than 100 wt%. Any remaining compositional gap between 100 wt% and this sum is considered ‘other’ LOI and likely includes H 2 O and CO 2 from the reported ‘amorphous’ phase (of unknown geochemical or mineralogical composition) as well as other volatile components present in soil. We validate the technique against a separate dataset from Australia where geochemical (including all major oxides) and mineralogical data exist on the same samples. The correlation between predicted and observed SiO 2 is linear, strong ( R 2 = 0.91) and homoscedastic. We also compare the estimated NASGL SiO 2 concentrations with a sparser, publicly available continental-scale survey over the conterminous USA, the ‘Shacklette and Boerngen’ dataset. This comparison shows the new data to be a reasonable representation of SiO 2 values measured on the ground over the conterminous USA. We recommend the approach of combining geochemical and mineralogical information to estimate missing SiO 2 and LOI by the recursive inversion approach in datasets elsewhere, with the caveat to always validate results.
A systematic approach to the evaluation of geochemical data involves the use of multivariate methods that identify processes. These processes are represented by element associations that reflect mineralogy. Processes may be linear or nonlinear, depending on the type of process. Different metrics can reflect different processes. Metrics with coordinates derived from principal component analysis, independent component analysis, and t-distributed stochastic embedding, to name a few, reflect different processes. The dominant components of these metrics can be used to enhance the signal/noise ratio in the data. An integral part of process discovery is the geospatial coherence of multivariate signatures. Models can be constructed by tagging the dominant components with attributes such as geology or mineral deposit information. These models can be tested using a range of multivariate classification/validation/prediction procedures from which probability-based measures of likelihood can be determined and displayed geospatially. The application of these techniques requires acknowledgment of the limitations inherent in the data.
Purpose Mehlich-3 extractable P, Al, Ca, and Fe combined with pH can be used to help explain soil chemical processes which regulate P retention, such as the role of Al, Ca, Fe, and pH levels in P fixation and buffering capacity. However, Mehlich-3 is not always the standard test used in agriculture. The objective of this study is to assess the most reliable conversion of Mehlich-3 Al, Ca, Fe, and P and pH into a commonly used soil P test, Morgan’s P, and specifically to predict values into decision support for fertiliser recommendations. Methods A geochemical database of 5631 mineral soil samples which covered the northern area of Ireland was used to model soil test P and P indices using Mehlich-3 data. Results A random forest machine learning algorithm produced an R 2 of 0.96 and accurately predicted soil P index from external validation in 90% of samples (with an error range of ± 1 mg L −1 ). The model accuracy was reduced when predicted Morgan’s P concentration was outside of the sampled area. Conclusions It is recommended that random forest is used to produce Mehlich-3 conversions, especially when data covers large spatial scales with large heterogeneity in soil types and regional variations. To implement conversion models into P testing regimes, it is recommended that representative soil types/geochemical attributes are present in the dataset. Furthermore, completion of a national scale geochemical survey is needed. This will enable accurate predictions of Morgan’s P concentration for a wider range of soils and geographical scale.
Soil mapping of phosphorus (P) pools at regional scale can be used to inform policy and land management strategies in agri-environmental systems. However, linking an element that is predominantly managed by fertiliser and organic inputs to regional scale geochemistry can be problematic. This study used a geological survey of the northern half of Ireland at <= 4 km(2) resolution to map total P (ICP aqua regia), plant available P (Morgan' P) and Labile P pools (ICP aqua regia and Mehlich-3 P). Spatial modelling was used to map P interactions with total and available aluminium (Al) (ICP aqua regia and Mehlich-3 Al), as predefined sorption indicators. With the aim to develop agri-environmental mapping approaches to link soil P sorption dynamics to geological controls that can assist with regional, catchment and national-scale modelling for policy development.Plant available P concentration showed no spatial continuity in Kriged variogram output and was regulated by fertiliser inputs which mask any underlying geological processes. Subsequently, making this exploration method redundant in highly managed landscapes. Available Al concentrations determined by Mehlich-3 Al (M3-Al) are however regulated by underlying geology. Moreover, they have a significant negative correlation with bioavailable P indices. Areas of high M3 Al concentrations (>= 702.5 mg kg(-1)) are dominant in the study area and show mainly low plant available P values (<5 mg l(-1)), however, potential labile P pools were indicated. Whereas areas of low Al concentrations (<= 697 mg kg(-1)) show moderate to high plant available P concentrations. Which could identify areas where P is likely to be more soluble and at greater risk to loss to water. Therefore, using available Al as a sorption indicator can identify regional areas of high P fixation and low soil P retention in soils. Furthermore, using sorption dynamics in geospatial modelling could benefit targeted management approaches.
Transported soils cause difficulties in the identification of geochemical anomalies. It has been demonstrated that the joint application of local singularity analysis (LSA) and principal component analysis (PCA) can identify geochemical anomalies effectively, especially in regolith-covered areas. However, more convincing evidence is needed to explain the reasons for this. In this study, a soil profile overlying several mineralized veins cutting through bedrock was analysed in situ using a portable X-ray fluorescence spectrometer. The patterns of two mineralization-related elements, Cu and Mo, were analysed. The results revealed that the element concentrations of the soil sharply decreased as the distance from the bedrock increased, and this relationship can be described by a power-law model. LSA enhanced several vein-like anomalies corresponding to the mineralization veins in the bedrock, and the presence of vertically elongated weak anomalies in the soil indicates the migration of ore elements originating from the underlying bedrock through the soil. The statistics show that the patterns of the local singularity index (LSI) are stable at different depths and in different media, whereas the concentration patterns are not. In addition, the mineralization-related elements have a higher correlation coefficient for the LSI than for the concentration. Since a previous simulation study determined that a mineralization-indicative first principal component prefers that the variables have a close relationship and that the variables have similar patterns in different geological objects, the patterns discovered in this study explain why LSA is effective in identifying geochemical anomalies, especially when combined with PCA. Supplementary material: the high-resolution photo, the element concentration data and the lithologic data of the profile are available at https://doi.org/10.6084/m9.figshare.c.5957122 Thematic collection: This article is part of the Applications of Innovations in Geochemical Data Analysis collection available at: https://www.lyellcollection.org/cc/applications-of-innovations-in-geochemical-data-analysis
Regional stream water geochemistry acquired as part of the Tellus programme in Ireland has been analysed to assess its potential for application to environmental assessment and mineral exploration. Interpolated geochemical maps and multivariate statistical analysis, including principal component analysis and random forest classification, demonstrate broad geogenic control of stream water chemistry, with both bedrock and subsoil contributing to the patterns observed. Surface water regulations set Environmental Quality Standard values for individual Priority Substances and Specific Pollutants that may depend on background concentrations and/or water hardness. The high resolution of Tellus stream water data and their location on low-order streams have allowed estimation of background concentrations and water hardness in the survey area, with significant implications for water monitoring programmes. Anthropogenic inputs to stream water in the survey area come mainly from agricultural sources and Tellus data suggest few catchments are unaffected. Comparison of Tellus stream water geochemistry with stream sediment and topsoil geochemistry suggest that stream water geochemistry has strong potential for use in mineral exploration, with the same base metal and gold pathfinder anomalies apparent in all three data sets. Cluster analysis indicates that base metals in stream water are associated with organic matter but statistical analysis may be employed to distinguish mineralization-related signatures. Supplementary material: Comparison of cation/anion associations using Piper plots and principal component analysis is available at https://doi.org/10.6084/m9.figshare.c.5683094 Thematic collection: This article is part of the Hydrochemistry related to exploration and environmental issues collection available at: https://www.lyellcollection.org/cc/hydrochemistry-related-to-exploration-and-environmental-issues
The Editorial of this special issue focuses on Innovative methods applied to processing and interpreting geochemical data, summarizing the papers published therein, reviewing the advances made in analyzing and interpreting geochemical datasets, and highlighting some challenges facing present-day geochemical data analysis. Decades of research on mineral exploration and environmental pollution have led to three significant breakthroughs in analyzing geochemical datasets, namely the compositional nature of geochemical data, fractal/multifractal-based decomposition of anomalous patterns, and artificial intelligence (AI)-based geochemical pattern recognition. Some challenges still face the use of AI-based methods for the recognition of geochemical anomalies linked to ore-forming processes. The papers published in this issue employ state-of-the-art techniques to fully exploit the potential of geochemical datasets for exploration and environmental geochemistry. A total of 25 papers were submitted to this special issue of which 15 papers were published. The published papers can be categorized in three broad groups, namely (i) delineation of anomalous geochemical patterns, (ii) integrating geochemical datasets to other geospatial tools for mineral exploration, and (iii) geochemistry for hydrocarbon exploration.
Precision and sustainable agriculture requires information about soil pH and organic matter (OM) content at higher spatial and temporal scales than current agronomic sampling and analytical methods allow. This study examined the accuracy of spectral models using high throughput screening (HTS) in diffuse reflectance mode in mid Infra-red (MIR)/DRIFT combined with machine learning algorithms to predict soil pH(CaCl2) and %OM in shallow and deeper topsoils compared to laboratory methods. Models were developed from an archive of samples taken on a 4 km(2) grid from the northern half of Ireland (Terra Soil project), which includes 18,859 samples (9,396 shallow + 9,463 deeper). The application of Cubist models showed that for different depths there are minor different spectral group associations with pH and %OM values. These differences resulted in a loss of accuracy in the extrapolation of the topsoil model to predict values from deeper topsoils or vice versa. Therefore we recommend the use of samples from both depths to build a calibration model. The proposed methodology was able to determine %OM and pH using a unique multivariate regression model for both depths, with RMSEP values of 1.12 and 0.89 %; RPIQ values of 42.34 and 38.48; R-val (2)of 0.9989 and 0.9993 for %OM determinations in shallow and deeper topsoils, respectively. For pH determinations the RMSEP values obtained were 0.25 and 0.34; RPIQ values of 6.04 and 4.94; R-val (2) 0.9385 and 0.8954. Both regression models are classified as excellent predictions models, yielding RPIQ values >4.05 for shallow and deeper top-soils. The results demonstrated the high potential of HTS-DRIFT combined with machine learning algorithms as a rapid, accurate, and cost-effective method to build large soil spectral libraries, displaying predicted results similar to two separate soil laboratory methods (pH and LOI).
In this study we apply multivariate statistical and predictive classification methods to interpret geochemical data from 8545 stream-sediment samples collected in southern British Columbia, Canada. Data for 35 elements were corrected for laboratory bias and adjusted for values reported below the lower limit of detection. Each sample site was attributed with the closest British Columbia MINFILE occurrence within 2.5 km. MINFILE occurrences were grouped into ‘GroupModels’ based on similarities between the British Columbia Geological Survey mineral deposit models and geochemical signatures. These data were used to create a training dataset of 474 observations, including 100 samples not attributed with a MINFILE occurrence. The training set was used to generate predictions for the mineral deposit models from which posterior probabilities were estimated for the remaining 8071 samples. The data underwent a centred log-ratio transformation and then characterization using either principal component analysis (PCA) or t-distributed stochastic neighbour embedding using 9 dimensions (t-SNE) prior to classification by random forests. The posterior probabilities generated from the t-SNE metric provide a slightly higher level of prediction accuracy compared to the posterior probabilities obtained using the PCA metric. The results are comparable to those obtained using a conventional catchment analysis approach and expert-driven model. The approach presented here provides a repeatable, consistent and defensible methodology for the identification of prospective mineralized terrains and mineral systems.
The isometric logratio transformation, in the form of what has been called a "balance", has been promoted as a way to contrast two groups of parts in a compositional data set by forming ratios of their respective geometric means. This transformation has attractive theoretical properties and hence provides a useful reference, but geometric means are highly affected by parts with small relative values. When a comparison between two groups of parts is required in practical applications, such as the investigation and construction of models, while making use of substantive domain knowledge, it is demonstrated that the logratio of two amalgamations serves as an alternative, interpretable form of balance. A geochemical data set is considered, which has been analyzed previously by transforming to a set of isometric logratio balances. An alternative approach, using a reduced set of pairwise logratios of parts, optionally involving prescribed amalgamations, is very close to optimal in accounting for the variance in this compositional data set. These simpler transformations also have an exact back transformation to the original parts. This approach highlights for this dataset which compositional parts are driving the data structure, using variables that are easy to interpret and that map well to research-driven objectives.