Reliable and harmonised soil information remains critically limited across Africa, constraining soil monitoring, climate-resilient agriculture, and evidence-based land management. Existing soil resources are often fragmented, spatially uneven, outdated, or derived from legacy observations, limiting their usefulness for contemporary continental-scale assessment. The Soils4Africa project implemented a coordinated field campaign across 33 African countries between 2022 and 2025 to establish a harmonised soil monitoring framework for agricultural lands. Using a hierarchical probabilistic sampling design, 24,951 soil samples were collected from 14,311 locations, supported by standardised field protocols, digital data capture, QR-based sample traceability, and centralised quality control. This paper presents the conceptual, operational, and data-management framework underpinning the survey and reports baseline field observations on farming systems, land management, vegetation structure, and soil physical constraints. The framework achieved more than 70% of planned sampling coverage despite major logistical, environmental, and security-related constraints. Baseline observations show that African agricultural landscapes remain dominated by smallholder systems, low external input use, limited soil and water conservation, and widespread dependence on rainfed production. Field indicators also reveal sparse woody vegetation cover and common physical constraints, including compaction, coarse fragments, shallow effective rooting depth, and subsoil barriers. Unlike earlier continental resources based largely on legacy profiles or site-based surveillance, Soils4Africa provides a contemporary, harmonised, spatially structured field-survey framework designed to support future laboratory-based soil assessment, digital soil mapping, land suitability analysis, and long-term soil monitoring. The study therefore provides a scalable model for coordinated soil monitoring across diverse African agroecosystems and establishes an operational baseline for subsequent analytical studies.
Digital Soil mapping (DSM) provides standardised information layers. The recent availability of global and continental remote sensing-derived products coupled with the ease-of-access to computational resources has made the production of such layers easier. It is ever to characterise and evaluate such DSM-derived products, in particular the type of actual information they can provide to users. DSM studies commonly assess prediction uncertainty using various approaches, including multiple simulations or quantile random forests. These studies provide measures of accuracy derived from statistical (cross-)validation, often based on non-probability and non-representative observations. However, these accuracy metrics and uncertainty assessments do not encompass all the potential elements that could be used to characterise a DSM product, and they do not directly address the needs of the users. We assessed maps based on area of applicability (i.e., the area in covariate space where the model learns about relationships based on the training data), the landscape heterogeneity both in the landscape itself and in covariate space, and the local influence of the covariates on the final products. We present examples of continental and global mapping products, highlighting main accuracy, uncertainty and interpretability aspects and how these influence their suitability for intended use by stakeholders, decision makers and users in general at the given resolution. The results permit some practical reflections on how to integrate all the above elements to identify regions where the confidence in the predictions is highest and the associated uncertainty lowest, but also, where the product is not considered fit for the intended use.
Digital Soil mapping (DSM) at continental and global scale provides standardised global information layers. It is also an important tool to create soil information layers for areas for which local soil survey information is lacking. The recent availability of global and continental remote sensing derived products coupled with the ease-of-access to computational resources has made the production of such layers easier across the globe. Therefore, it is ever more important to assess the quality of DSM-derived products, in particular the type of information they can actually provide to users (i.e., fitness for intended use). DSM studies commonly assess prediction uncertainty using various approaches, including multiple simulations or quantile random forests. However, this does not encompass all the potential elements that could be used to characterise the uncertainty of a DSM product. In this study we are going to assess maps based also on area of applicability (i.e., the area in covariate space where the model learns about relationships based on the training data) and the landscape heterogeneity both in the landscape itself and in covariate space. We present examples of continental and global mapping products, highlighting main uncertainty-related issues and how these influence suitability for intended use by stakeholders, decision makers and users in general at the given resolution. The examples come from a range of projects with different aims and goals. The results permit some practical reflections on how to integrate all the above elements to identify regions where the confidence in the predictions is highest and the associated uncertainty lowest. We will integrate the practical reflections with information collected from a user survey on requirements and usability of continental and global DSM products.
Many Digital Soil Modelling (DSM) products have been generated for diverse regions, countries and continents. While most of these products provide some accuracy metrics, a few also incorporate assessments of uncertainty. The current uncertainty estimates for DSM products often fail to represent elements that are essential for evaluating the suitability of a map for a specific application. For instance, different models can have very similar accuracy metrics but produce different soil-landscape patterns. It is important to be able to evaluate the accuracy of the patterns as well, current accuracy metrics do not do this. Additional metrics could be defined that are able to do this such the ‘area of applicability’, i.e. the area in covariate space where the model learns about relationships based on the training data) and the landscape heterogeneity both in the landscape itself and in covariate space. This study delves into the integration of the aforementioned elements into an assessment of DSM uncertainty at the continental scale. Europe was used as test area, incorporating input observations from EU-LUCAS datasets. A covariate space encompassing the soil forming factors as defined by the SCORPAN model served as a basis for fitting the necessary models for soil products.. We characterized the spatial heterogeneity of both the landscape and the covariate space by employing commonly employed landscape metrics. The findings offer practical insights on how to integrate these components to produce more reliable products for stakeholders.
The commonly-used scorpan approach to Digital Soil Mapping is purely correlative, between observations at points and the values of covariates at those points. An-often-used approach is of maximum complexity: different quality observations are thrown together with the largest possible number of covariates, along with the most complex model (e.g., ensembles of many models). The resulting products are almost always evaluated with point-wise metrics with sometimes not large differences or improvement between models. Spatial patterns are rarely compared with soil geography. Further work should be done to include more the local and spatial structure component, the ‘n = neighbourhood’ of scorpan. This work addresses how pedological knowledge could be included and the potential advantages and disadvantages of doing so.
Digital Soil Mapping (DSM) is an established methodology to create maps of soil properties at different resolutions and extents. It establishes a statistical relationship between the measured values at point observations and environmental covariates selected to describe the soil forming factors and to explain the spatial variability of the soil properties. These relationships are then used to map the target soil properties across the area of interest. In this example, we used 1423 measurements on soil organic carbon and pH for the 0-20 cm soil layer from a Mongolian soil survey. This survey was organised within the framework of the “National Program to Combat Desertification” to determine the primary soil quality indicators for desertification assessment in Mongolia and it was conducted based on the state network of the Meteorological and Environmental Research Agency starting in 2012. The samples are collected from 1500 monitoring points every 5 years. We used data from the monitoring round between 2012 and 2015 We used two sets of covariates for modelling predictive relationships. The first is the set used in SoilGrids at 250 m resolution with over 400 layers available of which about 180 were used for the modelling, after de-correlation. The second is a reduced set of about 40 covariates at 100 m resolution derived mainly from Sentinel (1 and 2) images, ERA5 for climate data and ALOS for morphological information. In this study we will compare the results of the two models, with both point-wise evaluation matrices and assessment of spatial patterns. In the evaluation the expertise of local partners will also be used.
A crucial decision in designing a spatial sample for soil survey is the number of sampling locations required to answer, with sufficient accuracy and precision, the questions posed by decision makers at different levels of geographic aggregation. In the Indian Soil Health Card (SHC) scheme, many thousands of locations are sampled per district. In this paper the SHC data are used to estimate the mean of a soil property within a defined study area, e.g., a district, or the areal fraction of the study area where some condition is satisfied, e.g., exceedence of a critical level. The central question is whether this large sample size is needed for this aim. The sample size required for a given maximum length of a confidence interval can be computed with formulas from classical sampling theory, using a prior estimate of the variance of the property of interest within the study area. Similarly, for the areal fraction a prior estimate of this fraction is required. In practice we are uncertain about these prior estimates, and our uncertainty is not accounted for in classical sample size determination (SSD). This deficiency can be overcome with a Bayesian approach, in which the prior estimate of the variance or areal fraction is replaced by a prior distribution. Once new data from the sample are available, this prior distribution is updated to a posterior distribution using Bayes' rule. The apparent problem with a Bayesian approach prior to a sampling campaign is that the data are not yet available. This dilemma can be solved by computing, for a given sample size, the predictive distribution of the data, given a prior distribution on the population and design parameter. Thus we do not have a single vector with data values, but a finite or infinite set of possible data vectors. As a consequence, we have as many posterior distribution functions as we have data vectors. This leads to a probability distribution of lengths or coverages of Bayesian credible intervals, from which various criteria for SSD can be derived. Besides the fully Bayesian approach, a mixed Bayesian-likelihood approach for SSD is available. This is of interest when, after the data have been collected, we prefer to estimate the mean from these data only, using the frequentist approach, ignoring the prior distribution. The fully Bayesian and mixed Bayesian-likelihood approach are illustrated for estimating the mean of log-transformed Zn and the areal fraction with Zn-deficiency, defined as Zn concentration <0.9 mg kg -1, in the thirteen districts of Andhra Pradesh state. The SHC data from 2015-2017 are used to derive prior distributions. For all districts the Bayesian and mixed Bayesian-likelihood sample sizes are much smaller than the current sample sizes. The hyperparameters of the prior distributions have a strong effect on the sample sizes. We discuss methods to deal with this. Even at the mandal (sub-district) level the sample size can almost always be reduced substantially. Clearly SHC over-sampled, and here we show how to reduce the effort while still providing information required for decision-making. R scripts for SSD are provided as supplementary material.
Using fertilisers is indispensable for closing yield gaps in Sub Saharan Africa. Current fertiliser recommendations, however, are often blanket recommendations which do not take spatial variation in soil conditions within a region or country into account. Soil maps can potentially support fertiliser recommendations at a higher spatial resolution. The QUantitative Evaluation of the Fertility of Tropical Soils (QUEFTS) model is a decision support tool that predicts crop yields as an indicator of soil fertility and can be used to evaluate yield responses to fertilisers. It was designed for field level output and runs on field-specific soil information. The aim of this study was to compare two methods for developing maps of QUEFTS output, i.e. maize yield and the yield-limiting nutrient, with Rwanda as a case study. We used a database containing soil analysis results of 999 samples collected across Rwanda. Transfer functions were applied to predict the required P-Olsen and Exchangeable K input for QUEFTS based on the soil data. For the "Calculate-then-Interpolate " (CI) method, transfer functions and QUEFTS were applied to point data, and the final output was then interpolated using random forest modelling. For the "Interpolate-then-Calculate " (IC) method, maps of the soil parameters were developed first, before applying calculations. Implications of the chosen method (i.e. CI or IC) on QUEFTS predictions on a national scale were evaluated using set-aside locations. Results showed low precision and accuracy of QUEFTS maize yield predictions across Rwanda. The CI method performed better in predicting QUEFTS yield and yield-limiting nutrient than the IC method. Correlations between mapped yield predictions and predictions on set-aside evaluation locations were similar for the CI (r = 0.444) and IC (r = 0.439) methods. The poorer performance of the IC method was mostly due to overestimation of yields, which was most likely caused by the effect of smoothing on the soil maps used as input for QUEFTS. We conclude that the CI method is the preferred method for spatial application of QUEFTS.
Multi-element soil extractions such as Mehlich 3 (M3) have gained popularity in recent years, but comparing outcomes to other soil testing methods is not always straightforward. In this study, extraction mechanisms of M3, Olsen and neutral 1 M ammonium acetate (AA) soil tests were explored and transfer functions were derived between P-Olsen and P-M3 as well as between K-AA and K-M3. Soils from tropical and temperate areas were used to derive these P and K transfer functions and were evaluated separately. The application of these transfer functions for tropical soils was evaluated by using them as input for the Quantitative Evaluation of the Fertility of Tropical Soils (QUEFTS). AA and M3 generally extracted similar amounts of K, but relations between K-AA and K-M3 were different for tropical and temperate soils. For tropical soils, the transfer function did not require additional parameters besides K-M3 to predict K-AA, but for temperate soils inclusion of clay content and pH was needed. This difference between tropical and temperate soils was explained by clay mineralogy. The relation between P-Olsen and P-M3 in tropical soils was found to be dependent on pH, Al-M3, Fe-M3 and Ca-M3. P-Olsen and K-AA values, calculated with their respective transfer functions, were used as input for QUEFTS. The yields predicted with measured P-Olsen and Exch. K were used as benchmark. For 63 out of 81 soil samples, predicted maize yields with transfer functions deviated less than 10% from the benchmark. The largest deviations from the benchmark were found for low P-Olsen and K-AA values, which corresponds to QUEFTS maize yield predictions up to 3000 kg ha(-1). We conclude that a M3 extraction results and soil pH can reliably be transferred to, and thus replace P-Olsen and K-AA determinations with the functions developed for tropical soils. The transfer functions can be used to generate input for the QUEFTS model with minor effects on yield predictions, thus expanding its applicability in cases where only M3 extraction results are available.
Relevant soil information at different scales would greatly help addressing many of the Sustainable Development Goals. Digital Soil Mapping is an established methodology to create maps of soil properties at different resolutions and extents. Many projects across the globe have provided information on primary soil properties, such as soil textural fractions, soil organic carbon content, cation exchange capacity and soil pH. For environmental modelling and assessment, maps of complex soil properties are also important. These can be defined as properties that cannot be measured directly in the laboratory but are derived from primary soil properties, for instance by simple calculations, pedotransfer functions or more advanced spatial analyses. Examples are available water capacity, soil carbon density and stocks, as well as soil erodibility. There are two main approaches to map complex properties: 1) “model first, interpolate later”, where the complex property is first calculated at point locations where the primary properties are known and then mapped; and 2) “interpolate first, model later”, where the complex property is calculated from maps of the primary properties contributing to it. We present and discuss these two approaches for global applications using legacy data with a non-uniformspatial distribution of observations and the SoilGrids workflow. We compare the results for available water capacity of the 0 to 100 cm depth interval and soil carbon densities for six depth layers. Both properties were derived from a combination of simple calculations for point locations where the input soil properties were available and pedotransfer functions for other point locations where basic soil properties were available. There were substantial differences between the “model first, interpolate later” and “interpolate first, model later” approaches, both in point-wise evaluation metrics and in landscape patterns.
Soils provide a variety of goods and services and are a key natural resource, non-renewable on the time scale of a human life-span, to realise several UN-Sustainable-Development–Goals, including zero hunger (SDG2), climate action (SDG13) and life on land (SDG15). Consistent soil information is required to underpin a large range of global assessments, such as soil and land degradation, sustainable land management, and environmental conservation. In this study, we present an application addressing the modelling of soil functions at global scale, using erosivity risk and soil carbon sequestration potential as examples. We used SoilGridsv2.0, as set of soil property maps at 250m resolution, derived from a large set of standardized soil profile observations (WoSIS), an updated set of environmental covariates, and improved machine learning models. It provides global assessments of prediction uncertainty, quantified with a 90% prediction interval, and considers an evaluation procedure that provides more realistic metrics of map accuracy. We used simplified models to derive soil functions (e.g. erosivity or carbon sequestration potential) from basic soil properties, meaningful for different pedo-climatic regions. We provide an indication of areas of low/high risk of soil degradation to support sustainable soil management planning. The uncertainty limits of the input soil layers were used to provide a preliminary assessment of the uncertainty of the derived layers. The present modelling framework, using soil properties maps to derive soil functions, offers great flexibility and may be applied to a diverse set of models to generate soil information products tailored to specific applications. We highlight some of the challenges of assessing soil functions at global scale.
Rice is a staple food and cash crop for smallholder farmers in sub-Saharan Africa; however, yields are very low, with indications that both macro and micro-nutrients may limit rice productivity in East Africa next to the need for good agronomic practices. Diagnostic on-farm experiments were conducted in Uganda and Tanzania between 2015 and 2017 to assess the contribution of macro, secondary and micro-nutrients on lowland rice yield and identify options by which smallholder farmers can increase productivity. All treatments included good agronomic practices combined with: zero fertilisation as a control, NPK fertilisation with and without secondary and micro-nutrients (B, Mn, Zn, Cu, Mg, S), and/or treatments where B, Mn and Zn were omitted one at a time from the NPK + secondary and micro-nutrient treatment. NPK fertilisation significantly (p < 0.05) increased grain yield under irrigated condition by ca. 32 and 29 % during 2015 and 2016, and 24 and 100 % during 2016 and 2017 in Tanzania and Uganda, respectively; however, inconsistent effects were observed under rainfed condition. Observed higher yields corresponded mainly to higher panicle number with an additive effect of grains per panicle indicating major effects were at earlier growth stages supporting higher sink size development. Adding secondary and micro-nutrients to NPK enhanced yield significantly (p < 0.05) under irrigated condition in Tanzania 2015 and 2016 by 7 and 11 %, respectively, while varying results were obtained under rainfed condition. In Uganda, no significant (p> 0.22) effects of secondary and micro-nutrients were observed in both years and growing conditions. This study indicates that the first step to improving lowland rice productivity is proper water management, under otherwise also good crop management in terms of timely transplanting and weeding, and further yield gains can be realised with NPK fertilisation. Secondary and micro-nutrients were effective only when NPK were applied and on the fluvisols of Tanzania, and were not co-limiting yield on the plintosols of Uganda.
SoilGrids produces maps of soil properties for the entire globe at medium spatial resolution (250 m cell size) using state-of-the-art machine learning methods to generate the necessary models. It takes as inputs soil observations from about 240 000 locations worldwide and over 400 global environmental covariates describing vegetation, terrain morphology, climate, geology and hydrology. The aim of this work was the production of global maps of soil properties, with cross-validation, hyper-parameter selection and quantification of spatially explicit uncertainty, as implemented in the SoilGrids version 2.0 product incorporating state-of-the-art practices and adapting them for global digital soil mapping with legacy data. The paper presents the evaluation of the global predictions produced for soil organic carbon content, total nitrogen, coarse fragments, pH (water), cation exchange capacity, bulk density and texture fractions at six standard depths (up to 200 cm). The quantitative evaluation showed metrics in line with previous global, continental and large-region studies. The qualitative evaluation showed that coarse-scale patterns are well reproduced. The spatial uncertainty at global scale highlighted the need for more soil observations, especially in high-latitude regions.
Abstract. SoilGrids produces maps of soil properties for the entire globe at medium spatial resolution (250 metres cell size) using state-of-the-art machine learning methods to generate the necessary models. It takes as inputs soil observations from about 240 000 locations worldwide and over 400 global environmental covariates describing vegetation, terrain morphology, climate, geology and hydrology. The aim of this work was the production of quality-assessed global maps of soil properties, with cross-validation, hyper-parameters selection and quantification of spatially explicit uncertainty, as implemented in the SoilGrids version 2.0 product incorporating state of the art practices and adapting them for global digital soil mapping with legacy data. The paper presents the evaluation of the global predictions produced for soil organic carbon content, total nitrogen, coarse fragments, pH(water), cation exchange capacity, bulk density and texture fractions at six standard depths (up to 200 cm). The quantitative evaluation showed metrics in line with previous global, continental and large regions studies. The qualitative evaluation showed that coarse scale patterns are well reproduced. The spatial uncertainty at global scale highlighted the need for more soil observations, especially in high latitude regions.
Soil information is fundamental for many global applications, such as food security, land degradation, water resources, hydrology, climate change and ecological conservation. To address these diverse needs, it is important to provide free, consistent, easily accessible and standardized soil information. SoilGrids meets these requirements being a global product supporting global modelling and providing complementary information for the development of regional and national products in data-poor areas. This presentation will focus on the methodological aspects for modelling and mapping of global soil information. We describe the selection of models for global mapping using quantile random forest and recursive feature elimination to obtain a parsimonious model. We also use a refined cross-validation procedure to account for bias caused by spatial differences in sampling density at different depths. SoilGrids also quantifies location-specific uncertainty at global level by computing 90% prediction interval limits.
SoilGrids maps soil properties for the entire globe at medium spatial resolution (250 m cell side) using state-of-the-art machine learning methods. The expanding pool of input data and the increasing computational demands of predictive models required a prediction framework that could deal with large data. This article describes the mechanisms set in place for a geo-spatially parallelised prediction system for soil properties. The features provided by GRASS GIS – mapset and region – are used to limit predictions to a specific geographic area, enabling parallelisation. The Slurm job scheduler is used to deploy predictions in a high-performance computing cluster. The framework presented can be seamlessly applied to most other geo-spatial process requiring parallelisation. This framework can also be employed with a different job scheduler, GRASS GIS being the main requirement and engine.
Geospatially explicit information of soil-landscape resources of Ethiopia is lacking or fragmented for much of the country. Recently, massive soil data were collected, however these are limited to properties related to soil fertility and valid for the topsoil only. Understanding the country's soil-landscape resources, including their qualities and constraints beyond the topsoil, remains key information for systematic and reliable scaling up of evidence-based agricultural best practices including soil fertility management recommendations: The objective of this study was to produce a coherent dataset of the major soil-landscape resources of 30 highland woredas (districts), contributing to the Agricultural Growth Program of the Government of Ethiopia. The study started with an exploratory survey to identify the major (most common) soils occurring across the landscapes followed by a full survey to assess the distribution of the identified major soils. Representative soil profiles were characterised from soil pits and classified as Reference Soil Groups (RSGs), with prefix qualifiers (PQs), according to the World Reference Base for soil resources (WRB). A large number of soil profiles was classified from auger observations. Observed soil classes at both RSG and RSG + PQ level were combined with spatial explanatory variables (covariates), representing the soil forming factors in the landscapes, and their relationships were modelled and validated by random forest. A multitude of tree models was trained using each profile for calibration in approximately two third and cross-validation in approximately one third of the models. Cross-validation showed that RSGs were predicted with a reasonable overall purity of 0.58 and RSGs + PQ were predicted with a purity of 0.48. The most relevant covariate in the models was the Geomorphology and Soils map of Ethiopia at 1: 1 M scale disaggregated into soil-landscape facets. Next models were used to predict soil classes across woredas which resulted in a 250 m resolution raster map of the most probable major soils. This raster map was generalised into a polygon map of major soil-landscape resources. The purity of this final map was estimated to be 0.54 for RSGs and 0.45 for RSGs + PQ. Soil properties relevant for agricultural interpretation, such as depth, drainage, texture, pH, CEC and organic carbon and nutrient contents, were mapped according to the RSGs depicted on the soil-landscape resources map with a RMSE/mean ratio of on average 42%. We conclude that soil expert knowledge and conventional soil-landscape survey combined with random forest modelling results in an attractive hybrid approach. The approach proves cost-effective and sufficiently accurate and can be used to inform scaling up of evidence-based agricultural best practices.