ABSTRACT Comprehensive soil profile datasets are essential for understanding and modelling soil and soil processes, yet remain scarce in Sub‐Saharan Africa. This is particularly true for full soil profiles combining root observations with soil water retention characteristics. Here, we present a dataset of 50 soil profiles from maize and wheat systems in central Ethiopia, comprising observations and measurements from soil pits described and sampled to depths of down to 2 m. The dataset integrates pedogenic horizons with morphological descriptions, detailed root distribution observations, and analytical data for physical and chemical soil properties obtained from fixed sampling layers. All profiles are classified according to the World Reference Base for Soil Resources. Field photographs are included to document soil morphology and field procedures. The dataset may support improved understanding and modelling of soil processes, soil hydrology, and the relationship between crop roots, soil morphology, and soil properties. It may also support applications in digital soil mapping, and the development of pedotransfer functions. The dataset is publicly accessible in the PANGAEA repository at https://doi.pangaea.de/10.1594/PANGAEA.995660.
The growing demand for high-resolution soil data has stimulated the development of visible, near and mid-infrared (MIR) spectroscopy. This study assessed four modeller’s choices that can influence spectral model performance: (1) model type, (2) application of a calibration transfer function, (3) size of the subset selected from a soil spectral library for model calibration, and (4) pre-processing method. The effect of measurement errors in reference wet and dry chemistry data, and MIR spectra on model performance was evaluated through a Monte Carlo approach. Models were calibrated to predict cation-exchange capacity (CEC), pHH2O, total carbon (C) and total nitrogen (N) in the Netherlands. Accurate quantification of measurement errors is crucial, as large errors substantially reduced model performance. Measurement errors in MIR spectra had a smaller impact than those in wet and dry chemistry data, as overall measurement error was smaller. The effect of modeller’s choices varied per soil property. PLSR outperformed RF for the prediction of CEC and pHH2O, while the opposite was observed for total C and N. Lower RF performance was linked to higher sensitivity to correlations between wavenumbers. Calibration transfer did not consistently improve model performance, indicating case-specific performance. Larger spectral subsets improved predictions for CEC, pHH2O, and total N, but not total C. Pre-processing strongly affected results, with the 1st order derivative yielding the best overall performance. The optimal combination of modeller’s choices varied by soil property and/or spectral library, highlighting the importance of treating them as hyperparameters for optimizing spectral models.
Context: Maize is a key staple crop in Ghana, yet yields remain low (20-40 % of potential). Although fertilizer is promoted to enhance productivity, adoption is limited by highly variable yield responses. Objective: This study analyzed spatial and environmental drivers of fertilizer effect heterogeneity using 2854 yield observations from randomized controlled trials and 10,916 pairwise absolute yield response to fertilizer (AR) estimates. Methods: Causal forest (CF) and boosted random forest (BRF) models to estimated fertilizer effects, with BRF performance evaluated via a 10 x 10 nested cross-validation and grid search. SHapley Additive exPlanations and Accumulated Local Effects analyses identified key drivers of fertilizer effect heterogeneity and quantified the magnitude of their influence on fertilizer yield effect. Results and conclusions: Fertilizer effect varied widely (-4.7-8.9 t ha(-1)), with the Sudan Savannah showing the highest median AR (2.8 t ha(-1)) and the Forest-Savannah Transition the lowest (0.9 t ha(-1)). BRF outperformed CF in predicting fertilizer effects (ME: -0.06-0.05 t ha(-1) vs. -0.19 t ha(-1), RMSE: 1.17-1.23 t ha(-1) vs. 1.3 t ha(- 1), MEC: 0.32-0.38 vs. 0.24 and CCC: 0.46-0.54 vs. 0.34). Key determinants of fertilizer effect heterogeneity included both climatic variables (Palmer Drought Severity Index [PDSI], vapor pressure deficit, rainfall) and soil properties (silt content, exchangeable aluminum). PDSI emerged as the dominant driver of fertilizer effect heterogeneity in the entire data set. However, the relative importance of soil versus climate varied spatially: soil properties were the main drivers of fertilizer effect in the Semi-Deciduous Forest and the Forest-Savannah Transition, whereas climatic variables played a stronger role in northern zones. Fertilizer yield effect increased by 0.4-1.6 t ha(-1) with increasing PDSI, indicating that improved moisture availability enhances fertilizer use efficiency. Overall, optimal moisture conditions (PDSI > -2.0), the use of hybrid seeds, and the application of briquette fertilizer all contributed to higher fertilizer effects, whereas drought conditions substantially reduced them. Furthermore, fertilizer effect decreased by 0.2-1.4 t ha(-1) as silt increased from 9 % to 30 %, and by 0.3-0.6 t ha(-1) as exchangeable aluminum increased from 36 to 221 mg kg(-1). Significance: This study presents the first large-scale, data-driven assessment of fertilizer yield effects heterogeneity in Ghana, integrating causal and predictive machine learning with explainable AI. Findings support tailored fertilizer strategies by agro-ecological zones to reduce farmer risk and promote sustainable intensification.
Context: Fertilizer management in smallholder maize production systems in sub-Saharan Africa faces major challenges due to high environmental variability, uncertain input-output prices, and diverse farmer risk preferences. Conventional fertilizer recommendations often fail to account for these factors, resulting in poor adoption rates and inefficient fertilizer use. Objective: This study aimed to develop a fertilizer recommendation framework that integrates yield uncertainty and farmer risk profiles to improve fertilizer application decision-making. Specifically, it examined how uncertainty and farmer risk preferences affect optimal nutrient application rates. Methods: Using 4496 maize field experimental observations, a Quantile Regression Forest model was trained to predict maize yield responses to nitrogen, phosphorus, and potassium and assessed their profitability in 14 representative sites across three agro-ecological zones in Ghana. A utility-based economic model was then applied to simulate farmer decisions under varying levels of risk aversion. Results: Nitrogen emerged as the key yield-limiting nutrient, with yield responses increasing steeply up to approximately 90 kg N ha-1 . Phosphorus and potassium showed minimal agronomic and economic benefits under prevailing farmer conditions. Profit-maximizing nitrogen rates generally ranged from 60 to 90 kg ha-1 but declined to 10-80 kg N ha-1 under risk-sensitive scenarios, particularly for strongly risk-averse farmers. Importantly, our approach revealed site-level heterogeneity in optimal strategies, even within the same agroecological zone. Conclusions: Incorporating uncertainty and farmer behaviour into fertilizer decision-making produces more realistic and relevant nutrient recommendations. Implications: By accounting for uncertainty and farmer behaviour, this study offers a framework to improve the relevance and efficiency of fertilizer recommendations. The findings provide new insights into adaptive agronomy within the context of sustainable intensification and support a shift toward yield-enhanced, site-specific, and risk-informed nutrient management strategies that better reflect the realities of smallholder decision-making under increasingly uncertain conditions.
Efficient fertilizer application is vital for enhancing maize production and profitability in Sub-Saharan Africa, where soil fertility varies widely across regions. This study aimed to develop a machine learning approach for generating site-specific fertilizer recommendations for maize production in Ghana and to evaluate its performance against conventional and semi-mechanistic approaches. A random forest machine learning model was trained on 482 maize yield experiments, consisting of 3136 yield observations collected from 1991 to 2020, to predict maize yield response to different fertilizer rates. The model incorporated multiple explanatory variables, including soil properties, climate conditions, and management practices, to generate fertilizer response curves from which fertilizer recommendations were derived for 14 sites across three agro-ecological zones in Ghana where field validation experiments were conducted. On these sites, the recommendations were compared with recommendations derived from the Quantitative Evaluation of the Fertility of Tropical Soils (QUEFTS), Conventional Fertilizer Dose Response (CFDR), and Updated Conventional Fertilizer Dose Response (UCFDR) approaches and validated through field experiments. The machine learning approach generally recommended lower rates of phosphorus and potassium than the other approaches, while nitrogen recommendations were comparable. In the Guinea Savanna zone, the recommendations from the machine learning approach outperformed those from the other approaches, producing higher mean yields for three out of the four sites in the zone. In the Forest-Savanna Transition (FST) zone, the machine learning model recommendations led to higher mean yields at four sites, while the approaches based on QUEFTS and UCFDR performed best at two other sites. In the Semi-deciduous Forest zone, the recommendations of the QUEFTS approach resulted in the highest mean yields at three sites, and CFDR at one site. Despite high input prices during the period of experimentation, the machine learning approach-based recommendations demonstrated higher net profit margins in the FST zone, suggesting cost-effectiveness in this zone. These findings indicate that site-specific fertilizer recommendations are more efficient than blanket recommendations and that machine learning approaches offer a promising and innovative approach for generating cost-effective, site-specific fertilizer recommendations in tropical climates.
Machine learning and geostatistics are two fundamentally different frameworks for the prediction and spatial mapping of soil properties. Geostatistics leverages the spatial structure of soil properties, whereas machine learning models capture the relationship between available environmental features and soil properties. We propose a hybrid framework that augments machine learning with spatial context through the engineering of 'spatial lag' features derived from ordinary kriging. We call this approach 'kriging prior regression' (KpR), as it reverses the logic of regression kriging by incorperating kriging outputs before and during the regression step. To evaluate this approach, we assessed both the point and probabilistic prediction performance of KpR, using TabPFN across six field-scale datasets from LimeSoDa. These datasets included soil organic carbon, clay content, and pH, along with features derived from remote sensing and in-situ proximal soil sensing. KpR with TabPFN demonstrated reliable uncertainty estimates and accurate predictions in comparison to several other spatial techniques (e.g., regression/residual kriging with TabPFN), as well as to established non-spatial machine learning algorithms (e.g., random forest and categorical boosting). Most notably, it improved the average R2 by approximate to 30% relative to machine learning algorithms without spatial context. This improvement was due to the strong prediction performance of the TabPFN algorithm itself and the complementary spatial information provided by KpR features. TabPFN is particularly effective for prediction tasks with small sample sizes, common in precision agriculture, whereas KpR can compensate for weak relationships between sensing features and soil properties when proximal soil sensing data are limited. We conclude that KpR with TabPFN is a robust and versatile modelling framework for digital soil mapping in precision agriculture.
Soils are the largest terrestrial carbon reservoir, with soil organic carbon (SOC) playing a critical role in maintaining soil quality and associated ecosystem services. Accurately estimating SOC stocks at high spatial and temporal resolution over large scales remains challenging, particularly in agricultural systems where carbon inputs are often uncertain or unavailable. In this study, we used the RothC model to simulate SOC stocks in Dutch agricultural mineral soils from 1986 to 2022, at 25 m & times; 25 m resolution. We examined the temporal and spatial variation of the total SOC stock and its distribution over RothC carbon pools and unravelled how livestock manure inputs and land use affect the observed trends. Averaged SOC stocks in the topsoil (0-30 cm) increased by 13.2% under grassland, decreased by 10.4% under cropland, and decreased by 3.9% in areas with changing land use. Carbon gains in grassland were linked to systematically higher manure inputs and accumulation in stable pools, whereas lower manure inputs and more intensive management led to declining labile SOC pools. Independent validation on three spatial datasets showed the highest model performance for point-based field data (model efficiency coefficient MEC = 0.32 in 1986 and 0.37 in 2022). Observed changes in SOC over time could be less well reproduced (MEC approximate to 0) across all datasets, but simulated spatiotemporal patterns were consistent with previous observational studies. The study illustrates the potential of RothC for national-scale SOC stock assessment and monitoring, while highlighting the need for improved input data and temporal validation data. Importantly, this modelling approach effectively captures SOC stock dynamics, which remains challenging for purely empirical, statistical models. Future work could benefit from hybrid modelling approaches that integrate RothC with machine learning, enhancing the ability to capture currently unexplained variability and improve simulation performance.
CONTEXT: China has the world's largest potato cultivation area, yet potato yield remains relatively low. A better understanding of current yield-limiting and reducing factors is essential to improve yield and input use efficiency. OBJECTIVE: This study aims to quantify yield gaps at the field level and uncover its main drivers for irrigated and rainfed potato farming systems across China. METHODS: We used 1836 field-year combinations from major potato-producing areas in China between 2019 and 2021. Potential yield (YP), water-limited yield (YW) and water-and nitrogen-limited yield (YWN) were simulated with World Food Studies model (WOFOST). The primary drivers of the yield gap in irrigated fields (Ygap_irri) and rainfed fields (Ygap_rain) were explored using a WOFOST-based yield gap decomposition scheme and random forest model-based covariate importance analysis. RESULTS AND CONCLUSIONS: There remains substantial potential to increase potato yields in China. The highest Ygap_irri and Ygap_rain were observed in the northern cultivation zone. The average Ygap_irri was 8.2 t dry matter (DM) ha-1, with inadequate amounts and frequencies of irrigation identified as the main drivers, explaining 48.8% of Ygap_irri. Optimising irrigation regimes and water-saving strategies are therefore necessary to close the yield gap. The average Ygap_rain was 9.9 t DM ha-1, and province-level socioeconomic covariates were the primary explanatory variables, particularly farmers' income and education level. This suggests that crop management can only be improved when farmers in rainfed areas have better access to financial resources, high-quality inputs and technical knowledge. Yield limitation due to insufficient rainfall was substantial (3.5 t DM ha-1), indicating a notable water deficit across rainfed fields. Both approaches consistently identified nitrogen as the least influential factor for yield gap because of the general oversupply of N fertiliser, with half of the fields receiving more than 210 kg ha-1. SIGNIFICANCE: This study provides valuable recommendations for management practices and policy interventions aimed at narrowing the potato yield gap in China. The findings contribute to enhanced potato productivity and improved resource-use efficiency.
Background Cassava is grown on over one million hectares in Tanzania, yet yields remain far below their potential, partly due to limited use of clean planting material. Tanzania has developed one of Africa's largest cassava seed systems, with roughly 600 Commercial Seed Entrepreneurs (CSEs) supplying clean and certified seed. Despite this progress, large spatial differences persist in CSE sales and market development. This study maps spatial variation in cassava seed market performance and identifies where seed system interventions could be most effective. Methods We focused on Tanzania's Lake Zone, where cassava production is concentrated and CSE density is highest. We developed spatial models for sell ratios (proportion of seed sold versus seed produced) using two approaches: a Random Forest (RF) model with 25 agricultural, climate, terrain, and human activity predictors, and a Structural Potential Index (SPI) combining cassava production, economic importance, and market accessibility. RF performance was assessed using nearest-neighbour distance matching leave-one-out cross-validation to produce realistic accuracy estimates given the spatial clustering of CSE observations. Results The RF model explained 25% of variance in sell ratios out-of-bag and 13% under spatially realistic cross-validation. Prediction intervals spanned nearly the full data range, indicating that individual CSE circumstances likely dominate over spatial conditions. SPI showed a moderate association with observed sell ratio at CSE locations. The RF-SPI cross-classification revealed two divergence categories: opportunity zones (19.7%, ~ 56,000 km²) where structural potential exists but sell ratios remain low, and emergent markets (19.7%) where entrepreneurs succeed despite weak structural conditions. The remaining 60.6% comprises agreement zones. Conclusions The four typology zones each suggest distinct policy responses: functioning markets for consolidation, opportunity zones for interventions addressing market-entry constraints (business-skills training, buyer-seller linkages, working capital), emergent markets for investigation of factors enabling success despite unfavourable conditions, and constrained zones as lower priority for cassava-specific investment.
High-resolution maps of climate and ecosystem variables are essential for supporting terrestrial carbon stocks and fluxes estimation, climate change mitigation, and ecosystem degradation assessment. These maps are usually created using remotely sensed data obtained from various types of imagery and sensors. The remote sensing data typically serve as covariates to deliver spatially explicit information using machine learning algorithms. Often the uncertainty associated with the maps is also quantified, for instance by prediction error variance maps or by maps of the lower and upper limits of a prediction interval. In addition, these products are often aggregated to regional, national, or global scales relevant to climate policy, natural resource inventory, and measurement, reporting, and verification (MRV) frameworks. Quantifying uncertainty in aggregated products is crucial as it is necessary to assess their value and evaluate whether changes and trends in aggregated estimates are statistically significant. However, we argue that such uncertainty is frequently inaccurately assessed due to the neglect of spatial correlation in map errors. This critical methodological issue has been overlooked in most large-scale mapping studies.
Maize as staple crop is essential for food security in Ghana, yet its yields remain low and highly variable despite increased fertilizer use. Understanding yield variability is crucial for effective agronomic strategies and reducing the risks of inappropriate fertilization strategies. This study quantifies maize yield variability, evaluates predictive model accuracy, and identifies key yield-influencing factors using machine learning (ML) models. A dataset of 5213 maize yield observations from 2000 to 2022 was analyzed using random forest (RF), gradient boosting (GB), and cubist models. Model performance was assessed through nested 20 x 10-fold cross-validation and grid search, using model error, concordance correlation coefficient, model efficiency coefficient (MEC), and mean absolute error. SHapley Additive exPlanations value and H-statistics determined variable importance and first- and second-order effects. Yield coefficient of variation was 55 %, with fertilizer use reducing variability on average by 15 %. ML models demonstrated high predictive accuracy, with RF, GB, and cubist achieving similar mean MCEs of 0.69. All three models identified nitrogen fertilizer (NF) and exchangeable soil magnesium (Mg) as the key drivers of maize yield. NF increased maize yields by an average of 181 kg ha(-1) (RF), 287 kg ha(-1) (GB), and 293 kg ha(-1) (Cubist) across a range of 0-240 kg NF ha(-1). Similarly, Mg contributed to yield increases on average of 161 kg ha(-1) (RF), 198 kg ha(-1) (GB), and 287 kg ha(-1) (cubist) over a range of 0.5-2.4 mEq 100 g(-1). Additionally, interactions between NF, soil moisture, and Mg were found to significantly influence maize yield. Our findings underscore the pivotal role of NF and Mg (a relatively understudied soil nutrient in Ghana) in influencing maize yield. The study advocates for Mg- contain fertilization and highlights the need for further research to deepen Mg impact on maize production in Ghana.
Smallholder farmers in Ethiopia generally do not have access to soil testing services for nutrient management planning decisions; as soil analysis is too costly for most farmers. Fertiliser advice is generally accessible via blanket recommendations at a national scale. Hence, an alternative approach is needed to estimate soil nutrient content across the diverse landscapes of Ethiopia. In this study, we propose using diagnostic features to estimate soil nutrient content, which could contribute to the development of fertiliser recommendations. To achieve this the following objectives were defined: (i) to estimate soil nutrient content as influenced by soil diagnostic features; and (ii) to elucidate the influence of environmental covariates and diagnostic features on the estimation of soil nutrient levels in the Ethiopian context. Data from 550 soil profiles, distributed across Ethiopia, were collected from a range of published sources, collated and harmonised. The data were cleaned, and 496 soil profiles were prepared for modelling. To identify which diagnostic characteristics were present across these soils we applied a presence/absence scoring method to identify dominant diagnostic features. Multiple linear regression analyses were used to predict soil chemical properties from the diagnostic features and diagnostic features along with environmental covariates. The performance of the models was evaluated by applying a 10fold cross-validation using mean error (ME), Lin's concordance correlation coefficient (LCCC), root mean square error (RMSE) and model efficiency coefficient (MEC). The MEC values for pH, TN, and CEC derived from a combination of diagnostic features and environmental covariates were 0.38, 0.33, and 0.38. The corresponding RMSE values were 0.78, 0.07 %, and 13 cmol kg- 1. Additionally, the LCCC values for pH, TN, and CEC were 0.62, 0.58, and 0.62, respectively. The cross-validation results for soil chemical properties showed that the model's performance improved when environmental covariates were added. Precipitation, temperature, geology and land cover were the most important environmental covariates for estimating nutrient content, along with diagnostic features of Ethiopian soils. In conclusion, the diagnostic approach offers a useful starting point for estimating soil nutrient content. However, the variation in nutrient content across the six diagnostic features was not adequately quantified, and the model's predictive performance remains insufficient for practical application at the local scale. Further expansion of the dataset is required to fully exploit the potential of these models for underpinning nutrient management decisions across Ethiopia and in other regions where access to soil test information is limited.
Africa's soil information landscape is characterized by a paradox: a vast body of historical data coexists with a critical lack of integrated, FAIR (Findable, Accessible, Interoperable, and Reusable) data required to address food security and climate resilience. Fragmented, project-based efforts have led to a landscape of siloed, noninteroperable datasets, hindering scientific progress and evidence-based policymaking. This paper argues that a fundamental shift is needed and proposes a comprehensive research agenda to guide the creation of a sustainable, federated African Soil Information System. The agenda is structured around 10 interconnected pathways that provide a holistic roadmap, moving from problem identification to scalable solutions. Key pathways address the rescue and quality assessment of legacy data; the adaptation of practical data standards; the development of a federated infrastructure that respects national data sovereignty; the cultivation of human and institutional capacity; and the establishment of a long-term governance and stewardship model. By providing a blueprint for collective action, this paper serves as an invitation for stakeholders to unite around a common vision, transforming Africa's soil data from a fragmented archive into a durable, shared intelligence resource for sustainable development.
The turnover time (τ) of global soil organic carbon is central to the functioning of terrestrial ecosystems. Yet our spatially explicit understanding of the depth-dependent variations and environmental controls of τ at a global scale remains incomplete. In this study, we combine multiple state-of-the-art observation-based datasets, including over 90 000 geo-referenced soil profiles, the latest root observations distributed globally, and large numbers of satellite-derived environmental variables, to generate global maps of apparent τ in topsoil (0–0.3 m) and subsoil (0.3–1 m) layers, with a spatial resolution of 30 arcsec (∼1 km at the Equator). We show that subsoil τ (385203485 years (mean, with a variation range from the 2.5th to 97.5th percentile)) is over 8 times longer than topsoil τ (1511137 years). The cross-validation shows that the fitted machine learning models effectively captured the variabilities in τ, with R2 values of 0.87 and 0.70 for topsoil and subsoil τ mapping, respectively. The prediction uncertainties of the τ maps were quantified for better user applications. The environmental controls on topsoil and subsoil τ were investigated at global, biome, and local scales. Our analyses illustrate the ways in which temperature, water availability, physio-chemical properties, and depth jointly exert impacts on τ. The data-driven approaches allow us to identify their interactions, thereby enriching our comprehension of mechanisms driving nonlinear τ–environment relationships at global to local scales. The distributions of dominating factors of τ at local scales were mapped for purposes of identifying context-dependent controls on τ across different regions. We further reveal that the current Earth system models may underestimate τ by comparing model-derived maps with our observation-derived τ maps. The resulting maps, with new insights, as demonstrated in this study, will facilitate future modelling efforts relating to carbon cycle–climate feedbacks and support effective carbon management. The dataset is archived and freely available at https://doi.org/10.5281/zenodo.14560239 (Zhang, 2025a).
Hand-feel soil texture observations (HFST) are less accurate, yet much more numerous, than laboratory measurements of soil texture (LAST). Therefore, it is tempting to incorporate both LAST and HFST information as calibration data in digital soil mapping (DSM) of particle-size distribution. We used about 1000 LASTs and 15,000 HFSTs over an area of about 6,800 km2. We incorporated the uncertainties of HFST and LAST calibration data in DSM and compared it with a case where measurement errors were ignored. We added progressively HFST calibration data to LAST data and ran predictions and k-fold validations keeping the same validation set for all experiments. We added HFSTs according to different strategies: either based on the most uncertain predicted areas from the LAST-only model, or those from the preceding LAST + nHFST model, or randomly. We discuss the pros and cons of these different strategies. Adding HFST data brought useful information for model calibration, but only if the uncertainty was accounted for. Various strategies for adding HFSTs led to different unbalanced samplings, maps and prediction intervals. We explain how these various unbalanced samplings sharpened or enlarged the predictive distribution of various clay content ranges. Adding a too large number of HFSTs led to an over-optimistic estimation of the 90 % prediction interval and to large homogeneous patterns, smoothing the spatial variation of clay content. Adding HFSTs using weights to acknowledge their uncertainty substantially improved DSM predictions, but the number of HFSTs and the strategy to add them must be carefully adapted.
Accurate soil property maps are essential for effective soil nutrient management. In Ethiopia, fertilizer applications often ignored spatial variability in soil properties, leading to inefficiencies. This study employed digital soil mapping to generate three-dimensional (3D) maps of five key soil fertility properties-total nitrogen (TotalN), extractable phosphorus (OlsenP), exchangeable potassium (ExchK), pH-H2O (pH), and organic carbon (OC)-at 100 m resolution across central Ethiopia. The objectives were to (1) develop maps at six depth intervals while assessing prediction uncertainty; (2) evaluate the integration of topsoil data with soil profile data for model training; and (3) compare the maps with Africa-SoilGrids, SoilGrids, and iSDAsoil maps. We used two datasets: soil profile (1,379 profiles with 4,179 layers) and topsoil (13,724 locations), harmonized with transfer functions. Quantile regression forest was used to generate maps with 90 % prediction intervals. Models were calibrated with 80 % of the dataset and 194 covariates, including depth, and evaluated with the remaining 20 %. Integrating topsoil data with the soil profile dataset improved prediction accuracy for the five soil fertility properties in the topsoil (0-20 cm), demonstrating near-zero bias, reduced root mean squared error, and a higher model efficiency coefficient (MEC), compared to only using the soil profile dataset. It also enhanced uncertainty quantification for pH, OlsenP, TotalN, and ExchK in the topsoil. However, these benefits diminished with depth, with slight improvements in the subsoil (20-50 cm) but none in the deeper layers (50-200 cm) where pH and OC predictions were even slightly biased. Among the combined dataset models, the highest performance was for pH (MEC = 0.80), while the lowest was for OlsenP (MEC = 0.13). The maps generated from the combined dataset models showed MEC improvements of 27 % to over 1,000 % compared to SoilGrids, Africa-SoilGrids, and iSDAsoil. Additionally, the prediction intervals were also a realistic representation of the prediction uncertainty, with prediction interval coverage probability (PICP) values close to their ideal value. This markedly outperformed SoilGrids and iSDAsoil, which had unrealistically low PICP values. We conclude that the 100 m resolution 3D maps from this study offer satisfactory accuracy and realistic uncertainty quantification, making them the currently best available resource for developing location-specific fertilizer recommendations in central Ethiopia.
Assessment of the leaching potential of pesticides and their metabolites is an important part of the authorization procedure for pesticides in Europe. To protect groundwater quality, it must be demonstrated that concentrations of active substances in the upper groundwater do not exceed 0.1 μg/L before a pesticide can be approved for use. For the purpose of exposure assessment, this concentration limit is imposed on the water leaching downward at 1 m depth in the soil profile. For a given substance and application pattern, this leaching concentration can vary in space by several orders of magnitude, due to variation in site conditions, most importantly soil properties and climate. Spatially distributed leaching modelling (SDLM) is a methodology for exposure assessment over large spatial extents, dealing with this spatial variability in a comprehensive way. It involves performing simulations for many parametrizations representative for a spatial region and can be used to generate maps or calculate spatio-temporal percentiles of leaching concentrations. While such tools are already used in exposure assessment at national level in several EU member states, no generally accepted SDLM tool is available at the European level. In 2020, a working group of Society of Environmental Toxicology and Chemistry (SETAC) was formed with the purpose to develop a harmonized framework for SDLM across Europe (EU27 + UK). A first version of an SDLM—referred to as GeoPEARL-EU—was built around the pesticide leaching model PEARL, a field-scale model of pesticide fate in the soil-plant system. PEARL mechanistically simulates pesticide behaviour in a 1D soil column based on explicit descriptions of transport in the liquid and gas phases, sorption to the solid phase, degradation, volatilisation, and plant-uptake. Soil moisture content and fluxes are provided by the SWAP hydrological model. PEARL is used in regulatory exposure assessment for groundwater and soil. Furthermore, a spatially distributed tool based on PEARL (GeoPEARL) is used for exposure assessment in the Netherlands. To apply PEARL to Europe, pan-European gridded datasets were collected for several variables, including soil texture, pH, soil organic carbon, weather, irrigation patterns and crop area. These datasets were used to develop a set of parametrizations covering the variability of climate and soil conditions in Europe. To this end, all 1x1 km grid cells for the EU27 + UK were partitioned into approximately 10,000 clusters using k-means clustering, based on several soil- and climate-related variables relevant for leaching vulnerability. Subsequently, a representative grid cell was selected for each cluster, which was used to obtain the data required to parameterize PEARL from the spatial data sets. Pedotransfer functions were used to derive soil hydraulic parameters. We will present results from GeoPEARL-EU for several test cases with specific attention to the effect of the spatial aggregation approach on the model predictions. Moreover, we discuss how the tool could be used in the tiered approach of the regulatory exposure assessment for groundwater in the EU.
Machine learning (ML) is increasingly being used to enhance yield predictions and optimize agronomic practices in sub-Saharan Africa. Yet, understanding how these models generalize across heterogenous ecological context remains unresolved. This study, conducted in Ghana, evaluates the predictive performance of four ML models, namely, random forest (RF), support vector machine (SVM), k-nearest neighbors (KNN), and extreme gradient boosting (XGBoost) for predicting maize yield and agronomic efficiency-defined as the increase in yield per unit of nutrient applied. It also compares variable importances identified by these models and how they influence yield and agronomic efficiency. The analysis used 4496 georeferenced maize trial datasets from various agroecological zones across Ghana, incorporating 35 variables related to soil properties, climate, topography, crop management, and fertilizer application. Model performance was assessed using three cross-validation techniques: leave-one-out, leave-site-out, and leave-agroecological-zone-out. Accuracy was measured using mean error, root mean square error (RMSE), and model efficiency coefficient. When evaluated under leave-one-out cross-validation, XGBoost consistently achieved the highest predictive accuracy with the lowest RMSE for yield (639.5 kg ha-1) and for agronomic efficiency of nitrogen (11.6 kg kg-1), which is moderate given the high variability in on-farm nutrient response. RF also performed well, while KNN and SVM showed poor extrapolation under stringent validation. Nitrogen application rate, rainfall, and crop genotype were consistently identified as the most influential explanatory variables across all models, providing insight into key drivers of productivity. These findings demonstrate the power of ML techniques in supporting agricultural planning and improving maize production in sub-Saharan Africa.
Edzer J. Pebesma合作论文数Institute for Geoinformatics, University of Muenster26