On-farm experimentation (OFE) is increasingly recognized as a pathway for developing sustainable and context-specific agricultural practices under real-world production conditions. However, OFE outcomes are difficult to interpret, compare, and aggregate when the environmental and biophysical context of each experiment is not described in a consistent and reproducible way. The main barrier is not only access to public geospatial data, but the repeated transformation of heterogeneous soil, terrain, weather, drought, and satellite products into traceable, analysis-ready covariates linked to agricultural fields and experimental time windows.This study presents the Geospatial Data Retrieval Tool (GeoDaRT), a web-based system that implements a reproducible contextual data workflow for OFE and related agricultural research. GeoDaRT organizes retrieval around experiment-oriented requests consisting of an area of interest, temporal parameters, selected data products, and processing options. The workflow standardizes access, clipping, file formatting, naming, metadata, packaging, and provenance while preserving native source resolution and coordinate reference information where appropriate.GeoDaRT was evaluated through a distributed field-scale analytical-readiness workload involving 200 corn and soybean fields in New York State and the U.S. Midwest over three growing seasons. The system produced 105,984 contextual files representing 600 site-years. These outputs supported an end-to-end multivariate analysis of 19,900 field pairs without custom reformatting. The analytical-readiness workload produced a coherent contextual feature space and demonstrated that GeoDaRT outputs could support residual-tail summaries and mapped interpretation. These results demonstrate that GeoDaRT makes contextual covariate generation more feasible, auditable, and reusable for distributed OFE.
Soil health (SH) is inherently dynamic, and traditional soil surveys do not capture it. Its spatiotemporal variability presents challenges to understanding its drivers. We applied machine learning (ML) models and digital soil mapping (DSM) techniques to integrate SH observations with remotely-sensed data products for regions in New York State, representing soil forming factors, cropping systems, and management. Four biological (soil organic matter, permanganate-oxidizable C, protein, and respiration) and two physical (water aggregate stability and available water capacity) soil indicators as well as a composite SH index were evaluated to 1) quantify the relationships among climate, inherent soil properties, and land use to SH indicators; 2) develop data-driven models for predicting and mapping SH indicators at regional scale; and 3) use predicted SH maps to estimate impacts from hypothetical regional land use change scenarios. Model performance varied among SH indicators and region, with R2 values ranging from 0.47 to 0.71 in the smaller domain and from 0.45 to 0.65 in the larger one. The models include conventional effects of inherent soil properties and climate on SH, but also prove the pivotal role of land use and cropping systems, explaining on average 63% of SH variation. Positive effects were associated with perennial forage crops, and adverse effects of intensive annual crop production. Also, biomass production and cycling strongly affect SH. Modeling SH responses to land use changes at a regional scale showed interactions between management and inherent properties that can affect SH benefits from alternative cropping systems. Overall, the geospatial application of ML models to mapping SH provides insights into its drivers that can support the design of informed policies and management interventions to improve SH.
Digital technologies have significantly improved nitrogen (N) fertilizer optimization in precision agriculture. A limitation of existing crop N detection approaches is that they primarily focus on above-canopy spectral measurements, overlooking the potential insights from lower canopy levels, which may more accurately reflect N stress through spectral reflectance associated with differential pigment expression. This study introduces high-resolution Red, Green, and Blue (RGB) under-canopy imaging of maize (Zea mays L.) to assess in-season N fertilizer application at high spatial resolution using a 30 frames/sec acquisition rate. Utilizing a purpose-built field robot, developed specifically for this study and equipped with a digital RGB camera, field trials were conducted across Minnesota and New York with varying N rates. Analysis using multiple thresholding methods for the (R-B)/(R+B) index from images captured during day and night revealed a strong correlation between under-canopy images and applied N rates. R2 values reached up to 0.78 in daytime and 0.92 under nighttime conditions. Semivariogram analysis indicated a range of influence of less than 6 m and showed that N stress spatial patterns are most pronounced with low N levels. Maps were generated based on 6-m field sections to represent field variability of N stress. These findings suggest that high-resolution under-canopy RGB imaging is a viable, lowcost method for detecting maize N status with very high spatial resolution, offering a new perspective for precision agriculture.
Digital Soil mapping (DSM) at continental and global scale provides standardised global information layers. It is also an important tool to create soil information layers for areas for which local soil survey information is lacking. The recent availability of global and continental remote sensing derived products coupled with the ease-of-access to computational resources has made the production of such layers easier across the globe. Therefore, it is ever more important to assess the quality of DSM-derived products, in particular the type of information they can actually provide to users (i.e., fitness for intended use). DSM studies commonly assess prediction uncertainty using various approaches, including multiple simulations or quantile random forests. However, this does not encompass all the potential elements that could be used to characterise the uncertainty of a DSM product. In this study we are going to assess maps based also on area of applicability (i.e., the area in covariate space where the model learns about relationships based on the training data) and the landscape heterogeneity both in the landscape itself and in covariate space. We present examples of continental and global mapping products, highlighting main uncertainty-related issues and how these influence suitability for intended use by stakeholders, decision makers and users in general at the given resolution. The examples come from a range of projects with different aims and goals. The results permit some practical reflections on how to integrate all the above elements to identify regions where the confidence in the predictions is highest and the associated uncertainty lowest. We will integrate the practical reflections with information collected from a user survey on requirements and usability of continental and global DSM products.
Far removed from the agricultural fire “hotspots” of Northwestern India, rice residue burning is on the rise in Eastern India with implications for regional air quality and agricultural sustainability. The underlying drivers contributing to the increase in burning have been linked to the adoption of mechanized (combine) harvesting but, in general, are inadequately understood. We hypothesize that the adoption of burning as a management practice results from a set of socio-technical interactions rather than emerging from a single factor. Using a mixed methods approach, a household survey (n = 475) provided quantitative insights into landscape and farm-scale drivers of burning and was complemented by an in-depth qualitative survey (n = 36) to characterize decision processes and to verify causal inferences derived from the broader survey. For communities where the combine harvester is present, our results show that rice residue burning is not inevitable. The decision to burn appears to emerge from a cascading sequence of events, starting with the following: (1) decreasing household labor, leading to (2) decreasing household livestock holdings, resulting in (3) reduced demands for residue fodder, incentivizing (4) adoption of labor-efficient combine harvesting and subsequent burning of loose residues that are both difficult to collect and of lower feeding value than manually harvested straw. Local demand for crop residues for livestock feeding plays a central role mediating transitions to burning. Consequently, policy response options that only consider the role of the combine harvester are likely to be ineffective. Innovative strategies such as the creation of decentralized commercial models for dairy value chains may bolster local residue demand by addressing household-scale labor bottlenecks to maintaining livestock. Secondary issues, such as timely rice planting, merit consideration as part of holistic responses to “bend” agricultural burning trajectories in Eastern India towards more sustainable practices.
Parent materials have a strong control on soil formation process and soil properties, understanding their provenance can provide important information for soil genesis, classification, and regionalization. The Songnen Plain in northeast China has large areas of Phaeozems and Chernozems, popularly known as “black soils”, but their parent material provenance and sedimentary processes are still unclear. Therefore, this paper analyzes the types and formation process of soil parent materials in the context of regional environmental changes. This analysis is based on the characteristics of grain size distribution and quartz particle morphology at the regional and profile scales. The results indicate that aeolian loess is the predominant parent material. For instance, the surface of quartz particles exhibits characteristics indicative of mechanical impact, which is produced during the process of wind transportation. The particle size distribution curve displays a bimodal pattern, and the soil particle size tends to become finer from west to east. However, in some areas, the soil is influenced by river or lake sediments. The main source areas of aeolian deposits are likely the Gobi and sandy land in the upwind direction of the study area, while the Songhua River alluvial deposits only provide source material in local areas. High-resolution grain size analysis and K-feldspar single-particle OSL chronology show that from the Last Glacial Maximum to the Early Holocene, far-source materials dominated the deposition process. In the Middle Holocene, climate warming increased the frequency of dust activities and accelerated the deposition process. In the late Holocene, climate fluctuations and intensified human activities led to more intense dust storms in the provenance area, which in turn promoted the continuous accumulation of parent materials for soils that developed into Phaeozems and Chernozems.
Many Digital Soil Modelling (DSM) products have been generated for diverse regions, countries and continents. While most of these products provide some accuracy metrics, a few also incorporate assessments of uncertainty. The current uncertainty estimates for DSM products often fail to represent elements that are essential for evaluating the suitability of a map for a specific application. For instance, different models can have very similar accuracy metrics but produce different soil-landscape patterns. It is important to be able to evaluate the accuracy of the patterns as well, current accuracy metrics do not do this. Additional metrics could be defined that are able to do this such the ‘area of applicability’, i.e. the area in covariate space where the model learns about relationships based on the training data) and the landscape heterogeneity both in the landscape itself and in covariate space. This study delves into the integration of the aforementioned elements into an assessment of DSM uncertainty at the continental scale. Europe was used as test area, incorporating input observations from EU-LUCAS datasets. A covariate space encompassing the soil forming factors as defined by the SCORPAN model served as a basis for fitting the necessary models for soil products.. We characterized the spatial heterogeneity of both the landscape and the covariate space by employing commonly employed landscape metrics. The findings offer practical insights on how to integrate these components to produce more reliable products for stakeholders.
Reducing methane (CH 4 ) emissions is increasingly recognized as an urgent greenhouse gas mitigation priority for avoiding ecosystem 'tipping points ' that will accelerate global warming. Agricultural systems, namely ruminant livestock and rice cultivation are dominant sources of CH 4 emissions. Efforts to reduce methane from rice typically focus on water management strategies that implicitly assume that irrigated rice systems are consistently flooded and that farmers exert a high level of control over the field water balance. In India most rice is cultivated during the monsoon season and hydrologic variability is common, particularly in the Eastern Gangetic Plains (EGP) where high but variable rainfall, shallow groundwater, and subtle differences in topography interact to create complex mosaics of field water conditions. Here, we characterize the hydrologic variability of monsoon season rice fields ( n = 207) in the Indian EGP ('Eastern India ') across two contrasting climate years (2021, 2022) and use the D e n itrification D e c omposition (DNDC) model to estimate GHG emissions for the observed hydrologic conditions. Five distinct clusters of field hydrology patterns were evident in each year, but cluster characteristics were not stable across years. In 2021, average GHG emissions (8.14 mt CO 2-eq ha-1 ) were twice as high as in 2022 (3.81 mt CO 2-eq ha-1 ). Importantly, intra-annual variability between fields was also high, underlining the need to characterize representative emission distributions across the landscape and across seasons to appropriately target GHG mitigation strategies and generate accurate baseline values. Simulation results were also analyzed to identify main drivers of emissions, with readily identified factors such as flooding period and hydrologic interactions with crop residues and nitrogen management practices emerging as important. These insights provide a foundation for understanding landscape variability in GHG emissions from rice in Eastern India and suggest priorities for mitigation that honor the hydrologic complexity of the region.
In 2023, the European Commission released a legislative proposal for a Directive on Soil Monitoring and Resilience which aims to define a legal framework to achieve healthy soils across the European Union (EU) by 2050. A key component of the initial Directive is the mandate for Member States to establish basic geographic soil governance units, referred to as soil districts, and appoint a district-specific authority to oversee the implementation of soil health assessments. This paper proposes an operational definition of the districts following the conditions outlined in the proposal for the Directive and discusses various attention points for their implementation. Tentative districts were developed for seven EU countries, considering soil type, climate, topography, and land cover factors, starting from the smallest existing administrative unit (i.e. municipalities). Experts were asked to report on the applicability of the proposed districts within well-known pedo-ecological regions and discuss the relevance of the districts for establishing an EU-wide monitoring network and reporting on soil health and degradation. The outcomes highlight the need for detailed soil maps to account for specific soil types when stratifying countries into soil districts. The soilscape approach allows for a consistent method to defining soil districts across Member States. This enables contrasting soils within a district to be managed in a similar manner, with soil degradation/health thresholds applied to each district based on land cover. However, it is unclear whether soil districts as currently formulated in the Directive are in fact the right tool to support local soil management and monitoring of soil health. Districts can help ensure that all soil conditions are covered in a monitoring system, but they may not provide support for soil management or monitoring at a local scale due to short-scale soil variability and threats affecting soil management within the same soilscape. Beyond the use of districts for designing a European/national scale monitoring system, the districts can help create animations and other educational tools to promote soil literacy and connectivity of users to soils locally.
Groundwater irrigation supports over 40% of global crop production and stabilizes yields amidst climatic change. Yet, over-abstraction can cause water scarcity, disrupt ecosystems, and increase greenhouse gas emissions. Governments and international financial institutions have made significant investments in sustainable groundwater irrigation but require enhanced spatial targeting to increase impact. In response, this study employs an agro-hydrological machine-learning approach to analyze spatial patterns of (i) crop yield responses to increased irrigation and (ii) groundwater sustainability in South Asia – characterized by smallholder farming, increasing groundwater dependence, and post-green revolution sustainability challenges. We show that modestly increasing irrigation intensity in groundwater-rich areas with high yield responses could boost rice production by 2.22Mt annually – sufficient to feed over 33 million people with little anticipated risks of groundwater depletion. However, current investments overlook these areas. Our approach can be globally applied to catalyze sustainable irrigation through integrated use of expanding agricultural and hydrological datasets.
In 2019, the Government of India launched the National Clean Air Program to address the pervasive problem of poor air quality and the adverse effect on public health. Coordinated efforts to prevent agricultural burning of crop residues in Northwestern IGP (Indo-Gangetic Plain) have been implemented, but the practice is rapidly expanding into the populous Eastern IGP states, including Bihar, with uncertain consequences for regional air quality. This research has three objectives: (1) characterize historical rice residue burning trends since 2002 over space and time in Bihar State, (2) project future burning trajectories to 2050 under ‘business as usual’ and alternative scenarios of change, and (3) simulate air quality outcomes under each scenario to describe implications for public health. Six future burning scenarios were defined as maintenance of the ‘status quo’ fire extent, area expansion of burning at ‘business as usual’ rates, and a Northwest IGP analogue, of which both current rice yields and plausible yield intensification were considered for each case. The Community Earth System Model (CESM v2.1.0) was used to characterize the mid-century air quality impacts under each scenario. These analyses suggest that contemporary Bihar State burning levels contribute a small daily average proportion (8.1%) of the fine particle pollution load (i.e. PM _2.5 , particles ⩽2.5 μ m) during the burning months, but up to as much as 62% on the worst of winter days in Bihar’s capital region. With a projected 142% ‘business as usual’ increase in burned area extent anticipated for 2050, Bihar’s capital region may experience the equivalent of 30 PM _2.5 additional exceedance days, according to the WHO standard (24 h; exceedance level: 15 µ g m ^−3 ), due to rice residue burning alone in the October to December period. If historical burning trends intensify and Bihar resembles the Northwest States of Punjab and Haryana by 2050, 46 d would exceed the WHO standard for PM _2.5 in Bihar’s capital region.
The High Atlas Mountains of Morocco are recognized as global hotspot for rapid environmental change, but there is limited information about how communities and households are responding to these changes. Rural livelihoods that are dependent on agriculture are highly vulnerable to intensifying climate extremes, especially when these stressors intersect with long-term socioeconomic trends including out-migration to urban centers. In 2022–2023, we carried out a household surveys and focus group discussions to understand the evolution of livelihood strategies in four Amazigh villages in Imegdal Commune in the western High Atlas. Results suggest that water shortages are causing cropping systems to simplify as households stop planting some crop species and reduce the area planted to others. Households are also reducing livestock numbers in response to the current multi-year drought and reductions in labor availability created by migration. Other natural resource-based activities, including beekeeping and collecting wild herbs, are being abandoned. This study suggests that decreasing precipitation is rapidly undermining the viability of agricultural activities in the High Atlas. In the absence of viable adaptation strategies, this could lead to a profound restructuring of rural livelihoods across the region.
In the Eastern Gangetic Plain (EGP) soil hydrology is a major determinant of land use and also governs the ecosystem services derived from cropping systems, particularly greenhouse gas (GHG) emissions from rice fields. To characterize patterns of soil hydrology in these, daily field monitoring of water levels was conducted during the monsoon (kharif) season in a comparatively wet (2021) and dry (2022) year with flooding depth and drainage tracked with field water tubes across 47 (2021) and 183 (2022) locations. Fields were clustered into hydrologic response types (HRT) which can then be used for land surface modelling, land use recommendations, and to target agronomic interventions that contribute to sustainable development outcomes. Clusters based on two methods of summarizing a single information source were compared. The information source was a time-series of field water-level observations, and the two methods were (1) the original time-series and their first differences and (2) a set of derived hydrologic descriptors that are conceptually related to greenhouse gas (GHG) emissions. Clustering was (1) by k-means with an optimization of cluster numbers and (2) by hierarchical clustering with the same number of clusters as identified by k-means. Hydrologic behaviour shifted dramatically between growing seasons, and it was not possible to identify consistent HRT's across years. The clusters had only a weak relation with soil properties, almost no relation with farmer perception of relative landscape position, and no relation with rice establishment method. Clusters based on time-series were moderately well predicted in the dry year 2022 by optimized random forest models, with the most important predictors being the number of irrigations, seasonal precipitation, pre-monsoon groundwater levels, seasonal groundwater level change, and pH, this latter as a surrogate for landscape position and other soil properties. In the wet year 2021 clusters were (poorly) predicted by just seasonal precipitation and pre-monsoon groundwater levels. This shows the complex relation of soil hydrology with landscape position and land management, as well as synoptic climate. By contrast, clusters based on the descriptors were not well-matched with those from the time-series, and could not be well predicted by random forest models. This shows that different clustering criteria may result in different interpretations of the landscape hydrology and thus different heuristics for anticipating the hydrology of a given field under different management choices.
The commonly-used scorpan approach to Digital Soil Mapping is purely correlative, between observations at points and the values of covariates at those points. An-often-used approach is of maximum complexity: different quality observations are thrown together with the largest possible number of covariates, along with the most complex model (e.g., ensembles of many models). The resulting products are almost always evaluated with point-wise metrics with sometimes not large differences or improvement between models. Spatial patterns are rarely compared with soil geography. Further work should be done to include more the local and spatial structure component, the ‘n = neighbourhood’ of scorpan. This work addresses how pedological knowledge could be included and the potential advantages and disadvantages of doing so.
The admixture of loess in soils formed in sandy parent materials has considerable impact on pedogenesis and soil ecological functions. This study aimed to evaluate these effects on soil hydrology by quantifying the relationships between loess content and soil hydrological properties in a sandstone landscape covered by Pleistocene periglacial slope deposits in central Europe. Studied properties were saturated soil water capacity, field capacity, permanent wilting point, available water capacity, macro-porosity, matrix porosity and saturated hydraulic conductivity. The effect of loess addition on soil hydrological properties differed between topsoil (pedogenic A horizons) and subsoil (pedogenic B and C horizons). In the subsoil, the studied soil hydrological properties were mainly controlled by loess content. By contrast, in the topsoil the effects of loess content on soil hydrological properties were largely masked by the effects of soil organic matter (reflected by soil organic carbon, SOC). The enhancing effect of SOC on soil hydrological properties was most prominent for saturated soil water capacity and was also significant for all other parameters except permanent wilting point. After the effects of SOC were accounted for, the residual effects of loess on soil hydrological properties were the same in both topsoils and subsoils, with one exception: Saturated hydraulic conductivity decreased with increasing loess content for subsoils, but not for topsoils. This study highlighted the ecohydrological significance of loess admixture in coarse-textured soils, especially for subsoils with low SOC contents. However, for coarse-textured topsoils, SOC content plays the dominant role in affecting soil hydrological properties.
We present methods to evaluate the spatial patterns of the geographic distribution of soil properties in the USA, as shown in gridded maps produced by digital soil mapping (DSM) at global (SoilGrids v2), national (Soil Properties and Class 100 m Grids of the USA), and regional (POLARIS soil properties) scales and compare them to spatial patterns known from detailed field surveys (gNATSGO and gSSURGO). The methods are illustrated with an example, i.e. topsoil pH for an area in central New York state. A companion report examines other areas, soil properties, and depth intervals. A set of R Markdown scripts is referenced so that readers can apply the analysis for areas of their interest. For the test case, we discover and discuss substantial discrepancies between DSM products and large differences between the DSM products and legacy field surveys. These differences are in whole-map statistics, visually identifiable landscape features, level of detail, range and strength of spatial autocorrelation, landscape metrics (Shannon diversity and evenness, shape, aggregation, mean fractal dimension, and co-occurrence vectors), and spatial patterns of property maps classified by histogram equalization. Histograms and variogram analysis revealed the smoothing effect of machine learning models. Property class maps made by histogram equalization were substantially different, but there was no consistent trend in their landscape metrics. The model using only national points and covariates was not substantially different from the global model and, in some cases, introduced artefacts from a lithology covariate. Uncertainty (5 %–95 % confidence intervals) provided by SoilGrids and POLARIS were unrealistically wide compared to gNATSGO/gSSURGO low and high estimated values and show substantially different spatial patterns. We discuss the potential use of the DSM products as a (partial) replacement for field-based soil surveys. There is no substitute for actually examining and interpreting the soil–landscape relation, but despite the issues revealed in this study, DSM can be an important aid to the soil surveyor.
A crucial decision in designing a spatial sample for soil survey is the number of sampling locations required to answer, with sufficient accuracy and precision, the questions posed by decision makers at different levels of geographic aggregation. In the Indian Soil Health Card (SHC) scheme, many thousands of locations are sampled per district. In this paper the SHC data are used to estimate the mean of a soil property within a defined study area, e.g., a district, or the areal fraction of the study area where some condition is satisfied, e.g., exceedence of a critical level. The central question is whether this large sample size is needed for this aim. The sample size required for a given maximum length of a confidence interval can be computed with formulas from classical sampling theory, using a prior estimate of the variance of the property of interest within the study area. Similarly, for the areal fraction a prior estimate of this fraction is required. In practice we are uncertain about these prior estimates, and our uncertainty is not accounted for in classical sample size determination (SSD). This deficiency can be overcome with a Bayesian approach, in which the prior estimate of the variance or areal fraction is replaced by a prior distribution. Once new data from the sample are available, this prior distribution is updated to a posterior distribution using Bayes' rule. The apparent problem with a Bayesian approach prior to a sampling campaign is that the data are not yet available. This dilemma can be solved by computing, for a given sample size, the predictive distribution of the data, given a prior distribution on the population and design parameter. Thus we do not have a single vector with data values, but a finite or infinite set of possible data vectors. As a consequence, we have as many posterior distribution functions as we have data vectors. This leads to a probability distribution of lengths or coverages of Bayesian credible intervals, from which various criteria for SSD can be derived. Besides the fully Bayesian approach, a mixed Bayesian-likelihood approach for SSD is available. This is of interest when, after the data have been collected, we prefer to estimate the mean from these data only, using the frequentist approach, ignoring the prior distribution. The fully Bayesian and mixed Bayesian-likelihood approach are illustrated for estimating the mean of log-transformed Zn and the areal fraction with Zn-deficiency, defined as Zn concentration <0.9 mg kg -1, in the thirteen districts of Andhra Pradesh state. The SHC data from 2015-2017 are used to derive prior distributions. For all districts the Bayesian and mixed Bayesian-likelihood sample sizes are much smaller than the current sample sizes. The hyperparameters of the prior distributions have a strong effect on the sample sizes. We discuss methods to deal with this. Even at the mandal (sub-district) level the sample size can almost always be reduced substantially. Clearly SHC over-sampled, and here we show how to reduce the effort while still providing information required for decision-making. R scripts for SSD are provided as supplementary material.
Fertilization decisions depend on the measurement of a large set of soil fertility indicators, usually through laboratory determination, which is costly and time-consuming. Visible and near-infrared (vis-NIR) spectroscopy combined with machine learning can simultaneously predict various soil fertility indicators. Spectroscopy is inherently less accurate than direct laboratory determination. However, in many fertilization recommendation contexts, farmers mainly fertilize according to classified fertility indicators, rather than by continuous soil property values. These classes have defined limits of property values. We hypothesized that the additional inaccuracy from spectroscopy may not be important for properties grouped into classes. This study compared the indirect and direct prediction of soil fertility classes. Indirectly, by (1) using vis-NIR spectra with machine learning to predict 20 soil fertility indicators (pH, soil organic matter (SOM), cation exchange capacity (CEC), total nitrogen (TN), total phosphorus (TP), total potassium (TK), alkali-hydrolyzable nitrogen (AN), available phosphorus (AP), available potassium (AK), calcium (Ca), magnesium (Mg), silicon (Si), sulfur (S), boron (B), iron (Fe), manganese (Mn), copper (Cu), Zinc (Zn), molybdenum (Mo) and chlorine (Cl)) and (2) allocating the indicators to soil fertility classes. Directly, by predicting soil fertility classes directly from vis-NIR spectra using machine learning. The prediction accuracy of these two methods were compared and the accuracies needed for the acceptable class allocation of the fertility indicators were determined. The example dataset is a soil spectral library from the Guizhou Province, southwest China. The model performance was evaluated by the overall allocation accuracy and tau index, which accounts for class imbalance. For direct allocation based on three fertility classes (low, medium and high), the overall allocation accuracy of eight properties (CEC, Cu, Si, Zn, S, Mn, Ca and Mg), nine properties (B, AN, TK, AK, SOM, TN, TP, Fe and Mo) and three properties (Cl, AP and pH) were within the range of 0.80–1.0, 0.60–0.80 and 0.40–0.60, respectively. For indirect allocation based on the same classes, the allocation accuracy of nine properties (TN, CEC, Cu, S, Zn, Si, Mn, Ca and Mg), nine properties (B, TK, pH, TP, AK, AN, Fe, Mo and SOM) and two properties (Cl and AP) were within the range of 0.80–1.0, 0.60–0.80 and 0.40–0.60, respectively. We conclude that vis-NIR spectroscopy was fairly successful for soil fertility class allocation for most of the soil properties, using either direct or indirect models. The advantage of indirect models is that both specific property values and soil fertility classes can be obtained at no increase in cost, while direct models are suggested when only soil fertility class information are available.