We developed models of suppression expenditures for individual extended attack fires in British Columbia using parametric and nonparametric machine-learning (ML) methods. Our models revealed that suppression expenditures were significantly affected by a fire’s size, proximity to the wildland–urban interface (WUI) and populated places, a weather based fire severity index, and the amount of coniferous forest cover. We also found that inflation-adjusted individual fire suppression expenditures have increased over the 1981 to 2014 study period. The ML and parametric models had similar predictive performance: the ML models had somewhat lower root mean squared errors but not on mean average errors. Better specification of fire priority as well as resource constraints might improve future model performance.
Soil property and class maps for the continent of Africa were so far only available at very generalised scales, with many countries not mappedat all. Thanks to an increasing quantity and availability of soil samples collected at field point locations by various government and/or NGOfunded projects, it is now possible to produce detailed pan-African maps of soil nutrients, including micro-nutrients at fine spatial resolutions. Inthis paper we describe production of a 30 m resolution Soil Information System of the African continent using, to date, the most comprehensivecompilation of soil samples (N ≈ 150, 000) and Earth Observation data. We produced predictions for soil pH, organic carbon (C) and totalnitrogen (N), total carbon, Cation Exchange Capacity (eCEC), extractable — phosphorus (P), potassium (K), calcium (Ca), magnesium (Mg),sulfur (S), sodium (Na), iron (Fe), zinc (Zn) — silt, clay and sand, stone content, bulk density and depth to bedrock, at three depths (0, 20 and50 cm) and using 2-scale 3D Ensemble Machine Learning framework implemented in the mlr (Machine Learning in R) package. As covariatelayers we used 250 m resolution (MODIS, PROBA-V and SM2RAIN products), and 30 m resolution (Sentinel-2, Landsat and DTM derivatives)images. Our 5–fold spatial Cross-Validation results showed varying accuracy levels ranging from the best performing soil pH (CCC=0.900) tomore poorly predictable extractable phosphorus (CCC=0.654) and sulphur (CCC=0.708) and depth to bedrock. Sentinel-2 bands SWIR (B11,B12), NIR (B09, B8A), Landsat SWIR bands, and vertical depth derived from 30 m resolution DTM, were the overall most important 30 mresolution covariates. Climatic data images — SM2RAIN, bioclimatic variables and MODIS Land Surface Temperature — however, remainedas the overall most important variables for predicting soil chemical variables at continental scale. The publicly available 30–m soil maps aresuitable for numerous applications, including soil and fertilizer policies and investments, agronomic advice to close yield gaps, environmentalprograms, or targeting of nutrition interventions.
Spatial autocorrelation in the residuals of spatial environmental models can be due to missing covariate information. In many cases, this spatial autocorrelation can be accounted for by using covariates from multiple scales. Here, we propose a data-driven, objective and systematic method for deriving the relevant range of scales, with distinct upper and lower scale limits, for spatial modelling with machine learning and evaluated its effect on modelling accuracy. We also tested an approach that uses the variogram to see whether such an effective scale space can be approximated a priori and at smaller computational cost. Results showed that modelling with an effective scale space can improve spatial modelling with machine learning and that there is a strong correlation between properties of the variogram and the relevant range of scales. Hence, the variogram of a soil property can be used for a priori approximations of the effective scale space for contextual spatial modelling and is therefore an important analytical tool not only in geostatistics, but also for analyzing structural dependencies in contextual spatial modelling.
In pedology, spatial context is relevant to soil-landscape systems on at least three different scales: i) the scale of quasi-local processes, which are independent of influence from the direct or wider neighborhood, ii) the scale of short-range processes for example on the local hillslope or catena, and iii) the scale of long-range processes, or teleconnected systems. We can represent the effects of teleconnections using existing tools and covariates, but we cannot easily infer or identify their controls, landscape processes or landscape units. We consider that an ability to identify the relevant controls in teleconnected systems would greatly improve pedological interpretation and understanding. Such understanding relates to the interaction of environmental factors and processes in the spatial context, which is relevant for environmental mapping generally. Here we show that teleconnected systems can be disassembled and interpreted using contextual modelling in such a way that the controls, i.e. the cause, can be localized in space. We present examples of how teleconnected systems can be deciphered. The methodology is based on the previously described ConMap approach in combination with Random Forest's measures of local feature importance. ConMap uses elevation differences, computed along multiple rays radiating out from a center grid cell, as predictors instead of complex surface derivatives or decomposed scales of a digital elevation model (DEM) or terrain attribute. Using synthetic and real-world data sets, we show how to identify and interpret teleconnections in soil environmental systems. In the synthetic example, elevation peaks are shown to produce larger values of soil properties, while, in contrast, a valley-mountain system is the main control of soil texture in the real-world example. Our analyses of teleconnected soil environmental systems illustrate that the stochastic component of the universal model of spatial variation is an integral but typically unresolved part of the deterministic component.
This study introduces a hybrid spatial modelling framework, which accounts for spatial non-stationarity, spatial autocorrelation and environmental correlation. A set of geographic spatially autocorrelated Euclidean distance fields (EDF) was used to provide additional spatially relevant predictors to the environmental covariates commonly used for mapping. The approach was used in combination with machine-learning methods, so we called the method Euclidean distance fields in machine-learning (EDM). This method provides advantages over other prediction methods that integrate spatial dependence and state factor models, for example, regression kriging (RK) and geographically weighted regression (GWR). We used seven generic (EDFs) and several commonly used predictors with different regression algorithms in two digital soil mapping (DSM) case studies and compared the results to those achieved with ordinary kriging (OK), RK and GWR as well as the multiscale methods ConMap, ConStat and contextual spatial modelling (CSM). The algorithms tested in EDM were a linear model, bagged multivariate adaptive regression splines (MARS), radial basis function support vector machines (SVM), Cubist, random forest (RF) and a neural network (NN) ensemble. The study demonstrated that DSM with EDM provided results comparable to RK and to the contextual multiscale methods. Best results were obtained with Cubist, RF and bagged MARS. Because the tree-based approaches produce discontinuous response surfaces, the resulting maps can show visible artefacts when only the EDFs are used as predictors (i.e. no additional environmental covariates). Artefacts were not obvious for SVM and NN and to a lesser extent bagged MARS. An advantage of EDM is that it accounts for spatial non-stationarity and spatial autocorrelation when using a small set of additional predictors. The EDM is a new method that provides a practical alternative to more conventional spatial modelling and thus it enhances the DSM toolbox.
We compared different methods of multi-scale terrain feature construction and their relative effectiveness for digital soil mapping with a Deep Learning algorithm. The most common approach for multi-scale feature construction in DSM is to filter terrain attributes based on different neighborhood sizes, however results can be difficult to interpret because the approach is affected by outliers. Alternatively, one can derive the terrain attributes on decomposed elevation data, but the resulting maps can have artefacts rendering the approach undesirable. Here, we introduce ‘mixed scaling’ a new method that overcomes these issues and preserves the landscape features that are identifiable at different scales. The new method also extends the Gaussian pyramid by introducing additional intermediate scales. This minimizes the risk that the scales that are important for soil formation are not available in the model. In our extended implementation of the Gaussian pyramid, we tested four intermediate scales between any two consecutive octaves of the Gaussian pyramid and modelled the data with Deep Learning and Random Forests. We performed the experiments using three different datasets and show that mixed scaling with the extended Gaussian pyramid produced the best performing set of covariates and that modelling with Deep Learning produced the most accurate predictions, which on average were 4–7% more accurate compared to modelling with Random Forests.
Using the term "Open data" has become a bit of a fashion, but using it without clear specifications is misleading i.e. it can be considered just an empty phrase. Probably even worse is the term "Open Science" — can science be NOT open at all? Are we reinventing something that should be obvious from start? This guide tries to clarify some key aspects of Open Data, Open Source Software and Crowdsourcing using examples of projects and business. It aims at helping you understand and appreciate complexity of Open Data, Open Source software and Open Access publications. It was specifically written for producers and users of environmental data, however, the guide will likely be useful to any data producers and user.
Compilation and harmonization of environmental covariates for PSM in Canada Developed for: Agriculture and Agri-Food Canada/Agriculture et Agroalimentaire Canada. Layers (see: CanSIS explanation of names): Landsata cloud free images, FAPAR images, Prepared by: Robert A. MacMillan (bobmacm@gmail.com; LandMapper Environmental Solutions Inc.) and Tom Hengl (EnvirometriX Ltd)
Compilation and harmonization of environmental covariates for PSM in Canada Developed for: Agriculture and Agri-Food Canada/Agriculture et Agroalimentaire Canada. Layers (see: CanSIS explanation of names): DTM and DTM derivatives at 100 m (200 m, 400 m and 800 m) Downloaded precipitation and snow probability maps, Land cover and admin units, Prepared by: Robert A. MacMillan (bobmacm@gmail.com; LandMapper Environmental Solutions Inc.) and Tom Hengl (EnvirometriX Ltd)
We present a contextual spatial modelling (CSM) framework, as a methodology for multiscale, hierarchical mapping and analysis. The aim is to propose and evaluate a practical method that can account for the complex interactions of environmental covariates across multiple scales and their influence on soil formation. Here we derived common terrain attributes from multiscale versions of a DEM based on up-sampled octaves of the Gaussian pyramid. Because the CSM approach is based on a relatively small set of scales and terrain attributes it is efficient, and depending on the regression algorithm and the covariates used in the modelling, the results can be interpreted in terms of soil formation. Cross-validation coefficient of determination modelling (R2), for predictions of clay and silt increased from 0.38 and 0.16 when using the covariates derived at the original DEM resolution to 0.68 and 0.63, respectively, when using CSM. These results are similar to those achieved with the hyperscale covariates of ConMap and ConStat. As with these hyperscale covariates, the multiscale covariates derived from the Gaussian scale space in CSM capture the observed spatial dependencies and interactions of the landscape and soil. However, some advantages of CSM approach compared to ConMap and ConStat are i) a reduced set of scales that still manage to represent the entire extent of the range of scales, ii) a reduced set of attributes at each scale, iii) more efficient computation, and iv) better interpretability of the important covariates used in the modelling and thus of the factors that affect soil formation.
This paper describes the technical development and accuracy assessment of the most recent and improved version of the SoilGrids system at 250m resolution (June 2016 update). SoilGrids provides global predictions for standard numeric soil properties (organic carbon, bulk density, Cation Exchange Capacity (CEC), pH, soil texture fractions and coarse fragments) at seven standard depths (0, 5, 15, 30, 60, 100 and 200 cm), in addition to predictions of depth to bedrock and distribution of soil classes based on the World Reference Base (WRB) and USDA classification systems (ca. 280 raster layers in total). Predictions were based on ca. 150,000 soil profiles used for training and a stack of 158 remote sensing-based soil covariates (primarily derived from MODIS land products, SRTM DEM derivatives, climatic images and global landform and lithology maps), which were used to fit an ensemble of machine learning methods-random forest and gradient boosting and/or multinomial logistic regression-as implemented in the R packages ranger, xgboost, nnet and caret. The results of 10-fold cross-validation show that the ensemble models explain between 56% (coarse fragments) and 83% (pH) of variation with an overall average of 61%. Improvements in the relative accuracy considering the amount of variation explained, in comparison to the previous version of SoilGrids at 1 km spatial resolution, range from 60 to 230%. Improvements can be attributed to: (1) the use of machine learning instead of linear regression, (2) to considerable investments in preparing finer resolution covariate layers and (3) to insertion of additional soil profiles. Further development of SoilGrids could include refinement of methods to incorporate input uncertainties and derivation of posterior probability distributions (per pixel), and further automation of spatial modeling so that soil maps can be generated for potentially hundreds of soil variables. Another area of future research is the development of methods for multiscale merging of SoilGrids predictions with local and/or national gridded soil products (e.g. up to 50 m spatial resolution) so that increasingly more accurate, complete and consistent global soil information can be produced. SoilGrids are available under the Open Data Base License.
80% of arable land in Africa has low soil fertility and suffers from physical soil problems. Additionally, significant amounts of nutrients are lost every year due to unsustainable soil management practices. This is partially the result of insufficient use of soil management knowledge. To help bridge the soil information gap in Africa, the Africa Soil Information Service (AfSIS) project was established in 2008. Over the period 2008-2014, the AfSIS project compiled two point data sets: the Africa Soil Profiles (legacy) database and the AfSIS Sentinel Site database. These data sets contain over 28 thousand sampling locations and represent the most comprehensive soil sample data sets of the African continent to date. Utilizing these point data sets in combination with a large number of covariates, we have generated a series of spatial predictions of soil properties relevant to the agricultural management-organic carbon, pH, sand, silt and clay fractions, bulk density, cation-exchange capacity, total nitrogen, exchangeable acidity, Al content and exchangeable bases (Ca, K, Mg, Na). We specifically investigate differences between two predictive approaches: random forests and linear regression. Results of 5-fold cross-validation demonstrate that the random forests algorithm consistently outperforms the linear regression algorithm, with average decreases of 15-75% in Root Mean Squared Error (RMSE) across soil properties and depths. Fitting and running random forests models takes an order of magnitude more time and the modelling success is sensitive to artifacts in the input data, but as long as quality-controlled point data are provided, an increase in soil mapping accuracy can be expected. Results also indicate that globally predicted soil classes (USDA Soil Taxonomy, especially Alfisols and Mollisols) help improve continental scale soil property mapping, and are among the most important predictors. This indicates a promising potential for transferring pedological knowledge from data rich countries to countries with limited soil data.
SoilGrids is a collection of updatable soil property and class maps of the world produced using state-of-the-art model-based statistical methods: 3D regression with splines for continuous soil properties and multinomial logistic regression for soil classes.
SoilGrids1km is a collection of updatable soil property and class maps of the world at a relatively coarse resolution of 1 km produced using state-of-the-art model-based statistical methods: 3D regression with splines for continuous soil properties and multinomial logistic regression for soil classes.
For the North American node of the GlobalSoilMap.net project, a pilot study area was identified in SW Manitoba and adjacent North Dakota to develop and test methods to create GlobalSoilMap. net-compliant gridded soil map products. To this end, for the Canadian portion of the study area, we implemented and evaluated several methods to predict soil properties utilizing legacy soil maps at scales from 1: 20,000 to 1: 40,000 and soil series records from the Canadian Soil Information System. We then validated these predictions against a relatively small set of pedon data collected during the period of active soil survey (1970's and 1980's). Our first method was to simply use weighted average values from polygon map unit components to produce what we called GSM v0.5. The second approach to polygon disaggregation used finer resolution classifications of landform elements based on expert knowledge and heuristic rules. We hypothesized that disaggregation of the area-class soil polygon maps using a knowledge-based fuzzy classification of a fixed set of landform classes based on 90 m SRTM DEM data was one of the few feasible and practical options for producing predictions of continuous variation in soil properties for this area (and indeed for most of Canada) according to GlobalSoilMap. net specifications. The prediction efficiency of the fuzzy classification disaggregation approach, as measured by RSME values against pedon measurements, exceeded that achieved by simple polygon map unit component averaging for only some attributes (soil organic C and % clay) while for other attributes (i.e. pH) it was markedly less effective.