Using the example of data from the article by E.A. Dmitriev et al., the estimation of the required number of soil samples to assess the SOC content in the forest biogeocenosis during monitoring studies is considered. Primary data on SOC content were obtained in the spruce forest at 166 points in layers 0-10, 10-20, and 20-30 cm af er removal of the litter. T e sampling was carried out at the nodes of a regular grid of equilateral triangles with 1 m side within a regular hexagon with a side of 7 m. T e SOC content was determined by the Tyurin method. T e original article presents statistics for three zones - near-stem, under-crown and inter-crown space. Spatial variation in all zones and at all depths is high, the coef cients of variation are about 50%. It is shown that the number of replicates required for estimating the average SOC content at a 95% conf dence level in the 0-10 cm layer is hundreds of samples and decreases to tens of samples in the 20-30 cm layer. Since the number of repetitions for testing hypotheses about the equality of means depends not only on the conf dence level, but also on the power of the criterion used, the required number of repetitions increases several times. Sampling with samples taken from the entire vertical layer of 0-30 cm and forming mixed samples from them reduces the number of required repetitions, however, careful observance of sample preparation, including primary mixing of samples, is required.
The FAO methodology within the Global Soil Nutrient and Nutrient Budget Maps (GSNmap) project was tested for the first time for mapping humus content with a spatial resolution of 250 meters per pixel in soils of the Russian Federation at the regional scale, using the Bryansk oblast as an example. The map was created in the R software environment using data from Agrochemical Service and remote sensing, global databases and soil maps. The centroids of the sites from which the composite samples were taken by Agrochemical Service were selected as sampling points. The set of predictors available under the FAO project was expanded by additional data, including soil maps and maps of soil-forming rocks. The importance of the predictors was assessed using the Boruta algorithm, which is usually used as an initial stage for a random forest. The model was created using the caret package with the quantile regression forest method. The modeling efficiency coefficient (MEC) was 55
The uncertainty sources in the assessment of organic carbon stocks were studied in a layer of 0–30 cm within the sampling site (100 × 100 m) set on soddy-podzolic cultivated soil (Albic Glossic Retisol (Aric, Loamic, Ochric)). In the experiment, two sampling methods were used – the classic sampling in pits, by 10-cm layers, and sampling with an auger to a depth of 0–30 cm. The soil bulk density was determined by the Kachinsky method and the carbon content was analyzed by the Tyurin method. Some samples were additionally analyzed at the Bryansk State Agrarian University. The uncertainties associated with natural variation, sample preparation and the analytical process proper were estimated. The analytical uncertainty of bulk density measurement did not depend on the sampling depth under the experimental conditions and amounted to about 6%. The analytical error of Tyurin’s method did not differ in two laboratories. Its contribution reached 5–9% of the total variation in soil organic carbon content at the test plot. The uncertainty of sample preparation ranged from 11 to 26%, natural variation—from 49 to 68% to the total variance, respectively. Determination of carbon content in the samples taken with an auger from the depth of 0–30 cm is preferable than layer-by-layer sampling as the number of intermediate operations is fewer, and the obtained results are comparable.
T e spatial variability of the content of granulometric fraction d>0,25 mm in the arable layer of soddy-podzolic cultivated soil on an area of 18 hectares was studied. T e soil formed in the loess-like loam, underlain by moraine deposits. Samples were taken from the upper (0-10 cm) and lower parts of the arable layer (10-20 cm) following a random-stratif ed sampling scheme. T e total number of samples was 350 cores. T e average fraction content was about 22%, the coef cient of variation was 39-41%, the distributions were approximated by a logarithmically normal distribution. T e Spearman correlation coef cient between values at dif erent depths equaled to 0,87. T e spatial distribution cartograms were constructed by the ordinary kriging method using a spherical variogram model. It is shown that preliminary censoring of high sample values (quantile 0,95) gives better results when constructing a cartogram than removing a linear spatial trend and logarithm of the original data. Spatial structures with reduced and increased values were found on the site, the average linear sizes of which was about 100 m. Presumably, they were associated with the heterogeneity of soil-forming material.
This article is a review of the scientific problems that the Honored Professor of the Moscow State University Evgeny Anatolyevich Dmitriev worked on. The main area of his scientific interests was the genesis of soil and land cover data. Numerous works carried out under his supervision convincingly demonstrated the influence of sampling methods on the results of determining certain soil properties. As a consequence, this influence also extends to the final conclusions, which can be highly distorted if the method of obtaining them is not taken into account. In an attempt to solve the question “What classifies soil classification?” he developed the concept of a “single soil,” which can be the object of soil classifications, not being directly a soil body, but a kind of standard element of sampling. Theoretical understanding of the influence of heterogeneity of soils and soil cover at all hierarchical levels on the features of its “life” is the subject of the second part of the scientific legacy of E.A. Dmitriev. He proposed the concept of soil bodies of different dimensions, showed the influence of the dimensional characteristics of these bodies on soil regimes and, in particular, on the water regime, and discussed the need to study soil regimes at different hierarchical levels. To fix the heterogeneity of the soil and soil cover, he developed special devices, with which studies were carried out on soddy-podzolic, gray forest, chestnut, and other soils.
The most common inaccuracies and errors in the application of statistical methods found in Russian publications on soil science are considered. When designating random variables and distribution parameters in Greek letters, it is necessary to designate those that refer to general populations, and Latin letters – to sampling ones. A detailed description of the experiment and what the replications relate to allows you to draw correct conclusions from the study. It is necessary to avoid pseudoreplication when results at closely located sampling points are considered as characteristics of soil variability over large distances. Expanding the list of descriptive statistics will allow you to use a specific study in meta-analysis. Calculating the confidence interval for the average using the Student's test at different significance levels expands the scope of possible values of the average, but this approach is justified only if the indicator does not differ too much from the normal distribution. When testing statistical hypotheses, it is necessary to pay attention not only to the level of significance, but also to the power of the criterion. The normality distribution hypothesis can be tested using various criteria. The success of applying the criterion depends not only on the validity of the null hypothesis (a truly normal distribution), but also on other reasons: on the sample size and on the alternatives for which the criterion tests the hypothesis. Any statement about the type of relationship between features based on the correlation coefficient (Pearson or Spearman) is meaningless without specifying the number of replicates, since it is the number of replicates that determines the significance of the difference between the correlation coefficient and zero. It is proposed that authors and reviewers pay closer attention to such errors.
Spatial variability of the arable horizon’s agrochemically valuable properties (pHKCl, hydrolytic acidity, base exchange capacity, degree of base saturation, humus content, content of mobile phosphorus and exchangeable potassium) has been evaluated. It is shown that division of the totality into partition subsets corresponding to classification units significantly reduces the variation degree of humus and physicochemical properties, having practically no effect on pH variability and the content of mobile phosphorus and exchangeable potassium. Discriminant analysis shows that the arable horizons of soddy medium-podzolic, grey forest, and dark grey forest soils are satisfactorily classified according to a given set of properties. Bogpodzolic and grey forest gleyed soils are poorly classified; light grey forest soils have an intermediate position, gravitating toward grey forest soils.
The analysis of soil maps for three districts of Bryansk oblast demonstrates that the rank distributions of polygon areas for low soil taxonomic units on these maps have a specific form that can be described as the splicing of several power distributions known as the Pareto laws. The form of the distributions is preserved in the three administrative districts, as well as in the individual soil-geographical areas, into which these districts may be divided. The analysis of variance performed for the logarithms of polygon sizes attests to the impacts of the genetic soil type, soil texture, and the specificity of parent material and underlying substrates on the size of the polygons. However, the ranking of the curves does not display any manifested aggregation of the predominant sizes of the polygons of certain taxonomic soil units.
Formica pratensis anthills, which occur in great numbers in the Bryansk Opolie fallow land areas covered with the grey-luvic phaeozems, change both the local microrelief and the soil properties. The soil pH KC L in their dome is higher than that in the surrounding soils by an average of 0.6 units. In the top 5-cm soil layer, increased contents of the organic matter and the mobile phosphorus and potassium are recorded, while the cation exchange capacity and the hydrolytic acidity are decreased. The natural radionuclide distribution under and outside the anthill is uniform. The technogenic 137 Cs in its maximum is found at a greater depth under the anthill. It may be explained by the ant pedoturbation activities.
Increasingly, soil surveys make use of a combination of legacy data, ancillary data and new field data. While combining the different sources of information, positional errors can play a large role. For example, the spatial discrepancy between remote sensing images and field data can depend on many factors, including the positioning accuracy of ground-based observations. The accuracy of GPS receivers for the territory of Russia is approximately 3–10 m. The aim of the study was to estimate the impact of sampling positioning accuracy on the relationship between soil organic carbon content and the infrared channel of the WorldView-2 satellite image and for mapping soil organic carbon contents in an agricultural field in the territory of the Bryansk Opolje (Russia). Intensive sampling of the topsoil took place. The positional accuracy was also measured and used to perturb the locations of the samples. The data were used to study: (i) the relationships between soil organic carbon and infrared reflectance, (ii) the variation in soil organic carbon through five different interpolation techniques, and (iii) the fraction of the fields with low soil organic matter contents. The study showed that the positional inaccuracies can have an important impact. The standardized methods to estimate the positional accuracy, perturb the locations and evaluate its impact seems to be an easy way to explore the quality of data.
Dynamics and composition of snow cover in the system of a geochemical landscape typical of southern taiga landscapes and the urbanized landscapes were studied during winter seasons of 2011 – 2016. It is shown that the maximum heights and water equivalent of snow cover in natural landscapes are usually found on flood plains and forest openings on the first terrace whereas minimum values are typical of forest ecosystems. The cluster analysis revealed that the chemical composition of snow in the MSU Arboretum and within landscapes in the vicinity of Moscow is closer unlike that of true urban landscapes.
Empirical Bayesian kriging (EBK) is a modern mapping method, which accounts for the uncertainty of parameter estimates in functions describing the changes in property variance with increasing the survey area (variograms). Cartograms plotted using ordinary kriging and EBK have been compared for the data on the content of organic carbon in an isolated land with agrogray soils (Greyzemic Phaeozems (Loamic, Aric)) located in the Bryansk Opol’e region. It is shown that the cartograms of EBK errors reveal the structure of the spatial variability of the property, which cannot be revealed by other methods. Thus, the EBK method can be recommended for revealing heterogeneities in disputable cases.
In the scheme of agrochemical surveys, sampling is one of the most costly stages. The use of a priori information about the distributions of properties (Bayesian approach) allows one to reduce the number of samples to be taken by 5–10% without deteriorating the accuracy of the estimation. The maintenance of regional soil-mapping databases is a necessary condition for applying the Bayesian approach.