Given the profound impact of wildfires on human society and the natural environment, as well as the challenges faced by traditional supervised learning methods when sample sizes are insufficient, this study aims to evaluate the reliability of model-based and model-free semi-supervised learning (SSL) approaches in generating pseudo-samples for wildfire susceptibility mapping. The key distinction between these SSL approaches lie in their reliance on specific predictive models, such as random forest, support vector machine, or logistic regression, during pseudo-sample generation. In this study, we employed two semi-supervised learning methods: self-training (representing the model-based approach) and label propagation (representing the model-free approach) to create wildfire susceptibility maps in Wanzai County, Jiangxi Province, China. The results show that self-training combined with random forest exhibited optimal performance in the reliability evaluation of pseudo-sample quality, achieving an overall accuracy (OA) of 80.1%. In addition, label propagation also demonstrated high reliability with 77.5% of the OA. Further indirect reliability evaluation confirmed that integrating pseudo-samples into the original sample set enhances the accuracy of wildfire susceptibility mapping, particularly when the number of pseudo-samples is limited to 1,000 or fewer. Moreover, this study explores the motivations behind the reliability evaluation of pseudo-samples, potential efficacy mechanisms, and the applicability.
Developing habitat selection models is crucial for predicting the potential distribution of rare and endangered insects; however, current modeling approaches largely ignore absolute abundance data from sampling records. Integrating absolute abundance as sample weights involves two contrasting ecological hypotheses: one assumes high abundance reflects superior habitat quality (direct weighting), while the other attributes high abundance to spatially biased sampling effort (inverse weighting). Both hypotheses are plausible for Cheirotonus jansoni (Jordan, 1898), a highly phototactic endangered species, complicating the selection of an appropriate weighting strategy. This study evaluated four weighting strategies (unweighted, linear, inverse, and inverse square) derived from these hypotheses. Results indicated that the sampling effort correction strategy (inverse weighting) outperformed habitat-quality-driven strategies (linear weighting) in habitat suitability mapping, achieving a higher AUC (93.41) compared to the linear weighting (91.32). Notably, the linear weighting strategy performed worse than the unweighted baseline, yielding a lower AUC (91.32) than the baseline’s 93.11. These findings provide methodological guidance for modeling the distribution of phototactic insects and offer scientific support for the in situ conservation, ex situ site selection, and supplementary field surveys of C. jansoni.
Spatial prediction of environmental suitability for vector-borne disease transmission is crucial for public health, with spatial connectivity being a critical factor. However, existing models often face a trade-off between model interpretability and the accurate representation of spatial connectivity. Simpler, interpretable models tend to oversimplify connectivity, while more advanced models can lack interpretability and often rely on difficult-to-obtain fine-grained mobility data. To address this challenge, this paper proposes an integrated framework (GES-SC) for the spatial prediction of environmental suitability for vector-borne disease transmission, which combines Geographical Environmental Similarity (GES) with quantified Spatial Connectivity (SC). The method operates on the principle that environmentally similar and highly spatially connected locations have similar transmission potential. It quantifies connectivity using available road network data and distance decay, providing an effective alternative to direct mobility data. A case study in Guangzhou, China, demonstrates that integrating this quantified spatial connectivity enhances prediction accuracy. Compared to Geographically Weighted Regression, Bayesian Conditional Autoregressive, XGBoost, and a Simple Neural Network, the GES-SC method performed better in cross-validation, achieving lower RMSE and MAE, and a higher R2. The proposed framework effectively addresses the accuracy-interpretability trade-off, improving spatial epidemiological prediction accuracy and providing an uncertainty measure for public health decisions.
Soil spatial information from digital soil mapping (DSM) inherently contains uncertainty due to uncertainty stemming from both input data (samples and covariate data) and from DSM models. This overview highlights the major advancements in quantifying uncertainty within these elements, as well as efforts to integrate these uncertainties to characterize the overall uncertainty in final DSM products. It also discusses research on how quantified uncertainty can be utilized. Three key areas have been identified as immediate challenges for the research community in elevating DSM uncertainty as a critical component of geospatial analysis and decision-making: 1) The spatial nature and the evolving availability of geospatial data for uncertainty quantification in DSM; 2) The assessment of impacts and risks that DSM uncertainty poses to decision-making; and 3) The standardized provision of uncertainty associated with DSM-derived soil spatial data. Among these, the standard provision of uncertainty associated with DSM products is considered the most urgent one.
Forest fires threaten global ecosystems, socio-economic structures, and public safety. Accurately assessing forest fire susceptibility is critical for effective environmental management. Supervised learning methods dominate this assessment, relying on a substantial dataset of forest fire occurrences for model training. However, obtaining precise forest fire location data remains challenging. To address this issue, semi-supervised learning emerges as a viable solution, leveraging both a limited set of collected samples and unlabeled data containing environmental factors for training. Our study employed the transductive support vector machine (TSVM), a key semi-supervised learning method, to assess forest fire susceptibility in scenarios with limited samples. We conducted a comparative analysis, evaluating its performance against widely used supervised learning methods. The assessment area for forest fire susceptibility lies in Dayu County, Jiangxi Province, China, renowned for its vast forest cover and frequent fire incidents. We analyzed and generated maps depicting forest fire susceptibility, evaluating prediction accuracies for both supervised and semi-supervised learning methods across various small sample scenarios (e.g., 4, 8, 12, 16, 20, 24, 28, and 32 samples). Our findings indicate that TSVM exhibits superior prediction accuracy compared to supervised learning with limited samples, yielding more plausible forest fire susceptibility maps. For instance, at sample sizes of 4, 16, and 28, TSVM achieves prediction accuracies of approximately 0.8037, 0.9257, and 0.9583, respectively. In contrast, random forests, the top performers in supervised learning, demonstrate accuracies of approximately 0.7424, 0.8916, and 0.9431, respectively, for the same small sample sizes. Additionally, we discussed three key aspects: TSVM parameter configuration, the impact of unlabeled sample size, and performance within typical sample sizes. Our findings support semi-supervised learning as a promising approach compared to supervised learning for forest fire susceptibility assessment and mapping, particularly in scenarios with small sample sizes.
Soil moisture (SM) is a crucial environmental variable, and it plays an important role in energy and water cycles. SM data retrieval based on microwave satellite remote sensing has garnered significant attention due to its spatial continuity, wide observational coverage, and relatively low cost. Validating the accuracy of satellite remote sensing SM products is a critical step in enhancing data credibility, which plays a vital role in ensuring the effective application of satellite remote sensing data across various fields. Firstly, this study focused on Henan Province and evaluated the accuracy of the SMAP Enhanced L3 Radiometer Global and Polar Grid Daily 9 km EASE-Grid Soil Moisture (SPL3SMP_E) product along with its application in agriculture. The evaluation was based on in situ SM data from 55 stations in Henan Province. The assessment metrics used in this study include mean difference (MD), root mean square error (RMSE), unbiased root mean square error (ubRMSE), and the Pearson correlation coefficient (R). The time span of this study is from 2017 to 2020. The evaluation results indicated that the SPL3SMP_E soil moisture product performs well, as reflected by an ubRMSE value of 0.045 (m3/m3), which was relatively close to the product’s design accuracy of 0.04 (m3/m3). Moreover, the accuracy of the product was unaffected by temporal factors, but the product exhibited strong spatial aggregation, which was closely related to land use types. Then, this study explored the response of the SPL3SMP_E product to irrigation signals. The precipitation and irrigation data from Henan Province were employed to investigate the response of the SPL3SMP_E soil moisture product to irrigation. Our findings revealed that the SPL3SMP_E soil moisture product was capable of capturing over 70% of irrigation events in the study area, indicating its high sensitivity to irrigation signals in this region. In this study, the SPL3SMP_E product was also employed for monitoring agricultural drought in Henan Province. The findings revealed that the collaborative use of the SPL3SMP_E soil moisture product and machine learning algorithms proves highly effective in monitoring significant drought events. Furthermore, the integration of multiple indices demonstrated a notable enhancement in the accuracy of drought monitoring. Such an evaluation holds significant implications for the effective application of satellite remote sensing SM data in agriculture and other domains.
Sampling design can significantly reduce the uncertainty in geospatial predictions. In this paper, we developed an adaptive uncertainty-guided stepwise sampling (AUGSS) method to select sampling locations to supplement existing legacy sample points whose representation should be improved. The proposed method selects supplemental samples in a stepwise manner as guided by an objective function with two weighted sub-objectives. One reduces the area with high prediction uncertainty, and the other minimizes the overall prediction uncertainty for the entire area. The method takes an adaptive approach to adjust weights for the two sub-objectives and to tune an uncertainty threshold controlling whether a location can be reliably predicted during the sampling procedure. A case study on soil property prediction shows that AUGSS outperforms the stratified random sampling (SRS) and the non-adaptive uncertainty guided sampling method (UGSS) in terms of RMSE and Lin's concordance correlation coefficient with different sample sizes. This study shows that the AUGSS method offers a potential for effectively adding supplemental samples to existing samples which are insufficient for spatial prediction. The adaptive strategy guided by predicted uncertainty provides an efficient support to improve the spatial pattern of samples, which plays a key role in the result accuracy of geospatial predictive mapping.
Global climate change is a serious threat to food and energy security. Crop growth modelling is an important tool for simulating crop food production and assisting in decision making. Planting date is one of the important model parameters. Larger-scale spatial distribution with high accuracy for planting dates is essential for the widespread application of crop growth models. In this study, a planting date prediction method based on environmental similarity was developed in accordance with the third law of geography. Spring maize planting date observations from 124 agricultural meteorological experiment stations in China over the years 1992–2010 were used as the data source. Samples spanning from 1992 to 2009 were allocated as training data, while samples from 2010 constituted the independent validation set. The results indicated that the root mean square error (RMSE) for spring maize planting date based on environmental similarity was 10 days, which is better than that of multiple regression analysis (RMSE = 13 days) in 2010. Additionally, when applied at varying scales, the accuracy of national-scale prediction was better than that of regional-scale prediction in areas with large differences in planting dates. Consequently, the method based on environmental similarity can effectively and accurately estimate planting date parameters at multiple scales and provide reasonable parameter support for large-scale crop growth modelling.
Multi-scale habitat selection modeling (HSM) has garnered attention due to its ability to incorporate scale dependence of species. The key of multi-scale HSM is to select the appropriate combination of scales for different resources or environmental conditions, and then construct a set of multi-scale environmental covariates as the features of HSM. However, the existing scale selection methods do not determine the combination of scales under a unified model. In this study, a combinatorial optimization approach is proposed. We regard the combination of different scales as a search space, and use a heuristic optimization algorithm to search for the best-fitting model to determine the optimal scales for each resource or environmental condition. In a case study conducted in Yancheng National Nature Reserve, the proposed approach is applied to model the habitat selection of the endangered red-crowned crane. We compare the proposed method with single-scale, random-scale and other multiscale approaches. The results show that the combination of scales selected based on the proposed method obtained the best accuracy in the spatial prediction of habitat suitability with a test AUC of 0.865 for the daytime scenario and 0.932 for the nighttime scenario. Moreover, the selected scales are utilized to generate response curves, providing suggestions for habitat restoration and management of the red-crowned crane population in the nature reserve.
Over the past decades, conventional soil maps of various scales have been produced and become available in digital form. Efforts have been made to update these maps through various data mining methods to provide more detailed and precise information on soil spatial patterns. Key questions that remain unclear are: (1) How does the accuracy of legacy soil maps impact the update results; (2) Is the accuracy of inferred soil maps always improved regardless of the accuracy of the legacy maps. The current study aims to investigate these questions. Two noise production simulation methods were developed to simulate errors caused by inclusion and boundary displacement in the conventional maps, to generate a series of source maps with different accuracies and spatial patterns. Moreover, the impacts of two training sample selection methods and three data mining models on the accuracies and spatial patterns of the inferred soil maps were also evaluated. A case study was conducted in a small region, Raffelson study area, a typical ridge and valley terrain in La Crosse County, Wisconsin, USA. Results indicated that if the accuracies of the source soil maps ranged from 35% to 75%, the inferred soil map accuracies would be improved. These findings have important implications for updating conventional soil maps through data mining methods and understanding the situation in which the method is effective.
Numerous machine learning models have been developed for constructing the relationship between soil classes or properties and its environmental covariates in digital soil mapping (DSM). Most machine learning models are trained with a supervised learning (SL) method based on training samples. However, the collected sample data is often limited in practice due to that field sampling is expensive and time-consuming. The insufficient samples may limit the learning ability of the model to a large extent. Semi-supervised machine learning, a new machine learning paradigm that makes use of both unsampled data and a small amount of sampled data in the learning process, can be a potential effective method for DSM. In this study, we present a self-training semi-supervised learning (SSL) method for DSM. Different with the SL method for machine learning models, the SSL method not only utilizes the sampled locations but also the abundant environmental covariate information at the unvisited locations. Its basic idea is to iteratively enlarge the training data set by adding the unsampled points with high prediction confidence from the unvisited locations until a stopping criterion reached. The proposed SSL method was applied in machine learning models for predicting soil classes in Heshan Farm of Nenjiang County in Heilongjiang Province, China. Three machine learning models, including multinomial logistic regression (MLR), k-nearest neighbor (KNN) and random forest (RF), were selected to evaluate the efficiency of the SSL method. The entropy threshold was an important parameter in the SSL method, and a sensitivity analysis on this parameter was conducted with using a series of entropy thresholds. The SSL method was compared with the SL method for the three machine learning models for soil prediction. A cross-validation was employed to evaluate the accuracy of the predicted soil class maps generated based on each method. The results showed that the prediction accuracies (the proportion of the correctly predicted samples over the total number of validation samples) of the SSL method were higher than those of the SL method for MLR, KNN, and RF by 5.9%, 12.2%, and 6.0%, respectively. RF-SSL was the most accurate model in the study area, followed by KNN-SSL. Meanwhile, the self-training SSL method for the KNN model had the largest improvement comparing with the other two models. Furthermore, the predicted soil maps using the SSL method showed a more reasonable spatial variation pattern of soil classes. In the study area, a suitable value of the entropy threshold was 0.8 similar to 1.0. We concluded that the SSL method improved the soil prediction accuracy compared with the SL method when applying machine learning models for DSM, and thus is a potential efficient method for DSM with limit sample data.
Monitoring the work cycles of earthmoving excavators is an important aspect of construction productivity assessment. Currently, the most advanced method for the recognition of work cycles is the “Stretching-Bending” Sequential Pattern (SBSP), which is based on fixed-carrier video monitoring (FC-SBSP). However, the application of this method presupposes the availability of preconstructed installation carriers to act as a surveillance camera as well as installed and commissioned surveillance systems that work in tandem with them. Obviously, this method is difficult to apply to projects with no conditions for a monitoring camera installation or which have a short construction time. This highlights the potential application of Unmanned Aerial Vehicle (UAV) remote sensing, which is flexible and mobile. Unfortunately, few studies have been conducted on the application of UAV remote sensing for the work cycle monitoring of earthmoving excavators. This research is necessary because the use of UAV remote sensing for monitoring the work cycles of earthmoving excavators can improve construction productivity and save time and costs, especially in post-disaster reconstruction projects involving harsh construction environments, and emergency projects with short construction periods. In addition, the challenges posed by UAV shaking may have to be taken into account when using the SBSP for UAV remote sensing. To this end, this study used application experiments in which stabilization processing of UAV video data was performed for UAV shaking. The application experimental results show that the work cycle performance of UAV remote-sensing-based SBSP (UAV-SBSP) for UAV video data was 2.45% and 5.36% lower in terms of precision and recall, respectively, without stabilization processing than after stabilization processing. Comparative experiments were also designed to investigate the applicability of the SBSP oriented toward UAV remote sensing. Comparative experimental results show that the same level of performance was obtained for the recognition of work cycles with the UAV-SBSP as compared with the FC-SBSP, demonstrating the good applicability of this method. Therefore, the results of this study show that UAV remote sensing enables effective monitoring of earthmoving excavator work cycles in construction sites where monitoring cameras are not available for installation, and it can be used as an alternative technology to fixed-carrier video monitoring for onsite proximity monitoring.
Counting the number of work cycles per unit of time of earthmoving excavators is essential in order to calculate their productivity in earthmoving projects. The existing methods based on computer vision (CV) find it difficult to recognize the work cycles of earthmoving excavators effectively in long video sequences. Even the most advanced sequential pattern-based approach finds recognition difficult because it has to discern many atomic actions with a similar visual appearance. In this paper, we combine atomic actions with a similar visual appearance to build a stretching–bending sequential pattern (SBSP) containing only “Stretching” and “Bending” atomic actions. These two atomic actions are recognized using a deep learning-based single-shot detector (SSD). The intersection over union (IOU) is used to associate atomic actions to recognize the work cycle. In addition, we consider the impact of reality factors (such as driver misoperation) on work cycle recognition, which has been neglected in existing studies. We propose to use the time required to transform “Stretching” to “Bending” in the work cycle to filter out abnormal work cycles caused by driver misoperation. A case study is used to evaluate the proposed method. The results show that SBSP can effectively recognize the work cycles of earthmoving excavators in real time in long video sequences and has the ability to calculate the productivity of earthmoving excavators accurately.
Effective conservation measures largely depend on knowledge of habitat selection of target species. Little is known about the scale characteristics and temporal rhythm of habitat selection of the endangered red-crowned crane, limiting the habitat conservation. Here, two red-crowned cranes were tracked with Global position system (GPS) for two years in Yancheng National Nature Reserve (YNNR). A multiscale approach was developed to identify the spatiotemporal pattern of habitat selection of red-crowned cranes. The results revealed that Red-crowned cranes preferred to select Scirpus mariqueter, ponds, Suaeda salsa, and Phragmites australis, and avoid Spartina alterniflora. In each season, habitat selection ratio for Scirpus mariqueter and ponds was the highest during the day and night, respectively. Further multiscale analysis showed that the percent coverage of Scirpus mariqueter at the 200-m to 500-m scale was the most important predictor for all habitat selection modeling, emphasizing the importance of restoring a large area of Scirpus mariqueter habitat for red-crowned crane population restoration. Additionally, other variables affect habitat selection at different scales, and their contributions vary with seasonal and circadian rhythm. Furthermore, habitat suitability was mapped to provide a direct basis for habitat management. The suitable area of daytime and nighttime habitat accounted for 5.4%–19.0% and 4.6%–10.2% of the study area, respectively, implying the urgency of restoration. The study highlighted the scale and temporal rhythms of habitat selection for various endangered species that depend on small habitats. The proposed multiscale approach applies to the restoration and management of habitats of various endangered species.
This study investigates sampling design for mapping soil classes based on multiple environmental features associated with the soil classes. Two types of sampling design for calibrating the prediction models are compared: conditioned Latin hypercube sampling (CLHS) and feature space coverage sampling (FSCS). Simple random sampling (SRS), which does not utilize the environmental features, is added as a reference design. The sample sizes used are 20, 30, 40, 50, 75, and 100 points, and at each sample size 100 sample sets were drawn using each of the three types of design. Each of these sample sets was then used to calibrate three prediction models: random forest (RF), individual predictive soil mapping (iPSM), and multinomial logistic regression (MLR). These sampling designs were compared based on the overall accuracy of predicted soil class maps obtained by these three prediction methods. The comparison was conducted in two study areas: Ammertal (Germany) and Raffelson (USA). For each of these two areas a detailed legacy soil class map is available. These soil class maps were used as references in a simulation study for the comparison. Results of both study areas show that on average FSCS outperforms CLHS and SRS for all three prediction methods. The difference in estimated medians of overall accuracy with CLHS and SRS was marginal. Moreover, the variation in overall accuracy among sample sets of the same size was considerably smaller for FSCS than that for CLHS. These results in the two study areas suggest that FSCS is a more effective sampling design.
•Our results suggested a rapid decline in Red-crowned habitat.•Human disturbance and land use and land cover together shaped the historical dynamics of habitat.•Integration of the Maxent and landscape ecology theory provides a reliable method for habitat restoration.
In addition to soil samples, conventional soil maps, and experienced soil surveyors, text about soils(e.g., soil survey reports) is an important potential data source for extracting soil–environment relationships. Considering that the words describing soil–environment relationships are often mixed with unrelated words, the first step is to extract the needed words and organize them in a structured way. This paper applies natural language processing(NLP) techniques to automatically extract and structure information from soil survey reports regarding soil–environment relationships. The method includes two steps:(1) construction of a knowledge frame and(2) information extraction using either a rule-based method or a statistic-based method for different types of information. For uniformly written text information, the rule-based approach was used to extract information. These types of variables include slope, elevation, accumulated temperature, annual mean temperature, annual precipitation, and frost-free period. For information contained in text written in diverse styles, the statistic-based method was adopted. These types of variables include landform and parent material. The soil species of China soil survey reports were selected as the experimental dataset. Precision(P), recall(R), and F1-measure(F1) were used to evaluate the performances of the method. For the rule-based method, the P values were 1, the R values were above 92%, and the F1 values were above 96% for all the involved variables. For the method based on the conditional random fields(CRFs), the P, R and F1 values for the parent material were, respectively, 84.15, 83.13, and 83.64%; the values for landform were 88.33, 76.81, and 82.17%, respectively. To explore the impact of text types on the performance of the CRFs-based method, CRFs models were trained and validated separately by the descriptive texts of soil types and typical profiles. For parent material, the maximum F1 value for the descriptive text of soil types was 90.7%, while the maximum F1 value for the descriptive text of soil profiles was only 75%. For landform, the maximum F1 value for the descriptive text of soil types was 85.33%, which was similar to that of the descriptive text of soil profiles(i.e., 85.71%). These results suggest that NLP techniques are effective for the extraction and structuration of soil–environment relationship information from a text data source.
Field sampling is an essential step for digital soil mapping and various sampling strategies have been designed for achieving desirable mapping results. Some unpredictable complex circumstances in the field, however, often prevent some samples from being collected based on the pre-designed sampling strategies. Such circumstances include inaccessibility of some locations, change of land surface types, and heavily-disturbed soil at some locations, among others. This may result in the missing of some essential samples, which could impact the quality of digital soil mapping. Previous studies have attempted to design alternative samples for the selected ones beforehand to address this issue. It cannot solve the problem completely as those pre-designed alternative samples could also be inaccessible. In this paper, we propose a dynamic method to recommend alternative samples for those unavailable samples in the field. The identification of alternative samples is based on the environmental similarity between an unavailable soil sample and its alternative candidates, as well as the spatial accessibility of these candidates. For the convenience of fieldwork, the proposed method was implemented to be a mobile application on the Android platform. A simulated soil sampling study in Xuancheng county, Anhui province of China was used to evaluate its performance. From a sample set for the study are, 30 samples were assumed to be inaccessible. For each of them, an alternative soil sample from the set was recommended using the proposed method. A deviation analysis of silt and sand content at the depth of 20 ~ 40 cm between the soil samples and their alternatives shows that the deviation on silt content is less than 20% for half of the soil samples. A larger deviation on sand content might be attribute to the limited alternative candidates in this virtual experiment. In a second experiment, we randomly selected a number of existing soil samples and replaced them with their corresponding alternative soil samples. This created 1000 hybrid sample sets. Each hybrid sample set was then used for digital soil mapping with iPSM. An evaluation using 59 independent soil samples indicated that the RMSE and MAE with the hybrid sample sets were close to that with the original sample set. The proposed method proved to be able to recommend effective alternative samples for those unavailable samples in the field.
Field sampling is an important way of collecting soil information for the modeling and evaluation steps during digital soil mapping (DSM). However, some predesigned samples may not be accessible in the field due to natural or anthropogenic reasons. Simply abandoning the inaccessible samples or casually selecting substitutes from other locations may affect the quality of the corresponding DSM. To address this issue, we propose a new method of dynamically recommending substitute locations for inaccessible samples, which was implemented in a prototype system on a smart phone platform. The proposed method takes into concern the original sampling strategy and recommends substitute sample locations based on a measure of suitability index. The suitability index is calculated to incorporate a substitutive degree as well as the sampling cost involved. The substitutive degree depicts to what extent a substitute location may replace the original sample in the context of soil mapping, while the sampling cost characterizes the travel expense to the substitute location following the overall fieldwork route arrangements. The proposed method currently supports four commonly used sampling strategies, i.e., simple random sampling, stratified random sampling, grid sampling, and purposive sampling based on environmental similarity. Two substitute sampling scenarios, instant sampling and subsequent sampling, are considered by the proposed method, to adapt to surveyors’ actual field sampling route arrangements when estimating the accessibility and sampling cost of potential substitute locations. Monte Carlo simulation experiments in a study area (about 5800 km2) located in Anhui province of China were conducted to use the proposed method to recommend substitute locations for two modeling sample sets designed based on purposive sampling strategy and stratified random sampling strategy respectively (59 points for each set) from other 224 previously obtained samples. Experimental results evaluated based on 57 independent evaluation samples showed that the proposed method was able to recommend substitute locations without affecting the performance of DSM, when less than 10% samples were replaced by substitute samples. A subsequent sampling scenario was revealed to incur lower sampling cost than an instant sampling scenario.