Spatial processes observed in various fields, such as climate and environmental science, often occur on a large scale and demonstrate spatial nonstationarity. Fitting a Gaussian process with a nonstationary Mat\'ern covariance is challenging. Previous studies in the literature have tackled this challenge by employing spatial partitioning techniques to estimate the parameters that vary spatially in the covariance function. The selection of partitions is an important consideration, but it is often subjective and lacks a data-driven approach. To address this issue, in this study, we utilize the power of Convolutional Neural Networks (ConvNets) to derive subregions from the nonstationary data. We employ a selection mechanism to identify subregions that exhibit similar behavior to stationary fields. In order to distinguish between stationary and nonstationary random fields, we conducted training on ConvNet using various simulated data. These simulations are generated from Gaussian processes with Mat\'ern covariance models under a wide range of parameter settings, ensuring adequate representation of both stationary and nonstationary spatial data. We assess the performance of the proposed method with synthetic and real datasets at a large scale. The results revealed enhanced accuracy in parameter estimations when relying on ConvNet-based partition compared to traditional user-defined approaches.
In the spatial autologistic model, the dependence parameter is often assumed to be a single value. To construct a spatial autologistic model with spatial heterogeneity, we introduce additional covariance in the dependence parameter, and the proposed model is suitable for the data with binary responses where the spatial dependency pattern varies with space. Both the maximum pseudo-likelihood (MPL) method for parameter estimation and the Bayesian information criterion (BIC) for model selection are provided. The exponential consistency between the maximum likelihood estimator and the maximum block independent likelihood estimator (MBILE) is proved for a particular case. Simulation results show that the MPL algorithm achieves satisfactory performance in most cases, and the BIC algorithm is more suitable for model selection. We illustrate the application of our proposed model by fitting the Bur Oak presence data within the driftless area in the midwestern USA.
In the last few decades, the size of spatial and spatio-temporal datasets in many research areas has rapidly increased with the development of data collection technologies. As a result, classical statistical methods in spatial statistics are facing computational challenges. For example, the kriging predictor in geostatistics becomes prohibitive on traditional hardware architectures for large datasets as it requires high computing power and memory footprint when dealing with large dense matrix operations. Over the years, various approximation methods have been proposed to address such computational issues, however, the community lacks a holistic process to assess their approximation efficiency. To provide a fair assessment, in 2021, we organized the first competition on spatial statistics for large datasets, generated by our ExaGeoStat software, and asked participants to report the results of estimation and prediction. Thanks to its widely acknowledged success and at the request of many participants, we organized the second competition in 2022 focusing on predictions for more complex spatial and spatio-temporal processes, including univariate nonstationary spatial processes, univariate stationary space-time processes, and bivariate stationary spatial processes. In this paper, we describe in detail the data generation procedure and make the valuable datasets publicly available for a wider adoption. Then, we review the submitted methods from fourteen teams worldwide, analyze the competition outcomes, and assess the performance of each team.
Although the maximum likelihood estimation (MLE) for the uncertain discrete models has long been an academic interest, it has yet to be proposed in the literature. Thus, this study proposes the uncertain MLE for discrete models in the framework of the uncertainty theory, such as the uncertain logistic regression model. We also generalize the estimation proposed by Lio and Liu and obtain the uncertain MLE for non-linear continuous models, such as the uncertain Box-Cox regression model. Our proposed methods provide a useful tool for making inferences regarding non-linear data that is precisely or imprecisely observed, especially data based on degrees of belief, such as an expert's experimental data. We demonstrate our methodology by calculating proposed estimates and providing forecast values and confidence intervals for numerical examples. Moreover, we evaluate our proposed models via residual analysis and the cross-validation method. The study enriches the definition of the uncertain MLE, thus making it easy to construct estimation and prediction methods for general uncertainty models.
Given the computational challenges involved in calculating the maximum likelihood estimates for large spatial datasets, there has been significant interest in the research community regarding approximation methods for estimation and subsequent predictions. However, prior studies examining the evaluation of these methods have primarily focused on scenarios where the data are observed on a regular grid or originate from a uniform distribution of locations. Nevertheless, non-uniformly distributed locations are commonplace in fields like meteorology and ecology. Examples include gridded data with missing observations acquired through remote sensing techniques. To assess the reliability and effectiveness of cutting-edge approximation methods, we have initiated a competition focused on estimation and prediction for large spatial datasets with non-uniformly distributed locations. Participants were invited to employ their preferred methods to generate corresponding confidence and prediction intervals for synthetic datasets of varying sizes and spatial configurations. This competition serves as a valuable opportunity to benchmark and compare different approaches in a controlled setting. We evaluated the submissions from 11 different research teams worldwide. In summary, the Vecchia approximation and the fractional SPDE methods were among the best performers for estimation and prediction. Furthermore, the nearest neighbors Gaussian process and the multi-resolution approximation exhibited excellent performance in predictive tasks. These findings provide valuable guidance for selecting the most appropriate approximation methods based on specific data characteristics. Supplementary materials accompanying this paper appear online.
The asymmetry of residuals about the origin is a severe issue in estimating a Box-Cox transformed model. In the framework of uncertainty theory, there are such theoretical issues regarding the least-squares estimation (LSE) and maximum likelihood estimation (MLE) of the linear models after the Box-Cox transformation on the response variables. Heretofore, only weighting methods for least-squares analysis have been available. This article proposes an uncertain alternative Box-Cox model to alleviate the asymmetry of residuals and avoid λ tending to negative infinity for uncertain LSE or uncertain MLE. Such symmetry of residuals about the origin is reasonable in applications of experts’ experimental data. The parameter estimation method was given via a theorem, and the performance of our model was supported via numerical simulations. According to the numerical simulations, our proposed ‘alternative Box-Cox model’ can overcome the problems of a grossly underestimated lambda and the asymmetry of residuals. The estimated residuals neither deviated from zero nor changed unevenly, in clear contrast to the LSE and MLE for the uncertain Box-Cox model downward biased residuals. Thus, though the LSE and MLE are not applicable on the uncertain Box-Cox model, they fit the uncertain alternative Box-Cox model. Compared with the uncertain Box-Cox model, the issue of a systematically underestimated λ is not likely to occur in our uncertain alternative Box-Cox model. Both the LSE and MLE can be used directly without constructing a weighted estimation method, offering better performance in the asymmetry of residuals.
Due to the well-known computational showstopper of the exact Maximum Likelihood Estimation (MLE) for large geospatial observations, a variety of approximation methods have been proposed in the literature, which usually require tuning certain inputs. For example, the recently developed Tile Low-Rank approximation (TLR) method involves many tuning parameters, including numerical accuracy. To properly choose the tuning parameters, it is crucial to adopt a meaningful criterion for the assessment of the prediction efficiency with different inputs. Unfortunately, the most commonly-used Mean Square Prediction Error (MSPE) criterion cannot directly assess the loss of efficiency when the spatial covariance model is approximated. Though the Kullback–Leibler Divergence criterion can provide the information loss of the approximated model, it cannot give more detailed information that one may be interested in, e.g., the accuracy of the computed MSE. In this paper, we present three other criteria, the Mean Loss of Efficiency (MLOE), Mean Misspecification of the Mean Square Error (MMOM), and Root mean square MOM (RMOM), and show numerically that, in comparison with the common MSPE criterion and the Kullback–Leibler Divergence criterion, our criteria are more informative, and thus more adequate to assess the loss of the prediction efficiency by using the approximated or misspecified covariance models. Hence, our suggested criteria are more useful for the determination of tuning parameters for sophisticated approximation methods of spatial model fitting. To illustrate this, we investigate the trade-off between the execution time, estimation accuracy, and prediction efficiency for the TLR method with extensive simulation studies and suggest proper settings of the TLR tuning parameters. We then apply the TLR method to a large spatial dataset of soil moisture in the area of the Mississippi River basin, and compare the TLR with the Gaussian predictive process and the composite likelihood method, showing that our suggested criteria can successfully be used to choose the tuning parameters that can keep the estimation or the prediction accuracy in applications.
With a continuous expansion of the number and activity of wild animals induced by ecological conservation and restoration efforts, the human-wildlife conflict is becoming more prominent across China. There have been frequent and severe incidents of crop damage caused by wildlife. In this paper, we investigate the crop losses caused by wildlife in the rural districts of Beijing, using a unique dataset of 31,573 observations from 2009 to 2017. Through statistical tests on the individual coefficients and the overall fitness, we find that a negative binomial generalized regression model describes the pattern of crop loss events more accurately, compared to an alternative Poisson model. The frequency of crop loss events is positively related to a village's distance from the river system but negatively associated with the distance from woodland, population density, and protective measures taken. The predicted frequencies of crop damage events are then used to correlate with the amounts of losses at the village level. Based on these results, we propose solutions for effective reduction of future crop losses and practical assessment of the likely damage compensation or insurance premium.
We consider the hypothesis testing problem for the smoothness parameter ν in a stationary isotropic Gaussian random field with Matérn covariance. For the data observed on a regular grid, we construct the rejection region for one-tailed tests, and starting from there, we develop a chain-like testing procedure, which can determine an interval containing the true value of ν. Such an interval can help improve the performance of various estimation methods for ν, such as restricting the parameter space or validating the assumptions for the asymptotic properties of the estimator. The test statistic is built on recursive applications of the Laplace operator to the observations. For this statistic, the fixed-domain asymptotic normality is established and the forms of asymptotic mean and variance are derived. Therefore, the proposed tests are guaranteed to have correct asymptotic size under certain conditions. Simulation studies indicate that our proposed methods are efficient for moderate sample sizes. As an application of the chain-like testing procedure, we provide a method of choosing the number of differencing for a local Whittle-likelihood type estimator of ν proposed by Wu, Lim, and Xiao, and show that it can avoid obtaining inconsistent estimates of ν via a numerical experiment.
Under the uncertain statistical framework by Liu [19], there is still a lack of an effective fitting method for uncertain linear models with Box-Cox transformation of response variables. For example, for the transformation parameter lambda, the uncertain least squares estimation will produce a severely low estimation result. In this paper, we propose uncertain Box-Cox regression analysis by utilizing the uncertainty theory to model the imprecise data and applying a generalized Box-Cox transformation indexed by its parameter to validate classic regression assumptions. We use rescaled least squares to estimate unknown parameters and provide an estimate for noises followed by residual analysis for these uncertain Box-Cox regression models. We also give the forecast values and confidence intervals and use a numerical example to demonstrate our methodology. Ourwork sets a uniformed framework for Box-Cox transformation on the uncertain regression, and extends such regression from linear to nonlinear cases, taking the Johnson-Schumacher growth model as an example.
Regression models are often called for to quantify relationships between the explanatory variables and the response variable. Mathematically, data should be collected and recorded as their true values, which turns out to be unrealistic for the real world. In this paper, we introduce uncertain variables to characterize such imprecise data, apply the most useful logarithmic, square root or reciprocal transformation to alleviate possible nonlinearity problems and estimate the disturbance terms for the obtained uncertain regression models, followed by confidence interval estimations and point predictions. For each type of models being proposed, namely the uncertain revised regression models, uncertain revised asymptotic regression models and uncertain revised Michaelis–Menten kinetics regression models, we give a numerical example, respectively, to illustrate our approach.
In this study, human health risk derived from radioactive pollution in drinking water of China was assessed based on gross alpha and beta. Considering the presence of numerous data under the detection limits, the left-censored handling methods were employed to deal with the non-detected values in gross alpha and beta radioactive concentrations. Results show that concentrations of gross alpha and beta range from 4.98 × 10−4 Bq/L to 0.49 Bq/L with a mean value of 0.029 Bq/L and 5.00 × 10−3 Bq/L to 1.26 Bq/L with a mean value of 0.091 Bq/L, respectively. With the average effective dose being 1.41 × 10−2 mSv/y, the annual cancer risk due to radioactive pollution in Chinese drinking water is 7.75 × 10−7 /y. This study aimed to provide an easier method to quantify the radioactive pollution in drinking water and give a scientific basis for making policy decisions on radioactive pollution management.