
Spatial expected shortfall regression (SESR) characterizes spatial dependence and tail losses at a given quantile level. However, high-dimensional small-sample data easily lead to parameter redundancy and estimation distortion. Although existing transfer learning can alleviate the small-sample problem, it relies on raw data sharing and source domain screening, incurring privacy leakage risks and high computational costs. To address these issues, this paper proposes model averaging transfer learning at the spatial quantile regression (SQR) stage. Different from traditional mean regression frameworks, the proposed method first fuses multi-source information to improve SQR estimation accuracy, and then constructs SESR based on the optimized SQR results, without raw data sharing or source domain screening, thereby achieving SESR parameter estimation under high-dimensional small-sample conditions. Numerical simulations, robustness tests, and empirical analysis on the London property prices dataset demonstrate that the proposed method outperforms conventional approaches, providing a reliable analytical framework for spatial tail risk measurement under high-dimensional small-sample settings with privacy constraints.
In the era of big data, spatial data has become pivotal for exploring geographic phenomena due to its spatial correlation and heterogeneity. Traditional spatial autoregressive models are constrained by linear assumptions and parameter estimation limitations, making them ill-suited for handling non-stationary data and extreme risk analysis. While quantile regression overcomes limitations in error distribution, existing spatial quantile models still fail to capture nonlinear spatial spillover effects of the dependent variable due to the assumption of linear endogenous effects. To address this, this paper proposes a robust estimation method for spatial quantile autoregressive models under nonparametric endogenous effects, innovatively incorporating deep neural networks to handle nonlinear relationships. To address quantile overlap issues, two models are constructed: one fits customized models for each quantile, while the other employs a multi-task model to simultaneously estimate multiple quantile functions, thereby avoiding overlap. Dedicated neural network frameworks and algorithms are designed for both approaches. Numerical simulations demonstrate both neural network approaches exhibit superior predictive performance and interpretability. Validation using U.S. San Diego housing rent data further confirms their practical utility and forecasting advantages.
In this paper, we investigate quantile regression estimation for the partially linear single-index spatial autoregressive model with covariates missing at random. This model combines linear interpretability and single-index flexibility while capturing spatial dependence. We construct an inverse probability weighting approach based on an iterative algorithm and local linear estimation method to estimate unknown parameters and the link function. Under some regularity conditions, we derive the asymptotic distributions of the proposed estimators. Monte Carlo simulations verify the finite sample performance of our method, and an application to the Boston housing price dataset demonstrates its effectiveness in real data analysis.
Human mobility trajectories are central to understanding social behavior, infrastructure use, and population response across a wide range of settings. Modeling these trajectories is inherently a spatial statistics problem, yet the data differ fundamentally from traditional spatial datasets. Unlike conventional areal or point-referenced observations, mobility trajectories introduce statistical challenges such as sparse and irregular sampling, temporally dependent movement decisions, spatial dependence shaped by transportation networks rather than Euclidean proximity, and heterogeneous semantic attributes of visited locations. These complexities exceed the assumptions of classical spatial statistical models and motivate the use of modern representation-learning tools.We introduce STAGE (Spatio-Temporal Attention and Graph Embedding), a general framework for modeling human mobility trajectories through spatially and temporally informed representation learning. STAGE integrates (i) graph-based embeddings of a spatial tessellation that capture connectivity patterns beyond nearest-neighbor adjacency, (ii) semantic embeddings of Points of Interest (POIs) summarizing the functional context of the built environment, and (iii) an attention-based autoencoder that explicitly models temporal dependence while accommodating sparsity through masked reconstruction. Together, these components yield latent representations that preserve key spatial, semantic, and temporal structure.The resulting latent space provides a low-dimensional, interpretable summary of mobility behavior suitable for diverse downstream statistical analyses. We further demonstrate the merit of the STAGE framework by clustering real mobility data following Hurricane Ian and identify five coherent behavioral regimes related to sheltering, medical access, evacuation, and localized recovery.
Marked point processes provide a natural framework for modeling events that occur in space together with associated attributes, denoted as marks. Applications span across various fields including, for instance, seismology, ecology, and wildfire analysis, where marks are often continuous and may exhibit heavy-tailed behavior. A key quantity of interest is the joint intensity function, which characterizes the occurrence of events across both space and marks. Existing estimation approaches typically rely on restrictive separability assumptions, encounter difficulties when applied to complex domains, and perform poorly in the presence of right-skewed mark distributions. We propose a fully nonparametric, penalized likelihood estimator for the joint intensity function of continuously marked point processes. The method avoids separability assumptions, naturally handles non-convex and curved spatial domains, and introduces an exponential mark-dependent smoothing parameter that improves estimation particularly under truncated heavy-tailed mark distributions. Theoretical analysis establishes uniqueness and asymptotic convergence of the estimator. To enable practical implementation, we employ a numerical procedure based on a tensor-product discretization combining finite elements and B-splines, together with an efficient optimization scheme and a data-driven selection of the smoothing parameters. Finally, the performance of the method is assessed through simulation studies benchmarked against existing approaches, and its practical utility is demonstrated through real-data applications to wildfire occurrences in California and to worldwide earthquake events.
Urban activity systems exhibit recurring spatio-temporal patterns, yet disentangling temporal reallocation from spatial persistence remains a key challenge for urban monitoring and mobility analysis. This paper introduces a reproducible statistical framework based on canonical hotspots – spatial regions identified from historical data and held fixed throughout the analysis – to provide stable references for comparing urban dynamics across days. Using a cleaned analytical sample of Manhattan, NY, taxi pickup records covering the period 2009–2015, we represent daily activity as normalized temporal profiles and spatial allocations over a fixed hotspot structure. Information-theoretic distances, permutation-based inference, and SIMPER-inspired contribution analyses are used to quantify temporal reallocation, spatial divergence, and economic dominance patterns. Results show that weekday variation is driven primarily by temporal reallocation. Spatial differences, although statistically significant, are highly concentrated: six operationally critical hotspots account for more than 82% of the observed weekday–weekend spatial divergence. Furthermore, spatial persistence is found to be a poor proxy for economic relevance. Revenue-weighted centroids diverge systematically from trip-based centroids, indicating that economic dominance emerges from transient spatio-temporal alignments rather than from persistent demand concentrations alone. Beyond the empirical findings, the framework provides a mechanism for separating enduring spatial organization from recurrent temporal redistribution. The results suggest that urban activity systems exhibit a layered structure in which stable spatial patterns coexist with systematic temporal reallocation. By combining stable spatial references with distributional analysis, the proposed framework provides a reproducible basis for comparing urban activity systems and characterizing spatio-temporal heterogeneity in urban mobility data.
Neural simulation-based inference is a recent approach involving training a neural surrogate model for a likelihood, posterior, or density ratio. It is useful for producing computationally efficient estimators of unknown or intractable likelihood models if data from such models can be simulated easily, such as for Neyman–Scott point processes which we study here. For this process model, we use, extend and compare simulation-based inference methods based on normalizing flow and a recently proposed neural classifier approach for likelihood estimation. In a simulation study, normalizing flow appears to yield the narrowest confidence bounds while achieving similar empirical coverages. For the neural classifier approach, a proposed Godambe adjustment leads to better calibrated confidence regions than Platt scaling, especially when less training data is available. Approximate Bayesian computation yields Neyman–Scott clustering parameter estimates with substantially higher RMSEs than the other methods, at higher computational cost because it is not amortized.We apply the methods to seismic event data over time at the Åknes rock slope in Norway. We investigate the association of water pressure and temperature with seismic activity to better understand how these variables influence geohazard risk. Three different models are applied and compared: non-homogeneous Poisson, Neyman–Scott, and Hawkes process models. Effects of covariates are comparable for the different models, but the Neyman–Scott and Hawkes processes capture additional clustering, making them more representative of the Åknes data.
Spatial heterogeneity is a fundamental property of spatial data, with critical implications for understanding spatial patterns and supporting decision-making across diverse fields. While change point detection has proven powerful in identifying structural shifts in time series, its extension to spatial domains remains challenging. Specifically, most existing change point methods rely on temporal ordering and unidirectional dependence, assumptions that are ill-suited for spatial data, which lack a natural order and exhibit complex spatial autocorrelation. To address this limitation, we propose a non-parametric spatial change points detection approach based on block discrepancy measures, which involves partitioning the spatial domain into row/column-wise blocks to preserve local spatial correlations. Within each block, observations are ordered by their spatial coordinates to convert the two-dimensional process into a one-dimensional sequence, upon which methods such as binary segmentation (BS) and wild binary segmentation (WBS) can be applied to detect local change points. Requiring only observed values and their coordinates, the proposed approach does not impose parametric assumptions on the underlying distribution or spatial dependence structure. It therefore provides a flexible way to apply change point detection ideas to spatial data. Simulation studies are conducted to evaluate a range of change point detection methods under various spatial heterogeneity scenarios. The results indicate that most methods achieve satisfactory performance in detecting boundaries between spatially heterogeneous subregions, although their accuracy and robustness vary across different sampling designs and spatial error structures. Furthermore, the methods are applied to air quality monitoring data in China to demonstrate their practical applicability.
Spatial autoregressive (SAR) models become difficult to use when the response is non-Gaussian, the target sample is limited, and the fitting procedure is required to include privacy-oriented randomization. This paper proposes the TranSAR_BC_DP framework, which combines joint Bayesian Box–Cox–SAR estimation, source-aware transfer learning, and a final Gaussian-perturbed fine-tuning stage with explicit Rényi differential privacy (RDP) accounting. The transformation parameter, spatial autoregressive coefficient, regression coefficients, and innovation variance are estimated jointly under positivity, full-rank, stability, and non-collinearity conditions. A practical local-identifiability condition is stated in terms of the rank of the joint sensitivity matrix for (λ,ρ,β), and representative fits are evaluated using the numerical rank and scaled smallest singular value of its column-normalized form. The use of proper priors yields a well-defined finite-sample posterior under the stated assumptions. Candidate source domains are screened by residual-bootstrap comparisons. The empirical support count is converted to a one-sided binomial sign-test p-value, and the Benjamini–Yekutieli step-up rule is used to control the false discovery rate under arbitrary dependence among candidate-source tests. Retained sources are weighted by the Spatial Autocorrelation Congruence Index (SACI), a bounded score that combines normalized spatial autoregressive strength with fixed-dimensional bounded descriptors of the spatial weight matrices. The final noisy-gradient kernel uses Poisson subsampling, per-record clipping, Gaussian perturbation, and explicit RDP composition. Its formal (ϵ,δ)-DP guarantee requires the transformation, feature-standardization constants, source-selection quantities, transfer summaries, spatial contextual inputs, and initialization either to be fixed independently of the protected records or to have been produced by separately privatized mechanisms whose privacy costs are composed. In the reported benchmark pipeline, several upstream quantities are estimated from the same target data; consequently, the reported ϵ̄ values quantify the accountant-calibrated final-stage perturbation and are not presented as an end-to-end DP guarantee for the complete implemented pipeline. Experiments on synthetic data and county-level COVID-19 mortality data illustrate the statistical and final-stage perturbation–utility behavior of the stage-wise procedure. In a representative ablation comparison, TranSAR_BC_DP reduces RMSE by 70.4% and residual Moran’s I by 78.6% relative to Standard SAR. In the California application, the finite-perturbation fit indexed by the nominal final-stage accountant level ϵ̄=0.5 attains a model-specific utility preservation rate of 95.3% relative to the matched non-perturbed version of the same architecture. These empirical values describe the evaluated benchmark configurations and are not interpreted as universal performance or legal-compliance claims.
Classification is a supervised machine learning method that predicts a categorical response variable using several explanatory variables. If observations are sampled from a spatial point process, then we can also use x-and y-coordinates as explanatory variables. If the observations are sampled from a known linear network instead of whole space, then the distance between two points is defined in a different manner, and we require a classifier for the linearly clustered data. In this study, we address the classification problem on a tree-shaped linear network. We select a point on the edges in the given linear network to split the space, and then construct a decision tree through recursive splits. We propose an adaptive boosting algorithm using this decision tree as a weak classifier. Finally, we provide some simulated examples and real data analysis, comparing with adaptive boosting based on decision trees constructed using Cartesian coordinates. The proposed method has better accuracy than the comparison method, when the observations are clustered on linear network.
Modeling and simulating space–time random fields while accounting for possibly misaligned covariates is crucial for many applications requiring uncertainty quantification and risk assessment. To achieve this, we propose a flexible framework that couples machine learning quantile regression with space–time Gaussian random fields. In this framework, the target variable is modeled as a combination of a latent Gaussian random field and transformed marginals obtained by machine learning quantile regression conditionally on a set of covariates. We illustrate the approach on a synthetic experiment and on a case study that considers daily maximum temperature over northwestern Switzerland conditional on seasonal cycle and large-scale geopotential height.
Data exhibiting negative correlation, particularly in temporal, spatial and spatio-temporal contexts, are frequently described by employing hole effect models or damped oscillation models. Nevertheless, these classical models suffer from severe limitations that reduce their adaptability. Specifically, models belonging to the Bessel class are criticised for consistently displaying a parabolic behaviour near the origin and an uncountable number of zeros. Furthermore, particular models within this class, such as the cosine model, suffer from additional drawbacks, including lack of admissibility in Euclidean spaces of dimension greater than one and not being strictly positive definite. Recently, wide classes of covariance functions, built through the difference of two correlation models, present distinct characteristics with respect to the traditional hole effect models and are considered to be sufficiently flexible, also in space–time. In this paper a further advance to construct correlation models, characterised by negative values, in a spatio-temporal domain is proposed as well as some examples and a case study are presented.
Spatially aggregated power forecasts are commonly used to manage the energy grid; however, uncertainty estimates around those predictions are often unavailable. A forecast error, the difference between the predicted and the actual power output, offers valuable information for quantifying this uncertainty. Previous research has shown the benefit of using historical forecast errors to characterise uncertainty in power estimates. Building on this idea, we propose a hierarchical Bayesian formulation that transforms point forecasts into probabilistic scenarios. This approach incorporates temporal autocorrelation, seasonality, ramp behaviour, and heteroscedasticity, which are key for generating realistic power trajectories. Using wind power data from the United States at the aggregate level, and spatially disaggregated data from Scotland, the proposed approach produces calibrated probabilistic forecasts. For the United States data, the method improved wind power scenario sharpness compared with a nonparametric density-estimation benchmark. Our Bayesian hierarchical framework offers a flexible and interpretable alternative for uncertainty quantification in renewable energy forecasting that can be used in spatiotemporal settings.
Gaussian processes represent a rich class of models for supervised learning, particularly for mapping natural variables in geostatistics. They are especially appealing in that they provide a quantification of the prediction accuracy at new locations. Traditional implementations assume second order stationarity. This hypothesis is generally violated when considering large datasets issued from modern sensor networks or satellite images. Various constructions have been proposed to overcome this limitation, such as spatially varying parameters in convolution models or in stochastic partial differential equations, but they remain difficult to handle. In this work, we consider the space deformation model where the estimation of the deformation is considered as a transport problem. In this regard, we model the deformation function with neural networks inspired by normalizing flows, which are particularly taylored to this problem. Normalizing flows classically learn to transform a given, generally easy to sample, distribution into the data distribution. Here, we consider instead learning a bijective transformation from the input geographic space to the deformed space, where stationarity and isotropy hold. The estimation procedure is based on gradient descent optimization over the negative loglikelihood, which is scaled up using symbolic kernel matrices. The method is illustrated both on synthetic data and a real dataset.
In the context of parameter estimation for Gibbs point processes, the state-of-the-art method is Takacs-Fiksel estimation, of which pseudolikelihood estimation is a special case. An alternative method is the recently proposed Point Process Learning approach, based on point process cross-validation and point process prediction errors. Since both Takacs-Fiksel estimation and Point Process Learning are motivated by the Georgii-Nguyen-Zessin formula, which defines Gibbs point processes, in this paper we study Point Process Learning in relation to Takacs-Fiksel estimation. We show that, upon applying appropriate scaling and letting the cross-validation regime tend to leave-one-out cross-validation in Point Process Learning, averages of prediction errors converge to the innovation-based loss function in Takacs-Fiksel estimation. We further provide an empirical risk formulation of Point Process Learning, which highlights the nature of our asymptotic results, and show that the underlying convergence mechanism can be partially understood through a conditional law of large numbers for statistics of conditionally independent thinnings. We finally illustrate our theoretical findings through simulations for a Strauss process, focusing on both convergence diagnostics and comparison of parameter estimation performance between the two approaches.
Regression models have long been widely used to understand spatial processes in human behavior and natural phenomena. Such events as landslides, disease outbreaks, or election victories are often collected as binary responses and operate at different scales. Traditional global binomial regression models, however, assume constant relationships across space, while multiscale geographically weighted regression (MGWR) is mainly restricted to Gaussian and Poisson responses. This paper introduces a novel Multiscale Geographically Weighted Binomial Regression (MGWBR) model, extending the existing MGWR framework to accommodate binomial response variables. By integrating a Local Scoring Algorithm (LSA), the proposed MGWBR model accounts for spatial non-stationarity in the processes leading to binary outcomes and enables the estimation of covariate-specific bandwidths and localized parameter estimates. This approach addresses key limitations of traditional global binomial models and standard MGWR models, which either assume relationships are constant over space or else cannot handle binary dependent variables. To evaluate the performance of MGWBR, the model is applied to two simulated datasets and two empirical datasets: one on landslide occurrences in Clearwater National Forest and one on voting behavior in the Southeast US during the 2020 presidential election. Results from both the simulated and real-world applications demonstrate that MGWBR provides improved model fit over traditional global binary regression models and effectively captures spatial heterogeneity in covariate effects, highlighting its potential as a robust tool for spatial analysis involving binary data.