
The increasing frequency and clustering of extreme environmental events, particularly wildfires, presents significant challenges for risk assessment and resource allocation. Traditional statistical models often fail to capture both the self-exciting nature of these events and their dependence on non-stationary climate drivers. We propose a novel Bayesian hierarchical framework that integrates climate covariates into a spatio-temporal Hawkes process model. Our approach decomposes wildfire risk into chronic (climate-forced) and acute (self-excited) components through a conditional intensity function where the background rate depends on slow-moving climate variables while the self-exciting component captures local clustering. We employ Hamiltonian Monte Carlo (HMC) for efficient posterior inference in this high-dimensional parameter space. Applied to historical wildfire data in the Western USA (1992-2023), our model demonstrates superior predictive performance compared to baseline Poisson and Cox process models. The framework provides interpretable risk decomposition and probabilistic forecasts under different climate scenarios, offering valuable insights for adaptive management and policy planning.
This study explores the prediction of a continuous target variable that is primarily measured as an ordinal outcome. Interest in such prediction may occur given data from double- or two-phase sampling designs, where an auxiliary variable is measured at all sampling units and a target and predictor variable, X, at a subset of those units. This study focuses on the case where X is nonnegative, is conditional on a zero process, and is sampled using a cluster sampling design. We model the ordinal outcome using cumulative logistic regression and the conditional regressor X, following log transformation, in a Bayesian setting. Estimation of model parameters and prediction of missing X is generally accurate and precise, particularly with sufficiently large Pr(X>0) and associated variances. This study is motivated by a need to accurately estimate the biomass of submersed aquatic vegetation on the Upper Mississippi River, where two types of measurements for biomass are available: inexpensive, ordinal biomass scores and expensive, continuous diver-harvested biomass data. Here, X represents plant biomass and Pr(X>0) is the usual species site occupancy measure. When fit using biomass data from the submerged aquatic vegetation species Vallisneria americana Michx, our model estimates parameters within the parameter space region where the model performed acceptably using synthetic data. We expect this model to appeal to investigators with ordered outcomes, cluster designs, and one or more continuous, positive predictors.
Near-infrared (NIR) spectra provide rich functional information for predicting sugarcane extraneous matter, but effective modeling is hindered by the functional nature of the predictors and the bounded, percentage-based response. We propose a Transformed Partial Functional Linear Model that integrates smooth functional predictors, categorical covariates, and a data-driven monotone transformation of the response, with all tuning parameters selected through a leave-group-out cross-validation scheme designed to ensure robustness to previously unseen sugarcane varieties and field conditions. The approach yields an interpretable spectral coefficient function while preserving inference on the original percentage scale via the inverse transformation. Extensive simulations demonstrate improved stability over partial least squares regression, generalized functional linear model, and modern machine learning methods, and the analysis of real sugarcane NIR data shows that explicitly accounting for both the functional structure of spectra and the bounded nature of the response leads to substantially improved predictive accuracy. Supplementary materials accompanying this paper appear online.
In quantitative genetics and breeding, phenotypes are measured to improve traits of interest. These phenotypes combine genetic and environmental components and are therefore modeled using mixed models. As genetic values are unobservable, they are treated as latent variables with a covariance structure defined by the pedigree. Multiple observed phenotypes are typically assumed to follow a joint Gaussian distribution. Estimation is then performed via restricted maximum likelihood (REML) or Bayesian methods under Gaussian assumptions. However, even if each phenotype’s component appears Gaussian, the joint distribution may exhibit lower tail dependence despite Gaussian marginals, due to complex dependencies between random variables. This can lead to biased estimates of variance components. We propose an extension of the standard genetic and environmental mixed model by incorporating copula functions to flexibly capture the joint distribution of phenotypes beyond Gaussian assumptions. We develop a stochastic gradient descent algorithm combined with a Monte Carlo Markov Chain step to estimate variance components and predict genetic values. The method is evaluated through simulations and applied to a real pig breeding dataset, where normality assumptions for the joint phenotype distribution seem inappropriate. Our results show that the copula-based approach provides more robust estimates compared to standard REML under non-Gaussian conditions.
Accurate forecasting of crop yield anomalies is essential for global food security and effective agricultural policy-making, as it enables the separation of climate-driven impacts from baseline regional productivity. This study proposes a transparent and lightweight ensemble forecasting pipeline, termed Lightweight Exponential-Weighted Regression (LEWR), for predicting such yield deviations. To account for time-invariant country-specific characteristics, a cross-sectional demeaning procedure is first applied to the panel dataset. The procedure then integrates heterogeneous regression models—including Random Forest, XGBoost, and Support Vector Regressors—through exponentially decaying weights derived from validation errors, emphasizing higher-performing predictors without employing a second-layer meta-learner. The performance of the proposed method is evaluated on a hold-out test set (2009–2013) drawn from a panel dataset covering multiple countries and years (1990–2013) and benchmarked not only against individual base models but also against established ensemble methods. Results show that the procedure achieves competitive accuracy in forecasting yield deviations compared to more complex methods, while offering notable advantages in simplicity and computational efficiency. Furthermore, it supports incremental updates and provides a straightforward mechanism for quantifying predictive uncertainty—a critical factor in modern environmental risk assessment. Supplementary materials accompanying this paper appear on-line
Whenever many hypothesis tests are performed simultaneously, the risks of falsely rejecting some null hypotheses must be accounted for in statistical analyses. Methods controlling the possibility of any false rejections become very conservative as the number of tests grows, which motivated the false discovery rate (FDR) control framework that limits the average proportion of rejections that are false. However, FDR control methods, such as Benjamini–Hochberg, often have undesirably low power when working in spatially correlated settings. Here we focus on regression problems in large climate datasets involving a univariate response variable and spatially gridded covariate data. Given that nearby locations are more likely to share a signal, we propose a method in which the covariates are spatially smoothed with locally fitted covariance functions before applying the hypothesis testing procedure. Simulation results show that by using Benjamini–Hochberg with this smoothed data, power is increased while still maintaining empirical FDR control. We then apply the technique to real sea surface temperature covariates with both simulated and real responses. The real data example, linking sea surface temperatures to July temperatures in the central US, demonstrates that the smoothing procedure allows for signals to be identified when none pass traditional Benjamini–Hochberg. The proposed procedure provides climate scientists with a simple yet powerful tool for extracting more information from their data, while correcting for multiple hypothesis testing.Supplementary materials accompanying this paper appear online.
Environmental count series often exhibit long runs of zeros punctuated by sporadic, operationally critical bursts. We develop a distribution-first Bayesian framework that learns the Poisson–Tweedie (PT) tail index while shrinking toward the Poisson-Inverse Gaussian (PIG) corner when evidence supports heavier extremes. The likelihood is implemented via a Poisson-Generalized Inverse Gaussian (Poisson-GIG; Sichel) bridge, yielding closed-form marginal probability mass functions (pmfs) and a continuous path between negative binomial (NB) and PIG behavior. We establish identifiability through a monotone mapping from a one-dimensional GIG tail parameter to the PT index p∈ [2,3] , give minimal conditions for posterior propriety under weak priors, and develop computation based on latent-rate augmentation (with exact conditionals at the inverse-Gaussian (IG) corner), a stable one-dimensional tail update, and a fast variational surrogate. Posterior prediction targets decision-relevant functionals: high-threshold exceedance probabilities, a threshold-weighted continuous ranked probability score (CRPS), exceedance probability integral transform (PIT) diagnostics, and a zero-mass check orthogonalized via an optional mechanistic hurdle for structural zeros. Simulations confirm accurate recovery of the tail index and superior calibration of far-right probabilities when extremes matter. In an application to US wildfire event frequencies from the Storm Events archive, empirical exceedances—after matching dispersion—often align more closely with the heavy-tailed PIG corner, and the proposed shrinkage yields calibrated high-threshold predictions without ad hoc zero inflation. The framework remains compatible with mean/dispersion regression and spatial-temporal random effects, providing a scalable, interpretable basis for exceedance-oriented risk assessment.
A fundamental task in ecological statistics is to estimate abundance and growth rate distributions from wildlife monitoring data to inform conservation management. Modeling time series of wildlife populations presents a number of challenges from both statistical and ecological perspectives, including discreteness; lack of replication; nonstationarity; and observation, demographic, and other phenomenological processes. Nonstationary dynamics are often exhibited by populations undergoing environmental stressors. Models must account for these characteristics to produce reliable estimates of abundance and trends, yet estimation can be challenging with unreplicated data. We propose nonstationary demographic state-space models using unreplicated counts for populations undergoing environmental stressors. A reduced growth rate model matches the complexity of the unreplicated count data, and a fecundity bound on growth rate distributions allows the separation of processes affecting growth rates like environmental stressors from those affecting abundance external to growth rates like migration. NDSSMs allow for the embedding of nonstationary model components, and we explore the use of changepoints, volatility clustering, and migration processes. We apply the proposed nonstationary models in case studies of herons affected by predator/competitor reestablishment and three bat species affected by a fungal pathogen causing white-nose syndrome. Nonstationary models outperform stationary models and generalized linear mixed effects models according to model scoring and visual inspection of predictions, and provide estimates more consistent with published values. Incorporating migration improves model fit universally, even with approximate one-way immigration, most likely because populations are extirpated, recolonized, and increase multiple-fold over the upper bound set by species fecundity. In addition, estimates of the timing and severity of the environmental stressor differed for models with migration. Including nonstationary and demographic components in a fecundity-bounded growth rate model improves inference and benefits interpretability of hyperparameters. In turn, this adjusts uncertainties in predictions of abundance and growth rates over time, providing the ingredients needed for informed conservation analysis and for directing future monitoring of at-risk species. Supplementary materials accompanying this paper appear online.
Accurate crop classification using hyperspectral data is essential for precision agriculture and environmental monitoring. This study presents a hybrid framework combining wavelet transform and parallel bidirectional long short-term memory (BiLSTM) networks. The wavelet transform is first employed to extract multi-resolution spatial features from the hyperspectral cube, which are then processed through parallel BiLSTM branches to capture forward and backward spectral dependencies. These outputs are combined and transmitted through a dense classification layer. The model was evaluated on five standard hyperspectral datasets; Indian Pines, Salinas, and Wuhan UAV borne hyperspectral image dataset series, which have three different datasets HanChuan, HongHu, and LongKou for crop classification, demonstrating high classification performance with overall accuracy of values ranging from 99.28
When crop insurance programs are introduced in areas with limited historical data, premium rates are often based on proxy data from neighboring regions. This study evaluates the credibility of such proxy data in rating US county-level cotton revenue insurance contracts. We apply flexible copula-based models and value at risk to estimate target (‘true’) premiums using complete yield and price data, and simulate missing completely at random scenarios where proxy data are used to fill gaps. Premiums based on pooled proxy data from two or more adjacent counties are then compared to target rates using root mean square error and diversification effects. Results reveal substantial (−37 to +189
We introduce a new Bayesian model for detecting and estimating changepoints in spatially correlated functional time series to study the impact of the Mt. Pinatubo eruption on June 15th, 1991 on global temperatures. We model the exact quadratic form of the ℓ ^2 -norm of the functional cumulative sum (CUSUM) statistic and leverage a spike-and-slab approach for automatically detecting changepoints while also measuring their magnitudes. Our approach, Spatial Quadratic Regression (SQuaRe), improves over existing approaches by enabling simultaneous detection and estimation while also allowing for spatially correlated changepoints. Extensive simulations demonstrate that our approach more precisely estimates change magnitudes than existing methods, while having comparable detection performance to frequentist approaches. We apply our method to temperature profiles in the upper troposphere and stratosphere to track the impact of the Mt. Pinatubo eruption on global temperatures. The spatial changepoint pattern shows an early impact of the eruption on low-latitude regions, and then the impact gradually moved to high latitudes where intrinsic variation becomes high, whereas there is little impact on the southern hemisphere.
In this paper, we examine a finite population distributed across a two-dimensional spatial domain, where a small subset of units exhibit substantially larger response values compared to the rest. Each unit is associated with a size variable, which is proportional to its response value and known prior to sampling. From this population, we develop an augmented sampling strategy involving two randomization stages: a spatially balanced sample (SBS) and a probability proportional-to-size (PPS) sample with unequal inclusion probabilities. First, we select an SBS of size n_1 , ensuring spatial balance across the population. Subsequently, a local pivotal method (LPM) sample of size n_2 is drawn from the remaining unsampled units, using unequal inclusion probabilities proportional to the size variable. This two-stage sampling approach combines spatial balance with targeted selection of large-size units. Our results show that the population mean estimator derived from the proposed sampling design outperforms those based on individual SBS or LPM samples. Furthermore, we provide explicit expressions for the first-order inclusion probabilities associated with a given size variable. To illustrate the method’s effectiveness, we apply the proposed design to a population of poppy fields in Kandahar province of Afghanistan. Supplementary materials accompanying this paper appear online.
Understanding how the correlation structure among water quality variables varies across space is critical for effective environmental monitoring and management. In this study, we propose a methodological framework to estimate and analyze spatially varying partial correlation matrices over a fluvial network, with a particular focus on the Piedmont region in northern Italy. We adopt the perspective of Object Oriented Data Analysis, treating correlation matrices as complex data objects defined over a graph-based domain. Given the limited number of repeated measurements available at each monitoring site, we develop a nonparametric, topology-aware kernel estimator to reconstruct and predict valid correlation matrices that respect the structure of the fluvial network. To summarize and compare these matrices, we further employ tools from Topological Data Analysis, such as Betti curves. Clustering these curves enables the identification of spatial patterns and seasonal dynamics in the interdependence of water quality variables. The integration of nonparametric smoothing, network-based prediction, and functional clustering offers an approach for detecting localized anomalies and guiding adaptive monitoring strategies in freshwater systems. As a result, the analysis reveals four clusters per season, potentially reflecting localized vulnerabilities driven by seasonal water availability and anthropogenic pressures such as water withdrawals.
In econometrics, the impact of climate change on agricultural yield has often been modeled using linear functional regression, where crop yield, a scalar response, is regressed on the temperature distribution over a given time period, treated as an ordinary functional parameter, along with other covariates. We explore alternative models that respect the distributional nature of the temperature distribution parameter. Replacing functional observations with the corresponding distributional ones is appropriate for phenomena that are insensitive to the temporal order of events. Since classical addition and scalar multiplication are unsuitable for density functions, alternative operations, spaces and corresponding regression techniques are required. Moreover, compositional data analysis suggests that such covariates should undergo appropriate log-ratio transformations before inclusion in the model. We compare a discrete approach, where temperature histograms are treated as compositional vectors, with a functional scalar-on-density regression using a Bayes space representation of temperature densities. We evaluate the strengths of each method in modeling rice yield in Vietnam, using data on daily temperature extremes. Additionally, we propose modeling climate change scenarios with perturbations of the initial density along a change direction curve informed by IPCC scenarios. The resulting rice yield marginal impact is then quantified using a simple inner product between the density covariate parameter and the change direction curve. Our results indicate that while both approaches yield coherent findings, the scalar-on-density model outperforms the scalar-on-composition with an enhanced ability to accurately gauge the phenomenon’s scale. Supplementary materials accompanying this paper appear on-line.
In this paper, we introduces a novel partially functional linear varying coefficient geographically weighted autoregressive model for analyzing spatial data with scalar responses, functional data and scalar covariates. The proposed methodology extends conventional geographically weighted regression (GWR) by incorporating functional data analysis, thereby providing a powerful framework for investigating spatial heterogeneity in regression relationships involving functional predictors. Through the integration of functional principal component analysis (FPCA) with local linear smoothing techniques, we derive consistent estimators for three key components: (1) parametric coefficients for scalar covariates and spatial lag, (2) slope functions for functional covariates, and (3) spatially varying coefficient functions. We further establish a hypothesis testing procedure to assess the spatial stationarity of regression coefficients. Under appropriate regularity conditions, we prove the asymptotic properties of our estimators and derive their convergence rates. Extensive Monte Carlo simulations demonstrate the superior finite sample performance of our approach. Finally, we present an empirical application to meteorological data that show the practical utility of our methodology in environmental research.
The theoretical frameworks of the Capability, Opportunity, Motivation-Behavior (COM-B) model and the Theoretical Domains Framework (TDF) are often used by social scientists to study interconnections between determinants of human behavior. In this study, we seek to empirically validate COM-B by learning network structures from data on behavioral determinants to food safety practices in Cambodian markets. To this end, we implemented a carefully tailored combination of statistical methods in a novel context that is uniquely aligned with the motivating problem. Specifically, we leveraged a search algorithm for network structures with methodological adaptations for ordinal survey data. Data consisted of responses from 169 participants to 18 survey items formulated according to TDF domains within COM-B constructs for food safety and measured on an ordinal 1-to-7 Likert scale. We adapted the Inductive Causation algorithm to learn network structure from ordinal data. Marginal and partial Spearman rank correlations for all pairs of survey items yielded an undirected dependency graph, followed by directed-separation to identify unshielded colliders. Results showed a dense network structure of survey items depicting closely interconnected behavioral determinants of food safety practices in Cambodian markets. Findings were consistent with theoretical expectations and provided empirical support for COM-B. Yet, only a limited number of network edges were oriented based on these data. Additional empirical and methodological work is warranted to further refine insight into a data-informed COM-B network, with implications for development of targeted interventions that promote behavioral change. Supplementary materials accompanying this paper appear on-line.
Standard spatial capture–recapture (SCR) models assume independence between detections of the same individual conditional on its activity center, but in many circumstances this is unrealistic. We develop a model for continuous-time camera trap surveys that explicitly model dependence induced by animal movement. We formulate capture rates as a weighted mixture of two bivariate Gaussians. One of these is a fixed home range while the other is centered where the individual was last detected. After a capture, our model increases the capture rate at nearby traps in a Hawkes-like manner. We also decrease rates at far away traps so that individuals are not automatically “trap-happy.” This is in contrast to standard models whose rates vary in space but not time. Thus, our model gives additional inference about temporal and spatial clustering of detections. We present both frequentist and data augmentation approaches to fitting our model, demonstrate by simulation that our model gives reliable estimates, and fit our model to a dataset of leopard (Panthera pardus) sightings. Compared to SCR, we obtain a lower population estimate and a higher home range size estimate alongside additional inference about animal movement. Supplementary materials accompanying this paper appear on-line.
Motivated by an application to study the impact of temperature, precipitation and irrigation on soybean yield, this article proposes a sparse semiparametric functional quantile model. The model is called “sparse” because the functional coefficients are only nonzero in the local time region where the functional covariates have significant effects on the response under different quantile levels. To tackle the computational and theoretical challenges in optimizing the quantile loss function added with a concave penalty, we develop a novel convolution-smoothing-based locally sparse estimation (CLoSE) method, to do three tasks in one step, including selecting significant functional covariates, identifying the nonzero region of functional coefficients to enhance the interpretability of the model and estimating the functional coefficients. We establish the functional oracle properties and simultaneous confidence bands for the estimated functional coefficients, along with the asymptotic normality for the estimated parameters. In addition, because it is difficult to estimate the conditional density function given the scalar and functional covariates, we propose the split wild bootstrap method to construct the confidence interval of the estimated parameters and simultaneous confidence band for the functional coefficients. We also establish the consistency of the split wild bootstrap method. The finite sample performance of the proposed CLoSE method is assessed with simulation studies. The proposed model and estimation procedure are also illustrated by identifying the active time regions when the daily temperature influences the soybean yield. Supplementary materials accompanying this paper appear online.
Semi-Latin squares have been extensively studied. They can be interpreted as a special case of latinized block designs where the number of columns is equal to the number of replicates in the design. Latinized row-column designs are frequently used in field and glasshouse trials when replicates are contiguous. These designs allow for the efficient adjustment of row and column effects within replicates. Here we define extended semi-Latin squares as a special case of latinized row-column designs and investigate optimality using the average efficiency factor.
Riverine flooding poses considerable risks. Developing strategies to manage flood risks requires flood projections at decision-relevant scales, often at high spatial resolutions, and with solid uncertainty characterization. However, obtaining high-resolution flood projections can be computationally prohibitive. To address this challenge, we propose a probabilistic downscaling approach that maps low-resolution flood projections onto higher-resolution grids. The existing literature presents two distinct types of downscaling approaches: (1) probabilistic methods which are versatile and applicable across various physics-based models and (2) deterministic downscaling methods specifically tailored for flood hazard models. Both types of downscaling approaches come with their own set of mutually exclusive advantages. Here we introduce a new approach, PDFlood, that combines the advantages of existing probabilistic and flood hazard downscaling approaches: model-based uncertainty quantification and accurate point estimates from approximating physical processes. PDFlood provides (1) narrower prediction intervals and more accurate flooding probabilities compared to a state-of-the-art probabilistic downscaling approach and (2) comparably accurate point estimates and improved classification of cells as flooded versus non-flooded relative to a state-of-the-art flood hazard downscaling approach. By allowing users to consider both accurate point estimates and uncertainties, PDFlood can better inform the design of risk management strategies. While we develop PDFlood for flood hazard models, the general concepts translate to other applications such as wildfire models.