
Discriminating among competing statistical models for experimental data is of interest for physical science. While right-skewed experimental measurements commonly approximate a lognormal distribution, among many others, substantial deviations occur at the extremes of the distribution, suggesting that alternative models may provide better fits in certain contexts. In many practical contexts, the choice of a single underlying distribution has significant implications, such as estimating the proportion of dwellings exceeding regulatory reference radon levels. We employ a Bayesian approach to model selection, utilizing Bayes factors and posterior model probabilities to compare suitable distributions applied to radon concentration data from Gran Canaria. Using non-informative priors appropriate for shape-scale family distributions, we calculate marginal densities and derive evidence metrics that transcend the limitations of conventional frequentist goodness-of-fit tests. We found strong evidence against the Gamma model and mild/weak evidence for the LogLogistic over the LogNormal distribution. We also offer a Bayesian model averaging strategy that weights the LogLogistic and LogNormal distributions by their posterior probability to create BMA mixture that improves dwelling exceedance frequency estimates across regulatory reference levels.
This paper explores variance or mean squared error estimation in the design-based inferential framework, focusing on a broad class of estimators and sampling schemes commonly employed in spatial surveys for both continuous populations and finite populations of areas. Specifically, the use of the pseudo-population bootstrap is proposed. It involves constructing a spatial pseudo-population from the sample data using the inverse distance weighting interpolator and then selecting bootstrap samples from the pseudo-population using the same sampling scheme adopted to select the original sample. The properties of the resulting variance or mean squared error estimators are theoretically investigated and empirically assessed through a simulation study. Two case studies are considered.
Caribbean islands are increasingly affected by rainfall variability, with alternating droughts and extreme precipitation events threatening water security and increasing flood risk. This study investigates the spatio-temporal variability of rainfall in the Guadeloupe archipelago using daily observations from 20 meteorological stations over the period of 2005–2014. A dynamic clustering approach highlights five areas of consistent rainfall linked to topography and exposure to trade winds, with average rainfall ranging from 3.3 to 9.8 mm day ^-1 . In addition, an algorithm specifically designed to preserve extreme rainfall values was constructed for generating data within the resulting clusters. To capture both regular fluctuations and extreme events, we develop an Adapted Black–Scholes Model with Jumps and Regimes (ABSMJR), formulated as a stochastic differential equation comprising a drift term, a diffusion term driven by Brownian motion, jump components accounting for extreme events, and a regime-switching mechanism governing structural changes in the dynamics. Model parameters are estimated by maximum likelihood method. Model performance is assessed using MASE, RMSE, Theil’s U statistic, and extreme-event metrics. Compared with SARIMA, SARIMA-GARCH, and GLM Gamma, the ABSMJR framework provides a more accurate representation of extreme rainfall events (F1 score = 0.608) while maintaining satisfactory predictive performance. The proposed approach offers a useful tool for reservoir management, agricultural planning, and flood-risk assessment and can be transferred to other tropical regions exhibiting similar rainfall variability.
The proportional hazards (PH) model is one of the most widely used models in survival analysis, typically assuming a log-linear relationship between covariates and the hazard function. However, in the context of spatial survival data, where the time-to-event variable is associated with a spatial location within a given domain, this assumption is often unrealistic in capturing spatial effects. Thus, this paper proposes modeling the location effect through a nonparametric function of spatial location. The function is approximated using finite element methods on a triangulated mesh to accommodate irregular domains. Estimation is carried out within the classical partial likelihood framework, with smoothness of the spatial effect enforced through differential penalization. Using sieve methods, we establish the consistency and asymptotic normality of the parametric component. Simulations and two empirical applications demonstrate superior performance compared to existing approaches.
Exposure to high air pollution levels, especially in urban contexts, is a major risk factor for human health. Most models in literature, however, focus on the bulk of the distribution, and only few address its extremes, such as the right tail. In this work, we apply a Bayesian spatio-temporal quantile regression (QR) framework to daily air pollution data (NO2, PM10, and PM2.5) in the Rome (Italy) municipality between 2011 and 2022. The model specification includes temporal, spatial, and spatio-temporal predictors, and a spatial Gaussian process (GP) to adjust intercept levels and capture spatial variability between monitoring sites. Models were evaluated through temporal and spatio-temporal cross-validation (CV), and sensitivity analyses were performed. Results highlighted that the majority of variability was captured by the GP. Spatial variability was captured especially for NO2; the same pollutant, however, was also the most difficult to predict in spatial CV. All pollutants showed good temporal CV results and proper in-sample calibration. Exposure surfaces for 2011 and 2022 highlighted an overall decreasing trend whilst preserving the same high-concentration hotspots. These quantile-based exposure surfaces may support decision-making and subsequent epidemiological studies.
Cylindrical time series, obtained from the observation over time of a variable measured as an angle on the unit circle and a real-valued variable, often exhibit patterns consistent with the presence of multiple regimes that alternate randomly over time. In this paper, we address the problem of forecasting future observations of such series. To this end, we consider models specifically designed to account for both the existence of multiple regimes and serial correlation within each regime. The models investigated belong to two main classes: hidden Markov models and threshold autoregressive models. For the former, we explore two alternative joint distributions for the circular and linear components at each time point and propose several estimation strategies. For the latter, we consider two types of partitions associated with the definition of the thresholds, based on both variables. The proposed methods are illustrated through two applications: one involving wind direction and speed measured at a given site, and another concerning the direction and velocity of insect movement.
Accurately predicting microalgal biofuel yield is critical for optimizing commercial bio-oil production and establishing sustainable alternative energy systems. While existing machine learning approaches often rely on narrow, isolated experimental datasets or default algorithmic parameters, this study introduces a novel approach by systematically benchmarking advanced global hyperparameter optimization strategies against a large-scale, heterogeneous database to maximize predictive robustness and interpretability. This study aims to develop a highly robust hybrid machine learning framework to map complex environmental and cultivation parameters to volumetric biofuel yield (L/m3 Culture/day) and to identify the most effective optimization strategy. A comprehensive dataset of 2950 empirical observations, aggregated and curated from diverse peer-reviewed literature sources, was utilized, encompassing inputs such as light intensity, nutrient concentrations, cultivation time, temperature, and pH. A Gradient Boosting Decision Tree (GBDT) algorithm was employed as the primary predictive engine, and its hyperparameters were rigorously tuned using four distinct global optimization algorithms: Gaussian Process Optimizer (GPO), Evolutionary Strategies (ES), Bayesian Probability Improvement (BPI), and Bayesian Batch Optimizer (BBO), evaluated via a fivefold cross-validation framework. Performance assessment revealed that while the GPO offered the fastest computational runtime, the ES optimization yielded the most accurate and generalizable predictive model. It achieved a superior testing coefficient of determination (R2) of 0.956 and the lowest mean squared error (MSE) of 2.000; given typical microalgal yield ranges, this low error magnitude represents a highly acceptable deviation for preliminary industrial bioprocess screening and decision-making. SHapley Additive exPlanations (SHAP) analysis further elucidated that cultivation time and species type are the most dominant parameters dictating yield magnitude, while nutrient availability and light intensity exhibit direct positive correlations with lipid accumulation. The developed hybrid framework successfully captures the complex non-linear biochemical dynamics of microalgal cultivation, offering a highly practical computational tool for scaling and monitoring sustainable biological energy systems.
Accurate multi-step air temperature forecasting in climatologically complex inland water environments remains a critical challenge for environmental monitoring and climate adaptation. This study benchmarks nine forecasting architectures (Naive Persistence, SARIMAX, Mamba, Kolmogorov–Arnold Networks, LightGBM, GRU, LSTM, TFT, and PatchTST) for multi-horizon hourly near-surface air temperature prediction over Lake Van, the world’s largest soda lake. Models were evaluated across five horizons (t + 1 to t + 168 h) using a rolling-origin framework and five metrics (MAE, RMSE, sMAPE, NSE, R2), with pairwise Diebold–Mariano (DM) tests to assess statistical significance. LightGBM achieved the lowest mean MAE and RMSE at all horizons (MAE: 1.078–3.001; NSE: 0.867–0.973), confirming the advantage of gradient-boosted ensembles on tabular environmental datasets. The results of the Diebold–Mariano (DM) tests indicate, however, that while LightGBM is superior in terms of average errors for each of the given standard metrics, it does not produce as consistently low an error variance as KAN did at all horizons. This result indicates that LightGBM’s superiority to the other models was based on both its average errors (point metric), which are often used to rank forecasts, and its lower variance, which is typically evaluated with a distributional test like the DM. Thus, these two different measures of forecast performance assess different, but equally important, aspects of how well a model can make predictions about future data. LightGBM minimizes average error while KAN delivers more consistently accurate predictions across the test period, demonstrating that relying on a single metric may lead to suboptimal model selection. Mamba demonstrated competitive short-horizon skill (NSE = 0.928 at t + 1) while remaining statistically comparable to LightGBM at extended lead times. LSTM substantially outperformed GRU, and PatchTST consistently outperformed TFT at all horizons. SARIMAX failed across all horizons (sMAPE > 191
Drought, the second most lethal natural hazard after flood, is frequently occurring in different parts of the globe with varying characteristics, duration, frequency, severity, intensity and peak. This study investigated future shifts in drought characteristics over the biggest province in the south-west of Pakistan. The data for the duration of 1985–2014 were used as reference for evaluating the performance of ensemble data in comparison to observed data. Ensemble data was developed by using the K-fold cross-validation technique. To evaluate projected shifts in drought characteristics, future climate projections were stratified into three-time horizons: F1 (near future): 2017–2044, F2 (mid future): 2045–2072, F3 (far future): 2073–2100. The assessment was performed under two Shared Socioeconomic Pathway (SSP) scenarios using an ensemble of climate simulations from the Coupled Model Intercomparison Project Phase 6 (CMIP6). The SPEI (Standardized Precipitation and Evotranspiration Index) has been used to extract drought characteristics under various drought categories. To assess future shifts in drought characteristics, percentage changes in projected drought metrics relative to the baseline were computed at each meteorological station. The results were spatially interpolated across the study domain, and maps were generated to assess spatiotemporal patterns. Maximum decrease (60
The Cau River Basin in Vietnam has recently faced increasingly frequent and severe floods, underscoring the urgent need for improved hydrological prediction and management strategies. To tackle challenges posed by limited observational data, this study proposes a mutual learning framework to enhance deep learning model performance. This innovative approach enables paired models to exchange knowledge, improving generalization, stability, and predictive accuracy. Comparative experiments show that the proposed mutual learning model substantially outperforms machine learning techniques (i.e., Linear Regression, SVM, Random Forest, and Decision Tree) as well as deep learning models (i.e., DLinear, iTransformer, PatchTST, TiDE, TimeXer, TSMixer, WPMixer, MLP, CNN, and LSTM). Using streamflow data from 1997 to 2013, the mutual learning model (MLP-MLP) had the highest performance with an NSE of 0.75, surpassing baseline models. Sensitivity analysis showed that temperature and precipitation are both essential predictors with precipitation having a stronger impact on model performance. The evaluation of input window sizes indicated that Window 5 yields the optimal configuration, while larger windows reduce accuracy. These results demonstrate the potential of mutual learning for improving streamflow prediction in data-scarce regions and highlight its applicability for hydrological forecasting.
Mitigating carbon emissions from industrial sectors necessitates a deeper comprehension of the socio-economic and energy determinants influencing sectoral carbon intensity. This research establishes a machine learning framework to analyze and explain the factors influencing value-added carbon intensity across several economic sectors. A panel dataset including 83 nations and 11 industrial sectors from 2000 to 2021 was utilized to construct sector-specific predictive models employing Extreme Gradient Boosting (XGBoost). To preserve the temporal structure of the panel dataset, a chronological train–validation–test split was implemented, while lagged explanatory variables and lagged carbon intensity terms were incorporated to capture temporal persistence. Model interpretability was further evaluated using SHapley Additive exPlanations (SHAP). The models included economic, energy, and structural indicators, and model performance was evaluated using root mean square error (RMSE) and the coefficient of determination (R2), alongside benchmarking against ordinary least squares linear regression (OLS) and Random Forest (RF) models. The findings indicate that XGBoost attained high predictive accuracy across all sectors, with test R2 values ranging from approximately 0.80 to 0.95. Although linear regression was competitive in certain domains, the machine learning methodology offered enhanced adaptability for identifying nonlinear associations among predictors. The SHAP analysis demonstrated that lagged carbon intensity consistently served as the most significant predictor across sectors, signifying robust temporal persistence in emissions intensity. Indicators of economic activity, including GDP per capita, industry value added, and per capita power consumption, were recognized as significant determinants of sectoral carbon intensity. These findings underscore the significance of economic structure and energy consumption patterns in influencing industrial emissions performance. The findings provide insights that can facilitate targeted decarbonization initiatives and guide policy formulation to mitigate industrial emissions while preserving economic output.
Ambient air pollution measurements from regulatory monitoring networks are routinely used to support epidemiologic studies and environmental policy decision-making. However, regulatory monitors are spatially sparse and preferentially located in areas with large populations. Numerical air pollution model output can be leveraged into the inference and prediction of air pollution data combining with measurements from monitors. Nonstationary covariance functions allow the model to adapt to spatial surfaces whose variability changes with location like air pollution data. In the paper, we employ localized covariance parameters learned from the numerical output model to knit together into a global nonstationary covariance, to incorporate in a fully Bayesian model. We model the nonstationary structure in a computationally efficient way to make the Bayesian model scalable.
This study investigates the forecasting of hourly relative humidity (RH) using temperature as an exogenous variable, based on high-frequency environmental data collected from an IoT-enabled weather station installed in a vineyard in Italy. For this purpose, we fit and compare the predictive performance of classical autoregressive moving average (ARMA) models tailored for random variables bounded within the standard unit interval, so-called unit ARMA models, and the classical autoregressive integrated moving average (ARIMA) models. Advances in the unit ARMA literature include the beta ARMA ( β ARMA), Kumaraswamy ARMA (KARMA), unit Burr XII ARMA (UBXII-ARMA), and unit Weibull ARMA (UWARMA) models. All models were fitted with temperature as a covariate, which showed a significant–predominantly negative–effect across all model classes. Among the fitted models, UWARMA with temperature emerged as the most competitive, combining strong forecasting performance with theoretical consistency for modeling bounded variables. These results highlight the suitability of unit ARMA models for accurate RH forecasting in environmental monitoring applications.
In several fields such as ecological, climate, and environmental studies, modeling three-dimensional (3D) spatial heterogeneity, including altitude, is essential for capturing complex geographic variation. Geographically and altitudinal weighted regression (GAWR) is widely used for this purpose, extending traditional geographically weighted regression into three dimensions. However, GAWR cannot perform local variable selection—it assumes all predictors are relevant at every location - which can make the resulting models less accurate and difficult to interpret when true local sparsity exists. To address this limitation, we propose L_0 -GAWR, a novel framework that integrates GAWR with an L_0 -norm penalty for strict location-specific variable selection. This framework simultaneously models 3D spatial effects, including altitude, while identifying regionally relevant predictors in a data-driven manner. Since optimizing the L_0 penalty involves a computationally challenging NP-hard problem, we employ an iterative adaptive splicing algorithm to search for near-optimal solutions with stable convergence efficiently. Simulation studies demonstrate that L_0 —GAWR outperforms conventional GAWR in estimation accuracy and interpretability by correctly identifying zero coefficients. Application to South Korean air temperature data shows that the proposed model effectively captures altitude-driven spatial variations and produces sparse, interpretable coefficient estimates, demonstrating its strength in analyzing complex 3D spatial heterogeneity.
Compound temperature–precipitation shocks shape hydro-climatic impacts in semi-arid regions but are not fully represented by linear correlation. Using monthly observations from Tulcea, Dobrogea (1965–2019), we quantify contemporaneous dependence by isolating independent and identically distributed innovation shocks and modeling their joint distribution with copulas. Temperature is filtered with a periodic autoregressive moving average of orders 2 and 1 and a periodic Generalized Autoregressive Conditional Heteroskedasticity variance model, and fitted with a Student-t marginal. For precipitation, we use a periodic autoregressive of first-order model with a Generalized Extreme Value marginal. Probability-integral transforms (PIT) yield uniform innovations. Candidate copulas are estimated by maximum likelihood, selected via Akaike/Bayesian criteria (AIC/BIC), and checked with White’s information-matrix test. The Frank copula is selected for all four compound configurations (warm–dry, cold–wet, warm–wet, cold–dry), implying a zero asymptotic tail dependence but non-trivial finite quantile ( t̃∈( 0,1)) co-occurrence. At t̃ = 0.90 (the 90th percentile), the warm–dry and cold–wet co-occurrence ( J̃ =0.0196) is nearly double the independence ( B̃ =0.0100) expectation (a 97 R̃ =1.969). By contrast, warm–wet and cold–dry are suppressed ( J̃ =0.0015 and 0.0030) and co-occur well below independence (an 85 and 70 R̃ = 0.151 and R̃ = 0.303, respectively). The approach yields station-specific multipliers for compound event risk and a generalizable procedure that disentangles within-series dynamics from bivariate dependence.
Monitoring dissolved oxygen (DO) levels is crucial for effective water management and conservation strategies. However, monitoring the water quality for stratified DO using in situ sensors requires substantial time and operational costs. Data-driven approaches, particularly machine learning (ML) models, offer a promising and cost-efficient alternative for predicting DO dynamics in aquatic systems. This research develops data-driven prediction models for stratified DO in Lake Maninjau using single-target regression (STR) and multiple-target regression (MTR) approaches. The results demonstrate that the STR approach, which leverages a long-term temporal dataset, yields superior predictive performance compared with the MTR approach that relies on instantaneous, concurrently recorded observations. In the STR framework, near-surface DO is predicted from multilayer water temperature profiles, whereas in the MTR framework, stratified DO at multiple depths is estimated from a set of multilayer water quality variables. Both approaches evaluate five regression models: multilinear, polynomial, support vector, random forest, and extreme gradient-boosting regression. To address the high dimensionality of the predictors, recursive feature reduction is applied to each model. The result indicates that the vertical DO structure can be effectively represented by an upper layer (0 m and 2 m depths) and a lower layer (21 m depth). However, the MTR approach exhibits reliable performance for one prediction target while failing to generalize adequately across all target depths. Validation indicates that tree-based models, random forest and extreme gradient boosting, predict near-surface DO well in STR-MTR. In MTR, multilinear regression best predicts DO at 2 m, while support vector regression (SVR) best predicts surface DO. All models identify water temperature at 2 m as the main driver in both approaches, with chlorophyll fluorescence and salinity also important near the surface in MTR. Future research should focus on the spatial-temporal interactions between these factors to improve our understanding of dissolved oxygen dynamics, which is critical for the health of aquatic ecosystems.
Species distribution models (SDMs) are vital tools in ecology and conservation. The integration of increasingly available citizen science data with planned survey data offers a significant opportunity to improve estimates of species distributions. Whilst integrated SDMs often combine presence-only and abundance data, the link between the two data types is still not well understood. This study proposes a Bayesian spatial fusion modelling framework to jointly analyse presence-only and abundance data for the African baobab in Benin. The aim was to understand and map the spatial variation in the species’ distribution. We briefly reviewed process-based models for count and point process data. We explored various data fusion strategies using Integrated Nested Laplace Approximations (INLA) and Stochastic Partial Differential Equations (SPDE) for fast Bayesian computation. The results revealed a heterogeneous baobab distribution across Benin, characterised by a spatial autocorrelation range of 34.4 km (95
Respiratory medicines are among the first lines of defence when weather and climate push vulnerable lungs past their limits. This study quantifies how atmospheric conditions shape weekly prescription volumes and examines what continued global warming implies for pharmaceutical planning in Greece. A national retail panel of prescription respiratory sales for 20 regions (2016–2023) is combined with high-resolution meteorological reanalysis to estimate two classes of models: a Spatial Lag of X panel with region and week-of-year fixed effects, and a climate-augmented fixed-effects distributed lag forecaster. The spatial specification shows that contemporaneous conditions in neighbouring regions, especially warmer temperatures and stronger winds, exert more systematic effects on local demand than purely local shocks—positive for temperature, negative for wind—with spillovers concentrated within distance bands of a few hundred kilometres. The forecasting model couples short distributed lags of climatic variables with autoregressive dynamics and attains one-year-ahead accuracy with mean absolute percentage errors around 11
In this study, we explore the utility of Generalized Ridge Generalized Additive Models (GAMs) for trend modeling and spatial prediction. Trend modeling is conducted through an iterative procedure that updates the covariance matrix of the residuals at each step. This matrix is also estimated using a simple GAM-based approach that relates the residuals at different locations through a model with a coefficient depending on the distance between each pair of points. Spatial prediction is achieved by adding to the trend model new terms whose coefficients depend on the distance of each target point to its k nearest neighbors. The performance of the proposed methodology is assessed using synthetic data. Furthermore, a case study is presented involving the estimation of the underlying spatial distribution of a heavy metal from a public dataset. In both scenarios – synthetic and real-world – a comparative analysis for spatial prediction with universal kriging is performed.
Environmental sustainability remains a central policy challenge in advanced economies, where environmental regulation and technological efficiency interact in complex ways. This study examines the relationship between environmental policy stringency (EPS), energy efficiency (EEI), and ecological footprint across OECD countries over the period of 1990–2023. While prior studies often rely on homogeneous panel estimators, this analysis explicitly accounts for slope heterogeneity using the Mean Group (MG) estimator and complements this with diagnostics and robustness checks for cross-sectional dependence. The results show that improvements in energy efficiency are consistently associated with reductions in ecological footprint across baseline, robustness, and dynamic specifications. In contrast, the effect of environmental policy stringency is sensitive to model specification and country characteristics. Regime-based analysis shows that EPS is insignificant in low-efficiency countries but becomes significant in high-efficiency regimes. Income-based GDP specifications provide partial evidence consistent with an Environmental Kuznets Curve, although the turning point varies across countries. Overall, the findings highlight the importance of cross-country heterogeneity and identify energy efficiency as a key channel for reducing environmental pressure in OECD economies.