This study evaluates the performance of 36 historical CMIP6 GCM trajectories (1979-2005) in reproducing atmospheric circulation over the Iberian Peninsula in the summer months (June-September) using the Lamb Weather Type (WT) classification scheme. Using ERA5 reanalysis as the observational reference, we introduce a methodological framework-applicable to any region worldwide-to evaluate GCM performance. This approach extends traditional daily frequency analysis by evaluating both the daily frequency distribution of WTs and their 24-hour dynamic evolution (i.e., transition probabilities and persistence). Model performance is quantified using the Overlap coefficient. A filtering process is applied where only trajectories that successfully reproduce both daily and conditional distributions with a minimum Overlap threshold t_sim across a set number of grid points are retained. The findings show that while several models can adequately reproduce daily WT frequencies (16 out of 36), some struggle to capture day-to-day atmospheric transitions. This leads to a final selection of 12 trajectories over the Iberian Peninsula. Model performance across the region is then evaluated using integrated metrics assessing daily reproduction, conditional reproduction, and transition dynamics. Overall, models from the ec earth3 family-specifically the ec earth3 aerchem trajectory-exhibit the best and most consistent performance across the region. Additionally, the results highlight a geographical performance gap: while models generally represent circulation well in the northwest, they face significant challenges in the central and southern Mediterranean regions of the Peninsula. Ultimately, this study establishes that assessing WT persistence and transitions provides a far more discriminative, objective tool for GCM selection than evaluating daily distributions alone.
Abstract. As the frequency of extreme temperature events increases, so does the need for robust tools to understand them. This work develops and applies a methodological framework to model the occurrence of Tx calendar-day records and their relationship with geopotentials. The analysis includes Tx data from 36 Spanish stations (1960–2023) and geopotentials at 300, 500, and 700 hPa. Exploratory analysis revealed a non-stationary trend in records, a higher frequency in the interior of the peninsula, and decreasing spatial co-occurrence with distance. A hierarchical spatio-temporal logistic regression algorithm prioritizing interpretability was designed. The approach involves: (1) fitting local models per station; (2) applying a spatial consensus filter to reduce initial 1620 parameters to 17 in a base model; and (3) incorporation of interaction terms. Among the tested models, a global model that enhances the base model with geodetic interactions was selected for optimal balance between predictive performance and complexity. Geopotentials at 700 hPa are most relevant for characterizing records, while 300 hPa dominates in the southern corners and 500 hPa in the northern corners. The model demonstrates high predictive accuracy at interior stations, good performance at coastal stations, and adequately reproduces the persistence of record runs and spatial co-occurrence.
The increasing frequency of extreme temperature events, such as daily maximum temperature (T_x) records, underscores the need for robust tools to understand their drivers and predict their occurrence. Previous studies have identified increasing and non-stationary trends in T_x records across the Iberian Peninsula, particularly during summer, the literature directly exploring their connection with upper-level atmospheric covariates remains limited. This work develops and applies an innovative methodological framework to model the occurrence of T_x records and their relationship with geopotential height fields. We used daily T_x data from 36 Spanish stations (1960-2023) provided by ECA D and geopotential height data at 300, 500, and 700 hPa from ERA5. Exploratory analysis revealed a non-stationary trend in records, a higher frequency in the interior of the peninsula, and decreasing spatial co-occurrence with distance. We designed a hierarchical spatio-temporal logistic regression algorithm prioritizing interpretability and high-dimensionality reduction. The approach involves: (1) fitting local models per station; (2) applying a spatial consensus filter based on statistical significance to reduce the initial 1620 covariates to 17 in a base model (M1); and (3) a controlled incorporation of interaction terms. Among the tested models, a global model (M2) that enhances M1 with geodetic interactions was selected for its optimal balance between predictive performance (AUC) and complexity. Model M2 demonstrates high predictive accuracy at interior stations and good performance at coastal stations. It also adequately reproduces key observed properties, including the persistence of record streaks and patterns of spatial co-occurrence. This study provides a novel tool for predicting upcoming record events with high accuracy while maintaining a concise and interpretable structure.
Recent findings showed that extreme events such as daily maximum temperature record-breaking are not following a stationary pattern, with trends associated to global warming [1]. But there are spatial variability identifying the change-point when this pattern is dated and, most interesting, there is not evidence of non-stationarity in North stations of Spain in spring and autumn. To this regard, it is essential to develop novel tools and models that represent the seasonal and spatial variability, and effectively capture the spatio-temporal dependence of covariates of interest. Effective detection will enable us to more accurately describe and predict which regions may be affected by stronger trends.In this framework, the utility of spatio-temporal models including relevant covariates for the occurrence of records is evident. However, achieving this objective requires the preliminary identification of the covariates and interaction terms that influence the occurrence of records. Thus, we propose a two-level generalized linear models (GLM) approach to detect spatio-temporal dependence of covariates of interest in daily maximum temperature record-breaking. To do so, we took daily maximum temperatures in the 1960-2022 period in 36 stations distributed over the peninsular Spain with low level of missing values (below 0.5%). We computed the calendar day record-breaking by binarizing the temporal series assigning a one only if a particular daily maximum on year ‘t’ is above all its previous years on the same date [2]. First, for each station a local logistic regression was applied setting the trend term as log(t-1); note that the trend term in the probability of a record, on the logit scale, is -log(t-1) under a stationary climate. Finally, all the estimated beta coefficients for the trend were gathered and correlated with spatial covariates such as latitude, longitude, altitude, and distance to the coast. We found that only log-altitude and log-distance showed a significant positive correlation with the trend coefficients, being the latter the one with a higher effect (r=0.59). Although preliminary, these results showed a straightforward approach to model the relationship between spatial covariates and the temporal trend in extreme events, in particular, record of maximum temperatures. In addition, we anticipate that this tool will be potentially useful to build models based on atmospheric covariates. References:[1] Castillo-Mateo, J., Cebrián, A. C., and Asín, J. (2023). Statistical analysis of extreme and record-breaking daily maximum temperatures in peninsular Spain during 1960–2021. Atmospheric Research, 106934. https://doi.org/10.1016/j.atmosres.2023.106934[2] Castillo-Mateo, J., Cebrián, A. C., and Asín, J. (2023). RecordTest: An R Package to Analyze Non-Stationarity in the Extremes Based on Record-Breaking Events. Journal of Statistical Software, 106, 1-28. https://doi.org/10.18637/jss.v106.i05
There is continuing interest in the investigation of change in temperature over space and time. For this analysis, we offer statistical tools to illuminate changes temporally, at desired temporal resolution, and spatially, using data generated from suitable space-time models. The proposed tools can be used with the output from any suitable model fitted to any set of spatially referenced time series data. The tools to assess space and time changes include spatial surfaces of probabilities and spatial extents for events defined by exceeding a threshold. The spatial surfaces capture the spatial variation in the probability or risk of an exceedance event, while the spatial extents capture the expected proportion of incidence of an event for a region of interest. This approach is used analyse the changes in daily maximum temperature in an inland Mediterranean region (NE of Spain) in the period 1956-2015. The area is very heterogeneous in orography and climate, including the central Ebro valley and part of the Pyrenees. We use a collection of daily temperature series obtained from simulation under a Bayesian daily temperature model fitted to 18 stations in that area. The results for the summer period show that, although there is an increasing risk in all the events used to quantify the effects of climate change, it is not spatially homogeneous, with the largest increase arising in the centre of the Ebro valley and the Eastern Pyrenees area. The risk of an increase in the average daily maximum temperature from 1966-1975 to 2006-2015 higher than 1degree celsius is higher than 0.5 over all of the region, and close to 1 in the previous areas. The extent of daily maximum temperature higher than the reference mean has increased 3.5% per decade. The mean of the extent indicates that 95% of the area under study has suffered a positive increment of the average temperature, and almost 70% an increment higher than 1degree celsius.
Regression is the most widely used modeling tool in statistics. Quantile regression offers a strategy for enhancing the regression picture beyond customary mean regression. With time-series data, we move to quantile autoregression and, finally, with spatially referenced time series, we move to spacetime quantile regression. Here, we are concerned with the spatiotemporal evolution of daily maximum temperature, particularly with regard to extreme heat. Our motivating data set is 60 years of daily summer maximum temperature data over Aragon in Spain. Hence, we work with time on two scales- days within summer season across years-collected at geocoded station locations. For a specified quantile, we fit a very flexible, mixed-effects autoregressive model, introducing four spatial processes. We work with asymmetric Laplace errors to take advantage of the available conditional Gaussian representation for these distributions. Further, while the autoregressive model yields conditional quantiles, we demonstrate how to extract marginal quantiles with the asymmetric Laplace specification. Thus, we are able to interpolate quantiles for any days within years across our study region.
In many applications, interest focuses on assessing relationships between covariates and the extremes of the distribution of a continuous response. For example, in climate studies, a usual approach to assess climate change has been based on the analysis of annual maximum data. Using the generalized extreme value (GEV) distribution, we can model trends in the annual maximum temperature using the high number of available atmospheric covariates. However, there is typically uncertainty in which of the many candidate covariates should be included. Bayesian methods for variable selection are very useful to identify important covariates. However, such methods are currently very limited for moderately high dimensional variable selection in GEV regression. We propose a Bayesian method for variable selection based on a stochastic search variable selection (SSVS) algorithm proposed for posterior computation. The method is applied to the selection of atmospheric covariates in annual maximum temperature series in three Spanish stations.
Quantile regression continues to increase in usage, providing a useful alternative to customary mean regression. Primary implementation takes the form of so-called multiple quantile regression, creating a separate regression for each quantile of interest. However, recently, advances have been made in joint quantile regression, supplying a quantile function which avoids crossing of the regression across quantiles. Here, we turn to quantile autoregression (QAR), offering a fully Bayesian version. We extend the initial quantile regression work of Koenker and Xiao (J Am Stat Assoc 101(475):980–990, 2006. https://doi.org/10.1198/016214506000000672 ) in the spirit of Tokdar and Kadane (Bayesian Anal 7(1):51–72, 2012. https://doi.org/10.1214/12-BA702 ). We offer a directly interpretable parametric model specification for QAR. Further, we offer a pth-order QAR(p) version, a multivariate QAR(1) version, and a spatial QAR(1) version. We illustrate with simulation as well as a temperature dataset collected in Aragón, Spain.
Acknowledging a considerable literature on modeling daily temperature data, we propose a multi-level spatiotemporal model which introduces several innovations in order to explain the daily maximum temperature in the summer period over 60 years in a region containing Aragón, Spain. The model operates over continuous space but adopts two discrete temporal scales, year and day within year. It captures temporal dependence through autoregression on days within year and also on years. Spatial dependence is captured through spatial process modeling of intercepts, slope coefficients, variances, and autocorrelations. The model is expressed in a form which separates fixed effects from random effects and also separates space, years, and days for each type of effect. Motivated by exploratory data analysis, fixed effects to capture the influence of elevation, seasonality, and a linear trend are employed. Pure errors are introduced for years, for locations within years, and for locations at days within years. The performance of the model is checked using a leave-one-out cross-validation. Applications of the model are presented including prediction of the daily temperature series at unobserved or partially observed sites and inference to investigate climate change comparison. Supplementary materials accompanying this paper appear online.
There is increasing evidence that global warming manifests itself in more frequent warm days and that heat waves will become more frequent. Presently, a formal definition of a heat wave is not agreed upon in the literature. To avoid this debate, we consider extreme heat events, which, at a given location, are well-defined as a run of consecutive days above an associated local threshold. Characteristics of extreme heat events (EHEs) are of primary interest, such as incidence and duration, as well as the magnitude of the average exceedance and maximum exceedance above the threshold during the EHE. Using approximately 60-year time series of daily maximum temperature data collected at 18 locations in a given region, we propose a spatio-temporal model to study the characteristics of EHEs over time. The model enables prediction of the behaviour of EHE characteristics at unobserved locations within the region. Specifically, our approach employs a two-state space–time model for EHEs with local thresholds where one state defines above threshold daily maximum temperatures and the other below threshold temperatures. We show that our model is able to recover the EHE characteristics of interest and outperforms a corresponding autoregressive model that ignores thresholds based on out-of-sample prediction.
There is increasing evidence that global warming manifests itself in more frequent warm days and that heat waves will become more frequent. Presently, a formal definition of a heat wave is not agreed upon in the literature. To avoid this debate, we consider extreme heat events, which, at a given location, are well-defined as a run of consecutive days above an associated local threshold. Characteristics of extreme heat events (EHEs) are of primary interest, such as incidence and duration, as well as the magnitude of the average exceedance and maximum exceedance above the threshold during the EHE. Using approximately 60-year time series of daily maximum temperature data collected at 18 locations in a given region, we propose a spatio-temporal model to study the characteristics of EHEs over time. The model enables prediction of the behaviour of EHE characteristics at unobserved locations within the region. Specifically, our approach employs a two-state space-time model for EHEs with local thresholds where one state defines above threshold daily maximum temperatures and the other below threshold temperatures. We show that our model is able to recover the EHE characteristics of interest and outperforms a corresponding autoregressive model that ignores thresholds based on out-of-sample prediction.
Evidence of global warming induced from the increasing concentration of greenhouse gases in the atmosphere suggests more frequent warm days and heat waves. The concept of an extreme heat event (EHE), defined locally based on exceedance of a suitable local threshold, enables us to capture the notion of a period of persistent extremely high temperatures. Modeling for extreme heat events is customarily implemented using time series of temperatures collected at a set of locations. Since spatial dependence is anticipated in the occurrence of EHE’s, a joint model for the time series, incorporating spatial dependence is needed. Recent work by Schliep et al. (J R Stat Soc Ser A Stat Soc 184(3):1070–1092, 2021) develops a space-time model based on a point-referenced collection of temperature time series that enables the prediction of both the incidence and characteristics of EHE’s occurring at any location in a study region. The contribution here is to introduce a formal definition of the notion of the spatial extent of an extreme heat event and then to employ output from the Schliep et al. (J R Stat Soc Ser A Stat Soc 184(3):1070–1092, 2021) modeling work to illustrate the notion. For a specified region and a given day, the definition takes the form of a block average of indicator functions over the region. Our risk assessment examines extents for the Comunidad Autónoma de Aragón in northeastern Spain. We calculate daily, seasonal and decadal averages of the extents for two subregions in this comunidad. We generalize our definition to capture extents of persistence of extreme heat and make comparisons across decades to reveal evidence of increasing extent over time.
Point processes are often used to model the occurrence times of different phenomena, such as heatwaves or spike trains. Many of those problems require to study the independence between nonhommogeneous point processes in time, and this work develops three families of tests to assess that hypothesis. They can be applied to different types of processes, and all together they cover a wide range of situations appearing in real problems. The first family includes two tests for Poisson processes. The second family is based on the close point distance, and the third one on cross dependence functions. An extensive simulation study of the size and power of the tests is carried out and some practical rules to select the most appropriate test in different cases, are provided. The proposed tests are demonstrated on a real data application about the occurrence of extreme heat events in three Spanish locations.
En este trabajo se compara la prediccion de ocurrencia de lluvia apreciable (mayor o igual que 0.1 litros/m2) obtenida con un modelo de regresion logistica que usa la informacion registrada en un observatorio, con las previsiones del modelo Hirlam-02, ambas con un horizonte de prediccion de 6 horas. Se ha construido un modelo util para cualquier dia del ano y modelos especificos para las tres epocas (fria, templada y calida) en las que se ha dividido este. Los modelos especificos son mas simples y faciles de interpretar pero no producen mejoras notables en la prediccion. La comparacion se realiza con los datos y en la celda en la que se ubica el observatorio del aeropuerto de Zaragoza. La prediccion que produce directamente HIRLAM-02 no corresponde con el regimen observado, ni en los valores medios ni en el ciclo anual, y la prevision de ocurrencia es peor que la del modelo estadistico. Se analiza tambien la prediccion que produce HIRLAM-02 utilizando umbrales distintos para cada una de las tres epocas consideradas, lo que le permite alcanzar porcentajes de dias correctamente predichos superiores a los del modelo estadistico. [EN]In this work, the prediction of rainfall occurrence (rainfall higher than 0.1 mm/m2) obtained from a logistic regression model based on data measured in an observatory is compared with the forecast obtained from the Hirlam-02 model, both of them using a 6 hours forecasting horizon. We have built a global model, useful for any day, and specific models for three periods defined in the year: cold, hot and mild. The specific models are simpler and easier to understand but they do not produce a noticeable improvement of forecasting results. The comparison is performed using data from the Zaragoza airport observatory and from the HIRLAM cell where it is located. The HIRLAM-02 forecast is not able to reproduce the observed average or the annual cycle and the forecasting results are worse than those obtained from the statistical model. When a different threshold is applied in each of the periods considered, the percentage of success of HIRLAM-02 forecast is highly improved and overcomes those of the statistical model.
This work proposes a new statistical modelling approach to forecast the hourly river level at a gauging station, under potential flood risk situations and over a medium-term prediction horizon (around three days). For that aim we introduce a new model, the switching regression model with ARMA errors, which takes into account the serial correlation structure of the hourly level series, and the changing time delay between them. A whole modelling approach is developed, including a two-step estimation, which improves the medium-term prediction performance of the model, and uncertainty measures of the predictions. The proposed model not only provides predictions for longer periods than other statistical models, but also helps to understand the physics of the river, by characterizing the relationship between the river level in a gauging station and its influential factors. This approach is applied to forecast the Ebro River level at Zaragoza (Spain), using as input the series at Tudela. The approach has shown to be useful and the resulting model provides satisfactory hourly predictions, which can be fast and easily updated, together with their confidence intervals. The fitted model outperforms the predictions from other statistical and numerical models, specially in long prediction horizons.
The occurrence of extreme heat events in maximum and minimum daily temperatures is modelled using a non homogeneous common Poisson shock process. It is applied to five Spanish locations, representative of the most common climates over the Iberian Peninsula. The model is based on an excess over threshold approach and distinguishes three types of extreme events: only in maximum temperature, only in minimum temperature and in both of them (simultaneous events). It takes into account the dependence between the occurrence of extreme events in both temperatures and its parameters are expressed as functions of time and temperature related covariates. The fitted models allow us to characterize the occurrence of extreme heat events and to compare their evolution in the different climates during the observed period. This model is also a useful tool for obtaining local projections of the occurrence rate of extreme heat events under climate change conditions, using the future downscaled temperature trajectories generated by Earth System Models. The projections for 2031-60 under scenarios RCP4.5, RCP6.0 and RCP8.5 are obtained and analysed using the trajectories from four earth system models which have successfully passed a preliminary control analysis. Different graphical tools and summary measures of the projected daily intensities are used to quantify the climate change on a local scale. A high increase in the occurrence of extreme heat events, mainly in July and August, is projected in all the locations, all types of event and in the three scenarios, although in 2051-60 the increase is higher under RCP8.5. However, relevant differences are found between the evolution in the different climates and the types of event, with a specially high increase in the simultaneous ones.
The main contribution of this work is a bootstrap test to check the independence between temporal nonhomogeneous Poisson processes. The test statistic is based on the close point relation, which adapts the crossed nearest neighbour distance ideas of spatial point processes to the case of nonhomogeneous time point processes. Since it is complicated to obtain the probability distribution of the test statistic under the null hypothesis and the parameters of the processes are usually unknown, the p value is obtained using a parametric bootstrap approach. A simulation study shows that the size of the test is close to the nominal one. The power is analyzed considering three approaches for generating dependent nonhomogeneous processes and different levels of dependence, and satisfactory results are obtained in all cases. Although the test was initially intended for Poisson processes, it can be applied to any type of point process in one dimension which can be simulated. This test is a valuable tool in the validation analysis of common Poisson shock models. For the bivariate case, the process can be decomposed into three independent Poisson processes, and the assumption of independence between them has to be checked. As an application, the joint modeling of the occurrence process of extreme heat events in daily maximum and minimum temperatures is described.
NHPoisson is an R package for the modeling of nonhomogeneous Poisson processes in one dimension. It includes functions for data preparation, maximum likelihood estimation, covariate selection and inference based on asymptotic distributions and simulation methods. It also provides specific methods for the estimation of Poisson processes resulting from a peak over threshold approach. In addition, the package supports a wide range of model validation tools and functions for generating nonhomogenous Poisson process trajectories. This paper is a description of the package and aims to help those interested in modeling data using nonhomogeneous Poisson processes.
A joint model is proposed for analyzing and predicting the occurrence of extreme heat events in two temperature series, these being daily maximum and minimum temperatures. Extreme heat events are defined using a threshold approach and the suggested model, a non-homogeneous common Poisson shock process, accounts for the mutual dependence between the extreme events in the two series. This model is used to study the time evolution of the occurrence of extreme events and its relationship with temperature predictors. A wide range of tools for validating the model is provided, including influence analysis. The main application of this model is to obtain medium-term local projections of the occurrence of extreme heat events in a climate change scenario. Future temperature trajectories from general circulation models, conveniently downscaled, are used as predictors of the model. These trajectories show a generalized increase in temperatures, which may lead to extrapolation errors when the model is used to obtain projections. Various solutions for dealing with this problem are suggested. The results of the fitted model for the temperature series in Barcelona in 1951–2005 and future projections of extreme heat events for the period 2031–2060 are discussed, using three global circulation model trajectories under the SRES A1B scenario.