We develop a new targeted maximum likelihood estimation method that provides improved forecasting for misspecified linear autoregressive models. The method weighs data points in the observed sample and is useful in the presence of data generating processes featuring structural breaks, complex nonlinearities, or other time-varying properties which cannot be easily captured by model design. Additionally, the method reduces to classical maximum likelihood when the model is well specified, which results in weights which are set uniformly to one. We show how the optimal weights can be set by means of a cross-validation procedure. In a set of Monte Carlo experiments we reveal that the estimation method can significantly improve the forecasting accuracy of autoregressive models. In an empirical study concerned with forecasting the U.S. Industrial Production, we show that the forecast accuracy during the Great Recession can be significantly improved by giving greater weight to observations associated with past recessions. We further establish that the same empirical finding can be found for the 2008-2009 global financial crisis, for different macroeconomic time series, and for the COVID-19 recession in 2020.
We propose a new treatment of nonlinear regression with serially correlated disturbances that incorporates autoregressive moving average structures into feedforward neural networks. The resulting model provides an alternative to modeling temporal dependence using lagged variables. In simulations, the proposed method accurately recovers regression functions of varying complexity and the underlying error dynamics across a range of time-series lengths and signal-to-noise ratios. Finite-sample properties and out-of-sample predictive performances are shown to be robust to model misspecification induced by omitted lagged variables and incorrect specification of the error dynamics. Cloud cover is an important factor in climate projections. In an empirical study of cloud cover prediction for a grid of locations within and around the Mediterranean Sea, our proposed model yields more accurate predictions than existing methods, including long short-term memory networks. Improvements are observed broadly and are particularly pronounced in mountain areas relative to linear models with serially correlated errors, consistent with the presence of stronger nonlinear effects in cloud composure in such regions.
In this study, we employ a newly developed time series econometric approach to investigate the development in crime rates in various members states of the European Union (EU) between 1968 and 2019. We propose a panel data model with stochastically time-varying factors that also includes country-specific effects. This model enables us to evaluate the existence of a common EU crime trend, including a crime drop, to describe how individual countries depart from this common trend, and to estimate its association with macroeconomic and demographic explanatory variables. To have an equivocal measure of crime over the countries for the period of interest, we use homicide rates based on the Mortality Database from the World Health Organization. Results confirm the presence of a crime drop in the EU, be it stronger in Western EU countries than in Eastern EU countries. We also find that economic conditions explain a small portion of the crime trends in the EU; with macroeconomic activity (economic growth) being more relevant for Eastern EU countries, and macroeconomic performance (welfare growth) for Western EU countries. The young adult ratio (share of 25- to 34-year-olds in the total population) substantially explains the crime trend and drop in Western EU countries only. Our findings illustrate how the new model can be used to analyze the trends in crime, the fit from explanatory variables, and the differences in countries.
We develop a score-driven time-varying parameter model where no particular parametric error distribution needs to be specified. The proposed method relies on a versatile spline-based density, that produces a score function in the form of a natural cubic spline. This flexible approach nests the Gaussian density as a special case. It can also represent asymmetric and leptokurtic densities that produce outlier-robust updating functions for the time-varying parameter and that are often used in empirical applications. As leading examples, we consider models where the time-varying parameters appear in the location or in the log-scale of the observations. The static parameter vector of the model can be estimated by the maximum likelihood method and we formally establish some of the asymptotic properties of such estimators. We illustrate the practical relevance of the proposed method in two empirical studies. We employ the location model to filter the mean of the U.S. monthly CPI inflation series and the scale model for volatility filtering of the full panel of daily stock returns from the S P 500 index. The results show a competitive performance of the method compared to a set of competing models that are available in the existing literature.
Ethane is the most abundant non-methane hydrocarbon in Earth’s atmosphere and acts as an indirect greenhouse gas, influencing the atmospheric lifetime of methane. Therefore, understanding the development of trends and identifying trend reversals in atmospheric ethane is crucial. Ethane abundance is measured at different ground-based stations worldwide using Fourier transform infrared remote sensing techniques. We compile a new dataset comprising 26 ethane time series from the Northern and Southern Hemispheres. We analyze their long-term trends using different econometric techniques capable of handling missing data and strong seasonal components present in the data. The resulting trend patterns are consistent across the different methods, with similar estimated trends at the various stations. In the Northern Hemisphere, the common trend across stations declined from the 1990s to 2005, gradually increased over the next decade, and then resumed a similar downward trajectory from 2015 onward. The estimated trends reveal a pronounced peak around 2014/2015, marking a reversal from an upward to a downward trend.
The Global Carbon Budget (GCB), the community reference dataset for the carbon cycle, is reissued annually. The 2025 release introduces several adjustments to the published series that we compare with prior releases starting in 2017. On a common 1959–2016 sample, the mean of the GCB budget imbalance jumps from within ± 0.17 GtC/yr of zero for every vintage 2017–2024 to +0.61 GtC/yr in 2025, the only vintage whose 95% confidence interval for the imbalance mean excludes zero. We document and explore this shift in two ways. First, we conduct a model-free analysis, where we attribute the shift to a new adjustment that places the published land sink 0.40 GtC/yr below its ensemble mean (the average of the underlying models), a smaller adjustment in the ocean sink in the opposite direction, and the removal of one model from the bookkeeping ensemble. Second, we consider the dynamic statistical GCB model of , augmented with climate covariates. Its parameters are estimated for every GCB vintage 2017–2025. The coefficients of atmospheric concentrations in the sink equations shift sharply on the 2025 issue in opposite directions, mirroring the model-free findings. There is a persistent drifting imbalance across the entire sample in the budget equation. In case the intended effect of the adjustments to the 2025 vintage is to reduce the mean of the budget imbalance on the window of the last ten years, our results show that this comes at the cost of increased budget imbalance over the whole sample and inconsistency of the data record. We argue that the costs of the adjustments outweigh the benefits of a narrow view on the last ten years and are detrimental to statistical analysis of the full GCB sample.
We propose a new statistical reduced complexity climate model. The centerpiece of the model consists of a set of physical equations for the global climate system which we show how to cast in non-linear state space form. The parameters in the model are estimated using the method of maximum likelihood with the likelihood function being evaluated by the extended Kalman filter. Our statistical framework is based on well-established methodology and is computationally feasible. In an empirical analysis, we estimate the parameters for a data set comprising the period 1959-2022. A likelihood ratio test sheds light on the most appropriate equation for converting the level of atmospheric concentration of carbon dioxide into radiative forcing. Using the estimated model, and different future paths of greenhouse gas emissions, we project global mean surface temperature until the year 2100. Our results illustrate the potential of combining statistical modelling with physical insights to arrive at rigorous statistical analyses of the climate system.
This study presents the extremum Monte Carlo filter as a data assimilation method. The method can be regarded as a variant of the variational approach (three‐ and four‐dimensional variational), where the state estimates are obtained by solving an optimization problem numerically over a space of prediction functions, instead of the state space itself. We discuss the general principle of the new technique and its use of machine‐learning methods, such as feed‐forward neural networks and tree‐based gradient boosting. We further provide the details for its computationally efficient implementation, in both computing time and storage. The extremum Monte Carlo filter is illustrated for the well‐known Lorenz‐63 and Lorenz‐96 models. A performance improvement is found relative to the ensemble Kalman filter for the Lorenz‐96 model.
This study presents a statistical time-domain approach for identifying transitions between climate states, referred to as breakpoints, using well-established econometric tools. We analyze a 67.1 million year record of the oxygen isotope ratio delta-O-18 derived from benthic foraminifera. The dataset is presented in Westerhold et al. (2020), where the authors use recurrence analysis to identify six climate states. Fixing the number of breakpoints to five, our procedure results in breakpoint estimates that closely align with those identified by Westerhold et al. (2020). By treating the number of breakpoints as a parameter to be estimated, we provide the statistical justification for more than five breakpoints in the time series. Further, our approach offers the advantage of constructing confidence intervals for the breakpoints, and it allows for testing the number of breakpoints present in the time series.
In economics and finance, speculative bubbles take the form of locally explosive dynamics that eventually collapse. We propose a test for the presence of speculative bubbles in the context of mixed causal-noncausal autoregressive processes. The test exploits the fact that bubbles are anticipative, that is, they are generated by an extreme shock in the forward-looking dynamics. In particular, the test uses both path-level deviations and growth rates to assess the presence of a bubble of a given duration and size, at any moment in time. We show that the distribution of the test statistic can be either analytically determined or numerically approximated, depending on the error distribution. Size and power properties of the test are analyzed in controlled Monte Carlo experiments. An empirical application is presented for a monthly oil price index. It demonstrates the ability of the test to detect bubbles and to provide valuable insights in terms of risk assessments in the spirit of Value-at-Risk.
This study presents a statistical time-domain approach for identifying transitions between climate states, referred to as breakpoints, using well-established econometric tools. Our approach offers the advantage of constructing time-domain confidence intervals for the breakpoints, and it includes procedures to determine how many breakpoints are present in the time series. We apply these tools to a 67.1 million-year-long compilation of benthic foraminiferal oxygen isotopes (delta 18O), which signify global temperature and ice volume throughout the Cenozoic. This foundational dataset is presented in , where the authors use recurrence analysis to identify five breakpoints that define six climate states. Fixing the number of breakpoints to five, our procedure results in breakpoint estimates that closely align with those identified by . By allowing the number of breakpoints to vary, we provide statistical justification for more than five breakpoints in the time series. Our method adds to our understanding of Cenozoic climate history in terms of the timing and rate of transitions between climate states and provides a tool for robustly assessing breakpoints in many other paleoclimate time series.
Time series analysis of delta-O-18 and delta-C-13 measurements from benthic foraminifera for purposes of paleoclimatology is challenging. The time series reach back tens of millions of years, they are relatively sparse in the early record and relatively dense in the later, the time stamps of the observations are not evenly spaced, and there are instances of multiple different observations at the same time stamp. The time series appear non-stationary over most of the historical record with clearly visible temporary trends of varying directions. In this paper, we propose a continuous-time state-space framework to analyze the time series. State space models are uniquely suited for this purpose, since they can accommodate all the challenging features mentioned above. We specify univariate models and joint bivariate models for the two time series of delta-O-18 and delta-C-13. The models are estimated using maximum likelihood by way of the Kalman filter recursions. The suite of models we consider has an interpretation as an application of the Butterworth filter. We propose model specifications that take the origin of the data from different studies into account and that allow for a partition of the total period into sub-periods reflecting different climate states. The models can be used, for example, to impute evenly time-stamped values by way of Kalman filtering. They can also be used, in future work, to analyze the relation to proxies for CO2 concentrations.
We introduce a novel simulation-based method for signal extraction in a general class of state space models. It can be used to estimate time-varying conditional means, modes, and quantiles, and to predict latent variables or forecast observations. The method consists of generating artificial datasets from the model and estimating the quantities of interest via extremum estimation. The approach is broadly applicable and its implementation is straightforward. The method is suited for signal extraction in cases of long time series, missing data, or high-dimensionality. Furthermore, we demonstrate its use in real-time filtering, where most of the computations can be performed in advance, and in fixed-interval smoothing. Conditions for the stability and convergence of the filtering method are discussed, and its key properties are illustrated by various applications, including nonlinear and high-dimensional models.
This article introduces conditional score residuals and proposes a general framework for the diagnostic analysis of time series models. Conditional score residuals encompass commonly used definitions of residuals in time series models, including ARMA residuals, squared residuals, and Pearson residuals. In particular, these residuals are special cases of conditional score residuals when the conditional distribution of the model belongs to the exponential family. On the other hand, conditional score residuals offer an alternative definition of residuals when the conditional distribution is not of the exponential type. A key feature of conditional score residuals is that they account for the shape of the conditional distribution. This feature leads to more reliable and powerful diagnostic tools for testing residual autocorrelation. Furthermore, they can be employed in complex models where it may not be clear how to define residuals. The asymptotic properties of the empirical autocorrelation function of conditional score residuals are formally derived. The practical relevance of the proposed framework is illustrated for heavy-tailed GARCH models. Monte Carlo and empirical results support the finding that conditional score residuals are more reliable in testing residual autocorrelation, when compared to squared residuals. Finally, it is shown how a diagnostic analysis can be designed for dynamic copula models.
Multivariate unobserved components time series models can accommodate dynamic relations between dependent variables by introducing nonzero correlations between the disturbances that drive the dynamics of the components. The usual assumption of time-invariant correlations can be rather strong for many applications. In this study, we introduce a time-varying correlation parameter that allows the dynamic relations to change over time. We treat the parameter as a dynamic unobserved component. The resulting nonlinear model requires nonlinear methods for parameter estimation and filtering. For parameter estimation, we propose an indirect inference approach based on an auxiliary model using a cubic spline for the time-varying correlation. For the filtering of the time-varying correlation, a bootstrap filter is employed jointly with the technique of Rao-Blackwellization. A Monte Carlo simulation study shows that our proposed methodology is successful in both estimation and filtering. In our empirical study, we explore the strong dependence between monthly unemployment in the labour force and claimant counts. We find strong evidence that the underlying correlations are time-varying.
We introduce a new and general methodology for analyzing vector autoregressive models with time-varying coefficient matrices and conditionally heteroskedastic disturbances. The proposed approach is transparent and simple to implement. It allows the derivation of well-defined impulse response functions that rely on the overall stability of the system. We present the finite sample properties of the model in a simulation study. In an empirical illustration we investigate the possibly time-varying relationships between U.S. industrial production, inflation, and bond spread. We empirically identify a time-varying linkage between economic and financial variables which are effectively described by a common dynamic factor. The impulse response analysis identifies substantial differences in the effects of financial shocks on output and inflation during crisis and non-crisis periods. The results also illustrate how the widely-used approach of fixing the VAR coefficients in the derivation of the impulse responses leads to a sizeable underestimation of the impact of a financial shock on output and inflation during some of the crises in our sample.
The equivalence of the Beveridge–Nelson decomposition and the trend-cycle decomposition is well established. In this paper we argue that this equivalence is almost immediate when a Gaussian score-driven location model is considered. We also provide a natural extension towards heavy-tailed distributions for the disturbances which lead to a robust version of the Beveridge–Nelson decomposition.
We propose a multiplicative dynamic factor structure for the conditional modelling of the variances of an N-dimensional vector of financial returns. We identify common and idiosyncratic conditional volatility factors. The econometric framework is based on an observation-driven time series model that is simple and parsimonious. The common factor is modeled by a normal density and is robust to fat-tailed returns as it averages information over the cross-section of the observed N-dimensional vector of returns. The idiosyncratic factors are designed to capture the erratic shocks in returns and therefore rely on fat-tailed densities. Our model is potentially of a high-dimension, is parsimonious and it does not necessarily suffer from the curse of dimensionality. The relatively simple structure of the model leads to simple computations for the estimation of parameters and signal extraction of factors. We derive the stochastic properties of our proposed dynamic factor model, including bounded moments, stationarity, ergodicity, and filter invertibility. We further establish consistency and asymptotic normality of the maximum likelihood estimator. The finite sample properties of the estimator and the reliability of our method to track the common conditional volatility factor are investigated by means of a Monte Carlo study. Finally, we illustrate our approach with two empirical studies. The first study is for a panel of financial returns from ten stocks of the S&P100. The second study is for the panel of returns from all S&P100 stocks.
The global fraction of anthropogenically emitted carbon dioxide (CO2) that stays in the atmosphere, the CO2 airborne fraction, has been fluctuating around a constant value over the period 1959 to 2022. The consensus estimate of the airborne fraction is around 44%. In this study, we show that the conventional estimator of the airborne fraction, based on a ratio of changes in atmospheric CO2 concentrations and CO2 emissions, suffers from a number of statistical deficiencies. We propose an alternative regression-based estimator of the airborne fraction that does not suffer from these deficiencies. Our empirical analysis leads to an estimate of the airborne fraction over 1959-2022 of 47.0% (+/- 1.1%; 1 sigma), implying a higher, and better constrained, estimate than the current consensus. Using climate model output, we show that a regression-based approach provides sensible estimates of the airborne fraction, also in future scenarios where emissions are at or near zero.
Through its atmospheric teleconnections, El Ni & ntilde;o-Southern Oscillation (ENSO) shifts and disrupts weather and climate patterns far beyond the equatorial Pacific where it occurs often resulting in catastrophic consequences in many countries of the world. It is also the largest source of seasonal and interannual climate predictability. Despite its huge importance, ENSO forecasting is still not performed operationally at longer leads than about 6 months ahead. At the same time, there is mounting scientific evidence that forecasts are possible even more than a year in advance. Early warning of ENSO events could substantially mitigate some of the most damaging impacts, such as floods, droughts, and harvest failure and help avoid famine, migration, and disease outbreaks. Here, we present forecasts from a statistical ENSO model of the next El Ni & ntilde;o predicted to occur in the winter of 2023/24, at lead times between 11 and 17 months ahead of an expected peak in December 2023. We use a statistical unobserved dynamic components model (EDCM) based on subsurface ocean temperatures as well as sea surface temperatures and zonal wind stress. EDCM has been previously validated through hindcasts of the major El Ni & ntilde;os since 1970 and through real-time forecasts of the 2015/16 and 2018/19 El Ni & ntilde;os. Our statistical framework and results indicate that there is potential for doubling the operational predictive lead time of ENSO to at least 12 months, with additional promise for even earlier anticipation of 19 months. Such longer-lead forecasts could be of high value, because decision-making and management in a number of key socioeconomic sectors could be greatly improved.