Bayesian cross-validation (CV) is a popular method for predictive model assessment that is simple to implement and broadly applicable. A wide range of CV schemes is available for time series applications, including generic leave-one-out (LOO) and K-fold methods, as well as specialized approaches intended to deal with serial dependence such as leave-future-out (LFO), h-block, and hv-block. Existing large-sample results show that both specialized and generic methods are applicable to models of serially-dependent data. However, large sample consistency results overlook the impact of sampling variability on accuracy in finite samples. Moreover, the accuracy of a CV scheme depends on many aspects of the procedure. We show that poor design choices can lead to elevated rates of adverse selection. In this paper, we consider the problem of identifying the regression component of an important class of models of data with serial dependence, autoregressions of order p with q exogenous regressors (ARX(p,q)), under the logarithmic scoring rule. We show that when serial dependence is present, scores computed using the joint (multivariate) density have lower variance and better model selection accuracy than the popular pointwise estimator. In addition, we present a detailed case study of the special case of ARX models with fixed autoregressive structure and variance. For this class, we derive the finite-sample distribution of the CV estimators and the model selection statistic. We conclude with recommendations for practitioners.
Cross-validation (CV) is a widely-used method of predictive assessment based on repeated model fits to different subsets of the available data. CV is applicable in a wide range of statistical settings. However, in cases where data are not exchangeable, the design of CV schemes should account for suspected correlation structures within the data. CV scheme designs include the selection of left-out blocks and the choice of scoring function for evaluating predictive performance. This paper focuses on the impact of two scoring strategies for block-wise CV applied to spatial models with Gaussian covariance structures. We investigate, through several experiments, whether evaluating the predictive performance of blocks of left-out observations jointly, rather than aggregating individual (pointwise) predictions, improves model selection performance. Extending recent findings for data with serial correlation (such as time-series data), our experiments suggest that joint scoring reduces the variability of CV estimates, leading to more reliable model selection, particularly when spatial dependence is strong and model differences are subtle.
Dynamic factor models are widely-used to characterize high-dimensional time series data. While many applications are based on reducing the model into a minimal number of factors, empirical evidence suggests many factors may be required to capture the complex dynamics observed in the data. However, permitting a large number of factors will often lead to statistical problems, particularly in limited data contexts. In this article, we develop a new sparse dynamic factor model that substantially reduces the number of parameters, and caters for high-dimensional time series data containing heterogeneous patterns including clustering. In particular, our method provides a data-driven classification of the clustered high-dimensional time series with a single-factor model within each class. While each factor drives its own intra-cluster co-movement, the joint evolution of the factors induces inter-cluster dependency. In an application to French mortality data, our proposed model demonstrates superior forecasting performances when compared to a class of dynamic factor models, providing valuable insights for longevity risk management.
Brute force cross-validation (CV) is a method for predictive assessment and model selection that is general and applicable to a wide range of Bayesian models. Naive or ‘brute force’ CV approaches are often too computationally costly for interactive modeling workflows, especially when inference relies on Markov chain Monte Carlo (MCMC). We propose overcoming this limitation using massively parallel MCMC. Using accelerator hardware such as graphics processor units, our approach can be about as fast (in wall clock time) as a single full-data model fit. Parallel CV is flexible because it can easily exploit a wide range data partitioning schemes, such as those designed for non-exchangeable data. It can also accommodate a range of scoring rules. We propose MCMC diagnostics, including a summary of MCMC mixing based on the popular potential scale reduction factor ( R ) and MCMC effective sample size ( ESS ) measures. We also describe a method for determining whether an R diagnostic indicates approximate stationarity of the chains, that may be of more general interest for applications beyond parallel CV. Finally, we show that parallel CV and its diagnostics can be implemented with online algorithms, allowing parallel CV to scale up to very large blocking designs on memory-constrained computing accelerators.
Statistical hypotheses are translations of scientific hypotheses into statements about one or more distributions, often concerning their centre. Tests that assess statistical hypotheses of centre implicitly assume a specific centre, e.g., the mean or median. Yet, scientific hypotheses do not always specify a particular centre. This ambiguity leaves the possibility for a gap between scientific theory and statistical practice that can lead to rejection of a true null. In the face of replicability crises in many scientific disciplines, significant results of this kind are concerning. Rather than testing a single centre, this paper proposes testing a family of plausible centres, such as that induced by the Huber loss function (the Huber family). Each centre in the family generates a testing problem, and the resulting family of hypotheses constitutes a familial hypothesis. A Bayesian nonparametric procedure is devised to test familial hypotheses, enabled by a novel pathwise optimization routine to fit the Huber family. The favourable properties of the new test are demonstrated theoretically and experimentally. Two examples from psychology serve as real-world case studies.
A widely adopted measure of housing affordability is that households should spend no more than 30% of their household income on housing. However, this normative threshold is an arbitrary Great Depression-era guideline and may not be relevant today. This paper proposes a subjective indicator of housing affordability by introducing a method commonly used in the medical sciences. It utilizes discrete information to estimate a subjective affordability ratio that discriminates between subjective house-poor and non-house-poor households. We apply the proposed method to household-level data collected in Selangor, Malaysia, and show that the optimal cut-off point is 23.5%. This estimated value suggests a higher prevalence of house-poor households than is implied by the regularly assumed 30% threshold. In addition, we perform a sensitivity analysis and find the bias in the estimated cut-off point is close to zero.
Variational Bayesian (VB) methods produce posterior inference in a time frame considerably smaller than traditional Markov Chain Monte Carlo approaches. Although the VB posterior is an approximation, it has been shown to produce good parameter estimates and predicted values when a rich classes of approximating distributions are considered. In this paper, we propose the use of recursive algorithms to update a sequence of VB posterior approximations in an online, time series setting, with the computation of each posterior update requiring only the data observed since the previous update. We show how importance sampling can be incorporated into online variational inference allowing the user to trade accuracy for a substantial increase in computational speed. The proposed methods and their properties are detailed in two separate simulation studies. Additionally, two empirical illustrations are provided, including one where a Dirichlet Process Mixture model with a novel posterior dependence structure is repeatedly updated in the context of predicting the future behaviour of vehicles on a stretch of the US Highway 101.
We conduct an extensive evaluation of price jump tests based on high-frequency financial data. After providing a concise review of multiple alternative tests, we document the size and power of all tests in a range of empirically relevant scenarios. Particular focus is given to the robustness of test performance to the presence of jumps in volatility and microstructure noise, and to the impact of sampling frequency. The paper concludes by providing guidelines for empirical researchers about which test to choose in any given setting.
We find that factors explaining bank loan recovery rates vary depending on the state of the economic cycle. Our modeling approach incorporates a two-state Markov switching mechanism as a proxy for the latent credit cycle, helping to explain differences in observed recovery rates over time. We are able to demonstrate how the probability of default and certain loan-specific and other variables hold different explanatory power with respect to recovery rates over `good' and `bad' times in the credit cycle. That is, the relationship between recovery rates and certain loan characteristics, firm characteristics and the probability of default differs depending on underlying credit market conditions. This holds important implications for modelling capital retention, particularly in terms of countercyclicality.
We investigate the impact of filter choice on forecast accuracy in state space models. The filters are used both to estimate the posterior distribution of the parameters, via a particle marginal Metropolis-Hastings (PMMH) algorithm, and to produce draws from the filtered distribution of the final state. Multiple filters are entertained, including two new data-driven methods. Simulation exercises are used to document the performance of each PMMH algorithm, in terms of computation time and the efficiency of the chain. We then produce the forecast distributions for the one-step-ahead value of the observed variable, using a fixed number of particles and Markov chain draws. Despite distinct differences in efficiency, the filters yield virtually identical forecasting accuracy, with this result holding under both correct and incorrect specification of the model. This invariance of forecast performance to the specification of the filter also characterizes an empirical analysis of S&P500 daily returns.
This paper provides an extensive evaluation of high frequency jump tests and measures, in the context of using such tests and measures in the estimation of dynamic models for asset price jumps. Specifically, we investigate: i) the power of alternative tests to detect individual price jumps, most notably in the presence of volatility jumps; ii) the frequency with which sequences of dynamic jumps are correctly identified; iii) the accuracy with which the magnitude and sign of a sequence of jumps, including small clusters of consecutive jumps, are estimated; and iv) the robustness of inference about dynamic jumps to test and measure design. Substantial differences are discerned in the performance of alternative methods in certain dimensions, with inference being sensitive to these differences in some cases. Accounting for measurement error when using measures constructed from high frequency data to conduct inference on dynamic jump models is also shown to have an impact. The sensitivity of inference to test and measurement construction is documented using both artificially generated data and empirical data on both the Su0026P500 stock index and the IBM stock price. The paper concludes by providing guidelines for empirical researchers who wish to exploit high frequency data when drawing conclusions regarding dynamic jump processes.
This paper provides an extensive evaluation of high frequency jump tests and measures, in the context of dynamic models for asset price jumps. Specifically, we investigate: i) the power of alternative tests to detect individual price jumps, including in the presence of volatility jumps; ii) the frequency with which sequences of dynamic jumps are identified; iii) the accuracy with which the magnitude and sign of sequential jumps are estimated; and iv) the robustness of inference about dynamic jumps to test and measure design. Substantial differences are discerned in the performance of alternative methods in certain dimensions, with inference being sensitive to these differences in some cases. Accounting for measurement error when using measures constructed from high frequency data to conduct inference on dynamic jump models would appear to be advisable.
In this paper, a Bayesian version of the exponential smoothing method of forecasting is proposed. The approach is based on a state space model containing only a single source of error for each time interval. This model allows an improvement to current practices in exponential smoothing by providing both point predictions and measures of the uncertainty surrounding them. The method proposed calculates posterior prediction and parameter distributions via Monte Carlo composition. We evaluate the method with a Monte Carlo simulation study and apply it to forecasting car part demand. The main advantage of the approach is that it produces exact, small sample prediction distributions. It also works very quickly on modern computing machines.
Dynamic jumps in the price and volatility of an asset are modelled using a joint Hawkes process in conjunction with a bivariate jump diffusion. A state space representation is used to link observed returns, plus nonparametric measures of integrated volatility and price jumps, to the specified model components; with Bayesian inference conducted using a Markov chain Monte Carlo algorithm. An evaluation of marginal likelihoods for the proposed model relative to a large number of alternative models, including some that have featured in the literature, is provided. An extensive empirical investigation is undertaken using data on the S&P500 market index over the 1996 to 2014 period, with substantial support for dynamic jump intensities - including in terms of predictive accuracy - documented.
This paper proposes a new Bayesian approach for analysing moment condition models in the situation where the data may be contaminated by outliers. The approach builds upon the foundations developed by Schennach (2005) who proposed the Bayesian exponentially tilted empirical likelihood (BETEL) method, justified by the fact that an empirical likelihood (EL) can be interpreted as the nonparametric limit of a Bayesian procedure when the implied probabilities are obtained from maximizing entropy subject to some given moment constraints. Considering the impact that outliers are thought to have on the estimation of population moments, we develop a new robust BETEL (RBETEL) inferential methodology to deal with this potential problem. We show how the BETEL methods are linked to the recent work of Bissiri, Holmes and Walker (2016) who propose a general framework to update prior belief via a loss function. A controlled simulation experiment is conducted to investigate the performance of the RBETEL method. We find that the proposed methodology produces reliable posterior inference for the fundamental relationships that are embedded in the majority of the data, even when outliers are present. The method is also illustrated in an empirical study relating brain weight to body weight using a dataset containing sixty-five different land animal species.
This paper develops a new non-linear model to analyse the business cycle by exploiting the relationship between the asymmetrical behaviour of the cycle and leading indicators. The model proposed is an innovations form of the structural model underlying simple exponential smoothing that is augmented by a latent Markov switching process. Furthermore, the probabilities that drive the Markov process vary with the growth of the leading indicator. The proposed model is used to analyse the Australian business cycle using the gross domestic product as a proxy and the industrial materials prices index as the exogenous leading indicator influencing the transition probabilities. Model parameters are estimated using a Gibbs sampling algorithm and subsequently used for forecasting purposes.
We investigate the systemic risk of the European sovereign and banking system during 2008–2013. We utilize a conditional measure of systemic risk that reflects market perceptions and can be intuitively interpreted as an entity’s conditional joint probability of default, given the hypothetical default of other entities. The measure of systemic risk is applicable to high dimensions and not only incorporates individual default risk characteristics but also captures the underlying interdependent relations between sovereigns and banks in a multivariate setting. In empirical applications, our results reveal significant time variation in systemic risk spillover effects for the sovereign and banking system. We find that systemic risk is mainly driven by risk premiums coupled with a steady increase in physical default risk.
This paper proposes new automated proposal distributions for sequential Monte Carlo algorithms, including particle filtering and related sequential importance sampling methods. The wrights for these proposal distributions are easily established, as is the unbiasedness property of the resultant likelihood estimators, so that the methods may be used within a particle Markov chain Monte Carlo (PMCMC) inferential setting. Simulation exercises, based on a range of state space models, are used to demonstrate the linkage between the signal-to-noise ratio of the system and the performance of the new particle filters, in comparison with existing filters. In particular, we demonstrate that one of our proposed filters performs well in a high signal-to-noise ratio setting, that is, when the observation is informative in identifying the location of the unobserved state. A second filter, deliberately designed to draw proposals that are informed by both the current observation and past states, is shown to work well across a range of signal-to noise ratios and to be much more robust than the auxiliary particle filter, which is often used as the default choice. We then extend the study to explore the performance of the PMCMC algorithm using the new filters to estimate the likelihood function, once again in comparison with existing alternatives. Taking into consideration robustness to the signal-to-noise ratio, computation time and the efficiency of the chain, the second of the new filters is again found to be the best-performing method. Application of the preferred filter to a stochastic volatility model for weekly Australian/US exchange rate returns completes the paper.