
Threshold models are set up so that there is a switch between regimes for the parameters of an unobserved components model. When Gaussianity is assumed, the model is handled by the Kalman filter. The switching depends on a component crossing a boundary, and, because the component is not observed directly, the error in its estimation leads naturally to a smooth transition mechanism. A prominent example motivating thresholds is that of a cyclical time series characterized by a downturn that is more, or less, rapid than the upturn. The situation is illustrated by fitting a model with three potentially asymmetric cycles, each with its own threshold, to observations on ice volume in Antarctica since 799,000 BCE. The model is able to produce multi-step forecasts with associated prediction intervals. A second example shows how a hidden threshold model is able to deal with the asymmetric cycle in monthly US unemployment.
Distributional shifts occur frequently, with detrimental impacts on forecast accuracy. Almost all econometric forecasting models are equilibrium correction so are susceptible to systematic forecast failure after equilibrium-mean shifts, which could be direct or induced. Unanticipated out-of-sample shifts are a well-known cause of forecast failure, but they later become in-sample shifts so require handling. Previous forecast-error taxonomies are extended to include shifts both in-sample and after the forecast origin. The taxonomy reveals which shifts do and do not lead to forecast failure facing both in-sample and post forecast-origin shifts, also highlighting what features are amenable to rapid correction.
A historically structured survey of panel data econometrics is provided, tracing its evolution from the error components models of the 1960s to the modern integration of causal inference, machine learning, and high-dimensional methods. The survey emphasizes the conceptual breakthroughs that shaped the field, including approaches to unobserved heterogeneity, the development of dynamic panel GMM estimators, nonlinear and limited dependent variable models, cross-sectional dependence, nonstationary panels, and recent advances in causal inference and policy evaluation.
A novel method to estimate social effect coefficients in the popular so-called linear-in-means regression model in the Social Sciences is presented here that utilizes non-experimental multidimensional network data. The procedure can accommodate social interactions that correlate with the error in the model by making use of a different set of network links among the same observations that are exogenous in the traditional sense. In particular, the full observability of a two-layered multiplex network data structure is assumed here to propose a new Generalized 3-Stage Least Squares (G3SLS) estimator that is consistent, asymptotically normally distributed, and also easy to implement using widely-used existing statistical software because of its closed-form definition. The underlying assumptions are general enough to accommodate common problems with observational data such as measurement error, simultaneity, and unobserved heterogeneity. Monte Carlo exercises confirm the good small sample performance of the proposed G3SLS estimator in these scenarios. An empirical application finds positive and significant peer effects in citations among research articles published in top general-interest journals in economics.
Predictability is defined as the extent to which the conditional expectation of future outcomes diverges from its unconditional counterpart. A random walk is, by this criterion, unpredictable, since the expectations coincide. However, predictability does emerge once the perspective is inverted: rather than forecasting the level attained at a given time, the time required to reach a given level, the hitting time, is forecast. Under this formulation, the conditional distribution differs from the unconditional distribution and thereby random walk satisfies the predictability condition. The approach rests upon the principle of the infinite divisibility of probability generating functions. The analysis employs hypergeometric functions, which yield the Beta–binomial form of the conditional forecast distribution. This method permits a rigorous statistical treatment, thereby uncovering the inherent predictability of a random walk.
The Spectral Omnibus test (SPECO) is introduced as a diagnostic for assessing departures from cross-sectional independence in panel model residuals. SPECO operates on the eigenvalue spectrum of the residual correlation matrix and aggregates six complementary spectral indicators—capturing dominance, separation, concentration, and disorder—into a single omnibus decision. For each indicator, empirical significance values are obtained from a Monte Carlo null cache indexed by panel dimension and combined using the Cauchy method, yielding reliable finite-sample inference without relying on large-sample edge approximations. Extended simulations spanning global (linear and nonlinear), structured (sparse and block), and robustness (temporal and non-Gaussian) dependence structures show that all procedures achieve nominal size after empirical calibration. In power comparisons, SPECO attains near-unit power under linear and monotonic dependence and delivers substantial gains under oscillatory, sign-varying alternatives, where standard moment-based and pairwise diagnostics can exhibit substantially reduced power. SPECO also remains stable under heavy-tailed errors, Gaussian mixtures, heterogeneous panels, and moderate temporal dependence. Overall, SPECO provides a computationally efficient, broadly applicable diagnostic when the form of cross-sectional dependence is unknown.
The asymptotic behavior of the indirect inference estimator for a conditionally Gaussian stochastic volatility-in-mean model with asymmetric effects is investigated. The auxiliary model is based on Gaussian QML estimation using misspecified volatility filters, such as the EGARCH and the GQARCH. Under general assumptions, the binding function from the data-generating process to the auxiliary model is shown to be injective. Leveraging ergodic optimization, strong consistency and the asymptotic Gaussianity of the estimator are established, along with a consistent estimator of the asymptotic variance. Monte Carlo experiments and an empirical application to financial data suggest that filters closely approximating the true volatility recursion yield superior performance.
Proper econometric analysis should be informed by data structure. Many forms of financial data are recorded in discrete-time and relate to products of a finite term. If the data is sampled from a financial trust, it will often be further subject to random left-truncation. The estimation of a distribution function from left-truncated data has been extensively addressed, but the case of discrete data over a known, finite number of possible values has not yet been thoroughly investigated. A precise discrete framework and suitable sampling procedure for the Woodroofe-type estimator for discrete data over a known, finite number of possible values is therefore established. Subsequently, the resulting vector of hazard rate estimators is proved to be asymptotically normal with independent components. Asymptotic normality of the survival function estimator is then established. Sister results for the left-truncating random variable are also proved. Taken together, the resulting joint vector of hazard rate estimates for the lifetime and left-truncation random variables is proved to be the maximum likelihood estimate of the parameters of the conditional joint lifetime and left-truncation distribution given the lifetime has not been left-truncated. A hypothesis test for the shape of the distribution function based on our asymptotic results is derived. Such a test is useful to formally assess the plausibility of the stationarity assumption in length-biased sampling. The finite sample performance of the estimators is investigated in a simulation study. Applicability of the theoretical results in an econometric setting is demonstrated with a subset of data from the Mercedes-Benz 2017-A securitized bond. (c) 2023 EcoSta Econometrics and Statistics. Published by Elsevier B.V. All rights reserved.
The identification, inference, and validation of linear panel data models are studied under a general framework in which both factors and factor loadings are characterized by a nonparametric function. It encompasses widely used models such as two-way fixed effects and interactive fixed effects. Under a conditional mean independence assumption between unobserved heterogeneity and covariates, consistent estimators of the parameters of interest are obtained at the optimal rate of convergence, for both fixed and large T. A specification test for the modeling assumption is also developed, based on conditional moment test methodology and nonparametric estimation techniques. Using degenerate and nondegenerate U-statistics theory, the convergence and asymptotic distribution of the test under the null hypothesis are established, and divergence under the alternative is shown to occur at a rate arbitrarily close to NT. Finite-sample inference relies on bootstrap procedures. The simulation results demonstrate excellent performance of the proposed methods and an empirical application is provided.
Under mild conditions, a least-squares local linear Fréchet curve predictor is derived for a response and a regressor evaluated in a separable Hilbert space. The conditions that allow the implementation of the local linear Fréchet functional predictor in the ambient L2-space of vector functions, with values in the time-varying tangent space of a compact Riemannian manifold, are established. An intrinsic local linear Fréchet curve predictor on such a manifold is then proposed, based on a weighted Fréchet mean approach. Its asymptotic optimality is proved. Simulations and a real-data application are considered to analyze the finite-sample performance of the empirical versions of both predictors, compared with a geodesic Nadaraya–Watson-type curve predictor. In the real-data application, the functional prediction of the time-varying spherical coordinates of the Earth’s magnetic field is addressed using observations through time of the geocentric latitude and longitude of the NASA MAGSAT spacecraft.
Monitoring statistics for structural changes in systems of cointegrating relationships are proposed. The approach is based on parameter estimation over a calibration period. In case of homogenous systems and cross-sectional independence the pooled fully modified OLS estimator takes into account the effects of error serial correlation and regressor endogeneity. Cross-sectional dependence is allowed by using the pooled fully modified GLS estimator for homogenous systems and the fully modified SUR estimator for inhomogenous systems. The detectors show decent behaviour under the null hypothesis with controlled rejection probabilities and power against two alternatives for different data generating processes. An empirical application investigates deviations from the arbitrage parity condition for exchange rate triplets including Bitcoin. The procedures detect breakpoints in May to August 2014 and in January to May 2015 indicating an instability in arbitrage parities. Following this, a promising portfolio trading strategy based on the breakdates is constructed.
A new test is proposed to detect whether break points are common in heterogeneous panel data models where the time series dimension T could be large relative to cross-section dimension N. The error process is assumed to be cross-sectionally independent. The test is based on the cumulative sum (CUSUM) of ordinary least squares (OLS) residuals. The asymptotic distribution of the detecting statistic is derived under the null hypothesis, while the test is shown to be consistent under the alternative. Monte Carlo simulations and an empirical example show good performance of the test. (c) 2023 EcoSta Econometrics and Statistics. Published by Elsevier B.V. All rights reserved.
Using a transformation of the autoregressive distributed lag model due to Bewley, a novel pooled Bewley (PB) estimator of long-run coefficients for dynamic panels with heterogeneous short-run dynamics is proposed. The PB estimator is directly comparable to the widely used Pooled Mean Group (PMG) estimator, and is shown to be consistent and asymptotically normal. Monte Carlo simulations show good small sample performance of PB compared to the existing estimators in the literature, namely PMG, panel dynamic OLS (PDOLS), and panel fully-modified OLS (FMOLS). Application of two bias-correction methods and a bootstrapping of critical values to conduct inference robust to cross-sectional dependence of errors are also considered. The utility of the PB estimator is illustrated in an empirical application to the aggregate consumption function.
The large heterogeneous panel data models are extended to the setting where the heterogenous coefficients are changing over time and the regressors are endogenous. Kernel-based non-parametric time-varying parameter instrumental variable mean group (TVP-IV-MG) estimator is proposed for the time-varying cross-sectional mean coefficients. The uniform consistency is shown and the pointwise asymptotic normality of the proposed estimator is derived. A data-driven bandwidth selection procedure is also proposed. The finite sample performance of the proposed estimator is investigated through a Monte Carlo study and an empirical application on multi-country Phillips curve with time-varying parameters.
It is well-known that real data often contain outliers. The term outlier typically refers to a case, that is, a row of the $n \times d$ data matrix. In recent times a different type has come into focus, the cellwise outliers. These are suspicious cells (entries) that can occur anywhere in the data matrix. Even a relatively small proportion of outlying cells can contaminate over half the rows, which is a problem for rowwise robust methods. In this article we discuss the challenges posed by cellwise outliers, and some methods developed so far to deal with them. We obtain new results on cellwise breakdown values for location, covariance and regression. We also propose a cellwise robust method for correspondence analysis, with real data illustrations. The paper concludes by formulating some points for debate.
The cumulant based normality test after outlier removal is analyzed. It is shown that the standard least squares normalizations can be misleading in this context. The sample cumulants should be standardized according to the truncation imposed at the removal stage and the estimation method being used. New standardizations that lead to chi-squared inference are derived.
The purpose is to enable inference in case of quantile regression with endogenous covariates and clustered data. It is proven that the instrumental variable quantile regression estimator is consistent where there is correlation of errors within clusters, and an asymptotic distribution for the estimator, which may be used for inference for a given quantile tau, is derived. As regards inference based on the entire instrumental variable quantile regression process, it is proven that cluster-based resampling of a statistic of a certain class offers a computationally tractable approach for implementing asymptotic tests. The theoretical results concerning the asymptotic properties of the instrumental variable quantile regression estimator for clustered data are supported by simulation analysis. An empirical illustration shows the use of the proposed technique in order to estimate the earning equations of US men and women where female labor supply is endogenous and subject to the shock of World War II. (c) 2023 EcoSta Econometrics and Statistics. Published by Elsevier B.V. All rights reserved.