We study nonlinear serial dependence tests based on portmanteau statistics with nonlinear autocovariances for non-Gaussian time series and residuals of dynamic models. A new test with an asymptotic chi 2 distribution is introduced for testing nonlinear serial dependence (NLSD) in time series. It stems from the Generalized Covariance (GCov) residual-based specification test with an asymptotic chi 2 distribution for semi-parametric dynamic models with i.i.d. non-Gaussian errors. We derive new asymptotic results under local alternative hypotheses on the parameters of a dynamic model, and extend the GCov test to an infinite set of nonlinear autocovariance conditions. A GCov bootstrap test is introduced for size adjustments in finite samples, or residual diagnostics in models estimated parametrically by a maximum likelihood method. The proposed inference method applies to time series exhibiting bubbles, such as commodity prices (e.g., oil, gold) and cryptocurrency rates. A simulation study shows that the tests perform well in applications to mixed causal-noncausal autoregressive models. The GCov specification test is used to assess the fit of a mixed causal-noncausal model of aluminum prices with locally explosive patterns, such as bubbles and spikes, between 2005 and 2024.
This paper proposes the Regularized Generalized Covariance (RGCov) estimator, a ridge-type extension of the Generalized Covariance (GCov) estimator for high-dimensional stationary time series. By regularizing the GCov objective function, which involves the inverse of the covariance matrix, RGCov improves numerical stability while preserving positive definiteness. Under suitable conditions, the new estimator is consistent, asymptotically normal, and semi-parametrically efficient. We also extend the GCov specification test and the nonlinear serial dependence (NLSD) test to their regularized versions, both of which are asymptotically chi-square distributed. Simulation studies confirm the reliability of the RGCov and associated tests in high-dimensional settings. In the empirical application, RGCov is used to estimate a mixed causal-noncausal Vector Autoregressive (VAR) model for green energy stocks in the RENIXX index and to construct two bubble-based investment strategies: bubble-riding and bubble-hedging, both of which outperform the benchmark index.
We show that the mixed causal-noncausal vector autoregressive (VAR) processes satisfy the Markov property in both calendar and reverse time. Based on that property, we introduce closed-form formulas of forward and backward predictive densities for point and interval forecasting and backcasting out-of-sample. The backcasting formula is used for adjusting the forecast interval to obtain a desired coverage level when the tail quantiles are difficult to estimate. A confidence set for the prediction interval is introduced for assessing the uncertainty due to estimation. We also define new nonlinear past-dependent innovations of mixed causal-noncausal VAR models for impulse response function analysis. Our approach is illustrated by simulations and an application to the joint analysis of oil prices and real gross domestic product (GDP) growth rates.
This paper examines how Canadian firms balance the benefits of technology adoption against the rising risk of cyber security breaches. We merge data from the 2021 Canadian Survey of Digital Technology and Internet Use and the 2021 Canadian Survey of Cyber Security and Cybercrime to investigate the trade-off firms face when pursuing digitalization to enhance productivity and efficiency, balanced against the potential increase in cyber security risk. The analysis explores the extent of digital technology adoption, differences across industries, the subsequent associations with efficiency, and associated cyber security vulnerabilities. We build aggregate variables, such as the Business Digital Usage Score and a cyber security incidence variable to quantify each firm's digital engagement and cyber security risk. A survey-weight-adjusted Lasso estimator is employed, and a debiasing method for high-dimensional logit models is introduced to identify the predictors of technological efficiency and cyber risk. The analysis reveals a digital divide linked to firm size, industry, and workforce composition. While rapid expansion of tools such as cloud services or artificial intelligence can raise efficiency, it simultaneously heightens exposure to cyber threats, particularly among larger enterprises.
We study the Functional PCA (FPCA) forecasting method in application to functions of intraday returns on Bitcoin. We show that improved interval forecasts of future return functions are obtained when the conditional heteroscedasticity of return functions is taken into account. The Karhunen-Loeve (KL) dynamic factor model is introduced to bridge the functional and discrete time dynamic models. It offers a convenient framework for functional time series analysis. For intraday forecasting, we introduce a new algorithm based on the FPCA applied by rolling, which can be used for any data observed continuously 24/7. The proposed FPCA forecasting methods are applied to return functions computed from data sampled hourly and at 15-minute intervals. Next, the functional forecasts evaluated at discrete points in time are compared with the forecasts based on other methods, including machine learning and a traditional ARMA model. The proposed FPCA-based methods perform well in terms of forecast accuracy and outperform competitors in terms of directional (sign) of return forecasts at fixed points in time.
This paper introduces a local-to-unity/small sigma model for stationary processes with longrange persistence and non-negligible long-run prediction and estimation risks. The model represents a process containing unobserved short and long-run components measured on different time scales. The short-run component is defined in calendar time, while the longrun component evolves in rescaled time with ultra-long units. We develop estimation and long-run prediction methods for time series with multivariate Vector Autoregressive (VAR) short-run components and reveal the impossibility of estimating consistently some of the longrun parameters, which causes significant estimation and prediction risks in the long run. A simulation study and an application to macroeconomic data illustrate the approach.
This paper introduces a new approach for bubble detection based on mixed causal and noncausal autoregressive processes and their tail process representation during an explosive episode. Departing from traditional definitions of bubbles as nonstationary and temporarily explosive processes, we adopt a perspective in which prices are assumed to follow a strictly stationary process, with the bubble considered an intrinsic component of its nonlinear dynamics. The proposed approach provides a bubble indicator for detecting bubbles and measuring their duration. We implement our strategy to investigate the phenomenon called the "green bubble" in the field of renewable energy investment.
We introduce a regularized Generalized Covariance (RGCov) estimator as an extension of the GCov estimator to high dimensional setting that results either from high-dimensional data or a large number of nonlinear transformations used in the objective function. The approach relies on a ridge-type regularization for high-dimensional matrix inversion in the objective function of the GCov. The RGCov estimator is consistent and asymptotically normally distributed. We provide the conditions under which it can reach semiparametric efficiency and discuss the selection of the optimal regularization parameter. We also examine the diagonal GCov estimator, which simplifies the computation of the objective function. The GCov-based specification test, and the test for nonlinear serial dependence (NLSD) are extended to the regularized RGCov specification and RNLSD tests with asymptotic Chi-square distributions. Simulation studies show that the RGCov estimator and the regularized tests perform well in the high dimensional setting. We apply the RGCov to estimate the mixed causal and noncausal VAR model of stock prices of green energy companies.
We introduce a new stochastic tree representation of a strictly stationary submartingale process for modelling, forecasting, and pricing speculative bubbles on commodity and cryptocurrency markets. The model is compared to other trees proposed in the literature on bubble asset modelling and stochastic volatility approximation. We show that the proposed model is an extension of the well-known Blanchard-Watson bubble. The model provides (quasi) closed-form pricing formulas for European options, which are derived and illustrated.
The bubbles and spikes in cryptocurrency prices increase considerably the risk on investments in these assets. In the traditional time series literature bubbles are viewed as nonstationary and non-estimable components of a process. In this paper, we adopt a different approach and consider the bubbles as inherent features of a strictly stationary causal-noncausal (mixed) Vector Autoregressive (VAR) process. This approach allows us to model and estimate the common bubbles and spikes in cryptocurrency prices. It also provides us linear combinations of cryptocurrencies that eliminate common bubbles analogously to the cointegrating vectors eliminating common trends in unit root processes. They are used to build cryptocurrency portfolios immune to the risk of common bubbles that ensure stable investment strategies. The mixed VAR model is estimated from the US Dollar prices of Bitcoin, Ethereum, Ripple, and Stellar over the period 2017–2019. We document the common bubbles and illustrate the behavior of bubble-free portfolios.
This paper investigates the performance of routinely used optimization algorithms in application to the Generalized Covariance estimator (GCov) for univariate and multivariate mixed causal and noncausal models. The GCov is a semi-parametric estimator with an objective function based on nonlinear autocovariances to identify causal and noncausal orders. When the number and type of nonlinear autocovariances included in the objective function are insufficient/inadequate, or the error density is too close to the Gaussian, identification issues can arise. These issues result in local minima in the objective function, which correspond to parameter values associated with incorrect causal and noncausal orders. Then, depending on the starting point and the optimization algorithm employed, the algorithm can converge to a local minimum. The paper proposes the Simulated Annealing (SA) optimization algorithm as an alternative to conventional numerical optimization methods. The results demonstrate that SA performs well in its application to mixed causal and noncausal models, successfully eliminating the effects of local minima. The proposed approach is illustrated by an empirical study of a bivariate series of commodity prices.
We introduce closed-form formulas of out-of-sample predictive densities for forecasting and backcasting of mixed causal-noncausal (Structural) Vector Autoregressive VAR models. These nonlinear and time irreversible non-Gaussian VAR processes are shown to satisfy the Markov property in both calendar and reverse time. A post-estimation inference method for assessing the forecast interval uncertainty due to the preliminary estimation step is introduced too. The nonlinear past-dependent innovations of a mixed causal-noncausal VAR model are defined and their filtering and identification methods are discussed. Our approach is illustrated by a simulation study, and an application to cryptocurrency prices.
As Canada and other major economies consider implementing "digital money" or Central Bank Digital Currencies, understanding how demographic and geographic factors influence public engagement with digital technologies becomes increasingly important. This paper uses data from the 2020 Canadian Internet Use Survey and employs survey-adapted Lasso inference methods to identify individual socio-economic and demographic characteristics determining the digital divide in Canada. We also introduce a score to measure and compare the digital literacy of various segments of Canadian population. Our findings reveal that disparities in the use of e.g. online banking, emailing, and digital payments exist across different demographic and socio-economic groups. In addition, we document the effects of COVID-19 pandemic on internet use in Canada and describe changes in the characteristics of Canadian internet users over the last decade.
This article develops statistical inference methods for a class of set-identified models, where the errors are known functions of observations and the parameters satisfy either serial or/and cross-sectional independence conditions. This class of models includes the independent component analysis (ICA), Structural Vector Autoregressive (SVAR), and multi-variate mixed causal-non-causal models. We use the Generalized Covariance (GCov) estimator to compute the residual-based portmanteau statistic for testing the error independence hypothesis. Next, we build the confidence sets for the identified sets of parameters by inverting the test statistic. We also discuss the choice (design) of these statistics. The approach is illustrated by simulations examining the under-identification condition in an ICA model and an application to financial return series.
This paper examines and compares intraday and intraweek patterns in hourly and daily prices, returns, volumes and volatility of native cryptocurrencies, stablecoins and tokens traded on Bitstamp. We show that native cryptocurrencies and tokens share common intraday periodicity determined by the operating times of the NYSE, LSE and Hang Seng stock exchange markets. Periodic patterns are also documented in the returns on cryptocurrency market portfolio approximated by the PCA applied to intraday and intraweek cross-sectional correlation matrices of cryptocurrency returns. Stablecoins have distinct dynamics and their daily and hourly returns are uncorrelated with one another and with the returns on other cryptocurrencies. We introduce a functional CAPM to accommodate the periodic patterns and estimate it by regressing the functions of intraday and intraweek cryptocurrency returns on the market portfolio. We show that the return functions on Bitcoin, Ether, and Link satisfy affine relationships with the return functions of the market portfolio and their functional betas display periodic intraday and intraweek patterns.
This paper introduces a local-to-unity/small sigma process for a stationary time series with strong persistence and non-negligible long run risk. This process represents the stationary long run component in an unobserved shortand long-run components model involving different time scales. More specifically, the short run component evolves in the calendar time and the long run component evolves in an ultra long time scale. We develop the methods of estimation and long run prediction for the univariate and multivariate Structural VAR (SVAR) models with unobserved components and reveal the impossibility to consistently estimate some of the long run parameters. The approach is illustrated by a Monte-Carlo study and an application to macroeconomic data.
This paper extends three Lasso inferential methods, Debiased Lasso, $C(\alpha)$ and Selective Inference to a survey environment. We establish the asymptotic validity of the inference procedures in generalized linear models with survey weights and/or heteroskedasticity. Moreover, we generalize the methods to inference on nonlinear parameter functions e.g. the average marginal effect in survey logit models. We illustrate the effectiveness of the approach in simulated data and Canadian Internet Use Survey 2020 data.
Canada and other major countries are investigating the implementation of ``digital money'' or Central Bank Digital Currencies, necessitating answers to key questions about how demographic and geographic factors influence the population's digital literacy. This paper uses the Canadian Internet Use Survey (CIUS) 2020 and survey versions of Lasso inference methods to assess the digital divide in Canada and determine the relevant factors that influence it. We find that a significant divide in the use of digital technologies, e.g., online banking and virtual wallet, continues to exist across different demographic and geographic categories. We also create a digital divide score that measures the survey respondents' digital literacy and provide multiple correspondence analyses that further corroborate these findings.
The parametric estimators applied by rolling are commonly used for the analysis of time series with nonlinear patterns, including time varying parameters and local trends. This paper examines the properties of rolling estimators in the class of temporally local maximum likelihood (TLML) estimators. We consider the TLML estimators of (a) constant parameters, (b) stochastic, stationary parameters and (c) parameters with the ultra-long run (ULR) dynamics bridging the gap between the constant and stochastic parameters. We show that the weights used in the TLML estimators have a strong impact on the inference. For illustration, we provide a simulation study of the epidemiological susceptible-infected-susceptible (SIS) model, which explores the finite sample performance of TLML estimators of a time varying contagion parameter.
We introduce the conditional Maximum Composite Likelihood (MCL) estimation method for the stochastic factor ordered Probit model of credit rating transitions of firms. This model is recommended for internal credit risk assessment procedures in banks and financial institutions under the Basel III regulations. Its exact likelihood function involves a high-dimensional integral, which can be approximated numerically before maximization. However, the estimated migration risk and required capital tend to be sensitive to the quality of this approximation, potentially leading to statistical regulatory arbitrage. The proposed conditional MCL estimator circumvents this problem and maximizes the composite log-likelihood of the factor ordered Probit model. We present three conditional MCL estimators of different complexity and examine their consistency and asymptotic normality when n and T tend to infinity. The performance of these estimators at finite T is examined and compared with a granularity-based approach in a simulation study. The use of the MCL estimator is also illustrated in an empirical application.