Understanding regional Consumer Price Index (CPI) dynamics is essential for timely and effective economic policymaking. However, traditional modeling procedures typically rely only on parametric panel modeling with low-frequency and high-cost macroeconomic indicators, which often fail to capture rapid market fluctuations and lead to inaccurate predictions. To this end, we propose a residual-joint-modeling framework that integrates large language model (LLM) analyses and social media narratives via a new deep neural network based panel modeling. Specifically, we construct a large narrative corpus from a newly collected Sina Weibo dataset, and develop a prompt-based GPT model and a series of fine-tuned BERT models to generate high-frequency LLM-induced surrogates for regional CPI. A novel joint modeling strategy is then advocated to transfer the information from these surrogates to the target regional CPI data and hence empower CPI prediction. To solve the joint objectives, we further introduce a new deep panel learning procedure with region-wise homogeneity pursuit, which has its own significance in panel data analysis literature. In addition, conformal-based panel prediction intervals are provided to quantify the uncertainty of the LLM-powered prediction. The proposed approach significantly reduces short-term forecasting errors and more effectively captures abrupt inflationary shifts compared to traditional econometric models. While demonstrated for regional CPI forecasting, the proposed framework is broadly applicable for incorporating insights from LLMs to enhance traditional statistical modeling.
ABSTRACT This paper utilises high‐frequency data from the Chinese stock market and a panel of individual stocks to compare the forecasting performance of machine learning with popular econometric models across different periods. For short‐term volatility forecasting, most machine learning models outperform econometric models with limited explanatory variables, though they do not exhibit a significant advantage over the econometric model incorporating all features. For medium‐ and long‐term volatility forecasting, Light gradient boosting machine (LGBM) in machine learning substantially dominates econometric models. We also explore a simple‐to‐implement forecast combination method that leverages the best machine learning model and the best econometric model to explore if model averaging leads to any improvement. Our findings indicate that over a longer forecasting horizon, this method achieves the best performance among all forecast combinations and dominates the best econometric model.
We consider the estimation and statistical inference for low dimensional parameters for a regression model with covariates whose dimension increases with sample size. We suggest a computationally simple one stage orthogonal projection approach to estimate the low dimensional parameters under strict or approximate sparsity conditions. The orthogonal projection approach is simple to implement and the inference for the low dimensional parameters is straightforward to derive whether the high dimensional function is linear or nonlinear. It also avoids the complicated regularization bias issues commonly associated with two stage estimation methods. Monte Carlo simulations and empirical applications are also conducted to investigate the finite sample performance of the proposed estimator vs the double/debiased estimator of Belloni et al. (2014) and Chernozhukov et al. (2018).
We consider a two-step estimation procedure to estimate the panel sample selection models with interactive effects. In the first step, we follow the Robinson (1988) procedure to remove the sample selection factors. In the second step, we control the interactive effects. When the cross-section dimension N is large, we propose to use the Pesaran (2006) common correlated effects approach, and when the time series dimension T is large and N is finite we propose to follow the Hsiao, Shi, and Zhou (2022) transformed estimation procedure to eliminate the interactive effects. We show that the resulting estimators are consistent and asymptotically normally distributed. A limited Monte Carlo study is conducted, showing our methods appear to work well in a finite sample. An empirical illustration on female wage rate determination shows that an extra year of work experience could raise the expected log wage rate by 0.1507 under our maintained hypothesis, while neglecting sample selection or interactive effects could lead to seriously biased estimates.
We argue that the fundamental issue of measuring treatment effects is to obtain good predictions of missing outcomes. However, there is no realized value to evaluate the quality of the predicted value generated by a particular method. The choice of prediction method has to rely on the compatibility of the observed data with the underlying assumptions of an approach and the predictive ability of the approach. In this paper, we selectively review some causal and non-causal approaches and discuss their limitations from this perspective.
This paper proposes a simple pooling prediction for a mixed panel via a panel autoregressive approximation (PAR) framework to predict returns and volatilities of stocks for 307 U.S. firms as well as 20 advanced and emerging economies. This mixed panel includes stationary I (0) and I (d) processes simultaneously. The new proposed PAR-forecasting approach does not require prior information on the exact form and fractional parameter of each series of the mixed panel. We also show this approach remains valid when pervasive or weak common factors are included in the mixed panel. Insights from our theoretical analyses are confirmed by a set of Monte Carlo experiments and empirical applications, throughout which we demonstrate that our approach is competitive with several existing forecasting methods, such as HAR of Corsi (2009) and PHAR of Bollerslev et al. (2018). In particular, when an individual unit is larger than the time dimension, it is found that the PAR-forecasting approach can generate smaller mean square prediction errors compared to HAR and PHAR.
The fundamental methodologies of machine learning and econometrics are reviewed. We also discuss the challenges of integrating the data-driven and model-based causal approaches and conjecture how it may yield new insights to empirical economic studies.
Summary We discuss methods of measuring the treatment effects of a unit through the use of other units in panel data by either the factor‐based (FB) approach or the linear projection (LP) approach under different sample configurations of cross‐sectional dimension and time series dimension . We show that the LP approach in general yields smaller mean square prediction error than the FB approach when either both and are large or fixed and or fixed and large. The Monte Carlo simulation and empirical example are also conducted to consider their finite sample performances.
We propose a method to conduct uniform inference for the (optimal) value function, that is, the function that results from optimizing an objective function marginally over one of its arguments. Marginal optimization is not Hadamard differentiable (that is, compactly differentiable) as a map between the spaces of objective and value functions, which is problematic because standard inference methods for nonlinear maps usually rely on Hadamard differentiability. However, we show that the map from objective function to an Lp functional of a value function, for 1≤p≤∞, are Hadamard directionally differentiable. As a result, we establish consistency and weak convergence of nonparametric plug-in estimates of Cramér–von Mises and Kolmogorov–Smirnov test statistics applied to value functions. For practical inference, we develop detailed resampling techniques that combine a bootstrap procedure with estimates of the directional derivatives. In addition, we establish local and uniform size control of one-sided tests which use the resampling procedure. Monte Carlo simulations assess the finite-sample properties of the proposed methods and show accurate empirical size and nontrivial power of the procedures. Finally, we apply our methods to the evaluation of a job training program using bounds for the distribution function of treatment effects.
The authors consider the quasi maximum likelihood (MLE) estimation of dynamic panel models with interactive effects based on the Ahn et al. (2001, 2013) quasi-differencing methods to remove the interactive effects. The authors show that the quasi-difference MLE (QDMLE) over time is inconsistent when N→∞ whether T is fixed or goes to infinity. On the other hand, the QDMLE is consistent and asymptotically unbiased if the difference is taken over individuals when T is large whether N is fixed or large. Monte Carlo studies are conducted to compare the performance of the QDMLE using different quasi-difference methods.
Most literature works on estimating treatment effects assume that the observed data are either under the specific “treatment” or not. However, in many cases, the observed data could be subject to multiple treatments. We propose to combine econometric methods developed for different purposes to disentangle the multiple treatment effects. We illustrate this strategy by considering the impact of global pandemic v.s. the strictest “lockdown” policy of Hubei, China implemented in January, 2020. We show that although the strictest “lockdown” policy quickly contained the spread of COVID-19, it also inflicted huge economic loss on Hubei economy. It lowered Hubei GDP by about 37% compared to the level had there been no “lockdown” under the pandemic. However, even though Hubei economy managed to recover from the “lockdown”, it could not escape the global impact of pandemic. Its economy is still about 90% of the level had there been no pandemic.
In 2016, the city of Shanghai increased the minimum down payment rate requirement for purchasing various types of properties. We study the treatment effect of this major policy change on Shanghai's housing market by employing panel data from March 2009 to December 2021. Since the observed data are either in the form of no treatment or under the treatment but before and after the outbreak of COVID-19, we use the panel data approach suggested by Hsiao et al. (J Appl Econ, 27(5):705-740, 2012) to estimate the treatment effects and a time-series approach to disentangle the treatment effects and the effects of the pandemic. The results suggest that the average treatment effect on the housing price index of Shanghai over 36 months after the treatment is -8.17%. For time periods after the outbreak of the pandemic, we find no significant impact of the pandemic on the real estate price indices between 2020 and 2021.
This paper proposes a nonparametric framework to estimate point-wise price elasticities using aggregate market data. We derive a new constructive price elasticity estimator based on the nonparametric control function framework (Newey et al., 1999), integrate the bootstrap averaged (bagged) nearest neighbors predictor (Demirkaya et al., 2022) into this framework, and implement the estimation using just-in-time compilation and parallel computation. A series of Monte Carlo simulations across a wide range of data-generating processes show that our elasticity estimator is fast to compute and achieves both precise estimates and robust inferences. In an empirical application, we demonstrate that this method can (1) flexibly estimate price response and substitution patterns and (2) directly inform optimal pricing and supply-side counterfactuals
This research demonstrates the success of the new CDAR-family global market integration indices in the prediction of equity returns, which have not been thoroughly investigated in the past. We comprehensively investigate the predictive ability of the new CDAR-family indices by means of commonly used spillover indices, financial and macroeconomic variables with existing predictive approaches. Empirical results reveal that (i) a world factor is necessary in the out-of-sample prediction of the U.S equity premium; (ii) the new indices improve the forecasting accuracy; and (iii) the new indices have strongly better predictive power during expansions than recessions.
Panel data provide the possibilities of estimating individual treatment effects for multiple individuals. Two issues are considered: (1) differences in the estimated individual treatment effects are due to heterogeneity or a chance mechanism? (2) what is the best way to estimate the average treatment effects? Testing and aggregation methods are suggested. Monte Carlo simulations are also conducted to shed light on these two issues. An empirical analysis on the involvement of underground organization in China’s Peer-to-Peer (P2P) activities through the “anti-gang” campaign is also provided.
This book provides a comprehensive, coherent, and intuitive review of panel data methodologies that are useful for empirical analysis. Substantially revised from the second edition, it includes two new chapters on modeling cross-sectionally dependent data and dynamic systems of equations. Some of the more complicated concepts have been further streamlined. Other new material includes correlated random coefficient models, pseudo-panels, duration and count data models, quantile analysis, and alternative approaches for controlling the impact of unobserved heterogeneity in nonlinear panel data models.
This paper uses a panel data approach to assess the evolution of economic consequences of the drastic lockdown policy in the epicenter of COVID-19-the Hubei Province of China during worldwide curbs on economic activity. We find that the drastic 76-day COVID-19 lockdown policy brought huge negative impacts on Hubei's economy. In 2020:q1, the lockdown quarter, the treatment effect on GDP was about 37% of the counterfactual. However, the drastic lockdown also brought the spread of COVID-19 under control in little more than two months. After the government lifted the lockdown in early April, the economy quickly recovered with the exception of passenger transportation sector which rebounded not as quickly as the rest of the general economy.
We propose a transformed estimator for the slope coefficients of panel models with interactive effects. The transformed estimation method does not require the prior knowledge of the dimension of factor structure. It is consistent and asymptotically normally distributed under fairly general conditions when N is fixed and T ->infinity or T is fixed and N ->infinity, or when both N and T are large and N/T -> a not equal 0 ainfinity. Extensive Monte Carlo simulations are also conducted to examine the finite sample performance of the transformed estimation method.
The observed data could be the joint outcomes due to several events. We propose a combined time series modeling and panel data program evaluation approach to separate the impacts of different events or interventions. As an illustration, we analyze the outcomes of the COVID-19 in Hubei, China in December 2019. We estimate the net impact of COVID-19 without the “lockdown” policy and the net impact of the strictest “lockdown” policy under the pandemic and its aftermath. We show that although the strictest “lockdown” policy quickly contained the spread of COVID-19, it also inflicted huge economic loss on Hubei economy. It lowered Hubei GDP by about 37% compared to the level had there been no “lockdown” under the pandemic. However, even though Hubei economy recovered from the “lockdown”, it could not escape the global impact of pandemic. Its economy is still about 90% of the level had there been no pandemic.
The authors use a reduced form state-dependent labor participation decision model to illustrate that parameter stability is achieved only if a model properly takes account the observed sample heterogeneity and unobserved sample heterogeneity provided (external) conditions of a model stay constant. Our analysis of the dynamic response path to a health shock using Australian HILDA panel data from 2002 to 2009 shows that experiencing an event by itself can only have a temporary effects. The long-run equilibrium condition is independent of initial conditions or shocks that do not last. In other words, if experiencing an event does not lead to changes in the response parameters such as the real business cycle (Kydland & Prescott, 1977, 1982) or dynamic stochastic general equilibrium model (DSGE, e.g., Sbordone et al., 2010) assumed, policy change may only change the short-run response path. There is no long-term impact for a policy change. On the other hand, if a policy change leads to changes in the decision rules (e.g., the recent US–China trade friction) as the Lucas critique (1976) implies, then there is no other way to evaluate the impact of a policy except to explicitly model how agents respond to the policy change.