In this paper, we demonstrate the numerical equivalence between the within estimator and the Mundlak estimator for widely used three-dimensional (3D) balanced panel data models. The Mundlak estimator is obtained from the OLS regression, including relevant sample averages as additional regressors. This suggests that the three estimation methods, namely, within, Mundlak, and least squares dummy variable (LSDV), produce numerically identical estimates for balanced 3D models.
This paper establishes the inferential theory for unbalanced panel data models with interactive fixed effects. We propose a two-step estimation algorithm with the first step obtaining an initial consistent estimator followed by an alternating maximization procedure. We prove that the alternating maximization procedure is a contractionary mapping and the final estimator is asymptotically normal, as long as the initial estimator is consistent. We also develop analytical bias corrections according to the derived asymptotic bias expressions and the observed missing pattern. Our results cover important missing patterns such as completely exogenous missing, selection on regressors/factors/loadings and block/staggered missing, and we also show that our results can be readily extended to cases with a Heckman correction term or more general settings. An empirical analysis of the U.S. state-level tax rates from 1951 to 2000 with missing data reveals persistence in tax rates, while state income influences different taxes in varying ways.
This paper studies estimation and variable selection in conditional factor models with high-dimensional instruments, where the coefficient matrix exhibits a low-rank and row-sparse structure. We propose a multi-stage estimation procedure that combines nuclear norm regularization and adaptive group LASSO regression to consistently estimate latent factors and row-sparse loading coefficients, while selecting relevant instrumental characteristics. We establish theoretical results for estimation consistency, selection consistency, and post-LASSO inference for estimators of factors and loading coefficients at multiple stages. Furthermore, we implement a singular value thresholding procedure to determine the number of factors. Simulation results demonstrate the effectiveness of our estimators in consistently estimating factor loadings, selecting the appropriate number of factors, and conducting inference. Finally, we apply the proposed method to an empirical study on asset return prediction, showcasing its practical utility in real-world applications.
We consider a common correlated effects estimator for panel quantile regression models featuring heterogeneous slope coefficients and interactive fixed effects. Using convolution smoothing to facilitate the derivation of asymptotic properties, our method allows the number of time periods ( T ) to grow at a much slower rate than the number of cross-sectional units ( N ). We establish that the mean group estimator converges to a limiting normal distribution, which may have a nonzero mean in our asymptotic framework. To address biases from the smoothed objective function and a relatively small T , we introduce a two-step bias correction procedure. Monte Carlo simulations demonstrate the method's validity in finite samples, and we apply it to analyze the sharp rise in demand for high-skilled workers in U.S. manufacturing.
This paper studies the least squares estimation of unbalanced heterogeneous panel data models with interactive fixed effects. The asymptotic properties are established in a unified framework that covers both static and dynamic panels and both random-type and block-type missing patterns, such as completely exogenous missing, selection on regressors/factors/loadings and block missing/staggered missing/mixed frequency. These results allow us to account for heterogeneous regression coefficients, outcome dynamics and missing indicator dynamics in factor extraction, mean-group coefficients estimation, matrix completion and average treatment effect estimation. Our results do not require the covariates to have a factor structure. Interestingly, we find that the identification conditions imposed in the literature to establish consistency also ensure the well-behavior of the Hessian. Empirically, we apply our method to examine the effects of welfare waivers on per-capita caseloads of the AFDC program.
We consider a correlated random coefficient panel data model with two-way fixed effects and interactive fixed effects in a fixed T framework. The model allows slope coefficients to be arbitrarily correlated with the regressors, accommodating flexible forms of heterogeneity. We propose a two-way mean group estimator for the expected value of the slope coefficient and propose a leave-one-out jackknife method for valid inference. We apply our new methods to examine the relationship between healthcare expenditure and income.
This paper proposes a factor model with many latent groups. The population is divided into multiple groups, where the number of groups is large, the group sizes are distinct (even growing at different rates), and the group memberships are unknown. Within each group, a latent factor structure exists, resulting in a model characterized by many weak factors. In this context, traditional principal component analysis (PCA) is not directly applicable. A key challenge lies in accurately determining the number of groups and identifying group memberships. To address this, we propose an innovative and computationally efficient method based on a network derived from pairwise correlation estimates. Under mild regularity conditions, the proposed method consistently determines the correct number of groups and their memberships. Monte Carlo simulations demonstrate that the estimation algorithm performs well in finite samples across a variety of specifications, highlighting its robustness and practical utility. An empirical application to asset pricing identifies six economically interpretable latent groups among 950 test assets.
The widely-used common correlated effects (CCE) estimator, pioneered by Pesaran (2006), is computed using least squares applied to auxiliary regressions where the observed regressors are augmented with cross-sectional averages of the dependent variable and regressors. However, the CCE estimator requires a crucial rank condition and becomes inconsistent when this condition is violated and the factor loadings of the x- and y -equations are correlated, causing an endogeneity issue. This paper proposes a generalized CCE (GCCE) estimator by augmenting the regression with both cross-sectional and time-series averages of the regressors. We argue that the time-series average can serve as “control variables” to address the endogeneity issue. We show that the GCCE and CCE estimators are asymptotically equivalent when the rank condition holds, and the GCCE estimator remains consistent even when the rank condition is violated under our “control variable” condition. Therefore, our GCCE estimator is doubly robust, achieving consistency under either the rank condition or the “control variable” condition. Furthermore, we propose a leave-one-out jackknife method to conduct valid inferences regardless of whether the rank condition holds. Monte Carlo simulations demonstrate excellent performance of our estimators and inference methods in finite samples. We apply our new methods to two datasets to estimate the production function and gravity equation.
This article considers a consistent test for serial correlation of unknown form in the residual of panel data models with interactive fixed effects and possibly lagged dependent variables. Following the spirit of Hong, we construct a test statistic based on the comparison of a kernel-based spectral density estimator and the null spectral density. Under the null hypothesis, our test statistic is asymptotically N(0, 1) as both N and T tend to infinity. In contrast to existing tests for serial correlation, there is no need to specify the order of serial correlation about the alternative. We further examine the local and global power properties of test. A simulation study shows that our test performs well in finite samples. In the empirical application, we apply the test to study the impact of the divorce law reform on divorce rate. We find strong evidence of serial correlation in the residual, and our results show that the divorce law reform has permanent positive effects on divorce rates.
This article considers unified estimation and inference in panel autoregressive (PAR) models. The PAR coefficients are assumed to contain a latent group structure that allows the degree of persistence for each time series to be heterogeneous and unknown. We propose a novel penalized weighted least squares approach to simultaneously identify the unknown group membership and consistently estimate the PAR coefficients, regardless of whether the underlying PAR process is stationary, unit-root, near-integrated, or even explosive. Theoretically, we establish the classification consistency, oracle properties, and unified asymptotic normal distributions for the proposed Lasso-type estimators. Empirically, we apply our data-driven method to uncover the existence of firm-level hidden bubbles in the U.S. stock market that have not been accounted for in previous studies.
This paper establishes the inferential theory for the least squares estimation of large factor models with missing data. We propose a unified framework for asymptotic analysis of factor models that covers a wide range of missing patterns, including heterogenous random missing, selection on covariates/factors/loadings, block/staggered missing, mixed frequency and ragged edge. We establish the average convergence rates of the estimated factor space and loading space, the limit distributions of the estimated factors and loadings, as well as the limit distributions of the estimated average treatment effects and the parameter estimates in the factor-augmented regressions. These results allow us to impute the unbalanced panel appropriately or make inference for the heterogenous treatment effects. For computation, we can use the nuclear norm regularized estimator as the initial value for the EM algorithm and iterate until convergence. Empirically, we apply our method to test the average treatment effects of partisan alignment on grant allocation in UK.
We propose a two-step procedure to estimate a high dimensional discrete choice panel with interactive fixed effects where the initial and final estimators are obtained via a nuclear-norm regularized (NNR) maximum likelihood estimation and post-NNR iterated estimation, respectively. We apply the method to make counterfactual predictions of choice probabilities. Simulations demonstrate nice finite sample performance in estimation and tests. An illustrative application highlights the practical usefulness of our approach, revealing that the stock return of Fantasia Holdings Group Company Limited did experience a significant directional change following the 2021 credit rating downgrade event.
In this paper, we propose an easy-to-implement residual-based specification testing procedure for detecting structural changes in factor models, which is powerful against both smooth and abrupt structural changes with unknown break dates. The proposed test is robust to the over-specified number of factors, and serially and cross-sectionally correlated error processes. A new central limit theorem is given for the quadratic forms of panel data with dependence over both dimensions, thereby filling a gap in the literature. We establish the asymptotic properties of the proposed test statistic, and accordingly develop a simulation-based scheme to select critical value in order to improve finite sample performance. Through extensive simulations and a real-world application, we confirm our theoretical results and demonstrate that the proposed test exhibits desirable size and power in practice.
This article considers a three-dimensional latent factor model in the presence of one set of global factors and two sets of local factors. We show that the numbers of global and local factors can be estimated uniformly and consistently. Given the number of global and local factors, we propose a two-step estimation procedure based on principal component analysis (PCA) and establish the asymptotic properties of the PCA estimators. Monte Carlo simulations demonstrate that they perform well in finite samples. An application to the dataset of international trade reveals the relative importance of different types of factors.
This paper introduces a time-varying (TV) panel data model with interactive fixed effects where both the coefficients and factor loadings are allowed to change smoothly over time. We propose a local version of the least squares and principal component method to estimate the TV coefficients, TV factor loadings, and common factors simultaneously. We provide a bias-corrected local least squares estimator for the TV coefficients and establish the limiting distributions and uniform convergence of the bias-corrected coefficient estimators, estimated factors, and factor loadings in the large N and large T framework. Based on the estimates, we propose three test statistics to gauge possible sources of TV features. We establish the limit null distributions and the asymptotic local power properties of our tests. Simulations are conducted to evaluate the finite sample performance of our estimates and tests. We apply our theoretical results to analyze the Phillips curve using the U.S. state-level unemployment rates and nominal wages, and document significant TV behavior in both the slope coefficient and factor loadings.
In this article, we investigate a functional coefficient vector autoregressive model for conditional quantiles, in which the interdependences among tail risks such as Value-at-Risk are allowed to vary smoothly with a variable of general economy. Methodologically, we develop an easy-to-implement two-stage procedure to estimate functionals in the dynamic network system based on the deep learning method of neural networks and the local linear smoothing technique. We establish the consistency and the asymptotic normality of the proposed estimator under geometrically beta-mixing time series settings. The simulation studies are conducted to show that our new methods work fairly well. The potential of the proposed estimation procedures is demonstrated by an empirical study of constructing and estimating a new type of nonparametric dynamic financial network.
This paper studies a factor model with latent group structures. Given initial PCA estimates, we propose to apply the sequential binary segmentation algorithm (SBSA) to estimate the group structures and re-estimate the factors and loadings. In the case of unknown number of groups, we propose to estimate it via an information criterion. We establish the asymptotic properties of the proposed estimators. Simulations demonstrate that our proposed estimator outperforms some competitive approaches in finite samples.