This article studies the principal component analysis (PCA) estimation of weak factor models with sparse loadings. We uncover an intrinsic near-sparsity preservation property for the PCA estimators of loadings, which comes from the approximately (block) upper triangular structure of the rotation matrix. It suggests an asymmetric relationship among factors: the sparsity of the rotated loadings for a stronger factor can be contaminated by the loadings from weaker ones, but the sparsity of the rotated loadings of a weaker factor is almost unaffected by the loadings of stronger ones. Then, we propose a simple alternative to the existing penalized approaches to sparsify the loading estimators by screening out the small PCA loading estimators directly, and construct consistent estimators for factor strengths. The proposed estimators perform well in finite samples, as shown by a set of Monte Carlo simulations.
This paper studies mutual fund performance evaluation, where non-random missing data is an important issue. Ignoring sample selection issues can lead to biased parameter estimates and invalid performance evaluation. Therefore, we propose a novel Beta selection model to characterize this non-random missingness mechanism. We can significantly reduce the bias in alpha estimates, by introducing our Beta selection mechanism into the Noise-Reduced Alpha (NRA) model of Harvey and Liu (2018). A computationally attractive Bayesian estimation procedure is also provided. Simulation results show that our proposed method achieves superior finite-sample performance compared to classical fund-by-fund ordinary least squares (OLS) and the baseline NRA model. Finally, we apply our method to Chinese mutual fund data, and find that: (i) our out-of-sample alpha forecasts exhibit greater predictive accuracy across all tested missingness rates; (ii) approximately 18% of funds exhibit statistically significant positive alphas, suggesting market-beating skill; and (iii) portfolios constructed using our method deliver significantly higher returns than those using alternative approaches.
This paper addresses the computational challenges of calculating post-processing posterior predictive p-values by introducing a novel approximation method using the asymptotic pivotal discrepancy function. Existing approaches usually have a heavy computational burden due to the adoption of resampling in calculation. Our study proposes an efficient alternative by employing a posterior-based Wald-type discrepancy function, which can eliminate the need for resampling and significantly reduce computational demands. Through simulations, we demonstrate that our method achieves comparable results to computationally intensive approaches while offering substantial computational efficiency gains. We further validate our approach using two real-world datasets: CEO compensation and firm performance (analyzed via linear regression) and daily Pound/Dollar exchange rates (modeled using stochastic volatility). Our findings highlight the method's adaptability and efficacy across diverse applications, advancing the practicality of Bayesian model evaluation and inference.
This article proposes a One-Covariate-at-a-time Multiple Testing (OCMT) approach to choose significant variables in high-dimensional nonparametric additive regression models. Similarly to Chudik, Kapetanios, and Pesaran, we consider the statistical significance of individual nonparametric additive components one at a time and take into account the multiple testing nature of the problem. Both one-stage and multiple-stage procedures are considered. The former works well in terms of the true positive rate only if the net effects of all signals are strong enough; the latter helps to pick up hidden signals that have weak net effects. Simulations demonstrate the good finite-sample performance of the proposed procedures. As an empirical illustration, we apply the OCMT procedure to a dataset extracted from the Longitudinal Survey on Rural Urban Migration in China. We find that our procedure works well in terms of out-of-sample root mean square forecast errors, compared with competing methods such as adaptive group Lasso (AGLASSO).
This paper studies the principal component (PC) method-based estimation of weak factor models with sparse loadings. We uncover an intrinsic near-sparsity preservation property for the PC estimators of loadings, which comes from the approximately upper triangular (block) structure of the rotation matrix. It implies an asymmetric relationship among factors: the rotated loadings for a stronger factor can be contaminated by those from a weaker one, but the loadings for a weaker factor is almost free of the impact of those from a stronger one. More importantly, the finding implies that there is no need to use complicated penalties to sparsify the loading estimators. Instead, we adopt a simple screening method to recover the sparsity and construct estimators for various factor strengths. In addition, for sparse weak factor models, we provide a singular value thresholding-based approach to determine the number of factors and establish uniform convergence rates for PC estimators, which complement Bai and Ng (2023). The accuracy and efficiency of the proposed estimators are investigated via Monte Carlo simulations. The application to the FRED-QD dataset reveals the underlying factor strengths and loading sparsity as well as their dynamic features.
This paper provides nonparametric specification tests for the commonly used homogeneous and stable coefficients structures in panel data models. We first obtain the augmented residuals by estimating the model under the null hypothesis and then run auxiliary time series regressions of augmented residuals on covariates with time-varying coefficients (TVCs) via sieve methods. The test statistic is then constructed by averaging the squared fitted values, which are close to zero under the null and deviate from zero under the alternatives. We show that the test statistic, after being appropriately standardized, is asymptotically normal under the null and under a sequence of Pitman local alternatives. A bootstrap procedure is proposed to improve the finite sample performance of our test. In addition, we extend the procedure to test other structures, such as the homogeneity of TVCs or the stability of heterogeneous coefficients. The joint test is extended to panel models with two-way fixed effects. Monte Carlo simulations indicate that our tests perform reasonably well in finite samples. We apply the tests to re-examine the environmental Kuznets curve in the United States, and find that the model with homogenous TVCs is more appropriate for this application.
This paper considers a probit model for panel data in which the individual effects vary over time by interacting with unobserved factors. In estimation we adopt a correlated random effects approach for individual effects to get around the incidental parameter problem. This allows us to construct (asymptotically) unbiased estimators for average marginal effects (AMEs), which are often the ultimate quantities of interest in many empirical studies. We derive the asymptotic distributions for the AME estimators as well as provide the consistent estimators for their asymptotic variances. Next, we design a specification test for detecting whether individual effects are time-varying or not, and establish the asymptotic distribution for the proposed test statistic under the null hypothesis of no time variation of individual effects. Monte Carlo simulations demonstrate satisfactory finite sample performance of our proposed method. An empirical application to study the effect of fertility on labour force participation (LFP) is provided. We find that fertility has a larger impact on female LFP in Germany than in the US during the 1980s. We also provide some new empirical evidence of a even stronger effect of fertility on LFP during the 2010s in Germany, which might call for a reconsideration of relevant policies recently enacted such as the subsidized child care programme.
A new Bayesian bandwidth selection procedure is proposed for nonparametric kernel estimates based on the sequential Monte Carlo method. Compared with the existing Bayesian bandwidth selector of Zhang et al. (2009), this new method can enhance the convergence to the global optimum with a substantially faster computation speed. In particular, the method offers an improved out-of-sample performance as shown by simulations. The bandwidth selector is applied to the option state price density, production function, and nonparametric relationship between oil and stock index returns; results indicate that our proposed method outperforms other methods in terms of the mean square error and log-likelihood in all applications.
We study the nonparametric estimation and specification testing for partially linear functional-coefficient dynamic panel data models, where the effects of some covariates on the dependent variable vary nonparametrically according to a set of low-dimensional variables. Based on the sieve approximation of unknown slope functions, we propose a sieve 2SLS procedure to estimate the model. The asymptotic properties of the estimators of both parametric and nonparametric components are established when sample size N and T tend to infinity jointly. A nonparametric specification test for the constancy of slopes is also proposed. We show that after being appropriately standardized, the test is asymptotically normally distributed under the null hypothesis. The asymptotic properties of the test is also studied under a sequence of local Pitman alternatives and global alternatives. A set of Monte Carlo simulations show that our sieve 2SLS estimators and specification test perform remarkably well in finite samples. We apply our method to study the impact of income on democracy, and find strong evidence of nonlinear/nonconstant effect of income on democracy.
In this paper, we consider the generalized method of moment (GMM) and simple instrumental variable (IV) type estimation of dynamic panel data models with both individualspeci?c e?ects and heterogeneous time trend. We consider the forward demeaning (FOD) proposed by Hayakawa et al (2017) and the double ?rst di?erence (FD) to remove both the individual-speci?c e?ects and heterogeneous trend. We establish the asymptotic properties of the GMM estimation of the lag coe?cient and ?nd that the GMM estimation using FOD is asymptotically biased of order square root of T/N, while the GMM using FD is asymptotically biased of order square root of T^3/N. We also establish the asymptotic unbiasedness of the simple IV estimation. Monte Carlo simulations con?rm our ?ndings in this paper.
We propose a simple and fast approach to identify and estimate the unknown group structure in panel models by adapting the M-estimation method. We consider both linear and nonlinear panel models where the regression coefficients are heterogeneous across groups but homogeneous within a group and the group membership is unknown to researchers. The main result of the paper is that under certain assumptions, our approach is able to provide uniformly consistent estimation as long as the number of groups used in estimation is not smaller than the true number of groups. We also show that, asymptotically, our method may partition some true groups into further subgroups, but cannot mix units from different groups. When the true number of groups is used in estimation, all units can be categorized correctly with probability approaching one, and we establish the limiting distribution for the estimators of the group parameters. In addition, we provide an information criterion to select the number of groups, and establish the consistency of the selection criterion under some mild conditions. Monte Carlo simulations are conducted to examine the finite sample performance of the proposed method. The findings in the simulation confirm our theoretical results in the paper. Applications to two real datasets also highlight the necessity to consider both individual heterogeneity and group heterogeneity in the model.
In this article, we describe the implementation of fitting partially linear functional-coefficient panel models with fixed effects proposed by An, Hsiao, and Li [2016, Semiparametric estimation of partially linear varying coefficient panel data models in Essays in Honor of Aman Ullah ( Advances in Econometrics, Volume 36)] and Zhang and Zhou (Forthcoming, Econometric Reviews). Three new commands xtplfc, ivxtplfc, and xtdplfc are introduced and illustrated through Monte Carlo simulations to exemplify the effectiveness of these estimators.
It is shown in the literature that the Arellano-Bond type generalized method of moments (GMM) of dynamic panel models is asymptotically biased (e.g., Hsiao & Zhang, 2015; Hsiao & Zhou, 2017). To correct the asymptotical bias of Arellano-Bond GMM, the authors suggest to use the jackknife instrumental variables estimation (JIVE) and also show that the JIVE of Arellano-Bond GMM is indeed asymptotically unbiased. Monte Carlo studies are conducted to compare the performance of the JIVE as well as Arellano-Bond GMM for linear dynamic panels. The authors demonstrate that the reliability of statistical inference depends critically on whether an estimator is asymptotically unbiased or not.
This paper introduces a novel forecasting method based on a time-varying diffusion index model, where both factor loadings and regression coefficients are allowed to be time-varying. We first obtain the local principal component analysis (PCA) estimators for the latent factors and then estimate the factor augmented forecasting regression with time-varying coefficients nonparametrically. A feasible forecast is proposed by combining the estimated factors and the nonparametric estimators of coefficients. A set of Monte Carlo simulations demonstrates better performance of our proposed method than the standard diffusion index forecasters based on rolling windows. An empirical application of forecasting US macroeconomic variables is provided.
The usual t test, the t test based on heteroskedasticity and autocorrelation consistent (HAC) covariance matrix estimators, and the heteroskedasticity and autocorrelation robust (HAR) test are three statistics that are widely used in applied econometric work. The use of these significance tests in trend regression is of particular interest given the potential for spurious relationships in trend formulations. Following a longstanding tradition in the spurious regression literature, this paper investigates the asymptotic and finite sample properties of these test statistics in several spurious regression contexts, including regression of stochastic trends on time polynomials and regressions among independent random walks. Concordant with existing theory (Phillips 1986, 1998; Sun 2004, 2014b) the usual t test and HAC standardized test fail to control size as the sample size n → ∞ in these spurious formulations, whereas HAR tests converge to well-defined limit distributions in each case and therefore have the capacity to be consistent and control size. However, it is shown that when the number of trend regressors K → ∞ , all three statistics, including the HAR test, diverge and fail to control size as n → ∞ . These findings are relevant to high-dimensional nonstationary time series regressions where machine learning methods may be employed.
A two-step estimation procedure is proposed to estimate the time-invariant effects, i.e., the slopes of the time-invariant regressors, in dynamic panel data models. In the first step, generalized method of moments (GMM) is used to estimate the time-varying effects, and the second step is to run cross-sectional OLS regression of the time series average of the residuals from the GMM estimation on the time-invariant regressors to estimate the time-invariant effects. It is shown that the OLS estimator of time-invariant effects is N-consistent and asymptotically normally distributed. A consistent estimator for the asymptotic variance of the estimator is also provided, which is robust to errors with heteroscedasticity and works well even if the errors are serially correlated. Monte Carlo simulations confirm the theoretical findings. Application to income dynamics highlights the importance of estimating time-invariant effects such as education, race and gender in return to schooling.
In this paper, we consider the estimation of dynamic panel data models. We establish the equivalence of the GMM estimator proposed by Alvarez and Arellano, which is based on the forward differenced model using all lagged variables as instruments, and the double filter instrumental variables estimator (DIV for short) proposed by Hayakawa, which uses the backward differenced lags as instruments. Since the DIV estimator is asymptotically unbiased, thus we suggest using the DIV estimator for estimation of dynamic panel models. Monte Carlo simulations confirm our findings in this paper.
In empirical finance, interest rate models have been widely used for modeling short-term interest rate. Under the framework of the hypothesis testing, this paper provides a Bayesian approach for comparing a range of alternative models. These compared models are nested in a general single-factor diffusion process for the short-term interest rate, with each alternative model indexed by the level effect parameter for the volatility. The performance of the developed procedure is illustrated by an empirical example of Eurodollar deposit rates.
In this paper, we extend the kink regression model with an unknown threshold in Hansen (2017) to the panel data framework, where the cross-sectional dimension (N) goes to infinity and the time period (T) is fixed. Following the literature of threshold regressions, we propose an estimator based on the within-group transformation. Under fixed threshold effect assumption, we establish that the slope and threshold estimators are jointly normally distributed with the same convergence rate OpN−1∕2 and a non-zero asymptotic covariance. We also suggest a sup-Wald test for the presence of kink effect, and derive its limiting distribution. A bootstrap procedure is proposed to obtain the bootstrap p-values to improve the finite sample performance of the test. Monte Carlo simulations show that the FE estimator and the sup-Wald test perform quite well in estimating the unknown parameters and testing for kink effect, respectively.
This paper provides a practical test for strict exogeneity in linear panel data models with fixed effects when the number of individuals N goes to infinity while the number of time periods T is fixed. The test is based on the supremum of a sequence of Wald test statistics. Under suitable conditions, we establish the asymptotic distribution of the test statistic and consistency of the test. A bootstrap procedure is proposed to improve the finite sample performance and the validity of the procedure is justified. We investigate the finite sample performance of the test via a small set of Monte Carlo simulations.