
Quantile regression is one of the most important methods to estimate heterogeneous effects on a variable of interest. In many applications there is a subset of covariates of interest, while the rest operate as controls in the regression equation. This work presents a straightforward empirical strategy for situations in which the control variables have a homogeneous effect on the conditional distribution. We develop the asymptotic theory for the proposed estimator and the corresponding inference procedures. An application using environmental pollution data illustrates the method by estimating the Environmental Kuznets Curve.
This paper develops a consistent series-based specification test for semiparametric panel data models with fixed effects. The test statistic resembles the Lagrange Multiplier (LM) test statistic in parametric models and is based on a quadratic form in the restricted model residuals. The use of series methods facilitates both estimation of the null model and computation of the test statistic. The asymptotic distribution of the test statistic is standard normal, so that appropriate critical values can easily be computed. The projection property of series estimators allows me to develop a degrees of freedom correction. This correction makes it possible to account for the estimation variance and obtain refined asymptotic results. It also substantially improves the finite sample performance of the test.
The QR.break package provides methods for detecting, estimating, and conducting inference on multiple structural breaks in linear quantile regression models, based on one or multiple quantiles and applicable to both time series and repeated cross-sectional data. The main function, rq.break() , returns testing and estimation results based on user specifications of the quantiles of interest, the maximum number of breaks allowed, and the minimum length of a single regime. This note outlines the underlying methods and explains how to use the main function with two datasets: a time series dataset on U.S. real GDP growth rates and a repeated cross-sectional dataset on youth drinking and driving behavior. Both datasets are included in the package available on CRAN.
Heteroskedasticity robust standard errors are often presented as a formula that is not directly related to classical standard errors derived under homoskedasticity. This short paper introduces a moment-based result relating these two estimators through the correlation between squared residuals and squared regressors. Though the result does not rely on normality, it admits a simple approximation when all variables are normally distributed. This representation can be useful both for pedagogical purposes in undergraduate courses that do not use matrix algebra and in highlighting the relative magnitude of robust to non-robust standard errors.
This paper examines the properties of the ordinary least squares (OLS) estimator when applied to a model with a non-linear relationship between outcome and a discrete regressor. I investigate what parameters OLS estimates in such a case, focusing on both level and incremental effects. The analysis reveals that the OLS estimand is a convex average of incremental effects, but weights can be negative for level effects and in the presence of neglected heterogeneity. An empirical application to a wage equation demonstrates these issues, highlighting the importance of using unrestricted models or carefully considering the limitations of OLS estimates in similar situations.
Abstract Unknown parameters, including regression coefficients, in state space models can be estimated by maximum likelihood. An alternative approach is to augment the state vector to include regression coefficients. However, the state estimator obtained by the Kalman filter is numerically different from the maximum likelihood estimator. We address the discrepancy by a novel method based on proper distributions returned by the ordinary Kalman filter without dependency on diffuse initialization. We prove that maximizing a low-dimensional objective function that combines the likelihood, the filtering mean and variance can reproduce the high-dimensional maximum likelihood results.
Abstract This paper deals with a detailed analysis of the first-order diagonal bilinear time series model, first proposed in Granger and Andersen (1978. An Introduction to Bilinear Time Series Models. Göttingen: Vandenhoeck & Ruprecht). This model allows for sequences of “outliers” in the data. We show that the model has a variety of features that we can observe in practice, while we also document that the bilinear features show up in just a limited number of observations. When the moment restrictions are close, parameter estimation becomes difficult. When the parameters are further away from the moment restrictions, parameter estimation is easy. Yet, in those latter cases, approximative linear models appear to generate equally accurate fit and forecasts. In sum, in cases of proper inference on a bilinear model, the model is barely relevant for forecasting.
This paper addresses computational challenges in estimating Quantile Regression with Selection (QRS). The estimation of the parameters that model self-selection requires the estimation of the entire quantile process several times. Moreover, closed-form expressions of the asymptotic variance are too cumbersome, making the bootstrap more convenient to perform inference. Taking advantage of recent advancements in the estimation of quantile regression, along with some specific characteristics of the QRS estimation problem, I propose streamlined algorithms for the QRS estimator. These algorithms significantly reduce computation time through preprocessing techniques and quantile grid reduction for the estimation of the copula and slope parameters. I show the optimization enhancements with some simulations. Lastly, I show how preprocessing methods can improve the precision of the estimates without sacrificing computational efficiency. Hence, they constitute a practical solutions for estimators with non-differentiable and non-convex criterion functions such as those based on copulas.
Abstract This paper applies the Dagum Type III model to measure household net wealth inequality in Indonesia utilising data from the Indonesian Family Life Survey (IFLS) 1993–2014. The results are the distribution of household net wealth in Indonesia is right-skewed, with long-and sparse-hand tails that reflect a large proportion of households that have very low net wealth and a small proportion of households that have very high net wealth. Further, the inequality of household net wealth in Indonesia declined, as shown by the decrease in the Gini coefficient.
Abstract This paper introduces a Stein-like shrinkage method for estimating slope coefficients and forecasting in first order dynamic regression models under structural breaks. The model allows for unit root and non-stationary regressors. The proposed shrinkage estimator is a weighted average of a restricted estimator that ignores the break in the slope coefficients, and an unrestricted estimator that uses the observations within each regime. The restricted estimator is the most efficient estimator but inconsistent when there is a break. However, the unrestricted estimator is consistent but not efficient. Therefore, the proposed shrinkage estimator balances the trade-off between the bias and variance efficiency of the restricted estimator. The averaging weight is proportional to the weighted distance of the restricted estimator, and the unrestricted estimator. We derive the analytical large-sample approximation of the bias, mean squared error, and risk for the shrinkage estimator, the unrestricted estimator, and the restricted estimator. We show that the risk of the shrinkage estimator is lower than the risk of the unrestricted estimator under any break size and break points. Moreover, we extend the results for the model with a unit root and non-stationary regressors. We evaluate the finite sample performance of our proposed method via extensive simulation study, and empirically in forecasting output growth.
This paper introduces an innovative approach to identifying and estimating the parameters of interest in the widely recognized linear-in-means regression model under conditions where the initial randomization of peers determines the observed network. We assert that peers who are initially randomized do not produce social effects. However, after randomization, agents can endogenously develop significant connections that potentially generate peer influences. We present a moment condition that compiles local heterogeneous identifying information for all agents within the population. Under the assumption of ψ-dependence in the endogenous network space, we propose a Generalized Method of Moments (GMM) estimator, which is proven to be consistent, asymptotically normally distributed, and straightforward to implement using commonly available statistical software due to its closed-form expression. Monte Carlo simulations demonstrate the GMM estimator's strong small-sample performance. An empirical analysis utilizing data from Hong Kong high school students reveals substantial positive spillover effects on math test scores among study partners in our sample, provided that their seatmates were exogenously assigned by their teachers.
Abstract I compare two popular methods of estimation for linear panel data models with unobserved factors: the first eliminates the factors with a parameterized quasi-long-differencing (QLD) transformation. The other, referred to as common correlated effects (CCE), uses cross-sectional averages of the data to proxy for the factor space. I show that the CCE assumptions imply unused moment conditions that can be exploited by the QLD transformation. I also derive new linear estimators that weaken identifying assumptions and have desirable theoretical properties. Unlike CCE, these estimators do not require the number of covariates to be less than the number of time periods. I provide the first proof of a fixed-T consistent mean group estimator for heterogeneous linear models with interactive fixed effects. I investigate the effects of per-student expenditure on standardized test performance using data from the state of Michigan.
Abstract In this paper, we review several estimators of the average treatment effect (ATE) that belong to three main groups: regression, weighting and doubly robust methods. We unify the exposition of these estimators within an M-estimation framework and we derive their variance estimators from the sandwich form variance-covariance matrix of the M-Estimator. Additionally, we re-estimate the causal return to higher education on earnings by the reviewed methods using the rich dataset provided by the British National Child Development Study (NCDS) as an empirical illustration.
Abstract Bunching estimation of distortions in a distribution around a policy threshold provides a means of studying behavioral parameters. Standard cross-sectional bunching estimators rely on identification assumptions about heterogeneity that I show can be violated by serial dependence of the choice variable or attrition related to the threshold. I propose a bunching estimation design that exploits panel data to obtain identification from relative within-agent changes in income and to estimate new parameters. Simulations using household income data demonstrate the benefits of the panel design. An application to charitable organizations demonstrates opportunities for estimating elasticity correlates, causal effects, and extensive-margin responses.
Abstract This paper describes a simple and interesting application of structural equation modeling for a single lecture in an undergraduate econometrics course to introduce students to the concept of using data to recover latent variables. The application centers around using hourly observations on ride wait times at Disney’s Magic Kingdom to infer how crowded it is at the theme park. Pedagogically, the material is presented in the context of the linear regression model, so the discussion works to enhance students’ understanding of core material, not to introduce new disparate methods. The application provides interesting economic-based insights, like which ride’s wait times are categorically most informative about how crowded it is at the park.
Abstract We investigate whether receiving health information changes human behavior by using a novel approach to inference in the fuzzy regression discontinuity design. The approach is robust to the strength of identification and allows for mean squared error optimal bandwidths as well as undersmoothing. It is based on the Anderson-Rubin test in the instrumental variable literature augmented with either robust bias correction or critical value adjustment. We find that the resulting confidence sets of the treatment effect are mostly wide or even unbounded. These findings indicate that we could not rule out most magnitudes of behavior change, including zero and non-zero ones.
Abstract This paper proposes a novel estimator for nonparametric instrumental regression while controlling for additive two-way fixed effects. In particular, the Landweber–Fridman regularization, to overcome the ill-posed inverse problem in the nonparametric instrumental regression procedure, is combined with the local-within two-ways fixed effect estimator presented by Lee, Y., D. Mukherjee, and A. Ullah. (2019. “Nonparametric Estimation of the Marginal Effect in Fixed-Effect Panel Data Models.” Journal of Multivariate Analysis 171: 53–67). Compared to other estimators in this context, an appealing feature is its flexible applicability with respect to different panel model specifications, i.e. models comprising either individual, temporal, or two-way fixed effects. The estimator’s performance is tested on simulated data, where a Monte Carlo study reveals good finite sample behaviour. Confidence intervals are provided by applying the wild bootstrap.
Abstract We investigate the properties of a systematic bias that arises in the synthetic control estimator in panel data settings with finite pre-treatment periods, offering intuition and guidance to practitioners. The bias comes from matching to idiosyncratic error terms (noise) in the treated unit and the donor units’ pre-treatment outcome values. This in turn leads to a biased counterfactual for the post-treatment periods. We use Monte Carlo simulations to evaluate the determinants of the bias in terms of error term variance, sample characteristics and DGP complexity, providing guidance as to which situations are likely to yield more bias. We also offer a procedure to reduce the bias using a direct computational bias-correction procedure based on re-sampling from a pilot model that can reduce the bias in empirically feasible implementations. As a final potential solution, we compare the performance of our corrections to that of an Interactive Fixed Effects model. An empirical application focused on trade liberalization indicates that the magnitude of the bias may be economically meaningful in a real world setting.
Abstract We study the Feasible Generalized Least-Squares (FGLS) estimation of the parameters of a linear regression model in the presence of heteroskedasticity of unknown form in the errors. We suggest a Lasso based procedure to estimate the skedastic function of the residuals. The advantage of using Lasso is that it can handle a large number of potential covariates, yet still yields a parsimonious specification. Using extensive simulation experiments, we show that our suggested procedure always provide some improvements in the precision of the parameter of interest (lower Mean-Squared Errors) when heteroskedasticity is present and is equivalent to OLS when there is none. It also performs better than previously suggested procedures. Since the fitted value of the skedastic function falls short of the true specification, we form confidence intervals using a bias-corrected version of the usual heteroskedasticity-robust covariance matrix estimator. These have the correct size and substantially shorter length than when using OLS. Our method is applicable to both cross-section (with a random sample) and time series models, though here we concentrate on the former.
Abstract When a sample combines data from two or more groups, multivariate regression yields a matrix-weighted average of the group-specific coefficient vectors. However, it is possible that the weighted average of a specific coefficient falls outside the range of the group-specific coefficients, and it may even have a different sign compared to both group-level coefficients, a manifestation of Simpson’s paradox. The result of the combined regression is then prone to misinterpretation. The purpose of this paper is to raise awareness of this problem and to state conditions under which such non-convex weighting or sign reversal can arise, for a model with two regressors and two groups. Two illustrative examples, an investment equation estimated with panel data, and a cross-sectional earnings equation for men and women, highlight the relevance of these findings for applied work.