The use of machine learning methods for predictive purposes has exponentially increased over the past two decades, but uncertainty quantification for predictive comparisons remains challenging and underdeveloped. This article addresses this gap by extending the classic inference theory for predictive ability in time series to modern machine learners, such as Lasso or Deep Learning. We investigate under which conditions such extensions are possible. For standard out-of-sample asymptotic inference to be valid with machine learning, two key properties must hold: (i) a zero-mean condition for the score of the prediction loss function and (ii) a "fast rate" of convergence for the machine learner. Absent any of these conditions, the estimation risk may be unbounded, and inferences invalid and very sensitive to sample splitting. For accurate inferences, we recommend an 80%-20% training-test splitting rule. We illustrate the wide applicability of our results with three applications: high-dimensional time series regressions with Lasso, Deep learning for binary outcomes, and a new out-of-sample test for the Martingale Difference Hypothesis (MDH). The theoretical results are supported by extensive Monte Carlo simulations and an empirical application evaluating the MDH of some major exchange rates.
We develop a systematic, omnibus approach to goodness-of-fit testing for parametric distributional models when the variable of interest is only partially observed due to censoring and/or truncation. In many such designs, tests based on the nonparametric maximum likelihood estimator are hindered by nonexistence, computational instability, or convergence rates too slow to support reliable calibration under composite nulls. We avoid these difficulties by constructing a regular (pathwise differentiable) Neyman-orthogonal score process indexed by test functions, and aggregating it over a reproducing kernel Hilbert space ball. This yields a maximum-mean-discrepancy-type supremum statistic with a convenient quadratic-form representation. Critical values are obtained via a multiplier bootstrap that keeps nuisance estimates fixed. We establish asymptotic validity under the null and local alternatives and provide concrete constructions for left-truncated right-censored data, current status data, and random double truncation; in particular, to the best of our knowledge, we give the first omnibus goodness-of-fit test for a parametric family under random double truncation in the composite-hypothesis case. Simulations and an empirical illustration demonstrate size control and power in practically relevant incomplete-data designs.
We study fuzzy regression discontinuity designs with covariates and characterize the weighted averages of conditional local average treatment effects (WLATEs) that are point identified. Any identified WLATE equals a Wald ratio of conditional reduced-form and first-stage discontinuities. We highlight the Compliance-Weighted LATE (CWLATE), which weights cells by squared first-stage discontinuities and maximizes first-stage strength. For discrete covariates, we provide simple estimators and robust bias-corrected inference. In simulations calibrated to common designs, CWLATE improves stability and reduces mean squared error relative to standard fuzzy RDD estimators when compliance varies. An application to Uruguayan cash transfers during pregnancy yields precise RDD-based effects on low birthweight.
We develop kernel-based specification tests for semiparametric conditional moment models with high-dimensional nuisance parameters, extending existing conditional moment tests---which typically require asymptotically linear nuisance estimators---to accommodate modern machine-learning methods. The proposed locally robust kernel tests combine Neyman-orthogonal moments, cross-fitting, and reproducing kernel Hilbert space methods, yielding inference that is first-order insensitive to nuisance estimation error. We establish oracle equivalence between the feasible and infeasible test processes under local alternatives and weak nuisance-rate conditions, and characterize the resulting local power. A fast multiplier bootstrap avoids nuisance re-estimation. Applications include specification testing in high-dimensional linear and logistic regression, significance testing with machine-learning regressions, and tests of constant conditional treatment effects. Monte Carlo simulations and an application to the National Supported Work program illustrate the finite-sample performance of the proposed tests.
This paper proposes a new class of nonparametric tests for the correct specification of generalized propensity score models. The test procedure is based on two different projection arguments, which lead to test statistics with several appealing properties. They accommodate high-dimensional covariates; are asymptotically invariant to the estimation method used to estimate the nuisance parameters and do not requite estimators to be root-n asymptotically linear; are fully data-driven and do not require tuning parameters, can be written in closed-form, facilitating the implementation of an easy-to-use multiplier bootstrap procedure. We show that our proposed tests are able to detect a broad class of local alternatives converging to the null at the parametric rate. Monte Carlo simulation studies indicate that our double projected tests have much higher power than other tests available in the literature, highlighting their practical appeal.
Developing robust inference for models with nonparametric Unobserved Heterogeneity (UH) is both important and challenging. We propose novel Debiased Machine Learning (DML) procedures for valid inference on functionals of UH, allowing for partial identification of multivariate target and high-dimensional nuisance parameters. Our main contribution is a full characterization of all relevant Neyman-orthogonal moments in models with nonparametric UH, where relevance means informativeness about the parameter of interest. Under additional support conditions, orthogonal moments are globally robust to the distribution of the UH. They may still involve other high-dimensional nuisance parameters, but their local robustness reduces regularization bias and enables valid DML inference. We apply these results to: (i) common parameters, average marginal effects, and variances of UH in panel data models with high-dimensional controls; (ii) moments of the common factor in the Kotlarski model with a factor loading; and (iii) smooth functionals of teacher value-added. Monte Carlo simulations show substantial efficiency gains from using efficient orthogonal moments relative to ad-hoc choices. We illustrate the practical value of our approach by showing that existing estimates of the average and variance effects of maternal smoking on child birth weight are robust.
This paper proposes a Gaussian process (GP) approach for testing conditional moment restrictions. Tests are based on squared Neyman orthogonal function-parametric processes integrated with respect to a GP distribution. This methodology leads to a general unified framework of kernel-based tests having the following properties: (i) bootstrap tests are easy to implement in the presence of nuisance parameters (they are simple quadratic forms, and there is no need to reestimate the nuisance parameters in each bootstrap replication); and (ii) the new tests are valid under general conditions, including higher-order conditional moments of unknown form, regularized estimators (e.g., Lasso) or parameters at the boundary of the parameter space. Novel applications include distance kernel tests for zero conditional treatment effects. The paper introduces Neyman orthogonal kernels, a new asymptotic theory and a detailed local power analysis. Monte Carlo experiments and a real data application illustrate the sensitivity of tests to the dimension of covariates and to the mean and covariance kernel of the GP.
Models with Conditional Moment Restrictions (CMRs) are popular in economics. These models involve finite and infinite dimensional parameters. The infinite dimensional components include conditional expectations, conditional choice probabilities, or policy functions, which might be flexibly estimated using Machine Learning tools. This paper presents a characterization of locally debiased moments for regular models defined by general semiparametric CMRs with possibly different conditioning variables. These moments are appealing as they are known to be less affected by first-step bias. Additionally, we study their existence and relevance. Such results apply to a broad class of smooth functionals of finite and infinite dimensional parameters that do not necessarily appear in the CMRs. As a leading application of our theory, we characterize debiased machine learning for settings of treatment effects with endogeneity, giving necessary and sufficient conditions. We present a large class of relevant debiased moments in this context. We then propose the Compliance Machine Learning Estimator (CML), based on a practically convenient orthogonal relevant moment. We show that the resulting estimand can be written as a convex combination of conditional local average treatment effects (LATE). Altogether, CML enjoys three appealing properties in the LATE framework: (1) local robustness to first-stage estimation, (2) an estimand that can be identified under a minimal relevance condition, and (3) a meaningful causal interpretation. Our numerical experimentation shows satisfactory relative performance of such an estimator. Finally, we revisit the Oregon Health Insurance Experiment, analyzed by Finkelstein et al. (2012). We find that the use of machine learning and CML suggest larger positive effects on health care utilization than previously determined.
We study identification and estimation in the Regression Discontinuity Design (RDD) with a multivalued treatment variable. We also allow for the inclusion of covariates. We show that without additional information, treatment effects are not identified. We give necessary and sufficient conditions that lead to identification of LATEs as well as of weighted averages of the conditional LATEs. We show that if the first stage discontinuities of the multiple treatments conditional on covariates are linearly independent, then it is possible to identify multivariate weighted averages of the treatment effects with convenient identifiable weights. If, moreover, treatment effects do not vary with some covariates or a flexible parametric structure can be assumed, it is possible to identify (in fact, over-identify) all the treatment effects. The over-identification can be used to test these assumptions. We propose a simple estimator, which can be programmed in packaged software as a Two-Stage Least Squares regression, and packaged standard errors and tests can also be used. Finally, we implement our approach to identify the effects of different types of insurance coverage on health care utilization, as in Card, Dobkin and Maestas (2008).
The Basel Committee and the Financial Stability Board require a consensus on the identification of characteristics that make a financial institution more prone than others to be severely hit by systemic shocks. This paper introduces a new tool to achieve this goal: a model for the Conditional Average Systemic Effects (CASE). The CASE quantifies the average effect of a system wide shock or market downturn on the profit and loss account of a bank, a firm or on the return of an asset. We propose a linear model for CASE with heterogeneous effects in observable characteristics. These models complement alternative measures of systemic risk and allow researchers to identify the determinants of the vulnerability of a given financial institution. We develop bootstrap inference that accounts for both estimation risk and model misspecification risk, and show the utility of our results in Monte Carlo simulations and an empirical application to 100 large U.S. financial firms.
One of the most important empirical findings in microeconometrics is the pervasiveness of heterogeneity in economic behaviour (cf. Heckman 2001). This paper shows that cumulative distribution functions and quantiles of the nonparametric unobserved heterogeneity have an infinite efficiency bound in many structural economic models of interest. The paper presents a relatively simple check of this fact. The usefulness of the theory is demonstrated with several relevant examples in economics, including, among others, the proportion of individuals with severe long term unemployment duration, the average marginal effect and the proportion of individuals with a positive marginal effect in a correlated random coefficient model with heterogenous first-stage effects, and the distribution and quantiles of random coefficients in linear, binary and the Mixed Logit models. Monte Carlo simulations illustrate the finite sample implications of our findings for the distribution and quantiles of the random coefficients in the Mixed Logit model.
Machine-learning (ML) methods now routinely generate regressors used in subsequent econometric analyses, for example, estimated propensity scores, control-function residuals, imputed covariates, learned proxies, or low-dimensional embeddings of high-dimensional data. As these ML-generated regressors become ubiquitous, the lack of general inference methods for models that use them has become a critical limitation. Standard plug-in and Double ML procedures ignore how generated regressors enter later stages, leading to large biases and invalid inference. We develop a three-step locally robust GMM framework for inference with ML generated regressors. A key new insight is downstream local robustness: by a functional chain rule, moment functions that are constructed to be orthogonal to the second step eliminate the complicated indirect (conditioning) effects from the ML-generated regressors. We show how to implement this automatically by estimating the associated Riesz representers through cross-fitted auxiliary regressions, allowing for generic non-Donsker ML in both early steps. In leading treatment-effect and counterfactual settings, simulations demonstrate severe bias in existing methods and reductions of 85-95
This paper proposes minimum distance inference for a structural parameter of interest, which is robust to the lack of identification of other structural nuisance parameters. Some choices of the weighting matrix lead to asymptotic chi-squared distributions with degrees of freedom that can be consistently estimated from the data, even under partial identification. In any case, knowledge of the level of under-identification is not required. We study the power of our robust test. Several examples show the wide applicability of the procedure and a Monte Carlo investigates its finite sample performance. Our identification-robust inference method can be applied to make inferences on both calibrated (fixed) parameters and any other structural parameter of interest. We illustrate the method's usefulness by applying it to a structural model on the non-neutrality of monetary policy, as in \cite{nakamura2018high}, where we empirically evaluate the validity of the calibrated parameters and we carry out robust inference on the slope of the Phillips curve and the information effect.
Locally Robust (LR)/Orthogonal/Debiased moments have proven useful with machine learning first steps, but their existence has not been investigated for general parameters. In this paper, we provide a necessary and sufficient condition, referred to as Restricted Local Non-surjectivity (RLN), for the existence of such orthogonal moments to conduct robust inference on general parameters of interest in regular semiparametric models. Importantly, RLN does not require either identification of the parameters of interest or the nuisance parameters. However, for orthogonal moments to be informative, the efficient Fisher Information matrix for the parameter must be non-zero (though possibly singular). Thus, orthogonal moments exist and are informative under more general conditions than previously recognized. We demonstrate the utility of our general results by characterizing orthogonal moments in a class of models with Unobserved Heterogeneity (UH). For this class of models our method delivers functional differencing as a special case. Orthogonality for general smooth functionals of the distribution of UH is also characterized. As a second major application, we investigate the existence of orthogonal moments and their relevance for models defined by moment restrictions with possibly different conditioning variables. We find orthogonal moments for the fully saturated two stage least squares, for heterogeneous parameters in treatment effects, for sample selection models, and for popular models of demand for differentiated products. We apply our results to the Oregon Health Experiment to study heterogeneous treatment effects of Medicaid on different health outcomes.
Many economic and causal parameters depend on nonparametric or high dimensional first steps. We give a general construction of locally robust/orthogonal moment functions for GMM, where first steps have no effect, locally, on average moment functions. Using these orthogonal moments reduces model selection and regularization bias, as is important in many applications, especially for machine learning first steps. Also, associated standard errors are robust to misspecification when there is the same number of moment functions as parameters of interest. We use these orthogonal moments and cross-fitting to construct debiased machine learning estimators of functions of high dimensional conditional quantiles and of dynamic discrete choice parameters with high dimensional state variables. We show that additional first steps needed for the orthogonal moment functions have no effect, globally, on average orthogonal moment functions. We give a general approach to estimating those additional first steps. We characterize double robustness and give a variety of new doubly robust moment functions. We give general and simple regularity conditions for asymptotic theory.
This article proposes a new identification strategy and a new estimation method for the hybrid New Keynesian Phillips curve (NKPC). Unlike the predominant Generalized Method of Moments (GMM) approach, which leads to weak identification of the NKPC with U.S. postwar data, our non-parametric method exploits nonlinear variation in inflation dynamics and provides supporting evidence of point-identification. This article shows that identification of the NKPC is characterized by two conditional moment restrictions. This insight leads to a quantitative method to assess identification in the NKPC. For estimation, the article proposes a closed-form Generalized Band Spectrum Estimator (GBSE) that effectively uses information from the conditional moments, accounts for nonlinear variation, and permits a focus on short-run dynamics. Applying the GBSE to U.S postwar data, we find a significant coefficient of marginal cost and that the forward-looking component and the inflation inertia are both equally quantitatively important in explaining the short-run inflation dynamics, substantially reducing sampling uncertainty relative to existing GMM estimates.
Equality of opportunity has emerged as an important ideal of distributive justice. Empirically, Inequality of Opportunity (IOp) is measured in two steps: first, an outcome (e.g., income) is predicted given individual circumstances; and second, an inequality index (e.g., Gini) of the predictions is computed. Machine Learning (ML) methods are tremendously useful in the first step. However, they can cause sizable biases in IOp since the bias-variance trade-off allows the bias to creep in the second step. We propose a simple debiased IOp estimator robust to such ML biases and provide the first valid inferential theory for IOp. We demonstrate improved performance in simulations and report the first unbiased measures of income IOp in Europe. Mother's education and father's occupation are the circumstances that explain the most. Plug-in estimators are very sensitive to the ML algorithm, while debiased IOp estimators are robust. These results are extended to a general U-statistics setting.
Many economic and causal parameters depend on nonparametric or high dimensional first steps. We give a general construction of locally robust/orthogonal moment functions for GMM, where moment conditions have zero derivative with respect to first steps. We show that orthogonal moment functions can be constructed by adding to identifying moments the nonparametric influence function for the effect of the first step on identifying moments. Orthogonal moments reduce model selection and regularization bias, as is very important in many applications, especially for machine learning first steps. We give debiased machine learning estimators of functionals of high dimensional conditional quantiles and of dynamic discrete choice parameters with high dimensional state variables. We show that adding to identifying moments the nonparametric influence function provides a general construction of orthogonal moments, including regularity conditions, and show that the nonparametric influence function is robust to additional unknown functions on which it depends. We give a general approach to estimating the unknown functions in the nonparametric influence function and use it to automatically debias estimators of functionals of high dimensional conditional location learners. We give a variety of new doubly robust moment equations and characterize double robustness. We give general and simple regularity conditions and apply these for asymptotic inference on functionals of high dimensional regression quantiles and dynamic discrete choice parameters with high dimensional state variables.
This paper provides a systematic approach to semiparametric identification that is based on statistical information as a measure of its "quality." Identification can be regular or irregular, depending on whether the Fisher information for the parameter is positive or zero, respectively. I first characterize these cases in models with densities linear in an infinite-dimensional parameter. I then introduce a novel "generalized Fisher information." If positive, it implies (possibly irregular) identification when other conditions hold. If zero, it implies impossibility results on rates of estimation. Three examples illustrate the applicability of the general results. First, I consider the canonical example of average densities. Second, I show irregular identification of the median willingness to pay in contingent valuation studies. Finally, I study identification of the discount factor and average measures of risk aversion in a nonparametric Euler equation with nonparametric measurement error in consumption.