
We nonparametrically identify Type I and Type II error rates when mistakes are unobserved. Our strategy reframes the setting as a misclassified binary choice model, using a special regressor to identify the misclassification probabilities. Identification requires an exclusion restriction and a large support condition. When the exclusion restriction is in doubt, we show how our estimands can be interpreted as bounds under a weaker monotonicity condition. Bounds are also provided when the large support condition is violated. We apply our method to estimate miscarriages of justice in Virginia. Among defendants who proceeded to trial, the estimated probability of wrongful acquittal ranges from 18 to 42%, depending on offense type, race, and gender. Our method also produces estimates of wrongful conviction probabilities; however, these are very high, potentially reflecting the restriction of our sample to trial defendants and/or violations of key assumptions.
Using age-specific measures as proxies for lifetime outcomes can introduce life-cycle bias into estimates of intergenerational associations. We extend the generalized errors-in-variables (GEiV) model to the estimation of intergenerational crime associations and implement a data-driven method to correct for this bias. The GEiV adjustment accounts for both the extensive and intensive margins and introduces an elasticity estimator that accommodates zeros. Our results show that intergenerational associations are underestimated by 30-50%, even when crime is measured during the teenage years, a period least affected by life-cycle bias. The GEiV-adjusted intergenerational elasticity between fathers and sons in New Zealand is approximately 0.7 for the likelihood of committing any crime and 0.5 for the number of crimes committed. Mother-child elasticities are smaller than father-child elasticities. The GEiV-adjusted estimates remain stable across ages, birth cohorts, number of aggregated years, and crime types and severity.
We propose a framework to disentangle the quantile treatment effect of a binary treatment at a specific rank into an indirect quantile treatment effect that operates through a mediator and an unmediated direct quantile treatment effect. We establish identification results for these effects under the sequential ignorability assumption and propose double/debiased machine learning estimators, based on the efficient influence functions of the cumulative distribution functions of potential outcomes. We demonstrate uniform consistency and asymptotic normality of our effect estimators under specific regularity conditions and propose a multiplier bootstrap for statistical inference. Finally, we apply our method to data from the National Job Corps Study to assess the direct effect of training on earnings and the indirect effect operating through work experience.
Given the cyclical nature of market volatility and the increasing complexity of global financial systems, developing effective strategies for large-scale portfolio optimization is of critical importance. In this work, we propose a novel Dantzig-type portfolio optimization (DPO) model designed to help investors navigate these challenges and optimize their portfolios effectively. The model separately incorporates & ell;1 and folded concave penalties, enabling the direct estimation of optimal portfolio weights while enforcing the sum constraint and accommodating both long and short positions. We establish the desired theoretical properties under mild regularity conditions, and introduce efficient parallel computing algorithms based on asset-splitting. Through extensive simulation studies, we investigate the superior effectiveness and efficiency of the DPO model and proposed algorithms. Furthermore, we illustrate the usefulness of the model by applying it to U.S. stock market datasets, including the constituent stocks of both S&P 500 and Russell 2000 indices.
Current forecast evaluation tests can only assess forecasts generated at the same frequency. However, in real-world scenarios, predictions for the same economic variable may be made at different frequencies. The existing literature lacks statistical tests for evaluating such forecasts. This paper introduces a new evaluation test for comparing these forecasts. We propose a two-sample t -type test designed to test the equality of means between two potentially correlated time series that are sampled at different frequencies. No current estimator can compute the variance needed to studentize the difference in the sample means, given the potential temporal and cross-correlations, alongside the mixed-frequency nature of the related data. To address this, we propose a block-average-based variance estimator for this purpose. We derive the asymptotic null distribution of our new two-sample t statistic and analyze the test's local power. Through extensive Monte Carlo simulations, our two-sample test demonstrates favorable size and power properties in finite samples. Notably, in cases with small sample sizes, our approach outperforms existing heuristic methods, which involve discarding data to align forecasts and applying conventional tests. Additionally, we uncover interesting connections with relevant methods in the literature.
We introduce shrinkage estimators of the sample mean vector in high dimension. Our estimators share desirable properties relative to existing methods: they are distribution-free, consider bona-fide target estimators, and use simple estimators of the shrinkage intensities independent of the precision matrix which are L2 consistent in high dimension. Unlike existing estimators that impose that the whitened data be an i.i.d. matrix, we only impose i.i.d. across sample observations. We require uniform boundedness of the first four moments, and the high-dimensional asymptotics are in a general Kolmogorov setting where N/T=O(1) as T ->infinity, with N the mean-vector dimension and T the sample size. We consider as a target estimator either zero or the grand mean. Simulations show that our shrinkage estimators are competitive with a range of benchmark estimators, both when the theoretical assumptions are satisfied and violated. Finally, we apply our estimators to constructing mean-variance portfolios of a large number of stocks, and find that they deliver robust out-of-sample Sharpe ratios.
We estimate finite-dimensional parameters in conditional moment restriction (CMR) models when at least one of the endogenous variables (outcomes and/or explanatory variables) in the model is missing for some individuals in the sample. We demonstrate that efficiency gains in estimation occur if and only if there is at least one endogenous variable-included in or excluded from the CMR model-that is nonmissing (observed for all individuals in the sample), which we show characterizes informative imputation. We propose a semiparametrically efficient estimator which is also "doubly robust." To illustrate the insights our estimator can provide in empirical applications with large sample sizes, we artificially induce missingness in the female labor supply model of Angrist and Evans. Despite medium levels of missingness in female labor income (the outcome) and a sample size exceeding 200,000 observations, the inverse propensity score weighted generalized method of moments (GMM) estimator finds only a statistically insignificant negative effect of having a third child (the endogenous regressor) on labor income. In contrast, our efficient estimator yields point estimates of this effect that are not only comparable to the GMM estimates but are also statistically significant.
We study the gradient wild bootstrap-based inference for instrumental variable quantile regressions in the framework of a small number of large clusters in which the number of clusters is viewed as fixed, and the number of observations for each cluster diverges to infinity. We illustrate the good finite-sample performance of the new inference methods using simulations and provide an empirical application to a well-known dataset about US local labor markets.
We study inference for threshold regression in the context of a large factor model with common stochastic trends. We develop a Least Squares estimator for the threshold level, deriving almost sure rates of convergence and proposing a novel way of constructing confidence intervals based on a randomized test. Our confidence intervals are constructed using the rates of convergence of the estimated threshold level, and no limiting distribution is required. We also develop a procedure to estimate the number of common trends in each regime, and investigate the properties of the Principal Component estimator for the loadings and common factors in both regimes. Our theoretical findings are corroborated through a comprehensive set of Monte Carlo experiments, and an application to equity prices and bond yields.