This paper provides identification results to characterize a fairness-accuracy (FA) frontier, and statistical inference tools to test hypotheses and build a confidence set for the FA-frontier, when outcomes are observed only for selected individuals. When the selection process is unrestricted but loss is measured in specific ways, we provide a characterization of the sharp identification region of the FA-frontier. Under an assumption of unconfoundedness conditional on observables (and unrestricted loss functions), we obtain point identification and propose a debiased machine learning estimator, derive its asymptotic distribution, and show how this can be used to carry out inference for the FA-frontier. In work in progress, we extend the partial identification results to a broader class of loss functions.
We test the null hypothesis that two parameters have the same sign, assuming that (asymptotically) normal estimators are available. Examples of this problem include the analysis of heterogeneous treatment effects, causal interpretation of reduced-form estimands, meta-studies, and mediation analysis. A number of tests were recently proposed. We recommend a test that is simple and rejects more often than many of these recent proposals. Like all other tests in the literature, it is conservative if the truth is near and therefore also biased. To clarify whether these features are avoidable, we also provide a test that is unbiased and has exact size control on the boundary of the null hypothesis, but which has counterintuitive properties and hence we do not recommend. We show how to improve -values in an existing paper from information contained in that paper's main text, and we revisit an empirical analysis of the effect of trade on voter behavior.
This paper proposes an information-based inference method for partially identified parameters in incomplete models that is valid both when the model is correctly specified and when it is misspecified.Key features of the method are: (i) it is based on minimizing a suitably defined Kullback-Leibler information criterion that accounts for incompleteness of the model and delivers a non-empty pseudotrue set; (ii) it is computationally tractable; (iii) its implementation is the same for both correctly and incorrectly specified models; (iv) it exploits all information provided by variation in discrete and continuous covariates; (v) it relies on Rao's score statistic, which is shown to be asymptotically pivotal.
We provide sufficient conditions for semi-nonparametric point identification of a mixture model of decision making under risk, when agents make choices in multiple lines of insurance coverage (contexts) by purchasing a bundle. As a first departure from the related literature, the model allows for two preference types. In the first one, agents behave according to standard expected utility theory with CARA Bernoulli utility function, with an agent-specific coefficient of absolute risk aversion whose distribution is left completely unspecified. In the other, agents behave according to the dual theory of choice under risk(Yaari, 1987) combined with a one-parameter family distortion function, where the parameter is agent-specific and is drawn from a distribution that is left completely unspecified. Within each preference type, the model allows for unobserved heterogeneity in consideration sets, where the latter form at the bundle level -- a second departure from the related literature. Our point identification result rests on observing sufficient variation in covariates across contexts, without requiring any independent variation across alternatives within a single context. We estimate the model on data on households' deductible choices in two lines of property insurance, and use the results to assess the welfare implications of a hypothetical market intervention where the two lines of insurance are combined into a single one. We study the role of limited consideration in mediating the welfare effects of such intervention.
We study rounding of numerical expectations in the Health and Retirement Study (HRS) between 2002 and 2014. We document that respondent-specific rounding patterns across questions in individual waves are quite stable across waves. We discover a tendency by about half of the respondents to provide more refined responses in the tails of the 0-100 scale than the center. In contrast, only about five percent of the respondents give more refined responses in the center than the tails. We find that respondents tend to report the values 25 and 75 more frequently than other values ending in 5. We also find that rounding practices vary somewhat across question domains and respondent characteristics. We propose an inferential approach that assumes stability of response tendencies across questions and waves to infer person-specific rounding in each question domain and scale segment and that replaces each point-response with an interval representing the range of possible values of the true latent belief. Using expectations from the 2016 wave of the HRS, we validate our approach. To demonstrate the consequences of rounding on inference, we compare best-predictor estimates from face-value expectations with those implied by our intervals.
The literature on stochastic programming typically restricts attention to problems that fulfill constraint qualifications. The literature on estimation and inference under partial identification frequently restricts the geometry of identified sets with diverse high-level assumptions. These superficially appear to be different approaches to closely related problems. We extensively analyze their relation. Among other things, we show that for partial identification through pure moment inequalities, numerous assumptions from the literature essentially coincide with the Mangasarian-Fromowitz constraint qualification. This clarifies the relation between well-known contributions, including within econometrics, and elucidates stringency, as well as ease of verification, of some high-level assumptions in seminal papers.
We propose a robust method of discrete choice analysis when agents' choice sets are unobserved. Our core model assumes nothing about agents' choice sets apart from their minimum size. Importantly, it leaves unrestricted the dependence, conditional on observables, between choice sets and preferences. We first characterize the sharp identification region of the model's parameters by a finite set of conditional moment inequalities. We then apply our theoretical findings to learn about households' risk preferences and choice sets from data on their deductible choices in auto collision insurance. We find that the data can be explained by expected utility theory with low levels of risk aversion and heterogeneous non‐singleton choice sets, and that more than three in four households require limited choice sets to explain their deductible choices. We also provide simulation evidence on the computational tractability of our method in applications with larger feasible sets or higher‐dimensional unobserved heterogeneity.
This paper proposes a method to conduct local linear regression smoothing in the presence of set-valued outcome data. The proposed estimator is shown to be consistent, and its mean squared error and asymptotic distribution are derived. A method to build error tubes around the estimator is provided, and a small Monte Carlo exercise is conducted to confirm the good finite sample properties of the estimator. The usefulness of the method is illustrated on a novel dataset from a clinical trial to assess the effect of certain genes' expressions on different lung cancer treatments outcomes.
We propose a robust method of discrete choice analysis when agents' choice sets are unobserved. Our core model assumes nothing about agents' choice sets apart from their minimum size. Importantly, it leaves unrestricted the dependence, conditional on observables, between choice sets and preferences. We first characterize the sharp identification region of the model's parameters by a finite set of conditional moment inequalities. We then apply our theoretical findings to learn about households' risk preferences and choice sets from data on their deductible choices in auto collision insurance. We find that the data can be explained by expected utility theory with low levels of risk aversion and heterogeneous non-singleton choice sets, and that more than three in four households require limited choice sets to explain their deductible choices. We also provide simulation evidence on the computational tractability of our method in applications with larger feasible sets or higher-dimensional unobserved heterogeneity.
IT IS POSSIBLE TO REPRESENT UNOBSERVED HETEROGENEITY in choice sets through additively separable disturbances. In a classic random utility model with Ui(c)=Wi(c)+ ic , one may let ic ∈ {−∞ 0} for each alternative c ∈ D and allow ic to be correlated with ic′ for any two alternatives c c′ ∈ D. One would then posit that: if κ = |D|, then ic = 0 for each alternative c ∈D; if κ = |D| − 1, then ic = −∞ for at most one alternative in D (the identity of which is left unspecified); if κ = |D| − 2, then ic = −∞ for at most two alternatives in D (the identities of which are left unspecified); and so forth. This model yields that alternative c is not chosen if ic = −∞, which is analogous to alternative c not being chosen when it is not contained in the agent’s choice set.
This paper is concerned with learning decision-makers' preferences using data on observed choices from a finite set of risky alternatives. We propose a discrete choice model with unobserved heterogeneity in consideration sets and in standard risk aversion. We obtain sufficient conditions for the model's semi-nonparametric point identification, including in cases where consideration depends on preferences and on some of the exogenous variables. Our method yields an estimator that is easy to compute and is applicable in markets with large choice sets. We illustrate its properties using a dataset on property insurance purchases.
As a consequence of missing data on tests for infection and imperfect accuracy of tests, reported rates of population infection by the SARS CoV-2 virus are lower than actual rates of infection. Hence, reported rates of severe illness conditional on infection are higher than actual rates. Understanding the time path of the COVID-19 pandemic has been hampered by the absence of bounds on infection rates that are credible and informative. This paper explains the logical problem of bounding these rates and reports illustrative findings, using data from Illinois, New York, and Italy. We combine the data with assumptions on the infection rate in the untested population and on the accuracy of the tests that appear credible in the current context. We find that the infection rate might be substantially higher than reported. We also find that the infection fatality rate in Italy is substantially lower than reported.
We investigate whether an insured's claims experience contains valuable information about its latent risk type. Using data on households' claims histories in auto and home insurance, we estimate the variance-covariance matrix of unobserved heterogeneity and utilize the estimates to update a priori predictions about the households' claim risk. The estimates reveal that unobserved heterogeneity is positively correlated across coverages. We then explore how households' demand for insurance would respond to experience rating under different theories of risky choice, and we discuss what our findings imply about the economic consequences of legal restrictions on experience rating.
This chapter reviews the microeconometrics literature on partial identification, focusing on the developments of the last thirty years. The topics presented illustrate that the available data combined with credible maintained assumptions may yield much information about a parameter of interest, even if they do not reveal it exactly. Special attention is devoted to discussing the challenges associated with, and some of the solutions put forward to, (1) obtain a tractable characterization of the values for the parameters of interest which are observationally equivalent, given the available data and maintained assumptions; (2) estimate this set of values; (3) conduct test of hypotheses and make confidence statements. The chapter reviews advances in partial identification analysis both as applied to learning (functionals of) probability distributions that are well-defined in the absence of models, as well as to learning parameters that are well-defined only in the context of particular models. A simple organizing principle is highlighted: the source of the identification problem can often be traced to a collection of random variables that are consistent with the available data and maintained assumptions. This collection may be part of the observed data or be a model implication. In either case, it can be formalized as a random set. Random set theory is then used as a mathematical framework to unify a number of special results and produce a general methodology to carry out partial identification analysis.
We develop a search-and-matching model where the magnitude of unemployment insurance benefits affects the likelihood that unemployed actually engage in active job search. To quan- titively discipline this relation we use administrative data of unemployed search audits. We use the model to quantify the effects of unemployment reforms. For small benefits' increases, the policymaker faces a trade-off between an uptick in the measure of unemployed actually searching and a fall in the unemployment exit-rate conditional on searching. For larger bene- fits' increases, an active search margin magnifies the benefits' disincentives, leading to a bigger drop in the employment rate than previously thought.
This paper provides inference methods for best linear approximations to functions which are known to lie within a band.It extends the partial identification literature by allowing the upper and lower functions defining the band to carry an index, and to be unknown but parametrically or non-parametrically estimable functions.The identification region of the parameters of the best linear approximation is characterized via its support function, and limit theory is developed for the latter.We prove that the support function can be approximated by a Gaussian process and establish validity of the Bayesian bootstrap for inference.Because the bounds may carry an index, the approach covers many canonical examples in the partial identification literature arising in the presence of interval valued outcome and/or regressor data: not only mean regression, but also quantile and distribution regression, including sample selection problems, as well as mean, quantile, and distribution treatment effects.In addition, the framework can account for the availability of instruments.An application is carried out, studying female labor force participation using data from Mulligan and Rubinstein ( 2008) and insights from Blundell, Gosling, Ichimura, and Meghir (2007).Our results yield robust evidence of a gender wage gap, both in the 1970s and 1990s, at quantiles of the wage distribution up to the 0.4, while allowing for completely unrestricted selection into the labor force.Under the assumption that the median wage offer of the employed is larger than that of individuals that do not work, the evidence of a gender wage gap extends to quantiles up to the 0.7.When the assumption is further strengthened to require stochastic dominance, the evidence of a gender wage gap extends to all quantiles, and there is some evidence at the 0.8 and higher quantiles that the gender wage gap decreased between the 1970s and 1990s.
We propose a bootstrap-based calibrated projection procedure to build confidence intervals for single components and for smooth functions of a partially identified parameter vector in moment (in)equality models. The method controls asymptotic coverage uniformly over a large class of data generating processes. The extreme points of the calibrated projection confidence interval are obtained by extremizing the value of the function of interest subject to a proper relaxation of studentized sample analogs of the moment (in)equality conditions. The degree of relaxation, or critical level, is calibrated so that the function of theta, not theta itself, is uniformly asymptotically covered with prespecified probability. This calibration is based on repeatedly checking feasibility of linear programming problems, rendering it computationally attractive. Nonetheless, the program defining an extreme point of the confidence interval is generally nonlinear and potentially intricate. We provide an algorithm, based on the response surface method for global optimization, that approximates the solution rapidly and accurately, and we establish its rate of convergence. The algorithm is of independent interest for optimization problems with simple objectives and complicated constraints. An empirical application estimating an entry game illustrates the usefulness of the method. Monte Carlo simulations confirm the accuracy of the solution algorithm, the good statistical as well as computational performance of calibrated projection (including in comparison to other methods), and the algorithm's potential to greatly accelerate computation of other confidence intervals.
Multiplicity of equilibria implies that the relationship between the outcome variable and the exogenous variables characterising a model is a correspondence rather than a function. This results in an incomplete econometric model. Incompleteness complicates identification and statistical inference on functionals of the probability distribution of the population of interest. This is because it implies that the sampling process and the maintained assumptions may be consistent with a set of values for these functionals, rather than with a single one. As a result, the econometric analysis of models with multiple equilibria needs to either: (1) rely on simplifying assumptions that shift focus to outcome features that are common across equilibria; or (2) augment the model with a "selection mechanism" that chooses the equilibrium played in the regions of multiplicity; or (3) maintain only minimal assumptions that partially identify the functionals of interest. Each of these approaches is reviewed, focusing on static game theoretic models.
A summary is not available for this content so a preview has been provided. Please use the Get access link above for information on how to access this content.
Ilya Molchanov合作论文数Department of Mathematical Statistics and Actuarial Science, University of Bern8