We develop a life-cycle model of labor supply and human capital formation that incorporates health shocks, health insurance, and medical treatment decisions. We use the model to study effects of health shocks on health, labor supply, earnings, and earnings inequality. We also simulate provision of public insurance to agents who lack employer-sponsored insurance. While this increases medical spending substantially, it creates positive labor supply incentives for low-skill workers while reducing costs of social insurance, Medicaid, and free care. The net program cost is modest, and all model agents are ex ante better off in a balanced budget simulation.
In the 1960 cohort, American men and women graduated from college at similar rates, and this was true for Whites, Blacks and Hispanics. But in more recent cohorts, women graduate at much higher rates than men. Gaps between race/ethnic groups have also widened. To understand these patterns, we develop a model of individual and family decision-making where education, labor supply, marriage and fertility are all endogenous. Assuming stable preferences, our model explains changes in education for the ‘60-‘80 cohorts based on three exogenous factors: family background, labor market and marriage market constraints. We find changes in parental background account for 1/4 of the growth in women’s college graduation from the ’60 to ’80 cohort. The marriage market accounts for 1/5 and the labor market explains the rest. Thus, parent education plays an important role in generating social mobility, enabling us to predict future evolution of college graduation rates due to this factor. We predict White women’s graduation rate will plateau, while that of Hispanic and Black women will grow rapidly. But the aggregate graduation rate will grow very slowly due to the increasing Hispanic share of the population. Institutional subscribers to the NBER working paper series, and residents of developing countries may download this paper without additional charge at www.nber.org.
We survey the weak instrumental variables (IV) literature with the aim of giving simple advice to applied researchers. This literature focuses heavily on the problem of size inflation in two-stage least squares (2SLS) two-tailed t -tests that arises if instruments are weak. A common standard for acceptable instrument strength is a first-stage F of 10, which renders this size inflation modest. However, 2SLS suffers from other important problems that exist at much higher levels of instrument strength. In particular, 2SLS standard errors tend to be artificially small in samples where the 2SLS estimate is close to ordinary least squares (OLS). This power asymmetry means the t -test has inflated power to detect false positive effects when the OLS bias is positive. The Anderson-Rubin (AR) test avoids this problem and should be used in lieu of the t -test even with strong instruments. We illustrate the practical importance of this issue in IV papers published in the American Economic Review from 2011 to 2023. Use of the AR test often reverses t -test results. In particular, IV estimates that are close to OLS and significant according to the t -test are often insignificant according to AR. We also show that for first-stage F in the 10–20 range there is a high probability that OLS estimates will be closer to the truth than 2SLS. Hence we advocate a higher standard of instrument strength in applied work.
We survey the weak instrumental variables (IV) literature with the aim ofgiving simple advice to applied researchers. This literature focuses heavilyon the problem of size inflation in two-stage least squares (2SLS) two-tailedt-tests that arises if instruments are weak. A common standard for acceptableinstrument strength is a first-stageFof 10, which renders this size inflationmodest. However, 2SLS suffers from other important problems that existat much higher levels of instrument strength. In particular, 2SLS standarderrors tend to be artificially small in samples where the 2SLS estimate is closeto ordinary least squares (OLS). This power asymmetry means thet-test hasinflated power to detect false positive effects when the OLS bias is positive.The Anderson-Rubin (AR) test avoids this problem and should be used inlieu of thet-test even with strong instruments. We illustrate the practicalimportance of this issue in IV papers published in theAmerican EconomicReviewfrom 2011 to 2023. Use of the AR test often reversest-test results. Inparticular, IV estimates that are close to OLS and significant according tothet-test are often insignificant according to AR. We also show that for first-stageFin the 10-20 range there is a high probability that OLS estimates willbe closer to the truth than 2SLS. Hence we advocate a higher standard ofinstrument strength in applied work.
In the 1960 cohort, American men and women graduated from college at the same rate, and this was true for Whites, Blacks and Hispanics. But in more recent cohorts, women graduate at much higher rates than men. To understand the emerging gender education gap, we formulate and estimate a model of individual and family decision-making where education, labor supply, marriage and fertility are all endogenous. Assuming preferences that are common across ethnic groups and fixed over cohorts, our model explains differences in all endogenous variables by gender/ethnicity for the '60-'80 cohorts based on three exogenous factors: family background, labor market and marriage market constraints. Changes in parental background are a key factor driving the growing gender education gap: Women with college educated mothers get greater utility from college, and are much more likely to graduate themselves. The marriage market also contributes: Women's chance of getting marriage offers at older ages has increased, enabling them to defer marriage. The labor market is the largest factor: Improvement in women's labor market return to college in recent cohorts accounts for 50% of the increase in their graduation rate. But the labor market returns to college are still greater for men. Women go to college more because their overall return is greater, after factoring in marriage market returns and their greater utility from college attendance. We predict the recent large increases in women's graduation rates will cause their children's graduation rates to increase further. But growth in the aggregate graduation rate will slow substantially, due to significant increases in the share of Hispanics a group with a low graduation rate in recent birth cohorts.
Recent controversy has surrounded the sanctioning, by regulatory authorities, of doctors for publicly expressing views on elements of the COVID-19 pandemic. This controversy fundamentally involves the limits of intellectual freedom doctors have within the constraints of Codes of Conduct. In this context, a recent unanimous High Court of Australia judgement gives an important window into how the Court considers intellectual freedom and attempts to curtail such freedom under the guise of conduct.
Two stage least squares (2SLS) has poor properties if instruments are exogenous but weak. But how strong do instruments need to be for 2SLS estimates and test statistics to exhibit acceptable properties? A common standard is that first-stage F≥10. This is adequate to ensure two-tailed t-tests have modest size distortions. But other problems persist: In particular, we show 2SLS standard errors are artificially small in samples where the estimate is most contaminated by the OLS bias. Hence, if the bias is positive, the t-test has little power to detect true negative effects, and inflated power to find positive effects. This phenomenon, which we call a “power asymmetry,” persists even if first-stage F is in the thousands. Robust tests like Anderson–Rubin perform better, and should be used in lieu of the t-test even with strong instruments. We also show how 2SLS test statistics typically suffer from very low power if first-stage F is only 10, leading us to suggest a higher standard of instrument strength in empirical practice.
We study the impact of child work on cognitive development in four Low‐ and Middle‐Income Countries. We advance the literature by using cognitive test scores collected regardless of school attendance. We also address a key gap in the literature by controlling for children's complete time allocation budget. This allows us to estimate effects of different types of work, like chores and market/farm work, relative to specific alternative time‐uses, like school or study or play/leisure. Our results show child work is more detrimental to child development to the extent that it crowds out school/study time rather than leisure. We also show the adverse effect of time spent on domestic chores is similar to time spent on market and farm work, provided they both crowd out school/study time. Thus, policies to enhance child development should target a shift from all forms of work toward educational activities.
We provide a simple survey of the weak IV literature, aimed at giving practical advice to applied researchers. It is well-known that 2SLS has poor properties if instruments are exogenous but weak. We clarify these properties, explain weak instrument tests, and examine how behavior of 2SLS estimates and test statistics depend on instrument strength. A common standard for acceptably strong instruments is a first-stage F of 10, which renders two-tailed t-test size distortion modest. However, we show that 2SLS standard errors tend to be artificially small in samples where the estimate is most contaminated by the OLS bias. This means the t-test has inflated power to detect false positive effects when the OLS bias is positive. Surprisingly, this problem persists even if the first-stage F is in the thousands. Robust tests like Anderson-Rubin perform better, and should be used in lieu of the t-test even with strong instruments. In many realistic settings a first-stage F well above 10 may be necessary to give high confidence that 2SLS will outperform OLS. For example, in the archetypal application of estimating returns to education, we argue one needs F of at least 50.
There is a long standing controversy over the magnitude of the Frisch labor supply elasticity. Macro economists using DSGE models often calibrate it to be large, while numerous micro data studies estimate it is near zero. A large literature has emerged that attempts to reconcile the micro and macro results. We o i¬€ er a new and simple explanation: Most micro studies estimate the Frisch using a 2SLS regression of hours changes on income changes. But the available instruments are typically weak. In that case, it is an inherent property of 2SLS that estimates of the Frisch will (spuriously) appear more precise when they are more shifted in the direction of the OLS bias, which is negative. As a result, Frisch elasticities near zero will (spuriously) appear to be precisely estimated, while large estimates will appear to be very imprecise. This will naturally bias micro data studies toward concluding the Frisch is small. We show how the use of a weak instrument robust hypothesis test, the Anderson-Rubin test, leads us to conclude the Frisch elasticity is large and signii¬ cant in the NLSY97 data. In contrast, a conventional 2SLS t-test would lead us to conclude it is not signii¬ cantly greater then zero.
that was held at University College 29-30 August 2017.The aim of the conference was to bring together researchers from applied microeconomics, computational economics and econometrics who work on structural dynamic models.It provided a forum at which new structural models with applications and new numerical/econometric methods for their implementation were presented and discussed.To disseminate the ideas and results presented at the conference, participants were invited to submit papers to this special issue on the same topic.The papers that appear in this special issue can roughly be divided into two groups: The first contains methodological contributions where either new structural models or new methods for analyzing them are developed.The second group contains novel applications of existing models and/or methods with emphasis on their practical implementation and usages in answering policy-relevant questions.Taken as a whole, the special issue provides the reader with a good overview of state-of-the-art methods and their implementations for structural dynamic models in economics.In ''Sufficient Statistics for Unobserved Heterogeneity in Structural Dynamic Logit Models'', Victor Aguirregabiria, Yao Luo and Jiaying Gu analyze which parameters can be identified in a class of structural dynamic discrete choice (DDC) models with fixed effects.The authors take the so-called fixed-effect conditional likelihood (FE-CML) approach where a set of sufficient statistics for the incidental parameters is derived and estimators are obtained by maximization of the likelihood conditional on these sufficient statistics.For non-structural (i.e., myopic) dynamic logit models with unobserved heterogeneity only in the intercept, Chamberlain (1985) and Honoré and Kyriazidou (2000) have shown that the FE-CML approach can identify the parameters of interest.Augirregabiria et al. show that this approach extends to a class of structural DDC models with logit errors and so is the first paper providing an econometric theory for fixed-effects structural dynamic models with all existing papers focusing on random-effects models.Nicholas Buchholz, Matt Shum and Haiqing Xu also provide an econometric analysis of a class of structural DDC models in ''Semiparametric Estimation of Dynamic Discrete Choice Models''.Their focus is on relaxing the fully parametric restrictions that are normally imposed on the unobserved components entering the models.Buchholz et al. show identification of a class of semiparametric structural DDC models in which the utility indices are parametrically specified but the shock distribution is left unspecified and treated as a nonparametric object.Their identification result is constructive in the sense that an estimator can be developed from it.The authors provide an asymptotic analysis of this estimator.Their novel results open up the door for more flexible DDC models in applied work.Related papers include Blevins (2014), Norets and Tang (2014) and Chen ( 2017), but in contrast to these Buchholz et al. not rely on exclusion restrictions, rather they exploit the optimality conditions to achieve identification.Most structural DDC models cannot be solved in closed form and so involve numerical solution methods.The standard method is to discretize the state space of the model and solve the resulting approximate model, but this comes with a built-in curse-of-dimensionality.To circumvent these issues, Dennis Kristensen, Jong Myun Moon, Patrick Mogensen and Bertel Schjerning propose a novel solution method that combines sieve methods with simulations in ''Solving Dynamic Discrete Choice Models Using Smoothing and Sieve Methods''.In contrast to existing methods (see, e.g., Rust, 1997;Pal and Stachurski, 2013), theirs takes into account the particular features of DDC models used in economics, including the presence of unobserved i.i.d.shocks.They utilize these features to obtain a numerical efficient solution method that works for a broad class of DDC models.Mixed hitting-time (MHT) models make up a class of duration models that can be applied to the analysis of optimal stopping decisions by heterogeneous agents; see Abbring (2012).In ''The likelihood of mixed hitting times'', Jaap Abbring
Most work in optimal tax theory relies on simple labor supply models that fail to incorporate insights from the modern labor supply literature. As a result, it may have reached misleading conclusions regarding the optimal tax structure. The recent work on labor supply that I review here emphasizes human capital investment and the participation margin. When the data is viewed through the lens of models that account for these features, it implies labor supply is more elastic than conventional wisdom suggests. Recent work also stresses how elasticities vary by age, education, gender and marital status. Here I explore the implications of these recent developments in the labor supply literature for the optimal design of the tax system. I also review some recent work in the optimal tax literature that does utilize more sophisticated labor supply models, and discuss how incorporating those features influences optimal tax calculations.
We structurally estimate a life-cycle model of consumption, labor supply and retirement, using data from the Australian HILDA panel. We use the model to evaluate effects of Australia’s Age Pension system and income tax policy on labor supply, consumption and retirement. Our model accounts for human capital, savings, uninsurable wage risk and credit constraints. We account for “bunching” of hours by assuming a discrete set of hours levels, and we investigate labor supply on both the intensive and extensive margins. Our model allows us to quantify the effects of anticipated and unanticipated tax and pension policy changes at different points of the life-cycle. Our results imply that Australia’s Age Pension system as currently designed is poorly targeted. Our simulations suggest that a doubling of taper rates, combined with a 5.9% reduction of income tax rates, would be budget neutral and Pareto improving.
SummaryPredicting the impact of climate change on crop yield is difficult, in part because the production function mapping weather to yield is high dimensional and nonlinear. We compare three approaches to predicting yields: (a) deep neural networks (DNNs), (b) traditional panel-data models, and (c) a new panel-data model that allows for unit and time fixed effects in both intercepts and slopes in the agricultural production function—made feasible by a new estimator called Mean Observation OLS (MO-OLS). Using U.S. county-level corn-yield data from 1950 to 2015, we show that both DNNs and MO-OLS models outperform traditional panel-data models for predicting yield, both in-sample and in a Monte Carlo cross-validation exercise. However, the MO-OLS model substantially outperforms both DNNs and traditional panel-data models in forecasting yield in a 2006–2015 holdout sample. We compare the predictions of all these models for climate change impacts on yields from 2016 to 2100.
We analyze how two key factors contribute to the high cost of healthcare in the US relative to the UK: (i) the higher private cost of medical education, and (ii) the higher risk of malpractice litigation. To assess the role of these factors we formulate, calibrate and simulate an equilibrium model of physician wages and the supply of medical graduates, work hours and treatment decisions of practicing physicians, and malpractice risk and malpractice insurance pricing. Consistent with prior work, we find direct costs of malpractice fines and insurance explain little of the high cost of healthcare in the US. However, the high private cost of medical education interacts with high malpractice risk in an interesting way: It leads doctors to (i) demand high wages and (ii) use excessive diagnostics to mitigate risk (“defensive medicine”). The agency problem that arises because patients cannot judge the efficacy of tests allows them to be over-prescribed. Together, these factors increase costs far more than direct malpractice costs. Specifically, physician salaries plus diagnostic tests comprise 4.04% of GDP in the US, compared to only 2.3% in the UK. The mechanisms emphasized in our model can largely explain the difference. Our policy simulations imply that more generous medical education subsidies would lead to both improved patient welfare and reduced overall health care costs in the US system (a Pareto improvement). We also find policies to (i) reduce malpractice risk, or (ii) induce doctors to internalize a small part of diagnostic costs, would have similar efficacious effects.
We study potential impacts of future climate change on U.S. agricultural productivity using county‐level yield and weather data from 1950 to 2015. To account for adaptation of production to different weather conditions, it is crucial to allow for both spatial and temporal variation in the production process mapping weather to crop yields. We present a new panel data estimation technique, called mean observation OLS (MO‐OLS) that allows for spatial and temporal heterogeneity in all regression parameters (intercepts and slopes). Both forms of heterogeneity are important: We find strong evidence that production function parameters adapt to local climate, and also that sensitivity of yield to high temperature declined from 1950–89. We use our estimates to project corn yields to 2100 using 19 climate models and three greenhouse gas emission scenarios. We predict unmitigated climate change will greatly reduce yield. Our mean prediction (over climate models) is that adaptation alone can mitigate 36% of the damage, while emissions reductions consistent with the Paris targets would mitigate 76%.
We develop an econometric model of consumer panic (or panic buying) during the COVID-19 pandemic. Using Google search data on relevant keywords, we construct a daily index of consumer panic for 54 countries from January 1st to April 30th 2020. We also assemble data on government policy announcements and daily COVID-19 cases for all countries. Our panic index reveals widespread consumer panic in most countries, primarily during March, but with significant variation in the timing and severity of panic between countries. Our model implies that both domestic and world virus transmission contribute significantly to consumer panic. But government policy is also important: Internal movement restrictions – whether announced by domestic or foreign governments – generate substantial short run panic that largely vanishes in a week to ten days. Internal movement restrictions announced early in the pandemic generated more panic than those announced later. Stimulus announcements had smaller impacts, and travel restrictions do not appear to generate consumer panic.