Over the past two decades, a steady decline in response rates on national face-to-face surveys has been documented, with steeper declines observed in recent years. The impact of nonresponse on survey estimates is inconsistent and depends on the correlation between response propensity and the survey estimates. To better understand the impact of declining response rates on the 2017-2018 National Health and Nutrition Examination Survey (NHANES), potential nonresponse bias (NRB) was investigated. NRB was assessed using three approaches: (a) studying variation within the respondent set; (b) benchmarking and comparisons to external data; and (c) comparing alternative weighting adjustments. Because NHANES only samples 30 counties in every 2-year cycle, the sample of counties in any given cycle may be an outlier on some characteristics. Such sampling variability may compound the effects of NRB. For this reason, the representativeness of the 2017-2018 NHANES counties was examined by comparing: (a) the characteristics of the 2017-2018 sampled counties with those from prior cycles; (b) each sampled county with the average of all the counties in the sampling stratum from which that county was selected; and (c) the 2017-2018 counties with 5,000 other samples that could have been drawn under the same sample design using a simulation study. The NRB analyses showed that the 2017-2018 NHANES sample had a lower proportion of college graduates and higher-income individuals compared with prior cycles. Additionally, the 2017-2018 NHANES counties had lower proportions of college graduates and lower mean incomes compared with counties from prior cycles and counties not selected in 2017-2018, which exacerbated the effects of NRB. Weighting adjustments used in prior cycles were not sufficient to address the bias in the 2017-2018 NHANES. Instead, enhanced weighting adjustments for education and income reduced the bias resulting from nonresponse and location sampling variability.
Small area estimation (SAE) uses explicit or implicit statistical models to estimate characteristics for geographic areas or other domains where the available survey data are insufficient to produce acceptably reliable direct estimates. In practice, most SAE procedures are now based on small area models that explicitly account for the error in predicting the target characteristic given the auxiliary data. These SAE procedures can be classified as to whether the model is expressed at the area level using direct survey estimates for the areas being modeled or at the unit level, typically at the level of the individual survey response. In a 2016 paper, Hidiroglou and You examined the performance of some unitand area-level SAE procedures for simulated samples from a known population, finding that unit-level estimators had a distinct advantage over arealevel ones. This paper expands their simulation to a broader set of circumstances and estimators in order to assess the generality of their findings.
In a series of papers, Little and Vartivarian (2003, 2005) argued that basing survey nonresponse adjustments on propensity to respond could increase the sampling variance of the estimates while not reducing bias, if the predictors of response were unrelated to key survey outcomes. Applying this idea to a 2014 military workforce survey, RAND researchers used machine learning approaches to develop a two-step method for nonresponse adjustment. The two step method comprises (1) a model for the key outcome variables based on respondents and (2) a response propensity model using the predicted key outcome variables as predictors for all sampled units. At the 2016 JSM, we presented simulation results assessing the predictive performance of competing machine learning algorithms in the first of the two steps. In this paper, we investigate the circumstances necessary for the two-step method to outperform nonresponse approaches in common practice, most of which can be regarded as single-step methods.
The National Crime Victimization Survey (NCVS) in the U.S. has provided estimates of violent and property crime for over four decades. Until recently, the survey has had almost exclusively a national focus. As recommended by a National Academy of Sciences panel, the Bureau of Justice Statistics (BJS) has undertaken a number of efforts to expand the geographic utility of the survey results. This paper is an outgrowth of research to provide small area estimates of key crime rates from the NCVS for states, large counties, and large metropolitan areas. The estimates are based a modified version of a time-series model proposed by Rao and Yu to take advantage of strong area-level correlations in the estimates over time. A multivariate version of the model was used to provide estimates for components of the crime rates by type of crime and by relationship to the perpetrator. BJS plans to make the estimates available to users on their website. This paper describes a hybrid model representing a further extension. These methods have potential applications to other situations in which the underlying characteristic exhibits strong stability over time.
Rao and Scott (Type) Tests† Robert E. Fay, Robert E. FaySearch for more papers by this author Robert E. Fay, Robert E. FaySearch for more papers by this author First published: 29 September 2014 https://doi.org/10.1002/9781118445112.stat00394 †This article was originally published online in 2006 in Encyclopedia of Statistical Sciences, © John Wiley & Sons, Inc. and republished in Wiley StatsRef: Statistics Reference Online, 2014. Read the full textAboutPDF ToolsRequest permissionExport citationAdd to favoritesTrack citation ShareShare Give accessShare full text accessShare full-text accessPlease review our Terms and Conditions of Use and check box below to share full-text version of article.I have read and accept the Wiley Online Library Terms and Conditions of UseShareable LinkUse the link below to share a full-text version of this article with your friends and colleagues. Learn more.Copy URL Share a linkShare onFacebookTwitterLinked InRedditWechat No abstract is available for this article. Wiley StatsRef: Statistics Reference OnlineBrowse other articles of this reference work:BROWSE BY TOPICBROWSE A-Z RelatedInformation
Within the context of probability-based sampling from a finite population, a number of schemes have been studied to maximize or minimize the overlap between two sample selections while maintaining the required probabilities of selection for each. For example, in redesigning an in-person survey, it may be desirable to overlap the sampling of primary sampling units between the designs. Optimum solutions in general require mathematically and computationally complex approaches, but Ohlsson proposed simpler methods involving permanent random numbers applicable in some situations. Although not optimal, the methods are easily implemented and typically realize much of the gain achieved by the optimal solution. Ernst extended Ohlsson’s methods for sequential methods such as Durbin/Brewer method, by a probabilistically correct retrospective assignment of permanent random numbers. This paper presents an extension of the Ernst approach when the first sample was selected by drawing more than one unit per stratum systematically and illustrates its efficiency with a simulation study.
Now almost 40 years old, the National Crime Victimization Survey (NCVS) provides annual estimates of the number of victimizations by different types of crime and estimates of the characteristics of the victims. The Bureau of Justice Statistics (BJS) is currently investigating several strategies to expand the usefulness of the survey. A major goal is to shift the almost exclusive focus of the NCVS as a national survey to a revised program capable of also providing subnational detail relevant to state and local governments. Until now, the NCVS has been conducted exclusively by the Census Bureau. Currently, BJS is investigating a strategy that includes retaining a core NCVS preserving many of the basic features of the current design, including address sampling and personal visit as the primary mode for an initial interview. The core NCVS would continue to be conducted by the Census Bureau. At the same time, BJS is supporting research to determine the possibility of integrating the core NCVS with auxiliary or supplemental data collected by quite different survey strategies, modes, and collection agents. This paper will detail an investigation of one aspect of BJS’s overall strategy: To what extent can strategic sample boosting and allocation of the core NCVS meet a key set of subnational estimation goals through direct estimation? Direct survey estimates of annual crime rates for each of the states would require an unrealistic expansion of the survey, but the analysis reported here examines to what extent defining a restricted set of areas, such as states of over 8 million population and using 3or 5-year period averages in a manner similar to the American Community Survey, could achieve a set of useful results to address the needs of many NCVS users.
The National Crime Victimization Survey (NCVS) has provided annual estimates of the number of victimizations for several types of crime since 1972, with an almost exclusive focus on national rates. Most of the programs to prevent or reduce crime are implemented locally, however. To respond to a resulting interest in subnational statistics, the Bureau of Justice Statistics (BJS) has been recently supporting research on a variety of approaches to produce subnational estimates. In this paper, we report on the potential application of model-based small area estimation methods based on the NCVS and auxiliary data, particularly the FBI’s Uniform Crime Reports, using empirical best linear unbiased estimation (EBLUP). We compare a time-series model introduced by Rao and Yu to a new variant, termed here the dynamic model. We will also indicate how the small area approach might be integrated with other approaches that BJS is currently considering, including possible expansion of the NCVS sample size to augment the survey’s capacity to produce direct estimates for some or all states.
The overall design of the National Crime Victimization Survey (NCVS) has been largely stable for over 30 years. Households in sampled housing units are interviewed for 7 waves, each time collecting data for the previous 6 months. Until recently, the 1st wave has been omitted from the published estimates to exclude reports outside of the intended 6-month reference period. Using publicly available data for 1998-2004, we report on a series of analyses to investigate more current effects of the bounding and find evidence of more general time-in-sample effects. We will also report on the effect of recency in the observed incident reports, where more crimes are reported in the 1st month preceding the interview date than each of the previous 5 months. These phenomena are important in considering a range of design options for the NCVS that would alter the reference period or panel design.
The National Crime Victimization Survey (NCVS) began full-scale data collection in 1972. Over time, increasing costs of data collection and a diminishing budget have forced a gradual reduction of the sample size from 72,000 households in 1972 to 38,600 in 2005, leaving the NCVS far less able to meet its initial goals. The sponsor of the NCVS, the Bureau of Justice Statistics (BJS), asked a panel of the National Academy of Sciences to review the survey. The panel’s interim report, released in 2008, recommended to BJS a series of actions, including a systematic review of a range of sample design options for the survey. Since the beginning of the NCVS, the Census Bureau has been the primary architect and data collector. The panel report addressed the question of whether BJS should consider an alternative collector, but it detailed reasons for and against continuing with the Census Bureau. The report recommended, however, that the Census Bureau provide BJS with detailed information on costs to allow external assessment of survey design options. The report also recommended that BJS consult external sources to evaluate design options. On that basis, BJS has funded the research to be reported here. This paper summarizes our research design for the study. The panel report provided two forms of guidance for this research. For one, it listed the essential components of the design of a national household survey—the design, stratification, and selection of primary sampling units; methods for selecting households; the choice between a rotating panel or cross-sectional design; and possible subsampling of individuals within households. The panel report also recommended consideration of specific alternative designs, including the design of the British Crime Survey. Our planned research responds to both forms of guidance, attempting to build on the panel’s thoughtful review of the current status of the NCVS.
In 2010, the U.S. Census Bureau will publish the first set of 5-year period estimates from the American Community Survey (ACS), based on data for 2005-2009. Published for small places, tracts, and other small areas, the 5-year ACS estimates will attempt to replace the long-form data from recent decennial censuses. Because the effective sample size for the ACS will be somewhat less than half that of previous censuses, users will face an increased challenge to distinguish true variation from sampling error. The concept of the false discovery rate has become increasingly useful in other disciplines confronted by large numbers of estimates, such as micro-array analysis in genetics and fMRI studies of the brain. The paper will review this concept and suggest its possible future application, based on a preliminary analysis of published data from the ACS Multiyear Estimates Study.
The American Community Survey (ACS) began full implementation in 2005. Estimation for ACS, as typical for other large Census Bureau surveys, uses a complex series of ratio estimation and other adjustments to the weights. Working originally from ACS data for 1999- 2001 in 36 test counties, previous research suggested that multiyear tract-level estimates could be improved by imbedding a step of model-assisted estimation, specifically generalized regression estimation (GREG), in the current ACS estimation. In particular, the GREG step incorporates administrative record data. Most data sets produced for the ACS Multiyear Estimates Study incorporate the GREG step, but alternative sets of 3- and 5-year estimates without the GREG step were produced for purposes of comparison. The paper will describe new refinements in the estimation approach and its potential future role in ACS estimation.
The American Community Survey (ACS) began full implementation in 2005 as a replacement for the decennial census long form. In 2010, ACS estimates will be released for the 5-year period 2005-2009; this release will be the first to offer ACS results at the geographic detail previously provided by Census 2000. A model-assisted approach has been proposed to reduce the variances of estimates for geographic units below the county level for both 3and 5-year period estimates. The approach has been investigated as part of the Multiyear Estimates Study, based on an ACS test in 34 counties during 1999-2005. The anticipated variance reductions were empirically confirmed, although substantial reductions occurred for a few key variables and more modest ones for others. Given the beneficial variance impact, the primary focus of this paper will be to address some of the remaining concerns potential ACS users may have about the model-assisted methods, such as their impact on bias and whether their introduction complicates analysis.
Full implementation of the American Community Survey (ACS) began in 2005. Among other purposes, the ACS will replace the decennial census long-form data, enabling a short-form census in 2010. A test implementation of the ACS in 36 test counties during 1999-2001 suggested that the initial ACS estimation procedure, while adequate at the county level, yields higher variances at the small-area levels of tract and block group relative to the decennial long form. Previously reported research argued the likely success of an approach combining administrative record data with a generalized regression estimator. The case for likely success was based on the R-square of the underlying regression. This paper reports detailed results on a full implementation in 34 test counties, showing the degree to which the predictions based on regression diagnostics have been confirmed.
In recent decades, U.S. censuses have produced relatively accurate population counts, but the considerable importance of the results has driven efforts to study and possibly correct the errors of coverage that occur. The 2000 Accuracy and Coverage Evaluation (A.C.E.) attempted to measure the net error of the census, originally with the intention of correcting the census counts for all purposes other than the apportionment of the House of Representatives. In turn, the accuracy of the A.C.E. was assessed by several evaluation studies, including a reinterview (the Evaluation Followup or EFU). Although many of the evaluations implied that the A.C.E. had been generally successful, the reinterview indicated that the A.C.E. had seriously underestimated some types of erroneous enumerations in the census, including persons who lived elsewhere on Census Day. In October 2001, the U.S. Census Bureau decided not to incorporate the A.C.E. findings into Census 2000 results. Despite extensive evaluation of the A.C.E., so far there has been little detailed assessment of how the questions in the follow-up and reinterview instruments contributed to the different results of the two surveys. In part, this reflects the way the interview data were actually used. Interviewers were encouraged to, and did, record extensive notes describing residence situations. The notes were heavily relied on by the analysts and clerks who determined final residence status. Information from questionnaire responses and interviewers’ notes was clerically integrated into summary codes for analysis. The summary codes were the basis for official A.C.E. coverage estimates and the analyses that informed the October 2001 decision. In this paper, we take a different approach, and analyze the recently available responses to individual questionnaire items designed to measure where people lived on Census Day, April 1, 2000. Our objective is to identify and clarify particular sources of error and bias, and to provide a basis for improving the design of future coverage measurement instruments. We link reports given in A.C.E. interviews (including follow-ups conducted as part of the A.C.E.) with the more detailed information provided in reinterviews (EFU) in order to examine the consistency of reporting. W e also introduce independent, auxiliary evidence of census duplications, availab le from a special study conducted to evaluate the quality of coverage estimates. We focus here on measurements of mobility; a more complete paper (available from the first author) includes additional analyses of other coverage measurements. We first briefly review sources of coverage error and the Census Bureau’s coverage measurement methods and describe the sources of our data.