
Objective:This report compares results between survey analysis software and the National Cancer Institute Joinpoint Regression Software for trend analysis of survey data. The "National Center for Health Statistics Guidelines for Analysis of Trends" recommends that analysts use record-level data and survey analysis software to fit desired trend models of survey data. When changes in a trend will be assessed using piecewise regression models, the guidelines recommend that analysts use the most recent version of the Joinpoint software with aggregated data to identify the number and location of joinpoints, then survey analysis software with record-level data to obtain final slope estimates and to conduct tests of hypothesis for the model identified by the Joinpoint software. In practice, the Joinpoint software sometimes produces results that appear incongruent with those generated by the survey analysis software. The purpose of this report is to provide guidance for conducting trend analyses on NCHS survey data to handle and explain these apparent inconsistencies. This report should be considered a supplement to the "National Center for Health Statistics Guidelines for Analysis of Trends." Methods:Cases were identified where apparent inconsistencies between the Joinpoint software and survey analysis software could occur. Plausible explanations for the apparent differences are discussed through text, examples, and frequently asked questions. Solutions are not provided; rather, recommendations, cautions, and additional information are provided to assist analysts in making the best decisions for their analysis and data. Results:Most frequently, inconsistencies occur when the prespecified piecewise regression model provided by the Joinpoint software is estimated using survey analysis software and successive slopes are not statistically significantly different from one another, resulting in one or more joinpoints being removed from the final model. Potential explanations are divided into five main categories.
Background:Synthetic data has been gaining popularity in many fields as an approach to retain data utility (the validity of inference using synthetic data) and protect confidentiality. However, creating synthetic data for complex surveys remains a challenge. Methods:This research compared three approaches to incorporate survey design information (stratification, clustering, and sampling weights) during the synthetic data-generating process using the Research and Development Survey (RANDS), a series of primarily web surveys conducted by the National Center for Health Statistics, Centers for Disease Control and Prevention. Both parametric (logistic and linear regression models) and nonparametric (classification and regression trees [CART]) methods were used to create synthetic data. Data utility and disclosure risk were evaluated via confidence interval overlap, propensity score measurement, and average matching probability for re-identification. Results:Using the original survey design information as predictors during the synthesis process improved data utility for the parametric method. However, the nonparametric method yielded results with better data utility but slightly higher disclosure risk.
Introduction:The Healthy People initiative provides science-based, 10-year public health objectives and targets for the U.S. population. As in the previous four initiatives, Healthy People 2030 established overarching goals and objectives (with targets) at the start of the decade and will be monitoring progress toward the attainment of targets and elimination of health disparities among population subgroups over the course of the decade. Objective:This report outlines Healthy People 2030 measurement practices for both progress toward target attainment and elimination of disparities and compares the 2030 measurement practices with those that were in place in 2020, highlighting strengths and limitations. Methods:Progress toward target attainment is assessed for the total population. The "percentage of targeted change achieved" quantifies movement toward targets, and the "percentage change from baseline" can be calculated for all core objectives. Based on the percentage of targeted change achieved or percentage change from baseline, as well as the statistical significance of these measures (when applicable), core objectives in Healthy People 2030 are classified into four mutually exclusive categories: TARGET MET OR EXCEEDED, IMPROVING, LITTLE OR NO DETECTABLE CHANGE, or GETTING WORSE. Disparities at a single timepoint are assessed by a suite of six measures: the between-group rate difference and ratio; summary rate difference and ratio; and maximal rate difference and ratio. To enable comparisons among those six measures, changes in disparities over time are assessed using the percentage change from baseline. Variability (standard errors and 95% confidence intervals) and statistical significance for all six measures, when applicable, are derived using a resampling/bootstrap procedure. Conclusion:Expanding and building on the approaches to measurement in previous decades, methods to measure progress toward target attainment and elimination of health disparities in Healthy People 2030 represent a further evolution of these methods and address methodological issues and limitations previously identified.
Objectives The Research and Development Survey (RANDS) is a series of web-based, commercial panel surveys that have been conducted by the National Center for Health Statistics (NCHS) since 2015. RANDS was designed for methodological research purposes,including supplementing NCHS' evaluation of surveys and questionnaires to detect measurement error, and exploring methods to integrate data from commercial survey panels with high-quality data collections to improve survey estimation. The latter goal of improving survey estimation is in response to limitations of web surveys, including coverage and nonresponse bias. To address the potential bias in estimates from RANDS,NCHS has investigated various calibration weighting methods to adjust the RANDS panel weights using one of NCHS' national household surveys, the National Health Interview Survey. This report describes calibration weighting methods and the approaches used to calibrate weights in web-based panel surveys at NCHS.
Background Linking health survey data to administrative records expands the analytic utility of survey participant responses, but also creates the potential for new sources of bias when not all participants are eligible for linkage. Residual differences-bias-can occur between estimates made using the full survey sample and the subset eligible for linkage. Objective To assess linkage eligibility bias and provide examples of how bias may be reduced by changes in questionnaire design and adjustment of survey weights for linkage eligibility. Methods Linkage eligibility bias was estimated for various sociodemographic groups and health-related variables for the 2000-2013 National Health Interview Surveys. Conclusions Analysts using the linked data should consider the potential for linkage eligibility bias when planning their analyses and use approaches to reduce bias, such as survey weight adjustments, when appropriate.
Background The purpose of the National Health and Nutrition Examination Survey (NHANES) is to produce national estimates representative of the total noninstitutionalized civilian U.S. population. The sample for NHANES is selected using a complex, four-stage sample design. NHANES sample weights are used by analysts to produce estimates of the health-related statistics that would have been obtained if the entire sampling frame (i.e., the noninstitutionalized civilian U.S. population) had been surveyed. Sampling errors should be calculated for all survey estimates to aid in determining their statistical reliability. For complex sample surveys, exact mathematical formulas for variance estimates that fully incorporate the sample design are usually not available. Variance approximation procedures are required to provide reasonable, approximately unbiased, and design-consistent estimates of variance. Objective This report describes the NHANES 2015-2018 sample design and the methods used to create sample weights and variance units for the public-use data files, including sample weights for selected subsamples, such as the fasting subsample. The impacts of sample design changes on estimation for NHANES 2015-2018 are described. Approaches that data users can use to modify sample weights when combining survey cycles or when combining subsamples are also included.
Over the past two decades, a steady decline in response rates on national face-to-face surveys has been documented, with steeper declines observed in recent years. The impact of nonresponse on survey estimates is inconsistent and depends on the correlation between response propensity and the survey estimates. To better understand the impact of declining response rates on the 2017-2018 National Health and Nutrition Examination Survey (NHANES), potential nonresponse bias (NRB) was investigated. NRB was assessed using three approaches: (a) studying variation within the respondent set; (b) benchmarking and comparisons to external data; and (c) comparing alternative weighting adjustments. Because NHANES only samples 30 counties in every 2-year cycle, the sample of counties in any given cycle may be an outlier on some characteristics. Such sampling variability may compound the effects of NRB. For this reason, the representativeness of the 2017-2018 NHANES counties was examined by comparing: (a) the characteristics of the 2017-2018 sampled counties with those from prior cycles; (b) each sampled county with the average of all the counties in the sampling stratum from which that county was selected; and (c) the 2017-2018 counties with 5,000 other samples that could have been drawn under the same sample design using a simulation study. The NRB analyses showed that the 2017-2018 NHANES sample had a lower proportion of college graduates and higher-income individuals compared with prior cycles. Additionally, the 2017-2018 NHANES counties had lower proportions of college graduates and lower mean incomes compared with counties from prior cycles and counties not selected in 2017-2018, which exacerbated the effects of NRB. Weighting adjustments used in prior cycles were not sufficient to address the bias in the 2017-2018 NHANES. Instead, enhanced weighting adjustments for education and income reduced the bias resulting from nonresponse and location sampling variability.
Objective This report compares five methods of waist circumference (WC) measurements: 1) the National Heart, Lung, and Blood Institute (NHLBI-WC); 2) the World Health Organization (WHO-WC); 3) the Multi-Ethnic Study of Atherosclerosis (MESA-WC) using Gulick II Plus tape; 4) the Multi-Ethnic Study of Atherosclerosis (MESA-WC) using Lufkin tape; and 5) assisted self-measurement over clothes (MESA-assisted). Method During 2016, measurements were obtained from 2,297 participants aged 20 and over, who participated in the National Health and Nutrition Examination Survey (NHANES). The mean differences and sensitivity and specificity for abdominal obesity (AO) were calculated between the NHLBI-WC (reference) and the other four WC measurements. Results The mean difference between NHLBI-WC and WHO-WC was 0.81 cm for men and 3.21 cm for women ( p ≤ 0.0125 for both); between NHLBI-WC and MESA-WC (Gulick) was -0.68 cm for men ( p ≤ 0.0125) and -0.89 cm for women; between NHLBI-WC and MESA-WC (Lufkin) was 0.02 cm for men and 0.08 cm for women; and between NHLBI-WC and MESA-assisted was -0.71 cm for men and 1.34 cm for women ( p ≤ 0.0125 for both). Sensitivity and specificity for AO, with NHLBI-WC as a reference, for men were greater than 90% for all methods; for women, sensitivity and specificity for AO for MESA-WC (Lufkin) were greater than 90%; for women, WHO-WC, MESAWC (Gulick), and MESA-assisted methods were greater than 85%.
Statistically reliable, abridged, period life tables were produced for 88.7% of U.S. census tracts (65,662). A battery of tests revealed that the census-tract life table functions followed expected patterns; their distribution about state and U.S. values showed no aberrations; and their weighted mean values compared well with state- and national-level estimates. The weighted mean life expectancy at birth for the 65,662 census tracts was 78.7 years compared with the official U.S. estimate of 78.8 years in midyear 2013. The results of this study concur with previous research showing that a minimum population size of 5,000 is acceptable, with the caveat that missing age-specific death counts cannot be ignored. The methodology developed for this study addressed the issues of small populations and zero deaths as robustly as possible, although it is not without error.
To describe methodological issues that arise in the construction and design-based estimation of multidimensional indices that aggregate state-specific inequalities in core health measures, using data from the National Health Interview Survey (NHIS).
Many reports present analyses of trends over time based on multiple years of data from National Center for Health Statistics (NCHS) surveys and the National Vital Statistics System (NVSS). Trend analyses of NCHS data involve analytic choices that can lead to different conclusions about the trends. This report discusses issues that should be considered when conducting a time trend analysis using NCHS data and presents guidelines for making trend analysis choices. Trend analysis issues discussed include: choosing the observed time points to include in the analysis, considerations for survey data and vital records data (record level and aggregated), a general approach for conducting trend analyses, assorted other analytic issues, and joinpoint regression. This report provides 12 guidelines for trend analyses, examples of analyses using NCHS survey and vital records data, statistical details for some analysis issues, and SAS and SUDAAN code for specification of joinpoint regression models. Several an lytic choices must be made during the course of a trend analysis, and the choices made can affect the results. This report highlights the strengths and limitations of different choices and presents guidelines for making some of these choices. While this report focuses on time trend analyses, the issues discussed and guidelines presented are applicable to trend analyses involving other ordinal and interval variables.
This report describes the methods used to create NHANES 2011-2014 sample weights and variance units for the public-use data files, including sample weights for selected subsamples, such as the fasting subsample. The impacts of sample design changes on estimation for NHANES 2011-2014 and the addition of the NHANES National Youth Fitness Survey (NNYFS) 2012 are described. Approaches that data users can employ to modify sample weights when combining survey cycles or when combining subsamples are also included.
Dietary recommendations are intended to be met based on dietary intake over long periods, as associations between diet and health result from habitual intake, not a single eating occasion or day of intake. Measuring usual intake directly is impractical for large population-based surveys due to the respondent burden associated with reporting habitual intake over longer periods. Therefore, analytical techniques were developed to estimate usual intake using as few as 2 days of 24-hour dietary recall data. With National Health and Nutrition Examination Survey (NHANES) data, this report demonstrates how to estimate usual intake using the National Cancer Institute (NCI). This report demonstrates how to estimate the usual intake of nutrients consumed daily or episodically using NHANES data. Means, percentiles, and the percentages above or below specified Dietary Reference Intake (DRI) values for given day, within-person mean (WPM), and estimates of usual intake are presented. Consistent with previous analyses, mean intakes were similar across methods. However, the distributions estimated by nonusual intake methods were wider compared with the NCI Method, which can lead to misclassification of the percentage of the population above or below certain DRIs. Use of NHANES data to examine the proportion of the population at risk of insufficiency or excess of certain nutrients, with methods like given day and WPM that do not address within-person variation, may lead to biased estimates.
The 2014 Native Hawaiian and Pacific Islander National Health Interview Survey (NHPI NHIS) is the first federal survey designed exclusively to measure the health of the noninstitutionalized civilian NHPI population of the United States.
The National Center for Health Statistics (NCHS) disseminates information on a broad range of health topics through diverse publications. These publications must rely on clear and transparent presentation standards that can be broadly and efficiently applied. Standards are particularly important for large, cross-cutting reports where estimates cannot be individually evaluated and indicators of precision cannot be included alongside the estimates. This report describes the NCHS Data Presentation Standards for Proportions. The multistep NCHS Data Presentation Standards for Proportions are based on a minimum denominator sample size and on the absolute and relative widths of a confidence interval calculated using the Clopper-Pearson method. Proportions (usually multiplied by 100 and expressed as percentages) are the most commonly reported estimates in NCHS reports.
Objective This report examines ways to improve National Ambulatory Medical Care Survey (NAMCS) data on practice and physician characteristics in multispecialty group practices. Methods From February to April 2013, the National Center for Health Statistics (NCHS) conducted a pilot study to observe the collection of the NAMCS physician interview information component in a large multispecialty group practice. Nine physicians were randomly sampled using standard NAMCS recruitment procedures; eight were eligible and agreed to participate. Using standard protocols, three field representatives conducted NAMCS physician induction interviews (PIIs) while trained ethnographers observed and audio recorded the interviews. Transcripts and field notes were analyzed to identify recurrent issues in the data collection process. Results The majority of the NAMCS items appeared to have been easily answered by the physician respondents. Among the items that appeared to be difficult to answer, three themes emerged: (a) physician respondents demonstrated an inconsistent understanding of "location" in responding to questions; (b) lack of familiarity with administrative matters made certain questions difficult for physicians to answer; and (c) certain primary care‑oriented questions were not relevant to specialty care providers. Conclusions Some PII survey questions were challenging for physicians in a multispecialty practice setting. Improving the design and administration of NAMCS data collection is part of NCHS' continuous quality improvement process.
Background California is the most populated state and Los Angeles County is the most populated county in the United States. National Health and Nutrition Examination Survey (NHANES) sample weights and variance units were developed for these places to obtain subnational estimates. Objective This report describes the California and Los Angeles County NHANES 1999-2006 and 2007-2014 samples, including the creation of the sample weights and variance units and descriptions of the resulting data files. Some analytic guidelines are provided. Results Eight years of NHANES data were combined for each data file to provide an adequate sample size and reduce disclosure risks. Because Los Angeles County has been a self-representing primary sampling unit, sample weights for Los Angeles County were relatively straightforward. However, a modelbased approach was used to create sample weights for California. The relatively large proportion of Mexican- American and other Hispanic persons in California, coupled with the different NHANES 1999-2014 sample design requirements for oversampling these groups within the small number of NHANES locations selected each cycle, led to a relatively large size of these groups in the California and Los Angeles County NHANES files. For example, 1,137 and 374 of the 3,353 Mexican-Americans persons in NHANES 2007-2014 were in the California and Los Angeles County samples, respectively. Conclusion The California and Los Angeles County NHANES 1999-2006 and 2007-2014 samples are available in the National Center for Health Statistics Research Data Center.
BACKGROUND:The National Ambulatory Medical Care Survey (NAMCS) is an annual, nationally representative sample survey of physicians and of visits to physicians. Two major changes were made to the 2012 NAMCS to support reliable state estimates. The sampling design changed from an area sample to a fivefold-larger list sample of physicians stratified by the nine U.S. Census Bureau divisions and 34 states. At the same time, the data collection mode changed from paper forms to laptop-assisted data collection and from physician or office staff abstraction of medical records to predominantly Census interviewer abstraction using automated Patient Record Forms (PRFs).OBJECTIVES:This report presents an analysis of potential nonresponse bias in 2012 NAMCS estimates of physicians and visits to physicians. This analysis used two sets of physician-based estimates: one measuring the completion of the physician induction interview and another based on completing any PRF. Evaluation of visit response was measured by the percentage of expected PRFs completed. For each type of physician estimate, response was evaluated by (a) comparing percent distributions of respondents and nonrespondents by physician characteristics available for all in-scope sample physicians, (b) comparing response rates by physician characteristics with the national response rate, and (c) analyzing nonresponse bias after adjustments for nonresponse were applied in survey weights. For visit estimates, response was evaluated by (a) comparing the percent distributions of expected visits and completed visits, (b) comparing visit response rates by physician characteristics with the national visit response rate, and (c) analyzing visit-level nonresponse bias after adjustments for nonresponse were applied in visit survey weights. Finally, potential bias in the two physician-level estimates was computed by comparing them with those from an external survey.
ObjectivesThis report presents the findings of an updated study of the validity of race and Hispanic-origin reporting on death certificates in the United States, and its impact on race- and Hispanic origin-specific death rates.MethodsThe latest version of the National Longitudinal Mortality Study (NLMS) was used to evaluate the classification of race and Hispanic origin on death certificates for deaths occurring in 1999–2011 to decedents in NLMS. To evaluate change over time, these results were compared with those of a study based on an earlier version of NLMS that evaluated the quality of race and ethnicity classification on death certificates for 1979–1989 and 1990–1998. NLMS consists of a series of annual Current Population Survey files (1973 and 1978–2011) and a sample of the 1980 decennial census linked to death certificates for 1979–2011. Pooled 2009–2011 vital statistics mortality data and 2010 decennial census population data were used to estimate and compare observed and corrected race- and Hispanic origin-specific death rates.ResultsRace and ethnicity reporting on death certificates continued to be highly accurate for both white and black populations during the 1999–2011 period. Misclassification remained high at 40% for the American Indian or Alaska Native (AIAN) population. It improved, from 5% to 3%, for the Hispanic population, and from 7% to 3% for the Asian or Pacific Islander (API) population. Decedent characteristics such as place of residence and nativity affected the quality of reporting on the death certificate. Effects of misclassification on death rates were large for the AIAN population but not significant for the Hispanic or API populations.
BACKGROUND:The National Health and Nutrition Examination Survey's 9NHANES) biospecimena program was formed to manage the collection of biospecimena (including serum, plasma, urine, and DNA) from NHANES cycles, the storage of biospecimens in NHANES biospecimens, accessing of biospecimens by researchers and the providing of resulting data to future researchers. Data from biospeceimen research can be combined with existing NHANES data.OBJECTIVE:This report provides background on the development of NHANES biorepositories and describes the collection, processing, and storing of biospecimens; ethical considerations and informed consent; and the proposal process for accessing biospecimens and resulting data. The number and types of biospecimens collected in each survey cycle from NHANES III (1988- 1994) through NHANES 1999-2014 are discussed so that researchers can understand what biospecimens are available if they are considering using NHANES biospecimens in their research.