Nearest neighbor imputation is a popular procedure used to complete missing records. However, the nearest neighbor imputation estimator suffers from a bias that increases as the dimension of the matching variable increases. We study estimators that are model unbiased and have the hot-deck property. Estimation for domains is discussed. In the simulation study for the population mean, a proposed estimator is less biased and more efficient than the nearest neighbor estimator. The suggested replication variance estimator does not require repeated imputation and is appropriate for vectors with different missing patterns. In the simulation, the variance estimator gives confidence intervals with coverage rates close to the nominal level.
The standard method of creating post-strata is to define the boundaries of the strata on the basis of population characteristics of auxiliary variables. Estimation treats the post-strata as strata for standard stratified-sample estimation. Samples often contain empty post-strata requiring adjustment to the estimation procedure. To avoid empty post-strata, we propose using the sample distribution function of an auxiliary variable to define the post-strata. We show that the large-sample efficiency of the sample-based post-stratification procedure is the same as that of the equivalently defined population-based procedure. In the simulation, the sample-based procedure was slightly more efficient than the classical procedure. The Monte Carlo coverage of a nominal 95% interval was approximately 95% for the sample-based procedure and approximately 94% for the classical procedure.
We present a sampling algorithm for an ordered population with elements that have a measure of size. The algorithm enables one to select a sample with specified probabilities and with efficiency for the estimated mean between that of one per stratum and that of two per stratum. The algorithm contains a design parameter for the efficiency of the estimated mean relative to the efficiency of the estimated variance of the estimated mean. For a variable highly correlated with the order, it is possible for both the efficiency of the estimated mean and the efficiency of the estimated variance for an intermediate design to be greater than that for the two-per-stratum design. For most studied populations, the variance of the estimated mean declines and the variance of the estimated variance increases as one moves from the two-perstratum design toward the one-per-stratum design. We illustrate the trade-off between the variance of the estimated mean and the variance of the estimated variance using an autoregressive process. An estimator of the variance of the estimated mean and a replication form for variance estimation are given.
For analyses based on nonlinear models, agencies and policy makers are often interested in prediction intervals for small area means. We give statistics for small area predictions that can be used to construct prediction intervals in the same way that standard errors and degrees of freedom are used to construct prediction intervals based on the Student-t distribution. In a simulation study, the new parametric bootstrap prediction interval has good coverage properties and much better coverage than the bootstrap percentile prediction interval. The methods are applied in a study of soil erosion and water runoff conducted by the US Department of Agriculture.
Small area estimation often involves constructing predictions with an estimated model followed by a benchmarking step. In the benchmarking operation, the predictions are modified so that weighted sums satisfy constraints. The most common constraint is the constraint that a weighted sum of predictions is equal to the same weighted sum of the original observations. Two benchmarking procedures for nonlinear models are proposed: a linear additive adjustment and a method based on an augmented model for the expectation function. Variance estimators for benchmarked predictors are presented and vetted through simulation studies. The benchmarking procedures are applied to county estimates of the proportion of area in cropland using data from the National Resources Inventory. The Canadian Journal of Statistics 46: 482–500; 2018 © 2018 Statistical Society of Canada
We discuss developments in sample survey theory and methods covering the past 100 years. Neyman’s 1934 landmark paper laid the theoretical foundations for the probability sampling approach to inference from survey samples. Classical sampling books by Cochran, Deming, Hansen, Hurwitz and Madow, Sukhatme, and Yates, which appeared in the early 1950s, expanded and elaborated the theory of probability sampling, emphasizing unbiasedness, model free features, and designs that minimize variance for a fixed cost. During the period 1960-1970, theoretical foundations of inference from survey data received attention, with the model-dependent approach generating considerable discussion. Introduction of general purpose statistical software led to the use of such software with survey data, which led to the design of methods specifically for complex survey data. At the same time, weighting methods, such as regression estimation and calibration, became practical and design consistency replaced unbiasedness as the requirement for standard estimators. A bit later, computer-intensive resampling methods also became practical for large scale survey samples. Improved computer power led to more sophisticated imputation for missing data, use of more auxiliary data, some treatment of measurement errors in estimation, and more complex estimation procedures. A notable use of models was in the expanded use of small area estimation. Future directions in research and methods will be influenced by budgets, response rates, timeliness, improved data collection devices, and availability of auxiliary data, some of which will come from “Big Data”. Survey taking will be impacted by changing cultural behavior and by a changing physical-technical environment.
Replication procedures have proven useful for variance estimation for large scale complex surveys. As an extension of bootstrap procedures to rejective samples, we define a bootstrap sample that is a rejective, unequal probability, replacement sample selected from the original sample. A modification of the bootstrap with improved performance is suggested for stratified samples with small stratum sizes. Simulations for Poisson and stratified rejective samples support the use of replicates in estimating the variance of the regression estimator for rejective samples.
Construction of small area predictors and estimation of the prediction mean squared error, given different types of auxiliary information are illustrated for a unit level model. Of interest are situations where the mean and variance of an auxiliary variable are subject to estimation error. Fixed and random specifications for the auxiliary variables are considered. The efficiency gains associated with the random specification for the auxiliary variable measured with error are demonstrated. A parametric bootstrap procedure is proposed for the mean squared error of the predictor based on a logit model. The proposed bootstrap procedure has smaller bootstrap error than a classical double bootstrap procedure with the same number of samples.
Errors in Variables with Emphasis on Theory† Wayne A. Fuller, Wayne A. FullerSearch for more papers by this author Wayne A. Fuller, Wayne A. FullerSearch for more papers by this author First published: 29 September 2014 https://doi.org/10.1002/9781118445112.stat03452 †This article was originally published online in 2006 in Encyclopedia of Statistical Sciences, © John Wiley & Sons, Inc. and republished in Wiley StatsRef: Statistics Reference Online, 2014. Read the full textAboutPDF ToolsRequest permissionExport citationAdd to favoritesTrack citation ShareShare Give accessShare full text accessShare full-text accessPlease review our Terms and Conditions of Use and check box below to share full-text version of article.I have read and accept the Wiley Online Library Terms and Conditions of UseShareable LinkUse the link below to share a full-text version of this article with your friends and colleagues. Learn more.Copy URL Share a linkShare onFacebookTwitterLinked InRedditWechat No abstract is available for this article. Wiley StatsRef: Statistics Reference OnlineBrowse other articles of this reference work:BROWSE BY TOPICBROWSE A-Z RelatedInformation
Physical activity measurements derived from self-report surveys are prone to measurement errors. Monitoring devices like accelerometers offer more objective measurements of physical activity, but are impractical for use in large-scale surveys. A model capable of predicting objective measurements of physical activity from self-reports would offer a practical alternative to obtaining measurements directly from monitoring devices. Using data from National Health and Nutrition Examination Survey 2003–2006, we developed and validated models for predicting objective physical activity from self-report variables and other demographic characteristics. The prediction intervals produced by the models were large, suggesting that the ability to predict objective physical activity for individuals from self-reports is limited.
Small Area Prediction of Proportions with Applications to the Canadian Labour Force Survey Get access Emily J. Berg, Emily J. Berg * *Address correspondence to Emily Berg, 1218 Snedecor, Ames, IA 50010; E-mail: emilyb@iastate.edu. Search for other works by this author on: Oxford Academic Google Scholar Wayne A. Fuller Wayne A. Fuller Search for other works by this author on: Oxford Academic Google Scholar Journal of Survey Statistics and Methodology, Volume 2, Issue 3, September 2014, Pages 227–256, https://doi.org/10.1093/jssam/smu011 Published: 01 September 2014
A parametric bootstrap procedure is proposed for the mean squared error of the predictor based on a unit level model. It is demonstrated that the proposed procedure has smaller bootstrap error than a classical double bootstrap procedure with the same number of samples. Applications to a logit model under different types of auxiliary information are discussed.
Direct estimates for small areas or subpopulations may not be reliable because of small sample sizes for such objects. Procedures based on implicit or explicit models have been used to construct better estimates for given small areas, by exploiting auxiliary information. In this paper we consider binary responses, and investigate predictors for situations with different amounts of available information. We use generalized linear mixed models and present bias and mean squared error results for different prediction methods. Procedures based on models have been used to construct estimates for small areas, by exploiting auxiliary information. In this paper, we study nested models with a binary response and random area effects. These models form a subclass of generalized linear mixed models. We also consider stochastic covariates. Survey data often contain auxiliary variables with good correlation with the variable of interest. However, area level auxiliary data may be incomplete. We consider three cases of auxiliary information, when the covariates have known mean, when the covariates have unknown distribution, and when the covariates have unknown random mean. For the last two cases, we describe estimation methods for the area mean of the auxiliary data. Because the response variable is binary and the auxiliary information is not fixed, estimation and prediction are not as straight forward as in linear mixed models. Mixed models with unit level auxiliary data have been used for small area estimation by a number of authors. Battese, Harter, and Fuller (1988) use a linear mixed model to predict the area planted with corn and soybeans in Iowa counties. Datta and Ghosh (1991) introduce the hierarchical Bayes predictor for general mixed linear models. Larsen (2003) compared estimators for proportions based on two unit level models, a simple model with no area level covariates and a model using the area level information. Malec (2005) proposes Bayesian small area estimates for means of binary responses using a multivariate binomial/multinomial model. Jiang (2007) reviews the classical inferential approach for linear and generalized linear mixed models and discusses the prediction for a function of fixed and random effects. Ghosh et al (2009) consider a small area model where covariates have unknown distribution. They assume the sample has been selected so that weights!ij are available satisfying P ni j=1 !ij = 1. They consider both hierarchical Bayes and EB estimators
Fractional hot deck imputation, considered in Fuller and Kim (2005), is extended to multivariate missing data. The joint distribution of the study items is nonparametrically estimated using a discrete approximation, where the discrete transformation also serves to define imputation cells. The procedure first estimates the probabilities for the cells and then imputes real observations for missing items. Calibration weighting is used to reduce the imputation variance. Replication variance estimation is discussed.
In this article the authors discuss and develop a statistical model that describes the relationship between an individual's short-term recall of his or her physical activity and usual activity over a long period of time. Modeling measurement error in recall data can be used to provide more accurate estimates of long-term activity behavior.
Background:Physical activity recall instruments provide an inexpensive method of collecting physical activity patterns on a sample of individuals, but they are subject to systematic and random measurement error. Statistical models can be used to estimate measurement error in activity recalls and provide more accurate estimates of usual activity parameters for a population.Methods:We develop a measurement error model for a short-term activity recall that describes the relationship between the recall and an individual’s usual activity over a long period of time. The model includes terms for systematic and random measurement errors. To estimate model parameters, the design should include replicate observations of a concurrent activity recall and an objective monitor measurement on a subsample of respondents.Results:We illustrate the approach with preliminary data from the Iowa Physical Activity Measurement Study. In this dataset, recalls tend to overestimate actual activity, and measurement errors greatly increase the variance of recalls relative to the person-to-person variation in usual activity. Statistical adjustments are used to remove bias and extraneous variation in estimating the usual activity distribution.Conclusions:Modeling measurement error in recall data can be used to provide more accurate estimates of long-term activity behavior.
Prediction for the mixed model requires estimates of covariance matrices. There is often a direct estimate of the ''within area'' covariance matrix, and for survey samples this is an estimate of the sampling covariance matrix. The estimated covariance matrix may have large sampling variance, suggesting parametric modeling for the matrix. The model can play a role at various points in the construction of predictions for proportions for small areas. Simulations demonstrate that efficiency for predictions is improved by using a model for the covariance matrix in the estimator of mean parameters and in constructing the coefficients in the predictor.