Protection against disclosure of survey respondents' identifiable and/or sensitive information is a prerequisite for statistical agencies that release microdata files from their sample surveys. Coarsening is one of popular methods for protecting the confidentiality of the data. Grouped data can be released in the form of microdata or tabular data. Instead of releasing the data in a tabular form only, having microdata available to the public with interval codes with their representative values greatly enhances the utility of the data. It allows the researchers to compute covariance between the variables and build statistical models or to run a variety of statistical tests on the data. It may be conjectured that the variance of the interval data is lower that of the ungrouped data in the sense that the coarsened data do not have the within interval variance. This conjecture will be investigated using the uniform and triangular distributions. Traditionally, midpoint is used to represent all the values in an interval. This approach implicitly assumes that the data is uniformly distributed within each interval. However, this assumption may not hold, especially in the last interval of the economic data. In this paper, we will use three distributional assumptions - uniform, Pareto and lognormal distribution - in the last interval and use either midpoint or median for other intervals for wage and food costs of the Statistics Korea's 2006 Household Income and Expenditure Survey(HIES) data and compare these approaches in terms of the first two moments.
According to the type of microdata, the various methods have been in use for masking microdata. Multiplicative noise is the one of popular schemes for masking continuous variables. In this paper, we introduce the method of masking based on multiplicative noise and show some results of the application on the 2006 Householder Income and Expenditure Survey(HIES) data. To create the multiplicative noise factor, we used the triangular distribution, truncated triangular distribution, trapezoidal distribution, and double triangular distribution. Also, formulas for the domain estimation for the data masked by the multiplicative noise are developed.
Poststratification is a common method of estimation in household surveys. Cells are formed based on characteristics that are known for all sample respondents and for which external control counts are available from a census or another source. The inverses of the poststratification adjustments are usually referred to as coverage ratios. Coverage of some demographic groups may be substantially below 100 percent, and poststratifying serves to correct for biases due to poor coverage. A standard procedure in poststratification is to collapse or combine cells when the sample sizes fall below some minimum or the weight adjustments are above some maximum. Collapsing can either increase or decrease the variance of an estimate but may simultaneously increase its bias. We study the effects on bias and variance of this type of dynamic cell collapsing theoretically and through simulation using a population based on the 2003 National Health Interview Survey. Two alternative estimators are also proposed that restrict the size of weight adjustments when cells are collapsed.
To reduce nonresponse bias in sample surveys, a method of nonresponse weighting adjustment is often used which consists of multiplying the sampling weight of the respondent by the inverse of the estimated response probability. The authors examine the asymptotic properties of this estimator. They prove that it is generally more efficient than an estimator which uses the true response probability, provided that the parameters which govern this probability are estimated by maximum likelihood. The authors discuss variance estimation methods that account for the effect of using the estimated response probability; they compare their performances in a small simulation study. They also discuss extensions to the regression estimator.
In poststratification, one of the cell collapsing criteria is a ratio criterion, where the ratio is the poststrafication factor or inverse coverage ratio. Quite often, if the ratio for a cell is greater than 2 or less than 1⁄2, then the cell is collapsed with another cell. However, this can introduce bias in a poorly covered group. Two censoring (or truncation) ratio approaches in collapsing were proposed and implemented in a simulation study (Kim, et al, 2005). The simulation study showed that the censoring approaches are better than the conventional approach mentioned above. In this paper, we propose four new collapsing strategies, two of which are based on the conditional bias and the other on the conditional mean square error.
The Behavioral Risk Factor Surveillance System (BRFSS) is a State telephone based survey of the civilian non-institutionalized adult (18 years and over) population residing in the United States. Consequently, the BRFSS final weights that are currently available in the data files are designed to produce unbiased estimates of socio-demographic and health characteristics for adults at the State level (Gonzalez, et al, 2005). In addition to State level BRFSS estimates, there is interest in the health status of adults residing in the 25 U.S. counties contiguous to the United StatesMexico Border region (Arizona, California, New Mexico, and Texas.) The purpose of this paper is to investigate alternative ways of arriving at poststratification factors (ratio adjustments) by collapsing the weighting matrix by age-sex-ethnicity/race for producing final weights/estimates for this border region. An optimal approach which minimizes local (cell) squared bias was applied to BRFSS data (25 contiguous counties). Then, a conditional mean square analysis was used to observe the effect of cell collapsing (in tandem with the optimal bias approach) on the absolute bias and variance estimators for several BRFSS socio-demographic and health characteristics.
A standard procedure in poststratification is to collapse or combine cells when the sample sizes fall below some minimum or the weight adjustments are above some maximum. Collapsing may decrease the variance of an estimate but may simultaneously increase its bias. We study the effects on bias and variance of this type of dynamic cell collapsing through simulation using a population based on the 2003 National Health Interview Survey.
The Behavioral Risk Factor Surveillance System (BRFSS) is a State telephone based survey of the civilian non-institutionalized adult (18 years and over) population residing in the United States. Consequently, the BRFSS final weights that are currently available in the data files are designed to produce unbiased estimates of sociodemographic and health characteristics for adults at the State level. In addition to State and national level BRFSS estimates, there is another geographical subpopulation of interest, that is, the border counties within the four United States-Mexico border States: Arizona, California, New Mexico, and Texas. The focus of this paper will be on the 44 counties which are within 100 kilometers (62 miles) of the border, as defined by the “Healthy Border 2010 Program,” United States-Mexico Border Health Commission. The purpose of this paper is to investigate alternative ways of arriving at poststratification factors (ratio adjustments) by collapsing rows or columns by age-sexethnicity/race for producing final weights/estimates for this border region. A conditional mean square analysis was used to observe the effect of cell collapsing on bias and variance estimators for several BRFSS sociodemographic and health characteristics.
Seppo Laaksonen Weighting for two-phase surveyed data . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 121 Pascal Ardilly and Pierre Lavallée Weighting in rotating samples: The SILC survey in France . . . . . . . . . . . . . . . 131 Jay J. Kim, Jianzhu Li and Richard Valliant Cell collapsing in poststratification . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 139 Fulvia Mecatti A single frame multiplicity estimator for multiple frame surveys . . . . . . . . . . . 151 David Haziza Variance estimation for a ratio in the presence of imputed data . . . . . . . . . . . . 159 James Chipperfield and John Preston Efficient bootstrap for business surveys . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 167 Jacob J. Oleson, Chong Z. He, Dongchu Sun and Steven L. Sheriff Bayesian estimation in small areas when the sampling design strata differ from the study domains . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 173 Enrico Fabrizi, Maria Rosaria Ferrante and Silvia Pacei Small area estimation of average household income based on unit level models for panel data . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 187 Anne Renaud Estimation of the coverage of the 2000 census of population in Switzerland: Methods and results. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 199