
National Statistical Institutes (NSIs) are directing resources into advancing the use of administrative data in official statistics. Administrative data, however, are not developed for the purpose of producing statistics rather as a result of an event or transaction relating to administrative procedures of organizations, public administrations and government agencies. Therefore, it is essential to check the quality of the administrative data with respect to sources of error, particularly representativeness to the target population. In this paper, we utilize the strength of probability-based reference samples or censuses that can be used to detect the lack of representativeness in administrative data and introduce quality indicators based on distance metrics and representativity indicators (R-indicators). We demonstrate their application with a simulation study and discuss a real application applied on a UK Office for National Statistics (ONS) administrative dataset.
The observed best prediction (OBP) under a nested-error regression (NER) model was previously proposed using a design-based mean squared prediction error (MSPE) as a tool to derive the best predictive estimator (BPE). A recent study showed the OBP under the NER model may suffer from numerical instability when computing the BPE. We propose several modifications of the OBP under the NER model, including ones using a model-based MSPE to derive the BPE, to improve the numerical stability and predictive performance. We compare the performance of the modified OBP strategies with the existing methods in a simulation study. A real-data example is discussed.
The class of generalized linear models (GLM) is a flexible generalization of ordinary least squares regression that allows the linear model to be related to the response variable via a link function and assumes the magnitude of the variance of each measurement to be a function of its predicted value. Multicollinearity in GLMs can inflate variances of the estimated coefficients and cause poor prediction in certain regions of the regression space. It may also cause a nonsignificant Wald statistic even when the predictors are highly predictive in a model of the family of GLMs. Little previous research has closely investigated the diagnostics of multicollinearity in GLMs, especially when complex survey data are used. In this paper, we develop variance inflation factors (VIFs) that measure the amount that the variance of a parameter estimator is increased due to multicollinearity in GLMs. We also extend VIFs and condition indexes to apply to complex survey data, accounting for design features, e.g. weights, clusters, and strata. Illustrations of these methods are given using data from a household survey of health and nutrition.
In this paper a model-based inference procedure based on a multivariate structural time series model is developed for the production of monthly figures about consumer confidence. The input for the model are five series of direct estimates for the indices that measure consumer confidence, which are derived from the Dutch Consumer Survey. The model improves the accuracy of the direct estimates, since it provides a better separation of measurement errors and sampling errors from estimated target parameters. The standard errors for the month-to-month changes are clearly smaller under the time series model. A second problem addressed in this paper is related to the transition to a new survey process in 2017. Structural time series models in combination with a parallel run are applied to estimate discontinuities induced by the redesign. An algorithm designed for the consumer confidence variables is developed to construct uninterrupted input series for the aforementioned structural time series model. This inference method facilitated a smooth transition to a new survey design and resulted in uninterrupted series about consumer confidence that date back to 1986. The method is implemented for the production of official monthly figures on consumer confidence in the Netherlands.
This article examines the methodological complexities associated with the design of business surveys, with particular emphasis on sampling strategies implemented by National Statistical Offices (NSOs). It addresses the inherent challenges posed by the dynamic nature of the business population, which necessitates continual updates to the sampling frame to ensure representativeness and relevance. Critical design considerations include the determination of optimal sample sizes, stratification across key dimensions such as industry, geographic region, and enterprise size, as well as the treatment of business births and the exclusion of inactive (or "dead") units. The article applies Bankier's (1988) power allocation method to a two-way stratification scheme defined by industry and geography, evaluating its performance by comparing the resulting coefficients of variation with those obtained via a raking algorithm applied to the marginal coefficients. Furthermore, the approach is extended to a multivariate context to accommodate multiple estimation domains. The discussion also encompasses practical issues related to sample rotation and coordination, which are critical for maintaining data quality and minimizing respondent burden over time.
This study examines interviewer effects on household nonresponse in three waves of the Household Finance and Consumption Survey (HFCS) in Austria using a multilevel model. Addressing nonresponse at its source is crucial for maintaining survey data quality and representativeness. Our findings indicate that the variation in response behavior explained by interviewer effects decreased from about one-third in the first wave to 7% in the third wave. Effective interviewers tend to have a university degree, be married, homeowners, and have a larger workload. Additionally, higher mean wages in the household's municipality negatively affect survey participation. These insights suggest targeted interviewer selection and training strategies to improve response rates.
We propose an approximate hierarchical Bayes approach that uses the Natural Exponential Family with Quadratic Variance Function (NEF-QVF) in combining information from multiple sources to improve traditional survey estimates of finite population means for small areas. Unlike other Bayesian approaches in finite population sampling, we do not assume a model for all units of the finite population and do not require linking sampled units to the finite population frame. We assume a model only for the finite population units in which the outcome variable is observed; because, for these units, the assumed model can be checked using existing statistical tools. We do not posit an elaborate model on the true means for unobserved units. Instead, we assume that population means of cells with the same combination of factor levels are identical across small areas, and that the population mean for a cell is identical to the mean of the observed units in that cell. We apply our proposed methodology to a real-life survey, linking information from multiple disparate data sources. We also provide practical ways of model selection that can be applied to a wider class of models under similar setting but for a diverse range of scientific problems.
Although probability samples have been regarded as the gold standard to collect information for population-based study, non-probability samples have been used frequently in practice due to low cost, convenience, and the lack of the sampling frame for the survey. Na & iuml;ve estimates based on non-probability samples without any adjustments may be misleading due to selection bias. Recently, a valid data integration approach that includes mass imputation, propensity score weighting, and calibration has been used to improve the representativeness of non-probability samples. The effectiveness of the mass imputation approach depends on the underlying model assumptions. In this paper, we propose using deep learning for the mass imputation in the combining of probability and non-probability samples and compare it with several modern machine learning-based mass imputation approaches, including generalized additive modeling, regression tree, random forest, and XGboosting. In the simulation study, deep learning-based approaches have been shown to be more robust and effective than other mass imputation approaches against the failure of underlying model assumptions under non-linearity scenarios.
In this paper, we study the performance of hierarchical Bayes (HB) small area estimators using noninformative and informative priors. We apply the Bayesian models of You and Chapman (2006) and You (2021) to the Canadian Labor Force Survey (LFS) data and evaluate the impact of the priors on the HB estimators. A Bayesian model comparison and simulation study are also conducted. Our results indicate that a correct informative prior can lead to very good results, and noninformative priors can also perform very well. Incorrect informative priors can lead to poor results in terms of large bias and large coefficient of variation (CV). Noninformative priors are recommended in practice for HB small area estimation unless correctly specified informative priors are available. Informative priors are particularly useful when the number of small areas is relatively small.
Classical design-based survey estimation relies on a properly specified sampling design for valid inference. We consider the properties of regression estimation under a misspecified sample design, in which the nominal and true inclusion probabilities do not necessarily match. This general misspecified sample design setting encompasses many challenges in the modern survey environment. Under this setting, an asymptotic analysis of the regression estimator, an expression of the bias, and an expression of the variance are presented. Further, a consistent variance estimator is derived and an expression which estimates the bias in-part or in-whole is discussed. This later expression may be used as an indicator of the presence of bias due to misspecification by a practitioner. A simulation study is conducted to support the presented theory.
Survey practitioners have increasingly embraced the benefits of modern machine learning techniques, including classification and regression tree algorithms, in the development of nonresponse adjustments. These methods, which do not require a predefined functional relationship between outcomes and predictors, offer a practical means of conducting variable selection and deriving interpretable structures that link response propensity with explanatory variables. However, when applying these algorithms to survey data, it is common to overlook crucial factors like sampling weights, as well as sample design features such as stratification and clustering. To bridge this shortcoming, we propose an extension of the Chi-square Automatic Interaction Detector (CHAID) approach, and we describe the design-based asymptotic properties of the resulting "survey CHAID" (sCHAID) method. To facilitate the practical use of sCHAID, we incorporate a Rao-Scott correction into the splitting criterion, accounting for the survey design. Using data from the U.S. American Community Survey, we illustrate the use of the method and evaluate its performance through comparisons with existing weighted and unweighted algorithms.
We introduce a novel approach to model-assisted calibration estimation in survey sampling using generalized entropy. The method builds upon recent work by Kwon, Kim and Qiu (2024) and extends it to a model-assisted framework. Unlike traditional calibration techniques, this approach employs a generalized entropy function as the objective for optimization and incorporates a debiasing calibration constraint to ensure design consistency. The proposed estimator is shown to be asymptotically equivalent to an augmented generalized regression (GREG) estimator. It allows for unequal model variance, potentially improving efficiency when the sampling design is informative. The paper presents both design-based and model-based justifications for the method, along with asymptotic properties and variance estimation techniques. Computational aspects are discussed, including an unconstrained optimization approach that facilitates implementation, especially for high-dimensional auxiliary variables. The method's performance is evaluated through a simulation study, demonstrating its effectiveness in improving estimation efficiency, particularly when the sampling design is informative.
In recent years, there has been a significant interest in machine learning in national statistical offices. Thanks to their flexibility, these methods may prove useful at the nonresponse treatment stage. In this article, we conduct an empirical investigation in order to compare several machine learning procedures in terms of bias and efficiency. In addition to the classical machine learning procedures, we assess the performance of ensemble approaches that make use of different machine learning procedures to produce a set of weights adjusted for nonresponse.
BigData users and the BigData research community are expanding rapidly, while statisticians at large are seemingly becoming divided between those who are enthusiastic and those who are concerned, if not downright hostile. Is BigData also a big step ahead, truly advancing our ability to extract meaningful information and actual knowledge from data? Is BigData underplaying traditional statistical inference as we know it, supplanting survey methodology as a low-cost futuristic option? In this paper I will attempt to unravel the multifaceted relationship bridging BigData to sampling methodology. Starting by reasoning why it should be interesting to look at BigData from a sampling statistician's perspective, I will delve deeper into the somewhat ambiguous definition of BigData and share some very personal considerations and views on the matter. In the process, several open questions will arise while discussing a personal selection of insights that are traceable through the vast body of statistical literature around BigData and sampling methodology. The discussion will take various angles explored across nine key points, and it will conclude with a forward-looking perspective on a main challenge for future research: addressing the strong assumptions needed to manage deviations from purely randomized data collection.
Rao (1999) summarized trends in sample survey theory and methods at the turn of the millenium. We provide an updated discussion of some current trends in survey design and estimation methods for the 50th anniversary of Survey Methodology. Recent innovations in survey design include research on anticipating nonsampling errors at the design stage and development of balanced and adaptive sampling designs to take advantage of detailed sampling frame information or data gathered during the survey process. Nonparametric and machine learning methods are increasingly used for data editing as well as for model-assisted estimation and nonresponse adjustments. Small area models have been expanded to incorporate spatial and time series information, increase the flexibility and robustness of the linking and variance models, benchmark to large-area direct estimators, and (for unit level models) account for informative sampling designs. The increasing availability of large administrative datasets, sensor and satellite data, and convenience samples has spurred research on how to use these sources. on their own and when integrated with probability samples. We conclude by discussing some frontiers for survey research.
This article confronts survey science with important notions in philosophy of science: progress, paradigm, research tradition, research programmes. The article is conceptual and exploratory, rather than mathematical/technical. This is against a background where survey science must evolve in unfamiliar and challenging conditions. Society is changing. Survey nonresponse is high. Probability sampling surveys are in question, considered too expensive. Low cost alternative data sources - big data and others - must, in the opinion of some, be incorporated in statistics production at the national statistical offices. A lively research tradition has brought progress in survey science over more than one hundred years. The article recalls some of that progress and tries to foresee how the tradition may survive and face the coming decades.
Tightened budgets, continuing decrease of response rates in traditional probability surveys and increasing pressure by users for more timely data, has stimulated research on the use of nonprobability sample data, such as administrative records, web scraping, mobile phone data and voluntary internet surveys, for inference on finite population parameters like means and totals. These data are often easier, faster and cheaper to collect than traditional probability samples. However, a major concern with the use of this kind of data for official statistics is their nonrepresentativeness due to possible selection bias, which if not accounted for properly, could bias the inference. In this article, we review and discuss methods considered in the literature to deal with this problem and propose new methods, distinguishing between methods based on integration of the nonprobability sample with an appropriate probability sample, and methods that base the inference solely on the nonprobability sample. Empirical illustrations, based on simulated data are provided.
In this paper, we derive a second-order unbiased (or nearly unbiased) mean squared prediction error (MSPE) estimator of the empirical best linear unbiased predictor (EBLUP) of a small area mean for a semi-parametric extension to the well-known Fay-Herriot model. Specifically, we derive our MSPE estimator essentially assuming certain moment conditions on both the sampling errors and random effects distributions. The normality-based Prasad-Rao MSPE estimator has a surprising robustness property in that it remains second-order unbiased under the non-normality of random effects when a simple Prasad-Rao method-of-moments estimator is used for the variance component and the sampling error distribution is normal. We show that the normality-based MSPE estimator is no longer second-order unbiased when the sampling error distribution has non-zero kurtosis or when the Fay-Herriot moment method is used to estimate the variance component, even when the sampling error distribution is normal. Interestingly, when the simple method-of moments estimator is used for the variance component, our proposed MSPE estimator does not require the estimation of kurtosis of the random effects. Results of a simulation study on the accuracy of the proposed MSPE estimator, under non-normality of both sampling and random effects distributions, are also presented.
We present and apply methodology to improve inference for small area parameters by using data from several sources. This work extends Cahoy and Sedransk (2023) who showed how to integrate summary statistics from several sources. Our methodology uses hierarchical global-local prior distributions to make inferences for the proportion of individuals in Florida's counties who do not have health insurance. Results from an extensive simulation study show that this methodology will provide improved inference by using several data sources. Among the five model variants evaluated the ones using horseshoe priors for all variances have better performance than the ones using lasso priors for the local variances.
Survey data collection often is plagued by unit and item nonresponse. To reduce reliance on strong assumptions about the missingness mechanisms, statisticians can use information about population marginal distributions known, for example, from censuses or administrative databases. One approach that does so is the Missing Data with Auxiliary Margins, or MD-AM, framework, which uses multiple imputation for both unit and item nonresponse so that survey-weighted estimates accord with the known marginal distributions. However, this framework relies on specifying and estimating a joint distribution for the survey data and nonresponse indicators, which can be computationally and practically daunting in data with many variables of mixed types. We propose two adaptations to the MD-AM framework to simplify the imputation task. First, rather than specifying a joint model for unit respondents' data, we use random hot deck imputation while still leveraging the known marginal distributions. Second, instead of sampling from conditional distributions implied by the joint model for the missing data due to item nonresponse, we apply multiple imputation by chained equations for item nonresponse before imputation for unit nonresponse. Using simulation studies with nonignorable missingness mechanisms, we demonstrate that the proposed approach can provide more accurate point and interval estimates than models that do not leverage the auxiliary information. We illustrate the approach using data on voter turnout from the U.S. Current Population Survey.