When inferring population characteristics from a nonprobability sample, it is crucial to correct the possible selection bias therein by, for example, pseudo-weighting. Many correction methods focus on estimating the population means of the target variable. However, often the quantities of subpopulations are also of interest. It is unclear whether pseudo-weights are suitable for domain estimation, since the weights unavoidably introduce variation and possibly even bias in the downstream estimation.To address this issue, modeling on the domain level may be an option. We evaluate two promising domain estimation methods on weighted nonprobability samples. The first one is iterative proportional fitting (IPF), where the margins are considered in the domain estimation, so that the marginal values may be fixed when improving the domain estimates. The other is a hierarchical Bayesian model, in which the pseudo-weights are included in the domain modeling process. This approach enjoys the flexibility of modeling when different types of information are available. We evaluate a range of modeling options for the two methods, and compare them in a simulation study. We also evaluate the methods with resampled real data sets to mimic the scenario where the relation between variables and the inclusion mechanism of the nonprobability samples are unknown to the researchers.We found that applying IPF to the unweighted table and the hierarchical Bayesian model improves the domain estimation in most cases. If both marginal and domain estimates are of interest, the estimated overall population total or mean should be considered in the domain modeling process.
Selection bias correction is generally crucial to have a reliable inference for a nonprobability sample. Often the correction is applied with the two-sample setup. That is, along with the nonprobability sample we are interested in, a probability sample sharing some common auxiliary variables is used for constructing correction weights for the nonprobability sample. The two-sample setup allows one to calculate weighted estimates for population parameters of interest based on the nonprobability sample. Since the nonprobability sample is usually easy to collect, we often end up with a large nonprobability sample and a small probability sample. The imbalance of the two samples may cause difficulties in modelling the propensity of units to be included in the nonprobability sample. This paper discusses some often-seen solutions for imbalanced samples in machine learning literature, i.e. undersampling, Synthetic Minority Oversampling Technique (SMOTE), and a mixture of both. A selection bias correction framework is adjusted to incorporate the imbalance solutions. Three evaluation studies across different types of data sets are shown. The results indicate that SMOTE has the potential to deal with imbalances in selection bias correction, while further study is needed to identify in which scenarios SMOTE works well.
Data linkage is often used to make inferences from distinct data sources. The data may be linked by a set of variables with high discrimination, for example, a social security number or some personal information. A good linking variable often implies a high disclosure risk in identifying the corresponding person. To prevent the disclosure risk, the data linker may choose to remove the linking variables before publishing the linked datasets. The design of the original data sources, the response pattern, and the unlinked part of the original data are often also unknown to the secondary user. However, the quality of the linked data is constrained by the original data sources and the linking process. Without considering the potential error or bias in linked datasets, naïvely treating them as error-free in secondary analysis may result in biased inference. To indicate the quality of the linked dataset without sacrificing privacy, we propose publishing correction weights alongside the linked datasets. The weights are generated given the information in the original data sources and the quality of the linkage. Both selection issues of the sample and measurement issues of the linking variables are addressed in the constructed weights, and we allow the possibility of having multiple potential links for a record. Secondary users may apply design-based estimators for subsequent analyses based on the correction weights, or apply sensitivity analysis given different sample inclusion criteria. An example is presented, and the option of secondary analysis given the constructed weights is discussed.
When estimating a population parameter by a nonprobability sample, that is, a sample without a known sampling mechanism, the estimate may suffer from sample selection bias. To correct selection bias, one of the often-used methods is assigning a set of unit weights to the nonprobability sample, and estimating the target parameter by a weighted sum. Such weights are often obtained with classification methods. However, a tailor-made framework to evaluate the quality of the assigned weights is missing in the literature, and the evaluation framework for prediction may not be suitable for population parameter estimation by weighting. We try to fill in the gap by discussing several promising performance measures, which are inspired by classical calibration and measures of selection bias. In this paper, we assume that the population parameter of interest is the population mean of a target variable. A simulation study and real data examples show that some performance measures have a strong positive relationship with the mean squared error and/or error of the estimated population mean. These performance measures may be helpful for model selection when constructing weights by logistic regression or machine learning algorithms.
Probability surveys are experiencing important drawbacks nowadays: costs are relatively high and participation rates are decreasing, which could yield less accurate estimates. Alternatively, nonprobability samples like administrative records are having a rise in popularity due to their convenience and low costs. Unfortunately, nonprobability samples are often selective and, as the underlying sampling design is unknown, estimators based on such samples are generally biased. Research is ongoing on how to deal with this selection bias. In this paper, a method is proposed that combines estimators from a probability and nonprobability sample on an aggregated level. Our estimator is constructed as a weighted mean of both estimators. The weight is chosen to minimize the expected value of the mean squared error (MSE) of the combined estimator under an assumed model for the bias in the estimator based on the nonprobability sample. Our method does not require any data on the level of the individual units in the samples. We performed simulation studies where two different methods of modeling the bias in the nonprobability sample were tested. We also applied one of these methods to a real dataset from Statistics Netherlands and showed that the MSE was indeed reduced in a real application.
This paper proposes analytical variance estimation formulas for the estimated population mean from an extended pseudo weighting method developed by [1] (LSdW). LSdW is meant to correct selection bias in a nonprobability sample, also when the nonprobability sample or the reference probability sample has a large inclusion fraction. Since samples with large inclusion fractions often require massive computation resources, having an analytical expression for the variance will be more time-efficient compared to resampling methods. In addition, we show that LSdW is a consistent estimator of the population mean under certain assumptions. To deal with different designs of the probability sample, probability proportional to size (PPS) sampling and simple random sampling (SRS) are considered, and the variance estimator formulas are given accordingly. The proposed formulas are evaluated by a simulation study and it shows that the proposed formulas give reasonable estimates in terms of relative bias and coverage of the confidence interval.
Estimation of output quality based on sample surveys is well established.It accounts for the effects of sampling and non-response errors on the accuracy of an estimator.When administrative data are used or combinations of administrative data with survey data, more error types need to be taken into account.Moreover, estimators in multisource statistics can be based on different ways of combining data sources.That partly affects the methodology that is needed to estimate output quality.This paper presents results of the ESSnet project Quality of Multisource Statistics that studied methods to estimate output quality.We distinguish three main groups of methods: scoring methods, (re)sampling methods and methods based on parametric modeling.Each of those is split into methods that can be used for both single and multisource statistics and methods that can be applied to multisource statistics only.We end the paper by discussing some of the main challenges for the near future.We argue that estimating output quality for multisource statistics is still more an art than a technique.
Within the field of official statistics, we are often interested in the relative change of statistical indicators through time. A typical example is economic growth measured as the ratio of gross domestic product at two consecutive time periods. In this article, we investigate the accuracy of such ratios when the numerator and denominator describe a subpopulation. Subpopulations are always identified by a classification variable, which might contain classification errors. Our main objective is to quantify the bias and variance of statistical estimates of ratios of statistical indicators that are affected by classification errors. Previous studies have already shown how to estimate the bias and variance of statistical indicators for a single subpopulation as affected by classification errors. For estimated ratios, such results are not yet available in the academic literature. In this article, we will provide those results. More specifically, we will investigate three situations: the classification errors for the numerator and denominator are (1) completely dependent, (2) completely independent, or (3) partially dependent. For these three cases, we derive estimators of the bias and variance using a Taylor expansion and a bootstrap approach. By means of a simulation study, we show that the Taylor expansion is valid unless the distribution of the target variable of the numerator or the denominator is highly skewed. The methods are applied to two widely differing case studies: one on classification errors made by a machine learning classifier (independent errors) and one on classification errors in register data (partially dependent errors). This illustrates that our method can be useful for a wide range of applications.
Being able to quantify the accuracy (bias, variance) of published output is crucial in official statistics. Output in official statistics is nearly always divided into subpopulations according to some classification variable, such as mean income by categories of educational level. Such output is also referred to as domain statistics. In the current paper, we limit ourselves to binary classification variables. In practice, misclassifications occur and these contribute to the bias and variance of domain statistics. Existing analytical and numerical methods to estimate this effect have two disadvantages. The first disadvantage is that they require that the misclassification probabilities are known beforehand and the second is that the bias and variance estimates are biased themselves. In the current paper we present a new method, a Gaussian mixture model estimated by an ExpectationMaximisation (EM) algorithm combined with a bootstrap, referred to as the EM bootstrap method. This new method does not require that the misclassification probabilities are known beforehand, although it is more efficient when a small audit sample is used that yields a starting value for the misclassification probabilities in the EM algorithm. We compared the performance of the new method with currently available numerical methods: the bootstrap method and the SIMEX method. Previous research has shown that for non-linear parameters the bootstrap outperforms the analytical expressions. For nearly all conditions tested, the bias and variance estimates that are obtained by the EM bootstrap method are closer to their true values than those obtained by the bootstrap and SIMEX methods. We end this paper by discussing the results and possible future extensions of the method.
Nonprobability samples, for example observational studies, online opt-in surveys, or register data, do not come from a sampling design and therefore may suffer from selection bias. To correct for selection bias, Elliott and Valliant (EV) proposed a pseudo-weight estimation method that applies a two-sample setup for a probability sample and a nonprobability sample drawn from the same population, sharing some common auxiliary variables. By estimating the propensities of inclusion in the nonprobability sample given the two samples, we may correct the selection bias by (pseudo) design-based approaches. This paper expands the original method, allowing for large sampling fractions in either sample or for high expected overlap between selected units in each sample, conditions often present in administrative data sets and more frequently occurring with Big Data.
Record linkage aims to bring records together from two or more files that belong to the same statistical entity. Naïvely treating a linked file as if there are no linkage errors may lead to biased inference. We present two general approaches for compensating for linkage error when calculating and analysing a two-way contingency table for categorical data, and study the following question: under what conditions can a compensation approach improve on the naïve approach, where linkage error is not compensated for? To this end, we compare estimation errors, bias, variance and mean square error for the naïve approach and two compensation approaches by means of an analytical study as well as a simulation study.
This article discusses methods for evaluating the variance of estimated frequency tables based on mass imputation. We consider a general set-up in which data may be available from both administrative sources and a sample survey. Mass imputation involves predicting the missing values of a target variable for the entire population. The motivating application for this article is the Dutch virtual population census, for which it has been proposed to use mass imputation to estimate tables involving educational attainment. We present a new analytical design-based variance estimator for a frequency table based on mass imputation. We also discuss a more general bootstrap method that can be used to estimate this variance. Both approaches are compared in a simulation study on artificial data and in an application to real data of the Dutch census of 2011.
When applying supervised machine learning algorithms to classification, the classical goal is to reconstruct the true labels as accurately as possible. However, if the predictions of an accurate algorithm are aggregated, for example by counting the predictions of a single class label, the result is often still statistically biased. Implementing machine learning algorithms in the context of official statistics is therefore impeded. The statistical bias that occurs when aggregating the predictions of a machine learning algorithm is referred to as misclassification bias. In this paper, we focus on reducing the misclassification bias of binary classification algorithms by employing five existing estimation techniques, or estimators. As reducing bias might increase variance, the estimators are evaluated by their mean squared error (MSE). For three of the estimators, we are the first to derive an expression for the MSE in finite samples, complementing the existing asymptotic results in the literature. The expressions are then used to compute decision boundaries numerically, indicating under which conditions each of the estimators is optimal, i.e., has the lowest MSE. Our main conclusion is that the calibration estimator performs best in most applications. Moreover, the calibration estimator is unbiased and it significantly reduces the MSE compared to that of the uncorrected aggregated predictions, supporting the use of machine learning in the context of official statistics.
Auditing is a widely used method for quality improvement, and many guidelines are available advising on how to draw samples for auditing. However, researchers or auditors sometimes find themselves in situations that are not straightforward and the standard sampling techniques are not sufficient, for example when a selective sample has initially been audited and the auditor desires to re-use as many cases as possible from this initial audit in a new audit sample that is representative with respect to some background characteristics. In this paper, we introduce a method that selects an audit sample that re-uses initially audited cases by considering the selection of a representative audit sample as a constrained minimization problem. In addition, we evaluate the performance of this method by means of a simulation study and we apply the method to draw an audit sample of establishments to evaluate the quality of an establishment registry used to produce statistics on energy consumption per type of economic activity.
1. In this paper, we discuss methods for evaluating the design-based variance of estimated frequency tables based on mass imputation. The motivating application for this study is the Dutch decennial virtual population census. Since 1981, the Dutch census tables have been estimated by re-using data from existing sources rather than collecting data with a dedicated questionnaire. Nowadays, most variables needed for the census are available from administrative sources with (near-)complete population coverage. An exception occurs for educational attainment, which is observed partly in education registers and partly in the Labour Force Survey (LFS). For about 7 million Dutch persons (of a total of 17 million), educational attainment is not observed.
Many National Statistical Institutes (NSIs), especially in Europe, are moving from single-source statistics to multi-source statistics. By combining data sources, NSIs can produce more detailed and more timely statistics and respond more quickly to events in society. By combining survey data with already available administrative data and Big Data, NSIs can save data collection and processing costs and reduce the burden on respondents. However, multi-source statistics come with new problems that need to be overcome before the resulting output quality is sufficiently high and before those statistics can be produced efficiently. What complicates the production of multi-source statistics is that they come in many different varieties as data sets can be combined in many different ways. Given the rapidly increasing importance of producing multi-source statistics in Official Statistics, there has been considerable research activity in this area over the last few years, and some frameworks have been developed for multi-source statistics. Useful as these frameworks are, they generally do not give guidelines to which method could be applied in a certain situation arising in practice. In this paper, we aim to fill that gap, structure the world of multi-source statistics and its problems and provide some guidance to suitable methods for these problems.
Statistical matching is a technique to combine variables in two or more nonoverlapping samples that are drawn from the same population. In the current study, the unobserved joint distribution between two target variables in nonoverlapping samples is estimated using a parametric model. A classical assumption to estimate this joint distribution is that the target variables are independent given the background variables observed in both samples. A problem with the use of this conditional independence assumption is that the estimated joint distribution may be severely biased when the assumption does not hold, which in general will be unacceptable for official statistics. Here, we explored to what extent the accuracy can be improved by the use of two types of auxiliary information: the use of a common administrative variable and the use of a small additional sample from a similar population. This additional sample is included by using the partial correlation of the target variables given the background variables or by using an EM algorithm. In total, four different approaches were compared to estimate the joint distribution of the target variables. Starting with empirical data, we show how the accuracy of the joint distribution is affected by the use of administrative data and by the size of the additional sample included via a partial correlation and through an EM algorithm. The study further shows how this accuracy depends on the strength of the relations among the target and auxiliary variables. We found that including a common administrative variable does not always improve the accuracy of the results. We further found that the EM algorithm nearly always yielded the most accurate results; this effect is largest when the explained variance of the separate target variables by the common background variables is not large.
The results of the project on the quality of the multisource statistics launched by the European Statistical System and accomplished by the National Statistical Institutes of eight European countries under the name KOMUSO are described. The work carried out consists of two main documents: the Quality Guidelines for Multisource Statistics supplemented with the collection of the quality measures for statistical output and examples to use them; and Quality Guidelines for Frames in Social Statistics with the list of quality measures and indicators followed by the examples to use them. An overview of the documents created in the project is presented in this paper.
The widely used formulas for the variance of the ratio estimator may lead to serious underestimates when the sample size is small; see Sukhatme (1954), Koop (1968), Rao (1969), and Cochran (1977, pages 163-164). In order to solve this classical problem, we propose in this paper new estimators for the variance and the mean square error of the ratio estimator that do not suffer from such a large negative bias. Similar estimation formulas can be derived for alternative ratio estimators as discussed in Tin (1965). We compare three mean square error estimators for the ratio estimator in a simulation study.
Robert Tijdeman合作论文数Mathematical Institute
Leiden University1