
The New Ways of Measuring (NWM) or ‘smart perceptions’ survey has been conducted in three countries (Italy, the Netherlands, and Slovenia) between September 2023 and February 2024. The aims of the survey were to find clues for improving nudge-to-smart fieldwork strategies, to inform legal officers, and to understand country differences in doing so. The NWM consisted of a two-step approach with a questionnaire on digital skills, hypothetical willingness and corresponding perceptions, and a ‘smart’ survey including four smart features. We find differences both between countries and between types of features, but also commonalities. We discuss the survey results and provide recommendations for future smart surveys.
Despite the potentials of data from mobile network operators (MNOs), there exist no official statistics today that are produced from aggregated and anonymized mobile phone signaling data, which in principle are accessible in every country. In the lack of MNO data, we develop and conduct proof-of-concept experiments in Norway and Italy using statistical registers and sample survey data. Viable methods and potential complications are explored and analyzed, such that we are now prepared to negotiate access to MNO data for building the production solutions. Both the estimation methodology and our experimental approach to non-survey big data can be useful to others with similar statistical interests, whether the data arise from mobile phones, transactions, or other sources.
This article presents a factorization of means of order r , which provides a method for numerical comparison among standard index numbers. We apply it to a family of superlative indices written by the quadratic mean of order r , elementary indices (the Carli, Dutot, and Jevons indices), and the Lowe, Young, and related indices. Descriptive measures are provided to assess the differences between index numbers. Empirical data are studied and discussed to evaluate the theoretical results of the proposed factorization.
To improve the timeliness of data products, survey researchers and practitioners have increasingly used web surveys and alike to collect information for population health research and dissemination. From the statistical inferential perspective, these data are often referred to as nonprobability samples due to the lack of a well-defined probability sampling structure, or they come from probability panel surveys yet are subject to high nonresponse and/or coverage errors. Certain statistical adjustments are therefore needed to make proper inferences using web surveys. With a high-quality reference probability survey available, one popular adjustment approach is to create pseudoweights that properly “weight” the web survey samples back to the target population underlying the reference survey in order to produce population-weighted estimates of the target of interest. When the variable of interest is collected in the web survey but not in the reference survey, the analytical question can also be framed as a missing data problem. Thus we propose to multiply impute the missing variable on the reference survey which combines the information from both data sources. We illustrate main features and performances of the multiple imputation strategy using a simulation study. The simulation results have shown that the multiple imputation analysis method is comparable to the pseudoweighting method, both of which can adequately account for the nonprobability feature of web surveys to improve the statistical inference. We also present a real data analysis based on the web-based Research and Development Survey and the interviewer-administered National Health Interview Survey. Our research shows that results from different imputation and pseudoweighting models can be compared to better understand features of the web survey data and improve the analysis.
The shift toward non-traditional data sources has made Mobile Positioning Data (MPD) a key tool for more flexible and scalable data collection. By utilizing SMS-to-web invitations from network providers, this method improves coverage, timeliness, and cost efficiency, though it still faces hurdles regarding response rates, respondent burden, and respondent trust. This study analyzes respondent feedback from the Digital Domestic Tourist Survey 2024 using sentiment analysis and topic modeling. The most frequently discussed aspects were survey application reliability (64.48%), web questionnaire presentation (63.53%), perceived ease of use (63.09%), and questionnaire content (59.80%). Positive sentiment dominated all aspects, accounting for at least 61% of responses, indicating an overall favorable perception. Topic modeling of survey question variables produced a coherence score of 0.4199. Higher coherence was observed for the positive sentiment dataset (0.5447) compared to the negative sentiment dataset (0.4869), suggesting a well-structured alignment with the study’s framework. Despite the generally positive evaluation, improvements are needed in application reliability, questionnaire design, incentive clarity, data collection methods, and communication strategies to reduce respondent burden, and enhance participation in MPD-based SMS-to-web surveys.
Proxy pattern-mixture models (PPMMs) have previously been proposed as a model-based framework for assessing the potential for nonignorable nonresponse in sample surveys and nonignorable selection in nonprobability samples. One defining feature of the PPMM is the single sensitivity parameter, , that ranges from 0 to 1 and governs the degree of departure from ignorability. While this sensitivity parameter is attractive in its simplicity, it may also be of interest to describe departures from ignorability in terms of how the odds of response (or selection) depend on the outcome being measured. In this paper, we re-express the PPMM as a selection model in order to better understand the underlying assumptions of the PPMM and the implied effect of the outcome on nonresponse (or selection). The selection model that corresponds to the PPMM is a quadratic function of the survey outcome and proxy variable, and the magnitude of the effect depends on the value of the sensitivity parameter, (missingness/selection mechanism), the differences in the proxy means and standard deviations for the respondent and nonrespondent populations, and the strength of the proxy as measured by the correlation between the outcome and the proxy in the respondent/selected sample. Large values of (beyond ) may result in unrealistic selection mechanisms, and the corresponding selection model can be used to establish more realistic bounds on . We illustrate the results using a home pricing dataset extracted from the China Family Panel Studies.
Artificial intelligence is entering official statistics through automated coding, imputation, nowcasting from alternative data sources, and small area estimation. This article examines AI adoption through the lens of methodological transitions in survey sampling—from design-based inference through model-assisted estimation—and addresses three questions. First, what is the relationship between algorithm-assisted and model-assisted inference? We show that generalized difference estimators can incorporate machine learning predictions in the same way they incorporate parametric working models. Second, what quality framework extensions are needed for operational deployment? Drawing on vignettes from European and North American statistical offices, we identify five areas requiring development: training data documentation, algorithmic transparency, validation protocols, uncertainty characterization, and reproducibility. We argue for prequential evaluation—assessing calibration and stability over successive production cycles—as an operational practice suited to algorithm-assisted systems. Third, what institutional challenges distinguish AI from earlier transitions? We examine the public-private asymmetry in AI development: whereas twentieth-century methodological innovations emerged largely from public institutions, contemporary AI capabilities are concentrated in private technology companies. European Statistical System initiatives illustrate governance responses, but dependency risks persist. We conclude that algorithm-assisted inference can succeed, but only if it can be made auditable, reproducible, and publicly defensible.
It is often useful to decompose an index number into the contribution of each product toward the total index, and consequently there are several well-known decompositions for bilateral indexes. In this note, I extend these decompositions to cases where bilateral indexes are made into multilateral GEKS indexes. Although the result is primarily of theoretical interest, it shows how decompositions based on a bilateral index can be extended to a multilateral index, and highlights the challenge of decomposing GEKS indexes.
National statistical agencies increasingly face budget constraints and shrinking sample sizes, while simultaneously gaining access to rich auxiliary data and powerful pre-trained machine learning (ML) and artificial intelligence (AI) models, including Large Language Models (LLMs). Traditional model-assisted estimation techniques, which fit models using survey sample data, are limited by small sample sizes, struggle to leverage complex non-linear relationships in auxiliary data, and cannot accommodate frontier pre-trained models. This work re-examines the use of pre-trained black-box models, fit independently of the survey sample, for design-based parameter estimation. Inspired by the Prediction-Powered Inference (PPI) framework, we introduce the Prediction-Powered Estimator (PPE), an unbiased estimator with an unbiased variance estimator for the survey design setting. We also formalize the use of pre-trained models with the classic difference estimator—which we term the Prediction-Powered Difference (PPD) estimator—and with the Generalized Regression Estimator via predicted values as covariates ( GREG y ^ ). Through LLM-based use-cases leveraging unstructured auxiliary data (images and text) and experiments with real-world survey data from Statistics Canada, complemented by simulation studies in the Supplemental Material , we demonstrate that these approaches consistently outperform standard baseline estimators across bias, mean absolute error, mean squared error, coverage, and confidence interval width. The results suggest that pre-trained models can yield more accurate and efficient estimates while potentially reducing survey sample sizes and respondent burden, and motivate expanding the survey methodologist’s toolbox to include pre-trained models and novel auxiliary data sources.
Respondent burden is complex and represents more than just the time spent completing the survey. With this research, we highlight the importance of looking beyond objective measures to understand the respondent’s perception of the survey-taking experience. First, we asked participants to tell us in their own words about the term “burden” and how surveys and survey questions can be burdensome. We identified themes in how respondents think about surveys that can inform targeted approaches to reducing respondent burden. We then tested the idea that perceptions of burden may not adhere to any objective measure of what it means for a survey to be burdensome. Through survey instructions, we presented different frames of reference to dissociate perceptions of survey length from actual survey length and analyzed the effect on ratings of burden. Our research suggests that factors like repetition, disorganization, and perceptions of pointlessness are key to respondents’ understanding of burden. Different frames of reference translated to significant differences in both perceptions of survey length and perceptions of burden, regardless of actual survey length. To improve respondents’ survey-taking experience, survey designers must go beyond survey length to consider perceived burden.
Geographical rezoning is a combinatorial optimization problem of assigning basic spatial units into regions for statistical analysis and the publication of official government statistics, such as the Census. We propose a novel approach by reframing this problem into a single-player, deterministic, fully-informative, finite combinatorial game; or simply, a single-player puzzle. This definition allows effective and systematic exploration of state-space, and offers perfect information guarantees, as well as decouples the problem definition from a specific solver algorithm. We take an existing combinatorial algorithm called HeLP (Hierarchical Land Parcel Aggregation), and combine its rezoning heuristic with a game-playing algorithm, Monte Carlo tree search, to create three new game-playing solvers for the geographical rezoning problem. A case study was conducted on a real-world data set in Canberra, Australia, on the Australian Statistical Geography Standard (ASGS). Our Monte Carlo implementations, HeLP-MCTS, HeLP-RAVE, and HeLP-RHEA, were tested on simulated data and have yielded a 50.25%, 28.77%, and 57.02% increase in the mean value of the partitioning heuristic function, respectively, compared to a combinatorial solver HeLP. We have also computationally improved the original HeLP algorithm, offering significant speed increases.
Introducing a web option in interviewer-administered surveys could increase response rates and reduce costs. However, this requires careful assessment of the effects of mixed-mode designs on data quality and key measures, especially across important sociodemographic subgroups. In 2018 and 2020, the Health and Retirement Study (HRS) experimentally introduced web in a sequential mixed-mode design for panelists assigned to the telephone mode. Initial analyses found a limited number of mode effects on key outcome distributions and data quality measures. This paper extends this initial analysis by assessing possible heterogeneity in these effects among sociodemographic subgroups defined by race/ethnicity, sex, and others. We interact mode with each sociodemographic indicator in statistical models for each outcome. Overall, we found limited evidence of heterogeneity in the mode effects, with 3% of the 204 interaction terms we tested emerging as significant. For example, previous work showed that more household roster changes are reported in the web-first group, and we found that this was more pronounced for females and those with some college education. Although some heterogeneity in mode effects was observed across subgroups, the effects were generally too small to cause data quality concerns. We conclude with a discussion of broader considerations for survey researchers.
Occupation coding encodes job titles into standard occupation labels, which is effective for data processing but tedious. Research proves that classic machine learning is effective, but accuracy needs further improvement. We construct a real data set with 881 occupation categories, including 41,297 pairs of job titles and corresponding labels. We design a hierarchy-aware heterogeneous graph neural network, combining prior knowledge from occupation category trees and synonyms. Results show our model outperforms other methods by 7.62% on micro-F1. It also alleviates the dependence on data as it achieves 52.28% on micro-F1 with only 30% of the original training data set.
Low response rates due to unit nonresponse have always been a ubiquitous problem in survey-based empirical research, and calibration is a popular method to adjust for bias caused by unit nonresponse. Typically, some external information on the true population quantities of margins for some calibration variables is available, and sometimes also of higher-order interactions. Weighting algorithms try to adjust the sample to these external benchmarks. It is generally assumed that even if the underlying missingness mechanism of the unit nonresponse is non-ignorable, weighting will at least alleviate the severity of the bias. We discuss data situations where weighting under a missing at random (MAR) assumption adjusts the sample correctly but still increases the bias for the analysis model, and we describe strategies for identifying auxiliary variables that are less susceptible to these unwanted effects.
In this note, we present a plausible structural mechanism by which over-parameterized deep learning models trained on real data may produce pseudo-synthetic data that constitute merely a different representation (or re-encoding) of the training data. We conjecture that, in principle, similar mechanisms may be learned by large-scale AI models even if they are not intentionally designed to do so. From there, we derive some cautionary warnings for potential adopters of pseudo-synthetic data generation tools based on deep learning. We claim that the burden of proof that no data re-encoding mechanism is at play in AI-based generation models rests with their proponents.
Estimating the risk of re-identification probabilistically is well-developed for the case of a random representative sample drawn from the general population, such as large-scale government surveys conducted regularly at National Statistical Institutes. Recent work extended this procedure to assess the risk of re-identification in non-probability subpopulation registers such as a cancer register. In this paper, we extend this work further to the case of samples drawn from registers or more generally to non-probability samples, such as those used in opt-in panels at survey organizations. The assumption is that membership to the subpopulation register is not known and the sampling mechanism is also unknown. We show how to assess the risk of re-identification for these types of non-probability samples using a probability-based reference sample to infer population parameters under the probabilistic modeling framework. We demonstrate with a simulation study and a real application on the 2021 Survey of Doctoral Recipients drawn from a subpopulation register of all PhD recipients from an accredited US institution.
The total poverty gap provides a straightforward information to policymakers as it measures the amount of income needed to get poor people out of poverty. Monitoring the change in total poverty gap is useful to examine poverty dynamics, however such a measure would be more informative if it were linked to other poverty indicators in a unified analysis framework. This paper suggests a decomposition that links the change in total poverty gap to those in poverty incidence, poverty depth, population size and composition. The decomposition is used to analyze the change in total poverty gap in Italy between 2010 and 2020. The total poverty gap decreased in the period considered and the change in poverty incidence was the main driver of such a reduction.
The paper sketches a proper statistical setting necessary to define the sampling design for Small Area Estimation (SAE). Since SAE techniques are commonly used in official statistics, relying on appropriate sampling designs to improve the quality of estimates becomes crucial. The sampling design is based on both allocation and sampling selection. The allocation step solves a non-standard problem necessary for finding the minimum-cost solution that controls the accuracy of the model-based small area estimator. The sampling selection ensures the planned sample sizes for each level of random effects affecting the variables of interest.
Price index decomposition allows the National Statistical Institute to break down aggregated price movements into contributions of individual commodities or product groups. In particular, decomposing multilateral indices, which combine many different time comparisons within the time window, can be useful in understanding and interpreting them by common users. This paper discusses multiplicative decompositions of multilateral GEKS-type indices, among which, currently, the GEKS-F and GEKS-T formulas are the most widely used. The paper also provides a multiplicative decomposition of the known GEKS-W index, as well as a decomposition of the GEKS-L, GEKS-GL, and GEKS-LM multilateral indices recently proposed in the literature. The paper also proposes normalized multiplicative decompositions of multilateral indices along with a relative commodity impact measure to enable comparisons of commodity contributions between products and across indices. The effects of these decompositions are demonstrated on two real scanner data sets from the food product segment.