
The European Survey Research Association (ESRA) has issued a statement expressing concerns over the declining methodological quality of opinion polls and the implications this has on public discourse and democratic processes. Recent technological advances have facilitated the proliferation of online surveys, leading to increased media coverage of opinion polls without adequate scrutiny of their methodological rigor. ESRA highlights the improper use of terms like "representative" in defending poor-quality polls, which often distort the reality of public opinion. This statement is primarily directed at journalists, who play a crucial role in shaping public opinion by reporting survey results. ESRA offers practical recommendations for assessing the quality of opinion polls, emphasizing the importance of probability-based sampling, transparency in data collection modes, motivation behind polls, and meticulous questionnaire design. The statement also warns against non-probability sampling methods, which are susceptible to manipulations by automated bots and are not suitable for validly describing the public opinion of general populations. Furthermore, the ESRA urges careful consideration of margins of error and the "File Drawer" problem, where extreme or surprising results may overshadow more moderate but accurate findings. By promoting these standards, ESRA aims to support journalists and users of survey data in critically evaluating opinion polls and safeguarding democratic integrity.
Previous research on probability-based online panels shows that providing the offline-population with an alternative survey mode reduces coverage bias. However, little is known about whether offering an offline mode to internet users who are unwilling to participate in panel surveys online pays off in lower nonresponse bias. This study uses data from the GESIS Panel, a probability-based web and mail panel, to investigate how including internet users in the mail mode affects nonresponse bias of population and model estimates. The results show that internet users included in the mail mode differ from non-internet users assigned to the mail mode and panelists responding in the web mode in demographic variables and internet usage after the recruitment. However, excluding internet users in the mail mode from the data set rarely introduces a higher bias in estimates of demographic characteristics compared to data from a reference sample. Further analyses examine how mean and model estimates of three studies published with GESIS Panel data would have been affected by excluding internet users from the sample who are unwilling to provide survey data online. Here the findings show potential nonresponse bias in several estimates of means from the model variables. However, the model estimates of all three studies are largely the same after removing those cases from the analyses. Therefore, authors' conclusions would likely have remained unchanged if internet users had not been included in the mail mode.
Self-rated health (SRH) is a frequently used health measure in (cross-)national surveys. It is usually assessed with a single item, which differs in the wording and response format across surveys. In this paper, we compare four German-language five-point scale versions of self-rated health that vary in response scale labeling using web probing. The web survey (N=1,710) was conducted in 2019. Combining qualitative and quantitative methods, we assess how response scale labels affect response distributions of the SRH item and which health factors respondents consider when answering questions about their health. The main finding is that respondents refer to similar health aspects independent of the scale version they answered the question with. Self-reported health, however, varied across scales which might introduce a comparative bias. Using an unbalanced scale with three response options indicating good health led to a more positive self-assessment of health compared to the balanced scales with two positive, one neutral, and two negative scale points. We discuss practical implementations.
The main advantage of postal surveys is that they allow for a universal initial sampling. It has become standard practice to include an online answering option in those surveys, but a concern remains that not all groups of the population are equally inclined to use this option. In this note, we compare the answering options that were used in two election surveys in Belgium, that were conducted in 2019 and 2024. As a larger proportion of all respondents now uses the internet option, some elements of stratification like gender and education level have indeed weakened. Differences with regard to age however, remain strong as especially older age groups continue to prefer the pen and paper answering option. Our recommendation therefore remains that if one is interested in reaching a representative sample of especially older respondents, providing the pen and paper option does remain important.
This editorial introduces ESRA's statement on opinion polls and explains why the Association considers it necessary in the current media landscape. It also invites ESRA members and national survey research organisations to promote the statement through authorised translations and closer collaboration with journalists.
Solidarity within and between societies is a widely discussed topic these days. A multitude of global challenges makes the issue of social togetherness more relevant than ever before. But can we truly measure global solidarity in a comparative way? Or is our understanding shaped by the cultural, demographic, economic, and political contexts in which individuals find themselves? This study addresses these questions using data from the European Values Study 2017, which covers 21 European Union member states. By applying multigroup and multilevel structural equation modelling, we show that achieving cross-national measurement invariance of global solidarity is indeed difficult and that accounting for contextual and cultural differences can help explain why certain measurements lack equivalence. The analysis reveals that the measurements are subject to varying culturally driven interpretations about how people interpret survey items on global solidarity, particularly regarding concern for Europeans and immigrants. Varying cultural views on what is meant by ‘Europeans’, or how immigrants should be treated, and factors such as immigrant influx, cultural orientation, prioritising equal treatment, or cultural predominance affect how individuals express concern for immigrants. These differences in cultural and national contexts contribute to measurement noninvariance.
Incentives are routinely used to increase response rates to surveys. Recent research demonstrates the effectiveness of offering a second prepaid incentive in the follow-up reminders. To extend the incentive research, we embedded a 2 by 2 experiment in a web survey to investigate the impact of offering an electronic gift card (vs. cash) as a prepaid incentive and the impact of splitting an initial $5 incentive to a combination of an initial incentive of $2 and a second incentive of $3 on response rate, sample composition, nonresponse bias in selected survey variables, and cost per completed interview. We found that offering an initial incentive of $2 and a second incentive of $3 led to a higher response rate and a lower nonresponse bias in survey estimates than offering an initial incentive of $5 and no second incentive when the incentive was offered as cash but not when the incentive was an Amazon.com Gift Card. Furthermore, the condition offering an initial cash incentive of $2 and a second cash incentive of $5 resulted in the largest share of respondents aged 18 to 49 and the lowest cost per complete. Implications of the findings are discussed.
Reliability in survey measurement is a crucial aspect of data quality. Yet surprisingly little attention has been given to assessing reliability for commonly used survey measures or evaluating how reliability may vary across different survey contexts. This study leverages the Comparative Panel File (CPF), a dataset compiled from long-running household panel surveys in seven countries—Australia, Germany, Russia, South Korea, Switzerland, the United Kingdom, and the United States—to estimate the reliability of 14 survey indicators from 2001 to 2020 using quasi-simplex models. We find that reliability is high and consistent across countries for factual items but substantially lower and more variable for self-assessment items (e.g., satisfaction with life, satisfaction with work, self-rated health). Additionally, we find that variations in question wording across the panels (e.g., asking about health currently versus in general) led to different reliabilities. We discuss the implications these findings have for measurement comparability in cross-national research.
Responses to questionnaire items can be influenced by various factors including sample design, interview mode and/or how questions are phrased. To analyse these aspects, this paper draws on the Bank of Italy's surveys of households and firms, which employ different survey modes or questions with different phrasings, response options, or graphical features for sub-samples of respondents. We exploit the potential of CUB (Combination of Uniform Discrete and shifted Binomial random variables) modelling for the analysis of ordinal data. CUB models are able to capture and identify the different components of the cognitive process behind the responses and to study how these are related to the relevant covariates (such as respondents' characteristics). The results show that in general, although diverse survey modes and a different phrasing or graphical representation of questions may yield somewhat different findings in terms of uncertainty, responses to relevant questions such as those on reported satisfaction or expectations did not produce pronounced differences in data reliability.
In a survey experiment, we analyzed how different versions of a response scale affect the distributional characteristics and quality of the resulting survey data. Toward that end, we compared four different German-language versions of a five-point agree-disagree (AD) response scale, randomly assigned to four groups of online access panelists for a total of 15 items. The response scales were taken from different studies and varied in polarity or scale option intensity. Comparisons of frequency distributions as well as of response quality (response styles, response differentiation, response times) did not show any systematic differences between the response scales. Although there were no systematic differences in the overall sample, the few significant effects we found appeared to be largely due to the responses of participants with lower levels of education. Further research is warranted using non-access panel respondents and their perception of differently worded AD response scales, experimentally modified response scales, and other languages beyond the English-German pair.
Conducting face-to-face surveys is increasingly challenging. The evolving technological and social landscape of survey data collection is prompting researchers to seek more sustainable solutions through new data collection techniques and a rethinking of traditional survey methods. Recent examples include the Push-to-Web (PtW) survey experiments initiated by the European Social Survey (ESS) research consortium, which use probability sampling to reach respondents via postal invitation and combine web- and paper-based self-completion questionnaires. The literature indicates that the data collection method may strongly influence who participates in a survey and how people answer questions. These two phenomena are captured by the concept of the mode effect. This paper evaluates the PtW surveys in Hungary from the perspective of mode effects by pairing two PtW surveys with face-to-face ESS surveys closest in time. The evaluation examines sample compositions and responses to two culturally and socially sensitive questions for which the survey mode is considered a highly relevant factor given Hungary's specific political context: attitudes toward (1) gay and lesbian people and (2) immigrants. Responses to these questions from both surveys are assessed using one-, two-, and multidimensional (GLM) analyses. The results show consistent patterns across both survey pairs. (1) PtW surveys yield similar response rates to face-to-face surveys but tend to appeal to more highly educated and older segments of the population. (2) Despite post-stratification to correct sample composition, the results still reflect different attitudes toward gay and lesbian people and immigrants across the two survey modes. (3) The survey mode has an independent impact on sensitive measurements: in GLM models, the survey mode significantly affects how respondents report their attitudes and overrides the original demographic correlations.
Survey research on minoritized citizens tends to categorize respondents from the top down, instead of allowing for identification from the bottom up. I surveyed 1864 respondents in Germany and the Netherlands, including 401 respondents with a background in T & uuml;rkiye. I find that those who identify as Turkish often hold significantly different attitudes than those who have a background in T & uuml;rkiye but who identify as German or Dutch. In fact, those who do not identify as Turkish, hold significantly different attitudes than when a researcher would categorize them as "having a migration background" and/or as "having a migration background in T & uuml;rkiye." I provide proof of this with attitudes towards topics often associated with citizens with a migration background. Beyond these empirical advantages to identification over categorization, this paper outlines additional theoretical, methodological and conceptual advantages to following an identification approach in designing surveys.
During the COVID-19 pandemic, there was a high demand for readily available but also reliable survey data. Overall, a comprehensive assessment of how survey data were collected and what their quality was remains missing. In this study, we provide a multi-dimensional quality assessment of social science surveys conducted during the COVID-19 pandemic in Germany (N=686). We assess survey quality based on three dimensions: accuracy (proxied by survey design), interpretability (referring to the quality of documentation), and accessibility of data (referring to the timely publication of results and data). Our results show that surveys varied considerably in these three quality dimensions over time. We found that surveys followed different purposes at different times of the pandemic: whereas early surveys focused on quickly producing results and traded other aspects of survey quality for this goal, later surveys were more focused on operating better-designed surveys and producing shareable data.
SRM wants to publish a Special Issue on "Transitioning to Self-Completion Modes of Data Collection in Social Surveys in Europe". This call for papers expands on the papers searched and the policies used.
While the advantages of voice input for answering open questions in web surveys seem clear, the challenge remains of maximizing voice inputs use while still giving respondents alternatives. This study experimentally explores three options for encouraging voice input in web surveys: a) PushDictation: respondents are asked to answer using dictation. If they try to skip the question, they are proposed to type in a textbox; b) PushRecording: respondents are asked to answer using voice recording. If they try to skip the question, they are proposed to type in a textbox; and c) Choice: respondents are offered three options to answer (dictation, voice recording or type in a textbox). These three options are compared to a Control group in which participants can only answer by typing in a textbox. Using data from two open questions in a survey about nursing homes implemented in February/March 2023 (N = 1,001) in an opt-in online panel in Spain (Netquest), we answer three research questions: (RQ1) What are the overall rates of response to the two open questions? (RQ2) What are the rates of use of voice input to these same questions? (RQ3) What is the overall quality of the data across the different conditions? Overall, response to the open questions was significantly lower when voice inputs were proposed, especially in the PushRecording group (RQ1). Furthermore, significant differences emerged in voice input usage between the experimental groups (RQ2). Regarding data quality (RQ3), the Control group exhibited the lowest proportion of valid answers, while the average numbers of themes and characters were in general higher in the push groups. Our results contribute to the growing but still limited literature about the use of voice input in web surveys, by adding new empirical evidence for several designs encouraging voice input to answer open questions.
The respondents’ interest in a survey’s topic is frequently used by survey researchers to explain and predict survey errors. Whether respondents are interested in a survey’s content relates to their participation and cognitive answering processes, consequently impacting nonresponse and measurement errors. The content of a survey is under control of the researchers who design and conduct the survey, thus, content could be varied to improve participation and answering behavior. Unfortunately, research is lacking on (i) the topic preferences in the general population, (ii) whether groups of respondents differ in their topic preferences, and (iii) how to measure these preferences. We address this research gap by presenting the findings of three experimental studies that we conducted. We found that topic preferences varied between samples and respondent subgroups. Moreover, we validated a measurement instrument to assess respondents’ topic interests. Based on our empirical findings, we derive practical recommendations for survey research and outline future research opportunities.
This paper addresses two methodological debates: first, it examines the consistency of responses in mixed-method designs; second, it does so with a particular focus on older people, both with and without cognitive impairments, living in care institutions. The empirical basis consists of 135 semi-structured interviews with individuals aged 50 to 92, conducted in residential homes for the elderly. We compare responses to open, narrative questions with those to closed questions on three topics: quality of life, decision to move, and cognitive decline. We conclude that responses on the same topic are generally consistent, although declining cognitive ability reduces this consistency. With this article we contribute to the discussion on the feasibility of surveying older (institutionalised) people with and without cognitive impairment, as well as the potential and benefits of combining qualitative and quantitative data.
Standardized surveys scale efficiently but sacrifice depth, while conversational interviews improve response quality at the cost of scalability and consistency. This study bridges the gap between these methods by introducing a framework for AI-assisted conversational interviewing. To evaluate this framework, we conducted a web survey experiment where 1,800 participants were randomly assigned to AI 'chatbots' which use large language models (LLMs) to dynamically probe respondents for elaboration and interactively code open-ended responses to fixed questions developed by human researchers. We assessed the AI chatbot's performance in terms of coding accuracy, response quality, and respondent experience. Our findings reveal that AI chatbots perform moderately well in live coding even without survey-specific fine-tuning, despite slightly inflated false positive errors due to respondent acquiescence bias. Open-ended responses were more detailed and informative, but this came at a slight cost to respondent experience. Our findings highlight the feasibility of using AI methods such as chatbots enhanced by LLMs to enhance open-ended data collection in web surveys.
The Survey Quality Predictor (SQP) is an open-access system to predict the quality, i.e., the reliability and validity, of survey questions based on the characteristics of the questions. The prediction is based on a meta-regression of many multitrait-multimethod (MTMM) experiments in which characteristics of the survey questions were systematically varied. The release of SQP 3.0 that is based on an expanded data base as compared to previous SQP versions raised the need for a new meta-regression. To find the best method for analyzing the complex data structure of SQP (e.g., the existence of various uncorrelated predictors), we compared four suitable machine learning methods in terms of their ability to predict both survey quality indicators: LASSO, elastic net, boosting and random forest. The article discusses the performance of the models and illustrates the importance of the individual item characteristics in the random forest model, which was chosen for SQP 3.0.
Declining response rates have remained a major worry for survey research in the 21st century. In the past decades, it has become harder to convince people to participate in surveys in virtually all Western nations. Worrisome, declining willingness to participate in surveys (i.e., response propensities) may increase the risk of extensive nonresponse bias. Therefore, a better understanding of which factors are associated with survey nonresponse and its impact on nonresponse bias is paramount for any survey researcher interested in accurate statistical inferences. Knowing which factors relate to low response propensities enables appropriate models of nonresponse weights and aids in identifying which groups to tailor efforts for turning nonrespondents into respondents. This manuscript draws on previous theories and research on nonresponse and investigates the risk of nonresponse bias, both cross-sectionally and over time, in two time series cross-sectional studies administered in Sweden (the National SOM Surveys 1993-2023 and the Swedish National Election Study 2022). Capitalizing on available registry data on all sampled persons and their corresponding neighborhood-level contextual data, a meta-analytical analysis of nine years of data collection finds that educational attainment, age, and country of birth are among the strongest predictors of response propensities. However, contextual factors-such as living in socially disadvantaged neighborhoods-also predict willingness to participate in surveys. Furthermore, utilizing the three decades of data, the growing nonresponse could be identified to be wholly attributable to a deteriorating survey climate rather than birth cohort replacement or immigration patterns.