
Longitudinal surveys play a central role in social science research and policy advice but increasing demands for timely data pose substantial challenges for managing time, resources, and data quality across survey waves. While the Survey Life Cycle (SLC) provides a well established framework for organizing survey production, it is primarily geared toward single wave studies and offers limited guidance for coordinating interdependent longitudinal processes. This article introduces the Longitudinal Survey Life Spirals (LSLS), a conceptual planning and reflection framework that reimagines the SLC in a spiral structure to explicitly account for cross wave interdependencies, overlapping project phases, and learning processes over time. By integrating considerations of timing, resource allocation, and feedback loops across waves, LSLS support more coordinated longitudinal survey management. The framework is illustrated using the Socio Economic Panel (SOEP), a large scale longitudinal survey infrastructure characterized by simultaneous work on multiple survey waves. The article concludes by outlining key implications for survey practice, highlighting how LSLS can help survey teams to improve planning, quality assurance, and the timely dissemination of high quality longitudinal data.
Probability sampling remains the standard for face-to-face public opinion research, yet it becomes hardest to implement precisely where reliable data are most needed. In such settings, a large share of survey error can enter before the first interview begins, when field teams are left to identify households without a stable selection rule. In Myanmar and China, the Asian Barometer Survey faced different versions of the same operational problem. In Myanmar, lower-level household information was too weak to support ordinary final-stage selection, and some areas were inaccessible because of conflict. In China, migration, hukou-based dataset undercoverage, and the political sensitivity of requesting household name lists made conventional final-stage sampling difficult to defend and implement. We therefore retained multistage probability-proportional-to-size sampling where higher-level population counts were usable but replaced the weakest stage with GIS/GPS-assisted area sampling based on gridded population data and digital boundaries. Selected half-minute grids were divided into smaller grids, uninhabited cells were removed before field release, interviewers enumerated addresses inside the sampled cell in a fixed order, and one adult was chosen within the household using a Kish procedure. Field records from Myanmar show that this approach supported national fieldwork, pre-planned replacements, and intensive quality control. Weighted benchmark comparisons from both countries show that the resulting samples tracked several external population margins closely. The main lesson is practical: when household lists cannot be used cleanly, a documented grid-based design can replace field discretion with an auditable household-identification rule and preserve probability logic better than field improvisation.
Visual aids are widely used in public opinion polls to enhance comprehension and respondent engagement (Petersen 2008). However, incorporating visuals into telephone surveys has traditionally been considered impracticable (Lavrakas 2008). This study presents a proof-of-concept case derived from the City of Aspen’s 2024 community survey, which integrated visual elements across multiple survey modalities, including postal, telephone, and online formats. The multimodal design provided all respondents with uniform visual resources, ensuring a consistent informational framework regardless of participation mode (Dillman and Christian 2005; Nathan 2001). The findings demonstrate the practical viability of incorporating visual aids into telephone surveys through the use of mailed postcards and online visual references. Although the limited telephone subsample (n = 29) constrains statistical power and precludes inferential comparisons, the results illustrate a promising approach for future research. This research note highlights the potential of multimodal survey designs to improve inclusivity and data quality by ensuring equitable access to visual information across respondent groups (Ciccione et al. 2025; Couper, Tourangeau and Kenyon 2004; Tourangeau, Rips, and Rasinski 2000).
This research note introduces a qualitative testing technique, called MicroTesting, intended to supplement a large-scale, multi-phase questionnaire (re)design initiative. The tests are “micro” because they are much shorter, involve fewer respondents, focus on a subset of an instrument, and can be conducted quickly. With this nimbleness, MicroTesting can be introduced as needed, adding confidence and quality to the larger rounds of testing.
Social media advertising has become an increasingly popular tool for recruiting respondents for surveys. This article documents a recruitment campaign using Instagram advertisements to recruit individuals aged 18 to 29 in Germany while aiming to adhere to predefined quotas for age, gender, and region of residence. Advertisements were managed via Meta’s advertising interface, allowing for demographic targeting, budget adjustments, and iterative interventions during the field period. Despite continuous monitoring and multiple adjustments to budgets, targeting configurations, and advertisement content, recruitment deviated substantially from the intended quotas, most notably with respect to gender, whereas quota alignment for age and region remained comparatively stable. Interventions intended to correct imbalances affected overall recruitment volume but did not reliably improve quota alignment. In addition, highly effective advertisements were associated with systematic distortions in substantively relevant variables, and their removal led to sharp declines in participation without eliminating these biases. The findings highlight the limited controllability of quota-based recruitment via Instagram advertising under real-world conditions and provide practical lessons for survey practitioners regarding expectations, monitoring strategies, and trade-offs inherent in social media–based survey recruitment.
Artificial intelligence is reshaping the operational workflow of survey research. The tasks that historically formed the apprenticeship pathway for new professionals are increasingly automated: question drafting, instrument programming, open-ended coding, fraud detection, and preliminary analysis. These changes improve efficiency, but they also alter the developmental conditions under which methodological judgment historically formed. This article argues that survey organizations must redesign staff development intentionally to preserve foundational competencies while cultivating new capabilities in AI-integrated diagnostic judgment and system stewardship. We propose an interconnected framework of future survey expertise organized around three mutually reinforcing domains: foundational methodological competence, AI-integrated diagnostic judgment, and system stewardship and governance. Automation redistributes expertise rather than eliminating it. Meeting that challenge requires three sequential organizational actions: redefining roles to reflect what AI has changed, redesigning the early-career development pathway so that the next generation builds genuine methodological judgment, and building the infrastructure that makes both sustainable.
If we want to measure change, we shouldn’t change our measures. Yet measurement changes do happen. Some result from global circumstances such as mode switches necessitated by the COVID-19 pandemic. Others are related to changes in the presentation and prevalence of a topic. The number of people who identify with no religion, also known as religious “nones,” has grown rapidly in many countries (Hackett et al. 2015; Stolz et al. 2025). We recently published an extended discussion of many measurement issues that may exaggerate the growth of religious “nones” (Conrad and Hackett 2026). In this short note, we highlight three common changes in the measurement of religious “nones” and call for survey and census organizations that make these changes to concurrently study how they affect the apparent prevalence of the “no religion” population.
One metric for measuring the accessibility of a translated survey instrument is evaluating the “uptake,” or rate at which respondents access the translated version of the survey, which has the goal of increasing overall response rate among Spanish speakers. This paper analyzes an experimental effort to increase the uptake rate of the Spanish version of the Household Pulse Survey online questionnaire by varying the method for accessing the Spanish language instrument. In the experiment, respondents were randomly assigned to either a control instrument or a treatment instrument. In both instruments, respondents could toggle or switch the language of the instrument using a dropdown in the upper right corner on each screen. In the treatment condition, an explicit question appeared on one of the initial instrument screens asking respondents to choose their preferred language for completing the survey (either English or Spanish), while in the control condition respondents were not given an explicit choice of language. Implementing the explicit language selection question, along with a revised contact protocol, resulted in a slight increase in the uptake of the Spanish instrument. As a result, the decision was made to implement the treatment design for all respondents in subsequent data collections.
Researchers and practitioners often have a need to summarize qualitative text data obtained from open-ended survey responses, case reports and various other sources. However, traditional methods using human coders are costly and time-consuming. Recent advances in artificial intelligence (AI) tools hold promise for more timely and cost-effective processing of such qualitative text data. The present study assesses ChatGPT as an aid in conducting thematic content analyses. We analyzed archival data provided by West Point military academy cadets writing about their reactions to the 9/11 terrorist attacks. Cadet responses were analyzed first by human coders using traditional approaches, and then by ChatGPT. The ChatGPT results corresponded well with those of the human coders, with some minor differences. While caution is required in the application of AI tools such as ChatGPT, the present study suggests that AI tools can provide a useful adjunct to human coding, speeding the process of identifying underlying themes in qualitative text data.
Research examining the individual impacts of wildfires has found that wildfires are associated with worse mental health, more physical illness, and increased financial strain. However, these results often rely on local surveys fielded after a wildfire occurs and there is little to no understanding of how post-wildfire differential non-response may bias these estimates of impact. If those individuals most impacted by the wildfire are unwilling or unable to respond to post-wildfire surveys, this would lead to underestimating wildfire impacts. If individuals who are less impacted are less willing to respond, wildfire impacts will be overestimated. We use monthly panel data from the Understanding America Study and its LABarometer subpanel to study differences in response rates among Los Angeles County residents before and after the January 2025 Palisades and Eaton wildfires. By matching survey respondents with geocoded data on wildfire exposure, we find some evidence that response rates declined in post-wildfire months among respondents residing in census tracts in closer proximity to the burn zone, under evacuation orders, or under evacuation warnings. However, these effects are of short and inconsistent duration.
During the COVID-19 pandemic, demand for rapid, high-quality, representative data surged. This was particularly acute for underrepresented groups in the UK. To address gaps in research on Britain’s ethnic and religious minority groups, the Centre on the Dynamics of Ethnicity (CoDE) launched the Evidence for Equality National Survey (EVENS), as a collaborative study with community partnership at its core, embedding Voluntary, Community and Social Enterprise (VCSE) organizations in both instrument design and sample recruitment. While such partnerships are often encouraged, little is reported on the outcomes of such partnerships for study engagement, notably for recruitment pathways. For the EVENS study, VCSEs were intentionally given a large degree of autonomy in their recruitment approaches, leveraging organizational expertise and community familiarity of their respective constituent communities to maximize outreach. However, this approach introduces ambiguity into the recruitment process itself as research teams do not have full oversight of activities and processes used in the recruitment stage. Studies which incorporate VCSEs into recruitment procedures may not be fully aware of the ramifications of the ‘black box’ created by including intermediaries into the recruitment process. We analyze recruitment outcomes when community organizations, representing under-represented groups, act as primary sampling agents within their own communities. Specifically, we cluster ethnic groups based on their self-reported recruitment sources to explore (dis)similarities across recruitment pathways. Our findings suggest that recruitment based on intermediary recruitment actors (i.e., VCSEs) produces discernible patterns in recruitment ‘profiles’ which are important to consider for reporting and analysis. These results will be of interest to practitioners interested in community co-production and implementation of survey research with underrepresented groups, particularly in terms of building-in diagnostic considerations for nonprobability, community-partnered approaches.
Pay-for data collection platforms are increasingly popular in health sciences, due to the ease of use, accessibility, and speed of data collection. However, data quality among platforms may differ. The purpose of this study was to compare data quality from three commonly used, pay-for data collection crowdsourcing platforms, using predefined quality control measures. Data were collected using three major platforms, Qualtrics (N = 705), SurveyMonkey (N = 576), and Amazon Mechanical Turk (N = 1,034), for the methodological purpose of assessing data quality for a scale used by public health practitioners to assess ankle activity level. Data were collected using attempted quota sampling for gender, age, ethnicity and region, for representativeness of the general US population. Twenty quality checks in five categories were utilized to identify poor-quality data: attention verification (e.g. ‘please select response B’); demographic verification (e.g. state and zip code match); illogical responses (e.g. respondents separately reporting higher difficulty walking up one flight of stairs than walking up four flights); honesty/reliability verification (e.g. male reported to be pregnant); and assessment of open-ended questions (rejected based on similarity of content or nonsensical responses). Of the 2,315 completed responses received from the three different platforms, 947 (40.91%) passed our quality checks and were considered good-quality responses. SurveyMonkey yielded the highest proportion of good-quality data (70.5%), followed by Qualtrics (61.8%) and Amazon Mechanical Turk (10.2%); however, open-ended responses could not be included in SurveyMonkey due to their 80-item limit, which may have inflated its proportion of good-quality responses. This study demonstrates that data quality issues may differ across platforms and a variety of quality control measures should be implemented for researchers utilizing pay-for data collection services from crowdsourcing platforms.
Declining student survey response rates pose challenges for higher education assessment research focused on improving undergraduate education. This study explored one potential approach to increase response—a split questionnaire design—with a nonrespondent population from the 2023 National Survey of Student Engagement administration. Using a randomized controlled experiment with 20 U.S. institutions, we offered approximately 63,000 students the opportunity to complete either one of five short SQD surveys or the full-length instrument with over 100 questions. Results indicate SQD surveys yield modest but operationally meaningful response rate advantages relative to the complete instrument, particularly among graduating seniors as well as some minoritized student populations. Though institutional variation exists, SQD appears to be a viable strategy to improve student survey participation in higher education settings.
In longitudinal studies, retaining participants over time is necessary to ensure sufficient sample sizes for analyses and reduce the potential for bias; however, efforts to improve response rates can increase the cost of data collection substantially. To explore methods that would best balance data quality and cost for a longitudinal study of veterans deployed during the Gulf War era, we conducted an experiment comparing two multi-mode designs which offered a higher incentive for completing the survey by web. Veterans were randomly assigned to receive one of two protocols: web offered initially with paper and the bonus web incentive introduced later (sequential Choice+), or both web and paper offered at the same time, with the bonus web incentive offered at the beginning (concurrent Choice+). Both protocols offered Computer Assisted Telephone Interviews to non-respondents. We examined the impact of each approach on response rates, sample representativeness, and cost. The response rate for the concurrent Choice+ group was 2.8 percentage points above that of the sequential Choice+ group (47.5 versus 44.7 percent); however, there were no differences in the characteristics of respondents in each group compared with the eligible sample or between respondents in each group on key survey items, an indicator that lower response did not result in biased estimates. The cost to implement the concurrent Choice+ design was substantially higher. These findings suggest that the less costly sequential Choice+ approach may provide a better tradeoff between quality and cost, particularly for longitudinal studies seeking cost efficiencies. Future research may benefit from exploring varying amounts for the web bonus and how a sequential Choice+ protocol could be used to increase response rates for multimode surveys that include web and CATI but not paper.
When developing a web questionnaire, respondent burden is a major concern. It may be tempting to condense multiple questions into a single grid or question to reduce the number of screens shown to the respondent. However, as time spent answering a survey question varies by the question’s complexity and the cognitive burden placed on the respondent, condensed questions may not actually save time in some circumstances. Using a question about which children in the household eat what school-provided meals, this experiment aims to test if more screens do take more time, even if the questions being asked are simpler to answer when kept separate. The decomposed version asks which children eat school-provided meals and then shows a simple grid listing only the relevant children and the two possible meals. The time to answer this series is compared to the time to answer a comparable “condensed” survey item which uses a single grid to list the children, school meals, and “No meals”. Findings indicate that the decomposed version appears to be faster, but this is partially due to the number of households with no children who eat school-provided meals. In this case, the first question of the decomposed version is significantly faster than the condensed question. However, the condensed question is significantly faster for respondents with at least one child eating a school meal. Comparing these two situations, the time saved in the decomposed version is greater than the time saved in the condensed version, leading to the overall efficiency of the decomposed question.
In a pilot survey aiming to inform Veteran suicide prevention, Veterans who were Black/African American, Hispanic, Asian/Native Hawaiian/Pacific Islander (AANHPI), female, and recently separated responded at lower rates. To increase response in a subsequent survey, an experiment was conducted to examine the effect of receiving tailored inserts (i.e., recruitment materials with images of Veterans from one’s demographic group), compared to a generic insert (i.e., images from a variety of demographic groups) or no insert. We hypothesized that receiving tailored inserts with images of Veterans from respondents’ demographic group would increase survey response. The experimental manipulation yielded non-significant results, except for significant findings in the opposite direction as hypothesized among Black/African American Veterans, for whom yield (i.e., percentage of sampled cases who returned a completed survey) was lowest among those who received the tailored insert. Conversely, while not statistically significant, yields were higher in the tailored insert group for AANHPI, Hispanic, and recently separated Veterans. Findings suggest that the impact of visual representation in survey recruitment materials may differ across Veteran groups. Alternately, more effective images may be needed to optimize tailoring of recruitment materials. Additional research is warranted to better understand whether tailored inserts can increase response among Veterans who are harder to engage in research.
In recent years, United States Postal Service (USPS) delivery delays have been reported, and proposed changes may result in an expanded delivery window, which poses challenges for survey data collection relying on the USPS. Delays increase the possibility that mailed reminders and follow-up outreach will cross paths with the completed and en route self-administered mail booklets. One approach to limit this possibility is to include unique Intelligent Mail barcodes (IMb) on courtesy reply or business reply envelopes (BRE). We describe how we incorporated inbound BRE tracking for a fall 2024 survey and report on the accuracy of the IMb tracking, the average number of days between the initial USPS scan and the delivery date, and the potential cost savings from suppressing mailed follow-up reminders for returned BREs that were en route via the USPS but had not been delivered.
Survey designers need to consider what categories of addresses to include as part of address frames, often defined by “delivery type,” e.g., residential, business, PO BOX, and others. Drop points can be a particularly challenging category because they are non-differentiated multi-unit addresses. Specifying a selected housing unit from the computerized delivery sequence file (CDS) is thus not directly possible, challenging common processes associated with address-based sample (ABS) designs, including vendor matching and demographic appends in addition to basic data collection. Our paper first reviews solutions that have been considered in the research industry to include drop points in multi-mode ABS designs over the past decade. We then demonstrate a variety of imputation methods utilized at NORC to assist in assigning unit labels to drop point addresses. We conclude by describing how each imputation method has coverage and representation advantages over previous approaches.
One increasing trend in demographic survey items is to include a “prefer not to respond” (PNR) option, particularly for items that might be considered sensitive information. This mixed-methods study uses data from the 2023 administration of the National Survey of Student Engagement (NSSE) to explore whether the frequency of selecting this option is related to substantive survey content, in this case other aspects of the student experience. Quantitative results suggest small but statistically significant negative relationships between PNR frequency and all Engagement Indicators. There are also negative relationships for belonging, satisfaction, self-reported grades, and intent to return, as well as a positive relationship for age. To explore potential motivation for PNR response, qualitative thematic analysis of final open-ended responses for students with high PNR was conducted. This revealed several themes, including: disapproval of demographic items themselves; hostility toward Diversity, Equity, and Inclusion (DEI); institutional distrust; and conversely also some praise for professors and certain institutional experiences. Potential reasons for these patterns, along with implications for institutional and survey research professionals, are also discussed.
The emergence of agentic artificial intelligence (AI) tools capable of autonomously interacting with web interfaces presents new challenges and opportunities for online survey research. This study evaluates the capabilities of OpenAI’s Operator, an agentic AI tool, in completing online surveys and investigates methods for detecting AI-assisted responses. Specifically, Operator was tasked with completing a variety of survey question types under different prompt conditions. The study identifies both the strengths and limitations of Operator in mimicking human respondents and assesses the effectiveness of existing and novel detection mechanisms. Results indicate that while Operator can successfully complete most survey tasks and evade some traditional fraud detection methods, several behavioral and metadata-based indicators can reliably identify AI-assisted responses.