
Abstract University branding is often included on envelopes used for mail recruitment materials for university-sponsored or administered surveys based on the assumption that signaling the university sponsorship improves the response rate. This postulation was experimentally tested by randomizing whether or not the university name and logo were included above the return address on the envelopes used to send recruitment materials for a statewide mixed-mode survey fielded with both an address-based sample (ABS) and an address-appended random digit dial (RDD) sample. For both samples, the statewide response rate was marginally lower when the university affiliation was included on the recruitment envelopes by 1.4 (ABS) and 1.7 (RDD) percentage points. Consistent across samples, the use of the university branding had significant negative effects on response rates in counties not geographically connected to the university. Closer to the university, however, the results were mixed, with the university branding significantly improving participation in the university’s home county by 7.9 percentage points for the ABS sample and having an insignificant effect in surrounding counties for both samples and in the university’s home county for the RDD sample.
Abstract The transition to online data collection in general population surveys has accelerated in recent years. Although high-quality web surveys now achieve response rates comparable to those of telephone surveys, they frequently display greater educational bias. This study analyzes complete employment biographies of both respondents and non-respondents to address three key questions: (1) How do different stages of the online panel recruitment process contribute to educational bias? (2) Are specific subgroups within the lower-educated population underrepresented in online surveys? (3) Are there interaction effects between education and other predictors of nonresponse, such as age, nationality, or employment status? In 2023, the Institute for Employment Research (IAB) in Germany launched the IAB-OPAL online panel survey using a push-to-web approach. The sample was drawn from a comprehensive administrative database containing social security, unemployment insurance, and basic income records, allowing for detailed analyses of employment histories. Leveraging detailed employment histories enables a granular assessment of how educational bias emerges across recruitment stages and whether response tendencies within educational groups vary due to typically unobserved factors such as benefit receipt, occupational history, or wages. The results indicate that educational bias increases at each stage of the recruitment process. Nonresponse is particularly common among individuals with lower educational attainment, especially those aged 50 and above. Although response probabilities for foreigners and Germans without an educational degree are similar, the disparity between these groups increases by at least a factor of four among individuals with a university degree. Additionally, men are less likely to participate unless they possess advanced degrees or lack formal qualifications.
Abstract Respondent-driven sampling (RDS) has been shown to be a promising network sampling strategy. However, it has also been shown to fail in some applications. Basing an RDS procedure on an existing panel study may, therefore, contain the risk of failure and might even alienate non-compliant seeds. This study therefore investigates the hypothetical willingness of general population panel participants to recruit their social network members using RDS. The study reveals that nearly half of the online participants of the panel survey selected for the study are willing to recruit. Reasons for unwillingness include discomfort in inviting others and the belief that they are already contributing sufficiently. Proportional odds ordinal logistic regression analyses reveal that most of the anticipated predictor variables for willingness have no effect. However, an increased willingness to recruit is observed among respondents of younger age, with an immigration background, and with politically extreme opinions. The findings suggest a substantial potential for RDS in online surveys but highlight the need for further empirical studies to evaluate actual recruitment behaviors and refine RDS strategies.
Abstract This paper provides evidence for the response-enhancing effect of a second, last-minute reminder near the end of an online business survey. Sending reminders to prospective respondents of an online business survey in Germany two days before the end of the field period increased the likelihood of them taking part by 8.9 percentage points compared with prospective respondents who were not reminded again. This effect persists to a certain extent, since the participation rate of firms that received a reminder in a given wave is also higher in the subsequent wave. Reminding the firms does not affect nonresponse bias or data quality.
Abstract Ecological Momentary Assessment (EMA) is an intensive longitudinal data collection method to capture in-the-moment experiences through frequent assessments. With the advent of mobile technologies, EMA’s applications have expanded across a number of disciplines. Despite its growing popularity, methodological issues, particularly regarding response quality, have yet to be explored. This study systematically evaluates how response quality changes over time in a 21-day EMA study with five daily assessments. Data were analyzed from 100 university students who completed surveys via a mobile application. Response quality was measured using a variety of satisficing behaviors, including speeding, nondifferentiation in grids, extreme rounding (i.e., reporting 0 or 60) on questions that asked the duration of a given activity during the past hour, and anchoring (i.e., providing the same answer to a question as in the previous assessment). The findings revealed a significant increase in speeding over the three weeks, suggesting a possible decrease in response quality. There were also increases in extreme rounding and anchoring responses in weeks 2 and 3. The nondifferentiation in grids mainly stayed the same across the weeks, which might be because the grids in this study only contained a few simple items to rate. The findings also showed variation between participants regarding their satisficing behaviors across the weeks. However, such variation could not be explained by participants’ demographic characteristics or motives for participation. Compared to the changes in satisficing behaviors across the weeks, the response quality differences across the five daily assessments were less systematic and inconsistent. These findings highlight the need for effective strategies to motivate participants to provide thoughtful answers as EMA data collection proceeds, and careful consideration of EMA design parameters to ensure high-quality data collection.
Abstract The Conservation Effects Assessment Project (CEAP) evaluates the environmental impact of the United States Department of Agriculture’s (USDA) conservation programs. This project aims to develop estimates of the total rangeland acreage in arbitrarily specified regions, which also have improved accuracy at the county level compared to traditional design-based estimators. To achieve this, we developed a novel model-based estimator that integrates information from both the National Resources Inventory (NRI) survey data and the satellite-derived Cropland Data Layer (CDL) data. A fine grid of square tiles was designed to capture the information of the CDL pixel near the NRI samples. Lasso, fused Lasso, and K-means clustering were employed for feature engineering and variable selection to identify optimal models for rangeland acreage estimation. Empirical results indicate that the new model-based estimator provides a reliable and effective alternative to the traditional design-based estimator.
Abstract We introduce a new simple novel method to measure uncertainty in an estimated ranking of K populations. First, form a family of joint confidence intervals for differences of population parameters, with familywise coverage 100 (1−α) percent. Then construct a 100 (1−α) percent joint confidence region for the overall true ranking of the K populations. With the assumption of normality for the estimators, we only need K estimates and their associated standard errors to produce the new joint confidence region, and we assume these are given (published). A theoretically based visual shows at once: (1) joint confidence region revealing uncertainty in the estimated ranking; (2) possible true rankings, beyond the estimated ranking; (3) a marginal confidence set for population k true rank; and (4) a marginal confidence set for each rank r.
Abstract Smartphones play a central role in everyday life and offer significant potential for survey data collection. However, their usefulness may be constrained by the smartphone’s operating system (OS) and its version, which determine app compatibility. This article investigates methods for accurately determining smartphone OS version, an important factor when evaluating device compatibility for data collection tools. We compare three approaches: (1) passive collection of paradata (user agent strings; UASs) from web respondents, (2) self-reported smartphone make and model matched to OS version information from an online database, and (3) direct self-reporting of OS version by respondents following step-by-step instructions. Using data from the UK Household Longitudinal Study COVID-19 web survey, a probability sample of UK households, we assess the completeness and accuracy of each method. Our findings suggest the paradata were in some respects inferior to the methods based on self-reports. First, the UAS data only provided valid smartphone OS version data for about 50 percent of respondents; the other 50 percent did not use a smartphone to complete the web survey (compared to 90 percent and 71 percent of valid cases based on methods (2) and (3)). Second, the subset of respondents who completed the survey using their smartphone was not representative of the full set of respondents, while the respondents for whom OS version data were obtained from self-report methods were broadly representative. Third, the UAS data under-represented older OS versions. Respondents with older smartphones seem unlikely to use them to complete the survey. Our findings also show differences between OSes, suggesting it was easier for iPhone users to provide relevant and valid information than for Android users: they were more likely to report a valid make and model that could be linked to technical data, to report the OS version, and to use their smartphone to complete the survey.
Abstract With declining survey response rates, researchers have been exploring effective strategies for participant recruitment. This study considered the household screening stage of recruitment in the 2022 Health and Retirement Study (HRS), a national panel study. The study investigates the effects of different contact protocols and invitation letter envelope types on completion outcomes and contact attempts, specifically focusing on the variation in effects among race/ethnicity and HRS cohorts. A total of 16,452 sampled households were divided randomly into two groups. One received a web-field protocol, where sampled individuals were invited to complete the questionnaire online and nonrespondents were eventually followed up FTF (face-to-face), and the other received a field-only (FTF) protocol. Within the former group, two random subgroups were created. One group received a standard HRS mailing envelope containing the invitation letter and a prepaid $2 bill cash incentive. The other received a visible cash envelope with a window showing the incentive. Logistic regression models were employed to analyze completion outcomes by screening protocol and envelope type. Interaction terms were subsequently included to explore differences in these effects among race/ethnicity and HRS cohorts. Linear regression and Poisson models were used to assess predictors of the number of contact attempts needed among completed cases. Findings show that the web-field protocol improved completion rates and reduced contact attempts among the youngest cohort, but had a negative effect among Hispanic households. The type of envelope also significantly influenced screener completion. Receiving a standard envelope increased response likelihood among non-Hispanic Black households. The youngest cohort also exhibited a higher likelihood of response with a standard envelope and generally received fewer contact attempts with both types before completion. However, visible cash did not present strong effectiveness. Overall, the findings highlight the effectiveness of employing different contact methods to enhance survey participation, specifically among different socio-demographic subgroups.
Abstract This study extends the classical Fay–Herriot model to a threefold hierarchical structure incorporating heteroscedastic random effects. Three submodels are also introduced within this framework. Small area best linear unbiased predictors are derived for linear indicators, and their mean squared errors (MSEs) are estimated using both analytical and parametric bootstrap methods. Model performance and reliability are evaluated through diagnostic tools and influence measures, specifically designed for small area estimation. Simulation experiments are conducted to analyze the empirical properties of the predictors and MSE estimators. The proposed approach is applied to data from the 2019–2021 Spanish Living Conditions Survey to estimate the proportion of women and men below the poverty line, disaggregated by province, age group, and year.
The present study investigated and assessed translation errors committed by Iranian researchers in the process of translating questionnaires from English into Farsi (Persian) (the official language of Iran) within the context of questionnaire validation studies. Employing a measurement equivalence/invariance analysis, this research applied the evaluation framework proposed by Behr (2023) to compare the original English versions of questionnaires with their translated Farsi counterparts, based on the premise that fewer errors correlate with higher translation quality. The dataset for evaluation was compiled through systematic database searches, focusing on questionnaires on women. A total of 14 questionnaires meeting the inclusion criteria, comprising 292 items, were identified. Two researchers independently compared and evaluated the quality of the source and translated versions of each questionnaire, followed by online sessions in which they engaged in dialogic discussions to share their coding and perspectives. The questionnaires were examined for translation errors, and the frequency of each of the seven error types outlined in the framework was quantified, exemplified, and analyzed. Findings revealed that the majority of errors (88 percent) occurred at the linguistic level, encompassing general translation issues such as semantic accuracy, stylistic appropriateness, and grammatical correctness, whereas fewer errors pertained to questionnaire-specific translation challenges, including cultural adequacy, terminological consistency, and layout/presentation. Additionally, the study discussed the possibility of introducing some refined manifestations of error types mentioned by the framework. Based on these findings, practical recommendations were provided for researchers involved in questionnaire translation and for journal reviewers and editors overseeing such work.
This article addresses the challenge of low statistical power in A/B testing with small sample data, such as Automatic Speech Recognition data. Traditional methods, such as the Welch t-test, often underperform in these scenarios. We introduce two novel testing methods, the stratified and post-stratified empirical likelihood ratio tests, which reduce group variance and enhance test sensitivity. Our theoretical analysis and comprehensive experiments on real and synthetic datasets show that these nonparametric empirical likelihood ratio methods outperform the Welch t-test with small-sample data, providing a more effective tool for detecting treatment effects and informing data-driven decision-making.
Abstract One can study any real graphs based on the subgraphs obtained by probability sampling. This is useful when it is either infeasible or too costly to process the whole graph due to various reasons. We consider probability snowball sampling (SBS) from graphs, where the initial node sample is selected with known probabilities, and each following wave of observation is carried out exactly as specified. The literature on design-based inference for probability SBS from graphs has limited scope, and there does not exist any design-unbiased strategy that generally makes use of units (or networks of units) obtained after the initial sample. In this paper, we adopt a unified framework for T-wave SBS from graphs, where the study units are not limited to the nodes in the graph but may be any finite-order subgraphs, say, triangles, cycles, or stars. We propose two practical design-unbiased strategies for estimating the corresponding graph totals, which considerably extend the previous approaches to probability SBS. The practitioners are thereby provided with richer choices to improve sampling efficiency, which we will demonstrate with an application to the actor-actor network from IMDb.
In multi-day diary surveys, participants decide daily whether to continue. The day-level nonresponse can introduce nonresponse errors, and underreporting tends to increase over the data collection period. To address these problems, we propose an adjustment by either the construction of person-day level survey weights or multiple imputation (MI). This study used data from the US National Household Food Acquisition and Purchase Survey (FoodAPS) and focused on four key outcomes in this survey: the occurrence and expenditure of daily food-at-home (FAH) and food-away-from-home (FAFH) events reported by individuals. We employed logistic regression models and conditional inference trees to predict daily response propensities and used the inverse of predicted response propensity as an adjustment factor for the existing FoodAPS household weights. As an alternative, we conducted MI for the key outcome variables using chained equations and fit zero-inflated imputation models for counts and semi-continuous variables to allow bounds and subset restrictions. The results show that raw estimates for the key variables are consistently smaller than those produced by any weighting or imputation techniques, indicating that some form of adjustment is necessary. Person-day level weights and MI show varying impacts across outcome variables, with MI offering a modest efficiency gain. We also performed a sensitivity analysis that employs an alternative definition of missingness, which yielded different MI estimates. This study offers valuable insights into addressing nonresponse errors in multi-day diary surveys and contributes to methodological approaches for conducting innovative weighting and imputation with complex data structures.
This study extends the scope of the paradata discussion to respondent-driven sampling (RDS). Unlike traditional sampling, RDS relies on existing social networks within a target population. This unique process provides opportunities to produce novel paradata. Specifically, this study examined two types of paradata in RDS: one based on interviewer observations and the other based on recruitment behaviors ascertained from tracking recruitment coupons. We implemented these paradata features in two independent RDS surveys. In an in-person RDS survey of persons who inject drugs in Southeast Michigan, we implemented an interviewer observation questionnaire. This included questions about interviewers' assessments of respondents' understanding of coupon distribution instructions, as well as their expectations regarding respondents' chances to recruit others and to return for a follow-up interview. These observations predicted recruitment success. In a Web-RDS study of Korean Americans, physical distance between linked respondents (such as a respondent and their recruiter) was determined by tracking recruitment coupons and geocoding respondent addresses. Greater geographic distance was associated with a higher likelihood of serious psychological distress. The results demonstrate that the unique features of RDS offer new avenues for utilizing paradata in both methodological and substantive research. These findings warrant further exploration and development of paradata specific to RDS.
Happiness is a widely used measure of subjective well-being, often assessed through a single survey item (e.g., "How happy are you these days?") in global social surveys. However, variations in response scale design across surveys complicate comparisons of happiness both within and across countries. To examine the impact of response scale design on happiness ratings, we analyzed data from 34 interviewer-administered (mostly face-to-face) national probability-based surveys conducted in China, spanning six survey programs and comprising approximately half a million responses. We focused on three key scale design features that varied across surveys: scale direction, presence or absence of a midpoint, and scale length. To enable meaningful comparisons across surveys with different scale lengths, we standardized self-reported happiness scores using z-score transformations within each survey. This approach allowed us to evaluate how scale features are associated with respondents' relative positions within each survey's distribution, rather than comparing raw mean happiness scores across studies. Our analyses revealed no differences in standardized happiness ratings across scale lengths, but scales without a midpoint resulted in higher ratings. The relationship between scale direction and standardized happiness ratings differed markedly across education subgroups. Specifically, among respondents with lower levels of education, descending scales (e.g., from "a lot of happiness" to "a lot of unhappiness") were associated with higher standardized happiness ratings compared to ascending scales. This pattern aligns with findings from Western research on scale direction effects. In contrast, among respondents with higher education-particularly those in the highest education category-descending scales were associated with lower standardized happiness ratings. These findings suggest that culturally specific mechanisms may interact with scale direction to influence how individuals report their happiness.
Recent guidelines on the measurement of gender in surveys recommend a two-step strategy that asks about sex assigned at birth followed by current gender identity using a small number of offered response categories (e.g., female, male, transgender) and an open response field to capture any residual categories ("not listed, please tell us"). However, more research on several fronts is needed to continue to refine our measurement strategies across contexts, and this work continues to be critically important with the dismantling of federally funded research regarding gender identity. In the current study, we examine the impact of response format on the measurement of gender identity. This study reports results from a between-subjects experiment embedded in a campus climate survey about inclusion and belonging at a large Midwestern university in Fall 2024. Over 16,000 students were asked, "What is your gender?" and subsequently randomly assigned to respond using one of two response formats. The first allowed respondents to "select-one" of the response options "woman, man, nonbinary, not listed (please tell us)," and was followed by the question "Are you transgender? Yes/No." The second included the response options "woman, man, nonbinary, transgender, not listed (please tell us)" and instructed respondents to "select all that apply." We examine the distribution of responses, item nonresponse, the number of gender categories reported, response times, and concurrent validity (in terms of the association between gender and survey outcomes about campus climate) across the two formats. While the results mainly show similarities in outcomes between the response formats, respondents are more likely to indicate they are transgender in the "select-one" format. By contrast, respondents are more likely to indicate gender expansiveness other than "transgender" or "nonbinary" with the "select-all" format. We discuss the implications of these findings for future research.
The integration of QR codes (quick response codes) in web surveys has become commonplace, aimed at simplifying survey access for respondents using mobile devices. While previous studies have presented mixed findings regarding the effectiveness of QR codes in enhancing survey participation, this research re-examines their impact on survey breakoff, identifying potential explanatory factors such as the screen resolution of the respondent's device. Drawing on data from the Digitize! Online Panel Survey, a probability sample of up to 1,228 respondents in Austria, this study uses an exploratory research design to investigate how screen resolution, QR code usage, and frequency of personal computer or mobile device use influence survey breakoff behavior. We find that QR code utilization alone does not significantly predict survey breakoff once other covariates are accounted for. However, the resolution of the device used for the survey emerges as a critical determinant, with higher resolutions associated with lower probability of survey breakoff. The study thus uncovers a nuanced relationship between screen resolution and survey breakoff, particularly noticeable when transitioning from mobile devices likely to be smaller, such as smartphones, to those with larger screens, like tablets. While QR codes offer convenience and accessibility advantages, their implementation requires careful consideration to mitigate potential drawbacks and optimize survey engagement. Overall, this study contributes valuable insights into the evolving landscape of digital survey methodologies, highlighting avenues for future research to enhance data validity and reliability.
The Fay-Herriot model has been the workhorse of area-level small area estimation. This model links direct survey estimates to area-level covariates through a linear mixed-effects framework. While effective, this model assumes a linear relationship between covariates and the true area-level quantities, limiting its flexibility in capturing complex patterns inherent in real-world data. In this work, we introduce an extension of the Fay-Herriot model that integrates Bayesian Additive Regression Trees (BART) to model the true area-level quantities as nonlinear functions of the covariates, complemented by an additive random effect to account for unexplained heterogeneity. The proposed BART Fay-Herriot (BART-FH) model leverages the nonparametric capabilities of BART to capture intricate nonlinear relationships and interactions among covariates, offering a more flexible alternative to traditional linear models. To evaluate the performance of the BART-FH model, we conduct an empirical simulation study comparing its estimation accuracy and predictive capabilities against the standard Fay-Herriot model. Furthermore, we apply the BART-FH model to household income data from the American Community Survey (ACS).
Satisficing is a behavior that often results in careless responding and has been defined as engagement in simplified question responding approaches to reduce mental effort, often at the cost of response quality. Our objective was to examine the possibility of using a new survey response time (RT) based metric, response time adjustment, in combination with average survey RT, to identify groups with different levels of satisficing behavior. The RT adjustment parameter reflects the extent to which individuals adjust the amount of time they spend on survey items as a function of how demanding an item is, with low adjustment along with a low average survey RT expected to be associated with a greater likelihood of satisficing. We estimated a mixture model with RT adjustment and average RT as inputs using three questionnaires (with sample sizes of 5,321, 1,616, and 4,093, respectively) from the Understanding America Study (UAS), a large US Internet-based longitudinal panel. Weak and non-satisficing groups were identified in all three studies, with the former making up 3 to 9 percent of the samples, but no strong satisficing group was detected. Evidence supporting a weak satisficing group was based on good accuracy on easy attention check items (indicating they likely were not careless) but low accuracy on more demanding attention check items involving carefully reading long blocks of instructional text. Additionally, consistent with prior literature on satisficing, this weak satisficing group generally had lower mean values on cognitive ability and motivation-related variables (e.g., conscientiousness) compared to other groups. In two of the three studies, the satisficing group produced notable bias in study results involving high time intensity items. RT regulation and average RT may be useful for identifying satisficing in surveys with a mix of high and low time intensity items such as some tests and surveys with vignette-based items.