Comparable questionnaire translation is essential for drawing valid conclusions in cross-cultural survey research. Sound translation methodology, including the use of adequate personnel, is seen as crucial for reaching this goal (Harkness 2003). Recommended methodology should be empirically backed and stay tuned to latest developments, such as machine translation. Against this backdrop, to investigate the potential effect of varied translators’ backgrounds and machine translation on the statistical properties of surveys, we conducted an experiment in which an English questionnaire was translated into German by 16 professional translators and 16 social scientists; translations were subsequently fielded in web surveys. We introduced two translation conditions: translation from scratch and post-editing (machine translation corrected by a human translator). To investigate the quality of the survey data from these 32 translation versions (approx. 250 responses each), we use standardized mean distance and Cohen’s d with the official translation as a benchmark. We have four key findings: First, the resulting statistical means of the survey items vary, sometimes substantially, across translations. Second, post-editing is associated with a reduced gap between the survey data from the experimental questionnaires and the official translation, and it also lowers the variability among different translations. Third, when translating from scratch, social scientists are more likely to produce translations leading to statistical outliers of survey data. Fourth, post-editing can lead to systematic bias for both social scientists and professional translators if translation errors made by the machine are not identified and corrected. This study highlights to what extent decisions concerning the choice of translators and the integration of machine translation can impact the statistical properties of survey data. We offer evidence to implement recommendations for good practices in translation protocols to enhance data comparability in cross-cultural studies.
Questionnaires in different languages are at the heart of cross-cultural survey research. To ensure their comparability and comprehensibility, translation procedures have continuously been optimized. However, the impact of a translator’s competencies on questionnaire translation and post-editing quality has, to date, not been systematically examined. With the aim to optimize team compositions, this study looks at the very first step of the TRAPD process – initial translations and the translators producing them. In an in-between translation experiment, 16 professional translators and 16 social scientists (with different levels of questionnaire experience) translated an experimental questionnaire containing 45 items partly from scratch and partly neural machine translation output was post-edited (i.e., corrected). To explore how a large language model (LLM) compares, additionally a raw LLM version was generated. After a blind expert assessment of the translations, error counts across different categories were compared and showed that overall error counts were comparable but professional translators, drawing on their translation expertise, produced significantly fewer linguistic errors (linguistic conventions, style and register) in translations produced from scratch. Experience with the questionnaire genre tended to help both professional groups to improve their results. Both professional groups performed similarly when post-editing machine translated text. The LLM-generated translation contained fewer errors overall than the median of human translations.
In a survey experiment, we analyzed how different versions of a response scale affect the distributional characteristics and quality of the resulting survey data. Toward that end, we compared four different German-language versions of a five-point agree-disagree (AD) response scale, randomly assigned to four groups of online access panelists for a total of 15 items. The response scales were taken from different studies and varied in polarity or scale option intensity. Comparisons of frequency distributions as well as of response quality (response styles, response differentiation, response times) did not show any systematic differences between the response scales. Although there were no systematic differences in the overall sample, the few significant effects we found appeared to be largely due to the responses of participants with lower levels of education. Further research is warranted using non-access panel respondents and their perception of differently worded AD response scales, experimentally modified response scales, and other languages beyond the English-German pair.
Self-rated health (SRH) is a frequently used health measure in (cross-)national surveys. It is usually assessed with a single item, which differs in the wording and response format across surveys. In this paper, we compare four German-language five-point scale versions of self-rated health that vary in response scale labeling using web probing. The web survey (N=1,710) was conducted in 2019. Combining qualitative and quantitative methods, we assess how response scale labels affect response distributions of the SRH item and which health factors respondents consider when answering questions about their health. The main finding is that respondents refer to similar health aspects independent of the scale version they answered the question with. Self-reported health, however, varied across scales which might introduce a comparative bias. Using an unbalanced scale with three response options indicating good health led to a more positive self-assessment of health compared to the balanced scales with two positive, one neutral, and two negative scale points. We discuss practical implementations.
A highly controlled experimental setting using a sample of questions from the European Social Survey (ESS) and European Values Study (EVS) was used to test the effects of integrating machine translation and post-editing into the Translation, Review, Adjudication, Pretesting, and Documentation (TRAPD) approach in survey translation. Four experiments were conducted in total, two concerning the language pair English-German and two in the language pair English-Russian. The overall results of this study are positive for integrating machine translation and post-editing into the TRAPD process, when translating survey questionnaires. The experiments show evidence that in German and Russian languages and for a sample of ESS and EVS survey questions, the effect of integrating machine translation and post-editing on the quality of the review outputs-with quality understood as texts output with the fewest errors possible-can hardly be distinguished from the quality that derives from the setting with human translations only.
It is generally taken for granted that comparability in comparative research hinges, among others, on the quality of questionnaire translations. However, what do slight differences in translation mean for respondents’ answers? In this article, we look at a combination of quantitative evidence from split-ballot experiments and qualitative evidence from additional probing questions for three items that were translated according to different translation methods, resulting in different translations, e.g., for “our national way of life.” Two of the three items do not show any quantitative differences between translation versions when implemented in split-ballot experiments. However, using open-ended probing questions we delved deeper into the effects of different translation versions. This allowed us to show that different translations do indeed change respondent understanding. We suggest mechanisms that may lead to different translations (not) having an impact on the data, and we also try to align the results to the notion of equivalence/comparability in translation. Ultimately, we showcase the usefulness of web probing for exploring different translation understandings.
When it comes to quality in questionnaire translation and hence comparability in comparative research, the chosen translation method is crucial for the outcome. Few empirical studies compare different translation methods-a fact which is often deplored in the research community. To fill the gap, in this study, the team translation approach is compared against a simple back-translation approach. The starting point in both cases was the initial English-German translations of ISSP (International Social Survey Program) questions. The final translations from both approaches were assessed, with a focus on how translation issues, such as mistranslations or wording issues identified in the initial translations were addressed. While none of the twenty-nine issues in the initial translation were present in the final team translation version, twenty-two of these issues were still present in the final version after the back-translation approach. For a selected number of items, we also ran a split-ballot experiment in a web survey. Only five out of fifteen items (33 percent) that went into the experiment showed significant differences between the translations, and only one could clearly be attributed to remaining errors in the back-translation version. In sum, the final translation from the team approach clearly outperformed the final translation from the back-translation approach when it comes to text-based criteria (in particular, accuracy and fluency). The quantitative test showed that many translation issues (those remaining in the translation after the back translation step) had no effect on the estimates. Nevertheless, we ask respondents to put effort into survey responding; in the same vein, we as researchers should put effort in the survey experience by providing questions that are clearly worded and free of errors, which puts the team approach ahead of the back-translation approach.
This review summarizes the current state of the art of statistical and (survey) methodological research on measurement (non)invariance, which is considered a core challenge for the comparative social sciences. After outlining the historical roots, conceptual details, and standard procedures for measurement invariance testing, the paper focuses in particular on the statistical developments that have been achieved in the last 10 years. These include Bayesian approximate measurement invariance, the alignment method, measurement invariance testing within the multilevel modeling framework, mixture multigroup factor analysis, the measurement invariance explorer, and the response shift-true change decomposition approach. Furthermore, the contribution of survey methodological research to the construction of invariant measurement instruments is explicitly addressed and highlighted, including the issues of design decisions, pretesting, scale adoption, and translation. The paper ends with an outlook on future research perspectives.
The method of web probing integrates cognitive interviewing techniques into web surveys and is increasingly used to evaluate survey questions. In a usual web probing scenario, probes are administered immediately after the question to be tested (concurrent probing), typically as open-ended questions. A second possibility of administering probes is in a closed format, whereby the response categories for the closed probes are developed during previously conducted qualitative cognitive interviews. Using closed probes has several benefits, such as reduced costs and time efficiency, because this method does not require manual coding of open-ended responses. In this article, we investigate whether the insights gained into item functioning when implementing closed probes are comparable to the insights gained when asking open-ended probes and whether closed probes are equally suitable to capture the cognitive processes for which traditionally open-ended probes are intended. The findings reveal statistically significant differences with regard to the variety of themes, the patterns of interpretation, the number of themes per respondent, and nonresponse. No differences in number of themes across formats by sex and educational level were found.
AbstractThis chapter examines the technical challenges involved in translating and adapting measurement instruments, i.e., questionnaires, for migration research. The first part outlines good practices in questionnaire translation. In line with the technology-based focus of this book, the second part focuses on computerized surveys and on the interplay between technology, language, and culture. Frameworks from the software localization field are consulted and transferred to the context of computerized multilingual surveys with respect to their impact on source questionnaire design and on translation and adaptation. Real-life examples come from our own experiences in international and migration research, as well as from a review of existing reports and research articles. The main goal of this chapter is to raise awareness of the additional technology layer that impacts translation and adaptation, with an ultimate goal to improve translation and adaptation processes, and the outcomes of migration research.
Methodological studies usually gauge response quality in narrative open-ended questions with the proportion of nonresponse, response length, response time, and number of themes mentioned by respondents. However, not all of these indicators may be comparable and appropriate for evaluating open-ended questions in a cross-national context. This study assesses the cross-national appropriateness of these indicators and their potential bias. For the analysis, we use data from two web surveys conducted in May 2014 with 2,685 respondents and in June 2014 with 2,689 respondents and compare responses from Germany, Great Britain, the United States, Mexico, and Spain. We assess open-ended responses for a variety of topics (e.g., national identity, gender attitudes, and citizenship) with these indicators and evaluate whether they arrive at similar or contradictory conclusions about response quality. We find that all indicators are potentially biased in a cross-national context due to linguistic and cultural reasons and that the bias differs in prevalence across topics. Therefore, we recommend using multiple indicators as well as items covering a range of topics when evaluating response quality in open-ended questions across countries.
Survey documentation is an integral part of methodically sound survey research. These guidelines aim at providing the persons coordinating survey translations (e.g., researchers responsible for survey translation in a larger study, or those wishing to translate and adapt an existing instrument for their own research) with a framework within which they can plan and document survey translations both for internal as well as for external purposes (publications or technical reports). It summarizes di erent aspects of translation documentation and reviews elements to be included in such a documentation.
In 2012, a new question was introduced into the International Social Survey Program (ISSP). It asks respondents to indicate what they consider the best division of labor between men and women. In this paper, we propose to assess the validity and cross-national comparability of this new ISSP question, using a mixed-methods approach that combines quantitative experimental data with qualitative probing data. We implemented our experiment in non-probability online surveys in five countries, in which half of the respondents received the original ISSP question and the other half a variant with an additional category saying "Each family should find the solution which works best for them." In addition, the understanding of "individual solutions" was probed. We report on the understanding of this category.
Self-rated health (SRH) and subjective life expectancy (SLE) are widely used for understanding health and predicting mortality. However, what these items measure remains unclear, due to the lack of conceptual frameworks. We administered a web survey across the United States, Great Britain, Germany, Spain, and Mexico. The questionnaire included SRH and SLE, each immediately followed by a question that probed respondents’ thought processes. We examined the relationship between SRH and SLE, the response difficulty, and attributes that respondents considered for forming responses. Overall, SRH and SLE were moderately related, eliciting different information and varying in difficulty. Compared to SLE, SRH was perceived as easier but covered a narrower information spectrum. While illness and health behaviors were dominant attributes of SRH responses, family longevity history, life situations, and lack of control were additionally considered for SLE. When combined, SRH and SLE may capture a fuller range of attributes germane to health and mortality.