Recently, a novel computational model was proposed to investigate the processes and biases involved in human random number generation (RNG). This two-parameter model includes a repetition parameter and a side-switching parameter representing influences of the immediately preceding number on the choice of the next number. We propose two changes to the model. First, we replace the side-switching parameter with a more general and less task-dependent distance parameter, which accounts for the tendency to select subsequent numbers that tend to be either closer to or further away from the previous number on the selected response pad. Second, we extend the computational model by adding a third parameter to account for the human tendency to select subsequent numbers with greater probability the longer the respective number has not been previously selected, following the pattern of the well-known gambler's fallacy. This new "cycling" parameter takes into account the most recent and all previous selections. The generalized distance parameter, and particularly the new cycling parameter, improved the fit of the model to human-generated sequences and the rate of successful predictions of the next choice from 14.09% to 26.48%, significantly exceeding the expected chance value of 1/9 = 11.1%. Model-driven simulations also showed that the extended three-parameter model could better account for systematic patterns that can be observed in human RNG tasks. The improved model could be useful in many contexts where human biases in RNG tasks are analyzed.
The finding that repeating a statement typically increases its perceived validity is referred to as the truth effect. Research on individual differences in the magnitude of the effect and its correlates is scarce and has yielded rather mixed results. However, any search for replicable relations between the truth effect and other cognitive or personality variables is bound to fail if the truth effect cannot be measured reliably at the individual level and if the effect is not a stable phenomenon. We conducted two experiments investigating the split-half reliability and test-retest stability of the truth effect. To operationalize the magnitude of the effect, Experiment 1 used the between-items criterion and Experiment 2 used the within-items criterion of the truth effect (Dechêne et al., 2010). In both experiments, the truth effect's test-retest stability was found to be very low, probably due to a highly insufficient reliability of the measures that were used. While there may be meaningful and stable individual differences in the truth effect, our findings raise concerns about the usefulness of established indices and standard measures of the truth effect for personality and individual difference research.
Abstract: Indirect questioning techniques aim to provide more valid prevalence estimates for sensitive attributes than conventional direct questions. Despite being an important prerequisite for high estimation validity, indirect questioning techniques’ retest stability has rarely been addressed. For temporally stable attributes, high stability of both prevalence estimates and individual responses is expected; however, insufficient understanding of the instructions and random response behavior may compromise retest stability. The present study is the first to assess the retest stability of the Extended Crosswise Model (ECWM), a recent indirect questioning technique, and to compare it to the retest stability of a conventional direct question (DQ). With a retest interval of approximately 10 days, we asked N = 2,317 mothers twice whether they had smoked during a previous pregnancy. In both ECWM and DQ conditions, prevalence estimates were virtually identical over time, and most respondents answered consistently (ECWM: 89%, DQ: 95%). In the ECWM condition, inconsistent response behavior was slightly more prevalent and negatively associated with respondents’ education. However, as these effects were small, the retest stability of both ECWM and DQ in surveys on sensitive attributes was evaluated as high.
Indirect questioning techniques aim to provide more valid prevalence estimates for sensitive attributes than conventional direct questions. Despite being an important prerequisite for high estimation validity, indirect questioning techniques' retest stability has rarely been addressed. For temporally stable attributes, high stability of both prevalence estimates and individual responses is expected; however, insufficient understanding of the instructions and random response behavior may compromise retest stability. The present study is the first to assess the retest stability of the Extended Crosswise Model (ECWM), a recent indirect questioning technique, and to compare it to the retest stability of a conventional direct question (DQ). With a retest interval of approximately 10 days, we asked N = 2,317 mothers twice whether they had smoked during a previous pregnancy. In both ECWM and DQ conditions, prevalence estimates were virtually identical over time, and most respondents answered consistently (ECWM: 89%, DQ: 95%). In the ECWM condition, inconsistent response behavior was slightly more prevalent and negatively associated with respondents' education. However, as these effects were small, the retest stability of both ECWM and DQ in surveys on sensitive attributes was evaluated as high.
Indirect questioning techniques such as the randomized response technique aim to control social desirability bias in surveys of sensitive topics. To improve upon previous indirect questioning techniques, we propose the new Cheating Detection Triangular Model. Similar to the Cheating Detection Model, it includes a mechanism for detecting instruction non-adherence, and similar to the Triangular Model, it uses simplified instructions to improve respondents' understanding of the procedure. Based on a comparison with the known prevalence of a sensitive attribute serving as external criterion, we report the first individual-level validation of the Cheating Detection Model, the Triangular Model and the Cheating Detection Triangular Model. Moreover, the sensitivity and specificity of all models was assessed, as well as the respondents' subjective evaluation of all questioning technique formats. Based on our results, the Cheating Detection Triangular Model appears to be the best choice among the investigated indirect questioning techniques.
Whether and how well people can behave randomly is of interest in many areas of psychological research. The ability to generate randomness is often investigated using random number generation (RNG) tasks, in which participants are asked to generate a sequence of numbers that is as random as possible. However, there is no consensus on how best to quantify the randomness of responses in human-generated sequences. Traditionally, psychologists have used measures of randomness that directly assess specific features of human behavior in RNG tasks, such as the tendency to avoid repetition or to systematically generate numbers that have not been generated in the recent choice history, a behavior known as cycling. Other disciplines have proposed measures of randomness that are based on a more rigorous mathematical foundation and are less restricted to specific features of randomness, such as algorithmic complexity. More recently, variants of these measures have been proposed to assess systematic patterns in short sequences. We report the first large-scale integrative study to compare measures of specific aspects of randomness with entropy-derived measures based on information theory and measures based on algorithmic complexity. We compare the ability of the different measures to discriminate between human-generated sequences and truly random sequences based on atmospheric noise, and provide a systematic analysis of how the usefulness of randomness measures is affected by sequence length. We conclude with recommendations that can guide the selection of appropriate measures of randomness in psychological research.
The Randomized Response Technique (Warner, Journal of the American Statistical Association, 60, 63-69, 1965) has been developed to control for socially desirable responses in surveys on sensitive attributes. The Crosswise Model (CWM; Yu et al., Metrika, 67, 251-263, 2008) and its extension, the Extended Crosswise Model (ECWM; Heck et al., Behavior Research Methods, 50, 1895-1905, 2018), are advancements of the Randomized Response Technique that have provided promising results in terms of improved validity of the obtained prevalence estimates compared to estimates based on conventional direct questions. However, recent studies have raised the question as to whether these promising results might have been primarily driven by a methodological artifact in terms of random responses rather than a successful control of socially desirable responding. The current study was designed to disentangle the influence of successful control of socially desirable responding and random answer behavior on the validity of (E)CWM estimates. To this end, we orthogonally manipulated the direction of social desirability (undesirable vs. desirable) and the prevalence (high vs. low) of sensitive attributes. Our results generally support the notion that the ECWM successfully controls social desirability bias and is inconsistent with the alternative account that ECWM estimates are distorted by a substantial influence of random responding. The results do not rule out a small proportion of random answers, especially when socially undesirable attributes with high prevalence are studied, or when high randomization probabilities are applied. Our results however do rule out that random responding is a major factor that can account for the findings attesting to the improved validity of (E)CWM as compared with DQ estimates.
In self-reports, socially desirable responding threatens the validity of prevalence estimates for sensitive personal attitudes and behaviors. Indirect questioning techniques such as the crosswise model attempt to control for the influence of social desirability bias. The crosswise model has repeatedly been found to provide more valid prevalence estimates than direct questions. We investigated whether crosswise model estimates are also less susceptible to deliberate faking than direct questions. To this end, we investigated the effect of "fake good" instructions on responses to direct and crosswise model questions. In a sample of 1,946 university students, 12-month prevalence estimates for a sensitive road traffic behavior were higher and thus presumably more valid in the crosswise model than in a direct question. Moreover, "fake good" instructions severely impaired the validity of the direct questioning estimates, whereas the crosswise model estimates were unaffected by deliberate faking. Participants also reported higher levels of perceived confidentiality and a lower perceived ease of faking in the crosswise model compared to direct questions. Our results corroborate previous studies finding the crosswise model to be an effective tool for counteracting the detrimental effects of positive self-presentation in surveys on sensitive issues.
Non-randomized response techniques (NRRTs) such as the crosswise model and the triangular model (CWM and TRM; Yu et al. Metrika, 67, 251-263, 2008) have been developed to control for socially desirable responding in surveys on sensitive personal attributes. We present the first study to directly compare the validity of the CWM and TRM and contrast their performance with a conventional direct questioning (DQ) approach. In a paper-pencil survey of 1382 students, we obtained prevalence estimates for two sensitive attributes (xenophobia and rejection of further refugee admissions) and one nonsensitive control attribute with a known prevalence (the first letter of respondents' surnames). Both NRRTs yielded descriptively higher prevalence estimates for the sensitive attributes than DQ; however, only the CWM estimates were significantly higher. We attribute the higher prevalence estimates for the CWM to its response symmetry, which is lacking in the TRM. Only the CWM provides symmetric answer options, meaning that there is no "safe" alternative respondents can choose to distance themselves from being carriers of the sensitive attribute. Prevalence estimates for the nonsensitive control attribute with known prevalence confirmed that neither method suffered from method-specific bias towards over- or underestimation. Exploratory moderator analyses further suggested that the sensitive attributes were perceived as more sensitive among politically left-oriented than among politically right-oriented respondents. Based on our results, we recommend using the CWM over the TRM in future studies on sensitive personal attributes.
The crosswise model is an indirect questioning technique designed to control for socially desirable responding. Although the technique has delivered promising results in terms of improved validity in survey studies of sensitive issues, recent studies have indicated that the crosswise model may sometimes produce false positives. Hence, we investigated whether an insufficient understanding of the crosswise model instructions might be responsible for these false positives and whether ensuring a deeper understanding of the model and surveying more highly educated respondents reduces the problem of false positives. To this end, we experimentally manipulated the amount of information respondents received in the crosswise model instructions. We compared a crosswise model condition with only brief instructions and a crosswise model condition with detailed instructions and additional comprehension checks. Additionally, we compared the validity of crosswise model estimates between a higher- and a lower-educated subgroup of respondents. Our results indicate that false positives among highly educated respondents can be reduced when detailed instructions and comprehension checks are employed. Since false positives can also occur in direct questioning, they do not appear to be a specific flaw of the crosswise model, but rather a more general problem of self-reports on sensitive topics. False negatives were found to occur for all questioning techniques, but were less prevalent in the crosswise model than in the direct questioning condition. We highlight the importance of comprehension checks when applying indirect questioning and emphasize the necessity of developing instructions suitable for lower-educated respondents.
The finding that repeating a statement typically increases its perceived truth has been referred to as the truth effect. Previous research has found that warning participants about the truth effect can successfully reduce, but not eliminate the effect. We used a multinomial modeling approach to investigate how warnings affect the cognitive processes that are assumed to underlie judgments of truth. In a laboratory experiment (N = 167), half of the participants were warned about the truth effect before judging the truth of repeated and new statements. Importantly, whereas half of the presented statements were of relatively unknown validity, participants could likely identify the correct truth status for the other half of the statements by drawing on stored knowledge. Multinomial modeling analyses revealed that warning instructions did not affect the retrieval of knowledge or participants' guessing behavior relative to a control condition. Instead, warned participants exhibited a significantly reduced tendency to rely on experiential information such as processing fluency when judging a repeated statement's truth. However, this was only the case for statements for which participants held relevant knowledge. These results are consistent with the notion that it is possible to discount metacognitive experiences such as processing ease when their informational value is questioned. Specifically, our findings suggest that people are less likely to base their judgments of truth on experiential information and metacognitive experiences induced by repetition if (a) they are warned about the deceptive power of repetition, and (b) other valid cues are available to inform their judgments.
Indirect questioning techniques such as the crosswise model aim to control for socially desirable responding in surveys on sensitive personal attributes. Recently, the extended crosswise model has been proposed as an improvement over the original crosswise model. It offers all of the advantages of the original crosswise model while also enabling the detection of systematic response biases. We applied the extended crosswise model to a new sensitive attribute, campus islamophobia, and present the first experimental investigation including an extended crosswise model, and a direct questioning control condition, respectively. In a paper-pencil questionnaire, we surveyed 1,361 German university students using either a direct question or the extended crosswise model. We found that the extended crosswise model provided a good model fit, indicating no systematic response bias and allowing for a pooling of the data of both groups of the extended crosswise model. Moreover, the extended crosswise model yielded significantly higher estimates of campus Islamophobia than a direct question. This result could either indicate that the extended crosswise model was successful in controlling for social desirability, or that response biases such as false positives or careless responding have inflated the estimate, which cannot be decided on the basis of the available data. Our findings highlight the importance of detecting response biases in surveys implementing indirect questioning techniques.
Multiple-choice tests are frequently used in personnel selection contexts to measure knowledge and abilities. Option weighting is an alternative multiple-choice scoring procedure that awards partial credit for incomplete knowledge reflected in applicants' distractor choices. We investigated whether option weights should be based on expert judgment or on empirical data when trying to outperform conventional number-right scoring in terms of reliability and validity. To obtain generalizable results, we used repeated random sub-sampling validation and found that empirical option weighting, but not expert option weighting, increased the reliability of a knowledge test. Neither option weighting procedure improved test validity. We recommend to improve the reliability of existing ability and knowledge tests used for personnel selection by computing and publishing empirical option weights.
Testwiseness may introduce construct-irrelevant variance to multiple-choice test scores. Presenting response options sequentially has been proposed as a potential solution to this problem. In an experimental validation, we determined the psychometric properties of a test based on the sequential presentation of response options. We created a strong validity criterion by providing participants with different levels of information on a domain about which they had no prior knowledge. Test takers’ capability of guessing the correct answer was strongly reduced by the sequential test format, whereas sequential test scores were as valid and reliable as multiple-choice test scores. We concluded that the sequential presentation of response options should be investigated more closely as a viable alternative to the traditional multiple-choice test format.
To avoid social disapproval in studies on prejudice against women leaders, participants might provide socially desirable rather than truthful responses. Using the Crosswise Model, an indirect questioning technique that can be applied to control for socially desirable responding, we investigated the prevalence of prejudice against women leaders in a German university community sample of 1529 participants. Prevalence estimates that were based on an indirect question that guaranteed confidentiality of responses were higher than estimates that were based on a direct question. Prejudice against women leaders was thus shown to be more widespread than previously indicated by self-reports that were potentially biased by social desirability. Whereas women showed less prejudice against women leaders, their responses were actually found to be more biased by social desirability, as indicated by a significant interaction between questioning technique and participants’ gender. For men, prejudice estimates increased only slightly from 36% to 45% when an indirect question was used, whereas for women, prejudice estimates almost tripled from 10% to 28%. Whereas women were particularly hesitant to provide negative judgments regarding the qualities of women leaders, prejudice against women leaders was more prevalent among men even when gender differences in social desirability were controlled. Taken together, the results highlight the importance of controlling for socially desirable responding when using self-reports to investigate the prevalence of gender prejudice.
For decades, sequential lineups have been considered superior to simultaneous lineups in the context of eyewitness identification. However, most of the research leading to this conclusion was based on the analysis of diagnosticity ratios that do not control for the respondent's response criterion. Recent research based on the analysis of ROC curves has found either equal discriminability for sequential and simultaneous lineups, or higher discriminability for simultaneous lineups. Some evidence for potential position effects and for criterion shifts in sequential lineups has also been reported. Using ROC curve analysis, we investigated the effects of the suspect's position on discriminability and response criteria in both simultaneous and sequential lineups. We found that sequential lineups suffered from an unwanted position effect. Respondents employed a strict criterion for the earliest lineup positions, and shifted to a more liberal criterion for later positions. No position effects and no criterion shifts were observed in simultaneous lineups. This result suggests that sequential lineups are not superior to simultaneous lineups, and may give rise to unwanted position effects that have to be considered when conducting police lineups.
In multiple-choice tests, the quality of distractors may be more important than their number. We therefore examined the joint influence of distractor quality and quantity on test functioning by providing a sample of 5,793 participants with five parallel test sets consisting of items that differed in the number and quality of distractors. Surprisingly, we found that items in which only the one best distractor was presented together with the solution provided the strongest criterion-related evidence of the validity of test scores and thus allowed for the most valid conclusions on the general knowledge level of test takers. Items that included the best distractor produced more reliable test scores irrespective of option number. Increasing the number of options increased item difficulty, but did not increase internal consistency when testing time was controlled for.
The number of respondents who access web surveys on a mobile device (smartphone or tablet) has been increasing rapidly over the last few years. Compared with desktop computers, mobile devices have smaller screens, different input options, and are used in a larger variety of locations and situations. The suspicion that the quality of data may suffer when online respondents use mobile devices has stimulated a growing body of research, which has mainly focused on paradata and web survey design. To investigate whether the respondents’ device affects the quality of web survey data, we examined the responses of 1,826 mobile-device and desktop participants in a political online survey that asked questions about the 2013 German federal election. To determine the reliability and validity of data submitted via mobile devices, we determined the consistency of the participants’ responses across questions and validated the responses against various internal and external criteria. Replicating previous findings, mobile-device respondents were younger and more likely to be female, and they produced higher dropout rates and longer completion times than desktop respondents. However, data produced by respondents using mobile devices were as consistent, reliable, and valid as data produced by respondents using desktop computers. These findings contradict the notion that mobile-device users compromise the reliability and validity of data collected online and suggest that researchers do not necessarily need to be afraid of the participation of mobile-device respondents in web surveys.