AbstractReplication studies are recognized as essential to the scientific process. Numerous measures have been developed to quantify replication success. Most measures were developed for post hoc replications, in which the primary study has been conducted and sometimes assumed to show a specific result (e.g. statistical significance). Consequently, methodological studies have focused on evaluating replication success measures for those replications. However, recent work emphasizes the value of prospective replications, in which primary and replication studies are planned simultaneously. Such replications allow researchers to control study characteristics and thus investigate which characteristics cause effect heterogeneity. This study provides replication success measures for prospective replications, and guidelines for choosing between them. We present a taxonomy of measures based on research questions they address and evaluate existing frequentist and Bayesian approaches for their applicability to prospective replications. We illustrate their application using an example from social psychology. In simulations, we compare the statistical properties of measures that aim at the same research question. Results indicate that there is almost always a trade-off between error types. Thus, no single measure emerged as always clearly superior. We highlight the assumptions and strengths of each measure and offer recommendations for choosing a measure based on replication goals.
Although previous research has described that intervention effects vary across replication studies, less effort has been devoted to identifying causes of this effect heterogeneity with regard to differences in study implementations. However, knowing in which way study characteristics (such as population, measurement instrument, setting, or treatment implementation) impact the study results may not only help to better infer the impact of research practices but also provide evidence for theory building. Causal effects can be easily identified if all study characteristics but the one under investigation are kept constant across two studies. This is, however, not always possible in practice and unintended differences between the studies to be compared may confound the relationship of the study characteristic of interest and the treatment effect. In this article, we present a statistical approach for identifying effects of study characteristics on study-specific treatment effects from randomized experiments in cases in which unintended differences in study implementation across studies cannot be prevented. We present formal definitions of the causal effects of interest, identification assumptions, and derive respective causal estimands. The assumptions can more likely be fulfilled in prospective replication studies or many-lab studies, where researchers have more control over design and measurement of covariates in both studies. We also provide ways to test the assumptions and illustrate consequences of not meeting the assumptions. The approach is illustrated using an empirical example on the imagined intergroup contact effect in social psychology. (PsycInfo Database Record (c) 2026 APA, all rights reserved).
We propose a latent trait model for the responses and response times on tests that separates capability from persistence. Core of the model is a race between a diffusion process and a censoring process. The diffusion process represents item-level cognitive processing and determines the processing time of a test taker. The censoring process sets the maximal time a test taker is willing to invest into an item. If the processing time is shorter than the maximal time, the response is generated by the diffusion process; otherwise, the response is generated differently. In the first version of the model, the response is generated by a random guess. In the second version of the model, the response is determined by the actual level of the diffusion process. We relate the diffusion process to the capability and response caution of a test taker and the censoring process to his willingness to invest time. Similar to models for rapid guessing, the proposed models take account of disengaged responding, but allow for individual differences in persistence and informed guessing. The model also provides a mathematical specification of persistence. In a simulation study, we investigate model fitting by marginal maximum likelihood estimation. We also apply the model to two empirical data sets.
Avoidance behavior is a central transdiagnostic coping strategy in psychopathology and contributes to the maintenance and severity of mental disorders. We present a general framework for the standardized assessment of actual avoidance behavior that may be applied to several clinical phenomena. This resource-efficient approach balances ecological validity with methodological rigor. Moreover, the framework integrates the focus on measuring natural behavior found in Behavioral Avoidance Tests (BATs) with systematic computerized response time measurement from Approach and Avoidance Tasks (AATs) along with recent developments in psychometric modeling that allow for a methodologically and theoretically robust assessment. Providing a comprehensive theory, assessment design, and analysis models, the framework is the first to explicitly assess reactive avoidance (escaping aversive stimuli) and automatic or reflective anticipating avoidance (preventing aversive exposure) in computerized testing. We illustrate and discuss how the framework can be adapted to different clinical settings and utilized for large-scale data collection.
Causal inference of the effect of a treatment on an outcome is usually done on the group or subgroup level. Although the typically reported average treatment effect may be positive, suggesting that the treatment is effective, at the level of individual participants, the treatment effect may be zero or even negative-the treatment may even harm some individuals. For making decisions on whether a specific person should take the treatment, information on the probability of benefiting or being harmed by the treatment for a single person is necessary. Estimating the probability of possible benefit or possible harm for a person involves counterfactual reasoning and thus strong assumptions about unobservable events. Precise statements about the causal effect of a treatment for an individual are only possible to a limited extent. This tutorial introduces the method of causal attribution to psychology that allows for estimating bounds in which the probability of benefit or harm lies. These bounds can be calculated using data at the group level, which can come from experimental or observational studies. The bounds can be narrowed by simultaneously using data from both randomized trials and observational studies and by using information from pretreatment covariates. R functions are provided for calculating these bounds from binary data and are illustrated with examples from basic laboratory research and clinical intervention research. (PsycInfo Database Record (c) 2026 APA, all rights reserved).
With transcranial direct-current stimulation's (tDCS) popularity both in motor learning research and as a commercial product, it is important that the quality of evidence on its effectiveness be evaluated. Special attention should be paid to meta-analyses, as they usually have a large impact on research and clinical practice. The aim of this verification report was to gain insight on the methodological quality of meta-analyses estimating the effect of tDCS on motor learning. To that end, we verified the methodology of three meta-analyses with respect to reproducibility as the main focus, and reporting quality and publication bias control as secondary aspects. The three meta-analyses we verified largely adhered to PRISMA reporting guidelines and reported the primary effect sizes and sampling variances/confidence intervals they calculated, enabling successful reproductions of pooled effect size estimates. However, akin to previous meta-research with similar aims, we found the methods and results sections of the meta-analyses to be severely underreported, which compromises the ability to judge the soundness of the methodological procedure adopted as well as its reproducibility. While publication bias detection methods were applied, the approaches chosen do not allow for well informed decisions about the presence or extent of publication bias. These results reemphasise the need to transparently report methods in meta-analyses and to meticulously evaluate their quality before and after publication.
Avoidance is central in selective sound intolerance (misophonia), yet lacks objective behavioral assessment. We present the Misophonic Behavioral Avoidance Test (M-BAT), a standardized computerized task with two subtests to measure reactive (sound endurance) and anticipating avoidance tendency via behavioral indicators and response times (RTs). Different than most behavioral avoidance tests, the M-BAT can be easily applied online to large samples. In a validation study on N = 381 adults with misophonic symptoms, we fitted Bayesian Item Response Theory (IRT) models to endurance times, avoidance indicators and RTs to evaluate reliability, measurement invariance (MI), and several facets of validity. Unidimensional models fitted subtests well with high reliability (ρEAP ≈ .88/.94), and MI across gender, misophonia severity and age was broadly supported, despite localized differential item functioning. Sound endurance and anticipating avoidance tendency were moderately related (r = -.30) and showed theory-consistent, moderate associations with self-reported misophonic symptoms including avoidance, distinguished clinical from subclinical misophonia groups, and predicted everyday behavior (daily escape and prevention). We found little evidence that contextual factors, such as the stimuli presentation or the individuality of stimuli influenced test responses, however, findings demonstrated influences due to participants high test motivation. Challenges in establishing discriminant, internal, and incremental validity remain. Overall, the M-BAT provides reliable, easily applicable, and objective behavioral measures that complement questionnaires and psychoacoustic ratings, supporting clinical research and treatment evaluation.
Valid assessments with achievement tests hinge on test-takers being motivated to take the test. Existing latent trait models attempt to disentangle competence and motivational influences, but have theoretical limitations. We propose a single-process accumulator model based on the idea that test-takers accumulate information to solve an item at a continuously decreasing rate. A correct response is generated once the information exceeds a solution threshold. The model incorporates disengagement which is governed by the solution process. Once the accumulation rate falls below a critical level, test-takers stop working on the item. Due to the computational intensity of an analytic solution, we compare maximum likelihood, neural network and Bayesian estimators that use a simulation-based likelihood in a simulation study. Using two empirical examples, the model demonstrates good fit to accuracy and response times of individual items and is able to capture various forms of dependencies between accuracy and response times, including non-linear dependencies.
OBJECTIVE:Blended care (BC), the integration of Internet-based interventions into psychotherapy (PT), is thought of as a promising approach to enhance PT's effectiveness and efficiency. This randomized controlled trial aimed to investigate the effectiveness as well as the implementation and usage of BC with transdiagnostic online modules compared to PT in routine care in Germany. Routine outpatient PT is delivered by licensed psychotherapists across different therapeutic orientations (cognitive behavioral therapy, psychodynamic, systemic), with variable treatment lengths and procedures. METHOD:Psychotherapists in routine outpatient care recruited 1,159 patients who were randomized to BC or PT. The primary outcome was self-reported mental distress (the composite of anxiety and depression); secondary outcomes included self-reported satisfaction with life, level of functioning, eating pathology, and drug and alcohol use, as well as therapist-rated severity and changes. Outcomes were measured at baseline, 6 weeks, 12 weeks, 6 months, and 12 months. We examined whether BC and PT groups changed differently over time using linear mixed models. We also investigated differences in sessions and terminations and report usage metrics of the BC platform. RESULTS:Contrary to our hypotheses, we did not find differences between BC and PT in outcomes, including anxiety, depression, satisfaction with life, level of functioning, eating pathology, alcohol and drug use, therapist-rated severity, and satisfaction with treatment at 6 months postrandomization (all p > .05). BC and PT did not differ in the number of sessions or terminations. Regarding usage of the BC platform, 534 patients (91.6%) received at least one online chapter, with M = 7.26 (SD = 7.01) of a total of 39 online chapters assigned on average, and patients logged in M = 19.73 (SD = 24.66) times and spent M = 367.14 (SD = 338.27) minutes on the platform. CONCLUSIONS:In this real-world application of BC, therapists had considerable flexibility in implementing BC and integrating Internet-based interventions with sessions. Our findings suggest that the benefits observed in more structured BC setups may not fully translate to a flexible and transdiagnostic BC setup in routine care, potentially due to variations in implementation and adherence. (PsycInfo Database Record (c) 2026 APA, all rights reserved).
This paper introduces a dataset from a validation study of two psychometric models, one on the intraindividual speed-ability relationship and the other one on persistence. It includes responses, response times, and action sequences from N = 1244 participants who completed a matrix reasoning test under two experimental conditions, one being speeded and one being non-speeded. Additionally, it includes measures on motivational disposition, current motivation, effort, and concentration. Collected online via Prolific, the data is freely available at OSF (https://osf.io/9j6hm/). This dataset may aid in the development and validation of psychometric models on response processes as well as the investigation of test-taking behavior.
In what aspects do replication studies differ from their primary studies? This question is central for providing insights into the reasons for the nonreplicability of psychological effects. So far, research on potential explanations for the nonreplicability of effects has mainly focused on publication bias and methodological challenges related to measurement error or statistical inference. The recently developed causal-replication framework directs attention toward controlling for differences in study characteristics, including variations in treatment conditions, outcome measures, recruitment, causal estimates, time, location, population, and setting. To contribute to this aim, we conducted a systematic literature review to investigate the design practices of current replication studies. We preregistered the assessment of study characteristics in a detailed review protocol and investigated the available information and intended or unintended variations across primary and replication studies. To do this, we compiled a database of studies that aimed to replicate a causal effect of a clearly stated primary study and that were published in impactful social- and cognitive-psychological journals between January 2017 and August 2022. Our review results highlight that compared with the primary study, authors of replication studies predominantly focus on controlling specific study characteristics in (i.e., methods, procedures, analysis) while often neglecting other study characteristics, such as population or setting. Furthermore, the results indicate that in most replication studies, multiple study characteristics are varied in the study comparison or are insufficiently reported. Accordingly, we discuss prevalent variations, reporting standards, and strategies for planning future replication studies.
Background Results on parental burden during the COVID-19 pandemic are predominantly available from nonrepresentative samples. Although sample selection can significantly influence results, the effects of sampling strategies have been largely underexplored. Objective This study aimed to investigate how sampling strategy may impact study results. Specifically, we aimed to (1) investigate if outcomes on parental health and child maltreatment during the COVID-19 pandemic from a convenience sample differ from those of a specific representative sample and (2) investigate reasons for differences in the results. Methods In 2020, we simultaneously conducted 2 studies: (1) a web-based survey using a convenience sample of 4967 parents of underage children, primarily recruited via social media, and (2) a study using a quota sample representative of the German adult population with underage children (N=1024), recruited through a combination of telephone interviews and computer-assisted web interviews. In both studies, the same questionnaire was used. To evaluate the impact of sampling, we compared the results on outcomes (parental stress, subjective health, parental mental health, general stress, pandemic-related stress, and the occurrence of child maltreatment) between the 2 samples. To explain differences in the results between the 2 studies, we controlled for sociodemographic data, parent-related risk factors, and COVID-19–related experiences. Results Compared to parents from the quota sample, parents from the convenience sample reported significantly more parental stress (η2=0.024); decreased subjective health (η2=0.016); more anxiety and depression symptoms (η2=0.055); more general stress (η2=0.044); more occurrences of verbal emotional abuse (VEA; φ=0.12), witnessing domestic violence (WDV; φ=0.13), nonverbal emotional abuse (NEA; φ=0.03), physical abuse (φ=0.10), and emotional neglect (φ=0.06); and an increase of child maltreatment (VEA: exp(B)=2.95; WDV: exp(B)=3.19; NEA: exp(B)=1.65). Sociodemographic data, parent-related risk factors, and COVID-19–related experiences explained the differences in parental stress (remaining difference between samples after controlling for covariates: η2=0.002) and subjective health (remaining difference between samples after controlling for covariates: η2=0.004) and partially explained differences in parental mental health (remaining: η2=0.016), general stress (remaining: η2=0.014), and child maltreatment (remaining: VEA: exp(B)=2.05 and WDV: exp(B)=2.02) between the 2 samples. The covariates could not explain the difference in NEA (exp(B)=1.70). We discuss further factors that may explain the unexplained differences. Conclusions Results of studies can be heavily impacted by the sampling strategy. Scientists are advised to collect relevant explaining variables (covariates) that are possibly related to sample selection and the outcome under investigation. This approach enables us to identify the individuals to whom the results apply and to combine findings from different studies. Furthermore, if data on the distribution of these explanatory variables in the population are available, it becomes possible to adjust for sample selection bias.
In this article, we propose a series of latent trait models for the responses and the response times on low stakes tests where some test takers respond preliminary without making full effort to solve the items. The models consider individual differences in capability and persistence. Core of the models is a race between the solution process and a process of disengagement that interrupts the solution process. The different processes are modeled with the linear ballistic accumulator model. Within this general framework, we develop different model variants that differ in the number of accumulators and the way the response is generated when the solution process is interrupted. We distinguish no guessing, random guessing and informed guessing where the guessing probability depends on the status of the solution process. We conduct simulation studies on parameter recovery and on trait estimation. The simulation study suggests that parameter values and traits can be recovered well under certain conditions. Finally, we apply the model variants to empirical data.
The speed-accuracy tradeoff (SAT), where increased response speed often leads to decreased accuracy, is well established in experimental psychology. However, its implications for psychological assessments, especially in high-stakes settings, remain less understood. This study presents an experimental approach to investigate the SAT within a high-stakes spatial ability assessment. By manipulating instructions in a within-subjects design to induce speed variations in a large sample (N = 1,305) of applicants for an air traffic controller training program, we demonstrate the feasibility of manipulating working speed. Our findings confirm the presence of the SAT for most participants, suggesting that traditional ability scores may not fully reflect performance in high-stakes assessments. Importantly, we observed individual differences in the SAT, challenging the assumption of uniform SAT functions across test takers. These results highlight the complexity of interpreting high-stakes assessment outcomes and the influence of test conditions on performance dynamics. This study offers a valuable addition to the methodological toolkit for assessing the intraindividual relationship between speed and accuracy in psychological testing (including SAT research), providing a controlled approach while acknowledging the need to address potential confounders. Future research may apply this method across various cognitive domains, populations, and testing contexts to deepen our understanding of the SAT's broader implications for psychological measurement.
The deployment of statistical models—such as those used in item response theory—necessitates the use of indices that are informative about the degree to which a given model is appropriate for a specific data context. We introduce the InterModel Vigorish (IMV) as an index that can be used to quantify accuracy for models of dichotomous item responses based on the improvement across two sets of predictions (i.e., predictions from two item response models or predictions from a single such model relative to prediction based on the mean). This index has a range of desirable features: It can be used for the comparison of non-nested models and its values are highly portable and generalizable. We use this fact to compare predictive performance across a variety of simulated data contexts and also demonstrate qualitative differences in behavior between the IMV and other common indices (e.g., the AIC and RMSEA). We also illustrate the utility of the IMV in empirical applications with data from 89 dichotomous item response datasets. These empirical applications help illustrate how the IMV can be used in practice and substantiate our claims regarding various aspects of model performance. These findings indicate that the IMV may be a useful indicator in psychometrics, especially as it allows for easy comparison of predictions across a variety of contexts.
Background Refugee populations have an increased risk for mental disorders, such as depression, anxiety, and posttraumatic stress disorders. Comorbidity is common. At the same time, refugees face multiple barriers to accessing mental health treatment. Only a minority of them receive adequate help. The planned trial evaluates a low-threshold, transdiagnostic Internet-based treatment. The trial aims at establishing its efficacy and cost-effectiveness compared with no treatment. Methods N = 131 treatment-seeking Arabic- or Farsi-speaking patients, meeting diagnostic criteria for a depressive, anxiety, and/or posttraumatic stress disorder will be randomized to either the intervention or the waitlist control group. The intervention group receives an Internet-based treatment with weekly written guidance provided by Arabic- or Farsi-speaking professionals. The treatment is based on the Common Elements Treatment Approach (CETA), is tailored to the individual patient, and takes 6–16 weeks. The control group will wait for 3 months and then receive the Internet-based treatment. Discussion The planned trial will result in an estimate of the efficacy of a low-threshold and scalable treatment option for the most common mental disorders in refugees. Trial registration German Registry for Clinical Trials DRKS00024154. Registered on February 1, 2021.
Free AccessEditorialInnovations in Exploring Sequential Process DataEsther Ulitzsch, Qiwei He, and Steffi PohlEsther UlitzschEsther Ulitzsch, University of Oslo, Centre for Educational Measurement, Gaustadalleen 21, 0349 OsloNorway[email protected]Centre for Educational Measurement, University of Oslo, NorwayCentre of Research on Equality in Education, University of Oslo, Norway, Qiwei HeData Science and Analytics Program, Georgetown University, Washington DC, USA, and Steffi PohlMethods and Evaluation/Quality Assurance, Freie Universität Berlin, GermanyPublished Online:April 24, 2024https://doi.org/10.1027/2151-2604/a000560PDF ToolsAdd to favoritesDownload CitationsTrack Citations Cite ShareShare onFacebookTwitterLinkedInRedditE-Mail SectionsMoreThe widespread employment of computerized psychological assessment allows to routinely store subjects' interactions with the administered items in log files. From these, researchers can extract process data that are indicative of how subjects engaged with questionnaire items and approached tasks in cognitive assessments. Examples for process data are response times, answer changes, but also more complex data types taking the form of time-stamped sequences of events such as keystrokes, clickstreams, mouse movements, or navigation behaviors, just to name a few.Process data allow to move from investigating which response options respondents chose on a questionnaire or whether examinees solved a given task to how they chose the response options and solved the tasks. Having data that signal differences in cognitive and behavioral processes underlying the final product of psychological assessments facilitates, among others, to identify aberrant responding such as cheating or disengagement (e.g., Pokropek, 2016; Ulitzsch et al., 2020; van der Linden & Belov, 2023), allows to uncover and understand group differences in response and solution processes (Eichmann et al., 2020; He & von Davier, 2016; Zhu et al., 2016), and can be a central aspect of validity arguments, supporting conclusions on whether or not respondents interacted with a given measurement instrument as intended (e.g., Engelhardt & Goldhammer, 2019; Greiff et al., 2015).Having a long-standing tradition going back way before the era of computerized assessment, so far, response times (i.e., the total time required to respond to a questionnaire item or to solve a task) have received by far the most attention in the psychometric and psychological literature (e.g., Baxter, 1941; Gates, 1924; Iseler, 1970; Sternberg, 1977; Thorndike, 1914; Thursthone, 1937; Tresselt & Mayzner, 1965). Sequential process data such as keystrokes, clickstreams, and mouse movements, in contrast, pose a comparably new source of data that is much less understood and the potential of which is yet not fully harnessed. These data support much more detailed documentation of response processes above and beyond the mere time required for providing a response or solving a task. It is, however, not straightforward to extract meaningful information from these usually vast and unstructured data.The last decade has seen a rapid increase of methodology to tame and leverage this rich but unstructured source of data (He & von Davier, 2016; He et al., 2021; Man et al., 2022; Tang et al., 2020, 2021; Ulitzsch et al., 2021; Ulitzsch et al., 2021, 2022, 2023; Wang et al., 2023; Vista et al., 2017; Zhu et al., 2016). This topical issue highlights the potential and plurality of this still-developing field of research. It illustrates the large variety of sequential process data that can be obtained from computerized questionnaire and test administrations, the broad array of research questions and measurement challenges that can be tackled by making use of their rich information and sequential nature, and how these data can be used for theory building and evaluation alike. It further showcases how these complex data types can be synthesized through inherently contrasting viewpoints; drawing on exploratory machine learning and data mining techniques on the one hand and perceiving them through the lens of cognitive theory on the other.First, the articles illustrate how sequential process data can be leveraged for addressing different research questions in psychological and educational assessment. Process data are shown to be very useful for identifying possible validity and data quality threats in psychological assessments, such as identifying anomalous test takers (Bulut et al., 2024) or inattentive responding (Pokropek et al., 2024). Sequential process data also facilitate a deeper understanding of the constructs to be measured. In the present topical issue, they are used for expanding the theoretical understanding of the VOTAT (vary-one-thing-at-a-time) strategy when assessing problem-solving (Stadler et al., 2024) as well as for investigating differences in problem-solving strategies between groups of test takers (Zhang et al., 2024).Second, the featured collection of articles focuses on different types of data, ranging from mouse movements in survey questionnaires (Pokropek et al., 2024) to action sequences from interactive tasks and simulated environments from low-stakes cognitive assessment (Bulut et al., 2024; Stadler et al., 2024; Zhang et al., 2024). Each data type carries unique information and conveys a different structure for which targeted analyses need to be developed.Third, the four articles of this topical issue highlight different sequence-based methods that can be employed to extract meaningful patterns from sequential process data. Given the structural similarities between languages and sequential process data, many of the employed methods originate in natural language processing and text mining. Zhang et al. (2024) employed multidimensional scaling (as in Tang et al., 2020) and sequence-to-sequence (seq2seq) autoencoders (as in Tang et al., 2021). Both techniques support extracting latent features from action sequences, providing a parsimonious summary of the sequences' essential information. Subsequently, the relationship between the extracted latent features and covariates can be inspected to understand group differences in essential aspects of problem-solving in interactive tasks and simulated environments. Bulut et al. (2024) transformed action sequences into contextual embeddings using the Bidirectional Encoder Representations from Transformers model. Contextual embeddings provide numerical representations of performed actions that take the context in which the action occurred (i.e., the preceding and subsequent actions) into account. Subsequently, Bulut et al. (2024) used unsupervised machine learning (i.e., Isolation Forest) to detect anomalous action sequences. Using experimental manipulation, Pokropek et al. (2024) generated survey conditions either fostering or curbing inattentive response behavior. Then, using the experimental conditions as labels, Pokropek et al. (2024) trained neural networks models designed for sequential data such as Gated Recurrent Unit (Cho et al., 2014) and Bidirectional Long Short-term Memory (BiLSTM; Schuster & Paliwal, 1997) to deep-learn how mouse cursor trajectories differ between attentive and inattentive respondents. Stadler et al. (2024) used Markov models for fine-grained investigations of the VOTAT strategy in problem-solving tasks.Finally, the articles show that both exploratory and confirmatory perspectives can be taken on sequential process data. While the majority of sequential process data research and articles in this collection (i.e., Bulut et al., 2024; Pokropek et al., 2024; Zhang et al., 2024) make use of inherently exploratory machine learning and natural language processing techniques to uncover patterns from data, Stadler et al. (2024) illustrate how sequential process data can be aggregated and behavioral patterns contextualized by drawing on cognitive theories on problem-solving. Conversely, sequential process data documenting key parts of the problem-solving process can be employed to evaluate such theories.On a general note, the articles highlight how psychological assessment may profit from blending theory-grounded psychometrics with data-driven techniques originating in machine learning, natural language processing, and text mining (von Davier, 2017). Future research will need to focus on sharpening the link between theory and data. The papers in this special issue take some steps towards this goal.ReferencesBaxter, B. (1941). An experimental analysis of the contributions of speed and level in an intelligence test. Journal of Educational Psychology, 32(4), 285–296. 10.1037/h0061115 First citation in articleCrossref, Google ScholarBulut, O., Gorgun, G., & He, S. (2024). Unsupervised anomaly detection in sequential process data: Insights from PIAAC-Problem-solving tasks. Zeitschrift für Psychologie, 232(2), 74–94. 10.1027/2151-2604/a000558 First citation in articleLink, Google ScholarCho, K., Van Merriënboer, B., Gulcehre, C., Bahdanau, D., Bougares, F., Schwenk, H., & Bengio, Y. (2014). Learning phrase representations using RNN encoder-decoder for statistical machine translation. arXiv: 1406.1078. 10.48550/arXiv.1406.1078 First citation in articleCrossref, Google ScholarEichmann, B., Goldhammer, F., Greiff, S., Brandhuber, L., & Naumann, J. (2020). Using process data to explain group differences in complex problem solving. Journal of Educational Psychology, 112(8), 1546–1562. 10.1037/edu0000446 First citation in articleCrossref, Google ScholarEngelhardt, L., & Goldhammer, F. (2019). Validating test score interpretations using time information. Frontiers in Psychology, 10, Article 1131. 10.3389/fpsyg.2019.01131 First citation in articleCrossref, Google ScholarGates, A. I. (1924). The relation of quality and speed of performance: A formula for combining the two in the case of handwriting. Journal of Educational Psychology, 15(3), 129–144. 10.1037/h0075446 First citation in articleCrossref, Google ScholarGreiff, S., Wüstenberg, S., & Avvisati, F. (2015). Computer-generated log-file analyses as a window into students' minds? A showcase study based on the PISA 2012 assessment of problem solving. Computers & Education, 91, 92–105. 10.1016/j.compedu.2015.10.018 First citation in articleCrossref, Google ScholarHe, Q., Borgonovi, F., & Paccagnella, M. (2021). Leveraging process data to assess adults' problem-solving skills: Using sequence mining to identify behavioral patterns across digital tasks. Computers & Education, 166, Article 104170. 10.1016/j.compedu.2021.104170 First citation in articleCrossref, Google ScholarHe, Q., & von Davier, M. (2016). Analyzing process data from problem-solving items with n-grams: Insights from a computer-based large-scale assessment. In Y. RosenS. FerraraM. Mosharraf (Eds.), Handbook of research on technology tools for real-world skill development (pp. 750–777). IGI Global. First citation in articleCrossref, Google ScholarIseler, A. (1970). Leistungsgeschwindigkeit und Leistungsgüte. Theoretische Analysen unter besonderer Berücksichtigung von Intelligenzleistung [Speed and accuracy. Theoretical analyses with special consideration of performance in intelligence assessment]. Beltz. First citation in articleGoogle ScholarMan, K., Harring, J. R., & Zhan, P. (2022). Bridging models of biometric and psychometric assessment: A three-way joint modeling approach of item responses, response times, and gaze fixation counts. Applied Psychological Measurement, 46(5), 361–381. 10.1177/01466216221089344 First citation in articleCrossref, Google ScholarPokropek, A. (2016). Grade of membership response time model for detecting guessing behaviors. Journal of Educational and Behavioral Statistics, 41(3), 300–325. 10.3102/1076998616636618 First citation in articleCrossref, Google ScholarPokropek, A., Żółtak, T., & Muszyński, M. (2024). Identifying careless responding in web-based surveys: Exploiting sequence data from cursor trajectories and approximate areas of interest. Zeitschrift für Psychologie, 232(2), 95–108. 10.1027/2151-2604/a000555 First citation in articleLink, Google ScholarSchuster, M., & Paliwal, K. K. (1997). Bidirectional recurrent neural networks. IEEE Transactions on Signal Processing, 45(11), 2673–2681. 10.1109/78.650093 First citation in articleCrossref, Google ScholarStadler, M., Pickal, A., Brandl, L., & Krieger, S. (2024). VOTAT in action: Exploring epistemic activities in knowledge-lean problem-solving processes. Zeitschrift für Psychologie, 232(2), 109–119. 10.1027/2151-2604/a000559 First citation in articleLink, Google ScholarSternberg, R. J. (1977). Component processes in analogical reasoning. Psychological Review, 84(4), 353–378. 10.1037/0033-295X.84.4.353 First citation in articleCrossref, Google ScholarTang, X., Wang, Z., He, Q., Liu, J., & Ying, Z. (2020). Latent feature extraction for process data via multidimensional scaling. Psychometrika, 85(2), 378–397. 10.1007/s11336-020-09708-3 First citation in articleCrossref, Google ScholarTang, X., Wang, Z., Liu, J., & Ying, Z. (2021). An exploratory analysis of the latent structure of process data via action sequence autoencoders. British Journal of Mathematical and Statistical Psychology, 74(1), 1–33. 10.1111/bmsp.12203 First citation in articleCrossref, Google ScholarThorndike, E. L. (1914). On the relation between speed and accuracy in addition. Journal of Educational Psychology, 5(9), 537–541. 10.1037/h0075308 First citation in articleCrossref, Google ScholarThurstone, L. L. (1937). Ability, motivation, and speed. Psychometrika, 2(4), 249–254. 10.1007/BF02287896 First citation in articleCrossref, Google ScholarTresselt, M. E., & Mayzner, M. S. (1965). Anagram solution times: A function of individual differences in stored digram frequencies. Journal of Experimental Psychology, 70(6), 606–610. 10.1037/h0022667 First citation in articleCrossref, Google ScholarUlitzsch, E., He, Q., & Pohl, S. (2022). Using sequence mining techniques for understanding incorrect behavioral patterns on interactive tasks. Journal of Educational and Behavioral Statistics, 47(1), 3–35. 10.3102/10769986211010467 First citation in articleCrossref, Google ScholarUlitzsch, E., He, Q., Ulitzsch, V., Molter, H., Nichterlein, A., Niedermeier, R., & Pohl, S. (2021). Combining clickstream analyses and graph-modeled data clustering for identifying common response processes. Psychometrika, 86, 190–214. 10.1007/s11336-020-09743-0 First citation in articleCrossref, Google ScholarUlitzsch, E., Ulitzsch, V., He, Q., & Lüdtke, O. (2023). A machine learning-based procedure for leveraging clickstream data to investigate early predictability of failure on interactive tasks. Behavior Research Methods, 55(3), 1392–1412. 10.3758/s13428-022-01844-1 First citation in articleCrossref, Google ScholarUlitzsch, E., von Davier, M., & Pohl, S. (2020). A hierarchical latent response model for inferences about examinee engagement in terms of guessing and item-level non-response. British Journal of Mathematical and Statistical Psychology, 73(1), 83–112. 10.1111/bmsp.12188 First citation in articleCrossref, Google Scholarvan der Linden, W. J., & Belov, D. I. (2023). A statistical test for the detection of item compromise combining responses and response times. Journal of Educational Measurement, 60(2), 235–254. 10.1111/jedm.12346 First citation in articleCrossref, Google ScholarVista, A., Care, E., & Awwal, N. (2017). Visualising and examining sequential actions as behavioural paths that can be interpreted as markers of complex behaviours. Computers in Human Behavior, 76, 656–671. 10.1016/j.chb.2017.01.027 First citation in articleCrossref, Google Scholarvon Davier, A. A. (2017). Computational psychometrics in support of collaborative educational assessments. Journal of Educational Measurement, 54(1), 3–11. 10.1111/jedm.12129 First citation in articleCrossref, Google ScholarWang, Z., Tang, X., Liu, J., & Ying, Z. (2023). Subtask analysis of process data through a predictive model. British Journal of Mathematical and Statistical Psychology, 76(1), 211–235. 10.1111/bmsp.12290 First citation in articleCrossref, Google ScholarZhang, S., Tang, X., He, Q., Liu, J., & Ying, Z. (2024). External correlates of adult digital problem-solving process: An empirical analysis of PIAAC PSTRE action sequences. Zeitschrift für Psychologie, 232(2), 120–136. 10.1027/2151-2604/a000554 First citation in articleLink, Google ScholarZhu, M., Shu, Z., & von Davier, A. A. (2016). Using networks to visualize and analyze process data for educational assessment. Journal of Educational Measurement, 53(2), 190–211. 10.1111/jedm.12107 First citation in articleCrossref, Google ScholarFiguresReferencesRelatedDetails Volume 232Issue 2April 2024ISSN: 2190-8370eISSN: 2151-2604 HistoryPublished onlineApril 24, 2024 Licenses & Copyright© 2024Hogrefe PublishingPDF download Funding: Dr. Qiwei He's work is supported by the Institute of Education Sciences, U.S. Department of Education, through Grant IES R305A210344 to Georgetown University.
Questionnaires are by far the most common tool for measuring noncognitive constructs in psychology and educational sciences. Response bias may pose an additional source of variation between respondents that threatens validity of conclusions drawn from questionnaire data. We present a mixture modeling approach that leverages response time data from computer-administered questionnaires for the joint identification and modeling of two commonly encountered response bias that, so far, have only been modeled separately-careless and insufficient effort responding and response styles (RS) in attentive answering. Using empirical data from the Programme for International Student Assessment 2015 background questionnaire and the case of extreme RS as an example, we illustrate how the proposed approach supports gaining a more nuanced understanding of response behavior as well as how neglecting either type of response bias may impact conclusions on respondents' content trait levels as well as on their displayed response behavior. We further contrast the proposed approach against a more heuristic two-step procedure that first eliminates presumed careless respondents from the data and subsequently applies model-based approaches accommodating RS. To investigate the trustworthiness of results obtained in the empirical application, we conduct a parameter recovery study.