
In this article, we address the high prevalence of false discoveries in recognition memory research. Using Monte Carlo simulations, our goal was to find a valid measure of performance that reliably separates the contribution of sensitivity (accuracy) from that of bias. Relevant to myriad tasks, most notably old–new recognition memory, the simulations revealed that common measures confound sensitivity with bias, a finding termed the “measurement crisis.” As a solution, we propose a version of d-sub-a (da). We ran comprehensive simulations to evaluate the validity of sensitivity measures, including Pr = HR − FAR, A′, d′, and AUCg, in addition to da. Memory “signals” were randomly sampled from lure and target distributions. Sensitivity measures generated from iso-sensitive conditions that differed in bias were compared using t-tests, across thousands of simulations. For bias-independent measures, the rate of significant results should be 5
Researchers have witnessed a rapid increase in attention to within-person dynamic processes. However, individuals with observed zero within-person variability (i.e., non-varying individuals) remain understudied. This study investigates how non-varying individuals affect the estimation of within-person dynamic processes and offers practical recommendations. A motivating example of daily stressors demonstrates that including non-varying individuals can change substantive conclusions for univariate autoregressive models. Two simulation studies further explored how different proportions of non-varying individuals affect parameter estimation. Study 1 (univariate) found that non-varying individuals induce systematic upward bias in autoregressive estimates. Study 2 (bivariate) showed that, although average (fixed-effect) cross-lagged estimates may appear accurate, person-specific cross-lagged effects are systematically distorted, with some overestimated and others underestimated. Within this widely used Gaussian autoregressive modeling framework, we therefore recommend estimating dynamic parameters using a subsample that excludes non-varying individuals. We also indicate when subsample estimates can be interpreted as approximations to the full population, versus when they should be interpreted as describing the subsample population (e.g., when the proportion of non-varying individuals reaches around 10
Age of acquisition (AoA) is a key psycholinguistic variable known to influence lexical and conceptual processing across a broad range of tasks. However, existing Russian datasets are small and mainly limited to picturable nouns and verbs. In this study, we collected AoA ratings for 30,849 Russian words that represent a broad vocabulary range. Ratings were collected from 2,201 adult respondents in both in-person and online sessions. Reliability assessed via bootstrap split-half correlations was high (mean r = 0.895, 95
Remote data collection has become more common in psychological research on human participants due to the large, diverse participant pool available at relatively inexpensive costs. This shift, however, comes at the expense of the oversight and carefully controlled setting found in the laboratory. The present study examined the off-task behavior of participants hired through Amazon’s Mechanical Turk (MTurk) and its impact on memory performance across two pre-registered experiments. In Experiment 1, we found roughly one third of participants went off task during the encoding portions of a memory study and that this off-task behavior predicted poorer performance on the memory test. In Experiment 2, we replicated this outcome and further demonstrated that an explicit warning to stay on task reduced page switching but did not reduce the proportion of time spent off task. These results suggest the ease of remote data collection comes at a cost: a small but sizable minority of MTurk participants voluntarily navigate away from their study tasks, which negatively impacts performance. Further research is necessary to better understand how researchers might mitigate such behavior.
Many computerized adaptive testing (CAT) systems treat item parameters as if they were known without error, relying on point estimates obtained during item pool calibration. This practice can underestimate uncertainty in ability estimates and affect when a variable-length CAT terminates. A fully Bayesian (FB) CAT algorithm addresses this issue by explicitly incorporating item parameter uncertainty into both ability estimation and item selection. This study investigated the performance of FB CAT in a variable-length setting and compared it with conventional CAT under three stopping rules: a standard error (SE) rule, a change-in- θ (CIT) rule, and a combined CIT+SE rule. Simulation studies were conducted across a range of calibration sample sizes and item pool sizes. Results showed that the FB algorithm generally improved estimation accuracy and produced interval coverage rates closer to nominal levels, especially when the calibration sample size was small. The combined CIT+SE rule reduced unnecessarily long tests that can arise when using only the SE rule, particularly at θ levels for which the remaining item pool provides limited additional information to further reduce SE. Overall, the findings indicate that FB variable-length CAT can enhance uncertainty quantification, and that the combined CIT+SE rule offers a practical balance between measurement precision and testing efficiency.
Despite the widespread use of mixture models in psychology, education, and the behavioral sciences, there is little consolidated guidance on how distal outcome differences should be tested, corrected for multiplicity, and reported once a latent class or profile solution has been selected. In practice, applied researchers often report omnibus tests without follow-up comparisons, conduct uncorrected pairwise tests, or omit effect sizes and confidence intervals altogether, limiting the interpretability and reproducibility of findings. Consistent with this concern, a targeted reporting-practice audit of recent applied person-centered studies showed that multiplicity adjustment, global distal outcome effect sizes, pairwise effect sizes, and confidence intervals for pairwise effects were rarely reported. The present paper synthesizes recommendations from the general statistical literature and adapts them to the context of mixture modeling, focusing primarily on continuous distal outcomes and comparisons of class-specific means. We propose a principled framework for (a) defining appropriate families of pairwise comparisons for distal outcomes, (b) selecting and implementing multiplicity corrections with Benjamini–Hochberg recommended as the default procedure for applied distal outcome comparisons, and (c) computing and reporting global and pairwise effect sizes and confidence intervals using quantities readily available from standard mixture modeling software (e.g., Mplus). Through worked examples and a software-agnostic Quarto/R supplement, we demonstrate how these practices can be implemented transparently and consistently across common auxiliary-variable approaches, including maximum likelihood (ML) three-step and Bolck–Croon–Hagenaars (BCH) methods.
The 20-item prosopagnosia index (PI20) is a highly practical tool for assessing lifelong difficulties in face processing across individuals (developmental prosopagnosia). Recent research suggests that the quality of the PI20 may vary depending on the item; however, the quantitative quality of each item has not been sufficiently examined. In this study, the item properties of the PI20 were investigated using classical test theory and item response theory. The analyses confirmed that the one-factor structure of the PI20 showed high reliability and validity. Item response theory further revealed that the items differed in measurement power, discrimination, and difficulty under the unidimensional graded response model. Items reflecting the additional psychological load and social consequences associated with face processing difficulty demonstrated high measurement precision. In contrast, items related to the recognition of distinctive or self-faces contributed to the construct to a lesser degree. These results indicate that while most items adequately capture the characteristics of developmental prosopagnosia, certain items may not be suitable for screening for developmental prosopagnosia. The present study provides open materials to facilitate further psychometric evaluation of the PI20 and offers insights for refining self-report assessment of developmental prosopagnosia.
Three comprehensive sets of naturalistic stimuli, each with pairwise similarity ratings and multidimensional scaling (MDS)-based feature representations are presented, freely available on the corresponding OSF repository of this project under a CC-BY-SA-NC-4.0 license. The primary contribution of this work is the provision of these well-curated naturalistic stimulus sets and their associated high-quality similarity data as a resource for the research community. The sets include 80 representative items from three domains – foods, mammals, and countries – that can be used across cognitive research areas such as categorization, multiple-cue probability learning, judgment, decision-making, memory, and metamemory. Based on over 280,000 similarity judgments from N = 1798 participants, we derived 14 dimensions for foods, ten for mammals, and ten for countries, which reconstructed the similarity spaces with high accuracy ( r ≥ .93 ). In a proof-of-concept study, these dimensions also explained substantial variance in participants’ numerical judgments of domain-specific criteria, underscoring their functional relevance beyond similarity tasks. By making these stimuli, similarity data, and derived feature spaces freely available, we provide researchers with tools for testing computational cognitive models in relevant real-world domains, each of which allows for examining a multitude of judgment or classification criteria.
Embodied theories of language emphasize the role of sensorimotor experience in linguistic knowledge. Central to testing these theories is the creation of large datasets of linguistic norms, which contain judgments about a word's sensorimotor associations and can be used to predict human behavioral or brain data - sometimes in contrast to competing variables, such as those derived from distributional language models. Yet many of these datasets contain judgments about words in isolation, despite the fact that most words are ambiguous, making it difficult to determine which meaning of a word is characterized by its rating (e.g., "wooden table" vs. "data table"). In the current work, we introduce a new lexical resource (directly inspired by the Lancaster sensorimotor for 112 English words, each rated in four different contexts (448 sentences total). We demonstrate: first, that these ratings encode overlapping but distinct information from the Lancaster sensorimotor norms; second, that decontextualized ratings likely reflect the more dominant meaning of ambiguous words; third, that homonyms have more distinct sensorimotor profiles than polysemes; fourth, that the contextualized sensorimotor distance between two uses of an ambiguous word predicts human judgments about semantic relatedness; and fifth, that ratings derived from GPT-4 align reasonably well with human judgments. We conclude by suggesting that contextualized ratings like these can be used both to inform competing theories of semantic representations and also to evaluate or "probe" the ability of LMs to recover sensorimotor information.
The ability to precisely and accurately measure the distance between an implement and target point (i.e., the radial error) is crucial to conducting rigorous motor behavior research. However, manually measuring this error after each trial may be time-consuming and error-prone. In this paper, we present PinPointer, an error measurement system that enables motor behavior researchers to quickly and easily calculate x-axis, y-axis, and radial error distances. PinPointer automatically calculates descriptive statistical measures for each error type which can be exported for additional analysis. Our inter-rater and intra-rater reliability evaluations reveal that PinPointer has perfect absolute agreement and perfect consistency across eight raters and near-perfect to perfect intra-rater absolute agreement between two ratings of the same rater. PinPointer also demonstrated perfect fidelity to real-world measurements. PinPointer's Python-based source code is available as free and open source for other researchers to use and easily modify for their own applications and requirements. In addition, the PinPointer executable can be run on Windows or MacOS without installation or knowledge of Python programming. We hope PinPointer will prove to be a useful tool for motor behavior researchers to improve the speed, accuracy, and reliability of their error measurements.
Grid sampling is widely used in vision science and computer vision to study relations between local image regions and global image structure. Regions are sampled from an image by overlaying a grid with predefined rows and columns onto an image and treating the content of a cell as an image region. By manipulating the grid properties, researchers can systematically control sampling properties such as density and the spatial arrangement of local samples. Despite its frequent use, however, grid sampling is typically implemented using ad hoc, study-specific code, limiting reproducibility, comparability across studies, and accessibility for new users. To address this gap, we introduce GridSamp, an open-source Python toolbox that standardizes and streamlines the grid sampling workflow into an intuitive and flexible workflow. GridSamp supports multiple grid types with user-defined parameters and starting positions, allows manipulation of the appearance and arrangement of sampled regions (e.g., region shape, size, swapping, and shuffling), and enables extraction of image regions either with or without surrounding context, as regions of interest or as reassembled mosaic images. We demonstrate the utility of GridSamp by reproducing stimuli used in previous experimental studies, illustrating how the toolbox facilitates transparent and reproducible stimulus generation. The source code and interactive Jupyter Notebook tutorials are freely available on GitHub, and GridSamp can be downloaded as a module from the Python Package Index.
Using large language models (LLMs) with persona-based prompt engineering, this study simulates realistic insufficient effort responding (IER) data under controlled conditions, overcoming the limitations of traditional methods in ecological validity and controllability. The core objective is to generate controlled, distinct IER and non-IER datasets, thereby improving further research on detection methods. Our strategy involved systematically manipulating persona attributes, such as behavioral descriptions and psychological attributes, to produce synthetic IER data under controlled conditions. We instructed the LLM to generate one persona condition (general) with attentive response and three IER-intended persona conditions using varying combinations of IER-associated personality traits and IER behavioral descriptions (CRDO, CRPO, CRDP). To validate this approach, we first examined differences in response patterns across persona conditions using descriptive statistics and correlation analysis. Furthermore, we conducted confirmatory factor analysis (CFA) and analyzed IER detection indices to confirm that the synthetic IER data exhibited statistical traits like real IER data. Results indicate the pattern distinction among the persona conditions. Specifically, the IER-intended conditions consistently demonstrated IER characteristics, including degraded CFA model RMSEA (CRDO: 0.09; CRPO: 0.13; CRDP: 0.12) and high mean IER detection rates (n = 60; CRDO: 52.33
In spite of the inherently dynamic nature of actions, most psycholinguistic studies investigating action naming have relied on static images. To address this gap for French, we present a new database comprising 134 action videos. For each video, we developed five key psycholinguistic variables: Name agreement (NA), H-statistic, naming latency, uniqueness naming point (i.e., the moment in the video when the action becomes unmistakably identifiable), and adjusted naming latency. All variables are provided both for the overall sample and separately for adults under 50 years and those aged 50 and above. We then examined the influence of these variables on response latencies collected from 394 French-speaking participants. NA, the H-statistic, and instrumentality significantly predicted response latencies. The video-based action naming database is freely available to clinicians and researchers. This study provides a new resource to explore action naming in both clinical and research contexts.
Competence-based test development is a novel method for constructing tests that are as informative as possible about the competence state (the set of skills an individual possesses) underlying item responses. If desired, the tests can also be minimal, meaning that no item can be removed without reducing their informativeness. Consider a set of competencies, each encompassing the skills required to solve a particular item. An individual masters an item if their competence state includes all the skills in the item’s associated competency. A reduct is defined by a competency set that is as informative about individuals’ competence states as a larger set, and from which no competency can be removed without reducing its informativeness. Test development can be based on the reduct, since including only one item for each competency in the reduct yields a test that is as informative as a test that contains additional items and is minimal with respect to this property. This work introduces the competency addition procedure, an iterative method that successively collects competencies until the resulting set forms a reduct. The competencies to be added can be selected either randomly or based on test developer preferences, thereby favoring competencies with more desirable characteristics. The procedure is illustrated in three real-life applications for the assessment of arithmetic skills, involving the construction of a test from scratch, the improvement of an existing test, and the shortening of an existing test.
Errors in measurement can arise in study and survey responses when there is a discrepancy between the intended and selected response. A significant portion of the scientific discourse has centered on the comparison of discrete and continuous response scales. In this study, we examined various continuous rating scales to ascertain which scale would yield the lowest measurement error. To this end, we compared visual analogue scales (VAS) with different slider scales in an online study (N = 222) that built upon the original work by Reips and Funke (2008). In this study, participants were asked to estimate where a percentage value would lie on a line ranging from 0 to 100
A novel dataset of 556 street art images is presented, accompanied by affective evaluations from 1,239 Portuguese and Brazilian participants. Artworks were selected to reflect themes associated with the United Nations Sustainable Development Goals. Using a stimulus-sampling design, each participant completed an online survey in which 10 randomly selected artworks were presented and reported their responses in terms of valence, arousal, and specific emotion labels (being moved, awe, inspiration, hope, sadness, fear, anger, emotional connection, reflection, awareness, and interest), as well as their interest in street art and sustainability consciousness. Multilevel analyses showed that higher interest in street art and greater sustainability consciousness were consistent predictors of more positive emotional responses to the artworks. In contrast, the effects of gender and age were negligible, and national differences emerged only for feeling moved and awe. Network analyses revealed a highly interconnected emotional structure, with three clusters: self-transcendent, epistemic, and negative emotions. Feeling moved and emotional connection occupied central bridging positions, showing both direct and indirect links across positive and negative emotion clusters. Overall, these findings indicate that street art evokes a broad range of interconnected self-transcendent, cognitive-epistemic, and negative emotions, highlighting the complexity of viewers’ responses to the artworks. The dataset provides a valuable resource for research on emotional responses to street art and can support broader investigations into visual perception, aesthetic processing, and the communication of sustainability-related themes.
This study evaluated the usefulness of AI-generated estimates of word familiarity for predicting word difficulty in Simplified Chinese, building on previous research in alphabetic languages. We found that familiarity estimates produced using large language models (LLMs) showed moderate-to-strong correlations with human familiarity ratings. These LLM estimates were the most effective predictors of both word naming and lexical decision times, surpassing traditional metrics such as word frequency and human familiarity ratings, while the latter still provided modest, non-overlapping variance. GPT-4o with English instructions produced superior results compared to the Chinese-centered models currently available. The results imply that LLM familiarity estimates are a valuable resource for Chinese psycholinguistics, supporting work across experimental design, modeling, and norming. We release familiarity estimates for 27,624 words for unrestricted research and educational use.
Traditional profile analyses summarize multivariate person data with overall mean levels and relative patterns, but existing methods often blur these sources of variation or reduce each individual to a single best-fitting profile. Segmented Profile Analysis (SEPA) offers a unified, ipsatized singular-value decomposition (SVD) framework that decomposes individual profiles into orthogonal level (LE) and pattern (PE) effects and, crucially, introduces plane-wise segment profiles (summaries of each person's response pattern within each variable-contrast dimension) as primary person-oriented objects. Within each low-dimensional plane, SEPA defines a projected response pattern (segment profile), domain–person cosines that index variable-by-variable alignment, and plane-fit correlations that summarize how closely an individual’s pattern follows the plane’s domain structure. Across planes, SEPA aggregates information via singular-value weighting to yield an overall segment profile while preserving contrast-specific signal. A practical workflow combines ipsatized SVD, parallel analysis for PE dimensionality, marker-domain rules, bootstrap confidence intervals, and subspace-stability diagnostics. Using multivariate cognitive data from the Woodcock–Johnson IV, SEPA identifies interpretable marker domains, reveals distinct pattern facets across planes, and provides person-oriented indices that can be carried into standard regression models. Simulation studies examine the stability of segment profiles and cosines under varying sample sizes and variance structures. SEPA thus supplies a reproducible, geometry-based foundation for person-centered measurement that connects Q-type factor analysis, biplot methods, and contemporary within-person profiling in multidomain assessments.
A growing convergence between social science and machine learning enables, in principle, large-scale analyses of complex social phenomena through text. Yet, approaches leveraging supervised text classification based on human-annotated data for statistical analysis often treat conceptual validity and technical performance as separate challenges, impairing measurement quality. We provide guidelines to bridge this gap in what we call computational social mixed methods pipelines across three stages: data annotation, model training, and statistical analysis. Building on best practices and our own methodological innovations, such as “Iterative Annotation” and “Training on Confident Examples”, we address recurring pitfalls like unbalanced training data or stagnant model performance. We also discuss when large language models constitute a viable alternative to transformer-based classifiers. Using a case study on countering online hate, we illustrate how consequently integrating social science and machine learning expertise improves the validity and comparability of computational social science.
No matter how angry, sad, or happy we are, eventually, we will feel different. Studying this ebb and flow of affective experience in daily life provides important insights into psychological functioning and well-being. We have developed a parsimonious formalized model of intraindividual variability in affect (MIVA), resting on the assumption that such affective changes reflect transactions between an individual and their proximal environment. We provide an outline of its theoretical background, scope, and mathematical formulation. We situate MIVA within the research field of affect dynamics and illustrate the models’ behavior under realistic conditions using a simulation study. We use simulation-based inference to train a custom neural network on MIVA simulations, which we employ to rapidly estimate the model’s parameters on a multitude of synthetic experiments with different configurations. Our simulation study demonstrates that the synthesis between a computational model and probabilistic neural networks results in an efficient and flexible tool for model-based inference of affect dynamics. Our simulation study also offers insights into the data requirements for a precise recovery of the model’s parameters and recommendations for future data collection. The potential of MIVA for providing insights into affect dynamics is discussed.