
Sample-size determination (SSD) is the procedure of determining the number of subjects necessary to achieve a desired level of statistical power and is essential in planning an experiment. Although open-access software for SSD via closed-form equations is readily available in the null-hypothesis-significance-testing framework, this approach has been subject to severe criticism in the past. As an alternative, Bayesian evaluation of informative hypotheses via the Bayes’s factor or posterior model probabilities has been proposed. However, available software packages for Bayesian SSD are either (a) limited to simpler models, such as analysis of variance and t test, and cannot handle longitudinal data or (b) unable to handle more than two treatment conditions. Current software also neglects participant attrition—a common occurrence that may substantially reduce the power of a longitudinal experiment. In the present work, we address this gap by introducing the open-access R package BayesSSD , which performs simulation-based Bayesian SSD for longitudinal trials with two or more treatment conditions. Through a simulation study, we show that not only the proportion of individuals dropping out but also the timing of dropout need to be considered when performing SSD. In the presented method, various patterns of expected attrition can be specified via parametric and nonparametric survival functions and accounted for in the SSD procedure. To facilitate adoption, we provide a tutorial with empirical data, illustrating each step of the Bayesian SSD process.
When cognitive variables show a positive manifold, mainstream psychometric theory suggests that their common variance should be modeled first. Treating correlated cognitive scores as if they were separately interpretable outcomes instead requires explicit theoretical justification. We comment on Grassi et al., an important multilab study on musicians and nonmusicians, and reexamine their cognitive results at the latent level. Using multigroup and multilevel structural equation models on the same data set, we ask whether the reported pattern can be parsimoniously understood as a difference in shared cognitive variance rather than as a set of task-specific advantages. The results suggest that under this latent-variable representation, the musician advantage is mostly captured by a latent common factor, whereas melody span retained a clear residual advantage and vocabulary showed, at most, a smaller task-specific deviation. The focus, therefore, is not just whether musicians differ from nonmusicians on a set of observed tasks, possibly controlled for each other, but also how such differences should be interpreted once those tasks are allowed to share variance. Grassi et al. provided an excellent example of open, collaborative work that the field should actively encourage. We suggest that this kind of enterprise should also be paired with equally strong and theoretically grounded measurement modeling.
For 2 decades, online research has relied on a quality heuristic: Careful, coherent responding is good data. That heuristic is no longer reliable. Autonomous artificial-intelligence (AI) agents can now pass nearly all conventional quality checks, and in text-rich crowd-work tasks, reported use of large language models approaches one third. When such consultation shapes the response process itself—not just its surface expression—the resulting data appear human-generated while embedding systematic, model-shaped distortions. I synthesize emerging evidence on how AI-mediated contamination varies across research settings in prevalence, mechanism, and inferential consequence; and distinguish three contamination pathways (full delegation, partial mediation, and spillover) and three vulnerability zones (text-rich tasks at highest risk, browser-based cognitive paradigms as an emerging vulnerability, and supervised or identity-vetted settings at lower risk). Even modest contamination can shift estimated public opinion, compress attitudinal extremes, and, over time, feed back into the training data for future models. Current platform countermeasures may raise the cost of contamination but have not been independently validated under adversarial conditions. I argue for a shift from ad hoc detection to infrastructure redesign: contamination-aware sensitivity analyses, explicit stratification of data collection by evidential role, transparency norms that balance open science with adversarial robustness, and a minimum reporting checklist for online studies in vulnerable settings. I close by asking when AI mediation should be treated not as contamination but as part of the ecological baseline of human responding—a question that requires the field to specify the target cognitive system in any given study.
Traditional cross-sectional network-centrality metrics fail to distinguish causal directions between symptoms, leading to biases in selecting potential intervention targets. The nodeIdentifyR algorithm (NIRA) addresses this issue by using simulation-based interventions to identify projected optimal intervention target in cross-sectional networks. However, existing applications of NIRA typically overlook several recommended validation steps, which may reduce the robustness of its results. Specifically, a critical prerequisite for applying NIRA, testing for moderation effects to ensure the invariance of edge weights during simulated intervention, is consistently ignored. Moreover, they lack statistical significance testing for simulated intervention effects through permutation tests and stability assessment of NIRA outcomes via repeated simulations. In this article, we introduce the extended R package NIRApost , which supplements NIRA with these three recommended complementary procedures. We provide a comprehensive R tutorial demonstrating the implementation of both NIRA and these validation steps. Researchers applying NIRA are advised to conduct moderation-effect testing as a prerequisite, followed by permutation tests and stability analyses to ensure robust and interpretable findings. Upon completing this tutorial, readers are capable of properly applying NIRA and its validation procedures in their own data analyses.
The 65th anniversary of Campbell and Fiske’s multitrait-multimethod (MTMM) framework provides a timely opportunity to revisit and modernize this foundational model for construct validation. Although structural-equation-modeling-based MTMM approaches have enhanced the field, their widespread application remains constrained by convergence problems, ambiguous trait-method distinctions, and a lack of consensus regarding optimal model specifications. We propose an extended Campbell-Fiske framework that resolves these limitations while preserving the original guidelines’ conceptual strengths. Our key innovation is to apply the MTMM logic to a fully latent multitrait-multidomain correlation matrix derived from a rigorously tested multiple-indicator measurement model. Our approach treats traits and methods (i.e., domains, occasions, informants, contexts, or another method facet) as fully symmetrical, substantive facets, eliminates reliance on manifest correlations, corrects for measurement error, and introduces formal asymptotic parameter comparisons to test each validity criterion. This framework provides a formatively heuristic structure that retains the original appeal of the MTMM logic for applied research while meeting current psychometric standards for transparency, reproducibility, and inferential rigor expected by leading academic journals. We illustrate the method using a large, multidimensional data set ( N = 18,047), but the framework generalizes across domains of psychological science. The extended framework offers applied researchers a flexible, powerful tool for evaluating convergent and discriminant validity, diagnosing trait-domain interactions, and clarifying measurement quality. By “keeping the baby” while refreshing the empirical implementation, our approach affirms the enduring value of the Campbell-Fiske logic while aligning it with the demands of modern research practice.
Whether it stems from participant attrition, nonresponse, unwillingness to disclose information, technical errors, or flawed collection methods, incomplete data pose significant challenges to researchers in psychology. Although a rich methodological literature exists, applied researchers often lack clear guidance for aligning missing-data methods with study design, assumptions, and analytic goals. In this article, I provide a practical, assumption-aware framework for reasoning about missing data in psychology, emphasizing how missingness operates as a selection process and how method choice depends on the underlying data-generating structure. I review commonly used approaches, including likelihood-based estimation, multiple imputation, Bayesian data augmentation, and pattern-mixture models, highlighting their assumptions, strengths, and limitations. To support implementation and pedagogy, I introduce DataPatch, an interactive tool that allows users to simulate missing-data mechanisms, apply alternative handling strategies, and examine their consequences for estimation and interpretation (davidmoreau.shinyapps.io/DataPatch/). Together, the conceptual framework and accompanying tool aim to promote more transparent, principled, and informed handling of missing data in psychological research.
Multiverse analysis offers a comprehensive response to a core vulnerability in empirical research: the uncertainty of scientific conclusions arising from defensible yet flexible data-processing and -analysis decisions. By systematically mapping and computing all or a sample of all plausible data-processing pipelines, multiverse analysis reports the robustness of findings across analytical flexibility and increases transparency in the research process. As its adoption grows across disciplines, so too does the need for clarity on how to design, report, and interpret multiverse results responsibly. In this article, we provide interdisciplinary guidance on key procedural considerations, including defensibility and equivalence evaluations, preregistration, and computational demands. We aim to harmonize terminology, promote best practices, and foster conceptual cohesion across fields, supported by reference to domain-specific resources when appropriate. By doing so, we contribute to the broader movement toward more robust, reproducible, and transparent science, one that not only reports results but also interrogates the analytical pipelines that produce them.
Language-based assessments (LBAs), quantitative estimates of scientific constructs based on language, have advanced methods in the psychological and social sciences for more than a decade. LBAs based on individuals’ prompted descriptions analyzed with large language models to produce scores of their psychological states and traits have shown strong convergence with the corresponding rating scales ( r > .80) and have often surpassed rating scales in predicting theoretically relevant behaviors (external criteria). Despite their high validity across numerous psychological outcomes and contexts, the broader adoption of LBA models (LBAMs) has been limited. Even when made available alongside research publications, these models often remain inaccessible because of technical complexities, inconsistent documentation, and the absence of a standardized repository. In this tutorial, we introduce a framework targeted to social and psychological scientists for accessible sharing models with others—the Language-Based Assessment Models (L-BAM) Library—and a toolkit for easily using LBAMs via the text package in R. L-BAM covers a wide range of models for assessing mental-health disorders (e.g., depression, anxiety), well-being (e.g., satisfaction with life, harmony in life), implicit motives (need for power, affiliation, and achievement), and more. The L-BAM Library aims to increase the availability and resource efficiency of LBAs of psychological constructs while encouraging replication, independent validation, and the broad application of preexisting LBAMs.
Journalists are often maligned for covering sensational or desirable research results at the expense of studies with stronger methods. In the present study, we aimed to test how journalists’ preferences shift when studies are selected based on their methods rather than results (results-blind selection). Practicing journalists and editors, journalism faculty, and journalism graduate students ( N = 413) read summaries of real social-psychology studies and rated their interest in reporting on them. Participants were randomly assigned to read either “traditional” summaries that included the results or “results-blind” summaries that excluded the results. Summaries varied on three within-subjects dimensions: replication status, preregistration status, and belief consistency. Participants expressed more interest in replicable (vs. not replicable) and preregistered (vs. nonpreregistered) studies regardless of whether they learned the results, suggesting that these studies have features that are valued by journalists. Meanwhile, results-blind selection showed potential for reducing confirmation bias, suggesting it may be worth further exploration if feasibility challenges can be addressed.
Preregistration can help to restrict researcher degrees of freedom and thereby ensure the integrity of research findings. However, its ability to restrict such flexibility depends on whether researchers specify their study plan in sufficient detail and adhere to this plan. Previous research indicates higher restrictiveness when preregistrations are based on structured versus unstructured template formats, although there is room for further improvement. In this study, we built on these findings and investigated the restrictiveness of preregistrations based on the Psychological Research Preregistration-Quantitative (PRP-QUANT) Template, an extensive template that aids the preregistration of quantitative studies in psychology. Preregistrations were sampled from PsychArchives and coded for their level of restrictiveness using the coding schemes of Bakker et al. and Heirene et al. We predicted that preregistrations based on the PRP-QUANT Template ( N = 103) are more restrictive than preregistrations based on the OSF Preregistration Template ( N = 52; Hypothesis 1). We also inspected whether peer review can contribute further to restricting flexibility using nested Wilcoxon-Mann-Whitney tests and predicted higher restrictiveness for peer-reviewed ( n = 29) than non-peer-reviewed preregistrations ( n = 74; Hypothesis 2). In addition, we examined adherence to the preregistered plans in the associated publications ( N = 19). In line with Hypothesis 1, PRP-QUANT preregistrations had significantly higher restrictiveness scores than OSF preregistrations. Moreover, consistent with Hypothesis 2, peer-reviewed preregistrations had significantly higher restrictiveness than non-peer-reviewed ones. Of the associated articles, 73.68% included undeclared deviations. We discuss the implications of our findings for the PRP-QUANT Template and structured templates in general.
Improving the generalizability of psychology findings to a culture requires sampling participants in that culture. Yet psychology studies rarely sample from Africa even though Africa represents 17% of the global population. Although Africans can leverage the credibility-revolution initiatives to increase rigor and global representation, capacity building might speed the spread of these initiatives. In this study, we investigated an African-wide replication study to test whether Rottman and Young’s “mere-trace” hypothesis of moral reasoning (that people are more sensitive to the dosage of harm-based transgressions than purity transgressions) extends to several African communities. We used a training method developed by the Collaborative Replication and Education Project to train 23 African collaborators. During this process, we conducted a paradigmatic replication of Rottman and Young’s test of the mere-trace hypothesis in 12 contributing African sites from Burkina Faso, Kenya, Morocco, Nigeria, and Tanzania that sampled 783 participants after exclusions. Consistent with the original claim using U.S. samples, our African participants judged severe harm transgressions as more wrong than less severe ones but were not as sensitive to severity for purity transgressions (Domain × Dose: b = −4.63; p < .01). Moreover, the effect of dosage was smaller than reported among the U.S. sample, and our African participants rated all transgression scenarios more wrong than the U.S. sample. Resource constraints limited our sample to five African countries and to Africans dwelling in urban communities. Moral psychology should transcend the moral issues prioritized in the original study to include those considered important in African societies.
Construal-level theory (CLT) proposes that psychological distance influences the level of abstraction at which something is mentally construed: Things perceived as less probable (likelihood) or further away from the here (spatial distance), now (temporal distance), or self (social distance) are thought about more abstractly. In this international multilab study, we tested four basic hypotheses derived from core assumptions of CLT and explore potential moderators and boundary conditions of the effects. Participants ( N = 11,775) from 27 countries and regions were randomly assigned to one of four experimental protocols focused on different types of psychological distance (temporal, spatial, social, or likelihood), and each experiment manipulated psychological distance (close vs. distant). The protocols for temporal distance ( n = 2,941) and spatial distance ( n = 2,973) were direct replications of Liberman and Trope (Study 1) and Fujita et al. (Study 1), respectively. The remaining two protocols were paradigmatic replications, applying to social distance ( n = 2,926) and likelihood ( n = 2,936). The effects of psychological distance on construal level for the four present studies were as follows (positive effects are consistent with hypotheses): temporal, d = 0.08, 95% confidence interval [CI] = [0.003, 0.16] (effect in original study: d = 0.92); spatial, d = 0.04, 95% CI = [−0.03, 0.11] (effect in original study: d = 0.55); social, d = −0.27, 95% CI = [−0.34, −0.19]; and likelihood, d = 0.03, 95% CI = [−0.05, 0.11]. Pretests indicated that valence and abstraction were confounded in response options on the outcome measure. Controlling for this confound eliminated the hypothesis-inconsistent effect of social distance, d = 0.006, 95% CI = [−0.05, 0.07]. These findings provide limited evidence for the predictions of the theory and present a critical challenge for CLT.
A common goal of researchers using intensive longitudinal data is to develop models that predict emotions or behaviors, often using passively collected data from smartphone sensors or wearable devices. A frequent use case for such models is the development of just-in-time adaptive interventions (JITAIs). However, real-world effectiveness depends on rigorous evaluation. Previous research has highlighted challenges in selecting appropriate evaluation methods. To address these, we review key pitfalls in predictive modeling and provide recommendations for avoiding them. We focus on a common problem: the mismatch between development, evaluation and application, and use simulations to illustrate three pitfalls. First, although models may perform well from applying group-level validation (area under the curve [AUC] = .82), they may lack the ability in predicting within-persons change (mean AUC = .54, SD = .13). For JITAIs, this will prevent the model from identifying intervention-delivery moments and will discriminate only between individuals. Second, ensuring adequate variability in the outcome variable is critical. If outcomes remain stable, frequent prediction may offer little practical benefit. Third, selecting appropriate baseline models is essential; models that appear effective may underperform compared with simple baselines (e.g., AUC = .82 vs. AUC = .96). To address these pitfalls, we present recommendations for matching validation and evaluation strategies to the intended use-case scenario and provide a tool that can help researchers investigate whether their strategy and goal are misaligned. This can help improve the effectiveness of predictive models and increase their utility in real-world applications.
Google searches have been described as the most important data set on the human psyche ever assembled. Google-search data—accessible through a tool called Google Trends—can provide new insights on topics as varied as stereotypes and prejudices, political attitudes, religious identity and belief, personality, motivations, psychological well-being, mental health, and culture. Google Trends can generate highly customized data sets: Users can compare the popularity of search terms across most of the world or access longitudinal data as far back as 2004, and they can do so with high geographical and temporal granularity. Notwithstanding these opportunities, Google Trends has significant limitations. Without appropriate caution, users can easily rely on data that are not meaningful or draw mistaken conclusions. We provide a comprehensive overview and tutorial covering (a) opportunities of Google Trends for psychological scientists; (b) how Google Trends scores are calculated, how reliable they are, and why some queries might yield low-quality data; (c) instructions with accompanying R code for creating custom data sets beyond what Google Trends provides by default; (d) example analyses for studies that could be done using Google Trends data; (e) an overview of common pitfalls; and (f) recommendations for safeguarding data quality and their interpretation.
Big-team science collaborations have been heralded as a solution to oversampling in a limited number of high-income countries. Despite early successes, there is insufficient involvement from the global community and unclear benefits to globalized science. The expansion of research from sites in North America and Europe to parts of the world where most people live can create the appearance of progress based on geographical diversity while neglecting the perspectives, problems, and knowledge specific to those populations. Here, we describe participatory open-research practices that bring global perspectives to open science. Participatory practices involve revising and transparently communicating worldviews, valuing humility over control, prioritizing team facilitation over management, and listening to versus instructing collaborators. We detail these concepts and their utility and provide recommendations for conducting robust, open, and culturally embedded research that will help realize the potential value of big-team science.
Computational cognitive models offer powerful means for testing competing theoretical frameworks. A central challenge is determining which model best explains observed data, balancing goodness of fit with parsimony. Several fruitful approaches to model comparison have been used in the areas of cognitive and mathematical psychology, but the most popular in practice remain Akaike information criterion (AIC) and Bayesian information criterion (BIC), which penalize model complexity as measured by the number of free parameters. Here, we revisit these conventional approaches to model selection on a sample case of the prototype and exemplar models of categorization. We highlight the limitations of parameter count-based complexity measures, showing that they may fail to capture a model’s true flexibility. We then introduce a Monte Carlo permutation-testing approach as an alternative that has a rich tradition in many areas but whose use for model selection is still trailing that of AIC/BIC. We demonstrate that permutation testing offers at least three advantages: more robust comparison of models with chance, more robust comparison between models with equal or differing numbers of parameters, and quantification of uncertainty in model selection. After demonstrating how permutation testing offers a more nuanced and principled framework for evaluating cognitive models, we conclude with practical considerations for implementing permutation-based model selection in cognitive-modeling research.
Generative artificial intelligence (AI) poses a significant threat to data integrity on crowdsourcing platforms, such as Prolific, which behavioral scientists widely rely on for data collection. Large language models (LLMs) allow users to generate fluent and relevant responses to open-ended questions, which can mask inattention and compromise experimental validity. To empirically estimate the prevalence of this behavior, we analyzed keystroke data from three studies ( N = 928) on Prolific between May and July 2025. Using an embedded JavaScript tool, we flagged participants who pasted text or whose keystroke count was anomalously low compared with their response length. For each flagged participant, we manually compared detected keystrokes with their final response to determine if the text could have been typed. This confirmed that despite deterrence measures, approximately 9% of participants submitted responses consistent with AI assistance or other forms of outsourced responding. These participants outperformed noncheaters (by up to 1.5 SD ), were more than twice as likely to share geolocations with other participants (suggesting possible proxy use), and exhibited lower internal consistency on questionnaire scales. Simulated power analyses indicate that this level of undetected cheating can diminish observed effect sizes by 10% and inflate required sample sizes by up to 30%. These findings highlight the urgent need for new detection methods, such as keystroke logging, which offers verifiable evidence of cheating that is difficult to obtain from manual review of LLM-generated text alone. As AI continues to evolve, maintaining data quality in crowdsourced research will require active monitoring, methodological adaptation, and communication between researchers and platforms.
We introduce a human-in-the-loop pipeline for creating context-aware (e.g., culture, sex, and age) affect-induction images and the initial Library of AI-Generated Affective Images. Current limitations in image-based research include weak to moderate emotional-elicitation effects, limited image diversity, and minimal cultural tailoring of images. Using generative artificial intelligence (AI) guided by existing data sets and emotion taxonomies, we generated 847 images and their corresponding descriptions across 12 discrete emotions and then iteratively refined them with local cultural experts. We validated the library through six studies ( N = 2,470; 58 countries). Participants rated five types of images: (a) images from existing affective databases, (b) AI-generated images without cultural adjustments, (c) AI-generated images adjusted to specific cultural contexts, (d) AI-generated images adjusted by sex (male, female), and (e) AI-generated images adjusted by age group (childhood, adulthood, older age). The AI-generated images were as effective in eliciting affective responses as the images from existing affective databases. Culturally adjusted images were slightly more effective than unadjusted counterparts in targeting intended emotions. Sex- and age-adjusted variants produced comparable responses with their base images, demonstrating controllability without loss of affective impact. Furthermore, we calculated the smallest subjectively experienced difference for affect-induction research ( d s = 0.05–0.29). This work demonstrates that researchers can now generate high-quality affect-induction stimuli cost-effectively and at scale and tailor them to diverse contexts—overcoming long-standing barriers and laying the groundwork for future AI-driven methodologies in affective science.
Free-text responses are a crucial part of psychological research, enabling participants to respond without bias toward a predefined set of answers. Unfortunately, many established methods for analyzing such responses require extensive manual coding, which is time- and resource-intensive. To address this issue, automatic-processing methods based on word embeddings and clustering techniques have been proposed. In this article, we introduce SCORES (Semantic Clustering of Open Responses via Embedding Similarity), a user-friendly, graphical tool that makes such automatic methods easy to use and understand for psychological researchers.