Science has been going through a sequence of public credibility tests which have made one gap visible: the difference between what science aspires to (transparency, reproducibility, self-correction) and what it can actually deliver and verify at scale. The reform movement of the last decade and a half gave the field shared norms and initial solutions, but change has been slow and its impact unclear. The recent rise in AI capability has, for the first time, the potential to narrow that gap. I discuss how AI is beginning to drive a change in the behavioral sciences and, by extension, science more broadly, by making feasible what was until recently very difficult or outright impossible, and to do so at scale: the ongoing systematic assessment, verification, updating, and improvement of evidence, relevance, and impact. The promise is that AI would empower researchers to ensure science credibility and maximize the value and impact of science for society.
Replicability is a cornerstone of scientific progress. Yet, replications are often undervalued, and are sometimes seen as redundant, unimportant, or lacking novelty. This impedes their broader adoption in research and beyond. In response, the credibility revolution calls for slower, more deliberate science and greater responsiveness to fallibility. In this perspective piece, we argue that (a) replications are essential for validating scientific claims, (b) replications need to be made more visible, recognized, and integrated into research and educational practices, and (c) we can change the way we view and judge replication results. We propose a framework where replication studies can be systematically tracked and normalized through the Replication Hub as part of the Framework for Open and Reproducible Research Training (FORRT) initiative, with the goal of enhancing the visibility, integration, and cumulative impact of replication research across disciplines.
In the past two decades, research on the motivational underpinnings of the Dark Triad has burgeoned. However, it is unclear how each Dark Triad trait may map onto the values circumplex, and whether the research conducted thus far indicates consistent effects. In this multi-level meta-analysis, we examined the relationship between Dark Triad personality traits and personal values. Across 34 studies conducted between 2000 and 2020, Dark Triad traits were positively associated with self-enhancement and openness-to-change value dimensions, and negatively associated with self-transcendence and conservation value dimensions. Shape consistency for the Dark Triad associations was stronger for the self-enhancement versus self-transcendence values tension than for the openness-to-change versus conservation values tension. We concluded that Dark Triad traits showed meaningful patterns of associations with personal values with some differences between the traits. Materials, data, and code are available on: https://doi.org/10.17605/OSF.IO/4Z27F.
Weiner et al. (1988) found that compared to physically based stigmas, mental-behavioral stigmas were perceived as more onset controllable, less stable (irreversible) and were therefore associated with less pity and liking, more anger, and less willingness to help. We conducted a replication and extension Registered Report of Experiment 2 by Weiner et al. (1988) with a US American online Amazon Mechanical Turk sample through CloudResearch (N = 804). We found support for the original findings on the associations between stigma origin (i.e., mental-behavioral or physically based), perceived stability, perceived onset controllability, emotional reactions (i.e., sympathy, liking, and anger), and willingness to help (all eta(2)(p)s >= .35). We also found support for higher perceived controllability leading to increased ratings on perceived responsibility, blame, and anger, and decreasing liking, sympathy, and tendencies to provide donations and assistance for stigmatized individuals (all eta(2)(p)s >= .03). Extending the replication, we tested the model with four new stigmas prevalent in the last decade. We also assessed the original study's untested categorizations of stigmas sources - physical versus mental-behavioral and found that 7 out of 10 categorizations matched with the original's. We concluded that we found support for the relationship between stigma source and the attribution-affect-help judgment model. Materials, data, and code are available on the OSF: https://doi.org/10.17605/OSF.IO/GWCBT. This Registered Report has been officially endorsed by Peer Community in Registered Reports: https://doi.org/10.24072/pci.rr.100885
Default effect refers to the phenomenon that people tend to choose the default option in a choice-set. McKenzie et al. (2006) showed that one mechanism for the effect is that defaults are perceived as an implicit endorsement and a recommended option by the policy makers that set the defaults. In a pre-registered experiment with a Prolific online sample from the United States (N = 602), we conducted replications and extensions of Studies 1, 2, and 3 from McKenzie et al. (2006). We concluded a successful replication of all three studies, with support for their findings showing that: 1) people who are more willing to be organ donors are more likely to hold views that others should as well, and to think that enrollment should be the default, and that in turn: 2) people make inferences from policymaker’s set defaults (Study 2: organ donations; Study 3: retirement plans) regarding the policymakers’ own preferences regarding enrollment, and their views regarding whether others should enroll. In our extension we added a control condition with no defaults, finding that 1) people still make inferences about choice architects’ attitudes, even when defaults were not set, and 2) that setting defaults to not act (not enroll) is perceived as closer to no defaults than setting defaults to act (to enroll). Materials, data, and code are available on: https://doi.org/10.17605/OSF.IO/E5K4J.
Construal-level theory (CLT) proposes that psychological distance influences the level of abstraction at which something is mentally construed: Things perceived as less probable (likelihood) or further away from the here (spatial distance), now (temporal distance), or self (social distance) are thought about more abstractly. In this international multilab study, we tested four basic hypotheses derived from core assumptions of CLT and explore potential moderators and boundary conditions of the effects. Participants ( N = 11,775) from 27 countries and regions were randomly assigned to one of four experimental protocols focused on different types of psychological distance (temporal, spatial, social, or likelihood), and each experiment manipulated psychological distance (close vs. distant). The protocols for temporal distance ( n = 2,941) and spatial distance ( n = 2,973) were direct replications of Liberman and Trope (Study 1) and Fujita et al. (Study 1), respectively. The remaining two protocols were paradigmatic replications, applying to social distance ( n = 2,926) and likelihood ( n = 2,936). The effects of psychological distance on construal level for the four present studies were as follows (positive effects are consistent with hypotheses): temporal, d = 0.08, 95% confidence interval [CI] = [0.003, 0.16] (effect in original study: d = 0.92); spatial, d = 0.04, 95% CI = [−0.03, 0.11] (effect in original study: d = 0.55); social, d = −0.27, 95% CI = [−0.34, −0.19]; and likelihood, d = 0.03, 95% CI = [−0.05, 0.11]. Pretests indicated that valence and abstraction were confounded in response options on the outcome measure. Controlling for this confound eliminated the hypothesis-inconsistent effect of social distance, d = 0.006, 95% CI = [−0.05, 0.07]. These findings provide limited evidence for the predictions of the theory and present a critical challenge for CLT.
The inaction inertia effect is a cognitive bias in which forgoing an attractive opportunity reduces the likelihood of accepting a subsequent, less attractive opportunity. This phenomenon was initially demonstrated by Tykocinski et al. (1995), and the present study aimed to replicate their findings and assess the effect in a Brazilian population. We used an online survey with a translated version of the original questionnaire, as well as an extension scenario. The final sample included 436 participants. We performed statistical analyses using the Kruskal-Wallis non-parametric test, comparing three conditions ("Small-Difference", "Large-Difference", and "Control") for the four original scenarios and one extension scenario. We also stratified the scenarios into two categories for a second extension investigation: those involving money (Hotel and Car) and those that did not (Frequent flyer and Fitness center), and evaluated them using linear mixed-effects analyses. We found evidence of the inaction inertia effect in 6 of the 12 comparisons between conditions conducted across the four original scenarios, indicating its presence in this Brazilian sample. All scenarios presented significant differences between conditions (p-values < 0.001; ϵ² ranging from 0.06 to 0.22), except for the Frequent flyer scenario (p-value = 0.073; ϵ² = 0.01). We also observed support for the Inaction Inertia Effect in the extension scenario, providing further evidence of its presence in our sample (p-value = 0.002; ϵ² = 0.03). Moreover, we found that money was a significant factor in participants’ decisions (p-values = 0.004 and 0.002), suggesting that different decision-making processes are involved when money is at stake. This replication-extension study provides evidence of the inaction inertia effect in the Brazilian population, indicating that passage of time and cultural differences may not significantly affect this cognitive bias, first described almost 30 years ago. Materials, data, and code were made available at https://doi.org/10.17605/OSF.IO/62NXB
Västfjäll et al. (2014) found support for “compassion fade”, that affective reactions and charitable donations decrease as the number of victims increases. In a pre-registered experiment with a U.S. sample recruited through CloudResearch Connect (N = 1207), we conducted a replication and extension of Studies 1a and 3 from Västfjäll et al. (2014). In the replication of Study 1a, we found no support for the impact of the number of victims (one vs. two children) on hypothetical donation amount, positive- or negative affect, perceived impact of donations (all |d|s ≤ 0.16) or for any mediation effects. In the replication of Study 3, we found no support for the impact of the number of victims or relatedness on donation amount, perceived need, positive affect, or negative affect of donations (all ηp2 ≤ 0.012). Bayesian analyses indicated moderate to strong evidence in favor of the null hypothesis for most effects, and equivalence testing indicated that seven of eight effects were smaller than our conservative smallest effect size of interest. Overall, our findings fail to replicate the original compassion fade effects and suggest limitations on their generalizability across various dependent measures and contexts. Materials, data, and code are available on: https://doi.org/10.17605/OSF.IO/A2TGB
The Side-Effect Effect (SEE) is the phenomenon that negative side-effects elicit stronger attributions of intent and blame than intent and praise for positive side-effects. There are similar documented asymmetries showing stronger free will attributions to negative than to positive, and stronger associations between free will attributions and blame for negative outcomes than associations between free will attributions and praise for positive outcomes. Together, these are two well-known paradigms in experimental philosophy that have thus far mostly been studied separately. Given that they both examine similar domains regarding agency, intent, and responsibility, we aimed to integrate the two paradigms to examine possible joint effects and interactions. We used the classic SEE scenario with within and between designs, manipulated free will by contrasting deterministic versus indeterministic universes, and measured free will attributions. In two experiments (overall N = 1520), we found support for side-effect effects regarding attributions of intentionality and knowledge (Study 1: d = 0.58-1.77; Study 2: d = 0.61-1.75). We found a strong association between blame/praise and free will attributions, even when controlling for intent and knowledge. Finally, we found that when participants were asked to imagine a counterfactual and report praise or blame based on the experimental condition, blame was more strongly attributed to hypothetical harmful outcomes than praise to helpful outcomes. We found no consistent support for an interaction between the two paradigms, suggesting that they uniquely affect attributions. All materials, data, and code are available on: https://osf.io/z3g6d/
Choice partitioning refers to the phenomenon when the same choice set yields different decision-making behaviour when they are grouped into sets (broadly bracketed) or evaluated separately (narrowly bracketed). In a Registered Report experiment with a US sample recruited online through Prolific (N = 896), we conducted a replication of seven studies reviewed in Read et al. (Read et al. 1999 J. Risk Uncertain. 19, 99 (doi:10.1023/A:1007879411489)). We concluded a mostly successful replication: out of the seven studies, we found support for six (Studies 1, 3, 4 and 6: Cramer’s V > 0.31; Studies 2 and 7: Cohen’s d > 0.29) and no empirical support for one (Study 5: Cramer’s V = 0.02). Extending the replication, we added new conditions in Studies 6 and 7, further expanding the manipulation’s scope range, yet failed to find any impact. In our replication, we came across many challenges, both conceptual and empirical, and we therefore call bracketing scholars to better define bracketing in relation to other phenomena in decision-making (joint versus separate mode, framing effects, mental accounting, etc.), with falsifiable hypotheses, examining overlap with other constructs, and clearer mapping between theory and empirics. Materials, data and code are available on: https://osf.io/vdqek/.
Newman et al . 2014 Value judgments and the true self. Personal. Soc. Psychol. Bull. 40 , 203–216. (doi: 10.1177/0146167213508791 ) demonstrated that behaviours that are more aligned with moral values are perceived as more strongly reflecting a person’s ‘true-self’, suggesting that morality plays an important role in how people perceive others’ essential self. In this Registered Report, we conducted a close replication of Newman et al. 2014 Value judgments and the true self. Personal. Soc. Psychol. Bull. 40 , 203–216. (doi: 10.1177/0146167213508791 )’s Studies 1 and 2 with an online US American sample recruited from Amazon Mechanical Turk using CloudResearch ( N = 803). We found support for Study 1’s findings that morally positive changes in others are perceived as more reflective of true-self than morally negative changes, in both the forced-choice (original: η²p = 0.39, 95% CI [0.25, 0.51]; replication: η²p = 0.20, 95% CI [0.16, 0.23]) and the continuous scale (original: η²p = 0.33, 95% CI [0.19, 0.45]; replication: η²p = 0.22, 95% CI [0.15, 0.25]) measures. We found support for Study 2’s findings that changes more aligned with observers’ political moral views are perceived as more reflective of true-self (original: η²p = 0.04, 95% CI [0.00, 0.11]; replication: η²p = 0.35, 95% CI [0.29, 0.41]). Extending the replication, we examined associations between true-self attributions and perceived social norms and found that social norms were positively associated with true-self attributions (Study 1: most r s ranged from 0.07 to 0.21; Study 2: r s = 0.10 to 0.30). Materials, data and analysis code are available on https://doi.org/10.17605/OSF.IO/9FVTQ . This Registered Report has been officially endorsed by Peer Community in Registered Reports: https://doi.org/10.24072/pci.rr.100372 .
Individuals who donate to charity may be affected by various biases and donate inefficiently. In a replication and extension registered report with a US Amazon Mechanical Turk sample using CloudResearch (N = 1403), we replicated studies 1 to 4 in Baron & Szymanska (Baron & Szymanska 2011 In The science of giving: experimental approaches to the study of charity (eds DM Oppenheimer, CY Olivola), pp. 215–235 (doi:10.4324/9780203865972-24)) with extensions on reputation and overhead funding. We found support for the effects of a preference for lower perceived waste (d = 0.70, 95% CI [0.41, 0.99]), lower past costs (d = 0.59, 95% CI [0.16, 1.02]), for the ingroup (d = 0.52, 95% CI [0.47, 0.58]), for having some diversification between charities (d = 0.63, 95% CI [0.47, 0.78] for single projects; d = 1.18, 95% CI [1.00, 1.36] for several projects versus one) and against forced charity (d = 0.29, 95% CI [0.21, 0.37]; nominally replicated, but has caveats regarding validity); as at least four of our five hypotheses were found to replicate, we conclude this as being a successful replication. Extending the replication, we found support for an unexpected preference for anonymity on donation allocation (opposite to our predictions; d = 0.54, 95% CI [0.46, 0.61]), and support for a preference towards paid-for overhead costs on donation allocation (d = 0.60, 95% CI [0.52, 0.68]). We discuss the implications and validity of these findings. All materials, data and code were made available on: https://doi.org/10.17605/OSF.IO/BEP78. This registered report has been officially endorsed by Peer Community in Registered Reports: https://doi.org/10.24072/pci.rr.100775.
The appraisal-tendency framework proposed that specific emotions predispose individuals to appraise future events corresponding to the core appraisal themes of the emotions. In a registered report with a U.S. American online Amazon Mechanical Turk CloudResearch sample (N = 780), we conducted an independent close replication of Experiments 1, 2, and 3 in Lerner and Keltner (2001). We found support for the appraisal-tendency framework for risk optimism in general, risk optimism for positive events, and risk optimism for ambiguous events, but not for risk preference and risk optimism for negative events. Extending the replication, we added hope as a positive-valence dispositional emotion with low certainty and control and failed to find support for the assumptions of the appraisal-tendency framework. (PsycInfo Database Record (c) 2026 APA, all rights reserved).
Mental accounting, the internal categorization system individuals adopt to manage their financial activities, may result in suboptimal decisions or decision-making not aligned with one's own goals. In a Registered Report with an online US sample recruited from Amazon Mechanical Turk using CloudResearch, we conducted a replication of 17 classic problems reviewed in Thaler 1999 J. Behav. Decis. Mak. 12, 183-206. (doi:10.1002/(SICI)1099-0771(199909)12:33.0.CO;2-F) (N = approx. 500 per problem; overall: N = 1007). We concluded a mostly successful replication: out of the 17 problems, we found empirical support for 11, mixed empirical support for three and no empirical support for three. Extending the replication, we provided an initial test of four untested predictions described in Thaler 1999 J. Behav. Decis. Mak. 12, 183-206. (doi:10.1002/(SICI)1099-0771(199909)12:33.0.CO;2-F), of which we found empirical support for two, mixed support for one and no support for one. Materials, data, and analysis code are available on: https://osf.io/v7fbj/. This Registered Report has been officially endorsed by Peer Community in Registered Reports: https://doi.org/10.24072/pci.rr.100375.
Jordan et al., 2011, demonstrated that people underestimated the prevalence of others' negative emotional experiences and that these were associated with higher well-being. We conducted a preregistered replication of Studies 1b and 3 by Jordan et al., 2011 (N = 594) with adjustments and added extensions. Building on their methodology, we examined both prevalence and intensity of emotional experiences, and our findings suggest a much more complex story with surprising effects. We found an underestimation of the prevalence of negative emotions, but also unexpectedly of an underestimation of the prevalence of positive emotions, with stronger effects for negative than for positive emotions. However, we found an opposite effect for emotional intensity; people overestimated the intensity of both positive and negative emotional experiences, again with stronger effects for negative. Surprisingly, associations between prevalence estimations and well-being were in the opposite direction to the target article's. Materials, data, and code: https://osf.io/bwmtr/.
Replications are essential for rigorous credible science yet are still grossly undervalued and very rare. The value of replications is directly tied to the value of the research they aim to replicate, and replications offer many benefits that go far beyond the mere testing of replicability, such as including verifications and error detection, promoting long-term reproducibility of all research outputs, clarifying theory, refining measurement, and testing generalizability. We need far more independent pre-registered well-powered direct replications to strengthen the credibility of scientific findings. Isager et al. (2025)’s aim to define a formula for the value of replications based on over-simplified metrics of citation count and sample size is misaligned, already misunderstood, and may backfire by hindering the pursuit and publication of replications.
The studies in Shafir (1993, Memory & Cognition 21, 546–556) examined the impact of decision frames (choosing vs. rejecting) on decision-making. Our replication—Chandrashekar et al. (2021, Judgment and Decision Making 16, 36–56)—revealed mixed results with only partial support for the original findings, concluding a successful replication of only 2 out of 8 scenarios. Our data from an exploratory extension suggested a pattern in support of an alternative theoretical mechanism aligning with Wedell’s (1997, Memory & Cognition 25, 873–887) accentuation hypothesis. Shafir and Cheek’s (2024) commentary criticized our approach to replications, and the value and importance of direct close replications overall, and shared their views regarding the theory and scope of the phenomenon, with new information about what they consider to be needed steps to empirically test the phenomenon. In our response, we clarify misunderstandings and address empirical findings shared in the commentary. We discuss and defend the value and importance of direct replications and the necessity for full transparency regarding the theoretical assumptions and the process of empirical investigations. Finally, we call for the implementation of open science more broadly, in conducting more direct close replications, sharing of all protocols, materials, data, and code, and implementing outcome-blind reviewing and Registered Reports. These would allow for stronger theoretical and empirical foundations, and a more credible and robust psychological science.