
This commentary examines whether the motives that currently guide replication practices align with normative proposals for selecting replication targets. Drawing on the preliminary analysis of 1,075 replication studies in psychology, we find that researchers most often conduct replications to self-replicate, extend previous findings, or test their generalizability, whereas motives related to a study’s influence or uncertainty are mentioned less frequently. These findings suggest a discrepancy between what researchers do replicate and what meta-scientific frameworks recommend should be replicated, highlighting the continued influence of novelty-oriented incentives even within the replication landscape.
The inaction inertia effect is a cognitive bias in which forgoing an attractive opportunity reduces the likelihood of accepting a subsequent, less attractive opportunity. This phenomenon was initially demonstrated by Tykocinski et al. (1995), and the present study aimed to replicate their findings and assess the effect in a Brazilian population. We used an online survey with a translated version of the original questionnaire, as well as an extension scenario. The final sample included 436 participants. We performed statistical analyses using the Kruskal-Wallis non-parametric test, comparing three conditions ("Small-Difference", "Large-Difference", and "Control") for the four original scenarios and one extension scenario. We also stratified the scenarios into two categories for a second extension investigation: those involving money (Hotel and Car) and those that did not (Frequent flyer and Fitness center), and evaluated them using linear mixed-effects analyses. We found evidence of the inaction inertia effect in 6 of the 12 comparisons between conditions conducted across the four original scenarios, indicating its presence in this Brazilian sample. All scenarios presented significant differences between conditions (p-values < 0.001; ϵ² ranging from 0.06 to 0.22), except for the Frequent flyer scenario (p-value = 0.073; ϵ² = 0.01). We also observed support for the Inaction Inertia Effect in the extension scenario, providing further evidence of its presence in our sample (p-value = 0.002; ϵ² = 0.03). Moreover, we found that money was a significant factor in participants’ decisions (p-values = 0.004 and 0.002), suggesting that different decision-making processes are involved when money is at stake. This replication-extension study provides evidence of the inaction inertia effect in the Brazilian population, indicating that passage of time and cultural differences may not significantly affect this cognitive bias, first described almost 30 years ago. Materials, data, and code were made available at https://doi.org/10.17605/OSF.IO/62NXB
This paper is an updated English version of a report filed by a commission that the German Psychological Society (DGPs) appointed in 2022. The commission’s task was a) to identify factors in the academic system that enable and/or promote unethical behavior, and b) to propose corrective measures. Based on expert interviews, a literature review, feedback from the community, and discussions within the commission, the following problematic issues were identified: (P1) negligent or fraudulent research practices, (P2) abuse of power, (P3) inadequate supervision of Early Career Researchers, (P4) poor quality of teaching, (P5) counterproductive incentives, (P6) overburdening of professors with tasks, (P7) fixed-term and short-term employment, (P8) unnecessarily strong power imbalance, (P9) an ineffective peer-review system, (P10) questionable assessment practices in hiring professors, (P11) lack of clarity of, and low commitment to, ethical standards, and (P12) weak control and sanctioning mechanisms. The commission offers concrete recommendations for changes to the academic system that should make the occurrence of unethical behavior less likely. Even though some of these recommendations might be rather specific for the field of psychology and for the German academic system, most of them may be valid for science in general.
The present study is a direct replication and extension of Fogel and Kovalenko’s (2013) work on the association between viewing reality television (TV) and engaging in one-night stands among college students. Using an online survey, a total of 686 Canadian university students provided comprehensive data on their reality TV consumption patterns, sexual attitudes, motivations, and behaviors. Our results replicated several findings from Fogel and Kovalenko (2013), such that participants who reported engaging in one-night stands within the past year demonstrated higher scores on measures of sexual empowerment, sexual permissiveness, and perceived realism of reality show content. Demographic factors such as ethnicity and relationship status were also associated with one-night stand engagement. Additionally, compared to viewers of dating reality TV, viewers of sexual reality TV reported stronger parasocial connections, greater interest in characters, and perceived the show as more realistic. Our results suggest that while reality TV consumption is associated with one-night stand engagement, there are additional factors associated with this outcome and the causal relationship cannot yet be established. We conclude that a broader perspective is needed when assessing reality TV viewership that includes individual and contextual factors. Keywords: reality television, sexual behaviors, replication, open science
The replication crisis heightened interest in methods for assessing the credibility of published research. One approach to evaluate published results is to estimate the average power of original studies based on observed data. Recent criticisms, however, have challenged the validity and usefulness of this approach, arguing that it involves a fundamental "ontological error," fails to predict replication outcomes accurately, and yields imprecise estimates. This article aims to address these critics. We clarify that using observed data to estimate true power is a standard inferential practice and does not constitute an ontological error. We argue that the primary purpose of average power estimation is not to predict the outcome of future replication studies, but the hypothetical outcome if original researchers had to replicate their studies with new samples. Lastly, we demonstrate that even when uncertainty is substantial, average power estimates provide valuable diagnostic information about the credibility of literatures, especially when selection for significance is present. An applied example using a Z-curve analysis of terror management research shows that seemingly strong evidence of over 800 significant results does not rule out the possibility that all results are false positives. We conclude that, despite limitations, average power estimation remains a valid and useful tool for evaluating the evidential value of published research.
Everyone hates acronyms. This is known. This essay focuses on a somewhat different issue, called acronym absurdity, which is how acronyms can have the power to constrain discovery and advancement, thereby shaping science in undesirable ways. It is called acronym absurdity because the use of acronyms can sometimes lead to rather strange outcomes that any reasonable, mildly detached observer can only deem absurd. This essay focuses on two broad classes of the problem: how catchy acronyms lead to bizarre and limiting choices, and how acronyms can lead to modifications that betray their original intention. It is plainly obvious that the general practice around acronyms is absurd and should largely be abandoned.
In times of threat people often turn to social groups to fulfill various needs. In situations where this threat is related to people’s personal sense of control, the model of group-based control provides a social-identity-based account of why thinking and acting in terms of group membership should become more likely to occur. We set out to extend this perspective to the perception and processing of information about norms. Our initial hypothesis, the norm vigilance hypothesis, was that threat to people’s personal sense of control should induce a state where people become more vigilant for information about relevant social norms. This should become evident in more accurate recall of specifically this type of information under conditions of threat. In a series of four studies, we investigated this hypothesis with different paradigms, but were unable to find convincing evidence for the notion of norm vigilance in terms of enhanced accuracy. In a fifth study, we investigated an alternative hypothesis of motivated intergroup distortion of information about social norms after threat, but did not find convincing evidence for this mechanism either.
During his campaign for the Republican Party nomination and for U.S. president, Donald Trump suggested that Hillary Clinton benefited from playing a “woman card”. The effect of exposure to Trump’s woman-card attack was investigated in the Cassese and Holman (2019) Political Psychology article “Playing the woman card: Ambivalent sexism in the 2016 U.S. presidential race”. However, neither Cassese and Holman (2019) nor a reanalysis of data analyzed in the article provided sufficient evidence for key claims in the article. Moreover, Cassese and Holman (2019) is unclear whether its Study 2 experimental data could be used to test claims made based on its Study 1 non-experimental data, providing an example of how journal policy requiring access to survey questionnaires could help peer reviewers and readers better assess reported research.
The use of journal impact factors and other metric indicators of research productivity, such as the h-index, has been heavily criticized for being invalid for the assessment of individual researchers and for fueling a detrimental "publish or perish" culture. Multiple initiatives call for developing alternatives to existing metrics that better reflect quality (instead of quantity) in research assessment. This report, written by a task force established by the German Psychological Society, proposes how responsible research assessment could be done in the field of psychology. We present four principles of responsible research assessment in hiring and promotion and suggest a two-phase assessment procedure that combines the objectivity and efficiency of indicators with a qualitative, discursive assessment of shortlisted candidates. The main aspects of our proposal are (a) to broaden the range of relevant research contributions to include published data sets and research software, along with research papers, and (b) to place greater emphasis on quality and methodological rigor in research evaluation.
For decades, waxing and waning, there has been an ongoing debate on the values and problems of the ubiquitously used null hypothesis significance test (NHST). With the start of the replication crisis, this debate has flared-up once again, especially in the psychology and psychological methods literature. Arguing for or against the NHST method usually takes place in essays and opinion pieces that cover some, but not all the qualities and problems of the method. The NHST literature landscape is vast, a clear overview is lacking, and participants in the debate seem to be talking past one another. To contribute to a resolution, we conducted a systematic review on essay literature concerning NHST published in psychology and psychological methods journals between 2011 and 2018. We extracted all arguments in defense of (20) and against (70) NHST, and we extracted the solutions (33) that were proposed to remedy (some of) the perceived problems of NHST. Unfiltered, these 123 items form a landscape that is prohibitively difficult to keep in one’s sights. Our contribution to the resolution of the NHST debate is twofold. 1) We performed a thematic synthesis of the arguments and solutions, which carves the landscape in a framework of three zones: mild, moderate, and critical. This reduction summarizes groups of arguments and solutions, thus offering a manageable overview of NHST’s qualities, problems, and solutions. 2) We provide the data on the arguments and solutions as a resource for those who will carry-on the debate and/or study the use of NHST.
The primary goal of our target article (Isager et al., 2021) is to give the research community an example of what a well-justified replication value metric could look like, and to encourage discussion of how replication value could be quantified in practice. Furthermore, in the target article we discuss practical hurdles to quantification and possible practical applications for RVCn and other metrics. As that article proposes a method for how to do research–in this case a method to select which claims in the literature need replication most–it is important to receive criticism, feedback, and viewpoints from a diverse range of authors interested in this topic. We are delighted to read the many thoughtful yet critical commentaries, several of which proposing adjustments or alternatives to the equations we have proposed in the target article. This is very encouraging to see, as our aim with initiating this call in Meta-Psychology was to create an open dialogue in the scientific record. RVCn is an efficient but limited metric. Its limitations should be laid bare, and we fully expect that improved metrics and selection procedures can be created in the future. We hope our target article and these commentaries together will inspire readers to continue the discussion of how to efficiently and transparently select studies for replication.
P. Isager et al. (2025) propose a fast approximation of replication value, RVCn, that relies on citation count and sample size. This approximation is simple, transparent, and easy to implement across many studies. It can potentially help metascientists evaluate large collections of studies. However, RVCn is not a precise statement of fact; it should not substitute for detailed substantive and methodological arguments. I make two counterclaims: (1) studies with few citations might be worth replicating and (2) studies with large samples might be worth replicating. While RVCn can helpfully supplement researchers’ judgments, it should not substitute for researchers’ judgments.
The current commentary is focused on the methods and conclusions drawn in a recent meta-analysis which evaluated the impact of standalone interventions in treating anxiety and depressive symptoms (Weisel et al., 2019). The current commentary discusses the large impact of methodological choices made to exclude transdiagnostic treatments and group heterogeneous treatments on study conclusions. Additionally, the current commentary evaluates these conclusions considering opposing from two additional meta-analytic findings. The current review concludes that more research is needed effects before drawing any definitive conclusions, but the current evidence base suggests that apps show promising early efficacy in treating anxiety and depressive symptoms.
The Replication Value (RVCn) metric was introduced to help researchers prioritize studies for replication based on expected utility. While we welcome the introduction of this straightforward and systematic replication decision approach, we identify two limitations of the RVCn. First, when testing the “repeatability” of a study or systematically incorporating replication into a research workflow, the RVCn may not always be the most suitable metric to guide decisions. Use cases should consider the scope conditions of the metric. Second, the RVCn shows limited sensitivity in distinguishing between studies with large sample sizes. To address this, we propose a simple adjustment: a log transformation of the sample size component. This modification improves the metric’s discriminatory power for high-N studies and better aligns the (RVCn) with its intended purpose: guiding efficient and meaningful replication efforts.
This commentary is a response to Isager, P. M., van ’t Veer, A. E. & Lakens, D. (2025): Replication value as a function of citation impact and sample size. It argues that, in assessing "the value of of being correct about the truth status of a claim," it is important to try to capture nonscientific impact. This commentary focuses in particular on the impact that original research can have in a legal context.
The authors (Isager et al., 2025) start with the main assumption that researchers’ efforts toward replications are constrained by resources, and they propose a simple, practically scalable framework of research replication value that guides the researchers and the scientific community at large intending to achieve bigger bang for the buck. Specifically, the authors propose a framework that combines citation impact and sample size of the original articles as a metric for assessing replication value. This implies that original studies with higher scores on this metric can be prioritized for replication efforts. We thoroughly agree with the authors’ assumption and indeed support the view of working towards an optimal framework that helps the community achieve maximum research impact from the replication efforts. In this commentary, we propose to discuss three important limitations that have to be considered before using such metrics. We thoroughly agree with the authors' assumption, and indeed support the view of working towards an optimal framework that helps the community achieve maximum research impact from the replication efforts. In this commentary, we propose to discuss three important limitations that have to be considered before using such metrics.
Replications are essential for rigorous credible science yet are still grossly undervalued and very rare. The value of replications is directly tied to the value of the research they aim to replicate, and replications offer many benefits that go far beyond the mere testing of replicability, such as including verifications and error detection, promoting long-term reproducibility of all research outputs, clarifying theory, refining measurement, and testing generalizability. We need far more independent pre-registered well-powered direct replications to strengthen the credibility of scientific findings. Isager et al. (2025)’s aim to define a formula for the value of replications based on over-simplified metrics of citation count and sample size is misaligned, already misunderstood, and may backfire by hindering the pursuit and publication of replications.
This comment argues that replications should not be prioritized as a function of citation count and sample size, but instead as a function of societal impact, test severity, and the information value of a replication.
Scientific theories reflect some of humanity's greatest epistemic achievements. The best theories motivate us to search for discoveries, guide us towards successful interventions, and help us to explain and organize knowledge. Such theories require a high degree of specificity, which in turn requires formal modeling. Yet, in psychological science, many theories are not precise and psychological scientists often lack the technical skills to formally specify existing theories. This problem raises the question: How can we promote formal theory development in psychology, where there are many content experts but few modelers? In this paper, we discuss one strategy for addressing this issue: a Many Modelers approach. Many Modelers consists of mixed teams of modelers and non-modelers that collaborate to create a formal theory of a phenomenon. Here, we report a proof of concept of this approach, which we piloted as a three-hour hackathon at the Society for the Improvement of Psychological Science conference in 2021. After surveying the participants, results suggest that (a) psychologists who have never developed a formal model can become (more) excited about formal modeling + and theorizing; (b) a division of labor in formal theorizing is possible where only one or a few team members possess the prerequisite modeling expertise; and (c) first working prototypes of a theoretical model can be created in a short period of time. These results show some promise for the many modelers approach as a team science tool for theory development.