Journal editors play a key role in shaping research practices by deciding whether journals adopt policies related to open science and transparency. Despite sustained advocacy for open science, adoption of these policies remains highly variable across social science journals. This proposed qualitative study will use semi-structured interviews to examine how social science journal editors make decisions about adopting open science-related policies at their journals. Guided by the COM-B model of behaviour change, we will explore editors’ perceived capability, opportunity, and motivation to implement open science related policies at their journals. The study will investigate twelve open science policies and practices, with primary focus on data and code sharing and computational reproducibility, open content peer review, and naming the handling editor on published articles. Additional practices include open identities in peer review, materials sharing, preregistration, and publishing registered reports, null results, direct replications, and post-publication critique. Editors will be recruited through an international online community for social science journal editors, and purposely sampled to capture diverse disciplines, journal economic models, and perspectives. Interview data will be analysed using deductive content analysis informed by the COM-B model alongside inductive thematic analysis. By providing rich descriptive evidence on editors’ experiences, reservations, and constraints, this study aims to inform approaches to promoting open science policy adoption at journals.
Code underlying published findings in the social sciences often fails to reproduce reported results when others re-run it on the original data. This Comment proposes four recommendations to strengthen computational reproducibility, facilitating the trustworthiness and reliability of research.
AbstractThe extent to which application claims are common in published articles can help provide insights into the norms and culture of a field. Estimating the prevalence of application claims, particularly in a literature that is not specifically an applied field, can help to shed light on the extent to which researchers consider their findings applicable to the real world, and/or are incentivized to connect their findings to real-world applications. To investigate how often social and personality psychologists make application claims in their published (mostly basic research) articles, we extracted and categorized application claims from 669 original empirical social and personality psychology articles published from 2010 to 2020 across seven journals. Just over one quarter of articles contained at least one application claim, and one in 20 articles contained an application claim in the Abstract. We did not find a significant association between having an application claim, and articles' (i) citation count or (ii) Altmetric score; however, we did not have sufficient evidence to conclude that there is no association. It is difficult to say whether one in four articles with application claims is too many or too few without diving deeper into the evidence behind the application claims in those articles. Nevertheless, these descriptive results can help inform future scholarship on how incentives and norms in the field shape researchers' practices.
The replication crisis triggered critical reflection on several issues, most prominently, statistical errors and bias (e.g., p-hacking, irreplicability, false positives). This raises the question of whether published statistical results changed throughout psychology’s replication crisis. One tool for assessing the credibility of a set of statistical results is p-curve. Focusing on social and personality psychology, we present descriptive results from two projects that were initially independent. Following the p-curve guide, both teams extracted information about experiments’ key hypotheses, sample sizes, and key statistical results (among other things). The first team (Study 1) coded 1,526 experiments in 335 articles published in the Journal of Social and Personality Psychology: Attitudes and Social Cognition between 2006 and 2015. The second team (Study 2) coded 1,074 experiments in 507 articles published in Social Psychological and Personality Science between 2010 and 2019. The descriptive results suggest improvement during this critical period, however, they are also compatible with a range of scenarios and we therefore refrain from drawing strong conclusions. We invite readers to explore the data and consider the many different underlying research practices that could have produced these results. Our results provide constraints for discussions about the effects of the replication crisis and reform efforts.
Pursuing replicability - independent evidence for previous claims - is important for creating generalizable knowledge(1,2). Here we attempted replications of 274 claims of positive results from 164 quantitative papers published from 2009 to 2018 in 54 journals in the social and behavioural sciences. Replications were high powered on average to detect the original effect size (median of 99.6%), used original materials when relevant and available, and were peer reviewed in advance through a standardized internal protocol. Replications showed statistically significant results in the original pattern for 151 of 274 claims (55.1% (95% confidence interval (CI) 49.2-60.9%)) and for 80.8 of 164 papers (49.3% (95% CI 43.8-54.7%)), weighed for replicating multiple claims per paper. We observed modest variation in replication rates across disciplines (42.5-63.1%), although some estimates had high uncertainty. The median Pearson's r effect size was 0.25 (95% CI 0.21-0.27) for original studies and 0.10 (95% CI 0.09-0.13) for replication studies, an 82.4% (95% CI 67.8-88.2%) reduction in shared variance. Thirteen methods for evaluating replication success provided estimates ranging from 28.6% to 74.8% (median of 49.3%). Some decline in effect size and significance is expected based on power to detect original effects and regression to the mean because we replicated only positive results. We observe that challenges for replicability extend across social-behavioural sciences, illustrating the importance of identifying conditions that promote or inhibit replicability(3,4).
The topics and questions that social and personality psychologists study attract a great deal of public interest. The level of evidence required to make an application claim should be high, in our opinion, to avoid misapplying evidence that could be harmful or ineffective, which could erode the public’s trust in the field. Our field’s challenges with reproducibility, common threats to construct, internal, external, and statistical conclusion validity, and the dearth of independent replications in the field, suggest that application claims should be rare. To investigate how often social and personality psychologists make application claims in their published articles, we extracted and categorized application claims from 669 original empirical social and personality psychology articles published from 2010 to 2020 across sevens. Just over one quarter of articles contained at least one application claim, and 1 in 20 articles contained an application claim in the Abstract. We did not find a significant association between having an application claim, and articles’ i) citation count or ii) Altmetric score, however, we did not have sufficient evidence to conclude that there is no association. The extent to which published application claims are warranted, and are actually applied, are open question for future research.
Self-knowledge plays a central role in contemporary psychological science across various domains, including interpersonal relationships, moral behaviour and health. Despite its importance, many fundamental questions remain. We conducted a pre-registered, expert-based consensus process to address four key gaps in research on self-knowledge: its conceptualization, measurement, outcomes and changeability. Seventeen experts from diverse subfields of psychology participated in a structured Delphi process guided by four facilitators and an external advisor. The panel developed a consensus definition of self-knowledge as the extent to which a person has accurate perceptions of their own relatively stable characteristics and momentary states. Experts further agreed that self-knowledge is largely domain-specific, context-dependent in its benefits, and malleable in principle but difficult to change in practice. Measurement was identified as a central challenge, and avenues for refinement in future work were proposed. Consensus was weaker regarding the existence of a domain-general factor of self-knowledge and shared underlying processes across domains. Overall, the findings clarify where experts converge, where debates persist and what should be prioritized in future research, providing a crucial foundation for advancing the study of self-knowledge across fields. Self-knowledge plays a central role in contemporary psychological science. However, unresolved conceptual and methodological issues have hindered theoretical integration and cumulative scientific progress. In this Consensus Statement, Thielmann et al. identify gaps in four key areas of self-knowledge research: its conceptualization, measurement, outcomes and changeability.
Psychological scientists are increasingly acknowledging the importance of transparency for research integrity. The present study examined one important facet of transparency: providing enough information about measures so that readers can evaluate aspects of construct validity. With a focus on social and personality psychology, we explored how often authors in one journal report a scale name, citation, example item, number of items, and reliability coefficient, as well as how often authors provide access to the study’s materials. We also investigated how measurement reporting practices have changed from 2010 to 2020, the decade encompassing the start of the “credibility revolution” in psychology. Across two samples, we coded 506 Social Psychological and Personality Science (SPPS) articles (N = 425 articles with at least one questionnaire measure; 1,198 questionnaire measures). Overall, ~31% of measures were reported with a name, ~53% a citation, ~66% an example item, ~76% the number of items, and ~78% of multi-item measures included some reliability information; approximately 22% of measures were single-item and 46% were ad hoc. We did not detect any apparent changes in the reporting practices examined from 2010 to 2020 in either sample, except for an increase in the availability of materials over time. Therefore, the replication crisis may have motivated increased access to studies’ materials in recent years but otherwise does not seem to be associated with more transparent reporting of measurement information for questionnaires in brief social and personality articles.
Journal-based peer review is a widely used form of gatekeeping in science, and publication in a peer-reviewed journal is often treated as a reliable indicator of credible research. In this paper, we argue that this view of journal-based peer review is misguided. Currently, journal-based peer review rarely prioritises checking accuracy, and journals typically do not require submitted research to be reported transparently enough to do so. Moreover, vetting research requires much more time and expertise than can be feasibly achieved by an editor and a small number of peer-reviewers. With more investment in accuracy checking, journal-based peer review could serve as a reliable initial quality filter. However, comprehensive quality vetting requires an expanded pool of experts, which, we argue, can only be achieved via post-publication review. At present, there is little infrastructure to support post-publication review and few incentives for researchers to engage in it. We propose a fully transparent system in which all manuscripts submitted for publication are posted to a central repository alongside editorial decisions, peer review reports, data, code, and as they become available, post-publication reviews. Post-publication review efforts could be directed toward the most promising, influential, or high-stakes research. Evidence syntheses, such as meta-analyses, offer a natural focal point for post-publication review, provided they prioritise rigorous assessment of the primary literature.
Inadequacies in the conduct and quality of research are well established across many research domains, including sport science and medicine. Metaresearch-the practice of performing research on research-is presented as a practical vehicle for improving research quality through evaluating the research processes. This article introduces the concept of metaresearch to sport as a new sub-field of sport science. The broad types of metaresearch are introduced, with a mapping of current sports metaresearch activity across these areas. Interdisciplinary centres aimed at improving scientific quality across other fields are also introduced to sport, and specific considerations for beginning metaresearch are provided for sport. This includes, for example, not performing metaresearch poorly, beginning evaluative metaresearch early to intervene before bad practice becomes normalised, leveraging required interdisciplinary expertise depending on the metaresearch question and undertaking an ethical approach for carrying out evaluation of research quality.
Women are widely assumed to be more talkative than men. Challenging this assumption, Mehl et al. (2007) provided empirical evidence that men and women do not differ significantly in their daily word use, speaking about 16,000 words per day (WPD) each. However, concerns were raised that their sample was too small to yield generalizable estimates and too age and context homogeneous to permit inferences beyond college students. This registered report replicated and extended the previous study of binary gender differences in daily word use to address these concerns. Across 2,197 participants (more than five-fold the original sample size), pooled over 22 samples (631,030 ambient audio recordings), men spoke on average 11,950 WPD and women 13,349 WPD, with very large individual differences (<100 to >120,000 WPD). The estimated gender difference (1,073 WPD; d = 0.13; 95% CrI [316, 1,824]) was about twice as large as in the original study. Smaller differences emerged among adolescent (513 WPD), emerging adult (841 WPD), and older adult (-788 WPD) participants, but a substantially larger difference emerged for participants in early and middle adulthood (3,275 WPD; d = 0.32). Despite the considerable sample size(s), all estimates carried large statistical uncertainty and, except for the gender difference in early and middle adulthood, provide inconclusive evidence regarding whether the two genders ultimately speak a practically equivalent number of WPD, based on the preregistered +/- 1,000 WPD regions of practical equivalence criterion. Experienced stress had no meaningful effect on the gender difference, and no clear pattern emerged as to whether the gender difference is accentuated for subjectively rated compared with objectively observed talkativeness.
Postpublication critique, such as letters to the editor, can contribute to the validity and trustworthiness of scientific research. We conducted a cross-sectional analysis of the policy and practice of postpublication critique in (a) randomly selected (N = 100) and (b) prominent (N = 100) psychology journals. In 2023, an explicit submission option for postpublication critique was available at 23% (95% confidence interval [CI] = [16%, 32%]) of randomly sampled psychology journals and 38% of the most prominent psychology journals. Journals sometimes imposed limits on the length and time allowed to submit critiques. We manually inspected two random samples of empirical articles published in 2020 (articles per sample: N = 101), estimating the prevalence of postpublication critique to be 0% (95% CI = [0%, 3.7%]) in psychology journals generally and 1% (95% CI = [0.2%, 5.4%]) in the most prominent psychology journals. The policy and practice of postpublication critique is seriously neglected in psychology journals.
Research articles published by the journal eLife are accompanied by short evaluation statements that use phrases from a prescribed vocabulary to evaluate research on two dimensions: importance and strength of support. Intuitively, the prescribed phrases appear to be highly synonymous (e.g., important/valuable, compelling/convincing) and the vocabulary’s ordinal structure may not be obvious to readers. We conducted an online repeated-measures experiment to gauge whether the phrases were interpreted as intended. We also tested an alternative vocabulary with (in our view) a less ambiguous structure. 301 participants with a doctoral or graduate degree used a 0-100% scale to rate the importance and strength of support of hypothetical studies described using phrases from both vocabularies. For the eLife vocabulary, most participants’ implied ranking did not match the intended ranking on both the importance ( n = 59, 20% matched, 95% confidence interval [15% to 24%]) and strength of support dimensions ( n = 45, 15% matched [11% to 20%]). By contrast, for the alternative vocabulary, most participants’ implied ranking did match the intended ranking on both the importance ( n = 188, 62% matched [57% to 68%]) and strength of support dimensions ( n = 201, 67% matched [62% to 72%]). eLife’s vocabulary tended to produce less consistent between-person interpretations, though the alternative vocabulary still elicited some overlapping interpretations away from the middle of the scale. We speculate that explicit presentation of a vocabulary’s intended ordinal structure could improve interpretation. Overall, these findings suggest that more structured and less ambiguous language can improve communication of research evaluations. ### Competing Interest Statement Simine Vazire is a member of the PLOS Board of Directors. All other authors declare they have no conflicts of interest.
MetaMelb Research Group, School of BioSciences, The University of Melbourne, Melbourne, Victoria, Australia Melbourne Medical School, Faculty of Medicine, Dentistry & Health Sciences, The University of Melbourne, Melbourne, Victoria, Australia The Sir Peter MacCallum Department of Oncology, The University of Melbourne, Melbourne, Victoria, Australia School of Public Health and Preventive Medicine, Monash University, Melbourne, Victoria, Australia Melbourne School of Psychological Sciences, The University of Melbourne, Melbourne, Victoria, Australia School of Historical and Philosophical Studies, The University of Melbourne, Parkville, Victoria, Australia
Background Scientists are increasingly concerned with making their work easy to verify and build upon. Associated practices include sharing data, materials, and analytic scripts, and preregistering protocols. This shift towards increased transparency and rigor has been referred to as a “credibility revolution.” The credibility of empirical legal research has been questioned in the past due to its distinctive peer review system and because the legal background of its researchers means that many often are not trained in study design or statistics. Still, there has been no systematic study of transparency and credibility-related characteristics of published empirical legal research. Methods To fill this gap and provide an estimate of current practices that can be tracked as the field evolves, we assessed 300 empirical articles from highly ranked law journals including both faculty-edited journals and student-edited journals. Results We found high levels of article accessibility (86%, 95% CI = [82%, 90%]), especially among student-edited journals (100%). Few articles stated that a study’s data are available (19%, 95% CI = [15%, 23%]). Statements of preregistration (3%, 95% CI = [1%, 5%]) and availability of analytic scripts (6%, 95% CI = [4%, 9%]) were very uncommon. (i.e., they collected new data using the study’s reported methods, but found results inconsistent or not as strong as the original). Conclusion We suggest that empirical legal researchers and the journals that publish their work cultivate norms and practices to encourage research credibility. Our estimates may be revisited to track the field’s progress in the coming years.
The replication crisis and subsequent credibility revolution in psychology have highlighted many suboptimal research practices such as p-hacking, overgeneralizing, and a lack of transparency. These practices may have been employed reflexively but upon reflection, they are hard to defend. We suggest that current practices for reporting and discussing study limitations are another example of an area where there is much room for improvement. In this article, we call for more rigorous reporting of study limitations in social and personality psychology articles, and we offer advice for how to do this. We recommend that authors consider what the best argument is against their conclusions (which we call the “steel-person principle”). We consider limitations as threats to construct, internal, external, and statistical conclusion validity (Shadish et al., 2002), and offer some examples for better practice reporting of common study limitations. Our advice has its own limitations — both our representation of current practices and our recommendations are largely based on our own metaresearch and opinions. Nevertheless, we hope that we can prompt researchers to write more deeply and clearly about the limitations of their research, and to hold each other to higher standards when reviewing each other's work.
Research articles published by the journal eLife are accompanied by short evaluation statements that use phrases from a prescribed vocabulary to evaluate research on 2 dimensions: importance and strength of support. Intuitively, the prescribed phrases appear to be highly synonymous (e.g., important/valuable, compelling/convincing) and the vocabulary's ordinal structure may not be obvious to readers. We conducted an online repeated-measures experiment to gauge whether the phrases were interpreted as intended. We also tested an alternative vocabulary with (in our view) a less ambiguous structure. A total of 301 participants with a doctoral or graduate degree used a 0% to 100% scale to rate the importance and strength of support of hypothetical studies described using phrases from both vocabularies. For the eLife vocabulary, most participants' implied ranking did not match the intended ranking on both the importance (n = 59, 20% matched, 95% confidence interval [15% to 24%]) and strength of support dimensions (n = 45, 15% matched [11% to 20%]). By contrast, for the alternative vocabulary, most participants' implied ranking did match the intended ranking on both the importance (n = 188, 62% matched [57% to 68%]) and strength of support dimensions (n = 201, 67% matched [62% to 72%]). eLife's vocabulary tended to produce less consistent between-person interpretations, though the alternative vocabulary still elicited some overlapping interpretations away from the middle of the scale. We speculate that explicit presentation of a vocabulary's intended ordinal structure could improve interpretation. Overall, these findings suggest that more structured and less ambiguous language can improve communication of research evaluations.