The credibility revolution in psychology and related sciences contributed to the adoption of large-scale research initiatives known as Big Team Science (BTS). BTS has made significant advances in addressing issues of replication, statistical power, and diversity through the use of larger samples and more representative cross-cultural data. However, while these collaborations hold great potential, they also introduce unique challenges related to their scale. Drawing on experiences from successful BTS projects, we identified and outlined key strategies for overcoming diversity, volunteering, and capacity challenges. We emphasize the need for clear role definitions, structured and preregistered workflows, centralized project management, and transparent decision documenting to prevent common pitfalls. Ultimately, we call for reflection on the strengths and limitations of BTS to enhance the quality, generalizability, and impact of research across disciplines. This work complements existing BTS guides by offering experientially-grounded, discipline-specific strategies and addressing underexplored logistical, ethical, and epistemological challenges in large-scale collaborations.
When processing and analyzing empirical data, researchers regularly face choices that may appear arbitrary (e.g., how to define and handle outliers). If one chooses to exclusively focus on a particular option and conduct a single analysis, its outcome might be of limited utility. That is, one remains agnostic regarding the generalizability of the results, because plausible alternative paths remain unexplored. A multiverse analysis offers a solution to this issue by exploring the various choices pertaining to data-processing and/or model building, and examining their impact on the conclusion of a study. However, even though multiverse analyses are arguably less susceptible to biases compared to the typical single-pathway approach, it is still possible to selectively add or omit pathways. To address this issue, we outline a novel, more principled approach to conducting multiverse analyses through crowdsourcing. The approach is detailed in a step-by-step tutorial to facilitate its implementation. We also provide a worked-out illustration featuring the Semantic Priming Across Many Languages project, thereby demonstrating its feasibility and its ability to increase objectivity and transparency. (PsycInfo Database Record (c) 2026 APA, all rights reserved).
This paper provides a step-by-step guide to developing a neuroscience themed escape room. We designed the escape room based on our introductory neuroscience learning outcomes which required students to remember key concepts while working together both as a group and individually to solve six neuroscience-themed challenges. Data include time to escape, as well as the results of a post-event survey that had individual students rate the value of the activity, their own personal effort, and their perceptions of instructor contribution. We found that students enjoyed this activity and that the amount of personal effort put in by the student was correlated with how fast they solved the six challenges in our escape room. We conclude that the escape room is a low cost, high impact event that can motivate student learning of neuroscience and promote retention.
Semantic priming has been studied for nearly 50 years across various experimental manipulations and theoretical frameworks. Although previous studies provide insight into the cognitive underpinnings of semantic representations, they have suffered from small sample sizes and a lack of linguistic and cultural diversity. In this Registered Report, we measured the size and the variability of the semantic priming effect across 19 languages (N = 25,163 participants analyzed) by creating the largest available database of semantic priming values based on an adaptive sampling procedure. We found evidence for semantic priming in terms of differences in response latencies between related word-pair conditions and unrelated word-pair conditions. Model comparisons showed that inclusion of a random intercept for language improved model fit, providing support for variability in semantic priming across languages. This study highlights the robustness and variability of semantic priming across languages and provides a rich, linguistically diverse dataset for further analysis.
Acute bouts of exercise have been shown to have measurable positive impacts on cognition. Here participants either watched a movie (control), walked (moderate exercise), or ran (vigorous exercise) on a treadmill for 30 min while their heart rate was measured before completing a paired associative learning task in which they learned 40 word pairs over the course of 10 trials. We defined learning rate as how fast the participants correctly learned the word pairs. Two days later, all participants were given a surprise recall task, and we defined long-term memory as the number of word pairs correctly recalled. We also measured working memory capacity, anxiety, and sleep quality. We found that while there was no difference between exercise conditions in the rate of learning, participants in the vigorous condition recalled more word pairs 2 days later. Analyses revealed that average heart rate and condition were the only significant predictors of long-term recall. Potential mechanisms to explain the benefits of the vigorous exercise condition on long-term retention, but not on short-term retention, are discussed.
The COVID-19 pandemic (and its aftermath) highlights a critical need to communicate health information effectively to the global public. Given that subtle differences in information framing can have meaningful effects on behavior, behavioral science research highlights a pressing question: Is it more effective to frame COVID-19 health messages in terms of potential losses (e.g., “If you do not practice these steps, you can endanger yourself and others”) or potential gains (e.g., “If you practice these steps, you can protect yourself and others”)? Collecting data in 48 languages from 15,929 participants in 84 countries, we experimentally tested the effects of message framing on COVID-19-related judgments, intentions, and feelings. Loss- (vs. gain-) framed messages increased self-reported anxiety among participants cross-nationally with little-to-no impact on policy attitudes, behavioral intentions, or information seeking relevant to pandemic risks. These results were consistent across 84 countries, three variations of the message framing wording, and 560 data processing and analytic choices. Thus, results provide an empirical answer to a global communication question and highlight the emotional toll of loss-framed messages. Critically, this work demonstrates the importance of considering unintended affective consequences when evaluating nudge-style interventions.
Finding communication strategies that effectively motivate social distancing continues to be a global public health priority during the COVID-19 pandemic. This cross-country, preregistered experiment (n = 25,718 from 89 countries) tested hypotheses concerning generalizable positive and negative outcomes of social distancing messages that promoted personal agency and reflective choices (i.e., an autonomy-supportive message) or were restrictive and shaming (i.e. a controlling message) compared to no message at all. Results partially supported experimental hypotheses in that the controlling message increased controlled motivation (a poorly-internalized form of motivation relying on shame, guilt, and fear of social consequences) relative to no message. On the other hand, the autonomy-supportive message lowered feelings of defiance compared to the controlling message, but the controlling message did not differ from receiving no message at all. Unexpectedly, messages did not influence autonomous motivation (a highly-internalized form of motivation relying on one’s core values) or behavioral intentions. Results supported hypothesized associations between people’s existing autonomous and controlled motivations and self-reported behavioral intentions to engage in social distancing: Controlled motivation was associated with more defiance and less long-term behavioral intentions to engage in social distancing, whereas autonomous motivation was associated with less defiance and more short- and long-term intentions to social distance. Overall, this work highlights the potential harm of using shaming and pressuring language in public health communication, with implications for the current and future global health challenges.
Inhibitory control is a key executive function and has been studied extensively using the stop signal task. By applying a simple race model that posits an independent race between a GO process responsible for initiation of responses and a STOP process responsible for inhibition of responses, one can estimate how long it takes an individual to inhibit an ongoing response, the stop signal reaction time. Here, we examined how stop signal reaction time can be affected by working memory. Participants engaged in a dual task; they completed a stop signal task under low and high working memory load conditions. Working memory capacity was also measured. We found that the STOP process was lengthened in the high, compared to the low, working memory load condition, as evidenced by differences in stop signal reaction time. The GO process was unaffected and working memory capacity could not account for differences across the load conditions. These results indicate that inhibitory control can be influenced by placing demands on working memory.
The COVID-19 pandemic has increased negative emotions and decreased positive emotions globally. Left unchecked, these emotional changes might have a wide array of adverse impacts. To reduce negative emotions and increase positive emotions, we tested the effectiveness of reappraisal, an emotion-regulation strategy that modifies how one thinks about a situation. Participants from 87 countries and regions (n = 21,644) were randomly assigned to one of two brief reappraisal interventions (reconstrual or repurposing) or one of two control conditions (active or passive). Results revealed that both reappraisal interventions (vesus both control conditions) consistently reduced negative emotions and increased positive emotions across different measures. Reconstrual and repurposing interventions had similar effects. Importantly, planned exploratory analyses indicated that reappraisal interventions did not reduce intentions to practice preventive health behaviours. The findings demonstrate the viability of creating scalable, low-cost interventions for use around the world. The stage 1 protocol for this Registered Report was accepted in principle on 12 May 2020. The protocol, as accepted by the journal, can be found at https://doi.org/10.6084/m9.figshare.c.4878591.v1 This Registered Report presents evidence from 87 countries and regions showing that brief emotion-regulation interventions consistently reduced negative emotions and increased positive emotions during the COVID-19 pandemic.
Replications in psychological science sometimes fail to reproduce prior findings. If replications use methods that are unfaithful to the original study or ineffective in eliciting the phenomenon of interest, then a failure to replicate may be a failure of the protocol rather than a challenge to the original finding. Formal pre-data collection peer review by experts may address shortcomings and increase replicability rates. We selected 10 replications from the Reproducibility Project: Psychology (RP:P; Open Science Collaboration, 2015) in which the original authors had expressed concerns about the replication designs before data collection and only one of which was “statistically significant” (p < .05). Commenters suggested that lack of adherence to expert review and low-powered tests were the reasons that most of these RP:P studies failed to replicate (Gilbert et al., 2016). We revised the replication protocols and received formal peer review prior to conducting new replications. We administered the RP:P and Revised protocols in multiple laboratories (Median number of laboratories per original study = 6.5; Range 3 to 9; Median total sample = 1279.5; Range 276 to 3512) for high-powered tests of each original finding with both protocols. Overall, Revised protocols produced similar effect sizes as RP:P protocols following the preregistered analysis plan (Δr = .002 or .014, depending on analytic approach). The median effect size for Revised protocols (r = .05) was similar to RP:P protocols (r = .04) and the original RP:P replications (r = .11), and smaller than the original studies (r = .37). The cumulative evidence of original study and three replication attempts suggests that effect sizes for all 10 (median r = .07; range .00 to .15) are 78% smaller on average than original findings (median r = .37; range .19 to .50), with very precisely estimated effects.
Replications in psychological science sometimes fail to reproduce prior findings. If replications use methods that are unfaithful to the original study or ineffective in eliciting the phenomenon of interest, then a failure to replicate may be a failure of the protocol rather than a challenge to the original finding. Formal pre-data collection peer review by experts may address shortcomings and increase replicability rates. We selected 10 replications from the Reproducibility Project: Psychology (RP:P; Open Science Collaboration, 2015) in which the original authors had expressed concerns about the replication designs before data collection and only one of which was “statistically significant” (p < .05). Commenters suggested that lack of adherence to expert review and low-powered tests were the reasons that most of these RP:P studies failed to replicate (Gilbert et al., 2016). We revised the replication protocols and received formal peer review prior to conducting new replications. We administered the RP:P and Revised protocols in multiple laboratories (Median number of laboratories per original study = 6.5; Range 3 to 9; Median total sample = 1279.5; Range 276 to 3512) for high-powered tests of each original finding with both protocols. Overall, Revised protocols produced similar effect sizes as RP:P protocols following the preregistered analysis plan (Δr = .002 or .014, depending on analytic approach). The median effect size for Revised protocols (r = .05) was similar to RP:P protocols (r = .04) and the original RP:P replications (r = .11), and smaller than the original studies (r = .37). The cumulative evidence of original study and three replication attempts suggests that effect sizes for all 10 (median r = .07; range .00 to .15) are 78% smaller on average than original findings (median r = .37; range .19 to .50), with very precisely estimated effects.
In a test of their global-/local-processing-style model, Förster, Liberman, and Kuschel (2008) found that people assimilate a primed concept (e.g., “aggressive”) into their social judgments after a global prime (e.g., they rate a person as being more aggressive than do people in a no-prime condition) but contrast their judgment away from the primed concept after a local prime (e.g., they rate the person as being less aggressive than do people in a no prime-condition). This effect was not replicated by Reinhard (2015) in the Reproducibility Project: Psychology. However, the authors of the original study noted that the replication could not provide a test of the moderation effect because priming did not occur. They suggested that the primes might have been insufficiently applicable and the scenarios insufficiently ambiguous to produce priming. In the current replication project, we used both Reinhard’s protocol and a revised protocol that was designed to increase the likelihood of priming, to test the original authors’ suggested explanation for why Reinhard did not observe the moderation effect. Teams from nine universities contributed to this project. We first conducted a pilot study ( N = 530) and successfully selected ambiguous scenarios for each site. We then pilot-tested the aggression prime at five different sites ( N = 363) and found that it did not successfully produce priming. In agreement with the first author of the original report, we replaced the prime with a task that successfully primed aggression (hostility) in a pilot study by McCarthy et al. (2018). In the final replication study ( N = 1,460), we did not find moderation by protocol type, and judgment patterns in both protocols were inconsistent with the effects observed in the original study. We discuss these findings and possible explanations.
Objective: In recent years, immersive videogame technologies such as virtual reality have been shown to affect psychological welfare in such way that they can be applied to clinical psychology treatments. However, the effects of videogaming with other immersive gaming apparatuses such as commercial electroencephalography (EEG)-based brain-computer interfaces (BCIs) on psychological welfare have not been extensively researched. Thus, we aimed at providing early insights into some of these effects by looking at how videogaming with a commercial EEG-based BCI would impact mood and physiological arousal. Materials and Methods: A total of 26 participants were sampled. Participants were randomly assigned to either a BCI condition or a traditional condition wherein they played an action videogame with a commercial EEG-based BCI or a standard keyboard and mouse interface for 20 minutes. In both conditions, participants filled out the profile of mood states to assess mood and the perceived stress scale to control for stress. We also measured heart rate, heart rate variability as measured by the root mean square of successive differences, and galvanic skin response (GSR) amplitude differences. Results: Participants in the BCI condition overall reported a significantly higher total mood disturbance (P < 0.05), tension (P < 0.05), confusion (P < 0.05), and significantly less vigor (P < 0.05). We also found that participants in the BCI condition had significantly lower GSR amplitude differences between gaming and baseline (P < 0.05). Conclusion: The results suggest that the use of commercial EEG-based BCIs for playing with videogames can induce greater frustration and negative moods than playing with a traditional keyboard and mouse interface, possibly limiting their use in clinical psychology settings.