The same dataset can be analysed in different justifiable ways to answer the same research question, potentially challenging the robustness of empirical science1-3. In this crowd initiative, we investigated the degree to which research findings in the social and behavioural sciences are contingent on analysts' choices. We examined a stratified random sample of 100 studies published between 2009 and 2018, in which, for one claim per study, at least five reanalysts independently reanalysed the original data. The statistical appropriateness of the reanalyses was assessed in peer evaluations, and the robustness indicators were inspected along a range of research characteristics and study designs. We found that 34% of the independent reanalyses yielded the same result (within a tolerance region of ±0.05 Cohen's d) as the original report; with a four times broader tolerance region, this indicator increased to 57%. Of the reanalyses conducted, 74% reached the same conclusion as the original investigation, 24% yielded no effects or inconclusive results and 2% reported the opposite effect. This exploratory study indicates that the common single-path analyses in social and behavioural research should not be simply assumed to be robust to alternative analyses4. Therefore, we recommend the development and use of practices to explore and communicate this neglected source of uncertainty.
Many-analysts studies explore how well an empirical claim withstands plausible alternative analyses of the same data set by multiple, independent analysis teams. Conclusions from these studies typically rely on a single outcome metric (e.g., effect size) provided by each analysis team. Although informative about the range of plausible effects in a data set, a single effect size from each team does not provide a complete, nuanced understanding of how analysis choices are related to the outcome. We used the Delphi consensus technique with input from 37 experts to develop an 18-item Subjective Evidence Evaluation Survey (SEES) to evaluate how each analysis team views the methodological appropriateness of the research design and the strength of evidence for the hypothesis. We illustrate the usefulness of the SEES in providing richer evidence assessment with pilot data from a previous many-analysts study.
Science is often perceived to be a self-correcting enterprise. In principle, the assessment of scientific claims is supposed to proceed in a cumulative fashion, with the reigning theories of the day progressively approximating truth more accurately over time. In practice, however, cumulative self-correction tends to proceed less efficiently than one might naively suppose. Far from evaluating new evidence dispassionately and infallibly, individual scientists often cling stubbornly to prior findings. Here we explore the dynamics of scientific self-correction at an individual rather than collective level. In 13 written statements, researchers from diverse branches of psychology share why and how they have lost confidence in one of their own published findings. We qualitatively characterize these disclosures and explore their implications. A cross-disciplinary survey suggests that such loss-of-confidence sentiments are surprisingly common among members of the broader scientific population yet rarely become part of the public record. We argue that removing barriers to self-correction at the individual level is imperative if the scientific community as a whole is to achieve the ideal of efficient self-correction.
To what extent are research results influenced by subjective decisions that scientists make as they design studies? Fifteen research teams independently designed studies to answer five original research questions related to moral judgments, negotiations, and implicit cognition. Participants from 2 separate large samples (total N > 15,000) were then randomly assigned to complete 1 version of each study. Effect sizes varied dramatically across different sets of materials designed to test the same hypothesis: Materials from different teams rendered statistically significant effects in opposite directions for 4 of 5 hypotheses, with the narrowest range in estimates being d = -0.37 to + 0.26. Meta-analysis and a Bayesian perspective on the results revealed overall support for 2 hypotheses and a lack of support for 3 hypotheses. Overall, practically none of the variability in effect sizes was attributable to the skill of the research team in designing materials, whereas considerable variability was attributable to the hypothesis being tested. In a forecasting survey, predictions of other scientists were significantly correlated with study results, both across and within hypotheses. Crowdsourced testing of research hypotheses helps reveal the true consistency of empirical support for a scientific claim. (PsycINFO Database Record (c) 2020 APA, all rights reserved).
Crowdsourcing research can balance discussions, validate findings and better inform policy, say Raphael Silberzahn and Eric L. Uhlmann.
The effects of exposure to violent video games on automatic associations with the self were investigated in a sample of 121 students. Playing the violent video game Doom led participants to associate themselves with aggressive traits and actions on the Implicit Association Test. In addition, self-reported prior exposure to violent video games predicted automatic aggressive self-concept, above and beyond self-reported aggression. Results suggest that playing violent video games can lead to the automatic learning of aggressive self-views.
Are current theories of moral responsibility missing a factor in the attribution of blame and praise? Four studies demonstrated that even when cause, intention, and outcome (factors generally assumed to be sufficient for the ascription of moral responsibility) are all present, blame and praise are discounted when the factors are not linked together in the usual manner (i.e., cases of "causal deviance"). Experiment 4 further demonstrates that this effect of causal deviance is driven by intuitive gut feelings of right and wrong, not logical deliberation.
An important consideration in judging the blameworthiness (or praiseworthiness) of an action is whether the agent had sufficient control over it. In three experiments, we investigated judgments of moral blame and praise elicited when individuals were presented with vignettes describing actions that were performed either carefully and deliberately or impulsively and uncontrollably. Experiment 1 uncovered an asymmetry between judgments of positive versus negative actions--negative impulsive actions elicited a discounting of moral blame, but positive impulsive actions did not elicit a discounting of moral praise. Experiments 2 and 3 showed that this asymmetry arises because individuals judge agents on the basis of their metadesires (the degree to which the agents embrace or reject the impulses leading to their actions). Individuals assume that an agent would embrace an uncontrollable positive impulse, and reject an uncontrollable negative impulse.
Two experiments examined the influence of skin color on American Hispanics' and Chileans' attitudes towards their ethnic ingroup and toward subgroups within their ingroup. When implicit attitudes were examined, both American Hispanics and Chileans expressed strong preference for the lighter complexioned subgroup (“Blanco” in Spanish) over the darker complexioned subgroup (“Moreno” in Spanish) within their ethnic ingroup. Implicit preference for Blancos was evident among self-identified Moreno as well as Blanco participants in both countries, suggesting that the desirability of light skin apparently supersedes national boundaries and can reverse the ubiquitious ingroup favoritism effect usually obtained in intergroup research. When participants' implicit attitudes towards Hispanics versus Caucasians were assessed, national differences emerged: Chileans expressed implicit preference for Caucasians over Hispanics whereas American Hispanics did not favor either group. Self-report measures of attitudes revealed less consistent evidence of prejudice and preference based on skin color.