Behavioral scientists make claims about humans, but most research relies on English-speaking Americans. Despite longstanding concerns about this narrow sampling, only limited progress has been made to address it. We evaluated Chinese online participant pools, a potentially valuable resource given China’s large population, global influence, and frequent comparison with Western societies. Across 16 preregistered studies covering 21 established phenomena from behavioral and cultural sciences, we collected data from 24,932 Chinese and American participants. Chinese online platforms produced data broadly comparable to established American platforms. Participants were geographically and demographically diverse, attentive, and showed high test–retest reliability and temporal stability. Psychological and cultural effects were robust to variation in translation and translator backgrounds. We also provide practical insights and guidance for researchers using Chinese online platforms. This work helps reduce logistical barriers to extending behavioral research beyond traditionally studied WEIRD populations.
Reality is fleeting, and any moment can only be experienced once. Rewatching a video, however, allows people to repeatedly observe the exact same moment. We propose that people may fail to fully distinguish between merely observing behavior again (through replay) from that behavior being performed again in the exact same way. Using an assortment of stimuli that included auditions, commercials, and potential trial evidence, we demonstrated through nine experiments (N = 10,412 adults in the United States) that rewatching makes a recorded behavior appear more rehearsed and less spontaneous, as if the actors were simply precisely repeating their actions. These findings contribute to an emerging literature showing that incidental video features, like perspective or slow motion, can meaningfully change evaluations. Replay may inadvertently shape judgments in both mundane and consequential contexts. To understand how a video will influence its viewer, one will need to consider not only its content, but also how often it is viewed.
Understanding how objective quantities are translated into subjective evaluations has long been of interest to social scientists, medical professionals, and policymakers with an interest in how people process and act on quantitative information. The theory of decision by sampling proposes a comparative procedure: Values seem larger or smaller based on how they rank in a comparison set, the decision sample. But what values are included in this decision sample? We identify and test four mechanistic accounts, each suggesting that how previously encountered attribute values are processed determines whether they linger in the sample to guide the subjective interpretation, and thus the influence, of newly encountered values. Testing our ideas through studies of loss aversion, delay discounting, and vaccine hesitancy, we find strongest support for one account: Quantities need to be subjectively evaluated-rather than merely encountered-for them to enter the decision sample, alter the subjective interpretation of other values, and then guide decision making. Discussion focuses on how the present findings inform understanding of the nature of the decision sample and identify new research directions for the longstanding question of how comparison standards influence decision making.
Prior theoretical and empirical work suggests that behavioral interventions designed to amplify image concerns promote generosity. Consumer elective pricing — where individuals choose how much to pay for products and services — provides a unique opportunity for evaluating the effectiveness of different interventions in the field. We report data from nine field and one lab experiments (N= 3,192) conducted in nonprofit and for-profit settings, where we test how different image concern manipulations that have been previously shown to influence prosocial behavior affect payments. For each of the 10 experiments, we report corresponding forecasts generated by a separate set of participants (N = 1,592) about how the different manipulations influence generosity. In line with the findings from prior literature, lay individuals presented expect large effects under image manipulations. Yet, in contrast with these predictions and with past literature, we find small or no effects of such manipulations. We discuss implications for policymakers and researchers, who may rely on prior findings to make predictions about the effect of behavioral interventions.
Failures to replicate evidence of new discoveries have forced scientists to ask whether this unreliability is due to suboptimal implementation of methods or whether presumptively optimal methods are not, in fact, optimal. This paper reports an investigation by four coordinated laboratories of the prospective replicability of 16 novel experimental findings using rigour-enhancing practices: confirmatory tests, large sample sizes, preregistration and methodological transparency. In contrast to past systematic replication efforts that reported replication rates averaging 50%, replication attempts here produced the expected effects with significance testing ( P < 0.05) in 86% of attempts, slightly exceeding the maximum expected replicability based on observed effect sizes and sample sizes. When one lab attempted to replicate an effect discovered by another lab, the effect size in the replications was 97% that in the original study. This high replication rate justifies confidence in rigour-enhancing methods to increase the replicability of new discoveries.
When consumers select bundles of goods, they may construct those sequentially (e.g., building a bouquet one flower at a time) or make a single choice of a prepackaged bundle (e.g., selecting an already-complete bouquet). Previous research suggested that the sequential construction of bundles encourages variety seeking. The present research revisits this claim and offers a theoretical explanation rooted in combinatorics and norm communication. When constructing a bundle, a consumer chooses among different choice permutations, but when selecting amongst prepackaged bundles, the consumer typically considers unique choice combinations. Because variety is typically overrepresented among permutations compared to combinations, certain consumers (in particular, those with similar attitudes toward items that could compose a bundle) are induced by these different numbers of pathways to variety to display more or less variety-seeking behavior. This is in part explained by the variety norms communicated by different choice architectures, cues most likely to be inferred and used by those who are indifferent between the potential bundle components and thus looking for guidance. Across 5 studies in the main text and 11 in the , this article tests this account and offers preliminary exploration of newly identified residual effects that the pathways-to-variety account cannot explain.
The linkage between abuse to artisanal cobalt miners—including children—in the Democratic Republic of the Congo (DRC) and use of cobalt in advanced batteries has prompted global supply chain reviews, responsible sourcing initiatives, and ...From 2000 through 2020, demand for cobalt to manufacture batteries grew 26-fold. Eighty-two percent of this growth occurred in China and China’s cobalt refinery production increased 78-fold. Diminished industrial cobalt mine production in the early-to-mid ...
Spending money on one's self, whether to solve a problem, fulfill a need, or increase enjoyment, often heightens one's sense of happiness. It is therefore both surprising and important that people can be even happier after spending money on someone else. We conducted a close replication of a key experiment from Dunn, Aknin, and Norton (2008) to verify and expand upon their findings. Participants were given money and randomly assigned to either spend it on themselves or on someone else. Although the original study (N = 46) found that the latter group was happier, when we used the same analysis in our replication (N = 133), we did not observe a significant difference. However, we report an additional analysis, focused on a more direct measure of happiness, that does show a significant effect in the direction of the original. Follow-up analyses shed new insights into people's predictions about their own and others' happiness and their actual happiness when spending money for themselves or others.
Meta-analysts’ practice of transcribing and numerically combining all results in a research literature can generate uninterpretable and/or misleading conclusions. Meta-analysts should instead critically evaluate studies, draw conclusions only from those that are valid and provide readers with enough information to evaluate those conclusions.
We identify 15 claims Pham and Oh (2020) make to argue against pre‐registration. We agree with 7 of the claims, but think that none of them justify delaying the encouragement and adoption of pre‐registration. Moreover, while the claim they make in their title is correct—pre‐registration is neither necessary nor sufficient for a credible science—this is also true of many our science’s most valuable tools, such as random assignment. Indeed, both random assignment and pre‐registration lead to more credible research. Pre‐registration is a game changer.
Empirical audit and review is an approach to assessing the evidentiary value of a research area. It involves identifying a topic and selecting a cross-section of studies for replication. We apply the method to research on the psychological consequences of scarcity. Starting with the papers citing a seminal publication in the field, we conducted replications of 20 studies that evaluate the role of scarcity priming in pain sensitivity, resource allocation, materialism, and many other domains. There was considerable variability in the replicability, with some strong successes and other undeniable failures. Empirical audit and review does not attempt to assign an overall replication rate for a heterogeneous field, but rather facilitates researchers seeking to incorporate strength of evidence as they refine theories and plan new investigations in the research area. This method allows for an integration of qualitative and quantitative approaches to review and enables the growth of a cumulative science.
Do people have an irrational dislike for risk? People pay less for uncertain prospects than their worst possible outcomes, and researchers have proposed that this effect occurs because people strongly dislike risk. We challenge this proposition across seven studies. Though people seem to irrationally dislike risky prospects when preference is assessed with open-ended pricing measures, such as willingness-to-pay, people display rational responses toward risky prospects when preference is assessed using rating measures, such as ratings of expected enjoyment. This discrepancy does not seem to arise because these measures (a) focus on different components of the uncertainty, (b) rely on context-dependent versus normed scales, or (c) involve voluntarily opting into an uncertain situation. Accordingly, we find that people also display rational responses toward risky prospects with time measures (i.e., willingness-to-wait and anticipated time usage) and choice. We discuss alternative explanations and crucial implications of our effects for both theory and application. This paper was accepted by Yuval Rottenstreich, judgment and decision making.
In this article, we (1) discuss the reasons why pre‐registration is a good idea, both for the field and individual researchers, (2) respond to arguments against pre‐registration, (3) describe how to best write and review a pre‐registration, and (4) comment on pre‐registration’s rapidly accelerating popularity. Along the way, we describe the (big) problem that pre‐registration can solve (i.e., false positives caused by p‐hacking), while also offering viable solutions to the problems that pre‐registration cannot solve (e.g., hidden confounds or fraud). Pre‐registration does not guarantee that every published finding will be true, but without it you can safely bet that many more will be false. It is time for our field to embrace pre‐registration, while taking steps to ensure that it is done right.
Empirical results hinge on analytical decisions that are defensible, arbitrary and motivated. These decisions probably introduce bias (towards the narrative put forward by the authors), and they certainly involve variability not reflected by standard errors. To address this source of noise and bias, we introduce specification curve analysis, which consists of three steps: (1) identifying the set of theoretically justified, statistically valid and non-redundant specifications; (2) displaying the results graphically, allowing readers to identify consequential specifications decisions; and (3) conducting joint inference across all specifications. We illustrate the use of this technique by applying it to three findings from two different papers, one investigating discrimination based on distinctively Black names, the other investigating the effect of assigning female versus male names to hurricanes. Specification curve analysis reveals that one finding is robust, one is weak and one is not robust at all.
When researchers choose to identify and exclude outliers from their data, should they do so across all the data, or within experimental conditions? A survey of recent papers published in the Journal of Experimental Psychology: General shows that both methods are widely used, and common data visualization techniques suggest that outliers should be excluded at the conditionlevel. However, I highlight in the present paper that removing outliers by condition runs against the logic of hypothesis testing, and that this practice leads to unacceptable increases in falsepositive rates. I demonstrate that this conclusion holds true across a variety of statistical tests, exclusion criterion and cutoffs, sample sizes, and data types, and show in simulated experiments that Type I error rates can be as high as 29%. I then replicate this result in the context of a recent paper excluding outliers per condition (Cao, Kong, and Galinsky, 2020). Using the authors’ original data, I show that excluding outliers at the condition level can bring the likelihood of a false-positive result up to 47%, and demonstrate that the exclusion strategy reported by the authors is associated with a 56% Type I error rate. I conclude with a list of alternatives to withincondition exclusions.
This editorial introduces the special issue on marketing science and field experiments. We compare the characteristics of the papers that were submitted and accepted for the special issue and provide several recommendations for researchers. In general, we find field experiment research is greater in the areas of advertising and pricing with digital being the most common channel. We suggest that, beyond the estimation of effects and tests of hypotheses, field experiments can complement structural models; help train targeting policies; and also contribute to the nascent area of real-time, adaptive experimentation. We also discuss how field experiment research with a marketing science orientation can enhance and contribute in the areas of behavioral research and marketing strategy.
People often make judgments about their own and others’ valuations and preferences. Across 12 studies (N=18,818), we find a robust bias in these judgments such that people overestimate the valuations and preferences of others. This overestimation arises because, when making predictions about others, people rely on their intuitive core representation of the experience (e.g., is the experience generally positive?) in lieu of a more complex representation that might also include countervailing aspects (e.g., is any of the experience negative?). We first demonstrate that the overestimation bias is pervasive for a wide range of positive (Studies 1-5) and negative experiences (Study 6). Furthermore, the bias is not merely an artifact of how preferences are measured (Study 7). Consistent with judgments based on core representations, the bias significantly reduces when the core representation is uniformly positive (Studies 8A-8B). Such judgments lead to a paradox in how people see others trade off between valuation and utility (Studies 9A-9B). Specifically, relative to themselves, people believe that an identically-paying other will get more enjoyment from the same experience, but paradoxically, that an identically-enjoying other will pay more for the same experience. Finally, consistent with a core representation explanation, explicitly prompting people to consider the entire distribution of others’ preferences significantly reduced or eliminated the bias (Study 10). These findings suggest that social judgments of others’ preferences are not only largely biased, but they also ignore how others make tradeoffs between evaluative metrics.
p-curve, the distribution of significant p-values, can be analyzed to assess if the findings have evidential value, whether p-hacking and file-drawering can be ruled out as the sole explanations for them. Bruns and Ioannidis (2016) have proposed p-curve cannot examine evidential value with observational data. Their discussion confuses false-positive findings with confounded ones, failing to distinguish correlation from causation. We demonstrate this important distinction by showing that a confounded but real, hence replicable association, gun ownership and number of sexual partners, leads to a right-skewed p-curve, while a false-positive one, respondent ID number and trust in the supreme court, leads to a flat p-curve. P-curve can distinguish between replicable and non-replicable findings. The observational nature of the data is not consequential.
Several researchers have relied on, or advocated for, internal meta-analysis, which involves statistically aggregating multiple studies in a paper to assess their overall evidential value. Advocates of internal meta-analysis argue that it provides an efficient approach to increasing statistical power and solving the file-drawer problem. Here we show that the validity of internal meta-analysis rests on the assumption that no studies or analyses were selectively reported. That is, the technique is only valid if (a) all conducted studies were included (i.e., an empty file drawer), and (b) for each included study, exactly one analysis was attempted (i.e., there was no p-hacking). We show that even very small doses of selective reporting invalidate internal meta-analysis. For example, the kind of minimal p-hacking that increases the false-positive rate of 1 study to just 8% increases the false-positive rate of a 10-study internal meta-analysis to 83%. If selective reporting is approximately zero, but not exactly zero, then internal meta-analysis is invalid. To be valid, (a) an internal meta-analysis would need to contain exclusively studies that were properly preregistered, (b) those preregistrations would have to be followed in all essential aspects, and (c) the decision of whether to include a given study in an internal meta-analysis would have to be made before any of those studies are run. (PsycINFO Database Record (c) 2019 APA, all rights reserved).