In this paper, we critically examine the essay by Shiffrin, Stigler, and Keil (2026) on scientific understanding. We argue that, despite its uncontroversial moral message, the paper fails to articulate a coherent thesis. In particular, it turns on a central notion of an "illusion of understanding" that is not only insufficiently specified but supported by ill-suited examples. Moreover, the analysis rests on a narrow conception of explanation that all but equates understanding with the issuance of causal claims. When invoking the linear regression model and a number of purportedly perplexing or counterintuitive results (e.g., Simpson's Paradox), the essay also fails to clarify what is actually at stake. In the specific case of 'regression to the mean', we find its treatment to be misconceived. We conclude by arguing that analyses of scientific understanding would benefit from a hierarchical framework of the kind proposed by Suppes (1966).
Although most research into risky decision making has focused on simple scenarios – where isolated choices are made independent of one another – many important decisions in life play out across sequences of interdependent events and actions. Despite the ubiquity and importance of such decision problems, we know relatively little about how people manage the complexities of dynamic, multistage decisions. Our work combines techniques from two research traditions to investigate how people handle the challenges of dynamic decision making. We use true-and-error models to estimate the distribution and stability of preference profiles, and the presence of errors. In a complementary analysis we use cognitive modeling based on the Decision Field Theory to investigate the psychological processes underlying dynamic decision making. Decision Field Theory provides a unified framework for testing competing hypotheses about how people collect information and plan for the future. Results from both sets of analyses identify distinct groups of individuals. We discuss the behavioral and cognitive factors distinguishing groups from one another, including degree of planning, strategy shifts, biased information sampling, and effort-saving information processing.
Signal Detection Theory has long served as a cornerstone of psychological research, particularly in recognition memory. Yet its conventional application hinges almost exclusively on the Gaussian assumption—an adherence rooted more in historical convenience than theoretical necessity that comes with a number of well-documented drawbacks. In this work, we critically examine these limitations and pursue a principled parametric alternative: SDT modeling based on extreme-value distributions of event minima. A key feature of this family of models is its grounding in a behavioral principle of invariance under uniform choice-set expansions, a prediction we empirically validate in a novel recognition-memory experiment. Based on this empirical success, we turn our attention to one particular member of this family, the Gumbel_min model, which has the convenient feature of representing discriminability as a shift in distribution. We benchmark the Gumbel_min model against its Gaussian counterpart across multiple recognition-memory tasks, including confidence-rating, ranking, forced-choice, and detection-plus-identification paradigms. Our findings highlight the model's parsimonious yet successful characterization of recognition-memory judgments, as well as the utility of its associated discriminability index, g', which can be directly computed from a single pair of hit and false-alarm rates.
The study of visual working memory has long centered on debates between Signal Detection Theory (SDT) and discrete-slots models. A notable limitation of these debates is the strong reliance on parametric assumptions and selective-influence manipulations that rarely receive direct test. Here we take a different approach by examining whether visual working-memory judgments satisfy the structural constraints implied by a random-scale representation—a general latent-variable framework from which both SDT and discrete-slots models can be derived. In Experiments 1a and 1b, multiple-alternative forced-choice judgments conformed to these constraints, allowing the reconstruction of single-item ROC functions without response-bias manipulations. The reconstructed ROC functions were curved and asymmetric, contradicting the linear predictions of discrete-slots accounts. Experiment 2 provided a complementary failure case, showing that random-scale constraints break down precisely when their theoretical conditions are not met. Experiment 3 extended the findings to item-location bindings in change detection, again yielding an asymmetric ROC function that neither equal-variance SDT nor discrete-slots models can accommodate. Together, the results show that visual working memory judgments respect the general assumptions underlying SDT while undermining the core commitments of discrete-slots theories.
In everyday life, people routinely make decisions that involve irredeemable risks such as death (e.g., while driving). Even though these decisions under extinction risk are common, practically important, and have different properties compared to the types of decisions typically studied by decision scientists, they have received little research attention. The present work advances the formal understanding of decision making under extinction risk by introducing a novel experimental paradigm, the Extinction Gambling Task (EGT). We derive optimal strategies for three different types of extinction and near-extinction events, and compare them to participants' choices in three experiments. Leveraging computational modelling to describe strategies at the individual level, we document strengths and shortcomings in participants' decisions under extinction risk. Specifically, we find that, while participants are relatively good in terms of the qualitative strategies they employ, their decisions are nevertheless affected by loss chasing, scope insensitivity, and opportunity cost neglect. We hope that by formalising decisions under extinction risk and providing a task to study them, this work will facilitate future research on an important topic that has been largely ignored.
Postdoctoral training is a career stage often described as a demanding and anxiety-laden time when many promising PhDs see their academic dreams slip away due to circumstances beyond their control. We use a unique dataset of academic publishing and ...
In recent years, discussions comparing high-threshold and continuous accounts of recognition-memory judgments have increasingly turned their attention toward critical testing. One of the defining features of this approach is its requirement for the relationship between theoretical assumptions and predictions to be laid out in a transparent and precise way. One of the (fortunate) consequences of this requirement is that it encourages researchers to debate the merits of the different assumptions at play. The present work addresses a recent attempt to overturn the dismissal of high-threshold models by getting rid of a background selective-influence assumption. However, it can be shown that the contrast process proposed to explain this violation undermines a more general assumption that we dubbed "single-item generalization." We argue that the case for the dismissal of these assumptions and the claimed support for the proposed high-threshold contrast account does not stand the scrutiny of their theoretical properties and empirical implications. (PsycInfo Database Record (c) 2025 APA, all rights reserved).
Measurement literacy is required for strong scientific reasoning, effective experimental design, conceptual and empirical validation of measurement quantities, and the intelligible interpretation of error in theory construction. This discourse examines how issues in measurement are posed and resolved and addresses potential misunderstandings. Examples drawn from across the sciences are used to show that measurement literacy promotes the goals of scientific discourse and provides the necessary foundation for carving out perspectives and carrying out interventions in science.
We develop alternative families of Bayes factors for use in hypothesis tests as alternatives to the popular default Bayes factors. The alternative Bayes factors are derived for the statistical analyses most commonly used in psychological research – for one-sample and two-sample t tests, for regression and ANOVA analyses. They possess the same desirable theoretical and practical properties as the default Bayes factors and satisfy additional theoretical desiderata while mitigating against two features of the default priors that we consider implausible. They can be conveniently computed via an R package that we provide. Furthermore, hypothesis tests based on Bayes factors and those based on significance tests are juxtaposed. This discussion leads to the new finding that default Bayes factors as well as the alternative Bayes factors are equivalent to test-statistic based Bayes factors as proposed by Johnson (2005). We highlight test-statistic based Bayes factors as a general approach to Bayes-factor computation that is applicable to many hypothesis-testing problems for which an effect-size measure has been proposed and for which test power can be computed.
Researchers have become increasingly aware that data-analysis decisions affect results. Here, we examine this issue systematically for multinomial processing tree (MPT) models, a popular class of cognitive models for categorical data. Specifically, we examine the robustness of MPT model parameter estimates that arise from two important decisions: the level of data aggregation (complete-pooling, no-pooling, or partial-pooling) and the statistical framework (frequentist or Bayesian). These decisions span a multiverse of estimation methods. We synthesized the data from 13,956 participants (164 published data sets) with a meta-analytic strategy and analyzed the magnitude of divergence between estimation methods for the parameters of nine popular MPT models in psychology (e.g., process-dissociation, source monitoring). We further examined moderators as potential sources of divergence. We found that the absolute divergence between estimation methods was small on average (<.04; with MPT parameters ranging between 0 and 1); in some cases, however, divergence amounted to nearly the maximum possible range (.97). Divergence was partly explained by few moderators (e.g., the specific MPT model parameter, uncertainty in parameter estimation), but not by other plausible candidate moderators (e.g., parameter trade-offs, parameter correlations) or their interactions. Partial-pooling methods showed the smallest divergence within and across levels of pooling and thus seem to be an appropriate default method. Using MPT models as an example, we show how transparency and robustness can be increased in the field of cognitive modeling.
This commentary argues against the indictment of current experimental practices such as piecemeal testing, and the proposed integrated experiment design (IED) approach, which we see as yet another attempt at automating scientific thinking. We identify a number of undesirable features of IED that lead us to believe that its broad application will hinder scientific progress.
Decisions about extinction risks are ubiquitous in everyday life and for our continued existence as a species. We introduce a new risky-choice task that can be used to study this topic: The Extinction Gambling Task. Here, we investigate two versions of this task: a Keep variant, where participants cannot accumulate any more earnings after the extinction event, and a Lose variant, where extinction also wipes out all previous earnings. We derive optimal solutions for both variants and compare them to behavioural data. Our findings suggest that people understand the difference between the two variants and their behaviour is qualitatively in line with the optimal solution. Further, we find evidence for risk-aversion in the Keep condition but not in the Lose condition. We hope that this task can facilitate further research on this vital topic.
Objectives Numerous theories exist regarding age differences in risk preference and related constructs, yet many of them offer conflicting predictions and fail to consider convergence between measurement modalities or constructs. To pave the way for conceptual clarification and theoretical refinement, in this preregistered study we aimed to comprehensively examine age effects on risk preference, impulsivity, and self-control using different measurement modalities, and to assess their convergence.Methods We collected a large battery of self-report, informant report, behavioral, hormone, and neuroimaging measures from a cross-sectional sample of 148 (55% female) healthy human participants between 16 and 81 years (mean age = 46 years, standard deviation [SD] = 19). We used an extended sample of 182 participants (54% female, mean age = 46 years, SD = 19) for robustness checks concerning the results from self-reports, informant reports, and behavioral measures. For our main analysis, we performed specification curve analyses to visualize and estimate the convergence between the different modalities and constructs.Results Our multiverse analysis approach revealed convergent results for risk preference, impulsivity, and self-control from self- and informant reports, suggesting a negative effect of age. For behavioral, hormonal, and neuroimaging outcomes, age effects were mostly absent.Discussion Our findings call for conceptual clarification and improved operationalization to capture the putative mechanisms underlying age-related differences in risk preference and related constructs.
Signal detection theory (SDT) provides a prominent framework for modelling recognition memory judgments. Although specific SDT model variants have been extensively tested in working memory, fundamental properties like random-scale representation and receiver operating characteristic (ROC) symmetry have not been critically examined in this domain. Here, we tested these core assumptions across two experiments using multi-alternative forced-choice tasks. In Experiment 1 (N = 123), participants viewed displays of eight images and made recognition judgments among varying numbers of alternatives. Results demonstrated that these decisions satisfied theoretical constraints required by random-scale representations, validating a key assumption of the broader SDT framework. In Experiment 2 (N = 304), we directly tested ROC symmetry by comparing performance between standard recognition tasks and a modified version where participants selected novel rather than studied items. Against common assumption, results provided clear evidence for ROC asymmetry, mirroring patterns previously observed in long-term memory. Together, these findings validate fundamental SDT properties in visual working memory and suggest important commonalities across mnemonic faculties.
The ability to distinguish between different explanations of human memory abilities continues to be the subject of many ongoing theoretical debates. These debates attempt to account for a growing corpus of empirical phenomena in item-memory judgments, which include the list strength effect, the strength-based mirror effect, and output interference. One of the main theoretical contenders is the Retrieving Effectively from Memory (REM) model. We show that REM, in its current form, has difficulties in accounting for source-memory judgments – a situation that calls for its revision. We propose an extended REM model that assumes a local-matching process for source judgments alongside source differentiation. We report a first evaluation of this model’s predictions using three experiments in which we manipulated the relative source-memory strength of different lists of items. Analogous to item-memory judgments, we observed a null list strength effect and a strength-based mirror effect in the case of source memory. In a second evaluation, which relied on a novel experiment alongside two previously published datasets, we evaluated the model’s predictions regarding the manifestation of output interference in item and lack of it in source memory judgments. Our results showed output interference severely affecting the accuracy of item-memory judgments but having a null or negligible impact when it comes to source-memory judgments. Altogether, these results support REM’s core notion of differentiation (for both item and source information) as well as the concept of local matching proposed by the present extension.
This paper examines the practice of piecemeal testing in psychological science. We appraise a recent line of argument against its use in practice based on the conclusions drawn from a thought experiment purporting to show that piecemeal testing, even when confirmatory, can reduce the scope of a scientific theory to the extent that it applies to no one at all. We show that the line of reasoning fails to mount a compelling case against piecemeal testing. In addition to being founded upon a naive and unworkable logico-mathematical conception of scientific theory, the argument effectively imposes a rigid idealization of uncertainty over statistical quantities, enforces an untenable relationship between prediction satisfaction and individual differences, and incorrectly treats instances where no prediction is made as predictive failures.
This paper describes data collected from a cross-sectional convenience sample of 200 healthy human volunteers between 16 and 81 years of age. We assembled an extensive battery of measures of risk preference, impulsivity, and self-control, as well as a range of demographic and cognitive measures, Crucially, we adopted different measure categories, including self-reports, informant reports, behavioral measures, and biological measures (hormones, brain function) to capture individual differences, and adopted a within-participant design. Data collection took place over multiple sessions. First, participants completed a laboratory session at the university during which we collected computer-assisted self-report measures (i.e., standardized questionnaires) as well as behavioral measures using computerized tasks. Second, participants independently completed a home session that included the completion of self-report measures, and the collection of saliva samples. In parallel, we acquired informant reports from up to three individuals nominated by the study participants. Third, participants completed a final session at the local hospital during which we collected structural and functional neuroimaging data, as well as further self-report measures. The data was collected to address questions concerning the developmental trajectories of risk preference and related constructs while assessing the impact of the assessment method; however, we invite fellow researchers to benefit from and further explore the data for research on decision-making under risk and uncertainty in general, and to apply novel analytical approaches (e.g., machine-learning applications to the neuroimaging data). Combining a large set of measures with a within-participant design affords a wealth of opportunities for further insights and a more robust evidence base supporting current theorizing on (age-related) differences in risk preference, impulsivity, and self-control.
Individuals' decisions under risk tend to be in line with the notion that "losses loom larger than gains". This loss aversion in decision making is commonly understood as a stable individual preference that is manifested across different contexts. The presumed stability and generality, which underlies the prominence of loss aversion in the literature at large, has been recently questioned by studies reporting how loss aversion can disappear, and even reverse, as a function of the choice context. The present study investigated whether loss aversion reflects a trait-like attitude of avoiding losses or rather individuals' adaptability to different contexts. We report three experiments investigating the within-subject context sensitivity of loss aversion in a two-alternative forced-choice task. Our results show that the choice context can shift people's loss aversion, though somewhat inconsistently. Moreover, individual estimates of loss aversion are shown to have a considerable degree of stability. Altogether, these results indicate that even though the absolute value of loss aversion can be affected by external factors such as the choice context, estimates of people's loss aversion still capture the relative dispositions towards gains and losses across individuals.
van Doorn et al. (2021) outlined various questions that arise when conducting Bayesian model comparison for mixed effects models. Seven response articles offered their own perspective on the preferred setup for mixed model comparison, on the most appropriate specification of prior distributions, and on the desirability of default recommendations. This article presents a round-table discussion that aims to clarify outstanding issues, explore common ground, and outline practical considerations for any researcher wishing to conduct a Bayesian mixed effects model comparison.