
Understanding how conflict dynamics work in conflict tasks is crucial for unraveling why we can effectively focus on target information while ignoring distractor information. One way to examine conflict dynamics between distractor and target information is to manipulate distractor-target stimulus onset asynchronies (SOAs), which influence the allocation of attention and subsequent adaptations to conflict in the next trial. Both the spatial Stroop and flanker paradigms have been used to investigate the effects of SOA manipulations, revealing distinct conflict dynamics across different SOA conditions. However, it remains unclear whether the underlying cognitive mechanisms responsible for these differing patterns of effects can be explained within a common attention shift framework. To investigate this research question, we integrated an experimental manipulation of the spatial Stroop and flanker tasks with a model-based approach using the revised diffusion model for conflict tasks Lee and Sewell (2024). Our examination of the spatial Stroop and flanker tasks revealed a shared mechanism for within-trial conflict adjustment, where evidence is extracted from the target faster when the target is preceded by a distractor. Furthermore, the results suggest that when distractor information appears first, both tasks share common attributes of conflict dynamics, including reduced urgency in attentional shifts during congruent trials. Based on the behavioral and model-based evidence, an increasing sequential modulation across SOAs was only observed in the spatial Stroop task but not in the flanker task. Notably, our modeling results suggest that a larger sequential effect when distracting information precedes target information for a longer duration is likely due to stronger attentional control following an incongruent trial. This further implies that conflict adaptation is proactive and selective with respect to the congruency of the previous trial.
Visual search is a fundamental cognitive ability. This study investigates whether Multimodal Large Language Models (MLLMs) exhibit human-like difficulty signatures in visual search tasks. We compared search performance of humans (n = 1,250) and MLLMs using identical 2D and 3D stimuli across different set sizes. Both groups showed efficient performance in feature searches, most clearly when the target had a unique color, but performance degradation in conjunction searches as set sizes increased. Additionally, we found strong correlations between human and MLLM error rates (ρ = 0.82), which suggests that MLLMs are sensitive to similar objective complexities, such as stimulus heterogeneity. However, differences were found as well: whereas humans invested extra search time to respond accurately on target-absent trials, MLLMs exhibited extreme present/absent response biases in complex searches. We conclude that MLLMs replicate high-level human performance signatures, yet their underlying computations differ significantly.
The benefits of routines for daily functioning are widely acknowledged, yet, despite their apparent importance, methods for quantifying routine maintenance and the causes of their disruption remain lacking. Here, we propose a novel means of defining and quantifying routines (transition entropy). Using the transition entropy, we show that routines can be robustly elicited on tasks that require searching through a grid of squares for a hidden target. Over two experiments (N=100 each), we show that use of routines—as quantified by transition entropy—is robustly perturbed by frequent switches between search grids, as locations specific to the currently irrelevant grid become competitive for selection. Using a normative model that tracks task dynamics, we show that disruption to routines can be attributed to altered sensitivity to the odds of success for completing a task. This suggests that routine maintenance may be disrupted by over-sensitivity to a lack of reward early in routine performance, or increased expectations regarding the utility of pursuing other tasks.
Episodic memory is a type of long-term memory that encodes and retrieves personal experiences associated with their context. Previous episodic memory studies showed that the context or preexisting knowledge about retrieved information may influence the performance of memory tasks. Therefore, studying the semantic proximity effect by comparing memory task performance with different levels of semantic relatedness becomes crucial. In natural language processing (NLP) studies, semantic relations can be successfully represented by learning word vectors in a large text corpus using neural networks. This study investigated the impact of semantic factors on delayed free recall tasks by creating lists that include semantically related and unrelated words obtained through pre-trained NLP models and showed how semantic proximity effect and presentation order influence recall performance. The fastText and Word2Vec models were used to obtain Turkish word embeddings, allowing for the organization of words according to their semantic relatedness. Human raters then validated the word lists. The effect of semantic relatedness on recall dynamics was later compared across four different conditions of list relatedness and embedding models used to create the lists (fastText-related, fastText-unrelated, Word2Vec-related, Word2Vec-unrelated). Our results showed a significant positive correlation between cosine similarity values and human judgment, later indicating how semantic proximity effect and presentation order influenced the recall probability and retrieval dynamics. Different levels of semantic relatedness and choice of word embeddings played a role in the likelihood of recall. Therefore, this study suggests that word embeddings obtained from neural networks can represent and manipulate semantic relations in memory studies and that semantic proximity effect and presentation order influence different levels of semantic relatedness recall dynamics.
Good fits of models can produce an “illusion of understanding” when the models lack grounded theoretical interpretations. This is characterized by the recent claims surrounding the good fit of the Target Confusability Competition (TCC) model of visual working memory (VWM). Despite that TCC appears to assume no capacity limits in VWM, we recently found that TCC can outperform capacity-limit models even when capacity limits are the ground truth, reflecting TCC’s strong flexibility. This flexibility renders the model unfalsifiable in practice and obscures its mechanistic interpretation. We further argue that the distinction between memory failure and degraded memory precision, captured by the capacity-limit models (e.g., the Mixture model), are supported by converging empirical evidence, highlighting the necessity of empirically testable model predictions for achieving genuine understandings of mechanisms.
Algorithmic recommendations have drastically expanded in recent years to aid human decision-making. In this paper, we seek to understand the users of these tools and when, where, and why they obtain algorithmic advice. We do so examining data from two behavioural decision-making experiments (N = 216) and applying the Timed Racing Diffusion Model (TRDM) across choices and response times. Our experiments find that people are sensitive to when algorithmic advice is worthwhile obtaining. Notably, our results privilege experience and show that opportunities to test the recommendation accuracy can be as useful as descriptive information stating the same. Our main finding, however, centers on the time-course of when individuals choose to obtain a recommendation. We find that over time, algorithmic advice is sought as a means to terminate difficult decisions that one cannot derive on one’s own. The TRDM proposes a unifying cognitive mechanism for this pattern of recommendation seeking based on decision urgency though our individual differences analyses identify a diversity of strategies adapted to the same decision environment. Overall, our findings characterise decision-makers as adept users of decision aid tools, and that despite the possibility of recommendation errors, individuals are capable of appreciating the utility of helpful, albeit imperfect, recommendations.
In the working memory (WM) field, it has been documented that additional free time improves immediate memory performance. Recently, Leproult and collaborators (2024) demonstrated that beyond the total amount of free time, its distribution also played a crucial role. Specifically, WM performance was enhanced when free time was provided in massed rather than distributed periods. In the present paper, we demonstrated that the Time-Based Resource-Sharing (TBRS) model and its computational version, relying solely on refreshing as a maintenance mechanism, were unable to account for these outcomes, even when considering the most recent updates in the implementation of refreshing. In line with the conclusions of Leproult et al. (2024), the aim of the present study was to examine whether incorporating a mechanism reproducing the semantic maintenance effect into the TBRS* model could address this gap. We extensively tested this new TBRS*-S+ model by evaluating its capacity to simultaneously reproduce well-established primacy and recency effects, the results observed across the four experiments reported in Leproult et al. (2024), and unpublished data investigating the interaction between concreteness and free-time distribution. Furthermore, the TBRS*-S+ model was challenged by assessing its ability to simulate the extra free-time advantage in simple span tasks. Encouragingly, results showed that, whereas the original TBRS* model consistently failed, the TBRS*-S+ model accurately reproduced all these aspects of human performance. Based on these results, new insights into the functioning of refreshing processes within WM were provided. Finally, we discussed alternative theoretical and computational frameworks that might also account for this free-time distribution effect.
A recent paper (van Rooij et al., 2024) claims to have proved that achieving human-like intelligence using learning from data is intractable in a complexity-theoretic sense. We point out that the proof relies on an unjustified assumption about the distribution of (input, output) tuples in the data. We briefly discuss that assumption in the context of two fundamental barriers to repairing the proof: the need to precisely define “human-like,” and the need to account for the fact that a particular machine learning system will have particular inductive biases that are key to the analysis. Another attempt to repair the proof, by focusing on subsets of the data, faces barriers in terms of defining the subsets.
When presented with a large array of possible alternatives consumers must quickly screen out undesirable options before more carefully deliberating over a smaller set. We consider the situation where these screening decisions are made sequentially and independently for each alternative. Understanding how attribute information is processed during these multi-attribute screening decisions provides useful insights into consumer decision making. Previous approaches to this problem have used equipment such as eye trackers, or intervened in the decision making process by obscuring information and requiring the participant reveal it. Here, we classify a broad set of decision strategies into higher-order classes based on whether those strategies assume processing of each attribute is complete or selective, and whether “good” attributes can compensate “bad” attributes. We then introduce a hierarchical latent mixture modelling approach that uses response times and choices to infer the higher-order decision class that best explains each individual’s screening decisions. We test the model against empirical data where the strategy decision makers ought to use was directly manipulated, demonstrating the model identifies the expected attribute processing strategy for all participants. In simulation, we demonstrate good recovery of the exhaustive set of decision classes we investigated, and extended this to show the model appropriately identifies different decision classes when different classes are present across a sample of participants. Our modelling approach thus provides a two-stage, principled solution to the challenge of identifying individual differences in preferential decision making: grouping a large set of candidate decision strategies into a smaller set of higher-order classes, and then discriminating between those higher-order classes in data with quantitative cognitive models of choices and response times. This allows the researcher to relax assumptions that a sample of participants adheres to the same decision strategy, and also capitalising on the benefits of hierarchical modelling to jointly estimate population parameters and random effects.
Researchers have begun using Bayesian hierarchical modeling to study semantic representations, for instance, in the context of natural language quantifiers such as most, few, and more than half. Building on previous work, we propose a Bayesian hierarchical model to disentangle three key semantic parameters: the meaning threshold of quantifiers, the vagueness surrounding meaning thresholds, and response noise. We use this model to test the stability of semantic representations over time and across different paradigms. To examine stability over time, we analyzed existing data ( n=63 ) from Ramotowska et al. (2023). Contrary to the conclusions drawn by the original authors, we found overwhelming evidence in favor of the hypothesis that semantic representations change over time ( BF > 10^304 ). At the same time, we found overwhelming evidence that the relative ordering of meaning thresholds within individuals remained stable ( BF = 4 × 10^24 ). Next, we conducted a new experiment ( n=178 ) to test stability across paradigms, specifically comparing a linguistic paradigm to a visual one. Here too, we found overwhelming support for differences in between-subject variability in meaning thresholds across paradigms ( BF = 7.48 × 10^30 ) and for differences in vagueness ( BF = 1.17 × 10^110 ). Our findings challenge the assumption that semantic representations of logical vocabulary have stable, fixed values, while suggesting that their relative ordering remains stable within individuals. The model we propose provides an effective framework for studying the semantics of quantifiers, detecting individual-level effects, and explicitly accounting for potential instability.
Conditional effects, or interaction effects, do not imply multiplicative effects. However, product terms are the default method for modeling such conditional effects in psychological research. As a result, theoretically plausible conditional effects may go undetected when the functional form is misspecified. Our study had two objectives: (1) evaluate the extent to which non-linear phenomena can be identified as spurious multiplicative (i.e., standard) interaction terms in linear models, (2) assess how well linear models capture stepwise conditional effects. In Study 1, we examined spurious interactions from non-linear main effects. We found that traditional interaction terms were associated with increased Type-I error rates and small effect sizes. Importantly, this was also the case when the predictors were uncorrelated, indicating a mechanism beyond collinearity. Additionally, we found that, if captured, the spurious interaction effects did reduce prediction error on the population level. In Study 2, we simulated genuine conditional effects, following a stepwise pattern. When effects were monotonic, product terms performed adequately, however if the conditional effect is non-monotonic a traditional interaction term in a linear model does not sufficiently capture such an effect. We conclude that relying solely on traditional interaction terms in linear models can be misleading and the failure to replicate interaction effects may partly reflect a specification crisis: Researchers default to one functional form (multiplication) while the underlying theory may dictate a different form, creating a systematic mismatch between theory and model. To validly investigate conditional effects, researchers should specify and justify the expected functional form a priori.
Self-referential processing captures how individuals perceive themselves. When negatively biased, self-referential processing is closely linked to depression. Increasingly, research has employed drift-diffusion modeling (DDM) to analyze the Self-Referential Encoding Task (SRET), but variations in DDM estimation methods and model complexity may influence findings. This meta-analysis examined (1) the overall difference between DDM parameters for negative versus positive self-trait endorsement, and (2) how DDM parameters for negative and positive self-trait endorsement relate to symptoms of depression. We then qualitatively reviewed variations in DDM methods that may impact findings from the overall-meta-analyses, including sample composition, DDM estimation method, model complexity, and criteria for handling reaction time outliers. We found a reliable positivity advantage for drift rate in self-trait endorsement at the latent, dynamic decision-making level. Moreover, we found that greater drift rate for negative and lower drift rate for positive self-trait endorsement were associated with greater symptoms of depression. Effects were present for bias and non-decision time, but were sparse and underpowered. As an emerging research area, the present study underscores the importance of carefully considering sample composition, model estimation methods, model complexity, and handling of reaction time outliers when applying drift-diffusion modeling to self-referential tasks.
According to Shiffrin et al. (2026), scientists incompletely understand the phenomena they study. I agree and expand on Shiffrin et al. with a focus on the necessity to make auxiliary assumptions as an additional reason for incomplete understanding.
Scientists indeed show all the strengths and illusions of other humans. Hence, we should expect mistakes understanding a procedure as complex and unnatural as linear regression. Scientists, like other humans, are not designed to work in a vacuum. They work best with other people, preferably in an adversarial process. Adversarial processes can turn errors due to bias and illusion into features that enable progress.
The target article “Illusions of Understanding in the Sciences” by Shiffrin et al. (Computational Brain Behavior, 2026) demonstrates how scientists tend to misinterpret correlations by declaring one of the variables as the cause of the other. In our commentary, we discuss that even under controlled experiments it is typically not possible to conclusively identify causal relationships beyond a superficial level.
Several forms of illusory understanding in science are identified by Shiffrin et al. (2026). We focus on one they emphasize: the illusion that predictive success implies causal insight. Their diagnosis, we argue, lacks a criterion for when incomplete understanding becomes illusory, conflating predictions that warrant confidence with those that do not. We propose that what separates warranted confidence from illusion is not causal completeness but the risk structure of the prediction: whether the model could have failed had the hypothesis been wrong. The question shifts from “how complete is the causal story?” to “how severe was the test?”.
Human-focused cybersecurity research is a growing area. The focus in this work is studying attacker thinking and reasoning. Oppositional Human Factors (OHF) posits that cognitive biases and heuristics used by attackers during their decision-making process may be exploited by defensive security teams. OHF is positively supported by a variety of experimental data. The current study experimentally tests exploiting attackers’ loss aversion. We manipulated the endowment effect in a cyber environment by giving attackers access, and then by threatening that access we were able to measure whether attacker performance was affected by this bias. A major contribution is the use of demonstrated expert participants combined with a realistic mission and a high-fidelity military cyber range. The results revealed degraded attacker performance across several measures including reduced steps taken in the kill chain, and increased time taken. Loss aversion is a relevant and provocative cyber attacker bias that can be exploited in OHF-based defenses.
Shiffrin, Stigler, and Keil argue that scientists often overestimate the depth of their understanding, even in simple statistical contexts. While their analysis of regression and related paradoxes shows compellingly that predictive success need not imply explanatory mastery, the present commentary questions if partial understanding should necessarily be epistemically problematic. I argue that incomplete explanation is a structural feature of scientific progress rather than necessarily an illusion. The challenge is to acknowledge the limits of models without undermining science’s epistemic status. In this vein, scientific education should emphasize nuanced explanation of the world while maintaining clear standards distinguishing well-grounded explanations from unfounded alternatives.
Script is a fundamental cultural technology for preserving and communicating thought. Mastering literacy is thus essential for accessing these thoughts and communicating them effectively. Hence, understanding the neuro-cognitive mechanisms underpinning the processes of learning to read is highly relevant. Here, we use two orthographic learning datasets from baboons and humans to investigate how they implement visual orthographic representations in a learning task of known and novel letter strings. We use two connectionist neural-network models (i.e., CORnet-Z and ResNet-18, CNNs) and two variants of a mechanistic neuro-cognitive model (i.e., the Speechless Reader model, SLR) specific to the reading domain to investigate orthographic learning and infer the underlying neuro-cognitive processes. The connectionist models employ neuronally plausible architectures and are used across the domain of higher vision. The SLR versions are transparent neuro-cognitive models of orthographic decision behavior. Central to the domain-specific SLR implementations are three types of prediction-error representations that we use for computational phenotyping (i.e., pixel-, letter-, and letter-sequence-level prediction-errors). This approach allows us to infer the underlying representations in orthographic decisions. First, we fit the models and simulate the datasets to compare their performance (i.e., all models see the same stimuli as humans and baboons did). Second, after comparing model performance, we evaluate how orthographic decisions have been implemented based on the representations used in the SLR models. We find that the SLR, especially on the trial-wise metrics, outperforms the CNNs in both datasets, with both connectionist models generating behavioral responses without a considerable overlap with individual human or baboon responses. Inspecting the SLR representations, we found that both species implemented the most informative representations that developed from visual to more complex orthographic representations with increased learning. Thus, we show that domain-specific neuro-cognitive mechanistic models are highly valuable in understanding complex behavior and how it is learned across species.