Integrating theoretical neuroscience, decision theory, and probabilistic inference offers a promising route to understanding human cognition, yet concrete methodological bridges between agentic AI models and behavioral data analysis remain formally underdeveloped. We advance this synthesis under the framework of agentic behavioral modeling (ABM), which treats artificial agents as latent, generative hypotheses about cognitive mechanisms and evaluates them by their statistical adequacy in explaining human behavior. After outlining its conceptual foundations, we apply the framework to two minimal laboratory paradigms: a binary perceptual contrast-discrimination task and a symmetric two-armed bandit learning task. We formalize each task-agent-data system as a joint probability model, derive explicit conditional log-likelihoods for behavioral inference, validate different model variants using model and parameter recovery simulations, and evaluate them in light of empirical data. Using these minimal examples, we provide an agent-centric interpretation of the psychometric function, derive optimal policies for both tasks, and show the equivalence between Rescorla-Wagner learning and Bayesian inference in symmetric bandits. More broadly, this work may serve as a conceptual and practical foundation for applying ABM to cognitive behavioral science.
Adaptive behavior requires acting on sensory evidence that only partially reveals relevant environmental states. Partially observable Markov decision processes (POMDPs) provide a unified framework for this problem by representing perceptual uncertainty as belief states and linking these beliefs to decision functions, policies, and learning. We review how POMDP concepts can advance computational cognitive neuroscience across two domains. First, studies of information sampling show how beliefs about latent states and environmental dynamics shape evidence accumulation and perceptual choices, while full POMDP formulations derive optimal policies or infer subjective objectives. Second, studies of reinforcement learning show how perceptual uncertainty modulates credit assignment and the updating of state and action values, often through belief-weighted approximations. We argue that POMDPs are a useful and burgeoning framework for integrating perception, decision-making, and learning, as well as for understanding the neurocognitive foundations of adaptive behavior under uncertainty.
Experiences shape preferences. This is particularly the case when they deviate from our expectations and thus elicit prediction errors. Here we show that prediction errors do not only occur in response to actual events – they also arise endogenously in response to merely imagined events. Specifically, people repeatedly chose between different acquaintances and then imagined interacting with them. Our results show that they acquired a preference for acquaintances with whom they had pictured unexpectedly pleasant events. This learning can best be accounted for by a computational model that calculates prediction errors based on these rewarding experiences. Using functional MRI, we show that the prediction error is mediated via striatal activity. This activity, in turn, seems to update preferences about the individuals by updating their cortical representations. Our findings demonstrate that imaginings can violate our own expectations and thus drive endogenous learning by coopting a neural system that implements reinforcement learning. Experiences guide our preferences, particularly when they violate expectations. Here, the authors show that prediction errors also arise endogenously as a consequence of merely imagined events by coopting mechanisms of reinforcement learning.
Abstract Difficulties in adapting learning to meet the challenges of uncertain and changing environments are widely thought to play a central role in internalizing psychopathology, including anxiety and depression. This view stems from findings linking trait anxiety and transdiagnostic internalizing symptoms to learning impairments in laboratory tasks often used as proxies for real-world behavioral flexibility. These tasks typically require learners to adjust learning rates dynamically in response to uncertainty, for instance, increasing learning from prediction errors in volatile environments. However, prior studies have produced inconsistent and sometimes contradictory findings regarding the nature and extent of learning impairments in populations with internalizing disorders. To address this, we conducted eight experiments (N = 820) using predictive inference and reversal learning tasks, and applied a bi-factor analysis to capture internalizing symptom variance shared across and differentiated between anxiety and depression. While we observed robust evidence for adaptive learning-rate modulation across participants, we found no convincing evidence of a systematic relationship between internalizing symptoms and either learning rates or task performance. These findings challenge prominent claims that learning difficulties are a hallmark feature of internalizing psychopathology and suggest that the relationship between these traits and adaptive behavior under uncertainty may be more subtle than previously thought.
Learning accurate beliefs about the world is computationally demanding but critical for adaptive behavior across the lifespan. Here, we build on an established framework formalizing learning as predictive inference and examine the possibility that age differences in learning emerge from efficient computations that consider available cognitive resources differing across the lifespan. In our resource-rational model, beliefs are updated through a sampling process that stops after reaching a criterion level of accuracy. The sampling process navigates a trade-off between belief accuracy and computational cost, with more samples favoring belief accuracy and fewer samples minimizing costs. When cognitive resources are limited or costly, a maximization of the accuracy-cost ratio requires a more frugal sampling policy, which leads to systematically biased beliefs. Data from two lifespan studies (N = 129 and N = 90) and one study in younger adults (N = 94) show that children and older adults display biases characteristic of a more frugal sampling policy. This is reflected in (a) more frequent perseveration when participants are required to update from previous beliefs and (b) a stronger anchoring bias when updating beliefs from an externally generated value. These results are qualitatively consistent with simulated predictions of our resource-rational model, corroborating the assumption that the identified biases originate from sampling. Our model and results provide a unifying perspective on perseverative and anchoring biases, show that they can jointly emerge from efficient belief-updating computations, and suggest that resource-rational adjustments of sampling computations can explain age-related changes in adaptive learning. (PsycInfo Database Record (c) 2025 APA, all rights reserved).
Learning allows humans and other animals to make predictions about the environment that facilitate adaptive behavior. Casting learning as predictive inference can shed light on normative cognitive mechanisms that improve predictions under uncertainty. Drawing on normative learning models, we illustrate how learning should be adjusted to different sources of uncertainty, including perceptual uncertainty, risk, and uncertainty due to environmental changes. Such models explain many hallmarks of human learning in terms of specific statistical considerations that come into play when updating predictions under uncertainty. However, humans also display systematic learning biases that deviate from normative models, as studied in computational psychiatry. Some biases can be explained as normative inference conditioned on inaccurate prior assumptions about the environment, while others reflect approximations to Bayesian inference aimed at reducing cognitive demands. These biases offer insights into cognitive mechanisms underlying learning and how they might go awry in psychiatric illness. Flexible learning requires humans to adjust their behaviour to uncertainty. While normative learning models explain many adaptive behaviours, systematic biases-arising from inaccurate assumptions or cognitive simplifications-reveal key mechanisms of learning and their potential dysfunctions in psychiatric disorders.
Learning to predict future outcomes is essential for successful decision-making. One importantmechanism governing such learning is the reward prediction error. In many real-worldscenarios, sensory information about stimuli and choice options is ambiguous, leading to uncertaintyabout the environment’s underlying states that guide learning and choice behavior.In such cases, learning from prediction errors should be modulated by the probabilities ofthese hypothetical states, known as the belief state. We hypothesized that prediction errorsmight be weighted by the belief state during learning under perceptual uncertainty, and thatthis modulation is governed by pupil-linked arousal systems. Combining pupillometry andan uncertainty-augmented reward-learning task (N = 47), we found that pupil responses tooutcomes scaled with prediction errors and were down-weighted under higher uncertainty.This suggests that the brain’s arousal systems combine newly arriving perceptual and rewardinformation to dynamically regulate how much to learn in an uncertain world.
Decision neuroscience examines the neurobiological and computational foundations underlying decision-making. Economic decision-making, for example, about which item to purchase, is thought to depend on internal representations of subjective values related to the expected reward or punishment associated with an option. Economic choices typically involve risk due to inherent unpredictability of outcomes. Perceptual decision-making concerns choices based on sensory information under perceptual uncertainty about stimuli and environmental states, such as whether to drive or stop at a traffic light. Decision-making also requires responding to systematic environmental changes, which increases uncertainty substantially. We present common computational models and review behavioral and neurobiological findings of studies on these important concepts in perceptual and economic decision-making, as well as how these two classes of decision-making interact in natural settings.
Learning should be adjusted according to the surprise associated with observed outcomes but calibrated according to statistical context. For example, when occasional changepoints are expected, surprising outcomes should be weighted heavily to speed learning. In contrast, when uninformative outliers are expected to occur occasionally, surprising outcomes should be less influential. Here we dissociate surprising outcomes from the degree to which they demand learning using a predictive inference task and computational modeling. We show that the P300, a stimulus-locked electrophysiological response previously associated with adjustments in learning behavior, does so conditionally on the source of surprise. Larger P300 signals predicted greater learning in a changing context, but less learning in a context where surprise was indicative of a one-off outlier (oddball). Our results suggest that the P300 provides a surprise signal that is interpreted by downstream learning processes differentially according to statistical context in order to appropriately calibrate learning across complex environments.
Perceptual uncertainty and salience both impact decision-making, but how these factors precisely impact trial-and-error reinforcement learning is not well understood. Here, we test the hypotheses that (H1) perceptual uncertainty modulates reward-based learning and that (H2) economic decision-making is driven by the value and the salience of sensory information. For this, we combined computational modeling with a perceptual uncertainty-augmented reward-learning task in a human behavioral experiment (N= 98). In line with our hypotheses, we found that subjects regulated learning behavior in response to the uncertainty with which they could distinguish choice options based on sensory information (belief state), in addition to the errors they made in predicting outcomes. Moreover, subjects considered a combination of expected values and sensory salience for economic decision-making. Taken together, this shows that perceptual and economic decision-making are closely intertwined and share a common basis for behavior in the real world.
Surprise is a key component of many learning experiences, and yet its precise computational role, and how it changes with age, remain debated.One major challenge is that surprise often occurs jointly with other variables, such as uncertainty, outcome magnitude and outcome probability. To assess how humans learn from surprising events, and whether aging affects this process, we studied choices while participants learned from stationary asymmetric outcome distributions, which decouple outcome magnitude, probability, uncertainty, and surprise.A total of 102 participants (51 older, aged 50 -- 73; 51 younger, 19 -- 30 years) chose between three bandits, one of which had a bimodal outcome distribution. Behavioral analyses showed that both age-groups learned the average of the bimodal bandit less well. A trial-by-trial analysis indicated that participants performed choice reversals immediately following large absolute prediction errors, consistent with heightened sensitivity to surprise.This effect was stronger in older adults.Computational models indicated that learning rates in younger as well as older adults were influenced by surprise, rather than uncertainty. Our work bridges between behavioral economics research that has focused on how outcomes with low probability affect choice in older adults, and reinforcement learning work that has investigated age differences in the effects of uncertainty and suggests that older adults overly adapt to surprising events, even when accounting for probability and uncertainty effects.
Memories are stored as ensembles of engram neurons and their successful recall involves the reactivation of these cellular networks. However, significant gaps remain in connecting these cell ensembles with the process of forgetting. Here, we utilized a mouse model of object memory and investigated the conditions in which a memory could be preserved, retrieved, or forgotten. Direct modulation of engram activity via optogenetic stimulation or inhibition either facilitated or prevented the recall of an object memory. In addition, through behavioral and pharmacological interventions, we successfully prevented or accelerated forgetting of an object memory. Finally, we showed that these results can be explained by a computational model in which engrams that are subjectively less relevant for adaptive behavior are more likely to be forgotten. Together, these findings suggest that forgetting may be an adaptive form of engram plasticity which allows engrams to switch from an accessible state to an inaccessible state.
Adaptive decision-making is governed by at least two types of memory processes. On the one hand, learned predictions through integrating multiple experiences, and on the other hand, one-shot episodic memories. These two processes interact, and predictions – particularly prediction errors – influence how episodic memories are encoded. However, studies using computational models disagree on the exact shape of this relationship, with some findings showing an effect of signed prediction errors and others showing an effect of unsigned prediction errors on episodic memory. We argue that the choice-confirmation bias, which reflects stronger learning from choice-confirming compared to disconfirming outcomes, could explain these seemingly diverging results. Our perspective implies that the influence of prediction errors on episodic encoding critically depends on whether people can freely choose between options (i.e., instrumental learning tasks) or not (Pavlovian learning tasks). The choice-confirmation bias on memory encoding might have evolved to prioritize memory representations that optimize reward-guided decision-making. We conclude by discussing open issues and implications for future studies.
Decisions that require taking effort costs into account are ubiquitous in real life. The neural common currency theory hypothesizes that a particular neural network integrates different costs (e.g., risk) and rewards into a common scale to facilitate value comparison. Although there has been a surge of interest in the computational and neural basis of effort-related value integration, it is still under debate if effort-based decision-making relies on a domain-general valuation network as implicated in the neural common currency theory. Therefore, we comprehensively compared effort-based and risky decision-making using a combination of computational modeling, univariate and multivariate fMRI analyses, and data from two independent studies. We found that effort-based decision-making can be best described by a power discounting model that accounts for both the discounting rate and effort sensitivity. At the neural level, multivariate decoding analyses indicated that the neural patterns of the dorsomedial prefrontal cortex (dmPFC) represented subjective value across different decision-making tasks including either effort or risk costs, although univariate signals were more diverse. These findings suggest that multivariate dmPFC patterns play a critical role in computing subjective value in a task-independent manner and thus extend the scope of the neural common currency theory.
Expectations can lead to prediction errors of varying degrees depending on the extent to which the information encountered in the environment conforms with prior knowledge. While there is strong evidence on the computationally specific effects of such prediction errors on learning, relatively less evidence is available regarding their effects on episodic memory. Here, we had participants work on a task in which they learned context/object-category associations of different strengths based on the outcomes of their predictions. We then used a reinforcement learning model to derive subject-specific trial-to-trial estimates of prediction error at encoding and link it to subsequent recognition memory. Results showed that model-derived prediction errors at encoding influenced subsequent memory as a function of the outcome of participants' predictions (correct vs. incorrect). When participants correctly predicted the object category, stronger prediction errors (as a consequence of weak expectations) led to enhanced memory. In contrast, when participants incorrectly predicted the object category, stronger prediction errors (as a consequence of strong expectations) led to impaired memory. These results highlight the important moderating role of choice outcome that may be related to interactions between the hippocampal and striatal dopaminergic systems.
Decisions that require taking prospective effort costs into account are ubiquitous in real life. The common currency theory hypothesizes that a neural network integrates different costs and rewards into a common scale to facilitate value comparison. Although there has been a surge of interest in the computational and neural basis of effort-reward integration, it is still under debate if the common currency theory could be applied to value integration in this context. Here, we comprehensively compared effort-based and risky decision-making using computational modeling, univariate and multivariate fMRI analyses, and data from two independent studies. We found that prospective outcomes were distinctively discounted by effort and risk. Moreover, although univariate fMRI analyses showed diverse results between tasks, multivariate decoding analyses indicated that the neural patterns of the dorsomedial prefrontal cortex (dmPFC) represented subjective value information across effort-based and risky decision-making. These findings suggest that the dmPFC plays a critical role in computing subjective value in a task-independent manner and thus extend the scope of the common currency theory.
Predictive processing accounts propose that our brain constantly tries to match top-down internal representations with bottom-up incoming information from the environment. Predictions can lead to prediction errors of varying degrees depending on the extent to which the information encountered in the environment conforms with prior expectations. Theoretical and computational models assume that prediction errors have beneficial effects on learning and memory. However, while there is strong evidence on the effects of prediction error on learning, relatively less evidence is available regarding its effects on memory. Moreover, most of the studies available so far manipulated prediction error by using monetary rewards, whereas in everyday life learning does not always occur in the presence of explicit rewards. We used a task in which participants leaned context/object-category associations of different strength based on the outcomes of their predictions. After learning these associations, participants were presented with trial-unique objects that could match or violate their predictions. Finally, participants were asked to complete a surprise recognition memory test. We used a reinforcement learning model to derive subject-specific trial-to-trial estimates of prediction error at encoding and link it to subsequent recognition memory. Results showed that model-derived prediction errors at encoding influenced subsequent memory as a function of the outcome of participants’ predictions (correct vs incorrect). When participants correctly predicted the object category, stronger prediction errors (as a consequence of weak expectations) led to enhanced memory. In contrast, when participants incorrectly predicted the object category, stronger prediction errors (as a consequence of strong expectations) led to impaired memory. These results reveal a computationally specific influence of prediction error on memory formation, highlighting the important moderating role of choice outcome that may be related to interactions between the hippocampal and striatal dopaminergic systems.
Influential theories emphasize the importance of predictions in learning: we learn from feedback to the extent that it is surprising, and thus conveys new information. Here, we explore the hypothesis that surprise depends not only on comparing current events to past experience, but also on online evaluation of performance via internal monitoring. Specifically, we propose that people leverage insights from response-based performance monitoring – outcome predictions and confidence – to control learning from feedback. In line with predictions from a Bayesian inference model, we find that people who are better at calibrating their confidence to the precision of their outcome predictions learn more quickly. Further in line with our proposal, EEG signatures of feedback processing are sensitive to the accuracy of, and confidence in, post-response outcome predictions. Taken together, our results suggest that online predictions and confidence serve to calibrate neural error signals to improve the efficiency of learning.
Across the lifespan, humans rely on the ability to learn from new experiences to adapt to uncertain and changing environments. Here we investigated age-related differences in the reliance on default-belief settings during learning in these environments. We collected behavioral data with a predictive-inference task in children, adolescents as well as younger and older adults. Using a Bayesian belief-updating model, we first showed that age-related learning differences might be due to a reduced ability to adjust learning according to dissociable normative factors. The results revealed a reduced consideration of uncertainty in older adults and increased perseveration in both children and older adults. Simulations indicated that interventions to reduce perseveration might lead to more similar performance levels between the age groups. In a follow-up experiment, we found that one such intervention which randomly distorted participants' initial predictions strongly reduced perseveration, but led to increased performance differences. This counter-intuitive effect resulted from an environmental control of learning in children and older adults through random information from the intervention to reduce perseveration. Across the two experiments, our findings show that age-related learning impairments can be explained with insufficient updating from a default belief. In stable environments, this results in perseverative behavior, while in the presence of random environmental information, it leads to environmental control. We formalized the emergence of these belief-updating behaviors with a model that updated beliefs only to an acceptable level of plausibility, suggesting that children and older adults are more quickly satisfied to report beliefs that reflect their default than younger adults. This model not only accounted for our own findings but might also provide a new perspective on a wide variety of previous findings in the developmental and aging literature.