A fundamental question in cognitive science is how information from internal memory is combined with external sensory input when making decisions. We hypothesized that previously learned and currently perceived information trade-off against each other, such that information from one source reduces the gathering and usage of information from the other. To test this hypothesis, we designed a two-armed bandit task where each arm is composed of both learned and perceived elements. We monitored participants' gathering of perceptual information using eye tracking. Participants' choices and gaze deployment showed a trade-off between the impact of learned and perceived information. The more a participant utilized internally stored learned information, the less they gathered perceptual information, and vice versa. To understand the factors underlying the trade-off, we developed a computational model of participants' information gathering. This showed that the trade-off results from the faster gathering of learned information, which makes it less valuable to invest effort in gathering additional perceptual information. Preliminary findings also suggested that an individual's tendency to primarily rely on one source of information is a stable individual trait. These findings contribute to the understanding of how humans use learning and perception in forming decisions.
Planning is an expensive computational operation that easily exhausts our cognitive resources. For this reason, it is critical to learn strategies to initiate planning wisely. To understand how humans adaptively initiate planning, we developed a task where humans can learn to delay planning at certain timepoints when all actions are equally likely to reach an instructed goal. Across two pre-registered studies, we show that humans adaptively delay planning, and improve in their ability to do so with experience. To explain this behavior, we formalize a model of the underlying meta-control computations. Our model demonstrates how strategies to delay planning are not learned from experienced outcomes, but rather, from searching a cognitive map of the task to determine at which points it is more valuable to delay control. We thus establish that humans can adaptively delay planning, and learn to do so by means of a cognitive map.
Abstract The ability to determine how much the environment can be controlled through our actions has long been viewed as fundamental to adaptive behavior. While traditional accounts treat controllability as a fixed property of the environment, we argue that real-world controllability often depends on the effort, time and money we are able and willing to invest. In such cases, controllability can be said to be elastic to invested resources. Here we propose that inferring this elasticity is essential for efficient resource allocation, and thus, elasticity misestimations result in maladaptive behavior. To test this hypothesis, we developed a novel treasure hunt game where participants encountered environments with varying degrees of controllability and elasticity. Across two pre-registered studies (N=514), we first demonstrate that people infer elasticity and adapt their resource allocation accordingly. We then present a computational model that explains how people make this inference, and identify individual elasticity biases that lead to suboptimal resource allocation. Finally, we show that overestimation of elasticity is associated with elevated psychopathology involving an impaired sense of control. These findings establish the elasticity of control as a distinct cognitive construct guiding adaptive behavior, and a computational marker for control-related maladaptive behavior.
Recent studies suggest that the human brain is equipped to learn not only expected outcomes, but entire distributions of possible outcomes. However, the role of this distributional learning in shaping decision-making remains unclear. To investigate this question, we designed two tasks where participants experienced different outcome distributions, estimated their properties, and reported their preferences. In a simplified observation task, participants were able to learn and report the outcome distributions they experienced. Notably, their risk preferences in this task resembled classic patterns from behavioral economics—patterns typically observed when information is provided by description rather than learned through experience. By contrast, in a more ecological choice task, distributional learning was constrained, and these preference patterns were absent. This led us to suggest that distributional learning may play a causal role in preference formation. Computational modeling supported this interpretation: Preferences were best explained by a model in which distributional learning enables the application of a utility function across possible outcomes. These findings suggest distributional learning may have a critical influence on preference formation, offering new insight into the computational foundations of human decision-making.
Planning is an expensive computational operation that easily exhausts our cognitive resources. For this reason, it is critical to learn strategies to initiate planning wisely. Although recent advances have shown where and how planning should be engaged to aid adaptive behavior, little is known about when planning should be engaged. Here, we investigate how individuals choose to engage or delay planning. For this purpose, we developed a task where, at certain decision points in a multistep decision problem, it is optimal to delay planning and relinquish control because all actions are equally likely to reach an instructed goal. Across two studies, we show that humans can optimally delay planning and improve in their ability to do so with experience. To explain this behavior, we formalize a model of the underlying meta-control computations. In doing so, we demonstrate that meta-policies to engage or delay planning are not simply learned from experienced outcomes, rather, they are constructed by searching a cognitive map of the task to determine at which points it is more valuable to delay control. We thus establish that humans can optimally delay planning, and learn to do so by means of a cognitive map.
Emotion dysregulation, and specifically emotional instability, characterizes adults with ADHD. This study utilized ecological momentary assessment (EMA) to track emotional states and examine patterns of emotional instability within individuals over different time scales. Specifically, it focused on two aspects: overall emotional variability over time, and emotional lability, reflected in emotional states fluctuations within and across days. We further examined the interaction of these emotional instability factors with the subjective experience of emotion regulation difficulties. Young adults with (n = 57) and without (HC; n = 54) ADHD diagnosis completed a self-report questionnaire for emotion regulation difficulties, followed by a 5-day EMA protocol of 5 emotion reports/day. Individuals with ADHD displayed significantly higher intra-individual emotional variability, but no group differences were found for emotional lability, both between and across days. This higher emotional variability was linked to self-reported emotion regulation difficulties in the ADHD group. Finally, using cluster analysis, we found a higher probability of individuals with ADHD being included in a cluster characterized by elevated emotional variability and emotion regulation difficulties. This study demonstrates that young adults with ADHD may experience a broader range of emotions in their daily lives, which may be related to the way they evaluate their challenges in emotion regulation. The findings highlight the need to address emotion dysregulation difficulties in clinical practice, as understanding these emotional dynamics could enhance personalized therapeutic strategies for ADHD, and help design interventions tailored to the breadth and intensity of emotional experiences in ADHD.
The tendency to embrace or avoid risk varies across and within individuals, with significant consequences for economic behavior and mental health. Such variations can partially be explained by differences in the relative weights given to potential gains and losses. Applying this insight to real-life decisions, however, is complicated because such decisions are often based on prior learning experiences. Here, we ask which cognitive process-decision-making or learning-determines the weighting of gains or losses? Over 28 days, 100 participants engaged in a longitudinal decision task wherein choices were based on prior learning. Computational modeling of participants' choices revealed that changes in risk-taking are primarily explained by changes in how learning, not decisions, weight gains and losses. Moreover, inferred changes in learning manifested in participants' neural and physiological learning signals in response to outcomes. We conclude that in experience-based decisions, learning plays a primary role in governing risk-taking behavior.
Emotions consistently shape learning and decision-making, yet their study remains challenging because they are internal states that cannot be directly observed. Recent theory offers a way to overcome this challenge by mapping two classes of emotions onto distinct reinforcement-learning computations, with environmental controllability determining which computation dominates. In controllable settings, emotions guide actions, motivating greater investment of effort and other resources following disappointing outcomes to improve performance. In uncontrollable settings, emotions track reward availability, with disappointing outcomes suppressing reward-seeking behavior and good outcomes amplifying it. We tested this model in a treasure-hunt task (N=509) and found that controllability modulated emotional responses to prediction errors, which in turn determined changes in resource investment. Applying the framework to professional tennis matches (N=6,715) revealed parallel effects: performance changes reflected the same interaction between prediction errors and controllability. Thus, in the laboratory and the real world, controllability arbitrates between distinct emotional responses that shape adaptive and maladaptive behavior.
Recent landmark studies show that the brain is equipped to learn not just average expected outcomes, but entire distributions of expected outcomes. Yet the role of such distributional learning in shaping human decision-making remains to be determined. To study this question, we designed two tasks where participants experienced different outcome distributions, provided their estimates of each, and reported their preferences among them. In one task, which facilitated distributional learning, participants' preferences significantly diverged from their own estimates, consistent with predictions of Prospect Theory. Conversely, in a task that hindered distributional learning, the divergence of preferences from estimates was eliminated. Computational modelling showed how distributional learning may be responsible for disassociating preferences from estimations by enabling the application of a utility function to different potential outcomes. Our findings offer a new understanding of when and how preferences deviate from normative decision-making, a fundamental question in the study of human rationality.
Social norms shape a vast range of human behaviors, from everyday interactions to major life choices. Yet, existing theories of norm emergence typically focus either on why certain norms arise (substantive properties) or on how they spread and persist (dynamical properties), often making conflicting assumptions. Here, we propose a unified account in which norms prescribing how one ought to act emerge naturally from the fundamental algorithms that guide learning-whether in social or nonsocial settings. Our account builds on recent advances in decision making and emotion research that have highlighted "actor-critic" models as a core mechanism of learning from feedback. We extend this mechanism to social settings by assuming that it is not only we who critique our actions; others critique our actions as well. By simulating this interindividual form of learning, we show that it uniquely produces group behavior that exhibits both substantive and dynamical properties of real-world social norms, including prosociality, ingroup bias, stickiness, S-shaped curves, and local conformity/global diversity. Our framework thus offers a uniquely parsimonious way to bridge the gap between individual learning and group behavior.
The formation of predictions is essential to our ability to build models of the world and use them for intelligent decision-making. Here we challenge the dominant assumption that humans form only forward predictions, which specify what future events are likely to follow a given present event. We demonstrate that in some environments, it is more efficient to use backward prediction, which specifies what present events are likely to precede a given future event. This is particularly the case in diverging environments, where possible future events outnumber possible present events. Correspondingly, in six preregistered experiments (n = 1,299) involving both simple decision-making and more challenging planning tasks, we find that humans engage in backward prediction in divergent environments and use forward prediction in convergent environments. We thus establish that humans adaptively deploy forward and backward prediction in the service of efficient decision-making. Models of decision-making assume we predict forward from an action to its potential outcomes. In six studies, Sharp and Eldar show that humans also predict backward from a desired outcome to its preceding actions, particularly in divergent environments.
Previous studies have shown that fixations on familiar stimuli tend to be longer than on unfamiliar stimuli, theorized to be a result of retrieval of information from memory. We hypothesize that extended fixations are due to a lesser need to explore an already familiar stimulus. Participant's gaze was tracked as they tried to encode or retrieve a familiar face displayed either alone or alongside other unfamiliar faces. Regardless of the memory task (encoding\retrieval), longer fixation durations were observed when a single familiar face was presented alone, and not when presented among unfamiliar ones. Thus, fixations were not prolonged when it was possible to explore other, unfamiliar stimuli. We conclude that prolonged fixations on familiar stimuli reflect a lesser need to explore an already familiar percept. The results underscore how memory representations influence active sensing, yielding fresh insights into efficient deployment of attention resources. We conclude that fixation durations could be used in applied memory detection tests, preferably together with other measures and when the familiar stimulus is presented alone.