As the statistics of sensory environments often change, neural sensory systems must adapt to maintain useful representations. Efficient coding prescribes that neuronal tuning curves should be optimized to the prior, but whether they can adapt rapidly is unclear. Empirically, tuning curves after repeated stimulus presentations exhibit 'adapter repulsion', whose underlying mechanism remains uncertain, and which contrasts with the 'prior attraction' expected under many efficient-coding models. We propose a gain-adaptive, recurrent sensory network model in which gains optimize an efficient-coding objective balancing accuracy and spiking cost. From the propagation of modulated gains throughout the network emerge quickly adaptive tuning curves. The model accounts for subtle adapter-repulsion effects under peaked priors and predicts fast prior attraction under broader distributions, for which we provide supporting behavioral evidence. Our framework reconciles seemingly contradictory adaptive phenomena, under a unified theoretical and mechanistic model of efficient coding mediated by gain modulation in recurrent circuits.
Abstract Weber’s law is a rare quantitative regularity in psychology, yet its origins remain debated. Here we provide causal evidence that it arises from the more fundamental principle of efficient coding. This principle posits that representational resources are allocated according to stimulus frequencies: distributions skewed toward smaller stimuli thus result in discriminability decreasing with magnitude, as in Weber’s law. Skewing frequencies in the other direction—making large magnitudes more frequent than small ones—enabled us to invert this pattern, and to break Weber’s law. In discrimination tasks with three different sensory modalities, human subjects’ discriminability across stimuli was sensitive to the stimulus distribution, and this adaptation improved task performance. These findings establish efficient coding as a dynamic, organizing principle, explaining when and why Weber’s law holds.
A suite of impressive scientific discoveries have been driven by recent advances in artificial intelligence. These almost all result from training flexible algorithms to solve difficult optimization problems specified in advance by teams of domain scientists and engineers with access to large amounts of data. Although extremely useful, this kind of problem solving only corresponds to one part of science - the "easy problem." The other part of scientific research is coming up with the problem itself - the "hard problem." Solving the hard problem is beyond the capacities of current algorithms for scientific discovery because it requires continual conceptual revision based on poorly defined constraints. We can make progress on understanding how humans solve the hard problem by studying the cognitive science of scientists, and then use the results to design new computational agents that automatically infer and update their scientific paradigms.
Recent advances in artificial neural networks for machine learning, and language modeling in particular, have established a family of recurrent neural network (RNN) architectures that, unlike conventional RNNs with vector-form hidden states, use two-dimensional (2D) matrix-form hidden states. Such 2D-state RNNs, known as Fast Weight Programmers (FWPs), can be interpreted as a neural network whose synaptic weights (called fast weights) dynamically change over time as a function of input observations, and serve as short-term memory storage; corresponding synaptic weight modifications are controlled or programmed by another network (the programmer) whose parameters are trained (e.g., by gradient descent). In this Primer, we review the technical foundations of FWPs, their computational characteristics, and their connections to transformers and state space models. We also discuss connections between FWPs and models of synaptic plasticity in the brain, suggesting a convergence of natural and artificial intelligence.
Modern reinforcement learning (RL) systems have demonstrated remarkable capabilities in complex environments, such as video games. However, they still fall short of achieving human-like sample efficiency and adaptability when learning new domains. Theory-based reinforcement learning (TBRL) is an algorithmic framework specifically designed to address this gap. Modeled on cognitive theories, TBRL leverages structured, causal world models - "theories" - as forward simulators for use in planning, generalization and exploration. Although current TBRL systems provide compelling explanations of how humans learn to play video games, they face several technical limitations: their theory languages are restrictive, and their planning algorithms are not scalable. To address these challenges, we introduce TheoryCoder, an instantiation of TBRL that exploits hierarchical representations of theories and efficient program synthesis methods for more powerful learning and planning. TheoryCoder equips agents with general-purpose abstractions (e.g., "move to"), which are then grounded in a particular environment by learning a low-level transition model (a Python program synthesized from observations by a large language model). A bilevel planning algorithm can exploit this hierarchical structure to solve large domains. We demonstrate that this approach can be successfully applied to diverse and challenging grid-world games, where approaches based on directly synthesizing a policy perform poorly. Ablation studies demonstrate the benefits of using hierarchical abstractions.
Large Language Models (LLMs) have demonstrated impressive real-world utility, exemplifying artificial useful intelligence (AUI). However, their ability to reason adaptively and robustly – the hallmarks of artificial general intelligence (AGI) – remains fragile. While LLMs seemingly succeed in commonsense reasoning, programming, and mathematics, they struggle to generalize algorithmic understanding across novel contexts. Our experiments with algorithmic tasks in esoteric programming languages reveal that LLM's reasoning overfits to the training data and is limited in its transferability. We hypothesize that the core issue underlying such limited transferability is the coupling of reasoning and knowledge in LLMs. To transition from AUI to AGI, we propose disentangling knowledge and reasoning through three key directions: (1) pretaining to reason using RL from scratch as an alternative to the widely used next-token prediction pretraining, (2) using a curriculum of synthetic tasks to ease the learning of a reasoning prior for RL that can then be transferred to natural language tasks, and (3) learning more generalizable reasoning functions using a small context window to reduce exploiting spurious correlations between tokens. Such a reasoning system coupled with a trained retrieval system and a large external memory bank as a knowledge store can overcome several limitations of existing architectures at learning to reason in novel scenarios.
Group stereotypes are difficult to change and drive discriminatory behavior across numerous consequential contexts. Across seven experiments, we test predictions made by a domain-general structure learning model to understand how people decide what "counts" as a group and how those group representations inform our beliefs-here, stereotypes about what a group believes-about constituent members. We have two central hypotheses. First, given low levels of deviance within a collective, participants will perceive a single group among all agents; however, as deviance of one "counter-stereotypical" agent increases, that agent will be subtyped out of the group, yielding two perceived clusters. Second, as deviance increases, confidence in one's beliefs about the group and, correspondingly, a novel group member should decrease; however, once the deviating agent is subtyped out, confidence in one's beliefs about the remaining agents and novel member should increase again. We found consistent evidence for the first prediction: As one agent's deviation from the group increased from 0% to 25%, the deviant was subgrouped. As deviation increased to 50% and more, the deviant was subtyped out of the group. We only observed support for the second prediction in two of the experiments using the confidence measure. However, an exploratory analysis of these experiments revealed a new way to index group stereotype precision-quantifying perceived similarity of all the nondeviating agents to one another. Using this measure of group-based beliefs, we see support for our second hypothesis in a majority of the experiments. (PsycInfo Database Record (c) 2025 APA, all rights reserved).
In complex environments, the space of possible plans is vast. Generating a good plan therefore requires judicious selection of which parts of the plan space to mentally explore. Drawing on past studies of human exploration, we propose that mental exploration might invoke similar mechanisms. In particular, we test the hypothesis that mental exploration during planning is uncertainty-driven, such that people will exhibit a tendency to explore parts of the plan space that have high epistemic uncertainty. We developed a route-planning task, displayed as a binary tree, where participants were instructed to collect as many treats (rewards) as possible by traversing the tree. By separating the planning and execution phases, we encouraged participants to externalize their planning process. We manipulated uncertainty by varying the number of potential future states available from each current state. Across two studies, the data suggest that people preferred to explore options with more successor states after controlling for value differences, supporting the uncertainty-driven planning hypothesis. We also found that uncertainty played a larger role during the planning phase than during the execution phase, consistent with the hypothesis that the uncertainty effect primarily reflects a property of human planning algorithms rather than an intrinsic preference for uncertainty.
Does learning of task-relevant representations stop when behavior stops changing? Motivated by recent work in machine learning and the intuitive observation that human experts continue to learn after mastery, we hypothesize that task-specific representation learning in cortex can continue, even when behavior saturates. In a novel reanalysis of recently published neural data, we find evidence for such learning in posterior piriform cortex of mice following continued training on a task, long after behavior saturates at near-ceiling performance ("overtraining"). We demonstrate that class representations in cortex continue to separate during overtraining, so that examples that were incorrectly classified at the beginning of overtraining can abruptly be correctly classified later on, despite no changes in behavior during that time. We hypothesize this hidden learning takes the form of approximate margin maximization; we validate this and other predictions in the neural data, as well as build and interpret a simple synthetic model that recapitulates these phenomena. We conclude by demonstrating how this model of late-time feature learning implies an explanation for the empirical puzzle of overtraining reversal in animal learning, where task-specific representations are more robust to particular task changes because the learned features can be reused.
Generalization from past experience is an important feature of intelligent systems. When faced with a new task, one efficient computational approach is to evaluate solutions to earlier tasks as candidates for reuse. Consistent with this idea, we found that human participants ( n = 38) learned optimal solutions to a set of training tasks and generalized them to novel test tasks in a reward-selective manner. This behavior was consistent with a computational process based on the successor representation known as successor features and generalized policy improvement (SF&GPI). Neither model-free perseveration or model-based control using a complete model of the environment could explain choice behavior. Decoding from functional magnetic resonance imaging data revealed that solutions from the SF&GPI algorithm were activated on test tasks in visual and prefrontal cortex. This activation had a functional connection to behavior in that stronger activation of SF&GPI solutions in visual areas was associated with increased behavioral reuse. These findings point to a possible neural implementation of an adaptive algorithm for generalization across tasks.
Associative learning depends on contingency, the degree to which a stimulus predicts an outcome. Despite its importance, the neural mechanisms linking contingency to behavior remain elusive. In the present study, we examined the dopamine activity in the ventral striatum-a signal implicated in associative learning-in a Pavlovian contingency degradation task in mice. We show that both anticipatory licking and dopamine responses to a conditioned stimulus decreased when additional rewards were delivered uncued, but remained unchanged if additional rewards were cued. These results conflict with contingency-based accounts using a traditional definition of contingency or a new causal learning model (ANCCR), but can be explained by temporal difference (TD) learning models equipped with an appropriate intertrial interval state representation. Recurrent neural networks trained within a TD framework develop state representations akin to our best 'handcrafted' model. Our findings suggest that the TD error can be a measure that describes both contingency and dopaminergic activity.
Limits on information processing capacity impose limits on task performance. We show that male and female mice achieve performance on a perceptual decision task that is near-optimal given their capacity limits, as measured by policy complexity (the mutual information between states and actions). This behavioral profile could be achieved by reinforcement learning with a penalty on high complexity policies, realized through modulation of dopaminergic learning signals. In support of this hypothesis, we find that policy complexity suppresses midbrain dopamine responses to reward outcomes. Furthermore, neural and behavioral reward sensitivity were positively correlated across sessions. Our results suggest that policy compression shapes basic mechanisms of reinforcement learning in the brain.
The gambler's fallacy is typically defined as the false belief that a random event is less likely to occur if it has occurred recently. Although forms of this fallacy have been documented numerous times, past work either has not actually measured probabilistic predictions but rather point predictions or used sequences that were not independent. To address these problems, we conducted a series of high-powered, preregistered studies in which we asked 750 adult Amazon Mechanical Turk workers from the United States to report probabilistic predictions for truly independent sequences. In contrast to point predictions, which generated a significant gambler's fallacy, probabilistic predictions were not found to lead to a gambler's fallacy. Moreover, the point predictions could not be reconstructed by sampling from the probability judgments. This suggests that the gambler's fallacy originates at the decision stage rather than in probabilistic reasoning, as posited by several leading theories. New theories of the gambler's fallacy may be needed to explain these findings.
Individual contributors to a collaborative task are often rewarded for going above and beyond-salespeople earn commissions, athletes earn performance bonuses, and companies award special parking spots to their employee of the month. How do we decide when to reward collaborators, and are these decisions closely aligned with how responsible they were for the outcome of a collaboration? In Experiments 1a and 1b (N = 360), we tested how participants give bonuses, using stimuli and an experiment design that has previously been used to elicit responsibility judgments (Xiang et al., 2023a). Past work has found that responsibility judgments are driven both by how much effort people actually contributed and how much they could have contributed (Xiang et al., 2023a). In contrast, here we found that participants allocated bonuses based only on how much effort agents actually contributed. In Experiments 2a and 2b (N = 358), we introduced agents who were instructed to exert a particular level of effort; participants still rewarded effort, but their rewards were more sensitive to the precise level of effort exerted when the agents decided how much effort to exert. Together, these findings suggest that people reward collaborators based on their willingness to exert effort, and point to a difference between decisions about how to assign responsibility to collaborators and how to incentivize them. One possible explanation for this difference is that responsibility judgments may reflect causal inference about past collaborations, whereas providing incentives may motivate collaborators to keep exerting effort in the future. Our work sheds light on the cognitive capacities that underlie collaboration.
Midbrain dopamine cells encode differences in predictive and expected value to support learning through reward prediction error. Recent findings have questioned whether reward prediction error can fully account for dopamine function and suggest a more complex role for dopamine in encoding detailed features of the reward environment. In this series of studies, we describe a novel role for dopamine in devaluing sensory features of reward. Mesencephalic dopamine cells activated during a mediated devaluation phase were later chemogenetically reactivated. This retrieval of the devalued reward memory elicited a reduction in the hedonic evaluation of sucrose reward. Through optogenetic and chemogenetic manipulations, we confirm dopamine cells are both sufficient and necessary for mediated devaluation, and retrieval of these memories reflected dopamine release in the nucleus accumbens. Consistent with our computational modeling data, our findings indicate a critical role for dopamine in encoding predictive representations of the sensory features of reinforcement. Overall, we elucidate a novel role for dopamine function in mediated devaluation and illuminate a more elaborate framework through which dopamine encodes reinforcement signals. This study reveals that dopamine is necessary for devaluing sensory memories of reward and thus plays a more complex role in reinforcement learning than traditionally considered.
People use various strategies to bolster the perception of their competence. One strategy is self-handicapping, by which people deliberately impede their performance in order to protect or enhance perceived competence. Despite much prior research, it is unclear why, when, and how self-handicapping occurs. We develop a formal theory that chooses the optimal degree of self-handicapping based on its anticipated performance and signaling effects. We test the theory’s predictions in two experiments (𝑁 = 400), showing that self-handicapping occurs more often when it is unlikely to affect the outcome and when it increases the perceived competence in the eyes of a naive observer. With sophisticated observers (who consider whether a person chooses to self-handicap), self-handicapping is less effective when followed by failure. We show that the theory also explains the findings of several past studies. By offering a systematic explanation of self-handicapping, the theory lays the groundwork for developing effective interventions.
Social learning is a powerful mechanism through which agents learn about the world from others. However, humans don’t always choose to observe others, since social learning can carry time and cognitive resource costs. How do people balance social and non-social learning? In this paper, we propose a rational mentalizing model of the decision to engage in social learning. This model estimates the utility of social learning by reasoning about the other agent’s goal and the informativity of their future actions. It then weighs the utility of social learn- ing against the utility of self-exploration (non-social learning). Using a multi-player treasure hunt game, we show that our model can quantitatively capture human trade-offs between social and non-social learning. Furthermore, our results indicate that these two components allow agents to flexibly apply social learning to achieve their goals more efficiently.
Human intelligence exhibits a remarkable capacity for rapid adaptation and effective problem-solving in novel and unfamiliar contexts. We argue that this profound adaptability is fundamentally linked to the efficient construction and refinement of internal representations of the environment, commonly referred to as world models, and we refer to this adaptation mechanism as world model induction. However, current understanding and evaluation of world models in artificial intelligence (AI) remains narrow, often focusing on static representations learned from training on massive corpora of data, instead of the efficiency and efficacy in learning these representations through interaction and exploration within a novel environment. In this Perspective, we provide a view of world model induction drawing on decades of research in cognitive science on how humans learn and adapt so efficiently; we then call for a new evaluation framework for assessing adaptive world models in AI. Concretely, we propose a new benchmarking paradigm based on suites of carefully designed games with genuine, deep and continually refreshing novelty in the underlying game structures – we refer to this class of games as novel games. We detail key desiderata for constructing these games and propose appropriate metrics to explicitly challenge and evaluate the agent's ability for rapid world model induction. We hope that this new evaluation framework will inspire future evaluation efforts on world models in AI and provide a crucial step towards developing AI systems capable of human-like rapid adaptation and robust generalization – a critical component of artificial general intelligence.
Large Language Models (LLMs) excel at in-context learning, the ability to use information provided as context to improve prediction of future tokens. Induction heads have been argued to play a crucial role for in-context learning in Transformer Language Models. These attention heads make a token attend to successors of past occurrences of the same token in the input. This basic mechanism supports LLMs' ability to copy and predict repeating patterns. However, it is unclear if this same mechanism can support in-context learning of more complex repetitive patterns with hierarchical structure. Natural language is teeming with such cases: The article "the" in English usually prefaces multiple nouns in a text. When predicting which token succeeds a particular instance of "the", we need to integrate further contextual cues from the text to predict the correct noun. If induction heads naively attend to all past instances of successor tokens of "the" in a context-independent manner, they cannot support this level of contextual information integration. In this study, we design a synthetic in-context learning task, where tokens are repeated with hierarchical dependencies. Here, attending uniformly to all successor tokens is not sufficient to accurately predict future tokens. Evaluating a range of LLMs on these token sequences and natural language analogues, we find adaptive induction heads that support prediction by learning what to attend to in-context. Next, we investigate how induction heads themselves learn in-context. We find evidence that learning is supported by attention heads that uncover a set of latent contexts, determining the different token transition relationships. Overall, we not only show that LLMs have induction heads that learn, but offer a complete mechanistic account of how LLMs learn to predict higher-order repetitive patterns in-context.
Making context-dependent decisions incurs cognitive costs. Cognitive control studies have investigated the nature of such costs from both computational and neural perspectives. In this paper, we offer an information-theoretic account of the costs associated with context-dependent decisions. According to this account, the brain’s limited capacity to store context-dependent policies necessitates “compression” of policies into internal representations with an upper bound on codelength, quantified by an information-theoretic measure (policy complexity). These representations are decoded into actions by sequentially inspecting each bit, such that longer codes take more time to decode. When a response deadline is imposed, the account predicts that policy complexity should increase with the deadline. Higher policy complexity is associated with several behavioral signatures: (i) higher accuracy; (ii) lower variability; and (iii) lower perseveration. Analyzing electroencephalograpy data from a rule-based action selection task, we found evidence supporting all of these predictions. We further hypothesized that complex policies require higher neural dimensionality (which constrains the code space). Consistent with this hypothesis, we found that policy complexity correlates with a measure of neural dimensionality in a rule-based decision task. This finding brings us a step closer to understanding the neural implementation of policy compression and its implications for cognitive control.