1 Midbrain dopamine neurons are thought to implement a temporal difference (TD) reward prediction error (RPE) that updates cached values stored in striatum. This has been challenged by evidence that dopamine “ramps up” to predictable rewards during goal-directed behaviour. Here, we propose that dopamine ramps are RPEs generated by a dual-process learning system in which values inferred using a world model train cached values via the RPE. Ramps arise because efficient training of cached values requires that inferred values contribute to the update target but not the prediction component of the RPE. The model reproduces key dopamine ramp phenomena, including learning dynamics on fast and slow timescales, global updates following changes in reward expectation, transient responses during unexpected state transitions, and sensitivity to state uncertainty manipulations. We therefore argue that dopamine ramps are a signature of interactions between inferred and cached values that revise the traditional dichotomy between model-based and model-free learning.
Action selection using predictive models of the environment plays a fundamental role in human and animal behaviour, yet is poorly understood at circuit and algorithmic levels. Spatial navigation is an attractive domain for characterising how world models guide action selection. However spatial behaviours are shaped by multiple control systems including habits, vector-navigation using a Euclidean model of spatial relationships, and route planning using models of environment structure. Understanding how world models support navigation requires assays that dissociate control systems and decorrelate behavioural variables, while generating large datasets that allow precise quantification of brain-behaviour relationships. Here we developed and computationally optimised a behavioural assay to quantify flexible navigation using knowledge of environment structure. Mice navigated to visually cued goals in complex mazes, with randomised start and goal locations on each trial, generating thousands of non-repetitive goal-directed navigation trajectories. They navigated efficiently - strongly favouring options on the shortest path-to-goal, and learnt rapidly - demonstrating knowledge of maze structure from their first sessions in new environments. We anticipate the assay will be useful for characterising how world models support flexible behaviour.
Planning is critical for adaptive behaviour in a changing world, because it lets us anticipate the future and adjust our actions accordingly. While prefrontal cortex is crucial for this process, it remains unknown how planning is implemented in neural circuits. Prefrontal representations were recently discovered in simpler sequence memory tasks, where different populations of neurons represent different future time points. We demonstrate that combining such representations with the ubiquitous principle of neural attractor dynamics allows circuits to solve much richer problems including planning. This is achieved by embedding the environment structure directly in synaptic connections to implement an attractor network that infers desirable futures. The resulting 'spacetime attractor' excels at planning in challenging tasks known to depend on prefrontal cortex. Recurrent neural networks trained by gradient descent on such tasks learn a solution that precisely recapitulates the spacetime attractor – in representation, in dynamics, and in connectivity. Analyses of networks trained across different environment structures reveal a generalisation mechanism that rapidly reconfigures the world model used for planning, without the need for synaptic plasticity. The spacetime attractor is a testable mechanistic theory of planning. If true, it would provide a path towards detailed mechanistic understanding of how prefrontal cortex structures adaptive behaviour.
Abstract Reinforcement learning theory formulates distinct decision-making strategies, including reactive model-free and deliberative model-based strategies. This study investigates how mice adjust their reinforcement learning strategies while learning decision-making in dynamic environments. Unlike previous studies that focused on behaviors after extensive training periods, we analyzed changes in learning strategies in the course of training of a two-step decision-making task with probabilistic state transition and fluctuating reward probabilities. Our statistical behavioral analysis showed that the stay-probability following common and rare transitions diverged with training, a signature of strategies that utilize knowledge of task structure. We fit various reinforcement learning strategies to behavioral data and found that structure-informed strategies became increasingly dominant in their behaviors during training. Whereas previous studies emphasized transition from goal-directed to habitual strategies after extensive training, which were often associated with model-based and model-free strategies, respectively, our results newly demonstrate a shift from model-free to structure-informed strategies in early training in mice. Author summary Reinforcement learning theory allows us to examine how we make decisions and what approaches we use to optimize rewards. Most previous research, however, has examined animal behavior only after extensive training. Here we analyzed how mice adjust their reinforcement learning strategies as they are trained in a two-step decision-making task. Initially, mice relied on reactive model-free strategies, but as training progressed, their behavior began to incorporate knowledge of task structure. While previous studies suggested transition from model-based to model-free strategies with extensive training, our study revealed the opposite in the early stage of training.
Although hippocampal place cells replay nonlocal trajectories, the computational function of these events remains controversial. One hypothesis, formalized in a prominent reinforcement learning account, holds that replay plans routes to current goals. However, recent puzzling data appear to contradict this perspective by showing that replayed destinations lag current goals. These results may support an alternative hypothesis that replay updates route information to build a "cognitive map." Yet no similar theory exists to formalize this view, it is unclear how such a map is represented or what role replay plays in computing it. We address these gaps by introducing a theory of replay that learns a map of routes to candidate goals, before reward is available or when its location may change. Replay is then focused on current goals (as with planning) and/or potential future goals (like a map), depending on the animal's expectations about future goal switching. Our work thus generalizes the planning account to capture a general map-building function for replay, reconciling it with data, and revealing an unexpected relationship between the seemingly distinct hypotheses. The theory offers a unifying explanation why data from tasks with different goal dynamics have seemingly supported different hypotheses for the function of replay, and suggests new predictions for experiments testing these effects.
Despite many empirical results about hippocampal replay, its computational function remains controversial. The "value" hypothesis contends that replay plans routes to current goals, while the "map" hypothesis holds that replay builds an abstract environmental representation, distinct from immediate goals. Data appear to support either view, though the planning hypothesis is particularly challenged by recent observations of replay lagging, rather than leading, animals learning to reach current goals. However, differentiating these ideas is difficult due to a lack of formal specificity, especially about the map hypothesis. We address these gaps by extending a prominent theory of planning to include routes to future as well as current goals: effectively a map. Whether replay prefers current goals, like planning, or others, like maps, then depends on their estimated likelihood of future relevance. This account reconciles both views with one another and with much data, revealing a deep relationship between the seemingly distinct hypotheses.
The claustrum is thought to be one of the most highly interconnected forebrain structures, but its organizing principles have yet to be fully explored at the level of single neurons. Here, we investigated the identity, connectivity, and activity of identified claustrum neurons in Mus musculus to understand how the structure’s unique convergence of input and divergence of output support binding information streams. We found that neurons in the claustrum communicate with each other across efferent projection-defined modules which were differentially innervated by sensory and frontal cortical areas. Individual claustrum neurons were responsive to inputs from more than one cortical region in a cell-type and projection-specific manner, particularly between areas of frontal cortex. In vivo imaging of claustrum axons revealed responses to both unimodal and multimodal sensory stimuli. Finally, chronic claustrum silencing specifically reduced animals’ sensitivity to multimodal stimuli. These findings support the view that the claustrum is a fundamentally integrative structure, consolidating information from around the cortex and redistributing it following local computations.
ABSTRACT To flexibly adapt to new situations, our brains must understand the regularities in the world, but also in our own patterns of behaviour. A wealth of findings is beginning to reveal the algorithms we use to map the outside world 1–6 . In contrast, the biological algorithms that map the complex structured behaviours we compose to reach our goals remain enigmatic. Here we reveal a neuronal implementation of an algorithm for mapping abstract behavioural structure and transferring it to new scenarios. We trained mice on many tasks which shared a common structure organising a sequence of goals, but differed in the specific goal locations. Animals discovered the underlying task structure, enabling zero-shot inferences on the first trial of new tasks. The activity of most neurons in the medial Frontal cortex tiled progress-to-goal, akin to how place cells map physical space. These “goal-progress cells” generalised, stretching and compressing their tiling to accommodate different goal distances. In contrast, progress along the overall sequence of goals was not encoded explicitly. Instead a subset of goal-progress cells was further tuned such that individual neurons fired with a fixed task-lag from a particular behavioural step. Together these cells implemented an algorithm that instantaneously encoded the entire sequence of future behavioural steps, and whose dynamics automatically retrieved the appropriate action at each step. These dynamics mirrored the abstract task structure both on-task and during offline sleep. Our findings suggest that goal-progress cells in the medial frontal cortex may be elemental building blocks of schemata that can be sculpted to represent complex behavioural structures.
The anterior cingulate cortex (ACC) causally influences cognitive control of goal-directed behaviour. However, it is unclear whether ACC directly encodes cognitive variables like attention or impulsivity, or implements goal-directed action selection mechanisms that are modulated by them. We recorded ACC activity with miniature endoscopic microscopes in mice performing the 5-choice-serial-reaction time task, and applied decoding and encoding analyses. ACC pyramidal cells represented specific actions before and during the behavioural response, whereas the response type (e.g. correct/incorrect/premature) – indicating the state of attentional and impulse control – could only be decoded during and after the response with high reliability. Devaluation and extinction experiments further revealed that action encoding depended on reward expectation. Our findings support a role for ACC in goal-directed action selection and monitoring, that is modulated by cognitive state, rather than in tracking levels of attention or impulsivity directly in individual trials. ### Competing Interest Statement The authors have declared no competing interest.
Dopamine is implicated in adaptive behavior through reward prediction error (RPE) signals that update value estimates. There is also accumulating evidence that animals in structured environments can use inference processes to facilitate behavioral flexibility. However, it is unclear how these two accounts of reward-guided decision-making should be integrated. Using a two-step task for mice, we show that dopamine reports RPEs using value information inferred from task structure knowledge, alongside information about reward rate and movement. Nonetheless, although rewards strongly influenced choices and dopamine activity, neither activating nor inhibiting dopamine neurons at trial outcome affected future choice. These data were recapitulated by a neural network model where cortex learned to track hidden task states by predicting observations, while basal ganglia learned values and actions via RPEs. This shows that the influence of rewards on choices can stem from dopamine-independent information they convey about the world’s state, not the dopaminergic RPEs they produce.
The anterior cingulate cortex (ACC) has been implicated in attention deficit hyperactivity disorder (ADHD). More specifically, an appropriate balance of excitatory and inhibitory activity in the ACC may be critical for the control of impulsivity, hyperactivity, and sustained attention which are centrally affected in ADHD. Hence, pharmacological augmentation of parvalbumin- (PV) or somatostatin-positive (Sst) inhibitory ACC interneurons could be a potential treatment strategy. We, therefore, tested whether stimulation of Gq-protein-coupled receptors (GqPCRs) in these interneurons could improve attention or impulsivity assessed with the 5-choice-serial reaction-time task in male mice. When challenging impulse control behaviourally or pharmacologically, activation of the chemogenetic GqPCR hM3Dq in ACC PV-cells caused a selective decrease of active erroneous-i.e. incorrect and premature-responses, indicating improved attentional and impulse control. When challenging attention, in contrast, omissions were increased, albeit without extension of reward latencies or decreases of attentional accuracy. These effects largely resembled those of the ADHD medication atomoxetine. Additionally, they were mostly independent of each other within individual animals. GqPCR activation in ACC PV-cells also reduced hyperactivity. In contrast, if hM3Dq was activated in Sst-interneurons, no improvement of impulse control was observed, and a reduction of incorrect responses was only induced at high agonist levels and accompanied by reduced motivational drive. These results suggest that the activation of GqPCRs expressed specifically in PV-cells of the ACC may be a viable strategy to improve certain aspects of sustained attention, impulsivity and hyperactivity in ADHD.
Fiber photometry is a key technique for characterizing brain-behavior relationships in vivo. Initially, it was primarily used to report calcium dynamics as a proxy for neural activity via genetically encoded indicators. This generated new insights into brain functions including movement, memory, and motivation at the level of defined circuits and cell types. Recently, the opportunity for discovery with fiber photometry has exploded with the development of an extensive range of fluorescent sensors for biomolecules including neuromodulators and peptides that were previously inaccessible in vivo. This critical advance, combined with the new availability of affordable “plug-and-play” recording systems, has made monitoring molecules with high spatiotemporal precision during behavior highly accessible. However, while opening exciting new avenues for research, the rapid expansion in fiber photometry applications has occurred without coordination or consensus on best practices. Here, we provide a comprehensive guide to help end-users execute, analyze, and suitably interpret fiber photometry studies.
Brains are composed of anatomically and functionally distinct regions performing specialized tasks, but regions do not operate in isolation. Orchestration of complex behaviors requires communication between brain regions, but how neural dynamics are organized to facilitate reliable transmission is not well understood. Here we studied this process directly by generating neural activity that propagates between brain regions and drives behavior, assessing how neural populations in sensory cortex cooperate to transmit information. We achieved this by imaging two densely interconnected regions—the primary and secondary somatosensory cortex (S1 and S2)—in mice while performing two-photon photostimulation of S1 neurons and assigning behavioral salience to the photostimulation. We found that the probability of perception is determined not only by the strength of the photostimulation but also by the variability of S1 neural activity. Therefore, maximizing the signal-to-noise ratio of the stimulus representation in cortex relative to the noise or variability is critical to facilitate activity propagation and perception.
Explicit information obtained through instruction profoundly shapes human choice behaviour. However, this has been studied in computationally simple tasks, and it is unknown how model-based and model-free systems, respectively generating goal-directed and habitual actions, are affected by the absence or presence of instructions. We assessed behaviour in a variant of a computationally more complex decision-making task, before and after providing information about task structure, both in healthy volunteers and in individuals suffering from obsessive-compulsive or other disorders. Initial behaviour was model-free, with rewards directly reinforcing preceding actions. Model-based control, employing predictions of states resulting from each action, emerged with experience in a minority of participants, and less in those with obsessive-compulsive disorder. Providing task structure information strongly increased model-based control, similarly across all groups. Thus, in humans, explicit task structural knowledge is a primary determinant of model-based reinforcement learning and is most readily acquired from instruction rather than experience. Healthy volunteers and patients with obsessive-compulsive disorder learning a task from experience alone tend to repeat actions that lead to rewards. They are poor at learning predictive models, but their use of these models is strongly increased when explicit information is provided.
Laboratory behavioural tasks are an essential research tool. As questions asked of behaviour and brain activity become more sophisticated, the ability to specify and run richly structured tasks becomes more important. An increasing focus on reproducibility also necessitates accurate communication of task logic to other researchers. To these ends, we developed pyControl, a system of open-source hardware and software for controlling behavioural experiments comprising a simple yet flexible Python-based syntax for specifying tasks as extended state machines, hardware modules for building behavioural setups, and a graphical user interface designed for efficiently running high-throughput experiments on many setups in parallel, all with extensive online documentation. These tools make it quicker, easier, and cheaper to implement rich behavioural tasks at scale. As important, pyControl facilitates communication and reproducibility of behavioural experiments through a highly readable task definition syntax and self-documenting features. Here, we outline the system’s design and rationale, present validation experiments characterising system performance, and demonstrate example applications in freely moving and head-fixed mouse behaviour.
Humans and other animals effortlessly generalize prior knowledge to solve novel problems, by abstracting common structure and mapping it onto new sensorimotor specifics. To investigate how the brain achieves this, in this study, we trained mice on a series of reversal learning problems that shared the same structure but had different physical implementations. Performance improved across problems, indicating transfer of knowledge. Neurons in medial prefrontal cortex (mPFC) maintained similar representations across problems despite their different sensorimotor correlates, whereas hippocampal (dCA1) representations were more strongly influenced by the specifics of each problem. This was true for both representations of the events that comprised each trial and those that integrated choices and outcomes over multiple trials to guide an animal's decisions. These data suggest that prefrontal cortex and hippocampus play complementary roles in generalization of knowledge: PFC abstracts the common structure among related problems, and hippocampus maps this structure onto the specifics of the current situation.
Dopamine is thought to carry reward prediction errors (RPEs), which update values and hence modify future behaviour. However, updating values is not always the most efficient way of adapting to change. If previously encountered situations will be revisited in future, inferring that the state of the world has changed allows prior experience to be reused when situations are reencountered. To probe dopamine’s involvement in such inference-based behavioural flexibility, we measured and manipulated dopamine while mice solved a sequential decision task using state inference. Dopamine was strongly influenced by the value of states and actions, consistent with RPE signalling, using value information that respected task structure. However, though dopamine responded strongly to rewards, stimulating dopamine at the time of trial outcome had no effect on subsequent choice. Therefore, when inference guides choice, rewards have a dopamine-independent influence on policy through the information they carry about the world’s state.
Rewards are thought to influence future choices through dopaminergic reward prediction errors (RPEs) updating stored value estimates. However, accumulating evidence suggests that inference about hidden states of the environment may underlie much adaptive behaviour, and it is unclear how these two accounts of reward-guided decision-making should be integrated. Using a two-step task for mice, we show that dopamine reports RPEs using value information inferred from task structure knowledge, alongside information about recent reward rate and movement. Nonetheless, although rewards strongly influenced choices and dopamine, neither activating nor inhibiting dopamine neurons at trial outcome affected future choice. These data were recapitulated by a neural network model in which frontal cortex learned to track hidden task states by predicting observations, while basal ganglia learned corresponding values and actions via dopaminergic RPEs. Together, this two-process account reconciles how dopamine-independent state inference and dopamine-mediated reinforcement learning interact on different timescales to determine reward-guided choices.