Learning where and when rewards like food and water are available is essential for survival1,2. In the simplest cases where resource availability is stable, animals can learn reward contingencies by integrating outcomes across repeated samples of each possible action. In more natural settings, however, reward availability is governed by structured higher-order rules such as depletion and repletion over time. To adapt flexibly to such changing environments, optimal choices require meta-learning wherein animals learn how to learn from external feedback, ultimately enabling them to infer the underlying reward structure from abstract, generalizable rules rather than relying solely on recent outcomes3,4. The existence of meta-learning in animal behavior is well established3-8, yet the neural circuits and computations that implement it remain poorly understood9-11. Here we investigated meta-learning using a spatial foraging task in which rats acquired a depletion-repletion rule that regulated reward availability, and carried out longitudinal, high-density recordings from the medial prefrontal cortex (mPFC). We show that meta-learning engages specific, systematic changes in mPFC neural dynamics that embed the learned rule and thereby alter how the network learns action values from reward outcomes. These dynamics are based on mixed coding of task structure and value in individual mPFC neurons. At the population level, this coding organizes into low-dimensional dynamical motifs that generalize across task conditions. As meta-learning progresses, these motifs are reshaped to instantiate both rule-guided inference of future states before outcome delivery and rule-based value updating during the outcome period. These results indicate that meta-learning sculpts pre-existing prefrontal dynamics to support the acquisition of new, generalizable reward-learning strategies.
Brain-Computer Interfaces (BCIs) are devices that translate brain activity into commands for control or communication that present many clinical applications. However, the ability to control a BCI remains a learned skill that a non-negligible proportion cannot acquire after several sessions. Identifying the causes of such inter-variability is still an open avenue. We share the Networks for BCI (NETBCI) dataset to help identifying brain networks reorganization underlying BCI training and to develop improved BCI systems. The NETBCI dataset contains magnetoencephalographic (MEG) and electroencephalographic (EEG) recordings obtained from 19 healthy subjects across four sessions performed on four different days. Each session comprises two resting-state recordings of three minutes each (eyes open), and six runs of BCI experiment consisting of performing either a sustained right hand motor imagery or of remaining at rest to control the position of a virtual cursor. NETBCI comprises also anonymized MRI and behavioral scores. We hope that NETBCI can be used for further analysis beyond the scope of the investigation of brain network reorganization during BCI training. We believe that the sample size and the number of modalities should be useful for the scientific community.
Children are adept statistical learners, capable of parsing streams of structured input into meaningful units, but the cognitive processes they engage during learning may differ from those of adults. To date, however, it is unclear how learners of different ages predict upcoming experience when navigating environments with complex structure, as well as how changes in predictive learning mechanisms influence structured knowledge acquisition. To address this question, we tested 106 children, adolescents, and adults, ages 8 - 22 years, on a predictive learning task, in which they experienced sequences of stimuli with a higher-order temporal structure. After an initial learning phase, participants’ explicit knowledge of the relations between stimuli was probed via two additional task measures. We used a recently introduced computational model to characterize participants’ response times during learning, and found that all participants relied on simple, recency-based prediction, anticipating that they would encounter stimuli they recently encountered in the past. With increasing age, however, participants demonstrated greater evidence of additionally relying on a more sophisticated learning mechanism, which captured a predictive representation of the conditional relations between stimuli. Though predictive learning changed with age, we found only weak evidence that these changes related to the acquisition of explicit knowledge of the environment. Our results suggest that the learning mechanisms through which people parse continuous streams of experience change with age, influencing their predictions about upcoming events.
Decisions in humans and other organisms depend, in part, on learning and using models that capture the statistical structure of the world, including the long-run expected outcomes of our actions. One prominent approach to forecasting such long-run outcomes is the successor representation (SR), which predicts future states aggregated over multiple timesteps. Although much behavioral and neural evidence suggests that people and animals use such a representation, it remains unknown how they acquire it. It has frequently been assumed to be learned by temporal difference bootstrapping (SR-TD(0)), but this assumption has largely not been empirically tested or compared to alternatives including eligibility traces ( SR-TD ( λ > 0 ) ). Here we address this gap by leveraging trial-by-trial reaction times in graph sequence learning tasks, which are favorable for studying learning dynamics because the long horizons in these studies differentiate the transient update dynamics of different learning rules. We examined the behavior of SR-TD λ on a probabilistic graph learning task alongside a number of alternatives, and found that behavior was best explained by a hybrid model which learned via SR-TD λ alongside an additional predictive model of recency. The relatively large λ we estimate indicates a predominant role of eligibility trace mechanisms over the bootstrap-based chaining typically assumed. Our results provide insight into how humans learn predictive representations, and demonstrate that people simultaneously learn the SR alongside lower-order predictions.
By building a mental model of how the world works and using it to forecast the outcomes of different actions, a learner can make flexible choices in changing environments. However, while children and adolescents readily acquire structured knowledge about their environments, relative to adults, they tend to demonstrate weaker signatures of leveraging this knowledge to plan actions. One explanation for these developmental differences is that using a mental model to prospectively simulate potential choices and their outcomes is computationally costly, taxing cognitive control and working memory mechanisms that continue to develop into adulthood. Here, we ask whether children might effectively leverage structured knowledge to make flexible choices by relying on two alternative strategies that do not require costly mental simulation at choice time. First, through offline replanning, models can be queried before the time of choice to generate possible scenarios and update the values of potential actions. Second, an abstracted predictive model, known as a Successor Representation, can be built and harnessed to enable simplified computation of long-run reward values of candidate actions, without requiring iterative simulation of multiple time steps. To assess whether children, adolescents, and adults aged 7 - 23 years similarly harness these learning strategies, we ran three experiments. In Experiments 1 and 2, we used a reward revaluation task in which we manipulated the opportunity for offline replanning during rest, and found that children flexibly updated their behavior by leveraging structured knowledge in an adult-like manner. Surprisingly, across age, rest did not mediate flexible replanning, raising the possibility that participants may have behaved adaptively by harnessing predictive representations online. In Experiment 3, we directly tested whether children use predictive representations. Here, we observed early-emerging use of the SR, providing a mechanistic account of how children use structured knowledge to guide choice without detailed model-based simulation.
From sequences of discrete events, humans build mental models of their world. Referred to as graph learning, the process produces a model encoding the graph of event-to-event transition probabilities. Recent evidence suggests that some networks are easier to learn than others, but the neural underpinnings of this effect remain unknown. Here we use fMRI to show that even over short timescales the network structure of a temporal sequence of stimuli determines the fidelity of event representations as well as the dimensionality of the space in which those representations are encoded: when the graph was modular as opposed to lattice-like, BOLD representations in visual areas better predicted trial identity and displayed higher intrinsic dimensionality. Broadly, our study shows that network context influences the strength of learned neural representations, motivating future work in the design, optimization, and adaptation of network contexts for distinct types of learning.
Humans naturally attend to patterns that emerge in our perceptual environments, building mental models that allow future experiences to be processed more effectively and efficiently. Statistical learning research shows that people extract structure from stimuli even when it is not explicitly observable. A growing line of work formalizes this implicit structure as graphs in which perceptual events correspond to nodes and statistical transitions to edges, revealing that behavior is sensitive to underlying graph topology. Yet little is known about how different topologies shape the neural dynamics of learning itself. Here, we used time-resolved network analyses of fMRI data collected during a visuomotor graph-learning task in which stimuli were presented according to random walks on either a modular or lattice graph. Participants responded faster to modular graphs early in learning, though this advantage diminished over runs, replicating prior behavioral findings. Neurally, task performance was characterized by a flexible visual system, relatively stable large-scale community structure, and increased cohesiveness (or within-system recruitment) of the dorsal attention, limbic, default-mode, and subcortical systems. Across runs, integration between visual and ventral-attention regions increased, while coupling between frontoparietal control and visual/dorsal-attention regions decreased—signaling a shift from top-down to bottom-up processing as learning progressed. These results show that the brain’s dynamic functional organization reflects the statistical topology of experience: stronger integration among limbic, default-mode, temporoparietal, and subcortical systems predicted faster responses for modular but not lattice graphs. By linking behavioral sensitivity to graph structure with its time-resolved neural correlates, this work advances understanding of how the brain represents and adapts to complex statistical environments.
How do people model the world's dynamics to guide mental simulation and evaluate choices? One prominent approach, the Successor Representation (SR), takes advantage of temporal abstraction of future states: by aggregating trajectory predictions over multiple timesteps, the brain can avoid the costs of iterative, multi-step mental simulation. Human behavior broadly shows signatures of such temporal abstraction, but finer-grained characterization of individuals' strategies and their dynamic adjustment remains an open question. We developed a task to measure SR usage during dynamic, trial-by-trial learning. Using this approach, we find that participants exhibit a mix of SR and model-based learning strategies that varies across individuals. Further, by dynamically manipulating the task contingencies within-subject to favor or disfavor temporal abstraction, we observe evidence of resource-rational reliance on the SR, which decreases when future states are less predictable. Our work adds to a growing body of research showing that the brain arbitrates between approximate decision strategies. The current study extends these ideas from simple habits into usage of more sophisticated approximate predictive models, and demonstrates that individuals dynamically adapt these in response to the predictability of their environment.
The cognitive ability to go beyond the present to consider alternative possibilities, including potential futures and counterfactual pasts, can support adaptive decision making. Complex and changing real-world environments, however, have many possible alternatives. Whether and how the brain can select among them to represent alternatives that meet current cognitive needs remains unknown. We therefore examined neural representations of alternative spatial locations in the rat hippocampus during navigation in a complex patch foraging environment with changing reward probabilities. We found representations of multiple alternatives along paths ahead and behind the animal, including in distant alternative patches. Critically, these representations were modulated in distinct patterns across successive trials: alternative paths were represented proportionate to their evolving relative value and predicted subsequent decisions, whereas distant alternatives were prevalent during value updating. These results demonstrate that the brain modulates the generation of alternative possibilities in patterns that meet changing cognitive needs for adaptive behavior.
Human experience is built upon sequences of discrete events. From those sequences, humans build impressively accurate models of their world. This process has been referred to as graph learning, a form of structure learning in which the mental model encodes the graph of event-to-event transition probabilities [1], [2], typically in medial temporal cortex [3]–[6]. Recent evidence suggests that some network structures are easier to learn than others [7]–[9], but the neural properties of this effect remain unknown. Here we use fMRI to show that the network structure of a temporal sequence of stimuli influences the fidelity with which those stimuli are represented in the brain. Healthy adult human participants learned a set of stimulus-motor associations following one of two graph structures. The design of our experiment allowed us to separate regional sensitivity to the structural, stimulus, and motor response components of the task. As expected, whereas the motor response could be decoded from neural representations in postcentral gyrus, the shape of the stimulus could be decoded from lateral occipital cortex. The structure of the graph impacted the nature of neural representations: when the graph was modular as opposed to lattice-like, BOLD representations in visual areas better predicted trial identity in a held-out run and displayed higher intrinsic dimensionality. Our results demonstrate that even over relatively short timescales, graph structure determines the fidelity of event representations as well as the dimensionality of the space in which those representations are encoded. More broadly, our study shows that network context influences the strength of learned neural representations, motivating future work in the design, optimization, and adaptation of network contexts for distinct types of learning over different timescales.
Schizophrenia is marked by deficits in facial affect processing associated with abnormalities in GABAergic circuitry, deficits also found in first-degree relatives. Facial affect processing involves a distributed network of brain regions including limbic regions like amygdala and visual processing areas like fusiform cortex. Pharmacological modulation of GABAergic circuitry using benzodiazepines like alprazolam can be useful for studying this facial affect processing network and associated GABAergic abnormalities in schizophrenia. Here, we use pharmacological modulation and computational modeling to study the contribution of GABAergic abnormalities toward emotion processing deficits in schizophrenia. Specifically, we apply principles from network control theory to model persistence energy – the control energy required to maintain brain activation states – during emotion identification and recall tasks, with and without administration of alprazolam, in a sample of first-degree relatives and healthy controls. Here, persistence energy quantifies the magnitude of theoretical external inputs during the task. We find that alprazolam increases persistence energy in relatives but not in controls during threatening face processing, suggesting a compensatory mechanism given the relative absence of behavioral abnormalities in this sample of unaffected relatives. Further, we demonstrate that regions in the fusiform and occipital cortices are important for facilitating state transitions during facial affect processing. Finally, we uncover spatial relationships (i) between regional variation in differential control energy (alprazolam versus placebo) and (ii) both serotonin and dopamine neurotransmitter systems, indicating that alprazolam may exert its effects by altering neuromodulatory systems. Together, these findings provide a new perspective on the distributed emotion processing network and the effect of GABAergic modulation on this network, in addition to identifying an association between schizophrenia risk and abnormal GABAergic effects on persistence energy during threat processing.
Animals frequently make decisions based on expectations of future reward ("values"). Values are updated by ongoing experience: places and choices that result in reward are assigned greater value. Yet, the specific algorithms used by the brain for such credit assignment remain unclear. We monitored accumbens dopamine as rats foraged for rewards in a complex, changing environment. We observed brief dopamine pulses both at reward receipt (scaling with prediction error) and at novel path opportunities. Dopamine also ramped up as rats ran toward reward ports, in proportion to the value at each location. By examining the evolution of these dopamine place-value signals, we found evidence for two distinct update processes: progressive propagation of value along taken paths, as in temporal difference learning, and inference of value throughout the maze, using internal models. Our results demonstrate that within rich, naturalistic environments dopamine conveys place values that are updated via multiple, complementary learning algorithms.
Large-scale interactions among multiple brain regions manifest as bursts of activations called neuronal avalanches, which reconfigure according to the task at hand and, hence, might constitute natural candidates to design brain-computer interfaces (BCIs). To test this hypothesis, we used source-reconstructed magneto/electroencephalography during resting state and a motor imagery task performed within a BCI protocol. To track the probability that an avalanche would spread across any two regions, we built an avalanche transition matrix (ATM) and demonstrated that the edges whose transition probabilities significantly differed between conditions hinged selectively on premotor regions in all subjects. Furthermore, we showed that the topology of the ATMs allows task-decoding above the current gold standard. Hence, our results suggest that neuronal avalanches might capture interpretable differences between tasks that can be used to inform brain-computer interfaces.
Humans are constantly exposed to sequences of events in the environment. Those sequences frequently evince statistical regularities, such as the probabilities with which one event transitions to another. Collectively, inter-event transition probabilities can be modeled as a graph or network. Many real-world networks are organized hierarchically and understanding how these networks are learned by humans is an ongoing aim of current investigations. While much is known about how humans learn basic transition graph topology, whether and to what degree humans can learn hierarchical structures in such graphs remains unknown. Here, we investigate how humans learn hierarchical graphs of the Sierpiński family using computer simulations and behavioral laboratory experiments. We probe the mental estimates of transition probabilities via the surprisal effect: a phenomenon in which humans react more slowly to less expected transitions, such as those between communities or modules in the network. Using mean-field predictions and numerical simulations, we show that surprisal effects are stronger for finer-level than coarser-level hierarchical transitions. Notably, surprisal effects at coarser levels of the hierarchy are difficult to detect for limited learning times or in small samples. Using a serial response experiment with human participants (n=100), we replicate our predictions by detecting a surprisal effect at the finer-level of the hierarchy but not at the coarser-level of the hierarchy. To further explain our findings, we evaluate the presence of a trade-off in learning, whereby humans who learned the finer-level of the hierarchy better tended to learn the coarser-level worse, and vice versa. Taken together, our computational and experimental studies elucidate the processes by which humans learn sequential events in hierarchical contexts. More broadly, our work charts a road map for future investigation of the neural underpinnings and behavioral manifestations of graph learning.