Abstract Studying how the brain represents value spans distinct methods and training histories, from neuroimaging in task-naïve humans to single-neuron recordings in extensively trained non-human primates. Similar findings across fields have encouraged the untested assumption that rapidly emerging and overtrained value representations are equivalent. Here we recorded single-neuron activity in anterior cingulate cortex (ACC) and orbitofrontal cortex (OFC) as macaques learned novel cue values and chose between novel and overtrained cues. Value responses emerged within 4-7 cue presentations, matching behavioural adaptation. Yet ACC and OFC used distinct codes. ACC encoded value in a common format that generalised across training history. OFC coding was more context-dependent, with distinct subpopulations recruited by choice experience. Choice novelty was also encoded independently of value before chosen value signals emerged. During learning, ACC responses shifted from overtrained secondary reinforcers to newly predictive cues. These findings establish the necessary (rapid acquisition) and sufficient (generalised format) conditions for comparing value representations across methods, species, and training histories.
The prefrontal cortex (PFC) is critical for our ability to rapidly and flexibly adapt our behavior in new environments based on our previous experience. Despite its importance, the neural substrates and mechanisms by which the PFC supports this function have long remained enigmatic. Recent advances, however, have begun to change this. An increasingly large body of work suggests the PFC represents structured relationships—both among states of the outside world and between internally generated actions. In this review, we describe work from rodents, nonhuman primates, and humans to draw attention to the breadth of such representations and how they support flexible behavior. Across species, the PFC appears to represent the relational structure of problems: how stimuli relate to one another in cognitive maps or how different behaviors relate to one another when pursuing a goal. These results have started to reveal shared computational principles for PFC that generalize from rodents to humans and have inspired formal computational models and simulations. By reviewing experimental work showing both correlation and causation through invasive and noninvasive methods, along with theoretical work using artificial neural networks, we aim to highlight similarities and differences between species and models to provide a common language for interpreting findings in PFC. This will move us closer to a mechanistic understanding of the PFC that scales across tasks and species.
The prefrontal cortex (PFC) is crucial for economic decision-making. However, how PFC value representations facilitate flexible decisions remains unknown. We reframe economic decision-making as a navigation process through a cognitive map of choice values. We found rhesus macaques represented choices as navigation trajectories in a value space using a grid-like code. This occurred in ventromedial PFC (vmPFC) local field potential theta frequency across two datasets. vmPFC neurons deployed the same grid-like code and encoded chosen value. However, both signals depended on theta phase: occurring on theta troughs but on separate theta cycles. Finally, we found sharp-wave ripples-a key signature of planning and flexible behavior-in vmPFC. Thus, vmPFC utilizes cognitive map-based computations to organize and compare values, suggesting an alternative architecture for economic choice in PFC.
Reasoning flexibly composes known elements to solve novel problems. Recent theories suggest the brain uses the axis of time to compose elements for reasoning. In this view, elements are packaged into fast neural sequences, with each sequence exploring the implications of a different composition. Using magnetoencephalography, we tested this idea while participants mentally executed programs. Each program contained a set of steps linked to operations, the implications of which had to be computed. We found that behavioral performance scaled with program complexity, and by the end of execution, inferred program solutions were represented in prefrontal and parietal cortices. We identified a possible mechanism by which these solutions were computed. During reasoning, representations of steps in the program reactivated in fast sequences, consistent with sampling candidate partial solutions. Further evidence suggested these reactivations were accompanied by representations of their operations, and were followed by neural patterns reflecting their computed implications. Together, these results suggest replaying sequences supports program execution and reveal a highly organized temporal micro-architecture of reasoning.
A striking feature of human cognition is an exceptional ability to rapidly adapt to novel situations. It is proposed this relies on abstracting and generalizing past experiences. While previous research has explored how humans detect and generalize single sequential processes, we have a limited understanding of how humans adapt to more naturalistic scenarios, for example, complex, multisubprocess environments. Here, we propose a candidate computational mechanism that posits compositional generalization of knowledge about subprocess dynamics. In two samples (N = 238 and N = 137), we combined a novel sequence learning task and computational modeling to ask whether humans extract and generalize subprocesses compositionally to solve new problems. In prior learning, participants experienced sequences of compound images formed from two graphs' product spaces (group 1: G1 and G2, group 2: G3 and G4). In transfer learning, both groups encountered compound images from the product of G1 and G3, composed entirely of new images. We show that subprocess knowledge transferred between task phases, such that in a new task environment each group had enhanced accuracy in predicting subprocess dynamics they had experienced during prior learning. Computational models utilizing predictive representations, based solely on the temporal contiguity of experienced task states, without an ability to transfer knowledge, failed to explain these data. Instead, behavior was consistent with a predictive representation model that maps task states between prior and transfer learning. These results help advance a mechanistic understanding of how humans discover and abstract subprocesses composing their experiences and compositionally reuse prior knowledge as a scaffolding for new experiences.
The prefrontal cortex is crucial for learning and decision-making. Classic reinforcement learning (RL) theories center on learning the expectation of potential rewarding outcomes and explain a wealth of neural data in the prefrontal cortex. Distributional RL, on the other hand, learns the full distribution of rewarding outcomes and better explains dopamine responses. In the present study, we show that distributional RL also better explains macaque anterior cingulate cortex neuronal responses, suggesting that it is a common mechanism for reward-guided learning.
The last decade has seen substantial advances in the capacity to record behaviour and neural activity in humans in real-world settings, to simulate real-world situations in laboratory settings and to apply sophisticated analyses to large-scale data. Along with these developments, the call for ecological validity (increased use of naturalistic materials and more real-world-like settings for experiments) has been renewed. Here we sketch a framework for real-world research where previous approaches are integrated into a cyclic process of “bringing the lab to the real world” (recording behavioural and neural responses in their real-world settings) and “bringing the real-world to the lab” (manipulating the environments in which behaviours occur in the lab) that allows for discovery and theory development.
Abstract A human ability to adapt to the dynamics of novel environments relies on abstracting and generalizing from past experiences. Previous research has focused on how humans generalize from isolated sequential processes, yet we know little about mechanisms that enable adaptation to more complex dynamics, including those that govern much everyday experience. Here, using a novel sequence learning task based on graph factorization, coupled with simultaneous magnetoencephalography (MEG) recordings, we asked how reuse of experiential “building blocks” enables inference and generalization. Behavioral evidence was consistent with participants decomposing task experience into subprocesses, involving abstracting dynamical subprocess structures away from their sensory specifics and transferring these to a new task environment. Neurally this transfer was underpinned by a representational alignment of abstract subprocesses across task phases, evident in an enhanced neural similarity among stimuli that adhered to the same subprocesses, a temporally evolving mapping between predictive representations of subprocesses and a generalization of the dynamic roles that stimuli occupied within graph structures. Decoding strength for dynamical role representations predicted behavioral success in transfer of subprocess knowledge, consistent with a role in supporting behavioral adaptation in new environments. Our findings reveal neural dynamics that support compositional generalization, consistent with a structural scaffolding mechanism that facilitates efficient adaptation within new contexts.
Abstract Our experiences of the world usually reflect the interaction of multiple dynamical subprocesses. For example, clothes' appearance follows a complicated trajectory that is composed of factors such as becoming dirtier with wear and the colors fading slowly over time. In principle, decomposing these subprocesses can enhance learning efficiency, reduce memory requirements, and facilitate compositional reuse in new environments. This is because identical subprocesses can appear in other contexts, for example, discovering that colors in printed photographs also fade over time. Here, we combined a novel sequence learning task with computational modeling to test whether humans (N = 238) extract subprocesses from their holistic experiences, abstract these away from mere sensory experience, and efficiently recompose this knowledge to solve new problems. In a prior learning phase, two groups of participants were each exposed to sequences of compound images drawn from the product space of two graphs: G1 and G2 for group 1, G3 and G4 for group 2. Subsequently, in a transfer learning phase, all participants experienced compound images that were the product of G1 and G3 but composed of entirely new images. We found that knowledge of subprocesses transferred between tasks such that in a new task environment each group made more accurate predictions pertaining to the structure they had experienced in prior learning. Computational models utilizing predictive representations, based solely on the temporal contiguity of experienced task states, could not explain these data. Instead, behavior was consistent with a model performing structural inference over a hypothesis space of graph structures. Our results provide support for the idea that humans discover and abstract subprocesses from dynamic environments and reuse this knowledge to meet the demands of new environments in an efficient, resource-rational manner.
Physical human-robot interaction (pHRI) enables a user to interact with a physical robotic device to advance beyond the current capabilities of high-payload and high-precision industrial robots. This paradigm opens up novel applications where a the cognitive capability of a user is combined with the precision and strength of robots. Yet, current pHRI interfaces suffer from low take-up and a high cognitive burden for the user. We propose a novel framework that robustly and efficiently assists users by reacting proactively to their commands. The key insight is to include context- and user-awareness in the controller, improving decision-making on how to assist the user. Context-awareness is achieved by inferring the candidate objects to be grasped in a task or scene and automatically computing plans for reaching them. User-awareness is implemented by facilitating the motion toward the most likely object that the user wants to grasp, as well as dynamically recovering from incorrect predictions. Experimental results in a virtual environment of two degrees of freedom control show the capability of this approach to outperform manual control. By robustly predicting user intention, the proposed controller allows subjects to achieve superhuman performance in terms of accuracy and, thereby, usability.
The sex hormone estradiol is hypothesized to play a key role in human cognition, and reward processing specifically, via increased dopamine D1-receptor signalling. However, the effect of estradiol on reward processing in men has never been established. To fill this gap, we performed a double-blind placebo-controlled study in which men (N = 100) received either a single dose of estradiol (2 mg) or a placebo. Subjects performed a probabilistic reinforcement learning task where they had to choose between two options with varying reward probabilities to maximize monetary reward. Results showed that estradiol administration increased reward sensitivity compared to placebo. This effect was observed in subjects' choices, how much weight they assigned to their previous choices, and subjective reports about the reward probabilities. Furthermore, effects of estradiol were moderated by reward sensitivity, as measured through the BIS/BAS questionnaire. Using reinforcement learning models, we found that behavioral effects of estradiol were reflected in increased learning rates. These results demonstrate a causal role of estradiol within the framework of reinforcement learning, by enhancing reward sensitivity and learning. Furthermore, they provide preliminary evidence for dopamine-related genetic variants moderating the effect of estradiol on reward processing.
We use our eyes to assess the value of objects around us and carefully fixate options that we are about to choose. Neurons in the prefrontal cortex reliably encode the value of fixated options, which is essential for decision making. Yet as a decision unfolds, it remains unclear how prefrontal regions determine which option should be fixated next. Here we show that anterior cingulate cortex (ACC) encodes the value of options in the periphery to guide subsequent fixations during economic choice. In an economic decision-making task involving four simultaneously presented cues, we found rhesus macaques evaluated cues using their peripheral vision. This served two distinct purposes: subjects were more likely to fixate valuable peripheral cues, and more likely to choose valuable options whose cues were never even fixated. ACC, orbitofrontal cortex, dorsolateral prefrontal cortex, and ventromedial prefrontal cortex neurons all encoded cue value post-fixation. ACC was unique, however, in also encoding the value of cues before fixation and even cues that were never fixated. This pre-saccadic value encoding by ACC predicted which cue was next fixated during the decision process. ACC therefore conducts simultaneous processing of peripheral information to guide information sampling and choice during decision making.
Active inference is a normative principle underwriting perception, action, planning, decision-making and learning in biological or artificial agents. From its inception, its associated process theory has grown to incorporate complex generative models, enabling simulation of a wide range of complex behaviours. Due to successive developments in active inference, it is often difficult to see how its underlying principle relates to process theories and practical implementation. In this paper, we try to bridge this gap by providing a complete mathematical synthesis of active inference on discrete state-space models. This technical summary provides an overview of the theory, derives neuronal dynamics from first principles and relates this dynamics to biological processes. Furthermore, this paper provides a fundamental building block needed to understand active inference for mixed generative models; allowing continuous sensations to inform discrete representations. This paper may be used as follows: to guide research towards outstanding challenges, a practical guide on how to implement active inference to simulate experimental behaviour, or a pointer towards various in-silico neurophysiological responses that may be used to make empirical predictions.
!!Introduction It has been consistently proven that most of our decisions are malleable through informational and normative influence that individuals exert over us. Social status is known to play an important role in this. Those with high social status may influence those in a low status position, but not vice versa [1]. role of hormones in this has so far been neglected. Herein the hormone testosterone seems particularly relevant, given its role in a host of social behaviors, via its proposed role in promoting an individual's motivation to seek and maintain a high social status (for a review, see [2]). Therefore, we would like to test whether a manipulation of social status (high, low), and of testosterone levels (via administration of testosterone or a placebo), will influence an individual's decision making in a delay discounting task; thus effectively modulating time preferences. Delay discounting tasks are widely used in cognitive science to determine how strongly a person devaluates future rewards by having participants decide as to whether they prefer a smaller immediate or a larger delayed reward. Research so far has shown that a hyperbolic discounting model is best fitting to actual decisions people make. Because we would like to modulate individuals' time preferences, and expected effects are assumed to be fairly small based on previous research, we will need to develop very precise measurement tools. For that purpose, we will adapt an existing procedure [3] that involves a Bayesian updating approach, which is expected to yield the required precision. We predict that participants receiving testosterone will be less strongly influenced in their choices compared to participants receiving placebo. We also predict that the hormone effect is most pronounced in individuals with a high status, and that the hormone effect is reversed in individuals with low status. !!Results This master thesis topic is at an very early stage and therefore no data has yet been gathered. Possible caveats of the design will be discussed in the talk. !!Literature [1] R. M. Leary, P. K. Jongman-Sereno & J. K. Diebels, The Pursuit of Status: A Self-presentational Perspective on the Quest for Social Value in Psychology of Social Status, 2014, ch. 8, pp. 169–178. [2] C. Eisenegger, J. Haushofer & E. Fehr, The role of testosterone in social interaction, Trends in Cognitive Sciences, 2011, 15(6), pp. 263–271. [3] M. M. Garvert, M. Outoussis, Z. Kurth-Nelson, T. E. J. Behrens & R. J. Dolan, Learning-Induced plasticity in medial prefrontal cortex predicts preference malleability, 2015, Neuron, 85(2), pp. 418–428.