Recent advances in generative and embodied AI have been driven by large-scale predictive learning over multimodal data. However, the resulting systems remain largely based on passive training regimes where linguistic regularities create the scaffold onto which information from other modalities is attached. Conversely, neuroscience and cognitive science suggest that biological intelligence is organized in the opposite way, where grounded world models acquired through interaction with the environment provide the semantic scaffold to which language is attached. Here, we illustrate five examples of neural circuits supporting grounded world modelling, which underlie navigation in physical and conceptual spaces, affordance-based perception and interaction with objects, active perception and exploratory learning, allostatic control and emotion, and the distinction between self- and world-generated outcomes. These examples highlight several features largely missing from current embodied AI, including the role of intrinsic dynamics as a foundation for learning, the centrality of action in aligning these dynamics with the external world, the prominence of autonomous experience and open-ended learning over passive assimilation of externally provided data, and the fact that early predictive and control mechanisms scaffold higher cognitive abilities such as reasoning, conceptual navigation, planning, imagination, understanding others' minds, and communication. Finally, we discuss whether and how principles derived from biological systems may inform future embodied AI, including training regimes based on social interaction to construct world models that are not only grounded but also socially shared and aligned with human norms and values.
Everyday tasks, such as selecting routes when driving or preparing meals, require making sequences of embodied decisions, in which planning and action processes are intertwined. Here, we address how people make sequential embodied decisions, requiring balancing between immediate affordances and long-term utilities of alternative action plans. We designed a virtually embodied task in which participants controlled an avatar tasked with "crossing rivers" by jumping across rocks. The task permitted us to assess how participants balanced between immediate jumping affordances ("safe" vs. "risky" jumps) and the utility (length) of the ensuing paths to the goal. Behavioral and computational analyses revealed that participants planned ahead their path to the goal rather than simply focusing on the immediate jumping affordances. Furthermore, spatial and embodied components of the task influenced participants' decision strategies, as participants' current direction of movement and momentum influenced their choice between safe and risky jumps. Additionally, participants showed (pre)planning before making the first jump, but they continued deliberating during it, with movement speed decreasing at decision points and when approaching them. Computational modeling indicates that farsighted participants who assigned greater weight to the utility of future jumps showed a better performance, highlighting the usefulness of planning in embodied settings. Finally, analyzing participants' performance indicates that during the experiment, they become faster in moving and deciding but they do not change their overall strategy. Our findings underscore the importance of studying decision-making in ecologically valid, embodied settings, providing insights into the interplay between action and cognition in real-world planning-while-acting scenarios.
Humans and other animals continuously make embodied decisions about ongoing or pending courses of action. Examples of embodied decisions include a hunting lioness's decision of which gazelle to chase and a soccer player's decision of which teammate to pass the ball to. The study of embodied decisions has recently gained tractions across several fields, including cognitive psychology, neuroscience, and sports science. Here, we summarize key insights from these studies and highlight that they imply a shift of perspective from viewing decision-making as a central cognitive process largely separated from perception and action dynamics to a more integrative perspective that recognizes its embodied and situated nature. We discuss how embodied decisions can be effectively conceptualized in terms of the parallel specification and selection between available (and future) affordances, i.e., as an “affordance competition” process. We discuss studies addressing various aspects of embodied decisions, which include the selection between courses of action, the involvement of motor processes in perceptual and cognitive tasks, motivational factors and the decision of how vigorously and urgently to act. Furthermore, we highlight current controversies in the field and open directions for future work – and their implications for the advancement of our understanding of the mind and the behavior of athletes.
Perceptual decision-making is the process by which sensory evidence is combined with prior knowledge and transformed into possible movement plans according to a rule or policy. Classic studies suggested that perceptual decisions emerge from a feedforward hierarchy of brain areas with distinct functions and fairly homogeneous neural representations. However, more recent findings argue that decisions emerge from distributed, recurrent computations across many brain areas (a “heterarchy”) with complex, heterogeneous representations. How can we make sense of these findings in a way that preserves the computational elegance of the conventional view? In this review, we describe how a new generation of studies is leveraging high-density electrophysiology, incisive task designs, causal manipulations (e.g., optogenetics) and statistical approaches for probing inter-area communication, and theoretical methods that connect population dynamics with representational geometry to build a modern framework for understanding perceptual decisions.
Accurate interaction with the environment relies on the integration of external information about the spatial layout of potential actions and knowledge of their costs and benefits. Previous studies have shown that when given a choice between voluntary reaching movements, humans tend to prefer actions with lower biomechanical costs. However, these studies primarily focused on decisions made before the onset of movement ("decide-then-act" scenarios), and it is not known to what extent their conclusions generalize to many real-life situations, in which decisions occur during ongoing actions ("decide-while-acting"). For example, one recent study found that biomechanical costs did not influence decisions to switch from a continuous manual tracking movement to a point-to-point movement, suggesting that biomechanical costs may be disregarded in decide-while-acting scenarios. To better understand this surprising result, we designed an experiment in which participants were faced with the decision between continuing to track a target moving along a straight path or changing paths to track a new target that gradually moved along a direction that deviated from the initial one. We manipulated tracking direction, angular deviation rate, and side of deviation, allowing us to compare scenarios where biomechanical costs favored either continuing or changing the path. Crucially, here the choice was always between two continuous tracking actions. Our results show that in this situation decisions clearly took biomechanical costs into account. Thus we conclude that biomechanics are not disregarded during decide-while-acting scenarios but rather that cost comparisons can only be made between similar types of actions. NEW & NOTEWORTHY In this study, we aim to shed light on how biomechanical factors influence decisions made during ongoing actions. Previous work suggested that decisions made during actions disregard biomechanical costs, in contrast to decisions made before movement. Our results challenge that proposal and suggest instead that the effect of biomechanical factors is dependent on the types of actions being compared (e.g., continuous tracking vs. point-to-point reaching). These findings contribute to our understanding of the dynamic interplay between biomechanical considerations and action choices during ongoing interactions with the environment.
Everyday tasks, such as selecting routes when driving or preparing meals require making sequences of embodied decisions, in which planning and action processes are intertwined. In this study, we address how people make sequential embodied decisions, requiring balancing between immediate affordances and long-term utilities of alternative action plans. We designed a novel game-like task in which participants controlled an avatar tasked with "crossing rivers", by jumping across rocks. The task permitted us to assess how participants balanced between immediate jumping affordances ("safe" versus "risky" jumps) and the utility (length) of the ensuing paths to the goal. Behavioral and computational analyses revealed that participants' planned ahead their path to the goal rather than simply focusing on the most immediate jumping affordances. Furthermore, embodied components of the task influenced participants decision strategies, as evident by the fact that participants' current direction of movement influenced their choice between safe and risky jumps. We also found that participants showed (pre)planning before making the first jump, but they continued deliberating during it, with movement speed decreasing at decision points and when approaching them. Finally, computational modeling indicates that farsighted participants who assigned greater weight to the utility of future jumps showed a better performance, highlighting the usefulness of planning in embodied settings. Our findings underscore the importance of studying decision-making and planning in ecologically valid, embodied settings, providing new insights into the interplay between action and cognition in real-world planning-while-acting scenarios. ### Competing Interest Statement The authors have declared no competing interest.
One of the most exciting new developments in systems neuroscience is the progress being made toward neurophysiological experiments that move beyond simplified laboratory settings and address the richness of natural behavior. This is enabled by technological advances such as wireless recording in freely moving animals, automated quantification of behavior, and new methods for analyzing large data sets. Beyond new empirical methods and data, however, there is also a need for new theories and concepts to interpret that data. Such theories need to address the particular challenges of natural behavior, which often differ significantly from the scenarios studied in traditional laboratory settings. Here, we discuss some strategies for developing such novel theories and concepts and some example hypotheses being proposed.
Prominent accounts of sentient behavior depict brains as generative models of organismic interaction with the world, evincing intriguing similarities with current advances in generative artificial intelligence (AI). However, because they contend with the control of purposive, life-sustaining sensorimotor interactions, the generative models of living organisms are inextricably anchored to the body and world. Unlike the passive models learned by generative AI systems, they must capture and control the sensory consequences of action. This allows embodied agents to intervene upon their worlds in ways that constantly put their best models to the test, thus providing a solid bedrock that is - we argue - essential to the development of genuine understanding. We review the resulting implications and consider future directions for generative AI.
Psychology and neuroscience are concerned with the study of behavior, of internal cognitive processes, and their neural foundations. However, most laboratory studies use constrained experimental settings that greatly limit the range of behaviors that can be expressed. While focusing on restricted settings ensures methodological control, it risks impoverishing the object of study: by restricting behavior, we might miss key aspects of cognitive and neural functions. In this article, we argue that psychology and neuroscience should increasingly adopt innovative experimental designs, measurement methods, analysis techniques and sophisticated computational models to probe rich, ecologically valid forms of behavior, including social behavior. We discuss the challenges of studying rich forms of behavior as well as the novel opportunities offered by state-of-the-art methodologies and new sensing technologies, and we highlight the importance of developing sophisticated formal models. We exemplify our arguments by reviewing some recent streams of research in psychology, neuroscience and other fields (e.g., sports analytics, ethology and robotics) that have addressed rich forms of behavior in a model-based manner. We hope that these "success cases" will encourage psychologists and neuroscientists to extend their toolbox of techniques with sophisticated behavioral models - and to use them to study rich forms of behavior as well as the cognitive and neural processes that they engage.
Abstract In this chapter, the authors describe a hypothesis on the functional organization of the thalamocortical and corticostriatal circuits of early mammals, constrained by inferences about ancestral amniotes and studies in modern mammalian species. Following the work of Graziano and others, they interpret the organization of sensorimotor neocortex in terms of “action maps” adapted to the needs of specific aspects of an animal’s behavioral repertoire, such as searching, feeding, handling, and defensive actions. The authors propose that selection between these activity types (e.g., search versus feed) is governed by the basal ganglia, while the particular details of actions within a type of activity (e.g., direction of locomotion) are selected within the cortical action map specialized for that activity. This proposal is supported by neuro-anatomical data from a wide range of species, as well as neural recording, microstimulation, and inactivation studies in rodents and primates, and leads to some testable predictions.
When choosing between options with multiple attributes, do we integrate all attributes into a unified measure for comparison, or does the comparison also occur at the level of each attribute, involving parallel processes that can dynamically influence each other? What happens when independent sensory features all carry information about the same decision factor, such as reward value? To investigate these questions, we asked human participants to perform a two-alternative forced choice reaching task in which the reward value of a target was indicated by two visual attributes-its brightness ("bottom-up," BU feature) and its orientation ("top-down," TD feature). If decisions always occur after the integration of both features, there should be no difference in the reaction time (RT) regardless of the attribute combinations that drove the choice. Counter to that prediction, RT distributions depended on the attribute combinations of given targets and the choices made by participants. RTs were shortest when both attributes were congruent or when the choice was based on the bottom-up feature, and longer when the attributes were in conflict (favoring opposite options). In conflict trials, nearly two-thirds of participants made faster decisions when choosing the option favored by the bottom-up feature than when choosing the top-down-favored option. We also observed mid-reach changes-of-mind in a subset of conflict trials, mostly changing from the bottom-up to the top-down-favored target. These data suggest that multi-attribute value-based decisions are better explained by a distributed process including competition among different features than by a competition based on a single, integrated estimate of value.NEW & NOTEWORTHY We show that during value-based decisions, humans do not always use all reward-related information to make their choice, but instead can "jump the gun" using partial information. In particular, when different sources of information were in conflict, early decisions were mostly based on fast bottom-up information, and sometimes followed by corrective changes-of-mind based on slower top-down information. This supports parallel decision processes among different information sources, as opposed to a single integrated "common currency."
When choosing between options with multiple attributes, do we decide by integrating all of the attributes into a unified measure for comparison, or does the comparison occur at the level of each attribute, involving independent competitive processes that can dynamically influence each other? What happens when independent sensory features all carry information about the same decision factor, such as reward value? To investigate these questions, we asked human participants to perform a two-alternative forced choice task in which the reward value of a target was indicated by two independent visual attributes – its brightness (“bottom-up” feature) and its orientation (“top-down” feature). If decisions always occur after integration of both features, there should be no difference in the reaction time (RT) distribution regardless of the attribute combinations that drove the choice. Counter to that prediction, almost two-thirds of the participants exhibited RT differences that depended on the attribute combinations of given targets. The RT was shortest when both attributes were congruent or when the choice was based on the bottom-up feature, and longer when the attributes were in conflict (favoring opposite options), especially when choosing the option favored by the top-down feature. We also observed mid-reach changes-of-mind in a subset of conflict trials, mostly changing from the bottom-up to the top-down-favored target. These data suggest that multi-attribute value-based decisions are better explained by a distributed competition among different features than by a competition based on a single, integrated estimate of choice value. New & Noteworthy This study showed that during value-based decisions, humans do not always take all information about reward value into account to make their choice, but instead can “jump the gun” using partial information. In particular, when different sources of information were in conflict, early decisions were mostly based on fast bottom-up information, and sometimes followed by corrective changes-of-mind based on slower top-down information. Our results suggest that parallel decision processes occur among different information sources, as opposed to between a single integrated “common currency”.
Humans and other animals are able to adjust their speed-accuracy trade-off (SAT) at will depending on the urge to act, favoring either cautious or hasty decision policies in different contexts. An emerging view is that SAT regulation relies on influences exerting broad changes on the motor system, tuning its activity up globally when hastiness is at premium. The present study aimed to test this hypothesis. A total of 50 participants performed a task involving choices between left and right index fingers, in which incorrect choices led either to a high or to a low penalty in 2 contexts, inciting them to emphasize either cautious or hasty policies. We applied transcranial magnetic stimulation (TMS) on multiple motor representations, eliciting motor-evoked potentials (MEPs) in 9 finger and leg muscles. MEP amplitudes allowed us to probe activity changes in the corresponding finger and leg representations, while participants were deliberating about which index to choose. Our data indicate that hastiness entails a broad amplification of motor activity, although this amplification was limited to the chosen side. On top of this effect, we identified a local suppression of motor activity, surrounding the chosen index representation. Hence, a decision policy favoring speed over accuracy appears to rely on overlapping processes producing a broad (but not global) amplification and a surround suppression of motor activity. The latter effect may help to increase the signal-to-noise ratio of the chosen representation, as supported by single-trial correlation analyses indicating a stronger differentiation of activity changes in finger representations in the hasty context.
Studies of neural population dynamics of cell activity from monkey motor areas during reaching show that it mostly represents the generation and timing of motor behavior. We compared neural dynamics in dorsal premotor cortex (PMd) during the performance of a visuomotor task executed individually or cooperatively and during an observation task. In the visuomotor conditions, monkeys applied isometric forces on a joystick to guide a visual cursor in different directions, either alone or jointly with a conspecific. In the observation condition, they observed the cursor's motion guided by the partner. We found that in PMd neural dynamics were widely shared across action execution and observation, with cursor motion directions more accurately discriminated than task types. This suggests that PMd encodes spatial aspects irrespective of specific behavioral demands. Furthermore, our results suggest that largest components of premotor population dynamics, which have previously been suggested to reflect a transformation from planning to movement execution, may rather reflect higher cognitive-motor processes, such as the covert representation of actions and goals shared across tasks that require movement and those that do not.
Recent theoretical models suggest that deciding about actions and executing them are not implemented by completely distinct neural mechanisms but are instead two modes of an integrated dynamical system. Here, we investigate this proposal by examining how neural activity unfolds during a dynamic decision-making task within the high-dimensional space defined by the activity of cells in monkey dorsal premotor (PMd), primary motor (M1), and dorsolateral prefrontal cortex (dlPFC) as well as the external and internal segments of the globus pallidus (GPe, GPi). Dimensionality reduction shows that the four strongest components of neural activity are functionally interpretable, reflecting a state transition between deliberation and commitment, the transformation of sensory evidence into a choice, and the baseline and slope of the rising urgency to decide. Analysis of the contribution of each population to these components shows meaningful differences between regions but no distinct clusters within each region, consistent with an integrated dynamical system. During deliberation, cortical activity unfolds on a two-dimensional “decision manifold” defined by sensory evidence and urgency and falls off this manifold at the moment of commitment into a choice-dependent trajectory leading to movement initiation. The structure of the manifold varies between regions: In PMd, it is curved; in M1, it is nearly perfectly flat; and in dlPFC, it is almost entirely confined to the sensory evidence dimension. In contrast, pallidal activity during deliberation is primarily defined by urgency. We suggest that these findings reveal the distinct functional contributions of different brain regions to an integrated dynamical system governing action selection and execution.
Abstract Making a good decision often takes time, and in general, taking more time improves the chances of making the right choice. During the past several decades, the process of making decisions in time has been described through a class of models in which sensory evidence about choices is accumulated until the total evidence for one of the choices reaches some threshold, at which point commitment is made and movement initiated. Thus, if sensory evidence is weak (and noise in the signal increases the probability of an error), then it takes longer to reach that threshold than if sensory evidence is strong (thus helping filter out the noise). Crucially, the setting of the threshold can be increased to emphasize accuracy or lowered to emphasize speed. Such accumulation-to-bound models have been highly successful in explaining behavior in a very wide range of tasks, from perceptual discrimination to deliberative thinking, and in providing a mechanistic explanation for the observation that neural activity during decision-making tends to build up over time. However, like any model, they have limitations, and recent studies have motivated several important modifications to their basic assumptions. In particular, recent theoretical and experimental work suggests that the process of accumulation favors novel evidence, that the threshold decrease over time, and that the result yields improved decision-making in real, natural situations.
This article outlines a hypothetical sequence of evolutionary innovations, along the lineage that produced humans, which extended behavioural control from simple feedback loops to sophisticated control of diverse species-typical actions. I begin with basic feedback mechanisms of ancient mobile animals and follow the major niche transitions from aquatic to terrestrial life, the retreat into nocturnality in early mammals, the transition to arboreal life and the return to diurnality. Along the way, I propose a sequence of elaboration and diversification of the behavioural repertoire and associated neuroanatomical substrates. This includes midbrain control of approach versus escape actions, telencephalic control of local versus long-range foraging, detection of affordances by the dorsal pallium, diversified control of nocturnal foraging in the mammalian neocortex and expansion of primate frontal, temporal and parietal cortex to support a wide variety of primate-specific behavioural strategies. The result is a proposed functional architecture consisting of parallel control systems, each dedicated to specifying the affordances for guiding particular species-typical actions, which compete against each other through a hierarchy of selection mechanisms. This article is part of the theme issue 'Systems neuroscience through the lens of evolutionary theory'.