The historical literature on sensorimotor learning has focused on how we adapt to a single perturbation. This framing yielded influential state-space formulations of adaptation, but fails to capture the ecological demands faced by biological learners: the need to construct, select, and maintain a repertoire of motor memories tailored to diverse tools and environments. Recent research has shifted the focus toward frameworks based on contextual inference, in which the governing context is not directly observable and must instead be inferred from sensory cues, motor errors, and the statistics of context transitions. This perspective provides a unified account of how motor memories are created, expressed, and updated over time, and it delineates key open problems for future work: elucidating how motor memories are organized into hierarchical or compositional structures, how continual learning can proceed without overwriting existing memories, and determining how different training curricula shape the acquisition and retention of motor memories.
Abstract Skilled action requires expressing motor memories as appropriate for the current context, but context is often uncertain. Theoretical models make conflicting proposals about memory expression under contextual uncertainty, predicting belief-weighted combination of memories versus the selection of the most probable memory. We tested these predictions by training human participants to reach in two opposing force fields cued by the direction of a random dot motion stimulus whose coherence varied. When participants moved before reporting dot direction, adaptation scaled with coherence: low-reliability cues produced partial expression of both memories. Fitting Bayesian observer models to behavior favored belief-weighted memory combination. In contrast, when participants reported their choice before moving, adaptation was independent of coherence and model fits favored categorical memory selection. Thus, sensorimotor memories are expressed as either a probabilistic combination or categorical selection, depending on whether participants’ contextual inference remains implicit or is made explicit at the time of memory expression.
The Simons Collaboration on Ecological Neuroscience (SCENE) seeks to uncover general principles of brain function through an ecological perspective: studying perception, cognition, and action in the context of the affordances available to different agents. Here, we introduce SCENE's goals, hypotheses, and approaches outlining a collaborative vision for the next decade.
Most skilled behaviour occurs within reachable space; however, how humans learn to reach around obstacles in this space remains almost entirely unexplored. Here, using a novel robotic maze task that captures the richness of naturalistic hand-object interaction, we show that humans adaptively integrate model-based and model-free reinforcement learning strategies to act in reachable space. Fitting hybrid models to reach trajectories revealed that participants shifted from model-based toward model-free strategies across learning. Specifically, model-free reliance increased with state familiarity and distance from the goal, and with exclusive haptic feedback. Across participants, greater model-free reliance was associated with faster movements, consistent with reduced planning demands. Critically, direct comparison with an analogous virtual navigation task revealed stronger model-free reliance in reachable space than in navigable space, demonstrating that the computational architecture governing spatial learning is shared across scales but calibrated to the costs and constraints of the specific effector system.
Head movements cause artifacts in infant neuroimaging that can often render acquired data unusable. Here, we present a protocol that harnesses head motion from fMRI and resting-state (rs)-fMRI NIMH Data Archive (NDA) data to obtain quantitative measures of sensorimotor function in young children and associate them with future cognitive and autism outcomes. We describe steps for downloading, organizing, and pre-processing fMRI data to yield data on in-scan head motion. We then detail procedures for preparing phenotypic data to link with sensorimotor data. For complete details on the use and execution of this protocol, please refer to Denisova and Wolpert.1.
When learning multiple tasks, the structure of practice, or curriculum, profoundly influences learning outcomes across domains, including motor learning, rule learning, perceptual learning, and machine learning. In multitask learning settings, there is often a trade-off between the speed of acquisition and long-term retention. For example, in motor learning, acquisition appears faster, but retention is substantially reduced with blocked training compared to randomly interleaved training. In machine learning, this effect is known as catastrophic forgetting. In contrast, perceptual and cognitive learning benefit from structured, predictable curricula such as blocked training. We propose contextual inference as a unifying framework to explain these effects, emphasizing the integration of task transition dynamics, contextual cues and observation noise during learning. Insights from this framework may allow mitigating catastrophic interference in machine learning by leveraging principles inspired by biological learning.
Modern virtual reality (VR) devices record six-degree-of-freedom kinematic data with high spatial and temporal resolution and display high-resolution stereoscopic three-dimensional graphics. These capabilities make VR a powerful tool for many types of behavioural research, including studies of sensorimotor, perceptual and cognitive functions. Here we introduce Ouvrai, an open-source solution that facilitates the design and execution of remote VR studies, capitalizing on the surge in VR headset ownership. This tool allows researchers to develop sophisticated experiments using cutting-edge web technologies such as WebXR to enable browser-based VR, without compromising on experimental design. Ouvrai’s features include easy installation, intuitive JavaScript templates, a component library managing front- and backend processes and a streamlined workflow. It integrates with Firebase, Prolific and Amazon Mechanical Turk and provides data processing utilities for analysis. Unlike other tools, Ouvrai remains free, with researchers managing their web hosting and cloud database via personal Firebase accounts. Ouvrai is not limited to VR studies; researchers can also develop and run desktop or touchscreen studies using the same streamlined workflow. Through three distinct motor learning experiments, we confirm Ouvrai’s efficiency and viability for conducting remote VR studies. The authors introduce Ouvrai, an open-source solution that facilitates the design and execution of remote virtual reality studies, capitalizing on the surge in virtual reality headset ownership.
Neurobiological investigations of perceptual decision-making have furnished the first glimpse of a flexible cognitive process at the level of single neurons. Neurons in the parietal and prefrontal cortex are thought to represent the accumulation of noisy evidence, acquired over time, leading to a decision. Neural recordings averaged over many decisions have provided support for the deterministic rise in activity to a termination bound. Critically, it is the unobserved stochastic component that is thought to confer variability in both choice and decision time. Here, we elucidate this drift-diffusion signal on individual decisions. We recorded simultaneously from hundreds of neurons in the lateral intraparietal cortex of monkeys while they made decisions about the direction of random dot motion. We show that a single scalar quantity, derived from the weighted sum of the population activity, represents a combination of deterministic drift and stochastic diffusion. Moreover, we provide direct support for the hypothesis that this drift-diffusion signal approximates the quantity responsible for the variability in choice and reaction times. The population-derived signals rely on a small subset of neurons with response fields that overlap the choice targets. These neurons represent the integral of noisy evidence. Another subset of direction-selective neurons with response fields that overlap the motion stimulus appear to represent the integrand. This parsimonious architecture would escape detection by state-space analyses, absent a clear hypothesis.
Deciding how difficult it is going to be to perform a task allows us to choose between tasks, allocate appropriate resources, and predict future performance. To be useful for planning, difficulty judgments should not require completion of the task. Here, we examine the processes underlying difficulty judgments in a perceptual decision-making task. Participants viewed two patches of dynamic random dots, which were colored blue or yellow stochastically on each appearance. Stimulus coherence (the probability, p blue , of a dot being blue) varied across trials and patches thus establishing difficulty, | p blue −0.5|. Participants were asked to indicate for which patch it would be easier to decide the dominant color. Accuracy in difficulty decisions improved with the difference in the stimulus difficulties, whereas the reaction times were not determined solely by this quantity. For example, when the patches shared the same difficulty, reaction times were shorter for easier stimuli. A comparison of several models of difficulty judgment suggested that participants compare the absolute accumulated evidence from each stimulus and terminate their decision when they differed by a set amount. The model predicts that when the dominant color of each stimulus is known, reaction times should depend only on the difference in difficulty, which we confirm empirically. We also show that this model is preferred to one that compares the confidence one would have in making each decision. The results extend evidence accumulation models, used to explain choice, reaction time, and confidence to prospective judgments of difficulty.
Weight prediction is critical for dexterous object manipulation. Previous work has focused on lifting objects presented in isolation and has examined how the visual appearance of an object is used to predict its weight. Here we tested the novel hypothesis that when interacting with multiple objects, as is common in everyday tasks, people exploit the locations of objects to directly predict their weights, bypassing slower and more demanding processing of visual properties to predict weight. Using a three-dimensional robotic and virtual reality system, we developed a task in which participants were presented with a set of objects. In each trial a randomly chosen object translated onto the participant's hand and they had to anticipate the object's weight by generating an equivalent upward force. Across conditions we could control whether the visual appearance and/or location of the objects were informative as to their weight. Using this task, and a set of analogous web-based experiments, we show that when location information was predictive of the objects' weights participants used this information to achieve faster prediction than observed when prediction is based on visual appearance. We suggest that by "caching" associations between locations and weights, the sensorimotor system can speed prediction while also lowering working memory demands involved in predicting weight from object visual properties.NEW & NOTEWORTHY We use a novel object support task using a three-dimensional robotic interface and virtual reality system to provide evidence that the locations of objects are used to predict their weights. Using location information, rather than the visual appearance of the objects, supports fast prediction, thereby avoiding processes that can be demanding on working memory.
Learning to predict rewards is a fundamental driver of adaptive behavior. Midbrain dopamine neurons (DANs) play a key role in such learning by signaling reward prediction errors (RPEs) that teach recipient circuits about expected rewards given current circumstances and actions. However, the algorithm that DANs are thought to provide a substrate for, temporal difference (TD) reinforcement learning (RL), learns the mean of temporally discounted expected future rewards, discarding useful information concerning experienced distributions of reward amounts and delays. Here we present time-magnitude RL (TMRL), a multidimensional variant of distributional reinforcement learning that learns the joint distribution of future rewards over time and magnitude using an efficient code that adapts to environmental statistics. In addition, we discovered signatures of TMRL-like computations in the activity of optogenetically identified DANs in mice during a classical conditioning task. Specifically, we found significant diversity in both temporal discounting and tuning for the magnitude of rewards across DANs, features that allow the computation of a two dimensional, probabilistic map of future rewards from just 450ms of neural activity recorded from a population of DANs in response to a reward-predictive cue. In addition, reward time predictions derived from this population code correlated with the timing of anticipatory behavior, suggesting the information is used to guide decisions regarding when to act. Finally, by simulating behavior in a foraging environment, we highlight benefits of access to a joint probability distribution of reward over time and magnitude in the face of dynamic reward landscapes and internal physiological need states. These findings demonstrate surprisingly rich probabilistic reward information that is learned and communicated to DANs, and suggest a simple, local-in-time extension of TD learning algorithms that explains how such information may be acquired and computed.
Nearly all tasks of daily life involve skilled object manipulation, and successful manipulation requires knowledge of object dynamics. We recently developed a motor learning paradigm that reveals the categorical organization of motor memories of object dynamics. When participants repeatedly lift a constant-density “family” of cylindrical objects that vary in size, and then an outlier object with a greater density is interleaved into the sequence of lifts, they often fail to learn the weight of the outlier, persistently treating it as a family member despite repeated errors. Here we examine eight factors (Similarity, Cardinality, Frequency, History, Structure, Stochasticity, Persistence, and Time Pressure) that could influence the formation and retrieval of category representations in the outlier paradigm. In our web-based task, participants (N = 240) anticipated object weights by stretching a virtual spring attached to the top of each object. Using Bayesian t -tests, we analyze the relative impact of each manipulated factor on categorical encoding (strengthen, weaken, or no effect). Our results suggest that category representations of object weight are automatic, rigid, and linear and, as a consequence, the key determinant of whether an outlier is encoded as a member of the family is its discriminability from the family members.
The brain can learn to generate actions, such as reaching to a target, using different movement strategies. Understanding how different variables bias which strategies are learned to produce such a reach is important for our understanding of the neural bases of movement. Here we introduce a novel spatial forelimb target task in which perched head-fixed mice learn to reach to a circular target area from a set start position using a joystick. These reaches can be achieved by learning to move into a specific direction or to a specific endpoint location. We find that mice gradually learn to successfully reach the covert target. With time, they refine their initially exploratory complex joystick trajectories into controlled targeted reaches. The execution of these controlled reaches depends on the sensorimotor cortex. Using a probe test with shifting start positions, we show that individual mice learned to use strategies biased to either direction or endpoint-based movements. The degree of endpoint learning bias was correlated with the spatial directional variability with which the workspace was explored early in training. Furthermore, we demonstrate that reinforcement learning model agents exhibit a similar correlation between directional variability during training and learned strategy. These results provide evidence that individual exploratory behavior during training biases the control strategies that mice use to perform forelimb covert target reaches.
Motor errors can have both bias and noise components. Bias can be compensated for by adaptation and, in tasks in which the magnitude of noise varies across the environment, noise can be reduced by identifying and then acting in less noisy regions of the environment. Here we examine how these two processes interact when participants reach under a combination of an externally imposed visuomotor bias and noise. In a center-out reaching task, participants experienced noise (zero-mean random visuomotor rotations) that was target-direction dependent with a standard deviation that increased linearly from a least-noisy direction. They also experienced a constant bias, a visuomotor rotation that varied (across groups) from 0 to 40 degrees. Critically, on each trial, participants could select one of three targets to reach to, thereby allowing them to potentially select targets close to the least-noisy direction. The group who experienced no bias (0 degrees) quickly learned to select targets close to the least-noisy direction. However, groups who experienced a bias often failed to identify the least-noisy direction, even though they did partially adapt to the bias. When noise was introduced after participants experienced and adapted to a 40 degrees bias (without noise) in all directions, they exhibited an improved ability to find the least-noisy direction. We developed two models-one for reach adaptation and one for target selection-that could explain participants' adaptation and target-selection behavior. Our data and simulations indicate that there is a trade-off between adaptation and selection. Specifically, because bias learning is local, participants can improve performance, through adaptation, by always selecting targets that are closest to a chosen direction. However, this comes at the expense of improving performance, through selection, by reaching toward targets in different directions to find the least-noisy direction.
Modern virtual reality (VR) devices offer 6 degree-of-freedom kinematic data with high spatial and tem-poral resolution, making them powerful tools for research on sensorimotor and cognitive functions. We introduce Ouvrai, an open-source solution that facilitates the design and execution of remote VR studies, capitalizing on the surge in VR headset ownership. This tool allows researchers to develop sophisticated experiments using cutting-edge web technologies like the WebXR Device API for browser-based VR, with-out compromising on experimental design. Ouvrai’s features include easy installation, intuitive JavaScript templates, a component library managing front- and back-end processes, and a streamlined workflow. It also integrates APIs for Firebase, Prolific, and Amazon Mechanical Turk and provides data processing utilities for analysis. Unlike other tools, Ouvrai remains free, with researchers managing their web hosting and cloud database via personal Firebase accounts. Through three distinct motor learning experiments, we confirm Ouvrai’s efficiency and viability for conducting remote VR studies.
Motor adaptation can be achieved through error-based learning, driven by sensory prediction errors, or reinforcement learning, driven by reward prediction errors. Recent work on visuomotor adaptation has shown that reinforcement learning leads to more persistent adaptation when visual feedback is removed, compared to error-based learning in which continuous visual feedback of the movement is provided. However, there is evidence that error-based learning with terminal visual feedback of the movement (provided at the end of movement) may be driven by both sensory and reward prediction errors. Here we examined the influence of feedback on learning using a visuomotor adaptation task in which participants moved a cursor to a single target while the gain between hand and cursor movement displacement was gradually altered. Different groups received either continuous error feedback (EC), terminal error feedback (ET), or binary reinforcement feedback (success/fail) at the end of the movement (R). Following adaptation we tested generalization to targets located in different directions and found that generalization in the ET group was intermediate between the EC and R groups. We then examined the persistence of adaptation in the EC and ET groups when the cursor was extinguished and only binary reward feedback was provided. Whereas performance was maintained in the ET group, it quickly deteriorated in the EC group. These results suggest that terminal error feedback leads to a more robust form of learning than continuous error feedback. In addition our findings are consistent with the view that error-based learning with terminal feedback involves both error-based and reinforcement learning.
Context is widely regarded as a major determinant of learning and memory across numerous domains, including classical and instrumental conditioning, episodic memory, economic decision-making, and motor learning. However, studies across these domains remain disconnected due to the lack of a unifying framework formalizing the concept of context and its role in learning. Here, we develop a unified vernacular allowing direct comparisons between different domains of contextual learning. This leads to a Bayesian model positing that context is unobserved and needs to be inferred. Contextual inference then controls the creation, expression, and updating of memories. This theoretical approach reveals two distinct components that underlie adaptation, proper and apparent learning, respectively referring to the creation and updating of memories versus time-varying adjustments in their expression. We review a number of extensions of the basic Bayesian model that allow it to account for increasingly complex forms of contextual learning.