Systematic intra-saccadic displacements of the saccade target object induce oculomotor learning. This type of manipulation remains undetected because of saccadic suppression of displacement. The subtle intra-saccadic manipulation causes the target object to appear at a retinal position that is inconsistent with the computed displacement of visual space during the saccade, leading to the computation of a postdicted motor error and corresponding adaptive changes to the saccade amplitude and the perceived location of objects in space. Adaptive changes to the saccade amplitude or direction have also been observed following obvious intra-saccadic manipulations, but very little is known about the characteristics of this learning process or its underlying mechanisms. The current study therefore compares oculomotor learning in response to subtle and obvious changes of the stimulus display. By measuring their effect on object localization and saccade gain, we were able to show that two different kinds of sensory error can induce oculomotor learning: a global spatial error and a local feature error. While experiencing the former leads to the typical adaptation-induced shift in perceived object location, experiencing the latter does not. On top of that, we demonstrate that adaptive adjustments to the oculomotor behavior do not necessarily rely on those sensory errors. Instead, the overt oculomotor behavior minimizes a task error.
For goal-directed movements like throwing darts or shooting a soccer penalty, the optimal location to aim depends on the endpoint variability of an individual. Currently, there is no consensus on whether people can optimize their movement planning based on information about their motor variability. Here, we tested the role of different types of feedback for movement planning under risk. We measured saccades toward a bar that consisted of a reward and a penalty region. Participants either received error-based feedback about their endpoint or reinforcement feedback about the resulting reward. We additionally manipulated the feedback schedule to assess the role of feedback frequency and whether feedback focusses on individual trials or a group of trials. Participants with trial-by-trial reinforcement feedback performed best. They were less loss-aversive, had the least endpoint deviation from optimality, and showed more consistent performance at the group level. This combination of reduced between-participant variability and the improved alignment with optimality suggests that reinforcement feedback about a single movement is particularly effective to optimize movement planning under risk.
Redirected walking (RDW) alters the relationship between physical locomotion and visual feedback to guide users along virtual paths that vary from their real-world trajectories. Over time these changes can lead to adaptation of the user and produce novel sensorimotor settings that users might retain across sessions and apply directly once they enter virtual reality (VR). Although short-term adaptation to altered sensorimotor contingencies is well established, it remains unclear whether such adaptation is retained across multiple days. We investigated long-term retention of RDW adaptation by repeatedly exposing ten participants to a fixed rightward curvature gain of π /30 across nine sessions over two weeks. Each session consisted of 200 walking repetitions with gain applied. Adaptation was assessed before and after each session using blind walking and a pointing task. Moreover, perceptual detection thresholds for curvature gains were measured before the first session (Baseline), a day after the ninth session (Final), and once in-between before the 5th adaptation session. Results showed clear retention of adapted locomotor behavior: during blind walking at the beginning of the sessions, when instructed to walk straight, participants consistently exhibited curved trajectories, which indicates that newly acquired sensorimotor contingencies can be retained over days and are immediately available upon re-entering VR. At the same time, pointing accuracy remained stable throughout the experiment and detection thresholds showed no consistent changes across sessions. In summary, our study provides evidence that adaptation to altered sensorimotor contingencies in VR can be retained across multiple sessions and days and can be available as soon as the user enters VR. This may be useful for many scenarios in which users repeatedly use VR tools over a long period of time. The complete data set, all supplemental materials and the preprint of the manuscript are available at https://osf.io/z973w.
Why do people differ in their perceptual judgment despite observing the same situation? According to “motivated perception”, a person’s motivation can alter how the brain interprets incoming sensory information. Yet, empirical support remains mixed, often due to methodological confounds. Here, we systematically tested whether motivation alters perception or whether it instead biases behavior. Across four experiments, we assessed the quality and quantity of motivation (self-concordance, value) and two key dimensions of perception: bias and sensitivity. Moreover, we tested two potentially mediating mechanisms (gaze position, spatial attention) as well as an implicit perceptual measure. Using smooth pursuit eye movements as an implicit measure of motion perception, we show that motivation biases responses without altering perception (Experiment 2). Changes in perceptual sensitivity (Experiment 1) and perceptual bias (Experiments 3–4) only arise when participants can freely select their gaze position in response to uncertain or ambiguous visual displays. Our findings therefore challenge the notion of “motivated perception”. Instead, they suggest that motivation shapes how we look and respond – but not how we perceive.
Within theories of agency perception and sensorimotor learning, predictive sensorimotor processing is essential for understanding how humans estimate and experience control over their actions and their environment. This study investigates the relationship between predictive eye movements and the subjective sense of control in a dynamic visuomotor task. Control is manipulated through a gain parameter that alters the correspondence between finger movement and visual feedback. Participants are required to locate the outcome of their finger movements via predictive saccades and to report perceived changes in control across trials within blocks of either stable or variable gain. Predictive gaze errors, calculated as the difference between the saccade landing point and the actual visual outcome location, are modeled alongside trial-wise perceived control judgments using Bayesian hierarchical logistic regression. This paradigm enables analysis of how control estimation evolves through learning and how prior experience with stable control conditions influences behavior under variable ones. The aim of this study is to establish oculomotor behaviour as a proxy for predictive sensorimotor processing during visual feedback monitoring, thereby enabling inference about an agent’s sense of control from eye-movements. Such a connection has implications for optimisation of human–machine interactions.
Intercepting a moving target requires prediction: knowing the current position is not enough due to sensorimotor delays. A moving “vortex” stimulus, which lacks a conventional velocity signal for smooth pursuit yet still elicits accurate saccade targeting, was used here to assess whether it also supports accurate reaching movements. In our experiments, participants reached to the vortex or to a rigid moving disk. Three manipulations probed prediction and online control: constant-speed interception, speed perturbations at reach onset, and target disappearance at reach onset. Reaction times when reaching to the vortex did not depend on target speed in the way they do for rigid targets. Moreover, when reaching to the vortex participants made larger undershoots that increased with target speed, and they compensated less for speed changes. Thus, the vortex provides motion-related information that is sufficient for accurate saccadic localization, but that is not effectively exploited for predictive manual interception. This dissociation suggests that motion information may not be used in the same manner across motor effectors.
Many physical objects in the visual world can move in deformable, non-rigid ways. A new study shows how such motion is perceived and finds striking similarities to the perception of human bodies in motion.
A considerable part of the performance of today's large language models (LLM's) and multimodal large language models (MLLM's) depends on their tokenization strategies. While tokenizers are extensively researched for textual and visual input, there is no research on tokenization strategies for gaze data due to its nature. However, a corresponding tokenization strategy would allow using the vision capabilities of pre-trained MLLM's for gaze data, for example, through fine-tuning. In this paper, we aim to close this research gap by analyzing five different tokenizers for gaze data on three different datasets for the forecasting and generation of gaze data through LLMs (cf. ). We evaluate the tokenizers regarding their reconstruction and compression abilities. Further, we train an LLM for each tokenization strategy, measuring its generative and predictive performance. Overall, we found that a quantile tokenizer outperforms all others in predicting the gaze positions and k-means is best when predicting gaze velocities.
The interplay between attention, alertness, and motor planning is crucial for our manual interactions. To investigate the neural bases of this interaction and challenge the views that attention cannot be disentangled from motor planning, we instructed human volunteers of both sexes to plan and execute reaching movements while attending to the target, while attending elsewhere, or without constraining attention. We recorded reaction times to reach initiation and pupil diameter and interfered with the functions of the medial posterior parietal cortex (mPPC) with online repetitive transcranial magnetic stimulation to test the causal role of this cortical region in the interplay between spatial attention and reaching. We found that mPPC plays a key role in the spatial association of reach planning and covert attention. Moreover, we have found that alertness, measured by pupil size, is a good predictor of the promptness of reach initiation only if we plan a reach to attended targets, and mPPC is causally involved in this coupling. Different from previous understanding, we suggest that mPPC is neither involved in reach planning per se, nor in sustained covert attention in the absence of a reach plan, but it is specifically involved in attention functional to reaching.
Optic flow, the retinal pattern of motion experienced during self-motion, contains information about one's direction of heading. The global pattern due to self-motion is locally confounded when moving objects are present, and the flow is the sum of components due to the different causal sources. Nonetheless, humans can accurately retrieve information from such flow, including the direction of heading and the scene-relative motion of an object. Flow parsing is a process speculated to allow the brain's sensitivity to optic flow to separate the causal sources of retinal motion in information due to self-motion and information due to object motion. In a computational model that retrieves object and self-motion information from optic flow, we implemented flow parsing based on heading likelihood maps, whose distributions indicate the consistency of parts of the flow with self-motion. This allows for concurrent estimation of heading, detecting and localizing a moving object, and estimating its scene-relative motion. We developed a paradigm that allows the model to perform all these estimations while systematically varying the object's contribution to the flow field. Simulations of that paradigm show that the model replicates many aspects of human performance, including the dependence of heading estimation on object speed and direction.
Gaze-based interaction techniques have created significant interest in the field of spatial interaction. Many of these methods require additional input modalities, such as hand gestures (e.g., gaze coupled with pinch). Those can be uncomfortable and difficult to perform in public or limited spaces, and pose challenges for users who are unable to execute pinch gestures. To address these aspects, we propose a novel, hands-free Gaze+Blink interaction technique that leverages the user's gaze and intentional eye blinks. This technique enables users to perform selections by executing intentional blinks. It facilitates continuous interactions, such as scrolling or drag-and-drop, through eye blinks coupled with head movements. So far, this concept has not been explored for hands-free spatial interaction techniques. We evaluated the performance and user experience (UX) of our Gaze+Blink method with two user studies and compared it with Gaze+Pinch in a realistic user interface setup featuring common menu interaction tasks. Study 1 demonstrated that while Gaze+Blink achieved comparable selection speeds, it was prone to accidental selections resulting from unintentional blinks. In Study 2 we explored an enhanced technique employing a deep learning algorithms for filtering out unintentional blinks.
Understanding and estimating body pose is becoming increasingly important for enhancing user experiences in Virtual Reality (VR). Eye gaze, in particular, plays a critical role in many VR and Augmented Reality (AR) applications. In this paper, we present GaMo, a novel dataset that integrates gaze data and human joint motion capture data during both inter-subject interactions and subject-environment engagements. This dataset provides a comprehensive foundation for advanced pose estimation, enabling the modeling of interactions between users and their surroundings. Based on this dataset, we present the PoseFusionNet model, composed of a Long Short-Term Memory (LSTM) module, and a Transformer Encoder module, focusing on the impact of gaze on body pose estimation. Our model utilizes data from a head-mounted display (HMD), left and right controllers, and 15 previous frames of gaze data to predict the current frame’s pose. Experimental results demonstrate that incorporating gaze data alongside detailed joint information significantly improves pose estimation accuracy.
In current computational models on oculomotor learning 'the' movement vector is adapted in response to targeting errors. However, for saccadic eye movements, learning exhibits a spatially distributive nature, i.e. it transfers to surrounding positions. This adaptation field resembles the topographic maps of visual and motor activity in the brain and suggests that learning does not act on the population vector but already on the level of the 2D population response. Here we present a population-based gain field model for saccade adaptation in which sensorimotor transformations are implemented as error-sensitive gain field maps that modulate the population response of visual and motor signals and of the internal saccade estimate based on corollary discharge (CD). We fit the model to saccades and visual target localizations across adaptation, showing that adaptation and its spatial transfer can be explained by locally distributive learning that operates on visual, motor and CD gain field maps. We show that 1) the scaled locality of the adaptation field is explained by population coding, 2) its radial shape is explained by error encoding in polar-angle coordinates, and 3) its asymmetry is explained by an asymmetric shape of learning rates along the amplitude dimension. Learning exhibits the highest peak rate, the widest spatial extension and a pronounced asymmetry in the motor domain, while in the visual and the internal saccade domain learning appears more localized. Moreover, our results suggest that the CD-based internal saccade representation has a response field that monitors only part of the ongoing saccade changes during learning. Our framework opens the door to study spatial generalization and interference of learning in multiple contexts.
We present a gaze-based augmented reality control interface for electric wheelchairs, addressing the challenges faced by individuals with mobility impairments. The development transitions through three stages: model training with offline evaluation, Virtual Reality (VR) simulations, and physical deployment. First, we trained deep learning models, comparing Transformers and LSTMs, to predict locomotion intentions based on gaze data. While gaze predicts steering intentions well, it sometimes diverges from locomotion goals. To tackle this, we classify gaze movements as either indicative of locomotor intention or not. This novel approach addresses the Midas Touch Problem of gaze. Datasets were collected in controlled VR environments featuring different tasks. We find that data sets with tasks that encouraged diverse navigation and gaze behaviors enable strong generalization. The online VR simulation evaluation phase enabled safe and immersive testing, allowing the assessment of system performance and the integration of feedback for user guidance. Our approach provided smoother navigation control compared to traditional “Where-You-Look-Is-Where-You-Go” methods. Feedback improved user ratings of the system. In the final stage, the system was deployed on a physical wheelchair equipped with an augmented reality (AR) device to provide feedback about the predictions to the user, allowing real-world evaluation. Despite differences in user behavior between VR and physical environments, the system successfully translated gaze inputs into precise and safe navigation commands. Users were able to steer the wheelchair solely using their eyes while simultaneously being able to look at destinations at the side of the path.
Marine mammal vision is often considered to only provide limited information, particularly underwater in low light levels and turbidity. However, when these animals move through turbid water optic flow is elicited. A past study has documented the harbour seal's (Phoca vitulina) ability to perceive deviations from heading from optic flow simulating movement through a volume of turbid water. Here, we asked whether harbour seals are also able to perceive and analyse surface optic flow. Thus, we simulated three optic flow environments and trained three harbour seals to determine the simulated heading. The harbour seals precisely indicated their heading with a mean (±s.d.) accuracy of 4.61±0.56 deg for volume optic flow, 4.96±0.74 deg for surface optic flow mimicking movement over a surface and 3.58±1.12 deg for surface optic flow mimicking movement underneath a surface. We conclude that harbour seals have access to and can thus rely on optic (flow) information whenever there is enough light for vision, thus refuting existing opinions about poor visual guidance in harbour seals or, more generally, in marine mammals. A detailed analysis of optic flow perception in (semi-) aquatic animals is expected to enhance our understanding of optic flow perception and vision in general.
Optic flow, the retinal pattern of motion that is experience during self-motion contains information about one's direction of heading. If the visual scene contains other moving objects optic flow becomes a combination of self-motion and independent object motion. The global pattern due to self-motion is locally confounded, and the locally restricted object flow is the sum of components due to these different causal sources of motion. Nonetheless, humans are able to retrieve information from such flow accurately, including the direction of heading and the scene-relative motion of an object. One way to handle such complex flow is flow parsing, a process speculated to allow the brain's sensitivity to optic flow to separate the causal sources of retinal motion in information due to self-motion and information due to object motion. In a computational model that retrieves object and self-motion information from optic flow, we implemented such a process of causal separation based on heading likelihood maps, whose distributions indicate the consistency of parts of the flow with self-motion alone. The flow parsing allows for concurrent estimation of heading, the detection and localization of and independently moving object, and the estimation of scene-relative motion of that object. We developed a paradigm that allows the model to perform all the different estimations while systematically varying how the object contributes to the flow field. Simulations of that paradigm showed that the model replicates many aspects of human performance, including the dependence of heading estimation on object speed and how different object movements bias that estimation. Regarding object detection and motion estimation, the model's results fit human behavioral data, the latter even for flow of reduced quality. ### Competing Interest Statement The authors have declared no competing interest.
Vision generates a stable representation of space by combining retinal input with internal predictions about the visual consequences of eye movements. We report a type of nonrigid motion that disrupts the connection between eye movements and perception, causing visual instability. This motion is accurately perceived during fixation, but it cannot be pursued. Catch-up saccades are accurately directed to the moving target but the motion stimulus appears to jump in space with each saccade. Our results reveal four major findings about perception and the visuomotor system: (i) Pursuit fails for certain types of motion; (ii) pursuit and catch-up saccades are independently controlled; (iii) prediction of saccade consequences is independent from saccade control; and (iv) the visual stability of moving objects relies on similar motion mechanisms as pursuit.
We present a comprehensive dataset comprising head- and eye-centred video recordings from human participants performing a search task in a variety of Virtual Reality (VR) environments. Using a VR motion platform, participants navigated these environments freely while their eye movements and positional data were captured and stored in CSV format. The dataset spans six distinct environments, including one specifically for calibrating the motion platform, and provides a cumulative playtime of over 10 h for both head- and eye-centred perspectives.The data collection was conducted in naturalistic VR settings, where participants collected virtual coins scattered across diverse landscapes such as grassy fields, dense forests, and an abandoned urban area, each characterized by unique ecological features. This structured and detailed dataset offers substantial reuse potential, particularly for machine learning applications.The richness of the dataset makes it an ideal resource for training models on various tasks, including the prediction and analysis of visual search behaviour, eye movement and navigation strategies within VR environments. Researchers can leverage this extensive dataset to develop and refine algorithms requiring comprehensive and annotated video and positional data. By providing a well-organized and detailed dataset, it serves as an invaluable resource for advancing machine learning research in VR and fostering the development of innovative VR technologies.