
With the rapid growth of digital textual content, users face increasing challenges in rediscovering information which they have encountered in the past. Traditional search engines lack the ability to prioritize content truly viewed by users over merely visible content. We propose a method that captures both text passages and corresponding gaze data while reading, storing them in a searchable knowledge base. We then explore boosting strategies to improve text retrieval by prioritizing passages that users actually viewed, improving the personalization and relevance of search results. To this end, we repurposed the recent g-Rel-READER dataset to evaluate various gaze-based boosting techniques and address the research gap caused by the lack of combined text, gaze, and relevance data. The evaluation demonstrates the potential of gaze data to serve as a boosting criterion for search, with mean average precision (MAP) increased by over 33% over a purely text-based retrieval.
Expanding an object’s region in motor space (selection area) using Voronoi tessellation can facilitate the selection of small targets in gaze-based interfaces. While visual feedforward, such as visualizing the selection area, has been shown to improve selection performance in hand-based interfaces, its influence on gaze-based interfaces remains unclear. To address this, we conducted two user studies to examine how visual feedforward affects fixation behavior and selection performance when the centers of a target and its selection area differ. The results reveal that visualizing either the selection area or its center effectively guides users’ gaze toward the center of the selection area, thereby improving selection accuracy. Furthermore, displaying a gaze cursor reduces errors by enabling users to adjust their fixation, even though it does not directly guide gaze toward the center of the selection area.
We present the Eye2Head Gesture (E2H), a gaze-based technique that integrates saccade gestures with eye-head coordination for target selection in virtual reality. It involves a back-and-forth saccadic gesture, aligning the gaze with the head forward direction and then returning to the target to complete the selection. The first study identified the parameters for implementing the E2H technique, including visualization of the relay zone (i.e., the area anchor to the head’s forward direction), relay zone size, and saccade gesture duration. Four variants of the E2H were developed, each differing in the shape and timing of the relay zone. The second study evaluated four E2H variants against Dwell and Eye&Head techniques, revealing that Region-shaped variant of E2H outperformed others in speed, error, and fatigue. This work contributes a gaze-based technique that harnesses eye-head coordination, using saccadic gestures to enhance the efficiency of target selection, particularly for those at large angular distances.
Recent advances in extended reality (XR) have enabled seamless access to immersive services. However, user authentication in these environments typically relies on conventional methods, such as PIN entry, which remains vulnerable to unauthorized use. This paper investigates gaze behavior as an implicit second authentication factor in XR. Using a Meta Quest Pro headset, participants performed legitimate and impostor login attempts while their gaze data were recorded. Temporal, spatial, and oculomotor metrics revealed distinctive and reproducible gaze dynamics between user roles. Tree-based machine learning models, particularly XGBoost, reliably distinguished legitimate from impostor sessions under user-independent validation (AUC =.84). Calibrating model thresholds further enabled adaptive balancing between security and usability. These findings demonstrate that gaze dynamics can unobtrusively enhance PIN authentication, introducing an adaptive layer that aligns authentication sensitivity with situational risk and user needs in immersive environments.
This eye-tracking study examines saccadic latency during subtitled video processing, i.e., the time to shift attention from the video image to the subtitle, based on subtitle speed (145 vs. 180 words per minute–wpm) and number of lines (one vs. two). We employed a mixed-design experiment with deaf (N=20) and hearing (N=20) participants, making use of a linear mixed-model regression for analysis. While the group difference was not statistically significant, the results revealed a significant interaction between subtitle speed and number of lines: at fast speeds (180 wpm), subtitle line count did not affect latency, yet at slow speeds (145 wpm), two-line subtitles incurred a significantly shorter latency. These findings indicate that two-line subtitles, especially at slower speeds, reduce processing latency, which may preserve more time for actual text comprehension.
This study investigates gaze behavior as a potential biomarker for individuals with Social Anxiety Disorder during immersive virtual reality socio-evaluative tasks. Thirty participants diagnosed with SAD and thirty matched healthy controls completed three VR tasks while eye movements were recorded. Gaze features were analyzed using a neural predictive model with PAC-Bayes estimating generalization and uncertainty. The model predicted group membership with accuracies of 87 % and 93 % for performance and social tasks, with certified PAC-Bayes lower bounds accuracy of 78 % and 84 %, respectively. Variations in social demand across tasks were associated with corresponding differences in predictive gaze features. Gender bias was investigated, but was not detected. Model predictions correlated strongly with standardized psychometric measures. These findings indicate that VR-based eye tracking, combined with machine learning models, provides a scalable, explainable, and objective method to assess SAD.
Hands-free interaction is essential when users’ hands are occupied with primary tasks. A key challenge is unifying precise, UI-independent selection and continuous dragging without explicit mode switching. This paper introduces two novel techniques: Aligner (spatial eye-head alignment) and Nodder (gestural decomposition of a nod); and evaluates them against an established Winker technique, a gestural clutch using single-eye closure. A user study with controlled Fitts’ law and dragging tasks, alongside a practical application, evaluated the techniques. Results show that Aligner offers superior selection precision but creates a visuomotor conflict during dragging; Nodder optimizes movement efficiency for large-amplitude targets despite higher physical effort; and Winker provides the fastest performance but is susceptible to accidental activation. These findings inform the design of future hands-free systems by highlighting the context-dependent nature of this interaction challenge.
Hands-free computer interaction provides an important alternative for users with motor impairments who struggle with conventional input devices. In this paper, we present LookAHead, a hybrid gaze and head-based interaction system that enables natural and hands-free computer control. The system uses gaze estimation to provide a coarse target location, refined by head movements for fast and convenient pointer control. To improve model's accuracy, facial landmarks are used not only for cursor control, 3D head pose estimation, and input normalization, but also to enhance the gaze estimation model. A revised gaze normalization process is proposed to improve stability under varying facial expressions. LookAHead operates without explicit calibration by employing an auto-calibration strategy that continuously adapts to users’ gaze. Experimental results demonstrate that the proposed method enables fast and convenient control, while reducing head movement threefold compared to conventional head-mouse, making it a practical step toward more accessible hands-free computing.
Microsurgery requires visuomotor coordination under high magnification, where depth and scale are distorted. Understanding how surgeons adapt eye-hand coordination under the surgical microscope is key to advancing microsurgical education. We investigated gaze behavior and performance differences between expert plastic surgeons and novices during microsurgical suturing. An ocular-mounted eye-tracking system magnetically attached to the microscope preserved naturalistic gaze patterns while maintaining authentic spatial and visual conditions. Suturing was decomposed into phases: needle alignment and tissue piercing, pull-through, and knot tying, for phase-specific analysis. Performance was evaluated using the UWOMSA, and perceived workload was measured with the SURG-TLX. Experts completed the overall task faster than novices, exhibited shorter mean fixations overall, and showed the largest event-level advantage during knot tying. Findings indicate that visual strategies in microsurgery are phase-dependent and expertise-specific. In conclusion, ocular-mounted eye-tracking offers a valuable framework for studying microsurgical visuomotor control and developing objective, data-driven approach for microsurgical training.
With the growing presence of AI-driven avatars in everyday applications, understanding how users perceive and respond to such agents becomes increasingly important. A key open question is whether gaze awareness, an avatar’s ability to react to the gaze of a user, contributes to more natural and socially engaging interaction, particularly on conventional 2D screens. We present a real-time system that integrates live speech recognition, LLM-generated responses, and gaze feedback to enable naturalistic unscripted conversations with a virtual avatar. We conducted a within-subjects study (N = 17) comparing a gaze-aware avatar with a nonreactive baseline. The results of the eye-tracking show that gaze awareness significantly altered user eye movement patterns, increasing mutual gaze episodes, and generating subconscious gaze regulation. Our findings demonstrate that realistic gaze behavior affects user interaction on a subconscious level, even outside immersive environments. We discuss implications for designing future conversational AI systems that foster natural social presence.
Smart eyewear is emerging as an always-on platform capable of perceiving the environment and inferring intent through eye movements, making pupil tracking essential for personalized interaction. However, reliable tracking on wearable hardware remains challenging due to strict power limits and scarce annotated data. Event-based cameras offer a low-power, microsecond-latency solution, but labeled recordings still remain limited. We address this issue with a training framework that combines limited annotated real data with synthetic events and unlabeled real recordings, learning event-based pupil trackers with strong real-world generalization. We pair the U2Eyes tool with the v2e event camera simulator to generate realistic event streams, showing that networks trained on these events exhibit smaller sim-to-real gaps than networks trained on synthetic images. Moreover, our training procedure further bridges this gap, enabling our models to outperform networks trained exclusively on real data across all benchmarks, advancing toward more robust event-based eye tracking on wearable platforms.
We introduce GazeMorph, a 3D UNet-style neural network that aligns measured fixation points to on-screen content to reduce eye-tracking measurement error in realistic settings. The model takes as input a time-stacked sequence of fixation maps and a content structure, and predicts a dense displacement field constrained by spatial and temporal smoothness that enforces coherent local motion and by a soft fold penalty discouraging folded or self-overlapping deformations. Trained on ZuCo-based pages where fixation maps are elastically distorted to emulate spatial noise of eye-trackers, GazeMorph amortizes fixation correction and consistently reduces the average fixation-position error across all noise levels. These results demonstrate that on-content alignment serves as an effective post-processing step for robust gaze correction, complementing hardware calibration and improving the reliability of eye-tracking data in practical applications.
This paper presents PyEtSimul, an open-source Python-based framework for simulating video-based eye trackers by generating synthetic eye features through geometric modeling. The framework allows flexible positioning of eyes, cameras, and light sources in 3D space, with controlled variation of eye anatomical features and camera properties. PyEtSimul generalizes corneal modeling by representing the cornea as a conic surface rather than the common sphere. It also supports non-circular pupil shapes, size-dependent pupil decentration, eyelid occlusion, and camera lens distortion. It supports systematic data generation and principled comparison of gaze estimation algorithms across calibrated and uncalibrated settings. These features enable analyses not possible with other available simulators. PyEtSimul facilitates controlled experiments with known parameters often latent in normal settings, enabling reproducible benchmarking and systematic exploration of hardware designs. By generating fully synthetic data, PyEtSimul removes privacy concerns and the need for costly hardware, making it practical for both educational and research applications.
Watching subtitled videos in a foreign language demands sustained visual attention, which can put viewers at risk of missing content due to distraction, such as checking notifications. In this work, we introduced a gaze-aware video player that adapts playback to support attention recovery. We evaluated three gaze-aware techniques: adaptive pausing, stacked subtitles, and audio language switching (dubbing). In a comparative study with 24 participants, we evaluated these techniques against a standard video player with subtitles. While adaptive pausing improved task performance and reduced distractions, stacked subtitles helped recover reading but occasionally slowed faster readers. The benefit of dubbing was limited, resulting in additional cognitive load during the process. Ultimately, all gaze-aware interventions outperformed the standard video player. This work highlights gaze-adaptive systems that seamlessly support attention recovery into everyday viewing experiences.
Cognitive load affects learning and task performance; specifically, increased cognitive load hinders an individual’s ability to process information. In augmented reality (AR) interfaces, distracting notifications can also heighten cognitive load. Integrated gaze tracking offers a non-intrusive way to monitor cognitive states and provides the opportunity to predict and adapt to changes in cognitive load. In this paper, we demonstrate how cognitive load prediction models can leverage built-in gaze-tracking data in AR to accurately predict cognitive load during search tasks. We collected gaze data from participants under both cognitively overloaded and non-overloaded states and analyzed gaze feature signatures to identify load-dependent patterns. We compared individual and group-trained models for predictive performance and generalizability. We initially used logistic regression, then tested tree-based ensemble models to improve performance. The best-performing XGBoost group model achieves a test AUC-ROC of 0.85. This work demonstrates robust cognitive load monitoring for AR tasks using built-in eye-tracking measurements.
Pupil size is a key eye-based indicator of mental processing and internal states. Nonetheless, pupil size is also influenced by light, causing challenges for monitoring internal states with eye tracking in dynamic and naturalistic settings. We investigated how convolution-based modeling can capture the influence of luminance and emotional arousal on pupil size, and conversely, how pupil dynamics inform arousal prediction. In a lab study, we analyzed data collected from 19 participants who watched travel-themed videos with different arousal levels and their pixelated counterparts. We present and evaluate a convolution-based approach for pupil size modeling that integrates low-level visual features and higher-order emotional factors, as previous work has shown effectiveness of such methods in modeling pupil light reflex. Our results show that incorporating luminance and contrast indeed enhances arousal prediction from pupil data, although performance varies due to different user behavior and stimuli.
Autism Spectrum Disorder (ASD) is a neurodevelopmental condition marked by impairments in social interaction and delayed language acquisition. Early and accurate identification is crucial for timely interventions that support cognitive and social development. Motivated by the subjectivity of traditional behavior-based assessments, computational methodologies offer more objective and cost-effective alternatives. Among these, eye-tracking stands out for capturing subtle attentional and perceptual patterns. This paper investigates the use of eye-tracking data for automatic ASD detection in children during audio-visual storytelling interactions, emphasizing traditional yet explainable machine learning methods. Although performance remains modest, our analyses reveal that fixation duration and revisit patterns to facial regions may serve as potential biomarkers. Further analyses highlight the impact of stimulus modality, suggesting that the inclusion of visual speech cues provides valuable discriminative information. These findings have the potential to support and guide the work of psychologists in the assessment of ASD within speech comprehension contexts.
This study presents the first data-driven review of research on the privacy of eye tracking, covering gaze, iris, and eye image data across immersive, mobile, and clinical contexts. The analysis examined 78 papers published between 2015 and 2025 using ensemble topic modeling with non-negative matrix factorization. Nine topics emerged, showing how normative and technical approaches address privacy across different stages of eye tracking data processing. Findings reveal that normative works emphasize inference of identity and personal traits as an unresolved risk, whereas technical studies treat privacy as a computational problem and mitigate risks within specific contexts. However, protections are often constrained by utility requirements, meaning eye tracking data cannot be sufficiently safeguarded without impairing function, leaving enough detail to enable inference. This persistent vulnerability suggests that eye tracking data may warrant regulatory protection comparable to other sensitive categories, requiring stronger governance and safeguards across its collection and use.
Gaze can be directed at will, or guided by objects that draw our attention. We propose classifying gaze shifts as either directed or guided, since these two forms of attention have different implications for HCI. Directed attention may serve as stronger indicator of users’ planning and intent, whereas guided attention reflects interface efficacy in guiding information acquisition. We introduce a method based on eye and head movement features during gaze shifts, using data collected in virtual reality to train and evaluate a machine learning model, which we then validate in application to visual search. Our results show that this classification is both feasible and practical, extending established uses of eye tracking in HCI. This is significant because it enables a new level of analysis of visual attention.