This paper presents PyEtSimul, an open-source Python-based framework for simulating video-based eye trackers by generating synthetic eye features through geometric modeling. The framework allows flexible positioning of eyes, cameras, and light sources in 3D space, with controlled variation of eye anatomical features and camera properties. PyEtSimul generalizes corneal modeling by representing the cornea as a conic surface rather than the common sphere. It also supports non-circular pupil shapes, size-dependent pupil decentration, eyelid occlusion, and camera lens distortion. It supports systematic data generation and principled comparison of gaze estimation algorithms across calibrated and uncalibrated settings. These features enable analyses not possible with other available simulators. PyEtSimul facilitates controlled experiments with known parameters often latent in normal settings, enabling reproducible benchmarking and systematic exploration of hardware designs. By generating fully synthetic data, PyEtSimul removes privacy concerns and the need for costly hardware, making it practical for both educational and research applications.
Previous work has reported that vision foundation models show promising zero-shot performance in eye image segmentation. Here we examine whether the latest iteration of the Segment Anything Model, SAM3, offers better eye image segmentation performance than SAM2, and explore the performance of its new concept (text) prompting mode. Eye image segmentation performance was evaluated using diverse datasets encompassing both high-resolution high-quality videos from a lab environment and the TEyeD dataset consisting of challenging eye videos acquired in the wild. Results show that in most cases SAM3 with either visual or concept prompts did not perform better than SAM2, for both lab and in-the-wild datasets. Since SAM2 not only performed better but was also faster, we conclude that SAM2 remains the best option for eye image segmentation. We provide our adaptation of SAM3's codebase that allows processing videos of arbitrary duration.
Understanding the quality of eye-tracking recordings, often characterized using accuracy, precision, and data loss, is crucial for the interpretation of eye tracking data. Eye-tracking data quality can furthermore place fundamental limits on what studies can be conducted with an eye tracker, and one may be required to report eye-tracking data quality when publishing a study. However, how does one determine the quality of eye-tracking data? This article provides an overview of operationalizations of accuracy, precision, and data loss and practical advice for determining eye-tracking data quality. Furthermore, the programming code for calculating various quality metrics for a segment of eye-tracking data is provided in MATLAB, Python, and R. Also provided is ETDQualitizer, a tool designed to enable anyone to easily determine the data quality of their recordings. We provide a version that is browser-based ( https://dcnieho.github.io/ETDQualitizer ) and enables determining eye-tracking data quality without installation or programming, while ensuring data privacy by running entirely locally. ETDQualitizer is further provided as a MATLAB, Python, and R library ( https://github.com/dcnieho/ETDQualitizer ) that can be integrated in one’s analysis scripts. We hope that this article enables any researcher to determine, critically evaluate, and report on eye-tracking data quality, and that it spurs researchers to adopt a data quality perspective in all their future eye-tracking studies.
Blink characteristics such as duration, amplitude, and eyelid velocity are widely used indicators of cognitive and physiological states. While early magnetic search coil studies suggested that vertical eye orientation relative to the head influences blink measurements, subsequent research has largely ignored this factor. No studies have investigated whether vertical eye orientation effects replicate in modern video-based methods or whether different video-based blink-detection approaches show similar sensitivities to changes in vertical eye orientation. In this study, we investigated how vertical eye orientation affects blink parameters estimated using both pupil-based and eyelid-based detection. We recorded pupil diameter and estimated eye openness from video data as seventeen participants performed voluntary blinks from three vertical eye orientations while keeping their heads stationary. Vertical eye orientation systematically influenced all measured blink parameters. Eye openness at blink onset and closing amplitude decreased with downward eye orientation. With more downward eye orientation, closing velocity increased, whereas opening velocity decreased. Crucially, pupil-based measurements of blink duration showed much larger vertical eye orientation effects than measurements of eye openness (32% vs. 8% increase of blink duration from upward to downward eye orientation), though eyelid-based estimates are sensitive to how blink onset and offset are derived. These results show that the vertical eye orientation is a systematic confounding factor in video-based blink measurement, with the measurement method influencing the magnitude of observed effects. The findings have important implications for studies investigating blink characteristics where vertical eye orientation varies, and we conclude with practical recommendations for study design and reporting.
This study investigated whether saccadic search exists on a micro scale. Using a high-speed, confocal retinal eye tracker (FET), participants searched for tiny targets within dense displays subtending less than 4 deg2. The participants exhibited goal-directed, exploratory scanning behavior with tiny saccades (median amplitude < 0.5°); saccade amplitude scaled with inter-element distance, fixation duration increased in denser displays, and consecutive saccades followed typical directional dependencies (saccadic momentum and facilitation of return). Search performance and scanpaths resembled those found in large-scale visual search tasks, suggesting functional equivalence across spatial scales. These findings support the view that microsaccades serve the same perceptual and attentional roles as larger saccades. We argue that the distinction between microsaccades and saccades is unnecessary when describing natural, task-driven visual scanning behavior.
Wearable eye tracking in unconstrained settings is often compromised by slippage, yet the direct relationship between physical frame displacement and gaze error remains unexplored. This study addresses this gap combining eye tracking and motion capture to evaluate the slippage robustness of Pupil Labs Neon, Tobii Pro Glasses 3 and ViewPointSystems Lite eye trackers. Data were collected as twelve participants performed tasks involving facial expressions, induced glasses motion and locomotion. Results indicate that the Tobii Pro Glasses 3 are the most slippage-robust, maintaining an average gaze error below 2.5° regardless of movement. Conversely, the ViewPointSystems Lite exhibited large errors scaling with displacement (up to 29.3° on average) during induced motion tasks. The Pupil Labs Neon demonstrated resilience against large errors (<3.5° on average). Our findings enable researchers to anticipate the gaze error in unconstrained environments, and push manufacturers to provide realistic specifications of slippage instead of claims of “slippage-robust“ eye tracking.
Researchers use area of interest (AOI) analyses to interpret eye-tracking data. This article addresses four key aspects of AOI use: 1) how to report AOIs to support replicable analyses, 2) how to interpret AOI-related statistics, 3) methods for generating both static and dynamic AOIs, and 4) recent developments and future directions in AOI use. The article underscores the importance of aligning AOI design with the study’s conceptual and methodological foundations. It argues that critical decisions, such as the size, shape, and placement of AOIs, should be made early in the experimental design process and should involve eye-tracking data quality, the research question, participant tasks, and the nature of the visual stimulus. It also evaluates recent advances in AOI automation, outlining both their benefits and limitations. The article’s main message is that researchers should plan AOIs carefully and explain their choices openly so others can replicate the work.
Previous research suggests a pattern of gaze avoidance in East Asian compared with Western cultures. Yet, recent eye-tracking studies of face-to-face conversation do not corroborate this. More generally, differences in nonverbal communication and analytic versus holistic strategies have been described between East Asian and Western cultures. Using wearable eye-tracking technology and an automated gaze-processing pipeline, we investigated cross-cultural differences in gaze behavior during unstructured conversation and collaborative interactions from an information-gathering and information-signaling perspective. We compared Japanese and Dutch individuals on gaze to faces, gaze-gesture coupling, and gaze-action coupling. Japanese participants consistently looked less at faces than their Dutch counterparts in all interactive scenarios. Additionally, Japanese individuals displayed fewer pointing gestures and kept their hands under the table longer. Although gaze coupling with manual actions and pointing gestures was similar across both groups, longer-term gaze-action patterns varied, reflecting potential differences in cultural strategies (holistic vs. analytic) and error orientation styles. These findings suggest that while visuomotor coordination is consistent, extended patterns of gaze in the context of collaboration diverge based on cultural context. Our study underscores the need to assess gaze behavior within the interaction rather than in isolation, integrating visuomotor behavior, nonverbal communication and cultural context. Additionally, our findings may aid in developing individually- and culturally-sensitive anthropomorphic virtual avatars and social robots.
In many tasks, participants are instructed to fixate a target. While maintaining fixation, the eyes nonetheless make small fixational eye movements, such as microsaccades and drift. Previous work has examined the effect of fixation point design on fixation stability and the amount and spatial extent of fixational eye movements. However, much of this work used video-based eye trackers, which have insufficient resolution and suffer from artefacts that make them unsuitable for this topic of study. Here, we therefore use a retinal eye tracker, which offers superior resolution and does not suffer from the same artifacts to reexamine what fixation point design minimizes fixational eye movements. Participants were shown five fixation targets in two target polarity conditions, while the overall spatial spread of their gaze position during fixation, as well as their microsaccades and fixational drift, were examined. We found that gaze was more stable for white-on-black than black-on-grey fixation targets. Gaze was also more stable (lower spatial spread, microsaccade, and drift displacement) for fixation targets with a small central feature but these targets also yielded higher microsaccade rates than larger fixation targets without such a small central feature. In conclusion, there is not a single best fixation target that minimizes all aspects of fixational eye movements. Instead, if one wishes to optimize for minimal spatial spread of the gaze position, microsaccade or drift displacements, we recommend using a target with a small central feature. If one instead wishes to optimize for the lowest microsaccade rate, we recommend using a larger target without a small central feature.
The problem: wearable eye trackers deliver eye-tracking data on a scene video that is acquired by a camera affixed to the participant’s head. Analyzing and interpreting such head-centered data is difficult and laborious manual work. Automated methods to map eye-tracking data to a world-centered reference frame (e.g., screens and tabletops) are available. These methods usually make use of fiducial markers. However, such mapping methods may be difficult to implement, expensive, and eye tracker-specific. The solution: here we present gazeMapper, an open-source tool for automated mapping and processing of eye-tracking data. gazeMapper can: (1) Transform head-centered data to planes in the world, (2) synchronize recordings from multiple participants, (3) determine data quality measures, e.g., accuracy and precision. gazeMapper comes with a GUI application (Windows, macOS, and Linux) and supports 11 different wearable eye trackers from AdHawk, Meta, Pupil, SeeTrue, SMI, Tobii, and Viewpointsystem. It is also possible to sidestep the GUI and use gazeMapper as a Python library directly.
The eyeball is not rigid and deforms during saccades. As a consequence, the saccade waveform recorded by an eye tracker may depend on which structure of the eye is used to estimate eyeball rotation. Here, we systematically describe and compare signals co-recorded from the retina, the cornea (corneal reflection, CR), the pupil, and the lens (fourth Purkinje reflection, P4) during saccades. We found that several commonly used parameters for saccade characterization differ systematically across the signals. For instance, saccades in the retinal signal had earlier onsets compared to saccades in the pupil and the P4 signals. The retinal signal had the smallest saccade amplitude and reached the peak saccade velocity earlier compared to the other signals. At the end of saccades, the retinal signal came to a stop faster than the other signals. We discuss possible explanations that may account for the relationship between the retinal signal and the other signals.
Deep learning methods have significantly advanced the field of gaze estimation, yet the development of these algorithms is often hindered by a lack of appropriate publicly accessible training datasets. Moreover, models trained on the few available datasets often fail to generalize to new datasets due to both discrepancies in hardware and biological diversity among subjects. To mitigate these challenges, the research community has frequently turned to synthetic datasets, although this approach also has drawbacks, such as the computational resource and labor-intensive nature of creating photorealistic representations of eye images to be used as training data. In response, we introduce “Light Eyes” (LEyes), a novel framework that diverges from traditional photorealistic methods by utilizing simple synthetic image generators to train neural networks for detecting key image features like pupils and corneal reflections, diverging from traditional photorealistic approaches. LEyes facilitates the generation of synthetic data on the fly that is adaptable to any recording device and enhances the efficiency of training neural networks for a wide range of gaze-estimation tasks. Presented evaluations show that LEyes, in many cases, outperforms existing methods in accurately identifying and localizing pupils and corneal reflections across diverse datasets. Additionally, models trained using LEyes data outperform standard eye trackers while employing more cost-effective hardware, offering a promising avenue to overcome the current limitations in gaze estimation technology.
Changes in pupil size can lead to apparent gaze shifts in data recorded with video-based eye trackers in the absence of physical eye rotation. This is known as the pupil-size artifact (PSA). While the PSA is widely reported in desktop eye trackers, it is unknown whether and to what extent it occurs in head-mounted eye trackers. In this paper, we examined the effects of pupil size variations on eye-tracking data quality in four head-mounted eye trackers: the Pupil Core, the Pupil Neon, the SMI ETG 2w, and the Tobii Pro Glasses 2, in addition to a widely used desktop eye tracker, the SR Research EyeLink 1000 Plus. Participants viewed a central target on a monitor while we systematically varied the screen brightness to induce controlled pupil size changes. All head-mounted eye trackers exhibited PSA, with apparent gaze shifts ranging from 0.94 ^∘ for the Pupil Neon to 3.46 ^∘ for the Pupil Core. Except for the Pupil Neon, all eye trackers exhibited a significant change in accuracy due to pupil size variations. Precision measures showed device-specific effects of pupil size changes, with some eye trackers performing better in the bright condition and others in the dark condition. These findings demonstrated that, just like desktop eye trackers, head-mounted video-based eye trackers exhibited PSA.
We explore the transformative potential of SAM 2, a vision foundation model, in advancing gaze estimation. SAM 2 addresses key challenges in gaze estimation by significantly reducing annotation time, simplifying deployment, and enhancing segmentation accuracy. Utilizing its zero-shot capabilities with minimal user input—a single click per video—we tested SAM 2 on over 14 million eye images from a diverse range of datasets, including the EDS challenge datasets and Labelled Pupils in the Wild. This is the first application of SAM 2 to the gaze estimation domain. Remarkably, SAM 2 matches the performance of domain-specific models in pupil segmentation, achieving competitive mIOU scores of up to 93% without fine-tuning. We argue that SAM 2 achieves the sought-after standard of domain generalization, with consistent mIOU scores (89.71%-93.74%) across diverse datasets, from virtual reality to "gaze-in-the-wild" scenarios. We provide our code and segmentation masks for these datasets to promote further research.
Researchers using eye tracking are heavily dependent on software and hardware tools to perform their studies, from recording eye tracking data and visualizing it, to processing and analyzing it. This article provides an overview of available tools for research using eye trackers and discusses considerations to make when choosing which tools to adopt for one's study.
Extensive studies have shown that humans process faces holistically, considering not only individual features but also the relationships among them. Knowing where humans and dogs fixate first and the longest when they view faces is highly informative, because the locations can be used to evaluate whether they use a holistic face processing strategy or not. However, the conclusions reported by previous eye-tracking studies appear inconclusive. To address this, we conducted an experiment with humans and dogs, employing experimental settings and analysis methods that can enable direct cross-species comparisons. Our findings reveal that humans, unlike dogs, preferentially fixated on the central region, surrounded by the inner facial features, for both human and dog faces. This pattern was consistent for initial and sustained fixations over seven seconds, indicating a clear tendency towards holistic processing. Although dogs did not show an initial preference for what to look at, their later fixations may suggest holistic processing when viewing faces of their own species. We discuss various potential factors influencing species differences in our results, as well as differences compared to the results of previous studies.
We explore whether SAM2, a vision foundation model, can be used for accurate localization of eye image features that are used in labbased eye tracking: corneal reflections (CRs), the pupil, and the iris. We prompted SAM2 via a typical hand annotation process that consisted of clicking on the pupil, CR, iris and sclera for only one image per participant. SAM2 was found to support better spatial precision in the resulting gaze signals for the pupil (>44% lower RMS-S2S), but not the CR and iris, than traditional image-processing methods or two state-of-the-art deep-learning tools. Providing more frames with prompts to initialize SAM2 did not improve performance. We conclude that SAM2's powerful zero-shot segmentation capabilities provide an interesting new avenue to explore in high-resolution labbased eye tracking. We provide our adaptation of SAM2's codebase that allows segmenting videos of arbitrary duration and prepending arbitrary prompting frames.