
Eye tracking technologies have become increasingly accessible through webcam-based solutions such as WebGazer, enabling large-scale and low-cost experimentation. However, their performance remains sensitive to several experimental and device-related parameters. This study presents a controlled experimental framework developed to systematically evaluate WebGazer under configurable experimental conditions. Within this framework, we design and implement a controlled experimental protocol to investigate the effect of key parameters, including (i) joint calibration/evaluation point-density configuration, (ii) evaluation-point distribution, (iii) target presentation duration, (iv) stimulus color, and (v) stimulus (target) size. WebGazer’s performance is assessed using both device-pixel spatial error metrics and angular error (delta), providing complementary measures of gaze accuracy. Results from 20 participants show that denser point configurations were associated with lower mean spatial error, from 1383 to 1009 device pixels, and lower mean angular error, from 13.556 to 10.322 degrees, within the tested conditions. Across the aggregated condition summaries, the joint point-density configuration showed the widest spatial-error range, whereas target presentation duration showed the widest angular-error range. Taken together, these two factors displayed the clearest descriptive differences, followed by evaluation-point distribution. Pseudo-random layouts were associated with moderately higher errors than fixed layouts. In contrast, visual properties such as color and size showed comparatively smaller and less consistent differences, with stimulus size displaying a non-monotonic pattern. Overall, temporal and calibration-related settings showed the clearest descriptive associations with WebGazer performance, while visual design choices showed smaller differences. This study provides a practical quantitative framework for systematically assessing webcam-based eye tracking systems under realistic usage conditions.
Generative artificial intelligence increasingly produces text, images, feedback, summaries, advertisements, synthetic faces, and audiovisual content evaluated alongside human-produced material. This systematic review synthesized comparative eye-tracking evidence on visual attention, cognitive processing, and engagement with AI-generated content. Searches of Scopus, Web of Science, and PubMed yielded 896 records; 23 studies met the eligibility criteria and contributed 778 participants in the review-relevant eye-tracking components. The evidence covered textual, static visual, audiovisual, and interactive outputs. Across heterogeneous designs and tasks, no modality-independent gaze pattern emerged. AI-generated material sometimes attracted more focal inspection, sometimes received less task-relevant attention, and often redistributed gaze across interface elements. Where supported by task characteristics or complementary outcomes, longer viewing was more often associated with processing difficulty, uncertainty, or checking than with preference. Generated summaries supported learning in some settings, whereas realistic synthetic media remained difficult to identify despite focused inspection. Methodological appraisal identified recurrent limitations in sampling, stimulus matching, confounder control, eye-tracking reporting, and documentation of model versions, prompts, generation settings, and output selection. Observed gaze differences were context-dependent and varied with modality, task, comparator, source belief, expertise, output quality, and measurement choices. Standardized reporting and stronger links between gaze and functional outcomes are needed for cumulative inference.
The Concealed Information Test (CIT) detects whether an individual recognizes crime-relevant information, and eye tracking has been proposed as a non-invasive method for detecting this recognition. It remains unclear, however, whether individuals who are motivated to conceal their knowledge can voluntarily control their gaze and thereby undermine the test’s validity. We developed a novel free-viewing, parallel-presentation eye-movement-based CIT and administered it to 75 participants randomly assigned to three groups: Innocent, Informed Innocent, and Guilty. For each participant, twelve eye-tracking measures were calculated as differential (crime-relevant minus control) scores and analyzed using four supervised classifiers under a participant-wise, leave-two-out cross-validation, with accuracy reported at both the trial and participant levels. Cumulative dwell- and duration-based measures, together with glance and fixation counts, carried most of the discriminative signal. The classifiers demonstrated high accuracy when distinguishing between the naive and knowledgeable but non-concealing individuals, as well as between knowledgeable non-concealing individuals and culpable individuals. However, classification performance fell to chance-level probability when distinguishing naive and culpable individuals. Thus, participants who were motivated to avoid detection markedly reduced their distinctive eye-movement signature with voluntary gaze control, becoming indistinguishable from truly naive individuals. We discuss methodological implications for designing eye-movement-based CITs that exhibit greater resilience against voluntary gaze modulation.
It is widely proposed that perceptual expertise is characterized by an increased reliance on peripheral visual information during search, as described by Holistic Visual Processing (HVP) theory. HVP predicts that experts make longer saccades to construct a holistic representation of the visual scene and facilitate target detection. Our previous work challenged this prediction by showing that expert architects made shorter, rather than longer, saccades during visual search. The present study investigated how this atypical search strategy develops by comparing expert architects, intermediate architects, and naïve observers performing the same visual search task. Intermediate architects exhibited shorter saccades than naïve observers, but not to the extent observed in experts, while demonstrating search efficiency intermediate to the two groups. Analyses of the temporal organization of search further revealed that intermediate architects had already adopted the overall structure of the expert search strategy despite not yet exhibiting expert-level reductions in saccadic amplitude. These findings suggest that perceptual expertise in architecture develops through a layered progression in which the organization of visual search precedes the full refinement of eye movement behavior. More broadly, these findings suggest that eye movement research should interpret gaze metrics such as saccadic amplitude within the context of domain-specific task demands and developmental stage, while considering the temporal organization of search as an additional window into the emergence of perceptual expertise.
Background: This prospective observational study assessed acute changes in distance ocular deviation after brief smartphone viewing in young adults without manifest strabismus and identified factors associated with the magnitude of change. Methods: Fifty-nine healthy young adults without strabismus (aged 20–30 years; mean ± standard deviation, 23.2 ± 2.5 years) wore a head-mounted eye-tracking system and viewed a smartphone for 30 min. Ocular deviation angles at near and distance fixation were measured immediately before and after smartphone viewing. During smartphone viewing, eye-movement parameters and smartphone distance were continuously recorded. Results: Before smartphone viewing, the mean deviation angles were −6.2 ± 6.8Δ and −2.0 ± 2.8Δ at near and distance fixation, respectively (negative values indicate exodeviation). After smartphone viewing, deviation angles shifted toward less exodeviation, measuring −4.3 ± 6.6Δ and −1.1 ± 2.4Δ at near and distance fixation, respectively. The magnitude of exodeviation reduction was negatively correlated with baseline ocular deviation, indicating that individuals with larger pre-viewing exodeviation exhibited greater alignment shifts after smartphone viewing. Eye-movement analysis revealed that gaze movements of <30°/s accounted for 57.8% of detected gaze-movement events. Conclusions: Thirty minutes of smartphone viewing was followed by a reduction in exodeviation at near and distance fixation in young adults without manifest strabismus. Greater baseline exodeviation was associated with a larger reduction, and low-velocity gaze movements predominated.
As generative AI chatbots become a primary information channel, users increasingly accept answers without verification, and citations can raise trust even when sources are irrelevant or fabricated. How source-attribution visualization shapes the visual preconditions of verification remains unknown: users can notice, read, or compare a source without clicking. This within-subjects eye-tracking study (N = 23; 92 trials) evaluated four attribution visualizations abstracted from commercial AI chatbots and rendered as simulated screens: inline component (sentence-end chips), card list (cards above the answer), side panel (adjacent panel), and raw hyperlink (bare URLs), combining gaze metrics, surveys, and interviews. Repeated-measures ANOVAs revealed strong layout effects on source discoverability and engagement, largely robust to sensitivity checks (the panel’s discovery latency was order-sensitive): the card list was discovered almost immediately, with the raw hyperlink last. Yet no self-reported measure differed detectably. The most frequently nominated format, the inline component, attracted about half the dwell time of the stand-alone formats, whose prolonged fixations suggested citation-to-text mapping cost rather than genuine engagement. This attention–preference gap means both must be measured jointly. We contribute a four-layout gaze-based comparison, a reproducible participant-level analysis workflow, and three design principles (pre-click identifiability, sentence-level claim–source mapping, and in situ preview) within a proposed two-stage attribution architecture.
This study investigates how strategic environments shape a decision maker’s reasoning processes by incorporating visual attention and cognitive engagement. Using eye-tracking methods within experimental guessing games featuring interior equilibria, we analyze how strategic complementarity and substitutability alter choice behavior and cognitive engagement. We find that while choices diverge further from equilibrium under strategic complementarity, eye gaze metrics—specifically fixation and transition measures—are consistent with more intensive cognitive engagement under strategic substitutability. Furthermore, by using a novel approach to measure strategic sophistication, we find that individuals with higher strategic aptitude exhibit higher levels of visual engagement with the payoff table. As strategic sophistication increases, the two distinct dimensions of attention build up sequentially rather than together, with attention to an individual’s own strategies strengthening prior to attention to their opponents’. These findings indicate that visual attention is shaped systematically by the strategic environment and are consistent with a sequential organization of reasoning that may underlie the strategic environment effect.
General gaze following aims to infer the region attended to by a person within a natural scene and offers a computational perspective on visual attention and social scene understanding. Existing direction-guided methods usually reduce gaze direction to a single vector or spatial mask. This deterministic treatment can obscure directional ambiguity, discard coexisting candidate directions, and propagate early estimation errors to gaze target localization when head cues are weak or multiple targets are plausible. To address these limitations, we propose an uncertainty-aware framework based on Circular Direction Distribution Learning (CDDL) and Probabilistic Gaze Geometry Modeling (PGGM). CDDL represents gaze direction as a 72-bin circular probability distribution under von Mises soft supervision, thereby preserving neighboring directional hypotheses before target localization. Rather than predicting a target space distribution directly, PGGM aggregates the direction distribution into 18 groups and projects the retained hypotheses into cone-like spatial probability fields, allowing spatial tolerance to expand with distance while preserving a channel-wise direction-to-region representation. The full probability volume is fused with the RGB scene image at the input level and processed by a ResNet-50-FPN network for gaze heatmap prediction. Controlled experiments on GazeFollow demonstrate competitive localization and support the contribution of direction-distribution learning and channel-wise geometric projection. Zero-shot transfer, dataset-level uncertainty and failure analyses, efficiency measurements, and qualitative results further indicate interpretable behavior together with domain-dependent and computational trade-offs.
Children with autism spectrum disorder (ASD) often exhibit atypical patterns of visual attention allocation and social-cue processing. Eye-tracking scanpath (ETSP) retains information about fixation points, saccade paths and their temporal changes in the form of images, providing an intuitive and computable data representation for analyzing ASD-related visual attention patterns. However, in ASD auxiliary identification studies, the same participant often generates multiple eye-tracking recordings or multiple visual representation samples. If participant independence is not properly considered during model evaluation, the training and test sets may share individualized eye-movement patterns from the same child. In such cases, the model may learn subject-specific characteristics rather than stable and transferable ASD-related visual attention features, leading to an overestimation of its recognition ability on unseen participants. To address this issue, we propose a Global–Local Collaborative Fusion Network (GLCF-Net) under a strict participant-independent splitting protocol. Specifically, the proposed method first maps ETSP images into patch token sequences through a shared Patch Embedding layer. A CNN-based local branch is then used to extract local trajectory morphology, path density, and spatial neighborhood structure, while a ViT-based global branch models cross-region gaze transitions and the overall attention distribution. Finally, a gated adaptive fusion module dynamically integrates local and global information to enhance the representation of stable visual attention features. In the primary repeated stratified five-fold participant-level evaluation, averaging the two out-of-fold probabilities for each participant yielded an Accuracy of 87.0% and a ROC-AUC of 93.7%; the original participant split, retained as a secondary analysis, yielded an Accuracy of 83.52% and a ROC-AUC of 90.27%. Under the reported frozen-backbone configurations, the model also showed a balanced pattern across Accuracy, Recall, and F1-score. These results characterize performance for unseen participants within the same dataset and acquisition conditions.
The current study examined the control of long-range regressive eye movements during sentence reading. Skilled readers were asked to read single-line sentences for comprehension. As a secondary task, they identified a probe word presented to the right of each sentence, and then went back to check the corresponding target word for spelling errors that had been added after reading. The regression target was located either close to or far from the probe. Consistent with prior work, initial regressions were larger for far than for near targets, and the finding of far targets required more eye movements and more search time. We replicated previous descriptions of visuomotor regression strategies but also discovered a novel strategy in which readers frequently send their initial regressions to the center of the line, followed by subsequent saccades. Apparently, this happened primarily when spatial knowledge on saccade target location was lacking. Spatial memory and reading skill made distinct contributions to regression targeting. Memory skills strongly determined primary regressions, especially to near targets, whereas reading speed influenced the time needed to attain targets with subsequent saccades.
Deaf children who acquire a signed language at home and at school prior to learning to read are a unique group of developing bilinguals. They are first-language users of a signed language (e.g., American Sign Language; ASL) and second-language users of the written form of the ambient spoken language (e.g., English). Because signed languages do not possess orthographies that are widely used, deaf readers typically read in their second language. This profile stands in contrast to hearing bilingual readers, who read in their first language and, sometimes, their second language. In studies of child readers, vocabulary knowledge is a strong predictor of reading comprehension for both monolinguals and bilinguals. We combine the unique linguistic backgrounds of young signing deaf readers with what is known about vocabulary knowledge and reading, guided by the following research question: Does signed (L1) and spoken (L2) vocabulary knowledge predict reading (eye-gaze) behaviors of deaf children when reading in their L2? We present results from an exploratory pilot eye-tracking study of nine deaf ASL-English bilingual children ages 9–11, examining whether ASL and English vocabulary knowledge predict different aspects of the reading paradigm. Results suggest that ASL vocabulary predicts where-decisions and measures of oculomotor control such as skipping and regressions, while English vocabulary predicts when-decisions and first-pass durations. We suggest these results from this small-scale study provide initial support for the language interdependence hypothesis because deaf child signers’ first-language vocabulary knowledge predicts some of the variance in second-language reading performance.
Background: Many individuals with traumatic brain injuries (TBIs) exhibit oculomotor dysfunctions that impact their daily functioning. As current clinical screening tools are limited, we have created and pilot-tested the Rehabilitation Oculomotor Screening Evaluation (ROSE) previously in a small sample of people with acquired brain injuries and neurotypical participants. The current study aims to validate ROSE in persons with TBI, focusing on mild TBI (mTBI). Methods: Participants with TBI (n = 25) completed different clinical scales, including ROSE, Sensory Organization Test (SOT) for standing balance, Reintegration to Normal Living Index (RNLI), Timed Up and Go (TUG) for mobility, and a visual analogue scale for the subjective perception of visual vertigo. Neurotypical individuals (n = 24) who were age- and sex-matched completed only ROSE. Results: The group with mTBI (n = 18) had significantly higher ROSE scores compared to the neurotypical group, with a large effect size. Significant correlation was found between ROSE and RNLI scores, but not with other clinical outcomes. Conclusions: Significant between-group difference in ROSE scores and their association with RNLI scores suggest that ROSE is a valid tool in detecting oculomotor dysfunction in TBI. Future studies should continue the validation of ROSE in other TBI and neurologic populations and in larger sample sizes.
Vertigo, dizziness, and oculomotor disturbances may occur as manifestations of immune-mediated disorders affecting the inner ear, central vestibular pathways, or multisystem autoimmune disease. Although uncommon, these conditions are clinically important because delayed recognition may lead to irreversible hearing loss, vestibular dysfunction, or neurological disability. This review summarizes the clinical presentation, diagnostic approach, and treatment of immune-mediated vestibular and oculomotor disorders. We suggest a practical classification into isolated immune-mediated inner ear disease, systemic autoimmune disorders with audio-vestibular involvement, and autoimmune disorders of the central or peripheral nervous system affecting balance and eye movements. Red flags for such conditions include bilateral or progressive symptoms, fluctuating audio-vestibular deficits, associated neurological signs, and accompanied autoimmune disease. Corticosteroids remain the main first-line treatment in many of these disorders, mainly due to missing data from controlled trials. Steroid-sparing immunosuppressants, biologics, and tumor-directed therapies are effective in many cases; however, because of the missing data, they are only used in selected entities without any other choice. A structured neuro-otological and immunological workup is essential to improve diagnostic accuracy and enable timely therapy.
This study investigated the relationship between fixation-frequency-based Shannon entropy and dwell-time-based entropy across two different visual task domains: emotional evaluation of automotive exterior designs and safety-critical monitoring of nuclear power plant emergency scenarios. Although gaze entropy has been widely used to explain emotional responses, task performance, and situation awareness, the relationship between entropy measures derived from fixation counts and fixation durations remains insufficiently examined. Eye-tracking data were analyzed from two experiments with different attentional characteristics. In the emotional visual task, 10 participants evaluated three automotive design images. In the safety-critical task, 20 participants performed four nuclear power plant emergency monitoring scenarios. Shannon entropy and dwell-time entropy were calculated using fixation count and fixation duration distributions across Areas of Interest, respectively. Pearson correlation and simple regression analyses were conducted within each task domain. The results showed strong positive associations between Shannon entropy and dwell-time entropy in both domains. The emotional task showed a correlation of r = 0.844, while the safety-critical task showed a correlation of r = 0.890. These findings suggest that fixation-frequency-based and dwell-time-based entropy measures exhibit substantial overlap across different visual task contexts. However, the observed associations may partly reflect mathematical dependency between fixation frequency and cumulative dwell-time, and the findings should be interpreted as exploratory evidence rather than proof of metric interchangeability. The study highlights that gaze entropy metrics should be interpreted in relation to task-dependent attentional contexts. Higher entropy may be associated with exploratory visual attention in emotional evaluation, whereas lower entropy may be associated with focused monitoring in safety-critical tasks.
Eye-movement studies of manual production often average gaze across an entire trial, obscuring how visual information use changes once actions begin. We separated the pre-writing and writing phases in a fixed progressive Chinese calligraphy task. Thirty-seven postgraduate students completed two style-guided transfer (SGT) pages, a worked example, and two evolution-based mapping (EBM) pages; 34 contributed usable gaze data. On SGT pages, reference allocation fell from 0.626 before writing to 0.131 during writing, whereas the share of reference viewing directed to diagnostic tokens rose from 0.473 to 0.601. On EBM pages, allocation to the cue-plus-context display fell from 0.825 to 0.447 after pen onset but remained substantial; cue share and context coverage also declined. Participant-level process blocks did not improve quality models. In exploratory page-level EBM analyses, greater pre-writing context coverage was associated with higher product quality. These findings identify pen onset as a useful boundary for analyzing visual information use in constrained production: external sampling is greatest before writing, and task-specific re-access persists during execution. Because the task order was fixed, page-family differences cannot be separated from practice or scaffolding. Phase-specific area-of-interest measures can therefore add process information to product scores without treating gaze as a direct measure of cognition.
Reproducible eye-movement research requires a documented path from device exports to structured, quality-checked, and reportable data objects. Gazepoint GP3 and Gazepoint Analysis provide accessible gaze, fixation, pupil, timing, media, and area-of-interest data, but their folder-based CSV exports are not automatically analysis-ready. Existing R tools support important stages of eye-tracking and pupillometry analysis, but they generally assume that data have already been organised into suitable structures; they do not provide a Gazepoint-aware workflow beginning with export-folder checks, all-gaze/fixation pairing, sampling and tracking-quality diagnostics, and preservation of preprocessing decisions through reporting. This article presents gp3tools, an open-source R package (R version 4.6.1) that converts Gazepoint GP3/Gazepoint Analysis exports into structured R objects, diagnostic summaries, preprocessing outputs, model-ready tables, interoperability objects, and reproducible reports. Rather than introducing a new statistical estimator, the package provides an export-aware workflow scaffold for import checking, quality control, pupil preprocessing, area-of-interest, fixation and transition summaries, model preparation, interoperability, and reporting. A synthetic Gazepoint-style demonstration dataset was used to evaluate workflow execution without exposing private participant data. The demonstration identified all expected file pairs and produced sample-level gaze/pupil tables, fixation tables, sampling-quality summaries, area-of-interest summaries, review flags, and reporting outputs. A small, private real-export compatibility check further showed that the workflow could process one empirical Gazepoint folder without manual restructuring. These results support a bounded software-evaluation claim: gp3tools executes the tested Gazepoint-style workflow and returns expected diagnostic and reporting objects, but the evidence does not establish hardware accuracy, preprocessing accuracy, computational scalability, general robustness across all Gazepoint exports, or substantive psychological or perceptual effects.
This study analyzes the relationship between verbal interaction and eye behavior among 112 primary school students in urban and rural classrooms in Chile, using wireless eye-tracking technology. The results reveal statistically significant differences based on socioeducational context and sex. Linear regression analyses show that gaze is a significantly more robust predictor of class participation in rural contexts (R2 adjusted = 0.671) than in urban contexts (R2 adjusted = 0.342). Furthermore, eye behavior explained 71% of the variance in male students, compared to 37.7% in female students. While female students focused their attention primarily on teachers, male students relied on a shared visual distribution between the teacher and peers to regulate their participation in class. In conclusion, the gaze acts as a differentiated scaffolding whose importance intensifies in boys and rural environments. These findings suggest distinct maturational trajectories that require teachers to implement visually intentional instructional strategies to ensure communicative efficiency in the classroom.
Brain-computer interface (BCI) technology has shown potential for future rehabilitation-related and assistive control applications. Nevertheless, single-modality electroencephalography-based motor imagery (EEG-MI) signals are susceptible to interference, whereas existing algorithmic models suffer from limited classification accuracy and insufficient actionable control commands for interactive devices, thereby impeding their practical deployment. To tackle these limitations, this study presents a multimodal human-computer interaction control scheme that integrates eye-movement command encoding with EEG motor imagery decoding. Self-collected EEG-MI and eye-movement datasets were established to support the proposed multimodal control framework. In this framework, eye movements are not used merely as auxiliary inputs, but are encoded as discrete commands for start, stop, grasp, and release, thereby reducing the command burden of EEG-MI decoding. The EEG-TransNet model is enhanced by integrating a time-frequency feature branch and replacing the original convolutional encoder with an adaptive multi-branch EEG feature gating module, strengthening the representation and fusion of multi-domain features. The model yields average classification accuracies of 86.96% and 88.73% on the BCI IV-2a dataset and the self-collected EEG dataset, respectively. Four independent SVM binary classifiers are adopted to identify four eye movement patterns. The EEG and eye movement classification results are binary-encoded to generate hardware-compatible control commands. Robotic-arm grasping experiments with healthy trained participants showed an average task completion time of 17 s, and the repeated grasping success-rate results further provide preliminary evidence for the real-time feasibility of the multimodal control framework under controlled laboratory conditions.
This study investigated how congruence between pictogram complexity and typographic complexity influences visual processing and perceived appropriateness in multimodal communication. Drawing on theories of visual communication, legibility and aesthetic perception, the study examined whether simple or complex pictograms harmonise more effectively with sans-serif or serif typefaces. Ninety participants viewed stimuli from three thematic categories (cobbler, herbal pharmacy and gluten-free restaurant), while their eye movements were recorded using a Tobii Pro Fusion eye-tracking device. Measures included reading time, fixation count and saccade count, together with subjective evaluations of pictogram-typeface suitability. The results show that reading time was the most sensitive indicator of formal congruence. In the cobbler and gluten-free restaurant categories, simple pictograms increased reading time with sans-serif typography but decreased it with serif typography. In the herbal pharmacy category, simple pictograms and sans-serif typography independently supported faster reading performance. Subjective evaluations showed no significant differences between combinations, indicating that participants perceived all pairings as similarly appropriate despite measurable differences in processing efficiency. The findings suggest that the effectiveness of pictogram-typeface combinations depends on both formal complexity and thematic context. Eye-tracking proved valuable for revealing subtle cognitive processing differences not reflected in subjective judgements.
Background: The popularity of Dr. Seuss books across generations and languages is worthy of introspective analysis in order to establish how well the translation adheres to the original essence that is the underlying reason for its enduring nature. This article analyses Horton Hears a Who! and its Afrikaans translation. Methods: A basic stylistic analysis of Dr. Seuss’ Horton Hears a Who! was conducted in order to determine whether its rhythmic nature and the meter used in the original text translate well into Afrikaans. Additionally, two extracts were used for an eye-tracking experiment whereby participants were asked to read both so that gaze behavior could be analyzed and compared between Afrikaans and English. Results: Eye-tracking results of both passages reveal that readers are comfortably within typical reading behavior and do not experience difficulty when reading the unique style of Dr. Seuss in either language. Stylistic meter does however change the position of elevated fixation durations, but made-up words and the underlying playful tone do not result in difficulty. However, text characteristics, such as length and number of syllables, and word positions, significantly affect gaze behavior. Conclusions: It is concluded that the very nature of Dr. Seuss books that makes the stories come alive for children over generations is successfully embodied in the Afrikaans translation with respect to the aspects tested and that both the serious, skillful writing and the more frivolous story-telling tone is present in the Afrikaans target text.