Eye movements are promising biomarkers for psychiatric and neurological disorders, yet conventional recording methods rely on bulky, expensive, and laboratory-bound devices. Deep-learning-based gaze-tracking technology now provides a practical means to capture these biomarkers outside the lab, with consumer-grade smartphones offering a particularly scalable platform. To evaluate the feasibility of this approach for psychiatric assessment, we conducted two complementary studies: a clinical investigation of hospitalized patients with schizophrenia and a normative study of healthy college students. In Study 1, we collected gaze data from individuals with clinically diagnosed schizophrenia (N = 134) and matched healthy controls (N = 130) using both an iPhone and a research-grade EyeLink eye-tracker. In Study 2, we used smartphone gaze-tracking data from a university cohort (N = 631) to classify depressive symptoms. Results demonstrated that smartphone-derived gaze metrics can distinguish psychiatric conditions. For schizophrenia detection, the smartphone model achieved an area under the receiver operating characteristic curve (AUC) of 87.00% and an accuracy of 83.33%, comparable to the EyeLink benchmark (AUC = 87.12%; accuracy = 86.67%). For classifying depressive symptoms, a free-viewing task on Android smartphones yielded an AUC of 75.54% and an accuracy of 75.79%. These findings highlight the potential of smartphone gaze-tracking as an accessible, privacy-preserving tool for real-world psychiatric assessment and treatment monitoring.
The Apple Vision Pro (AVP) represents a shift in mixed reality (MR) interaction by eliminating controllers in favor of gaze- and gesture-based input. However, limited information is available about the AVP eye-tracking specifications because raw gaze data are inaccessible to developers. This study evaluates the eye-tracking accuracy of the AVP using two testing pipelines across MR and virtual reality (VR) environments. This study also examines the relationship between eye-tracking accuracy and user-perceived usability. The tests revealed an overall eye-tracking accuracy of 0.45°–1.06° across the two pipelines, within a field of view (FOV) of approximately 34° × 18°. The usability and learnability scores of the AVP, measured using the System Usability Scale (SUS), were 76.05 and 69.30, respectively. No statistically significant correlation was observed between eye-tracking accuracy and usability scores, suggesting that the AVP was sufficiently accurate across participants to avoid affecting usability during the testing session. Beyond these findings, the testing methods reported in this work provide a practical guide for evaluating eye-tracking accuracy in closed MR/AR systems.
Gaze-tracking has a wide range of applications across scientific and industrial fields, and recent computer vision and deep-learning advances have made gaze-tracking with standard webcams possible. However, current solutions only offer suboptimal performance and lack flexibility. This paper introduces GazeFollower, an accessible system for webcam gaze-tracking in Python. GazeFollower stands out for its customizability, allowing researchers to quickly develop and adapt algorithms to meet their needs. At its core, GazeFollower estimates gaze with a model trained on 32 million face images, ensuring robust gaze tracking. A benchmark test on a sizeable sample (N=31) shows that the tracking performance of GazeFollower is on par with or better than budget commercial eye trackers. With calibration, GazeFollower has an accuracy of 1.11 cm and a precision of 0.11 cm, and personalized model fine-tuning further enhances these metrics to 0.92 cm and 0.08 cm. These results suggest that GazeFollower holds potential for real-world applications.
Eye-tracking is widely used to measure human attention in research, commercial, and clinical applications. With the rapid advancements in artificial intelligence and mobile computing, deep learning algorithms for computer vision-based eye tracking have become feasible for smartphones. This paper presents a real-time smartphone eye-tracking system built upon a deep neural network trained on a dataset of 7.4 million facial images. The tracking performance of the system was benchmarked against an industrial gold-standard EyeLink eye tracker using a reasonably large sample (N = 32). The benchmark test showed that, while the smartphone eye-tracking system was less precise (0.177° vs. 0.028°), its tracking accuracy was comparable to the EyeLink tracker (1.32° vs. 1.20°). To evaluate whether the smartphone eye-tracking system is sensitive enough for real-world application, a field test involving 98 volunteers assessed depressive symptoms using three simple visual tasks on a smartphone: fixation stability, free-viewing, and smooth pursuit. The results showed that using the smartphone eye-tracking system can achieve an accuracy of 76.67
The invisible protective bubbles surrounding us delineate the boundary of personal space we anticipate others to respect. This study examined the impact of collaborative robots on this protective bubble in the context of human-robot teaming, utilizing the widely used Godspeed questionnaire and pupil recording. The questionnaire revealed heightened feelings of perceived unsafety, coupled with increased discomfort when the robot was situated behind and head-aligned with its human teammate. The pupil recordings further supported these observations, revealing a more pronounced pupil dilation when the robot stepped in place behind its human teammate, especially when the robot was heading in the same direction. These findings underscore the pivotal role of the pose of robots in human-robot teams. In human-robot collaborations, it is advisable to avoid situating the robot behind and aligning its head with the human teammate.
In human communication, people often turn and gaze at a specific person in a crowd to signal their intention to interact with them. Similarly, it has been proposed that robots should also use social cues such as facing their human interaction partner during human–robot interaction tasks. This study introduces an initiatively interactive pose control (IPC) framework that allows a robot to face its task-relevant human interaction partner and proactively use this social cue while carrying out desired actions based on the ongoing task state. The IPC framework integrates a task planning module and a 3D identity recognition module. The task planning module can generate task states that include information about the desired human interaction partner’s name and the expected actions of robots, including social cues. The 3D identity recognition module implemented in the IPC framework can identify potential human interaction partners and estimate their pose in relation to the robot. The relative pose serves as the control parameter for orienting the robot toward the selected human interaction partner. The experimental results show that the IPC framework achieves a relative pose estimation error ranging from 0.04 to 6.56 degrees, which signifies a substantial enhancement compared to traditional sound source localization methods. Moreover, experiments also demonstrate that the robot can proactively turn toward the interaction partner and execute expected actions using the IPC framework. In conclusion, this paper introduces a new protocol for interactive pose control, enabling robots to actively select their human interaction partners and exhibit social cues and associated interaction actions.
Eye tracking has emerged as a valuable tool for both research and clinical applications. However, traditional eye-tracking systems are often bulky and expensive, limiting their widespread adoption in various fields. Smartphone eye tracking has become feasible with advanced deep learning and edge computing technologies. However, the field still faces practical challenges related to large-scale datasets, model inference speed, and gaze estimation accuracy. The present study created a new dataset that contains over 3.2 million face images collected with recent phone models and presents a comprehensive smartphone eye-tracking pipeline comprising a deep neural network framework (MGazeNet), a personalized model calibration method, and a heuristic gaze signal filter. The MGazeNet model introduced a linear adaptive batch normalization module to efficiently combine eye and face features, achieving the state-of-the-art gaze estimation accuracy of 1.59 cm on the GazeCapture dataset and 1.48 cm on our custom dataset. In addition, an algorithm that utilizes multiverse optimization to optimize the hyperparameters of support vector regression (MVO-SVR) was proposed to improve eye-tracking calibration accuracy with 13 or fewer ground-truth gaze points, further improving gaze estimation accuracy to 0.89 cm. This integrated approach allows for eye tracking with accuracy comparable to that of research-grade eye trackers, offering new application possibilities for smartphone eye tracking.
Although Severe Acute Respiratory Syndrome Coronavirus 2 infection (SARS-CoV-2) is primarily recognized as a respiratory disease, mounting evidence suggests that it may lead to neurological and cognitive impairments. The current study used three eye-tracking tasks (free-viewing, fixation, and smooth pursuit) to assess the oculomotor functions of mild infected cases over six months with symptomatic SARS-CoV-2 infected volunteers. Fifty symptomatic SARS-CoV-2 infected, and 24 self-reported healthy controls completed the eye-tracking tasks in an initial assessment. Then, 45, and 40 symptomatic SARS-CoV-2 infected completed the tasks at 2- and 6-months post-infection, respectively. In the initial assessment, symptomatic SARS-CoV-2 infected exhibited impairments in diverse eye movement metrics. Over the six months following infection, the infected reported overall improvement in health condition, except for self-perceived mental health. The eye movement patterns in the free-viewing task shifted toward a more focal processing mode and there was no significant improvement in fixation stability among the infected. A linear discriminant analysis shows that eye movement metrics could differentiate the infected from healthy controls with an accuracy of approximately 62%, even 6 months post-infection. These findings suggest that symptomatic SARS-CoV-2 infection may result in persistent impairments in oculomotor functions, and the employment of eye-tracking technology can offer valuable insights into both the immediate and long-term effects of SARS-CoV-2 infections. Future studies should employ a more balanced research design and leverage advanced machine-learning methods to comprehensively investigate the impact of SARS-CoV-2 infection on oculomotor functions.
Most commercially available eye-tracking devices rely on video cameras and image processing algorithms to track gaze. Despite this, emerging technologies are entering the field, making high-speed, cameraless eye-tracking more accessible. In this study, a series of tests were conducted to compare the data quality of MEMS-based eye-tracking glasses (AdHawk MindLink) with three widely used camera-based eye-tracking devices (EyeLink Portable Duo, Tobii Pro Glasses 2, and SMI Eye Tracking Glasses 2). The data quality measures assessed in these tests included accuracy, precision, data loss, and system latency. The results suggest that, overall, the data quality of the eye-tracking glasses was lower compared to that of a desktop EyeLink Portable Duo eye-tracker. Among the eye-tracking glasses, the accuracy and precision of the MindLink eye-tracking glasses were either higher or on par with those of Tobii Pro Glasses 2 and SMI Eye Tracking Glasses 2. The system latency of MindLink was approximately 9 ms, significantly lower than that of camera-based eye-tracking devices found in VR goggles. These results suggest that the MindLink eye-tracking glasses show promise for research applications where high sampling rates and low latency are preferred.
With built-in eye-tracking cameras, the recently released Apple Vision Pro (AVP) mixed reality (MR) headset features gaze-based interaction, eye image rendering on external screens, and iris recognition for device unlocking. One of the technological advancements of the AVP is its heavy reliance on gaze- and gesture-based interaction. However, limited information is available regarding the technological specifications of the eye-tracking capability of the AVP, and raw gaze data is inaccessible to developers. This study evaluates the eye-tracking accuracy of the AVP with two sets of tests spanning both MR and virtual reality (VR) applications. This study also examines how eye-tracking accuracy relates to user-reported usability. The results revealed an overall eye-tracking accuracy of 1.11 and 0.93 in two testing setups, within a field of view (FOV) of approximately 34 x 18. The usability and learnability scores of the AVP, measured using the standard System Usability Scale (SUS), were 75.24 and 68.26, respectively. Importantly, no statistically reliable correlation was found between eye-tracking accuracy and usability scores. These results suggest that eye-tracking accuracy is critical for gaze-based interaction, but it is not the sole determinant of user experience in VR/AR.
Drowsiness poses a serious challenge to road safety and various in-cabin sensing technologies have been experimented with to monitor driver alertness. Cameras offer a convenient means for contactless sensing, but they may violate user privacy and require complex algorithms to accommodate user (e.g., sunglasses) and environmental (e.g., lighting conditions) constraints. This paper presents a lightweight convolution neural network that measures eye closure based on eye images captured by a wearable glass prototype, which features a hot mirror-based design that allows the camera to be installed on the glass temples. The experimental results showed that the wearable glass prototype, with the neural network in its core, was highly effective in detecting eye blinks. The blink rate derived from the glass output was highly consistent with an industrial gold standard EyeLink eye-tracker. As eye blink characteristics are sensitive measures of driver drowsiness, the glass prototype and the lightweight neural network presented in this paper would provide a computationally efficient yet viable solution for real-world applications.
In this paper,Substep crystalization of 2,5-cresol was studied,thus 2,5-cresol(≥99.0%)was got after substep crystalization.The effects of substep crystalization were studied for optimizing the crystalization conditions,then the conditions with high yield and low cost were successfully applied in the test,the yield of 2,5-cresol reached 55.0%。
The refinement of 3,5-cresol by fractional crystallization was studid.The content of 3,5-cresol was up to 98.0%,and the crystallization yield was 44.0%.
In this paper, step crystallization of 3,4-dimethylphenol was studied to give 3,4-dimethylphenol (≥98.0%). The effects of step crystallization were studied for optimizing the crystallization conditions, then the conditions with high yield and low cost were successfully applied in the test, the yield of 3,4-dimethylphenol reached 44.0%.