Immersive virtual reality (IVR) provides interactive and experiential environments that enhance learning beyond traditional classrooms. This study examined engineering students’ perceptions of presence, usability, and acceptance while engaging in a chemistry laboratory implemented through IVR. Twentyeight students from ten engineering programs participated in virtual laboratory sessions using Meta Quest 3 headsets. The virtual environment replicated real-world experimental procedures, enabling students to manipulate materials and make decisions that influenced experimental outcomes. Data were collected using validated scales to assess the presence, ease of use, graphical interface, and perceived usefulness in a pretest–posttest design, complemented by interviews. No significant differences emerged in perceived usefulness, ease of use, or presence (p > .30). However, a small-to-moderate improvement appeared in the graphical interface (dAKP = –.36). Students reported high realism, intuitive interaction, and motivation to continue using IVR, considering the IVR laboratory a meaningful complement to physical laboratories. Findings suggest that a well-designed, intuitive IVR environment can enhance students’ engagement and perceived usability, even without prior virtual lab experience. These results support IVR’s potential as a valuable pedagogical complement in STEM education. Future studies should further explore its integration into broader instructional strategies and its relationship with cognitive and affective learning outcomes.
Egocentric 3D human pose estimation remains challenging due to severe perspective distortion, limited body visibility, and complex camera motion inherent in first-person viewpoints. Existing methods typically rely on single-frame analysis or limited temporal fusion, which fails to effectively leverage the rich motion context available in egocentric videos. We introduce AG-EgoPose, a novel dual-stream framework that integrates short- and long-range motion context with fine-grained spatial cues for robust pose estimation from fisheye camera input. Our framework features two parallel streams: A spatial stream uses a weight-sharing ResNet-18 encoder-decoder to generate 2D joint heatmaps and corresponding joint-specific spatial feature tokens. Simultaneously, a temporal stream uses a ResNet-50 backbone to extract visual features, which are then processed by an action recognition backbone to capture the motion dynamics. These complementary representations are fused and refined in a transformer decoder with learnable joint tokens, which allows for the joint-level integration of spatial and temporal evidence while maintaining anatomical constraints. Experiments on real-world datasets demonstrate that AG-EgoPose achieves state-of-the-art performance in both quantitative and qualitative metrics. Code is available at: https://github.com/Mushfiq5647/AG-EgoPose.
Assessing cognitive load continuously and at low latency would help adaptive human-computer interaction, but it remains hard because labeled data are scarce and models generalize poorly across subjects. Recent ECG foundation models, pre-trained on millions of clinical diagnostic ECG recordings, yet they do not apply directly to wearable devices when the sensor configuration and the task both differ. We present CogAdapt, a framework that adapts a clinical ECG foundation model to wearable cognitive load assessment. CogAdapt has two parts. LeadBridge is a learnable adapter that maps 3-lead wearable signals to a 12-lead-compatible representation. ProFine is a progressive fine-tuning strategy that unfreezes encoder layers in stages while limiting representational drift in the pre-trained model. On two public datasets (CLARE and CL-Drive) under leave-one-subject-out cross-validation, CogAdapt reaches macro-F1 of 0.626 and 0.768, improving over from-scratch baselines by 11.2 and 16.1 percentage points. The results show that a clinical ECG pretraining can support subject-independent cognitive load assessment from wearable sensors.
Cybersickness (CS) is a significant challenge to the usability of virtual reality (VR) applications, with its subjective nature complicating real-time detection. Due to individual differences, many VR users do not experience CS, leading to imbalanced datasets dominated by non-CS cases, which can compromise the machine learning-based models. Addressing this issue requires diverse and comprehensive data, but collecting such datasets through user studies is time-consuming and resourceintensive. To mitigate this, data augmentation techniques can be applied to expand datasets, improving model performance without requiring extensive new data. In this paper, we investigated various data augmentation techniques, including time seriesbased methods, encoder-decoder models, and generative adversarial networks (GANs), to enhance CS detection performance using the publicly available CS datasets. The comprehensive performance analysis shows that cGAN based data augmentation improves F1 score of cybersickness detection by 6.49% for the Simulations 2021 (Sim21) dataset and by 3.85% for the VRwalking dataset. In comparison, the time series-based augmentation method SMOTE, achieved a modest F1 score improvement of up to 3.9%. Our findings offer important implications for VR developers and researchers, highlighting how data augmentation techniques can enhance the accuracy of cybersickness prediction models while reducing the time and effort required for additional data collection.
Despite the rapid growth of Virtual Reality (VR) technology across diverse applications, cybersickness remains a significant barrier to its widespread adoption. To move toward effective mitigation, current efforts have concentrated on predicting cybersickness, particularly through the integration of multimodal data sources. However, current cybersickness prediction models face several limitations that impede progress toward robust, generalizable systems: most studies are confined to single datasets due to substantial technical challenges in integrating heterogeneous VR sensor data from different hardware platforms, and existing approaches fail to appropriately model the ordinal nature of cybersickness severity scales, treating Fast Motion Sickness (FMS) scores as either continuous variables or arbitrary categorical classifications that discard valuable ordering information. We developed a comprehensive statistical feature abstraction framework combined with dual-head ordinal regression to enable cybersickness prediction across heterogeneous VR datasets. Our approach transforms diverse sensor modalities from VR.net (Meta Quest Pro proprietary systems), SIM21, and VRWalking (HTC VIVE/Tobii) platforms into standardized statistical descriptors that capture cybersickness-relevant behavioral patterns while remaining invariant to study design and data collection methodologies. The dual-head architecture combines ordinal-aware predictions with direct regression outputs through weighted averaging. Using the VR.net and VRWalking datasets, we achieved cross-dataset generalization with an RMSE of 0.8341 ± 0.0731 and MAE of 0.6254 ± 0.0389 through rigorous 10-fold participant-aware cross-validation, along with 97.7% ± 0.87% accuracy in severity classification despite heterogeneous hardware configurations and scale differences. Our statistical feature abstraction approach enables effective cross-platform cybersickness prediction by creating domain-agnostic representations that preserve essential physiological patterns, advancing the development of universal cybersickness monitoring systems capable of operating across diverse VR platforms.
Time perception in virtual reality (VR) is influenced by various factors, including avatar embodiment. This study investigates how avatar type (human vs. Godzilla) and bodily proportions (human-like vs. Godzilla-like) affect time perception during passive (waiting) and active (walking) tasks. Nineteen participants embodied four avatars, with different combinations of these features, in two virtual environments—a biomechanics lab and an urban cityscape—while their embodiment levels and perception of time were assessed. Results reveal that avatar proportions influenced the estimation of time during the waiting phase only, with human-like proportions leading to greater underestimations. These findings highlight the distinct roles of avatar characteristics in shaping user experiences in VR, particularly their influence on temporal judgment and embodiment.
Cybersickness and cognitive overload present significant barriers to the widespread adoption of virtual reality (VR) across various domains. These conditions can disrupt the user experience and impair learning outcomes in VR-based training simulations. Existing research has explored innovative mitigation strategies, typically employing a structured approach: collecting multimodal data alongside labels for cybersickness and cognitive load, training predictive models to classify these states with high accuracy, and dynamically applying mitigation techniques. While such approaches have provided valuable insights, prior studies indicate that cybersickness and cognitive overload often evolve, making real-time prediction less effective. Previous attempts at forecasting cybersickness and cognitive load perform well up to 60-90 seconds into the future, but their accuracy significantly drops shortly thereafter. In this study, we propose using foundational time-series models, pre-trained on large amounts of time-series data, for forecasting cybersickness and cognitive load, rather than training models from scratch. Our findings demonstrate that, even with limited labeled data from a 14 minute session per participant, forecasting cybersickness and cognitive load remains feasible. Using just 57% of the labeled data, both models achieve strong predictive performance and remain reasonably accurate with as little as 28.5% labeled data. This enables us to use labels for the first half of the session and forecast the second half with high precision, enabling the early and timely application of mitigation strategies. This research contributes to the field of VR by providing a robust forecasting framework for cybersickness and cognitive overload, paving the way for adaptive mitigation strategies.
Real-time cognitive load assessment from eye-tracking signals could potentially enable adaptive human-centered-AI such as safety-critical applications such as driver vigilance monitoring or automated flight deck assistance, yet two challenges persist: handling frequent data missingness from blinks and tracking failures, and efficiently modeling long-range temporal dependencies. We propose MambaGaze, a framework that addresses these challenges through 1) XMD encoding, which augments raw features with observation masks and time-deltas to explicitly model data uncertainty, and 2) bidirectional Mamba-2, which captures temporal dependencies with linear computational complexity. Experiments on CLARE and CL-Drive datasets under leave-one-subject-out evaluation show that MambaGaze achieves 76.8
Cybersickness, a motion sickness like discomfort, is a major barrier to the usability of virtual reality (VR) systems. While prior work has focused mainly on predicting cybersickness severity, practical mitigation requires not only detecting how sick a user feels but also deciding whether a countermeasure is beneficial and determining its appropriate intensity. In this paper, we propose a two phase multitask learning framework that jointly models cybersickness severity, blur effectiveness, and blur intensity. In Phase 1, we pretrain temporal deep learning backbones on two single label datasets with only severity annotations. In Phase 2, we pro-gressively finetune the models on a multi-label dataset containing severity, blur effectiveness, and blur level labels. We evaluate three backbone architectures a Time-Series Transformer, Deep Temporal Convolutional Network, and TS-Mamba under a 10-fold block aware cross validation scheme. Results show that two phase training significantly outperforms single phase baselines, with the Time Series Transformer achieving best performance (FMS MAE = 0.57, R2 = 0.87; Blur Level MAE = 0.49, 2 = 0.95; Blur Preference ACC = 99.5%). Unlike prior rule based reduction frameworks that rely on static heuristics, our approach provides a data-driven "detect-decide-dose" pipeline that adapts blur mitigation dynamically to individual users. This demonstrates that single label pre-training is an effective strategy for developing multitask VR safety models under limited labeled data. To our knowledge, this is the first framework that unifies cybersickness prediction and adaptive reduction in a single model.
Egocentric 3D human pose estimation from head-mounted stereo cameras is challenging due to fisheye distortion, severe self-occlusion, and frequent truncation of body joints outside the camera field of view. Recent stereo egocentric methods have improved performance through heatmap lifting, stereo correspondence, and transformer-based refinement, but they often rely heavily on frame-local evidence or use temporal information only as auxiliary pose-level context. This limits robustness when current-frame stereo cues are weak, occluded, or ambiguous. We propose TSR-Ego, a temporally guided stereo framework that couples short-term motion evidence with projection-guided feature sampling. The model first enriches dense stereo feature maps using a causal depthwise-separable temporal convolution, allowing past visual evidence to influence the feature space before deformable cross-attention. A single-stage causal stereo decoder then refines learned 3D joint queries through temporal self-attention, joint self-attention, and fisheye deformable stereo cross-attention, using the evolving pose estimate to generate 2D sampling references. Unlike methods that apply temporal reasoning mainly after pose prediction, TSR-Ego uses motion context to shape both the sampled stereo features and the joint representations while preserving online inference without future frames. Experiments on UnrealEgo2 and UnrealEgo-RW show state-of-the-art performance, with especially strong gains on real-world sequences.
Cybersickness could negatively impact user comfort, immersion and the long-term adoption of virtual reality technologies. Past literature suggests the benefits of accurate cybersickness prediction for early interventions; however, the labeled data scarcity poses significant challenges to the development of an accurate predictive machine learning approach in this problem domain. In particular, traditional fully-supervised methods require large labeled datasets, which are expensive and timeconsuming to collect. This, in turn, could have significant adverse effects on the effectiveness of such fully-supervised approaches in this application domain. Accordingly, this study investigates a few-shot learning approach based on prototypical networks to predict cybersickness symptoms (introduced in the Simulator Sickness Questionnaire) under the label scarcity challenge, and compares its performance with a fully-supervised single-task learning approach. Our empirical evaluation shows that, on average, the best-performing Prototypical Network achieves a mean AUC of 0.54 across all symptoms, compared to 0.45 for the singletask baseline. In particular, on several symptoms, the Prototypical Network outperforms the baseline by more than 0.2 in terms of AUC. However, some exceptions have also been observed with respect to the superiority of Prototypical Networks to the baseline on certain symptoms. Therefore, these results suggest that fewshot Prototypical Networks deliver average improvements in cybersickness symptom prediction while highlighting a need for further research to consistently achieve a high accuracy across all symptoms.
Walking in immersive virtual reality (VR) environments is often disrupted by gait (i.e., walking patterns) instability, a challenge that is further exacerbated for individuals with mobility impairments. This study examines the effectiveness of multimodal feedback by integrating auditory, vibrotactile, and visual cues in improving walking performance within VR. A total of 68 participants, equally divided between those with mobility impairments and those without, completed walking tasks under multiple feedback conditions. Walking velocity was the primary performance metric, supplemented by subjective assessments of mental workload, fatigue, presence, usability, and simulator sickness. Results revealed that multimodal feedback significantly enhanced walking velocity compared to unimodal and bimodal conditions, with statistical analysis confirming strong improvements (p < .001). Participants also reported a favorable user experience under multimodal conditions despite slightly increased cognitive demand. These findings highlight the potential of integrated sensory feedback to mitigate gait disturbances in VR, promoting safer and more accessible locomotion for users with and without mobility impairments.
Ensuring a safe virtual reality (VR) experience requires systems that can predict and respond when users lose their balance. Although prior work has examined fall prediction and motion sickness, many approaches are regression-based and postural state classification remains less explored. This study compares machine learning (ML) and deep learning (DL) models for classifying postural states in VR under visual perturbations. We used a multimodal dataset containing kinematic, electromyographic (EMG), and electrodermal activity (EDA) signals. The data were prepared for a binary task to distinguish balanced from imbalanced postural states, and participant-wise downsampling addressed class imbalance. All models were evaluated with Leave-One-Participant-Out (LOPO) cross-validation to test generalization to unseen participants. Among the models, the Mamba-inspired CNN (MI-CNN) achieved the highest accuracy of 96.76
With the rapid advancement of virtual reality (VR) technology, its adoption across domains such as healthcare, education, and entertainment has grown significantly. However, the persistent issue of cybersickness, marked by symptoms resembling motion sickness, continues to hinder widespread acceptance of VR. While recent research has explored multimodal deep learning approaches leveraging data from integrated VR sensors like eye and head tracking, there remains limited investigation into the use of video-based features for predicting cybersickness. In this study, we address this gap by utilizing transfer learning to extract high-level visual features from VR gameplay videos using the InceptionV3 model pretrained on the ImageNet dataset. These features are then passed to a Long Short-Term Memory (LSTM) network to capture the temporal dynamics of the VR experience and predict cybersickness severity over time. Our approach effectively leverages the time-series nature of video data, achieving a 68.4
This research aims to examine the effects of various vibrotactile feedback techniques on gait (i.e., walking patterns) in virtual reality (VR). Prior studies have demonstrated that gait disturbances in VR users are significant usability barriers. However, adequate research has not been performed to address this problem. In our study, 39 participants (with mobility impairments: 18, without mobility impairments: 21) performed timed walking tasks in a real-world environment and identical activities in a VR environment with different forms of vibrotactile feedback (spatial, static, and rhythmic). Within-group results revealed that each form of vibrotactile feedback improved gait performance in VR significantly compared to the no vibrotactile condition in VR for individuals with and without mobility impairments. Moreover, spatial vibrotactile feedback increased gait performance significantly in both participant groups compared to other vibrotactile conditions.
As VR technology advances, the demand for multitasking within virtual environments escalates. Negotiating multiple tasks within the immersive virtual setting presents cognitive challenges, where users experience difficulty executing multiple concurrent tasks. This phenomenon highlights the importance of cognitive functions like attention and working memory, which are vital for navigating intricate virtual environments effectively. In addition to attention and working memory, assessing the extent of physical and mental strain induced by the virtual environment and the concurrent tasks performed by the participant is key. While previous research has focused on investigating factors influencing attention and working memory in virtual reality, more comprehensive approaches addressing the prediction of physical and mental strain alongside these cognitive aspects remain. This gap inspired our investigation, where we utilized an open dataset - VRWalking, which included eye and head tracking and physiological measures like heart rate(HR) and galvanic skin response(GSR). The VRwalking dataset has timestamped labeled data for physical and mental load, working memory, and attention metrics. In our investigation, we employed straightforward deep learning models to predict these labels, achieving noteworthy performance with 91%, 96%, 93%, and 91% accuracy in predicting physical load, mental load, working memory, and attention, respectively. Additionally, we conducted SHAP (SHapley Additive exPlanations) analysis to identify the most critical features driving these predictions. Our findings contribute to understanding the overall cognitive state of a participant and effective data collection practices for future researchers, as well as provide insights for virtual reality developers. Developers can utilize these predictive approaches to adaptively optimize user experience in real-time and minimize cognitive strain, ultimately enhancing the effectiveness and usability of virtual reality applications.
Deep learning is widely used to forecast cybersickness to inform mitigation efforts. However, prior work mostly focused on predicting the Fast Motion Sickness score, which does not provide detailed information on cybersickness symptoms. To fill this gap, this paper focuses on forecasting cybersickness symptoms and evaluates the effectiveness of multi-task learning. The results show that multitask learning does not always outperform the traditional single-task learning approach in this domain.
Cybersickness, characterized by discomforts such as dizziness, nausea, and eye strain, remains a significant barrier to the widespread adoption of virtual reality (VR). Recent research have proposed supervised machine learning models to predict the onset of cybersickness; however, these approaches depend heavily on labeled datasets. Acquiring labeled datasets typically necessitates time-consuming and resource-intensive user studies, limiting the feasibility of these supervised methods for consumer-level VR applications where obtaining labeled user data during use is impractical. Moreover, due to individual differences, often these datasets are not generalizable in consumer VR use. To address these limitations, we propose a novel semi-supervised learning framework for predicting cybersickness (i.e., Fast Motion Sickness (FMS)) using eye tracking, heart rate, and galvanic skin response data. Our proposed semi-supervised approach uses pseudo-labeling techniques (i.e., Self-training, Label Propagation, and Label Spreading) fused with temporal deep learning models (i.e., DeepTCN, CNN-LSTM, Transformers, LSTM). We evaluated our approach on three public cybersickness datasets (i.e., Bumpy Ride, Simulation 21, Maze) and our proposed semi-supervised approach demonstrates strong cybersickness predictive performance using only $1-5 {\%}$ labeled data (i.e., $95-99 {\%}$ data remains unlabeled). Notably, the self-training approach with a DeepTCN model achieved an accuracy of $\mathbf{7 5. 8 6 \%}$ in FMS prediction, outperforming the other models and pseudolabeling approaches. Our findings establish the viability of semisupervised learning for cybersickness prediction with minimally labeled datasets, paving the way for more practical and potentially generalizable cybersickness prediction systems in consumer VR applications.
This study explores perspectives on diversity, equity, inclusion, and accessibility (DEIA) from researchers in the IEEE VR community. Fourteen participants expressed sustained commitment to DEIA, noting underrepresentation of women, BIPOC, individuals with disabilities, and researchers from developing countries or low socioeconomic backgrounds. Participants perceived high costs, lack of accessibility features (e.g., subtitles, mobility support), and insufficient family resources (e.g., childcare) as barriers. Among factors that could deter support of DEIA listed were fear of retaliation, financial constraints, structural challenges, and difficulties identifying diverse representatives. Strategies participants supported included reporting demographic statistics, strategic planning, and promoting DEIA in speakers and attendees.
Cybersickness remains a significant challenge in virtual reality (VR), with various machine learning (ML) and deep learning (DL) methods proposed for its detection. In recent works, Explainable AI (XAI) has emerged as an effective solution in detecting and explaining cybersickness features while reducing the feature space for efficient VR deployment. This paper argues that since XAI can identify dominant features in a DL model for predicting cybersickness, this XAI capability can also guide the selection of cybersickness mitigation strategies. This is in contrast to adopting a fixed/static cybersickness mitigation in state-of-the-art literature. Towards this, we present ExciteVR, an XAI-guided interactive framework for predicting, explaining, and mitigating cybersickness. Our system uses an ensemble ML model deployed on a consumer VR headset (HTC Vive Pro Eye), leveraging eye and head-tracking data for cybersickness detection. Upon detection, a mitigation engine uses XAI-guided feature importance to select the appropriate technique from a library of three mitigation methods. We further develop a voice-based dialogue system powered by large language models to help users understand detection outcomes and choose mitigation strategies. A user study with a custom VR roller coaster simulation demonstrates that ExciteVR effectively reduces cybersickness with minimal impact on immersion, and 91% of participants found the framework easy to use and highly effective.
B Lok合作论文数Computer and Information Science and Engineering Department
College of Engineering
University of Florida3