
SMART VR is an intelligent training interface that combines immersive virtual reality with real-time biofeedback to support stress regulation in firefighter simulations. The system employs machine learning algorithms to interpret biometric data and adjust scenario parameters dynamically, creating a responsive and personalized training experience. Evaluation with professional trainees shows improved stress awareness and performance under pressure. This work contributes to the field of intelligent user interfaces by showcasing a novel application of physiological computing in adaptive training systems.
We present a learning-support system that integrates physical three-dimensional models with augmented reality (AR) to facilitate the acquisition of spatial-figure concepts. The system is implemented on an iPad, which functions as both input and output device. Following instructions displayed on the screen, users mentally construct the spatial figure that satisfies a displayed problem and then place vertex markers along the edges of a physical cube model. The iPad's rear camera tracks these markers, and the system superimposes the user-defined spatial figure onto the model in real-time AR, enabling the learner to examine the figure from multiple viewpoints. While interacting with the system, data-including responses, completion times, and pre- and post-task questionnaire scores-are recorded. These data are subsequently analyzed to evaluate learning outcomes and to inform iterative refinement of instructional content and exhibit design. An evaluation experiment was conducted at a university festival. Nineteen participants completed questionnaires assessing ease of use, perceived effectiveness for learning spatial figures, and interest. The results indicated generally positive evaluations. Although the system did not yield immediate, statistically significant learning gains-likely due to operational difficulty and the inherent challenge of the material-the integration of AR with tangible models appears promising for stimulating learner engagement in the comprehension of spatial figures.
Accurately predicting head motion is essential for reducing motion-to-photon (MTP) latency in XR systems. However, standard evaluation practices often conflate true prediction error with residual calibration artifacts-such as global rotation bias, lever-arm offsets, or sub-frame time skews-even after alignment. These small but realistic miscalibrations can dominate reported errors and obscure the actual performance of learningbased predictors, creating a gap in fair and robust benchmarking. This paper introduces a calibration-robust learning and evaluation pipeline for XR head-motion prediction that explicitly addresses these residuals. By rethinking both the training targets and evaluation protocol, our approach disentangles model accuracy from calibration artifacts, enabling more meaningful comparisons across devices and algorithms. Experiments on the public ADVIO dataset and our HMD-MoCap sessions show that body-frame targets consistently reduce calibration sensitivity for data-driven prediction models-yielding lower calibration sensitivity slopes (CSS) and higher robust area under the curve (AUC)-while maintaining short-horizon rotational accuracy compared to conventional world-frame approaches. Our findings establish a protocol for robust, reproducible benchmarking in XR pose prediction and demonstrate the feasibility of ground-truth-level head-motion prediction for EdgeXR.
As immersive technologies become more prevalent in entertainment, education and industry, the manual creation of 3D content remains a major bottleneck, demanding time, specialized labor and significant computational resources. This challenge is directly related to Sustainable Development Goal (SDG) 12, which calls for more sustainable consumption and production patterns, including more efficient use of materials, energy and human effort. In this work, we present an AI-based XR pipeline that automatically generates interactive VR scenes from a single RGB image, reducing the need for extensive manual 3D modeling. Our approach combines state-of-the-art 3D object detection, single-image mesh reconstruction and virtual reality visualization on consumer hardware (Meta Quest 3). Using images from the SUN RGB-D dataset and models such as Total3DUnderstanding, Implicit3DUnderstanding, Unique3D and Hunyuan3D-v2, we demonstrate how everyday environments can be rapidly digitized into VR-ready objects and scenes. Preliminary results indicate that recent reconstruction models substantially improve mesh quality and geometric fidelity, enabling reusable, editable assets that support sustainable XR applications, such as virtual product exploration, sustainable tourism experiences, and learning environments focused on consumption and waste. Our evaluation is qualitative and exploratory in nature, focusing on visual fidelity, reconstruction plausibility, and end-to-end feasibility rather than on numerical benchmarking. We discuss how this type of pipeline can contribute to SDG 12 by lowering the content-production barrier, encouraging the reuse of digital assets, and enabling end-users to transform their own spaces into immersive XR experiences without the need for large-scale production pipelines.
This paper presents a prototype of a mixed reality (MR) museum artifact preview system, enabling visitors to view virtual artifacts on real pedestals before entering the exhibit. The system uses MR devices to display 3D artifact previews with guide avatars and AI-generated introductory text, enhancing pre-visit engagement.
This paper investigates in an exploratory manner small group interactions in Social VR classroom conversations using eye-tracking data (group eye-gaze configurations, blink rate and blink duration) in terms of task engagement. Task engagement has been previously studied in educational settings as it is an important construct for learning. Prior research showed that eye signals in groups can assess rapport, attention, agreement or communication engagement. However, little research focused on whether eye-tracking data can predict task engagement in educational small group Social VR conversations. We conducted an exploratory study with 52 pedagogy university students (8 dyads, 12 triads) engaging in free-flow conversations on a given topic. Results indicated that eye-tracking data could predict task engagement using multiple linear regression analysis. However, group eye-gaze configurations were the significant predictors in dyads, whereas the blink rate was the strongest predictor associated with task engagement in triads. Based on these findings, we discuss the higher complexity in triadic gaze dynamics and why the blink rate was the task engagement predictor. On the other hand, the simplicity of dyadic gaze configurations might as well explain why it was found as a predictor. Furthermore, we argue that future work can explore machine learning detectors of task engagement dependent of group size. The results from this exploratory study help understanding non-verbal group interactions in VR and contribute to further work on engagement models for more socially-abled virtual experiences in Social VR applications.
Object detection models have recently become lightweight enough for deployment on resource-constrained devices such as smartphones and Augmented Reality (AR) headsets. Techniques such as few-shot, zero-shot, and incremental learning have moved toward lower training times and faster inference for detection models, but most frameworks still require substantial training or retraining to dynamically add new object classes in real-time for practical use with AR. In this paper, we present a novel AR-integrated object detection system that leverages Visual Retrieval Augmented Generation (ARVRag for short), eliminating much of the lengthy and expensive training processes associated with traditional object detection models. By vectorizing a smaller set of training images and conducting a local similarity search, we achieve enough accuracy for practical use with AR in significantly less time than state-of-the-art models like YOLO. Moreover, we have built this process into support systems for maintenance and industrial operations. End-users can select an object using a 2 D image capture interface that automatically retrieves semantically similar items from a local corpus (database), and the resulting detection provides metadataenhanced object-specific responses from a large language model (LLM). This work demonstrates a practical and scalable pathway for LLM-supported object detection in AR applications, especially for manufacturing and repair.
Virtual Reality representations of distant planets, stars, and galaxies are a driving factor for research, scientific dissemination, and public outreach. Immersive environments can be used to train astronauts, plan future missions, or study planetary features. Moreover, the dissemination of these endeavors through exhibits in museums and planetariums makes it possible to effectively communicate the latest research findings to a wider audience. To correctly visualize outer space, in particular the surfaces of celestial bodies, virtual environments must rely on scientifically accurate datasets derived from telescopes and spacecraft, which often present missing or degraded values. The current work analyzes the processing pipeline for extraterrestrial terrain data representation in VR, highlighting the most frequent errors that result in data degradation, along with their negative impacts on visualization and usability. The current work presents an approach for restoring missing terrain data using generative diffusion models. Results show visually consistent 3D reconstructions, suitable for representation in Virtual Reality, also supported by positive preliminary quantitative assessments (39.4495 in PSNR and 0.9660 in SSIM).
Although commercial metaverse platforms have advanced social implementation, they struggle to retain users. Most new users abandon these platforms shortly after their initial experience, which hinders the realization of the metaverse's social and economic potential. Although AI-agents could promote user interaction and retention by acting as social catalysts, existing research has focused on laboratory-based short-term validation. This approach provides limited evidence of the longterm effectiveness of agents in commercial environments. This study examines the impact of Large Language Model (LLM)-based AI-agents on user continuation behavior by operating AI-agents for 31 days on the commercial metaverse platform Cluster and observing the natural usage behavior of $\mathbf{5, 0 2 0}$ unique users. The analysis used two complementary approaches: (1) an aggregate effect analysis at the weekly habit formation level and (2) an within-user effect analysis at the daily decision-making level with individual difference controls. The results revealed that interaction with the AI-agent produces lasting changes in user behavior, especially among new users. Furthermore, the cumulative relationship-building process through continuous contact opportunities rather than single impressive experiences with AI-agents was found to decisively affect users' continued metaverse usage.
Debriefings are a crucial component of simulation-based learning, as they allow participants to critically reflect on their performance and translate experiences into knowledge. With the rise of artificial intelligence (AI), large language models (LLMs) are increasingly being explored as conversational agents in debriefings. However, concerns have been raised that AI debriefers may display sycophantic tendencies, characterized by overly positive sentiment and a lack of critical feedback. Such tendencies could affect learners' self-evaluations and limit the educational value of AI-supported debriefings. In this study, 45 university students participated in a virtual reality (VR) counseling simulation, followed by either an AI-guided or a human-led debriefing. Sentiment analysis of 958 debriefing statements (350 AI, 608 human) was conducted using a German sentiment classification model based on Google's BERT. Results show that AI debriefers expressed significantly more positive sentiment and substantially fewer negative statements compared to human debriefers, corroborating the assumption of a sycophancy effect. However, sentiment was not found to predict self-rated counseling competence, or self-efficacy. These findings highlight both the promise and the limitations of AI-supported debriefings in counseling education. While AI debriefers provide consistent encouragement, their lack of critical feedback raises questions about their long-term effectiveness for professional skill development. Future research should explore hybrid approaches that combine AI scalability with human expertise in constructive reflection.
Training robots in real-world is costly and time-consuming. Consequently, recent research has increasingly used simulator-based virtual environments for robot training. 3D Gaussian Splatting, which can generate fast and photorealistic 3D scenes from multi-view images, is particularly suitable for constructing such virtual environments. However, due to the inherent characteristics of splats, artifacts such as floaters and noise frequently arise, it is unsuitable to make the raw representation for simulators that rely on physics engines and collision detection. To address these limitations, we propose a pipeline composed of YOLO-based floor segmentation, planar estimation, and mesh reconstruction. The proposed method reliably identifies and flattens the floor region in the 3D Gaussian Splatting model and converts only the non-floor regions into meshes. This approach effectively reduces surface irregularities and collision errors that commonly occur in conventional gaussian-to-mesh conversion methods. When we deployed in a Unity-based simulation environment, the reconstructed mesh significantly reduced floor-related noise and improved the stability of mobile robot navigation compared to the baseline.
We built an immersive virtual reality environment where preservice nurses practice with generative artificial intelligence-driven virtual standardized patients (AI-VSPs). We ran mixed-methods study (experimental $n=23$; control $n=34$). Usability score was $81.6 / 100$, indicating excellent usability. A single session increased self-efficacy from 3.74 to $4.11(t(22)=3.82, p<0.001$). In the standardized-patient (SP) encounter performance assessment, scores for AI-VSP users did not differ from controls $(t(55)=-0.134, p=0.894)$. AI-VSP encounter performance did not relate to engagement, presence, or immersion, although these experience measures were strongly intercorrelated. Interviews highlighted helps and hindrances. Overall, AI-VSPs are a usable, confidence-building supplement to SP encounter preparation. We also outline practical design steps: cap reply length, enable reliable barge-in, improve speech quality and latency, and keep scaffolding minimal.
As personal digital media accumulate across devices and cloud services, they become abundant yet intangible, weakening the rituals that once supported remembering. Graspable Memories is an AI-based XR system built on Embodied Projected Mixed Reality (EPMR) that explores how projected interaction and bodily occlusion can restore tangibility to digital memories while advancing United Nations Sustainable Development Goal 12 on responsible consumption and production. Using a fixed projector-camera setup with AI-based hand tracking, the system projects personal media onto a tabletop, where they can be selected by occluding them with an open palm, moved through a volumetric interaction space, and transferred with a grasp gesture to a wearable pendant with a reconfigurable display. This dematerialized workflow offers an alternative to physical photo printing and disposable keepsakes, directly engaging SDG 12 targets on waste reduction and more sustainable consumption. An exploratory study with twelve participants, using the User Experience Questionnaire and semi-structured interviews, examines usability, emotional engagement, and the symbolism of hand-topendant transitions. Findings highlight strong perceived novelty and personal resonance, along with current tracking limitations, and point to design implications for human-centered, sustainable Industry 5.0 XR experiences that treat the body and occlusion as primary elements of the interface.
Navigating unfamiliar indoor environments poses a significant challenge for people with visual impairments, especially in areas with unpredictable or overhead obstacles that cannot be detected by traditional aids such as canes or guide dogs. In this paper, we present a low-cost portable sensory substitution device (SSD) that enables users to perceive spatial depth using audio. Our method uses a novel real-time sonification technique based on the Hilbert curve, a space-filling curve that maps twodimensional depth images into a one-dimensional audio signal. Unlike prior approaches that sonify images through sequential scanning, our system represents entire scenes simultaneously by assigning frequencies and amplitudes to pixels based on position and depth. We show that this approach enables real-time audio rendering on resource-constrained hardware such as the Raspberry Pi 5 through an efficient multithreaded implementation. Our system fills a critical perception gap between obstacles at the waist and head level, offering a new form of spatial awareness to visually impaired users. In addition, we propose a virtual training environment to ease the steep learning curve of smart assistive devices. With a user study, we show that the virtual training environment may reduce the barriers to adoption. Our contributions lay the groundwork for more intuitive and efficient SSDs that can enhance independent navigation without costly hardware.
Mixed Reality Navigation (MRN) has been proposed as a low-cost and intuitive alternative to conventional navigation systems, yet its adoption remains limited by registration inaccuracy and strong user dependence. This work presents a Hybrid Conical-Linear Laser Registration (HCLR) method that combines a conical projection onto five coplanar fiducials with a single linear-laser constraint to recover all six degrees of freedom analytically. A systematic evaluation of 57 marker layouts identified three robust configurations, which were tested in simulated scenarios using clinical CT and MRI datasets from 19 patients. Under realistic noise conditions, HCLR achieved a mean target registration error (TRE) of 1.06 mm compared with 2.57 mm for standard fiducial-based registration, and nearly eliminated rotation errors. These findings demonstrate that HCLR can improve registration accuracy, reduce operator dependence, and provide a compact, hardware-light solution for MRN. With further physical and clinical validation, HCLR has the potential to advance cost-effective neurosurgical navigation.
Stroke often leads to significant upper-limb motor impairments, necessitating long-term rehabilitation. However, traditional methods depend on therapist supervision or complex robotic systems, which limit their accessibility and scalability. To address these issues, this study proposes an AI-driven virtual rehabilitation system comprising real-time posture tracking, a virtual doctor with voice interaction, and a plug-in game to enhance engagement. Specifically, the proposed system architecture leverages a ZED 2i Stereo camera for real-time motion tracking and incorporates a DeepSeek assistant module to provide personalized feedback and adjust game difficulty levels. A proof-of-concept experiment involving a healthy participant was conducted to validate the system's architectural integrity, realtime latency, and interaction logic. Preliminary results confirm the system's technical feasibility and suggest it holds promise for future clinical use, pending further validation and optimisation.
This work investigates the detection of digitally manipulated faces in an interactive Virtual Reality (VR) environment, extending previous work on static, screen-based tasks. Participants $(N=81)$ completed a face manipulation detection task in VR while their head movement and gaze data were recorded. Contrary to our hypothesis, overall accuracy (56.6%) was lower than in the static-lab benchmark, and neither head movements nor the frequency of gaze shifts between images correlated with performance. However, gaze density analysis revealed a clear strategic difference between high- and low-performing participants. High-performers tended to employ a holistic scanning strategy, focusing on the central T-zone (eyes, nose, and mouth). These findings suggest that the effectiveness of manipulation detection is less dependent on the interactivity of the viewing modality and more on the observer's underlying visual search strategy, offering a promising research direction for both human and algorithm training.
We introduce DreamAble, a novel framework that uses diffusion models to generate 3D avatars for individuals with upper limb differences and amputations in augmented and virtual reality. Existing systems often overlook these users, forcing them to choose between normative body models or completely avoiding self-expression. DreamAble addresses this gap through a novel end-to-end pipeline that combines text-guided structural editing with skeletal-guided generation. Experimental evaluations demonstrate that DreamAble is competitive with existing methods in semantic alignment, achieving an average CLIPScore of 29.05. It accurately represents everyday users and professional athletes with limb differences, supporting both participation in Metaverse communities and the development of digital Paralympic games.
Real-time breathing input in Virtual Reality (VR) offers new possibilities for accessible, hands-free interaction and immersive biofeedback. However, existing sensing mechanisms often struggle to combine comfort, reliability, and seamless VR integration to produce valuable input data. This work introduces a novel respiratory input method using thermal imaging and a thin medium aligned with the user's exhaled airflow. The system captures and processes cross-sectional heat signatures of the user's exhale to identify distinct “exhale gestures”, enabling realtime, contact-free breath-based control in virtual environments. This approach is particularly promising for interactive breathing exercises and accessible controls, which we demonstrate with a VR maze navigable solely using exhale gestures as input.
It is known that a user's personality influences the genres of games they prefer, but the relationship between personality and game mechanics has been less investigated. However, discovering such relationships is important for understanding how adjusting game difficulty based on users' personalities can enhance their intrinsic motivation. This paper investigates the relationship between the Big Five personality traits and difficulty adaptation methods constructed using static difficulty curves in a VR Kendama game (a cup-and-ball game). The results reveal trend-level correlations between specific personality traits within the Big Five factors and indicators of intrinsic motivation—namely, enjoyment and self-efficacy—across three different difficulty adaptation methods. This paper also suggests that when users with specific personality traits play the game, appropriately selecting difficulty adaptation methods may enhance both their enjoyment and self-efficacy through the findings of these correlations.