![International Symposium on Mixed and Augmented Reality : (ISMAR) [proceedings]](https://originalfileserver.aminer.cn/sys/aminer/magazine.png)
We are witnessing the dawn of Distributed AI, where intelligence is no longer confined to centralized clouds but extends to the edge of networks. Edge AI now encompasses Physical AI that interacts with the real world and Agentic AI that acts autonomously, bringing intelligence closer to people, devices, and environments. Concurrently, Cloud AI provides the scale, coordination, and learning capacity necessary to connect these edge systems. This keynote will highlight how the rise of Distributed AI signals the beginning of a new era—one in which intelligence emerges through the interplay of Edge and Cloud, with connectivity as the key enabler of progress.
In the era of multi-modal Artificial Intelligence (AI), the convergence of AI and eXtended Reality (XR) is revolutionizing immersive technologies. In this talk, I will review the technical challenges in the XR industry over the last decade, and show how we have overcome those by exploring the pivotal role of on-device AI, multi-modal AI and computer vision in advancing XR experiences, especially on various XR form-factors including Project Moohan - the first device utilizing a new AndroidXR platform, an initiative bridging AI and XR to redefine user-centric real-world XR applications. The talk will underscore how multi-modal AI integrates diverse data streams-visual, spatial, and contextual- to create seamless, intelligent, and immersive XR experiences while synergies between software and hardware innovations are also emphasized. Attendees will gain insights into Samsung's vision for next-generation devices and the transformative potential of these technologies across various industries.
CALIVERSE is a next-generation AI metaverse platform that seamlessly combines immersive shopping, live performances, gaming, and Web3 user-generated content into a unified virtual experience. Built on cutting-edge technologies such as Unreal Engine 5, AI-powered lighting, VR AI enhancement, real time live-action rendering integration, and 3D AI conversion, CALIVERSE delivers an incredibly lifelike virtual world. The platform supports multiple environments, including PC, VR HMD, and glasses-free 3D displays (Screen Protector), delivering high-fidelity 4K+ virtual experiences with premium quality and deep immersion. In particular, CALIVERSE seamlessly composites live-action footage with real-time rendered graphics and utilizes AI lighting and image enhancement technologies to deliver natural and visually coherent virtual performances. It can depict audiences of over 80,000, setting it apart from conventional metaverse platforms designed for younger users. By integrating Korea's advanced server networking technology, CALIVERSE enables massively scaled simultaneous connections, setting a new standard for ultra-immersive next generation metaverse platforms.
This work addresses the limitations of traditional 2D collaboration tools in automotive design, where complex 3D model reviews require immersive spatial interaction. We developed a metaverse-based collaborative platform for automotive design processes using Unreal Engine 5, addressing three research directions: (1) core requirements for automotive design collaboration, (2) effectiveness of immersive technologies in such professional workflows, and (3) an empirical user study of test scenarios comparing the platform as a conventional PC application to an immersive Virtual Reality (VR) experience. Xstudio enables real-time interaction with 3D models and advanced functionalities for design teams. We evaluated the effectiveness of the tool through user studies and technical performance analysis. The expert study of 11 employees, who were part of the engineering department of BMW (a European car manufacturer), was designed to evaluate the usability of the features using the System Usability Score (SUS), which scored a promising value of 76 (VR) and 74 (PC). The NASA Task Load Index (NASA-TLX) resulted in 40.1 for VR and 37.7 in PC mode. A set of quantitative analyses with factors such as task completion time (TCT), error rates (ER), and user interaction metrics (UIM) were applied in the comparison of VR and PC modes. The result shows that the VR mode has higher TCT than the PC mode, but in contrast, VR has a higher interaction rate to complete the task, input precision, and the multi-step manipulation can be optimized. We summarize the findings in guidelines for future research in the collaborative car design process. Key technical contributions include a modular Entity-Component-System architecture for dynamic tool customization and hybrid VR/PC interactions. In the domain-specific automotive engineering space, both application types complement each other, offering efficiency (PC) or immersion (VR) for specific tasks. Future work will focus on UI/UX refinements with rigorous validation in follow-up studies. Specific improvements will include streamlined menus, snap-based selection aids, and pointer-locking tools to address precision challenges.
Shared synchronised activities, such as dancing and singing, can promote social bonding in both physical and virtual settings, but are susceptible to network latencies in mediated environments, such as social virtual reality (VR). Temporal delays can hinder interpersonal entrainment, causing users to feel out of sync, and reduce presence. While prior work has examined how stimulus type and frequency affect synchrony perception, the role of rhythmic interaction, central to many interactive synchronised experiences, remains unexplored. We conducted a within-subject study with 32 participants to systematically investigate how rhythmic interaction shapes synchrony perception of rhythmic audiovisual stimuli in networked VR environments. Motor behaviour and subjective synchrony ratings ($\mathrm{n}=32$) are analysed, complemented by eye-tracking metrics from a quality-controlled subset ($\mathrm{n}=25$). Through controlled stimulus onset asynchronies (SOAs) between audio and video stimuli, we simulate network latency scenarios wshere (1) audio precedes video (as in VR dance environments where music playback is synchronised to a global clock while remote dancer movements appear delayed), and (2) video precedes audio (as in orchestra experiences where a conductor's movements are perceived first, followed by a delayed auditory response from the orchestra). Participants either synchronised hand movements with a virtual avatar or passively observed it across two stimulus frequencies (0.5 Hz and 1 Hz), for both audio- and video-leading offsets. Our results show that rhythmic interaction significantly increases tolerance to audiovisual offsets, shifting the left decision criterion (LDC) by up to 66 ms, the point of subjective synchrony (PSS) by up to 31 ms, and widening the window of subjective synchrony (WSS) by up to 59 ms. These effects suggest that rhythmic interaction can reduce sensitivity to asynchronies in audio-leading scenarios, such as VR dance environments. We also found that temporal offsets influence rhythmic interaction behaviour and confirm that recent findings on the influence of stimulus frequency on PSS and temporal offsets on pupil dilation can to a certain extent be replicated in interactive VR.
The multi-directional tapping task has long served as a foundational tool for evaluating pointing performance in human-computer interaction research. However, its transition from 2D interfaces to virtual reality (VR) raises challenges, especially in standardizing target depth. This study explores how target depth influences performance in VR, focusing on two common techniques: Raycasting and Virtual Hand. We conducted two controlled experiments (each with $\mathrm{n}=20$) to isolate depth effects. In Experiment 1, fixed target size led to visual angle (VA) shifts across depths, affecting performance. Both techniques performed best when VA was between $1-4^{\circ}$; Raycasting peaked at 2 m, Virtual Hand at $0.4-0.5 ~\mathrm{m}$. In Experiment 2, we controlled VA to isolate depth itself. Raycasting remained stable beyond 2 m but degraded at close range due to biomechanical limits. Virtual Hand remained sensitive to depth despite fixed VA, but differences were smaller, with throughput unaffected. These results suggest VA should be the primary parameter for standardizing the task in VR. Depth-specific evaluation remains necessary, except for Raycasting beyond 2 m. We provide depth-aware guidelines to improve standardization and comparability while aligning with ISO protocols.
Multi-projector displays show significant variation of brightness and black offset. Prior work use single-camera based methods to create a seamless display. However, the limited field-of-view (FOV) of single-camera based solutions limit their applicability to large, non-flat display shapes such as surround displays or domes. Although using a single camera at multiple locations can provide some ad-hoc corrections, it lacks both automation and accuracy. Further, automated black offset seamlessness is arguably still one of the most challenging problems in multi-projector displays. In this paper, we present the first scalable brightness and black offset seamlessness method for multi-projector displays using multiple uncalibrated cameras that scale to a large number of projectors on developable display geometries, and can operate in the presence of camera vignetting effects and varying projector brightness. Using a display partitioning method, we relate the brightness response for every pixel of the display to a common reference via appropriate camera correction factors. Then, we adapt a prior single-camera-based brightness seamlessness method to achieve seamlessness across the entire display. Next, we show that the same technique can be used to provide an automated method to achieve black offset seamlessness across the display. We demonstrate the scalability of our method to displays with upto 32 projectors on complex developable projection surfaces with cameras that have different color properties. We also provide quantitative results to demonstrate the significantly positive impact of the brightness and black offset seamlessness methods on the quality of the display.
Managing negative emotions through emotion regulation (ER) is key to mental well-being. While virtual reality (VR) shows promise for supporting ER, prior work has primarily focused on selfregulation. This paper introduces a virtual agent that helps users manage emotions through conversation-based ER strategies. We compared three conditions: no agent, an agent with non-supportive responses, and an agent with ER-supportive responses. Results showed that the ER-supportive agent significantly improved users' emotional states and overall experience. Building on this, we conducted a second experiment to examine how the agent's appearance (realistic vs. cartoon) and voice tone (emotional vs. neutral) affect ER. Results indicated that an emotional voice tone improved users' ability to regulate emotions. Although a realistic appearance did not directly improve ER, it increased users' trust and sense of social presence. This paper contributes to VR and human-agent interaction by demonstrating the potential of virtual agents to support ER and offering design implications for future ER-supportive agents.
Video pass-through (VPT) techniques often cause visual distortions like the ‘telescope eye' effect, which leads to less immersion and spatial awareness in Extended Reality (XR). This study investigates the potential of 3D Gaussian Splatting (3DGS) as a rendering alternative to video pass-through, particularly focusing on its ability to mitigate common perceptual artifacts and enhance immersion. To evaluate the efficacy of 3DGS in an XR context, we conducted a user study comparing two visualization types: (1) conventional VPT and (2) 3DGS-based scene representation. Three distinct scenes with high detail, natural depth perception, and reflective surfaces were included in the study. An empirical study was conducted with a structured questionnaire that inquired about the level of immersion and the impact of the visual stimuli. Our results provide insights into the scalability and hardware considerations for implementing 3DGS in realtime interactive environments. The evaluation results indicated that users with a 3DGS visualization experienced a reduction in visual distortion and an enhancement in the perceived realism of objects. The findings show that 3DGS can serve as an effective alternative to VPT, potentially improving realism and reducing perceptual artifacts in XR applications.
The study of cognitive load (CL) has been an active field of research across disciplines such as psychology, education, and computer graphics and visualization for decades. In the context of Virtual Reality (VR), understanding mental demand becomes particularly relevant, as immersive experiences increasingly integrate multisensory stimuli that require users to distribute their limited cognitive resources. In this work, we investigate the effects of cognitive load during a search task in VR, combining objective and subjective measurements, including physiological signals and validated questionnaires. We designed an experiment in which participants performed a visual search task under two cognitive load conditions (either alone or while responding to a concurrent auditory task) and across two visual search areas (90° and 360°). We collected a rich dataset comprising task performance, eye tracking, electrocardiogram (ECG), electrodermal activity (EDA), photoplethysmography (PPG), and inertial measurements, along with subjective assessments (NASA-TLX questionnaires). Our analysis shows that increased cognitive load hinders visual search performance and affects multiple physiological markers, offering a solid foundation for future research on cognitive load in multisensory virtual environments.
Active reading is crucial for the acquisition of information and the comprehension of literature. However, existing digital media still encounter several challenges in facilitating active reading, such as the limited contextual space caused by constrained screen size, which hinders effective information integration. Furthermore, the rationality of the layout of various types of literature resources also significantly influences the coherence of the reading experience. To address these challenges, we propose DocVision, a seamless, cross-device immersive active reading framework. To reduce the difficulty of information integration, DocVision incorporates a lightweight resource acquisition framework, which seamlessly integrates resources such as diagrams and literature references into the mixed reality (MR) environment. Furthermore, we propose a dynamic layout method for MR to improve reading continuity in multi-resource contexts. Experimental results demonstrate that DocVision substantially enhances the coherence of readers' reading experience and significantly alleviates cognitive load, offering valuable insights for developing future active reading approaches based on digital media.
As Virtual Reality (VR) devices become increasingly shared among users, there is a pressing need for authentication methods that balance security, usability, and privacy while accommodating VR's unique interaction constraints. This paper presents cLock, a novel single-handed two-factor authentication technique in VR that allows users to enter PINs with multiple cursors on a virtual circular numpad through wrist rotation and finger tapping. We first optimized the UI design of cLock by comparing participants' input performance with different UI parameters. We then extracted spatiotemporal behavioral features of both fingers and palm during PIN entry, which facilitated cLock's authentication algorithm. In the usability evaluation with four input postures, cLock achieved significantly faster authentication speed than laser and touch-based baselines, without sacrificing accuracy. Meanwhile, it was most preferred by participants in terms of privacy, social acceptance and physical effort. A following evaluation of security demonstrated that cLock achieved a deciphering rate of only $1 / 8$ of the baselines against shoulder-surfing within 1 m. Even in scenarios of password leakage, cLock could still achieve an FAR of 2.3% and FRR of 3.2% with 20 registered users. A final 11-day study verified the longitudinal stability of cLock.
Cybersickness, characterized by discomforts such as dizziness, nausea, and eye strain, remains a significant barrier to the widespread adoption of virtual reality (VR). Recent research have proposed supervised machine learning models to predict the onset of cybersickness; however, these approaches depend heavily on labeled datasets. Acquiring labeled datasets typically necessitates time-consuming and resource-intensive user studies, limiting the feasibility of these supervised methods for consumer-level VR applications where obtaining labeled user data during use is impractical. Moreover, due to individual differences, often these datasets are not generalizable in consumer VR use. To address these limitations, we propose a novel semi-supervised learning framework for predicting cybersickness (i.e., Fast Motion Sickness (FMS)) using eye tracking, heart rate, and galvanic skin response data. Our proposed semi-supervised approach uses pseudo-labeling techniques (i.e., Self-training, Label Propagation, and Label Spreading) fused with temporal deep learning models (i.e., DeepTCN, CNN-LSTM, Transformers, LSTM). We evaluated our approach on three public cybersickness datasets (i.e., Bumpy Ride, Simulation 21, Maze) and our proposed semi-supervised approach demonstrates strong cybersickness predictive performance using only $1-5 {\%}$ labeled data (i.e., $95-99 {\%}$ data remains unlabeled). Notably, the self-training approach with a DeepTCN model achieved an accuracy of $\mathbf{7 5. 8 6 \%}$ in FMS prediction, outperforming the other models and pseudolabeling approaches. Our findings establish the viability of semisupervised learning for cybersickness prediction with minimally labeled datasets, paving the way for more practical and potentially generalizable cybersickness prediction systems in consumer VR applications.
The convergence of artificial AI and XR technologies (AI XR) promises innovative applications across many domains. However, the sensitive nature of data (e.g., eye-tracking) used in these systems raises significant privacy concerns, as adversaries can exploit these data and models to infer and leak personal information through membership inference attacks (MIA) and re-identification (RDA) with a high success rate. Researchers have proposed various techniques to mitigate such privacy attacks, including differential privacy (DP). However, AI XR datasets often contain numerous features, and applying DP uniformly can introduce unnecessary noise to less relevant features, degrade model accuracy, and increase inference time, limiting real-time XR deployment. Motivated by this, we propose a novel framework combining explainable AI (XAI) and DP-enabled privacy-preserving mechanisms to defend against privacy attacks. Specifically, we leverage post-hoc explanations to identify the most influential features in AI XR models and selectively apply DP to those features during inference. We evaluate our XAI-guided DP approach on three state-of-the-art AI XR models and three datasets: cybersickness, emotion, and activity classification. Our results show that the proposed method reduces MIA and RDA success rates by up to 43% and 39%, respectively, for cybersickness tasks while preserving model utility with up to 97% accuracy using Transformer models. Furthermore, it improves inference time by up to ~2x compared to traditional DP approaches. To demonstrate practicality, we deploy the XAI-guided DP AI XR models on an HTC VIVE Pro headset and develop a user interface (UI), namely PrivateXR, allowing users to adjust privacy levels (e.g., low, medium, high) while receiving real-time task predictions, protecting user privacy during XR gameplay.
Augmented Reality (AR) surgical navigation systems are emerging as the next generation of intraoperative surgical guidance, promising to overcome limitations of traditional navigation systems. However, known issues with AR depth perception due to vergenceaccommodation conflict and occlusion handling limitations of the currently commercially available display technology present acute challenges in surgical settings where precision is paramount. This study presents a novel methodology for utilizing AR guidance to register anatomical targets and provide real-time instrument navigation using placement of simulated external ventricular drain catheters on a phantom model as the clinical scenario. The system registers target positions to the patient through a novel surface tracing method and uses real-time infrared tool tracking to aid in catheter placement, relying only on the onboard sensors of the Microsoft HoloLens 2. A group of intended users performed the procedure of simulated insertions under two AR guidance conditions: static in-situ visualization, where planned trajectories are overlaid directly onto the patient anatomy, and real-time tool-tracking guidance, where live feedback of the catheter's pose is provided relative to the plan. Following the insertion tests, computed tomography scans of the phantom models were acquired, allowing for evaluation of insertion accuracy, target deviation, angular error, and depth precision. System Usability Scale surveys assessed user experience and cognitive workload. Tool-tracking guidance improved performance metrics across all accuracy measures and was preferred by users in subjective evaluations.
Augmented reality (AR) applications present virtual elements that are connected to the real world. While these may be aware of a user's geographical or visual context, the real sounds in a user's environment are rarely used in the AR experience. We investigate audio augmented reality (AAR) with sound as the primary output. We present the first evaluation of sonic linking, where an AAR application uses real sounds (bird calls and car engines) in a user's surroundings to drive the interaction. We developed two AAR applications to investigate how to design and use such links: a game where entities are spawned based on real-world sounds and a music player with sound-reactive filtering and volume adjustment. Design variations are compared to cover different AAR scenarios, types of sonic link, and existing unlinked equivalents. The results show that sonic linking can create a more augmented, engaging AAR experience, and may alter a user's relationship with their real-world surroundings, enabling new types of augmented reality applications.
Hand-tracking based 3D object manipulation in Extended Reality (XR) typically employs a pinch gesture for acquisition and manipulation through a direct mapping from 6-degrees-of-freedom (DOF) hand movement to that of the object. In this work, we investigate the effect of separating this mapping to concurrent 3DOF controls (DOF-Separation) of translation and rotation of the virtual object using the position and orientation of the hand independently. We aim to understand how DOF-Separation could ease manipulation for different techniques with varying requirements for hand position and orientation during acquisition, including Virtual Hand, Hand Ray, and Gaze&Pinch. Through a user study that features a docking task in VR, we found that DOF-Separation significantly improves the manipulation performance of Hand Ray, while improving that of Virtual Hand only in difficult tasks of combined translation and rotation. We suggest future XR systems to adopt DOF-Separation for input in manipulation-heavy applications, such as 3D design.
In 3D environments, designing efficient notifications is crucial for capturing user attention. While visual notifications -such as objectattached and fixed-position ones- are commonly used in virtual environments, they often require users to shift their gaze away from task-relevant areas, which can interrupt workflow and delay responses. To address these limitations, we designed two gaze-based notification techniques to provide responsive and intuitive notification in a virtual reality (VR) cooking task. We evaluated four different notification types: two world-fixed notifications (onObject and onDock) and two gaze-based methods (GazeCue and Gaze+Dock) with 16 participants. Our results show that participants performed better in using gaze-based notifications compared to world-fixed ones. Questionnaire results indicated higher usability, greater perceived presence, and lower cognitive load for gaze-based notifications. These results highlight the effectiveness of gaze-based notification techniques in VR by reducing the need for visual search and minimizing cognitive load. Our findings provide insights for developers or engineers to design more intuitive and responsive user-interfaces for 3D environments.
In the tourism industry, Extended Reality (XR) technologies are evolving beyond simple marketing tools into sophisticated platforms that deliver immersive experiences to visitors. While these technologies offer significant potential for enhancing tourism experiences, current research primarily examines the tourist perspective, with limited attention to the professional guides who facilitate these interactions. This study investigates how expert tourist guides adopt and integrate XR technology into their professional practice through a novel process for Expert Adoption of XR technology that theoretically grounds the participatory development process in TAM's acceptance mechanisms. Through a co-design process involving experienced guides, we developed an XR application aimed at complementing, rather than replacing, the guide-tourist relationship. Utilizing the Technology Acceptance Model (TAM) for qualitative research, we identified specific concerns and considerations of tourist guides regarding XR adoption and outlined strategies for addressing these challenges during both development and implementation phases. Our findings reveal task-related, individual, and organizational factors for successful XR integration, as well as misconceptions that must be addressed. We suggest that initial scepticism toward XR can be mitigated by ensuring a positive first encounter with technology. Additionally, recognizing stakeholder preferences for specific XR features and prioritizing the most accepted options can enhance adoption. All relevant components of the task, in our case guided tour, must be considered during development, as neglecting any component can impact the overall acceptance. Furthermore, city guides often feel personally responsible for technological failures, highlighting the need for reliable and stable XR applications. These insights provide a methodology applicable to other domains where expert adoption is essential.
Autonomous robot operations in the industry are becoming increasingly complex. It is therefore a significant challenge to comprehend the fundamental processes and to gain an understanding of the status of these systems. The RoboCup Logistics League (RCLL) represents a small smart factory environment with several workstations and operating robots. Despite its small scale the processes that occur within the league are very complex. Even with live commentary, observers have difficulties to follow the processes and game progress. This results in low interest in the RCLL and only few of visitors at competitions. To address this, we want to present RCLL-AR, an augmented reality (AR) solution visualizing highly relevant information of the RoboCup Logistics League. By using RCLL-AR, spectators of the game can see the current progress of the game, receive additional information about different workstations and understand future robot movements. To gain insights into the benefits of RCLL-AR for different stakeholders, we conducted expert interviews, a novice user study and an HMD study. Our findings showcase challenges AR faces in complex autonomous systems but also indicate benefits for novices and experts.