Volumetric medical imaging offers great potential for understanding complex pathologies. Yet, traditional 2D slices provide little support for interpreting spatial relationships, forcing users to mentally reconstruct anatomy into three dimensions. Direct volumetric path tracing and VR rendering can improve perception but are computationally expensive, while precomputed representations, like Gaussian Splatting, require planning ahead. Both approaches limit interactive use. We propose a hybrid rendering approach for high-quality, interactive, and immersive anatomical visualization. Our method combines streamed foveated path tracing with a lightweight Gaussian Splatting approximation of the periphery. The peripheral model generation is optimized with volume data and continuously refined using foveal renderings, enabling interactive updates. Depth-guided reprojection further improves robustness to latency and allows users to balance fidelity with refresh rate. We compare our method against direct path tracing and Gaussian Splatting. Our results highlight how their combination can preserve strengths in visual quality while re-generating the peripheral model in under a second, eliminating extensive preprocessing and approximations. This opens new options for interactive medical visualization.
Cross Reality (CR) is a new emerging field based on the current developments in Mixed Reality hardware, especially supported by the broad market penetration of video-based see-through Head-Mounted Displays. It refers to applications that span across different stages (real, Augmented Reality, Augmented Virtuality, Virtual Reality) of the reality-virtuality continuum, where users are interconnected between different stages and/or are able to transition between these stages. This publication follows the concept of other grand challenges publications and reflects the discussion of various researchers invested in CR. After an initial discussion at the 1st Joint Workshop on Cross Reality at IEEE ISMAR 2023, six topic groups have been identified, leading to 22 challenges, which were discussed in groups over the period of multiple months. The discussion of these challenges should act as a road map for future research in the area of CR.
Autonomous robots in unknown indoor environments require both reliable collision avoidance and object-level understanding. Classical representations such as TSDF support safe planning but lack semantics, while photorealistic methods like Gaussian Splatting (GS) provide rich appearance yet suffer from soft geometry, limiting precise obstacle avoidance. We present LiftNav, a hybrid navigation framework built on GSFusion's TSDF+GS dual map, augmented with a real-time pipeline of YOLO-based detection, TSDF-based 3D lifting, and B-spline trajectory optimization. This design enables flexible semantic navigation without dense 3D embeddings. We further introduce a hinge-loss-based collision penalty that improves trajectory smoothness and safety. We evaluate our approach in a simulation using the Replica dataset. Compared against a state-of-the-art radiance field baseline we show a 100
Multi-camera dynamic Augmented Reality (AR) applications require a camera pose estimation to leverage individual information from each camera in one common system. This can be achieved by combining contextual information, such as markers or objects, across multiple views. While commonly cameras are calibrated in an initial step or updated through the constant use of markers, another option is to leverage information already present in the scene, like known objects. Another downside of marker-based tracking is that markers have to be tracked inside the field-of-view (FoV) of the cameras. To overcome these limitations, we propose a constant dynamic camera pose estimation leveraging spatiotemporal FoV overlaps of known objects on the fly. To achieve that, we enhance the state-of-the-art object pose estimator to update our spatiotemporal scene graph, enabling a relation even among non-overlapping FoV cameras. To evaluate our approach, we introduce a multi-camera, multi-object pose estimation dataset with temporal FoV overlap, including static and dynamic cameras. Furthermore, in FoV overlapping scenarios, we outperform the state-of-the-art on the widely used YCB-V and T-LESS dataset in camera pose accuracy. Our performance on both previous and our proposed datasets validates the effectiveness of our marker-less approach for AR applications. The code and dataset are available on https://github.com/roth-hex-lab/IEEE-VR-2026-MultiCam.
Efficient and robust 3D scene representation is crucial in autonomous driving, robotics, and related fields. While RGB images provide valuable content for 3D reconstruction, other modalities like thermal or depth can enable additional information on the environment. Lately, novel view synthesis methods like 3D Gaussian Splatting have started using multiple modalities to further boost their performance. But fusing or combining multimodal data can make the process slower and can bring in additional challenges. Therefore, our project aims to use single modality based on thermal infrared domain, by removing the reliance on visible light as much as possible. This single modality can be expected to be faster as it does not rely on multimodal data. We propose a method, Thermal-to-Depth Gaussian Splatting (TDg), that uses only thermal images and depth estimation in its architecture to derive the radiance fields. Our TDg method outperforms the MSMG (Multiple Single-Modal Gaussians) baseline in most cases on our test datasets, RGBT-Scenes and ThermalMix. On average, the rendering quality metrics such as learned perceptual image patch similarity (LPIPS), structural similarity index measure (SSIM), and peak signal-to-noise ratio (PSNR) of TDg are 1.12
Psychological factors such as ownership, agency, and trust are critical to the acceptance and effective use of prosthetic devices, yet their relationship to control reliability remains underexplored. We investigated how induced delays and artificial malfunctions influence these factors during prosthesis simulator use in a fully immersive virtual reality environment. A Pasta Box Task was implemented in Unity with MuJoCo physics simulation, using surface electromyography myocontrol and integrated eye tracking to measure subjective and visuomotor responses. Thirty non-disabled participants completed six within-participant conditions crossing two control delay and three artificial malfunction levels. Validated questionnaires assessed ownership, agency, and trust, while gaze metrics quantified fixation percent, target-locking strategy, and eye arrival and leaving latencies. Both delay and malfunction significantly reduced psychometric scores, with artificial malfunctions exerting the largest overall effect, while delay particularly diminished agency. Artificial malfunctions also increased fixations on the prosthesis and altered gaze strategies, suggesting compensatory behavior. Delay primarily affected eye-arrival latency and the number of fixations, whereas artificial malfunctions influenced target-locking strategy and eye-leaving latency, indicating distinct visuomotor adaptations to each reliability factor. Weak but significant correlations emerged between gaze behavior and psychometric measures. The results of the experiment highlight the value of immersive, physics-accurate virtual reality as an early-stage platform for the controlled evaluation of myocontrol and prosthesis behavior and for capturing relevant psychometric and visuomotor indicators relevant to user-centered design.
A convincing sense of embodiment in virtual reality (VR) is crucial for creating immersive and engaging experiences, as it shapes how users perceive and interact with their virtual bodies. The sense of embodiment is thereby, among others, affected by the shape, appearance, and fidelity of the virtual body. However, achieving convincing avatar appearance remains a challenge for VR applications. One promising solution is Pass-Through Embodiment (PTE), which enables users to see their real bodies in VR. PTE combines depth-based segmentation with the pass-through video stream of video-see-through displays to effectively visualize photon-captured representations of their own bodies. Despite the source-fidelity of the representation, the resolution of integrated depth sensors in Head Mounted Displays (HMD) can produce artifacts at segmentation boundaries, leading to visible aliasing. The perceptual impact of these artifacts on the VR experience remains unexplored. Therefore, in this paper we compare three edge-rendering techniques designed to reduce artifacts without compromising performance. Aside from a soft gradient, we introduce two new methods with a hard and dithered edge. The latter aims to balance the sharpness of hard masks with the smoothness of gradient transitions, without relying on alpha blending. To evaluate those methods in a PTE context, we conducted a within-subjects study that introduces a novel real-mirror paradigm, using an actual physical mirror as reference for reflection. We found significant results in measured presence and embodiment. Subsequent analysis revealed that our dithered cutout approach significantly outperforms hard masks, while no significant difference was found between soft condition. These results suggest a perceptual continuum where dithering and soft blending both effectively reduce visual artifacts through gradient representation. Together with high overall ratings on presence and embodiment across all conditions, these findings confirm PTE as a robust method for supporting embodiment and presence, while highlighting the potential of dithering as a computationally efficient yet perceptually comparable alternative to smooth blending.
The handling and assembly of instruments during surgery imposes high cognitive demands on scrub nurses, particularly when instruments are unfamiliar. We present a supporting guidance system for surgical instrumentation that combines multi-camera 6D pose estimation with augmented reality in-situ visualization on a head-mounted display without the requirement for additional markers. Pose estimation and consecutive camera calibration are achieved through known objects. The 6D pose estimation network is trained purely on synthetic data, aiming for better generalizability and real-world applicability. The AR guidance displays tooltip localization cues and step-wise assembly animations. Via gaze-based selection and a foot pedal, users can switch between assembly steps in intraoperative use. In a technical evaluation, our approach outperforms state-of-art 6D pose estimation. A user study with 29 scrub nurses was conducted in a surgical simulation of knee arthroplasty, comparing the system against a paper manual. AR guidance significantly reduced the perceived workload compared. Objectively, AR guidance reduced task completion time by 21.3% (4.76 minutes). Specifically, scrub nurses less experienced with the instrument set benefited when using the system. Error frequencies were comparable between conditions. Qualitative feedback highlighted improved process clarity, reduced information overload, and perceived independence. To summarize, our marker-free multi-camera AR guidance approach for surgical instruments can, subjectively and objectively, improve intraoperative instrumentation performance, particularly for untrained scrub nurses.
We present a task-conditioned refinement for 3D Gaussian Splatting (GS) that enables robots or human operators to selectively extract task-relevant regions of a learned scene. Given a pre-trained GS map, our approach supports local region-of-interest (ROI) refinement, preserving a global map consistency while meeting close to real-time constraints required for interactive robotic perception. The framework decouples semantic ROI selection from initial GS optimization, allowing flexible integration with external and novel perception models. We evaluate our approach on indoor and outdoor data (TUM RGB-D, MipNeRF360), demonstrating a higher novel view syn-thesis quality compared to the state-of-the-art, reduced artifacts, and bounded latency suitable for human-in-the-loop operation.
Immersive virtual reality learning environments (IVRLEs) are increasingly used in medical education, yet the role of environmental fidelity—particularly scene design—remains underexplored. This study examines how varying levels of fidelity and contextualization affect motivational and cognitive outcomes. Eighty-seven medical students were randomly assigned to one of three scene conditions: a minimalistic “Blank Scene,” a “Reconstructed Classroom Scene”, or an “Inside-Human Scene”. All students used a custom-developed application to learn about embryonic heart development. We measured virtual presence, intrinsic motivation, cognitive load, learning outcomes, and usability. Results showed that scene design influenced virtual presence, selected aspects of intrinsic motivation, cognitive load, and learning outcomes. The Inside-Human Scene elicited higher physical and self-presence as well as higher comprehension scores compared to the Reconstructed Classroom Scene. The Reconstructed Classroom Scene was associated with higher extraneous cognitive load. Intrinsic cognitive load was rated higher in the Inside-Human Scene, while germane cognitive load did not differ between conditions. No significant differences were found for task performance or factual recall. Overall, the findings indicate that scene design in IVRLEs affects how learners engage with complex content and may support deeper understanding when perceptual and contextual properties are coherent, while visually detailed environments may increase extraneous cognitive demands without improving learning.
Embodiment is fundamental to immersive VR, yet traditional avatars often suffer from perceptual mismatches. Pass-Through Embodiment (PTE) addresses this by integrating a stereoscopic live video feed of the user's body into the virtual environment. By combining depth-based segmentation with video-see-through streams, PTE provides a high-fidelity, photon-captured representation. We present the first public PTE demonstration, allowing users to perceive themselves while interacting with diverse virtual scenes. Participants can evaluate how their body, different edge rendering, and environmental contexts affect visual coherence in real-time. This demonstrates how PTE facilitates a naturally anchored sense of embodiment, bridging physical and virtual worlds through natural interaction.
VR simulations are becoming essential for staff training in clinical care, specifically in environments that cannot be trained well during regular operation, such as critical care. This paper presents a novel system that aims at further closing the gap between simulation and reality by integrating real patient data into immersive training scenarios in the context of acute care. Focused on sepsis recognition, the VR simulation introduces trainees to clinical routines and to make informed decisions while observing the evolving patient conditions. By engaging with dynamic disease progression, it fosters understanding of critical conditions in a time-sensitive context. Preliminary feedback from a pilot assessment with nursing professionals highlighted its value for trainees and potential to enhance preparedness and decision-making skills in real-world scenarios.
Mobile reconstruction has the potential to support time-critical tasks such as tele-guidance and disaster response, where operators must quickly gain an accurate understanding of the environment. Full high-fidelity scene reconstruction is computationally expensive and often unnecessary when only specific points of interest (POIs) matter for timely decision making. We address this challenge with CoRe-GS, a semantic POI-focused extension of Gaussian Splatting (GS). Instead of optimizing every scene element uniformly, CoRe-GS first produces a fast segmentation-ready GS representation and then selectively refines splats belonging to semantically relevant POIs detected during data acquisition. This targeted refinement reduces training time to 25% compared to full semantic GS while improving novel view synthesis quality in the areas that matter most. We validate CoRe-GS on both real-world (SCRREAM) and synthetic (NeRDS 360) datasets, demonstrating that prioritizing POIs enables faster and higher-quality mobile reconstruction tailored to operational needs.
Neue Werkzeuge der maschinellen und künstlichen Intelligenz (KI) erlauben personalisierte Behandlungen, präzise Diagnosen, optimierte Planungen und effektives Monitoring. Durch große KI-Modelle und die Fusion von Daten aus z. B. Bildgebung, Wearables und Textdaten entstehen neue Möglichkeiten für präventive Risikobewertung, individualisierte Therapieentscheidungen und intraoperative Assistenzsysteme. Diese könnten zukünftig Personal entlasten und die Versorgungsqualität entlang des Behandlungspfades verbessern. In diesem Beitrag beschreiben wir einige aktuelle Technologieentwicklungen und zukünftige Perspektiven und ordnen diese ein.
In medical image visualization, path tracing of volumetric medical data like computed tomography (CT) scans produces lifelike three-dimensional visualizations. Immersive virtual reality (VR) displays can further enhance the understanding of complex anatomies. Going beyond the diagnostic quality of traditional 2D slices, they enable interactive 3D evaluation of anatomies, supporting medical education and planning. Rendering high-quality visualizations in real-time, however, is computationally intensive and impractical for compute-constrained devices like mobile headsets. We propose a novel approach utilizing Gaussian Splatting (GS) to create an efficient but static intermediate representation of CT scans. We introduce a layered GS representation, incrementally including different anatomical structures while minimizing overlap and extending the GS training to remove inactive Gaussians. We further compress the created model with clustering across layers. Our approach achieves interactive frame rates while preserving anatomical structures, with quality adjustable to the target hardware. Compared to standard GS, our representation retains some of the explorative qualities initially enabled by immersive path tracing. Selective activation and clipping of layers are possible at rendering time, adding a degree of interactivity to otherwise static GS models. This could enable scenarios where high computational demands would otherwise prohibit using path-traced medical volumes.
Creating a compelling sense of presence and embodiment can enhance the user experience in virtual reality (VR). One method to accomplish this is through self-representation with embodied personalized avatars or video self-avatars. However, these approaches require external hardware and primarily evaluate hand representations in VR across various tasks. We therefore present in this paper an alternative approach: video Pass-Through Embodiment (PTE), which utilizes the per-eye real-time depth map from Head-Mounted Displays (HMDs) traditionally used for Augmented Reality features. This method allows the user's real body to be cut out of the pass-through video stream and be represented in the VR environment without the need for additional hardware. To evaluate our approach, we conducted a between-subjects study involving 40 participants who completed a seated object sorting task using either PTE or a customized avatar. The results show that PTE, despite its limited depth resolution that leads to some visual artifacts, significantly enhances the user's sense of presence and embodiment. In addition, PTE does not negatively affect task performance, cognitive load, or cause VR sickness. These findings imply that video pass-through embodiment offers a practical and efficient alternative to traditional avatar-based methods in VR.
The accurate reconstruction of dynamic scenes with neural radiance fields is significantly dependent on the estimation of camera poses. Widely used structure-from-motion pipelines encounter difficulties in accurately tracking the camera trajectory when faced with separate dynamics of the scene content and the camera movement. To address this challenge, we propose Dynamic Motion-Aware Fast and Robust Camera Localization for Dynamic Neural Radiance Fields (DynaMoN). DynaMoN utilizes semantic segmentation and generic motion masks to handle dynamic content for initial camera pose estimation and statics-focused ray sampling for fast and accurate novel-view synthesis. Our novel iterative learning scheme switches between training the NeRF and updating the pose parameters for an improved reconstruction and trajectory estimation quality. The proposed pipeline shows significant acceleration of the training process. We extensively evaluate our approach on two real-world dynamic datasets, the TUM RGB-D dataset and the BONN RGB-D Dynamic dataset. DynaMoN improves over the state-of-the-art both in terms of reconstruction quality and trajectory accuracy. We plan to make our code public to enhance research in this area.
Objective: Extended reality (XR) teleconsultation is used in surgery and medical emergencies, employing various technological approaches that differ in accuracy, timeliness, and user preference.Methods: We conducted a systematic literature review following PRISMA. We searched the databases IEEE Xplore, Springer Link, ACM and added an additional manual search. In total, we found 187 studies and included 14 in our review.Conclusion: Our findings highlight the widespread use of video-based streaming and 3D reconstruction based on static RGB-D sensor. We found limitations in the reconstruction quality, where existing work would benefit from high-quality rendering. Interaction via annotations is common, addressing key usability needs for various surgeries and emergency situations. A standardized evaluation for interaction techniques would be beneficial for comparability. Our findings hold significant implications for improving teleconsultation and evaluation of XR telemedicine approaches.
In Augmented Reality (AR), virtual objects interact with real objects. However, the lack of physicality of virtual objects leads to the absence of natural sonic interactions. When virtual and real objects collide, either no sound or a generic sound is played. Both lead to an incongruent multisensory experience, reducing interaction and object realism. Unlike in Virtual Reality (VR) and games, where predefined scenes and interactions allow for the playback of pre-recorded sound samples, AR requires real-time sound synthesis that dynamically adapts to novel contexts and objects to provide audiovisual congruence during interaction. To enhance real-virtual object interactions in AR, we propose a framework for context-aware sounds using methods from computer vision to recognize and segment the materials of real objects. The material's physical properties and the impact dynamics of the interaction are used to generate material-based sounds in real-time using physical modelling synthesis. In a user study with 24 participants, we compared our congruent material-based sounds to a generic sound effect, mirroring the current standard of non-context-aware sounds in AR applications. The results showed that material-based sounds led to significantly more realistic sonic interactions. Material-based sounds also enabled participants to distinguish visually similar materials with significantly greater accuracy and confidence. These findings show that context-aware, material-based sonic interactions in AR foster a stronger sense of realism and enhance our perception of real-world surroundings.