Volumetric medical imaging offers great potential for understanding complex pathologies. Yet, traditional 2D slices provide little support for interpreting spatial relationships, forcing users to mentally reconstruct anatomy into three dimensions. Direct volumetric path tracing and VR rendering can improve perception but are computationally expensive, while precomputed representations, like Gaussian Splatting, require planning ahead. Both approaches limit interactive use. We propose a hybrid rendering approach for high-quality, interactive, and immersive anatomical visualization. Our method combines streamed foveated path tracing with a lightweight Gaussian Splatting approximation of the periphery. The peripheral model generation is optimized with volume data and continuously refined using foveal renderings, enabling interactive updates. Depth-guided reprojection further improves robustness to latency and allows users to balance fidelity with refresh rate. We compare our method against direct path tracing and Gaussian Splatting. Our results highlight how their combination can preserve strengths in visual quality while re-generating the peripheral model in under a second, eliminating extensive preprocessing and approximations. This opens new options for interactive medical visualization.
Accurate needle placement in spine interventions is critical for effective pain management, yet it depends on reliable identification of anatomical landmarks and careful trajectory planning. Conventional imaging guidance often relies both on CT and X-ray fluoroscopy, exposing patients and staff to high dose of radiation while providing limited real-time 3D feedback. We present an optical see-through augmented reality (OST-AR)-guided robotic system for spine procedures that provides in situ visualization of spinal structures to support needle trajectory planning. We integrate a cone-beam CT (CBCT)-derived 3D spine model which is co-registered with live ultrasound, enabling users to combine global anatomical context with local, real-time imaging. We evaluated the system in a phantom user study involving two representative spine procedures: facet joint injection and lumbar puncture. Sixteen participants performed insertions under two visualization conditions: conventional screen vs. AR. Results show that AR significantly reduces execution time and across-task placement error, while also improving usability, trust, and spatial understanding and lowering cognitive workload. These findings demonstrate the feasibility of AR-guided robotic ultrasound for spine interventions, highlighting its potential to enhance accuracy, efficiency, and user experience in image-guided procedures.
(1) The complete resection of a solid tumor is a vital part of the current oncological therapy of head and neck squamous cell carcinoma (HNSCC) in a complex anatomical region. Studies show that surgical resections of head and neck tumors are scarce or incomplete in up to 30-40%. This study investigates the feasibility of generating 3D models of the surface topography of the resected tumor using a mobile phone-based approach and implementing them into visualization software to improve communication between surgeons and pathologists during frozen section (FS) pathology. (2) Materials and Methods: We conducted a pilot study on 10 head and neck cancer patients undergoing surgical resection. A digital 3D model of the surface topography of fresh surgical specimens was created using a mobile phone with photogrammetry software and implemented into software to annotate regions of interest. Utility was assessed using a 0-5 Likert questionnaire completed by four pathologists and four head and neck surgeons; inter-rater reliability was quantified using average pairwise quadratic-weighted Cohen's kappa. (3) Results: 3D model generation was feasible in 10/10 cases with a median scanning time of 4.2 min. A soft-tissue deformation experiment demonstrated good agreement in lateral dimensions but larger deviations in height consistent with soft-tissue compression. Both pathologists and surgeons rated the approach favorably, with an overall utility median of 4.2 (pathologists) and 4.8 (surgeons). (4) Discussion: Smartphone-based photogrammetry enables rapid 3D surface modeling of fresh head and neck specimens and can be integrated into dedicated software for orientation and communication during FS pathology.
Recent robotics advancements have enabled novel applications in medicine, such as automating ultrasound acquisition to support sonographers and enable remote operation. However, a key challenge is the lack of haptic and perceptual transparency regarding the robot’s decisions and actions, which may undermine trust and user acceptance. In this study, we propose a multisensory feedback system to enhance physician–robot interaction during robotic ultrasound procedures. The system provides real-time feedback on the force exerted by the robot on the patient and was developed following a user-centered approach, informed by clinical domain expertise, to support usability and clinical relevance. In a simulated Extended Reality (XR) environment, we evaluated the impact of different feedback conditions—including two sonification strategies, one visual feedback method, and their combinations. A user study with 30 participants, including three practicing clinicians, was conducted to assess these modalities. Results showed that multisensory feedback significantly reduced cognitive load (NASA-TLX: A1+V vs.V, p<0.001; A2+V vs. V, p < 0.01), improved usability (SUS: A1+V vs. V, p < 0.001; A2+V vs. V, p < 0.001) and perceived ease of use (SEQ: A1+V vs. V, p < 0.01; A2+V vs. V, p < 0.01). While visual feedback alone yielded better performance in force monitoring, the multisensory conditions supported better focus on the primary task of image analysis. These findings suggest that multisensory feedback can enhance user experience and support interaction factors associated with trust formation, such as reduced workload and improved awareness of robot behavior, especially with further training and adaptation. Given the predominantly non-clinical sample and the use of abstract visual targets, these results should be read as an exploratory, preclinical validation of multisensory human–robot interaction rather than as evidence of clinical efficacy.
PurposeThis study compares two augmented reality (AR)-guided imaging workflows, one based on ultrasound shape completion and the other on cone-beam computed tomography (CBCT), for planning and executing lumbar needle interventions. The aim is to assess how imaging modality influences user performance, usability, and trust during AR-assisted spinal procedures.MethodsBoth imaging systems were integrated into an AR framework, enabling in situ visualization and trajectory guidance. The ultrasound-based workflow combined AR-guided robotic scanning, probabilistic shape completion, and AR visualization. The CBCT-based workflow used AR-assisted scan volume planning, CBCT acquisition, and AR visualization. A between-subject user study was conducted and evaluated in two phases: (1) planning and image acquisition, and (2) needle insertion.ResultsPlanning time was significantly shorter with the CBCT-based workflow, while SUS, SEQ, and NASA-TLX were comparable between modalities. In the needle insertion phase, the CBCT-based workflow yielded marginally faster insertion times, significantly lower overall placement error, and better subjective ratings with higher Trust. The ultrasound-based workflow achieved adequate accuracy for facet joint insertion, but showed larger errors for lumbar puncture, where reconstructions depended more heavily on shape completion.ConclusionThe findings indicate that both AR-guided imaging pipelines are viable for spinal intervention support. CBCT-based AR offers advantages in efficiency, precision, usability, and user confidence during insertion, whereas ultrasound-based AR provides adaptive, radiation-free imaging but is limited by shape completion in deeper spinal regions. These complementary characteristics motivate hybrid AR guidance that uses CBCT for global anatomical context and planning, augmented by ultrasound for adaptive intraoperative updates.
BACKGROUND:Telementoring and teleconsultation are increasingly employed for collaboration within the healthcare system. The ArtekMed alliance project has developed a mixed reality (MR) teleconsultation system for intensive care units (ICU) using virtual reality (VR) and augmented reality (AR), facilitating real-time interaction between the real world and its reconstructed virtual model, shared by two or more coworkers. OBJECTIVE:We aimed to explore the feasibility and user acceptance of the ArtekMed MR teleconsultation system in a critical care setting and compare it to a standard teleconsultation system using a simulated video call. METHOD:A randomized cross-over study was conducted in a local simulation center: A remote expert (VR user) solved four clinical scenarios, each involving the treatment of an ICU patient with respiratory failure in collaboration with a local practitioner as facilitator (AR user). They used either the MR system (intervention) or a simulated video call (control). A mixed-methods approach was followed to explore structured pre- and post-trial interviews with qualitative and quantitative analyses including standardized usability scores (NASA Task Load Index, System Usability Scale SUS). RESULTS:Twenty-five professionals with intensive care experience completed 100 simulated scenarios. The ArtekMed system achieved an average SUS score of 66, while the simulated video call system was rated almost excellent (SUS score: 84). In three out of four scenarios, the perceived workload using the MR teleconsultation system did not significantly differ from the workload using the standard video call. Most users rated working with both teleconsultation systems positively and anticipated increased efficiency and feasibility with greater familiarity with the MR system. Common issues included visual impairment due to insufficient graphical resolution and unfamiliarity with handling the equipment. 80% of the participants expressed willingness to incorporate the system into their ICU work. CONCLUSION:Collaboration in the ICU using a real-time MR teleconsultation system was rated as a promising technology by the majority of the participants for future use. Technical imperfections seem to prevent further implementation at this stage. Thus, the MR reconstruction needs improvement before clinical implementation.
Operating rooms (ORs) are complex, high-stakes environments requiring precise understanding of interactions among medical staff, tools, and equipment for enhancing surgical assistance, situational awareness, and patient safety. Current datasets fall short in scale, realism and do not capture the multimodal nature of OR scenes, limiting progress in OR modeling. To this end, we introduce MM-OR, a realistic and large-scale multimodal spatiotemporal OR dataset, and the first dataset to enable multimodal scene graph generation. MM-OR captures comprehensive OR scenes containing RGB-D data, detail views, audio, speech transcripts, robotic logs, and tracking data and is annotated with panoptic segmentations, semantic scene graphs, and downstream task labels. Further, we propose MM2SG, the first multimodal large vision-language model for scene graph generation, and through extensive experiments, demonstrate its ability to effectively leverage multimodal inputs. Together, MM-OR and MM2SG establish a new benchmark for holistic OR understanding, and open the path towards multimodal scene analysis in complex, high-stakes environments. Our code, and data is available at https://github.com/egeozsoy/MM-OR.
Robotic ultrasound (RUS) presents new opportunities for remote diagnostics but introduces challenges related to perceptual transparency and operator trust. To address this, we explore the use of multisensory feedback by combining visual and auditory cues within an Extended Reality (XR) interface to support physicians during robot-assisted ultrasound procedures. We developed a prototype system that visualizes probe force and provides real-time sonification using both musical and physics-based approaches. A user study in a simulated XR environment evaluated five feedback conditions across concurrent imaging, force monitoring, and patient interaction tasks. Initial findings suggest that audiovisual feedback may reduce cognitive load and enhance task coordination. This work offers early insights into the design of multimodal interfaces for clinical human–robot interaction.
Robotic ultrasound systems have the potential to improve medical diagnostics, but patient acceptance remains a key challenge. To address this, we propose a novel system that combines an AI-based virtual agent, powered by a large language model (LLM), with three mixed reality visualizations aimed at enhancing patient comfort and trust. The LLM enables the virtual assistant to engage in natural, conversational dialogue with patients, answering questions in any format and offering real-time reassurance, creating a more intelligent and reliable interaction. The virtual assistant is animated as controlling the ultrasound probe, giving the impression that the robot is guided by the assistant. The first visualization employs augmented reality (AR), allowing patients to see the real world and the robot with the virtual avatar superimposed. The second visualization is an augmented virtuality (AV) environment, where the real-world body part being scanned is visible, while a 3D Gaussian Splatting reconstruction of the room, excluding the robot, forms the virtual environment. The third is a fully immersive virtual reality (VR) experience, featuring the same 3D reconstruction but entirely virtual, where the patient sees a virtual representation of their body being scanned in a robot-free environment. In this case, the virtual ultrasound probe, mirrors the movement of the probe controlled by the robot, creating a synchronized experience as it touches and moves over the patient's virtual body. We conducted a comprehensive agent-guided robotic ultrasound study with all participants, comparing these visualizations against a standard robotic ultrasound procedure. Results showed significant improvements in patient trust, acceptance, and comfort. Based on these findings, we offer insights into designing future mixed reality visualizations and virtual agents to further enhance patient comfort and acceptance in autonomous medical procedures.
Objective: Extended reality (XR) teleconsultation is used in surgery and medical emergencies, employing various technological approaches that differ in accuracy, timeliness, and user preference.Methods: We conducted a systematic literature review following PRISMA. We searched the databases IEEE Xplore, Springer Link, ACM and added an additional manual search. In total, we found 187 studies and included 14 in our review.Conclusion: Our findings highlight the widespread use of video-based streaming and 3D reconstruction based on static RGB-D sensor. We found limitations in the reconstruction quality, where existing work would benefit from high-quality rendering. Interaction via annotations is common, addressing key usability needs for various surgeries and emergency situations. A standardized evaluation for interaction techniques would be beneficial for comparability. Our findings hold significant implications for improving teleconsultation and evaluation of XR telemedicine approaches.
Accurate labels of surgical procedures such as image segmentations or interaction labels are paramount for many of today's medical image computing tasks. Creating a dataset with these labels requires a great deal of manual work and relies on the involvement of medical experts, which is very time-consuming and costly. We propose a pathway for the automatic generation of such labels utilizing the spatial and temporal registration between a patient, the anatomical model, tracked surgical instruments, and the surgeon's view of the patient. These requirements for the automatic generation of labels are identical to the requirements of many navigated and augmented reality (AR) enabled surgeries. The AR system, through 3D registration, has the defining ability to accurately overlay real objects with their virtual counterparts. Our approach collects the complete raw data (e.g. video, tracking data, calibrations etc.) that feeds a live laparoscopic AR system for later analysis. By converting these complete recordings of the surgery into different representations, the AR system generates valuable datasets as mere by-products. Additionally, as our approach does not rely on visual input alone but on additional 3D information, the system can create labels even if the visual input is occluded or a tool interacts with tissue outside of the view of the laparoscopic camera. In this paper, we present a realization of this concept, then evolve this foundational idea into an interactive system that assists users in annotating surgical data. Finally, we gather and analyse feedback from six participants to evaluate the efficacy and user-friendliness of our system.
In Augmented Reality (AR), virtual objects interact with real objects. However, the lack of physicality of virtual objects leads to the absence of natural sonic interactions. When virtual and real objects collide, either no sound or a generic sound is played. Both lead to an incongruent multisensory experience, reducing interaction and object realism. Unlike in Virtual Reality (VR) and games, where predefined scenes and interactions allow for the playback of pre-recorded sound samples, AR requires real-time sound synthesis that dynamically adapts to novel contexts and objects to provide audiovisual congruence during interaction. To enhance real-virtual object interactions in AR, we propose a framework for context-aware sounds using methods from computer vision to recognize and segment the materials of real objects. The material's physical properties and the impact dynamics of the interaction are used to generate material-based sounds in real-time using physical modelling synthesis. In a user study with 24 participants, we compared our congruent material-based sounds to a generic sound effect, mirroring the current standard of non-context-aware sounds in AR applications. The results showed that material-based sounds led to significantly more realistic sonic interactions. Material-based sounds also enabled participants to distinguish visually similar materials with significantly greater accuracy and confidence. These findings show that context-aware, material-based sonic interactions in AR foster a stronger sense of realism and enhance our perception of real-world surroundings.
The advancement and maturity of large language models (LLMs) and robotics have unlocked vast potential for human-computer interaction, particularly in the field of robotic ultrasound. While existing research primarily focuses on either patient-robot or physician-robot interaction, the role of an intelligent virtual sonographer (IVS) bridging physician-robot-patient communication remains underexplored. This work introduces a conversational virtual agent in Extended Reality (XR) that facilitates real-time interaction between physicians, a robotic ultrasound system(RUS), and patients. The IVS agent communicates with physicians in a professional manner while offering empathetic explanations and reassurance to patients. Furthermore, it actively controls the RUS by executing physician commands and transparently relays these actions to the patient. By integrating LLM-powered dialogue with speech-to-text, text-to-speech, and robotic control, our system enhances the efficiency, clarity, and accessibility of robotic ultrasound acquisition. This work constitutes a first step toward understanding how IVS can bridge communication gaps in physician-robot-patient interaction, providing more control and therefore trust into physician-robot interaction while improving patient experience and acceptance of robotic ultrasound. The code is available at https://github.com/stytim/IVS .
In mixed reality (MR) telepresence applications, the differences between participants' physical environments can interfere with effective collaboration. For asymmetric tasks, users might need to access different resources (information, objects, tools) distributed throughout their room. Existing intersection methods do not support such interactions, because a large portion of the telepresence participants' rooms become inaccessible, along with the relevant task resources. We propose MRUnion, a Mixed Reality Telepresence pipeline for asymmetric task-aware 3D mutual scene generation. The key concept of our approach is to enable a user in an asymmetric telecollaboration scenario to access the entire room, while still being able to communicate with remote users in a shared space. For this purpose, we introduce a novel mutual room layout called Union. We evaluated 882 space combinations quantitatively involving two, three, and four combined remote spaces and compared it to a conventional Intersect room layout. The results show that our method outperforms existing intersection methods and enables a significant increase in space and accessibility to resources within the shared space. In an exploratory user study (N=24), we investigated the applicability of the synthetic mutual scene in both MR and VR setups, where users collaborated on an asymmetric remote assembly task. The study results showed that our method achieved comparable results to the intersect method but requires further investigation in terms of social presence, safety and support of collaboration. From this study, we derived design implications for synthetic mutual spaces.
Medical doctors rely on images of the human anatomy, such as magnetic resonance imaging (MRI), to localize regions of interest in the patient during diagnosis and treatment. Despite advances in medical imaging technology, the information conveyance remains unimodal. This visual representation fails to capture the complexity of the real, multisensory interaction with human tissue. However, perceiving multimodal information about the patient's anatomy and disease in real-time is critical for the success of medical procedures and patient outcome. We introduce a Multimodal Medical Image Interaction (MMII) framework to allow medical experts a dynamic, audiovisual interaction with human tissue in three-dimensional space. In a virtual reality environment, the user receives physically informed audiovisual feedback to improve the spatial perception of anatomical structures. MMII uses a model-based sonification approach to generate sounds derived from the geometry and physical properties of tissue, thereby eliminating the need for hand-crafted sound design. Two user studies involving 34 general and nine clinical experts were conducted to evaluate the proposed interaction framework's learnability, usability, and accuracy. Our results showed excellent learnability of audiovisual correspondence as the rate of correct associations significantly improved (p < 0.001) over the course of the study. MMII resulted in superior brain tumor localization accuracy (p < 0.05) compared to conventional medical image interaction. Our findings substantiate the potential of this novel framework to enhance interaction with medical images, for example, during surgical procedures where immediate and precise feedback is needed.
Researching novel user experiences in medicine is challenging due to limited access to equipment and strict ethical protocols. Extended Reality (XR) simulation technologies offer a cost-and time-efficient solution for developing interactive systems. Recent work has shown Extended Reality Prototyping (XRP)’s potential, but its applicability to specific domains like controlling complex machinery needs further exploration. This paper explores the benefits and limitations of XRP in controlling a mobile medical imaging robot. We compare two XR visualization techniques to reduce perceived latency between user input and robot activation. Our XRP validation study demonstrates its potential for comparative studies, but identifies a gap in modeling human behavior in the analytic XRP validation framework.
The utilization of augmented reality (AR) in medical robotics offers significant advancements in enhancing procedural accuracy and patient safety. This paper investigates novel AR visualization techniques designed to depict in-contact force applied by a robotic ultrasound probe, aiming to optimize the control practitioners have over probe force for ultrasound procedures, thereby enhancing both image quality and patient comfort. We developed and evaluated four distinct AR visualization techniques through a comprehensive user study conducted in a clinical setting. The study assessed the efficiency and user experience associated with each technique. The findings revealed notable differences in user performance and preferences, indicating that specific visualizations significantly improve the precision of force application and could lead to better procedural outcomes. The results underscore the potential of AR visualizations to transform robotic-assisted medical procedures by improving the interface between clinicians and robotic systems. Moreover, these advancements foster a deeper trust and acceptance of robotic technologies among healthcare professionals and patients. This study not only highlights the immediate benefits of AR in enhancing robotic ultrasound but also sets the stage for further research into AR's expansive role in complex medical robotics scenarios.
Mixed Reality (MR) is proven in the literature to support precise spatial dental drill positioning by superimposing 3D widgets. Despite this, the related knowledge about widget's visual design and interactive user feedback is still limited. Therefore, this study is contributed to by co-designed MR drill tool positioning widgets with two expert dentists and three MR experts. The results of co-design are two static widgets (SWs): a simple entry point, a target axis, and two dynamic widgets (DWs), variants of dynamic error visualization with and without a target axis (DWTA and DWEP). We evaluated the co-designed widgets in a virtual reality simulation supported by a realistic setup with a tracked phantom patient, a virtual magnifying loupe, and a dentist's foot pedal. The user study involved 35 dentists with various backgrounds and years of experience. The findings demonstrated significant results; DWs outperform SWs in positional and rotational precision, especially with younger generations and subjects with gaming experiences. The user preference remains for DWs (19) instead of SWs (16). However, findings indicated that the precision positively correlates with the time trade-off. The post-experience questionnaire (NASA-TLX) showed that DWs increase mental and physical demand, effort, and frustration more than SWs. Comparisons between DWEP and DWTA show that the DW's complexity level influences time, physical and mental demands. The DWs are extensible to diverse medical and industrial scenarios that demand precision.