Abstract Background Current imaging assessment for pancreatic cancer resectability demonstrates problematic inter-observer variability, with only fair-to-moderate agreement among experienced raters. Virtual reality technology offers stereoscopic three-dimensional visualization that may improve diagnostic accuracy and agreement. However, optimal visualization strategies for clinical adoption remain unclear. Methods Ten hepatopancreatobiliary surgeons from two high-volume centers were randomized 1:1 to assess twelve contrast-enhanced CT cases using either VR volumetric rendering or CSI. Primary outcomes included inter-rater agreement, diagnostic accuracy against expert reference standard, assessment time, and surgeon confidence. Statistical analysis employed Fleiss’ κ for inter-rater agreement and two-sided Mann–Whitney U tests on surgeon-level summary measures for between-group comparisons. Results CSI display on 2D screens achieved substantial inter-rater agreement for resectability assessment (κ = 0.609) while VR demonstrated only slight agreement (κ = 0.127). Diagnostic accuracy was superior with CSI (84.7% vs. 79.7%), with the most pronounced difference in resectability determination (83.3% vs. 58.3%, p = 0.033). VR users reported significantly lower confidence (4.85 ± 1.15 vs. 6.32 ± 0.77, p = 0.028). Assessment times were comparable between groups (median 313.5 s vs. 327.5 s, p = 1.00). Conclusions In this preliminary investigation, our VR visualization strategy demonstrated lower diagnostic accuracy and inter-rater agreement than CSI. However, prior studies suggest that VR systems employing alternative, hybrid visualization approaches may improve inter-rater agreement, indicating that visualization strategy, rather than VR technology per se, is the primary determinant of utility. Trial registration DRKS00033932 (German Clinical Trials Register), registered prospectively.
Robotic-assisted surgery provides superior fine motor control of instruments but typically deprives surgeons of haptic cues. We present the Haptic Interaction Toolkit, a mixed-reality robotic console that integrates dual force-feedback devices, and a Unity-based software pipeline, along with details of its configuration and implementation. The toolkit is a reproducible technology platform that enables real/virtual registration, configurable visuo-haptic interactions, and replication via source code. Three exemplar experimental scenarios (discrimination of liver stiffness, protection of critical structures, and fatty tissue dissection next to sensitive structures) demonstrate how the system can be used to prototype and evaluate novel haptic interaction concepts for future robotic systems.
Precise spatial understanding of complex anatomy is critical for preoperative planning in hepatobiliary surgery. Traditional CT and MRI imaging require mental reconstruction of anatomy from 2D slices, imposing substantial cognitive load. Although 3D reconstructions improve spatial understanding, they are typically displayed on 2D screens, limiting true depth perception. Virtual Reality (VR) visualization offers both stereoscopic depth and embodied interaction to improve spatial-anatomical understanding, yet its quantitative advantage over standard desktop visualization remains uncertain, especially regarding task complexity. In this randomized crossover study, 58 medical students analyzed 3D liver models of varying complexity using both VR and desktop visualization. Performance on lesion/vessel relations and lesion segment allocation tasks served as a measure of spatial-anatomical understanding, while visuospatial ability was assessed with the Mental Rotations Test. In complex models, VR significantly improved performance compared with desktop visualization (28.0 ± 3.3 vs. 26.4 ± 3.6; p = 0.002, d = 0.46), whereas results for simpler models were comparable. The VR advantage scaled with task complexity and correlated with higher visuospatial ability (r = 0.31, p = 0.018). These findings indicate that VR is associated with measurable advantages under higher task complexity, supporting its potential role in surgical education and preoperative planning, although the present design cannot isolate which immersive features drive this benefit.
Robotic-assisted surgery (RAS) transforms traditional surgical practice by mediating the surgeon’s actions through visual interfaces, significantly altering sensory experiences, particularly through the loss of direct haptic feedback. This paper presents the Haptic Interaction Toolkit, a Virtual Reality (VR)-enhanced physical robotic surgery console designed to investigate and prototype new modes of haptic interaction in digitally mediated surgical environments. Drawing on insights from expert interviews, the toolkit offers two experimental scenarios, “Detecting Tissue Stiffness” and “Protecting Critical Structures,” to examine how tactile cues can support surgical precision, embodiment, and spatial awareness. The research emphasizes the conceptual challenges of integrating haptic perception into VR simulations and highlights the need for human-centered frameworks that prioritize sensory feedback as a critical component of surgical decision-making. By advancing a multisensory, embodied approach to human–robot interaction, the toolkit aims to inform the design of next-generation surgical systems and contribute to safer, more intuitive practices in digital surgery.
Surgical education requires the combination of declarative knowledge with spatial, procedural, and sensorimotor skills, yet most training materials rely predominantly on two-dimensional media and passive instruction. This paper presents a mixed reality (MR) training system that enables learners to physically manipulate real surgical instruments while receiving spatially registered, context-aware digital guidance. The system is designed to support embodied learning and to improve the transition between training and clinical practice. In a controlled study $(\mathrm{N}= 42)$, MR-based training was compared to conventional text- and image-based instruction. Results show that MR significantly improved recall of spatially embedded instrument knowledge, such as component names, while no overall advantage was observed for purely symbolic information. Both groups reported comparable gains in perceived learning success. User experience measures indicate high engagement, immersion, and minimal discomfort in the MR condition. These findings suggest that MR is particularly effective for learning tasks involving spatial relations and interaction, and may serve as a complementary modality to existing instructional approaches in early-stage surgical training.
Robot-assisted surgery (RAS) has raised concerns within human-computer interaction, particularly regarding the socio-material configurations of robotic systems and their impact on surgical practices. To investigate these configurations in the context of the da Vinci robotic system, we used a research-through-design approach. Through this approach, we developed six Visual RAS (ViRAS) scenarios derived from RAS observations. These scenarios represent different configurations of interacting surgical team members and the robotic system in distinct adverse RAS events. In ViRAS-guided interviews with experienced RAS surgeons, we found that ViRAS scenarios help reflect on surgical practices and the specific material and spatial properties of RAS. Our findings indicated that the team’s cognitive engagement during surgery could be improved by providing sensory augmentation to facilitate task perception and individual skill development. Through our research, we show how ViRAS scenarios, as a tool for reflection, can reveal opportunities for designing socio-material configurations in RAS and beyond.
This study presents a mixed reality training concept designed to enhance medical students’ acquisition of surgical knot-tying skills, a fundamental component of surgical training critical for effective wound closure and tissue healing. Utilizing a virtual reality headset with video passthrough functionality, the system provides adaptive visual instructions tailored to the user’s hand movements during the knot-tying process. A prototype was developed based on the concept, featuring three-dimensional videos in which virtual instructor hands demonstrate each step of the procedure. The training concept was derived from an iterative, user-centered process encompassing requirement analysis, prototype development, and evaluation. Key functionalities include the ability to display thread tension and tensile strength, dynamically adapt learning speed to the user’s progress, and deliver personalized feedback by visually augmenting the hands and fingers. Evaluation results indicate that spatial and tangible interactions facilitated by the mixed reality training prototype support the acquisition of practical skills, bridging the gap between digital and physical simulation training.
Robotic-assisted surgery (RAS) has fundamentally transformed surgical practice by expanding the boundaries of human motion, offering increased precision and dexterity. It has opened new opportunities for autonomous and remote surgery beyond traditional operating room settings. The repertoire of possibilities is currently limited by the absence of haptic feedback, challenges in intraoperative communication, and ergonomic constraints. In this study, we gathered insights from clinically active surgeons across multiple specialties to identify current limitations and unmet needs in RAS. We hereby aim to bridge real-world surgical challenges by facilitating interdisciplinary collaboration and advancing the field through focused technological and clinical development. This study underscores the critical need for enhanced haptic feedback, improved intraoperative communication tools, and ergonomic refinements in RAS. Addressing these issues is essential to optimize surgical performance and patient outcomes. Future developments of robotic-assisted systems should prioritize multimodal concepts of sensory integration to improve the sense of embodiment within the surgical environment.
The anatomical complexity of the thyroid region presents significant challenges in surgical training, particularly regarding the identification and preservation of the recurrent laryngeal nerve and parathyroid glands. We present a prototype of a virtual reality simulator designed to support thyroidectomy training by enabling the immersive, interactive exploration of CT-derived, deformable anatomical models in a photorealistic operating room environment. Structures not detectable in CT, such as nerves and glands, were manually integrated. The simulator was evaluated qualitatively by three surgeons using a structured questionnaire. Feedback indicated high usability, visual realism, and potential for improving anatomical recognition skills. Limitations include the absence of instrument interaction, haptic feedback, and full procedural simulation. This prototype demonstrates feasibility and outlines a clear development roadmap toward a high-fidelity, scalable training platform for endocrine surgery.
IntroductionVolumetric video production in commercial studios is predominantly produced using a multi-view stereo process that relies on a high two-digit number of cameras to capture a scene. Due to the hardware requirements and associated processing costs, this workflow is resource-intensive and expensive, making it unattainable for creators and researchers with smaller budgets. Low-cost volumetric video systems using RGBD cameras offer an affordable alternative. As these small, mobile systems are a relatively new technology, the available software applications vary in terms of workflow and image quality. In this paper we provide an overview of the technical capabilities of sparse camera volumetric video capture applications and assess their visual fidelity and workflow.Materials and methodsWe selected volumetric video applications that are publicly available, support capture with multiple Microsoft Azure Kinect cameras and run on consumer-grade computer hardware. We compared the features, usability, and workflow of each application and benchmarked them in five different scenarios. Based on the benchmark footage, we analyzed spatial calibration accuracy, artifact occurrence and conducted a subjective perception study with 19 participants from a game design study program to assess the visual fidelity of the captures.ResultsWe evaluated three applications, Depthkit Studio, LiveScan3D and VolumetricCapture. We found Depthkit Studio to provide the best experience for novel users, while LiveScan3D and VolumetricCapture require advanced technical knowledge to be operated. The footage captured by Depthkit Studio showed the least amount of artifacts by a larger margin, followed by LiveScan3D and VolumetricCapture. These findings were confirmed by the participants who preferred Depthkit Studio over LiveScan3D and VolumetricCapture.DiscussionBased on the results, we recommend Depthkit Studio for the highest fidelity captures. LiveScan3D produces footage of only acceptable fidelity but is the only candidate that is available as open-source software. We therefore recommend it as a platform for research and experimentation. Due to the lower fidelity and high setup complexity, we recommend VolumetricCapture only for specific use-cases where its ability to handle a high number of sensors in a large capture volume is required.
Purpose:Virtual reality (VR) technology has emerged as a promising tool for physicians, offering the ability to assess anatomical data in 3D with visuospatial interaction qualities. The last decade has witnessed a remarkable increase in the number of studies focusing on the application of VR to assess patient-specific image data. This systematic review aims to provide an up-to-date overview of the latest research on VR in the field of surgical planning.Approach:A comprehensive literature search was conducted based on the preferred reporting items for systematic reviews and meta-analyses covering the period from April 1, 2021 to May 10, 2023. It includes research articles reporting on preoperative surgical planning using patient-specific medical images in virtual reality using head-mounted displays. The review summarizes the current state of research in this field, identifying key findings, technologies, study designs, methods, and potential directions for future research.Results:The selected studies show a positive impact on surgical decision-making and anatomy understanding compared to other visualization modalities. A substantial number of studies are reporting anecdotal evidence and case-specific outcomes. Notably, surgical planning using VR led to more frequent changes in surgical plans compared to planning with other visualization methods when surgeons reassessed their initial plans. VR demonstrated benefits in reducing planning time and improving spatial localization of pathologies.Conclusions:Results show that the application of VR for surgical planning is still in an experimental stage but is gradually advancing toward clinical use. The diverse study designs, methodologies, and varying reporting hinder a comprehensive analysis. Some findings lack statistical evidence and rely on subjective assumptions. To strengthen evaluation, future research should focus on refining study designs, improving technical reporting, defining visual and technical proficiency requirements, and enhancing VR software usability and design. Addressing these areas could pave the way for an effective implementation of VR in clinical settings.
A significant challenge in image-guided surgery is the accurate measurement task of relevant structures such as vessel segments, resection margins, or bowel lengths. While this task is an essential component of many surgeries, it involves substantial human effort and is prone to inaccuracies. In this paper, we develop a novel human-AI-based method for laparoscopic measurements utilizing stereo vision that has been guided by practicing surgeons. Based on a holistic qualitative requirements analysis, this work proposes a comprehensive measurement method, which comprises state-of-the-art machine learning architectures, such as RAFT-Stereo and YOLOv8. The developed method is assessed in various realistic experimental evaluation environments. Our results outline the potential of our method achieving high accuracies in distance measurements with errors below 1 mm. Furthermore, on-surface measurements demonstrate robustness when applied in challenging environments with textureless regions. Overall, by addressing the inherent challenges of image-guided surgery, we lay the foundation for a more robust and accurate solution for intra- and postoperative measurements, enabling more precise, safe, and efficient surgical procedures.
Recent advances in synthetic imaging open up opportunities for obtaining additional data in the field of surgical imaging. This data can provide reliable supplements supporting surgical applications and decision-making through computer vision. Particularly the field of image-guided surgery, such as laparoscopic and robotic-assisted surgery, benefits strongly from synthetic image datasets and virtual surgical training methods. Our study presents an intuitive approach for generating synthetic laparoscopic images from short text prompts using diffusion-based generative models. We demonstrate the usage of state-of-the-art text-to-image architectures in the context of laparoscopic imaging with regard to the surgical removal of the gallbladder as an example. Results on fidelity and diversity demonstrate that diffusion-based models can acquire knowledge about the style and semantics in the field of image-guided surgery. A validation study with a human assessment survey underlines the realistic nature of our synthetic data, as medical personnel detects actual images in a pool with generated images causing a false-positive rate of 66%. In addition, the investigation of a state-of-the-art machine learning model to recognize surgical actions indicates enhanced results when trained with additional generated images of up to 5.20%. Overall, the achieved image quality contributes to the usage of computer-generated images in surgical applications and enhances its path to maturity.
Background Surgical training is primarily carried out through observation during assistance or on-site classes, by watching videos as well as by different formats of simulation. The simulation of physical presence in the operating theatre in virtual reality might complement these necessary experiences. A prerequisite is a new education concept for virtual classes that communicates the unique workflows and decision-making paths of surgical health professions (i.e. surgeons, anesthesiologists and surgical assistants) in an authentic and immersive way. For this project, media scientists, designers and surgeons worked together to develop the foundations for new ways of conveying knowledge using virtual reality in surgery. Materials and method A technical workflow to record and present volumetric videos of surgical interventions in a photorealistic virtual operating room was developed. Situated in the virtual reality demonstrator called VolumetricOR, users can experience and navigate through surgical workflows as if they are physically present. The concept is compared with traditional video-based formats of digital simulation in surgical training. Results VolumetricOR let trainees experience surgical action and workflows (a) three-dimensionally, (b) from any perspective and (c) in real scale. This improves the linking of theoretical expertise and practical application of knowledge and shifts the learning experience from observation to participation. Discussion Volumetric training environments allow trainees to acquire procedural knowledge before going to the operating room and could improve the efficiency and quality of the learning and training process for professional staff by communicating techniques and workflows when the possibilities of training on-site are limited.
Medien und interaktiven Anwendungen wird zunehmend Prozessorund Sensortechnik verbaut, die es ermöglicht, Bilder an ihre Umwelt anzupassen und dabei auf Eingaben und Situationen in Echtzeit zu reagieren. Bild, Körper und Raum werden miteinander verschaltet und synchronisiert, mit langfristigen Folgen für die menschliche Wahrnehmung, für Handlungen und Entscheidungen. Die erweitert en Möglichkeiten bedingen neue Abhängigkeiten von Technologien und von den ästhetischen und operativen Vorgaben jener, die diese Technologien gestalten und bereitstellen.
Digital images increasingly determine the way people interact with physical space. Combined imaging and sensing technologies register, process, and transmit information about the physical world in real time and make it possible to continuously adapt such images to specific spatio-temporal settings and in relation to motion and perspective. With the ability to integrate situative and customised information in media, like digital maps or virtual reality applications, images also gain in importance for perception and interpretation. Such integration of image, action, and space heralds a new type of visual media described as adaptive images. Based on cases from industrial production, medicine, and psychotherapy as well as from sports and entertainment, the paper addresses their aesthetic, spatial, and operational conditions, and provides a typological survey of adaptive images as a phenomenon, including their respective challenges and implications for image and media theory.