
Laser tissue soldering (LTS) represents a promising alternative to conventional suturing methods in minimally invasive surgery (MIS). This paper presents a novel robotic control framework that aims to automate the LTS approach. The method involves attaching a soldering laser to a laparoscopically-guided robotic arm. The pose of the external laparoscope is estimated online, while detecting the soldering paste simultaneously utilizing a segmentation algorithm. The presented vision-based control shows promising results, as it is able to track the soldering paste in real-time with a maximum average deviation from the target path of 5mm, while reacting to camera and tissue movements.
Correct pedicle screw placement is essential in scoliosis surgery to ensure patient safety and surgical success. We propose a method to compute screw placements from vertebrae CT scans by mapping the entry point from a dictionary and then optimize the placement with particle swarm optimization. Our approach was tested on T6, T12, and L4 vertebrae and evaluated using the Gertzbein-Robbins classification (GR), the Zdichavsky grading system (ZG). Furthermore, a specialized surgeon assessed the placements. The results indicate that 89.58% of the total placements were valid. Among these valid placements, 46.51% were rated an A according to the GR and the other 53.49% were rated as B. All valid placements were Ia according to ZG. The medical assessment showed that 41.86% and 20.93% were optimal placements and good placements respectively.
The analysis of the liver, including its vessels and lesions, is essential for pre-operative intervention planning. Automated deep learning-based segmentation provides a time-efficient method for consistently localizing and visualizing these structures. However, the integration of deep learning algorithms into clinical practice necessitates thorough testing and validation. It is crucial not only to assess quantitative and volumetric metrics but also to evaluate these algorithms clinically through experienced radiologists and surgeons. In this study, we used an independent test set of 107 computed tomography scans to conduct a multi-user, multi-disciplinary analysis of a deep learning-based segmentation results. Six experienced clinicians-three surgeons, one surgical resident and two radiologists-evaluated the predictions made by the deep learning models. While liver segmentation can be considered largely solved, vessel segmentation still requires improvement. By manually identifying missing lesions (false negatives) as well as accepting and rejecting proposed lesions, performance metrics can be calculated for the lesion segmentation without needing a reference segmentation. Notably, lesion segmentation achieved an F1-score of 0.82 for lesions larger than 1 cm in diameter.
The angiography suite for endovascular interventions has seen rapid medical and technical innovation in the past decade. In this highly complex work environment, context-aware support systems could simplify the interaction with the X-ray imaging system and therefore might increase efficiency. To learn about the context of endovascular interventions and recognize workflow phases, we collected and annotated a dataset of mechanical thrombectomy interventions. A classification model was trained for different tasks on several compositions of the dataset to gain an understanding for the challenges that arise when expanding the research field of surgical workflow detection to neuroradiological X-ray data. Implications were found that, despite limited visibility, device positioning is highly relevant and learned implicitly during model training. This paves the way to further enhance workflow phase recognition on interventional X-ray data by the integration of temporal context.
Augmented Reality (AR) has the potential to assist in neurosurgical procedures by providing real-time visual guidance. Previous work superimposes the cranial ventricular system (obtained from pre-operative data) on the patient’s head using AR. For this, pre-operative data needs to be registered to the intra-operative patient head. But currently, this registration is performed manually, which is a time-consuming task.We present preliminary work on an automated Computed Tomography (CT) image preprocessing and 3D registration pipeline for AR-assisted ventricular punctures. Our approach combines an automatic CT image segmentation and preprocessing with the Iterative Closest Point (ICP) algorithm to perform the registration. Experiments demonstrate a promising registration accuracy of 1.833 ± 0.395 mm for points of interest inside the ventricular system. Moreover, the execution duration of 0.8265 ± 0.002 s addresses the manual and timeconsuming registration used in previous work (100.5 s).
In radioguided surgery, G-probes are used intraoperatively to localize targets marked by radionuclides. However, the interpretation of G-probe measurements is challenging due to background activity from surrounding organs. This work investigated whether deep neural networks can localize radiation sources in such environments and how different background activities impact this task. A physics-guided forward model simulated G-probe measurements for different intra-abdominal distributions, including an anatomically inspired bladder activity. A convolutional neural network was trained on simulated measurements to predict 3D target source positions. Results indicated that prediction accuracy improved with more Gprobe measurements and degraded with increased background activity. In particular, proximity to high-activity regions like the bladder significantly reduced accuracy. This study demonstrates the need to consider background activity distributions for target localization and that a convolutional neural network could solve this task.
This study evaluates the use of vibroacoustic (VA) sensing to detect material transitions during needle insertions into layered synthetic phantoms. Experiments were conducted using foams with different pore densities and configurations including internal cavities. Time-domain and wavelet analyses showed that denser foams produced more frequent vibrational events, while less dense foams resulted in sparser signals. In setups with elastic membranes and cavities, differences in time-frequency patterns were observed: high-frequency activity appeared after membrane rupture but was absent in cavity insertions. These results support the use of VA sensing to identify mechanical changes during minimally invasive procedures.
Time needed for navigation plays an important role in success of thrombectomies and other endovascular interventions. Thus, precise control of the catheter is crucial for the intervention’s success. We experimentally compare different interactive control strategies for endovascular robots with and without haptic feedback on one common input device while navigating a guidewire through an aortic arch. Using both a manual-catheter-manipulation-inspired control method and haptic feedback of the instruments tip force led to up to 33 % shorter intervention time and 49 % reduced task load, measured using the NASA TLX, compared to joystick control without haptic feedback.
Mitral valve insufficiency can be surgically corrected by minimally invasive mitral valve surgery. This approach is difficult to learn and training opportunities are still limited. Therefore, we built a physical mitral valve flow simulator to accommodate patient-individual valve models. The models must fulfill certain material properties to replicate adequate dynamic behavior, while being surgically modifiable. We developed two approaches for manufacturing complex anisotropic physiological mitral valve tissue, the first one being a thermoformed thermoplastic urethan (TPU) film and the second one being a mesh reinforced silicone cast into a 3D printed mold. The valve models were mounted into a pulsatile flow simulator, on which their hemodynamic performance before and after ring annuloplasty was evaluated by common flow measurements. Out of the presented approaches, the anisotropic mesh reinforced valve performed best, creating larger coaptation height and thereby enabling higher endsystolic pressures and significantly larger cardiac outputs.
This work demonstrates the detection and identification of pulsatile activity using proximal vibroacoustic sensing in laparoscopic instruments. A controlled setup simulating vascular pulsations was used to evaluate two signal processing methods: Continuous Wavelet Transform and autocorrelation. Both approaches identified pulsations up to 2 cm deep with high accuracy. These findings support the exploration of proximal sensing for intraoperative identification of vascular structures in robot-assisted surgery, where direct haptic feedback is absent and visual cues may be insufficient.
In the literature, the proximal and distal interphalangeal joints (PIP and DIP) are usually described as singleaxis hinge joints, whereas the metacarpophalangeal (MCP) joint is typically described as a two-axis joint. Biomechanical research studies revealed that due to the anatomical incongruence between the articular head (condyle) and the concave base of the adjacent phalanx, there exist additional rotational and translational movements. These findings indicate that the human finger joint is not a single-axis joint but rather a dimeric joint system. This study focuses on the accurate kinematic model to represent the configuration space and the biomechanical forces of the dimeric system. The configuration space of the single-axis hinge joint model is a subspace of the dimeric system’s configuration space. Vice versa, the dimeric system’s configuration space is a manifold over the single-axis system’s configuration space, enabling an uncountable set of configurations for each configuration of the subspace. This study enhances the understanding of the human finger kinematics as a basis for developing more realistic biomechanical models.
The Digital Imaging and Communications in Medicine (DICOM) standard has evolved from a radiological data exchange protocol into a foundational element for surgical informatics. DICOM Working Group 24 - “Surgery” aims to facilitate intra-operative usage of the DICOM standard to optimize surgical procedures. This paper presents a structured overview of DICOM-based use cases in surgery, spanning preoperative planning, intraoperative guidance, and postoperative documentation. Based on a comprehensive user story schema, we categorize and analyze over 20 real-world scenarios that demonstrate how DICOM services-such as segmentation (DICOM-SEG), structured reporting (DICOMSR), and Unified Procedure Steps (UPS)-enable interoperability, precision, and automation in surgical environments. We highlight the benefits of standardized data exchange for surgical planning, intraoperative decision support, and quality assurance, while also addressing integration with clinical information systems and data protection. The findings underscore DICOM’s potential to serve as a unifying framework for data-driven, context-aware surgical workflows.
Neurosurgical preoperative planning and training require high-fidelity models that accurately replicate the complexity of brain anatomy. Although commercial phantoms offer high visual and mechanical accuracy, they remain prohibitively expensive for training purposes.We present a method for matching tissue densities to suitable agarose gel concentrations and demonstrate a simple procedure for creating a multi-tissue, low-cost, and easily adjustable brain phantom. We also investigated storage stabilities for various agarose gel concentrations, finding that water-based gels at concentrations of 1% or higher exhibited minimal mass loss and maintained structural integrity over time when stored submerged in regular water.
MAVERICKAI transforms medical computing through touchless 3D interaction, adaptive AI integration, and immersive visualization. Designed for clinical and educational use, it enables sterile, intuitive manipulation of anatomical data and streamlines workflows via intelligent automation. This multifunctional platform enhances precision, efficiency, and training across disciplines-especially in surgery and otorhinolaryngology.
Vision-language models represent an emerging paradigm that leverages natural language to train vision systems with broad capabilities. Recently, the use of surgical lecture videos has emerged as a promising method for developing models capable of understanding surgical scenes. In this work, we aim to translate these developments to the domain of cardiac surgery, which is marked by heterogeneity and complexity of surgical cases. To this end, we curate a dataset of cardiac surgery lecture videos and augment the training dataset by using a Large Language Model (LLM) to extract procedural steps for each surgery. Preliminary results suggest that this form of data augmentation can enhance model performance on text-based video retrieval tasks.
Accurate left ventricular (LV) wall thickness (WT) estimation can identify arrhythmogenic ventricular tachycardia (VT) channels, but Laplace-based method is limited by computational inefficiency. This study proposes the Morphological Sphere Propagation (MSP) framework, a rapid, spherebased approach for LV WT mapping. MSP employs isotropic resampling followed by sequential endocardial dilation and epicardial contraction using a spherical structuring element, with thickness derived from propagation iterations scaled to image resolution. The proposed method was validated against a Laplace-based solver using synthetic phantoms with known ground-truth thickness, assessing computation time and mean absolute error (MAE). Additional validation was performed on clinical cardiac CT datasets from 10 ventricular tachycardia (VT) patients by evaluating CT channel agreement. Results demonstrated comparable accuracy (mean absolute error: MSP 0.49 ± 0.01 mm vs. Laplace 0.37 ± 0.03 mm) but 25× faster computation (MSP: 1.8 ± 0.1 s vs. Laplace: 47.0 ± 19.3 s). MSP showed strong agreement with Laplace in CT channel detection (Sensitivity = 0.85, PPV = 0.85). MSP offers a practical, accessible alternative for rapid arrhythmia substrate characterization without sacrificing accuracy.
Background: In the past decades, robots have transformed the healthcare sector by supporting clinical staff in various tasks. Applications of range from robot-supported surgical interventions and imaging, via healthcare logistics to cleaning and social robots. With the availability of natural language processing (NLP) methods based on large-language models (LLMs), the next generation of healthcare robots will in the next couple of years be able to communicate with humans in a completely new manner. Objective: To evaluate such NLPbased robots in various healthcare scenarios, adequate evaluation methods must be defined and implemented. Methods: For the evaluation of NLP-based robots a multi-dimensional framework is proposed, consisting of four dimensions and a living lab: D1: confidence, correctness and certifiability of LLMs for speech-based robots; D2: usability, acceptance and specifics of human-robot interaction (HRI); D3: the potentials to relieve clinical staff, and D4: the prospective technologyreflective analysis of HRI with respect to ethical, legal and social implications (ELSI) and its normative design. Additionally, LL: a living lab resp. a regulatory sandbox serves as a central hub for testing and validating future NLP-based robotic technologies under realistic conditions. Resume: Through this concept for evaluation, we intend to optimize the impact in the field of speech-based and no-code/low-code robotics.
Augmented Reality (AR) in surgery relies heavily on visual clarity, which is influenced by the interplay of physical and virtual illumination. This study evaluates how different color combinations of ambient room light and virtual illumination of the hologram affect the visibility, detail perception, and comfort of volume-rendered medical CT data on a HoloLens 2. Fifteen medically skilled participants assessed 21 lighting setups in a realistic operation room (OR) environment using a remote-rendered abdominal CT dataset. Results show that virtual light color significantly impacts perception, with white and orange lights performing best across all metrics. The physical light setting showed no significant influence. These findings support the optimization of virtual lighting of AR application in the OR.
Ultrasound (US)-guided needle interventions require precise coordination between imaging and instrument handling, which can be ergonomically challenging with conventional displays. This work presents a mixed reality (MR) prototype that integrates live US and pre-interventional CT data into a video passthrough head-mounted display (HMD), enabling a spatially registered in-situ view and interactive virtual displays. Unlike prior work relying on optical see-through HMDs, our system combines immersive guidance and interactive tools to support image-guided needle insertion. A preliminary user study with six radiologists assessed usability and clinical potential. Results indicate slightly above-average usability, with participants valuing field-of-view integration. User feedback highlighted the need for features such as interactive planning and extended visualizations, e.g. patientspecific anatomy.With these refinements, the prototype shows strong potential to improve ergonomics and spatial awareness during US-guided procedures.
Point clouds are a common 3D representation being increasingly used in various fields. However, they are rarely applied in medical applications. In video bronchoscopy, point cloud data can help to guide lung procedures more accurately by using 3D shape information instead of solely relying on RGB image data. This is especially useful because the appearance of tissues can vary greatly and thus becomes difficult to interpret. Current methods often require patient-specific CT scans and electromagnetic tracking, which limits their use in places such as intensive care units. In this work, we present a deep learning method that predicts airway labels directly from point cloud data without using CT scans or tracking systems. This makes our method more flexible and reliable, even in difficult conditions. Using 3D data, we reduce the impact of lighting, camera quality, and patient differences, promising better results during medical procedures.