
Volumetric liver ultrasound (US) plays an important role in clinical diagnosis but remains highly dependent on operator expertise, particularly during target view localization. This study aims to develop an automatic probe guidance framework that reduces reliance on manual demonstrations and additional sensing hardware, while supporting robust volumetric liver US acquisition. We propose an image-based imitation learning framework that learns probe guidance policies from a virtual expert in a simulated US scanning environment. A simulation pipeline is constructed using cross-modal medical images and a hybrid US simulator that combines physics-based ray casting with generation-based image synthesis to produce anatomically consistent and acoustically realistic US images. Optimal scanning trajectories are generated based solely on target views typically available in clinical practice. To improve robustness, pose-level and image-level data augmentations are introduced, and US observations are encoded into an anatomy-aware state representation for intercostal liver scanning. Experimental results in simulation and real clinical data demonstrate that the proposed framework achieves accurate and stable target view localization for volumetric liver US acquisition. Compared with baseline and ablated models, the method shows improved localization accuracy, increased liver coverage, and reduced rib interference, while maintaining robustness across different anatomical conditions. This work presents a data-efficient and clinically practical solution for automatic probe guidance in volumetric liver US. By leveraging realistic simulation, virtual expert demonstrations, and anatomy-aware image representations, the proposed framework enables effective learning of probe movements without requiring manual trajectory annotations or additional sensors. The results suggest strong potential for integration into computer-assisted and robotic US systems.
The ongoing 6G standardization process is largely driven by technical advancements, but with limited consideration of sector-specific needs, particularly in health care. The current scientific literature primarily focuses on technical aspects, offering minimal empirical input from healthcare professionals, patient representatives, or medical device manufacturers. This neglects clinical acceptance and applicability as a critical factor in technology adoption. The objective of this study is to explore the hopes and concerns of healthcare stakeholders in the context of 6G standardization, to support improved alignment between technological development and clinical practice. To address this gap, we conducted semi-structured interviews with key healthcare stakeholders in Germany and derived recommendations for technology providers. The findings highlight insufficient interoperability, fragmented systems, and inadequate network coverage as central challenges in today’s daily practice. Manual documentation and data transfers further burden staff under conditions of shortage and time pressure. A robust infrastructure, particularly in rural areas where telemedicine is growing, is considered essential. Concerns persist regarding data protection and security, alongside unresolved liability for automated systems and skepticism about increasing technological dependence. Nevertheless, stakeholders see substantial opportunities in 6G for simplifying workflows, enabling seamless device integration, and supporting use cases such as real-time imaging, remote monitoring, and cross-institutional data sharing. This study contributes to the 6G standardization by integrating the healthcare perspectives of medtech stakeholders and clinical end users, thereby supporting the development of patient-centered and clinically relevant communication solutions. The findings provide practical guidance for aligning future 6G solutions with real-world clinical requirements.
This paper presents the design, development and evaluation of a novel robotic platform for endoscopic ultrasound-guided fine-needle biopsy of liver lesions. The system combines a four degrees-of-freedom (DoF) two-segment tendon-driven continuum robot (TDCR) endoscope with a two DoF superelastic nickel–titanium bevel-tip steerable needle. Needle path planning is achieved using NeedleNav, a soft actor–critic (SAC) deep reinforcement learning (DRL) model that generates collision-free trajectories to deep-seated lesions. This represents one of the first integrated systems combining a TDCR, steerable needle and DRL-based navigation, and the first application of a SAC to liver lesion targeting. Evaluation of the TDCR through tip tracking of circular, diamond-shaped and arc trajectories demonstrated a mean absolute error (MAE) of 13.19 mm. NeedleNav converged to obstacle avoidance trajectories in ∼ 2500 training episodes. Needle curvature was augmented by hand-fabricating notches on its distal section. Two needles with a 3-cm and 8-cm notched section were evaluated in a gelatine liver phantom, achieving an MAE of 21.78 mm and 14.86 mm, respectively, for obstacle avoidance trajectories. The system demonstrated observable path deflection compared to obstacle-free trajectories for the same targets. Together, these findings suggest the feasibility of our proposed solution, expanding the reach of endoscopic needle interventions to deep-seated lesions in the right lobe.
Achieving fine-grained understanding of surgical gestures remains a fundamental challenge in computer vision, due to the subtle and temporally overlapping nature of surgical motions. Gesture boundaries, where transitions between surgical actions occur, present challenges for precise temporal localization. We propose a temporal boundary analysis framework that improves overall surgical gesture segmentation by explicitly modeling transitions between actions. While most existing methods rely on both RGB and kinematic data, our approach operates on RGB-only video, without requiring additional annotations or computational overhead at inference. We introduce a temporal boundary distillation module (TBDM) that leverages privileged information during training to learn boundary-aware features. TBDM employs cross-attention between class-present and class-absent temporal regions derived from ground-truth annotations, explicitly encoding transition information. A lightweight projection layer learns boundary-aware features through knowledge distillation from TBDM, supervised by classification and distillation loss (MSE). At inference, only the trained projection layer is required, resulting in no additional computational cost. We evaluated TBDM on CholecT50 and RARP-45 surgical datasets. TBDM consistently improved baseline models across all metrics, achieving up to +8.5 edit score improvement on CholecT50. On RARP-45, our approach achieved state-of-the-art edit score (81.4) and F1@50 (77.9), demonstrating effectiveness across different architectures and datasets. TBDM provides a generalized, plug-and-play framework for fine-grained surgical gesture segmentation, using RGB-only data. By explicitly modeling temporal boundaries, it achieves consistent improvement across multiple architectures and surgical datasets without increasing inference complexity. Code is publicly available at https://github.com/ezemsuraekmekci/TBDM-Surgical-Gesture-Segmentation .
Proper management of hand fractures requires accurate reduction of bone fragments and stabilization with Kirschner wire (K-wire) fixation. Limited case exposure in residency and a shortage of feedback-rich skills assessment constrain deliberate practice. This work describes a mixed reality hand fracture simulator that integrates a realistic physical model with instrument tracking, real-time virtual image guidance, and quantitative performance feedback for training percutaneous K-wire fixation. We developed a benchtop mixed reality (MR) simulator integrating an electromagnetic (EM) tracker, a three-dimensional (3D)-printed hand model with fracture patterns, and a microcontroller-enabled K-wire driver. To ensure high physical-to-virtual geometric correspondence, the 3D-printed components were optically scanned to serve as virtual representations. A custom application provides real-time 3D visualization, guidance, and simulated virtual X-rays. To evaluate system accuracy, five users performed repeated paired-point rigid registrations. Accuracy was assessed by calculating fiducial registration error (FRE) and target registration error (TRE) across 25 attempts using a pre-calibrated stylus and the tracked K-wire driver. Image guidance capabilities were evaluated by measuring positional and angular alignment errors of the tracked K-wire driver against predefined ideal trajectories. The overall mean FRE was 0.69 ± 0.11 mm for the hand model (11 fiducials) and 0.60 ± 0.09 mm for the K-wire driver (8 fiducials). Target localization to six ground-truth targets yielded a mean TRE of 0.69 ± 0.10 mm with the stylus and 0.82 ± 0.16 mm with the driver. Trajectory alignment demonstrated maximum positional variance during parallel approaches to the bone (K-wire 1), highlighting specific procedural challenges. The MR hand fracture simulator platform enables repeated practice of fracture reduction and K-wire placement with interactive guidance, simulated X-rays, and objective metrics. The system is able to achieve millimeter-scale alignment suitable for procedural training. Future work will evaluate learning outcomes and skill transfer in surgical trainees.
Ablation therapies are a treatment option for cancer patients, particularly for conditions such as spinal metastases and liver tumors. Precisely delineating ablation zones is essential for accurately assessing treatment success. However, research in MRI-guided interventions remains limited for automated segmentation approaches and quantitative analysis. We developed a framework for automated segmentation of ablation zones following thermal interventions. The performance was tested on two representative types of clinical cases from different clinical sites: post-ablative liver lesions and spinal metastases. Four leading neural networks (nnUNet, TransUNet, SwinUNETR, and SwinUNETR-V2) were evaluated for their segmentation accuracy in segmenting necrotic tissue. Additionally, a statistical analysis was performed to investigate the influence of an optimized image ROI selection on the segmentation performance. The nnUNet achieved the highest segmentation performance, with a Dice Similarity Coefficient of 83.3 ± 13.2
The transition toward competency-based medical education requires scalable, objective surgical skill assessment. While sensor-based approaches remain burdensome, video-based deep learning models often lack the interpretability required for formative feedback. We introduce a fully automated, markerless framework for assessing vascular open surgery suturing skills providing meaningful interpretation. Our pipeline leverages 3D hand tracking coupled to an LSTM-attention network for precise temporal segmentation. By merging explicit kinematic metrics with latent deep features, we develop a voting ensemble classifier categorizing surgeons into three skill levels. We further evaluated the impact of ground-truth reliability by comparing models trained on single-assessor versus multi-assessor consensus labels. The framework achieved 96.4
The Critical View of Safety (CVS) is a clinically mandated criterion for reducing bile duct injury during laparoscopic cholecystectomy; however, its automated recognition from endoscopic video remains challenging due to visual occlusion, depth ambiguity, and the inherently progressive nature of CVS formation across time. This work aims to develop a robust and clinically interpretable framework for automated multi-label CVS recognition from monocular surgical video, without relying on auxiliary segmentation supervision or stereo hardware. We propose a novel spatio-temporal CVS recognition framework that integrates depth-augmented visual representations with efficient temporal state-space modeling. The approach employs a DINOv2-based vision foundation model extended to RGB-D input for robust self-supervised spatial representation learning, alongside a Swin Transformer for hierarchical and anatomically localized feature encoding. To capture the temporal persistence and geometric consistency required for CVS identification, a depth-aware Spatio-temporal Mamba module is introduced to model long-range temporal dependencies with linear computational complexity. The network produces multi-label predictions corresponding to individual CVS criteria. Experimental evaluations demonstrate that the proposed RGB-D spatio-temporal architecture consistently outperforms RGB-only and attention-based temporal baselines. The method shows improved robustness to visual occlusion, viewpoint variation, and dataset shift across evaluation scenarios. The proposed RGB-D spatio-temporal framework demonstrates that integrating weight-inflated foundation model adaptation, hierarchical spatial encoding, and depth-conditioned selective state-space temporal modeling yields reliable, scalable, and clinically interpretable CVS recognition. In contrast to prior video SSM approaches that process homogeneous RGB token sequences, the depth-aware gating and dual-encoder fusion introduced here provide targeted inductive biases suited to the geometric and temporal demands of surgical safety verification.
Reliable 3D reconstruction of colonoscopic scenes can facilitate navigation and enhance lesion assessment by indicating the regions that have been inspected and revealing the geometric structure of polyps. However, accurate depth estimation in colonoscopy remains highly challenging due to specular highlights and low-texture areas. This study aims to improve monocular depth estimation in colonoscopy by incorporating surface normal information. We propose a self-supervised deep learning method for relative depth estimation that integrates surface normal information through a cross-attention mechanism. Surface normals are first predicted by an existing normal estimation model and then used as auxiliary geometric priors. The depth estimation network employs cross-attention to adaptively fuse normal features with image features, enabling spatially selective geometric refinement. Quantitative and qualitative comparisons show that the proposed method outperforms state-of-the-art self-supervised methods on SimCol3D, C3VD, and real colonoscopy data. The proposed method generates more realistic depth maps, preserving mucosal folds and lumen geometry. Ablation studies verify that the cross-attention module is crucial for effectively exploiting normal information for accurate depth estimation. By leveraging cross-attention to fuse predicted surface normals with image features, the proposed method enhances monocular depth estimation accuracy and robustness in colonoscopy. This approach provides a general and lightweight strategy for incorporating geometric priors into self-supervised frameworks, offering potential benefits for downstream tasks such as 3D reconstruction and endoscopic navigation.
Handheld laparoscopic robots may enhance surgical quality while reducing medical resource use. These types of devices primarily combine the portability of traditional instruments with the flexibility of advanced surgical robotics. However, challenges related to the usability and intuitiveness of wrist-controlled operation modes in handheld robots have raised concerns. This study introduces a novel modular handheld electric surgical robot equipped with a joystick, a multi-layer cross-shaped Hooke hinge end effector, and a quick-change mechanism. The joystick exhibits a master-slave isomorphism with the continuum end effector. We conducted kinematic modeling of the end effector using continuous curvature models and proposed a joint incremental proportional control method. To evaluate the device’s performance, we constructed an optical 3D motion capture system to monitor the master-slave trajectory of the electric robotic prototype. The end effector achieved a combined yaw and pitch range of ± 90°, robot can execute multi-axis circular movements with excellent consistency. The consumable single quick-change time was less than 7 s. The maximum trajectory error was 3.43 mm, with a master-slave control delay of 0.304 s. Suturing operation with a laparoscopic simulator showed the robot can grasp and suture tissues more easily. This electric surgical device showed advantages in master-slave trajectory consistency, usability, and finger-controlled operation mode. The proposed control approach can also be applied to other rope-driven continuum robotics.
This study investigates the use of electrocardiogram (ECG) as a surrogate signal to model the liver’s respiratory-induced motion. A learning-based model was trained to predict respiratory-induced liver motion by relying exclusively on ECG data, without requiring additional imaging. A correspondence model based on an encoder–decoder architecture was defined to map internal liver motion from ECG signals. Experimental validation was conducted through a human subject study involving eight participants performing various breathing patterns. The mean absolute error during normal breathing was 2.83 mm, while the overall error considering all breathing patterns was 4.02 mm with correlation coefficients above 0.90. More than 90
Mixed reality (MR) simulations are increasingly used in medical training to provide safe, controllable, and repeatable practice environments. While visual fidelity (VF) and interaction fidelity (IF) are known to influence user experience (UX), their respective effects on performance and UX in complex, precision-based medical procedures remain insufficiently understood. We conducted a within-subjects user study investigating the effects of VF (high vs. low) and IF (tangible vs. virtual interaction) in a precise needle insertion task. A non-immersive setup served as a reference condition. Performance metrics, perceived realism, workload, and UX were assessed across conditions. Both fidelity factors successfully modulated perceived realism. Tangible interaction significantly improved insertion depth control, reduced workload, and enhanced UX. High VF primarily increased spatial presence and involvement, with only limited benefits for task performance. Interaction effects revealed that tangible interaction provided greater advantages under low VF, particularly in terms of task completion time and needle alignment. Visual and interaction fidelity play distinct and complementary roles in MR-based medical training. Tangible interaction predominantly supports procedural accuracy and efficiency, whereas high visual fidelity mainly enhances immersive engagement. MR system design should therefore align fidelity decisions with specific training objectives.
Selecting an optimal C-arm working view is critical in endovascular coiling of intracranial aneurysms, influencing procedural safety and efficiency. Current practice relies on operator experience and manual angulation adjustments, leading to inter-operator variability and increased radiation exposure. In this preliminary feasibility study, we investigate whether a generative adversarial framework can learn to predict clinically meaningful C-arm viewing directions from 3D rotational angiography (3DRA) data. We developed a generator–discriminator framework in which the generator predicts a 3D unit view vector from segmented vascular and aneurysm volumes, and the discriminator evaluates differentiable 2D projections along the predicted view. Three projection strategies (soft first-hit ray casting, digitally reconstructed radiographs (DRR), and label-aware maximum-intensity projections with soft label encoding (MIP-OR)) were tested with two generator architectures (CNN and U-Net). Performance was assessed on a held-out test set using absolute dot product (ADP) error between predicted and expert-annotated views, complemented by qualitative assessment from two experienced interventional neuroradiologists. Experiments on a preliminary cohort of 18 patients demonstrate that the CNN generator combined with MIP-OR projections achieves the lowest mean ADP error and high expert scores, demonstrating improved alignment with clinically acceptable working views. Training analysis indicated that low ADP values alone do not guarantee clinical usability, highlighting the need for combined geometric and expert-based evaluation. Adversarial learning with anatomically informed projections is a promising approach for automated C-arm view prediction. Future work will integrate procedural ground-truth views, larger datasets, and additional endovascular procedures to evaluate reliability and generalizability.
Image-guided surgical navigation can enhance surgical precision and safety. The rise of robot-assisted surgeries has increased the need of incorporating navigation systems. Our goal is to evaluate the integration of three novel techniques on registration, instrument tracking and visualization with navigation using the robot-assisted sentinel lymph node biopsy (SLNB). A prospective feasibility study was performed at the Netherlands Cancer Institute (November 2023–July 2025). In total, 30 patients scheduled for SLNB participated. First, registration and localization accuracy were analyzed. Secondly, surgeons answered questionnaires to evaluate the ultrasound registration, instrument tracking and enhanced visualization. Virtual reality (VR) and augmented reality (AR) were compared in the last ten patients. Feasibility was determined with the sentinel nodes (SN) percentage that were successfully localized by navigation and validated ex vivo with the gamma probe. Per-protocol, the first ten patients were used for workflow optimization; the subsequent 20 patients were included in the analysis. Ultrasound registration was fast (7 min), accurate (0.1 cm) and intuitive. Tracked instruments were well-integrated and helped to localize precisely the SN (0.3 cm). Surgeons preferred VR while opening the resection path to find the SN and AR in proximity for exact localization. Using navigation, successful identification was achieved in 90
Accurate three-dimensional representations of lumbar vertebral anatomy are essential for spinal research and clinical decision-making, particularly biomechanical analyses and the development of patient-specific interventions. However, in many practical settings, only partial information is available, making it difficult to obtain complete vertebral geometries. Statistical Shape Models (SSMs) provide a powerful way to characterize population-level anatomical variability, while Gaussian Process Regression (GPR) enables the reconstruction of full three-dimensional shapes from limited surface information. To reconstruct complete vertebral geometries from partial anatomical information using SSM and GPR and determine the minimum partial information needed for clinically acceptable reconstruction. Thirteen high-resolution CT datasets of healthy adult lumbar spines were segmented. A two-step registration framework was implemented: rigid registration followed by 3D-3D embedded deformation non-rigid registration. Principal Component Analysis (PCA) was used to generate SSMs of the lumbar spine. GPR was then employed for shape reconstruction from partial input data. Reconstruction performance was assessed with a leave-one-out cross-validation method. In the full lumbar spine SSM, the first eight principal modes captured 92.7
Efficient patient flow management is a key factor for optimizing workflows in cost-intensive hospital units such as operating rooms. While existing prediction models can accurately estimate surgical durations, uncertainties in patient transport and availability of resources continue to affect daily schedules. Real-time location systems offer the potential to address this gap by providing information on patient and device locations. However, commercially available solutions are often costly and closed by providers. In this work, we present a scalable, research-oriented real-time location system, designed for low-power and low-cost deployment in hospital environments. Our system uses Bluetooth Low Energy beacons for device tracking and OpenThread as a transport layer to a database. A custom localization algorithm enables room-level assignment and triangulation-based localization in terms of coordinates. This setup is designed for low-power and low-cost deployment for research purposes. Preliminary results demonstrate the feasibility of the proposed network for room-level assignment and position estimation of beacons while revealing notable variability in accuracy. Triangulation-based localization achieved a mean positioning error of 3.53 m, while room-level assignment showed error rates ranging from 0 to 40
Esophageal involvement is a common and early manifestation of systemic sclerosis (SSc). This involvement is mainly evaluated by esophageal manometry, but this test is invasive and provides limited morphological information. This study aimed to identify CT-based radiomics features related to esophageal involvement and to evaluate whether the presence or absence of esophageal involvement can be classified with high accuracy. Chest CT images from 64 subjects were analyzed, including 30 patients with SSc and 34 normal subjects. The esophageal region was automatically segmented, and radiomics features were extracted from ellipse-fitted regions of interest on axial slices. Features showing significant differences were selected, and feature combinations were evaluated using exhaustive search. Classification was performed using XGBoost. First-order statistical features at levels corresponding to the third to fifth thoracic vertebrae were significantly associated with the presence of esophageal involvement. Classification using these features showed high accuracy when five to six features were used, indicating that the presence or absence of esophageal involvement can be identified accurately even with a small number of features. CT-based radiomics analysis may be a useful noninvasive method for evaluating esophageal involvement. Future studies incorporating pathological correlation and surrounding tissue information may further improve clinical applicability.
The purpose of this study is to propose a method for automatically estimating the position and angle of the acoustic window required to acquire the apical four-chamber view in echocardiography. The apical four-chamber view is clinically important, but it requires accurate probe placement near the cardiac apex. Acquisition of the apical four-chamber view is technically challenging because the acoustic window available for transthoracic ultrasound transmission is restricted by the intercostal space, and the cardiac apex and endocardial borders can be difficult to delineate due to rib shadowing and acoustic attenuation caused by the chest wall and subcutaneous tissue. Therefore, this study proposes a method for automatically acquiring cross sectional images without directly detecting the apex. The proposed method estimates the apical four-chamber view through a three-step procedure. First, starting from the parasternal long-axis view, the probe is swept obliquely along the body surface while maintaining the mitral valve near the center of the ultrasound image and gradually rotating the probe toward the apical region. This maneuver is continued through the acoustic window where the cardiac apex can be visualized until the mitral valve leaflets are no longer visible. Second, the probe is rotated and angled to visualize the tricuspid valve, mitral valve, and interventricular septum in a single imaging plane. Finally, the probe is tilted and rocked to optimize the apical four-chamber view. Proof-of-concept experiments demonstrated that the proposed framework could estimate probe positions and orientations for acquiring apical four-chamber views without direct apex detection. Across five subjects, the system achieved an image quality score of 68.0
Computer-integrated surgical navigation systems are often run in the OR with the assistance of a technician for controlling the user interface and advising on technical details of the system. In lower-resource healthcare settings, limited access to additional OR staff and technical training for operating navigation systems can represent a barrier to sustainably deploying a low-cost surgical navigation solution. Recent advancements in locally deployable large language models (LLMs) have improved their ability to answer technical questions based on source materials and safely perform limited tasks on behalf of a user. The objective of this paper is to explore the feasibility of using a network of local LLM-based agents to act as a natural language interface for a low-cost surgical navigation system, facilitating hands-free manipulation of the user interface and providing documentation-grounded technical guidance. We propose the navigation offline virtual agent (NOVA), an end-to-end architecture that integrates distinct local LLM-based agents for knowledge tasks, action tasks, and delegation between the agents. Two semisynthetic benchmark datasets were generated for ablation studies of individual agents, and a prototype was built which integrates these agents into NousNav, an open-source neuronavigation system. The agents based on local LLMs were found to perform comparably to closed-source, commercially hosted LLMs. In a user study with nine participants, NOVA facilitated hands-free patient registration with a mean end-to-end latency of 10.4± 2.4 s and an 81