Surgical video understanding is a crucial prerequisite for advancing Computer-Assisted Surgery. While vision-language models (VLMs) have recently been applied to the surgical domain, existing surgical vision-language datasets lack in capturing and evaluating complex, interleaved spatial-temporal dynamics. Creating large scale datasets that accurately represent fine-grained spatial-temporal relationships in surgical videos is challenging due to costly manual annotations or error-prone generation using large language models. To address this gap, we introduce the SurgSTU-Pipeline, a deterministic generation pipeline featuring temporal and spatial continuity filtering to reliably create surgical datasets for fine-grained spatial-temporal multimodal understanding. Applying this pipeline to publicly available surgical datasets, we create the SurgSTU dataset, comprising 7515 video clips densely extended with 150k fine-grained spatial-temporal question-answer samples. Our comprehensive evaluation shows that while state-of-the-art generalist VLMs struggle in zero-shot settings, their spatial-temporal capabilities can be improved through in-context learning. A fine-tuned VLM on the SurgSTU training dataset achieves highest performance among all spatial-temporal tasks, validating the dataset's efficacy to improve spatial-temporal understanding of VLMs in surgical videos. Code will be made publicly available.
Corrections are done in [1] in the section authors affilations. Vidas Raudonis affilation where corrected from “$^{4}$SustAInLivWork Center of Excellence, Kaunas, Lithuania” to “$^{4}$SustAInLivWork Center of Excellence, Kaunas, Lithuania and $^{5}$Faculty of Electrical and Electronics Eng., Kaunas University of Technology, Kaunas, Lithuania”
Endoscopic sinus surgery requires careful preoperative assessment of the skull base anatomy to minimize risks such as cerebrospinal fluid leakage. Anatomical risk scores like the Keros, Gera and Thailand-Malaysia-Singapore score offer a standardized approach but require time-consuming manual measurements on coronal CT or CBCT scans. We propose an automated deep learning pipeline that estimates these risk scores by localizing key anatomical landmarks via heatmap regression. We compare a direct approach to a specialized global-to-local learning strategy and find mean absolute errors on the relevant anatomical measurements as low as 0.506 mm for the Keros, 4.516 degrees for the Gera and 0.802 mm / 0.777 mm for the TMS classification, each corresponding to the best-performing model for the respective measurement.
The Pelvic Rosetta Classification (PRC) project aimed to develop an interdisciplinary, landmark-based pelvic lymph node map for patients with prostate cancer to improve communication between imaging specialists and urologists. Methods: After an intense development phase, we conducted 3 evaluation rounds including 19 clinical experts having consensus meetings after each evaluation round. Experts contoured lymph node areas (LNA) for 2 patients with prostate cancer. Contours were assessed qualitatively and quantitatively. The PRC was further validated by assignment of 30 prostate-specific membrane antigen PET/CT-positive lesions to LNAs. The interrater reliability was calculated using Fleiss κ. Based on the final PRC, a complete contour and a 3-dimensional model were created. Results: Eight pelvic (external iliac, cranial/caudal obturator fossa, dorsal internal iliac, vesico-prostatic pedicle, mesorectal/perirectal, presacral, preprostatic/retropubic) and 4 extrapelvic (common iliac, intercommon, sigmoid, inguinal) LNAs were defined using anatomic landmarks which are consistently recognizable on imaging and intraoperatively. Strong consensus between experts existed for smaller, well-defined LNAs (e.g., preprostatic/retropubic, mesorectal/perirectal LNAs) compared with regions with proportionally large borders (e.g., obturator fossa, vesico-prostatic pedicle LNAs). Overall, moderate agreement (κ = 0.53) was observed during validation. Discrepancies were mostly encountered for lesions adjacent to borders between LNAs. The final contour and 3-dimensional model were approved by all experts. Conclusion: The PRC project showed fair reproducibility and validity. Further external validation is needed to assess its influence on interdisciplinary communication and treatment outcomes.
Point tracking in surgery is crucial to enable applications in downstream tasks such as segmentation, 3D reconstruction, virtual tissue landmarking, autonomous probe-based scanning, and subtask autonomy. This paper introduces the 2025 iteration of a point tracking challenge to address this, wherein participants submit their algorithms for quantification. Their algorithms are evaluated using a dataset named surgical tattoos in infrared (STIR), with the challenge named the STIR Challenge 2025 (STIRC2025). The STIR Challenge 2025 comprises two quantitative components: accuracy and efficiency. The accuracy component tests the accuracy of algorithms on in vivo and ex vivo sequences. The efficiency component tests algorithm inference latency. The challenge was conducted as a part of MICCAI EndoVis 2025, and seven teams participated in this challenge. In this paper we summarize the challenge results and participant methods. The challenge dataset is available at: https://zenodo.org/records/20191078, and the code for baseline models and metrics calculation is available here: https://github.com/athaddius/STIRMetrics
Studying tissue samples obtained during autopsies is the gold standard when diagnosing the cause of death and for understanding disease pathophysiology. Recently, the interest in post mortem minimally invasive biopsies has grown which is a less destructive approach in comparison to an open autopsy and reduces the risk of infection. While manual biopsies under ultrasound guidance are more widely performed, robotic post mortem biopsies have been recently proposed. This approach can further reduce the risk of infection for physicians. However, planning of the procedure and control of the robot need to be efficient and usable. We explore a virtual reality setup with a digital twin to realize fully remote planning and control of robotic post mortem biopsies. The setup is evaluated with forensic pathologists in a usability study for three interaction methods. Furthermore, we evaluate clinical feasibility and evaluate the system with three human cadavers. Overall, 132 needle insertions were performed with an off-axis needle placement error of 5.30+-3.25 mm. Tissue samples were successfully biopsied and histopathologically verified. Users reported a very intuitive needle placement approach, indicating that the system is a promising, precise, and low-risk alternative to conventional approaches.
Optical coherence tomography (OCT) is a non-invasive volumetric imaging modality with high spatial and temporal resolution. For imaging larger tissue structures, OCT probes need to be moved to scan the respective area. For handheld scanning, stitching of the acquired OCT volumes requires overlap to register the images. For robotic scanning and stitching, a typical approach is to restrict the motion to translations, as this avoids a full hand-eye calibration, which is complicated by the small field of view of most OCT probes. However, stitching by registration or by translational scanning are limited when curved tissue surfaces need to be scanned. We propose a marker for full six-dimensional hand-eye calibration of a robot mounted OCT probe. We show that the calibration results in highly repeatable estimates of the transformation. Moreover, we evaluate robotic scanning of two phantom surfaces to demonstrate that the proposed calibration allows for consistent scanning of large, curved tissue surfaces. As the proposed approach is not relying on image registration, it does not suffer from a potential accumulation of errors along a scan path. We also illustrate the improvement compared to conventional 3D-translational robotic scanning.
The manual assessment of brain Magnetic Resonance Imaging (MRI) scans can be labor-intensive and time-consuming for radiologists. Deep Learning methods have demonstrated the potential to aid this process. However, their effectiveness relies on the availability of large, annotated data sets. Unsupervised Anomaly Detection (UAD) presents a promising alternative, offering the potential to identify and localize anomalies without per-pixel annotations. Instead, a normative distribution is learned using healthy data, enabling the identification of abnormalities as deviations. This allows UAD methods to detect abnormalities that were unseen during training. This appealing feature has led to numerous studies proposing innovations and novel approaches. In this work, we provide a review of the literature and systematically collect and compare the proposed approaches. We observe that UAD has made significant advancements in brain MRI analysis. However, individual approaches are often evaluated in different contexts, i.e., changes in acquisition parameters, pre- and post-processing, and anomaly scoring. This variability makes it challenging to assess which models perform best, underscoring the need for comprehensive comparative studies concerning the specific context of MRI scans. Our collection, featuring public data sets, research studies, and open implementations, is available at our GitHub repository https://github.com/FinnBehrendt/Unsupervised-Anomaly-Detection-in-Brain-MRI.
Zebrafish serve as a key pre-clinical model for disease research. Studying disease progression requires long-term tracking of fish motion, particularly when considering the function of the musculoskeletal apparatus. While vision-based tracking has been studied, current approaches typically require substantial fine tuning and often lack robustness, especially when multiple fish need to be tracked and motion occurs along all three spatial dimensions.We propose a simplified setup and a 3D tracking method using only two cameras and two convolutional neural networks (CNN) to localize and track multiple fish robustly. We separate the localization of fish and their classification as an individual. This allows for simple training data acquisition by placing fish into the basin individually, without the need for manual annotation.We evaluate the approach on sets of two, five, and ten fish resulting in multiple object tracking accuracies (MOTA) of 0.997, 0.994, and 0.966, respectively. The mean MOTA over all scenarios is 0.99 compared to a mean MOTA of 0.89 for a state-of-the-art tracking algorithm (idtracker.ai). Finally, we present comprehensive data on the motion activity of individual fish in a 30-minute video sequence, demonstrating robust tracking and unique activity profiles of each fish.
Recently, Vision Large Language Models (VLMs) have demonstrated high potential in computer-aided diagnosis and decision-support. However, current VLMs show deficits in domain specific surgical scene understanding, such as identifying and explaining anatomical landmarks during Complete Mesocolic Excision. Additionally, there is a need for locally deployable models to avoid patient data leakage to large VLMs, hosted outside the clinic. We propose a privacy-preserving framework to distill knowledge from large, general-purpose LLMs into an efficient, local VLM. We generate an expert-supervised dataset by prompting a teacher LLM without sensitive images, using only textual context and binary segmentation masks for spatial information. This dataset is used for Supervised Fine-Tuning (SFT) and subsequent Direct Preference Optimization (DPO) of the locally deployable VLM. Our evaluation confirms that finetuning VLMs with our generated datasets increases surgical domain knowledge compared to its base VLM by a large margin. Overall, this work validates a data-efficient and privacy-conforming way to train a surgical domain optimized, locally deployable VLM for surgical scene understanding.
KiMeKo (KI-Med-Kollaborationsplattform) is a publically funded collaborative research project that develops a sustainable AI-Med ecosystem for AI-based medical device development. The project runs from July 2024 to December 2027 and joins seven Northern German research institutions. KiMeKo addresses the complete development trajectory, from concept and data acquisition to validation, regulatory evidence generation, and approval-oriented documentation. The project contributes a practical toolchain and platform capabilities for non-experts and experts, including structured innovation support, uncertaintyaware sensor-data fusion, hybrid expert-system modeling, and workflow-guided data acquisition and anonymization. This paper summarizes project objectives, expected outputs, relevance to IEEE COMPSAC 2026 themes, and current progress. In particular, KiMeKo aligns with Applied AI and Smart & Connected Health by combining AI engineering, privacy-conscious data processing, and regulation-aware medical software development.
Comprehensive documentation of internal tissue structures plays an important role in legal medicine. Although ultrasound is widely used in medical diagnostics as a powerful modality for injury detection and documentation, its integration into legal medicine workflows remains limited. Potential reasons include physical strain and infection risks due to manual probe placement at different locations along the body. Robotic ultrasound systems have been proposed to address such challenges. However, current systems are typically restricted to specific target regions and are limited in mobility and autonomy, complicating their clinical integration and fully autonomous, task-level operation. We therefore propose a mobile robotic system for autonomous ultrasound scanning across the body. The system localizes the body, estimates internal target structures, moves to reach the target structure, and performs ultrasound scans. We evaluate the system in three stages. First, we assess the accuracy and robustness of target localization and scan path planning across different body types. Second, we conduct experiments with a body phantom to evaluate the repeatability of covering target structures. Third, we demonstrate the system in a real legal medicine setting. Our experiments demonstrate that the system can approach the body, define the scan path, and acquire reproducible ultrasound images of target structures. These results demonstrate the feasibility of fully autonomous ultrasound imaging across multiple anatomical regions using a mobile robotic system. Beyond legal medicine, the proposed mobile system offers perspectives for autonomous imaging in broader clinical settings such as telemedicine and infectious disease settings, supporting remote diagnostics while reducing infection risks.
Continuous electrocardiogram (ECG) monitoring via wearables offers significant potential for early cardiovascular disease (CVD) detection. However, deploying deep learning models for automated analysis in resource-constrained environments faces reliability challenges due to inevitable Out-of-Distribution (OOD) data. OOD inputs, such as unseen pathologies or noisecorrupted signals, often cause erroneous, high-confidence predictions by standard classifiers, compromising patient safety. Existing OOD detection methods either neglect computational constraints or address noise and unseen classes separately. This paper explores Unsupervised Anomaly Detection (UAD) as an independent, upstream filtering mechanism to improve robustness. We benchmark six UAD approaches, including Deep SVDD, reconstruction-based models, Masked Anomaly Detection, normalizing flows, and diffusion models, optimized via Neural Architecture Search (NAS) under strict resource constraints (at most 512k parameters). Evaluation on PTB-XL and BUT QDB datasets assessed detection of OOD CVD classes and signals unsuitable for analysis due to noise. Results show Deep SVDD consistently achieves the best trade-off between detection and efficiency. In a realistic deployment simulation, integrating the optimized Deep SVDD filter with a diagnostic classifier improved accuracy by up to 21 percentage points over a classifier-only baseline. This study demonstrates that optimized UAD filters can safeguard automated ECG analysis, enabling safer, more reliable continuous cardiovascular monitoring on wearables.
Medical diagnosis from Computed Tomography (CT) often suffers from artifacts caused by high-density materials like metal implants. Metal Artifact Reduction (MAR) remains a challenge which has been addressed using deep learning. Supervised learning methods for MAR traditionally require paired scans with and without metal a requirement rarely met in practice. Unsupervised approaches such as the Artifact Disentanglement Network (ADN) learn from unpaired data, while Variational Autoencoders (VAE) reconstruct artifact-free images via latent representations. Dual-domain methods have been proposed, combining image and sinogram data to exploit the localization of metal artifacts in Radon domain for improved reduction. This study evaluates (1) using VAE as a baseline, (2) performing disentanglement of ADN solely in Radon domain, and (3) integrating Cartesian and Radon features into ADN. Experiments on synthetic and clinical data show ADN outperforms VAE, and our dual-domain approach RadonADN further improves results, achieving an SSIM of 0.92 versus 0.91 without Radon domain information, demonstrating the benefit of incorporating the sinogram representation.
Wearable cardiovascular sensor patches promise continuous, unobtrusive monitoring, but their tight energy, memory, and compute budgets make it unclear whether physiological signals should be analyzed on the device or streamed to the cloud for processing. We study this inference-versus-transmission trade-off for a resource-constrained patch that records synchronized electrocardiogram (ECG) and phonocardiogram (PCG) signals. We propose an end-to-end, multi-modal convolutional neural network (CNN) with early fusion that classifies the two modalities directly on the device, without hand-crafted features. Trained and validated on the PhysioNet/Computing in Cardiology Challenge 2016 dataset, the floating-point model attains an accuracy of 0.975, which is competitive with the best reported results. At the same time, it reduces the parameter count and computational cost by approximately three orders of magnitude. We deploy an 8-bit integer version of the model on a microcontroller with an integrated neural processing unit (NPU) and measure its inference energy. We also benchmark the energy required for Bluetooth Low Energy (BLE) communication on a representative evaluation kit across a range of payload sizes. NPU inference consumes approximately one-seventh of the energy required for CPU inference. For realistic per-second payloads, local inference is also several times more energy efficient than continuous raw-data streaming. These results show that on-device intelligence, rather than constant transmission, is the more energy-efficient basis for always-on wearable cardiovascular monitoring at the edge.
Intraoperative fluorescent cardiac imaging enables quality control following coronary bypass grafting surgery.We can estimate local quantitative indicators, such as cardiac perfusion, by tracking local feature points. However, heart motion and significant fluctuations in image characteristics caused by vessel structural enrichment limit traditional tracking methods. We propose a particle filtering tracker based on cyclicconsistency checks to robustly track particles sampled to follow target landmarks. Our method tracks 117 targets simultaneously at 25.4 fps, allowing real-time estimates during interventions. It achieves a tracking error of (5.00 ± 0.22 px) and outperforms other deep learning trackers (22.3 ± 1.1 px) and conventional trackers (58.1 ± 27.1 px).
Minimally invasive surgery presents challenges such as dynamic tissue motion and a limited field of view. Accurate tissue tracking has the potential to support surgical guidance, improve safety by helping avoid damage to sensitive structures, and enable context-aware robotic assistance during complex procedures. In this work, we propose a novel method for markerless 3D tissue tracking by leveraging 2D Tracking Any Point (TAP) networks. Our method combines two CoTracker models, one for temporal tracking and one for stereo matching, to estimate 3D motion from stereo endoscopic images. We evaluate the system using a clinical laparoscopic setup and a robotic arm simulating tissue motion, with experiments conducted on a synthetic 3D-printed phantom and a chicken tissue phantom. Tracking on the chicken tissue phantom yielded more reliable results, with Euclidean distance errors as low as 1.1 mm at a velocity of 10 mm/s. These findings highlight the potential of TAP-based models for accurate, markerless 3D tracking in challenging surgical scenarios.
Comprehensive legal medicine documentation includes internal and external examination of the corpse. Typically, this documentation is conducted manually during conventional autopsy. Systematic digital documentation would be desirable, especially for external wound examination, which is becoming more relevant for legal medicine analysis. For this purpose, RGB surface scanning has been introduced. While manual full-surface scanning using a handheld camera is time-consuming and operator-dependent, floor or ceiling-mounted robotic systems require specialized rooms. Hence, we consider whether a mobile robotic system can be used for external documentation. We develop a mobile robotic system that enables full-body RGB-D surface scanning. Our work includes a detailed configuration space analysis to identify the environmental parameters that must be considered for a successful surface scan. We validate our findings through an experimental study in the lab and demonstrate the systems application in legal medicine. Our configuration space analysis shows that a good trade-off between coverage and time is reached with three robot base positions, leading to a coverage of 94.96 96.90 ± 3.16 92.45 ± 1.43
OBJECTIVE:Automatic segmentation and detection of vestibular schwannoma (VS) in MRI by deep learning is an upcoming topic. However, deep learning faces generalization challenges due to tumor variability even though measurements and segmentation of VS are essential for growth monitoring and treatment planning. Therefore, we introduce a novel model combining two Convolutional Neural Network (CNN) models for the detection of VS by deep learning aiming to improve performance of automatic segmentation. METHODS:Deep learning techniques have been employed for automatic VS tumor segmentation, including 2D, 2.5D, and 3D UNet-like architectures, which is a specific CNN designed to improve automatic segmentation for medical imaging. Specifically, we introduce a sequential connection where the first UNet's predicted segmentation map is passed to a second complementary network for refinement. Additionally, spatial attention mechanisms are utilized to further guide refinement in the second network. RESULTS:We conducted experiments on both public and private datasets containing contrast-enhanced T1 and high-resolution T2-weighted magnetic resonance imaging (MRI). Across the public dataset, we observed consistent improvements in Dice scores for all variants of 2D, 2.5D, and 3D CNN methods, with a notable enhancement of 8.86% for the 2D UNet variant on T1. In our private dataset, a 3.75% improvement was reported for 2D T1. Moreover, we found that T1 images generally outperformed T2 in VS segmentation. CONCLUSION:We demonstrate that sequential connection of UNets combined with spatial attention mechanisms enhances VS segmentation performance across state-of-the-art 2D, 2.5D, and 3D deep learning methods. LEVEL OF EVIDENCE:3 Laryngoscope, 135:1301-1308, 2025.
The application of supervised models to clinical screening tasks is challenging due to the need for annotated data for each considered pathology. Unsupervised Anomaly Detection (UAD) is an alternative approach that aims to identify any anomaly as an outlier from a healthy training distribution. A prevalent strategy for UAD in brain MRI involves using generative models to learn the reconstruction of healthy brain anatomy for a given input image. As these models should fail to reconstruct unhealthy structures, the reconstruction errors indicate anomalies. However, a significant challenge is to balance the accurate reconstruction of healthy anatomy and the undesired replication of abnormal structures. While diffusion models have shown promising results with detailed and accurate reconstructions, they face challenges in preserving intensity characteristics, resulting in false positives. We propose conditioning the denoising process of diffusion models with additional information derived from a latent representation of the input image. We demonstrate that this conditioning allows for accurate and local adaptation to the general input intensity distribution while avoiding the replication of unhealthy structures. We compare the novel approach to different state-of-the-art methods and for different data sets. Our results show substantial improvements in the segmentation performance, with the Dice score improved by 11.9%, 20.0%, and 44.6%, for the BraTS, ATLAS and MSLUB data sets, respectively, while maintaining competitive performance on the WMH data set. Furthermore, our results indicate effective domain adaptation across different MRI acquisitions and simulated contrasts, an important attribute for general anomaly detection methods. The code for our work is available at https://github.com/FinnBehrendt/Conditioned-Diffusion-Models-UAD.