INTRODUCTION:Technological innovation is rapidly transforming surgical care, yet disparities in access and adoption persist. This study introduces a flexible, population-based model designed to assess and potentially guide the equitable implementation of advanced technologies in visceral (gastrointestinal) surgery, ranging from laparoscopy and basic imaging up to robotic-assisted surgery and artificial intelligence. MATERIAL AND METHODS:We conducted a survey among technologically experienced visceral surgeons to classify current and emerging technologies by perceived complexity and requirement. This classification was used to query a self-developed translation model, which was integrated with demographic, hospital, and geographic data from official German sources. The model quantifies technological maturity and accessibility by calculating travel-to-treat times from the population's residence to hospitals offering specific technologies. We applied the model to the entire German population and all 16 counties, enabling regional and national comparisons. RESULTS:Based on the survey results, we constructed a layered classification for grouping surgical technologies with two dimensions, S and D, for surgical procedures and available devices and increasing levels of complexity (1-3). The model identified substantial regional disparities in access to advanced technology-assisted visceral surgery across Germany. While basic technologies (S1-S2, D1-D2) were broadly accessible nationwide, higher-level technologies such as robotic surgery (S3) were concentrated in urban areas. However, some advanced diagnostics (D3) were also available in remote regions. The model effectively quantified these differences from the population's perspective and demonstrated its potential for objective, population-based assessment of technological translation and accessibility. CONCLUSION:Our model highlights persistent regional disparities in advanced technology adoption and demonstrates potential for guiding equitable implementation strategies. Its population-based, data-driven approach can be extended to other countries, healthcare settings, and technologies, supporting political decision-making and improved patient access.
Diagnosing esophageal motility disorders, including dysphagia, pose significant challenges due to the complexity of high-resolution impedance manometry (HRIM) data and variability in clinical interpretation. Thiswork explores the feasibility of a multimodal machine learning (ML)-based classification approach that combines HRIM recordings with patient-specific information and incorporates a graph-based modeling of esophageal physiology. We analyze HRIM recordings with corresponding patient information from 104 patientswith esophageal motility disorders collected atTUMUniversity Hospital. Patient data include demographic, clinical, and symptom information extracted from structured questionnaires and free-text notes using keyword detection and large language model-based processing. HRIM data are represented as spatiotemporal graphs, where nodes correspond to pressure values along the esophagus and edges encode spatial adjacency and impedance dynamics.Agraph neural network (GNN) is applied to learn physiologically meaningful representations,which are fused with patient embeddings for multi-category, multi-class classification of swallow events. The impact of patient features and graphbased modeling is evaluated by ablation studies and comparison to vision-based classifier baselines. The proposed multimodal approach, incorporating patient-specific information, indicates improvements over models that rely solely on HRIM-derived features across all classification categories. Additionally, the graph-based modeling provides gains compared to vision-based baselines. Our experiments systematically assess the complementary contribution of multiple modalities, as well as demonstrate the feasibility of our proposed graph-based approach. Our initial findings demonstrate that integrating patient-level data with graph-based representations of HRIM signals appears to be a promising direction for more accurate classification of esophageal motility disorders. For further validation, future studies should include larger and more representative datasets to confirm these trends and ensure generalizability.
BACKGROUND:This work aims to develop and evaluate a robotic end effector capable of performing needle decompression for tension pneumothorax, enabling life-saving intervention in the absence of on-site medical personnel. METHODS:A compact modular system was designed featuring a single-actuator control for relative needle-catheter motion, automated sterile pickup, positioning via ultrasound and a quick-switch interface for tool exchange. Experimental validation focused on mechanical performance through critical load, friction, and needle-handling tests using ex vivo porcine thoracic tissue with a simulated pleural pressure model. RESULTS:The end effector, including its needle-actuating gripper, was capable of delivering the estimated 30 N insertion force required for thoracic decompression. The gripper design further demonstrated structural robustness, withstanding peak axial loads exceeding 290 N. In realistic needle-handling experiments, 10 out of 11 valid attempts (90.9%) successfully achieved complete insertion, simulated decompression, and correct needle retraction, with the catheter remaining in place throughout the procedure. CONCLUSION:This paper proposes the first end effector design capable of performing robotic needle decompression for tension pneumothorax, enabling emergency intervention without on-site medical personnel. The proposed design demonstrates reproducible performance and procedural feasibility for robotic tension pneumothorax decompression.
PURPOSE:We explore an approach to develop telemedical Graphical User Interfaces (GUIs) using User-centered Design (UCD) methods. In contrast with prior work that emphasizes system integration, we center on the GUI as the clinician's primary interface. METHODS:In a user-centered, two-stage process, we developed a modular and generalized GUI architecture and adapted it to a Robotic Ultrasound System (RUSS) and a robotic Tension Pneumothorax (tPTX) system. The GUIs were iteratively co-designed and evaluated with physicians using mixed methods (System Usability Scale (SUS), end-user-adapted heuristics, eye-tracking). RESULTS:The results of the study confirm the benefits of UCD methods for the development of GUIs for telemedical robotic systems. Repeated exposure increased efficiency across tasks and scenarios, indicating the usability (training effect) and consistency (cross-scenario learning transfer) of a modular GUI architecture. The mixed-methods design provided complementary insights via triangulation, helped to analyze outliers, and revealed additional design issues not captured by one method alone. CONCLUSION:A user-centered GUI development process with multiple evaluation rounds can improve the usability of telemedical robotic systems for medical professionals. Our findings show that a modular architecture can facilitate the transfer of training effects across implementations and clinical use cases by preserving high-level structures and interaction concepts.
Accurate grasping point prediction is a key challenge for autonomous tissue manipulation in minimally invasive surgery, particularly in complex and variable procedures such as colorectal interventions. Due to their complexity and prolonged duration, colorectal procedures have been underrepresented in current research. At the same time, they pose a particularly interesting learning environment due to repetitive tissue manipulation, making them a promising entry point for autonomous, machine learning-driven support. Therefore, in this work, we introduce attachment anchors, a structured representation that encodes the local geometric and mechanical relationships between tissue and its anatomical attachments in colorectal surgery. This representation reduces uncertainty in grasping point prediction by normalizing surgical scenes into a consistent local reference frame. We demonstrate that attachment anchors can be predicted from laparoscopic images and incorporated into a grasping framework based on machine learning. Experiments on a dataset of 90 colorectal surgeries demonstrate that attachment anchors improve grasping point prediction compared to image-only baselines. There are particularly strong gains in out-of-distribution settings, including unseen procedures and operating surgeons. These results suggest that attachment anchors are an effective intermediate representation for learning-based tissue manipulation in colorectal surgery.
Deep learning methods are commonly used to generate context understanding to support surgeons and medical professionals. By expanding the current focus beyond the operating room (OR) to postoperative workflows, new forms of assistance are possible. In this article, we propose a novel multi-target multi-camera tracking (MTMCT) architecture for postoperative phase recognition, location tracking, and automatic timestamp generation. Three RGB cameras were used to create a multi-camera data set containing 19 reenacted postoperative patient flows. Patients and beds were annotated and used to train the custom MTMCT architecture. It includes bed and patient tracking for each camera and a postoperative patient state module to provide the postoperative phase, current location of the patient, and automatically generated timestamps. The architecture demonstrates robust performance for single- and multi-patient scenarios by embedding medical domain-specific knowledge. In multi-patient scenarios, the state machine representing the postoperative phases has a traversal accuracy of 84.9 ± 6.0% , 91.4 ± 1.5% of timestamps are generated correctly, and the patient tracking IDF1 reaches 92.0 ± 3.6% . Comparative experiments show the effectiveness of using AFLink for matching partial trajectories in postoperative settings. As our approach shows promising results, it lays the foundation for real-time surgeon support, enhancing clinical documentation and ultimately improving patient care.
The explainability of deep learning models remains a significant challenge, particularly in the medical domain where interpretable outputs are critical for clinical trust and transparency. Path attribution methods such as Integrated Gradients rely on a baseline input representing the absence of relevant features ("missingness"). Commonly used baselines, such as all-zero inputs, are often semantically meaningless, especially in medical contexts where missingness can itself be informative. While alternative baseline choices have been explored, existing methods lack a principled approach to dynamically select baselines tailored to each input. In this work, we examine the notion of missingness in the medical setting, analyze its implications for baseline selection, and introduce a counterfactual-guided approach to address the limitations of conventional baselines. We argue that a clinically normal but input-close counterfactual represents a more accurate representation of a meaningful absence of features in medical data. To implement this, we use a Variational Autoencoder to generate counterfactual baselines, though our concept is generative-model-agnostic and can be applied with any suitable counterfactual method. We evaluate the approach on three distinct medical data sets and empirically demonstrate that counterfactual baselines yield more faithful and medically relevant attributions compared to standard baseline choices.
BACKGROUND:High-resolution manometry (HRM) is the gold standard for diagnosing esophageal motility disorders. However, its short-term laboratory setting often fails to capture intermittent abnormalities. Long-term HRM (LTHRM, up to 24h) provides richer insights into swallowing behavior, but the resulting data volume is immense. Manual analysis by medical experts is laborious, time-consuming, and prone to errors, limiting its clinical feasibility. METHODS:We propose a deep learning-based approach for automatic analysis of LTHRM data. Our method detects both swallow events and secondary non-deglutitive motility disorders with high accuracy. Detected swallows are then clustered into distinct classes of similar events, creating a structured overview of motility patterns and their frequency. This reduces the analytical burden by allowing clinicians to focus on a small number of representative swallows rather than manually reviewing thousands of individual events. We evaluate our pipeline on 25 LTHRMs that were meticulously annotated, resulting in a dataset of more than 23,000 expert-labeled events. RESULTS:Our approach is able to detect more than 94% of all relevant events in LTHRM sequences, while the subsequent clustering is able to capture and group all relevant events into distinct swallow groups. To evaluate the overall approach, we conduct a user study with medical experts, demonstrating its effectiveness and positive clinical impact. CONCLUSIONS:Our findings demonstrate that deep learning-based approaches to analyze LTHRM examinations are capable of providing a more reliable and efficient diagnostic process, ultimately making LTHRM assessments more feasible in clinical care.
Teamwork is fundamental to medical practice and relies on seamless collaboration among professionals with different tasks. Integrating robotic systems into this environment demands smooth interactions. Human action recognition, which infers a person’s state without explicit input, can support this. We focus on handovers between medical staff, using the actions as implicit cues for robotic assistance to replace the giving party in such scenarios. Skeletal information processed with differing machine learning algorithms makes it possible to derive actions out of sequential image data. Transferred to the medical context, we aim to infer actions defined for each situation in two datasets, a surgery in the operating room and a care intervention in the patient ward, depicting a handover between staff. We aim to abstract movement patterns across individuals through skeletal representation, leveraging the spatiotemporal information of medical handovers to enable future robotic systems to interact based on implicit cues. We report an F1 score of 0.736 ± 0.045 for the OR dataset with ST-GCN and an F1 score of 0.941 ± 0.009 for the Ward dataset with the SkateFormer human action recognition. The defined actions showed distinction in the confusion matrix with limitations on actions with a rapid transition like approach and reach as well as the handover actions in the OR. The handover phases in two medical contexts, a minimally invasive surgery and a wound dressing on the patient station, are recognized with the proposed framework. This lays a first step for the integration of robotic assistance in the handover of medical material or instruments.
Video-based intra-abdominal instrument tracking for laparoscopic surgeries is a common research area. However, the tracking can only be done with instruments that are actually visible in the laparoscopic image. By using extra-abdominal cameras to detect trocars and classify their occupancy state, additional information about the instrument location, whether an instrument is still in the abdomen or not, can be obtained. This can enhance laparoscopic workflow understanding and enrich already existing intra-abdominal solutions. A data set of four laparoscopic surgeries recorded with two time-synchronized extra-abdominal 2D cameras was generated. The preprocessed and annotated data were used to train a deep learning-based network architecture consisting of a trocar detection, a centroid tracker and a temporal model to provide the occupancy state of all trocars during the surgery. The trocar detection model achieves an F1 score of 95.06± 0.88% . The prediction of the occupancy state yields an F1 score of 89.29± 5.29% , providing a first step towards enhanced surgical workflow understanding. The current method shows promising results for the extra-abdominal tracking of trocars and their occupancy state. Future advancements include the enlargement of the data set and incorporation of intra-abdominal imaging to facilitate accurate assignment of instruments to trocars.
Surgical documentation has many implications. However, its primary function is to transfer information about surgical procedures to other medical professionals. Thereby, written reports describing procedures in detail are the current standard, impeding comprehensive understanding of patient-individual life-spanning surgical course, especially if surgeries are performed at a timely distance and in diverse facilities. Therefore, we developed a novel model-based approach for documentation of visceral surgeries, denoted as 'Surgical Documentation Markup-Modeling' (SDM-M). For scientific evaluation, we developed a web-based prototype software allowing for creating hierarchical anatomical models that can be modified by individual surgery-related markup information. Thus, a patient's cumulated 'surgical load' can be displayed on a timeline deploying interactive anatomical 3D models. To evaluate the possible impact on daily clinical routine, we performed an evaluation study with 24 surgeons and advanced medical students, elaborating on simulated complex surgical cases, once with classic written reports and once with our prototypical SDM-M software. Leveraging SDM-M in an experimental environment reduced the time needed for elaborating simulated complex surgical cases from 354 ± 85 s with the classic approach to 277 ± 128 s. (p = 0.00109) The perceived task load measured by the Raw NASA-TLX was reduced significantly (p = 0.00003) with decreased mental (p = 0.00004) and physical (p = 0.01403) demand. Also, time demand (p = 0.00041), performance (p = 0.00161), effort (p = 0.00024), and frustration (p = 0.00031) were improved significantly. Model-based approaches for life-spanning surgical documentation could improve the daily clinical elaboration and understanding of complex cases in visceral surgery. Besides reduced workload and time sparing, even a more structured assessment of individual surgical cases could foster improved planning of further surgeries, information transfer, and even scientific evaluation, considering the cumulative 'surgical load.' Life-spanning model-based documentation of visceral surgical cases could significantly improve surgery and workload.
Purpose Even though workflow analysis in the operating room has come a long way, current systems are still limited to research. In the quest for a robust, universal setup, hardly any attention has been given to the dimension of audio despite its numerous advantages, such as low costs, location, and sight independence, or little required processing power.Methodology We present an approach for audio-based event detection that solely relies on two microphones capturing the sound in the operating room. Therefore, a new data set was created with over 63 h of audio recorded and annotated at the University Hospital rechts der Isar. Sound files were labeled, preprocessed, augmented, and subsequently converted to log-mel-spectrograms that served as a visual input for an event classification using pretrained convolutional neural networks.Results Comparing multiple architectures, we were able to show that even lightweight models, such as MobileNet, can already provide promising results. Data augmentation additionally improved the classification of 11 defined classes, including inter alia different types of coagulation, operating table movements as well as an idle class. With the newly created audio data set, an overall accuracy of 90%, a precision of 91% and a F1-score of 91% were achieved, demonstrating the feasibility of an audio-based event recognition in the operating room.Conclusion With this first proof of concept, we demonstrated that audio events can serve as a meaningful source of information that goes beyond spoken language and can easily be integrated into future workflow recognition pipelines using computational inexpensive architectures.
High-resolution manometry (HRM) is the gold standard in diagnosing esophageal motility disorders. As HRM is typically conducted under short-term laboratory settings, intermittently occurring disorders are likely to be missed. Therefore, long-term (up to 24h) HRM (LTHRM) is used to gain detailed insights into the swallowing behavior. However, analyzing the extensive data from LTHRM is challenging and time consuming as medical experts have to analyze the data manually, which is slow and prone to errors. To address this challenge, we propose a Deep Learning based swallowing detection method to accurately identify swallowing events and secondary non-deglutitive-induced esophageal motility disorders in LTHRM data. We then proceed with clustering the identified swallows into distinct classes, which are analyzed by highly experienced clinicians to validate the different swallowing patterns. We evaluate our computational pipeline on a total of 25 LTHRMs, which were meticulously annotated by medical experts. By detecting more than 94 relevant swallow events and providing all relevant clusters for a more reliable diagnostic process among experienced clinicians, we are able to demonstrate the effectiveness as well as positive clinical impact of our approach to make LTHRM feasible in clinical care.
BACKGROUND:Machine learning and robotics technologies are increasingly being used in the healthcare domain to improve the quality and efficiency of surgeries and to address challenges such as staff shortages. Robotic scrub nurses in particular offer great potential to address staff shortages by assuming nursing tasks such as the handover of surgical instruments. METHODS:We introduce a robotic scrub nurse system designed to enhance the quality of surgeries and efficiency of surgical workflows by predicting and delivering the required surgical instruments based on real-time laparoscopic video analysis. We propose a three-stage deep learning architecture consisting of a single frame-, temporal multi frame-, and informed model to anticipate surgical instruments. The anticipation model was trained on a total of 62 laparoscopic cholecystectomies. RESULTS:Here, we show that our prediction system can accurately anticipate 71.54% of the surgical instruments required during laparoscopic cholecystectomies in advance, facilitating a smoother surgical workflow and reducing the need for verbal communication. As the instruments in the left working trocar are changed less frequently and according to a standardized procedure, the prediction system works particularly well for this trocar. CONCLUSIONS:The robotic scrub nurse thus acts as a mind reader and helps to mitigate staff shortages by taking over a great share of the workload during surgeries while additionally enabling an enhanced process standardization.
PurposeDecision support systems and context-aware assistance in the operating room have emerged as the key clinical applications supporting surgeons in their daily work and are generally based on single modalities. The model- and knowledge-based integration of multimodal data as a basis for decision support systems that can dynamically adapt to the surgical workflow has not yet been established. Therefore, we propose a knowledge-enhanced method for fusing multimodal data for anticipation tasks.MethodsWe developed a holistic, multimodal graph-based approach combining imaging and non-imaging information in a knowledge graph representing the intraoperative scene of a surgery. Node and edge features of the knowledge graph are extracted from suitable data sources in the operating room using machine learning. A spatiotemporal graph neural network architecture subsequently allows for interpretation of relational and temporal patterns within the knowledge graph. We apply our approach to the downstream task of instrument anticipation while presenting a suitable modeling and evaluation strategy for this task.ResultsOur approach achieves an F1 score of 66.86% in terms of instrument anticipation, allowing for a seamless surgical workflow and adding a valuable impact for surgical decision support systems. A resting recall of 63.33% indicates the non-prematurity of the anticipations.ConclusionThis work shows how multimodal data can be combined with the topological properties of an operating room in a graph-based approach. Our multimodal graph architecture serves as a basis for context-sensitive decision support systems in laparoscopic surgery considering a comprehensive intraoperative operating scene.
A proposed solution to address increasing scarcity of operating room (OR) nurses is the deployment of robotic scrub nurses (RSNs). One of the primary responsibilities of these robots is the handover of surgical instruments to the surgeon. Despite this tasks’ apparent simplicity, it requires great precision and timing, while seamless execution is essential for the surgical workflow. This paper outlines the unique challenges of the instrument handover process between a RSN and a surgeon taking place during laparoscopic surgery. After customizing existing solutions from different domains to meet the specific needs of laparoscopic surgery, we introduced these solutions to medical professionals in a simulated OR scenario, and performed a utility analysis to identify the most suitable solutions.
Due to the ongoing shortage of qualified surgical assistants and the drive for automation, the deployment of robotic scrub nurses (RSN) is being investigated. As such robotic systems are expected to fulfill all indirect and direct forms of surgical assistance currently provided by human operating room (OR) assistants, they must also be capable of performing intraoperative cleaning of laparoscopic instruments, which are prone to contamination when using electrosurgical techniques during minimally invasive procedures. We present a cleaning station for robotic scrub nurse systems which provides intraoperative cleaning of laparoscopic instruments during minimally invasive procedures. The system uses deep learning to decide autonomously on the need of intraoperative cleaning to preserve instrument functions. We performed configuration and durability tests to determine an optimal set of system parameters and to verify the system performance in an application context. The results of the configuration tests indicate that the use of hard brushes in combination with a sodium chloride cleaning solution and a sequence of 3 s cleaning intervals provides the best cleaning performance with a minimal total cleaning time. The results of the durability tests show that the cleaning function is in principle guaranteed for the duration of a surgical intervention. Our evaluation tests have shown that our deep learning assisted cleaning station for robotic scrub nurse systems is capable of performing autonomous intraoperative cleaning of laparoscopic instruments, providing a further step towards the integration of robotic scrub nurse systems into the OR.