Purpose Vision-based navigation systems rely on the registered camera poses in the CT space to guide surgeons. However, while it is possible to provide an approximate initialization, this registration becomes outdated as the endoscopic camera leaves and reenters the anatomy. Endoscopic camera relocalization is the process of determining the position of an endoscope relative to an anatomical reference after reinsertion. However, accurately reidentifying the global surgical scene and estimating camera pose have proven challenging due to the varying appearance of endoscopic sequences. Methods We present a training-free approach to accurately reidentify the region of interest (ROI) and estimate the camera position of a query image after reinsertion. This method utilizes previously observed images with known poses and a CT scan. By combining advanced foundation models with classical techniques, we globally reidentify a prior image of the ROI, which is then used for image-based feature matching and pose recovery via the Perspective-n-Point algorithm. Results We conducted experiments on eight sequences from three cadaver studies. Our results show that our method accurately reidentifies when the endoscope reaches the ROI and identifies suitable image pairs for PnP-based pose estimation. It achieves an average translation error of 1.74 mm and a rotational error of 0.09 radians, making it suitable for reinitialization in image-based navigation without human intervention. Conclusion Our work presents a training-free approach for detecting when the endoscope reenters the ROI and estimating the camera's pose after reinsertions. The approach demonstrates promising results contributing toward enabling pose reinitialization for vision-based surgical applications.
Minimally invasive procedures performed within confined anatomical spaces depend on continuous endoscopic visualization. Current robotic endoscope systems can stabilize or reposition an endoscope, but they do not possess relevant context to provide effective visualization assistance. We present EndoNav, an anatomy-grounded natural-language framework that translates high-level surgeon commands into autonomous endoscopic visualization behaviors within patient-specific sinonasal anatomy. Spoken surgeon commands are transcribed and interpreted by an endoscopic viewpoint agent conditioned on a patient-specific anatomical scene representation. Rather than generating robot motion directly, the viewpoint agent generates structured visualization objectives that are converted into target viewpoints and inspection trajectories, which are then executed through geometry-constrained endoscope motion planning and joint-space control. We evaluate EndoNav using a structured three-pass sinus examination across three CT-derived anatomical models. For one cadaveric specimen, autonomous visualization is compared with sinus examinations performed by two resident surgeons. EndoNav achieved mean visualization IoUs of 87.04
Large language model-based (LLM) agents are emerging as a powerful enabler of robust embodied intelligence due to their capability of planning complex action sequences. Sound planning ability is necessary for robust automation in many task domains, but especially in surgical automation. These agents rely on a highly detailed natural language representation of the scene. Thus, to leverage the emergent capabilities of LLM agents for surgical task planning, developing similarly powerful and robust perception algorithms is necessary to derive a detailed scene representation of the environment from visual input. Previous research has focused primarily on enabling LLM-based task planning while adopting simple yet severely limited perception solutions to meet the needs for bench-top experiments, but lacks the critical flexibility to scale to less constrained settings. In this work, we propose an alternate perception approach – a digital twin (DT)-based machine perception approach that capitalizes on the convincing performance and out-of-the-box generalization of recent vision foundation models. Integrating our DT representation and LLM agent for planning with the dVRK platform, we develop an embodied intelligence system and evaluate its robustness in performing peg transfer and gauze retrieval tasks. Our approach shows strong task performance and generalizability to varied environmental settings. Despite a convincing performance, this work is merely a first step towards the integration of DT representations. Future studies are necessary for the realization of a comprehensive DT framework to improve the interpretability and generalizability of embodied intelligence in surgery.
Reliable volumetric representation of the nasal cavity is crucial for enabling quantitative assessment in Functional Endoscopic Sinus Surgery (FESS), yet for most patients the evaluation of their anatomy remains largely qualitative and subjective. While computed tomography (CT) scans can provide 3D anatomical information, their routine use is impractical due to radiation exposure concerns and cost constraints, underscoring the need for a non-invasive alternative. Computer vision methods offer a promising solution for reconstructing sinus anatomy from routine endoscopic video. Current methods rely on Structure-from-Motion (SfM), however, this relies on point correspondences that struggle with photometric inconsistencies inherent to endoscopic imaging, reducing robustness and generalizability. Several sinus reconstruction approaches attempt to mitigate this through learning-based and patient-specific approaches, but suffer from error propagation, leading to inaccurate 3D representations. Optimization-based approaches further introduce excessive training times, limiting their practicality. In this work, we revisit simpler techniques for sinus reconstruction and augment them with track-any-point foundation models to develop a training-free, vision-based 3D reconstruction method. Our approach leverages SfM poses and local point-tracks to generate depth information, recovering a globally consistent structure without fine-tuning requirements. We evaluate our method on six pre-operative endoscopic sequences with respect to the ground-truth CT scan. Our results show that this method improves global geometric accuracy by reducing both point-to-point and pose errors from prior work. Our vision-based approach improves spatial consistency and accuracy in sinus 3D reconstruction, enabling non-invasive postoperative monitoring and seamless clinical integration, offering physicians data-driven insights for improved surgical decision-making.
Subretinal injection is a highly delicate procedure that demands micron-level precision to avoid irreversible retinal damage. Current robotic systems achieve accurate positioning but remain limited by retinal motion and the lack of tip-force feedback. We present the first adaptive tip-force compensation framework for robotic subretinal injection, fusing intraoperative optical coherence tomography (iOCT) vision with fiber Bragg grating (FBG) force sensing. Our architecture integrates a finite-state machine (FSM) for surgical phase coordination, a Long Short-Term Memory (LSTM) enhanced residual Kalman filter for real-time motion prediction, and an adaptive compliance estimator for safe force regulation. Compared to previous vision-only and force-only method, ex vivo experiments on porcine eyes demonstrate robust improvements: the root-mean-square tracking error reduced by 40% (to 18.5μm), the maximum absolute error lowered by 2.5 times, and 96.7% of tip forces maintained within ± 0.7mN. Control delays were minimized to 0.25s, enabling low-latency corrections beyond freehand capabilities. Our system enhances precision and safety in fragile retinal tissues, advancing the potential for reliable robot-assisted surgeries for retinal diseases.
PURPOSE:Preoperative CT provides a detailed anatomical basis for navigation in sinus surgery, guiding surgeons through intricate nasal structures while providing 3D awareness of critical anatomy. However, this static representation lacks the capacity to adapt to intraoperative tissue ablation. Efforts toward endoscope-based navigation reconstruct anatomy from visual cues alone, often without incorporating the CT prior that anchors surgical context. We propose a method to ground both CT and endoscope in a unified representation that bridges pre- and intraoperative domains, enabling dynamic updates of the surgical scene. METHODS:We initialize the CT as a volumetric signed distance field (SDF), reformulating the update process with non-projective SDFs and accumulating camera-space observations with dense depth estimated from endoscopic video. These volumes are reconciled to produce intraoperative SDFs for updated surface mesh extraction. RESULTS:We validated our method on three multi-step cadaveric sinus datasets with paired preoperative CT and intraoperative endoscopic video of a simulated sinus surgery. We extract intraoperative meshes at each surgical step and compare to ground-truth intraoperative CT. Our vision-guided update method demonstrates improved surface alignment of the intraoperative representations throughout surgical progression, achieving submillimeter geometric accuracy. CONCLUSIONS:We show that integrating CT priors with intraoperative vision enables anatomically consistent updates to patient-specific models. This method provides a foundation for CT-consistent endoscopic reconstruction, and future work aimed at integrating camera localization toward a fully realized vision-guided surgical navigation system.
In the real world, robots frequently make errors, yet little is known about people's social responses to errors outside of lab settings. Prior work has shown that social signals are reliable and useful for error management in constrained interactions, but it is unclear if this holds in the real world - especially with a non-social robot in repeated and group interactions with successive or propagated errors. To explore this, we built a coffee robot and conducted a public field deployment (N = 49). We found that participants consistently expressed varied social signals in response to errors and other stimuli, particularly during group interactions. Our findings suggest that social signals in the wild are rich (with participants volunteering information about the interaction), but "noisy." We discuss lessons, benefits, and challenges for using social signals in real-world HRI.
In endoscopic surgery, surgeons continuously locate the endoscopic view relative to the anatomy by interpreting the evolving visual appearance of the intraoperative scene in the context of their prior knowledge. Vision-based navigation systems seek to replicate this capability by recovering camera pose directly from endoscopic video, but most approaches do not embody the same principles of reasoning about new frames that makes surgeons successful. Instead, they remain grounded in feature matching and geometric optimization over keyframes, an approach that has been shown to degrade under the challenging conditions of endoscopic imaging like low texture and rapid illumination changes. Here, we pursue an alternative approach and investigate a policy-based formulation of endoscopic camera pose recovery that seeks to imitate experts in estimating trajectories conditioned on the previous camera state. Our approach directly predicts short-horizon relative motions without maintaining an explicit geometric representation at inference time. It thus addresses, by design, some of the notorious challenges of geometry-based approaches, such as brittle correspondence matching, instability in texture-sparse regions, and limited pose coverage due to reconstruction failure. We evaluate the proposed formulation on cadaveric sinus endoscopy. Under oracle state conditioning, we compare short-horizon motion prediction quality to geometric baselines achieving lowest mean translation error and competitive rotational accuracy. We analyze robustness by grouping prediction windows according to texture richness and illumination change indicating reduced sensitivity to low-texture conditions. These findings suggest that a learned motion policy offers a viable alternative formulation for endoscopic camera pose recovery.
The ability to synthesize realistic X-ray images has catalyzed the development of AI models for X-ray image-guided procedures, which otherwise suffer from a lack of available annotated data. Prior work has demonstrated the effectiveness of mechanistic simulation of digitally reconstructed radiographs (DRRs) as a training data source for a myriad of tasks, including segmentation and anatomical landmark detection, with comparable or superior performance to real data training. However, mechanistic DRR synthesis still relies on the availability of annotated high-resolution anatomical models. Deriving these from CT images of real patients or specimens imposes an undesirable bottleneck on data quantity and variability. In this work, we explore two methods for synthesizing training data: (1) a 3D conditional latent diffusion model that generates CT volumes to use as inputs for mechanistic DRR generation without real, 3D anatomical models, and (2) a view-conditioned 2D diffusion model that produces synthetic X-rays. In controlled experiments, we demonstrate that synthetic 2D diffusion-based X-rays can be used to train an anatomical landmark detection model that generalized to real X-ray images with performance rivaling that of a model trained on real X-ray images. Thus, we provide preliminary evidence that synthetic, 2D diffusion-based training data can substitute for real X-ray data, identifying a promising avenue towards generating large, diverse datasets for training robust AI models in interventional X-ray imaging.
Artificial intelligence, imaging, and large language models have the potential to transform surgical practice, training, and automation. Understanding and modeling of basic surgical actions (BSA), the fundamental unit of operation in any surgery, is important to drive the evolution of this field. In this paper, we present a BSA dataset comprising 10 basic actions across 6 surgical specialties with over 11,000 video clips, which is the largest to date. Based on the BSA dataset, we developed a new foundation model that conducts general-purpose recognition of basic actions. Our approach demonstrates robust cross-specialist performance in experiments validated on datasets from different procedural types and various body parts. Furthermore, we demonstrate downstream applications enabled by the BAS foundation model through surgical skill assessment in prostatectomy using domain-specific knowledge, and action planning in cholecystectomy and nephrectomy using large vision-language models. Multinational surgeons' evaluation of the language model's output of the action planning explainable texts demonstrated clinical relevance. These findings indicate that basic surgical actions can be robustly recognized across scenarios, and an accurate BSA understanding model can essentially facilitate complex applications and speed up the realization of surgical superintelligence.
Surgical data science (SDS) is rapidly advancing, yet clinical adoption of artificial intelligence (AI) in surgery remains limited, with inadequate validation emerging as an important contributing factor. In fact, existing validation practices often neglect the temporal and hierarchical structure of intraoperative videos, producing misleading, unstable, or clinically irrelevant results. In a pioneering, consensus-driven effort, we introduce a comprehensive catalog of validation pitfalls in AI-based surgical video analysis that was derived from a multi-stage Delphi process with 92 international experts. The collected pitfalls span three categories: (1) data (e.g., incomplete annotation, spurious correlations), (2) metric selection and configuration (e.g., neglect of temporal stability, mismatch with clinical needs), and (3) aggregation and reporting (e.g., clinically uninformative aggregation, failure to account for frame dependencies in hierarchical data structures). A systematic review of surgical AI papers reveals that these pitfalls are widespread in current practice, with the majority of studies failing to account for temporal dynamics or hierarchical data structure, or relying on clinically uninformative metrics. Experiments on real surgical video datasets provide empirical evidence that ignoring temporal and hierarchical data structures can substantially understate uncertainty, obscure critical failure modes, and even alter algorithm rankings. To address these shortcomings, we provide a catalogue of best practices compiled in a multi-stage Delphi process. Together, this work provides an evidence-based framework to inform more rigorous validation of surgical video analysis algorithms and to guide future efforts in benchmarking, reporting, regulatory review, and clinical translation.
Precise osteotomies are vital in maxillofacial procedures such as the bilateral sagittal split osteotomy (BSSO) where surgical accuracy and precision directly impacts patient outcomes. Conventional freehand drilling can lead to unfavorable splits, negatively impacting surgical outcome. This paper presents the development work of a cooperatively controlled robot system designed to enhance the efficacy of osteotomies during BSSO. The system features two assistive modes for the execution of a patient-specific surgical plan: (1) a Haptic guidance mode that helps the surgeon align the surgical drill with the planned cutting plane to improve surgical accuracy of the cut and (2) an Active constraint mode that restricts deviations from the cutting plane to enhance surgical precision during drilling. We validated the system through feasibility experiments involving 36 mandible phantoms and a cadaveric specimen, with a surgeon, a surgical resident, and a medical student performing osteotomies freehand and with robotic assistance. Additionally, NASA TLX surveys were conducted to assess the perceived ease of use of the robotic system. Compared to freehand methods, the robotic system improved the efficacy of the cut from 2.16± 0.98 to 0.71± 0.53 mm for the med student, 1.74± 0.95 to 0.53± 0.35 mm for the resident, and 1.64± 0.85 to 0.63± 0.24 mm for the surgeon while reducing the task load. Our experimental results demonstrate that the proposed robotic system can enhance the precision of surgical drilling in the BSSO compared to a freehand approach. These findings indicate the potential of robotic systems to reduce errors and enhance patient outcomes in maxillofacial surgery.
Significance:Despite advances in perinatal medicine over decades, perinatal hypoxic-ischemic encephalopathy (HIE) remains a significant cause of fetal cerebral palsy and can lead to other severe medical sequelae or death. Therefore, it is highly desirable to effectively detect brain hypoxia during labor and postnatally for HIE management. Aim:We recently validated the feasibility of transcranial photoacoustic (PA) imaging for oxyhemoglobin saturation measurement at the superior sagittal sinus ( O 2 Sat ss ) in the neonatal piglet brain, at which overall oxygen supply status can be reflected as a primary collective vein. We aim to automate the PA-based workflow of at-risk subject detection and enable fully autonomous and continuous perinatal monitoring. Approach:We proposed a two-step algorithm that focuses on the most informative region of the brain for oxygenation status, the superior sagittal sinus (SSS). First, a convolutional neural network (U-Net) is trained to detect the location of SSS in the coronal cross-section PA images. Then, an optimized region of interest patch around the predicted SSS location is cropped from the spectral unmixed image and averaged as the O 2 Sat ss measurement. A confidence score can be computed for the measurement via Monte Carlo dropout (MCD), which infers the prediction uncertainty for better clinical decision-making. Results:The algorithm was evaluated on an in vivo piglet brain imaging dataset containing 84 independent experimental settings from 10 piglet subjects. A 10-fold leave-one-subject-out cross-validation experiment reports 85.2% sensitivity and 93.3% specificity for healthy/hypoxia classification with an R -squared value of 0.708 and a confidence score of 94.06% based on MCD computation, well agreed with our ground-truth given by blood gas measurements. Conclusions:The proposed automatic O 2 Sat ss monitoring solution demonstrated a hypoxia detection capability comparable to the human expert manual annotation on the same task. We concluded with high feasibility for a noninvasive PA-based continuous monitoring of the perinatal brain.
Subretinal injection is a critical procedure for delivering therapeutic agents to treat retinal diseases such as inherited retinal diseases (IRD) and age-related macular degeneration (AMD). However, retinal motion caused by physiological factors such as respiration and heartbeat significantly impacts precise needle positioning, increasing the risk of retinal pigment epithelium (RPE) damage. This paper presents a fully autonomous robotic subretinal injection system that integrates intraoperative optical coherence tomography (iOCT) imaging and deep learning-based motion prediction to synchronize needle and retinal motion. A Long Short-Term Memory (LSTM) neural network is used to predict internal limiting membrane (ILM) motion, outperforming a Fast Fourier Transform (FFT)-based baseline model. Additionally, a real-time registration framework aligns the needle tip position with the robot's coordinate frame. Then, a dynamic proportional speed control strategy ensures smooth and adaptive needle insertion. Experimental validation in both simulation and ex vivo open-sky porcine eyes demonstrates precise motion synchronization and successful subretinal injections. The experiments achieve a mean tracking error below 16.4 μm in pre-insertion phases. These results show the potential of AI-driven robotic assistance to improve the safety and accuracy of retinal microsurgery.
Cancer resection surgery is unsuccessful if tumor tissue is left behind in the surgical cavity. Identifying the residual cancer requires additional imaging or postoperative histological analysis. Photoacoustic imaging can be used to image both the surface and depths of the resection cavity; however, its performance hinges on consistent probe placement and stable acoustic and optical coupling. As intra-cavity deployment of photoacoustic imaging is largely uncharted, several potential embodiments warrant rigorous investigation. We address this need with an open-source robotic testbed for intraoperative tumor-bed inspection using photoacoustic imaging. The platform integrates the da Vinci Research Kit, depth imaging, and electromagnetic tracking to automate cavity scanning and maintain repeatable probe trajectories. Using tissue-mimicking phantoms, we (i) demonstrate a novel imaging embodiment for photoacoustic tumor-bed inspection and (ii) show how this testbed can be used to investigate and optimize tumor bed inspection strategies and configurations. This study establishes the feasibility of detecting and mapping residual cancer within a simulated surgical cavity. The primary contribution is the testbed itself, designed for integration with existing surgical navigation workflows and rapid prototyping. This testbed serves as an essential foundation for systematic evaluation of photoacoustic, robot-assisted strategies for improving intraoperative margin assessment.
Arthroscopy is a minimally invasive surgical procedure used to diagnose and treat joint problems. The clinical workflow of arthroscopy typically involves inserting an arthroscope into the joint through a small incision, during which surgeons navigate and operate largely by relying on their visual assessment through the arthroscope. However, the arthroscope's restricted field of view and lack of depth perception pose challenges in navigating complex articular structures and achieving surgical precision during procedures. Aiming at enhancing intraoperative awareness, we present a robust pipeline that incorporates simultaneous localization and mapping, depth estimation, and 3D Gaussian splatting to realistically reconstruct intra-articular structures solely based on monocular arthroscope video. Extending 3D reconstruction to Augmented Reality (AR) applications, our solution offers AR assistance for articular notch measurement and annotation anchoring in a human-in-the-loop manner. Compared to traditional Structure-from-Motion and Neural Radiance Field-based methods, our pipeline achieves dense 3D reconstruction and competitive rendering fidelity with explicit 3D representation in 7 minutes on average. When evaluated on four phantom datasets, our method achieves RMSE = 2.21mm reconstruction error, PSNR = 32.86 and SSIM = 0.89 on average. Because our pipeline enables AR reconstruction and guidance directly from monocular arthroscopy without any additional data and/or hardware, our solution may hold the potential for enhancing intraoperative awareness and facilitating surgical precision in arthroscopy. Our AR measurement tool achieves accuracy within 1.59 +/- 1.81mm and the AR annotation tool achieves a mIoU of 0.721.
Language promptable X-ray image segmentation would enable greater flexibility for human-in-the-loop workflows in diagnostic and interventional precision medicine. Prior efforts have contributed task-specific models capable of solving problems within a narrow scope, but expanding to broader use requires additional data, annotations, and training time. Recently, language-aligned foundation models (LFMs) – machine learning models trained on large amounts of highly variable image and text data thus enabling broad applicability – have emerged as promising tools for automated image analysis. Existing foundation models for medical image analysis focus on scenarios and modalities where large, richly annotated datasets are available. However, the X-ray imaging modality features highly variable image appearance and applications, from diagnostic chest X-rays to interventional fluoroscopy, with varying availability of data. To pave the way toward an LFM for comprehensive and language-aligned analysis of arbitrary medical X-ray images, we introduce FluoroSAM, a language-promptable variant of the Segment-Anything Model, trained from scratch on 3M synthetic X-ray images from a wide variety of human anatomies, imaging geometries, and viewing angles. These include pseudo-ground truth masks for 128 organ types and 464 tools with associated text descriptions. FluoroSAM is capable of segmenting myriad anatomical structures and tools based on natural language prompts, thanks to the novel incorporation of vector quantization (VQ) of text embeddings in the training process. We demonstrate FluoroSAM’s performance quantitatively on real X-ray images and showcase on several applications how FluoroSAM is a key enabler for rich human-machine interaction in the X-ray image acquisition and analysis context. Information on data, weights, and code is available at https://github.com/arcadelab/fluorosam .
Surgical robots capable of autonomously performing various tasks could enhance efficiency and augment human productivity in addressing clinical needs. Although current solutions have automated specific actions within defined contexts, they are challenging to generalize across diverse environments in general surgery. Embodied intelligence enables general-purpose robot learning with applications for daily tasks, yet its application in the medical domain remains limited. We introduced an open-source surgical embodied intelligence simulator for an interactive environment to develop reinforcement learning methods for minimally invasive surgical robots. Using such embodied artificial intelligence, this study further addresses surgical task automation, enabling zero-shot transfer of simulation-trained policies to real-world scenarios. The proposed method encompasses visual parsing, a perceptual regressor, policy learning, and a visual servoing controller, forming a paradigm that combines the advantages of data-driven policy and classic controller. The visual parsing uses stereo depth estimation and image segmentation with a visual foundation model to handle complex scenes. Experiments demonstrated autonomy in seven game-based skill training tasks on the da Vinci Research Kit, with a proof-of-concept study on haptic-assisted skill training as a practical application. Moreover, we conducted automation of five surgical assistive tasks with the Sentire surgical system on ex vivo animal tissues with various scenes, object sizes, instrument types, and illuminations. The learned policies were also validated in a live-animal trial for three tasks in dynamic in vivo surgical environments. We hope this open-source infrastructure, coupled with a general-purpose learning paradigm, will inspire and facilitate future research on embodied intelligence toward autonomous surgical robots.
INTRODUCTION: The ablative nature of surgery means that pre-operative imaging studies lose correspondence as a case progresses, which can be problematic when accurate intraoperative navigation is required. Accurate 3D surface reconstruction from endoscopic video is a potential strategy for real-time intraoperative imaging updates without additional equipment. We have previously used traditional computational models to generate skull base reconstructions. However, they are time-consuming and require technical skills to process the video. Recent foundational AI models, like DUST3R, are an opportunity for timely, generalizable reconstructions of surgical anatomy. METHODS: We compared our previously described three-step reconstruction process with DUST3R to generate a water-tight 3D mesh of cadaveric skull base anatomy visualized using a Karl Storz Image 1 Hub HD Video Camera fitted with a 0° rigid endoscope. For DUST3R, we selected four video frames and did not perform any training or fine-tuning. RESULTS: Without endoscope calibration and using only the four input frames, DUST3R created 3D surface reconstructions in less than two minutes. This is compared to the three-step reconstruction process, which requires 8 to 12 hours to reconstruct. CONCLUSIONS: Our findings show that DUST3R, a publicly available foundational AI model, can rapidly generate 3D anatomical reconstructions from a limited set of video frames. Models like DUST3R illustrate the rapidly evolving potential of computer vision. With fine-tuning, they may represent a path toward foundational AI models that generalize across surgical procedures.