Introduction Acquisition of technical skills in surgery requires both precise motor coordination and efficient cognitive control. While simulation platforms and benchtop dexterity tests have demonstrated utility in differentiating performance levels, their predictive validity remains variable. Concurrently, physiologic metrics of cognitive load have emerged as promising tools for quantifying mental effort during surgical tasks. Pupillometry, in particular, offers high temporal resolution, noninvasiveness, and task-synchronized measurement capabilities. However, few studies have integrated objective dexterity metrics with real-time physiologic workload assessment in a unified, scalable framework. To address this gap, we developed a benchtop testing protocol combining standardized fine-motor tasks with pupillometry to evaluate technical performance and cognitive load at different levels of surgical training. Methods In a pilot cross-sectional observational study, 79 participants, including 49 medical students, 17 nonmedical controls, and 13 experts, completed 2 standardized dexterity tasks: the O’Connor Finger Dexterity and O’Connor Tweezer Dexterity Tests, followed by a knot-tying task. Performance metrics included completion time for each task. Cognitive load was indexed by baseline-corrected pupil diameter (Δ-BCPD), continuously captured and rescaled to normalized task time. Cognitive-load dynamics were compared across groups and performance tiers. Results Finger dexterity scores differed significantly by role (p = 0.048) and correlated with knot-tying time (R² = 0.472), supporting their construct-specific validity. Tweezer-task scores showed no significant relationship with procedural performance (R² = 0.004). Pupillometry showed directionally consistent group-level gradients in Δ-BCPD, with experts exhibiting lower mean task-evoked dilation. High-performing participants demonstrated reduced cognitive load, particularly during the finger task, indicative of greater automation. Trends in Δ-BCPD were directionally consistent across groups but did not reach statistical significance. Conclusions Brief dexterity tasks paired with time-normalized pupillometry may provide complementary information about motor performance and cognitive workload. This exploratory, rater-independent framework showed promise for distinguishing performance differences and characterizing task demands, supporting further validation in larger cohorts.
Visual reasoning, the capability to interpret visual input in response to implicit text query through multi-step reasoning, remains a challenge for deep learning models due to the lack of relevant benchmarks. Previous work in visual reasoning has primarily focused on reasoning segmentation, where models aim to segment objects based on implicit text queries. This paper introduces reasoning visual tasks (RVTs), a unified formulation that extends beyond traditional video reasoning segmentation to a diverse family of visual language reasoning problems, which can therefore accommodate multiple output formats including bounding boxes, natural language descriptions, and question-answer pairs. Correspondingly, we identify the limitations in current benchmark construction methods that rely solely on large language models (LLMs), which inadequately capture complex spatial-temporal relationships and multi-step reasoning chains in video due to their reliance on token representation, resulting in benchmarks with artificially limited reasoning complexity. To address this limitation, we propose a novel automated RVT benchmark construction pipeline that leverages digital twin (DT) representations as structured intermediaries between perception and the generation of implicit text queries. Based on this method, we construct RVTBench, a RVT benchmark containing 3,896 queries of over 1.2 million tokens across four types of RVT (segmentation, grounding, VQA and summary), three reasoning categories (semantic, spatial, and temporal), and four increasing difficulty levels, derived from 200 video sequences. Finally, we propose RVTagent, an agent framework for RVT that allows for zero-shot generalization across various types of RVT without task-specific fine-tuning. Dataset and code are available at https://doi.org/10.5281/zenodo.19697191 and https://github.com/yiqings/rvt.
Purpose Vision-based navigation systems rely on the registered camera poses in the CT space to guide surgeons. However, while it is possible to provide an approximate initialization, this registration becomes outdated as the endoscopic camera leaves and reenters the anatomy. Endoscopic camera relocalization is the process of determining the position of an endoscope relative to an anatomical reference after reinsertion. However, accurately reidentifying the global surgical scene and estimating camera pose have proven challenging due to the varying appearance of endoscopic sequences. Methods We present a training-free approach to accurately reidentify the region of interest (ROI) and estimate the camera position of a query image after reinsertion. This method utilizes previously observed images with known poses and a CT scan. By combining advanced foundation models with classical techniques, we globally reidentify a prior image of the ROI, which is then used for image-based feature matching and pose recovery via the Perspective-n-Point algorithm. Results We conducted experiments on eight sequences from three cadaver studies. Our results show that our method accurately reidentifies when the endoscope reaches the ROI and identifies suitable image pairs for PnP-based pose estimation. It achieves an average translation error of 1.74 mm and a rotational error of 0.09 radians, making it suitable for reinitialization in image-based navigation without human intervention. Conclusion Our work presents a training-free approach for detecting when the endoscope reenters the ROI and estimating the camera's pose after reinsertions. The approach demonstrates promising results contributing toward enabling pose reinitialization for vision-based surgical applications.
Reasoning Segmentation (RS) aims to delineate objects based on implicit text queries, the interpretation of which requires reasoning and knowledge integration. Unlike the traditional formulation of segmentation problems that relies on fixed semantic categories or explicit prompting, RS bridges the gap between visual perception and human-like reasoning capabilities, facilitating more intuitive human-AI interaction through natural language. Our work presents the first comprehensive survey of RS for image and video processing, examining 26 state-of-the-art methods together with a review of the corresponding evaluation metrics, as well as 29 datasets and benchmarks. We also explore existing applications of RS across diverse domains and identify their potential extensions. Finally, we identify current research gaps and highlight promising future directions.
Autonomous medical robots hold promise to improve patient outcomes, reduce provider workload, democratize access to care, and enable superhuman precision. However, autonomous medical robotics has been limited by a fundamental data problem: existing medical robotic datasets are small, single-embodiment, and rarely shared openly, restricting the development of foundation models that the field needs to advance. We introduce Open-H-Embodiment, the largest open dataset of medical robotic video with synchronized kinematics to date, spanning more than 49 institutions and multiple robotic platforms including the CMR Versius, Intuitive Surgical's da Vinci, da Vinci Research Kit (dVRK), Rob Surgical BiTrack, Virtual Incision's MIRA, Moon Surgical Maestro, and a variety of custom systems, spanning surgical manipulation, robotic ultrasound, and endoscopy procedures. We demonstrate the research enabled by this dataset through two foundation models. GR00T-H is the first open foundation vision-language-action model for medical robotics, which is the only evaluated model to achieve full end-to-end task completion on a structured suturing benchmark (25
Large language model-based (LLM) agents are emerging as a powerful enabler of robust embodied intelligence due to their capability of planning complex action sequences. Sound planning ability is necessary for robust automation in many task domains, but especially in surgical automation. These agents rely on a highly detailed natural language representation of the scene. Thus, to leverage the emergent capabilities of LLM agents for surgical task planning, developing similarly powerful and robust perception algorithms is necessary to derive a detailed scene representation of the environment from visual input. Previous research has focused primarily on enabling LLM-based task planning while adopting simple yet severely limited perception solutions to meet the needs for bench-top experiments, but lacks the critical flexibility to scale to less constrained settings. In this work, we propose an alternate perception approach – a digital twin (DT)-based machine perception approach that capitalizes on the convincing performance and out-of-the-box generalization of recent vision foundation models. Integrating our DT representation and LLM agent for planning with the dVRK platform, we develop an embodied intelligence system and evaluate its robustness in performing peg transfer and gauze retrieval tasks. Our approach shows strong task performance and generalizability to varied environmental settings. Despite a convincing performance, this work is merely a first step towards the integration of DT representations. Future studies are necessary for the realization of a comprehensive DT framework to improve the interpretability and generalizability of embodied intelligence in surgery.
The discovery and evolution of medical imaging technologies have enabled non-invasive visualization of internal anatomy that has be- come essential for supporting diagnosis, monitoring, and treatment. However, because medical imaging relies on complex physical processes and contrast mechanisms for image formation, imaging alone is not sufficient to enable humans to leverage the resulting in- formation fully. In addition, traditional methods to visualize the resulting information use two-dimensional displays to present three-dimensional anatomical structures. The introduction of Augmented and Mixed Reality (AR/MR) technologies offers an opportunity to provide valuable paradigms for medical imaging visualization, allowing users to observe, explore, and interact with anatomical infor- mation in more spatially intuitive ways. However, naive implementation without careful design considerations can lead to perceptual inconsistencies, potentially compromising utility and effectiveness. In this paper, we present a structured taxonomy of medical AR/MR visualization strategies aimed at providing clearer insight into how visualization design varies across clinical use cases. The taxonomy organizes techniques based on four core design components: image modality, data dimensionality, display technology, and clinical application. In addition, we introduce two critical dimensions that are often overlooked in the literature: visualization anchoring (the spatial relationship between virtual content and the physical world), and perceptual awareness (the use of visual cues to support spatial interpretation). Together, these components form a comprehensive taxonomy, offering a detailed framework for selecting appropriate visualization techniques in medical applications.
Text-to-video retrieval in operating rooms (OR) is an enabling technology for OR safety, as it allows stakeholders to retrieve and inspect recordings of specific events. However, because the most safety-critical events may not follow the common structure, to unlock its full potential text-to-video retrieval must be able to handle implicit queries that require reasoning to identify the right video (e.g., the step right before clipping). However, existing methods rely on global embeddings that cannot reason over such queries. We propose OR3, a text-to-video retrieval method that converts clips into action-driven digital twins (ActDTs), grouping concurrent subject-action-object triplets under non-overlapping temporal intervals. Moreover, rather than cross-modal matching through paired encoders, OR3 performs imagination-based retrieval where an LLM generates hypothetical ActDTs from queries. This enables intra-modal matching via a single encoder trained with ActDT-tailored hard negatives. Finally, evidence-grounded refinement revises imagined ActDTs based on discrepancies with top candidates to capture procedure-specific patterns. We construct a benchmark from MM-OR with 276 implicit queries across four reasoning categories over 386 clips from robotic knee procedures. OR3 achieves 57.6 R@1 and 77.3 R@5, outperforming the strongest baseline. These results demonstrate that OR3 enables fine-grained discrimination between visually similar OR video clips through temporal action reasoning.
Replicating human-level intelligence in embodied task execution remains challenging due to the unconstrained nature of real-world environments. Recent use of large language models (LLMs) for task planning seeks to address the previously intractable state/action space of complex planning tasks, but hallucinations limit their viability. Additionally, the prompt engineering required for adequate system performance lacks transparency and repeatability. In contrast, symbolic planning methods offer strong reliability and repeatability guarantees, but struggle to scale to the complexity of real-world tasks. We introduce a new planning method that augments LLM planners with symbolic planning oversight to improve reliability and repeatability, and provide a transparent approach to defining hard constraints with more clarity than traditional prompt engineering. We demonstrate our approach in simulated environments, outperforming current state-of-the-art methods. Deployment of our method to a real-world quadruped robot resulted in 75% task success compared to 50% and 14.3% for pure LLM and symbolic planners across several embodied tasks, many of which required complex reasoning and interaction with humans in realistic scenarios. Our approach presents an effective strategy to enhance the reliability, repeatability, and transparency of LLM-based robot planners while retaining their key strengths: flexibility and generalizability to complex real-world environments.
Humanoid robots have become a focal point of technological ambition, with claims of surgical capability within years in mainstream discourse. These projections are aspirational yet lack empirical grounding. To date, no humanoid has assisted a surgeon through an actual procedure, let alone performed one. The work described here breaks this new ground. Here we report a proof of concept in which a teleoperated Unitree G1 provided endoscopic visualization while an attending otolaryngologist performed a cadaveric sphenoidectomy. The procedure was completed successfully, with stable visualization maintained throughout. Teleoperation allowed assessment of whether the humanoid form factor could meet the physical demands of surgical assistance in terms of sustenance and precision; the cognitive demands were satisfied -- for now -- by the operator. Post-procedure analysis identified engineering targets for clinical translation, alongside near-term opportunities such as autonomous diagnostic scoping. This work establishes form-factor feasibility for humanoid surgical assistance while identifying challenges for continued development.
Purpose Imitation learning-based robot control policies are enjoying renewed interest in video-based robotics. However, it remains unclear whether this approach applies to X-ray-guided procedures, such as spine instrumentation, with sparse inputs. We examine the feasibility, opportunities and challenges for imitation policy learning in bi-plane-guided cannula insertion. Method We develop an in silico sandbox for scalable, automated simulation of X-ray-guided spine procedures with a high degree of realism. We curate a dataset of correct trajectories and corresponding bi-planar X-ray sequences that emulate the stepwise alignment of providers. We then train imitation learning policies for planning and open-loop control that iteratively align a cannula solely based on visual information. This precisely controlled setup offers insights into limitations and capabilities of this method. Results Our policy succeeded on the first attempt in 68.5% of cases, maintaining safe intra-pedicular trajectories across diverse vertebral levels. The policy transferred to complex anatomy, including fractures, as well as varied anatomies and initializations. Rollouts on real X-ray indicate that partial sim-to-real transfer with plausible trajectories is possible. Conclusion While these preliminary results are promising, we also identify limitations, especially in entry-point precision. The current results present a clear benchmark for future efforts, while with more robust priors and domain knowledge, such models may provide a foundation for future efforts toward lightweight and CT-free robotic intra-operative spinal navigation.
Central airway obstruction (CAO) procedures require precise, repetitive dual-arm coordination under progressively degraded visibility from smoke and charring. Model-based control struggles with deformable-tissue variability and real-time constraints, while existing data-driven methods emphasize short, single-skill tasks rather than cyclic workflows. We investigate whether imitation learning can achieve fully autonomous task-level and high-level for CAO tumor resection on da Vinci Research Kit. We propose a hierarchical framework combining low-level action chunking with transformers (ACT) policies for autonomous task-level execution with finite state machine (FSM)-based high-level coordination for procedure-level sequencing. The workflow decomposes into five recurring tasks. Policies are trained on synchronized multi-view video and kinematics with hybrid-relative actions. We also analyze task-level policies with different data regimes through an ablation study. Low-level policies achieved fully autonomous execution with 0.938mm RMSE without intervention, comparable to an operator (0.968 mm) and surgeon (1.086 mm). High-level coordination enabled supervised autonomous full procedure execution with minimal interventions resulting in 1.232mm accuracy. Ablation studies showed balanced task coverage with minimal data (3
Accurate vision-based navigation in monocular endoscopy is difficult due to limited depth cues, weak tissue texture, non-rigid deformation, and substantial appearance variation across domains, all of which complicate pose estimation, depth prediction, and image-to-anatomy alignment. Although recent vision foundation models have shown promise, their learned representations often remain insufficiently geometry-consistent, hindering stable feature correspondence and limiting their reliability for downstream navigation tasks. We propose a unified framework for learning geometry-consistent and domain-robust image representations for monocular endoscopy. The framework combines a synthetic data pipeline that provides accurate geometric supervision with Hierarchy-Aware Geometry-Semantic Adaptation, a structured alternative to standard LoRA that inserts low-rank adapters selectively across the transformer hierarchy and couples them with layer-wise training objectives to encourage geometric correspondence in intermediate features and semantic consistency in deeper features. Experiments on public and proprietary datasets show improved geometric and semantic representation quality, leading to better performance on downstream navigation tasks including pose estimation and monocular depth estimation. The learned representations show favorable synthetic-to-real transfer on clinical bronchoscopy and provide a useful initialization for adaptation to sinus endoscopy and colonoscopy under limited supervision. The framework also shows favorable scaling with model size and training data. These results support hierarchy-aware, geometry-guided adaptation as a practical approach for endoscopic representation learning.
Patient-specific anatomical models provide individualized context for surgical planning, image-guided intervention, and algorithm development. However, most CT-derived models are static: they preserve the body configuration captured at scan time, but cannot represent how the same anatomy would appear after patient repositioning. This limitation is especially important for radiographic imaging, where appearance depends jointly on imaging geometry and patient pose. We present a proof-of-concept for constructing a patient-specific articulated digital twin from a single full-body CT scan. The method fits a parametric human body model (SMPL) to obtain a patient-aligned kinematic scaffold, binds segmented bones and organs to an anatomy-aware rig, and retargets body-pose changes while preserving skeletal geometry. On three full-body CT subjects, the fitted scaffold achieved 15.8 ± 4.0 mm chamfer distance and 95.9 ± 1.8
Reliable volumetric representation of the nasal cavity is crucial for enabling quantitative assessment in Functional Endoscopic Sinus Surgery (FESS), yet for most patients the evaluation of their anatomy remains largely qualitative and subjective. While computed tomography (CT) scans can provide 3D anatomical information, their routine use is impractical due to radiation exposure concerns and cost constraints, underscoring the need for a non-invasive alternative. Computer vision methods offer a promising solution for reconstructing sinus anatomy from routine endoscopic video. Current methods rely on Structure-from-Motion (SfM), however, this relies on point correspondences that struggle with photometric inconsistencies inherent to endoscopic imaging, reducing robustness and generalizability. Several sinus reconstruction approaches attempt to mitigate this through learning-based and patient-specific approaches, but suffer from error propagation, leading to inaccurate 3D representations. Optimization-based approaches further introduce excessive training times, limiting their practicality. In this work, we revisit simpler techniques for sinus reconstruction and augment them with track-any-point foundation models to develop a training-free, vision-based 3D reconstruction method. Our approach leverages SfM poses and local point-tracks to generate depth information, recovering a globally consistent structure without fine-tuning requirements. We evaluate our method on six pre-operative endoscopic sequences with respect to the ground-truth CT scan. Our results show that this method improves global geometric accuracy by reducing both point-to-point and pose errors from prior work. Our vision-based approach improves spatial consistency and accuracy in sinus 3D reconstruction, enabling non-invasive postoperative monitoring and seamless clinical integration, offering physicians data-driven insights for improved surgical decision-making.
The segmentation of pelvic fracture fragments in CT and X-ray images is crucial for trauma diagnosis, surgical planning, and intraoperative guidance. However, accurately and efficiently delineating the bone fragments remains a significant challenge due to complex anatomy and imaging limitations. The PENGWIN challenge, organized as a MICCAI 2024 satellite event, aimed to advance automated fracture segmentation by benchmarking state-of-the-art algorithms on these complex tasks. A diverse dataset of 150 CT scans was collected from multiple clinical centers, and a large set of simulated X-ray images was generated using the DeepDRR method. Final submissions from 16 teams worldwide were evaluated under a rigorous multi-metric testing scheme. The top-performing CT algorithm achieved an average fragment-wise intersection over union (IoU) of 0.930, demonstrating satisfactory accuracy. However, in the X-ray task, the best algorithm achieved an IoU of 0.774, which is promising but not yet sufficient for intra-operative decision-making, reflecting the inherent challenges of fragment overlap in projection imaging. Beyond the quantitative evaluation, the challenge revealed methodological diversity in algorithm design. Variations in instance representation, such as primary-secondary classification versus boundary-core separation, led to differing segmentation strategies. Despite promising results, the challenge also exposed inherent uncertainties in fragment definition, particularly in cases of incomplete fractures. These findings suggest that interactive segmentation approaches, integrating human decision-making with task-relevant information, may be essential for improving model reliability and clinical applicability.
PurposeAccurate intra-operative localization of the endoscope tip relative to the anatomy remains a major challenge in bronchoscopy due to respiratory motion, anatomical variability, and CT-to-body divergence, which cause deformation and misalignment between intra-operative views and pre-operative CT. Existing vision-based methods often struggle to generalize across domains and patients, limiting robustness and leaving residual alignment errors. This work aims to establish a generalizable foundation for bronchoscopy navigation through a robust vision-based framework and a new synthetic benchmark dataset that enables standardized evaluation and reproducible development.MethodsWe propose a vision-based pose optimization framework for frame-wise 2D-3D registration between intra-operative endoscopic views and pre-operative CT anatomy. A fine-tuned modality- and domain-invariant encoder enables direct similarity measurements between real endoscopic RGB images and CT-rendered depth maps, while differentiable rendering refines camera poses through depth consistency. To enhance reproducibility, we introduce the first public synthetic benchmark dataset for bronchoscopy navigation to address the lack of publicly available paired CT-endoscopy data.ResultsTrained solely on synthetic data distinct from the benchmark, our model attains an average translational error of 2.65 mm and a rotational error of 0.19 rad, demonstrating high localization accuracy and stability. Qualitative results on real patient data further confirm strong cross-domain generalization, achieving consistent frame-wise 2D-3D alignment without domain-specific adaptation.ConclusionThe proposed framework achieves robust, domain-invariant bronchoscopy localization through iterative vision-based optimization, offering a scalable solution toward reliable vision-based bronchoscopy localization. The introduced synthetic benchmark dataset provides a valuable resource for standardized evaluation on bronchoscopy navigation.
Participation in crowd-sourced user studies is often driven by monetary incentives. However, standard payment schemes that reward completion unless responses are of poor quality may not invoke sufficient accountability. By compromising user engagement, a lack of accountability can affect data quality and the study’s ecological validity. Here, we investigate alternative compensation strategies that manipulate payment framing and evaluate their impact on engagement through task effort, outcomes, and perception. We compared a standard scheme with implicit rejection risk to a reinforced accountability condition with explicit performance-linked deductions, and two dynamic conditions that unexpectedly switched strategies. In a study with 106 Prolific participants on an image captioning task, we found that only shifting from implicit risk to reinforced accountability significantly increased engagement, likely due to loss aversion after participants had already invested time. The reverse shift decreased effort as observed in the standard group. Our results highlight the importance of carefully designing compensation schemes.