Multi-object tracking (MOT) has been a subject of intensive research for decades. Multiple standard datasets and benchmarks have been set up, and several evaluation metrics, such as MOTA, IDF1 and HOTA. These metrics have become the de facto standard for comparing and ranking trackers on standardized datasets to measure progress. In this paper, we focus on MOTA and HOTA, and present a study of cases where these metrics’ behaviors may not be desirable. In addition, we demonstrate how they might not be ideal when used as a tool to inspect a tracker’s failure cases. We point out that these issues are related to the sizes of the context windows in which they measure association quality, where MOTA is too nearsighted while HOTA can be too holistic depending on the task settings.In this paper, we rethink the familiar notion of identity switches (IDSw) proposed in MOTA, and propose a generalized version of it by introducing a context window when evaluating the ID assignment choice for each detection. We show that the proposed metric, CAST, mitigates the limitations of MOTA and HOTA, and demonstrate its usefulness when diagnosing model failures through examples. Our code and toolkit will be made available at https://github.com/bkkm78/cast.
In the last ten years, medical robotics has moved from the margins to the mainstream. Since the Engineering Research Center for Computer-Integrated Surgical Systems and Technology was Launched in 1998 with National Science Foundation funding, medical robots have been promoted from handling routine tasks to performing highly sophisticated interventions and related assignments. The CISST ERC has played a significant role in this transformation. And thanks to NSF support, the ERC has built the professional infrastructure that will continue our mission: bringing data and technology together in clinical systems that will dramatically change how surgery and other procedures are done. The enhancements we envision touch virtually every aspect of the delivery of care: - More accurate procedures - More consistent, predictable results from one patient to the next - Improved clinical outcomes - Greater patient safety - Reduced liability for healthcare providers - Lower costs for everyone - patients, facilities, insurers, government - Easier, faster recovery for patients - Effective new ways to treat health problems - Healthier patients, and a healthier system The basic science and engineering the ERC is developing now will yield profound benefits for all concerned about health care - from government agencies to insurers, from clinicians to patients to the general public. All will experience the healing touch of medical robotics, thanks in no small part to the work of the CISST ERC and its successors.
Humanoid robots have become a focal point of technological ambition, with claims of surgical capability within years in mainstream discourse. These projections are aspirational yet lack empirical grounding. To date, no humanoid has assisted a surgeon through an actual procedure, let alone performed one. The work described here breaks this new ground. Here we report a proof of concept in which a teleoperated Unitree G1 provided endoscopic visualization while an attending otolaryngologist performed a cadaveric sphenoidectomy. The procedure was completed successfully, with stable visualization maintained throughout. Teleoperation allowed assessment of whether the humanoid form factor could meet the physical demands of surgical assistance in terms of sustenance and precision; the cognitive demands were satisfied -- for now -- by the operator. Post-procedure analysis identified engineering targets for clinical translation, alongside near-term opportunities such as autonomous diagnostic scoping. This work establishes form-factor feasibility for humanoid surgical assistance while identifying challenges for continued development.
Reliable volumetric representation of the nasal cavity is crucial for enabling quantitative assessment in Functional Endoscopic Sinus Surgery (FESS), yet for most patients the evaluation of their anatomy remains largely qualitative and subjective. While computed tomography (CT) scans can provide 3D anatomical information, their routine use is impractical due to radiation exposure concerns and cost constraints, underscoring the need for a non-invasive alternative. Computer vision methods offer a promising solution for reconstructing sinus anatomy from routine endoscopic video. Current methods rely on Structure-from-Motion (SfM), however, this relies on point correspondences that struggle with photometric inconsistencies inherent to endoscopic imaging, reducing robustness and generalizability. Several sinus reconstruction approaches attempt to mitigate this through learning-based and patient-specific approaches, but suffer from error propagation, leading to inaccurate 3D representations. Optimization-based approaches further introduce excessive training times, limiting their practicality. In this work, we revisit simpler techniques for sinus reconstruction and augment them with track-any-point foundation models to develop a training-free, vision-based 3D reconstruction method. Our approach leverages SfM poses and local point-tracks to generate depth information, recovering a globally consistent structure without fine-tuning requirements. We evaluate our method on six pre-operative endoscopic sequences with respect to the ground-truth CT scan. Our results show that this method improves global geometric accuracy by reducing both point-to-point and pose errors from prior work. Our vision-based approach improves spatial consistency and accuracy in sinus 3D reconstruction, enabling non-invasive postoperative monitoring and seamless clinical integration, offering physicians data-driven insights for improved surgical decision-making.
In endoscopic surgery, surgeons continuously locate the endoscopic view relative to the anatomy by interpreting the evolving visual appearance of the intraoperative scene in the context of their prior knowledge. Vision-based navigation systems seek to replicate this capability by recovering camera pose directly from endoscopic video, but most approaches do not embody the same principles of reasoning about new frames that makes surgeons successful. Instead, they remain grounded in feature matching and geometric optimization over keyframes, an approach that has been shown to degrade under the challenging conditions of endoscopic imaging like low texture and rapid illumination changes. Here, we pursue an alternative approach and investigate a policy-based formulation of endoscopic camera pose recovery that seeks to imitate experts in estimating trajectories conditioned on the previous camera state. Our approach directly predicts short-horizon relative motions without maintaining an explicit geometric representation at inference time. It thus addresses, by design, some of the notorious challenges of geometry-based approaches, such as brittle correspondence matching, instability in texture-sparse regions, and limited pose coverage due to reconstruction failure. We evaluate the proposed formulation on cadaveric sinus endoscopy. Under oracle state conditioning, we compare short-horizon motion prediction quality to geometric baselines achieving lowest mean translation error and competitive rotational accuracy. We analyze robustness by grouping prediction windows according to texture richness and illumination change indicating reduced sensitivity to low-texture conditions. These findings suggest that a learned motion policy offers a viable alternative formulation for endoscopic camera pose recovery.
The most exciting frontiers of contemporary endoscopic video processing include tissue and instrument tracking and scene reconstruction together with their applications in surgical navigation, mixed reality visualization, and surgical automation. These tasks are enabled by task-specific techniques such as segmentation, object tracking, structure from motion, simultaneous localization and mapping (SfM and SLAM), and neural rendering. While these techniques vary in purpose and methodology, fundamentally, they rely on the ability to reliably and robustly establish point correspondences across video frames. However, point tracking in endoscopic scenes is hard due to the lack of distinct yet repetitive features, varying illuminations, and continuously changing visual appearance of corresponding points. A dense point-tracking model capable of reliably establishing point correspondences across video frames could catalyze endoscopic video processing and its downstream applications, and drive significant advancement in surgical data science. While any point-tracking foundation models have recently been proposed, they are trained on large simulated and natural scene videos and need to be adapted to endoscopic scenes to address domain-specific challenges. In this work, we present a dense point tracking foundation model for endoscopic scenes by fine-tuning a public foundation model on a large custom dataset comprising 13k endoscopic video sequences (314k frames). We present three benchmark datasets with ground-truth point correspondences to quantitatively evaluate the point tracking performance in variable endoscopic scenes. Through quantitative analysis, we find that models fine-tuned on endoscopic scenes outperform out-of-the-box models, especially when considering conservative thresholds for tracking success, suggesting improved suitability for downstream tasks due to error propagation.
In this paper, we report our discovery of a gaze behavior called Quiet Eye (QE) in minimally invasive surgery. The QE behavior has been extensively studied in sports training and has been associated with higher level of expertise in multiple sports. We investigated the QE behavior in two independently collected data sets of surgeons performing tasks in a sinus surgery setting and a robotic surgery setting, respectively. Our results show that the QE behavior is more likely to occur in successful task executions and in performances of surgeons of high level of expertise. These results open the door to use the QE behavior in both training and skill assessment in minimally invasive surgery.
INTRODUCTION: The ablative nature of surgery means that pre-operative imaging studies lose correspondence as a case progresses, which can be problematic when accurate intraoperative navigation is required. Accurate 3D surface reconstruction from endoscopic video is a potential strategy for real-time intraoperative imaging updates without additional equipment. We have previously used traditional computational models to generate skull base reconstructions. However, they are time-consuming and require technical skills to process the video. Recent foundational AI models, like DUST3R, are an opportunity for timely, generalizable reconstructions of surgical anatomy. METHODS: We compared our previously described three-step reconstruction process with DUST3R to generate a water-tight 3D mesh of cadaveric skull base anatomy visualized using a Karl Storz Image 1 Hub HD Video Camera fitted with a 0° rigid endoscope. For DUST3R, we selected four video frames and did not perform any training or fine-tuning. RESULTS: Without endoscope calibration and using only the four input frames, DUST3R created 3D surface reconstructions in less than two minutes. This is compared to the three-step reconstruction process, which requires 8 to 12 hours to reconstruct. CONCLUSIONS: Our findings show that DUST3R, a publicly available foundational AI model, can rapidly generate 3D anatomical reconstructions from a limited set of video frames. Models like DUST3R illustrate the rapidly evolving potential of computer vision. With fine-tuning, they may represent a path toward foundational AI models that generalize across surgical procedures.
Objective: Transventricular approach to deep-brain targets offers direct visualization but also imparts deformation that challenges accurate neuronavigation. 3D reconstruction and registration of the endoscopic view could provide up-to-date, realtime guidance. We develop and evaluate a self-supervised feature detection method for 3D reconstruction and navigation in neuroendoscopy. Methods: Unlabeled neuroendoscopic video data from 15 clinical cases yielding 11,527 video frames yielding 11,527 video frames were used to train a self-supervised learning method (R2D2-E) with 5-fold cross validation integrated into a simultaneous localization and mapping (SLAM) pipeline for 3D reconstruction. A series of experiments guided nominal hyperparameters selection and evaluated performance in comparison to SIFT, SURF and SuperPoint in terms of the accuracy of feature matching and 3D reconstruction. Results: R2D2-E demonstrated a superior performance in feature matching and 3D reconstruction. R2D2-E features achieved a median projected error of 0.64 mm compared to 0.90 mm, 0.99 mm and 0.83 mm error for SIFT, SURF and SuperPoint, respectively. The method also improved F1 score by 14%, 25% and 22% compared to SIFT, SURF and SuperPoint, respectively. Conclusion: The proposed feature detection approach enables accurate, real-time 3D reconstruction in neuroendoscopy, offering robust feature detection in the presence of endoscopic artifacts and provides up-to-date navigation following soft-tissue deformation. Significance: The self-supervised feature detection method advances capabilities for vision-based guidance and augmented visualization of target structures in neuroendoscopic procedures. The approach could enhance the accuracy and precision of neurosurgery to improve patient outcomes.
Purpose Monocular SLAM algorithms are the key enabling technology for image-based surgical navigation systems for endoscopic procedures. Due to the visual feature scarcity and unique lighting conditions encountered in endoscopy, classical SLAM approaches perform inconsistently. Many of the recent approaches to endoscopic SLAM rely on deep learning models. They show promising results when optimized on singular domains such as arthroscopy, sinus endoscopy, colonoscopy or laparoscopy, but are limited by an inability to generalize to different domains without retraining. Methods To address this generality issue, we propose OneSLAM a monocular SLAM algorithm for surgical endoscopy that works out of the box for several endoscopic domains, including sinus endoscopy, colonoscopy, arthroscopy and laparoscopy. Our pipeline builds upon robust tracking any point (TAP) foundation models to reliably track sparse correspondences across multiple frames and runs local bundle adjustment to jointly optimize camera poses and a sparse 3D reconstruction of the anatomy. Results We compare the performance of our method against three strong baselines previously proposed for monocular SLAM in endoscopy and general scenes. OneSLAM presents better or comparable performance over existing approaches targeted to that specific data in all four tested domains, generalizing across domains without the need for retraining. Conclusion OneSLAM benefits from the convincing performance of TAP foundation models but generalizes to endoscopic sequences of different anatomies all while demonstrating better or comparable performance over domain-specific SLAM approaches. Future research on global loop closure will investigate how to reliably detect loops in endoscopic scenes to reduce accumulated drift and enhance long-term navigation capabilities.
We present a straightforward, flexible method to enhance the accuracy and quality of action detection by expressing temporal and structural relationships of actions in the loss function of a deep network. We describe ways to represent otherwise implicit structure in video data and demonstrate how these structures reflect natural biases that improve network training. Our experiments show that our approach improves both accuracy and edit-distance of action recognition and detection models over a baseline. Our framework leads to improvements over prior work and obtains state-of-the-art results on multiple benchmarks. The code is available here.
Machine learning approaches for multi-view geometric scene understanding in endoscopic surgery often assume temporal consistency across the frames to limit challenges that algorithms contend with. However, in monocular scenarios where multiple views are acquired sequentially rather than simultaneously, the static scene assumption is too strong because surgical tools move during the procedure. To enable multi-view models despite tool motion, masking these temporally inconsistent tool regions is a feasible solution. However, manual tool-masking requires a prohibitive effort, given that endoscopic video can contain thousands of frames. This underscores the need for (semi-)automated techniques to 1) automatically mask the tools and/or 2) semi-automatically annotate large datasets such that algorithms for 1) may be developed. To facilitate semi-automated annotation, any solution must be both generalizable, such that it can be used out-of-the-box on diverse datasets, and easy to use. Recent methods for surgical tool annotation require either fine-tuning on domain-specific data or excessive user interaction, limiting their application to new data. Our work introduces GSAM+Cutie, a surgical tool annotation process relying on a combination of two recent foundation models for text-based image segmentation and video object segmentation. We show that a combination of Grounded-SAM and Cutie models provides good generalization for robust text-prompt-based video-level binary segmentation on endoscopic video, streamlining the mask annotation task. Through quantitative evaluation on two datasets, including a proprietary in-house dataset and EndoVis, we show that GSAM+Cutie outperforms similar ensembles, like SAM-PT, for video object segmentation. We also discuss the limitations and future research directions that GSAM+Cutie can motivate. Our code is available at https://github. com/arcadelab/cutie_plus_gsam
Accurate, unbiased, and reproducible assessment of skill is a vital resource for surgeons throughout their career. The objective in this research is to develop and validate algorithms for video-based assessment of intraoperative surgical skill. Algorithms to classify surgical video into expert or novice categories provide a summative assessment of skill, which is useful for evaluating surgeons at discrete time points in their training or certification of surgeons. Using a spatial-temporal neural network architecture, we tested the hypothesis that explicit supervision of spatial attention supervised by instrument tip locations improves the algorithm’s generalizability to unseen dataset. The best performing model had an area under the receiver operating characteristic curve (AUC) of 0.88. Augmenting the network with supervision of spatial attention improved specificity of its predictions (with small changes in sensitivity and AUC) and led to improved measures of discrimination when tested with unseen dataset. Our findings show that explicit supervision of attention learned from images using instrument tip locations can improve performance of algorithms for objective video-based assessment of surgical skill.
Purpose: Navigating deep-brain structures in neurosurgery, especially under deformation from CSF egress, remains challenging due to the limitations of current robotic systems relying on rigid registration. This study presents the initial steps towards vision-based navigation leveraging Neural Radiance Fields (NeRF) to enable 3D neuroendoscopic reconstruction on the Robot-Assisted Ventriculoscopy (RAV) platform. Methods: An end-to-end 3D reconstruction and registration method using posed images was developed and integrated with the RAV platform. The hyperparameters for training the dual-branch network were first identified. Further experiments were conducted to evaluate reconstruction accuracy using projected error (PE) while varying the volume density threshold parameter. Results: A 3D volume was reconstructed using a simple linear trajectory for data acquisition with 300 frames and corresponding camera poses. The density volume threshold was varied to obtain an optimal value of 96.55 percentile, with a corresponding PE of 0.65 mm. Conclusions: Initial methods for end-to-end neuroendoscopic video reconstruction were developed in phantom studies. Experiments identified the optimal parameters, yielding a geometrically accurate reconstruction along with fast network convergence runtime of < 30 s. The method is highly promising for future clinical translation in realistic neuroendoscopic scenes. Future work will also develop a direct surface-to-volume registration method for improving reconstruction accuracy and runtime.
OBJECTIVE:To estimate and adjust for rater effects in operating room surgical skills assessment performed using a structured rating scale for nasal septoplasty. METHODS:We analyzed survey responses from attending surgeons (raters) who supervised residents and fellows (trainees) performing nasal septoplasty in a prospective cohort study. We fit a structural equation model with the rubric item scores regressed on a latent component of skill and then fit a second model including the rating surgeon as a random effect to model a rater-effects-adjusted latent surgical skill. We validated this model against conventional measures including the level of expertise and post-graduation year (PGY) commensurate with the trainee's performance, the actual PGY of the trainee, and whether the surgical goals were achieved. RESULTS:Our dataset included 188 assessments by 7 raters and 41 trainees. The model with one latent construct for surgical skill and the rater as a random effect was the best. Rubric scores depended on how severe or lenient the rater was, sometimes almost as much as they depended on trainee skill. Rater-adjusted latent skill scores increased with attending-estimated skill levels and PGY of trainees, increased with the actual PGY, and appeared constant over different levels of achievement of surgical goals. CONCLUSION:Our work provides a method to obtain rater effect adjusted surgical skill assessments in the operating room using structured rating scales. Our method allows for the creation of standardized (i.e., rater-effects-adjusted) quantitative surgical skill benchmarks using national-level databases on trainee assessments. LEVEL OF EVIDENCE:N/A Laryngoscope, 134:3548-3554, 2024.
PURPOSE:Develop and evaluate the performance of a deep learning model (DLM) that forecasts eyes with low future visual field (VF) variability, and study the impact of using this DLM on sample size requirements for neuroprotective trials. DESIGN:Retrospective cohort and simulation study. METHODS:We included 1 eye per patient with baseline reliable VFs, OCT, clinical measures (demographics, intraocular pressure, and visual acuity), and 5 subsequent reliable VFs to forecast VF variability using DLMs and perform sample size estimates. We estimated sample size for 3 groups of eyes: all eyes (AE), low variability eyes (LVE: the subset of AE with a standard deviation of mean deviation [MD] slope residuals in the bottom 25th percentile), and DLM-predicted low variability eyes (DLPE: the subset of AE predicted to be low variability by the DLM). Deep learning models using only baseline VF/OCT/clinical data as input (DLM1), or also using a second VF (DLM2) were constructed to predict low VF variability (DLPE1 and DLPE2, respectively). Data were split 60/10/30 into train/val/test. Clinical trial simulations were performed only on the test set. We estimated the sample size necessary to detect treatment effects of 20% to 50% in MD slope with 80% power. Power was defined as the percentage of simulated clinical trials where the MD slope was significantly worse from the control. Clinical trials were simulated with visits every 3 months with a total of 10 visits. RESULTS:A total of 2817 eyes were included in the analysis. Deep learning models 1 and 2 achieved an area under the receiver operating characteristic curve of 0.73 (95% confidence interval [CI]: 0.68, 0.76) and 0.82 (95% CI: 0.78, 0.85) in forecasting low VF variability. When compared with including AE, using DLPE1 and DLPE2 reduced sample size to achieve 80% power by 30% and 38% for 30% treatment effect, and 31% and 38% for 50% treatment effect. CONCLUSIONS:Deep learning models can forecast eyes with low VF variability using data from a single baseline clinical visit. This can reduce sample size requirements, and potentially reduce the burden of future glaucoma clinical trials. FINANCIAL DISCLOSURE(S):Proprietary or commercial disclosure may be found in the Footnotes and Disclosures at the end of this article.
In this work, we introduce the Virtual In-Hand Eye Transformer (VIHE), a novel method designed to enhance 3D manipulation capabilities through action-aware view rendering. VIHE autoregressively refines actions in multiple stages by conditioning on rendered views posed from action predictions in the earlier stages. These virtual in-hand views provide a strong inductive bias for effectively recognizing the correct pose for the hand, especially for challenging high-precision tasks such as peg insertion. On 18 manipulation tasks in RLBench simulated environments, VIHE achieves a new state-of-the-art, with a 12% absolute improvement, increasing from 65% to 77% over the existing state-of-the-art model using 100 demonstrations per task. In real-world scenarios, VIHE can learn manipulation tasks with just a handful of demonstrations, highlighting its practical utility. Videos and code implementation can be found at our project site: https://vihe-3d.github.io.
Image-based reinforcement learning (RL) faces significant challenges in generalization when the visual environment undergoes substantial changes between training and deployment. Under such circumstances, learned policies may not perform well leading to degraded results. Previous approaches to this problem have largely focused on broadening the training observation distribution, employing techniques like data augmentation and domain randomization. However, given the sequential nature of the RL decision-making problem, it is often the case that residual errors are propagated by the learned policy model and accumulate throughout the trajectory, resulting in highly degraded performance. In this paper, we leverage the observation that predicted rewards under domain shift, even though imperfect, can still be a useful signal to guide fine-tuning. We exploit this property to fine-tune a policy using reward prediction in the target domain. We have found that, even under significant domain shift, the predicted reward can still provide meaningful signal and fine-tuning substantially improves the original policy. Our approach, termed Predicted Reward Fine-tuning (PRFT), improves performance across diverse tasks in both simulated benchmarks and real-world experiments. More information is available at project web page: https://sites.google.com/view/prft.
Deploying machine learning algorithms for robot tasks in real-world applications presents a core challenge: overcoming the domain gap between the training and the deployment environment. This is particularly difficult for visuomotor policies that utilize high-dimensional images as input, particularly when those images are generated via simulation. A common method to tackle this issue is through domain randomization, which aims to broaden the span of the training distribution to cover the test-time distribution. However, this approach is only effective when the domain randomization encompasses the actual shifts in the test-time distribution. We take a different approach, where we make use of a single demonstration (a prompt) to learn policy that adapts to the testing target environment. Our proposed framework, PromptAdapt, leverages the Transformer architecture's capacity to model sequential data to learn demonstration-conditioned visual policies, allowing for in-context adaptation to a target domain that is distinct from training. Our experiments in both simulation and real-world settings show that PromptAdapt is a strong domain-adapting policy that outperforms baseline methods by a large margin under a range of domain shifts, including variations in lighting, color, texture, and camera pose. Videos and more information can be viewed at project webpage: https://sites.google.com/view/promptadapt.
Purpose: Preoperative imaging plays a pivotal role in sinus surgery where CTs offer patient-specific insights of complex anatomy, enabling real-time intraoperative navigation to complement endoscopy imaging. However, surgery elicits anatomical changes not represented in the preoperative model, generating an inaccurate basis for navigation during surgery progression. Methods: We propose a first vision-based approach to update the preoperative 3D anatomical model leveraging intraoperative endoscopic video for navigated sinus surgery where relative camera poses are known. We rely on comparisons of intraoperative monocular depth estimates and preoperative depth renders to identify modified regions. The new depths are integrated in these regions through volumetric fusion in a truncated signed distance function representation to generate an intraoperative 3D model that reflects tissue manipulation. Results: We quantitatively evaluate our approach by sequentially updating models for a five-step surgical progression in an ex vivo specimen. We compute the error between correspondences from the updated model and ground-truth intraoperative CT in the region of anatomical modification. The resulting models show a decrease in error during surgical progression as opposed to increasing when no update is employed. Conclusion: Our findings suggest that preoperative 3D anatomical models can be updated using intraoperative endoscopy video in navigated sinus surgery. Future work will investigate improvements to monocular depth estimation as well as removing the need for external navigation systems. The resulting ability to continuously update the patient model may provide surgeons with a more precise understanding of the current anatomical state and paves the way toward a digital twin paradigm for sinus surgery.
A Stephen Morse合作论文数Department of Electrical Engineering, School of Engineering and Applied Science, Yale University13