To understand the value and experience of using the da Vinci 5 (dV-5) robotic surgical system among early-adopting surgeons in the United States. In March 2024, da Vinci 5 was released with new features such as a haptic technology (called Force Feedback) and Case Insights, a tool leveraging artificial intelligence (AI) to deliver video recordings of cases with objective metrics of performance. Few studies have assessed the value and challenges experienced by surgeons using this new system. Twenty-three semi-structured qualitative interviews were completed with surgeon-participants over video conferencing software representing a selection of surgical specialties, case volumes, and practice types. Interviews were recorded, transcribed verbatim, and deidentified. Results were analyzed by one reviewer using an inductive-deductive thematic approach and further verified by another reviewer. Themes were mapped to value domains and challenges with adoption. Among the participants, there were a higher proportion of males, surgeons who practiced at community hospitals, and those with medium to high volumes of robotic cases. Thematic analysis revealed two main themes with seven subthemes exploring either the value beliefs or barriers/challenges with adoption to the new system. Participants found the most value in dV-5 with its ergonomic comfort, ability to support future training of surgeons, and an economic benefit in reducing operative time. There were mixed findings around its impact for improving clinical outcomes given the early system maturity. Thematic analysis of interviews with early-adopting surgeons indicated that the dV5 system supports better ergonomics, surgeon training and may have implications for some key surgical metrics such as tissue tearing and operative time.
Vision-language pre-training (VLP) offers unique advantages for surgery by aligning language with surgical videos, enabling workflow understanding and transfer across tasks without relying on expert-labeled datasets. However, progress in surgical VLP remains constrained by the limited scale, procedural diversity, semantic quality, and hierarchical structure of existing datasets. In this work, we present SurgLaVi, the largest and most diverse surgical vision-language dataset to date, comprising nearly 240k clip-caption pairs from more than 200 procedures, and featuring hierarchical levels at coarse-, mid-, and fine-level. At the core of SurgLaVi lies a fully automated pipeline that systematically generates fine-grained transcriptions of surgical videos and segments them into coherent procedural units. To ensure high-quality annotations, it applies dual-modality filtering to remove irrelevant and noisy samples. Within this framework, the resulting captions are enriched with contextual detail, producing annotations that are both semantically rich and easy to interpret. To ensure accessibility, we release SurgLaVi-β, an open-source derivative of 113k clip-caption pairs constructed entirely from public data, which is over four times larger than existing surgical VLP datasets. To demonstrate the value of the SurgLaVi datasets, we introduce SurgCLIP, a CLIP-style video-text contrastive framework with dual encoders, as a representative base model. SurgCLIP achieves consistent improvements across phase, step, action, and tool recognition, surpassing prior state-of-the-art methods, often by large margins. These results validate that large-scale, semantically rich, and hierarchically structured datasets directly translate into stronger and more generalizable representations, establishing SurgLaVi as a key resource for developing surgical foundation models.
The Parkland Grading Scale (PGS) is widely used to quantify operative difficulty in cholecystectomy, with higher grades associated with worse post-operative outcomes. However, consistent, scalable PGS assessment is limited by the reliance on two manual steps: determining where to look in the surgical video for key evidence, and assigning a grade. Previous machine learning approaches have either depended on manual selection of where to look, or approximated it with fixed-duration video segments, leaving it unclear whether models can accurately predict PGS without explicit guidance on where to look. To address this, we evaluate 287 robotic cholecystectomy videos annotated with PGS and a standardized key-segment. Using a temporal convolution network and attention-based framework, we compare the performance of a fully automated model using full surgical videos without key-segment supervision to a model provided with the key-segment (where to look). Providing the key-segment yields substantial performance gains (weighted F1 +0.25 and Krippendorff’s α (KA) +0.29). We further introduce ParkNet _LEARN , which learns to where to look and predicts PGS from full surgical videos, achieving significant improvements over the no-supervision automation (weighted F1 +0.18 and KA +0.23), and a KA = 0.60–within 0.06 of the model with key-segment provided. These findings highlight the importance of attending to where to look for automating operative difficulty assessment, and is a valuable step toward supporting large-scale research on surgical performance and post-operative outcomes.
Surgeons don't just see – they interpret. When an expert observes a surgical scene, they understand not only what instrument is being used, but why it was chosen, what risk it poses, and what comes next. Current surgical AI cannot answer such questions, largely because training data that explicitly encodes surgical reasoning is immensely difficult to annotate at scale. Yet surgical video lectures already contain exactly this – explanations of intent, rationale, and anticipation, narrated by experts for the purpose of teaching. Though inherently noisy and unstructured, these narrations encode the reasoning that surgical AI currently lacks. We introduce SUREON, a large-scale video QA dataset that systematically harvests this training signal from surgical academic videos. SUREON defines 12 question categories covering safety assessment, decision rationale, and forecasting, and uses a multi-agent pipeline to extract and structure supervision at scale. Across 134.7K clips and 170 procedure types, SUREON yields 206.8k QA pairs and an expert-validated benchmark of 354 examples. To evaluate the extent to which this supervision translates to surgical reasoning ability, we introduce two models: SureonVLM, a vision-language model adapted through supervised fine-tuning, and SureonVLM-R1, a reasoning model trained with Group Relative Policy Optimization. Both models can answer complex questions about surgery and substantially outperform larger general-domain models, exceeding 84
Objective performance indicators (OPIs) derived from robotic surgery are showing potential for automated skill assessment, but their high dimensionality, data sparsity, and lack of functional context limit their clinical utility and interpretability. Here, we introduce and validate a semantic taxonomy that automatically classifies surgical instruments into functional roles—such as ‘Dominant’, ‘Active Retractor’, and ‘Passive Retractor’—based on their kinematic signatures. Applied to 462 cholecystectomies, hernia repairs, and sleeve gastrectomies, this framework drastically reduced data dimensionality. In predictive modeling for surgical experience and task efficiency, taxonomy-structured OPIs achieved superior performance to conventional metrics while requiring substantially fewer features to reach optimal results (mean, 12.4 vs. 19.5; P = 0.025). By providing functional context, this approach streamlines kinematic analysis, creating a more scalable and interpretable foundation for objective skill assessment, actionable feedback, and data-driven surgical training, ultimately enhancing surgical quality and safety.