Purpose: Robotic-assisted surgery (RAS) generates vast amounts of video and robotic data, presenting opportunities for machine learning. Video-based models, in particular, that can temporally segment frames by ontological categories such as procedure type, phase, steps, actions, etc., are needed. Training separate models for each category neglects statistical dependencies between categories and can yield incompatible predictions. Training large multi-category models may help, but increases complexity while reducing model modularity and interpretability. Methods: We present a model fusion alternative: an effectively zero-free-parameter Bayesian model fusion technique. Incorporating the empirical conditional dependencies across categories and time, we combine predictions from multiple segmentation models into one joint Bayesian inference. The result is a Bayes' optimal distribution over all categories evolving over time with accumulated evidence. Results: On a large test set of hundreds of surgical cases, of nearly eight million frames of annotated data, we found that fused predictions from the joint Bayesian model provide clear benefits over the individual models, correcting inconsistent and inaccurate predictions, and even forming accurate beliefs when evidence was absent. Conclusion: The model we present is a lightweight, principled alternative to machine learning-based model fusion. A sufficiently complex model could be trained to produce the same results, but would effectively trade explainable predictions with minimal overheard for computational complexity and transparency. We end by discussing how the same approach can be used to encompass larger more sophisticated models within the same conceptual framework.
Objective Performance Indicators (OPIs) quantify surgical activity by measuring critical aspects of a case (e.g., surgeon movements, interaction forces, etc.) at clinically meaningful steps. They are particularly useful in robotic-assisted surgery, where the information required to compute them can be harvested automatically, and early studies have found that they help quantify surgeon skill. However, since surgical procedures can be highly variable and the boundaries between surgical steps often subject to interpretation, OPIs become similarly variable. Yet, OPIs should be insensitive to this variability to remain informative and unambiguous. Previously we proposed a method for computing expected OPI values in a probabilistic framework that has theoretical support for being robust to this variability. Here we validate this method and compare it against the conventional approach on a large clinical dataset. Using 5408 annotated clinical steps from 1,016 clinical cases across seven procedure types, we performed an extensive comparison of this new method alongside the conventional approach. A series of analyses compared the resulting OPI values focusing on potential biases and the variability of their values in the presence of annotator noise. We found that the probabilistic method resulted in OPI values that were less often biased than their conventional OPI counterparts (49.37%, and 56.63% respectively), and the bias magnitudes were smaller. Finally, we found that the probabilistic method resulted in significantly less variable OPI values in 40.56% of comparison, whereas 8.52% showed the opposite. Our analysis suggests that the expected OPIs are more robust to annotator variability, with less bias and smaller variance. We suggest that this probabilistic approach holds promise for the computational analysis, and perhaps the conceptual interpretation, of surgical data.
Owing to recent advances in machine learning and the ability to harvest large amounts of data during robotic-assisted surgeries, surgical data science is ripe for foundational work. We present a large dataset of surgical videos and their accompanying labels for this purpose. We describe how the data was collected and some of its unique attributes. Multiple example problems are outlined. Although the dataset was curated for a particular set of scientific challenges (in an accompanying paper), it is general enough to be used for a broad range machine learning questions. Our hope is that this dataset exposes the larger machine learning community to the challenging problems within surgical data science, and becomes a touch-stone for future research. The videos are available at https://storage.googleapis.com/isi-surgvu/surgvu24_videos_only.zip, the labels at https://storage.googleapis.com/isi-surgvu/surgvu24_labels_updated_v2.zip, a validation set for tool detection problem at https://storage.googleapis.com/isi-surgvu/cat1_test_set_public.zip, and a sample set of question answer pairs dataset for surgical visual question answering at https://storage.googleapis.com/isi-surgvu/SURGVU25_cat_2_sample_set_public.zip.
The ability to automatically detect and track surgical instruments in endoscopic videos can enable transformational interventions. Assessing surgical performance and efficiency, identifying skilled tool use and choreography, and planning operational and logistical aspects of OR resources are just a few of the applications that could benefit. Unfortunately, obtaining the annotations needed to train machine learning models to identify and localize surgical tools is a difficult task. Annotating bounding boxes frame-by-frame is tedious and time-consuming, yet large amounts of data with a wide variety of surgical tools and surgeries must be captured for robust training. Moreover, ongoing annotator training is needed to stay up to date with surgical instrument innovation. In robotic-assisted surgery, however, potentially informative data like timestamps of instrument installation and removal can be programmatically harvested. The ability to rely on tool installation data alone would significantly reduce the workload to train robust tool-tracking models. With this motivation in mind we invited the surgical data science community to participate in the challenge, SurgToolLoc 2022. The goal was to leverage tool presence data as weak labels for machine learning models trained to detect tools and localize them in video frames with bounding boxes. We present the results of this challenge along with many of the team's efforts. We conclude by discussing these results in the broader context of machine learning and surgical data science. The training data used for this challenge consisting of 24,695 video clips with tool presence labels is also being released publicly and can be accessed at https://console.cloud.google.com/storage/browser/isi-surgtoolloc-2022.
When learning new movements some people make larger kinematic errors than others, interpreted as a reduction in motor-learning ability. Consider a learning task where error-cancelling strategies incur higher effort costs, specifically where subjects reach to targets in a force field. Concluding that those with greater error have learned less has a critical assumption: everyone uses the same error-canceling strategy. Alternatively, it could be that those with greater error may be choosing to sacrifice error reduction in favor of a lower effort movement. Here, we test this hypothesis in a dataset that includes both younger and older adults, where older adults exhibited greater kinematic errors. Utilizing the framework of optimal control theory, we infer subjective costs (i.e., strategies) and internal model accuracy (i.e., proportion of the novel dynamics learned) by fitting a model to each population's trajectory data. Our results demonstrate trajectories are defined by a combination of the amount learned and strategic differences represented by relative cost weights. Based on the model fits, younger adults could have learned between 65-90% of the novel dynamics. Critically, older adults could have learned between 60-85%. Each model fit produces trajectories that match the experimentally observed data, where a lower proportion learned in the model is compensated for by increasing costs on kinematic errors relative to effort. This suggests older and younger adults could be learning to the same extent, but older adults have a higher relative cost on effort compared to younger adults. These results call into question the proposition that older adults learn less than younger adults and provide a potential explanation for the equivocal findings in the literature. Importantly, our findings suggest that the metrics commonly used to probe motor learning paint an incomplete picture, and that to accurately quantify the learning process the subjective costs of movements should be considered.
Formalizing surgical activities as triplets of the used instruments, actions performed, and target anatomies is becoming a gold standard approach for surgical activity modeling. The benefit is that this formalization helps to obtain a more detailed understanding of tool-tissue interaction which can be used to develop better Artificial Intelligence assistance for image-guided surgery. Earlier efforts and the CholecTriplet challenge introduced in 2021 have put together techniques aimed at recognizing these triplets from surgical footage. Estimating also the spatial locations of the triplets would offer a more precise intraoperative context-aware decision support for computer-assisted intervention. This paper presents the CholecTriplet2022 challenge, which extends surgical action triplet modeling from recognition to detection. It includes weakly-supervised bounding box localization of every visible surgical instrument (or tool), as the key actors, and the modeling of each tool-activity in the form of ‹instrument, verb, target› triplet. The paper describes a baseline method and 10 new deep learning algorithms presented at the challenge to solve the task. It also provides thorough methodological comparisons of the methods, an in-depth analysis of the obtained results across multiple metrics, visual and procedural challenges; their significance, and useful insights for future research directions and applications in surgery.
Timely and effective feedback within surgical training plays a critical role in developing the skills required to perform safe and efficient surgery. Feedback from expert surgeons, while especially valuable in this regard, is challenging to acquire due to their typically busy schedules, and may be subject to biases. Formal assessment procedures like OSATS and GEARS attempt to provide objective measures of skill, but remain time-consuming. With advances in machine learning there is an opportunity for fast and objective automated feedback on technical skills. The SimSurgSkill 2021 challenge (hosted as a sub-challenge of EndoVis at MICCAI 2021) aimed to promote and foster work in this endeavor. Using virtual reality (VR) surgical tasks, competitors were tasked with localizing instruments and predicting surgical skill. Here we summarize the winning approaches and how they performed. Using this publicly available dataset and results as a springboard, future work may enable more efficient training of surgeons with advances in surgical data science. The dataset can be accessed from https://console.cloud.google.com/storage/browser/isi-simsurgskill-2021.
Reaches in experimental settings are commonly found to be straight. This straightness is robust to physical, but not visual, perturbations. Here, we question whether typical visual feedback contributes to this finding by implicitly promoting straight movements. To do so, we replaced the conventional feedback depicting the hand's location with feedback depicting the limb's orientation. Reaching movements with three different visual feedback conditions were examined. In the final condition, the subject's arm was depicted as two rotating links, and targets were depicted as two links indicating a desired arm posture. We found that by replacing standard cursor feedback, reaches became curved and arched to the target. Our findings further demonstrate that depicted feedback influences movements, and feedback depicting the limb, in particular, may elicit curved reaches.
Identifying and quantifying the activities that compose surgery is essential for effective interventions, computer-aided analyses and the advancement of surgical data science. For example, recent studies have shown that objective metrics (referred to as objective performance indicators, OPIs) computed during key surgical tasks correlate with surgeon skill and clinical outcomes. Unambiguous identification of these surgical tasks can be particularly challenging for both human annotators and algorithms. Each surgical procedure has multiple approaches, each surgeon has their own level of skill, and the initiation and termination of surgical tasks can be subject to interpretation. As such, human annotators and machine learning models face the same basic problem, accurately identifying the boundaries of surgical tasks despite variable and unstructured information. For use in surgeon feedback, OPIs should also be robust to the variability and diversity in this data. To mitigate this difficulty, we propose a probabilistic approach to surgical task identification and calculation of OPIs. Rather than relying on tasks that are identified by hard temporal boundaries, we demonstrate an approach that relies on distributions of start and stop times, for a probabilistic interpretation of when the task was performed. We first use hypothetical data to outline how this approach is superior to other conventional approaches. Then we present similar analyses on surgical data. We find that when surgical tasks are identified by their individual probabilities, the resulting OPIs are less sensitive to noise in the identification of the start and stop times. These results suggest that this probabilistic approach holds promise for the future of surgical data science.
When learning a new motor behavior, e.g. reaching in a force field, the nervous system builds an internal representation. Examining how subsequent reaches in unpracticed directions generalize reveals this representation. Although often studied, it is not known how this representation changes across training directions, or how changes in reach direction and the corresponding changes in limb impedance, influence these measurements. We ran a force field adaptation experiment using eight groups of subjects each trained on one of eight standard directions and then tested for generalization in the remaining seven directions. Generalization in all directions was local and asymmetric, providing limited and unequal transfer to the left and right side of the trained target. These asymmetries were not consistent in either magnitude or direction, even after correcting for changes in limb impedance. Relying on a standard model for generalization the inferred representations inconsistently shifted to one side or the other of their respective training direction. A second model that accounted for limb impedance and variations in baseline trajectories explained more data and the inferred representations were centered on their respective training directions. Our results highlight the influence of limb mechanics and impedance on psychophysical measurements and their interpretations for motor learning.
While we can readily observe and model the dynamics of our limbs, analyzing the neurons that drive movement is not nearly as straightforward. As a result, their role in motor behavior (e.g., forward models, state estimators, controllers, etc.) remains elusive. Computational explanations of electrophysiological data often rely on firing rate models or deterministic spiking models. Yet neither can accurately describe the interactions of neurons that issue spikes, probabilistically. Here we take a normative approach by designing a probabilistic spiking network to implement LQR control for a limb model. We find typical results: cosine tuning curves, population vectors that correlate with reaching directions, low-dimensional oscillatory activity for reaches that have no oscillatory movement, and changes in neuron's tuning curves after force field adaptation. Importantly, while the model is consistent with these empirically derived correlations, we can also analyze it in terms of the known causal mechanism: an LQR controller and the probability distributions of the neurons that encode it. Redesigning the system under a different set of assumptions (e.g. a different controller, or network architecture) would yield a new set of testable predictions. We suggest this normative approach can be a framework for examining the motor system, providing testable links between observed neural activity and motor behavior.
Behavioral studies consistently find that subjects move their hand along straight paths despite considerations that suggest reaches should be curved. Literature on this topic makes it clear that the experimentally displayed feedback influences how subjects reach. Could the standard visual feedback, a displayed cursor, explain the lack of path curvature in experimental results? To address this question, we conducted three experiments to examine reach behavior in the absence of the standard visual feedback. In the first experiment, we found significant increases in curvature as visual feedback was progressively extinguished across groups. A second experiment revealed that practiced reaches became curved after the standard visual feedback was removed. A final experiment found that subjects' reaches made before and after a brief display of visual feedback were similar, indicating a preference for specific curved trajectories. Our results suggest that the consistently straight reaches often observed could be due to a bias to move the displayed cursor straight, which when removed reveal subject-specific preferences for reaches that are often curved.
Subjects in laboratory settings exhibit straight hand paths—typified by the minimum jerk path—even in the presence of a learned but disturbing force field. At the same time it is known that in this setting, visual feedback strongly influences reaches, biasing them to be straight. Here we examine whether or not this bias can account for the straightness of movements made in a force field. We ran three curl field experiments to investigate how the lack of visual feedback influences adapted reaches. In a first experiment, hand position was displayed at the beginning and at the end of each trial, but extinguished during movement, and the hand was passively brought back to the home location. In the second experiment, visual feedback of neither the hand nor the target was provided, and targets were haptically rendered as “dimples.” In order to provide extended practice, a third experiment was run with a single target and an active reach back to the home location. In all three cases we found minor changes in the adapted reaches relative to control groups that had full visual feedback. Our subjects adopted trajectories that were better explained by minimum jerk paths over those that minimize effort. The results indicate that for point-to-point reaching movements the visual feedback, or lack there of, cannot explain why reaches appear to be straight, even after adapting to a perturbing force field.
Background: We control the movements of our body and limbs through our muscles. However, the forces produced by our muscles depend unpredictably on the commands sent to them. This uncertainty has two sources: irreducible noise in the motor system's processes (i.e., motor noise) and variability in the relationship between muscle commands and muscle outputs (i.e., model uncertainty). Any controller, neural or artificial, benefits from estimating these uncertainties when choosing commands. Methods: To examine these benefits, we used an experimental preparation of the rat hindlimb to electrically stimulate muscles and measure the resulting isometric forces. We compare a functional electric stimulation (FES) controller that represents and compensates for uncertainty in muscle forces with a standard FES controller that neglects uncertainty. Results: Accounting for uncertainty substantially increased the precision of force control. Conclusion: Our study demonstrates the theoretical and practical benefits of representing muscle uncertainty when computing muscle commands. Significance: The findings are relevant beyond FES as they highlight the benefits of estimating statistical properties of muscles for both artificial controllers and the nervous system.
How to move efficiently is an optimal control problem, whose computational complexity grows exponentially with the horizon of the planned trajectory. Breaking a compound movement into a series of chunks, each planned over a shorter horizon can thus reduce the overall computational complexity and associated costs while limiting the achievable efficiency. This trade-off suggests a cost-effective learning strategy: to learn new movements we should start with many short chunks (to limit the cost of computation). As practice reduces the impediments to more complex computation, the chunking structure should evolve to allow progressively more efficient movements (to maximize efficiency). Here we show that monkeys learning a reaching sequence over an extended period of time adopt this strategy by performing movements that can be described as locally optimal trajectories. Chunking can thus be understood as a cost-effective strategy for producing and learning efficient movements.
The two-alternative forced-choice (2AFC) task is the workhorse of psychophysics and is used to measure the just-noticeable difference, generally assumed to accurately quantify sensory precision. However, this assumption is not true for all mechanisms of decision making. Here we derive the behavioral predictions for two popular mechanisms, sampling and maximum a posteriori, and examine how they affect the outcome of the 2AFC task. These predictions are used in a combined visual 2AFC and estimation experiment. Our results strongly suggest that subjects use a maximum a posteriori mechanism. Further, our derivations and experimental paradigm establish the already standard 2AFC task as a behavioral tool for measuring how humans make decisions under uncertainty.
The motor system generates time-varying commands to move our limbs and body. Conventional descriptions of motor control and learning rely on dynamical representations of our body's state (forward and inverse models), and control policies that must be integrated forward to generate feedforward time-varying commands; thus these are representations across space, but not time. Here we examine a new approach that directly represents both time-varying commands and the resulting state trajectories with a function; a representation across space and time. Since the output of this function includes time, it necessarily requires more parameters than a typical dynamical model. To avoid the problems of local minima these extra parameters introduce, we exploit recent advances in machine learning to build our function using a stacked autoencoder, or deep network. With initial and target states as inputs, this deep network can be trained to output an accurate temporal profile of the optimal command and state trajectory for a point-to-point reach of a nonlinear limb model, even when influenced by varying force fields. In a manner that mirrors motor babble, the network can also teach itself to learn through trial and error. Lastly, we demonstrate how this network can learn to optimize a cost objective. This functional approach to motor control is a sharp departure from the standard dynamical approach, and may offer new insights into the neural implementation of motor control.