In many endoscopic surgical procedures, the surgical team must identify and remove pathological tissue while avoiding critical structures such as arteries and nerves. Augmented reality (AR) offers potential support by overlaying visual information about the location of pathology and critical structures directly onto the operative field, enhancing spatial awareness and surgical navigation. However, limited research has evaluated how best to design and present AR overlays in ways that align with surgical workflow and perception. This study investigates surgeons’ preferences across three key AR overlay dimensions: Design (how anatomy is visualised: outlines, heatmaps, masks, or centroids), Trigger (how and when overlays are activated: always visible, activated by the user, or triggered by instrument position), and Placement (where the overlay appears: above or below the surgical instrument). We take endoscopic pituitary adenoma surgery as a high-risk exemplar. Using a web-based prototype, 38 neurosurgeons ranked options and provided qualitative feedback. Surgeons preferred outline designs for clarity, with a trend towards user-activated triggers for control of information flow and distraction minimisation, and below-instrument placement for better spatial awareness. Preferences were consistent across experience levels and emphasised the importance of balancing visual saliency with cognitive load, to facilitate surgical navigation without distraction or disruption. These findings inform AR interface design, but require evaluation for impact on surgical performance and safety in further physical simulation and clinical studies.
Abstract Introduction Postoperative meningioma recurrence poses a significant clinical challenge. Extent of resection (EOR) is a key predictor, and the Simpson Grading (SG) scale has historically been utilised to estimate such. No consensus exists for which of SG, postoperative radiology report (RG), and residual tumour volume (RTV) best predicts recurrence. Aims (1) Evaluate the concordance between intraoperative and postoperative EOR and (2) determine which estimate has the greatest prognostic value for predicting tumour recurrence. Methods In this retrospective review of 270 patients, intraoperative EOR was compared to EOR from postoperative imaging. Agreement was assessed using Cohen’s Kappa and absolute agreement. Machine learning models (logistic and Cox regression) were developed to compare the prognostic performance of SG, RG, and RTV for predicting recurrence. Results Agreement between surgical and radiological EOR was substantial (κ = 0.704). Machine learning models demonstrated that imaging-based metrics consistently outperformed SG. RTV performed best for 3- and 5-year prediction (AUROC = 0.787, 0.709). RTV Cox models also performed best (C-index = 0.752). No significant recurrence risk difference was found between SG 1,2, and 3. Conclusions Objective, imaging-based metrics are superior to the subjective Simpson Grade for predicting meningioma recurrence, with quantitative residual tumour volume offering the most robust prognostic value. To our knowledge, this is the first study to directly compare these EOR estimators within a machine learning framework. These findings support a shift away from a reliance on the Simpson Grade and towards the integration of quantitative radiological data for more accurate postoperative risk stratification and patient management.
Pituitary adenoma resection via the endoscopic transsphenoidal approach is technically demanding, with outcomes influenced by surgical skill. However, the association between technique and outcomes remains poorly defined. Existing workflow analyses focus on broad procedural steps and phases, but a more detailed, action-level approach is needed to capture skill-related variation. While AI shows promise in automating workflow analysis, its use at the action level is limited. This study develops and validates a reproducible action-level classification ontology for endoscopic pituitary adenoma resection, establishing the structured annotation foundation required for future AI-based workflow and skill analysis. Endoscopic videos of primary pituitary adenoma resections were collected from two high-volume international pituitary centres. A multi-disciplinary panel of neurosurgeons and data scientists iteratively reviewed and annotated surgical actions to establish a standardized classification system. Actions were categorized into triplets (instrument, target, verb), with additional temporal annotations. To evaluate framework reliability, an independent annotator followed a structured annotation guide, and inter-annotator agreement was measured using Cohen’s Kappa. A consensus-based classification ontology was developed, comprising 9 verbs, 12 instruments, and 7 targets from the review of 18 endoscopic pituitary adenoma resections (9 microadenomas, 9 macroadenomas). Action distribution differed between micro- and macroadenomas, with grasping being the predominant action in microadenomas (72
Introduction: Precise anatomical navigation is fundamental to safe endoscopic pituitary surgery, a high-stakes procedure characterised by a challenging learning curve. While traditional navigation systems often rely on workflow-disrupting probes or static preoperative imaging, advancements in computer vision AI (CVAI) now enable dynamic, real-time anatomical segmentation directly from live surgical video1-3. Our group has previously conducted a series of preclinical human-computer interaction studies to refine the system's design, alongside digital and high-fidelity physical simulations demonstrating the benefit of AI assistance in improving overall performance, training, and safety4-8. Building on this foundation, the current study represents a first-in-human application of real-time CVAI assistance in the neurosurgical operating room, serving to assess feasibility and safety, and to iteratively improve the system. Method: Guided by DECIDE-AI and IDEAL frameworks, this single-centre evaluation comprises an initial proof-of-concept phase (n=6) for endoscopic transsphenoidal pituitary surgeries. The AI model utilised a DINOv3-derived vision transformer architecture, deployed via a high-performance edge computing unit to achieve low-latency, real-time inference without reliance on cloud infrastructure2. Given the high-risk nature of the procedure and the early stage of clinical AI integration, the system was initially deployed as an educational adjunct on a secondary monitor, ensuring the primary surgical feed remains uncompromised. Functionality and safety were assessed via structured questionnaire, prospective observation, and blinded retrospective review of the recordings of the endoscopic surgical video feed and wider operating room environment. Continuous multi-stakeholder feedback through validated human factors surveys drove iterative technical refinements between cases. Results: Six patients with pituitary adenomas were enrolled. The CVAI system was successfully deployed in four cases, demonstrating acceptable real-time sella segmentation accuracy. Deployment failed pre-operatively in two cases owing to a single recurring system reboot bug. Iterative refinement between cases were driven by our experience and surgical team feedback. This resulted in the integration of additional anatomical structure segmentations (e.g., carotid arteries), enhanced model accuracy via training dataset expansion, and hardware firmware upgrades. Multi-stakeholder surveys demonstrated satisfactory system feasibility, usability, and acceptability among the surgical team. Both prospective observation and retrospective video review confirmed the absence of adverse events, including no significant distraction to the primary surgeon, and there were no AI-related clinical complications. Conclusion: This first-in-human early clinical evaluation demonstrates the feasibility, safety and iterative development of real-time, CVAI-based anatomical navigation during high-stakes neurosurgery. Future work will include a larger single-centre case series (IDEAL Stage 2a) with more surgical teams to further iterate the system and explore its impact on training and workflow. As the underpinning technology improves, deployment will transition to direct intra-operative decision support and integration with other intra-operative navigational technologies. ### Competing Interest Statement HJM is employed by and hold shares in Panda Surgical. DS holds shares in Panda Surgical, Odin Vision, and is employed by TouchSurgery, Medtronic. ### Clinical Trial NCT07568366 ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: Ethical approval has been granted (IRAS 127474 via University College London Hospitals) and the study is registered on a clinical trials database ([NCT07568366][1]). I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes Upon reasonable request. EPSRC, EP/Y01958X/1, EP/W00805X/1, UKRI145, 203145/Z/16/Z NIHR, Doctoral Fellowship, Biomedical Research Centre at University College London Google, PhD Fellowship Amethyst Radiotherapy Group, Clinical Research Fellowship [1]: /lookup/external-ref?link_type=CLINTRIALGOV&access_num=NCT07568366&atom=%2Fmedrxiv%2Fearly%2F2026%2F06%2F11%2F2026.06.11.26355205.atom
Image-guided surgery demands adaptive, real-time decision support, yet static AI models struggle with structured task planning and providing interactive guidance. Large language models (LLMs)-powered agents offer a promising solution by enabling dynamic task planning and predictive decision support. Despite recent advances, the absence of surgical agent datasets and robust parameter-efficient fine-tuning techniques limits the development of LLM agents capable of complex intraoperative reasoning. In this paper, we introduce Surgical AI Copilot, an LLM agent for image-guided pituitary surgery, capable of conversation, planning, and task execution in response to queries involving tasks such as MRI tumor segmentation, endoscope anatomy segmentation, overlaying preoperative imaging with intraoperative views, instrument tracking, and surgical visual question answering (VQA). To enable structured agent planning, we develop the PitAgent dataset, a surgical context-aware planning dataset covering surgical tasks like workflow analysis, instrument localization, anatomical segmentation, and query-based reasoning. Additionally, we propose DEFT-GaLore, a Deterministic Energy-based Fourier Transform (DEFT) gradient projection technique for efficient low-rank adaptation of recent LLMs (e.g., LLaMA 3.2, Qwen 2.5), enabling their use as surgical agent planners. We extensively validate our agent's performance and the proposed adaptation technique against other state-of-the-art low-rank adaptation methods on agent planning and prompt generation tasks, including a zero-shot surgical VQA benchmark, demonstrating the significant potential for truly efficient and scalable surgical LLM agents in real-time operative settings.
Vision-Language Models (VLMs) in visual question answering (VQA) offer a unique opportunity to enhance intra-operative decision-making, promote intuitive interactions, and significantly advance surgical education. However, the development of VLMs for surgical VQA is challenging due to limited datasets and the risk of overfitting and catastrophic forgetting during full fine-tuning of pretrained weights. While parameter-efficient techniques like Low-Rank Adaptation (LoRA) and Matrix of Rank Adaptation (MoRA) address adaptation challenges, their uniform parameter distribution overlooks the feature hierarchy in deep networks, where earlier layers, that learn general features, require more parameters than later ones. This work introduces PitVQA++ with an Open-ended PitVQA dataset and vector matrix-low-rank adaptation (Vector-MoLoRA), an innovative VLM fine-tuning approach for adapting GPT-2 to pituitary surgery. Open-Ended PitVQA comprises 109,173 frames from 25 procedural videos with 795,270 question-answer sentence pairs, covering key surgical elements such as phase and step recognition, context understanding, tool detection, localization, and interactions recognition. Vector-MoLoRA incorporates the principles of LoRA and MoRA to develop a matrix-low-rank adaptation strategy that employs rank vectors to allocate more parameters to earlier layers, gradually reducing them in the later layers. Our approach, validated on the Open-Ended PitVQA and EndoVis18-VQA datasets, effectively mitigates catastrophic forgetting while significantly enhancing performance over recent baselines. Performance-rejection analysis further highlights Vector-MoLoRA’s enhanced reliability and trust-worthiness in handling uncertain predictions. Our source code and dataset is available at https://github.com/ HRL-Mike/PitVQA-Plus.
Introduction: Vestibular schwannoma (VS) is the most common pathology in the lateral skull base, with a rising prevalence resulting in over 3,500 cases per 100,000 per year in the United States between 2004 and 2016. Management varies by tumor size; smaller tumors are typically managed with serial imaging, while larger, symptomatic tumors often require surgery. The primary goal of VS surgery is maximal tumor removal while preserving neurological function. Facial nerve palsy remains a significant concern, with large VS (> 30 mm in diameter) significantly more likely to result in facial paralysis compared to small tumors.
Endoscopic pituitary surgery is a minimally invasive technique to remove pituitary tumours through the nose. Currently, image guidance may be used in the form of a tracked pointer to help surgeons navigate the region and avoid damage to critical structures. However, the pointer method is mentally demanding as the pointer location is displayed in a different modality and disrupts the surgical workflow due to the setup time and frequent tool removal. We propose an Augmented Reality (AR) system where information from the pre-operative scan is displayed directly onto the endoscopic video. Our system features an on-board tracking system, allowing for the registration process to be performed automatically. We evaluated the accuracy of our system and compared it to an AR system that uses an infrared (IR) camera to track an endoscope with reflective markers. Our system gave an accuracy of 1.1 (± 0.4) mm, compared to 2.4 (± 0.9) mm in the IR-tracked endoscope approach. Our Augmented Reality system is a more compact and transportable setup which outperformed the IR-tracked endoscope. The automatic registration method can save time in the operating room as well as increase AR overlay accuracy, improving the translation of these technologies.
BACKGROUND:The use of simulation in neurosurgery is a widespread and popular means of training worldwide. However, little is known about patient and public acceptability of simulation in neurosurgical training and the potential consequences of this for future simulation development. METHODS:A two-stage questionnaire strategy was utilized, the first gathering insights from neurosurgical inpatients, and the second from the general public. These questionnaires assessed general understanding of the concept of simulation in neurosurgery, the relative importance of factors affecting simulation training, and acceptability of different simulation modalities and means of providing feedback to trainees. RESULTS:Seventeen inpatients responded to the first-stage survey, and 192 members of the public responded to the second-stage survey. Familiarity with the concept of simulation training in neurosurgery was generally lacking. Fidelity was established as the most important element of simulation training by the public, with cadavers and physical models the most acceptable form of simulation training. Augmented reality solutions were least popular among the public. There was enthusiasm for both artificial intelligence and telementoring as training feedback solutions. CONCLUSIONS:Patients and the public are accepting of the use of simulation training in neurosurgery. Future development should focus on improving access to high-fidelity simulation and exploring the use of artificial intelligence and telementoring in providing trainee feedback.
BACKGROUND AND OBJECTIVES: Endoscopic skull base surgery aims to reduce surgical morbidity by minimizing tissue manipulation and exposure. However, the anatomic constraints posed by the narrow surgical corridors and constrained operative workspace present technical challenges due to reduced dexterity. This study evaluates the applicability of a novel dexterity-enhancing handheld robot for endoscopic skull base approaches. METHODS: The robotic system is comprised of interchangeable articulated end-effectors coupled to a handheld controller. Two attending skull base neurosurgeons and 2 neurosurgery residents performed 8 skull base approaches on cadaveric specimens, spanning anterior, anterolateral, lateral, posterolateral, and posterior approaches. Conventional instruments were used to expose anatomic landmarks, followed by intraoperative tasks using the handheld robot. Participants were interviewed during the procedures to assess the robot's feasibility (ability to safely reach and perform its objective of manipulating tissue at the operative site) and usefulness (ability to perform desired objectives well). RESULTS: The handheld robotic system was tested across 8 endoscopic skull base approaches, achieving feasibility in all cases. Superior workspace reach compared with standard instruments was demonstrated in 6 of 8 approaches. Tissue manipulation was satisfactory in all approaches. All surgeons reported that the current or a future device prototype could be useful across all 8 approaches. The most frequently cited advantage was the expanded dextrous workspace reach provided by the articulated end-effectors, particularly in approaches with long working channels, such as the endonasal approach. However, the robot encountered difficulties in transcranial approaches (trans-sylvian and subtemporal) due to the lack of shorter, curved shafts, which impaired visualization. CONCLUSION: The handheld robotic system demonstrated applicability across various endoscopic skull base approaches, offering increased dextrous workspace and effective tissue manipulation capabilities. Overall, this study supports the potential of handheld robots in endoscopic skull base surgery while highlighting the need for iterative development to optimize instrument design and functionality.
Timely care in a specialised neuro-intensive therapy unit (ITU) reduces mortality and hospital stays. Planned admissions to ITU following surgery are safer than unplanned ones. However, post-operative care decisions remain subjective. This study used artificial intelligence (AI), specifically natural language processing (NLP) to analyse electronic health records (EHRs) of elective neurosurgery patients from University College London Hospital (UCLH) and predict ITU admissions. Using a refined CogStack-MedCAT NLP model, we extracted clinical concepts from 2268 patient records and trained AI models to classify admissions into ward and ITU. The Random Forest model achieved a recall of 0.87 (CI 0.82-0.91) for ITU admissions, reducing the proportion of unplanned ITU cases missed by human experts from 36% to 4%. Interpretability analysis confirmed the use of clinically relevant concepts. The study highlights the opportunity for AI to aid in allocating resources for neurosurgical patients but requires further research and integration into practice.
PURPOSE:Prognostication of surgical complexity is crucial for optimizing decision-making and patient counseling in pituitary surgery. This study aimed to develop a clinical score to predict gross-total resection (GTR) in non-functioning pituitary adenomas (NFPAs) using externally validated machine-learning (ML) models. METHODS:Clinical and radiological data were collected from two tertiary medical centers. Patients had pre- and postoperative structural T1-weighted MRI with gadolinium and T2-weighted preoperative scans. Three ML classifiers were trained on the National Hospital for Neurology and Neurosurgery dataset and tested on the Foundation IRCCS Ca' Granda Polyclinic of Milan dataset. Feature importance analyses and hierarchical-tree inspection identified predictors of surgical complexity, which were used to create the grading score. The prognostic performance of the proposed score was compared to that of the state-of-the art TRANSSPHER grade in the external dataset. Surgical morbidity was also analyzed. RESULTS:All ML models accurately predicted GTR, with the random forest classifier achieving the best performance (weighted-F1 score of 0.87; CIs: 0.71, 0.97). Key predictors-Knosp grade, tumor maximum diameter, consistency, and supra-sellar nodular extension-were included in the modified (m)-TRANSSPHER grade. The ROC analysis showed superior performance of the m-TRANSSPHER grade over the TRANSSPHER grade for predicting GTR in NFPAs (AUC 0.85 vs. 0.79). CONCLUSIONS:This international multi-center study used validated ML algorithms to refine predictors of surgical complexity in NFPAs, yielding the m-TRANSSPHER grade, which demonstrated enhanced prognostic accuracy for surgical complexity prediction compared to existing scales.
OBJECTIVE:To identify cognitive biases and heuristics experienced by surgeons in operative settings and the impact these biases and heuristics have on patient care. BACKGROUND:Cognitive biases and heuristics are systematic errors in thinking that can affect clinical decisions. Both are noted in surgical settings and are a risk to patient safety. METHODS:This review was conducted in accordance with Preferred Reporting Items for Systematic Reviews and Meta-Analyses guidelines and PROSPERO registered (CRD42023432099). Five major databases were searched from inception to August 28, 2022, with an updated search on January 27, 2024. Original primary research studies in English were included, with relevant risk of bias tools employed for each study. RESULTS:Twenty-one papers were included. Thirty-eight biases were identified across 6 experiments, 5 analyses, and 10 survey studies. Confirmation bias, anchoring, risk aversion, and overconfidence bias were the most represented. Risk of bias was moderate across most studies. Cognitive biases and heuristics were found to influence surgical outcomes and 6 studies cited a negative impact on patient care, with one associating biases with fatal outcomes. CONCLUSIONS:Biases and heuristics contribute to surgical errors and never events, and will continue to do so until they are recognised and addressed. Implementing debiasing strategies, such as mindfulness training and deliberate reflection, was found to reduce surgical errors in 2 studies. This review highlights the need for experimental studies, which are essential for understanding how and why biases lead to negative outcomes and for evaluating further debiasing interventions. We propose directions for future research and system changes.
BACKGROUND AND OBJECTIVES:Machine learning (ML) in surgical video analysis offers promising prospects for training and decision support in surgery. The past decade has seen key advances in ML-based operative workflow analysis, though existing applications mostly feature shorter surgeries (<2 hours) with limited scene changes. The aim of this study was to develop and evaluate a ML model capable of automated operative workflow recognition for retrosigmoid vestibular schwannoma (VS) resection. In doing so, this project furthers previous research by applying workflow prediction platforms to lengthy (median >5 hours duration), data-heavy surgeries, using VS resection as an exemplar. METHODS:A video dataset of 21 microscopic retrosigmoid VS resections was collected at a single institution over 3 years and underwent workflow annotation according to a previously agreed expert consensus (Approach, Excision, and Closure phases; and Debulking or Dissection steps within the Excision phase). Annotations were used to train a ML model consisting of a convolutional neural network and a recurrent neural network. 5-fold cross-validation was used, and performance metrics (accuracy, precision, recall, F1 score) were assessed for phase and step prediction. RESULTS:Median operative video time was 5 hours 18 minutes (IQR 3 hours 21 minutes-6 hours 1 minute). The "Tumor Excision" phase accounted for the majority of each case (median 4 hours 23 minutes), whereas "Approach and Exposure" (28 minutes) and "Closure" (17 minutes) comprised shorter phases. The ML model accurately predicted operative phases (accuracy 81%, weighted F1 0.83) and dichotomized steps (accuracy 86%, weighted F1 0.86). CONCLUSION:This study demonstrates that our ML model can accurately predict the surgical phases and intraphase steps in retrosigmoid VS resection. This demonstrates the successful application of ML in operative workflow recognition on low-volume, lengthy, data-heavy surgical videos. Despite this, there remains room for improvement in individual step classification. Future applications of ML in low-volume high-complexity operations should prioritize collaborative video sharing to overcome barriers to clinical translation.
Introduction: Timely care in a specialised neuro-intensive therapy unit (ITU) reduces mortality and hospital stays, with planned admissions being safer than unplanned ones. However, post-operative care decisions remain subjective. This study used artificial intelligence (AI), specifically natural language processing (NLP) to analyse electronic health records (EHRs) and predict ITU admissions for elective surgery patients. Methods: This study analysed the EHRs of elective neurosurgery patients from University College London Hospital (UCLH) using NLP. Patients were categorised into planned high dependency unit (HDU) or ITU admission; unplanned HDU or ITU admission; or ward / overnight recovery (ONR). The Medical Concept Annotation Tool (MedCAT) was used to identify SNOMED-CT concepts within the clinical notes. We then explored the utility of these identified concepts for a range of AI algorithms trained to predict ITU admission. Results: The CogStack-MedCAT NLP model, initially trained on hospital-wide EHRs, underwent two refinements: first with data from patients with Normal Pressure Hydrocephalus (NPH) and then with data from Vestibular Schwannoma (VS) patients, achieving a concept detection F1-score of 0.93. This refined model was then used to extract concepts from EHR notes of 2,268 eligible neurosurgical patients. We integrated the extracted concepts into AI models, including a decision tree model and a neural time-series model. Using the simpler decision tree model, we achieved a recall of 0.87 (CI 0.82 - 0.91) for ITU admissions, reducing the proportion of unplanned ITU cases missed by human experts from 36 accuracy, has proven its efficiency in extracting relevant concepts, providing a reliable basis for predictive AI models to use in clinically valid applications.
Automated detection of papilloedema using artificial intelligence (AI) and retinal images acquired through an ophthalmoscope for triage of patients with potential intracranial pathology could prove to be beneficial, particularly in resource-limited settings where access to neuroimaging may be limited. However, a comprehensive overview of the current literature on this field is lacking. We conducted a systematic review on the use of AI for papilloedema detection by searching four databases: Ovid MEDLINE, Embase, Web of Science, and IEEE Xplore. Included studies were assessed for quality of reporting using the Checklist for AI in Medical Imaging and appraised using a novel 5-domain rubric, 'SMART', for the presence of bias. For a subset of studies, we also assessed the diagnostic test accuracy using the 'Metadta' command on Stata. Nineteen deep learning systems and eight non-deep learning systems were included. The median number of images of normal optic discs used in the training set was 2509 (IQR 580-9156) and in the testing set was 569 (IQR 119-1378). The number of papilloedema images in the training and testing sets was lower with a median of 1292 (IQR 201-2882) in training set and 201 (IQR 57-388) in the testing set. Age and gender were the two most frequently reported demographic data, included by one-third of the studies. Only ten studies performed external validation. The pooled sensitivity and specificity were calculated to be 0.87 [95% CI 0.76-0.93] and 0.90 [95% CI 0.74-0.97], respectively. Though AI model performance values are reported to be high, these results need to be interpreted with caution due highly biased data selection, poor quality of reporting, and limited evidence of reproducibility. Deep learning models show promise in retinal image analysis of papilloedema, however, external validation using large, diverse datasets in a variety of clinical settings is required before it can be considered a tool for triage of intracranial pathologies in resource-limited areas.
Image-guided surgery demands adaptive, real-time decision support, yet static AI models struggle with structured task planning and providing interactive guidance. Large language models (LLMs)-powered agents offer a promising solution by enabling dynamic task planning and predictive decision support. Despite recent advances, the absence of surgical agent datasets and robust parameter-efficient fine-tuning techniques limits the development of LLM agents capable of complex intraoperative reasoning. In this paper, we introduce Surgical AI Copilot, an LLM agent for image-guided pituitary surgery, capable of conversation, planning, and task execution in response to queries involving tasks such as MRI tumor segmentation, endoscope anatomy segmentation, overlaying preoperative imaging with intraoperative views, instrument tracking, and surgical visual question answering (VQA). To enable structured agent planning, we develop the PitAgent dataset, a surgical context-aware planning dataset covering surgical tasks like workflow analysis, instrument localization, anatomical segmentation, and query-based reasoning. Additionally, we propose DEFT-GaLore, a Deterministic Energy-based Fourier Transform (DEFT) gradient projection technique for efficient low-rank adaptation of recent LLMs (e.g., LLaMA 3.2, Qwen 2.5), enabling their use as surgical agent planners. We extensively validate our agent's performance and the proposed adaptation technique against other state-of-the-art low-rank adaptation methods on agent planning and prompt generation tasks, including a zero-shot surgical VQA benchmark, demonstrating the significant potential for truly efficient and scalable surgical LLM agents in real-time operative settings.
BACKGROUND:Identifying patients eligible for clinical trials through eligibility screening is time and resource-intensive. Natural Language Processing (NLP) models may enhance clinical trial screening by extracting data from Electronic Health Records (EHRs). OBJECTIVE:We aimed to determine whether an NLP model can extract brain tumor diagnoses from outpatient clinic letters and link this with ongoing clinical trials. METHODS:This retrospective cohort study reviewed outpatient neuro-oncology clinic letters, to detect brain tumor diagnoses. We used an NLP model to perform a Named Entity Recognition + Linking algorithm that identified medical concepts in free text and linked them to a Systematized Nomenclature of Medicine Clinical Terms ontology, which we used to search a clinical trials database. Human annotators reviewed the accuracy of the concepts extracted and the relevance of recommended clinical trials. Search results were shown on a notification dashboard accessible by clinicians and patients on the EHR. We report the model's performance using precision, recall, and F1 scores. RESULTS:The model recognized 399 concepts across 196 letters with macro-precision = 0.994, macro-recall = 0.964, and macro-F1 = 0.977. Linking the model results with a clinical trials database identified 1417 ongoing clinical trials; of these, 755 were highly relevant to the individual patient, who met the eligibility criteria for trial recruitment. CONCLUSIONS:NLP can be used effectively to extract brain tumor diagnoses from free-text EHR records with minimal additional training. The extracted concepts can then be linked to ongoing clinical trials. While further analysis is required to assess the impact on clinical outcomes, these findings suggest a potential application for integrating NLP algorithms into clinical care.