Objective: To determine whether large language models (LLMs) can automatically extract organ-level disease involvement to populate the Surgical Findings section of the European Society of Gynaecological Oncology (ESGO) Operative Report for advanced ovarian cancer. Methods: We retrospectively collected 300 operative notes from cytoreductive surgeries performed at a tertiary ESGO-accredited center. Each note was interrogated to identify disease involvement across 35 pre-defined ESGO anatomical sites. For each site, LLMs were tasked with classifying whether disease was present. Their accuracy was compared with expert annotations using F1 scores. Four modern models were selected based on their state-of-the-art performance and suitability for clinical text interpretation. Operative notes were converted into sets of binary (yes/no) questions corresponding to each anatomical site. Models were tested both in their basic form and after targeted enhancement strategies to reduce common errors. These enhancements included adding a clinical terminology list, providing clearer task instructions, and showing a small number of examples. Results: The models showed good baseline accuracy, with the two top-performing systems achieving F1 scores of 0.851 (95% confidence interval [CI] 0.841 to 0.861) and 0.864 (95% CI 0.854 to 0.873). Following optimization strategies, accuracy increased further, reaching 0.897 (95% CI 0.888 to 0.906) and 0.875 (95% CI 0.866 to 0.884). Performance was highest for key clinical sites, including the omentum, right diaphragm (95%), and ovaries (92%). Lower accuracy was observed for complex anatomical sites such as bowel (small bowel 73%, large bowel 61%) and peritoneal sites (pouch of Douglas 82%, abdominal wall 68%). Frequent errors involved laterality, overlapping anatomical regions, and ambiguous abbreviations. Optimization strategies improved distinction between closely related sites (rectosigmoid vs large bowel/mesentery) and reduced left/right errors. Conclusions: With enhancement strategies, LLMs demonstrated near-human performance in extracting ESGO-compliant operative information. Integrating model-assisted extraction into surgical workflows may reduce reporting time, improve completeness, and help standardize operative documentation.
Background: The release of Large Language Models (LLMs) has introduced numerous benefits across the healthcare domain. This study evaluated the responses of 11 LLMs from the Claude, Mistral, Llama, and GPT families to Frequently Asked Questions (FAQs) regarding ovarian cancer with regards to three domains: (a) ease of understanding, (b) accuracy, and (c) empathy. Methods: Fifteen FAQs were sourced from the Ovarian Cancer Action (OCA) website comprising (a) anticipated questions and (b) actual questions. Responses from each of the 11 LLMs were blinded and then evaluated by three Gynaecological Oncology Surgical Fellows using a 5-point Likert scale. Inter-observer agreement was calculated for each response, and LLMs were compared across the three domains using Friedman’s test (p<0.05). Finally, all LLM responses were compared with the ones from the OCA website using the same evaluation criteria. Results: Varying levels of inter-observer agreement were observed. Claude 3 Opus produced the easiest-to-understand answers (average score 4.38), followed by Mistral Large (4.36) and GPT-4o (4.33). GPT-4o scored highest for accuracy (average score 4.24) and showed strongest performance in empathy (average of 3.87). Compared with the OCA responses, GPT-4o outperformed all models in accuracy (4.24) and empathy (3.87), with 50% of its responses rated more accurate and 70% more empathetic than the OCA content. Claude 3 Opus, Mistral Large, and Mixtral 8x7B surpassed OCA in clarity for one-third of responses, while Claude 3 Sonnet achieved the highest readability gains (40%). Conclusion: The study informs the development of LLMs suitable for patient-facing ovarian cancer communication with Claude 3 Opus and GPT-4o excelling in different metrics. While improvements in emotional intelligence remain necessary, our findings pave the way for developing a specialized LLM for ovarian cancer, using domain-specific text to provide comprehensive and empathetic information.
Background/Objectives: Malnutrition is common among women undergoing gynaecologic oncology (GO) surgery and is associated with increased morbidity, prolonged hospitalisation, and reduced survival. Nevertheless, the optimal nutritional screening tools remain uncertain. Methods: We conducted a narrative review of commonly used nutritional screening and assessment tools in surgical GO patients. To highlight practical challenges in accurately identifying at risk individuals, we incorporated findings from our clinical audit. Results: There was a considerable variation between tools. While many tools were associated with adverse outcomes, their clinical value in this population was unclear. The presence of ascites and rapid deterioration in oral intake may contribute to under-recognition of at-risk patients, as illustrated by our audit findings. Emerging strategies including determining body composition from routine pre-operative Computed Tomography (CT) scans, which have shown statistical associations with survival and toxicity in observational studies, but their clinical utility is not yet established. Conclusions: Although several screening tools were statistically associated with adverse outcomes, robust data on their clinical utility are lacking. Current tools may inadequately factor for the specific considerations when screening this group. Consequently, nutritionally vulnerable surgical GO patients requiring nutritional intervention may be missed. As no gold standard currently exists for this population, bespoke, objective approaches and prospective studies are urgently needed to address disease-specific nutritional considerations.
Background/Objectives: Nutritional risk screening is critical in the management of gynaecologic oncology (GO) surgical patients. Malnutrition is prevalent in this population and is associated with poorer surgical outcomes, including increased morbidity, prolonged hospital stays, and reduced survival rates. Nevertheless, the optimal nutritional screening tools for this patient group remain undefined. Methods: We conducted a narrative review to critically appraise commonly used nutritional screening and assessment tools in surgical GO patients. To highlight practical challenges in accurately identifying at-risk individuals, we incorporated findings from our recent clinical audit. Results: Several nutritional screening and assessment tools were identified. The results varied considerably between tools. The presence of ascites and rapid deterioration in oral intake were frequently overlooked, leading to under-recognition of malnutrition. These issues were corroborated by our audit findings. Emerging strategies including determining body composition from routine preoperative CT scans show promise. Conclusions: Accurate nutritional assessment is imperative to improve surgical outcomes in surgical GO patients. As currently no gold standard currently exists for this population, bespoke approaches to address disease-specific nutritional considerations are urgently needed to identify those at risk and allow for timely nutritional interventions. Integrating CT-based body composition analysis can provide an objective solution, thus requiring further investigation.
Background/Objectives: The incidence of acute kidney injury (AKI) following advanced epithelial ovarian cancer (EOC) surgery has not been extensively studied. This study aimed to investigate the incidence of AKI and identify preoperative and intraoperative predictors in patients undergoing advanced EOC cytoreduction using both traditional statistics and Artificial Intelligence (AI) modelling. Methods: Retrospective data were collected for 134 patients with a suspected or confirmed diagnosis of advanced EOC (FIGO Stage III–IV) who underwent surgical cytoreduction between January 2021 and December 2022 at a UK tertiary referral centre. AKI was diagnosed according to the KDIGO criteria. Data on 22 patient variables were extracted, including age, Charlson Comorbidity Index (CCI), procedure length, surgical complexity, and length of hospital stay. Logistic regression analysis was used for feature selection to identify AKI predictors, and an extreme gradient boost (XGBoost) model was applied to all variables related to AKI events. Results: The incidence of postoperative AKI was 6.72% (n=9). Predictive factors for AKI included younger age (OR = 0.942, p=0.037), lower CCI (OR = 0.415, p=0.015), longer procedure duration (OR = 1.006, p=0.019), and greater surgical effort (OR = 1.427, p=0.007). Patients with perioperative AKI experienced a doubling in the length of hospital stay (p=0.008). Mortality rates were similar between patients with and without AKI. AI-driven algorithms highlighted the complexity of AKI prediction and provided individual risk profiles, enabling future stratification and prompting different frequencies of AKI monitoring following cytoreduction. Conclusions: Predicting AKI is a complex task. This study found a lower-than-expected incidence of AKI following advanced EOC cytoreductive surgery. AKI is linked to heightened surgical risk-taking, underscoring the need for improved guidelines focusing on postoperative monitoring for targeted patients. Artificial Intelligence offers the potential for personalized AKI prediction.
BACKGROUND:The kynurenine pathway is a key immunosuppressive mechanism implicated in resistance to immune checkpoint inhibitors (ICIs). This study investigated expression of tryptophan-metabolising enzymes (IDO1, TDO2, IL4I1) and their relationship with the immune microenvironment across molecular subtypes of endometrial cancer (EC). METHODS:A cohort of 570 ECs was classified as mismatch repair-deficient (MMRd), p53-mutant (p53mut), or no specific molecular profile (NSMP). Expression of IDO1, TDO2, IL4I1, PD-L1, kynurenine, and immune markers (CD8, FOXP3, CD68, CD163) was assessed by immunohistochemistry and quantified in tumour and stromal compartments using QuPath. Associations with disease-specific survival (DSS) were analysed using correlation testing, Kaplan-Meier, and Cox regression. RESULTS:IDO1, TDO2, and IL4I1 correlated strongly with immune infiltrate density. High IDO1 tumour expression was linked to improved DSS in NSMP and p53mut tumours (p < 0.05), consistent with an "inflamed" phenotype. In contrast, high TDO2 stromal expression predicted reduced DSS in NSMP patients (p < 0.05). Elevated IL4I1 tumour expression was associated with improved DSS in MMRd tumours (p < 0.05). A high CD163:CD8 ratio independently predicted worse DSS in p53mut tumours (p < 0.05). Both TDO2 and IL4I1 were highly expressed high-risk tumours, particularly p53mut cases. CONCLUSIONS:Tryptophan-kynurenine enzymes shape the immune landscape of EC in a subtype-specific manner. High IDO1 was linked to favourable outcomes in p53mut and NSMP cases, whereas TDO2 predicted poor prognosis. The CD163:CD8 ratio emerged as an independent marker of poor survival. These findings support therapeutic strategies combining dual IDO1/TDO2 inhibition or targeting the IL4I1- aryl hydrocarbon receptor (AhR) axis to enhance immunotherapy efficacy in EC.
Background:Endobronchial metastasis from primary ovarian cancer (OC) is very rare. To enhance our understanding of this disease, we present a case report and retrospective analysis of a patient with a bronchial tumor as a manifestation of primary OC recurrence. Case Description:A 51-year-old woman presented with a history of intermittent cough and expectoration over 3 months by suffocating pneumonia for 3 weeks. Chest X-ray revealed multiple nodular masses at the right upper lobe, soft tissue thickening with bronchial invasion in the left upper lobe, enlargement of the right and left upper hilar, spreading mediastinum, and elevated right septum. Bronchoscopy identified stenosis in the right main bronchus opening with obstruction of the apical, middle, and posterior segmental bronchi in the opening of left main bronchus by a visible neoplasm. Biopsy of the endobronchial lesion was akin to metastatic OC. Indeed, the patient was previously treated for advanced OC with enlarged left supraclavicular nodules [International Federation of Gynecology and Obstetrics (FIGO) stage 4B]. The treatment includes surgical resection of the uterus, fallopian tubes, ovaries, omentum, and left supraclavicular lymph nodes, as well as chemotherapy before and after surgery. Unfortunately, further chemotherapy was discontinued due to intolerance. Rapid disease progression occurred leading to her late self-referral and admission, decision for palliation, ultimately resulting in her demise. Conclusions:Flexible bronchoscopy combined with imaging and immunohistochemistry tests proves to be an effective diagnostic strategy for identifying endobronchial metastasis in OC patients. Endobronchial intervention, radiotherapy, and chemotherapy emerge as viable treatment modalities for these patients. The prognosis of OC patients with an endobronchial metastasis as a manifestation of recurrent disease should be considered in the context of their advanced disease despite available active treatment modalities.
Background/Objectives: The advancement of natural language processing (NLP) technologies has transformed various sectors. However, their application in the healthcare domain, particularly for analysing clinical notes, remains underdeveloped. We investigated the use of deep neural networks, specifically transformer-based models, to predict intraoperative and post-operative outcomes related to advanced-stage epithelial ovarian cancer cytoreduction (aEOC) using unstructured surgical notes. Methods: We evaluated the performance of RoBERTa, a general-purpose language model, and GatorTron, a domain-specific model, across eight binary classification tasks using the same dataset. The dataset consisted of 560 surgical records from patients with aEOC who underwent cytoreductive surgery at a tertiary UK reference centre. Predictive outcomes were converted into binary features to facilitate classification tasks. To enhance the contextual information available to the models, textual data from “operative findings” and “operative notes” were concatenated. Results: Our findings highlight the tangible benefits of employing domain-specific language models for clinical text analysis. GatorTron generally outperformed RoBERTa across most predictive tasks, underscoring the advantages of domain-specific pretraining for understanding medical terminology and context. Both models struggled to predict certain outcomes, particularly those involving post-operative events like major complications and length of hospital stay, despite adjustments in hyperparameters and training strategies. This limitation suggests that operative text alone may not sufficiently capture the complexities of post-operative recovery. Conclusions: These findings have valuable implications for developing medical AI systems to improve the delivery of modern aEOC healthcare.