We evaluated whether a compact open-source large language model (LLM), Llama3.2-3B-instruct ( base model), can be systematically optimized with low-rank adaptation (LoRA) for clinical information extraction and downstream 5year survival prediction in bladder cancer while remaining suitable for local, HIPAA-compliant deployment. From 781 radical cystectomy cases, 156 patients met inclusion criteria and were split into training/validation/test sets of 82/10/64. The LLM extraction task targeted post-surgery pathologic T stage, pathologic node stage, and lymphovascular invasion for generation of input to a nomogram. LoRA was applied to query, key, value, and output projections and trained with causal language modeling (batch size 1, learning rate 0.0002, and 1500 epochs). We selected one checkpoint based on validation stability and maximal clinical accuracy. On independent testing, extraction accuracy reached 96% (pathologic T stage 97%) for the finetuned Llama model, compared to 94% (pathologic T stage 91%) for GPT-4, and 89% (pathologic T stage 55%) for the base model. Nomogram AUCs were 0.82 +/- 0.05 (Finetuned Llama), 0.82 +/- 0.05 (GPT-4), 0.82 +/- 0.06 (manual), and 0.79 +/- 0.06 (base model). Multimodal Clinical-Radiomic-Deep model AUCs were 0.88 +/- 0.05 for the finetuned Llama model, 0.88 +/- 0.05 for GPT-4, 0.89 +/- 0.04 for manual extraction, and 0.85 +/- 0.05 for the base model. These results show that GPT-4- level prognostic performance can be achieved by finetuning a resource-efficient local model.
INTRODUCTION:Enfortumab vedotin (EV)-based regimens have become standard treatments for advanced urothelial carcinoma (aUC). However, clinical trials excluded clinically significant peripheral neuropathy or uncontrolled diabetes mellitus. Therefore, we characterized real-world outcomes of patients with aUC and pre-existing peripheral neuropathy and/or diabetes mellitus receiving EV-based therapies. PATIENTS AND METHODS:In the multicenter retrospective UNITE database, patients with documented baseline neuropathy and diabetes mellitus were identified. Observed response rate (ORR) and duration of response (DOR) were compared using Fisher's exact test. Progression-free survival (PFS) and overall survival (OS) were estimated by Kaplan-Meier method. Multivariable Cox models were used to adjust for confounding variables, including treatment duration. RESULTS:Comparing 268 patients with baseline neuropathy to 586 patients without neuropathy, median PFS was 7.2 versus 6.3 months (HR 0.87 [0.73-1.04]; P = .11), and median OS 14.4 versus 13.0 months (HR 0.88 [0.73-1.07]; P = .21), Although discontinuation rate due to intolerance was significantly higher for patients with baseline neuropathy (26% vs. 18%; P = .008), these patients did not have shorter survival. Comparing 157 patients with baseline diabetes mellitus to 697 patients without diabetes mellitus, median PFS was 5.3 versus 6.7 months (HR 1.24 [1.01-1.52]; P = .04), and median OS 13.7 versus 13.3 months (HR 1.06 [0.84-1.34]; P = .62). CONCLUSION:Patients with known baseline neuropathy and/or diabetes mellitus did not have worse survival outcomes receiving EV-based regimens. Patients with baseline neuropathy who discontinued EV early due to intolerance also did not exhibit worse survival. Our findings suggest despite these relevant underlying comorbidities, patients with aUC can derive benefit from EV-containing regimens, with very close monitoring.
Accurate extraction of clinical information from unstructured medical records is essential for developing robust predictive models for decision support in clinical tasks. In this study, we investigate the variability of large language models (LLMs) in retrieving key clinical variables for bladder cancer survival prediction. Building on our previous work, where we developed a multi-modal survival prediction model (CRD) integrating clinical data, radiomics, and deep-learning-based image analysis, we assessed the consistency of four LLMs-Dolly-v2-12b, Vicuna-v1.3-13b, Llama-v2.0-13b, and GPT-4.0-in extracting three critical pathological variables from 10 representative electronic medical records (EMRs). Experiments were conducted to evaluate (1) model effect, (2) model evolution over time, and (3) input length and order effects. Agreement across repeated extractions was measured using Fleiss' Kappa. The results showed that Llama and GPT-4 exhibited high reproducibility (Kappa > 0.80), whereas Dolly and Vicuna demonstrated moderate agreement (Kappa = 0.60-0.70). GPT-4 maintained excellent consistency across multiple time points, though minor variability was observed in a later test. Input length and order had no apparent impact on retrieval accuracy. These findings underscore the potential of LLMs in consistent and reliable data extraction from unstructured medical records, paving the way for their integration into multi-modal prognostic modeling for bladder cancer patients.
4553 Background: AS is a recognized strategy in select pts with mRCC to maximize QOL and delay potential toxicity of systemic therapy (ST); however, only 2 prospective studies of AS in mRCC have been published: only 1 with patient-reported outcomes (PRO) and none in the IO-TKI era. Methods: ODYSSEY is a prospective observational study of 500 US pts with mRCC. Eligible pts must have mRCC (any histology), no prior ST, age ≥19. Pts were excluded if treated for non-mRCC cancers or if not followed at a PCORnet study site. Pts completed QOL surveys at baseline, by phone every 3 months for 2 years and then every 6 months until end of follow up. The primary objective is to determine patterns of change in QOL and symptom burden of pts with mRCC. Minimally important differences (MID) are 3 points for FKSI-19 total score, 1 point for the FKSI-Disease Related Symptoms (DRS) subscale and 7 points for FACT-G. Here we report baseline pt characteristics in AS pts compared to ST pts, and baseline QOL differences between these cohorts. Results: As of 1/6/25, 392 pts were enrolled of whom 299 were managed with ST, and 93 pts deferred ST; of these, 53 pts (57%) were classified as AS. Pts on AS are median age 68 yrs, 66% male, 94% white, 81% clear cell, 50% favorable risk, 44% intermediate risk. Compared with ST pts, AS pts were more likely to have undergone nephrectomy (91% vs 53%), favorable risk profile (50% vs 15%), pancreatic metastasis (15% vs 5%), and longer time since RCC diagnosis (median 58 vs 3.3 months); and less likely to have bone, brain, or liver metastasis. After median 8.8 months follow-up (IQR 2.9, 16.2), 2 pts (4%) on AS had died compared with 45 pts (15%) on ST. One pt (2%) on AS started first-line therapy and 45 pts (15%) discontinued ST. Mean baseline QOL (FKSI-19 total, DRS and FACT-G) for AS and ST ODYSSEY pts is shown in the Table (higher score indicates better QOL), with RCT data for reference (NA, not assessed). ODYSSEY pts on AS had higher QOL for all measures compared with ODYSSEY pts on ST. FSKI-DRS was the same or lower for pts on AS compared to the pivotal trials, while ODYSSEY pts on ST had both lower FKSI-19 total and DRS. Conclusions: In our large prospective cohort from ODYSSEY, pts on AS had higher median QOL scores than pts on ST, but similar to those included in RCTs. These results suggest that some RCT pts could have benefitted from AS. Further follow up is needed to determine long term outcomes in pts on AS and how they respond to deferred ST. Clinical trial information: NCT04919122 . Instrument, Mean (SD) ODYSSEY AS (N=53) ODYSSEY ST (N=299) ODYSSEY Difference, AS vs ST (95% CI) CheckMate 214 (N=425) KEYNOTE 426 (N=402) CheckMate 9ER (N=323) CLEAR (N=351) FKSI-19 total 63.4 (9.2) 56.3 (12.5) 7.0(4.0, 10.1) 60.1 (9.8) NA 58.7 (10.6) NA FKSI-DRS 30.8 (4.2) 27.9 (6.0) 2.9(1.5, 4.7) 30.7 (4.5) 32 (4.2) 30.2 (5.2) 31.3 (4.4) FACT-G 87.9 (14.0) 82.4 (16.7) 5.6(1.1, 10.0) 82.6 (15.0) NA NA NA
Outcomes from chemotherapy initiation in maintenance-eligible patients with locally advanced urothelial carcinoma (la/mUC) are of interest given new/emerging upfront immunotherapy-based treatment options. This exploratory analysis of KEYNOTE-361 by retrospective eligibility for maintenance therapy suggests that a majority of patients with untreated la/mUC who received chemotherapy alone may have been considered maintenance eligible; these patients had favorable survival outcomes relative to those considered maintenance ineligible. Introduction: The phase 3 KEYNOTE-361 trial of first-line pembrolizumab with or without chemotherapy versus chemotherapy alone in patients with locally advanced or metastatic urothelial carcinoma (la/mUC) completed enrollment before the approval of postchemotherapy maintenance avelumab for patients without progressive disease. This post hoc analysis evaluated the outcomes of patients who received chemotherapy alone in KEYNOTE-361 by retrospective eligibility for subsequent maintenance therapy. Patients and Methods: Patients in the chemotherapy alone arm were retrospectively categorized as maintenance eligible (received >= 4 cycles of chemotherapy and did not die or experience disease progression within 10 weeks of chemotherapy completion), maintenance ineligible (received < 4 cycles of chemotherapy or had progressive disease or died within 0-10 weeks after completion of >= 4 cycles of chemotherapy), and indeterminate eligibility for maintenance therapy (if neither maintenance eligible or ineligible). End points included progression-free survival per Response Evaluation Cr iter ia in Solid Tumors version 1.1 by blinded independent central review and overall survival from randomization (start of chemotherapy). Results: Median follow-up was 31.7 months (range, 22.0-42.3). Among 342 patients who received chemotherapy alone, 172 (50.3%) were maintenance eligible, 108 (31.6%) were maintenance ineligible, and 62 (18.1%) had indeterminate eligibility for maintenance therapy. The median progression-free survival was 9.0 months (95% CI 8.4-10.4) in maintenance-eligible patients, 5.1 months (4.2- 6.0) in maintenance-ineligible patients, and 2.3 months (1.9-3.8) in the indeterminate group; median overall survival was 23.3 months (95% CI 19.4-26.1), 10.2 months (9.1-11.6), and 5.5 months (3.7-8.5), respectively. Conclusion: This post hoc analysis suggests that a majority of patients with untreated la/mUC who initiated chemotherapy in a clinical trial may have been considered eligible for maintenance therapy and had favorable survival outcomes compared with those considered maintenance ineligible.
220 Background: Prostate cancer (PC) is genetically heterogeneous, and genomic alterations may impact prognosis and therapy response. We created CAPSTONE, a database of lethal PC that integrates comprehensive genomic sequencing with deep clinical phenotyping to explore the clinical implications of castrate resistant prostate cancer (CRPC) evolution. Methods: PC patients underwent tissue collection (4/05-7/21) for tumor RNA-sequencing and tumor/normal whole exome sequencing (HUM00046018, HUM00048105, HUM00067928, SU2C). Sequencing was processed using Turnkey Precision Oncology. We analyzed somatic and germline mutations, gene fusions, copy number alterations, and chromosomal instability (CIN) along with transcriptomic signatures and pathways. We collected clinical data (05/21-01/22) including overall survival from time of castrate resistant prostate cancer (OScrpc) and from time of biopsy (OSb). Patients were split into discovery and validation cohorts. We used cox proportional hazard models to evaluate OS. Results: Data was available for 454 men (n=192, n=262). Median follow up from CRPC was 32.1 (IQR: 14.7-50.1) and 33.3 (IQR: 20.9-55.6) months respectively. Median age at CRPC was 66 (IQR: 60-73) and 67 (IQR: 62-72) years. Of the 1,581 recurrently altered genes (>2%), 72 had a significant (p<0.01) univariate association with OS in the discovery cohort. From these, 3 (RB1, TP53, CDKN1B) were significantly associated with OS in the validation cohort (T1). Subgroup analysis revealed AR mutations but not amplifications were associated with improved OS (T1). AR amplifications were enriched in samples with TP53 alterations (p<1.0e-3, p<1.0e-3) and high CIN (p<1.0e-3, p=8.7e-3). Finally, using discovery data, we generated a gene signature associated with OS independent of TP53, RB1, and CDKN1B and validated this signature in the validation cohort (T1). Conclusions: RB1, TP53, and CDKN1B were recurrently altered genes independently associated with worse OS in CRPC. Further, we identified a gene signature associated with poor OS in CRPC independent of these alterations. Association with OSb in CRPC. Discovery Validation (n=192) (n=262) Univariate HR (95% CI) [p] Multivariate HR (95% CI) [p] Univariate HR (95% CI) [p] Multivariate HR (95% CI) [p] OSb RB1 2.74 (1.69-4.43) [4.3e-5] 2.5 (1.47-4.24) [0.00067] 2.37 (1.53-3.68) [0.00012] 2.33 (1.49-3.63) [2e-04] TP53 2.23 (1.57-3.11) [5.9e-6] 2.55 (1.78-3.66) [3.8e-07] 1.52 (1.14-2.03) [0.0046] 1.48 (1.1-1.98) [0.0094] CDKN1B 3.39 (1.42-8.13) [0.0061] 3.46 (1.47-8.15) [0.0046] 2.47 (1.32-4.6) [0.0045] 2.86 (1.51-5.42) [0.0013] Gene signature 2.08 (1.72-2.5) [1.5e-14] 1.88 (1.56-2.28) [5e-11] 1.36 (1.17-1.57) [4.3e-05] 1.32 (1.13-1.53) [0.00035] AR amplification 1.03 (0.74-1.44) [0.86] - 1.15 (0.86-1.54) [0.34] - AR mutation 0.55 (0.3-1) [0.049] - 0.64 (0.4-1.02) [0.063] -
This study explores the impact of physician experience, specialty, and institutional background on the performance of AI-assisted assessment of treatment responses in bladder cancer patients. Utilizing pre- and post-chemotherapy CTU scans from 123 patients, 17 physicians with varying levels of experience and from different specialties and institutions assessed 157 lesion pairs. The lesion pairs were divided into easy and difficult cases to evaluate the AI system's effectiveness in different scenarios. The study revealed that AI assistance significantly improved diagnostic accuracy in easy cases for both experienced and inexperienced physicians, with a great benefit observed in radiologists and oncologists. In difficult cases, the AI's impact was present but less pronounced, indicating that while AI can enhance performance in challenging situations, its effectiveness is more limited in complex cases. Additionally, the study found that institutional background influenced the effectiveness of AI assistance, suggesting that certain training or cultural factors may affect physicians' trust in AI recommendations. The findings underscore the potential of AI to support clinical decision-making in bladder cancer treatment response assessment, particularly in less complex cases. However, they also highlight the need for tailored implementation and user training of AI systems to maximize their effectiveness across different medical specialties and institutions. By aligning AI tools with the specific needs and expertise of physicians, their confidence and efficacy in using AI in complex medical scenarios can be enhanced.
BACKGROUND:Immune Checkpoint Inhibitors (ICIs) are used for advanced urothelial carcinoma (aUC) in different settings. Most patients have pure UC (PUC) but about one-third have UC mixed with histology subtypes (HS). We examined outcomes in patients with HS aUC treated with ICI. MATERIALS AND METHODS:We included patients from 26 centers with PUC and any HS treated with ICI as 1st line (1L) upfront, maintenance avelumab (mAV), and ≥2nd line [2+L] therapy. We calculated overall and progression-free survival (OS, PFS) and observed response rate (ORR) from ICI start. RESULTS:We included 1511 patients; 752 1L, 609 2+L, 150 mAV. 1L: median OS was 15 (95% CI, 12-17) months for patients with PUC (n = 518), 15 (95% CI, 8-23) months for squamous UC (n = 85) (HR = 1.2, [95% CI, 0.8-1.6]), 11 (95% CI, 6-17) months for micropapillary UC (n = 46) (HR = 1.2, [95% 0.8-1.8]), and 21 (95% CI, 12-30) months in patients with UC mixed with ≥2 HS (n = 30), (HR = 0.9, [95% CI, 0.5-1.4]). 2+L: median OS was 9 (95% CI, 8-10) months for patients with PUC (n = 441), 9 (95% CI, 1-12) months for squamous UC (n = 60) (HR = 1.1, (95% CI, 0.8-1.6]), 6 (95% CI, 1-11) months for micropapillary UC (n = 37) (HR = 1.1, [95% 0.7-1.6]), and 7 (95% CI, 4-10) months in patients with UC mixed with ≥2 HS (n = 17), (HR = 1.6, [95% CI, 0.9-3.1]). CONCLUSION:We found no significant OS difference between PUC and HS in patients with aUC treated with ICI monotherapy. Limitations include retrospective design, small sample size in several subsets, lack of randomization, no central imaging or pathology review, selection and confounding biases. Results are hypothesis-generating and need prospective validation.
479 Background: Cabozantinib is approved as a subsequent therapy for patients with metastatic renal cell carcinoma (mRCC) based on the METEOR trial. However, only 5% of patients in this trial received prior immunotherapy. Methods: We identified patients with mRCC from the IMDC who were treated with cabozantinib in the second-line (2L) setting from 2010 to 2023. These patients were stratified by IMDC risk groups and first-line (1L) treatment. We analyzed overall response rate (ORR), time to next treatment (TTNT), treatment duration (TD), overall survival (OS) and performed a multivariable analysis adjusted by IMDC criteria at 2L. Results: A total of 603 patients were identified. Baseline characteristics are summarized in the table. For the entire cohort, the ORR was 25.7%, TTNT was 10.1 months (mo), TD was 8.9 mo and mOS was 19 mo. Among patients treated with 1L ipilimumab/nivolumab (n=190), anti-PD1 + TKI (n=148), and TKI alone (n=207), cabozantinib showed an ORR of 27.2%, 26.4%, and 25%, respectively; a median TTNT of 9.9, 10.3, and 9.7 mo; a median TD of 9.4, 8.2, and 8.3 mo. Median OS was 18.6, 17.6, and 21.3 mo, respectively. A multivariable analysis was unable to demonstrate that first-line ORR (CR/PR vs SD vs PD) or TTNT (< 12 vs ≥ 12 mo) predicts for second-line cabozantinib ORR in the overall cohort and by first-line therapy type. Specifically, patients with stable disease or with partial and complete response in 1L were associated with an OR for a response of 0.99 (95% CI 0.52-1.92) or 1.23 (95% CI 0.61-2.49), respectively. Similarly, a first-line TTNT of ≥12 months had an OR for response of 1.04 (95% CI 0.59-1.82). Conclusions: This study demonstrates that cabozantinib maintains efficacy comparable to that observed in the METEOR trial in a real-world setting, including in patients with prior immunotherapy combination therapies. Efficacy of 1L treatment does not predict efficacy of 2L cabozantinib. Baseline characteristics. Variable Overall (N = 603) IO-IO (N = 190) IO-TKI (N = 148) TKI Alone (N = 207) Other (N = 58) p-value Non clear cell histology 107 (17.7) 33 (17.4) 26 (17.6) 28 (13.5) 20 (34.5) 0.003 Nephrectomy 416 (69.0) 97 (51.1) 110 (74.3) 166 (80.2) 43 (74.1) <0.001 1st line IMDC Risk Fav/Int/Poor 83 (13.8)/296 (49.1)/101 (16.7) 10 (5.3)/100 (52.6) /46 (24.2) 37 (25)/62 (41.9) /20 (13.5) 29 (14)/99 (47.8)/27 (13) 7 (12.1)/35 (60.3)/8 (13.8) <0.001 2nd line IMDC risk Fav/Int/Poor 56 (9.3)/269 (44.6)/108 (17.9) 7 (3.7)/90 (47.4)/45 (23.7) 25 (16.9)/63 (42.6)/24 (16.2) 19 (9.2)/87 (42.0)/32 (15.5) 5 (8.6)/29 (50.0)/7 (12.1) 0.002 Greater than 1 site of Metastasis 473 (78.4) 146 (76.8) 114 (77.0) 161 (77.8) 45 (77.6) 0.722 Brain Metastasis 36 (6.0) 18 (9.5) 5 (3.4) 13 (6.3) 0 (0.0) 0.022 Bone Metastasis 221 (36.7) 74 (38.9) 60 (40.5) 70 (33.8) 17 (29.3) 0.326 Liver Metastasis 106 (17.6) 32 (16.8) 26 (17.6) 39 (18.8) 9 (15.5) 0.926
Immune checkpoint inhibitors (ICI) improve overall survival in advanced urothelial carcinoma (aUC), but response rates as monotherapy remain modest. In this retrospective cohort we compared outcomes between patients with and without tumor genomic alterations ( FGFR 2/3, MTAP, ERBB2). Our results suggest that specific genomic alterations may have association with response and survival with ICI in aUC and support further validation. Background: FGFR2/3, MTAP and ERBB2 genomic alterations have treatment targets in advanced urothelial carcinoma (aUC). These alterations may affect tumor microenvironment and outcomes with immune checkpoint inhibitors (ICIs) in aUC. Patients and Methods: We identified patients with available genomic data in our multi-institution cohort of patients with aUC treated with ICI. Outcomes (observed response rate [ORR], progression-free and overall survival [PFS, OS]) with ICI were compared between patients with and without FGFR 2/3, MTAP, ERBB2 alterations. We compared ORR using logistic regression and PFS/OS using Cox proportional hazards. Results: Out of 1,514 patients, 276 (18%), 174 (11%) and 208 (14%) patients had known FGFR2/3, MTAP and ERBB2 alteration status, respectively. and were treated with ICI in 1L or 2 + L. In patients with (vs. without) FGFR2/3 alteration, ORR with ICI was 21% vs. 32% (OR 0.54; [95%CI 0.32-0.91]), PFS was significantly shorter in patients with FGFR2/3 alterations (HR = 1.36 [95%CI 1.03-1.80]; P = 0 . 03); OS was not significantly different (HR = 1.22 [95%CI 0.86-1.47]). In patients with (vs. without) MTAP alteration, ORR with ICI was 25% versus 40% (OR 0.52 [95%CI 0.20-1.38]); PFS and OS were nonsignificantly different. In patients with (vs. without) ERBB2 alteration, ORR with ICI was similar (37% vs. 35%; OR 1.06; 95%CI 0.57-1.97); PFS and OS were significantly longer in patients with ERBB2 alteration [HR 0.63 (95%CI 0.41-0.95); P = 0 . 03; HR 0.66, [95% CI 0.44-0.97]), respectively. Conclusion: Our results support further evaluation of FGFR2/3, MTAP and ERBB2 alterations as putative biomarkers in patients with aUC treated with ICI.
BACKGROUND:177Lu-PSMA-617 is approved for patients with metastatic castration-resistant prostate cancer (mCRPC). Although treatment is associated with improved outcomes, not all patients benefit and response is heterogeneous. We aim to characterize genomic alterations associated with benefit to 177Lu-PSMA-617. MATERIALS AND METHODS:This study used the Prostate Cancer Precision Medicine Multi-Institutional Collaborative Effort (PROMISE) clinical-genomic database (n = 2445). The primary endpoint was ≥50% PSA decline (PSA50) from baseline with 177Lu-PSMA-617 in molecular subgroups. Secondary endpoints included 90% PSA decline (PSA90). Associations were assessed using Fisher's exact test and Cox regression in multivariable analysis. RESULTS:We identified 183 mCRPC patients treated with 177Lu-PSMA-617. Median number of prior lines of mCRPC therapy was 3. Overall, PSA50 was 49%, median progression-free survival was 7.6 months, and median overall survival was 13.9 months. NF1 (n = 8) and FOXA1 alterations (n = 5) were associated with increased PSA50 (88% vs 47%, P = .03 for NF1; 100% vs 47%, P = .03 for FOXA1). Among CRPC sequenced tumors (n = 119), androgen receptor (AR) alterations (n = 58) were associated with lower PSA50 (38% vs 60%, P = .03). While any tumor suppressor genes (TSG) (PTEN, TP53, RB1) (n = 109) or TP53 (n = 83) alteration were associated with lower PSA90 (P = .02 for both), NF1 (n = 8), and FOXA1 alterations were associated with higher PSA90 (P = .03 and P = .003, respectively). CONCLUSIONS:This analysis identifies potential genomic predictors of response to 177Lu-PSMA-617, with NF1 and FOXA1 alterations associated with favorable outcomes and AR and TSG alterations with diminished response. These hypothesis-generating findings suggest genomic profiling may inform selection for PSMA-targeted therapy and warrant prospective validation in larger cohorts.
Confidence estimation of machine learning (ML) and artificial intelligence (AI) model outputs is crucial for guiding decision-making. This study proposes two confidence estimation methods to enhance the trustworthiness of ML and AI models for classifying complete and non-complete responders for bladder cancer following neoadjuvant chemotherapy using CT urography: Segmentation Variability-Based Confidence Estimation (SVCE) and Classification Model Variability-Based Confidence Estimation (CMVCE). SVCE assesses prediction confidence based on the variability in predictions from models with different tumor segmentations, and thus extracted feature values, but the same classification model, while CMVCE evaluates confidence based on the variability in predictions from models with the same segmentations but different classification models. Our results demonstrate that cases with lower prediction variability consistently achieve better classification performance, with area under the receiver operating characteristic curve (AUC) values exceeding those of cases with higher variability. This study highlights that our proposed variability-based methods have the potential to provide an indicator of trustworthiness of the ML/AI prediction, which may enhance clinicians' confidence in adopting ML/AI models for decision support.
867 Background: In trials testing enfortumab vedotin and pembrolizumab (EVP), prior treatment with immune checkpoint inhibitors (ICIs) was not permitted, resulting in a knowledge gap regarding efficacy of EVP in patients (pts) previously treated with ICIs. We hypothesized that EVP would have efficacy in pts with prior ICI exposure. Methods: In the retrospective UNITE study, we identified all pts treated with ICI prior to receiving EVP. The observed response rate (ORR) was assessed in evaluable pts who had imaging after ≥1 cycle of EVP. The following factors were evaluated to assess effect on EVP outcomes: type of ICI received (PD-1 vs PD-L1), prior pembrolizumab vs other ICI, time from last ICI to start of EVP, duration on prior ICI treatment, whether patient had clinical benefit (SD/PR/CR) to prior ICI, whether ICI was the immediate prior therapy line and whether it was the only prior line. ORR for these categories was compared using logistic regression, while progression-free survival (PFS) and overall survival (OS) from EVP start were analyzed using the log-rank test and Cox proportional hazards model. Results: Among 220 pts treated with EVP across 10 US sites, 43 (20%) had previously received ICI. Median age was 69 years; 34 (79%) were men, 40 (93%) were Caucasian, and 27 (63%) had pure urothelial carcinoma, 9 (21%) liver mets and 31 (72%) ECOG PS 0/1. Of 43 pts, 4 received ICI in the peri-operative setting (3 nivolumab, 1 pembrolizumab) and 39 in the metastatic setting (19 pembrolizumab, 8 avelumab maintenance, 6 nivolumab, 2 pembrolizumab maintenance and 1 each of atezolizumab, ipilimumab/nivolumab, durvalumab/tremelimumab, nivolumab/NKTR214). ORR was 48% (95% CI: 31 - 66) in 33 evaluable pts, with 13 (39%) PR and 3 (9%) CR; 30% had SD as best response (DCR 79%), and 21% PD. After median follow-up of 14 mos, median PFS was 6.9 mos (95% CI: 3.91 – 12.2) and median OS 15.4 mos (95% CI: 8.7 – NR). Outcomes by group are shown in the Table. Conclusions: Pts treated with EVP after prior ICI experienced high ORR and disease control rate. Results from this multi-site retrospective study are hypothesis-generating and require prospective validation in larger cohorts. Groups ORR PFS OS OR (95% CI) p-value HR (95% CI) p-value HR (95% CI) p-value Type of ICI (PD-1 vs PD-L1) 2.15 (0.36 – 17.49) 0.4 0.72 (0.32 – 1.61) 0.4 0.59 (0.23 -1.52) 0.3 Prior pembrolizumab vs not 0.54 (0.12 – 2.23) 0.4 0.55 (0.26 – 1.16) 0.1 0.40 (0.15 – 1.02) 0.05 Time from prior ICI* 1.01 (0.98 – 1.05) 0. 5 0.99 (0.98 – 1.01) 0.9 0.98 (0.96 – 1.02) 0.3 Time on prior ICI* 0.87 (0.74 - 0.98) 0.05 0.99 (0.94 – 1.05) 0.9 0.97 (0.89 – 1.05) 0.4 Clinical benefit to prior ICI (DCR vs PD) 0.50 (0.08 – 2.75) 0.4 0.70 (0.27 – 1.81) 0.5 0.89 (0.29 – 2.73) 0.8 Multiple prior lines vs ICI as only prior line 0.51 (0.11 – 2.29) 0.4 1.40 (0.57 – 3.44) 0.5 1.87 (0.54 – 6.43) 0.3 ICI as immediate prior line vs not 0.93 (0.15 -5.82) 0.9 0.56 (0.20 – 1.57) 0.3 0.53 (0.17 – 1.64) 0.3 *Continuous.
4558 Background: In the past 7 years, 4 immuno-oncology (IO) based combinations have been approved for mRCC. However, QOL data on these combinations are limited to trials in which collection and reporting was not standardized, which further limits cross trial comparisons. Real-world QOL data with multiple treatment regimens are needed to understand how these regimens are tolerated in practice. Methods: ODYSSEY is a multi-site, prospective observational study of 500 US pts with mRCC. Pts must have mRCC (any histology), no prior ST, and follow up at a PCORnet study site. Exclusion criteria include treatment for cancer except mRCC. The primary objective is to determine patterns of change in QOL and symptom burden of pts with mRCC. Minimally important differences (MID) are 3 points for FKSI-19 total score, 1 point for the FKSI-Disease Related Symptoms (DRS) subscale and 7 points for FACT-G. Results: As of 1/6/25, 392 pts were enrolled of whom 299 were managed with ST. Of pts on ST, 114 were treated with IO-IO, 108 with IO-TKI, 33 with IO alone, 27 with Other, and 18 with TKI alone. Median follow up for all pts is 8.8 months (IQR 2.9, 16.2). IO-IO pts are median age 64 yrs, 81% male, 84% white, 56% prior nephrectomy, 84% clear cell; IO-TKI pts median age 66 yrs, 76% male, 92% white, 43% prior nephrectomy, 66% clear cell; IMDC risk profiles are similar. IO-IO pts are more likely to be KPS 100; whereas, IO-TKI pts have a higher median number of metastatic sites and are more likely to have bone or liver metastasis. With median follow-up of 6.0 (IQR 1.8, 14.9) and 8.8 months (4.5, 17.8) in the IO-IO and IO-TKI cohorts, 17 (15%) and 21 (19%) pts had died with a median time to death of 5.4 (2.5, 6.4) and 5.7 months (4.4, 12.8), respectively. 20 (18%) and 16 (15%) pts on IO-IO and IO-TKI had discontinued therapy at a median of 3.0 (IQR 2.0, 4.0) and 6.4 months (2.4, 13.0), respectively. Rates of discontinuation for disease progression and toxicity are similar. Baseline PROs for pts on ST are shown in the Table (higher score indicates better QOL), with RCT data for reference (NA, not assessed). Conclusions: In our prospective multi-center ODYSSEY study, we demonstrate that real world pts treated with contemporary ST have worse baseline QOL than those enrolled on pivotal RCTs. One-third of IO-IO and IO-TKI pts died or discontinued therapy within 6 months of initiation. Our data on real world vs RCT differences in baseline QOL may partially explain the limitations of current IO combination regimens in practice and support development of alternative treatment approaches. Instrument, Mean (SD) ODYSSEYIO-IO(N=114) CheckMate214(N=425) ODYSSEYIO-TKI(N=108) KEYNOTE426(N=402) CheckMate9ER(N=323) CLEAR(N=351) FKSI-19 total 56.0 (12.5) 60.1 (9.8) 53.8 (11.8) NA 58.7 (10.6) NA FKSI-DRS 27.5 (6.3) 30.7 (4.5) 27.0 (5.8) 32 (4.2) 30.2 (5.2) 31.3 (4.4) FACT-G 82.4 (16.7) 82.6 (15.0) 78.5 (16.0) NA NA NA
Advances in natural language processing (NLP) and machine learning could assist human users in clinical data extraction from unstructured electronic medical records (EMRs). This study investigates the accuracy and consistency of several Large Language Models (LLMs) - including Dolly, Vicuna, Llama, and GPT-4 - in extracting critical clinical information pertinent to bladder cancer survival prediction. Using EMRs from 163 bladder cancer patients, we assessed the impact on LLM performance by factors such as differences in the trained models, model evolution, input text length, and sequencing of case inputs. GPT-4 demonstrated superior performance with Fleiss' Kappa values exceeding 0.97, accuracy consistently above 93%, and survival prediction metrics closely aligned with ground truth (AUC ± 0.02). Among offline models, Llama-2.0-13b and Llama-3.3-70b exhibited the highest reliability in both information extraction and survival prediction. This study underscores the potential of LLMs to automate clinical data extraction for predictive modeling while highlighting the challenges related to LLM variability and reliability.
Prostate cancer is the second leading cause of cancer-related death among American men, with a new diagnosis made every 2 min in the United States. Advanced cases are commonly treated with androgen deprivation therapy (ADT). Despite its effectiveness, treatment failure remains inevitable for many patients, necessitating better predictive tools for clinical management of disease. This study presents a data-driven mathematical modeling approach that integrates patient-specific prostate-specific antigen (PSA) time-course data with experimentally measured PSA expression rates to improve the prediction of ADT failure. Our findings suggest that post-nadir PSA dynamics, rather than initial decline, hold greater prognostic value and can inform PSA monitoring schedules. By employing virtual clones of individual patients, our model integrates routinely collected PSA measurements to dynamically predict ADT failure probabilities at future clinic visits. If implemented in clinical practice, this personalized framework could empower oncologists to make proactive, informed treatment decisions and guide timely interventions.