11138 Background: The majority of patients with non-small cell lung cancer (NSCLC) will have multiple organ sites of metastasis and poor overall survival (rwOS). Specific organ involvement influences patient outcomes, but the impact of overarching patterns of metastasis on rwOS and response to systemic therapy (ST) is poorly understood. Prior studies have been limited in their scale, availability of detailed data on sites of metastasis, and inclusion of patients receiving immunotherapy (IO) or targeted therapy (TT). We leveraged a large-scale, real-world database to characterize and assess the prognostic value of patterns of metastasis on rwOS and time to next therapy (rwTTNT) in patients with metastatic NSCLC receiving ST. Methods: This retrospective study utilized the Flatiron Health Research Database of patients with NSCLC treated with first-line (1L) ST (n = 327,044). Patient, tumor, and treatment variables were summarized descriptively and compared across patterns of metastasis. rwOS and rwTTNT from 1L start were estimated via Kaplan Meier method and compared with logrank test. Bernoulli mixture models were used to cluster patients by metastatic sites at 1L start. Results: Sites of disease were evaluated for 75,304 patients. Data from 5 metastatic sites revealed 6 clusters reflective of real-world patterns of metastasis: bone (n = 33,882), pleural (n = 12,123), liver (n = 6,133), brain (n = 11,441), adrenal avid (n = 6,975), and high metastatic burden (HMB, ≥3 sites with frequency ≥0.5, n = 4,750). While HMB (7.2 months median rwOS, 4.8 months median rwTTNT) and liver avid disease (9.4, 5.2 months) had poor outcomes as expected, we also found that bone (11.0, 5.9 months) and adrenal (11.8, 6.0 months) avid clusters had worse median rwOS and shorter rwTTNT compared to pleural (15.7, 7.6 months) and brain (15.3, 7.4 months, p < 0.001). Outcomes associated with a specific site of metastasis varied based on patient cluster assignment. Compared to the cluster equivalent (e.g., brain avid), individual metastatic site (e.g., brain) as a prognostic factor was associated with worse median rwOS by 1.0-2.9 months and shorter median rwTTNT by 0.1-0.9 months. Compared to individual sites and site-avid clusters, HMB more often had EGFR mutations (20.7% vs 13.6% overall). IO was associated with greater median rwOS over TT in the adrenal cluster (17.7 vs 15.6 months), and IO + CTX had better rwOS vs IO alone in pleural (15.7 vs 12.9 months) and liver clusters (10.7 vs 9.4 months, p < 0.001). Conclusions: The survival outcomes and treatment response associated with defined clusters of metastases are distinct from those associated with individual organ sites or widespread disease alone. A risk-stratification framework that uses “risk phenotypes” representative of real-world patterns of metastasis may have better prognostic and predictive significance in patients with multiple sites or higher burden of disease.
BACKGROUND:Cyclin-dependent kinase (CDK) 4/6 inhibitors plus endocrine therapy (ET) are more effective than ET alone in diverse populations of patients with HR+/HER2- metastatic breast cancer (MBC). However, real-world effectiveness data for CDK4/6 inhibitors plus ET combination in bone-only MBC are limited. METHODS:This retrospective, real-world study compared clinical outcomes of first-line (1L) palbociclib (PAL) plus an aromatase inhibitor (AI) with AI alone in bone-only MBC using data from the US-based, nationwide Flatiron Health Research database. Eligible patients had HR+/HER2- bone-only MBC, were ≥18 years of age, and started 1L PAL + AI or AI alone from February 2015 to June 2022. Stabilized inverse probability of treatment weighting (sIPTW) was used to balance baseline characteristics. Overall survival (OS), real-world progression-free survival (rwPFS), and time to chemotherapy (TTC) were evaluated. RESULTS:Of 974 eligible patients with bone-only MBC, 538 received PAL + AI and 436 received AI alone. After sIPTW, baseline patient characteristics were balanced between treatment groups. Median follow-up was 31.9 months for the PAL + AI group and 35.8 months for the AI group. After sIPTW, the PAL + AI group compared with the AI group had a significantly longer OS (median 63.4 vs 51.3 months; HR = 0.78; 95% CI, 0.64-0.97, P = 0.0221), rwPFS (median 23.0 vs 18.2 months; HR = 0.72; 95% CI, 0.59-0.87, P = 0.0008), and TTC (median 47.9 vs 40.6 months; HR = 0.79; 95% CI, 0.66-0.96, P = 0.016). Consistent results were observed in unadjusted and sensitivity analyses. CONCLUSION:This real-world study demonstrates that first-line PAL + AI versus AI alone was associated with prolonged OS, rwPFS, and TTC in patients with HR+/HER2- bone-only MBC. TRIAL REGISTRATION NUMBER:NCT06495164.
e18007 Background: First-line treatment for recurrent or metastatic head and neck squamous cell carcinoma (R/M HNSCC) includes pembrolizumab or pembrolizumab plus chemotherapy, guided by PD-L1 when available. Although both were established as standards in KEYNOTE-048, they were not directly compared, leaving optimal patient selection unclear. We hypothesized that machine learning (ML)–predicted baseline prognosis modifies chemotherapy benefit, with higher-risk patients more likely to benefit from chemotherapy’s rapid clinical effects. Methods: Using the Flatiron Health Research Database, we identified patients with R/M HNSCC treated with first-line pembrolizumab or pembrolizumab plus chemotherapy and positive or unknown PD-L1. A gradient-boosted survival model predicted 6-month survival from baseline clinical variables using cross-validation, then calibrated via isotonic regression. Heterogeneity of absolute treatment benefit was evaluated using overlap-weighted regression with 2-year restricted mean survival time (RMST) pseudo-observations. We summarized the continuous treatment-effect function using a crossover point, defined as the baseline 6-month survival probability at which the estimated RMST benefit of adding chemotherapy reached a clinically meaningful magnitude (≥30 days). Patients were stratified by this survival probability, and survival was compared between treatments using inverse probability treatment weighting (IPTW). Results: Among 1,736 patients, 1,095 received pembrolizumab and 641 received pembrolizumab plus chemotherapy. Median age was 68 years, 78% were male, median follow-up was 24 months, and PD-L1 was positive in 17.9%, with 82.1% unknown. The model achieved 6-month AUC 0.75 with good calibration (Brier 0.17). Chemotherapy benefit increased as predicted survival worsened: for every 10 percentage-point decrease in predicted 6-month survival, patients gained 24 days in 2-year RMST with combination therapy (p<0.001). The crossover point corresponded to a baseline 6-month survival probability of 64%. Patients below the crossover survival probability (31.2%)—characterized by worse ECOG, weight loss, lower HPV positivity, bone metastases, and hypoalbuminemia—derived significant benefit from adding chemotherapy in the IPTW-adjusted survival analysis, while those above (68.8%) showed no benefit (Table 1). Conclusions: An ML model trained on nationally-representative data identified subgroups of patients with R/M HNSCC who benefit from adding chemotherapy to pembrolizumab. RMST differences stratified by crossover survival probability. 6-month Survival <64% 6-month Survival ≥64% 1-year RMST Δ 52.5 (22.2-80.1) 1.0 (-13.1-15.3) 2-year RMST Δ 84.4 (27.0-142.3) -17.3 (-51.6-18.2) RMST differences (Pembro+Chemo - Pembro) in days; 95% CI in parentheses.
Large language models (LLMs) are increasingly used to extract clinical data from electronic health records, offering significant improvements in scalability and efficiency for real-world data (RWD) curation in oncology. However, the adoption of LLMs introduces new challenges in ensuring the reliability, accuracy, and fairness of extracted data, which are essential for research, regulatory, and clinical applications. Existing quality assurance frameworks for RWD and artificial intelligence (AI) do not fully address the unique error modes and complexities associated with LLM-extracted data. In this paper, we propose a comprehensive framework for evaluating the quality of clinical data extracted by LLMs. The framework integrates variable-level performance benchmarking against expert human abstraction, verification checks for internal consistency and plausibility, and replication analyses comparing LLM-extracted data to human-abstracted data sets or external standards. This multidimensional approach enables the identification of variables most in need of improvement, systematic detection of latent errors, and confirmation of data set fitness-for-purpose in real-world research. Additionally, the framework supports bias assessment by stratifying across demographic subgroups. By providing a rigorous and transparent method for assessing LLM-extracted RWD, this framework advances industry standards and supports the trustworthy use of AI-powered evidence generation in oncology research and practice.
Background Comparative cyclin-dependent kinase 4/6 inhibitor (CDK4/6i) effectiveness data to date are limited to real-world progression-free survival (PFS) and overall survival. Additional effectiveness measures such as time from initial treatment to progression on first subsequent therapy (rwPFS2) and real-world tumor response (rwTR) provide full understanding of CDK4/6i real-world effectiveness. This analysis compared rwPFS2 and rwTR in patients with HR+/HER2− mBC receiving first-line (1L) CDK4/6i plus an aromatase inhibitor (AI) in US routine clinical practice. Patients and methods P-VERIFY was a retrospective Flatiron Health Research Database analysis of adult patients with HR+/HER2− mBC who started 1L CDK4/6i (palbociclib; ribociclib; abemaciclib) plus AI between February 2015 and November 2023 (N=9146). rwPFS2 analysis included all 9146 patients while rwTR analysis included 8010 (87.6%) patients who had ≥1 tumor response assessment. Stabilized inverse probability of treatment weighting (sIPTW) was used to balance patient baseline characteristics. A multivariable Cox proportional hazards model was used for sensitivity analysis. Results After sIPTW, no significant differences were observed across all pairwise treatment comparisons between the CDK4/6i for rwPFS2 and rwTR (all p> 0.05). Findings were generally consistent across subgroups and sensitivity analyses. Findings were also consistent in a separate analysis of patients treated from 2017 onward, corresponding with commercial availability of all three CDK4/6i in the US. Conclusions This real-world study demonstrated no statistically significant differences in rwPFS2 and rwTR in patients with HR+/HER2− mBC receiving 1L palbociclib, ribociclib, or abemaciclib, in combination with an AI, in routine clinical practice in the US. ClinicalTrials gov Identifier NCT06495164
e20541 Background: The non-random distribution of metastases (mets) to distant organ sites, known as organotropism, has been observed in patients with non-small cell lung cancer (NSCLC). Most patients will develop mets in two or more organ sites, with select sites and high met burden having prognostic and predictive significance. However, mechanisms driving site-specific and widespread metastatic potential are not well understood, and real-world overall survival (rwOS) remains poor. We used a database containing real-world clinical and genomic information to characterize and evaluate the prognostic value of temporal patterns of metastatic progression on rwOS in patients with NSCLC receiving systemic therapy (ST). Methods: This retrospective study utilized the Flatiron Health-Foundation Medicine Clinico-Genomic Database of patients with metastatic NSCLC treated with first-line (1L) ST. Patient, tumor, and treatment variables were summarized descriptively and compared with chi-squared tests. rwOS was estimated via Kaplan Meier method and compared with logrank test. Bernoulli mixture models were used to cluster patients by met sites at time of metastatic diagnosis, 1L therapy start, and last follow-up. Results: Data from 18 met sites was evaluated for 11,527 patients. Common sites at diagnosis were bone (22.7%), lung (17.4%), pleura (16.6%), and brain (13.0%). 72.9% of patients developed mets in two or more organ sites, and certain initial sites were associated with the non-random distribution of secondary mets. For example, patients with initial adrenal mets were more likely to develop a secondary brain met (OR 1.43, p < 0.001) compared to another site. Two phenotypes of patients with high met burden (HMB, ≥3 sites with frequency ≥0.5) at last follow-up were identified. HMB 1 (n = 1,556, median 4 sites) had frequent brain, bone, lung, and liver mets, whereas HMB 2 (n = 545, median 5 sites) had bone, distant lymph node, adrenal, soft tissue, and liver mets. HMB 1 patients more frequently had bone or liver avid disease at 1L start, whereas HMB 2 patients often had already developed HMB by this time. HMB 1 patients had better median rwOS (11.8 months) than HMB 2 patients (10.2 months), as well as compared to patients that had a lower met burden with bone (11.6 months) or liver avid disease (9.8 months, p < 0.05). Mutations in select genes that regulate interferon-gamma response and cytoplasm organization (KIF5B, CD74, ETV6, NCOA4) were associated with HMB 1, but not HMB 2 or site-specific disease. Conclusions: Definable patterns of metastasis are associated with rwOS in NSCLC. This study is the most comprehensive evaluation of the impact of longitudinal patterns of organ-specific spread and drivers of overall metastatic potential on survival outcomes in patients with NSCLC to date. This novel framework could serve as a crucial counseling and clinical management tool when caring for patients with advanced NSCLC.
1583 Background: Landmark trials define standard-of-care (SOC), yet the translation of trial evidence into routine practice remains poorly characterized. LLMs enable the systematic extraction of nuanced clinical discourse from unstructured EHR data previously infeasible at scale. We performed an LLM-based thematic analysis of past ASCO Plenary Session (PS) trial discussions to better understand the nature of shared decision-making in routine care and characterize themes influencing SOC adoption. Methods: We designed a retrospective study of clinician-documented discussions with patients about PS trials presented at ASCO Annual Meetings 2021-2025. 2 trials (negative study; non-therapeutic trial) were excluded. Patient records from the Flatiron Health Database with mention of a relevant trial/NCT number within 2 years of presentation were selected. We prompted an LLM (Claude 4.5 Sonnet) to summarize documentation of clinician-patient trial discussions and extract: 1) context of discussion, 2) clinician sentiment toward trial (positive, neutral, negative) and 3) documented receipt of trial therapy (yes/no; if no, reason why not). To characterize the gap in SOC adoption, a thematic analysis was performed using a multi-model consensus approach (Gemini 2.5pro, GPT-4o, Claude 4.5 Sonnet) to identify EHR-documented discussion themes associated with trial treatment omission. LLMs provided supporting evidence for all answers; a human-in-the-loop approach was used to verify concordance between LLM-extracted themes and expert clinical interpretation. LLMs were hosted on Flatiron Health private servers and HIPAA compliant. Results: The study included 8650 patients and 21 trials discussed across 15 cancer types. Common discussion topics were: efficacy (78%), eligibility (78%), patient education (63%), and toxicity (28%). Clinician sentiment across studies was favorable (47%-93% positive), yet real-world trial therapy initiation occurred in only 64% of cases. Main themes associated with trial treatment omission were: clinical factors (biomarker ineligibility, comorbidities), systemic barriers (insurance, transportation, incarceration), and patient preferences (prioritizing quality of life, preserving life roles). Conclusions: This is the largest thematic analysis of clinician-patient trial discussions to date. We uncovered rich qualitative insights into how clinicians engage in shared decision-making with their patients, balancing positive evidence and sentiment with practical and patient-centered factors that define individualized care. Novel LLM methods can transform clinician-patient narratives into actionable evidence to mitigate disparities and optimize real-world SOC adoption. Future work will examine thematic differences by cancer type and prevalence and correlate findings with outcomes.
e20523 Background: First-line treatment options for advanced non-small cell lung cancer (aNSCLC) with PD-L1 TPS ≥50% include pembrolizumab or pembrolizumab plus platinum-doublet chemotherapy. Without head-to-head randomized data, optimal patient selection remains unclear. We hypothesized that machine learning–predicted baseline prognosis modifies chemotherapy benefit, with higher-risk patients more likely to benefit from chemotherapy’s rapid clinical effects. Methods: Using the Flatiron Health Research Database, we identified patients with aNSCLC and PD-L1 CPS TPS ≥50% treated with first-line pembrolizumab or pembrolizumab plus chemotherapy. A gradient-boosted survival model predicted 6-month survival from baseline clinical variables (demographics, cancer features, ECOG, laboratory values, and comorbidities) using cross-validation, then calibrated via isotonic regression. Heterogeneity of absolute treatment benefit was evaluated using overlap-weighted regression with 2-year restricted mean survival time (RMST) pseudo-observations. We summarized the continuous treatment-effect function using a crossover point, defined as the baseline 6-month survival probability at which the estimated RMST benefit of adding chemotherapy reached a clinically meaningful magnitude (≥30 days). Patients were stratified by this survival probability, and survival was compared between treatments using inverse probability of treatment weighting (IPTW). Results: Among 1,434 eligible patients, 930 received pembrolizumab and 504 received pembrolizumab plus chemotherapy. Median age was 71 years, 53% were male, and median follow-up was 38 months. The model achieved 6-month AUC 0.76 with good calibration (Brier score 0.17). Chemotherapy benefit increased as baseline predicted survival worsened: for every 10 percentage-point decrease in predicted 6-month survival, patients gained 17 days in 2-year RMST with combination therapy (p = 0.051). The crossover point corresponded to a baseline 6-month survival probability of 70%. Patients below the crossover survival probability (30.8%)—characterized by worse ECOG, weight loss, bone and liver metastases, anemia and hypoalbuminemia—derived significant benefit from adding chemotherapy in the IPTW-adjusted survival analysis, while those above (69.2%) showed no benefit (Table 1). Conclusions: A machine learning model trained on nationally-representative data identified subgroups of patients with PD-L1-high aNSCLC who benefit from adding chemotherapy to pembrolizumab. RMST differences stratified by crossover survival probability. 6-month Survival <70% 6-month Survival ≥70% 1-year RMST Δ 57.9 (22.7-88.7) 2.0 (-15.7-20.4) 2-year RMST Δ 93.7 (22.9-155.1) 3.6 (-36.2-44.3) RMST differences (Pembro+Chemo - Pembro) in days; 95% CI in parentheses.
e20517 Background: Digital twin models (DTMs) developed on real world data can estimate a patient’s response to treatments they did not receive. These counterfactual (CF) predictions can help reduce a trial’s sample size without compromising its statistical power. Validation of DTM-based CF predictions remains limited. We evaluated CF outcome predictions in NSCLC by benchmarking DTM estimates against observed outcomes from IMpower131 and 132. 1 We then assessed potential trial sample size reductions enabled by DTMs. Methods: We trained machine learning (ML) models (penalized logistic regression [LR], XGBoost [XGB], MLP, GAT) to predict real-world overall survival (rwOS) for patients with stage IV NSCLC initiating 1L platinum chemotherapy (chemo) 2011-2016 from the Flatiron Health Research Database. Features were generated using structured and LLM-extracted clinical details (e.g. Charlson comorbidity index [CCI]; sites of metastases [SOM]). Models were internally validated on a held-out test set and externally validated (MAD 2 , rwOS and hazard ratio [HR] comparisons) using patient-level data from IMpower131 and 132. Predictions were made for stage IV patients in 1) the control arm to compare model predictions with observed trial outcomes (Table, Control Arm Validation) and 2) the experimental arm to simulate their outcomes as if they had instead received control-arm therapy (Table, CF Treatment Effect). We estimated trial sample size reduction based on the C-index from Cox Proportional Hazards models with model-based prognostic scores for covariate adjustment. Results: Key features across models included ECOG status, albumin, Brain+liver SOM, and CCI. LR and XGB performed best for IMpower131 and 132 respectively, with MAD 2 < 5% and predicted (pred) control arm rwOS similar to those observed (obs) in the trials 3 . CF predictions yielded HRs consistent with trial results (Table). Estimated sample size reduction was 9-15% and 15-21% for IMpower131 and 132 respectively. Conclusions: Validated DTMs demonstrate the feasibility of accurately predicting CF outcomes with advanced ML. This approach can enable efficient, well-powered trials with smaller sample sizes and supports a path toward broader validation and regulatory adoption. IMpower 131 IMpower 132 Control Arm Validation Pred vs Obs Median rwOS (months) 10 vs 12.6 3 15 vs. 13.1 3 MAD 2 3.9% 1.8% C-Index 0.61 0.67 HR (Pred:Obs; 1 is perfect prediction) 0.97 [0.86-1.00] 4 1.03 [1.00-1.06] 4 CF Treatment Effect Obs trial HR 0.86 (0.73-1.02) 3 0.82 (0.66-1.01) 3 DTM-derived HR 0.81 (0.79-0.84) 4 0.85 (0.83-0.86) 4 1 IMpower131 & IMpower132: atezolizumab plus chemo vs chemo alone in squamous and non-squamous NSCLC respectively. 2 Mean absolute difference between pred & obs OS curves. 3 Observed results differ slightly from the published trials as this analysis restricted to subset of stage IV patients. 4 95% CI based on 1000 bootstrapped samples of trial data.
Background: Palbociclib (PAL), the first CDK4/6i, in combination with endocrine therapy (ET) was approved for HR+/HER2- advanced/metastatic breast cancer (MBC) in 2015. Two additional CDK4/6is, Ribociclib (RIB) and Abemaciclib (ABE), were approved in 2017. CDK4/6i combination therapy has become standard of care for 1st line HR+/HER2- MBC. Randomized clinical trials (RCT) demonstrated that the 3 CDK4/6is plus ET vs ET plus placebo all significantly prolonged patients’ progression free survival (PFS, primary endpoint). However, the 3 CDK4/6is have inconsistent 1st line RCT overall survival findings (OS, secondary endpoint). In the absence of head-to-head RCTs, real-world data (RWD) is an important complementary source of evidence. Several small RWD studies have evaluated the relative effectiveness between CDK4/6is, and their findings are inconsistent. Large RWD studies are needed to understand the effectiveness of the 3 CDK4/6is. This study compared OS of 1st line PAL vs RIB and ABE plus AI for HR+/HER2- MBC in routine US clinical practice. Methods: We conducted a retrospective comparative effectiveness study of CDK4/6is plus AI in HR+/HER2-MBC using the US nationwide Flatiron Health electronic health record (EHR)-derived deidentified Panoramic database, comprised of >650k patients with breast cancer. Patients included had HR+/HER2- MBC, were ≥18 years, started index treatment (PAL+AI, RIB+AI, or ABE+AI) as 1st line therapy within 90 days of MBC diagnosis between February 2015 and September 2023 (index period), and did not participate in clinical trials. Patients were assessed from start of index treatment to March 2024, death, or last medical activity, whichever came first. OS was defined as months from start of index treatment to death. Patients were balanced via exact 1:1 matching on age, gender, race and ethnicity, practice type, disease stage at initial diagnosis, ECOG, time from initial to MBC diagnosis, visceral disease, bone-only disease, and number of metastatic sites. Kaplan-Meier method and Cox proportional hazard regression model were used to analyze OS. Results: Of 9770 patients eligible for the analysis, 7563, 1130, and 1077 patients received PAL+AI, RIB+AI, and ABE+AI, respectively. Median follow-up was 32.9 months for PAL+AI, 16.5 months for RIB+AI, and 21.3 months for ABE+AI treated patients. Compared with RIB and ABE groups, PAL group was 1-2 years older and had a lower proportion of patients with ECOG=0. After 1:1 matching, baseline demographics and clinical characteristics were well balanced between PAL vs RIB pairs (n=942) and PAL vs ABE pairs (n=857). In PAL-RIB pairs, 2- and 3-year OS rates were 81.3% and 67.6% for PAL group vs 79.8% and 69.7% for RIB group. In PAL-ABE pairs, 2- and 3-year OS rates were 76.8% and 63.0% for PAL group vs 75.7% and 67.4% for ABE group. Compared with PAL+AI, RIB+AI was not significantly associated with prolonged OS (unadjusted HR=0.90, 95%CI=0.80-1.02, p=0.097; matched HR=0.98, 95%CI=0.84-1.15, p=0.823). Similarly, ABE+AI vs PAL+AI was not significantly associated with prolonged OS (unadjusted HR=0.95, 95%CI=0.84-1.07, p=0.376; matched HR=0.90, 95%CI=0.77-1.06, p=0.212). Further analyses with additional follow-up, treatment duration, subsequent therapies, and time to chemotherapy will be reported. Conclusions: Our findings suggest that there is no significant OS superiority of first-line RIB+AI and ABE +AI compared to PAL+AI for HR+/HER2- MBC patients in routine clinical practice in the US. Although the sample size precludes a formal non-inferiority analysis and short follow up may limit interpretation, this study represents the largest real-world comparative analysis of OS between the CDK 4/6is in combination with AI conducted to date. Citation Format: Hope S. Rugo, Rachel M. Layman, Filipa Lynce, Xianchen Liu, Benjamin Li, Lynn McRoy, Aaron B. Cohen, Melissa Estevez, Giuseppe Curigliano, Adam Brufsky.Comparative overall survival of CDK4/6is plus an aromatase inhibitor (AI) in HR+/HER2- MBC in the US real-world setting [abstract]. In: Proceedings of the San Antonio Breast Cancer Symposium 2024; 2024 Dec 10-13; San Antonio, TX. Philadelphia (PA): AACR; Clin Cancer Res 2025;31(12 Suppl):Abstract nr PS2-03.
Large language models (LLMs) are increasingly used to extract clinical data from electronic health records (EHRs), offering significant improvements in scalability and efficiency for real-world data (RWD) curation in oncology. However, the adoption of LLMs introduces new challenges in ensuring the reliability, accuracy, and fairness of extracted data, which are essential for research, regulatory, and clinical applications. Existing quality assurance frameworks for RWD and artificial intelligence do not fully address the unique error modes and complexities associated with LLM-extracted data. In this paper, we propose a comprehensive framework for evaluating the quality of clinical data extracted by LLMs. The framework integrates variable-level performance benchmarking against expert human abstraction, automated verification checks for internal consistency and plausibility, and replication analyses comparing LLM-extracted data to human-abstracted datasets or external standards. This multidimensional approach enables the identification of variables most in need of improvement, systematic detection of latent errors, and confirmation of dataset fitness-for-purpose in real-world research. Additionally, the framework supports bias assessment by stratifying metrics across demographic subgroups. By providing a rigorous and transparent method for assessing LLM-extracted RWD, this framework advances industry standards and supports the trustworthy use of AI-powered evidence generation in oncology research and practice.
e13625 Background: Real-world evidence (RWE) is increasingly used to complement clinical trial data in oncology, providing rapid insights to inform study design and drug development. Using deep learning, natural language processing (NLP)-based machine learning models, we developed a real-world response (rwR) approach.This study evaluates the concordance between clinical trial and real-world (rw) end points in patients with stage IV non–small cell lung cancer (NSCLC) treated with first-line platinum plus pemetrexed chemotherapy. Methods: This retrospective study compared response-based outcomes generated from patients included in the control arm of IMpower132 with a trial-aligned cohort of rw patients selected from the US-nationwide Flatiron Health electronic health record (EHR)-derived deidentified database. Rw patients were aligned to key trial inclusion/exclusion criteria and further adjusted using propensity score weighting on selected baseline characteristics including demographics (eg, age, race) and clinical factors (eg, Eastern Cooperative Oncology Group [ECOG] performance status, metastatic sites). rwR was generated using NLP-based machine learning models trained on expert human-abstracted data (training set N ~12 000 patients) to extract clinician-documented change in disease burden (ie, complete response, partial response, stable disease, progressive disease, unknown) at each imaging-based disease assessment timepoint. Trial response data were captured according to a RECIST-based trial protocol. End points included response rates (rwRR vs objective response rate [ORR]), duration of response (rwDOR vs DOR), and progression-free survival (rwPFS vs PFS). Concordance was evaluated using logistic regression for response rates and Cox regression for DOR and PFS. Results: The rw cohort (N = 494) was well aligned with the clinical trial cohort (N = 275) after weighting, with standardized mean differences below 0.1 across all selected baseline characteristics. The rwRR was 34%, compared with the 38% ORR observed in the trial cohort (OR, 0.83 [95% CI, 0.54-1.27]). The median rwDOR was 5.7 months (95% CI, 4.3-7.7), compared with 6.9 months (95% CI, 4.5-8.3) for DOR in the trial cohort (HR, 1.18 [95% CI, 0.86-1.62]). rwPFS was 5.5 months (95% CI, 4.5-7.1), closely aligned with the 5.4 months (95% CI, 4.3-5.7) observed in the trial cohort (HR, 0.96 [95% CI, 0.79-1.16]). Conclusions: NLP-based ML models enable scalable and reliable generation of response-based end points from EHRs. Concordance between trial and rw end points highlights the utility of ML-driven approaches to advance RWE in oncology, particularly when leveraging large-scale, clinically rich, and well-curated training datasets.
Accurate identification of cancer progression events from electronic health records (EHRs) can help enable promising oncology applications such as predicting disease trajectory, assessing treatment efficacy, and generating real-world evidence. These use cases require both large-scale and high-quality data but manual abstraction of real-world progression (rwP) is time-intensive, difficult to scale, and inherently challenging given the varied and unstructured ways it can be documented across cancer types. Large language models (LLMs) offer a scalable alternative, but their accuracy relative to expert human abstractors is unclear. We evaluated the ability of LLMs to extract rwP events and dates across 7 cancer types and assessed how using LLM-extracted data impacted real-world progression-free survival (rwPFS) estimates compared to using human-abstracted data. We applied LLM-based extraction techniques to unstructured EHR text for 7 cancer types from the Flatiron Health Research Database: bladder (N=377), breast (N=1000), colorectal (N=564), hepatocellular (N=217), renal cell (N= 229), non–small cell lung (N=1000), and small cell lung (N=955). Prompt engineering strategies including zero-shot, few-shot, and chain-of-thought were tested to optimize performance. We measured agreement between the LLM and abstractor on the presence of rwP (Y/N) and first rwP date (± 30 days) in the first-line (1L) setting. To contextualize the LLM’s ability to extract rwP relative to an expert human abstractor, we evaluated the difference in F1 scores between both curation approaches using a duplicate human-abstracted reference dataset. We also compared rwPFS calculated from LLM-curated data versus human abstractor-curated data for 1000 patients in each cancer type, indexed to 1L start date. Across all cancer types, agreement between the LLM and abstractor on the presence of at least 1 rwP event ranged from 86%-90% while first rwP date agreement ranged from 80%-92%. The difference in F1 score between the LLM and human abstraction was within 3-8 points across cancer types. A comparison of rwPFS between the LLM and human abstractors showed <1 month difference in median rwPFS and overlapping 95% confidence intervals across all cancer types. LLMs extracted rwP with high performance, achieving F1 scores similar to expert human abstraction. Across 7 distinct cancers, agreement with human-abstracted data aligned with published inter-abstractor reliability benchmarks, and rwPFS estimates were nearly identical across curation approaches, demonstrating both the generalizability and validity of the approach. These results highlight the potential of LLMs to extract high-quality clinical endpoints at scale, helping to advance research, enhance applications such as predictive algorithms, and ultimately supporting more personalized and effective cancer care. Aaron B. Cohen, Konstantin Krismer, Kelly Magee, James Gippetti, Aaron Dolor, Tori Williams, Erin Fidyk, Hank Kim, Qianyu Yuan, Melissa Estevez. Using large language models for scalable extraction of real-world progression events across multiple cancer types [abstract]. In: Proceedings of the AACR Special Conference in Cancer Research: Artificial Intelligence and Machine Learning; 2025 Jul 10-12; Montreal, QC, Canada. Philadelphia (PA): AACR; Clin Cancer Res 2025;31(13_Suppl):Abstract nr B006.
555 Background: Emerging evidence from prospective studies underscores the prognostic potential of ctDNA to inform risk stratification and clinical decision-making in EBC (stage I-III). However, its use and correlation with outcomes in routine clinical practice remains less understood. We describe tumor-informed ctDNA testing trends and the association of test results with recurrence risk among pts with EBC in the rw setting. Methods: Pts with an EBC diagnosis (dx) after 1/1/2018 who had documented HR/HER2 status at initial dx, surgery, and ≥1 commercial ctDNA test in the early stage setting were selected using machine learning models from the Flatiron Health US-nationwide EHR-derived, deidentified, longitudinal database of >750 000 pts with BC (data cutoff 8/31/2024). ctDNA positivity (ctDNA+) was defined as having ≥1 positive test in EBC. Baseline characteristics were stratified by EBC subtype and ctDNA status. To examine the association of ctDNA status with recurrence, unadjusted Kaplan-Meier (KM) plots and adjusted Cox proportional hazards models were performed to assess recurrence-free survival (RFS) among ctDNA-tested pts, controlling for age, race/ethnicity, stage, ECOG status, dx year, insurance status, practice type, and neoadjuvant/adjuvant treatment. Recurrence was indexed to surgery date and defined as locoregional or metastatic recurrence or death. Results: In a cohort of 195 279 pts with EBC, 14 496 ctDNA tests were performed in 4639 pts (median 2 per pt) with most in stage I (43.3%) and II (37.1%). Testing prevalence was highest in HR-/HER2- (4.9%), followed by HR-/HER2+ (3.5%), HR+/HER2+ (2.9%), and HR+/HER2- (1.9%). Testing increased from 1.6% (n = 450) of EBC pts dx in 2020 to 4.25% (n = 1278) in 2023 with a decrease in median time to first test pre- and post-2022 (35 vs 8 months respectively). Among tested pts, 921 (19.9%) had ≥1 positive test and were more likely to be younger (58 vs 64 years) and have stage III disease compared to non-tested pts. ctDNA+ patients had a worse 3-year overall survival (OS) as well as a strong association with recurrence (Table). Conclusions: In the largest rw study of ctDNA testing in EBC to date, pts with ctDNA+ disease across all subtypes were more likely to recur, highlighting the potential prognostic value of ctDNA testing to inform pt counseling, monitoring and treatment strategies. These rw results, coupled with findings from prospective randomized trials, support the case for ctDNA+ as a distinct risk category in the management of EBC. Unadjusted 3-year RFS probability (95% CI) EBC Subtype ctDNA- pts ctDNA+ pts Adjusted HR (95% CI) All with P <0.01 HR+/HER2-N = 2786 0.98 (0.97-0.99) 0.76 (0.7-0.82) 10.7 (7.08-16.1) HR-/HER2-N = 1002 0.96 (0.95-0.98) 0.61 (0.52-0.72) 10.7 (6.34-18.1) HR+/HER2+N = 592 0.97 (0.95-0.99) 0.85 (0.77-0.94) 11.8 (4.54-30.8) HR-/HER2+N = 259 0.97 (0.94-1) 0.78 (0.62-0.97) 8.94 (1.72-46.4)
Background: The suitability of artificial intelligence (AI) and large language models (LLMs) to assist in curating real-world data (RWD) from electronic health records (EHR) for research holds transformative potential. Programmed death-ligand 1 (PD-L1) biomarker testing guides cancer treatment decisions, but results are hard to access because lab reports are unstructured and require clinical expertise to interpret. Additionally, results vary by cancer type, and documentation patterns have changed over time. This study explored the ability of LLMs to rapidly extract PD-L1 biomarker details from the EHR. Materials and Methods: We applied open-source LLMs (Llama-2-7B and Mistral-v0.1-7B) to extract seven biomarker details relating to PD-L1 testing from the Flatiron Health US nationwide EHR-derived database: collection/receipt/report date, cell type, percent staining, combined positive score, and staining intensity. Two approaches were used: zero-shot experiments (no fine-tuning) exploring a range of prompts and fine-tuning on manually curated answers from 500, 1000, and 1500 documents. In both cases, we validated performance using 250 human-abstracted answers across 10 cancer types. Additionally, we compared the LLM's ability to extract PD-L1 percent staining to a deep learning model baseline trained on >10,000 examples. Results: We successfully used LLMs to extract biomarker testing details from EHR documents. Fine-tuned outputs consistently conformed to the desired RWD structure. In contrast, zero-shot outputs were frequently invalid and exhibited hallucinations. Fine-tuning performance improved with additional training examples. F1 scores ranged from 0.8 to 0.95, and date accuracy (within 15 days) ranged from 0.85 to 0.9. Fine-tuned LLMs exceeded the performance of the deep learning model baseline (Delta F1 = 0.05) despite the significant difference in training data. Conclusion: LLMs, fine-tuned with high-quality labeled data, accurately extracted complex PD-L1 test details from EHRs despite considerable variability in cancer type, documentation, and time. In contrast, zero-shot prompt extraction was not effective at the model scale examined here. Validation required access to high-quality data labeled by experts with access to the source EHR.