
PURPOSE:Most low-dose computed tomography (LDCT) scans in lung cancer (LC) screening are negative, yet contain prognostic information. We developed a fully automated 3D deep learning model for thoracic body composition (BC) quantification from LDCT and evaluated its prognostic value for overall and cause-specific mortality. MATERIALS AND METHODS:Baseline and first follow-up LDCTs from the National Lung Screening Trial were analyzed. A 3D convolutional neural network was trained for full thoracic volumetric segmentation of subcutaneous adipose tissue (SAT) and skeletal muscle (SM) and extraction of attenuation values. BC metrics were categorized into sex-specific percentile groups (<20%, 20%-80%, >80%). Kaplan-Meier and Cox regression assessed the association between baseline BC/1-year BC changes and all-cause mortality (primary end point) and atherosclerotic cardiovascular disease (ASCVD) and LC mortality (secondary end points) adjusted for age, sex, ethnicity, BMI, smoking status, diabetes, hypertension, and history of stroke or heart disease. RESULTS:Among 23,195 participants (mean age 61.4 ± 5 years; 41.5% women), 1,601 (6.9%) deaths occurred over 6.5 years. Low SM volume and density were independently associated with increased all-cause mortality (adjusted hazard ratio [aHR], 1.40 [95% CI, 1.25 to 1.57]; aHR, 1.80 [95% CI, 1.60 to 2.02]; both P < .001). Similar associations were found for low SAT volume, high SAT density, and low SM density with ASCVD and LC mortality. A >10% 1-year decline in SAT or SM measures further increased mortality risk, strongest for reduced SAT density (aHR, 2.13 [95% CI, 1.54 to 2.95]; P < .001). CONCLUSION:Automated 3D thoracic BC quantification from LDCT enables opportunistic, artificial intelligence-driven risk stratification beyond traditional factors, supporting personalized prevention in high-risk populations.
PURPOSE:Radiology and pathology reports are important in breast surgical oncology planning, but are often unstructured and variable, requiring time-intensive previsit review and preparation. Here we evaluate the accuracy, completeness, and safety of retrieval-augmented large language model (LLM)-generated structured preconsult summaries compared with clinical staff-authored summaries. METHODS:In this single-institution retrospective validation study, 200 randomly sampled new breast surgical oncology consultations were included. LLM summaries were generated and compared with human summaries documented in consultation notes. Field-level accuracy, omission rate, and fabrication rate among extracted radiology and pathology variables were measured. Error analysis included patient level and clinically significant fabrication rates. RESULTS:Artificial intelligence (AI)-generated summaries demonstrated high field-level accuracy, with fabrication rates ≤2%. Accuracy rates were significantly higher than those observed in staff-authored notes across multiple domains, including nodal status (95% v 72%, P < .001) and documentation of invasive tumor component on biopsy (94% v 30%, P < .001). In contrast, staff-authored summaries were more accurate for receptor status (96% v 86%, P = .002). At the patient level, fabrication occurred in 16 AI-generated summaries (8%), including six clinically significant cases (3%). Staff-authored summaries demonstrated fabrication in six cases (3.0%), with one clinically significant instance (0.5%). CONCLUSION:In this retrospective validation study, structured retrieval-augmented LLM-generated preconsult summaries demonstrated field-level accuracy comparable with or exceeding staff-authored documentation across most radiologic and pathologic variables, with low per-field fabrication rates. Clinically significant errors occurred in a small proportion of summaries and were largely confined to identifiable high-risk variables that are amenable to targeted verification. With appropriate human oversight and safeguards, LLM-based structured information extraction may support documentation standardization and improved workflow efficiency in breast surgical oncology.
PURPOSE:Machine learning models that predict survival time for patients are increasingly used for clinical decision support. Validating model performance in deployment is important but challenging because the only timely source of follow-up/death data is the electronic medical record (EMR), which is known to undercapture deaths, resulting in informative censoring. We examined whether model performance evaluation using EMR data alone can distinguish between low- and high-quality models using high-quality cancer registry data combined with EMR data as a comparison to calculate model performance. MATERIALS AND METHODS:This was a retrospective study of 3,330 patients with metastatic cancer diagnosed from 2008 to 2018. We used regularized Cox proportional hazards regression trained on features from the EMR to predict length of survival after diagnosis. We trained models by varying the number of features to span a range of baseline discrimination performance and validated performance first using higher-quality EMR and cancer registry outcome data as a reference standard, followed by EMR data alone to simulate the scenario when real-time validation is performed and only EMR data are available. RESULTS:The model with all features had a C-index of 0.66 (95% CI, 0.65 to 0.68) and an integrated Brier score (IBS) of 0.17 (95% CI, 0.16 to 0.18) when validating with reference standard data, compared with a C-index of 0.67 (95% CI, 0.65 to 0.69) and an IBS of 0.16 (95% CI, 0.15 to 0.17) with EMR data only. When using fewer features, performance estimates dropped similarly in both scenarios. When using reference standard data, the model was found to be reasonably calibrated, but with EMR data, the model was incorrectly found to systematically underpredict survival. CONCLUSION:EMR data were useful for validation of model discrimination, but not calibration.
PURPOSE Predictors of optimal or improved health-related quality-of-life (HRQOL) in adult survivors of childhood cancer are understudied. METHODS This cohort study used data from 4,755 survivors in the Childhood Cancer Survivor Study who completed baseline (T0) and two follow-up surveys (T1, T2) between 1992 and 2016. HRQOL was assessed using SF-36 Physical and Mental Component Summaries (PCS; MCS), classified as optimal (≥40) or suboptimal (<40), with improvement being suboptimal at T1 to optimal at T2. RESULTS At T1, 88.4% and 82.5% had optimal physical and mental HRQOL; among those suboptimal, 51.3% and 63.9% improved by T2. Higher education, physical activity, income, mental health, and absence of chronic conditions predicted optimal or improved HRQOL using multivariable logistic regression with backward selection. Model performance was acceptable for optimal PCS (0.81), MCS (0.72), and improved models (0.69 each). CONCLUSION Findings may inform targeted interventions addressing education, lifestyle, mental health, and chronic conditions to enhance well-being in childhood cancer survivors.
Despite vast investments in data collection, generation, and sharing from childhood cancer studies, for example, Gabriella Miller Kids First Research Program, Childhood Cancer Data Initiative, and other efforts, secondary use of data remains a challenge especially for rare diseases. The National Cancer Institute (NCI) Office of Data Sharing (ODS) aimed to promote FAIR (Findable, Accessible, Interoperable, Reusable) data sharing practices to enhance data utility, accelerate discovery, and foster interdisciplinary collaboration. NCI ODS launched its inaugural Data Jamboree focusing on childhood cancer alongside its third Annual Symposium in September 2025. Over 120 participants with diverse backgrounds coalesced into 23 multidisciplinary teams to solve real-world problems in a collaborative setting. The Jamboree built a diverse pool of data users and catalyzed pediatric cancer research, with several project findings being prepared for publication. Moreover, key feedback and recommendations for areas of improvement were provided to NCI in real time. The event highlighted the power of team science, data sharing, use, and reuse to accelerate childhood cancer research and improve diagnostic and therapeutic outcomes.
PURPOSE:AML is a highly aggressive hematologic cancer. During induction chemotherapy, up to 40% of patients experience complications resulting in unplanned readmissions or early death. This study aimed to identify the reasons for postinduction unplanned readmissions and to develop predictive models for unplanned readmissions or early death using structured and unstructured electronic health record (EHR) data to identify patients at highest risk. METHODS:We retrospectively analyzed 1,111 inpatient encounters from 305 adult patients with AML treated at a Midwestern university hospital between 2006 and 2021. Inclusion criterion was adults with AML undergoing induction chemotherapy; exclusion criteria included acute promyelocytic leukemia, stem cell transplant recipients, and confirmed chronic myeloid leukemia. Adverse events-unplanned readmissions or mortality within 30 days of discharge from the initial induction hospitalization-were identified through chart review using predefined rules. Multiple logistic regression was used to identify risk factors. Variable selection was performed using the least absolute shrinkage and selection operator. RESULTS:Within 30 days postdischarge, 22% of patients experienced an unplanned readmission. The most frequent reasons were fever/infection (53.5%), metabolic/GI/renal complications (16.2%), and pain/discomfort (14.1%). Mortality within 30 days was 22%. Predictive performance improved when comorbidities and symptom frequency from clinical notes were added (AUC increased from 0.67 to 0.74). Statistically significant predictors of increased adverse event risk included solid neoplasm (odds ratio [OR], 2.34), higher cardiopulmonary symptom burden per day (OR, 1.16), and a chemotherapy intensity moderated by age (OR, 0.90). CONCLUSION:Unplanned readmissions or early mortality occurred in 39.7% of patients with AML within 30 days postdischarge. Integrated risk stratification tools leveraging structured and unstructured EHR data may inform timely interventions to improve survival and reduce avoidable hospitalizations.
PURPOSE:Tumor budding (TB) is an independent prognostic biomarker for colorectal cancer (CRC), yet its clinical adoption is limited by labor-intensive and subjective visual scoring. Computational pathology (CPath) algorithms using deep learning offer potential to improve efficiency and reproducibility. Using an international Delphi study, we established consensus requirements to guide validation and clinical adoption of CPath algorithms, aiming to enhance diagnostic accuracy and patient care. METHODS:A two-round international Delphi process was conducted, involving international experts, to reach consensus on predefined statements. In round 1, baseline characteristics were collected and participants were asked to rank statements representing the minimal requirements for the implementation of a CPath algorithm for TB assessment. After completion of this round, participants received a personalized feedback report summarizing interim results for items that lacked consensus. In round 2, participants re-evaluated their initial responses to statements without consensus, in the light of anonymized group feedback provided in their personalized report. RESULTS:Fifty-nine pathologists participated in round 1, with a 90% response rate in round 2. Consensus was reached for 21 of 29 (72%) minimal requirements. All reached consensus by agreement. Eight statements (28%) remained without consensus after round 2. CONCLUSION:Agreement was reached on the technical, organizational, ethical, and legal aspects, although some requirements were not agreed upon. This highlights the importance of evaluations that take specific contexts into account, as well as the need for continuous stakeholder engagement throughout the development and implementation phases.
PURPOSE:To evaluate open-source large language models (LLMs) for extracting cancer-specific phenotypic data, benchmark their performance against GPT4 models, and assess the impact of fine-tuning with training data sizes. METHODS:Open-source LLMs (Mistral, LLaMa, MAMBA, BioMistral) were evaluated in zero-/one-shot and fine-tuned setups against GPT4-turbo/GPT4o to extract the cancer presence, progression, response, and metastatic sites from radiology impressions of patients with solid tumors treated at Dana-Farber Cancer Institute. Performance metrics (accuracy, precision, recall, F1-score) were computed. McNemar's odds ratio (OR), measuring which model is more likely to be correct when they disagree, was computed with 95% CI. Statistical significance was assessed using the alpha of .000139. RESULTS:This study included 2,623 patients (25,273 radiology impressions). In zero-/one-shot settings, GPT4-turbo/GPT4o outperformed open-source LLMs. However, fine-tuned open-source LLMs achieved higher F1-scores than GPT4 models. Compared with the best-performing GPT4 model, fine-tuned Mistral0.2-7.3B (OR, 0.27 [95% CI, 0.20 to 0.36]; P < .00001), Mistral0.3-7.3B (OR, 0.26 [95% CI, 0.19 to 0.36]; P < .00001), LLaMa2-6.7B (OR, 0.30 [95% CI, 0.22 to 0.40]; P < .00001), LLaMa3.1-8B (OR, 0.37 [95% CI, 0.28 to 0.48]; P < .00001), and MAMBA-2.8B (OR, 0.32 [95% CI, 0.24 to 0.42]; P < .00001) showed significantly better performance in ascertaining disease progression. Performance was consistently better for inferring overall response, any evidence of cancer, and sites of metastases, with no significant differences among fine-tuned open-source LLMs. Fine-tuning gains plateaued at 25% of training data (5,718 impressions) and remained comparable at 5% (1,144 impressions). CONCLUSION:Open-source LLMs, when fine-tuned using labeled data, can effectively automate the ascertainment of key radiophenotypic variables using only the impression section of radiology reports, without the full report text. Their consistent performance in small training sets suggests that these models may provide a scalable approach for phenotypic characterization of patients with cancer in real-world clinical settings.
PURPOSE:Guidelines recommend germline testing in advanced prostate cancer (PCa) to inform treatment and personal/familial cancer risk, yet Black patients are less likely than White patients to complete testing. Artificial intelligence (AI) tools are increasingly used in pretest education to improve access but remain unevaluated for racial bias. MATERIALS AND METHODS:We developed ProGene, a secure, generative AI chatbot for PCa germline testing education in Black patients. We prompted ProGene with seven questions about testing types, personal benefits, family benefits, drawbacks, logistics, costs, and privacy. Each question was asked nine times across three patient vignettes (Black, non-Hispanic White, race-agnostic), in triplicate. Two blinded reviewers assessed responses across five domains: (1) comprehensiveness (0%-100%), (2) accuracy (presence/absence of inaccuracies), (3) readability (grade level via Simple Measure of Gobbledygook [SMOG] and Flesch-Kincaid), (4) actionability (0%-100% via Patient Education Materials Assessment Tool), and (5) quality (1-18 via DISCERN-AI). Outcomes were compared by race and question using two-sample t-tests or Wilcoxon rank-sum tests for continuous measures and chi-square or proportion tests for categorical measures; ANOVA was used for question-level comparisons. RESULTS:The mean comprehensiveness was 67% and did not vary by race, although responses for question-1 (genetic testing types) were less comprehensive for Black versus race-agnostic vignette (60% v 93%, P < .01). Inaccuracies appeared in 32% of responses, primarily related to sample collection and cost/insurance, and did not vary by race. The mean readability were 10th (SMOG) and 13th (Flesch-Kincaid) grades, actionability 92%, and DISCERN score 14/18 (good quality); none varied by race. CONCLUSION:We identified no statistically significant racial disparities across the five evaluation domains in ProGene. Responses were generally good quality and actionable. Although AI may facilitate equity in PCa genetic education, deficiencies in comprehensiveness, accuracy, and readability highlight the need for refinement and/or human oversight with implementation.
PURPOSE Several in silico models concluded that ultra-high dose rate (UHDR) radiotherapy could spare large quantities of circulating lymphocytes. However, preclinical studies failed to show a reduction in radiation-induced lymphopenia after UHDR irradiation compared with conventional dose rates (CONV). This study aims to investigate the influence of fractionation and beam sequencing on absorbed dose to circulating lymphocytes during CONV and UHDR irradiations. MATERIALS AND METHODS The LymphoDose framework was applied to a cohort of 162 patients treated for brain tumors with 3D conformal radiotherapy. Four scenarios of UHDR treatment were compared: (S1) one fraction with all beams delivered simultaneously, (S2) one fraction with sequentially delivered beams, (S3) three fractions with all beams delivered simultaneously, and (S4) three fractions with sequentially delivered beams. RESULTS UHDR fractionation and the beam delivery scheme had a significant impact on irradiated blood volume (1.3% ± 0.1% for S1 v 13% ± 3.2% for S4) and lymphocyte pool (12.5% ± 0.1% for S1 v 32.8% ± 0.2% for S4). UHDR scenarios primarily irradiate lymphocytes through exposure of head-and-neck lymph node rather than circulating blood. CONV and UHDR result in distinct temporal dose patterns for lymphocytes, characterized by either numerous low-dose fractions or a few high-dose pulses. CONCLUSION The doses delivered to lymphoid organs account for a substantial portion of the total dose received by the lymphocyte pool, showing only a limited difference between UHDR and CONV in terms of lymphocyte exposure. The fractionation strategy of UHDR irradiation beams could play an important role in successfully translating UHDR treatments into clinical practice.
PURPOSE:Enzalutamide or abiraterone, combined with androgen-deprivation therapy, is standard of care for metastatic castration-sensitive prostate cancer (mCSPC). However, no trials have compared these drugs. This study compared clinical outcomes in patients with mCSPC treated with enzalutamide or abiraterone. METHODS:This retrospective cohort study included patients with mCSPC initiating enzalutamide or abiraterone between January 1, 2020, and December 31, 2023, within the nationwide US Veterans Affairs health care system. Inverse probability of treatment weighting was used to balance baseline characteristics. Restricted mean survival time (RMST) differences in overall survival (OS), time to treatment switch or death (TTS), and prostate cancer survival (PCS) were evaluated. RESULTS:The study included 5,135 patients with mCSPC treated with enzalutamide (1,803) or abiraterone (3,332). The median age was 74.33 years; 58.0% were non-Hispanic White, 28.2% non-Hispanic Black, and 5.6% Hispanic. After weighting, baseline characteristics were well balanced. The median follow-up was 18.74 months for abiraterone and 24.76 months for enzalutamide. Outcomes were similar overall: for OS, the 3-year RMST difference was 0.72 months (95% CI, -0.06 to 1.50); for TTS, the 3-year RMST difference was 0.53 months (95% CI, -0.45 to 1.51), and for PCS, the 1-year RMST difference was -0.12 months (95% CI, -0.35 to 0.11). In subgroup analysis, enzalutamide was associated with improved OS among patients age 75 years and older (3-year RMST difference: 1.65 months, 95% CI, 0.41 to 2.89), but not among younger patients (3-year RMST difference: 0.13 months, 95% CI, -0.86 to 1.11). CONCLUSION:In this nationwide cohort study, enzalutamide and abiraterone yielded comparable OS, TTS, and PCS outcomes overall, although a small but statistically significant OS benefit was observed for enzalutamide among older patients (≥75 years). These real-world findings from the largest integrated US health care system may provide guidance for selecting mCSPC treatments, although residual confounding cannot be fully excluded.
PURPOSE:Accurate identification and segmentation of abdominal lymph nodes on computed tomography (CT) is crucial for cancer staging and treatment planning but remains a challenging and time-consuming radiologic task. We developed and validated ConLymphNet, a deep learning-based approach for automated abdominal lymph node segmentation using vertebral landmarks as anatomic reference points to standardize the region of interest. MATERIALS AND METHODS:We implemented a novel preprocessing pipeline incorporating vertebral landmark-based region selection followed by a 3D nnUNet segmentation model. The model was trained on 481 contrast-enhanced CT scans from patients with gallbladder cancer (GBC) and validated on external data sets including GBC (n = 54), public lymph node (n = 81), and multicancer data sets (n = 606). Performance was compared with three radiologists of varying experience levels. Interobserver agreement among radiologists and the incremental value of artificial intelligence (AI)-assisted reading were also assessed. To support reproducibility and future research, we also contribute 45 expert-verified lymph node segmentations for cases from public multicancer data sets. RESULTS:ConLymphNet achieved mean Dice coefficients of 0.696 ± 0.179, 0.631 ± 0.187, and 0.639 ± 0.176 and low false-positive rates of 0.78 ± 0.35, 0.81 ± 0.33, 0.93 ± 0.47/volume on internal cross-validation, external GBC, and public lymph node data sets, respectively. Performance varied by lymph node location (P = .0080) and size (P = .0444), with larger nodes yielding better results. The model's detection rate (89.6%) was comparable with experienced residents (89.2%) but with reduced processing time (28.7 ± 12.4s v 165.3 ± 47.8 s). AI assistance improved radiologists' performance significantly, particularly for less experienced readers (detection rate +20.2%, segmentation accuracy +13.0%, P < .001), while reducing false positives by 33.3%. CONCLUSION:ConLymphNet demonstrates reasonable performance across diverse cancer types and imaging protocols, offering segmentation accuracy comparable with radiologists with substantially faster processing times.
PURPOSE:To develop a large language model (LLM) (Truveta Language Model Oncology [TLM-Oncology]) to extract real-world oncology staging data across multiple cancer types from clinical documentation with high precision. METHODS:We selected patients from a large integrated health system with a bladder, cervical, colorectal, breast, or prostate cancer diagnosis in their structured data. We identified relevant notes using note metadata and keywords and annotated overall stage; T, N, and M; associated timeframe; and cancer diagnosis on a sample of 700 notes as ground truth. Of the 700 notes, 450 were divided equally between training, validation, and test sets for bladder, cervical, and colorectal cancers; 150 were used for targeted error-pattern training on these cancers; and the remaining 100 were split equally between breast and prostate cancer test sets. We started with a pretrained LLM and applied supervised fine-tuning to adapt the model to structured clinical information extraction. Model performance was measured using precision, recall, and F1 scores at the relation level and individual attribute level. RESULTS:We extracted over 2.5 million staging records for 217,768 patients from over two million notes. Relation-level precision across the six attributes ranged from 0.77 to 1.0 for the first three cancers and, without further training, 0.83 to 1.0 for two additional cancers. CONCLUSION:TLM-Oncology extracted detailed cancer staging information for five cancers from a variety of clinical documentation within a single integrated health system with high precision and turned data that were previously inaccessible into a valuable resource for downstream use. We are currently evaluating TLM-Oncology on other solid tumors within three additional health systems to assess its generalizability.
PURPOSE Survivorship care models that extend beyond recurrence surveillance to ones that also address treatment-related symptoms are needed. Using data obtained in routine clinical care, we aimed to examine patient-reported symptom burden among adult thyroid cancer survivors. METHODS Between September 2019 and September 2022, adults were electronically administered the MDASI-Thy, a patient-reported outcome measure that measures symptom severity and interference, within 7 days before their visit at a dedicated thyroid cancer survivorship clinic. The MDASI-Thy generates (1) core symptom severity, (2) thyroid-specific symptom severity, and (3) symptom interference scores, where lower is better. High alert values (HAVs) were defined for four symptoms: Distress (Upset), Pain, Sad, and Shortness of Breath. Multivariable generalized linear models examined associations of patient, cancer, and treatment factors with scores and any HAV. RESULTS Among 1,557 thyroid cancer survivors, 864 (55.5%) responded. Respondents were a median of 5 years from diagnosis (IQR, 4-8) and predominantly female (79.1%), and most had papillary thyroid carcinoma (92.5%). Mean (standard deviation) scores were 1.20 (1.34) for core severity, 0.99 (1.26) for thyroid-specific severity, and 1.07 (1.80) for interference. Fatigue (11.5%) and Disturbed Sleep (11.3%) were the most common severe symptoms. HAVs occurred in 72 survivors (8.3%), of whom 54 (75%) had a documented plan addressing the HAV. Higher symptom burden was associated with female sex, Black race, greater comorbidity, active smoking, and total thyroidectomy. CONCLUSION Routine patient-reported symptom screening in thyroid cancer survivorship identified generally low symptom burden but meaningful variations, with a subset reporting severe symptoms, functional interference, and HAVs requiring action.
PURPOSE JAVELIN Renal 101 revealed that avelumab plus axitinib improved progression-free survival (PFS) compared with sunitinib in advanced renal cell carcinoma. However, Japanese subgroup estimates were imprecise. In this study, we evaluated how Bayesian borrowing assumptions translate into posterior probabilities for this subgroup. PATIENTS AND METHODS Using published summary data from JAVELIN Renal 101, we modeled PFS on the log hazard-ratio scale with a normal-normal framework, comparing a weakly informative neutral prior with informative borrowing priors centered on the global final estimate (with sensitivity to the third interim estimate). Outputs were posterior HRs, 95% credible intervals (CrIs), P (hazard ratio [HR] <1), and P (HR < 0.8). Sensitivity analyses evaluated alternative borrowing specifications. Adverse events (AEs) were assessed with a no-borrowing beta-binomial framework using noninformative Jeffreys priors for Japanese and global data sets, reporting Japanese posterior means, 95% CrIs, and probabilities of enrichment for any grade and grade ≥3. RESULTS Under the neutral prior, evidence for PFS benefit in the Japanese subgroup was moderate ( P [HR < 1] = .83; P [HR < 0.8] = .64), whereas informative borrowing priors yielded substantially higher probabilities. Borrowing in PD-L1 + Japanese patients narrowed CrIs while preserving the favorable direction of effect. Any-grade AEs in Japanese patients were most often enriched in hepatic/laboratory, thyroid, infusion-related, and hand-foot domains. For grade ≥3, enrichment mainly involved hepatic/laboratory abnormalities, while rare severe events yielded imprecise posteriors. CONCLUSION A paired Bayesian strategy, graded priors for efficacy and no-borrowing priors for safety, produced transparent posterior probability summaries for the Japanese subgroup. This framework illustrates how borrowing can narrow uncertainty when exchangeability is plausible, while clarifying toxicity domains that may warrant closer monitoring, using published trial data.
PURPOSE Recurrent cancers are not captured in a standardized way by US tumor registries, making it difficult to conduct research on risk factors for cancer recurrence. We developed rule-based algorithms to be used with electronic health data to identify recurrent cases of diffuse large B-cell lymphoma (DLBCL) and follicular lymphoma (FL). METHODS Incident DLBCL and FL cases (2000-2018) were identified in tumor registry data at two health plan study sites. We captured pharmacy and procedure codes to indicate first-line treatment initiation. Recurrent cases were defined as those who completed first-line treatment followed by ≥6 months with no treatment-related codes, but who later restarted treatment. The baseline algorithm was built using a claims-based database from Fallon Health (FH; Massachusetts) and tested using electronic health records and claims data at Henry Ford Health (Michigan). Results were validated by chart review at Henry Ford, and measures of validity calculated overall and by subtype. The algorithm was subsequently revised to reduce the false-positive rate. RESULTS FH identified 137 DLBCL and 88 FL eligible cases; 42 patients met the baseline algorithm-defined criteria for recurrent disease. Henry Ford identified 246 DLBCL and 146 FL cases. The baseline algorithm identified 115 recurrent cases with a 54% false-positive rate; the revised algorithm (R2D-non-Hodgkin lymphoma [NHL]) identified 60 recurrent cases, with a 10% false-positive rate. Following chart review, the R2D-NHL algorithm had a sensitivity of 74%, specificity of 90%, negative predictive value of 83%, and positive predictive value of 83%. Measures varied slightly between subtypes. CONCLUSION We developed a rule-based algorithm that can be applied to electronic health data for population-based research requiring the identification of recurrence for two common but dissimilar NHL subtypes.
PURPOSE:We evaluated whether offering access to a multicomponent mHealth app improves quality of life (QoL) and psychosocial outcomes among breast cancer survivors under pragmatic, nonprescriptive conditions. METHODS:In this single-center, randomized, controlled trial at Hospital Clínic de Barcelona, women age ≥18 years, disease-free after breast cancer treatment, were recruited (December 2020-December 2021) and randomly assigned 1:1 to usual follow-up plus app access or usual follow-up alone. The app provided CTCAE v4.03-aligned symptom tracking with self-care guidance, educational content, an events calendar, and gamified smartphone-based step counting; no protocolized clinician monitoring or feedback was provided. Outcomes were assessed at baseline and 3, 6, 9, and 12 months using European Organisation for Research and Treatment of Cancer-Quality of Life Questionnaire (QLQ)-C30/BR23, Hospital Anxiety and Depression Scale (HADS), and Three-Item Loneliness Scale (TILS). The primary end point was the difference in QLQ-C30 Global Health Status/QoL at 3 months. Analyses followed intention-to-treat using mixed models for repeated measures adjusted for baseline values. RESULTS:Of 124 women assessed, 121 were randomized (intervention n = 60; control n = 61). Patient-reported outcome measures were available for 106 of 121 (87.6%) at 3 months and 95 of 121 (78.5%) at 12 months. At 3 months, there was no significant difference in Global Health Status/QoL (adjusted mean difference [Intervention-Control], -2.24 [95% CI, -9.29 to 4.81]; P = .53); estimates at later time points were similarly imprecise. No significant between-group difference were observed for QLQ-BR23 domains, HADS anxiety/depression, or TILS. Exploratory subgroup analyses suggested possible heterogeneity in TILS by hormonal-treatment category; this was descriptive and hypothesis-generating only. App engagement was the highest in months 0-3 (48/60 [80.0%] with any use) and declined thereafter; 12 of 60 (20.0%) never used the app. CONCLUSION:In a pragmatic, nonprescriptive survivorship trial, offering access to a multicomponent mHealth app without closed-loop clinical integration did not show a statistically significant between-group differences in QoL or psychosocial outcomes; confidence intervals were compatible with meaningful harm and did not exclude small benefit depending on the threshold used to define clinical relevance.
PURPOSE The rapidly evolving breast cancer treatment landscape creates significant information synthesis challenges for clinicians. We evaluated whether small open-source large language models (LLMs) augmented with retrieval-augmented generation (RAG) could match proprietary model performance for clinical guideline queries. METHODS We developed a domain-specialized RAG pipeline using HTML-structure-preserving chunking of 1,356 ASCO breast cancer guideline documents. Five LLMs were each evaluated with and without RAG: GPT-4-turbo, GPT-3.5-turbo, Qwen2.5-14B (14 billion parameters), LLaMA3-8B, and OpenBioLLM-8B. Performance was assessed using 98 expert-curated question-answer-context triplets across seven breast cancer categories. Evaluation used both rubric-based scoring (six metrics: fluency, relevance, reliability, consistency, clarity, and clinical impact) and exhaustive pairwise ranking by GPT-4-turbo as judge. Human validation was conducted with 15 practicing oncologists on a 10-query subset. RESULTS RAG-enhanced Qwen2.5-14B achieved mean rubric scores of 3.77 versus 3.96 for GPT-4-turbo and pairwise ranking performance of 0.72 versus 0.81 (normalized scale). Although absolute rubric gains were modest (0.02-0.05 on a five-point scale), relative improvements in head-to-head win rates ranged from 16% to 46%. Human expert scores confirmed RAG superiority but were consistently more conservative than LLM judge scores (mean 3.81 v 4.12 across all metrics). Optimal retrieval used top-5 contexts; performance degraded sharply at higher context volumes. CONCLUSION Small open-source LLMs with optimized RAG can approach state-of-the-art proprietary model performance for clinical decision support. This approach enables scalable, cost-effective, privacy-preserving deployment without recurrent fine-tuning, suggesting potential for real-world clinical implementation on single-graphics processing unit infrastructure under expert supervision.