PURPOSE:Biliary tract cancer (BTC) is the leading cause of death in patients with primary sclerosing cholangitis (PSC). PSC-related BTC is poorly understood, and the risks and benefits of conventional and immunotherapy treatments are unknown. We aimed to characterize clinical outcomes and genomes of PSC-related BTCs. EXPERIMENTAL DESIGN:This was a retrospective cohort study of patients with BTC with underlying PSC treated at MD Anderson Cancer Center (N = 46) and Princess Margaret Cancer Centre (N = 16), which were contrasted to patients with non-PSC-related BTC (N = 146). We compared outcomes between PSC and non-PSC, and PSC treated with and without immunotherapy. A combination of targeted sequencing (N = 139), whole-genome sequencing (WGS; N = 27), and WGS with paired RNA sequencing (N = 33) delineated the genomic and transcriptomic landscape of PSC-associated BTCs. RESULTS:In PSC-related BTC, the addition of immunotherapy to chemotherapy was associated with improved first-line progression-free survival (PFS; N = 22 vs. 11; median PFS, 12.2 vs. 4.7 months; P = 0.01). Immune-related adverse events were rare (N = 2, 12.5%) and improved after treatment discontinuation. Classic actionable genomic alterations, including IDH1 mutations and FGFR2 fusions, were absent in PSC-related BTCs. PSC tumors had a 2.6-fold higher tumor mutational burden (P = 3.28e-05) compared with non-PSC tumors. Transcriptomic profiling revealed a subset of PSC tumors displaying RNA signatures of immunotherapy response. CONCLUSIONS:Immunotherapy in PSC-associated BTCs seemed safe, with a potential signal of effectiveness. Given the sample size and retrospective design, these results are hypothesis-generating. Together, these results demonstrate the unique biology underlying PSC-associated BTCs, highlighting the need for prospective trials and the development of specialized treatment strategies.
Culturally and linguistically diverse (CALD) populations affected by cancer experience challenges within healthcare systems leading to inequities and disparities in care. This study aimed to describe factors that affect care for gynecologic cancer patients from CALD backgrounds by gathering perspectives from patients, interpreters, and cancer care professionals. This questionnaire-based study was conducted in the Gynecologic Oncology Clinics at Princess Margaret Cancer Center, Toronto, Canada. Study-specific questionnaires were administered to CALD patients, professional interpreters, and cancer care professionals with domains including demographics, clinic experiences, and perceived barriers and facilitators to care. Descriptive statistics summarized survey results and content analysis of free-text answers was performed. Between May 2022 and December 2023, 23 patients, 10 interpreters, and 11 cancer care professionals completed surveys. Twenty (87
Pancreatic Ductal Adenocarcinoma (PDAC) is a deadly cancer with a 5-year survival rate of ~ 13%. PDAC treatment options are limited and are often given based on patient fitness rather than molecular characteristics. Transcriptomics analysis characterizes PDAC into molecular subtypes which could help guide patient treatment based on tumour characteristics. The consensual classical and basal-like PDAC subtypes are characterized by the expression of pancreatic developmental genes or epithelial to mesenchymal transition genes, respectively. However, while basal-like tumors are more aggressive than classical tumors, these transcriptomic signatures may not translate to ex vivo analysis and also do not reflect the continuum of phenotypes observed in vivo. In this study, we focused on DNA methylation, a heritable and stable epigenetic mark that can be used to track cell of origin. DNA methylation patterns were assessed in 62 patient-derived organoids, originating from two clinical trial studies, using a targeted analysis involving methylation arrays. The analysis confirmed hypermethylation of basal-like PDAC compared to classical PDAC and identified four DNA methylation-based clusters independent of transcriptional subtypes. Methylation profiles aligned with biological processes linked to differentiated or poorly-differentiated subtypes and was able to further characterize PDAC heterogeneity. The identification of these DNA methylation signatures may uncover unique sensitivities for treatment and possible biological mechanisms that drive PDAC biology.
Abstract In pancreatic cancer, 50% of the patients are diagnosed at the metastatic stage due to a lack of symptoms. The overall survival rate dramatically reduces from 44% in early-stage resectable patients to 3% in metastatic cases. However, even in early-stage patients, >75% of them still recur after resection and adjuvant therapies. Therefore, there is an urgent need to develop a sensitive strategy to detect this disease earlier before it spreads. Characterizing circulating tumor DNA (ctDNA) in plasma is an effective approach to detect and monitor cancer. Whole-genome sequencing (WGS) tracks all the tumor mutations simultaneously, and has a higher sensitivity than targeted approaches, especially in samples with extremely low tumor burden. In this study, to investigate the ctDNA dynamics and early dissemination in pancreatic cancer, we established a cohort of 1,013 samples from 277 donors, including plasma WGS at 20-60x, alongside germline DNA WGS, tissue WGS and transcriptomic sequencing. Using a tumor-guided approach, we found that the early-stage resectable cases shed very little ctDNA (<1% tumor fraction, TFx). Plasma TFx levels were elevated at the metastatic stage, but were also dependent on metastatic tissue site, where patients with liver metastases had higher TFx than the ones with only non-hepatic lesions. We further discovered the tumor-intrinsic features that were related to increasing ctDNA shedding, such as whole-genome duplication (WGD), high cell cycle activity and non-glandular morphology, as well as extrinsic features related to reduced shedding, including a reactive microenvironment and B cell immunity. Our analysis on tumor clonal architecture revealed that subclonal mutations were more frequently detected than the clonal ones in plasma samples with low ctDNA levels or from early-stage patients, which strongly suggests that dissemination from early-stage primary tumors mostly derived from subclones. In the longitudinal plasma, we observed that subclones seeded metastasis years before imaging diagnosis. This study with a unique large cohort of paired tumor and plasma sequencing data provided a comprehensive insight on ctDNA dynamics, disease monitoring and early detection in pancreatic cancer. Citation Format: Yuanchang Fang, Michelle Chan-Seng-Yue, Karen Ng, Amy Zhang, Tuan Hoang, Gun Ho Jang, Sabiq Chaudhary, Catia Gaspar, Eugenia Flores-Figueroa, Daniela Bevacqua, Stephanie Ramotar, Ayelet Borgida, Shawn Hutchinson, Anna Dodd, Barbara Grünwald, Julie Wilson, Robert Grant, Erica Tsang, George Zogopoulos, Masoom Haider, Jennifer Knox, Steven Gallinger, Faiyaz Notta. Pervasive early dissemination in pancreatic cancer uncovered by tissue-paired plasma whole-genomes [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2026; Part 1 (Regular Abstracts); 2026 Apr 17-22; San Diego, CA. Philadelphia (PA): AACR; Cancer Res 2026;86(7 Suppl):Abstract nr 1126.
PURPOSE:Use of chemotherapy at the end of life (EOL) is discouraged, but evidence to guide decisions on the use of novel systemic anticancer treatment (SACT) agents is lacking. We examined trends of use among SACT types and association with health services use at the EOL. MATERIALS AND METHODS:We analyzed Canadian Ontario Cancer Registry data for adults diagnosed with solid tumors or hematologic malignancies within 5 years of death who received SACT between March 2015 and March 2021. Receipt of SACT in the last 30 days of life was categorized as chemotherapy alone, chemotherapy and immunotherapy, immunotherapy alone, and targeted therapy alone. Outcomes included high health services use, including multiple (≥2) emergency department (ED) visits, multiple (≥2) hospitalizations, or any (≥1) intensive care unit admission, and hospital deaths. Segmented linear regression estimated monthly trends; multivariable logistic regression estimated adjusted odds ratios (aORs) of outcomes for various SACT types. RESULTS:Among 68,963 patients, 18,337 (26.6%) received SACT at the EOL. From March 2015 to March 2020, use of SACT at the EOL increased (0.072% per month; P < .001), mainly driven by increased use of immunotherapy alone (0.064% per month; P < .001). Adjusted odds of high health services use and hospital death were more than two-fold greater among patients receiving SACT at the EOL (vs. none); individual aORs of high health services use and hospital death were 2.20 and 2.72 for chemotherapy alone, 2.36 and 3.10 for chemotherapy and immunotherapy, 1.92 and 2.27 for immunotherapy alone, and 1.75 and 2.37 for targeted therapy alone, respectively. CONCLUSION:Use of SACT at the EOL increased significantly over time, driven by increased use of immunotherapy. SACT use at the EOL, regardless of its type, was associated with high health services use and hospital death. Guidelines on the use of SACT at the EOL should include novel cancer treatments.
Accurate predictions of future cancer risk can increase early detection through selecting high risk individuals for screening. Existing risk prediction tools have limited predictive performance and are underutilized due to workflow disruption and reliance on structured data. We investigated whether large language models (LLMs) can predict cancer risk directly from routine free-text clinical notes recorded by primary care physicians. The dataset used for this study was from individuals aged over 18 years living in Ontario (2010–2016), obtained from ICES. ICES is an independent, non-profit research institute whose legal status under Ontario’s health information privacy law allows it to collect and analyze health care and demographic data, without consent, for health system evaluation and improvement. The development dataset consisted of 1,080 individuals selected from a cohort of 109,378 patients in Southern Ontario using stratified sampling. The testing dataset included 1,080 individuals from 135,894 patients in Toronto. A pipeline was developed using source-available LLMs to estimate cancer risk. Prompt and sampling methods were optimized on the development dataset. Prompts were designed to elicit probabilistic cancer risk estimates from progress notes. The study employed a fixed exclusion window of 365 days prior to cancer diagnosis and evaluated performance across lookahead windows of up to ten years. Lung cancer diagnosis was used as a case example to optimize the system, which was tested against the current screening eligibility guidelines based on age and pack-years of smoking. We evaluated generalizability to other geographic regions (1,080 individuals from 16,130 patients in Northern Ontario) and adapted the prompting pipeline for other cancer types (200 individuals from Southern Ontario). The best-performing model was a Mistral Small model using a multiple prompting strategy (chain-of-thought, emotion, persona). The model achieved an area under the receiver operating characteristic curve (AUROC) of 0.787 (95% CI 0.749 - 0.824) in the testing dataset in Toronto for lung cancer prediction over a 10 year lookahead period. At the equivalent specificity to current lung cancer screening eligibility (0.941, 95% CI 0.920-0.961), predictions from the system had increased sensitivity (0.421 [95% CI 0.362-0.479] compared to 0.207 [95% CI 0.158-0.254], p<0.05). The system generalized to Northern Ontario (AUROC 0.807, 95% CI 0.771-0.844). When extended to other cancer types, the system achieved AUROC of 0.780 (95% CI 0.702–0.841) for pancreatic cancer and 0.729 (95% CI: 0.728–0.730) for prediction of any cancer diagnosis. LLMs can predict future cancer risk directly from routine clinical progress notes, offering a scalable, privacy-preserving, and generalizable alternative to structured-data-based models. This approach could be readily integrated into clinical workflows to improve cancer screening and early detection. Daniel Mau, Karl Everett, Ning Liu, Jason Chai-Onn, Liisa Jaakkimainen, Anna Dodd, Spring Holter, Steven Gallinger, Rahul G. Krishnan, Kelvin Chan, Robert Grant. Large language models to predict cancer risk from free-text clinical notes [abstract]. In: Proceedings of the AACR Special Conference in Cancer Research: Artificial Intelligence and Machine Learning; 2025 Jul 10-12; Montreal, QC, Canada. Philadelphia (PA): AACR; Clin Cancer Res 2025;31(13_Suppl):Abstract nr B015.
High-throughput drug screening enables rapid testing of numerous drugs on tumor samples and has achieved breakthroughs in treatment selection for fatal diseases such as acute myeloid leukemia. This study leverages high-throughput screening to identify molecular vulnerabilities in pancreatic ductal adenocarcinoma (PDAC), a highly lethal cancer with a 5-year survival rate of just 13%. Using PDAC patient-derived organoids (PDOs), we developed a platform to screen approximately 3, 000 compounds at a single dose and 600 compounds across six doses on 120 PDOs. Additionally, we designed a tailored focused screening of 22 drugs at 20 doses, prioritizing Health Canada-approved drugs identified as promising in earlier screenings. To ensure robust data analysis, we developed an optimized bioinformatics pipeline incorporating a three-tier evaluation: Area Under the Curve (AUC), Drug Sensitivity Score (DSS), and maximum efficacy (Emax). Results from each PDO were systematically compared against the entire cohort, ensuring that sensitivity profiles reflect reliable outcomes. Additionally, we integrated drug response data with genomics and transcriptomics obtained through whole-genome and whole-exome sequencing of PDOs. We identified highly potent compounds capable of suppressing tumor growth in over 55% of PDOs, prioritizing clinically approved drugs to facilitate rapid translation to patient care. Drug sensitivity patterns from single-dose screenings were corroborated through six-dose validation and a focused screening platform. By integrating drug response data with multi-omics analyses, we uncovered patient-specific drug response profiles linked to molecular vulnerabilities. Notably, we validated a gene-drug association involving anagrelide, a selective agent that induces cytotoxicity in cancer cells with elevated phosphodiesterase PDE3A levels, highlighting the robustness of our approach. Preliminary findings demonstrate a strong concordance between PDO-derived and patient drug responses, establishing a foundation for actionable clinical insights in the ADOPT trial to guide personalized treatment strategies. In conclusion, our high-throughput drug screening platform, integrated with multi-omics analysis and robust bioinformatics, enables the identification of patient-specific molecular vulnerabilities in PDAC. This approach advances the potential for precision oncology, providing a foundation for tailored therapeutic strategies in this highly lethal cancer. Nikta Feizi, Eugenia Flores-Figueroa, Karen Ng, Zhen-Mei Liu, Farnoosh Abbas Aghababazadeh, Gun Ho Jang, Daniela Bevacqua, Stephanie Ramotar, Shawn Hutchinson, Anna Dodd, Julie Wilson, Robert Grant, Ronan Mclaughlin, Erica Tsang, Jennifer Knox, Steven Gallinger, Benjamin Haibe-Kains, Faiyaz Notta. High throughput drug screening to uncover molecular vulnerabilities in pancreatic ductal adenocarcinoma [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2025; Part 1 (Regular Abstracts); 2025 Apr 25-30; Chicago, IL. Philadelphia (PA): AACR; Cancer Res 2025;85(8_Suppl_1):Abstract nr 3161.
The cause/s of the increasing incidence of early-onset colorectal cancer (EOCRC) are unknown. Tumor mutational signatures provide a powerful genomic tool for discovering mutational processes associated with known or unknown etiologies. The aim was to identify subgroups of EOCRCs based on their tumor mutational signature profiles, then validate and genomically characterize these novel subgroups of EOCRC. Whole exome sequencing (WES) was performed on tumor and matched blood-derived DNA from 275 non-hereditary, mismatch repair proficient (MMRp) EOCRCs from the ANGELS and CCFR studies (age at diagnosis groups: 18-35yrs n=102; 36-45yrs n=128; 46-55yrs n=45). Single base substitution (SBS), insertion/deletion (ID) and doublet (DBS) tumor mutational signatures were calculated using COSMIC v3.4. An independent dataset comprising 1,716 whole-genome sequenced MMRp CRCs from the Genomics England (GEL)(including 240 EOCRCs) served as validation. Unsupervised dimensionality reduction in conjunction with hierarchical clustering was applied to identify signature-based clusters associated with common mutational processes, without reference to clinical features. Hierarchical clustering identified nine subgroups in 275 EOCRCs and five in 1,716 CRCs from GEL (when compared with the largest subgroup from each study). A subgroup defined by dominant SBS89 and DBS8 signatures was present in both EOCRCs and GEL CRCs. In the ANGELS/CCFR EOCRCs, this subgroup comprised 14% and was associated with a younger age at diagnosis (aged 18-35 (24%) vs 36-45 (12%) vs 46-55 (0%) (p=0.006)), with those born in recent decades, (<1960 (0%) vs1960-1979 (9%) vs ≥1980 (24%) (p=0.001)), with a proximal location (proximal (24%), distal (12%), rectal (11%) (p=0.02)), and with the co-occurrence of multiple polyps at diagnosis (p=1x10-9). In the GEL CRCs, the SBS89/DBS8 subgroup was similarly significantly associated with younger age of diagnosis (≤45, p=4x10-10) and patients born more recently (≥1980, p=5x10-7). The SBS89/DBS8 EOCRC subgroup was associated with BRAF p.V600E mutation (55% vs 9%, p=1x10-8)) and high doublet somatic mutation count (mean=4.5 ± 3.0 vs 0.9 ±1.5; p=1x10-13) when compared with EOCRCs without SBS89/DBS8. These results were also observed in the GEL CRCs. Genomic characterization of the SBS89/DBS8 positive CRCs in GEL demonstrated significant differences in large-scale variants, including increased loss of heterozygosity events (17.7 ± 17.3 vs 10.8 ± 13.8; p=0.004). Tumor mutational signature profiling identified a distinct subgroup of CRCs associated with young age at diagnosis, more recent birth year and unique genomic features, that is characterized by high proportions of SBS89 and DBS8 mutational patterns, both of which currently have an unknown etiology. This distinct subgroup may contribute to the recent rising incidence in EOCRC and warrants further investigation to elucidate its underlying mechanisms. Peter Georgeson, Alysha Prisc, Jihoon E. Joo, Khalid Mahmood, Romy Walker, Mark Clendenning, Julia Como, Natalie Diepenhorst, Julie McDonald, Steven Gallinger, Robert Grant, Dylan E. O'Sullivan, Darren R. Brenner, Finlay A. Macrae, Christophe Rosty, Ingrid M. Winship, Mark A. Jenkins, Daniel D. Buchanan. Mutational signature profiling identifies a distinct subgroup of early-onset colorectal cancer associated with younger age at diagnosis, recent birth year and specific genomic features [abstract]. In: Proceedings of the AACR Special Conference in Cancer Research: The Rise in Early-Onset Cancers—Knowledge Gaps and Research Opportunities; 2025 Dec 10-13; Montreal, QC, Canada. Philadelphia (PA): AACR; Clin Cancer Res 2025;31(23_Suppl):Abstract nr PR004.
Pancreatic ductal adenocarcinoma (PDA) is a highly aggressive disease with a 5-year survival rate of 12%. Patients are typically diagnosed at the metastatic stage due to lack of symptoms and current treatment options do not effectively control the disease. There is an urgent need to develop sensitive strategies to detect PDA at an early stage before it spreads. Characterizing circulating tumor DNA (ctDNA) in plasma is an effective approach to detect and monitor cancer. Typically, cancer genomes contain thousands of mutations, but targeted sequencing approaches only focus on pre-determined driver genes, which limits their performance in detecting ctDNA. Whole-genome sequencing (WGS) however, tracks all the tumor mutations simultaneously, and has a higher sensitivity than targeted approaches, especially in low tumor burden samples. In this study, we aim to use plasma WGS to build a liquid biopsy platform for PDA detection and characterization. To date, we have established the largest plasma WGS (20-60x) PDAC cohort which consists of over 250 baseline and longitudinal plasma samples from 205 PDA patients, with an additional 45 plasma samples from age-matched healthy individuals. Paired tumor tissue WGS and RNAseq are also available. We developed and validated a bioinformatics pipeline to detect cancer mutations in plasma and estimate plasma tumor fraction (TFx) for tumor burden. In the PDA cohort, plasma TFx increased with clinical stages (medians of 0.5% in resectable, 0.8% in locally advanced and 5% in metastatic), as did the ctDNA detection rate which was measured by whether the plasma TFx was significantly higher than controls (87%, 91% and 95%). Plasma TFx positively correlated with liver metastasis size (R=0.68, P<0.001), but not with the primary pancreatic tumor size. These results suggest that in PDA, metastases are the major site to shed ctDNA, but not primary tumors. We also found that liver metastasis size had a stronger correlation with plasma TFx than with serum CA19-9 (R=0.32, P=0.002). Baseline plasma TFx was associated with shorter overall survival (OS) in both resectable (P=0.013) and metastatic (P=0.022) cases. By analyzing the paired tumor WGS and RNAseq, we found that higher plasma TFx was associated with higher cell cycle, lower immunity and polyploid genomes in the tumors. No difference was found in plasma TFx between basal-like and classical PDA subtypes. In summary, the current study provides insights into the interplay between ctDNA dynamics, tumor biology and clinical features in PDA. Alongside with these results, we will provide a freely accessible computational platform for ctDNA genomic profiling, PDA early detection and progression monitoring using plasma. Yuanchang Fang, Michelle Chan-Seng-Yue, Karen Ng, Amy Zhang, Tuan Hoang, Gun Ho Jang, Sabiq Chaudhary, Eugenia Flores-Figueroa, Daniela Bevacqua, Stephanie Ramotar, Ayelet Borgida, Shawn Hutchinson, Anna Dodd, Julie Wilson, Barbara Grünwald, Robert Grant, Erica Tsang, George Zogopoulos, Jennifer Knox, Masoom Haider, Steven Gallinger, Faiyaz Notta. Plasma whole-genome sequencing to monitor pancreatic cancer [abstract]. In: Proceedings of the AACR Special Conference in Cancer Research: Advances in Pancreatic Cancer Research—Emerging Science Driving Transformative Solutions; Boston, MA; 2025 Sep 28-Oct 1; Boston, MA. Philadelphia (PA): AACR; Cancer Res 2025;85(18_Suppl_3):Abstract nr B072.
Pancreatic ductal adenocarcinoma (PDA) is a highly aggressive disease with a 5-year survival rate of 12%. Patients are typically diagnosed at the metastatic stage due to lack of symptoms and current treatment options do not effectively control the disease. There is an urgent need to develop sensitive strategies to detect PDA at an early stage before it spreads. Characterizing circulating tumor DNA (ctDNA) in plasma is an effective approach to detect and monitor cancer. Typically, cancer genomes contain thousands of mutations, but targeted sequencing approaches only focus on pre-determined diver genes, which largely limits its performance in detecting ctDNA. Whole-genome sequencing (WGS) however, tracks all the tumor mutations simultaneously, and has a higher sensitivity than targeted approaches, especially in low tumor burden samples. In this study, we aim to use plasma WGS to build a liquid biopsy platform for PDA detection and characterization. To date, we have established the largest plasma WGS (20-60x) cohort for PDA which consists of over 200 baseline and longitudinal plasma samples from patients, with additional 45 plasma samples from age-matched healthy individuals. Paired tumor tissue WGS and RNAseq are also available. We have developed and validated a bioinformatics tool to detect cancer mutations in plasma and estimate plasma tumor fraction (TFx) for tumor burden. In the PDA cohort, the median of plasma TFx was 1% with a wide range between statistically undetectable and 70%. Plasma TFx increased with clinical stages, from a median of 0.5% in resectable PDA to 5% in metastatic cases. Plasma TFx was positively correlated with liver metastasis size, but not with the primary pancreatic tumor size. These results suggest that in PDA, metastases are the major site to shed ctDNA, but not the primary tumors. A positive correlation between serum CA19-9 and plasma TFx was found in metastatic but not in resectable PDA. Higher baseline plasma TFx was associated with worse prognosis, which was more evident in the metastatic PDA. Using the paired tumor WGS and RNAseq, we found that higher ctDNA shedding was associated with higher cell cycle, lower immunity and polyploid genomes in the tumors. No difference was found in plasma TFx between Basal-like and Classical PDA subtypes. In summary, the current study provides insights of the interplay between ctDNA dynamics and clinical features in PDA. We aspire to deliver a freely accessible computational platform for cfDNA genomic profiling, PDA early detection and progression monitoring using plasma. Yuanchang Fang, Michelle Chan-Seng-Yue, Karen Ng, Amy Zhang, Nguyen Hoang, Gun Ho Jang, Sabiq Chaudhary, Eugenia Flores-Figueroa, Daniela Bevacqua, Stephanie Ramotar, Ayelet Borgida, Shawn Hutchinson, Anna Dodd, Julie Wilson, Barbara Grünwald, Robert Grant, Erica Tsang, George Zogopoulos, Jennifer Knox, Masoom Haider, Steven Gallinger, Faiyaz Notta. Plasma whole-genome sequencing to monitor pancreatic cancer [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2025; Part 1 (Regular Abstracts); 2025 Apr 25-30; Chicago, IL. Philadelphia (PA): AACR; Cancer Res 2025;85(8_Suppl_1):Abstract nr 3261.
Profiling colorectal cancers (CRCs) for tumor mutational signatures (TMS) offers new opportunities to characterize molecular subtypes. Recently, COSMIC published a comprehensive set of experimental mutational signatures that directly link specific environmental exposures to mutational patterns observed in human cancers. Environmental exposures are recognized as major contributors to CRC development and have been hypothesized to drive the recent rise in early-onset CRC (EOCRC), but the molecular fingerprints of these exposures in EOCRCs have not been systematically characterized. We performed whole exome sequencing (WES) on tumor and matched blood-derived DNA from 324 non-hereditary, mismatch repair proficient early-onset samples (diagnosed <55 years of age) comprising 277 CRCs and 47 pre-malignant polyps. Mutational signatures were calculated using a novel approach that combined COSMIC v3.4 signatures previously observed in CRC (n=26) and a curated set of experimental mutational signatures (n=50). Environmental signatures were filtered to human iPSC-derived signatures, excluding those with a negative AMES test, those marked as controls, and signatures not seen previously in CRC. Signature definitions with >95% cosine similarity were merged. On average, 24.7% ± 10.1% (mean ± s.d, range 0.1%-60.1%) of somatic mutations were assigned to environmental signatures. Across the cohort, 14% (46/324) exhibited a dominant environmental signature, with N-nitrosopyrrolidine being the most prevalent (6.2%, 20/324). Premalignant lesions showed higher rates of dominant environmental signatures (21%, 10/47) compared to invasive cancers, suggesting environmental exposures may be a key component in early carcinogenesis. This study provides a comprehensive view of the landscape of mutational processes in non-hereditary mismatch repair proficient EOCRC and early-onset polyps through assessment of both tumor mutational signatures and experimental mutational signatures. Environmental exposures represent a significant component of the mutational landscape in early-onset colorectal neoplasia, with enhanced prevalence in premalignant lesions. These findings support the likely role of environmental drivers in the rising incidence of EOCRC. Peter Georgeson, Alysha Prisc, Jihoon Joo, Khalid Mahmood, Romy Walker, Mark Clendenning, Julia Como, Natalie Diepenhorst, Julie McDonald, Steven Gallinger, Robert Grant, Dylan E. O’Sullivan, Darren R. Brenner, Finlay Macrae, Christophe Rosty, Ingrid M. Winship, Mark A. Jenkins, Daniel D. Buchanan. Novel insights from the investigation of experimental mutational signatures in early-onset colorectal cancer and colonic polyps [abstract]. In: Proceedings of the AACR Special Conference in Cancer Research: The Rise in Early-Onset Cancers—Knowledge Gaps and Research Opportunities; 2025 Dec 10-13; Montreal, QC, Canada. Philadelphia (PA): AACR; Clin Cancer Res 2025;31(23_Suppl):Abstract nr C020.
Pancreatic cancer is typically detected at an incurable stage because symptoms are often absent during the early stages, and the current screening guidelines lack sensitivity and specificity. Existing risk prediction tools require structured data, which hinders deployment. We evaluated whether large language models (LLMs) could predict pancreatic cancer risk, using only free-text clinical notes. We used routine free-text general practitioner clinical notes from individuals in Ontario aged >18 years, collected through ICES between 2010 and 2016. Pancreatic cancer patients were matched with controls using a nested case–control design with metrics adjusted using inverse probability weighting (IPW). Two approaches were explored: Reasoning-based LLM prediction: Source-available reasoning LLMs (DeepSeek-R1, QwQ) were prompted to simulate step-by-step clinical reasoning using raw clinical notes. Ensemble prediction: To minimize computational requirements at deployment, we tested several lightweight LLMs using different ensembling techniques, such as sampling with various decoding parameters (min-p, top-k, top-p) and LLMs, with samples aggregated using different strategies. We developed both methods using a development cohort of 200 patients (1:1 cases to controls) in Southwestern Ontario and subsequently tested them with a cohort of 750 patients (1:5) in Toronto. Look-ahead windows of five years were evaluated with a one-year exclusion period preceding diagnosis to focus on future risk and exclude patients undergoing a diagnostic work-up. The median (range) number of characters per note was 390 (20-8000), and the median number of notes per patient was 20. In the reasoning-based approach, the best-performing model in the development cohort achieved an area under the receiver operating characteristic curve (AUROC) of 0.77 (95% CI: 0.70–0.84) for predicting pancreatic cancer in the five years after each clinical note in the test cohort. Lightweight models in ensembles exhibited variable performance, depending on the strategy used. When sampling a single model with different decoding parameters, performance reached an AUROC of 0.70. Ensembling multiple models and selecting the minimal predicted score across samples yielded an AUROC of 0.75. Using the most frequently predicted score across models with different decoding parameters improved the AUROC to 0.77. A simulated screening strategy that selected the top 0.5% highest-risk individuals resulted in a relative risk of 28.1×, a specificity of 0.991, a sensitivity of 0.192, a positive predictive value of 0.025, and a negative predictive value of 0.999. LLMs can predict pancreatic cancer risk directly from clinical notes years before diagnosis, without structured inputs or pre-processing. This approach provides a scalable, generalizable, and interpretable framework for future risk prediction, potentially supporting novel population-based approaches to pancreatic cancer screening. Daniel Mau, Karl Everett, Ning Liu, Jason Chai-Onn, Liisa Jaakkimainen, Anna Dodd, Spring Holter, Steven Gallinger, Rahul G. Krishnan, Kelvin Chan, Robert Grant. Predicting pancreatic cancer risk from clinical notes using large language models [abstract]. In: Proceedings of the AACR Special Conference in Cancer Research: Advances in Pancreatic Cancer Research—Emerging Science Driving Transformative Solutions; Boston, MA; 2025 Sep 28-Oct 1; Boston, MA. Philadelphia (PA): AACR; Cancer Res 2025;85(18_Suppl_3):Abstract nr B073.
Artificial intelligence (AI) continues to advance oncology research, yet inconsistent development pipelines impair reproducibility and introduce bias. A lack of transparency in AI research1—including undisclosed preprocessing, omitted hyperparameters, and unshared code—undermines validation and trust. To address these issues, we present JARVAIS (Just A Really Versatile AI Service), an open-source Python package that standardizes machine learning (ML) workflows for oncology, improving reproducibility and mitigating bias through modular automation of data preprocessing, model training, and interpretability. JARVAIS (pmcdi.github.io/jarvais/) is composed of three core interoperable modules: (1) Analyzer for data quality audits and bias detection, (2) Trainer for automated feature selection, model training, and hyperparameter optimization, and (3) Explainer for interpretability and post-hoc fairness audits. We assessed JARVAIS by comparing it to AIM2REDUCE-ED (Emergency Department), a manually developed model predicting 30-day ED visits among gastrointestinal cancer patients on systemic therapy at Princess Margaret Cancer Centre, currently in silent deployment with strong performance2,3. JARVAIS was used to generate a model predicting the same outcome from the longitudinal retrospective de-identified Electronic Health Record data used to develop the AIM2REDUCE-ED model. The Analyzer module in JARVAIS automatically handled missing value imputation and outlier trimming (0.01–0.99 quantile) and the Trainer module used 5-fold cross validation, consistent with prior manual strategies. The JARVAIS model achieved an AUROC=0.70 and an AUPRC=0.23 on a held-out test set, closely aligning with the original manually engineered model (AUROC=0.69, AUPRC=0.20). In the prospective cohort, the JARVAIS developed model obtained an (AUROC=0.67,AUPRC=0.11) as compared to AIM2REDUCE-ED model (AUROC=0.69,AUPRC=0.16). The Explainer module also flagged patients on treatment regimens (CISPFU+TRAS, CISPCAPE+TRAS) for potential performance bias. These groups were previously identified in manual bias analysis and had worse performance in ED prediction with AUROC=0.27. Despite matching predictive accuracy, JARVAIS achieved results in under 30 minutes with minimal manual effort—automating data processing, model selection, hyperparameter tuning, validation and bias analysis. JARVAIS demonstrates strong potential as a scalable, reproducible platform for rapid ML development in oncology. Benchmarked in a real-world clinical setting, it matched the performance of a human-developed model while accelerating development and ensuring consistency. By embedding explainability and fairness into its workflow, JARVAIS supports clinician understanding of model outputs and fosters trust in downstream use, enabling more transparent and equitable AI integration in clinical practice. 1. Haibe-Kains, B. et al. Nature 586, E14–E16 (2020). 2. Grant, R. et al. JCO 41, 1557-1557 (2023). 3. Kabir, M. et al. JCO Oncol Pract 20, 407-407(2024). Joshua Siraj, Muammar Kabir, Sejin Kim, Jiang Chen. He, Baijiang Yuan, Wayne Uy, Tirth Patel, Benjamin Grant, Sharon Narine, Monika Krzyzanowska, Tran Truong, Geoffrey Liu, Clare McElcheran, Robert Grant, Mattea Welch. A modular framework to standardize machine learning workflows and accelerate reproducible AI in oncology – benchmarking against a human-developed model for predicting emergency department visits during cancer treatment [abstract]. In: Proceedings of the AACR Special Conference in Cancer Research: Artificial Intelligence and Machine Learning; 2025 Jul 10-12; Montreal, QC, Canada. Philadelphia (PA): AACR; Clin Cancer Res 2025;31(13_Suppl):Abstract nr B001.
Background Wearable digital health technologies and mobile apps (personal digital health technologies [DHTs]) hold great promise for transforming health research and care. However, engagement in personal DHT research is poor. Objective The objective of this paper is to describe how participant engagement techniques and different study designs affect participant adherence, retention, and overall engagement in research involving personal DHTs. Methods Quantitative and qualitative analysis of engagement factors are reported across 6 unique personal DHT research studies that adopted aspects of a participant-centric design. Study populations included (1) frontline health care workers; (2) a conception, pregnant, and postpartum population; (3) individuals with Crohn disease; (4) individuals with pancreatic cancer; (5) individuals with central nervous system tumors; and (6) families with a Li-Fraumeni syndrome affected member. All included studies involved the use of a study smartphone app that collected both daily and intermittent passive and active tasks, as well as using multiple wearable devices including smartwatches, smart rings, and smart scales. All studies included a variety of participant-centric engagement strategies centered on working with participants as co-designers and regular check-in phone calls to provide support over study participation. Overall retention, probability of staying in the study, and median adherence to study activities are reported. Results The median proportion of participants retained in the study across the 6 studies was 77.2% (IQR 72.6%-88%). The probability of staying in the study stayed above 80% for all studies during the first month of study participation and stayed above 50% for the entire active study period across all studies. Median adherence to study activities varied by study population. Severely ill cancer populations and postpartum mothers showed the lowest adherence to personal DHT research tasks, largely the result of physical, mental, and situational barriers. Except for the cancer and postpartum populations, median adherences for the Oura smart ring, Garmin, and Apple smartwatches were over 80% and 90%, respectively. Median adherence to the scheduled check-in calls was high across all but one cohort (50%, IQR 20%-75%: low-engagement cohort). Median adherence to study-related activities in this low-engagement cohort was lower than in all other included studies. Conclusions Participant-centric engagement strategies aid in participant retention and maintain good adherence in some populations. Primary barriers to engagement were participant burden (task fatigue and inconvenience), physical, mental, and situational barriers (unable to complete tasks), and low perceived benefit (lack of understanding of the value of personal DHTs). More population-specific tailoring of personal DHT designs is needed so that these new tools can be perceived as personally valuable to the end user.
PURPOSE OF REVIEW:Using transplant oncology principles, selected patients with intrahepatic cholangiocarcinoma (iCCA) may achieve long-term survival after liver transplantation. Strategies for identifying and managing these patients are discussed in this review.RECENT FINDINGS:Unlike initial reports, several modern series have reported positive outcomes after liver transplantation for iCCA. The main challenges are in identifying the appropriate candidates and graft scarcity. Tumor burden and response to neoadjuvant therapies have been successfully used to identify favorable biology in unresectable cases. New molecular biomarkers will probably predict this response in the future. Also, new technologies and better strategies have been used to increase graft availability for these patients without affecting the liver waitlist.SUMMARY:Liver transplantation for the management of patients with unresectable iCCA is currently a reality under strict research protocols. Who is a candidate for transplantation, when to use neoadjuvant and locoregional therapies, and how to increase graft availability are the main topics of this review.