BACKGROUND:Real-world data (RWD) in unstructured electronic health records (EHRs) is crucial for understanding complex diseases like cancer, but extracting structured information is challenging due to linguistic variability, semantic complexity, and privacy concerns. This study evaluates the performance of four small, locally deployable language models for information extraction from Italian EHRs. METHODS:We examine three prompting strategies (zero-shot, few-shot, and annotated few-shot) across English and Italian, involving clinicians with varying expertise to assess the impact of prompt design on accuracy. We evaluate the performance of four open-source small language models (SLMs) for clinical information extraction from Italian electronic health records (EHRs) in the APOLLO 11 trial on non-small cell lung cancer (NSCLC). The extraction protocol involves four steps: problem definition, data preprocessing, Large Language Model (LLM)-based information extraction, and output evaluation. RESULTS:We show that general-purpose models (e.g., LLaMA 3.1 8B) outperform biomedical models in most tasks, particularly in extracting binary features. Multiclass variables such as TNM (Tumor, Node, Metastasis) staging, PD-L1 (Programmed death-ligand 1), and ECOG-PS (Eastern Cooperative Oncology Group-Performance Status) are more difficult due to implicit language and lack of standardization. Few-shot prompting and native-language inputs significantly improve performance and reduced hallucinations. Clinical expertise enhances consistency in the extraction, particularly among students using annotated examples. CONCLUSIONS:The study confirms that privacy-preserving SLMs can be deployed locally for efficient and secure cancer data extraction. Findings highlight the need for hybrid systems combining SLMs with expert input and underline the importance of aligning clinical documentation practices with SLM capabilities. This is the first study to benchmark SLMs on Italian EHRs and investigate the role of clinical expertise in prompt engineering, offering valuable insights for the future integration of SLMs into real-world clinical workflows.
Supplementary Table 1 lists HCC case counts per UKB assessment center used in the study. Supplementary Table 2 lists the assessment-visit variables extracted and harmonized in UKB. Supplementary Table 3 reports PheWAS results used to derive disease-level features in UKB. Supplementary Table 4 lists the diagnosis-based inclusion criteria defining the PAR cohort in UKB. Supplementary Table 5 lists all single ICD code features used across UKB and AOU. Supplementary Table 6 lists grouped ICD (phecode/category) features used across UKB and AOU. Supplementary Table 7 maps self-reported conditions to ICD-based features in UKB. Supplementary Table 8 summarizes at-risk subset definitions and corresponding HCC/control sample sizes in UKB. Supplementary Table 9 provides the UKB ICD code summary representations. Supplementary Table 10 lists blood count and biochemistry features used for modeling in UKB and AOU. Supplementary Table 11 reports UKB blood count/biochemistry summary statistics stratified by HCC status and sex. Supplementary Table 12 lists genomic features (variants, coding, and availability) used in UKB modeling. Supplementary Table 13 lists metabolomics features used in UKB modeling. Supplementary Table 14 reports UKB metabolomics summary statistics stratified by HCC status. Supplementary Table 15 details the center-based train/test splits applied in UKB. Supplementary Table 16 provides UKB–AOU variable mappings and conversion summaries (min/max/median/mean and unit conversion where applicable). Supplementary Table 17 reports UKB threshold-independent performance metrics (AUROC, AUPRC) for all models and benchmarks. Supplementary Table 18 reports UKB threshold-dependent performance metrics across predefined operating points. Supplementary Table 19 reports UKB DeLong test results for AUROC comparisons between models/benchmarks. Supplementary Table 20 reports feature importances for Models A–E and TOP75–TOP15 in UKB. Supplementary Table 21 provides confusion matrices for selected thresholds in UKB and AOU (All and PAR where applicable). Supplementary Table 22 provides the AOU ICD code summary representations. Supplementary Table 23 reports AOU blood count/biochemistry summary statistics stratified by HCC status and sex. Supplementary Table 24 reports AOU threshold-independent performance metrics (AUROC, AUPRC) for all models and benchmarks. Supplementary Table 25 reports AOU threshold-dependent performance metrics across predefined operating points. Supplementary Table 26 reports AOU DeLong test results for AUROC comparisons between models/benchmarks. Supplementary Table 27 compares predicted risk distributions and related summary statistics across UKB and AOU. Supplementary Table 28 summarizes missingness (NA) patterns across features in AOU. Supplementary Table 29 provides an overview of all trained models, including modalities, feature sets, and evaluation cohorts (UKB/AOU).
Independent classifier validation and representative bidirectional counterfactual transformations in a multiclass setting. A, External classifier responses to counterfactual morphing. Counterfactual images were generated at increasing morphing amplitudes (α) with MoPaDi and then encoded with three foundation models (UNI2, CONCH, and Virchow2). Independent classifiers were trained on the corresponding encoders’ features extracted from all TCGA-CRC tiles and then used to predict the target probability P(target) on both the original and counterfactual tiles. The resulting change ΔP(target) reflects how strongly the morphing affected class evidence. Bars show the median ΔP(target) with percentile-based variability across tiles. B, Representative examples of bidirectional counterfactual explanations for MSIL patients. We defined MSIL by fitting a two-component Gaussian mixture model to the log-transformed distribution of total MSI events and using the intersection of the two components as the cutoff separating MSIL from MSIH samples.
Rapid technological progress now enables large-scale generation of single-cell data. Many laboratories can produce single-cell transcriptomic profiles from diverse tissues. A key step in single-cell analysis is unsupervised clustering followed by cell-type annotation, yet there is no agreement on marker genes, and annotation is typically done manually, making it irreproducible and poorly scalable. Privacy constraints in human datasets further complicate data sharing. There is a need for standardized, automated, and privacy-preserving cell-type annotation across datasets. We developed SwarmMAP, which applies Swarm Learning to train machine-learning models for cell-type classification in a decentralized setting without exchanging raw data between centers. SwarmMAP achieves F1-scores of 0.93, 0.98, and 0.88 in heart, lung, and breast datasets, respectively. Swarm Learning models reach an average performance of 0.907, comparable to models trained on centralized data (p-val = 0.937, Mann-Whitney U Test). Increasing the number of datasets improves prediction accuracy and supports classification across broader cell-type diversity. These results show that Swarm Learning provides an effective approach for automated cell-type annotation. SwarmMAP is available at https://github.com/hayatlab/SwarmMAP.
Abstract Background: HER2-targeted therapies are guided by HER2 protein overexpression or ERBB2 amplification. However, these biomarkers may imperfectly reflect functional dependence on HER2 signaling. We performed an integrative multi-omic analysis to quantify the discordance between clinical HER2 status and its corresponding functional biology, assessing both downstream pathway activation in HER2-positive (HER2+) tumors and inactivity in HER2-negative (HER2−) cases. Methods: Analysis was performed on the TCGA breast cancer dataset. HER2+ (n=106) and HER2− (n=595) cohorts were defined using clinical HER2 status. Transcriptomic and proteomic over-/under-expression states were derived for clinically relevant driver genes by thresholding RNA/protein expression values. For each tumor, we assessed activation of HER2-associated pathways (e.g., direct HER2 targets, PI3K-AKT-mTOR, MAPK, RTK crosstalk) by comparing expression states against the expected direction of activation. A pathway was considered 'validated' when at least half of the tested states aligned with expected activation patterns. Samples with no available data were excluded from the analysis, reducing the evaluable cases to 85 and 457 for HER2+ and HER2−, respectively. Results: Among HER2+ tumors, proximal HER2 signaling showed strong concordance with clinical status, with direct targets being active in 80.0% of cases. In contrast, canonical downstream pathways were often inactive, with PI3K-AKT-mTOR and MAPK activity validated in only 29.4% and 15.3% of tumors, respectively. This heterogeneity suggests that many HER2+ tumors may rely less on canonical HER2-driven signaling than expected. In HER2− tumors, proximal HER2 signaling and MAPK pathway activity were appropriately low as validated in >96% cases. However, the PI3K-AKT-mTOR pathway was active in 132/457 (28.9%) HER2− tumors, revealing a sizable subset with HER2-independent PI3K activation. Conclusions: Our findings highlight biological mechanisms that may underlie therapeutic resistance and reveal alternative targets for expanding precision treatment strategies across breast cancer subtypes. The limited canonical signaling in most HER2+ tumors suggests pathway-independent drivers, feedback inhibition, or gaps in transcriptomic/proteomic measurements that miss critical phosphorylation events. In contrast, activation signatures in a subset of HER2− tumors indicate therapeutically relevant HER2-low biology. The results support complementing standard testing with biology-aware signatures to (i) flag HER2+ tumors unlikely to respond to combination therapies and (ii) stratify HER2− patients into actionable groups, including those who may benefit from PI3K/mTOR inhibitors. Future work will incorporate phospho-proteomic data and link pathway signatures to treatment response to refine patient selection. Citation Format: Salim Arslan, Julian Schmidt, Cher Bass, Foivos Ntelemis, J Carl Barrett, Oscar Maiques, Jakob Nikolas Kather, Pahini Pandya. Functional pathway analysis reveals discordance between clinical HER2 status and downstream effector activation in breast cancer [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2026; Part 1 (Regular Abstracts); 2026 Apr 17-22; San Diego, CA. Philadelphia (PA): AACR; Cancer Res 2026;86(7 Suppl):Abstract nr 57.
‘Black box’ deep learning models for medical image interpretation limit clinical trust and analysis of performance degradation. Here we introduce Concept-Level Embeddings for Auditable Radiology (CLEAR), an auditable foundation model based on clinical concepts. Trained on over 0.87 million image–report pairs from 239,391 patients, CLEAR learns a visual representation and projects chest X-rays into a semantically rich space defined by large language model embeddings, making every prediction decomposable into weighted contributions from individual radiological observations. External validation on four large, physician-annotated datasets from the United States, Europe and Asia shows that CLEAR not only achieves state-of-the-art classification performance but also enables applications: auditable zero-shot pathology detection, systematic identification of radiological confounders and the creation of expert-level concept bottleneck models from data-driven concepts. By integrating clinical knowledge directly into its reasoning process, CLEAR offers a framework for robust model auditing, safer deployment and enhanced physician–AI collaboration, advancing towards trustworthy medical AI. CLEAR is an auditable foundation model for chest X-rays that leverages the collective knowledge of the radiological community to offer improved interpretability, performance and new applications.
Background Increasing evidence shows that intermuscular adipose tissue (IMAT) and lean muscle mass (LMM) influence cardiometabolic health; however, their independent and/or combined associations with cardiovascular risk in individuals without pre-existing conditions remain unclear. Purpose To assess whether IMAT and LMM are associated with cardiometabolic risk factors in individuals without pre-existing conditions. Materials and Methods A total of 11 348 participants (6460 [56.9%] men; median age, 43.0 years; IQR, 33.5-52.5 years) without any known pre-existing conditions underwent whole-body 3-T MRI as part of a prospective multicenter population study (German National Cohort, or NAKO). LMM and IMAT were quantified on MRI-based paraspinal muscle segmentations with a deep learning model. Cardiometabolic risk factors (hypertension, dysglycemia, and atherogenic dyslipidemia) were defined on the basis of laboratory test results and clinical examinations. Age- and sex-corrected z scores of LMM and IMAT were calculated. Associations of LMM and IMAT percentage with physical activity and cardiometabolic risk factors were examined with univariable and multivariable analyses. Results The percentage of IMAT increased with age and was greater in women, whereas LMM decreased with age and was lower in women. After adjustments for age, sex, and study site, increased IMAT was associated with increased odds of hypertension (odds ratio [OR], 1.67; 95% CI: 1.49, 1.86; P < .001), atherogenic dyslipidemia (OR, 1.82; 95% CI: 1.65, 2.00; P < .001), and dysglycemia (OR, 0.51; 95% CI: 0.35, 0.76; P = .009) in both sexes, whereas increased LMM was associated with decreased odds of all risk factors (dysglycemia: OR, 0.51; 95% CI: 0.35, 0.76; P = .009; atherogenic dyslipidemia: OR, 0.49; 95% CI: 0.39, 0.62; P < .001; hypertension: OR, 0.34; 95% CI: 0.24, 0.48; P < .001) in male participants only. Across z score combinations, participants with higher IMAT and lower LMM showed the highest prevalence of cardiometabolic risk factors. Conclusion IMAT and LMM, assessed on MRI scans, were independently associated with cardiometabolic risk factors in individuals without pre-existing conditions. © RSNA, 2026 Supplemental material is available for this article. See also the editorial by Hu in this issue. See also the editorial by Mohajer and Bari in this issue.
Spatial biomarkers are critical for precision oncology but remain challenging to systematically discover due to the complexity of whole-slide images. We present PathPrism, an interpretable AI framework for spatial biomarker discovery and virtual experimentation. Unlike black-box models, PathPrism encodes tissue architecture into pathologically informed spatial features, enabling transparent modeling of prognosis, molecular alterations, and therapy response. Applied to 7,000 patients with colorectal cancer across 11 cohorts, PathPrism uncovered hundreds of biomarkers predictive of survival, MSI, BRAF, and TP53 mutations, and stratified chemotherapy benefit in stage II/III disease. Building on these interpretable findings, PathPrism uses large language models as auxiliary tools to generate hypotheses grounded in spatial semantics. We further introduce VirtualWSI, a platform for semantic perturbation within an interpretable spatial biomarker atlas. PathPrism provides a scalable and interpretable framework for spatial biomarker discovery.
Large language models (LLMs) have transformative potential in radiology, including textual summaries, diagnostic decision support, proofreading, and image analysis. However, the rapid increase in studies investigating these models, along with the lack of standardized LLM-specific reporting practices, affects reproducibility, reliability, and clinical applicability. To address this, reporting guidelines for LLM studies in radiology were developed using a two-step process. First, a systematic review of LLM studies in radiology was conducted across PubMed, IEEE Xplore, and the ACM Digital Library, covering publications between May 2023 and March 2024. Of 511 screened studies, 57 were included to identify relevant aspects for the guidelines. Then, in a Delphi process, 20 international experts developed the final list of items for inclusion. Items consented as relevant were summarized into a structured checklist containing 32 items across six key categories: general information and data input; prompting and fine-tuning; performance metrics; ethics and data transparency; implementation, risks, and limitations; and further/optional aspects. The final FLAIR (Framework for LLM Assessment in Radiology) checklist aims to standardize reporting of LLM studies in radiology, fostering transparency, reproducibility, comparability, and clinical applicability to enhance clinical translation and patient care. © The Author(s) 2026. Published by the Radiological Society of North America under a CC BY 4.0 license. Supplemental material is available for this article.
Abstract Background: Metabolic reprogramming is a hallmark of cancer and a potential therapeutic target, yet clinical assessment remains challenging. We hypothesized that functional metabolic states could be inferred from routinely collected H&E-stained whole-slide images (WSIs), offering a scalable approach to biomarker discovery and patient stratification. We developed a multi-omics platform integrating histomorphology with transcriptomic and metabolic modelling data, applying it across 21 cancer types to identify digital biomarkers predictive of pathway-specific metabolic activity. Methods: We derived binary biomarkers for 32 pathways using personalized genome-scale metabolic models and transcriptomic data from The Cancer Genome Atlas (TCGA). Pathway activity was inferred by comparing sample-specific gene expression profiles to a generalized metabolic network, with statistical significance determined by t-tests and Benjamini-Hochberg correction (FDR < 0.05). Deep learning models were trained to learn the morphological features predictive of these metabolic states, with cohort sizes ranging from 41 to 887 WSIs per biomarker. Model performance was validated using three-fold cross-validation and evaluated by the area under the curve (AUC), with ± showing standard deviation across folds. Results: The approach identified robust morphological patterns predictive of key metabolic pathways across diverse tumor types. High predictability was achieved for nucleotide metabolism in testicular germ cell tumors (AUC = 0.85 ±0.13) and pancreatic adenocarcinoma (AUC = 0.79 ±0.08). Strong morphological signals of nuclear transport were observed in colon adenocarcinoma (AUC = 0.79 ±0.05), skin cutaneous melanoma (AUC = 0.75 ±0.06), and lung squamous cell carcinoma (AUC = 0.71 ±0.07). Pathways related to fatty acid beta-oxidation showed consistent predictability (AUC > 0.65) across thyroid carcinoma (AUC = 0.73 ±0.07), liver hepatocellular carcinoma (AUC = 0.70 ±0.05), pancreatic adenocarcinoma (AUC = 0.69 ±0.02), stomach adenocarcinoma (AUC = 0.68 ±0.04), and ovarian serous cystadenocarcinoma (AUC = 0.67 ±0.05). Conclusions: This study demonstrates the potential of a multi-omics platform for decoding functional metabolic states from routine pathology slides. The platform can identify patient populations with specific metabolic dependencies (e.g., nucleotide metabolism) or pathway dysregulations (e.g., nuclear transport), thereby enabling targeted patient stratification for clinical trials. By translating complex histomorphology from standard H&E slides into actionable molecular insights, this technology offers a scalable and cost-effective solution to accelerate biomarker-driven drug discovery in precision oncology. Citation Format: Salim Arslan, Julian Schmidt, Cher Bass, Foivos Ntelemis, Oscar Maiques, Vishali Sharma, Jakob Nikolas Kather, Pahini Pandya. A translational deep learning platform predicts metabolic pathway activity from routine H&E slides to enable patient stratification [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2026; Part 1 (Regular Abstracts); 2026 Apr 17-22; San Diego, CA. Philadelphia (PA): AACR; Cancer Res 2026;86(7 Suppl):Abstract nr 6726.
Abstract Deep learning can extract predictive and prognostic biomarkers from histopathology whole-slide images. However, explainable artificial intelligence approaches widely used in digital pathology, such as attention heatmaps and class activation mapping, provide limited insight into the image features associated with classifier outputs. In this study, we developed Morphing histoPathology Diffusion (MoPaDi), a framework for generating counterfactual explanations for histopathology images that help identify morphologic or stain-related features linked to model predictions. MoPaDi combined diffusion autoencoders with task-specific multiple instance learning classifiers to manipulate images and induce prediction shifts by modifying classifier-associated features. The framework was evaluated on multiple datasets spanning colorectal, breast, liver, and lung cancers, including tasks for tissue type, cancer subtype, and biomarker [microsatellite instability (MSI)] classification. MoPaDi generated perceptually realistic counterfactual histopathology images, enabling pathologists to identify morphologic features associated with changes in model predictions, complementing the conventional inspection of highly attended regions in digital pathology. In the MSI status prediction task, MoPaDi highlighted morphologic features linked to classifier predictions, including mucinous differentiation, altered glandular architecture, and lymphocytic infiltration, consistent with prior literature. Analyses separating stain-related from morphology-related components suggested that in this setting, prediction changes were predominantly associated with morphology-related rather than stain-related alterations. Overall, MoPaDi is a practical framework for counterfactual explanations in computational pathology that supports the evaluation of model-specific decision cues and hypothesis generation. Significance: MoPaDi is a diffusion-based tool for counterfactual image generation in cancer histopathology that reveals features associated with deep learning classifier predictions and supports transparent auditing of computational models in biomedical research.
Style–morphology decomposition for disentangling structural and staining effects on MSI prediction changes. A, Representative examples of counterfactual manipulation between MSIH and non-MSIH classes. For each original image x and its counterfactual xcf, style-hybrid (xstyle) and morphology-hybrid (xmorph) images were generated using Vahadane stain transfer. Each hybrid isolates the effect of either stain or morphology while controlling for the other. The right-hand bars show the Shapley-style decomposition of the logit change (Δf) into stain (φstyle) and morphology (φmorph) contributions, demonstrating that morphologic differences dominate the model’s predictions. Grad-CAM visualizations below provide region-level attribution under MIL. In contrast, MoPaDi produces class-directed “what-if” edits that offer a complementary view of candidate morphologic and style changes associated with prediction shifts. Scale bar applies to all images within the panel. B, Decomposition results across test-set patients, showing median contributions of φstyle, φmorph, and total (Δf) for manipulations toward (↑) and away from (↓) each class. C, Scatter plot of morphology versus style contributions per patient, illustrating consistent dominance of morphologic effects across both manipulation directions and classes.
The rapid adoption of transformer-based models in computational pathology has enabled prediction of molecular and clinical biomarkers from H E whole-slide images, yet interpretability has not kept pace with model complexity. While attribution- and generative-based methods are common, feature visualization approaches such as class visualizations (CVs) and activation atlases (AAs) have not been systematically evaluated for these models. We developed a visualization framework and assessed CVs and AAs for a transformer-based foundation model across tissue and multi-organ cancer classification tasks with increasing label granularity. Four pathologists annotated real and generated images to quantify inter-observer agreement, complemented by attribution and similarity metrics. CVs preserved recognizability for morphologically distinct tissues but showed reduced separability for overlapping cancer subclasses. In tissue classification, agreement decreased from Fleiss k = 0.75 (scans) to k = 0.31 (CVs), with similar trends in cancer subclass tasks. AAs revealed layer-dependent organization: coarse tissue-level concepts formed coherent regions, whereas finer subclasses exhibited dispersion and overlap. Agreement was moderate for tissue classification (k = 0.58), high for coarse cancer groupings (k = 0.82), and low at subclass level (k = 0.11). Atlas separability closely tracked expert agreement on real images, indicating that representational ambiguity reflects intrinsic pathological complexity. Attribution-based metrics approximated expert variability in low-complexity settings, whereas perceptual and distributional metrics showed limited alignment. Overall, concept-level feature visualization reveals structured morphological manifolds in transformer-based pathology models and provides a framework for expert-centered interrogation of learned representations across label granularities.