Precision therapy for liver cancer necessitates accurately delineating liver sub-regions to protect healthy tissue while targeting tumors, which is essential for reducing recurrence and improving survival rates. However, the segmentation of hepatic segments, known as Couinaud segmentation, is challenging due to indistinct sub-region boundaries and the need for extensive annotated datasets. This study introduces LiverFormer, a novel Couinaud segmentation model that effectively integrates global context with low-level local features based on a 3D hybrid CNN-Transformer architecture. Additionally, a registration-based data augmentation strategy is equipped to enhance the segmentation performance with limited labeled data. Evaluated on CT images from a development cohort with additional external validation, LiverFormer demonstrated high accuracy and strong concordance with expert annotations across multiple metrics. These results establish LiverFormer as a reliable and standardized Couinaud segmentation framework, providing a methodological basis for future studies related to surgical and radiotherapy planning.
Recent breakthroughs in artificial intelligence through foundation models and agents have accelerated the evolution of computational pathology. Demonstrated performance gains reported across academia in benchmarking datasets in predictive tasks such as diagnosis, prognosis, and treatment response have ignited substantial enthusiasm for clinical application. Despite this development momentum, real world adoption has lagged, as implementation faces economic, technical, and administrative challenges. Beyond existing discussions of technical architectures and comparative performance, this review considers how these emerging AI systems can be responsibly integrated into medical practice by connecting deployable clinical relevance with downstream analytical capabilities and their technical maturity, operational readiness, and economic and regulatory context. Drawing on perspectives from an international group, we provide a practical assessment of current capabilities and barriers to adoption in patient care settings.
Neuroblastoma is a leading cause of childhood cancer mortality, presenting persistent management challenges due to its biological heterogeneity and the limited accessibility of molecular profiling in routine practice. Here we present NEVA (NEuroblastoma Vision-language AI), a multimodal foundation model designed to address these barriers. Unlike conventional approaches that rely on frozen encoders and multiple instance learning, NEVA implements a pathologist-inspired hierarchical workflow with end-to-end optimization. Developed and evaluated in a large multi-institutional cohort of 1,238 patients across multiple centers, NEVA outperforms ten representative foundation models, including TITAN, UNI, and Virchow, across the majority of the 11 clinical tasks evaluated. The model demonstrates diagnostic capability, achieving Area Under the Receiver Operating Characteristic curves of 0.916 for subtype classification, 0.823 for Shimada classification, and 0.806 for risk group stratification. Furthermore, NEVA predicts key molecular alterations from routinely available pathology data, reaching an AUROC of 0.924 for NMYC amplification and 0.830 for 1p36 deletion, while enabling prognostic stratification for progression-free and overall survival across multiple test cohorts. By integrating interpretable attention maps that localize histologically relevant regions, NEVA establishes a scalable framework for neuroblastoma risk stratification and clinical decision support.
Neuroblastoma treatment planning in children relies on accurate risk stratification, but access to specialized pathology assessment and molecular testing remains limited in low- and middle-income countries. We present ANCHOR, a neuroblastoma-adapted pathology foundation model trained by self-supervised adaptation on more than 15 million H&E image patches and evaluated in 1,551 patients from six centers across two continents. ANCHOR outperformed general-purpose pathology foundation models across clinically relevant diagnostic and biomarker tasks, including histological subtype classification, Shimada classification, mitosiskaryorrhexis index grading, and prediction of risk-associated molecular alterations such as MYCN amplification, 1p36 deletion, and 11q23 deletion. Further, ANCHOR improved INRG pretreatment risk assessment and prognostication by integrating H&E-derived features with diagnostic predictions, molecular biomarker predictions, and patient age, achieving up to a 12% improvement over H&E-only features. In a randomized crossover reader study involving 15 pathologists from six centers, AI assistance increased balanced accuracy by 5%, reduced mean diagnostic time by 42%, and improved inter-rater agreement. Together, these results position ANCHOR as a scalable pathology-support framework for neuroblastoma, with potential to improve diagnostic consistency, inform molecular testing priorities, and strengthen risk-informed treatment planning from routine H&E slides.
Cancer survival prediction through multimodal learning that combines histopathology images with genomic data represents a promising research direction. However, current approaches still suffer from two key limitations. First, most methods operate in a Euclidean feature space, which makes it difficult to capture the intrinsic hierarchies in histopathology, where information is organized from patches to whole-slide images to patients, and in genomics, where it progresses from genes to pathways to patients. Second, they typically discretize survival times into coarse risk intervals, neglecting fine-grained ordinal relationships among samples within the same interval and thus failing to capture the continuous ranking characteristics of survival outcomes.To address these issues, we propose \ourmethod, a hyperbolic hierarchical multimodal learning framework for survival prediction. H2-SurvNet first employs a hyperbolic hierarchical information modeling (H2IM) module that maps multimodal features into a shared hyperbolic space and explicitly encodes intra-modal and inter-modal hierarchies across patches, WSIs, patients, genes, and pathways. On top of this representation, we design a Temporal Ordinal Contrastive learning (TOCL) module that models the temporal progression of survival outcomes by enforcing ordinal risk ordering through contrastive objectives, thereby promoting continuity in the learned risk scores.Extensive experiments on heterogeneous cohorts from TCGA, CPTAC, and NLST demonstrate that H2-SurvNet consistently outperforms state-of-the-art multimodal survival prediction methods and exhibits strong robustness and generalization across diverse data distributions. Source code will be released upon acceptance.
In locally advanced gastric cancer, a substantial number of patients relapse with liver metastasis months after an apparently curative operation, and standard tumor staging offers little warning of who is at risk. Here, we develop the Radiopathomics-Clinical Stratification Assessment (RCSA), an interpretable model that integrates three complementary sources of information: radiomic features from preoperative computed tomography, pathomic features from routine hematoxylin and eosin tumor slides, and conventional clinical features. Trained on patients from one hospital and then tested on separate internal, external, public, and prospective trial groups (NCT02555358), RCSA consistently separates high- and low-risk patients, with area under the curves between 0.862 and 0.909. Tumors it labels low-risk carry a notably more active immune environment, indicating that these patients are the ones most likely to gain from added immunotherapy. RCSA therefore turns existing hospital data into individualized guidance for postoperative follow-up and treatment. Accurate prediction of metachronous liver metastasis (MLM) after curative surgery for locally advanced gastric cancer (LAGC) remains challenging. Here, the authors develop a multimodal model to predict MLM in LAGC patients by integrating clinical, imaging, and pathology data across multi-centre cohorts, also revealing patients who could be more responsive to adjuvant immunotherapy.
Spatial proteomics enables high-resolution mapping of protein expression and can transform our understanding of biology and disease. However, major challenges remain for clinical translation, including cost, complexity and scalability. Here we present H&E to protein expression (HEX), an AI model designed to computationally generate spatial proteomics profiles from standard histopathology slides. Trained and validated on 819,000 histopathology image tiles with matched protein expression from 382 tumor samples, HEX accurately predicts the expression of 40 biomarkers encompassing immune, structural and functional programs. HEX demonstrates substantial performance gains over alternative methods for protein expression prediction from H&E images. We develop a multimodal data integration approach that combines the original H&E image and AI-derived virtual spatial proteomics to enhance outcome prediction. Applied to six independent non-small-cell lung cancer cohorts totaling 2,298 patients, HEX-enabled multimodal integration improved prognostic accuracy by 22% and immunotherapy response prediction by 24-39% compared with conventional clinicopathological and molecular biomarkers. Biological interpretation revealed spatially organized tumor-immune niches predictive of therapeutic response, including the co-localization of T helper cells and cytotoxic T cells in responders, and immunosuppressive tumor-associated macrophage and neutrophil aggregates in non-responders. HEX provides a low-cost and scalable approach to study spatial biology and enables the discovery and clinical translation of interpretable biomarkers for precision medicine.
Early postoperative recurrence is a major cause of treatment failure in patients with locally advanced gastric cancer (LAGC), yet current staging systems inadequately capture the biological heterogeneity that underlies recurrence risk. Here, we introduce a clinically interpretable multimodal prediction model, Recurrence Stratification and Assessment (RSA), which integrates deep learning-derived histopathological features from routine hematoxylin and eosin slides with conventional clinical variables. The model was developed using a retrospective multicenter cohort (n = 1,763) and rigorously validated across two internal cohorts, two geographically distinct external cohorts, and an exploratory post-hoc analysis of a prospective clinical trial population (NCT01516944), demonstrating robust and generalizable performance (area under the curves ranging from 0.843 to 0.887). Shapley Additive Explanations-based interpretation identifies key histological features contributing to recurrence risk. To explore biological underpinnings, we perform transcriptomic sequencing and immune profiling on tumor specimens, revealing immune-enriched microenvironments and elevated checkpoint gene expression in the RSA-defined low-risk group. These findings suggest differential immunological activity may influence recurrence dynamics. This study demonstrates the application of digital pathology-based artificial intelligence for recurrence risk prediction in LAGC, offering not only a high-performance and biologically informed tool, but also a transparent framework for clinical deployment. The RSA model may support risk-adapted postoperative surveillance and provides a biologically informed framework for exploring the potential utility of immune checkpoint inhibitors.
Various forms of degradation, including noise, blur, and adverse weather conditions (e.g., rain, snow, and fog), significantly compromise video quality and system reliability across critical domains ranging from surveillance and medical imaging to entertainment. Previous research mainly focuses on network models tailored to specific degradation types, while recent unified frameworks and foundation models still face critical challenges in temporal consistency, automated degradation recognition, and detail preservation. Despite recent advances in foundation models, current approaches rely heavily on predefined degradation labels and remain focused on image-level operations, limiting their generalization to real-world scenarios and struggling with preserving fine-grained details. To address these challenges, we propose Grid Splicing Diffusion Model (GSDiff), a general framework for video reconstruction that leverages a novel grid splicing execution alongside instruction-tuned Large Language Model (LLM). GSDiff introduces three key innovative modules: (1) a LLM-driven degradation recognition module that enables automatic and fine-grained restoration guidance through zero-shot degradation analysis, (2) a Grid Splicing Module that organizes multiple frames into a unified grid structure to facilitate spatiotemporal feature processing, and (3) a Detail Preservation Module integrated with a Tail Refine Network to enhance fine-grained details during diffusion and post-processing. Extensive experiments demonstrate that GSDiff delivers state-of-the-art performance across a wide range of reconstruction tasks, including deraining, desnowing, denoising, and deblurring, propelling advancements in medical diagnostics and smart city applications.
Multimodal artificial intelligence, albeit showing great potential in computational pathology, remains limited to isolated patch-level interpretation and often fails to analyze gigapixel-scale whole-slide images (WSIs) essential for clinical utility. Here we present SlideChat, a multimodal generative artificial intelligence assistant for whole-slide computational pathology across cancer types. SlideChat integrates patch-level and slide-level pathology encoders with a pretrained large language model. Using 274,233 multimodal instruction samples, SlideChat is trained to learn the associations between WSIs and diagnostic reports and interpret complex queries in clinical practice. Evaluated on 8,836 closed-ended questions, 129 open-ended questions and 3,149 WSI reports from five cohorts spanning 31 cancer types, SlideChat outperformed leading baselines by 19.1% in closed-ended accuracy and by 7.7% in report-generation Metric for Evaluation of Translation with Explicit Ordering score and received the highest expert ratings across five dimensions in open-ended question answering, showing the potential to enhance diagnostic workflows, medical education and clinical decision-making.
Recent advancements in computational pathology have produced patch-level Multi-modal Large Language Models (MLLMs), but these models are limited by their inability to analyze whole slide images (WSIs) comprehensively and their tendency to bypass crucial morphological features that pathologists rely on for diagnosis. To address these challenges, we first introduce WSI-Bench, a large-scale morphology-aware benchmark containing 180k VQA pairs from 9,850 WSIs across 30 cancer types, designed to evaluate MLLMs' understanding of morphological characteristics crucial for accurate diagnosis. Building upon this benchmark, we present WSI-LLaVA, a novel framework for gigapixel WSI understanding that employs a three-stage training approach: WSI-text alignment, feature space alignment, and task-specific instruction tuning. To better assess model performance in pathological contexts, we develop two specialized WSI metrics: WSI-Precision and WSI-Relevance. Experimental results demonstrate that WSI-LLaVA outperforms existing models across all capability dimensions, with a significant improvement in morphological analysis, establishing a clear correlation between morphological understanding and diagnostic accuracy.
Accurate prognosis prediction is essential for guiding cancer treatment and improving patient outcomes. While recent studies have demonstrated the potential of histopathological images in survival analysis, existing models are typically developed in a cancer-specific manner, lack extensive external validation, and often rely on molecular data that are not routinely available in clinical practice. To address these limitations, we present PROGPATH, a unified model capable of integrating histopathological image features with routinely collected clinical variables to achieve pancancer prognosis prediction. PROGPATH employs a weakly supervised deep learning architecture built upon the foundation model for image encoding. Morphological features are aggregated through an attention-guided multiple instance learning module and fused with clinical information via a cross-attention transformer. A router-based classification strategy further refines the prediction performance. PROGPATH was trained on 7999 whole-slide images (WSIs) from 6,670 patients across 15 cancer types, and extensively validated on 17 external cohorts with a total of 7374 WSIs from 4441 patients, covering 12 cancer types from 8 consortia and institutions across three continents. PROGPATH achieved consistently superior performance compared with state-of-the-art multimodal prognosis prediction models. It demonstrated strong generalizability across cancer types and robustness in stratified subgroups, including early- and advanced-stage patients, treatment cohorts (radiotherapy and pharmaceutical therapy), and biomarker-defined subsets. We further provide model interpretability by identifying pathological patterns critical to PROGPATH’s risk predictions, such as the degree of cell differentiation and extent of necrosis. Together, these results highlight the potential of PROGPATH to support pancancer outcome prediction and inform personalized cancer management strategies.
Drug repurposing significantly reduces development costs and shortens research cycles, making it a critical strategy in drug discovery. An emerging class of drug repurposing approaches applies deep learning to structural data. However, these methods often depend on static representations of molecular and protein structures, which may not fully capture the dynamic character of compound-protein interactions. To address these challenges and enhance the accuracy of compound-protein interaction predictions, we introduce an innovative prompt-based multimodal representation learning framework that dynamically encodes task-specific contextual information for drug repurposing. Specifically, the framework includes a dynamic prompt generation module that adaptively creates receptor-specific prompts and a prompt calibration module for effective multimodal feature integration and optimization. When applied to identifying FDA-approved drug candidates targeting G-protein-coupled receptors, our method achieved a 7.4% improvement in mean absolute error compared with state-of-the-art methods, with up to a 25.1% improvement for specific target-of-interest. By demonstrating potential in repurposing non-opioid treatments without the risk of addiction for safe pain management, our method has the capacity to advance drug discovery and meet a wide range of therapeutic needs.
Applying deep learning to predict patient prognostic survival outcomes using histological whole-slide images (WSIs) and genomic data is challenging due to the morphological and transcriptomic heterogeneity present in the tumor microenvironment. Existing deep learning-enabled methods often exhibit learning biases, primarily because the genomic knowledge used to guide directional feature extraction from WSIs may be irrelevant or incomplete. This results in a suboptimal and sometimes myopic understanding of the overall pathological landscape, potentially overlooking crucial histological insights. To tackle these challenges, we propose the CounterFactual Bidirectional Co-Attention Transformer framework. By integrating a bidirectional co-attention layer, our framework fosters effective feature interactions between the genomic and histology modalities and ensures consistent identification of prognostic features from WSIs. Using counterfactual reasoning, our model utilizes causality to model unimodal and multimodal knowledge for cancer risk stratification. This approach directly addresses and reduces bias, enables the exploration of 'what-if' scenarios, and offers a deeper understanding of how different features influence survival outcomes. Our framework, validated across eight diverse cancer benchmark datasets from The Cancer Genome Atlas (TCGA), represents a major improvement over current histology-genomic model learning methods. It shows an average 2.5% improvement in c-index performance over 18 state-of-the-art models in predicting patient prognoses across eight cancer types.
Spatially resolved transcriptomics enable comprehensive measurement of gene expression at subcellular resolution while preserving the spatial context of the tissue microenvironment. While deep learning has shown promise in analyzing SCST datasets, most efforts have focused on sequence data and spatial localization, with limited emphasis on leveraging rich histopathological insights from staining images. We introduce GIST, a deep learning-enabled gene expression and histology integration for spatial cellular profiling. GIST employs histopathology foundation models pretrained on millions of histology images to enhance feature extraction and a hybrid graph transformer model to integrate them with transcriptome features. Validated with datasets from human lung, breast, and colorectal cancers, GIST effectively reveals spatial domains and substantially improves the accuracy of segmenting the microenvironment after denoising transcriptomics data. This enhancement enables more accurate gene expression analysis and aids in identifying prognostic marker genes, outperforming state-of-the-art deep learning methods with a total improvement of up to 49.72%. GIST provides a generalizable framework for integrating histology with spatial transcriptome analysis, revealing novel insights into spatial organization and functional dynamics.
The scarcity of well-annotated diverse medical images is a major hurdle for developing reliable AI models in healthcare. Substantial technical advances have been made in generative foundation models for natural images. Here we develop `ChexGen', a generative vision-language foundation model that introduces a unified framework for text-, mask-, and bounding box-guided synthesis of chest radiographs. Built upon the latent diffusion transformer architecture, ChexGen was pretrained on the largest curated chest X-ray dataset to date, consisting of 960,000 radiograph-report pairs. ChexGen achieves accurate synthesis of radiographs through expert evaluations and quantitative metrics. We demonstrate the utility of ChexGen for training data augmentation and supervised pretraining, which led to performance improvements across disease classification, detection, and segmentation tasks using a small fraction of training data. Further, our model enables the creation of diverse patient cohorts that enhance model fairness by detecting and mitigating demographic biases. Our study supports the transformative role of generative foundation models in building more accurate, data-efficient, and equitable medical AI systems.
Histopathology is essential for cancer diagnosis and treatment selection, and pathology foundation models learn visual representations from whole-slide images (WSIs). However, existing foundation models are trained on disparate datasets using varying strategies, leading to inconsistent performance and limited generalizability. Here, we introduce ELF (Ensemble Learning of Foundation models), which integrates five pretrained pathology foundation models into unified slide-level representations. Trained on 53,699 WSIs spanning 20 anatomical sites, ELF leverages ensemble learning to capture complementary information across models. ELF's slide-level architecture is designed for data-efficient downstream evaluation, including settings with limited data such as therapeutic response prediction. We evaluate ELF for disease classification, biomarker detection, as well as anticancer and immunotherapy response prediction across multiple cancer types. ELF achieves higher performance than the evaluated constituent and slide-level foundation models across the tested tasks, supporting further evaluation of ensemble learning for pathology applications in oncology.
Accurately recognizing rice seed varieties poses significant challenges due to their diverse morphological characteristics and complex classification requirements. Traditional image recognition methods often struggle with both accuracy and efficiency in this context. To address these limitations, this study proposes the Deep Space and Channel Residual Network with Double Attention Mechanism (RSCD-Net) to enhance the recognition accuracy of 36 rice seed varieties. The core innovation of RSCD-Net is the introduction of the Space and Channel Feature Extraction Residual Block (SCR-Block), which improves inter-class differentiation while minimizing redundant features, thereby optimizing computational efficiency. The RSCD-Net architecture consists of 16 layers of SCR-Blocks, structured into four convolutional stages with 3, 4, 6, and 3 units, respectively. Additionally, a Double Attention Mechanism (A2Net) is incorporated to enhance the network's global receptive field, improving its capacity to distinguish subtle variations among seed types. Experimental results on a self-collected dataset demonstrate that RSCD-Net achieves an average accuracy of 81.94%, surpassing the baseline model by 4.16%. Compared with state-of-the-art models such as InceptionResNetV2, ConvNeXt, MobileNetV3, and Swin Transformer, RSCD Net has improved by 1.17%, 3%, 24.72%, and 13.22%, respectively, showcasing its superior performance. These findings confirm that RSCD-Net provides an effective and efficient solution for rice seed classification, offering a promising reference for addressing similar fine-grained recognition challenges in agricultural applications.
Artificial intelligence (AI) holds significant promise in transforming medical imaging, enhancing diagnostics, and refining treatment strategies. However, the reliance on extensive multicenter datasets for training AI models poses challenges due to privacy concerns. Federated learning provides a solution by facilitating collaborative model training across multiple centers without sharing raw data. This study introduces a federated attention-consistent learning (FACL) framework to address challenges associated with large-scale pathological images and data heterogeneity. FACL enhances model generalization by maximizing attention consistency between local clients and the server model. To ensure privacy and validate robustness, we incorporated differential privacy by introducing noise during parameter transfer. We assessed the effectiveness of FACL in cancer diagnosis and Gleason grading tasks using 19,461 whole-slide images of prostate cancer from multiple centers. In the diagnosis task, FACL achieved an area under the curve (AUC) of 0.9718, outperforming seven centers with an average AUC of 0.9499 when categories are relatively balanced. For the Gleason grading task, FACL attained a Kappa score of 0.8463, surpassing the average Kappa score of 0.7379 from six centers. In conclusion, FACL offers a robust, accurate, and cost-effective AI training model for prostate cancer pathology while maintaining effective data safeguards.
Cervical cancer is the fourth most common cancer in women and its subtyping requires examining histopathological slides or digital images, such as whole slide images (WSIs). However, manually inspecting WSIs with gigapixel sizes can be laborious and prone to errors for pathologists. To address this issue, computer-aided approaches based on weakly-supervised learning techniques have been proposed. These methods can predict disease types directly from WSIs and highlight diagnosis-relevant regions, which can help pathologists achieve faster and more accurate diagnoses. WSIs are divided into overlapping patches using a sliding window approach, and these patches are subsequently screened in a sequential zig-zag pattern to identify spatiotemporal dependencies. These dependencies are further analyzed to generate predictions at the WSI level. Therefore, effective patch feature learning and spatiotemporal aggregation are two key issues in the weakly-supervised WSI classification (WSWC) task. In this paper, we present a label-efficient WSWC method called spatiotemporal aggregation for cervical WSIs (SAC-Net), which jointly performs online feature extraction and feature aggregation to infer the WSI-level prediction in an end-to-end manner. The online feature extractor helps to learn cervical-cancer-specific features and obtain more accurate patch representations. The feature aggregator uses an online instance clustering method to learn proper weight parameters for each cluster, which generates the WSI embedding with enhanced spatiotemporal aggregation. SAC-Net is developed and evaluated on a public cervical WSI dataset (TissueNet) containing 1015 WSIs, which are also externally tested on three independent cervical WSI datasets. Our results demonstrate that SAC-Net achieves state-of-the-art classification performance and is robust. SAC-Net has the potential to be a useful tool for clinical cervical cancer detection.