Digital pathology has seen significant advancements in artificial intelligence (AI) applications. However, challenges persist in integrating these solutions into digital pathology platforms for human and AI collaborations. We introduce I-Viewer, an online AI Copilot framework designed to facilitate real-time human-AI and human-human collaboration for digital pathology analysis. The I-Viewer platform enables precise annotations and descriptions from tissue to the nuclei level through an Agentic-Retrieval Augmented Generation (RAG) system. By leveraging agents' outputs as reference points, aggregating information through the RAG system, and incorporating Large Language Models (LLM) for human feedback and refinement, I-Viewer sets a new standard for collaborative and accurate digital pathology analysis. We demonstrate I-Viewer's effectiveness on different pathology tasks using three datasets across different types of cancers, including non-small cell lung cancer, breast cancer, and colorectal cancer. The results show that I-Viewer achieves significant improvements in annotation speed and accuracy for pathology tasks, such as detecting cell morphology, cellular structures, and tumor growth patterns, outperforming current individual foundation models. Through its advanced AI agents, collaborative features, and LLM integrations, I-Viewer optimizes diagnostic workflows in clinical care and biomedical research.
Sup. Fig. 1 Simulation analyses show accurate dissection of bulk tumor RNA-Seq data by DisHet. Sup. Fig. 2 Validating immune/stroma component expression dissected by DisHet. Sup. Fig. 3 Comparison of dissection performance of DeMix and DisHet. Sup. Fig. 4 Principal Component Analysis (PCA) shows the clustering of TCGA pan-RCC patients tumors by the tumor-specific genes and eTME-specific genes, respectively. Sup. Fig. 5 CD4 and CD8 immunohistochemistry staining in the Six2-Cre;VhlF/F;Pbrm1F/F and Six2-Cre;VhlF/F;Bap1F/+. Supl. Fig. 6 Clinical manifestations of eTME-NIS patients.
The emerging field of spatially resolved transcriptomics (SRT) has revolutionized biomedical research. SRT quantifies expression levels at different spatial locations, providing a new and powerful tool to interrogate novel biological insights. An essential question in the analysis of SRT data is to identify spatially variable (SV) genes; the expression levels of such genes have spatial variation across different tissues. SV genes usually play an important role in underlying biological mechanisms and tissue heterogeneity. Currently, several computational methods have been developed to detect such genes; however, there is a lack of unbiased assessment of these approaches to guide researchers in selecting the appropriate methods for their specific biomedical applications. In addition, it is difficult for researchers to implement different existing methods for either biological study or methodology development. Furthermore, currently available public SRT datasets are scattered across different websites and preprocessed in different ways, posing additional obstacles for quantitative researchers developing computational methods for SRT data analysis. To address these challenges, we designed Spatial Transcriptomics Arena (STAr), an open platform comprising 193 curated datasets from seven technologies, seven statistical methods, and analysis results. This resource allows users to retrieve high-quality datasets, apply or develop spatial gene detection methods, as well as browse and compare spatial gene analysis results. It also enables researchers to comprehensively evaluate SRT methodology research in both simulated and real datasets. Altogether, STAr is an integrated research resource intended to promote reproducible research and accelerate rigorous methodology development, which can eventually lead to an improved understanding of biological processes and diseases. STAr can be accessed at https://lce.biohpc.swmed.edu/star/ .
Gene ontology analysis shows enriched GO terms of the 904 eTME genes (20x cutoff) (first sheet), the 778 and the 610 eTME genes (20x cutoff) that are not captured by the Immunome alone and Immunome+ESTIMATE+Winslow gene signatures (second and third sheets).
Supplemental Figure S4 Whole-slide image nuclei segmentation and classification by HD-Staining model.
Supplemental Figure S5 Validation of lymphocyte detection results in whole-slide image nuclei segmentation and classification by HD-Staining model.
Motivation Spatial transcriptomics (ST) enables a high-resolution interrogation of molecular characteristics within specific spatial contexts and tissue morphology. Despite its potential, visualization of ST data is a challenging task due to the complexities in handling, sharing and visualizing large image datasets together with molecular information. Results We introduce ScopeViewer, a browser-based software designed to overcome these challenges. ScopeViewer offers the following functionalities: (1) It visualizes large image data and associated annotations at various zoom levels, allowing for intricate exploration of the data; (2) It enables dual interactive viewing of the original images along with their annotations, providing a comprehensive understanding of the context; (3) It displays spatial molecular features with optimized bandwidth, ensuring a smooth user experience; and (4) It bolsters data security by circumventing data transfers. Availability ScopeViewer is available at: https://datacommons.swmed.edu/scopeviewer Contact Xiaowei.Zhan@UTSouthwestern.edu , Guanghua.Xiao@UTSouthwestern.edu Supplementary information Supplementary data are available at Bioinformatics online.
Microscopic examination of pathology slides is essential to disease diagnosis and biomedical research. However, traditional manual examination of tissue slides is laborious and subjective. Tumor whole-slide image (WSI) scanning is becoming part of routine clinical procedures and produces massive data that capture tumor histologic details at high resolution. Furthermore, the rapid development of deep learning algorithms has significantly increased the efficiency and accuracy of pathology image analysis. In light of this progress, digital pathology is fast becoming a powerful tool to assist pathologists. Studying tumor tissue and its surrounding microenvironment provides critical insight into tumor initiation, progression, metastasis, and potential therapeutic targets. Nucleus segmentation and classification are critical to pathology image analysis, especially in characterizing and quantifying the tumor microenvironment (TME). Computational algorithms have been developed for nucleus segmentation and TME quantification within image patches. However, existing algorithms are computationally intensive and time consuming for WSI analysis. This study presents Histology-based Detection using Yolo (HD-Yolo), a new method that significantly accelerates nucleus segmentation and TME quantification. We demonstrate that HD-Yolo outperforms existing WSI analysis methods in nucleus detection, classification accuracy, and computation time. We validated the advantages of the system on 3 different tissue types: lung cancer, liver cancer, and breast cancer. For breast cancer, nucleus features by HD-Yolo were more prognostically significant than both the estrogen receptor status by immunohistochemistry and the progesterone receptor status by immunohistochemistry. The WSI analysis pipeline and a real-time nucleus segmentation viewer are available at https://github.com/impromptuRong/hd_wsi.
Supplemental Figure S3 Confusion matrix between true labels and predicted labels, in 5 major lung adenocarcinoma subtypes respectively.
Supplemental Figure S6. Validation of macrophage detection results in whole-slide image nuclei segmentation and classification by HD-Staining model.
Immunohistochemistry (IHC) is a well-established and commonly used staining method for clinical diagnosis and biomedical research. In most IHC images, the target protein is conjugated with a specific antibody and stained using diaminobenzidine (DAB), resulting in a brown coloration, whereas hematoxylin serves as a blue counterstain for cell nuclei. The protein expression level is quantified through the H-score, calculated from DAB staining intensity within the target cell region. Traditionally, this process requires evaluation by 2 expert pathologists, which is both time consuming and subjective. To enhance the efficiency and accuracy of this process, we have developed an automatic algorithm for quantifying the H-score of IHC images. To characterize protein expression in specific cell regions, a deep learning model for region recognition was trained based on hematoxylin staining only, achieving pixel accuracy for each class ranging from 0.92 to 0.99. Within the desired area, the algorithm categorizes DAB intensity of each pixel as negative, weak, moderate, or strong staining and calculates the final H-score based on the percentage of each intensity category. Overall, this algorithm takes an IHC image as input and directly outputs the H-score within a few seconds, significantly enhancing the speed of IHC image analysis. This automated tool provides H-score quantification with precision and consistency comparable to experienced pathologists but at a significantly reduced cost during IHC diagnostic workups. It holds significant potential to advance biomedical research reliant on IHC staining for protein expression quantification.
Supplemental Figure S9 Example of gene set enrichment analysis (GSEA) result correlating mRNA expression with image-derived stroma cell number and lymphocyte number, while patient IDs are randomly shuffled twice.
List of UTSW KCP patients with RNA-Seq and exome-seq. Sheet 1 contains information of 35 patients whose matched trio RNA-Seq data were used in the primary DisHet analyses. Sheet 2 contains information of patients whose bulk tumors were analyzed by RNA-Seq and whose follow up data are available. There are overlap between patients from these two sheets. Sheet 3 contains the information of tumorgraft data used in Fig. 5bc.
PURPOSE Osteosarcoma research advancement requires enhanced data integration across different modalities and sources. Current osteosarcoma research, encompassing clinical, genomic, protein, and tissue imaging data, is hindered by the siloed landscape of data generation and storage. MATERIALS AND METHODS Clinical, molecular profiling, and tissue imaging data for 573 patients with pediatric osteosarcoma were collected from four public and institutional sources. A common data model incorporating standardized terminology was created to facilitate the transformation, integration, and load of source data into a relational database. On the basis of this database, a data commons accompanied by a user-friendly web portal was developed, enabling various data exploration and analytics functions. RESULTS The Osteosarcoma Explorer (OSE) was released to the public in 2021. Leveraging a comprehensive and harmonized data set on the backend, the OSE offers a wide range of functions, including Cohort Discovery, Patient Dashboard, Image Visualization, and Online Analysis. Since its initial release, the OSE has experienced an increasing utilization by the osteosarcoma research community and provided solid, continuous user support. To our knowledge, the OSE is the largest (N = 573) and most comprehensive research data commons for pediatric osteosarcoma, a rare disease. This project demonstrates an effective framework for data integration and data commons development that can be readily applied to other projects sharing similar goals. CONCLUSION The OSE offers an online exploration and analysis platform for integrated clinical, molecular profiling, and tissue imaging data of osteosarcoma. Its underlying data model, database, and web framework support continuous expansion onto new data modalities and sources.
Supplemental Figure S1 Illustration of 100 patches extraction from the Region of Interest (ROI) circled by pathologists.
Supplemental Table S2. Image features that are significantly associated with patient overall survival outcome in univariate survival analysis in the National Lung Screening Trial (NLST) dataset. Features are dichotomized by the median value. P values are calculated using Ward test. CI, confidence interval; HR, hazard ratio.
Supplemental Figure S8 Volcano plots of gene set enrichment analysis results correlating mRNA expression level with stroma nuclei density and karyorrhexis density, respectively.
Supplemental Figure S2 Illustration of instance segmentation using the developed HD-Staining model.