Pathology underpins modern diagnosis and cancer care, yet its most valuable asset, the accumulated experience encoded in millions of narrative reports, remains largely inaccessible. Although institutions are rapidly digitizing pathology workflows, storing data without effective mechanisms for retrieval and reasoning risks transforming archives into a passive data repository, where institutional knowledge exists but cannot meaningfully inform patient care. True progress requires not only digitization, but the ability for pathologists to interrogate prior similar cases in real time while evaluating a new diagnostic dilemma. We present PathoScribe, a unified retrieval-augmented large language model (LLM) framework designed to transform static pathology archives into a searchable, reasoning-enabled living library. PathoScribe enables natural language case exploration, automated cohort construction, clinical question answering, immunohistochemistry (IHC) panel recommendation, and prompt-controlled report transformation within a single architecture. Evaluated on 70,000 multi-institutional surgical pathology reports, PathoScribe achieved perfect Recall@10 for natural language case retrieval and demonstrated high-quality retrieval-grounded reasoning (mean reviewer score 4.56/5). Critically, the system operationalized automated cohort construction from free-text eligibility criteria, assembling research-ready cohorts in minutes (mean 9.2 minutes) with 91.3
Diagnostic errors in pathology remain a significant contributor to patient harm despite advances in laboratory automation and quality management systems. Failure Mode and Effects Analysis (FMEA) provides a structured approach to the prospective identification and prioritization of workflow risks, while artificial intelligence (AI) offers emerging capabilities in image analysis, data validation, and workflow monitoring. In this study, we surveyed 43 pathologists to characterize commonly encountered diagnostic errors across the pathology workflow and found that errors were most prevalent in the pre-examination phase. These findings informed an FMEA-based risk assessment that identified high-priority failure modes. Building on these insights, we propose a conceptual framework that integrates agentic AI with FMEA-driven quality management. In this model, AI systems can enhance error detection, reduce failures, and enable continuous recalibration of risk through dynamic data feedback. This work is intended as a hypothesis-generating and conceptual contribution. Prospective validation, standardized performance metrics, and real-world implementation studies will be necessary to determine the clinical and operational impact of such integrated systems.
Diagnostic errors in pathology remain an important contributor to patient morbidity and healthcare inefficiency despite increasing laboratory automation and quality management initiatives. Failure Mode and Effects Analysis (FMEA) offers a structured framework for prospective identification and prioritization of workflow risks across the total testing process. Concurrently, artificial intelligence (AI) has demonstrated substantial capability in image analysis, data validation, and workflow optimisation in both anatomical and clinical pathology. We present our findings from a survey among 43 pathologists who identified 266 errors i.e., (183, 69%) within pre-analytical phase, 59 (22%) errors within analytical phase and 22 (9%) errors within post-analytical phase. Then we present the results of the FMEA along with the most suitable AI approaches to mitigate the errors identified. In the end, we present a conceptual framework to integrate AI particularly agentic AI with FMEA-driven quality governance. Agentic AI enables dynamic monitoring, predictive error detection, and autonomous workflow orchestration across pre-analytical, analytical, and post-analytical phases. Integrating agentic AI with FMEA may enable continuous recalibration of risk models, transition pathology quality programs from reactive to predictive paradigms, and support scalable diagnostic safety ecosystems. Implementation considerations including algorithmic transparency, human–AI interaction, and regulatory governance are discussed. This integrated approach represents a potential next-generation model for improving diagnostic accuracy, operational efficiency, and patient safety across pathology services.
The 5th edition of the World Health Organization Classification of Tumours (WCT) serves as a foundation for global diagnostic standards in tumour pathology. Similar to other volumes in this series, the Skin Tumours (Skin5) edition follows a standardized approach. This edition introduces two new chapters: 'Tumours of the nail unit' and 'Metastases to skin', along with new entities across relevant chapters. This review article provides an overview of the updates in Skin5 based on currently published evidence, with emphasis on newly introduced chapters and newly described entities that involve the skin, in particular, epidermal, melanocytic and appendageal tumours.
This review focuses on the purported applications of multimodal Gen-AI models for anatomic pathology image analysis and interpretation to predict future directions. A scoping review was conducted to explore the applications of multimodal Gen-AI models in advancing histopathology image analysis. A comprehensive search was conducted using electronic databases for relevant articles published within the past year (July 1, 2023 to June 30, 2024). The selected articles were critically analyzed to identify and summarize the applications of multimodal Gen-AI in anatomic pathology image analysis. Multimodal Gen AI models reported in the literature claim moderate to high accuracy on tasks including image classification, segmentation, and text-to-image retrieval. This review demonstrates the potential of multimodal Gen AI models for useful applications in pathology, including assisting with diagnoses, generating data for education and research, and detection of molecular features from anatomic pathology images. These models use data from a few academic institutions thus they require validation on diverse real-world data. There is an urgent need to build consensus models for optimal model performance through multicenter collaboration using a federated learning approach and the use of carefully curated synthetic anatomic pathology data. These models also need to achieve reliability, generalizability and meet the standards required for clinical use. Despite the rigorous need for evaluation and the need to address genuine concerns, multimodal GenAI models present a promising perspective for the advancement and scalability of anatomic pathology.
Context.— Generative artificial intelligence (GAI) is a promising new technology with the potential to transform communication and workflows in health care and pathology. Although new technologies offer advantages, they also come with risks that users, particularly early adopters, must recognize. Given the fast pace of GAI developments, pathologists may find it challenging to stay current with the terminology, technical underpinnings, and latest advancements. Building this knowledge base will enable pathologists to grasp the potential risks and impacts that GAI may have on the future practice of pathology. Objective.— To present key elements of GAI development, evaluation, and implementation in a way that is accessible to pathologists and relevant to laboratory applications. Data Sources.— Information was gathered from recent studies and reviews from PubMed and arXiv. Conclusions.— GAI offers many potential benefits for practicing pathologists. However, the use of GAI in clinical practice requires rigorous oversight and continuous refinement to fully realize its potential and mitigate inherent risks. The performance of GAI is highly dependent on the quality and diversity of the training and fine-tuning data, which can also propagate biases if not carefully managed. Ethical concerns, particularly regarding patient privacy and autonomy, must be addressed to ensure responsible use. By harnessing these emergent technologies, pathologists will be well placed to continue forward as leaders in diagnostic medicine.
High levels of H2A.Z promote melanoma cell proliferation and correlate with poor prognosis. However, the role of the two distinct H2A.Z histone chaperone complexes SRCAP and P400-TIP60 in melanoma remains unclear. Here, we show that individual subunit depletion of SRCAP, P400, and VPS72 (YL1) results in not only the loss of H2A.Z deposition into chromatin but also a reduction of H4 acetylation in melanoma cells. This loss of H4 acetylation is particularly found at the promoters of cell cycle genes directly bound by H2A.Z and its chaperones, suggesting a coordinated regulation between H2A.Z deposition and H4 acetylation to promote their expression. Knockdown of each of the three subunits downregulates E2F1 and its targets, resulting in a cell cycle arrest akin to H2A.Z depletion. However, unlike H2A.Z deficiency, loss of the shared H2A.Z chaperone subunit YL1 induces apoptosis. Furthermore, YL1 is overexpressed in melanoma tissues, and its upregulation is associated with poor patient outcome. Together, these findings provide a rationale for future targeting of H2A.Z chaperones as an epigenetic strategy for melanoma treatment.
Context.— Machine learning applications in the pathology clinical domain are emerging rapidly. As decision support systems continue to mature, laboratories will increasingly need guidance to evaluate their performance in clinical practice. Currently there are no formal guidelines to assist pathology laboratories in verification and/or validation of such systems. These recommendations are being proposed for the evaluation of machine learning systems in the clinical practice of pathology. Objective.— To propose recommendations for performance evaluation of in vitro diagnostic tests on patient samples that incorporate machine learning as part of the preanalytical, analytical, or postanalytical phases of the laboratory workflow. Topics described include considerations for machine learning model evaluation including risk assessment, predeployment requirements, data sourcing and curation, verification and validation, change control management, human-computer interaction, practitioner training, and competency evaluation. Data Sources.— An expert panel performed a review of the literature, Clinical and Laboratory Standards Institute guidance, and laboratory and government regulatory frameworks. Conclusions.— Review of the literature and existing documents enabled the development of proposed recommendations. This white paper pertains to performance evaluation of machine learning systems intended to be implemented for clinical patient testing. Further studies with real-world clinical data are encouraged to support these proposed recommendations. Performance evaluation of machine learning models is critical to verification and/or validation of in vitro diagnostic tests using machine learning intended for clinical practice.
Background The integration of large language models (LLMs) like ChatGPT in diagnostic medicine, with a focus on digital pathology, has garnered significant attention. However, understanding the challenges and barriers associated with the use of LLMs in this context is crucial for their successful implementation. Methods A scoping review was conducted to explore the challenges and barriers of using LLMs, in diagnostic medicine with a focus on digital pathology. A comprehensive search was conducted using electronic databases, including PubMed and Google Scholar, for relevant articles published within the past four years. The selected articles were critically analyzed to identify and summarize the challenges and barriers reported in the literature. Results The scoping review identified several challenges and barriers associated with the use of LLMs in diagnostic medicine. These included limitations in contextual understanding and interpretability, biases in training data, ethical considerations, impact on healthcare professionals, and regulatory concerns. Contextual understanding and interpretability challenges arise due to the lack of true understanding of medical concepts and lack of these models being explicitly trained on medical records selected by trained professionals, and the black-box nature of LLMs. Biases in training data pose a risk of perpetuating disparities and inaccuracies in diagnoses. Ethical considerations include patient privacy, data security, and responsible AI use. The integration of LLMs may impact healthcare professionals’ autonomy and decision-making abilities. Regulatory concerns surround the need for guidelines and frameworks to ensure safe and ethical implementation. Conclusion The scoping review highlights the challenges and barriers of using LLMs in diagnostic medicine with a focus on digital pathology. Understanding these challenges is essential for addressing the limitations and developing strategies to overcome barriers. It is critical for health professionals to be involved in the selection of data and fine tuning of the models. Further research, validation, and collaboration between AI developers, healthcare professionals, and regulatory bodies are necessary to ensure the responsible and effective integration of LLMs in diagnostic medicine.
High levels of H2A.Z promote melanoma cell proliferation and correlate with poor prognosis. However, the role of the two distinct H2A.Z histone chaperone complexes, SRCAP and P400-TIP60, in melanoma remains unclear. Here, we show that individual depletion of SRCAP, P400, and VPS72 (YL1) not only results in loss of H2A.Z deposition into chromatin, but also a striking reduction of H4 acetylation in melanoma cells. This loss of H4 acetylation is found at the promoters of cell cycle genes directly bound by H2A.Z and its chaperones, suggesting a highly coordinated regulation between H2A.Z deposition and H4 acetylation to promote their expression. Knockdown of each of the three subunits downregulates E2F1 and its targets, resulting in a cell cycle arrest akin to H2A.Z depletion. However, unlike H2A.Z deficiency, loss of the shared H2A.Z chaperone subunit YL1 induces apoptosis. Furthermore, YL1 is overexpressed in melanoma tissues, and its upregulation is associated with poor patient outcome. Together, these findings provide a rationale for future targeting of H2A.Z chaperones as an epigenetic strategy for melanoma treatment.
BACKGROUND:Desmoplastic melanoma is a rare subtype of melanoma mainly appearing on sun-exposed skin. Clinically, it is many times non-pigmented and therefore the diagnosis is often not suspected.METHODS:Review article.RESULTS:In this paper we review the main histopathological, immunohistochemical, and molecular features of desmoplastic melanoma, as well as the top 10 morphologic differential diagnoses which should be considered in most cases. The histopathological pattern can be many times deceptive, mimicking a scar, a fibrous reaction, a fibrohistiocytic tumor such as a dermatofibroma, a vascular tumor such as angiosarcoma, a smooth muscle tumor such as leiomyosarcoma, or a neural tumor. Although an overlying atypical junctional melanocytic proliferation may be seen in most cases, it is absent in a significant percentage (up to 30%) of cases, making the diagnosis even more difficult in those instances. The range of diagnostic pitfalls is wide, which may present disastrous prognostic consequences.CONCLUSION:Desmoplastic melanoma is often a difficult diagnosis to make, as it frequently shows nonspecific clinical findings and overlapping histologic features with many other tumors. However, the potential clinical and prognostic consequences of misdiagnosis as another entity are great. Therefore, this diagnosis must always be kept in mind when encountering spindle cell tumors affecting the head and neck area.
Purpose: Validating artificial intelligence algorithms for clinical use in medical images is a challenging endeavor due to a lack of standard reference data (ground truth). This topic typically occupies a small portion of the discussion in research papers since most of the efforts are focused on developing novel algorithms. In this work, we present a collaboration to create a validation dataset of pathologist annotations for algorithms that process whole slide images. We focus on data collection and evaluation of algorithm performance in the context of estimating the density of stromal tumor-infiltrating lymphocytes (sTILs) in breast cancer. Methods: We digitized 64 glass slides of hematoxylin- and eosin-stained invasive ductal carcinoma core biopsies prepared at a single clinical site. A collaborating pathologist selected 10 regions of interest (ROIs) per slide for evaluation. We created training materials and workflows to crowdsource pathologist image annotations on two modes: an optical microscope and two digital platforms. The microscope platform allows the same ROIs to be evaluated in both modes. The workflows collect the ROI type, a decision on whether the ROI is appropriate for estimating the density of sTILs, and if appropriate, the sTIL density value for that ROI. Results: In total, 19 pathologists made 1645 ROI evaluations during a data collection event and the following 2 weeks. The pilot study yielded an abundant number of cases with nominal sTIL infiltration. Furthermore, we found that the sTIL densities are correlated within a case, and there is notable pathologist variability. Consequently, we outline plans to improve our ROI and case sampling methods. We also outline statistical methods to account for ROI correlations within a case and pathologist variability when validating an algorithm. Conclusion: We have built workflows for efficient data collection and tested them in a pilot study. As we prepare for pivotal studies, we will investigate methods to use the dataset as an external validation tool for algorithms. We will also consider what it will take for the dataset to be fit for a regulatory purpose: study size, patient population, and pathologist training and qualifications. To this end, we will elicit feedback from the Food and Drug Administration via the Medical Device Development Tool program and from the broader digital pathology and AI community. Ultimately, we intend to share the dataset, statistical methods, and lessons learned.
Stromal tumor-infiltrating lymphocytes (sTILs) are important prognostic and predictive biomarkers in triple-negative (TNBC) and HER2-positive breast cancer. Incorporating sTILs into clinical practice necessitates reproducible assessment. Previously developed standardized scoring guidelines have been widely embraced by the clinical and research communities. We evaluated sources of variability in sTIL assessment by pathologists in three previous sTIL ring studies. We identify common challenges and evaluate impact of discrepancies on outcome estimates in early TNBC using a newly-developed prognostic tool. Discordant sTIL assessment is driven by heterogeneity in lymphocyte distribution. Additional factors include: technical slide-related issues; scoring outside the tumor boundary; tumors with minimal assessable stroma; including lymphocytes associated with other structures; and including other inflammatory cells. Small variations in sTIL assessment modestly alter risk estimation in early TNBC but have the potential to affect treatment selection if cutpoints are employed. Scoring and averaging multiple areas, as well as use of reference images, improve consistency of sTIL evaluation. Moreover, to assist in avoiding the pitfalls identified in this analysis, we developed an educational resource available at www.tilsinbreastcancer.org/pitfalls .