Artificial intelligence models have been increasingly used in the analysis of tumor histology to perform tasks ranging from routine classification to identification of molecular features. These approaches distill cancer histologic images into high-level features, which are used in predictions, but understanding the biologic meaning of such features remains challenging. We present and validate a custom generative adversarial network—HistoXGAN—capable of reconstructing representative histology using feature vectors produced by common feature extractors. We evaluate HistoXGAN across 29 cancer subtypes and demonstrate that reconstructed images retain information regarding tumor grade, histologic subtype, and gene expression patterns. We leverage HistoXGAN to illustrate the underlying histologic features for deep learning models for actionable mutations, identify model reliance on histologic batch effect in predictions, and demonstrate accurate reconstruction of tumor histology from radiographic imaging for a “virtual biopsy.”
Deployment and access to state-of-the-art diagnostic technologies remains a fundamental challenge in providing equitable global cancer care to low-resource settings. The expansion of digital pathology in recent years and its interface with computational biomarkers provides an opportunity to democratize access to personalized medicine. Here we describe a low-cost platform for digital side capture and computational analysis composed of open-source components. The platform provides low-cost ($200) digital image capture from glass slides and is capable of real-time computational image analysis using an open-source deep learning (DL) algorithm and Raspberry Pi ($35) computer. We validate the performance of deep learning models’ performance using images captured from the open-source workstation and show similar model performance when compared against significantly more expensive standard institutional hardware.
10023 Background: Rapid and accurate identification of morphologic features of neuroblastic tumors (NTs) is critical for risk stratification and therapeutic decision making. The prognostic value of features like neuroblast differentiation, mitosis-karyorrhexis index (MKI), and Schwannian stromal presence is well established. Deep learning permits objective histopathological analysis, streamlining workflows for pathologists, notably in rare cancers. In rare cancers, our method minimizes bias and optimizes limited data using transfer and self-supervised learning (SSL) for feature extraction, with improved explainability. Here, we used an artificial intelligence-based model to morphologically classify NT tumors and MYCN-amplification. Methods: Annotated H&E-stained slides of diagnostic NT tumor biopsies from the University of Chicago and the Children’s Oncology Group were digitalized. Pathologists defined three binarized measures including diagnostic category (ganglioneuroblastoma/neuroblastoma), grade (differentiating/poorly differentiating), and MKI (low and intermediate/high). MYCN status was abstracted from patient records (amplified/non-amplified). Using Slideflow, our open-source pipeline, we developed an attention-based multiple instance learning model with features extracted by CTransPath, a SSL model pretrained on pan-cancer images from The Cancer Genome Atlas. For each measure, model performance was evaluated using 5-fold cross validation by aggregating k-fold model predictions across multiple metrics. Patients were excluded from a model if the measure of interest was unknown. Feature significance was assessed visually using Class Activation Mapping (Grad-CAM). Results: The mean age of the study cohort (n = 172) was 3.66 years. Of patients with clinical information, 84 of 138 (60.2%) had metastatic disease and 94 of 133 (70.7%) were high-risk. Of the 148 tumors with a diagnostic category of neuroblastoma, 93.2% were poorly differentiated and 25% had high MKI. Of the 135 tumors with known MYCN status, 40 were amplified (29.6%). The final models excelled across all outcomes, performing best for diagnostic category, grade, and MYCN status (Table 1). Physician review of the attention-based heatmaps for all measures highlighted biologically relevant regions such as neuropil. Conclusions: We created a deep learning pipeline for auto-characterization of digitized H&E-stained NT pathology slides. Our approach may also aid in identifying molecular features including MYCN-amplification. Review of heatmaps showed pertinent biological tissue, boosting model reliability.[Table: see text]
Deep learning methods have emerged as powerful tools for analyzing histopathological images, but current methods are often specialized for specific domains and software environments, and few open-source options exist for deploying models in an interactive interface. Experimenting with different deep learning approaches typically requires switching software libraries and reprocessing data, reducing the feasibility and practicality of experimenting with new architectures. We developed a flexible deep learning library for histopathology called Slideflow, a package which supports a broad array of deep learning methods for digital pathology and includes a fast whole-slide interface for deploying trained models. Slideflow includes unique tools for whole-slide image data processing, efficient stain normalization and augmentation, weakly-supervised whole-slide classification, uncertainty quantification, feature generation, feature space analysis, and explainability. Whole-slide image processing is highly optimized, enabling whole-slide tile extraction at 40x magnification in 2.5 s per slide. The framework-agnostic data processing pipeline enables rapid experimentation with new methods built with either Tensorflow or PyTorch, and the graphical user interface supports real-time visualization of slides, predictions, heatmaps, and feature space characteristics on a variety of hardware devices, including ARM-based devices such as the Raspberry Pi.
A deep learning model using attention-based multiple instance learning (aMIL) and self-supervised learning (SSL) was developed to perform pathologic classification of neuroblastic tumors and assess MYCN-amplification status using H&E-stained whole slide images from the largest reported cohort to date. The model showed promising performance in identifying diagnostic category, grade, mitosis-karyorrhexis index (MKI), and MYCN-amplification with validation on an external test dataset, suggesting potential for AI-assisted neuroblastoma classification.
As quantum technology advances, it holds immense potential to accelerate oncology discovery through enhanced molecular modeling, genomic analysis, medical imaging, and quantum sensing.
90 Background: Premalignant lesions within the oral cavity demonstrate varied degrees of dysplasia. Oral stratified squamous epithelium is sensitive to carcinogenic insults and lesions progress to cancer at variable rates. The current diagnostic standard involves visual and tactile inspection followed by a biopsy. However, thus far, there is clinical ambiguity about which lesions may progress to carcinoma and there are insufficient prognostic tools available to guide clinical decision-making for follow-up. Deep learning is a form of artificial intelligence which can be applied to images. Self-supervised learning (SSL) and attention-based multiple instance learning (aMIL), are deep learning extensions which eliminate the need for pathologist annotations of histology and are more generalizable as the model learns general representations of unlabeled data, compared to traditional tile-based methods. We developed a novel artificial intelligence pipeline to predict progression to oral carcinoma using histological images from oral premalignant lesions. Methods: Digital histopathology images and clinical data cohorts were obtained from the US National Cancer Institute (n=561) and the University of Iowa (n=193) for training, and from the University of Chicago (n=214) for external validation. Patients were defined as progressing if initially non-malignant lesions had documented progression to carcinoma within 12 months of follow-up. Model training and evaluation was done in Python using our publicly available open-source Slideflow platform. We developed a deep learning pipeline utilizing SimCLR as an SSL method for feature extraction paired with an aMIL head for classification. The SimCLR feature generator was trained on only the training cohort, then used to generate features for both training and validation data. The aMIL model was trained on SimCLR-generated features. Trained model performance was evaluated on the external University of Chicago cohort and measured via areas under the receiver operating characteristic curve and precision recall curve (AUROC and AUPRC, respectively). During training and validation, model reliability was interpreted using Class Activation Mapping (Grad-CAM) and generative adversarial network (GAN) methods. Results: On training data, the model achieved average 3-fold cross validation AUROC of 0.86 with AUPRC of 0.83. In the external validation data from UC, the model achieved an AUROC of 0.80 with an AUPRC of 0.79. Deployed as a “rule-in” test for early progression prediction, the model achieved a validation data set specificity of 91% with a sensitivity of 40%. Conclusions: We developed a deep learning model to automatically predict cancer progression from premalignant lesions histopathology using a multi-institutional cohort, with promising external validation performance.
Large language models such as ChatGPT can produce increasingly realistic text, with unknown information on the accuracy and integrity of using these models in scientific writing. We gathered fifth research abstracts from five high-impact factor medical journals and asked ChatGPT to generate research abstracts based on their titles and journals. Most generated abstracts were detected using an AI output detector, 'GPT-2 Output Detector', with % 'fake' scores (higher meaning more likely to be generated) of median [interquartile range] of 99.98% 'fake' [12.73%, 99.98%] compared with median 0.02% [IQR 0.02%, 0.09%] for the original abstracts. The AUROC of the AI output detector was 0.94. Generated abstracts scored lower than original abstracts when run through a plagiarism detector website and iThenticate (higher scores meaning more matching text found). When given a mixture of original and general abstracts, blinded human reviewers correctly identified 68% of generated abstracts as being generated by ChatGPT, but incorrectly identified 14% of original abstracts as being generated. Reviewers indicated that it was surprisingly difficult to differentiate between the two, though abstracts they suspected were generated were vaguer and more formulaic. ChatGPT writes believable scientific abstracts, though with completely generated data. Depending on publisher-specific guidelines, AI output detectors may serve as an editorial tool to help maintain scientific standards. The boundaries of ethical and acceptable use of large language models to help scientific writing are still being discussed, and different journals and conferences are adopting varying policies.
Artificial intelligence methods including deep neural networks (DNN) can provide rapid molecular classification of tumors from routine histology with accuracy that matches or exceeds human pathologists. Discerning how neural networks make their predictions remains a significant challenge, but explainability tools help provide insights into what models have learned when corresponding histologic features are poorly defined. Here, we present a method for improving explainability of DNN models using synthetic histology generated by a conditional generative adversarial network (cGAN). We show that cGANs generate high-quality synthetic histology images that can be leveraged for explaining DNN models trained to classify molecularly-subtyped tumors, exposing histologic features associated with molecular state. Fine-tuning synthetic histology through class and layer blending illustrates nuanced morphologic differences between tumor subtypes. Finally, we demonstrate the use of synthetic histology for augmenting pathologist-in-training education, showing that these intuitive visualizations can reinforce and improve understanding of histologic manifestations of tumor biology.
The histopathological phenotype of tumors reflects the underlying genetic makeup. Deep learning can predict genetic alterations from pathology slides, but it is unclear how well these predictions generalize to external datasets. We performed a systematic study on Deep-Learning-based prediction of genetic alterations from histology, using two large datasets of multiple tumor types. We show that an analysis pipeline that integrates self-supervised feature extraction and attention-based multiple instance learning achieves a robust predictability and generalizability.
Machine learning methods have been growing in prominence across all areas of medicine. In pathology, recent advances in deep learning (DL) have enabled computational analysis of histological samples, aiding in diagnosis and characterization in multiple disease areas. In cancer, and particularly endocrine cancer, DL approaches have been shown to be useful in tasks ranging from tumor grading to gene expression prediction. This review summarizes the current state of DL research in endocrine cancer histopathology with an emphasis on experimental design, significant findings, and key limitations.
A model's ability to express its own predictive uncertainty is an essential attribute for maintaining clinical user confidence as computational biomarkers are deployed into real-world medical settings. In the domain of cancer digital histopathology, we describe a clinically-oriented approach to uncertainty quantification for whole-slide images, estimating uncertainty using dropout and calculating thresholds on training data to establish cutoffs for low- and high-confidence predictions. We train models to identify lung adenocarcinoma vs. squamous cell carcinoma and show that high-confidence predictions outperform predictions without uncertainty, in both cross-validation and testing on two large external datasets spanning multiple institutions. Our testing strategy closely approximates real-world application, with predictions generated on unsupervised, unannotated slides using predetermined thresholds. Furthermore, we show that uncertainty thresholding remains reliable in the setting of domain shift, with accurate high-confidence predictions of adenocarcinoma vs. squamous cell carcinoma for out-of-distribution, non-lung cancer cohorts.
10039 Background: Metaiodobenzylguanidine (MIBG) scans are a radionucleotide imaging modality used to evaluate neuroblastoma stage at diagnosis and also determine disease response following therapy. Curie scoring is used to semi-quantitatively assess disease burden from an MIBG scan on a scale from none (0) to widespread throughout the body (30). While a Curie score ≤2 after six cycles of induction chemotherapy has been shown to be prognostic of outcome, there is no established correlation between diagnostic Curie score and outcome. Deep learning models, such as convolutional neural networks (CNN), have been shown to learn generalizable patterns within images for successful classification of metastases and detection of multiple adult cancers. We hypothesized a CNN could be developed to predict response to induction chemotherapy, a proxy for outcome, using diagnostic MIBG scans. Methods: DICOM MIBG scans and associated clinical data from a Children’s Oncology Group (COG) pilot study for children diagnosed with high-risk neuroblastoma (ANBL12P1; NCT1798004) were deidentified and linked to clinical data by the Pediatric Cancer Data Commons and obtained from the International Neuroblastoma Risk Group Data Commons. Patients were defined as having a poor response to induction chemotherapy if their Curie score after four cycles of induction chemotherapy was ≥2. An independent external validation cohort was comprised of 29 images from 26 high-risk patients treated at the University of Chicago with clinically-annotated diagnostic and post-cycle six induction DICOM MIBG scans. The CNN was trained using 2D whole body MIBG scans obtained at diagnosis. We developed the CNN using a transfer learning approach using the Xception architecture as the base layer. Hyperparameter optimization was performed using an 80%-20% train-validation strategy. Model performance was evaluated using area under the receiver operating characteristic curve (AUROC). Results: Among 146 patients with high-risk neuroblastoma enrolled on ANBL12P1, 104 had available diagnostic and end-induction MIBG scans. There were no differences in clinical or biological characteristics between included and excluded patients. The base model CNN was able to predict which patients had a poor response to induction chemotherapy with an AUROC of 0.72 in the validation set from the ANBL12P1 cohort. Additionally, the CNN was able to predict patient response to therapy with an AUROC of 0.64 in an independent external dataset from University of Chicago. Conclusions: Our study suggests it is feasible to apply machine learning of diagnostic MIBG scans to predict response to chemotherapy for high-risk neuroblastoma patients. Given these promising results, further work to improve AUROC and performance within larger datasets is ongoing.
PURPOSE:Metaiodobenzylguanidine (MIBG) scans are a radionucleotide imaging modality that undergo Curie scoring to semiquantitatively assess neuroblastoma burden, which can be used as a marker of therapy response. We hypothesized that a convolutional neural network (CNN) could be developed that uses diagnostic MIBG scans to predict response to induction chemotherapy. METHODS:We analyzed MIBG scans housed in the International Neuroblastoma Risk Group Data Commons from patients enrolled in the Children's Oncology Group high-risk neuroblastoma study ANBL12P1. The primary outcome was response to upfront chemotherapy, defined as a Curie score ≤ 2 after four cycles of induction chemotherapy. We derived and validated a CNN using two-dimensional whole-body MIBG scans from diagnosis and evaluated model performance using area under the receiver operating characteristic curve (AUC). We also developed a clinical classification model to predict response on the basis of age, stage, and MYCN amplification. RESULTS:Among 103 patients with high-risk neuroblastoma included in the final cohort, 67 (65%) were responders. Performance in predicting response to upfront chemotherapy was equivalent using the CNN and the clinical model. Class-activation heatmaps verified that the CNN used areas of disease within the MIBG scans to make predictions. Furthermore, integrating predictions using a geometric mean approach improved detection of responders to upfront chemotherapy (geometric mean AUC 0.73 v CNN AUC 0.63, P < .05; v clinical model AUC 0.65, P < .05). CONCLUSION:We demonstrate feasibility in using machine learning of diagnostic MIBG scans to predict response to induction chemotherapy for patients with high-risk neuroblastoma. We highlight improvements when clinical risk factors are also integrated, laying the foundation for using a multimodal approach to guiding treatment decisions for patients with high-risk neuroblastoma.
PURPOSE:There is a need for an improved understanding of clinical and biologic risk factors in pediatric cancer to improve patient outcomes. Machine learning (ML) represents the application of computational inference from advanced statistical methods that can be applied to increasing amount of data available for study in pediatric oncology. The goal of this systematic review was to systematically characterize the state of ML in pediatric oncology and highlight advances and opportunities in the field. METHODS:We conducted a systematic review of the Embase, Scopus, and MEDLINE databases for applications of ML in pediatric oncology. Query results from all three databases were aggregated and duplicate studies were removed. RESULTS:A total of 42 unique articles that examined the applications of ML in pediatric oncology met inclusion criteria for review. We identified 20 studies of CNS tumors, 13 of solid tumors, and nine of leukemia. ML tasks included classification, prediction of treatment response, and dose optimization with a variety of methods being used including neural network, k-nearest neighbor, random forest, naive Bayes, and support vector machines. Strengths of the identified studies included matching or outperforming physician comparators via automated analysis and predicting therapeutic response. Common limitations included significant heterogeneity in reporting standards, clinical applicability, small sample sizes, and missing external validation cohorts. CONCLUSION:We identified areas where ML can enhance clinical care in ways that may not otherwise be achievable. Although ML promises enormous potential in improving diagnostics, decision making, and monitoring for children with cancer, the field remains in early stages and future work will be aided by standards and guidelines to ensure rigorous methodologic design and maximizing clinical utility.
The urgency of the COVID-19 pandemic has increased attention on the need to include patients traditionally underrepresented in research, especially people with comorbid diseases, racial/ethnic minorities, pregnant/lactating women, and children.1–4 We characterized the inclusion and exclusion criteria of US COVID-19 treatment trials to understand their applicability to highly affected populations and groups traditionally underrepresented in research.
Infectious disease outbreaks related to outdoor sporting events are an underacknowledged environmental health risk. We reviewed documented instances of sporting events that led to outbreaks of illness due to interactions of athletes with environmental pathogen reservoirs such as soil or water. We note that aspects of outdoor athletic activities can mediate a suppression in immune function. The implications of this review are of particular interest to environmental health professionals and healthcare professionals, as they show that populations of young, otherwise healthy adult athletes and other members of their communities are a contextually at-risk group to be aware of in order to improve and speed ability to cluster individuals in outbreaks, and for diagnosis and prevention for individual patients. This review highlights an opportunity where environmental health professionals can provide a critical linkage between patients and the environment.
ABSTRACT: Introduction: Stable low-grade osteochondritis dissecans (OCD) lesions in adolescents are typically treated with an initial trial of conservative management including bracing, activity limitation, and physical therapy. While there are some tools to help predict healing rates, overall only about 2/3rd heal with non-surgical management. Surgical treatment, on the other hand, has an overall higher success rate for healing, over 90%, in a shorter time frame. We sought to better understand the factors that influence the treatment choice by investigating the tradeoff between initial conservative treatment and early surgical management through a cost-effectiveness analysis (CEA). Methods: A …
Background As sports specialization continues to grow amongst youth, the prevalence of shoulder instability continues to increase. Bankart repair exists as a primary method to combat these issues, yet a third or more of patients experience recurrent instability following surgery. Currently most studies examining Bankart repair look at surgical technique (open vs arthroscopic) and the original treatment employed; however, there has been a lack of research examining whether demographic, injury history, and perioperative risk factors exist as well. The purpose of this prognostic study was to examine Bankart repair in the pediatric population and identify a wider range of risk factors associated with poor functional outcomes. Methods We conducted a retrospective review of all patients who received Bankart repair between 01/01/2010 and 12/31/2015 at the Children’s Hospital of Philadelphia. A range of demographic, injury history, clinical exam, operative, and perioperative variables were abstracted to identify risk factors associated with limited functional outcome as defined by inability to return to baseline sport or activity. Chi square analyses were used to find significant associations between dichotomous variables, and point-biserial analyses were used to find strength of correlations for continuous variables. Results 118 patients were included with an average age of 15.92 ±2.11 at the time of surgery. The mean OR time was 206.98 ±55.88 minutes. Point-biserial correlations for total OR time, BMI, and age with inability to return to baseline sport or activity were 0.32, .09, and .04 respectively. Chi square analyses found no significant associations for prior non-operative treatment, sulcus sign, and insurance type. For individuals who presented with prior instability, 92/119 (77.3%) with and 27/119 (22.7%) without, this was significantly associated with inability to return to baseline sport or activity (p = 0.04). For presence of a Hill Sachs Lesion 70/145 (48.2%) had a lesion, 75/145 (51.8%) did not, and this was significantly associated as well (p = 0.02). For injury type, 71/98 (72.4%) were contact, 27/98 (27.6%) were noncontact, and this was significantly associated as well (p = 0.01). Conclusion Many clinicians have informally found in practice that certain patient characteristics seem to be associated with poor outcomes after Bankart repair. Here we find that a history of prior instability, Hill Sachs lesion presence, and a contact injury mechanism are significantly associated with inability to return to baseline sport or activity post-operatively. The risk factors identified here can be categorized into injury history risk factors and perioperative risk factors. Clinicians can use these identified risk factors to better inform patients receiving Bankart repair of their likelihood of returning to baseline sport or activity.