Intestinal ultrasound (IUS) is increasingly used for real-time assessment of inflammatory bowel disease (IBD) but quantitative analysis with artificial intelligence (AI) remains limited to bowel wall thickness (BWT) and single-frame segmentation, failing to exploit the temporal information in cineloops (CLs). The aim here was to evaluate the accuracy of an AI model compared to expert central readers using IUS CLs. We developed a novel model using a semi-supervised algorithm combined with a temporal module, trained on 680,000 IUS images from 1572 CLs (327 patients from multiple centres). 1205 frames from 403 CLs on 94 patients were labelled by 8 domain experts (203 validation, 1002 training images). The CLs had various segments per patient with active IBD, and included axial, longitudinal and oblique sections. An external test set (27 patients, 58 CLs) of active Crohn’s patients (ascending, transverse, descending colon), centrally read by 3 expert readers, was used for initial evaluation. The model was optimized to segment 9 categories including bowel wall (BW), psoas, iliac vessels (IV), and inflamed mesentery (IM). Each output was based on a current frame and context representations from previous frames. The model continuously segments and tracks anatomy, generating temporally coherent delineations and data that feed intra- and inter-frame measurement algorithms. On the validation set, we calculated the Dice similarity coefficient (DSC), an overlap-based metric, from the expert annotations. Intraclass correlation coefficient (ICC) scores from ground truth measurements were made for every labelled image. Qualitative evaluation by domain experts confirmed consistency and anatomical precision of the outputs across a range of bowel regions and disease activity. Quantitatively, on the validation set our model achieved DSC of 0.73 [0.70 - 0.75], 0.26 [0.23 - 0.28], 0.62 [0.55 - 0.68], 0.63 [0.53 - 0.72], 0.72 [0.64 - 0.79] and 0.41 [0.32 - 0.49] for BW, peritoneal lining, rectus, psoas, IV and IM. Based on BW segmentations, the measurement algorithm achieved an ICC of 0.83 [0.77 - 0.87] between predictions and per frame expert measurements. On the external test dataset, the ICC between an aggregated AI-driven measurement per CL and the mean of all reader measurements was 0.83 [0.73 - 0.9]. Unlike prior approaches focused on static frames or derived BWT estimation, this method preserves frame-to-frame continuity, enabling reproducible analysis of CLs with no manual pre-selection of measurement areas. Our model provides identification of many anatomical structures in addition to BWT, essential to moving AI beyond this measure and towards fully automated disease activity monitoring. Conflict of interest: Novak, Kerri L.: Research Grants: Helmsley Trust, Pfizer, Janssen Adboard, consulting fees: Abbvie, Janssen, Pfizer, Pendopharm, Takeda, Elli Lilly, Celltrion, Bristal Myers Non financial support (ultrasound machine) McKesson Pharmcy. Pfizer, Celltrion Mendel, Robert: No conflict of interest Dolinger, Michael: Personal Fees: Michael Dolinger is a consultant for Neruologica., a subsidiary of Samsung Electronics Co., Ltd. Maaser, Christian: Speaker honaria and/or Advisory honaria Abbvie, Alfasigma, Biogen, Falk Foundation, Galapagos, Gilead, J & J, MSD Sharp & Dome, Pfizer, Roche, Samsung, Takeda Gecse, Krisztina B.: Grant: Abbvie, Pfizer Inc, Celltrion and Galapagos/Alfasigma Personal Fees: Consultancy fees from AbbVie, Galapagos, Gilead, Immunic Therapeutics, Janssen Pharmaceuticals, Pfizer Inc., and Takeda and speaker’s honoraria from Celltrion, Eli Lilly, Janssen Pharmaceuticals, Pfizer Inc. and Takeda. Nash, Carla: No conflict of interest Smyth, Matthew: No conflict of interest Sagami, Shintaro: Shintaro Sagami has served as an advisory board member, consultant, or speaker for AbbVie, Alimentiv, Bristol Myers Squibb, Celltrion, EA Pharma, Eli Lilly, Ferring Pharmaceuticals, Gilead Sciences, Janssen Pharmaceuticals, Kyorin Pharmaceutical, Mitsubishi Tanabe Pharma, Mochida Pharmaceutical, Nippon Kayaku, Pfizer, Takeda, and Zeria Pharmaceutical, and has received research funding from Bristol Myers Squibb, EA Pharma, Gilead Sciences, Helmsley Charitable Trust, JIMRO, Kyorin Pharmaceutical, Miyarisan, Mochida Pharmaceutical, Nippon Kayaku, Pfizer, Sekisui Medical, Samsung, Takeda, and Zeria Pharmaceutical. Nylund, Kim: Other: PI in clinical trial (Takeda) Ellis, Edward: No conflict of interest Sanghera, Daljinder: No conflict of interest Dr. Flegg, Daniel: No conflict of interest Torisu, Misaki: No conflict of interest Fu, Y. Nancy: No conflict of interest Ernest-Suárez, Kenneth: Consulting/Advisory Board fees: Abbvie, AstraZeneca, Johnson & Johnson, Pfizer, Ferring, Sandoz, SatisfAI, Takeda Johannessen, Solveig: No conflict of interest Gurm, Sunny: No conflict of interest Panaccione, Remo: No conflict of interest Wilkens, Rune Levring: Personal Fees: Janssen, Takeda Denmark, AbbVie, Pfizer Denmark, Alimentiv Byrne, Michael: Founder and shareholder, Dova Health Intelligence
Acquiring and annotating large datasets in ultrasound imaging is challenging due to low contrast, high noise, and susceptibility to artefacts. This process requires significant time and clinical expertise. Self-supervised learning (SSL) offers a promising solution by leveraging unlabelled data to learn useful representations, enabling improved segmentation performance when annotated data is limited. Recent state-of-the-art developments in SSL for video data include V-JEPA, a framework solely based on feature prediction, avoiding pixel level reconstruction or negative samples. We hypothesise that V-JEPA is well-suited to ultrasound imaging, as it is less sensitive to noisy pixel-level detail while effectively leveraging temporal information. To the best of our knowledge, this is the first study to adopt V-JEPA for ultrasound video data. Similar to other patch-based masking SSL techniques such as VideoMAE, V-JEPA is well-suited to ViT-based models. However, ViTs can underperform on small medical datasets due to lack of inductive biases, limited spatial locality and absence of hierarchical feature learning. To improve locality understanding, we propose a novel 3D localisation auxiliary task to improve locality in ViT representations during V-JEPA pre-training. Our results show V-JEPA with our auxiliary task improves segmentation performance significantly across various frozen encoder configurations, with gains up to 3.4% using 100% and up to 8.35% using only 10% of the training data.
Background:While artificial intelligence (AI) shows high potential in decision support for diagnostic gastrointestinal endoscopy, its role in therapeutic endoscopy remains unclear. Third-space endoscopic procedures pose the risk of intraprocedural bleeding. Therefore, we aimed to develop an AI algorithm for intraprocedural blood vessel detection. Methods:Using a test dataset of 101 standardized video clips containing 200 predefined submucosal blood vessels, 19 endoscopists were evaluated for vessel detection rate (VDR) and time (VDT) with and without support of an AI algorithm. Endoscopists were grouped according to experience in endoscopic submucosal dissection. Results:With AI support, endoscopist VDR increased from 56.4% (95%CI CI 54.1–58.6) to 72.4% (95%CI CI 70.3–74.4). Endoscopist VDT dropped from 6.7 seconds (95%CI 6.2–7.1) to 5.2 seconds (95%CI 4.8–5.7). False-positive readings appeared in 4.5% of frames and were marked for a significantly shorter time than true positives (0.7 seconds [95%CI 0.55–0.87] vs. 6.0 seconds [95%CI 5.28–6.70]). Conclusions:AI improved the VDR and VDT of endoscopists during third-space endoscopy. While these data need to be corroborated by clinical trials, AI may prove to be an invaluable tool for improving safety and speed of endoscopic interventions.
Objective:Despite high stand-alone performance, studies demonstrate that artificial intelligence (AI)-supported endoscopic diagnostics often fall short in clinical applications due to human-AI interaction factors. This video-based trial on Barrett's esophagus aimed to investigate how examiner behavior, their levels of confidence, and system usability influence the diagnostic outcomes of AI-assisted endoscopy. Methods:The present analysis employed data from a multicenter randomized controlled tandem video trial involving 22 endoscopists with varying degrees of expertise. Participants were tasked with evaluating a set of 96 endoscopic videos of Barrett's esophagus in two distinct rounds, with and without AI assistance. Diagnostic confidence levels were recorded, and decision changes were categorized according to the AI prediction. Additional surveys assessed user experience and system usability ratings. Results:AI assistance significantly increased examiner confidence levels (p < 0.001) and accuracy. Withdrawing AI assistance decreased confidence (p < 0.001), but not accuracy. Experts consistently reported higher confidence than non-experts (p < 0.001), regardless of performance. Despite improved confidence, correct AI guidance was disregarded in 16% of all cases, and 9% of initially correct diagnoses were changed to incorrect ones. Overreliance on AI, algorithm aversion, and uncertainty in AI predictions were identified as key factors influencing outcomes. The System Usability Scale questionnaire scores indicated good to excellent usability, with non-experts scoring 73.5 and experts 85.6. Conclusions:Our findings highlight the pivotal function of examiner behavior in AI-assisted endoscopy. To fully realize the benefits of AI, implementing explainable AI, improving user interfaces, and providing targeted training are essential. Addressing these factors could enhance diagnostic accuracy and confidence in clinical practice.
Background and study aims: While artificial intelligence (AI) shows high potential in decision support for diagnostic gastrointestinal endoscopy, its role in therapeutic endoscopy remains unclear. Third space endoscopic procedures pose the risk of intraprocedural bleeding. Therefore, we aimed to develop an AI algorithm for intraprocedural blood vessel detection. Patients and Methods: Using a test dataset with 101 standardized video clips containing 200 predefined submucosal blood vessels, 19 endoscopists were evaluated for the vessel detection rate (VDR) and time (VDT) with and without support of an AI algorithm. Test subjects were grouped according to experience in ESD. Results: With AI support, endoscopists VDR increased from 56.4% [CI 54.1–58.6] to 72.4% [CI 70.3–74.4]. Endoscopists‘ VDT dropped from 6.7sec [CI 6.2-7.1] to 5.2sec [CI 4.8-5.7]. False positive (FP) readings appeared in 4.5% of frames and were marked significantly shorter than true positives (6.0sec [CI 5.28-6.70] vs. 0.7sec [CI 0.55-0.87]). Conclusions: AI improved the vessel detection rate and time of endoscopists during third space endoscopy. While these data need to be corroborated by clinical trials, AI may prove to be an invaluable tool for the improvement of endoscopic interventions.
In semi-supervised segmentation, the strong-weak augmentation scheme has gained significant traction. Typically, a teacher model predicts a pseudo-label or consistency target from a weakly augmented image, while the student is tasked with matching the prediction when given a strong augmentation. However, this approach, popularized in self-supervised learning, is constrained by the model's current state. Even though the approach has led to state-of-the-art improvements as part of various algorithms, the inherent limitation, being confined to what the teacher model can predict, remains. In Sinkhorn Output Perturbations, we introduce an algorithm that adds structured pseudo-label noise to the training, extending the strong-weak scheme to perturbations of the output beyond just input and feature perturbations. Our strategy softens the inherent limitations of the student-teacher methodologies by constructing noisy yet plausible pseudo-labels. Sinkhorn Output Perturbations impose no specific architectural requirements and can be integrated into any segmentation model and combined with other semi-supervised strategies. Our method achieves state-of-the-art results on Cityscapes and presents competitive performance on Pascal VOC 2012, further improved upon combining our with another recent algorithm. The experiments also show the efficacy of the reallocation algorithm and provide further empirical insights into pseudo-label noise in semi-supervised segmentation. Code is available at:
Real-time computational speed and a high degree of precision are requirements for computer-assisted interventions. Applying a segmentation network to a medical video processing task can introduce significant inter-frame prediction noise. Existing approaches can reduce inconsistencies by including temporal information but often impose requirements on the architecture or dataset. This paper proposes a method to include temporal information in any segmentation model and, thus, a technique to improve video segmentation performance without alterations during training or additional labeling. With Motion-Corrected Moving Average, we refine the exponential moving average between the current and previous predictions. Using optical flow to estimate the movement between consecutive frames, we can shift the prior term in the moving-average calculation to align with the geometry of the current frame. The optical flow calculation does not require the output of the model and can therefore be performed in parallel, leading to no significant runtime penalty for our approach. We evaluate our approach on two publicly available segmentation datasets and two proprietary endoscopic datasets and show improvements over a baseline approach.
Considering the increase in the number of the Barrett's esophagus (BE) in the last decade, and its expected continuous increase, methods that can provide an early diagnosis of dysplasia in BE-diagnosed patients may provide a high probability of cancer remission. The limitations related to traditional methods of BE detection and management encourage the creation of computer-aided tools to assist in this problem. In this work, we introduce the unsupervised Optimum-Path Forest (OPF) classifier for learning visual dictionaries in the context of Barrett's esophagus (BE) and automatic adenocarcinoma diagnosis. The proposed approach was validated in two datasets (MICCAI 2015 and Augsburg) using three different feature extractors (SIFT, SURF, and the not yet applied to the BE context A-KAZE), as well as five supervised classifiers, including two variants of the OPF, Support Vector Machines with Radial Basis Function and Linear kernels, and a Bayesian classifier. Concerning MICCAI 2015 dataset, the best results were obtained using unsupervised OPF for dictionary generation using supervised OPF for classification purposes and using SURF feature extractor with accuracy nearly to 78% for distinguishing BE patients from adenocarcinoma ones. Regarding the Augsburg dataset, the most accurate results were also obtained using both OPF classifiers but with A-KAZE as the feature extractor with accuracy close to 73%. The combination of feature extraction and bag-of-visual-words techniques showed results that outperformed others obtained recently in the literature, as well as we highlight new advances in the related research area. Reinforcing the significance of this work, to the best of our knowledge, this is the first one that aimed at addressing computer-aided BE identification using bag-of-visual-words and OPF classifiers, being the application of unsupervised technique in the BE feature calculation the major contribution of this work. It is also proposed a new BE and adenocarcinoma description using the A-KAZE features, not yet applied in the literature.
Limitations in computer-assisted diagnosis include lack of labeled data and inability to model the relation between what experts see and what computers learn. Even though artificial intelligence and machine learning have demonstrated remarkable performances in medical image computing, their accountability and transparency level must be improved to transfer this success into clinical practice. The reliability of machine learning decisions must be explained and interpreted, especially for supporting the medical diagnosis. While deep learning techniques are broad so that unseen information might help learn patterns of interest, human insights to describe objects of interest help in decision-making. This paper proposes a novel approach, DeepCraftFuse, to address the challenge of combining information provided by deep networks with visual-based features to significantly enhance the correct identification of cancerous tissues in patients affected with Barrett’s esophagus (BE). We demonstrate that DeepCraftFuse outperforms state-of-the-art techniques on private and public datasets, reaching results of around 95% when distinguishing patients affected by BE that is either positive or negative to esophageal cancer.
Background This study evaluated the effect of an artificial intelligence (AI)-based clinical decision support system on the performance and diagnostic confidence of endoscopists in their assessment of Barrett's esophagus (BE). Methods 96 standardized endoscopy videos were assessed by 22 endoscopists with varying degrees of BE experience from 12 centers. Assessment was randomized into two video sets: group A (review first without AI and second with AI) and group B (review first with AI and second without AI). Endoscopists were required to evaluate each video for the presence of Barrett's esophagus-related neoplasia (BERN) and then decide on a spot for a targeted biopsy. After the second assessment, they were allowed to change their clinical decision and confidence level. Results AI had a stand-alone sensitivity, specificity, and accuracy of 92.2%, 68.9%, and 81.3%, respectively. Without AI, BE experts had an overall sensitivity, specificity, and accuracy of 83.3%, 58.1%, and 71.5%, respectively. With AI, BE nonexperts showed a significant improvement in sensitivity and specificity when videos were assessed a second time with AI (sensitivity 69.8% [95%CI 65.2%-74.2%] to 78.0% [95%CI 74.0%-82.0%]; specificity 67.3% [95%CI 62.5%-72.2%] to 72.7% [95%CI 68.2%-77.3%]). In addition, the diagnostic confidence of BE nonexperts improved significantly with AI. Conclusion BE nonexperts benefitted significantly from additional AI. BE experts and nonexperts remained significantly below the stand-alone performance of AI, suggesting that there may be other factors influencing endoscopists' decisions to follow or discard AI advice.
Even though artificial intelligence and machine learning have demonstrated remarkable performances in medical image computing, their accountability and transparency level must be improved to transfer this success into clinical practice. The reliability of machine learning decisions must be explained and interpreted, especially for supporting the medical diagnosis. For this task, the deep learning techniques' black-box nature must somehow be lightened up to clarify its promising results. Hence, we aim to investigate the impact of the ResNet-50 deep convolutional design for Barrett's esophagus and adenocarcinoma classification. For such a task, and aiming at proposing a two-step learning technique, the output of each convolutional layer that composes the ResNet-50 architecture was trained and classified for further definition of layers that would provide more impact in the architecture. We showed that local information and high-dimensional features are essential to improve the classification for our task. Besides, we observed a significant improvement when the most discriminative layers expressed more impact in the training and classification of ResNet-50 for Barrett's esophagus and adenocarcinoma classification, demonstrating that both human knowledge and computational processing may influence the correct learning of such a problem.
Semantic segmentation is an essential task in medical imaging research. Many powerful deep-learning-based approaches can be employed for this problem, but they are dependent on the availability of an expansive labeled dataset. In this work, we augment such supervised segmentation models to be suitable for learning from unlabeled data. Our semi-supervised approach, termed Error-Correcting Mean-Teacher, uses an exponential moving average model like the original Mean Teacher but introduces our new paradigm of error correction. The original segmentation network is augmented to handle this secondary correction task. Both tasks build upon the core feature extraction layers of the model. For the correction task, features detected in the input image are fused with features detected in the predicted segmentation and further processed with task-specific decoder layers. The combination of image and segmentation features allows the model to correct present mistakes in the given input pair. The correction task is trained jointly on the labeled data. On unlabeled data, the exponential moving average of the original network corrects the student’s prediction. The combined outputs of the students’ prediction with the teachers’ correction form the basis for the semi-supervised update. We evaluate our method with the 2017 and 2018 Robotic Scene Segmentation data, the ISIC 2017 and the BraTS 2020 Challenges, a proprietary Endoscopic Submucosal Dissection dataset, Cityscapes, and Pascal VOC 2012. Additionally, we analyze the impact of the individual components and examine the behavior when the amount of labeled data varies, with experiments performed on two distinct segmentation architectures. Our method shows improvements in terms of the mean Intersection over Union over the supervised baseline and competing methods. Code is available at https://github.com/CloneRob/ECMT.
We investigate contrastive learning in a multi-task learning setting classifying and segmenting early Barrett’s cancer. How can contrastive learning be applied in a domain with few classes and low inter-class and inter-sample variance, potentially enabling image retrieval or image attribution? We introduce a data sampling strategy that mines per-lesion data for positive samples and keeps a queue of the recent projections as negative samples. We propose a masking strategy for the NT-Xent loss that keeps the negative set pure and removes samples from the same lesion. We show cohesion and uniqueness improvements of the proposed method in feature space. The introduction of the auxiliary objective does not affect the performance but adds the ability to indicate similarity between lesions. Therefore, the approach could enable downstream auto-documentation tasks on homogeneous medical image data.
Celiac disease is an autoimmune disorder caused by gluten that results in an inflammatory response of the small intestine.We investigated whether celiac disease can be detected using endoscopic images through a deep learning approach. The results show that additional clinical parameters can improve the classification accuracy. In this work, we distinguished between healthy tissue and Marsh III, according to the Marsh score system.We first trained a baseline network to classify endoscopic images of the small bowel into these two classes and then augmented the approach with a multimodality component that took the antibody status into account.
The endoscopic features associated with eosinophilic esophagitis (EoE) may be missed during routine endoscopy. We aimed to develop and evaluate an Artificial Intelligence (AI) algorithm for detecting and quantifying the endoscopic features of EoE in white light images, supplemented by the EoE Endoscopic Reference Score (EREFS). An AI algorithm (AI-EoE) was constructed and trained to differentiate between EoE and normal esophagus using endoscopic white light images extracted from the database of the University Hospital Augsburg. In addition to binary classification, a second algorithm was trained with specific auxiliary branches for each EREFS feature (AI-EoE-EREFS). The AI algorithms were evaluated on an external data set from the University of North Carolina, Chapel Hill (UNC), and compared with the performance of human endoscopists with varying levels of experience. The overall sensitivity, specificity, and accuracy of AI-EoE were 0.93 for all measures, while the AUC was 0.986. With additional auxiliary branches for the EREFS categories, the AI algorithm (AI-EoE-EREFS) performance improved to 0.96, 0.94, 0.95, and 0.992 for sensitivity, specificity, accuracy, and AUC, respectively. AI-EoE and AI-EoE-EREFS performed significantly better than endoscopy beginners and senior fellows on the same set of images. An AI algorithm can be trained to detect and quantify endoscopic features of EoE with excellent performance scores. The addition of the EREFS criteria improved the performance of the AI algorithm, which performed significantly better than endoscopists with a lower or medium experience level.
In this study, we aimed to develop an artificial intelligence clinical decision support solution to mitigate operator-dependent limitations during complex endoscopic procedures such as endoscopic submucosal dissection and peroral endoscopic myotomy, for example, bleeding and perforation. A DeepLabv3-based model was trained to delineate vessels, tissue structures and instruments on endoscopic still images from such procedures. The mean cross-validated Intersection over Union and Dice Score were 63% and 76%, respectively. Applied to standardised video clips from third-space endoscopic procedures, the algorithm showed a mean vessel detection rate of 85% with a false-positive rate of 0.75/min. These performance statistics suggest a potential clinical benefit for procedure safety, time and also training.
Even though artificial intelligence and machine learning have demonstrated remarkable performances in medical image computing, their level of accountability and transparency must be provided in such evaluations. The reliability related to machine learning predictions must be explained and interpreted, especially if diagnosis support is addressed. For this task, the black-box nature of deep learning techniques must be lightened up to transfer its promising results into clinical practice. Hence, we aim to investigate the use of explainable artificial intelligence techniques to quantitatively highlight discriminative regions during the classification of early-cancerous tissues in Barrett's esophagus-diagnosed patients. Four Convolutional Neural Network models (AlexNet, SqueezeNet, ResNet50, and VGG16) were analyzed using five different interpretation techniques (saliency, guided backpropagation, integrated gradients, input × gradients, and DeepLIFT) to compare their agreement with experts' previous annotations of cancerous tissue. We could show that saliency attributes match best with the manual experts' delineations. Moreover, there is moderate to high correlation between the sensitivity of a model and the human-and-computer agreement. The results also lightened that the higher the model's sensitivity, the stronger the correlation of human and computational segmentation agreement. We observed a relevant relation between computational learning and experts' insights, demonstrating how human knowledge may influence the correct computational learning.