Background:Single-wavelength endoscopy (SWE) has shown promising results in assessing histological disease activity in ulcerative colitis. Our objective was to validate the real-time performance of a bedside prototype of SWE computer-aided diagnosis (CAD) as proof of concept. Methods:A bedside module for real-time use evaluated histological disease activity when endoscopy was performed in the rectum and sigmoid based on white-light endoscopy and SWE (410 nm monochromatic light). Biopsies were taken for reference and scored using the Geboes score (GBS). SWE-CAD displayed a blue or red indicator, corresponding to histological remission or non-remission, respectively. Simultaneous scoring using the Mayo Endoscopic Score was performed by the endoscopist. Results:In 36 patients, histological disease activity was automatically scored using SWE-CAD. At the section level, SWE-CAD showed an accuracy of 96.4%, sensitivity of 96.1%, and specificity of 85.5%. When differentiating disease activity into mild, moderate, and severe, accuracy was 97.7%, 62.8%, and 95.0%, respectively. At the per-patient level, overall diagnostic accuracy remained high at 94.4%, with only 2/36 underestimations compared with GBS. Conclusion:A novel SWE-CAD system demonstrated clinical accuracy of 94.4% at the per-patient level, potentially aiding physicians in interpreting subtle endoscopic abnormalities and paving the way to more individualized patient management.
Artificial intelligence (AI) holds significant potential for enhancing quality of gastrointestinal (GI) endoscopy, but the adoption of AI in clinical practice is hampered by the lack of rigorous standardisation and development methodology ensuring generalisability. The aim of the Quality Assessment of pre-clinical AI studies in Diagnostic Endoscopy (QUAIDE) Explanation and Checklist was to develop recommendations for standardised design and reporting of preclinical AI studies in GI endoscopy.The recommendations were developed based on a formal consensus approach with an international multidisciplinary panel of 32 experts among endoscopists and computer scientists. The Delphi methodology was employed to achieve consensus on statements, with a predetermined threshold of 80% agreement. A maximum three rounds of voting were permitted.Consensus was reached on 18 key recommendations, covering 6 key domains: data acquisition and annotation (6 statements), outcome reporting (3 statements), experimental setup and algorithm architecture (4 statements) and result presentation and interpretation (5 statements). QUAIDE provides recommendations on how to properly design (1. Methods, statements 1-14), present results (2. Results, statements 15-16) and integrate and interpret the obtained results (3. Discussion, statements 17-18).The QUAIDE framework offers practical guidance for authors, readers, editors and reviewers involved in AI preclinical studies in GI endoscopy, aiming at improving design and reporting, thereby promoting research standardisation and accelerating the translation of AI innovations into clinical practice.
The current landscape of machine learning models in GI endoscopy is fraught with considerable variability in methodologies and quality, posing challenges for validation and generalization. To ensure the effective integration of AI in clinical practice, it is crucial to develop and validate models rigorously across diverse and representative datasets. This involves standardizing reference standards, ensuring thorough external validation, using representative patient populations, and incorporating a range of image qualities. Addressing these methodological discrepancies will enhance the reliability and robustness of AI models, thereby facilitating their adoption and improving patient care in GI endoscopy.
Background and Aims Ulcerative colitis (UC) management employs a strategy targeting histological and endoscopic remission. Correlation of white light endoscopy (WLE) scores with histological activity is limited. Single-wavelength endoscopy (SWE), addressing microvascular changes reflecting histological disease activity, may better assess histological remission. Our goal was to assess the accuracy of a computer-aided diagnosis (CAD) system for histological activity estimation in UC, based on either WLE or SWE. Methods We collected 6926 sets of corresponding WLE and SWE frames in 112 patients with UC, using a prototype endoscopic system enabling both imaging methods (FUJIFILM, Tokyo, Japan). Histological remission (Geboes score <= 2B.0) assessed at the location of imaging was annotated for all frames, and separate WLE-CAD and SWE-CAD models were trained using deep learning for automated detection of histological remission with either imaging modality. Results Initial training of both models on the same subset of 42 patients resulted in SWE-CAD outperforming WLE-CAD with a mean sensitivity of 88.0% vs 73.9% (p < 0.001), a mean specificity of 71.7% vs 65.6% (p = 0.45), and a diagnostic accuracy of 83.3% vs 67.5% (p < 0.005), respectively. Consecutive training of the SWE-CAD model on the entire dataset (112 patients) resulted in an accuracy of 95.2%, sensitivity of 96.4%, and specificity of 92.9% on a section level. Conclusions By utilizing automated CAD based on non-magnifying SWE for enhanced capillary visibility vs WLE, histological remission was detected with 95.2% diagnostic accuracy in patients with UC, offering stable objectivity and helping to exclude inter-reader variability.
Background and aimRandomised trials show improved polyp detection with computer-aided detection (CADe), mostly of small lesions. However, operator and selection bias may affect CADe’s true benefit. Clinical outcomes of increased detection have not yet been fully elucidated.MethodsIn this multicentre trial, CADe combining convolutional and recurrent neural networks was used for polyp detection. Blinded endoscopists were monitored in real time by a second observer with CADe access. CADe detections prompted reinspection. Adenoma detection rates (ADR) and polyp detection rates were measured prestudy and poststudy. Histological assessments were done by independent histopathologists. The primary outcome compared polyp detection between endoscopists and CADe.ResultsIn 946 patients (51.9% male, mean age 64), a total of 2141 polyps were identified, including 989 adenomas. CADe was not superior to human polyp detection (sensitivity 94.6% vs 96.0%) but outperformed them when restricted to adenomas. Unblinding led to an additional yield of 86 true positive polyp detections (1.1% ADR increase per patient; 73.8% were <5 mm). CADe also increased non-neoplastic polyp detection by an absolute value of 4.9% of the cases (1.8% increase of entire polyp load). Procedure time increased with 6.6±6.5 min (+42.6%). In 22/946 patients, the additional detection of adenomas changed surveillance intervals (2.3%), mostly by increasing the number of small adenomas beyond the cut-off.ConclusionEven if CADe appears to be slightly more sensitive than human endoscopists, the additional gain in ADR was minimal and follow-up intervals rarely changed. Additional inspection of non-neoplastic lesions was increased, adding to the inspection and/or polypectomy workload.
Katholieke Universiteit Leuven, Belgium; Katholieke Universiteit Leuven Universitaire Ziekenhuizen Leuven, Belgium.
Katholieke Universiteit Leuven, Belgium; Krankenhaus Barmherzige Bruder Regensburg, Germany; AZ Maria Middelares vzw, Belgium; Klinikum Lippe Standort Bad Salzuflen, Germany; Poliambulatorio Nuovo Regina Margherita, Italy; Universitair Ziekenhuis Gent, Belgium; Hopital Erasme, Belgium; Narodowy Instytut Onkologii im Marii Sklodowskiej-Curie Panstwowy Instytut Badawczy w Warszawie, Poland; AZ Delta vzw, Belgium; Katholieke Universiteit Leuven Universitaire Ziekenhuizen Leuven, Belgium.
This ESGE Position Statement defines the expected value of artificial intelligence (AI) for the diagnosis and management of gastrointestinal neoplasia within the framework of the performance measures already defined by ESGE. This is based on the clinical relevance of the expected task and the preliminary evidence regarding artificial intelligence in artificial or clinical settings. MAIN RECOMMENDATIONS:: (1) For acceptance of AI in assessment of completeness of upper GI endoscopy, the adequate level of mucosal inspection with AI should be comparable to that assessed by experienced endoscopists. (2) For acceptance of AI in assessment of completeness of upper GI endoscopy, automated recognition and photodocumentation of relevant anatomical landmarks should be obtained in ≥90% of the procedures. (3) For acceptance of AI in the detection of Barrett's high grade intraepithelial neoplasia or cancer, the AI-assisted detection rate for suspicious lesions for targeted biopsies should be comparable to that of experienced endoscopists with or without advanced imaging techniques. (4) For acceptance of AI in the management of Barrett's neoplasia, AI-assisted selection of lesions amenable to endoscopic resection should be comparable to that of experienced endoscopists. (5) For acceptance of AI in the diagnosis of gastric precancerous conditions, AI-assisted diagnosis of atrophy and intestinal metaplasia should be comparable to that provided by the established biopsy protocol, including the estimation of extent, and consequent allocation to the correct endoscopic surveillance interval. (6) For acceptance of artificial intelligence for automated lesion detection in small-bowel capsule endoscopy (SBCE), the performance of AI-assisted reading should be comparable to that of experienced endoscopists for lesion detection, without increasing but possibly reducing the reading time of the operator. (7) For acceptance of AI in the detection of colorectal polyps, the AI-assisted adenoma detection rate should be comparable to that of experienced endoscopists. (8) For acceptance of AI optical diagnosis (computer-aided diagnosis [CADx]) of diminutive polyps (≤5 mm), AI-assisted characterization should match performance standards for implementing resect-and-discard and diagnose-and-leave strategies. (9) For acceptance of AI in the management of polyps ≥ 6 mm, AI-assisted characterization should be comparable to that of experienced endoscopists in selecting lesions amenable to endoscopic resection.
BACKGROUND : Artificial intelligence (AI) research in colonoscopy is progressing rapidly but widespread clinical implementation is not yet a reality. We aimed to identify the top implementation research priorities. METHODS : An established modified Delphi approach for research priority setting was used. Fifteen international experts, including endoscopists and translational computer scientists/engineers, from nine countries participated in an online survey over 9 months. Questions related to AI implementation in colonoscopy were generated as a long-list in the first round, and then scored in two subsequent rounds to identify the top 10 research questions. RESULTS : The top 10 ranked questions were categorized into five themes. Theme 1: clinical trial design/end points (4 questions), related to optimum trial designs for polyp detection and characterization, determining the optimal end points for evaluation of AI, and demonstrating impact on interval cancer rates. Theme 2: technological developments (3 questions), including improving detection of more challenging and advanced lesions, reduction of false-positive rates, and minimizing latency. Theme 3: clinical adoption/integration (1 question), concerning the effective combination of detection and characterization into one workflow. Theme 4: data access/annotation (1 question), concerning more efficient or automated data annotation methods to reduce the burden on human experts. Theme 5: regulatory approval (1 question), related to making regulatory approval processes more efficient. CONCLUSIONS : This is the first reported international research priority setting exercise for AI in colonoscopy. The study findings should be used as a framework to guide future research with key stakeholders to accelerate the clinical implementation of AI in endoscopy.
The number of publications in endoscopic journals that present deep learning applications has risen tremendously over the past years. Deep learning has shown great promise for automated detection, diagnosis and quality improvement in endoscopy. However, the interdisciplinary nature of these works has undoubtedly made it more difficult to estimate their value and applicability. In this review, the pitfalls and common misconducts when training and validating deep learning systems are discussed and some practical guidelines are proposed that should be taken into account when acquiring data and handling it to ensure an unbiased system that will generalize for application in routine clinical practice. Finally, some considerations are presented to ensure correct validation and comparison of AI systems.
Introduction Expert labelling of each frame in a polyp video is the most robust way for constructing a training set for deep learning, but this is very time-consuming and currently represents a major barrier for widespread implementation of AI in endoscopy. In this study, two alternative approaches are evaluated, an innovative semi-automated labelling tool and trained medical students providing annotations. Methods 20 unique polyp white light videos containing 6282 frames (14 adenomas and 6 sessile serrated lesions confirmed by histopathology, mean size 7mm, Olympus) were annotated with bounding boxes by a clinical expert. These annotations are used as the gold standard for comparison. Two cheaper annotation methods were then applied to evaluate their validity and relative performance: (1) a semi-automated labelling technique – this tool only requires 3 manually annotated video frames, from which a representation of the polyp is learned and transferred automatically to all the other frames in the video; (2) independent manual labelling of each video by three medical students – following a training module with polyp images and videos. Results The mean and standard deviation of the frame-level sensitivity, positive predictive value (PPV) and adjudicated PPV (for borderline low-quality frames) over all videos are provided in table 1. The semi-automated method significantly outperforms all three students on sensitivity and annotation time (paired t-test, p-value < 0.05), while also achieving the highest value for PPV, both before and after adjudication. Conclusions A semi-automated labelling tool is a faster, more efficient and valid approach for polyp detection. It outperforms three medical students, specifically trained for polyp recognition and is comparable to clinical expert performance.
In many medical imaging and classical computer vision tasks, the Dice score and Jaccard index are used to evaluate the segmentation performance. Despite the existence and great empirical success of metric-sensitive losses, i.e. relaxations of these metrics such as soft Dice, soft Jaccard and Lovász-Softmax, many researchers still use per-pixel losses, such as (weighted) cross-entropy to train CNNs for segmentation. Therefore, the target metric is in many cases not directly optimized. We investigate from a theoretical perspective, the relation within the group of metric-sensitive loss functions and question the existence of an optimal weighting scheme for weighted cross-entropy to optimize the Dice score and Jaccard index at test time. We find that the Dice score and Jaccard index approximate each other relatively and absolutely, but we find no such approximation for a weighted Hamming similarity. For the Tversky loss, the approximation gets monotonically worse when deviating from the trivial weight setting where soft Tversky equals soft Dice. We verify these results empirically in an extensive validation on six medical segmentation tasks and can confirm that metric-sensitive losses are superior to cross-entropy based loss functions in case of evaluation with Dice Score or Jaccard Index. This further holds in a multi-class setting, and across different object sizes and foreground/background ratios. These results encourage a wider adoption of metric-sensitive loss functions for medical segmentation tasks where the performance measure of interest is the Dice score or Jaccard index.
Aims Polyp size is directly correlated with risk of future CRC and growth invasiveness. Despite its big impact, endoscopists typically provide a visual size estimation, but several studies have reported low accuracies. We aim to enable more accurate in-vivo polyp size measurements and ultimately reduce clinical mis-sizing by endoscopists. Therefore, we developed an AI system that can objectively infer polyp size based on a reference tool in the endoscopic image.
In many countries, colonoscopy is part of the screening process for colorectal cancer (CRC). The recommended surveillance interval for patients depends on the findings during the procedure: the number of polyps found, histology and the size of these polyps are all considered when determining surveillance interval. Indeed, studies have shown that size matters as it is directly correlated with the risk of future CRC and growth invasiveness. Polyp size measurement is thus an important factor in the clinical decision-making.
Last year, we presented a deep learning framework for automated polyp detection. Contrary to classical CNNs, we use 'memory cells' enabling more accurate predictions. Little evidence is currently available on the performance of AI systems for polyp detection in real clinical practice. Additionally, studies have shown that 25% of all polyps are missed during colonoscopy, but it is unknown how many of these misses are due to failure of polyp recognition and how many due to suboptimal mucosal exposure.
Recent research on COVID-19 suggests that CT imaging provides useful information to assess disease progression and assist diagnosis, in addition to help understanding the disease. There is an increasing number of studies that propose to use deep learning to provide fast and accurate quantification of COVID-19 using chest CT scans. The main tasks of interest are the automatic segmentation of lung and lung lesions in chest CT scans of confirmed or suspected COVID-19 patients. In this study, we compare twelve deep learning algorithms using a multi-center dataset, including both open-source and in-house developed algorithms. Results show that ensembling different methods can boost the overall test set performance for lung segmentation, binary lesion segmentation and multiclass lesion segmentation, resulting in mean Dice scores of 0.982, 0.724 and 0.469, respectively. The resulting binary lesions were segmented with a mean absolute volume error of 91.3 ml. In general, the task of distinguishing different lesion types was more difficult, with a mean absolute volume difference of 152 ml and mean Dice scores of 0.369 and 0.523 for consolidation and ground glass opacity, respectively. All methods perform binary lesion segmentation with an average volume error that is better than visual assessment by human raters, suggesting these methods are mature enough for a large-scale evaluation for use in clinical practice.
Artificial intelligence (AI) and its application in medicine has grown large interest. Within gastrointestinal (GI) endoscopy, the field of colonoscopy and polyp detection is the most investigated, however, upper GI follows the lead. Since endoscopy is performed by humans, it is inherently an imperfect procedure. Computer‐aided diagnosis may improve its quality by helping prevent missing lesions and supporting optical diagnosis for those detected. An entire evolution in AI systems has been established in the last decades, resulting in optimization of the diagnostic performance with lower variability and matching or even outperformance of expert endoscopists. This shows a great potential for future quality improvement of endoscopy, given the outstanding diagnostic features of AI. With this narrative review, we highlight the potential benefit of AI to improve overall quality in daily endoscopy and describe the most recent developments for characterization and diagnosis as well as the recent conditions for regulatory approval.
Objective treatment targets are required in a treat-to-target approach for ulcerative colitis (UC). An automated endoscopic system that correlates with histology can be an objective predictor for sustained remission in UC. The infiltration of neutrophils is associated with irregularities of the pericryptal capillaries. For this, we aimed to develop an objective automated endoscopic tool to assess histological remission based on the evaluation of the morphology of the pericryptal capillaries during endoscopy. We used a prototype endoscopic system with short-wave-length monochromatic light from a LED system. This enables to evaluate in real-time the superficial (<200µm) mucosal architecture (crypts, pericryptal capillaries) and mucosal bleeding, Figure 1. An image analysis algorithm was applied to provide a score that quantifies the specific morphology of the mucosal capillaries. The algorithm included two steps. First, bleeding (mucosal/luminal) was assessed by pattern recognition. Samples with bleedings were automatically classified as non-remission. In case of non-bleeding, the degree of congestion of the capillaries was measured (maximal localised density estimation after morphological hessian based vessel recognition) to assess an ideal cut off value that identifies histological remission (Geboes score (GBS) <2B.1; no neutrophils in the lamina propria). Consecutive patients with UC were evaluated with the Mayo endoscopic subscore (MES), ulcerative colitis endoscopic index of severity (UCEIS) and the automated image analysis algorithm. To test the reliability of the algorithm and scores, the results were correlated with the GBS. Biopsies were taken in the matching area of the endoscopic evaluation. Fifty-eight patients with UC (53% male, median (IQR) age 41y (38–56), disease duration 7.1y (2.4–16.4)) with 113 evaluable segments (89% rectum or sigmoid) were included. The correlation between GBS and MES, UCEIS was good (r = 0.76, 0.75 respectively). The automated image analysis algorithm (Figure 2) detected histological remission with a higher performance (sens 0.79, spec 0.90) compared with UCEIS (sens 0.95, spec 0.69) and MES (sens 0.98, spec 0.61), resulting in a positive predictive value of 0.83, 0.65 and 0.59 for the automated image analysis algorithm, UCEIS and MES, respectively. The algorithm detected histological remission with high accuracy (86%). Mucosal capillary pattern recognition based on an automated image analysis with short-wave-length monochromatic light detects histological remission with high accuracy in UC. This technique provides an objective and quantitative tool to assess histological remission in UC, and excludes inter-reader variability.
Aims The management of Ulcerative Colitis (UC) requires objective targets. Automated endoscopic systems that correlate with histology can be objective and predictive for sustained remission. The infiltration of neutrophils is associated with irregularities of the pericryptal capillaries. We aimed to develop an objective automated endoscopic tool to assess histological remission based on the evaluation of the morphology of the pericryptal capillaries during endoscopy.