INTRODUCTION:The American College of Gastroenterology assembled a multidisciplinary task force to evaluate the current state and future direction of artificial intelligence (AI) in gastroenterology, hepatology, and endoscopy leading to the development of consensus-based recommendations for responsible AI integration in clinical practice. METHODS:A total of 32 subject-matter experts and 12 industry partners, representing diverse practice settings and expertise, conducted subgroup literature reviews across 5 key areas (endoscopy, practice management clinical applications, training and education, inflammatory bowel disease and liver disease, ethics and equity). Draft statements were developed and rated on a 5-point Likert scale using a modified Delphi process. A consensus was set at ≥70% combined agreement. Nonconsensus items were revised and revoted electronically. RESULTS:A total of 43 statements, 40 (93%) reached consensus in round 1 and the remaining 3 achieved consensus after round 2. Evidence supports computer-aided detection improving adenoma detection rate and miss rate in controlled studies, with mixed real-world impact and insufficient long-term outcomes (e.g., interval colon cancer rate). Recommendations emphasize thorough validation and reduction of bias by heterogeneous Data sets. Outside endoscopy, ambient AI scribes, natural language processing (NLP)-enabled coding, workflow optimization, and previous authorization support show potential. Training recommendations endorse a structured AI curriculum while preserving independent procedural competence to avoid deskilling. In inflammatory bowel disease and hepatology, AI could help improve diagnostic accuracy, help predict risk of disease progression, and help guide therapy. Equity, governance, and reimbursement statements call for chain-of-custody data protections, specialty-society oversight, and payment models that reward quality and cost reduction. DISCUSSION:This consensus outlines how AI can augment rather than replace clinical expertise while promoting safety, transparency, interoperability, and equity. Priorities include pragmatic and prospective trials, multi-institutional data-sharing consortia, bias mitigation, and workforce training to enable trustworthy and clinically impactful AI adoption in gastroenterology, liver, and endoscopy care.
Artificial intelligence (AI) is rapidly transforming the management landscape of inflammatory bowel disease (IBD). While early applications in endoscopy, digital pathology and cross-sectional imaging drew substantial attention, next-generation AI systems that enable deeper disease understanding, personalized treatment and streamlined clinical workflows are now emerging. These advances encompass the multimodal integration of endoscopic, histological and molecular data ('endo-histo-omics'); AI-assisted assessment of the intestinal barrier; remote monitoring via wearables; and the incorporation of large language models for decision-making support and patient interactions. This Perspective traces the evolution of AI in IBD from domain-specific tools to foundational platforms supporting data-driven precision medicine. We highlight validated AI applications across diagnosis, monitoring, outcome prediction and neoplasia surveillance. We also explore the expectations of key stakeholders, including clinicians, patients, regulatory bodies and industry, and discuss unresolved challenges such as explainability, integration into workflows, reimbursement and environmental sustainability. By aligning innovation with ethical and clinical priorities, AI holds the potential to redefine IBD care. Its future will be shaped by collaboration, transparency and responsible implementation, ushering in a new era of personalized, efficient and equitable care for individuals with IBD.
Artificial intelligence (AI), encompassing various methods that enable computers to process complex data and perform tasks previously requiring human input, is set to play a pivotal role in the future of inflammatory bowel diseases (IBDs) health care and research. This review examines AI's role in IBD and outlines a roadmap of priorities for developments in the field. To fully harness AI's potential for advancing high-quality health care and scientific discovery in IBD, large and diverse data sets are essential, along with AI tools that undergo rigorous and continuous validation. In addition, because AI becomes increasingly integrated into IBD care, it will be essential to study and optimize clinician-AI interactions and to provide education for health care professionals on the effective use of AI.
BACKGROUND & AIMS:Assessing endoscopic activity is integral in the management of postoperative Crohn's disease (CD). We aimed to comprehensively characterize the reliability and responsiveness of different endoscopic instruments when used to assess postoperative CD activity. METHODS:Ileocolonoscopy videos (n = 70) from the PREVENT (Prospective, Multicenter, Randomized, Double-Blind, Placebo-Controlled Trial Comparing REMICADE ® [infliximab] and Placebo in the Prevention of Recurrence in Crohn's Disease Patients Undergoing Surgical Resection Who Are at an Increased Risk of Recurrence) trial were reviewed by 3 blinded central readers. Disease activity was assessed using the Rutgeerts and modified Rutgeerts scores, POCER (postoperative Crohn's endoscopic recurrence) index, REMIND (groupe de REcherche sur les Maladies INflammatoires Digestives) score, Simple Endoscopic Score for Crohn's Disease (SES-CD), and the Crohn's Disease Endoscopic Index of Severity (CDEIS). Reliability was quantified by the intraclass correlation coefficient (ICC). Responsiveness was quantified using the win probability (WinP) defined as the probability that a patient in the treatment (infliximab) group had a better score than a patient in the placebo group. The neoterminal ileum, anastomosis, and distal colon were scored separately. RESULTS:Interrater reliability was substantial for the Rutgeerts and modified Rutgeerts scores, ileal REMIND score, SES-CD, and CDEIS (ICC 0.74-0.80), moderate for the POCER index (ICC 0.49), and fair for the anastomotic REMIND score (ICC 0.30). A large degree of responsiveness was observed for the Rutgeerts and modified Rutgeerts scores, ileal REMIND score, SES-CD, and CDEIS (WinP 0.75-0.83). The degree of responsiveness for the POCER index and the anastomotic REMIND score was small (WinP 0.54 and 0.53, respectively). Estimates of index reliability and responsiveness were consistently lower when assessed at the anastomosis or distal colonic segment compared with the neoterminal ileum. CONCLUSIONS:Existing endoscopic indices are reliable and responsive for assessing postoperative CD activity in the neoterminal ileum, although are suboptimal for evaluation in the anastomosis or distal colonic segment.
BACKGROUND & AIMS:Endoscopic scoring of Crohn's disease (CD) is challenging, as mucosal disease is patchy with highly variable morphology, size, and severity. Computer vision may help quantify disease activity with similar performance as standard instruments like the Simple Endoscopic Score for Crohn's Disease (SES-CD). METHODS:Colonoscopy videos from the STARDUST and SEAVUE phase 3 clinical trials underwent post-hoc computer vision endoscopic (CVE) assessment to quantify CD mucosal ulceration and injury. A segmentation model was trained on hand annotations of images performed by 2 gastroenterologists, predicting ulcer area, severity, and relative size. Using complete endoscopic video, predicted ulceration and general mucosal injury were then spatially mapped to the ileum and colon to quantify CD burden. CVE ulceration and general injury values were compared with the SES-CD in terms of disease quantification, localization, and agreement with end-of-study clinical remission (Crohn's Disease Activity Index [CDAI] <150). RESULTS:Ulcer semantic segmentation models matched the performance of gastroenterologist annotators (Dice similarity coefficient, 0.591 vs 0.462), with neither performing better on qualitative review of disagreements. CVE measures were highly correlated with SES-CD scores (r = 0.73-0.85; P < .0001), although there was expected poor correlation with the degree of stenosis (r = 0.12-0.21). CVE ulcer measurements (28.4 vs 52.3; P = .0012) and SES-CD (6.3 vs 9.0; P = .0193) separated end-of-study clinical remission status, although CVE measures had greater effect size than SES-CD (g = 0.416 vs g = 0.290). Performance of CVE for CD was similar in the SEAVUE validation cohort. CONCLUSIONS:CVE provides a means for automated ulceration and mucosal injury quantitation that shows conceptual agreement with SES-CD. CVE offers new capabilities to improve the granularity and personalization of endoscopic disease assessment in CD.
Introduction. The long-term goal of our studies is to determine if, and to what extent, a multi-mineral product (Aquamin) could have beneficial impact on individuals with ulcerative colitis (UC). As a step toward achieving that goal, we carried out a 180-day biomarker trial in patients with UC in remission or at the mild stage. Approach. A total of 28 subjects were included in the study. Each was randomized to receive either Aquamin for 180 days or placebo for the first 90 days. At day-90, placebo subjects crossed over to Aquamin for the final 90 days. At days-0, -90 and -180, serum samples were assessed for alkaline phosphatase (ALP), intestine-specific ALP (ALPI), C-reactive protein (CRP) and for biomarkers of bone turnover (osteocalcin, TRAP5b and bone-specific ALP [e.g., BALP]). Stool specimens were assessed for fecal calprotectin at the same time points and colon biopsies were examined histologically. Each subject underwent DEXA scanning (day-0 and -180 only). In addition, a mass spectrometry-based proteomic assessment was performed using colon biopsy specimens obtained at each time point. Results. Subjects receiving Aquamin for the complete 180-day period (a total of 12) demonstrated improvement in all biomarkers; this was not seen in the placebo group (16 subjects). Subjects who received Aquamin for 90-days were intermediary in their responses. Subjects receiving Aquamin for 180-days also demonstrated increases in bone mineral density (BMD) and bone mineral content (BMC) resulting in a statistically-significant increase in the hip strength index over the period of treatment. This was accompanied by increases in osteocalcin and TRAP5b and by a decrease in BALP. The proteomic screen demonstrated up-regulation of multiple gut barrier proteins, cell surface transporter molecules and certain proteins with anti-inflammatory potential in response to Aquamin. Aquamin treatment also led to down-regulation of several proteins associated with the pro-inflammatory state. Conclusion. These findings suggest the potential value of multi-mineral intervention (Aquamin) as a low-cost, non-toxic adjuvant therapy for mild UC or for individuals with UC in remission. ### Competing Interest Statement The authors have declared no competing interest. ### Clinical Trial NCT03869905 ### Funding Statement This investigator-initiated trial was supported through discretionary funds (JV) provided by Marigot Inc. as a gift to the University of Michigan, as well as the University of Michigan Pandemic Research Recovery (PRR) funding awarded to MA, and funding from the American Society for Investigative Pathology (ASIP) Summer Research Opportunity Program in Pathology (SROPP) to MA. None of these entities played any role in or had any influence on the research activities (i.e., study design, recruitment, data collection, data interpretation, or data dissemination). This study also utilized services at the University of Michigan supported by NIH funding (UM1TR004404 to the Michigan Institute for Clinical and Health Research). ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: The interventional study was conducted with FDA approval of Aquamin as an Investigational New Drug (IND#141600) and with oversight by the Institutional Review Board at the University of Michigan Medical School (IRBMED). I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes - All data produced in the present study are available upon reasonable request to the authors - All data produced in the present work are contained in the manuscript - All generated proteomic data will be made publicly available in an online repository following the manuscript's publication.
INTRODUCTION:Ulcerative colitis (UC) disease assessment considers both mucosal healing and symptoms which can disagree. We compared the artificial intelligence (AI) endoscopic cumulative disease score (CDS) with the Mayo endoscopic score (MES) for agreement with symptoms and health related quality of life (HR-QoL) in UC. METHODS:Endoscopic video was obtained from TrueNorth, a 52-week phase 3 trial comparing ozanimod (OZA) with placebo (PBO) for UC. End of maintenance CDS was compared with endoscopically inactive (MES 0, 1) vs active (MES 2, 3) groups by partial Mayo remission (PMS, ≤2) and treatment using the Mann-Whitney U test. CDS and MES were compared with EuroQol 5 Dimension HR-QoL measures using the Kruskal-Wallis test. Low CDS cutoff for PMS remission used the Youden index, with agreement to PMS assessed using Cohen's κ. RESULTS:In 387 subjects, among endoscopically active subjects (MES 2, 3), the CDS was lower in those achieving PMS symptomatic remission (83.8 vs 186.9, P < 0.0001). EuroQol 5 Dimension had better agreement with CDS than MES (κ = 0.53 vs 0.44) and detected QoL differences within MES 2, 3 subjects for patient-reported health dimensions. Compared with low endoscopic maximum intensity (MES 0, 1), low-cumulative disease burden (CDS <40, area under the curve 0.85) had better agreement with PMS remission (κ = 0.57 vs 0.72). When defining adequate endoscopic criteria as PMS remission plus either low MES or low CDS, the difference between OZA vs PBO remission increased from 22.6% to 28.6%. DISCUSSION:AI-enabled endoscopic scoring has better agreement with symptomatic remission and HR-QoLs compared with MES and could quantitatively redefine adequate mucosal healing.
Artificial intelligence (AI) will fundamentally improve how we perform clinical trials by addressing issues of standardizing disease scoring, improving the sensitivity and precision of activity and phenotype assessments, and automating laborious and time-consuming study functions. Progress in AI image analysis is quickly proving to replicate expert judgment in endoscopy, histology, and cross-sectional imaging with speed, reproducibility, and reduced bias. However, AI analytics offer the ability to quantify disease characteristics with more detail and precision than human experts. Large language models and generative AI are automating the collection of high-quality data from electronic records and improving our ability to predict patient outcomes. This narrative review will focus on AI tools available today, their expected implementation, and future-facing opportunities for AI to reimagine inflammatory bowel disease clinical trials.
BACKGROUND:We previously identified circulating and MRI biomarkers associated with the surgical management of Crohn's disease (CD). Here we tested associations between these biomarkers and ileal resection inflammation and collagen content. METHODS:Fifty CD patients undergoing ileal resection were prospectively enrolled at 4 centers. Circulating CD64, extracellular matrix protein 1 (ECM1), GM-CSF autoantibodies (GM-CSF Ab), and fecal calprotectin were measured by ELISA. Ileal 3-dimensional magnetization transfer ratio (3D MTR), modified Look-Locker inversion recovery (MOLLI) T1 relaxation, diffusion-weighted intravoxel incoherent motion (IVIM), and the simplified magnetic resonance index of activity (sMaRIA) were measured by MRI. Ileal resection specimen acute inflammation was graded, and collagen content was measured quantitatively using second harmonic imaging microscopy. Associations between biomarkers and ileal collagen content were tested. RESULTS:Median (interquartile range [IQR]) age was 19.5 (16-33) years. We observed an inverse relationship between ileal acute inflammation and collagen content (r = -0.39 [95% confidence interval {CI}: -0.61, -0.10], P = .008). Most patients (33 [66%]) received biologics, with no variation in collagen content with treatment exposures. In the univariate analysis, CD64, GM-CSF Ab, fecal calprotectin, and sMaRIA were positively associated with acute inflammation and negatively associated with collagen content (P < .1). The multivariable model for ileal collagen content (R2 = 0.31 [95% CI: 0.11, 0.52]) included log CD64 (β = -.27; P = .19), log ECM1 (β = .47; P = .06), log GM-CSF Ab (β = -.15; P = .01), IVIM f (β = .29, P = .10), and IVIM D* (β = 1.69, P = .13). CONCLUSIONS:Clinically available and exploratory circulating and MRI biomarkers are associated with the degree of inflammation versus fibrosis in CD ileal resections. With further validation, these biomarkers may be used to guide medical and surgical decision-making for refractory CD.
INTRODUCTION:Assessing the cumulative degree of bowel injury in ileal Crohn's disease (CD) is difficult. We aimed to develop machine learning (ML) methodologies for automated estimation of cumulative ileal injury on computed tomography-enterography (CTE) to help predict future bowel surgery. METHODS:Adults with ileal CD using biologic therapy at a tertiary care center underwent ML analysis of CTE scans. Two fellowship-trained radiologists graded bowel injury severity at granular spatial increments along the ileum (1 cm), called mini-segments. ML segmentation methods were trained on radiologist grading with predicted severity and then spatially mapped to the ileum. Cumulative injury was calculated as the sum (S-CIDSS) and mean of severity grades along the ileum. Multivariate models of future small bowel resection were compared with cumulative ileum injury metrics and traditional bowel measures, adjusting for laboratory values, medications, and prior surgery at the time of CTE. RESULTS:In 229 CTE scans, 8,424 mini-segments underwent analysis. Agreement between ML and radiologists injury grading was strong (κ = 0.80, 95% confidence interval 0.79-0.81) and similar to inter-radiologist agreement (κ = 0.87, 95% confidence interval 0.85-0.88). S-CIDSS (46.6 vs 30.4, P = 0.0007) and mean cumulative injury grade scores (1.80 vs 1.42, P < 0.0001) were greater in CD biologic users that went to future surgery. Models using cumulative spatial metrics (area under the curve = 0.76) outperformed models using conventional bowel measures, laboratory values, and medical history (area under the curve = 0.62) for predicting future surgery in biologic users. DISCUSSION:Automated cumulative ileal injury scores show promise for improving prediction of outcomes in small bowel CD. Beyond replicating expert judgment, spatial enterography analysis can augment the personalization of bowel assessment in CD.
INTRODUCTION:A significant proportion of patients with acute severe ulcerative colitis (ASUC) require colectomy. METHODS:Patients with ASUC treated with upadacitinib and intravenous corticosteroids at 5 hospitals are presented. The primary outcome was 90-day colectomy rate. Secondary outcomes included frequency of steroid-free clinical remission, adverse events, and all-cause readmissions. RESULTS:Of the 25 patients with ASUC treated with upadacitinib, 6 (24%) patients underwent colectomy, 15 (83%) of the 18 patients with available data and who did not undergo colectomy experienced steroid-free clinical remission (1 patient did not have complete data), 1 (4%) patient experienced a venous thromboembolic event, while 5 (20%) patients were readmitted. DISCUSSION:Upadacitinib along with intravenous corticosteroids may be an effective treatment for ASUC.
PURPOSE:Qualitative findings in Crohn's disease (CD) can be challenging to reliably report and quantify. We evaluated machine learning methodologies to both standardize the detection of common qualitative findings of ileal CD and determine finding spatial localization on CT enterography (CTE). MATERIALS AND METHODS:Subjects with ileal CD and a CTE from a single center retrospective study between 2016 and 2021 were included. 165 CTEs were reviewed by two fellowship-trained abdominal radiologists for the presence and spatial distribution of five qualitative CD findings: mural enhancement, mural stratification, stenosis, wall thickening, and mesenteric fat stranding. A Random Forest (RF) ensemble model using automatically extracted specialist-directed bowel features and an unbiased convolutional neural network (CNN) were developed to predict the presence of qualitative findings. Model performance was assessed using area under the curve (AUC), sensitivity, specificity, accuracy, and kappa agreement statistics. RESULTS:In 165 subjects with 29,895 individual qualitative finding assessments, agreement between radiologists for localization was good to very good (κ = 0.66 to 0.73), except for mesenteric fat stranding (κ = 0.47). RF prediction models had excellent performance, with an overall AUC, sensitivity, specificity of 0.91, 0.81 and 0.85, respectively. RF model and radiologist agreement for localization of CD findings approximated agreement between radiologists (κ = 0.67 to 0.76). Unbiased CNN models without benefit of disease knowledge had very similar performance to RF models which used specialist-defined imaging features. CONCLUSION:Machine learning techniques for CTE image analysis can identify the presence, location, and distribution of qualitative CD findings with similar performance to experienced radiologists.
With increased application of natural language processing (NLP) in medicine, many NLP models are being developed for uncovering relevant clinical features from electronic health records. Temporal information plays a key role in understanding the context, significance, and interpretation of medical concepts extracted from clinical notes. This is particularly true in situations where the behavior, value, or status of a medical concept changes over time. In this paper, we introduce a systematic framework, NLP annotation-Relaxation-Generation (NRG). NRG compiles incidents of medical concept changes from status annotations and timestamps of multiple clinical notes. We demonstrate the effectiveness of the NRG pipeline by applying it to two medical concepts related to patients with inflammatory bowel disease: extra-intestinal manifestations and medications. We show that the NRG pipeline offers not only insights into medical concept changes over time, but can help convey longitudinal changes in clinical features at both individual and population level.
Rationale and ObjectivesWe present a machine learning and computer vision approach for a localized, automated, and standardized scoring of Crohn’s disease (CD) severity in the small bowel, overcoming the current limitations of manual measurements CT enterography (CTE) imaging and qualitative assessments, while also considering the complex anatomy and distribution of the disease.Materials and MethodsTwo radiologists introduced a severity score and evaluated disease severity at 7.5 mm intervals along the curved planar reconstruction of the distal and terminal ileum using 236 CTE scans. A hybrid model, combining deep-learning, 3-D CNN, and Random Forest model, was developed to classify disease severity at each mini-segment. Precision, sensitivity, weighted Cohen’s score, and accuracy were evaluated on a 20% hold-out test set.ResultsThe hybrid model achieved precision and sensitivity ranging from 42.4% to 84.1% for various severity categories (normal, mild, moderate, and severe) on the test set. The model’s Cohen’s score (κ = 0.83) and accuracy (70.7%) were comparable to the inter-observer agreement between experienced radiologists (κ = 0.87, accuracy = 76.3%). The model accurately predicted disease length, correlated with radiologist-reported disease length (r = 0.83), and accurately identified the portion of total ileum containing moderate-to-severe disease with an accuracy of 91.51%.ConclusionThe proposed automated hybrid model offers a standardized, reproducible, and quantitative local assessment of small bowel CD severity and demonstrates its value in CD severity assessment.