Objective: To report real-world data on hypothermic oxygenated perfusion (HOPE) and normothermic machine perfusion (NMP) in liver transplantation (LT). Summary Background Data: Real-world comparisons between HOPE and NMP are limited and methodologically challenging due to heterogeneity in donor and recipient risk profiles and regional differences in practice patterns. Methods: This international cohort study analyzed consecutive NMP-preserved LTs performed at 15 predominantly North American centers between 2021 and 2025. Outcomes were compared with the European HOPE-REAL cohort, comprising HOPE-treated LTs from 22 centers between 2012 and 2021. Risk-adjusted analyses were performed, stratified by graft type and risk category. Imbalances in baseline characteristics were addressed using entropy balancing. Results: A total of 954 NMP-treated and 1202 HOPE-treated grafts were analyzed, revealing substantial differences in donor risk. Extended-criteria DBD grafts accounted for 30% versus 64%, and futile DCD grafts for 10% versus 30%, in the NMP and HOPE cohorts, respectively. In the NMP cohort, death-censored graft survival at 1, 2, and 3 years exceeded 96% for DBD grafts and 94% for DCD grafts. Comparable outcomes were observed in the HOPE cohort, with 93% survival in DBD and 87% in DCD grafts at up to 3 years, despite significantly higher donor risk in the HOPE-DCD cohort. After risk adjustment, death-censored graft survival remained similar between both modalities across graft types and risk categories. Conclusions: Real-world data on HOPE-treated and NMP-treated LT demonstrate excellent outcomes. Nevertheless, compared with HOPE, further high-quality evidence and longer preservation time is needed to substantiate the clinical benefits of NMP in high-risk grafts.
To establish international benchmark values for relevant outcome parameters in robotic Whipple. For safe adoption of surgical innovation, robust quality control is essential. Benchmarking is a validated tool for assessing surgical performance. Recent international consensus identified establishing benchmark values for robotic Whipple as top priority. We analyzed consecutive patients undergoing robotic Whipple between 2020-2023 with a minimum one-year follow-up. Reference centers were required to perform ≥15 cases/year, be scientifically active in the field, and maintain a prospective database. Benchmark criteria included benign or resectable malignant disease without neoadjuvant therapy, arterial resection, major co-morbidities, or significant previous abdominal surgery. Benchmarks were established for 13 outcome parameters. The benchmark cohort comprised 418 patients from 12 centers across four continents. Benchmark values were: conversion rate ≤4.3%, transfusion rate ≤2.1%, 6-month mortality ≤2.2%, major complications ≤23.2%, and CCI® ≤20.9. Clinically relevant pancreatic fistula (grade B/C) and hemorrhage (grade B/C) rates were ≤23.6% and ≤12.7%, respectively. For pancreatic ductal adenocarcinoma (n=123), the benchmark for lymph node yield was ≥20. Higher surgical difficulty was associated with increased overall postoperative morbidity (R 2 =0.38, P =0.019), higher center caseload with reduced pancreas-specific complications (R 2 =0.28, P =0.044). Independent POPF predictors included duct diameter ≤4 mm (OR 1.37, 95% CI: 1.03, 1.82), anticoagulation (OR 2.45, 95% CI: 1.47, 3.99), and indication other than PDAC (OR 2.33, 95% CI: 1.68, 3.27). This study establishes the first international benchmarks for robotic Whipple, demonstrating oncologic outcomes and morbidity comparable to open surgery with the benefits of minimally invasive surgery.
OBJECTIVE:To report real-world data on hypothermic oxygenated perfusion (HOPE) and normothermic machine perfusion (NMP) in liver transplantation (LT). SUMMARY BACKGROUND DATA:Real-world comparisons between HOPE and NMP are limited and methodologically challenging due to heterogeneity in donor and recipient risk profiles and regional differences in practice patterns. METHODS:This international cohort study analyzed consecutive NMP-preserved LTs performed at 15 predominantly North American centers between 2021 and 2025. Outcomes were compared with the European HOPE-REAL cohort, comprising HOPE-treated LTs from 22 centers between 2012 and 2021. Risk-adjusted analyses were performed, stratified by graft type and risk category. Imbalances in baseline characteristics were addressed using entropy balancing. RESULTS:A total of 954 NMP-treated and 1202 HOPE-treated grafts were analyzed, revealing substantial differences in donor risk. Extended-criteria DBD grafts accounted for 30% versus 64%, and futile DCD grafts for 10% versus 30%, in the NMP and HOPE cohorts, respectively. In the NMP cohort, death-censored graft survival at 1, 2, and 3 years exceeded 96% for DBD grafts and 94% for DCD grafts. Comparable outcomes were observed in the HOPE cohort, with 93% survival in DBD and 87% in DCD grafts at up to 3 years, despite significantly higher donor risk in the HOPE-DCD cohort. After risk adjustment, death-censored graft survival remained similar between both modalities across graft types and risk categories. CONCLUSIONS:Real-world data on HOPE-treated and NMP-treated LT demonstrate excellent outcomes. Nevertheless, compared with HOPE, further high-quality evidence and longer preservation time is needed to substantiate the clinical benefits of NMP in high-risk grafts.
OBJECTIVE:To establish international benchmark values for relevant outcome parameters in robotic Whipple. SUMMARY BACKGROUND DATA:For safe adoption of surgical innovation, robust quality control is essential. Benchmarking is a validated tool for assessing surgical performance. Recent international consensus identified establishing benchmark values for robotic Whipple as top priority. METHODS:We analyzed consecutive patients undergoing robotic Whipple between 2020-2023 with a minimum one-year follow-up. Reference centers were required to perform ≥15 cases/year, be scientifically active in the field, and maintain a prospective database. Benchmark criteria included benign or resectable malignant disease without neoadjuvant therapy, arterial resection, major co-morbidities, or significant previous abdominal surgery. Benchmarks were established for 13 outcome parameters. RESULT:The benchmark cohort comprised 418 patients from 12 centers across four continents. Benchmark values were: conversion rate ≤4.3%, transfusion rate ≤2.1%, 6-month mortality ≤2.2%, major complications ≤23.2%, and CCI® ≤20.9. Clinically relevant pancreatic fistula (grade B/C) and hemorrhage (grade B/C) rates were ≤23.6% and ≤12.7%, respectively. For pancreatic ductal adenocarcinoma (n=123), the benchmark for lymph node yield was ≥20. Higher surgical difficulty was associated with increased overall postoperative morbidity (R2=0.38, P=0.019), higher center caseload with reduced pancreas-specific complications (R2=0.28, P=0.044). Independent POPF predictors included duct diameter ≤4 mm (OR 1.37, 95% CI: 1.03, 1.82), anticoagulation (OR 2.45, 95% CI: 1.47, 3.99), and indication other than PDAC (OR 2.33, 95% CI: 1.68, 3.27). CONCLUSIONS:This study establishes the first international benchmarks for robotic Whipple, demonstrating oncologic outcomes and morbidity comparable to open surgery with the benefits of minimally invasive surgery.
OBJECTIVE:To estimate the minimal important difference (MID) of the Comprehensive Complication Index (CCI ® ) in patients undergoing abdominal surgery. BACKGROUND:The CCI ® is a validated metric that quantifies cumulative surgical morbidity. While the CCI ® is a sensitive endpoint to detect treatment effects, a statistically significant effect does not necessarily translate into clinical relevance. Relevant differences from the patients' perspective are best captured by the MID. METHODS:Individual patient data were extracted from surgical studies reporting CCI ® at 30 days and using patient-reported outcome measures with established MIDs at baseline and 30 days. To determine the MID for the CCI ® , we used an anchor-based approach as recommended by methods guidelines. A patient-reported outcome measure was selected as an anchor only if the Spearman correlation coefficient between its change in score (baseline to 30 days postoperative) and the CCI ® was ≥|0.30|. We used linear regression to estimate the MID of the CCI ® across different anchors, and triangulation to determine a single MID. RESULTS:Data were extracted from 3 published randomized controlled trials and 1 prospective observational study (n = 1583 patients) in major abdominal surgery. In colorectal surgery cohorts, 2 subscores of the Short Form-36, 2 subscores of the Multidimensional Fatigue Inventory-20, the EuroQol-5-Dimension Index Score, and the EuroQol Visual Analog Scale showed a correlation with the CCI ® of ≥|0.30|. This resulted in MID estimates for the CCI ® ranging from 6.1 to 22.2. In hepato-pancreato-biliary surgery, 1 subscore of the Short Form-36, and 2 subscores of the Patient Reported Outcome Measure Information System-29 questionnaire qualified as anchors providing MID estimates ranging from 6.2 to 13.8. CONCLUSIONS:We propose a mean difference of 12 points in the CCI ® between treatment groups as a relevant difference in patients undergoing abdominal surgery. This MID provides an important foundation for sample size calculations and interpretation of randomized controlled trials and large real-world observational studies.
Robotic Whipple holds the promise to overcome safety concerns associated with laparoscopy, paving the way for widespread implementation of minimal-invasive surgery in this complex procedure. However, randomized data comparing robot vs. open Whipple demonstrate more pancreas-specific complications and R1-resections in the robotic arm. Recent international consensus identified establishing benchmarks as critical to ensure safe adoption of the robot. Benchmarking is a validated quality improvement tool, enabling comparison of surgical performance. The aim was to define benchmarks for outcome parameters in robotic Whipple. We analyzed consecutive patients undergoing robotic Whipple from January 2020 until December 2023 in 11 centers across 4 continents, with a minimum one-year follow-up. Centers had to perform ≥15 cases/year and have mounted their learning curve. Benchmark criteria included benign or resectable malignant disease without neoadjuvant therapy, arterial resection, major co-morbidities, or significant previous abdominal surgery. Medians across centers represented benchmark cutoffs. Eleven centers performed 1’037 Whipple procedures, of which 603 (58%) were benchmark cases. One third (n=192) were pancreatic ductal adenocarcinoma (PDAC) patients. Key benchmarks at 6 months included ≤1.2% mortality, ≤24.2% major complications, and ≤ 8.7 points Comprehensive Complication Index®. Pancreas-specific cutoffs included ≤13.0% postoperative pancreatic fistula (POPF) B/C and ≤3.4% post-pancreatectomy hemorrhage B/C, with 100% R0-resection and ≥19 harvested lymph nodes in PDAC patients. One-year actuarial overall and recurrence-free survival was 87% and 77%. In the entire cohort POPF B/C occurred in 16% (n=195). Independent POPF predictors included duct diameter ≤4mm (OR 1.79 95%CI [1.27-2.55]), anticoagulation (OR 3.68 95%CI [2.14-6.24]), and indication other than PDAC (OR 3.17, 95%CI [2.13-4.85]). This study establishes benchmarks for key outcomes in robotic Whipple, demonstrating oncologic adequacy and morbidity comparable to open surgery. Risk factors for POPF in open surgery also hold true in the robotic approach.
To estimate the minimal important difference (MID) of the Comprehensive Complication Index (CCI®) in patients undergoing abdominal surgery. The CCI® is a validated metric that quantifies cumulative surgical morbidity. While the CCI® is a sensitive endpoint to detect treatment effects, a statistically significant effect does not necessarily translate into clinical relevance. Relevant differences from the patients’ perspective are best captured by the MID. Individual patient data were extracted from surgical studies reporting CCI® at 30 days and using patient-reported outcome measures with established MIDs at baseline and 30 days. To determine the MID for the CCI®, we used an anchor-based approach as recommended by methods guidelines. A patient-reported outcome measure was selected as an anchor only if the Spearman correlation coefficient between its change in score (baseline to 30 days postoperative) and the CCI® was ≥|0.30|. We used linear regression to estimate the MID of the CCI® across different anchors, and triangulation to determine a single MID. Data were extracted from 3 published randomized controlled trials and 1 prospective observational study (n = 1583 patients) in major abdominal surgery. In colorectal surgery cohorts, 2 subscores of the Short Form-36, 2 subscores of the Multidimensional Fatigue Inventory-20, the EuroQol-5-Dimension Index Score, and the EuroQol Visual Analog Scale showed a correlation with the CCI® of ≥|0.30|. This resulted in MID estimates for the CCI® ranging from 6.1 to 22.2. In hepato-pancreato-biliary surgery, 1 subscore of the Short Form-36, and 2 subscores of the Patient Reported Outcome Measure Information System-29 questionnaire qualified as anchors providing MID estimates ranging from 6.2 to 13.8. We propose a mean difference of 12 points in the CCI® between treatment groups as a relevant difference in patients undergoing abdominal surgery. This MID provides an important foundation for sample size calculations and interpretation of randomized controlled trials and large real-world observational studies.
Oncologic right hemicolectomy (rHC) remains the only curative treatment for right-sided colon cancer. Despite its increasing complexity, this procedure is not centralized in many countries, underscoring the need for rigorous assessment and continuous improvement in surgical quality. Benchmarking is a validated quality improvement tool. By defining best achievable outcomes as reference (i.e. benchmarks), it enables centers to evaluate their performance and identify weaknesses or areas for improvement. This analysis aimed to establish benchmarks for outcome parameters in minimal-invasive rHC. We analyzed data from consecutive patients with adenocarcinoma of the colon who underwent minimal-invasive rHC between July 2017 and June 2022 at 19 expert centers across five continents. Ideal cases were defined as elective surgeries for cT1-T3 tumors without distant metastases, major comorbidities, or significant prior abdominal surgeries. Benchmarks were derived for 19 clinically relevant surgical outcomes, including perioperative and oncological parameters, procedure-specific complications, overall morbidity, and mortality. Benchmarks were set at the 75th percentile for negative outcomes and the 25th percentile for positive outcomes across all centers’ medians. Among 3154 patients, 686 (22%) qualified as ideal. The proportion of ideal cases varied widely across centers (range: 2 – 51%). Key benchmarks at 3 months were overall morbidity ≤38%, major (Clavien-Dindo ≥3a) complications ≤8%, and 0% mortality. Procedure-specific benchmarks were anastomotic leak ≤3%, and deep surgical site infections ≤6%. Finally, oncologic benchmarks included R0 resection rates 100% and ≥12 lymph nodes harvested ≥96.9%. Ideal compared to non-ideal patients and centers performing ≥500 cases compared to <500 cases annually demonstrated superior outcomes. This study demonstrates that, despite its complexity, minimally invasive rHC can be performed with low morbidity and high oncological accuracy. The established benchmarks provide a reference for centers striving to achieve excellence in this procedure.
To develop a standardized, clinically applicable methodology for comparing surgical outcomes with established benchmark cut-offs and guiding structured quality improvement. Benchmarking compares clinical outcomes with defined performance thresholds to identify areas for improvement. While benchmark values are available for various surgical procedures, there is no standardized methodology that allows for direct application in clinical practice. This limits their use in routine quality management and continuous improvement processes. A structured quality improvement cycle was developed to compare own data with surgical benchmark cut-offs for ideal (and non-ideal patients). The approach includes periodic comparison of clinical outcomes with benchmark cut-offs, root cause analysis for deviations, and implementation of targeted interventions. The method was applied to a cohort of 98 patients undergoing low anterior resections between 2018 and 2023. Outcomes were analyzed over overlapping 18-month rolling windows, updated every 6 months, to track trends and assess adherence to benchmarks. The analysis revealed deviations from benchmark targets, especially in readmission rates due to ileostomy-related complications. Root cause analysis identified gaps in postoperative care and patient education. In response, targeted measures were implemented, including multimedia-based ostomy education, improved nutrition protocols, enhanced outpatient support, and structured follow-up. These interventions aim to reduce deviations and improve outcomes in future assessment cycles. This methodology allows comparison of surgical outcomes with established benchmark cut-offs to guide structured quality improvement. It enables surgical teams to identify outcome gaps, implement data-driven interventions, and foster continuous quality improvement. The framework is adaptable to various procedures with existing benchmarks and promotes evidence-based surgical excellence.
OBJECTIVE:To assess the reliability and construct validity of the CCI®️ following pancreatic surgery. SUMMARY BACKGROUND DATA:The Comprehensive Complication Index (CCI®️) is the only validated metric that quantifies cumulative morbidity, with a continuous score ranging from 0 (no complications) to 100 (death). METHODS:To address construct validity, we assessed patients undergoing elective pancreatic surgery for any disease at five Italian centers enrolled in a randomized controlled trial (NCT04438447) and a prospective cohort study (NCT04431076). The severity of 90-day complications was assessed using the CCI®️. We tested 10 a priori construct validity hypotheses through linear regression. Regression coefficients represented the between-group mean difference in CCI®️, with an effect size ≥0.2 considered potentially meaningful. Validity was deemed adequate if >75% of the hypotheses were supported. To address reliability, three independent raters among six centers assessed the CCI®️ from 100 anonymous case vignettes to evaluate inter-rater and inter-center reliability through intraclass correlation coefficient (ICC) and standard error of measurement (SEM). RESULTS:797 patients were included (66±11 y, 50% female, 60% malignancy). The construct validity was supported by data, with 9/10 a priori hypotheses confirmed (90%). The CCI®️ showed excellent inter-rater (ICC=0.96, 95%CI: 0.95-0.97), high inter-center reliability (ICC >0.75 in each center), with a SEM ranging from 2.73 to 6.38. CONCLUSIONS:This study supports CCI®️as a valid and reliable measure of morbidity after pancreatic surgery, supporting its use in both clinical practice and comparative effectiveness research.
Background:Studies examining machine perfusion (MP) in liver transplantation (LT) report variable post-transplant outcomes, making comparison difficult. Core outcome sets (COS) with standardized definitions and specified time points have been proposed to mitigate this variability. This study explores the quality of outcome reporting in studies with MP in LT. Methods:We conducted a systematic review examining outcome reporting within MP LT studies. PubMed, Embase and Ovid Medline were queried for studies examining perfusion techniques from January 1st 2018 to June 1st 2025. Articles that reported clinical outcomes after LT with current perfusion techniques, i.e., normothermic (NMP), hypothermic (HMP) and normothermic regional perfusion (NRP) with >10 cases were included. Risk of bias was assessed with Rob2 and Newcastle Ottawa scores. Median number of COS were reported by perfusion technique. The percentage of studies reporting each measured outcomes was recorded and classified by perfusion technique. The study was registered on PROSPERO (ID: CRD42024590000). Findings:2789 records were screened, and 166 articles met inclusion criteria. HMP, including oxygen per-sufflation, was examined by 35 studies, NMP by 55 and NRP by 34 studies, respectively. Forty-two studies included a combination of perfusion techniques. The median number of COS measured was highest in studies examining HMP at 9.0 (IQR 6.0-11.0). NMP studies had the lowest median number of COS measured per article (5.0; IQR 3.0-9.5). Of the 13 COS metrics, patient and graft survival were the most frequently reported, at 88.4% and 87.1% followed by primary-non-function (72.3%), length of stay (68.4%) and non-anastomotic biliary strictures (61.3). Weighted and cumulative complications metrics, i.e., the Clavien-Dindo-Classification and Comprehensive Complication Index were poorly assessed (28.4% and 19.4%, respectively). Only 4 studies examined all 13 COS parameters (2.6%), and 24 reported >10 COS (15.5%). There was significant heterogeneity in outcome parameter definitions and time-points of assessment, especially amongst biliary complications. Interpretation:There is significant variability in post-transplant complications reporting by LT with MP studies. Uniform definitions and standardized guidelines are critical to allow rigorous comparison of different perfusion techniques and other future innovations. Funding:None.
OBJECTIVE:To provide improved guidance for the consistent application of the Clavien-Dindo classification (CDC) and Comprehensive Complication Index (CCI ® ) in challenging clinical scenarios. BACKGROUND:Standardized outcome reporting is key for the proper assessment of surgical procedures. A recent consensus conference recommended the CDC and the CCI ® for assessing postoperative morbidity. Several challenging scenarios for grading complications still require evidence-based guidance, and the use of the 2 metrics in randomized controlled trials (RCTs) remains unexplored. METHODS:We assessed the use of the CDC and CCI ® as an outcome measure in a systematic literature search. In addition, we asked 163 international surgeons to critically evaluate and independently grade complications in 20 complex clinical scenarios. Finally, a Core Group of 5 experts used this information to develop consistent recommendations. RESULTS:Until July 2023, 1327 RCTs selected the CDC and/or CCI ® to assess morbidity. Annual use was steadily increasing with now over 200 new RCTs per year. However, only a third (n = 335) of published RCTs provided the complete range of CDC grades, including all subgrades. Eighty-nine out of 163 surgeons (response rate: 55%) completed the questionnaire that served as a basis for the recommendations: repetitive interventions that are required to treat one complication, complications followed by further complications, complications occurring before referral, and expected and unrelated complications to the original procedure should all be counted separately and included in the CCI ® . Invasive blank diagnostic interventions should not be considered a complication. CONCLUSIONS:The increasing use of the CDC and CCI ® in RCTs highlights the importance of their standardized application. The current consensus on various difficult scenarios may offer novel guidance for the consistent use of the CDC and CCI ® , aiming to improve complication reporting and better quality control, ultimately benefiting all health care stakeholders and, first and foremost, all patients.
The comparison of outcomes in liver transplantation (LT) is hampered by using clinically nonrelevant surrogate endpoints and considerable variability in reported relevant posttransplant outcomes. Such variability stems from nonstandard outcome measures across studies, variable definitions of the same complication, and different timing of reporting. The Clavien-Dindo classification was established to improve the rigor of outcome reporting but is nonspecific to an intervention, and there are unsolved dilemmas specifically related to LT. Core outcome sets (COSs) have been used in other specialties to standardize outcomes research, but have not been defined for LT. Thus, we use the 5 major benchmarking studies published to date to define a 10-measure COS for LT using previously validated metrics. We further provide standard definitions for each of the 10 measures that may be used in international research on the topic. These definitions also include standard time points for recording to facilitate between-study comparisons and future meta-analysis. These 10 outcomes are paired with 3 validated, procedure-independent metrics, including the Clavien-Dindo Classification and the Comprehensive Complications Index. The Clavien scale and Comprehensive Complications Index are specifically reviewed to enhance their utility in LT, and their use, along with the COS, is explored. We encourage future studies to employ this COS along with the Clavien-Dindo grading system and Comprehensive Complications Index to improve the reproducibility and generalizability of research concerning LT.
OBJECTIVES:To assess the current quality of surgical outcome reporting in the medical literature and to provide recommendations for improvement. BACKGROUND:In 1996, The Lancet labeled surgery as a "comic opera" mostly referring to the poor quality of outcome reporting in the literature impeding improvement in surgical quality and patient care. METHODS:We screened 3 first-tier and 2 second-tier surgical journals, as well as 3 leading medical journals for original articles reporting on results of surgical procedures published over a recent 18-month period. The quality of outcome reporting was assessed using a prespecified 12-item checklist. RESULTS:Six hundred twenty-seven articles reporting surgical outcomes were analyzed, including 125 randomized controlled trials. Only 1 (0.2%) article met all 12 criteria of the checklist, whereas 356 articles (57%) fulfilled less than half of the criteria. The poorest reporting was on cumulative morbidity burden, which was missing in 94% of articles (n=591) as well as patient-reported outcomes missing in 83% of publications (n=518). Comparing journal groups for the individual criterion, we found moderate to very strong statistical evidence for better quality of reporting in high versus lower impact journals for 7 of 12 criteria and strong statistical evidence for better reporting of patient-reported outcomes in medical versus surgical journals ( P <0·001). CONCLUSIONS:The quality of outcomes reporting in the medical literature remains poor, lacking improvement over the past 20 years on most key end points. The implementation of standardized outcome reporting is urgently needed to minimize biased interpretation of data thereby enabling improved patient care and the elaboration of meaningful guidelines.
Abstract Background Standardized outcome reporting is key for proper assessment of surgical procedures. A recent consensus conference recommended the Clavien-Dindo classification (CDC) and the Comprehensive Complication Index (CCI®) for assessing postoperative morbidity. However, their use in randomized controlled trials (RCTs) has not been assessed, and several challenging scenarios for grading complications require consensus-based guidance. Aims The aim of this study was to assess the use of the CDC and CCI® in RCTs and to provide guidance on their standardized and consistent application. Methods We identified all RCTs that used the CDC or CCI® as a primary or secondary outcome. In addition, we asked 163 international surgeons to independently grade complications of 20 clinical cases covering seven challenging scenarios. Finally, a core group of five experts used this information do develop consistent recommendations. Results Up to July 2023, 1424 RCTs used the CDC or CCI® to assess postoperative morbidity. Annual use was steadily increasing with now over 200 new RCTs per year. Eighty-nine (55%) surgeons completed the survey. Complications requiring multiple interventions, complications of complications, complications occurring prior to referral, and expected and unrelated complications should all be counted as separate complications and included in the CCI®. Invasive diagnostics without findings should not be considered as a complication since purely diagnostic. Conclusion We observed an extensive and steadily increasing use of CDC and CCI® in RCTs, highlighting the importance of their consistent application. Provided by the original developers of the CDC and CCI® and based on an international survey of their frequent users, the current consensus offers much-needed guidance for challenging scenarios. This will further improve the consistency and accuracy of complication reporting, leading to higher quality RCTs, improved cost estimations, and better quality control, ultimately benefiting all stakeholders.