Individual human health is inextricably linked to planetary health, which refers to the health of human civilization and the natural systems on which it depends. These systems are captured by the nine planetary boundaries which together define a safe operating space for humanity. The health sector contributes to the transgression of these boundaries, thereby jeopardizing both individual and planetary health. However, there is a growing interest in the health sector in incorporating the preservation of planetary boundaries into health guidelines and health technology assessments (HTAs), which synthesize evidence and guide healthcare decisions. This scoping review aims to describe methodological guidance and considerations for incorporating planetary health in health guidelines and HTAs. We conducted a scoping review adhering to the JBI methodology including articles from several databases since inception until September 2023. We used narrative synthesis and descriptive statistics to describe eligible studies. Of the 38 included studies, 14 (37%) were commentaries, six (16%) were methodological papers, and four (10·5%) were health guidelines. Included studies focused primarily on greenhouse gas emissions as an outcome. The included health guidelines were rated at a median of 52% (range: 16% to 67%) on AGREE II for methodological quality. Key findings pertain to the normative significance of planetary health in guidelines and HTAs, scope and quality of included studies, inconsistencies in methods, and challenges in implementation. These findings will inform future Grading of Recommendations, Assessment, Development, and Evaluations (GRADE) guidance on planetary health. A protocol of this scoping review was registered in Open Science Framework ( https://osf.io/3jmsa ) and published in the journal Systematic Reviews (10.1186/s13643-024-02577-2).
BACKGROUND Preoperative fasting aims to reduce pulmonary aspiration risk during anaesthesia. However, actual fasting durations often exceed recommendations, leading to patient discomfort, thirst and anxiety. Liberal fasting protocols may improve patient well-being without increasing aspiration risk, while gastric ultrasound may help identify patients with a higher risk of aspiration. OBJECTIVES This guideline aims to provide evidence-based recommendations on preoperative fasting duration for clear fluids and the use of gastric ultrasound assessment in adults undergoing elective procedures. METHODS A multidisciplinary task force will address two questions: (1) Should adults follow liberal (0–2 h) instead of standard (2–4 h) fasting for clear fluids before elective anaesthesia? (2) Should ultrasound be routinely used to assess gastric content before elective anaesthesia? Two systematic reviews will be conducted, as well as systematic searches to identify evidence on patient values and contextual factors. Evidence certainty will be rated using GRADE, and recommendations will be developed through evidence-to-decision frameworks considering benefits, harms, patient values, resources, acceptability, feasibility and equity. CONCLUSIONS This guideline aims to provide evidence-based recommendations on the duration of preoperative fasting for clear fluids and the use of gastric ultrasound to assess gastric content before anaesthesia for elective procedures in adults.
Background Aplastic anemia is a rare, life-threatening condition marked by pancytopenia and bone marrow hypocellularity. Despite therapeutic advances, clinical practice remains variable, and uncertainties persist regarding diagnosis and optimal management. To address these gaps, the American Society of Hematology (ASH) developed evidence-based guidelines to provide standardized, patient-centered recommendations. Objective The recommendations are intended to support patients, clinicians, and other health care professionals in their decisions about the management and diagnosis of severe and very severe immune-acquired aplastic anemia. Methods ASH formed a multidisciplinary guideline panel of content experts, methodologists, and a patient representative. An evidence synthesis team supported the guideline development process by conducting systematic evidence reviews. The panel prioritized clinical questions and used the grading of recommendations assessment, development, and evaluation approach, including the evidence-to-decision frameworks, to assess evidence and make recommendations, which were subject to public comment. Results The panel agreed on 33 recommendations and 4 good practice statements addressing the use of diagnostic tests, treatment strategies and supportive care. Additional recommendations covered the incorporation of eltrombopag into immunosuppressive regimens and the use of antimicrobial prophylaxis for patients at high risk. For most clinical questions, the certainty in the evidence was rated as low or very low, largely due to the reliance on small, nonrandomized studies. Conclusions Recommendations emphasize prioritizing hematopoietic cell transplantation for younger individuals with an available matched sibling or unrelated donor and as a second-line option after failure of immunosuppressive therapy. The panel also recommended adding eltrombopag to immunosuppressive regimens and using antibiotic and antifungal prophylaxis for patients with neutropenia.
Human health and natural systems are intrinsically linked-stable natural systems enable healthy human life. Health systems aim to promote, restore, and maintain health. Health systems may promote human health while having detrimental effects on natural systems, contributing to the transgression of planetary boundaries, such as biosphere integrity, climate change, and the introduction of new entities like microplastics. To date, the health guideline field lacks methods to assess the impacts of health interventions on planetary boundaries. The GRADE (Grading of Recommendations Assessment, Development and Evaluation) Working Group established the Planetary Health Project Group in 2023 to develop formal GRADE guidance for integrating planetary health into guideline recommendations to address this gap. Guided by the concepts of planetary health and planetary boundaries and following established methods for GRADE guidance development, the project group conducted iterative case study analyses, expert workshops, and a 2-round global Delphi consensus process. Four case studies were selected for application of this guidance before recommendations were finalized. The GRADE Working Group approved the official guidance. The Planetary Health Project Group presents 7 domains of guidance for incorporating planetary health aspects into the guideline development process, including highly desirable items and optional items. Highly desirable items include formally addressing planetary health in public health and health system guidelines and explicitly justifying its exclusion where it is not addressed. Judgments within the evidence-to-decision (EtD) framework should systematically integrate included evidence across the prioritized planetary boundaries and equity. This guidance aims to support guideline developers and policymakers in making evidence-based, trustworthy recommendations to protect individual and planetary health, while maintaining thoroughness and feasibility for guideline developers within the GRADE approach.
Introduction:The assessment of the certainty of evidence using GRADE requires evaluating inconsistency and imprecision. These domains are difficult to judge in a fully independent way. We performed a meta-research study to propose an approach to disentangle inconsistency and imprecision based on decision thresholds. Methods:We evaluated systematic reviews published in the Cochrane Database of Systematic Reviews. We included all meta-analyses published in these reviews (i) with at least five primary studies, (ii) which did not correspond to subgroup analyses, and (iii) which had provided their results as continuous outcomes or as dichotomous outcomes with previously established decision thresholds. We recalculated each included meta-analysis using the fixed- and the random-effects models, retrieving metrics potentially expressing inconsistency or imprecision. This includes inconsistency indices, variables related to the confidence intervals (CI) of the meta-analytical estimate and prediction intervals. Based on these variables, we performed factor analysis, hierarchical cluster analysis, and decision tree analysis. Results:We assessed 8703 meta-analyses. The factor analysis resulted in the identification of three factors: (i) one factor associated with inconsistency metrics, (ii) one factor associated with imprecision metrics, and (iii) one factor of metrics not clearly related to inconsistency or imprecision. Taken together, the analysis suggests an approach based on the number of decision thresholds crossed by the CI of the fixed- and random-effects meta-analytical estimates: the number of thresholds crossed by the CI of the fixed-effect model informs about imprecision, while the difference in the number of thresholds crossed by the CIs of the random- versus the fixed-effect models reflects inconsistency. Prediction intervals and inconsistency indices, such as the I 2, can help detect cases in which inconsistency may be underestimated. Conclusion:Our results may help those assessing the certainty in evidence to independently rate inconsistency and imprecision.
ABSTRACT Introduction Artificial intelligence (AI) may support several processes of the health guideline enterprise. This article describes the development of an extension of the Guidelines International Journal (GIN)‐McMaster Guideline Development Checklist (GDC) for integrating AI in the guideline enterprise. This development has been led by the GIN‐AI Working Group. Methods We started by prompting a large language model (LLM) for items related to the use of AI in each of the steps of the original GDC. Subsequently, the members of the working group engaged in a set of iterative discussions, resulting in item refinement and in a consensus first version of the extension. We retrospectively applied this first version to a case use of guidelines incorporating AI in their development (Allergic Rhinitis and its Impact on Asthma [ARIA] 2024‐2025 guidelines), leading to further refinement and to the approval of the final extension tool. Results Prompting LLMs resulted in the generation of 149 items. Of those, 117 were removed and 19 were modified by members of the working group. On the other hand, 17 new items were added during the iterative discussion process. The retrospective application of the extension led to changes in the wording of four items. The final version of the checklist extension has been approved with 49 items modifying or adding to the original GDC. Discussion We have developed an extension of the GIN‐McMaster GDC that encompasses a set of conduct standards that are intended to facilitate the comprehensive and transparent integration of AI in the health guideline enterprise. Clinical Trial Registration Not applicable. This study is not a clinical trial.
Background and Objectives Health guidelines play a central role in informing clinical practice, public health measures and health policy. But their trustworthiness may be undermined by factors such as insufficient methodological rigor, lack of transparency, conflicts of interest, and inconsistent application of established standards. Existing appraisal tools address selected aspects of guideline quality but do not comprehensively assess the trustworthiness of individual recommendations, nor do they adequately reflect recent advances in guideline methodology, including living guidelines, Grading of Recommendations, Assessment, Development, and Evaluation, adaptation, and the use of artificial intelligence (AI). This study aims to develop and validate Transparent, Rigorous, Useable, Standardized, and Trustworthy Guide (TRUSTGUIDES), a globally applicable, flexible set of tools to assess the trustworthiness of health guideline recommendations. We define trustworthiness as distinct from methodological quality: it encompasses not only rigorous methods but also transparency, independence, and applicability, which together determine whether a recommendation merits user confidence. Methods TRUSTGUIDES will be developed through a multistep, mixed-methods process. First, a scoping review and expert consultation will identify existing guideline appraisal tools and inform domains and items generation. Using deductive and inductive approaches, domains and items will be generated and may be refined through focus groups and selected through iterative Delphi surveys involving an international, multidisciplinary working group. TRUSTGUIDES will be validated by assessing internal consistency, inter-rater reliability, content validity, and construct validity, including comparisons with established instruments such as the Grading of Recommendations, Assessment, Development, and Evaluation certainty domains, AGREE II, and PANELVIEW. Psychometric properties will be examined using factor analysis and, as necessary, item response theory models. AI will be integrated both as an object of assessment and as methodological support for tool application, with large language models evaluated against a human reference standard. Conclusion TRUSTGUIDES will be designed to evaluate the trustworthiness of individual guideline recommendations across key factors, including transparency and credibility, and to address relevant domains such as the certainty of evidence, strength of recommendations, conflicts of interest, applicability, adaptability, currency, certification, and the appropriate use of AI. TRUSTGUIDES addresses critical gaps in current guideline appraisal by offering a comprehensive, recommendation-level assessment of trustworthiness aligned with the World Health Organization guideline standard methodology. By integrating AI, our tools will support efficient, transparent, and future-ready guideline evaluation within an evolving health evidence ecosystem.
INTRODUCTION:One of the challenges in managing patients with hantavirus infection is accurately identifying individuals who are at risk of developing severe disease. Prompt identification of these patients can facilitate critical decisions, such as early referral to an intensive care unit. The identified prognostic factors could be of utility in guiding medical care to enhance the management of hantavirus infection. OBJECTIVE:To identify and evaluate prognostic factors associated with mortality in hantavirus infection. METHODS:We conducted a Preferred Reporting Items for Systematic Reviews and Meta-Analyses-reported systematic review following Cochrane guidance adapted for prognosis. We searched PubMed/MEDLINE, Cochrane Central Register of Controlled Trials (CENTRAL), Biblioteca Virtual de Saúde or Lilac and EMBASE, from 1 January 1993 to 2 October 2025. We included studies evaluating individual prognostic factors or risk assessment models of New World hantavirus infections, with no restrictions on study design, publication status or language. When feasible, we conducted meta-analyses for prognostic factors using the inverse variance-based method with random effect model. We assessed the certainty of the evidence using the Grading of Recommendations Assessment, Development and Evaluation approach. RESULTS:We included 25 studies with a total of 7284 participants. We identified the following prognostic factors for which we found moderate to high certainty that are associated with increased mortality: age over 18 years, female sex, rural residence, elevated creatinine levels, increased haematocrit, signs of bleeding and the presence of infiltrates on chest radiographs. DISCUSSION:Our systematic review identified prognostic factors for mortality in patients with New World hantavirus infection. These factors can inform clinicians in making more informed management decisions. Furthermore, our findings lay the groundwork for the future development of a clinical prognostic model, potentially enhancing patient care and outcomes. PROSPERO REGISTRATION NUMBER:CRD42021225823.
BACKGROUND AND OBJECTIVE:The Grading of Recommendation Assessment, Development and Evaluation (GRADE) approach supports evidence syntheses, practice guidelines, health technology assessments, and health policies. As the number, type, and scope of GRADE publications have expanded and sciences advances, users have faced growing challenges in identifying official GRADE materials and utilizing the most current methods. To address this problem, we developed the interactive GRADE Living Map (GLM), a hierarchical, interactive, and continuously updated repository of official outputs. METHODS:We identified eligible publications through (i) the Cochrane training website on GRADE, (ii) an OVID-MEDLINE search, (iii) iterative review by the writing group, and (iv) feedback from the GRADE Guidance Group (G3). We designed the taxonomic hierarchy of the map branching following a combined deductive-inductive approach. All records were iteratively reviewed and classified by currency, publication type, and official standing as GRADE content. RESULTS:The GLM categorizes articles into the GRADE Repository (current material) and the GRADE Archive (for historical traceability and transparency). As of May 2026, the GLM contains 154 unique GRADE citations organized into six major thematic areas (topics): ''Principles of GRADE,'' ''Assessment of Certainty in the Evidence,'' ''Evidence-to-Decision Process,'' ''Presentation and Dissemination of Evidence,'' ''Methodological Foundations,'' and ''Application to Specific Contexts.'' Each topic further branches into multiple hierarchical levels. We also classified 88 GRADE-related articles. CONCLUSION:The GLM is a comprehensive and dynamic systematization of GRADE's and related output. By distinguishing current from outdated materials and mapping content across themes, the GLM serves as both a navigational tool and a quality safeguard. Integrated with the GRADE Book and the INGUIDE (International Guideline Training and Certification Programme) program, the GLM strengthens the accessibility of official GRADE material.
OBJECTIVES:The Grading of Recommendations Assessment, Development and Evaluation (GRADE) Working Group is developing GRADErater (GRADE Rating Automation Through Enhanced Reasoning), an official, automated tool for evaluation of the certainty of evidence (CoE). In this article, we describe the principles and methods underlying its development and state how the GRADE Working Group (GWG) will use automation to rate the CoE of intervention effects. METHODS:We followed the GRADE methods for registering a project group on GRADE and artificial intelligence (AI). The project group established that the automated tool should be developed according to the following principles: (i) compliance with the current GRADE guidance, (ii) transparency, (iii) human oversight, (iv) implementation of decision rules, (v) ease of use, (vi) understandability, and (vii) continuous improvement. We developed a set of decision rules to appraise each domain of the CoE in pairwise and network meta-analysis. These rules were developed based on official GRADE sources and validated by GRADE experts. Based on these rules, we created a first version (https://gradeai.med.up.pt/) for GRADErater. We are now evaluating this version in terms of (i) its underlying rules and (ii) its user interface. These assessments will allow for a refinement of the rules and of the first version. The modified version will be appraised and presented to the GWG with input from internal and external interest-holders before we seek formal approval. Once launched, the automated tool will be continuously evaluated and refined by incorporating feedback from end users. We will add new features (including generative AI-based functionalities), integrate this tool with GRADEpro, and develop versions in languages other than English. CONCLUSION:This project will follow a transparent methodology to create GRADErater, an official tool endorsed by the GWG that will support humans in applying the most current GRADE methods to rate the CoE.
In 2017, the GRADE (Grading of Recommendations Assessment, Development and Evaluation) working group defined the certainty of evidence as the certainty that the true effect lies on one side of a threshold or in a particular range. This definition has proved useful as the basis for rating certainty, facilitating the interpretation of the results for the target audience. However, the categorization of suggested thresholds and ranges as levels of contextualization led to inconsistencies between the initial and subsequent papers and has proved confusing for some GRADE users. Although considering context in choosing thresholds remains worthwhile, the GRADE working group will no longer use the categorization of contextualization. It will instead refer simply to chosen thresholds or ranges for determining the target of certainty rating.
BACKGROUND:COVID-19-related critical and acute illness is associated with an increased risk of venous thromboembolism (VTE). These evidence-based recommendations of the American Society of Hematology (ASH) are intended to support patients, clinicians, and other health care professionals in decisions about using anticoagulation for thromboprophylaxis for patients with COVID-19-related critical illness; patients with COVID-19-related acute illness; and those being discharged from the hospital, who do not have suspected or confirmed VTE. METHODS:ASH formed a multidisciplinary panel, including patient representatives. The Michael G. DeGroote Cochrane Canada and MacGRADE Centres at McMaster University supported guideline development, including performing systematic reviews (up to June 2023). The panel prioritized clinical questions and outcomes according to their importance for clinicians and patients. The panel used the Grading of Recommendations Assessment, Development, and Evaluation (GRADE) approach to assess certainty in the evidence and make recommendations. RESULTS:This is an executive summary of 3 updated recommendations that have been published, which concludes the living phase of the guidelines. For patients with COVID-19-related critical illness, the panel issued conditional recommendations suggesting (a) prophylactic-intensity over therapeutic-intensity anticoagulation and (b) prophylactic-intensity over intermediate-intensity anticoagulation. For patients with COVID-19-related acute illness, conditional recommendations were suggested (a) prophylactic-intensity over intermediate-intensity anticoagulation, and (b) therapeutic-intensity over prophylactic-intensity anticoagulation. The panel issued a conditional recommendation suggesting against the use of postdischarge anticoagulant thromboprophylaxis. CONCLUSIONS:These conditional recommendations were made based on low or very low certainty in the evidence, underscoring the need for additional, high-quality, randomized controlled trials for patients with COVID-19.
Over the past decade, clinical guideline development has undergone significant advancements, with an increased focus on rigour, transparency and stakeholder engagement. In this study, the first in a series of four articles, the European Society of Anaesthesiology and Intensive Care (ESAIC) outlines key principles that guide the development of high-quality guidelines within ESAIC. Starting in 2025, ESAIC will progressively adopt the GRADE approach to enhance the quality and applicability of its guidelines. Core methodological principles include the use of systematic reviews, a multidisciplinary development process, conflict-of-interest management, a transparent link between evidence and recommendations, and ensuring that recommendations are specific and actionable. This study provides an overview of the roles and processes involved in guideline production and details the transition of ESAIC from its traditional grading systems to the updated GRADE standards. Strong recommendations indicate that one option offers a clear net benefit and should be implemented with all or nearly all patients. In contrast, conditional recommendations reflect uncertainty in the balance between benefits and harms, emphasising the need for careful consideration of clinical contexts and patient values. Although specific methodological decisions may vary, the principles outlined here provide a foundation for rigorous, evidence-based guideline development within ESAIC.
Users of GRADE (Grading of Recommendations Assessment, Development and Evaluation) make judgments about the size of intervention effects on desirable and undesirable people-important health outcomes or on benefits and harms. Benchmarking effect sizes by using decision thresholds (DTs) can help to facilitate these judgments and the process. This article provides GRADE guidance for use of DTs for judgments about the magnitude of desirable and undesirable health effects, such as in a health guideline or health technology assessment. Through iterative discussions and refinement in in-person and online meetings of a GRADE project group and through e-mail communication, the authors developed guidance for using DTs in Evidence-to-Decision (EtD) frameworks. The authors applied the approach and used these examples from guidelines and the results of a randomized methodological study to develop official GRADE guidance. Several alternatives for determining and using DTs are presented. In the first main approach, outcome-specific DTs for trivial, small, moderate, and large effects are determined through a calculation using empirically derived generic coefficients and the outcome's utility value and are compared with the effect estimate obtained from an evidence synthesis. In the second main approach, outcome-specific DTs are also determined, but through direct surveying of decision makers to explicitly assign thresholds for the prioritized health outcomes. The article also describes how these approaches can be combined. The suggested approaches provide transparency for judgments in EtD frameworks that are based on findings from evidence syntheses.
ABSTRACT:Antithrombotic therapy can prevent recurrent deep vein thrombosis (DVT) and pulmonary embolism (PE). It is, however, associated with an increased risk for major bleeding. This meta-analysis systematically reviewed the evidence regarding the duration of antithrombotic therapy to assess benefits and harms. We systematically searched for randomized controlled trials (RCTs) that compared shorter (3-6 months) with longer (>6 months) courses of anticoagulation for the primary treatment of venous thromboembolism (VTE) or that compared discontinued with indefinite antithrombotic therapy for the secondary prevention of VTE. Pairs of reviewers screened the eligible trials and collected data. This study included 22 RCTs (11 617 participants). Pooled estimates showed that, for the primary treatment of unprovoked VTE, VTE provoked by chronic risk factors or transient risk factors, treating patients with a longer course (>6 months) of anticoagulation, as opposed to a shorter course (3-6 months), probably reduced recurrent PE (risk ratio [RR], 0.66; 95% confidence interval [CI], 0.42-1.02) and DVT (RR, 0.85; 95% CI, 0.63-1.14), but it was associated with increased mortality (RR, 1.43; 95% CI, 0.85-2.41) (moderate certainty) and a higher risk for major bleeding (RR, 2.02; 95% CI, 1.02-3.98; high certainty). For the secondary prevention of unprovoked VTE and VTE provoked by chronic risk factors, when compared with discontinuing treatment, indefinite anticoagulation therapy was associated with decreased mortality (RR, 0.54; 95% CI, 0.36-0.81), a reduction in recurrent PE (RR, 0.25; 95% CI, 0.16-0.41) and DVT (RR, 0.15; 95% CI, 0.10-0.21), and an increase in the risk for bleeding (RR, 1.98; 95% CI, 1.18-3.30), all supported by high certainty. Indefinite antiplatelet therapy may be associated with decreased mortality (RR, 0.95; 95% CI: 0.53-1.68; low certainty), probably a reduction in recurrent PE (RR, 0.65; 95% CI, 0.41-1.03) and DVT (RR, 0.44; 95% CI, 0.17-1.13) (moderate certainty), and may increase the risk for bleeding (RR, 1.28; 95% CI, 0.48-3.41; low certainty). In summary, for the primary treatment of all types of VTE, shorter (3-6 months) duration of anticoagulation is more beneficial. For the secondary prevention of unprovoked VTE or VTE provoked by chronic risk factors, indefinite antithrombotic treatment is more beneficial.
Users of GRADE (Grading of Recommendations Assessment, Development and Evaluation) make judgments about the size of intervention effects on desirable and undesirable people-important health outcomes or on benefits and harms. Benchmarking effect sizes by using decision thresholds (DTs) can help to facilitate these judgments and the process. This article provides GRADE guidance for use of DTs for judgments about the magnitude of desirable and undesirable health effects, such as in a health guideline or health technology assessment. Through iterative discussions and refinement in in-person and online meetings of a GRADE project group and through e-mail communication, the authors developed guidance for using DTs in Evidence-to-Decision (EtD) frameworks. The authors applied the approach and used these examples from guidelines and the results of a randomized methodological study to develop official GRADE guidance. Several alternatives for determining and using DTs are presented. In the first main approach, outcome-specific DTs for trivial, small, moderate, and large effects are determined through a calculation using empirically derived generic coefficients and the outcome's utility value and are compared with the effect estimate obtained from an evidence synthesis. In the second main approach, outcome-specific DTs are also determined, but through direct surveying of decision makers to explicitly assign thresholds for the prioritized health outcomes. The article also describes how these approaches can be combined. The suggested approaches provide transparency for judgments in EtD frameworks that are based on findings from evidence syntheses.
BACKGROUND AND OBJECTIVES:Ideally, guideline developers and health technology assessment authors base intervention decisions on randomized controlled trials (RCTs). However, relying solely on RCTs is uncommon, especially for public health interventions and harms assessment. In these situations, nonrandomized studies of interventions (NRSIs) can provide valuable information. This article presents Grading of Recommendations Assessment, Development, and Evaluation (GRADE) guidance for integrating bodies of evidence RCT and NRSI in evidence syntheses of health interventions. METHODS:Following standard GRADE methods, we developed this guidance through iterative discussions and examples with experts from the GRADE NRSI project group in multiple dedicated meetings. We presented findings of the group discussions for feedback at GRADE Working Group meetings in September 2023 and May 2024. RESULTS:The resulting GRADE guidance outlines a structured approach: (1) assessing the certainty of evidence (CoE) after defining the number of decision thresholds and the target of the certainty rating; (2) evaluating congruency of effect estimates between RCTs and NRSIs; (3) identifying which GRADE domains are affected by certainty ratings to inform complementariness between RCTs and NRSIs and the overall CoE; and (4) deciding whether and how to use one or both types of studies. CONCLUSION:This GRADE guidance offers a structured and practical approach for integrating or not integrating RCTs and NRSIs in evidence syntheses. By addressing the interplay between affected GRADE domains and assessing the congruency of effects, it helps GRADE users determine when and how NRSIs can meaningfully complement or replace RCT evidence to inform certainty ratings and decision-making.