ABSTRACT Aim To synthesise quantitative evidence on the values and preferences of people living with Motor Neurone Disease (MND), caregivers, and genetic carriers regarding health‐related outcomes to inform the Australian MND Guideline. Methods A systematic review was conducted following Cochrane and GRADE guidance, informed by an a priori protocol. Major electronic databases (including MEDLINE, Embase, CENTRAL) and trial registries were searched to identify studies that met the eligibility criteria. Risk of bias of the studies that met the eligibility criteria was assessed using the Risk of Bias in Studies of Values and Utilities (ROBVALU) tool. Data on health state utility values were synthesised using meta‐analysis where appropriate, while other quantitative data deemed inappropriate for meta‐analysis were synthesised narratively. The certainty of the evidence for each outcome was assessed using the GRADE approach. Results Twenty‐four studies ( n = 10,397) were included. Overall health‐related quality of life (hrQoL) utility values for adults with MND varied significantly based on the regional preference set utilised (mean EQ‐5D utility ranging from 0.57 in UK cohort (high certainty in the evidence) to 0.72 in Chinese cohorts (low certainty in the evidence). Utility values declined consistently with increasing disease severity across multiple staging systems, such as King's and MiToS. Narrative synthesis identified clear preferences across physical, psychosocial, and healthcare domains regarding both current and hypothetical treatment strategies. Conclusion This review provides a comprehensive synthesis of the values and preferences of the broad MND community. This review has been conducted following rigorous best‐practice methodology, to directly inform the selection and prioritisation of outcomes for the development of the Australian MND Guideline, ensuring the guideline adheres to a patient‐centred approach. Standardisation of preference elicitation methods from the MND community, and the development of a core outcome set for future MND research are key future priorities.
Evidence-based decision making in health often requires comparison of multiple options for a given condition. The GRADE (Grading of Recommendations Assessment, Development and Evaluation) evidence-to-decision (EtD) framework provides a structured approach for moving from evidence to decisions but was originally designed for pairwise comparisons. Hence, there is a need to accommodate decision making based on multiple comparisons, especially with the increasing use of systematic reviews and network meta-analyses in guideline development. Furthermore, since the original EtD framework was developed, further relevant GRADE guidance has been developed. The aim of this work was to develop a new EtD framework to accommodate multiple comparisons and reflect current GRADE guidance. The new EtD framework was revised and developed through iterative discussion, feedback, and refinement by the GRADE EtD Project Group and the GRADE Working Group. Experiences and examples from guideline developers, methodological experts, and other stakeholders informed improvements in its structure and usability for multiple comparisons and were subsequently approved by the GRADE Working Group. This article describes the new EtD framework, which now includes 2 corresponding parts for reviews of pairwise and multiple comparisons. The authors describe application to a review with multiple comparisons for the different parts of the EtD framework: the question definition, which now includes the presentation of values of health outcomes and decision thresholds; the assessment section, where the new "net effect" criterion has been included; and the conclusion section, which includes an adaptation for multiple comparisons. The article provides examples and suggestions for presentation of findings. The framework does have limitations, in that its usability has not been tested across a broad spectrum of guideline development contexts.
ObjectiveTo generate HODs for rheumatoid arthritis (RA) using a customized chatbot and evaluate their quality.MethodsWe developed a Retrieval Augmented Generation (RAG) chatbot in ChatGPT 4.0 to generate HODs for 11 RA outcomes, in standard and short formats (22 in total). The chatbot incorporated four frameworks for improving language complexity, structure, and comprehensibility, and custom, iteratively refined prompts. In a web-based survey, patients, clinicians and researchers from four international groups rated HODs across five quality attributes using Likert scale responses, with acceptable quality defined a priori as ≥70% agreement. Responses were collected in a cross-sectional web-based survey and reported following Checklist for Reporting Results of Internet E-surveys (CHERRIES) guidelines.ResultsThirty panelists completed the survey. Seven of eleven standard HODs (64%), and all eleven (100%) short HODs met the predefined acceptability threshold (≥70% agreement) across all five quality attributes. Standard HODs not meeting the acceptability threshold fell below the criterion in only one attribute, most commonly appropriateness for individuals with low health literacy, where agreement ranged from 50% to 91%. The attribute evaluating technical accuracy had an average rate of agreement of 86% for standard HODs, and 87% for short HODs. The attribute evaluating appropriateness for patients with low literacy had an average rate of agreement of 74% for standard HODs, and 89% for short HODs.ConclusionsHODs generated using a customized GPT4.0 chatbot were rated highly by experts for quality. This structured AI-assisted approach can facilitate the generation of HODs that are acceptable to clinicians and researchers, but still require human oversight and proof-reading.
ABSTRACT Introduction Artificial intelligence (AI) may support several processes of the health guideline enterprise. This article describes the development of an extension of the Guidelines International Journal (GIN)‐McMaster Guideline Development Checklist (GDC) for integrating AI in the guideline enterprise. This development has been led by the GIN‐AI Working Group. Methods We started by prompting a large language model (LLM) for items related to the use of AI in each of the steps of the original GDC. Subsequently, the members of the working group engaged in a set of iterative discussions, resulting in item refinement and in a consensus first version of the extension. We retrospectively applied this first version to a case use of guidelines incorporating AI in their development (Allergic Rhinitis and its Impact on Asthma [ARIA] 2024‐2025 guidelines), leading to further refinement and to the approval of the final extension tool. Results Prompting LLMs resulted in the generation of 149 items. Of those, 117 were removed and 19 were modified by members of the working group. On the other hand, 17 new items were added during the iterative discussion process. The retrospective application of the extension led to changes in the wording of four items. The final version of the checklist extension has been approved with 49 items modifying or adding to the original GDC. Discussion We have developed an extension of the GIN‐McMaster GDC that encompasses a set of conduct standards that are intended to facilitate the comprehensive and transparent integration of AI in the health guideline enterprise. Clinical Trial Registration Not applicable. This study is not a clinical trial.
Background and Objectives Health guidelines play a central role in informing clinical practice, public health measures and health policy. But their trustworthiness may be undermined by factors such as insufficient methodological rigor, lack of transparency, conflicts of interest, and inconsistent application of established standards. Existing appraisal tools address selected aspects of guideline quality but do not comprehensively assess the trustworthiness of individual recommendations, nor do they adequately reflect recent advances in guideline methodology, including living guidelines, Grading of Recommendations, Assessment, Development, and Evaluation, adaptation, and the use of artificial intelligence (AI). This study aims to develop and validate Transparent, Rigorous, Useable, Standardized, and Trustworthy Guide (TRUSTGUIDES), a globally applicable, flexible set of tools to assess the trustworthiness of health guideline recommendations. We define trustworthiness as distinct from methodological quality: it encompasses not only rigorous methods but also transparency, independence, and applicability, which together determine whether a recommendation merits user confidence. Methods TRUSTGUIDES will be developed through a multistep, mixed-methods process. First, a scoping review and expert consultation will identify existing guideline appraisal tools and inform domains and items generation. Using deductive and inductive approaches, domains and items will be generated and may be refined through focus groups and selected through iterative Delphi surveys involving an international, multidisciplinary working group. TRUSTGUIDES will be validated by assessing internal consistency, inter-rater reliability, content validity, and construct validity, including comparisons with established instruments such as the Grading of Recommendations, Assessment, Development, and Evaluation certainty domains, AGREE II, and PANELVIEW. Psychometric properties will be examined using factor analysis and, as necessary, item response theory models. AI will be integrated both as an object of assessment and as methodological support for tool application, with large language models evaluated against a human reference standard. Conclusion TRUSTGUIDES will be designed to evaluate the trustworthiness of individual guideline recommendations across key factors, including transparency and credibility, and to address relevant domains such as the certainty of evidence, strength of recommendations, conflicts of interest, applicability, adaptability, currency, certification, and the appropriate use of AI. TRUSTGUIDES addresses critical gaps in current guideline appraisal by offering a comprehensive, recommendation-level assessment of trustworthiness aligned with the World Health Organization guideline standard methodology. By integrating AI, our tools will support efficient, transparent, and future-ready guideline evaluation within an evolving health evidence ecosystem.
Introduction: Guideline development is both time-consuming and resource-intensive, and duplication of efforts contribute to research waste. Although adopting existing guidelines is an efficient alternative, it often fails to account for contextual factors. Adapting trustworthy, previously developed relevant guidelines offers a promising alternative. Thus, we aimed to identify, describe and evaluate formal frameworks for the adaptation of health-related guidelines.Methods: We updated a previously published systematic survey of guideline adaptation frameworks by searching MEDLINE (Ovid) and Embase (Ovid) from January 1st, 2015, to October 31st, 2025. Eligible studies described an adaptation framework for health guidelines in sufficient detail to allow reproducibility. We excluded reviews, guideline adaptations, and implementation reports. We extracted in duplicate and independently data on framework characteristics and presented results in both narrative and tabular formats.Results: Our search identified 20 guideline adaptation frameworks, up from eight identified in the previous survey. These frameworks included five to 24 steps (median: 13). Each framework followed one of four primary approaches: ADAPTE-based (6/20); GRADE-ADOLOPMENT-based (3/20); hybrid approaches (2/20); and non-systematic approaches (9/20). ADAPTE-based frameworks follow a comprehensive three-phase, 24-step process to assess and modify entire source guidelines. In contrast, GRADE-ADOLOPMENT-based frameworks focus on recommendation-level contextualization using evidence-to-decision frameworks. The remaining frameworks primarily utilize non-systematic methods. The most frequently reported challenges were high resource and expertise requirements, dependence on the quality and availability of source guidelines, contextual and implementation barriers, and limited capacity for new evidence synthesis.Discussion: Guideline adaptation frameworks vary substantially in scope and methods, reflecting a balance between methodological rigor and feasibility. ADAPTE- and GRADE- ADOLOPMENT-based approaches provide structured methods, while more recent frameworks prioritize flexibility. This synthesis will help guideline developers select approaches that best fit their context and resources.
The COVID-19 pandemic highlighted the crucial role of evidence-based practice guidelines (EBPGs) in healthcare systems. Reliable and timely guidelines are essential in health emergencies. This study aims to identify and comprehensively understand the determinants that influence the development and implementation of EBPGs during health emergencies from the perspective of guideline developers and implementers. In this article, we describe the study protocol. We will conduct an exploratory sequential mixed-methods study composed of four phases: (I) a qualitative descriptive study to document the experiences of guideline developers and implementers and use this data to generate a list of determinants that influence the development and implementation of EBPG; (II) exploratory integration for survey development and validation using the data of phase I; (III) cross-sectional study to administer the survey to a large sample and estimate the frequency and perceived impact of those determinants; (IV) interpretation of the results: a combination of qualitative and quantitative findings through the narrative description and joint displays to generate a comprehensive understanding of those determinants. We will explore if results vary across subgroups of interest (participants' roles, gender, and type of organization). We will include people worldwide from five groups participating in an EBPG development and implementation process during the pandemic. This study will provide insights into the determinants that influence the development and implementation of EBPGs during health emergencies. Its potential impact includes improving future health emergency preparedness, addressing equity gaps, and enhancing guideline development and implementation.
OBJECTIVES:The Grading of Recommendations Assessment, Development and Evaluation (GRADE) Working Group is developing GRADErater (GRADE Rating Automation Through Enhanced Reasoning), an official, automated tool for evaluation of the certainty of evidence (CoE). In this article, we describe the principles and methods underlying its development and state how the GRADE Working Group (GWG) will use automation to rate the CoE of intervention effects. METHODS:We followed the GRADE methods for registering a project group on GRADE and artificial intelligence (AI). The project group established that the automated tool should be developed according to the following principles: (i) compliance with the current GRADE guidance, (ii) transparency, (iii) human oversight, (iv) implementation of decision rules, (v) ease of use, (vi) understandability, and (vii) continuous improvement. We developed a set of decision rules to appraise each domain of the CoE in pairwise and network meta-analysis. These rules were developed based on official GRADE sources and validated by GRADE experts. Based on these rules, we created a first version (https://gradeai.med.up.pt/) for GRADErater. We are now evaluating this version in terms of (i) its underlying rules and (ii) its user interface. These assessments will allow for a refinement of the rules and of the first version. The modified version will be appraised and presented to the GWG with input from internal and external interest-holders before we seek formal approval. Once launched, the automated tool will be continuously evaluated and refined by incorporating feedback from end users. We will add new features (including generative AI-based functionalities), integrate this tool with GRADEpro, and develop versions in languages other than English. CONCLUSION:This project will follow a transparent methodology to create GRADErater, an official tool endorsed by the GWG that will support humans in applying the most current GRADE methods to rate the CoE.
BACKGROUND AND OBJECTIVE:Post-COVID-19 condition (PCC) is a complication of acute COVID-19, which often presents with a variety of symptoms. It can also impact individuals' overall well-being, capacity to carry out daily activities, engage in physical exercise, maintain employment, and their general quality of life. Therefore, this review aimed to examine the role of patient-reported outcome measures (PROMs) questionnaires in assessing patients with PCC through prevalence of abnormal tests. METHODS:We searched three databases. Two reviewers independently screened articles using Laser Al and extracted relevant data using a piloted Google Sheets. We performed a meta-analysis using OpenMeta and RevManWeb and conducted a subgroup analysis based on the setting of the patients during their acute COVID-19 infection. We assessed the risk of bias using a modified ROBINS-I tool and the certainty using the Grading of Recommendations Assessment, Development, and Evaluation approach. RESULTS:No studies reported diagnostic test accuracy measures for questionnaires in PCC. However, 23 comparative studies reported on the prevalence of abnormal questionnaire results in patients with PCC. Outcomes showed that patients with PCC have higher abnormal results than controls, regardless of their setting during acute COVID-19 infection. The overall certainty of the evidence was low due to the high risk of bias, indirectness, and imprecision. CONCLUSION:This review sheds light on the importance of testing PROMs in patients with PCC using these questionnaires and the need for further testing their validity in this condition.
The delivery of healthcare in the out‐of‐hospital setting by paramedics, emergency medical technicians, and first responders is informed by various health guidelines that cover a myriad of medical presentations. These guidelines often score poorly on quality assessment tools and do not routinely meet accepted standards for guideline development. To address this, we propose the development of an extension to the Guideline International Network (GIN)‐McMaster Guideline Development Checklist (GDC) to assist developers of health guidelines on out‐of‐hospital care. The aim of this project are to: (i) identify the current methodologies, approaches and processes used by developers of guidelines on out‐of‐hospital care and determine how they align with the original GIN‐McMaster GDC; (ii) understand the overarching barriers and enablers of guideline development in this setting; (iii) generate a fit‐for‐purpose extension through an iterative consensus‐based approach; and (iv) disseminate this document to target users. This protocol outlines the proposed development of an extension to the GIN‐McMaster GDC, designed specifically for developers of health guidelines on out‐of‐hospital care. The development of this extension will be informed by understanding current practices and the determinants that influence guideline development in this setting. This will involve reviewing all items within the original checklist to determine if modifications and/or new items are required.
Background: Post-COVID-19 condition (PCC) is a complication following acute COVID-19 infection, which may lead to long-term cardiac abnormalities. This review aimed to assess the prevalence of structural/functional deviations in echocardiography in individuals with PCC compared to patients without PCC. Methods: We searched three databases. Two reviewers independently screened articles using LASER Al and extracted relevant data using a piloted Excel sheet. We performed meta-analysis using OpenMeta and RevManWeb and a subgroup analysis based on patients’ settings during acute COVID-19. We assessed the risk of bias using the Hoy et al. tool and the certainty using the Grading of Recommendations Assessment, Development, and Evaluation (GRADE) approach. Results: We included 16 studies that reported on differences in echocardiographic findings in patients with or without PCC. Individuals with PCC were more likely to have structural/functional deviations in echocardiographic readings of unclear clinical significance, particularly those who were hospitalized during acute COVID-19. The overall certainty of the evidence was very low due to the high risk of bias, indirectness, and imprecision. Conclusions: This review provides insight into the use of echocardiograms and the frequency of test deviations in individuals with PCC. Despite existing evidence, there is a need for future studies to assess the diagnostic test accuracy of echocardiograms in PCC.
BACKGROUND AND OBJECTIVES:Determining the types of contributions to guideline development, as well as acknowledging these contributions groups, are critical steps in the guideline development process. The objective of this study was to describe types of contributions to guideline development and authorship policies of guideline-producing organizations as described in their guidance documents on guideline development. METHODS:We conducted a descriptive summary of guidance documents on guideline development. Using multiple sources, we initially compiled a list of guideline-producing organizations and then searched for their publicly available guidance documents on guideline development (eg, guideline handbooks). Authors abstracted data in duplicate and independently on the organizations' characteristics, types of contributions to guideline development, and authorship policies. RESULTS:We identified 133 guideline-producing organizations with publicly available guidance documents, of which the majority were professional associations (59%) from the clinical field (84%). Types of contributions to guideline development described by the organizations could be categorized as related to: management; content expertise; technical expertise; or dissemination, implementation, and quality measures. Commonly reported specific contributions included panel membership (99%), executive (83%), evidence synthesis (86%), and peer review (92%). A minority of organizations mentioned entities specifically dedicated to conflict-of-interest management (20%) and to dissemination, implementation, and quality measures (24%). For most organizations, panelists were involved in either supporting or conducting the evidence synthesis (73%). Sixty percent of organizations mentioned that panels should be multidisciplinary, and 44% mentioned that they should be balanced according to at least one characteristic (eg, geographical region) (44%). A minority of organizations had a guideline authorship policy (38%). Out of those, a majority specified types of contributions eligible for authorship (76%), a minority specified criteria for exclusion from authorship (18%), and rules for authorship order (27%). CONCLUSION:Guidance documents of guideline-developing organizations consistently describe four types of contributions (panel membership, executive, evidence synthesis, and peer review), while others are less commonly described. They also lack important details on authorship policies.
BACKGROUND:COVID-19-related critical and acute illness is associated with an increased risk of venous thromboembolism (VTE). These evidence-based recommendations of the American Society of Hematology (ASH) are intended to support patients, clinicians, and other health care professionals in decisions about using anticoagulation for thromboprophylaxis for patients with COVID-19-related critical illness; patients with COVID-19-related acute illness; and those being discharged from the hospital, who do not have suspected or confirmed VTE. METHODS:ASH formed a multidisciplinary panel, including patient representatives. The Michael G. DeGroote Cochrane Canada and MacGRADE Centres at McMaster University supported guideline development, including performing systematic reviews (up to June 2023). The panel prioritized clinical questions and outcomes according to their importance for clinicians and patients. The panel used the Grading of Recommendations Assessment, Development, and Evaluation (GRADE) approach to assess certainty in the evidence and make recommendations. RESULTS:This is an executive summary of 3 updated recommendations that have been published, which concludes the living phase of the guidelines. For patients with COVID-19-related critical illness, the panel issued conditional recommendations suggesting (a) prophylactic-intensity over therapeutic-intensity anticoagulation and (b) prophylactic-intensity over intermediate-intensity anticoagulation. For patients with COVID-19-related acute illness, conditional recommendations were suggested (a) prophylactic-intensity over intermediate-intensity anticoagulation, and (b) therapeutic-intensity over prophylactic-intensity anticoagulation. The panel issued a conditional recommendation suggesting against the use of postdischarge anticoagulant thromboprophylaxis. CONCLUSIONS:These conditional recommendations were made based on low or very low certainty in the evidence, underscoring the need for additional, high-quality, randomized controlled trials for patients with COVID-19.
Users of GRADE (Grading of Recommendations Assessment, Development and Evaluation) make judgments about the size of intervention effects on desirable and undesirable people-important health outcomes or on benefits and harms. Benchmarking effect sizes by using decision thresholds (DTs) can help to facilitate these judgments and the process. This article provides GRADE guidance for use of DTs for judgments about the magnitude of desirable and undesirable health effects, such as in a health guideline or health technology assessment. Through iterative discussions and refinement in in-person and online meetings of a GRADE project group and through e-mail communication, the authors developed guidance for using DTs in Evidence-to-Decision (EtD) frameworks. The authors applied the approach and used these examples from guidelines and the results of a randomized methodological study to develop official GRADE guidance. Several alternatives for determining and using DTs are presented. In the first main approach, outcome-specific DTs for trivial, small, moderate, and large effects are determined through a calculation using empirically derived generic coefficients and the outcome's utility value and are compared with the effect estimate obtained from an evidence synthesis. In the second main approach, outcome-specific DTs are also determined, but through direct surveying of decision makers to explicitly assign thresholds for the prioritized health outcomes. The article also describes how these approaches can be combined. The suggested approaches provide transparency for judgments in EtD frameworks that are based on findings from evidence syntheses.
ABSTRACT:Antithrombotic therapy can prevent recurrent deep vein thrombosis (DVT) and pulmonary embolism (PE). It is, however, associated with an increased risk for major bleeding. This meta-analysis systematically reviewed the evidence regarding the duration of antithrombotic therapy to assess benefits and harms. We systematically searched for randomized controlled trials (RCTs) that compared shorter (3-6 months) with longer (>6 months) courses of anticoagulation for the primary treatment of venous thromboembolism (VTE) or that compared discontinued with indefinite antithrombotic therapy for the secondary prevention of VTE. Pairs of reviewers screened the eligible trials and collected data. This study included 22 RCTs (11 617 participants). Pooled estimates showed that, for the primary treatment of unprovoked VTE, VTE provoked by chronic risk factors or transient risk factors, treating patients with a longer course (>6 months) of anticoagulation, as opposed to a shorter course (3-6 months), probably reduced recurrent PE (risk ratio [RR], 0.66; 95% confidence interval [CI], 0.42-1.02) and DVT (RR, 0.85; 95% CI, 0.63-1.14), but it was associated with increased mortality (RR, 1.43; 95% CI, 0.85-2.41) (moderate certainty) and a higher risk for major bleeding (RR, 2.02; 95% CI, 1.02-3.98; high certainty). For the secondary prevention of unprovoked VTE and VTE provoked by chronic risk factors, when compared with discontinuing treatment, indefinite anticoagulation therapy was associated with decreased mortality (RR, 0.54; 95% CI, 0.36-0.81), a reduction in recurrent PE (RR, 0.25; 95% CI, 0.16-0.41) and DVT (RR, 0.15; 95% CI, 0.10-0.21), and an increase in the risk for bleeding (RR, 1.98; 95% CI, 1.18-3.30), all supported by high certainty. Indefinite antiplatelet therapy may be associated with decreased mortality (RR, 0.95; 95% CI: 0.53-1.68; low certainty), probably a reduction in recurrent PE (RR, 0.65; 95% CI, 0.41-1.03) and DVT (RR, 0.44; 95% CI, 0.17-1.13) (moderate certainty), and may increase the risk for bleeding (RR, 1.28; 95% CI, 0.48-3.41; low certainty). In summary, for the primary treatment of all types of VTE, shorter (3-6 months) duration of anticoagulation is more beneficial. For the secondary prevention of unprovoked VTE or VTE provoked by chronic risk factors, indefinite antithrombotic treatment is more beneficial.
To systematically review the values and preferences of people with lived experience of motor neurone disease (MND), including those living with MND, caregivers and genetic carriers, regarding their health-related outcomes. MND is a devastating neurodegenerative disease that significantly impacts those living with the disease, their caregivers, and their families. Understanding the values and preferences of those affected by MND is crucial for providing patient-centred care and developing trustworthy guidelines. Studies will be included if they report on the values and preferences of adults living with MND; caregivers and families of those diagnosed with MND; clinicians as a proxy for adults living with MND; or asymptomatic genetic carriers of MND. Peer-reviewed studies utilising either quantitative or qualitative methodologies (or both) will be eligible for inclusion in this review. A comprehensive search of electronic databases will be conducted to identify relevant studies. Data extraction, risk of bias assessment (of quantitative studies), and assessment of methodological limitations (of qualitative studies) will be performed independently by two reviewers. Quantitative data will be pooled using meta-analysis where appropriate, and qualitative data will be synthesised following a modified meta-aggregative approach. The relevant GRADE approach will be used to assess the certainty of evidence. CRD420250653287.
Users of GRADE (Grading of Recommendations Assessment, Development and Evaluation) make judgments about the size of intervention effects on desirable and undesirable people-important health outcomes or on benefits and harms. Benchmarking effect sizes by using decision thresholds (DTs) can help to facilitate these judgments and the process. This article provides GRADE guidance for use of DTs for judgments about the magnitude of desirable and undesirable health effects, such as in a health guideline or health technology assessment. Through iterative discussions and refinement in in-person and online meetings of a GRADE project group and through e-mail communication, the authors developed guidance for using DTs in Evidence-to-Decision (EtD) frameworks. The authors applied the approach and used these examples from guidelines and the results of a randomized methodological study to develop official GRADE guidance. Several alternatives for determining and using DTs are presented. In the first main approach, outcome-specific DTs for trivial, small, moderate, and large effects are determined through a calculation using empirically derived generic coefficients and the outcome's utility value and are compared with the effect estimate obtained from an evidence synthesis. In the second main approach, outcome-specific DTs are also determined, but through direct surveying of decision makers to explicitly assign thresholds for the prioritized health outcomes. The article also describes how these approaches can be combined. The suggested approaches provide transparency for judgments in EtD frameworks that are based on findings from evidence syntheses.
BACKGROUND:A decision threshold (DT) reflects the point at which a decision or judgment changes, leading to the selection of an action or a commitment for one of several alternatives. Thresholds have always played a role in decision-making. Very small effects may achieve statistical significance yet remain not important to patients or the public. Judgments shift, for instance, from "no or trivial effect" to "small, moderate, and large benefit" with direct implications for decision-making. However, in guideline panels and other clinical or policy decisions, these thresholds are often applied subconsciously when interpreting effect estimates from studies and likely to vary across panel members. STUDY DESIGN AND SETTING:In this commentary, inspired by the concepts leading to recent publications by the Grading of Recommendations Assessment, Development and Evaluation (GRADE) working group and its members, we argue that the use of DTs has many advantages. RESULTS:DTs are the basis for an interpretation of results that is not centered on "statistical significance." In addition, DTs are useful for other aspects of evidence synthesis. The certainty of evidence ratings using the GRADE approach (https://book.gradepro.org/) are centered on DTs, including the determination of the target of the certainty rating, with advantages for transparency, objectivity, and simplicity. For example, judging imprecision is informed by DTs. Specifically, the number of DTs crossed by the plausible effect sizes, as indicated by its confidence interval, helps determine the degree of uncertainty assigned in a GRADE assessment of imprecision, including the number of levels of certainty a user rates down. DTs have also altered the way how users can transparently integrate bodies of evidence from both nonrandomized and randomized studies. Once determined, DTs can be used to validate automated judgments about the certainty of evidence. Beyond these developments, DTs can be useful for designing primary research. For example, sample size calculations could use standardized DTs for large effects when there are known harms that the intended benefits need to outweigh. CONCLUSIONS:DTs have many roles in interpretation, certainty assessments and research planning and design. PLAIN LANGUAGE SUMMARY:Decision thresholds are the points where a decision changes-for example, when evidence shifts our judgment from "moderate benefit" to "large benefit." Unlike statistical significance, decision thresholds focus on what matters to citizens and decision-makers. In health guidelines, these thresholds often influence judgments unconsciously, but making them explicit improves transparency and consistency. The GRADE approach uses decision thresholds to judge how certain we are about evidence, helping to make these judgments clearer and more objective. They can also guide research design, such as calculating sample sizes. Although identifying thresholds takes some effort, it ultimately makes evidence assessment and decision-making more efficient and structured.