OBJECTIVES:This study aims to address the status of children's and caregivers' participation in the development of paediatric core outcome sets (COS). METHODS:We included all paediatric COS from a previous systematic review and searched the Core Outcome Measures in Effectiveness Trials database to 26 February 2024 for recent paediatric COS. We used descriptive and thematic analysis methods to present the characteristics of the included COS and to describe children's and caregivers' participation in the development, including any facilitators and barriers. We assessed the degree of participation of children and caregivers in two steps: by rating whether their views were considered in forming the outcome list (yes/no) and then whether their views were integrated in determining the most important outcomes (fully integrated/partially integrated/not integrated). RESULTS:A total of 114 paediatric COS were included. 60 (53%) COS involved children and caregivers in the development process. 29 (48%) of the 60 COS considered children's and caregivers' views in forming the initial outcome list, which was most often conducted by interview (n=12 of 29, 41%). Regarding determining the most important outcomes, 35 (58%) of the 60 COS fully integrated children's and caregivers' views, and the most common method was the Delphi survey with consensus meeting (n=29 of 35, 83%); the youngest child participants were aged 7 years. The most frequently mentioned facilitator of children's and caregivers' participation was the engagement of patient groups or organisations. CONCLUSION AND RELEVANCE:We evaluated the degree of children's and caregivers' participation in the development of COS and found that strategies to promote children's and caregivers' participation should be constructed.
BACKGROUND:Although guidelines focus predominantly on individual diseases, some dedicated guidelines and recommendations exist for common combinations of comorbidities. Using diabetes, coronary heart disease (CHD) and stroke as an example, we aimed to assess the current status and quality of such guidelines. METHODS:We systematically searched literature databases, Google, guideline platforms, and websites of relevant organizations. We included guidelines published between 2020 and 2024 that addressed at least two of the following diseases: diabetes, CHD, and stroke; and included drug therapy interventions. We extracted the recommendations on drug therapy and sources of the supporting evidence, and assessed the quality of the guidelines using AGREE II and RIGHT with the help of a large language model based tool. RESULTS:We identified 82 guidelines: 64 focused on one disease with recommendations on comorbidities, and 18 specifically addressed a combination of diseases. China was the most frequent country of origin (n = 50, 61.0%). The methodological and reporting quality of these guidelines was moderate on average. For most guidelines (n = 54, 65.9%), the primary focus was on diabetes. We grouped the recommended drug therapies for patients with combinations of these diseases into four main categories: anti-diabetic therapy, antihypertensive therapy, lipid-lowering therapy, and anticoagulant therapy. Most recommendations were supported by RCTs, but only a third of the guidelines referred to studies done in multimorbid patients. CONCLUSION:Most recommendations related to multimorbidity were found in guidelines focusing on a single target disease, and supporting evidence from multimorbid patients was rare. More primary evidence from multimorbid patients is needed.
Background:Large language models (LLMs) have attracted increasing attention in medical research and clinical practice and have been applied to processes related to evidence-based medicine (EBM). However, the extent of their integration with evidence-based Chinese medicine (CM) remains unclear. Methods:We systematically searched PubMed, Web of Science, China National Knowledge Infrastructure (CNKI), and Wanfang Data from 30 November 2022 to 31 January 2026, with supplementary searches conducted in Google Scholar. Studies were included if they applied LLMs to EBM processes within a CM context or investigated LLMs in CM using established evidence-based research designs. Descriptive analysis summarized study characteristics, and findings were mapped according to the evidence ecosystem framework. Results:A total of 12 studies published between 2023 and 2025 were included. Most studies integrated LLMs into different stages of the EBM workflow within a CM context. At the evidence generation stage, studies explored the role of LLMs in identifying research priorities. At the evidence synthesis stage, LLM performance was evaluated in literature screening, data extraction, and risk-of-bias assessment. At the evidence translation stage, studies evaluated the performance of LLMs in guideline-related question answering and recommendation generation. At the evidence implementation stage, LLMs combined with knowledge graphs or retrieval-augmented generation were used to develop intelligent question-answering systems based on CM guidelines or standards. Conclusion:Existing studies suggest that LLMs have begun to be explored across multiple stages of evidence-based CM research and show potential for improving evidence synthesis efficiency and supporting knowledge translation and application. Protocol registration:Open Science Framework (https://osf.io/ztbd5/overview).
BACKGROUND:The management of multimorbidity often requires applying multiple disease-specific guidelines, which may lead to harmful disease and drug interactions. The objective of the study was to collect the opinions of leading experts on factors influencing the use of guidelines in decision-making in multimorbidity management from a multidisciplinary perspective, and provide empirical insights to inform guideline-based decision support. METHODS:We conducted a series of semi-structured interviews with experts in multimorbidity research and practice and guideline development to participate. Interviewees were selected using purposive sampling. Our interviews covered the current use of clinical guidelines in multimorbidity management, key barriers and facilitators, and priority action areas. Thematic analysis was performed using NVivo 12.0. RESULTS:Fifteen experts from nine countries were interviewed. The lack of multimorbidity-specific guidelines and recommendations may lead healthcare providers to either rely solely on personal experience, or follow guidelines without considering the needs and comorbidities of individual patients. Major barriers to guideline application include insufficient evidence and limited awareness and training among clinicians in both multimorbidity management and guideline implementation. The interviewees emphasized flexible use of guidelines, prioritizing patient values and preferences, and individualized management goals informed by (rather than strictly adhering to) guidelines. CONCLUSIONS:Applying guidelines in multimorbidity care faces major evidence and applicability gaps. Particularly the generation of high-quality clinical evidence specific to multimorbidity, enhancing the methodological rigor in guideline development, and improvement of the contextual relevance and usability of guidelines in real-world settings need more attention.
OBJECTIVES:The application of large language models (LLMs) to systematic review tasks is rapidly expanding, yet the transparency and methodological rigor of these evaluations remain unclear. We aimed to assess reporting transparency, methodological quality, and how authors frame claims and caveats in studies applying LLMs to systematic review tasks. STUDY DESIGN AND SETTING:We conducted a cross-sectional meta-epidemiological study by searching PubMed, Embase, Web of Science Core Collection, IEEE Xplore, and five other databases from inception to December 1, 2025, for peer-reviewed articles and preprints. We included empirical studies evaluating generative transformer-based LLMs (eg, Gemini) for core systematic review tasks (eg, screening, data extraction) against a reference standard. We assessed reporting transparency using an adapted Chatbot Assessment Reporting Tool and methodological quality using an adapted Quality Assessment of Diagnostic Accuracy Studies 2 tool. We also analyzed the frequency and strength of claims and caveats mentioned by the authors. The study is registered with the Open Science Framework (https://osf.io/8edhb). RESULTS:We identified and included 229 studies comprising 440 empirical tasks. Reporting transparency was moderate, with a mean item score of 0.52 (standard deviation 0.30) on a 0-1 scale, where higher values indicate more complete reporting. We observed substantial gaps in reproducibility-essential domains, including protocol information (mean score 0.12) and model details (0.30). Although 60.6% of assessments were rated as having a low risk, key safeguards against overfitting and data leakage were rarely reported; for example, locking the test set before prompt optimization, a basic protection against information leakage, was not reported in 99.8% of tasks. We identified 837 claims and 693 caveats. Authors framed claims weakly more often than strongly (66.8% vs 33.2%). Performance superiority over a comparator was the most common claim (64.8% of tasks). Readiness for practical use was claimed in 47.5% of tasks, almost always in qualified terms (93.3%). CONCLUSION:Studies applying LLMs to systematic review tasks are reported with moderate transparency but often omit reproducibility-critical details necessary to assess leakage and overfitting. Although authors frequently make claims about performance and practice readiness, these are typically expressed cautiously. Improved reporting standards and clearer safeguards are urgently needed before routine use of LLMs in evidence synthesis can be recommended. PLAIN LANGUAGE SUMMARY:LLMs, such as ChatGPT, are increasingly used to help carry out parts of systematic reviews, which summarize evidence to inform healthcare decisions. We examined 229 studies that tested LLMs on tasks such as screening articles, extracting data, and assessing study quality, covering 440 evaluations in total. On average, these studies reported their methods with moderate clarity, but often omitted information needed to repeat the work or judge whether the results were trustworthy. Notably, almost none described safeguards to ensure that the test data had not already influenced how the model was set up, a key step for avoiding overly optimistic results. Authors frequently described LLMs as performing well and nearly ready for practical use, though usually in cautious terms. Clearer reporting standards and stronger safeguards are needed before LLMs can be routinely relied upon in evidence synthesis.
BACKGROUND:Global population aging exacerbates the challenges of multimorbidity and polypharmacy in older adults. Clinical practice guidelines are essential for addressing these issues. This systematic review aims to evaluate the quality of existing guidelines and synthesize their recommendations based on the Ariadne principles, to inform future guideline development and clinical practice. METHODS:We searched nine databases and five guideline repositories (e.g., PubMed, Web of Science, Cochrane Library, CNKI, WHO) up to August 2025. Guidelines and consensus documents focusing on multimorbidity or polypharmacy in older adults, published in English or Chinese, were included. Each guideline was evaluated using four validated tools: AGREE II (methodological quality), RIGHT (reporting completeness), AGREE-REX (recommendation credibility and applicability), and GLIA (implementation feasibility). Recommendations were categorized and synthesized according to the Ariadne principles, with independent screening and data extraction and consensus resolution of discrepancies. RESULTS:The multidimensional appraisal of the 21 included guidelines revealed consistent weaknesses. According to AGREE II, the domains of Scope and Purpose (81.9 %) and Clarity of Presentation (61.1 %) demonstrated the highest median scores, whereas Rigor of Development (16.7 %) and Applicability (8.3 %) scored the lowest. Based on the RIGHT checklist, overall reporting completeness was 43.2 %, with the Evidence (0.0 %) and Quality Assurance (0.0 %) domains being particularly underreported. AGREE-REX evaluation indicated limited implementability at the individual recommendation level (12.5 %), and GLIA, while suggesting moderate implementability at the guideline level (65.4 %), identified frequent barriers in the domains of Measurable Outcomes (100.0 %) and Innovation Requirements (66.7 %). Thematically, most guidelines addressed interaction assessment (n = 15, 71.4 %), but far fewer incorporated patient preferences (n = 9, 42.9 %) or monitoring strategies (n = 9, 42.9 %). Only three guidelines (14.3 %) fully adhered to all five steps of Ariadne principles. CONCLUSION:Current guidelines for older adults with multimorbidity or polypharmacy exhibit substantial weaknesses in methodological rigor, reporting completeness, and implementation feasibility. Synthesis based on the Ariadne principles revealed an imbalanced pattern of recommendations, with a predominant focus on medication safety rather than patient-centered and longitudinal care management. Future guideline development should strengthen methodological processes, systematically integrate patient perspectives, and co-design practical implementation strategies to better support personalized care for an aging population.
Abstract Background Despite the growing global burden of multimorbidity, the patterns of disease combinations, have not been extensively categorized. We aimed to explore the predictors, health consequences, and patterns of discordant and concordant multimorbidity. Methods We used the 2018 China Health and Retirement Longitudinal Study (CHARLS), a representative database of adults aged > 45 years from China. We conducted logistic regression analyses to assess the likelihood of having discordant (conditions from different disease systems) versus concordant (only cardiometabolic, or only respiratory diseases) multimorbidity, and to compare the health status and healthcare utilization between patients with discordant and concordant multimorbidity. Latent class analysis (LCA) was applied to both the entire sample and to patients with discordant multimorbidity to identify clusters of disease combinations. Results The sample included 1668 patients with concordant (mainly cardiometabolic), and 7306 patients with discordant, multimorbidity. Female patients, patients living in rural settings, former and current smokers, and patients engaging in high-intensity physical activity, were more likely to have discordant instead of concordant multimorbidity. Depression, limitations in daily activities, poor self-reported health, and frequent healthcare use were more common in patients with discordant than concordant multimorbidity. The LCA identified five clusters when all multimorbid patients were included (cardiometabolic, arthritis-digestive, respiratory, multisystem, and arthritis-hypertension classes), and four clusters when restricted to discordant multimorbidity (digestive, arthritis-cardiometabolic, respiratory, and multisystem classes). Conclusion Discordant multimorbidity is associated with poorer health and increased use of healthcare. Cardiometabolic diseases, arthritis, and digestive diseases have a central role in defining disease patterns.
Background: As part of the development of practice guidelines, developers need to systematically engage relevant interest-holder groups. The Reporting Items for practice Guidelines in HealThcare (RIGHT) checklist for reporting practice guidelines lacks detailed guidance on how to report the engagement process. Objective: To develop a standardized checklist for reporting interest-holder engagement in practice guidelines, named the RIGHT-MuSE checklist. Design: The development process followed the methods recommended by the Enhancing the QUAlity and Transparency Of health Research (EQUATOR) Network, as well as lessons learned from the development of the RIGHT statement and its extensions. Key steps included developing the protocol, registering the project, establishing a working group, conducting background work, generating an initial list of items, conducting a consensus survey, holding panel discussions, and creating the final RIGHT-MuSE checklist and an accompanying explanation and elaboration document. Setting: International collaboration. Participants: 25 panelists from various guideline development interest-holder groups. Measurements: Consensus agreement on checklist items. Results: The final RIGHT-MuSE checklist consists of 11 items covering the guidance used, methods of engagement, characteristics of interest-holders, evaluation of engagement, and management of conflicts of interest. The checklist is supplemented with a glossary of key terms and detailed explanations for each item to facilitate its use. Limitation: The RIGHT-MuSE checklist has not been evaluated in a broader context or widely applied in real-world guideline development processes. Conclusion: Guideline developers can use the RIGHT-MuSE checklist to comprehensively report on interest-holder engagement. Primary Funding Source: The Vincent and Lily Woo Foundation.
Background: Childhood pneumonia remains a leading driver of global mortality, yet its diagnosis in primary care is plagued by the absence of consistent, universally applied diagnostic criteria across facilities. This systematic review evaluates the diagnostic performance of clinical features for childhood pneumonia in primary care, stratified by country income level and age. Methods: We searched MEDLINE, Embase, Web of Science, CNKI, and Wanfang Data from inception to December 31, 2025. Prospective cohort studies assessing clinical features' diagnostic accuracy against chest radiography were included. Methodological quality was assessed using the Quality Assessment of Diagnostic Accuracy Studies (QUADAS-2) tool. Random-effects meta-analyses were performed to pool sensitivity, specificity, positive likelihood ratio (PLR), negative likelihood ratio (NLR), and area under the curve (AUC). The study was registered with PROSPERO (CRD420251179854). Findings: Fifty studies (n=48,138) met inclusion criteria, with 22 conducted in low- and middle-income countries (LMICs) and 27 in high-income countries (HICs). In LMICs, WHO age-specific tachypnea exhibited a weak association with pneumonia (PLR 2.00, 95% CI 1.10–3.70); however, fixed respiratory rate thresholds demonstrated stronger associations [PLR 2.5 (95% CI 1.6–4.0) for >40 breaths/min, 2.4 (95% CI 1.5–3.8) for >50 breaths/min, and 2.2–28.22 for >60 breaths/min], as did chest indrawing (PLR 2.5, 95% CI 1.3–4.9). In HICs, oxygen saturation ≤94% was the strongest predictor (PLR 3.31, 95% CI 2.45–4.47). The WHO pneumonia criteria demonstrated suboptimal performance (AUC 0.62, 95% CI 0.57–0.66). Multivariable models incorporating biomarkers improved discrimination (AUC 0.79–0.83) but frequently lacked external validation. Interpretation: The diagnostic accuracy of individual clinical features for childhood pneumonia is limited overall. Context-specific indicators (e.g., elevated respiratory rates and chest indrawing in LMICs, hypoxemia in HICs) and externally validated multivariable prediction tools are warranted to refine diagnostic guidelines and mitigate antibiotic overuse.
Background:Pediatric rare diseases often cause a prolonged diagnostic odyssey. AI, including machine learning, deep learning, large language models (LLMs), and multimodal systems, may support diagnosis, but these applications in children have not been systematically mapped. Objective:The aim of the study is to map diagnostic applications, data modalities, validation strategies, and evidence maturity of AI methods for pediatric rare diseases. Methods:We conducted a scoping review following Joanna Briggs Institute methodology and reported it according to PRISMA-ScR (Preferred Reporting Items for Systematic Reviews and Meta-Analyses extension for Scoping Reviews). On June 26, 2026, we searched PubMed, Scopus, Web of Science Core Collection, Embase, China National Knowledge Infrastructure (CNKI), Wanfang Data, and the Cochrane Library for records published from January 1, 2015, through June 1, 2026. We additionally searched medRxiv and arXiv and hand-searched the reference lists of included studies and relevant reviews. Eligibility was defined using the population-concept-context framework: pediatric rare diseases, diagnostic AI, and any clinical or research setting. JZ and JL independently screened titles and abstracts and assessed potentially eligible full-text reports. JZ charted the data, and JL verified every field. Findings were synthesized descriptively according to disease focus, AI technology, input modality, diagnostic task, validation strategy, and evidence maturity. Results:Database searches identified 2557 records; 2063 remained after deduplication. Of 106 full-text reports assessed, 77 database studies and 4 studies from hand searching and preprint servers were included, yielding 81 studies. Studies were published from 2016 through 2026, with 55 of 81 (67.9%) published from 2024 through 2026. Using a mutually exclusive primary technology classification, classical machine learning accounted for 38 (46.9%) studies, facial AI for 18 (22.2%), deep learning for 15 (18.5%), LLMs for 6 (7.4%), and multimodal AI for 4 (4.9%). Electronic health records, claims, clinical text, or structured clinical vignettes were used in 25 (30.9%) studies, facial images in 17 (21%), and other medical imaging in 13 (16%). Evidence remained mainly retrospective and internally validated: 62 (76.5%) studies included a retrospective component and 76 (93.8%) reported internal validation, whereas 21 (25.9%) included external validation and 12 (14.8%) included a prospective component. Conclusions:Research on AI-assisted diagnosis of pediatric rare diseases has expanded rapidly, but evidence maturity has not kept pace. Most studies established technical feasibility rather than generalizable clinical benefit, and performance should be interpreted by task, inputs, reference standard, and validation design rather than used to rank technologies. Evidence for LLMs and multimodal AI remains limited. Future research should prioritize multicenter validation, reproducible task-specific benchmarks, prospective evaluation, and assessment of incremental clinical value.
Introduction The rapid expansion of systematic reviews (SRs) has led to increasing concerns about research duplication and redundancy. While several tools exist to assess the quality and risk of bias of SRs, no standardized instrument specifically evaluates the degree of duplication between SRs on similar topics. This protocol describes the development and validation of the systematic review duplication (SRD) tool, which is designed to systematically assess and quantify duplication between intervention-based SRs in healthcare. Methods The development process follows established guidelines for creating reporting and quality assessment tools and comprises three phases: (Ⅰ) preparatory work, including team formation and tool conceptualization; (Ⅱ) tool development, involving initial item generation through a systematic literature review and analysis of existing tools, item validation through pilot testing with 40 SRs across 17 disease categories, modified Delphi surveys, and consensus meetings; and (Ⅲ) dissemination through academic channels and professional networks. The SRD tool will evaluate duplication across four domains: research topic, research methods, research results, and research quality. Expert consensus will be achieved through a two-round modified Delphi process with an agreement threshold of at least 70%. Inter-rater reliability will be assessed using the intraclass correlation coefficient and Kendall’s W coefficient. Expected outcomes: The SRD tool is expected to provide a standardized, user-friendly instrument for distinguishing necessary replication from redundant duplication in SRs. It will be available as both web-based and Excel-based applications. Discussion The SRD tool is intended to fill a critical gap in evidence synthesis methodology by enabling the systematic assessment of duplication and supporting informed decision-making by researchers, journal editors, guideline developers, and clinicians.
Background: With the increasing variety of pediatric diagnostic procedures, a growing number of children require sedation for diagnostic examinations. Appropriate sedation protocols guarantee the safety of children during sedation and improve its efficiency. Currently, there are significant differences in pediatric sedation practices across countries, regions, and medical institutions. Methods: The Pediatric Anesthesiology Group of the Chinese Society of Anesthesiology organized experts from multiple countries to develop the "Chinese Society of Pediatric Anesthesiology Guideline for Pediatric sedation (2025)" based on evidence-based medicine, considering the safety and efficacy, the preferences of children and parents, and drug accessibility. Results: The guideline provided 28 recommendations addressing 12 clinical issues related to pre-sedation assessment, safety measures for sedation personnel and facilities, sedation protocols, sedation monitoring, and post-sedation recovery. Conclusions: The guideline is planned to be disseminated globally through multiple channels, aiming to standardize the management of pediatric sedation and improve its safety and efficacy.
BackgroundEvidence briefs for policy (EBPs) are effective tools for delivering research evidence to policymakers and other stakeholders by highlighting high-priority issues, outlining options and considering implementation strategies. However, policymakers' demands for evidence and policy-relevant information across different fields have led to variability in the terminology used to describe EBPs, and the methodological quality of these EBPs remains unclear. This study aims to (1) identify organizations whose definitions of EBPs contain the three key components of problem, options and implementation considerations, (2) assess the methodological quality of EBPs that incorporate these three key components and (3) identify existing evaluation/assessment tools of EBPs.MethodsA two-stage documentary analysis approach was used. First, we identified documents that were produced by organizations/institutions to inform policymakers and that contained the three key components (Problem, Options and Implementation considerations). Second, the methodological quality of the documents was assessed from the perspectives of the evidence supply side (that is, evidence synthesis) and the evidence demand side (that is, mapping of and engagement between both policymakers and stakeholders).ResultsIn 22 organizations, the term policy brief was the most commonly used, accounting for 50% of organizations, while other terms varied. Issue briefs were used by three organizations (13.6%) and evidence briefs were used by two organizations (9.1%). In total, 50 individual documents from nine different organizations were included to evaluate components and methodology. (1) From the supply-side perspective: 17 (34%) documents described the search resources, 10 (20%) documents described evidence certainty and 15 (30%) assessed the methodological quality of the research evidence. (2) From the demand-side perspective: 30 (60%) documents were developed in response to demand-side needs, while 27 (54%) included both stakeholder mapping and engagement.ConclusionsMethodological shortcomings were identified in the EBPs from both the supply-side and demand-side perspectives, highlighting the need to validate and better implement existing tools and to complement existing guidelines.
Pediatric rare diseases are highly heterogeneous and are frequently associated with missed or delayed diagnosis, creating substantial burden for patients, families, and clinicians. Although artificial intelligence (AI), including large language model–enabled approaches, has shown potential for diagnostic support, translation into real-world pediatric care remains limited. A key gap is the mismatch between metric-centric evidence reporting and clinician-defined implementation needs. To address this, we integrated published evidence with clinician perspectives to derive an implementation-oriented evidence-to-requirements framework for AI-assisted pediatric rare-disease diagnosis. We used a convergent multimethod design with two complementary evidence sources. First, we conducted a PRISMA-ScR scoping review of four databases (PubMed, Embase, Web of Science, and Scopus) from inception to December 2025 and included 28 original studies on AI-assisted pediatric rare-disease diagnosis. Second, we conducted semi-structured interviews with 21 pediatric clinicians from 15 departments at a tertiary children’s hospital in Chongqing, China, and analyzed transcripts using inductive thematic analysis. We then integrated findings side-by-side to identify convergences, divergences, and translational gaps. The scoping review showed rapid movement toward multimodal and LLM-enabled approaches across several diagnostic task types, including screening or cohort identification, phenotyping, differential diagnostic support, and variant or gene prioritization. Translation-oriented evidence remained uneven, with limited prospective evaluation and inconsistent reporting of fairness, safety, and deployment context. Interview analysis identified four recurrent themes: diagnosis as time-pressured puzzle-solving; AI as a cognitive extender rather than replacement; trust dependent on traceable evidence and transparent reasoning; and demand for structured, actionable outputs with low workflow burden. Integrated analysis revealed a persistent implementation gap between metric-centric publication practices and clinician-defined requirements for real-world adoption. This scoping review and qualitative interview study does not establish clinical effectiveness of AI-assisted diagnosis. Instead, it identifies implementation requirements that may guide future development and evaluation, including representative multicenter data, prospective validation, evidence traceability, actionability, safety, fairness, and workflow fit. Main limitations include restriction to English-language studies, reliance on umbrella rare-disease terminology, possible missed studies among unscreened records after ASReview-assisted screening, and a single-institution clinician interview sample.
Background:Perimenopausal Syndrome (PS) results from estrogen fluctuations due to ovarian dysfunction, leading to autonomic and neuropsychological symptoms. Existing therapeutic methods have certain limitations, and integrated Chinese-Western medicine (ICWM) treatment can achieve complementary advantages. Methods:We defined 16 clinical questions and outcomes for PS and searched CNKI, PubMed, the Cochrane Library, and other databases and relevant websites from inception to August 29, 2025. Relevant clinical practice guidelines/consensus statements (CPGs/CSs), systematic reviews (SRs), and randomized controlled trials (RCTs) were included, and their methodological quality was apprised using AGREE II, AMSTAR 2, and ROBUST-RCT, respectively. Data synthesis and visualization were performed using Microsoft Excel 2021 and R. Results:280 studies (5 CPGs/CSs, 38 SRs, 237 RCTs) were included, of which 84.6% were Chinese publications. Acupuncture and oral Chinese herbal medicine (CHM) had abundant evidence, some Chinese Medicine interventions had little or no evidence, and study quality varied across types. Evidence summaries showed acupuncture versus conventional Western therapy and integrated oral CHM and western medication (WM) versus WM alone were superior or equivalent in efficacy and safety for PS. Conclusions:This study mainly assesses the available evidence on acupuncture and integrated oral CHM and WM interventions for PS. Both treatment strategies demonstrate potential benefits for PS. Nevertheless, inconsistencies in the evidence base across different interventions, methodological shortcomings, and limited international applicability weaken the overall quality of current findings. Future research should target these gaps to advance evidence-based, personalized ICWM strategies, improving outcomes for women with PS worldwide.
Evidence-based medicine (EBM), formalized in the 1990s, has redefined clinical practice by advocating the integration of research evidence, clinical expertise, and patient values. This paradigm has introduced methodological rigor through randomized controlled trials (RCTs) to establish causality, systematic reviews to synthesize findings, and the GRADE approach to evaluate evidence based on risk of bias, inconsistency, indirectness, imprecision, and publication bias. These advancements have shaped clinical guidelines, reduced practice variability, and influenced medical education toward evidence-based inquiry. Despite its contributions, EBM faces challenges in the evolving landscape of modern medicine. The lengthy process of evidence generation, often requiring years for trials and guideline updates, limits responsiveness to emerging health needs, as observed during the COVID-19 pandemic. The external validity of RCT results is constrained by strict inclusion criteria, posing difficulties in applying findings to diverse patient populations with comorbidities. Additionally, the siloed nature of evidence complicates comprehensive care for multifactorial conditions, while the annual influx of over one million medical publications overwhelms traditional synthesis methods. Artificial intelligence (AI) presents a promising avenue to address these issues, leveraging capabilities in processing heterogeneous data. Natural language processing may enhance literature analysis, machine learning could identify patterns in complex datasets, and causal inference might improve the reliability of observational data insights. These technologies hold potential to accelerate evidence development and tailor it to individual needs. This paper proposes digital intelligent evidence-based medicine (i-EBM) as a conceptual evolution of EBM, designed for the AI era. i-EBM envisions a three-layered framework. The data foundation layer aims to integrate structured evidence from RCTs, domain knowledge such as biomedical ontologies and traditional Chinese medicine principles, and multi- modal patient data, including electronic health records, genomics, and wearable device outputs. Knowledge graphs are proposed to link these elements into a unified, computable knowledge network. The intelligent processing layer seeks to apply AI for evidence retrieval, data extraction, quality assessment, and synthesis, potentially using large language models to assist these processes. The knowledge service layer intends to provide dynamic guidelines and individualized predictions, supported by ongoing human-machine collaboration to ensure clinical relevance and ethical considerations. i-EBM has the potential to mitigate EBM's limitations by facilitating real-time evidence updates, reducing knowledge fragmentation through integrated data, and offering personalized decision support. For instance, it may support precision medicine by connecting diverse data sources, with applications possibly extending to fields like oncology or traditional Chinese medicine. Future research could explore autonomous AI systems, optimized clinical workflows, and governance frameworks to address data privacy, bias, and global standardization. In conclusion, i-EBM offers a theoretical framework to extend EBM principles, harnessing AI's potential alongside human expertise to advance medical research and practice. Meanwhile, for issues such as the quantitative study of the complex intervention characteristics and syndrome differentiation patterns of traditional Chinese medicine, i-EBM can provide methodological support in data integration, pattern recognition, and causal inference, offering potential tools and insights for uncovering the intrinsic regularities of TCM evidence and optimizing its evaluative framework.
BACKGROUND:Conflict of interest (COI) management is critical for ensuring the scientific integrity and fairness of clinical practice guidelines (CPGs). Large language models (LLMs) have great potential in strengthening COI management, particularly in information collection, assessment, and supporting guideline development groups. OBJECTIVE:To explore LLMs' role in COI management during CPG development, focusing on applications, challenges, and future directions. METHODS:We examined how LLMs can support COI management by designing and testing a set of simulated COI scenarios based on established management principles. RESULTS:LLMs can improve efficiency in data collection (e.g., in analyzing disclosures), objectivity in risk assessment, and transparency in reporting. However, privacy risks (e.g., data breaches) and technical issues (e.g., model bias) hinder the adoption of LLM based approaches. Setting up policy frameworks, research collaboration, and enhanced security, such as differential privacy levels, can enhance reliability. CONCLUSION:LLMs can support COI management in CPG development if ethical issues are adequately considered, but validation in real-world settings is still needed.