
Introduction:Global guidance provides evidence-informed recommendations on total fat intake to prevent unhealthy weight gain in adults and children, however, evidence regarding total fat intake and other health outcomes, particularly non-communicable diseases (NCDs), also exists. This scoping review aimed to map and identify gaps in this body of evidence. Methods:Four databases were searched from 2014 to 2024 for published systematic reviews (SRs), in any population, assessing the effects or associations of reduced total fat intake (e.g., ≤ 30% total energy [TE]), compared with usual fat intake (e.g., > 30%TE), on pre-specified health outcomes other than measures of unhealthy weight gain or acute post-prandial outcomes. We screened independently and in duplicate. Data were extracted, checked, and narratively summarized inductively by health topic. Results:Sixty-seven SRs were included. Eight were fully aligned (all included studies addressed the scoping review question) and reported on health topics related to NCDs, for example, cancers (gastrointestinal, hematological), bone and joint health, gastrointestinal and neurological disorders, and cardiometabolic outcomes. Fifty-nine SRs partially addressed the scoping review question, reporting additional topics, for example, cancers (breast, gynecological, urological, liver, skin), and respiratory and liver disorders. Almost one third of included SRs (n = 21) reported outcomes related to cardiometabolic risk and cardiovascular disease, obtained mostly from randomized controlled trial (RCT) evidence. Evidence for other topics was mostly observational. Overall, evidence was predominantly from high-income countries, and in adults. Conclusions:This scoping review maps available evidence on the relationships between total fat intake and health outcomes other than unhealthy weight gain and acute postprandial effects and demonstrated a heterogenous body of evidence, with several evidence gaps, predominantly relating to NCDs.
Introduction:The assessment of the certainty of evidence using GRADE requires evaluating inconsistency and imprecision. These domains are difficult to judge in a fully independent way. We performed a meta-research study to propose an approach to disentangle inconsistency and imprecision based on decision thresholds. Methods:We evaluated systematic reviews published in the Cochrane Database of Systematic Reviews. We included all meta-analyses published in these reviews (i) with at least five primary studies, (ii) which did not correspond to subgroup analyses, and (iii) which had provided their results as continuous outcomes or as dichotomous outcomes with previously established decision thresholds. We recalculated each included meta-analysis using the fixed- and the random-effects models, retrieving metrics potentially expressing inconsistency or imprecision. This includes inconsistency indices, variables related to the confidence intervals (CI) of the meta-analytical estimate and prediction intervals. Based on these variables, we performed factor analysis, hierarchical cluster analysis, and decision tree analysis. Results:We assessed 8703 meta-analyses. The factor analysis resulted in the identification of three factors: (i) one factor associated with inconsistency metrics, (ii) one factor associated with imprecision metrics, and (iii) one factor of metrics not clearly related to inconsistency or imprecision. Taken together, the analysis suggests an approach based on the number of decision thresholds crossed by the CI of the fixed- and random-effects meta-analytical estimates: the number of thresholds crossed by the CI of the fixed-effect model informs about imprecision, while the difference in the number of thresholds crossed by the CIs of the random- versus the fixed-effect models reflects inconsistency. Prediction intervals and inconsistency indices, such as the I 2, can help detect cases in which inconsistency may be underestimated. Conclusion:Our results may help those assessing the certainty in evidence to independently rate inconsistency and imprecision.
Introduction:Climate change has major health impacts to which researchers and policymakers need to respond. The volume and multidisciplinary nature of climate-health evidence pose challenges to its comprehensive identification. Search filters are currently not available for this topic. Our aim was to empirically develop sensitivity-maximizing climate-health search filters for two interfaces of the MEDLINE database. Methods:Climate change impacts health via several exposure pathways involving distinct mechanisms: extreme weather events, heat stress, air quality, water quality and quantity, food supply and safety, vector distribution and ecology, and social factors. Using the relative recall method, we established a gold standard by extracting included studies of 110 evidence syntheses and classifying them into exposure pathways. From this set, we empirically derived and validated filters per pathway. Results:Based on 1572 primary studies indexed in PubMed, we developed filters with a sensitivity of 95%, 97% and 99% for six of the seven major climate-health exposure pathways. Conclusion:We designed the first empirically derived search filters for climate-health pathways, enabling sensitive retrieval despite acknowledged limitations. Our filters can be applied to different types of research questions at the intersection of climate and health.
ABSTRACT Background The need for social care, in the form of practical assistance and personal care, is increasing alongside a growing older population with long‐term conditions and places those with limited or no access to publicly funded care at risk of increasing levels of unmet need. Aim This systematic review sought to understand the role of social and health conditions in the need, demand, utilization, and expenditure on social care in the United Kingdom. Methods We searched Medline, CINAHL, EconLit, ASSIA, and the Campbell Collaboration from January 1, 2009 to April 14, 2025, and gray literature on August 21, 2024, for randomized trials, cohort studies, case control studies, and interrupted time series, cross‐sectional, and modeling studies of populations aged ≥ 60 years old, in a UK setting, that examined the association of social and health conditions with the need and demand for, use of and expenditure on social care. We conducted risk of bias assessments using ROBINS‐E for longitudinal studies, and intended to use Risk of Bias 2 for randomized trials. We applied GRADE to assess the certainty of evidence. Results were synthesized narratively, and presented in an Evidence Gap Map. Results We included 10 longitudinal cohort and eight cross sectional studies, study sample size ranged from 526 to > 400,000 participants. Five of the cohort studies were assessed as high risk of bias, and the cross sectional studies as moderate concern or high risk of bias. Older age was associated with increased unmet need for care in most studies, and living alone with increased unmet need for care and use of residential care. Results for ethnicity, deprivation and sex varied across studies, with studies using different measures of deprivation that limited comparability of findings. A range of long‐term health conditions were assessed, some studies indicated an association with an increased need and use of care; there were mixed findings for cost. For populations with cognitive impairment or dementia, the association with age and the use and cost of care varied, there was no clear association for sex, ethnicity and living alone; a previous hospital admission or ongoing health problems that included severity of dementia were associated with increased use of care. Conclusion Our review of UK evidence indicates that there is insufficient evidence to accurately identify populations age > 60 years that are at increasing risk for needing care due to their social and health conditions.
Introduction:Rapid reviews aim to deliver timely evidence for decision-makers when full systematic reviews are not possible or practical. Efficient selection of studies is challenging when questions are complex, or the evidence base is diffuse. Methods:In a rapid review on trial informativeness, our team used EPPI-Reviewer, a web-based systematic review platform that supports document management, screening, and machine learning prioritization, to conduct title and abstract screening. We developed a machine learning classifier model within the platform to rank records by predicted relevance based on coding structures aligned with predefined criteria. Results:The classifier model correctly concentrated relevant studies in the higher probability bands, which allowed most eligible records to be identified early. As screening progressed to lower probability bands, the number of newly identified records declined, indicating effective prioritization. Real-time collaboration and a clear audit trail supported consistent decision-making across reviewers. Limitations included the initial effort to train the model and potential subscription costs. Conclusion:Classifier assisted screening in EPPI-Reviewer improved the feasibility of conducting a rapid review on a complex topic within a limited timeframe. Although the risk of missed citations remains, this is inherent to any review method. With appropriate training and support, classifier models and platforms like EPPI-Reviewer can enhance both efficiency and transparency in rapid evidence synthesis.
ABSTRACT Introduction The Preferred Reporting Items for Systematic reviews and Meta‐Analyses (PRISMA) is a reporting guideline aiming to facilitate transparent and accurate reporting in systematic reviews. The latest version, PRISMA 2020, specifies that it is intended for guiding the reporting, not the conduct, of systematic reviews. We examined how often systematic review authors reported having used PRISMA 2020 outside its scope. Methods We randomly sampled 200 systematic reviews citing PRISMA 2020 and extracted all sentences describing how PRISMA 2020 was used. Two authors independently assessed the extracted sentences and categorised the use of PRISMA 2020 as one of the following options: adequate (i.e., as a reporting guideline), inconclusive (e.g., unclear or ambiguous description) or inadequate (e.g., to guide review conduct). Results The systematic reviews were published between 2021 and 2024, and 88.0% were published in biomedical journals. More than half (55.5%) of the systematic reviews described using PRISMA 2020 inadequately, typically as a guide for review conduct; 24.0% were categorised as ‘inconclusive’, and only 20.5% were categorised as ‘adequate’. Conclusion More than half of the systematic reviews described using PRISMA 2020 outside its intended scope. We encourage systematic review authors to use PRISMA 2020 as a reporting guideline. Other resources are available to guide the conduct of systematic reviews, e.g., the Cochrane Handbook for Systematic Reviews of Interventions.
This tutorial describes how to explore across-trial heterogeneity in meta-analysis using trial-level characteristics. It describes how to carry out a subgroup analysis for a trial-level characteristic with more than two categories and provides approaches for analysing continuous trial-level characteristics. The tutorial introduces meta-regression, outlines its intended applications, and explains how to interpret meta-regression results. Common pitfalls when undertaking trial-level subgroup analysis and meta-regression are also discussed. A complementary micro-learning module is provided to support additional learning of these concepts.
ABSTRACT This tutorial introduces the concept of trial‐level subgroup analysis in meta‐analysis. We describe the rationale for conducting subgroup analyses, distinguish trial‐level from participant‐level characteristics, outline how to properly interpret effects within subgroups, and illustrate how to perform a trial‐level subgroup analysis, including formal statistical testing for differences between subgroups. A complementary micro‐learning module is provided to support additional learning of these concepts.
ABSTRACT Background and Objective Systematic reviews (SRs) and meta‐analyses (MAs) are central to evidence‐based practice (EBP). Despite their growing volume in low‐ and middle‐income countries, especially those produced by nursing‐affiliated institutions in North Africa (NA), no empirical meta‐research has systematically assessed their methodological quality, reporting practices, and research integrity indicators. This study investigated methodological quality, reporting practices, and scientific integrity of SRs and MAs conducted in this context over the past decade. Methods We conducted a meta‐research synthesis of SRs and MAs in nursing sciences published between 2015 and 2025 using the PubMed database. Eligible reviews used explicit SR/MA methods, had a first or last author affiliated with a NA institution, and were identified through structured searches. Methodological and reporting quality were assessed using PRISMA and MOOSE guidelines, complemented by an exploratory 10‐point composite quality score. Associations between review characteristics and quality indicators were examined. Results We included 38 SRs/MAs, of which 15 included quantitative synthesis. Only 50% reported PROSPERO registration, with several registrations being retrospective or inaccurate. While eligibility criteria, database searches, and quality appraisal were commonly reported, dual data extraction and gray‐literature searches were less common. Validated appraisal tools were inconsistently applied, and reporting guidelines were sometimes inappropriately used. The composite score indicated overall modest methodological quality; with PROSPERO registration was the only factor associated with higher quality scores (p = 0.008). Heterogeneity was usually reported, whereas sensitivity analyses and publication‐bias assessments were inconsistently conducted. Conclusions This study highlights important methodological and integrity‐related deficiencies in SRs and MAs in NA. Strengthening methodological expertise in this central field of clinical practice is critical.
ABSTRACT Objectives Open science practices, including the preregistration of protocols and the sharing of data and code, are increasingly promoted to enhance transparency, accessibility and the reuse of research outputs. However, systematic reviews, which are essential for informing clinical guidance and policy, often provide limited access to the materials needed for independent verification and reuse. Our objective was to estimate the prevalence of key open science indicators in interventional systematic reviews and to assess temporal changes by comparing our 2024 cohort with previously published cohorts from 2014 to 2020. Design Cross‐sectional meta‐research study. Data Sources MEDLINE, Embase, and CINAHL were searched in June 2024 for English‐language records indexed between 1 January and 1 February 2024. Eligibility Criteria Interventional systematic reviews including human participants, a clearly defined PICO question, a systematic search strategy, a risk‐of‐bias assessment, and at least one meta‐analysis. Methods Following our publicly registered protocol and STROBE guidance, records were randomly sampled and screened in duplicate until 300 eligible reviews were included. Data extraction was conducted in duplicate. We assessed protocol registration and availability, data sharing, code sharing, and preprint dissemination. Preprints were identified through extensive manual searches across more than 40 preprint servers and through author contact. Temporal changes were examined by comparing our 2024 cohort with previously published cohorts from 2014 to 2020 using pairwise risk ratios (RRs) with predefined equivalence ranges. Results Among the 300 included reviews, cancer was the most common health condition (17%), and pharmacological or biological interventions were the most common intervention type (35%). The median number of primary studies included in each systematic review was 14 (IQR 9–23). Most reviews (73%) were published in journals with data and/or code‐sharing policies. In the 2024 cohort, 70% of systematic reviews reported a registered protocol or protocol availability and 61% included a data availability statement. However, only 24% shared underlying data, and analytical code was shared in 1% of reviews. Preprints were rare, occurring in 2% of the reviews. Reporting of protocol registration or availability increased substantially over time (2024 vs. 2014: RR 2.03, 95% CI 1.55–2.65; 2024 vs. 2020: RR 1.67, 95% CI 1.43–1.95), as did data availability statements (2024 vs. 2020: RR 1.98, 95% CI 1.63–2.40). Data sharing showed a nonlinear pattern, decreasing from 2014 to 2020 (RR 0.22, 95% CI 0.13–0.36) and increasing from 2020 to 2024 (RR 3.71, 95% CI 2.30–6.00), but remained uncommon overall. Code sharing was rare and showed no detectable change (OR 1.00, 95% CI 0.15–6.51). In the 2024 cohort, 3% of protocols were registered after manuscript submission. Conclusions Several transparency indicators became more common over time, although trends differed across indicators. Yet the sharing of reusable research materials remains uncommon. Reusable data, analytic code, and preprints are rare outside the context of pandemics. These findings highlight a persistent gap between declarative openness and operational accessibility, emphasizing the necessity for more robust, verifiable mechanisms from journals and funders to reduce data distance and support independent verification and reuse of systematic review evidence. This gap between declared openness and practical accessibility continues to limit independent verification and reuse of systematic review evidence.
ABSTRACT Background Publication bias impacts the direction and magnitude of summary effect estimates in systematic reviews and the confidence in the review findings. It remains unclear to what degree publication bias is considered in the certainty assessments in systematic reviews. Our research aims are to generate an overview of and estimate how frequent rating down occurs for the GRADE domain “publication bias” across Cochrane reviews. Study Design and Setting This is a descriptive study. We included Cochrane reviews published between January 2021 and October 2023 that included at least one outcome for which the certainty of evidence was rated down for publication bias. We collected data from the reviews' full text, analyses, and figures. We analyzed data with descriptive statistical analysis. We used a deductive qualitative approach for analyzing qualitative data. Results We identified 81 reviews. These reviews included between 1 and 328 individual studies. All but three reviews included randomized‐controlled trials only. Data from less than 10 studies supported the summary of findings of 106 outcomes. The most commonly used methods for detecting publication bias were funnel plots. 220 outcomes were rated down by one level for publication bias. Four outcomes were rated down by two levels. The most commonly reported reasons were indications of publication bias in funnel plots. Conclusions Rating down for publication bias was rare in Cochrane reviews despite the high prevalence of publication bias in original research. Guidance might not suffice to inform assessments of publication bias in reviews with less than 10 studies.
ABSTRACT This tutorial provides an overview of the steps involved in conducting rapid reviews of interventions based on recent methodological guidance and empirical studies from the Cochrane Rapid Reviews Methods Group (RRMG). This tutorial provides practical guidance on methods, key decisions, and strategies to optimize the balance between timeliness and rigor and is intended as a practical introduction for researchers, clinicians, and policymakers interested in conducting or using rapid reviews.
ABSTRACT Background The Cochrane Handbook's I2 categorization system (0%–25% “low”, 25%–75% “moderate”, ≥ 50% “high” heterogeneity) defines the standard approach to interpreting heterogeneity in meta‐analysis and informs thousands of systematic reviews each year. Despite its widespread use, its logical coherence and its role as an analytical decision tool have received limited formal examination. Objective To examine whether the Cochrane Handbook's I2 categorization system includes overlapping category definitions that create ambiguous classification, and to propose context‐specific frameworks that preserve I2 as a decision tool for heterogeneity exploration. Methods We conducted a structured genealogical analysis of key methodological sources related to the Cochrane Handbook's I2 categorization. We examined whether individual I2 values can satisfy more than one category. We analysed the categorization using principles from formal logic, philosophy of science, and statistical theory. We traced the development of I2 interpretation from its original formulation to current Cochrane guidance. We developed context‐specific frameworks based on patient‐important outcome categories. Results The I2 categorization system includes overlapping definitions in which identical values satisfy more than one category (e.g., I2 = 50% corresponds to both “moderate” [25%–75%] and “high” [≥ 50%]). This structure assigns single values to multiple categories and departs from principles that require consistent and mutually exclusive classification. The primary literature provides limited explicit theoretical or empirical justification for the selected thresholds. The current approach uses I2 as an interpretive endpoint rather than as a decision tool for heterogeneity exploration. These features reduce interpretive clarity and obscure the role of clinical context in heterogeneity assessment. Conclusions The current I2 categorization system introduces ambiguity and leads to inconsistent analytical decisions across outcome contexts. Evidence synthesis requires context‐specific frameworks in which interpretation reflects outcome type and expected variability. We propose the PIOHA framework to align heterogeneity assessment with clinical relevance while preserving I2 as an analytical decision tool. Clinical Relevance Systematic reviews inform clinical guidelines and patient care. Context‐specific heterogeneity assessment supports analytical decisions that reflect the clinical importance and expected variability of patient‐important outcomes.
ABSTRACT Introduction Carrying out a systematic review (SR) of the literature entails a high workload and encompasses a variety of very different tasks. The emergence of artificial intelligence tools has brought further opportunities to improve the efficiency and reliability of SRs. SR processes can be optimised to the extent that integration and interoperability of software tools across production stages are progressively implemented. A key stage is data extraction, which can be challenging due to the large amounts of data items to consider and the variability of studies reporting styles, which heavily complicates data processing and analyses. We report the development of a software platform that integrates processes across all types of SR tasks, including overviews of SRs, is open source, and addresses the challenges of data extraction through the standardisation of data structures: the “Open‐Source SYstematic Reviews Integrated System” (OSSYRIS). Methods We established a series of criteria to select the software integrated in OSSYRIS: few applications, covering all SRs production processes, inter‐operable and open source. After several trials, we selected Zotero as reference manager, KoboToolbox XLSForms for screening and data extraction and R for analyses and reporting. We integrated all components using Application Programming Interfaces (API) in R. For the data extraction form, we identified content items from our own experience and from the Cochrane handbook. OSSYRIS has been piloted and used in several SRs and overviews carried out by the authors. Results In OSSYRIS, references are manually imported in Zotero and are integrated into XLSForms in KoboToolbox, which are used for online screening by reviewers. R automatically downloads the screening results from KoboToolbox and updates the status of the references in Zotero as ‘irrelevant’, ‘included’, ‘excluded,’ and ‘unclear’. R automatically produces the figure with the PRISMA flow of studies and references lists by status, for reporting. Data extraction is manually done using another XLSForm structured in sections: study characteristics, participants, intervention or exposure, outcomes, results and conclusion. Data extraction is standardised by using pre‐coded data items, filtering data items according to relevance criteria and modularising data structures. Results of studies are entered using a data structure consistent with the information on the type of outcomes, in a form preceding section. Items that require a decision based on certain criteria, such as which is the type of study or the risk of bias assessments, are not filled in by reviewers; rather reviewers enter the criteria and OSSYRIS internal algorithms issue the specific type of study design or the risk of bias assessments, based on those criteria. XLSForms provide additional functionalities to ensure data integrity. R automatically produces the characteristics of included studies and other analytical outputs for reporting. Standardisation and modularity facilitate adapting the form for different types of SR. Conclusions OSSYRIS provides an open source, integrated system to carry out SRs. Our work may support the promotion of open source and free tools to conduct SRs bringing together a community of practice to further improve it, within Cochrane and beyond.
ABSTRACT Introduction Social determinants of health (SDOH), which include economic stability, education access and quality, community and social context, healthcare access and quality, and neighborhood and built environment, are known to be related to variation in health outcomes across a variety of health conditions. Ocular neoplasms are among the health conditions that have been shown to be affected by SDOH, and retinoblastoma is the most prevalent form in children. Our objective was to systematically review documented associations between SDOH and retinoblastoma in the US. Methods We followed a pre‐established protocol. We searched Medline, Embase, and Web of Science from 2000 to November 2023 using a mix of keywords and controlled vocabulary. We included primary studies of any design that evaluated one or more relationships between SDOH and retinoblastoma outcomes, such as survival, mortality, disease staging, types of care, and incidence. We extracted data on study design, population characteristics, SDOH domains and indicators, and estimates of associations with retinoblastoma outcomes. We assessed risk of bias using a modified Newcastle‐Ottawa Scale and performed narrative syntheses. Results We included 26 studies that had reported a total of 552 associations. Social and community context, notably race and ethnicity, was the most commonly examined domain, suggesting evidence of worse survival outcomes for Black and Hispanic patients. The most clinically and policy‐relevant patterns were concentrated in social and community context, economic stability, and healthcare access and quality SDOH domains, where measures such as race and ethnicity, poverty, unemployment, and insurance status were repeatedly associated with stage at diagnosis, enucleation, and survival. Conclusion SDOH were frequently associated with retinoblastoma diagnosis, treatment, and survival in US studies. The most clinically relevant findings relate to insurance‐related access barriers, socioeconomic disadvantage, and social and community factors, while environmental and occupational exposures were more informative for disease cause and prevention. Our findings underscore the need for more consistent SDOH measurement and adjustment methods to improve study comparability and to inform targeted interventions to address retinoblastoma health outcomes.
ABSTRACT Background Manual abstract screening is a primary bottleneck in evidence synthesis. Emerging evidence suggests that large language models (LLMs) can automate this task, but their performance when processing multiple references simultaneously in “batches” is uncertain. Objectives To evaluate the classification performance of four state‐of‐the‐art LLMs (Gemini 2.5 Pro, Gemini 2.5 Flash, GPT‐5, and GPT‐5 mini) in predicting reference eligibility across a wide range of batch sizes for a systematic review of randomized controlled trials. Methods We used a gold‐standard dataset of 790 references (93 considered relevant) from a published Cochrane Review on stem cell treatment for acute myocardial infarction. Using the public APIs for each model, batches of 1 to 790 references were submitted to classify each as “Include” or “Exclude.” Performance was assessed using sensitivity and specificity, with internal validation conducted through 10 repeated runs for each model‐batch combination. Results Gemini 2.5 Pro was the most robust model, successfully processing the full 790‐reference batch. In contrast, GPT‐5 failed at batches ≥400, while GPT‐5 mini and Gemini 2.5 Flash failed at the 790‐reference batch. Overall, all models demonstrated strong performance within their operational ranges, with two notable exceptions: Gemini 2.5 Flash showed low initial sensitivity at batch 1, and GPT‐5 mini's sensitivity degraded at higher batch sizes (from 0.88 at batch 200 to 0.48 at batch 400). At a practical batch size of 100, Gemini 2.5 Pro achieved the highest sensitivity (1.00, 95% CI 1.00–1.00), whereas GPT‐5 delivered the highest specificity (0.98, 95% CI 0.98–0.98). Conclusion State‐of‐the‐art LLMs can effectively screen multiple abstracts per prompt, moving beyond inefficient single‐reference processing. However, performance is model‐dependent, revealing trade‐offs between sensitivity and specificity. Therefore, batch size optimization and strategic model selection are important parameters for successful implementation.
ABSTRACT With the increasing number of research priority setting (RPS) exercises, systematic reviews synthesising their findings have also grown in prevalence. While these reviews offer a structured way to compare methodologies, identify underrepresented stakeholder groups, and guide funding decisions, conventional systematic review methodologies, designed primarily for clinical and health research, often fail to capture the complexity, contextual nuance, and participatory nature of RPS. In this commentary, we critically examine these limitations and propose methodological adaptations to enhance the relevance and utility of systematic reviews of RPS. Beyond knowledge generation, we highlight the broader implications of RPS, including its role in stakeholder engagement, research funding allocation, and policy translation, as well as its impact on how these exercises are synthesised. By re‐evaluating how systematic reviews of RPS are conducted, we advocate for context‐sensitive methodologies that better reflect the dynamic and iterative nature of research priority setting.
ABSTRACT Introduction Systematic reviews occupy a central position in evidence hierarchies, providing structured syntheses intended to inform clinical decision‐making and health policy. However, the rapid expansion of artificial intelligence (AI) tools in literature searching, screening, data extraction, and manuscript drafting is transforming how these reviews are produced. Concurrently, the number of prospectively registered systematic reviews has grown substantially, with recent increases in PROSPERO registrations highlighting an accelerating output of evidence syntheses. While technological advances promise efficiency and scalability, they also raise concerns regarding methodological rigor, redundancy, and transparency. Methods This viewpoint argues that the current reporting and governance frameworks for systematic reviews remain largely anchored in pre‐AI workflows. Results Ongoing updates to reporting standards, including PRISMA revisions, have yet to fully address key challenges introduced by AI‐assisted methodologies, such as algorithmic bias, auditability, reproducibility limitations of proprietary models, and the need to document human oversight. The absence of explicit guidance for reporting AI use creates a critical transparency gap, potentially undermining confidence in systematic reviews and increasing the risk of superficial or duplicated syntheses. Conclusion We propose that the evidence‐synthesis ecosystem requires urgent adaptation, including the development of a PRISMA‐AI extension, strengthened metadata requirements in registries such as PROSPERO, and updated editorial policies for AI‐assisted reviews. Safeguarding rigor in the age of automated science is essential to maintain the credibility and clinical utility of systematic reviews.