
OBJECTIVE:To determine the efficacy of ginger (either stand-alone or as an adjunct treatment) for the treatment and management of migraine headaches. METHODS:We searched PubMed, Embase, CINAHL, PsycINFO, Epistemonikos and Cochrane CENTRAL from inception to 26 March 2025 which was updated on 16 April 2026 and conducted a backwards and forwards citation search of included studies. We included randomised controlled trials (RCTs) that compared the effect of ginger (either stand-alone or as an adjunct) to placebo or usual active control (such as Depakine (valproic acid), amitriptyline or sumatriptan). Two authors independently screened articles, extracted data, assessed risk of bias using the Cochrane Risk of Bias 2 Tool (ROB-2) tool. Primary outcomes included headache severity or pain scores post-treatment, proportion of patients pain free or with pain relief post-treatment, frequency and duration of migraine attacks and migraine-associated symptoms. Random-effects meta-analyses were done and the certainty of evidence was assessed using Grading of Recommendations Assessment, Development and Evaluation (GRADE) methodology. RESULTS:Eight RCTs involving 594 patients were included. For acute treatments, ginger with ketoprofen compared with placebo with ketoprofen may have little to no difference on headache severity at 2 hours post (mean difference (MD) 1.27 lower, 95% CI 1.46 lower to 1.08 lower; 1 study, 60 participants; low certainty evidence). It was very uncertain if ginger combined with feverfew versus placebo or ginger alone compared with sumatriptan improved proportions of those with pain relief, pain free or the mean headache severity at 2 hours post. For preventative treatment, ginger in addition to an active drug compared with the same active drug with or without placebo, may result in an important reduction in headache pain severity at 3 months (pooled MD -3.18, 95% CI -3.94 to -2.42; 2 studies, 183 participants; low certainty evidence). However, it was very uncertain if ginger combined with herbs compared with active drugs of both Depakine and amitriptyline lowered headache frequency, duration or headache severity at 3 months. The overall certainty of evidence for all outcome measures ranged from low to very low. CONCLUSIONS:The volume of the existing evidence is limited and combined with either a low or very low certainty of evidence, it is currently too uncertain if ginger is effective either as an acute or preventative treatment. Therefore, additional acute and preventative trials are required to help raise the certainty of evidence. TRIAL REGISTRATION NUMBER:https://www.crd.york.ac.uk/PROSPERO/view/CRD420251025768.
Understanding the risks associated with healthcare interventions is essential for clinical decision-making and for the regulatory evaluation of new drugs. Yet, safety analyses in randomised controlled trials are often conducted without clearly specifying (a) the estimand of interest in the presence of intercurrent events (IEs) (ie, events such as treatment discontinuation or use of rescue medication that occur after randomisation and complicate outcome interpretation) or (b) the estimand implicitly targeted by the statistical estimator used in the analysis. This lack of clarity produces ambiguous results and complicates their interpretation.This paper provides practical guidance on applying the estimand framework from the ICH E9(R1) addendum to safety outcomes in phase III randomised trials conducted for regulatory submissions. We describe how IEs influence the interpretation of treatment effects in safety assessment and how commonly used safety estimators align with different estimands, making explicit both the clinical question and the assumptions required to estimate it.Although the primary emphasis is on regulatory trials, where these decisions are especially consequential, the principles outlined are equally applicable to other randomised interventions and non-interventional studies. As safety results from pivotal regulatory trials are increasingly published using estimand terminology, these concepts are essential not only for researchers such as statisticians, clinicians and regulators involved in clinical development but also for clinicians who prescribe newly approved drugs and need to interpret safety results communicated using estimand terminology to inform their clinical decision-making. This paper aims to make these concepts accessible to a broader medical audience.
In clinical research, statistically significant effects do not necessarily indicate that an intervention provides benefits that are meaningful to patients. This is particularly important for patient-reported outcomes, where thresholds used to interpret clinical relevance are often derived from within-person changes, such as the minimal clinically important difference, and then inappropriately applied to between-group effects in randomised trials and meta-analyses. This article clarifies the conceptual distinction between within-group change and between-group effects, and argues that the latter should be the focus when judging the comparative value of healthcare interventions. We introduce the smallest worthwhile effect (SWE) as a patient-centred, intervention-specific construct representing the smallest between-group effect of an intervention over a comparator that patients consider worthwhile when weighed against harms, costs and other inconveniences. We describe the main methods used to estimate the SWE, including benefit-harm trade-off studies and discrete choice experiments, and illustrate its application by reinterpreting a randomised trial of physiotherapy for low back pain and a meta-analysis of discectomy versus non-surgical care for sciatica. We also show how the SWE can inform judgements of imprecision within the Grading of Recommendations Assessment, Development and Evaluation (GRADE) framework, offering a patient-derived threshold for distinguishing trivial from worthwhile effects. By focusing on patient-derived thresholds rather than arbitrary statistical or within-group approaches, the SWE construct can help researchers, clinicians and other stakeholders make more transparent and meaningful judgements about the worthwhileness of compared healthcare interventions from a patient perspective.
Ethnicity data are essential in understanding how different groups are impacted by different treatments/healthcare interventions, but more work is needed in how these data should be collected, handled and reported in trials to address distrust among marginalised ethnic groups towards healthcare research. To do this, researchers need effective training and opportunities to engage with communities from a diverse range of ethnicities. Developing guidance on how to meaningfully involve diverse ethnic groups in all stages of a trial including planning the collection of ethnicity data is one way to improve the impact of clinical trial results and ensure these groups can access equitable healthcare.Our collaborative work between health researchers and diverse ethnic communities produced 13 initial recommendations aimed at academic health researchers to change their practice involving diverse ethnic communities in areas from trial design through to reporting. The first four recommendations are concerned with the planning stage of the trial, the fifth with how researchers interact with funding bodies. Recommendations 6, 7 and 8 are relevant to data collection during trial set up, and recommendations 9, 10 and 11 focus on training and responsibility during the trial set up and the trial conduct stage. Recommendations 12 and 13 cover reporting.The recommendations will benefit from piloting but have the potential to help trial teams make their research more inclusive with regard to ethnicity, which in turn will support more representative, equitable and impactful research.
OBJECTIVE:To clarify the perspectives of different interest-holders on the definition of 'evidence'. DESIGN AND SETTING:This qualitative study employed purposeful sampling to select respondents from various disciplines and fields worldwide for semistructured interviews. The initial draft of the interview guide was developed based on a scoping review of definitions of 'evidence' and relevant documents. A pilot test was conducted to refine the interview guide. Interviews were conducted both online and in person until data saturation was achieved. Interview transcripts were analysed in NVivo V.12 using thematic analysis, which involved generating initial codes, searching for and reviewing themes and agreeing on the final set of themes. RESULTS:Interviews were conducted with 25 interest-holders from 16 countries, including policymakers, researchers, journal editors, clinicians and representatives of patients and the public. Findings show that most respondents believe there is a lack of clear, precise definitions of evidence and that the term is used variably and informally. Although respondents identified multiple barriers to establishing a standardised definition, some questioned whether a new definition is necessary, argued that existing definitions are sufficient or noted the practical difficulties of developing one. Moreover, views differed on which interest-holders should be involved and considered in developing such a definition. Based on these discussions, 13 of the 25 respondents explicitly provided personal definitions of 'evidence' for users' reference. CONCLUSIONS:The unclear definitions of evidence and their inconsistent usage not only lead researchers to misunderstand and overuse the term but also contribute to the misuse or mistrust of evidence by users, especially patients and the public, and even to criticism of evidence-based practice. We urge users to consult existing definitions of the term 'evidence' before employing it, to ensure standardised and accurate usage. If no accepted definition is available, we recommend clearly defining the term prior to use, so that users can accurately comprehend its meaning.
ObjectiveShared decision-making (SDM) is a key component of patient-centred care and major driver of healthcare policy, clinical practice and research globally. Widespread implementation of SDM is hindered by uncertainty about how to best measure decision-making processes. A seminal methodological review of SDM instruments published in 2018 identified widespread deficiencies in instrument quality and provided recommendations for improvement. This systematic review and COnsensus-based Standards for the selection of health status Measurement INstruments (COSMIN) quality appraisal aims to provide an update SDM instrument inventory, re-evaluated progress in the field and provided future direction for researchers, funders and policymakers. This systematic review and COnsensus-based Standards for the selection of health status Measurement INstruments (COSMIN) quality appraisal aims to provide an update SDM instrument inventory, re-evaluated progress in the field and provided future direction for researchers, funders and policymakers. DESIGN:A systematic search of six databases (September 2017-December 2024) identified studies of instruments measuring the process of SDM. Instrument appraisal followed COSMIN guidelines in a three-step process: appraisal of the methodological quality of instruments and of the quality of the measurement properties of newly identified measures (steps i, ii). This informed the comprehensive best-evidence synthesis which included the data from the original review (step iii). RESULTS:A total of 127 studies were included describing the development and/or evaluation of 105 unique instruments to measure the process of SDM. Some 61 new instruments were identified since the last review, of which 35 (57%) were translations/revisions to existing ones. The best-evidence synthesis revealed positive results for internal consistency (72%), structural validity (75%) and intrarater reliability (50%), but also negative or unknown results for test-retest reliability (58%), content validity (54%) and intrarater reliability (50%), where evaluated. No single instrument consistently showed positive evidence across all relevant measurement properties. Evidence gaps for many properties remain. CONCLUSION:This updated inventory guides researchers, clinicians and policymakers to the most appropriate instrument for an intended context/population. There was a proliferation of new, primarily patient-reported measurement instruments, but no single instrument had sufficient evidence of measurement quality. Methods were aligned with the original review and applied the 2011 COSMIN checklist which was updated in 2024, although using a previous checklist version is unlikely to affect overall conclusions. We recommend an international consensus with key interest holders on preferred instrument(s) for further validation and future core measurement set for efficient evidence synthesis. REGISTRATION:Prospero: CRD42024485655.
Background Proton pump inhibitors (PPIs) have been widely used for over 35 years. However, in recent decades, numerous adverse drug reactions (ADRs) have been reported, with evidence often inconsistent and heterogeneous across studies. Objective To conduct an overview of systematic reviews/meta-analyses (SR/MAs) to provide a contemporary review of the evidence for the safety of PPIs in patients using them for treating or prophylaxis, to summarise the outcome data, assess the methodological quality and rate the certainty of the evidence and to provide references for clinical decision-making and the subsequent formulation of evidence. Methods We conducted an overview of reviews following the Preferred Reporting Items for Overviews of Reviews guidelines. We identified SR/MAs regarding PPI safety through a search of multiple databases, including the China National Knowledge Infrastructure, Chinese Science and Technology Journal Database and Wanfang Database (WanFang) on 6 December 2025 as well as Embase, PubMed and the Cochrane Library database on 7 December 2025. The participants were human populations who had been administered PPIs; the intervention groups used PPI-based regimens and the control group received different PPI regimens, no-PPIs, H2-receptor antagonists (H2RAs), potassium-competitive acid blockers (P-CAB) or placebo; the outcomes mainly included the specific ADRs associated with PPI use. We used the A Measurement Tool to Assess Systematic Reviews-2 (AMSTAR-2) tool to assess the methodological quality of the included SR/MAs and employed the Grading of Recommendations Assessment, Development and Evaluation (GRADE) framework to evaluate the quality of evidence for the reported outcomes. The focus of the data presentation was descriptive, featuring detailed tabular presentations of characteristics and results at both the review level and the primary studies level. Main results A total of 940 studies were retrieved from various databases according to the search strategies. After eliminating duplicates and the screening process, 36 SR/MAs were finally included. Overall, there was a slight overlap (corrected covered area, CCA of 0.91%) in 483 primary studies included in the 36 SR/MAs. AMSTAR-2 evaluation results showed that 34 SR/MAs were of low quality, and the other two were of critically low quality. The GRADE evaluation results indicated that the certainty of evidence for all outcomes was low or very low . The use of PPIs might be associated with an increased risk of acute kidney injury (AKI) (relative risk (RR) 1.75, 95% CI 1.40 to 2.19), chronic kidney disease (CKD) (RR 1.35, 95% CI 1.15 to 1.56), gastric cancers (GC) (RR 1.67, 95% CI 1.39 to 2.00) and community-acquired pneumonia (CAP) (OR 1.37, 95% CI 1.22 to 1.53). PPI therapy was associated with a higher recurrence rate of Clostridioides difficile infection (CDI) (24% vs 18%) and might increase CDI risk in renal transplant recipients (RR 2.33, 95% CI 1.07 to 5.07). The use of PPIs might heighten the risk of fractures in young patients (aged <29 years) (RR 1.20, 95% CI 1.12 to 1.29). PPI therapy was also linked to hypomagnesaemia in haemodialysis patients (1/36, 2.78%) and renal transplant recipients (1/36, 2.78%). Conclusions This overview evaluates the safety profile of PPIs, noting associated risks including renal impairment, GC, fractures and infections. Most evidence comes from observational studies, resulting in low or very low certainty evidence that limits definitive causal conclusions. Clinicians and pharmacists may consider greater vigilance with PPI use, and these ADRs may not require de-escalation but de-implementation when inadequate (eg, long-term use or special group use). Future research should focus on high-quality prospective studies and investigate underlying mechanisms to better establish causality and quantify risks. PROSPERO registration number CRD420251059575.
OBJECTIVES:Migraine headaches are common and potentially disabling disorders, with several interventions available for prevention and symptom reduction. We explored the comparative effectiveness and tolerability of pharmacological prophylaxis for migraine through a network meta-analysis of randomised trials (RCTs). DESIGN:Our study design was a systematic review and network meta-analysis (PROSPERO registration CRD42023456915). ELIGIBILITY CRITERIA:We included randomised controlled trials of prophylactic pharmacological interventions that enrolled adults diagnosed with chronic and/or episodic migraine headaches. DATA SOURCES:Medline, Embase, Cochrane Central, PsycINFO, Web of Science and Scopus from inception to 15 January 2026. RISK OF BIAS AND CERTAINTY EVIDENCE:Risk of bias was assessed using the modified Cochrane risk-of-bias tool 2.0 and the certainty evidence was evaluated by using the Grading of Recommendations Assessment, Development and Evaluation approach. SYNTHESIS OF RESULTS:We performed a frequentist network meta-analysis using a random-effects model to compare the efficacy of interventions. RESULTS:We included 199 RCTs (47 420 participants). Overall, 29 trials (14.6%) were at low risk of bias; an adequate random allocation sequence generation was reported in 92 trials (46.2%), and missing outcome data was the most common limitation (110 trials, 55.3%). Compared with placebo, calcium channel blockers (mean difference (MD) -1.78 (95% CI -2.96 to -0.60), moderate certainty), calcitonin gene-related peptide (CGRP)-targeted therapies (MD -1.69 (95% CI -2.16 to -1.23), high certainty) and beta-blockers (MD -1.50 (95% CI -2.54 to -0.47), moderate certainty) were the most effective in reducing monthly migraine days. Moderate certainty evidence suggests beta-blockers (MD -1.31 (95% CI -1.76 to -0.85)), calcium channel blockers (MD -1.11 (95% CI -1.65 to -0.57)), anticonvulsants (MD -1.12 (95% CI -1.66 to -0.58)) and CGRP-targeted therapies (MD -0.76 (95% CI -1.49 to -0.02)) probably reduce monthly migraine attacks. However, moderate to high certainty evidence found that patients were more likely to discontinue calcium channel blockers (relative risk (RR) 1.40, 95% CI 1.04 to 1.88) and anticonvulsants (RR 1.14, 95% CI 1.01 to 1.29), compared with placebo. CONCLUSIONS:When restricted to moderate or high certainty evidence, beta-blockers and CGRP-targeted therapies probably reduce migraine frequency and may be well-tolerated prophylactic options for migraine. Calcium channel blockers and anticonvulsants may also be effective for reducing migraine frequency but are less well tolerated by some patients. PROSPERO REGISTRATION NUMBER:CRD42023456915.
OBJECTIVES:Probability or risk outcome estimates tailored to individual patient characteristics can facilitate more personalised care, for example, regarding treatment. Yet, the uptake of prediction models in clinical guidelines is lacking. It is unknown whether documents or handbooks for clinical guideline development provide any specific guidance, similar to clinical treatments, for how to assess and incorporate prediction models in clinical guidelines. We performed this systematic review of clinical guideline development documents and handbooks to explore the methodology regarding searching, selecting, critically appraising and implementing prediction models in clinical guidelines to enhance their usability in decision-making. DESIGN AND SETTING:We conducted a comprehensive systematic search for clinical guideline development documents of organisations included in the Guidelines International Network members directory and identified additional handbooks via PubMed and the TRIP database in October 2024. Data extracted from the retrieved documents included organisational details, methods and recommendations and guidance for prediction model inclusion. A narrative synthesis was performed. RESULTS:A total of 84 guidance documents were included in the final extraction, with eight (10.5%) guidance documents mentioning specific guidance for prediction models or related terms such as prognostic models or risk prediction. The included guidance related to searching, selecting and critically appraising prediction models for inclusion in clinical guidelines. No guidance on model performance or selection between competing models was provided. CONCLUSIONS:The minority of identified handbooks included guidance regarding appraising and deciding whether and how to include prediction models in clinical guidelines. The absence of such guidance underscores the need for more extensive and standardised methodology for incorporating prediction models in clinical guidelines. This could help facilitate implementation of models into clinical guidelines, clinical practice and eventually impact decision making and subsequently patient outcomes.
OBJECTIVES:To evaluate the performance of large language models (LLMs) in risk of bias assessment and to examine whether prompt engineering improves their accuracy and alignment with expert reasoning. METHODS:We analysed 158 randomised controlled trials from 10 dental systematic reviews and their risk of bias assessments were reviewed and revised to serve as the reference standard. Two LLMs (DeepSeek-V3 and GPT-5) were evaluated under four prompting strategies, including direct command, command with reference, constrained output and formula-constrained output. The direct command served as the blank control group, simulating the approach commonly used by clinicians, whereas the other three groups employed different prompt engineering. The performance of LLMs across the seven domains of RoB-1 was evaluated using accuracy and agreement. The reasoning process of the LLMs was expressed in the form of syllogisms and its similarity to expert reasoning was assessed using MMD2. RESULTS:LLMs showed limited capability in risk of bias assessment under the blank control condition, with mean accuracies of 0.72 for DeepSeek-V3 and 0.65 for GPT-5. With formula-constrained prompting, the performance of both LLMs improved significantly, and the overall accuracy increased to 0.85 for both DeepSeek-V3 and GPT-5 (both vs the blank control group, p<0.001). Agreement metrics showed a similar pattern, with higher agreement under formula-constrained prompting than under the other prompting strategies (p<0.001 for both models). In addition, the syllogistic output format provided a clear representation of the reasoning process underlying risk of bias assessment. Compared with constrained output, formula-constrained prompting also produced reasoning that was more closely aligned with the reference answers, as indicated by lower MMD² values (DeepSeek-V3: 0.0765 vs 0.1239; GPT-5: 0.0548 vs 0.1068). CONCLUSION:Prompt engineering substantially improved the performance of LLMs in risk of bias assessment. Although LLMs cannot currently replace human reviewers, they may serve as efficient and transparent tools to support this process.
Objective To examine the potential errors of a general large language model (LLM) (ie, Claude 3.5 Sonnet) on data extraction from randomised controlled trials (RCTs).Design and setting An empirical study comparing Claude 3.5 Sonnet extractions against a human-performed verification dataset. The extraction tasks for Claude 3.5 Sonnet were based solely on original RCT portable document format (PDF) files. For PDFs that could not be directly extracted by Claude 3.5 Sonnet, optical character recognition was employed to convert them into text format before extraction.Participants A random sample of 664 trials was selected from a well-established trial bank and a final data pool was established based on rigorous manual cross-checking as a reference standard.Data sources PubMed, EMBASE, Scopus, Web of Science (all databases) and the Cochrane Central Register of Controlled Trials (CENTRAL) up to February 2023.Eligibility criteria for selecting studies RCTs on children involving medication and adverse events.Main outcome measures Claude 3.5 Sonnet was applied to extract the basic information (eg, trial design, population information and source of funding) and adverse outcomes (ie, name of adverse events, number of events). Claude 3.5 Sonnet outputs were compared against the final data pool and all errors were recorded. Results are presented as error rates and with 95% CI, estimated using a generalised linear mixed model.Results For the 664 trials, a total of 23 069 data cells were extracted via Claude 3.5 Sonnet, with 10 624 for basic information and 12 445 for adverse outcomes. The overall error rate for data extraction was 6.6% (95% CI 5.4% to 8.2%), with 5.7% (95% CI 5.2% to 6.1%) in basic information and 7.6% (95% CI 4.9% to 11.8%) in adverse outcomes. When stratified the 1542 total errors by error types, misallocation (assigning data to incorrect fields; 57.1%, 881/1542) and missed or omitted data (incomplete extraction of available data; 23.2%, 357/1542) accounted for the two most frequent errors, with misallocation occurring more in basic information (53.3%, 470/881), while missed or omitted data occurred more in adverse outcomes (96.1%, 343/357). Post hoc analysis examining the association between trial reporting quality (assessed using Consolidated Standards of Reporting Trials (CONSORT) 2025 and LLM data extraction error rates indicated that higher CONSORT adherence was associated with lower extraction error rates.Conclusions The data extraction error of Claude was relatively low, but it alerts LLM applications in evidence synthesis. Detailed checking for LLM outputs should be the primary consideration for evidence synthesisers.
OBJECTIVES:Observing Patient Involvement in Decision Making (OPTION)-12 and OPTION-5 assess the extent to which observers score healthcare professionals' (HCPs) involvement of patients in shared decision-making (SDM). We systematically reviewed studies measuring the extent to which HCPs involve patients in the decision-making process using the OPTION instrument. DESIGN:Informed by Preferred Reporting Items for Systematic Reviews and Meta-Analyses, we updated a previous systematic review and included new studies reporting OPTION-12 or OPTION-5 scores from recordings of real-world clinical encounters, involving patients and HCPs making healthcare-related decisions. Searches were conducted across PubMed, EMBASE, Cochrane Central Register of Controlled Trials (CENTRAL), and Web of Science databases (2012-2025), supplemented by citation screening and outreach to professional networks. We extracted study characteristics, OPTION version, psychometric data and item-level score details. We also assessed the study quality using the reports of rating procedures and conducted meta-analyses, subgroup analyses using a priori hypotheses and completed meta-regressions. RESULTS:In total, 174 studies were included, comprising almost 20 000 clinical consultations: 102 studies used only OPTION-12 and 64 used only OPTION-5, while four studies reported using both scales. Mean OPTION-12 and OPTION-5 score for studies unaffected by interventions were 25.1 (95% CI 22.1 to 28.2, k=76, I2=99.71%) and 31.8 (95% CI 26.6 to 37.1, k=42, I2=99.55%), respectively. Subgroup analyses revealed significantly higher scores in studies with postintervention OPTION-scores for both OPTION-12 (38.4 vs 25.1, p<0.001, k=91, I2=99.55%) and OPTION-5 (47.7 vs 31.8, p<0.001, k=65, I2=99.39%). In univariable meta-regression, longer consultation duration and female patient percentage (only for OPTION-12) were associated with higher scores. However, multivariable meta-regression revealed that clinical setting was the sole independent predictor for OPTION-12 (p=0.007), whereas consultation duration remained the primary independent predictor for OPTION-5 (p=0.003). CONCLUSIONS:Since the 2015 previous review, little overall improvement has been observed. This limited progress raises important questions about how we interpret changes in observed SDM. Specifically, it remains unclear what degree of change in OPTION-12 scores reflects a meaningful improvement. Our multivariable findings provide a more nuanced perspective: while consultation duration remains the primary independent predictor for patient involvement when measured with OPTION-5, clinical setting emerges as a more critical independent driver for OPTION-12. These results suggest that the influence of time is not uniform across assessment tools and that structural barriers in different clinical environments must also be addressed to foster SDM effectively. PROSPERO REGISTRATION NUMBER:CRD42022332231.
OBJECTIVES:To investigate the effectiveness and clinical relevance of kinesio taping (KT) in musculoskeletal disorders (MSDs) at different follow-ups. DESIGN:Overview of systematic reviews (SRs) and evidence mapping. INFORMATION SOURCES:Ten electronic databases were searched for SRs published from inception to 31 December 2024, and updated on 15 October 2025. ELIGIBILITY CRITERIA:SRs with and without meta-analysis of randomised controlled trials (RCTs) were eligible for inclusion if they compared KT with interventions other than KT (eg, active interventions, no tape, placebo/sham KT) in participants with MSDs. MAIN OUTCOME MEASURES:The primary outcomes were pain intensity, function/disability, range of motion, muscle strength, quality of life and disease-specific symptoms. The secondary outcome was adverse events (AEs). RESULTS:A total of 128 SRs (73 published SRs and 55 registered yet unpublished SRs) involving 15 812 participants from 310 unique RCTs were included. Substantial SRs were focused on lower extremity conditions (45%) and reported pain intensity (89%). Most SRs were evaluated as critically low (78%) in methodological quality and low (58%) in risk of bias, with a median total compliance rate of 75.6% in reporting quality. Findings from new meta-analyses indicated that KT may reduce pain intensity in the immediate (Hedges' g -0.69, 95% CI -0.81 to -0.57) and short (Hedges' g -0.57, 95% CI -0.77 to -0.37) term and improve function/disability (Hedges' g -0.54, 95% CI -0.69 to -0.40) in the immediate term. These effect estimates may achieve the predefined minimal clinically important difference of 0.5 SD (medium effect size). KT may show little to no effect on pain intensity in the medium term, function/disability in the short and medium term, muscle strength, range of motion, disease-specific symptoms at all follow-ups. The effects of KT may vary across subgroups or conditions, and its impact on quality of life is unclear. AEs related to KT mainly included skin irritation (number needed to harm (NNH) 173) and pruritus (NNH 356). All evidence was highly inconclusive due to very low certainty (Grading of Recommendations Assessment, Development and Evaluation), non-significant level (evidence level) and unstable clinical relevance across most outcomes. CONCLUSIONS:Current evidence is very uncertain regarding the clinical effects of KT on MSDs. Considerable heterogeneity, unclear clinical relevance and potential AEs may limit its application in clinical practice. Further high-quality, well-reported RCTs and SRs are warranted to address the uncertainty regarding overall effects along with comprehensive consideration of heterogeneity in KT usage. PROSPERO REGISTRATION NUMBER:CRD42024517528.