We show that the covariance matrix of the treatment effect estimates in a network meta-analysis can be obtained without matrix inversion using a geometric series of diffusion matrices. This property extends to the hat matrix and provides a connection between parameter estimation in regression analysis and random walks on the network graph. We also provide a number of visualization tools implemented in R.
BACKGROUND: Network meta-analysis (NMA) is a widely used method for synthesizing evidence from multiple interventions for a medical condition. However, NMA applications typically ignore the crucial role of drug dosage on intervention effects. Traditional NMAs either consider each intervention dose as an independent node or ignore the intervention dose, which may impact heterogeneity, inconsistency, or sparsity. METHODS: This paper introduces a novel frequentist approach, termed dose-response network meta-analysis (DR-NMA), which explicitly models the dose-response relationships across multiple interventions. The DR-NMA approach incorporates both linear and nonlinear dose-response relationships, including exponential, quadratic, fractional polynomials, and restricted cubic splines. DR-NMA allows for dose-dependent estimation and prediction of treatment effects across dose ranges, even in disconnected networks if common agents exist. The proposed methods are implemented in the R package netdose, enhancing accessibility and reproducibility. We illustrate the approach using clinical datasets on postoperative nausea and vomiting, as well as antidepressant treatments. RESULTS: Our findings indicate that some dose-response NMA models yield substantially different results compared to standard NMA, emphasizing the critical importance of dose-response function selection in model performance. CONCLUSIONS: DR-NMA provides valuable insights into the dose-dependent effects of interventions, enhancing decision-making and offering perspectives beyond traditional methods.
BACKGROUND/AIM:We conducted a systematic review with network meta-analyses (NMA) summarizing the effects and safety of lifestyle interventions containing nutrition (NUT; e.g., calorie restriction), exercise (EX; e.g., aerobic/resistance exercise) and behavior change interventions (BCI; e.g., behavioral therapy) on physical function, body composition, quality of life, psychosocial outcomes, health and adverse events in community-dwelling older adults with obesity. METHODS:We used the methodology proposed by Cochrane and searched six databases and one trial registry for eligible randomized controlled trials (RCTs; intervention duration ≥ 12 weeks) up to May 2022 with a full new search in MEDLINE and a re-assessment of previously identified eligible trial registry entries in October 2025. Random-effects NMA ((standardized) mean difference ((S)MD), 95% confidence intervals) were conducted if possible. RESULTS:We included 72 RCTs (n = 6716) for descriptive summaries and 54 RCTs (n = 4249) for NMA. NUT+EX+BCI improved physical function (performance batteries) compared to control (SMD 3.37 [1.76;4.97]; high certainty of evidence). NUT+EX+BCI may reduce body (MD -8.69 [-13.14;-4.25]) and fat mass (MD -6.58 [-10.44;-2.73]) while not negatively affecting fat-free mass (MD -1.38 [-3.52;0.76]) or bone mineral density (MD -0.01 [-0.05;0.02]) (evidence very uncertain). Other interventions (single/combined) may also be effective; however, effects were often imprecise. For psychosocial outcomes, quality of life, and health events, data were insufficient or too heterogeneous to derive clear results. CONCLUSION:The evidence suggests that NUT+EX+BCI interventions are most suitable for the management of obesity in older adults. Nevertheless, further RCTs-especially in frail populations and on patient-relevant outcomes-are needed.
A key output of network meta-analysis (NMA) is the relative ranking of treatments; nevertheless, it has attracted substantial criticism. Existing ranking methods often lack clear interpretability and fail to adequately account for uncertainty, overemphasizing small differences in treatment effects. We propose a novel framework to estimate treatment hierarchies in NMA using a probabilistic model, focusing on a clinically relevant treatment-choice criterion (TCC). Initially, we define a TCC based on smallest worthwhile differences (SWD), converting NMA relative treatment effects into treatment preference format. These data are then synthesized using a probabilistic ranking model, assigning each treatment a latent "ability" parameter, representing its propensity to yield clinically important and beneficial true treatment effects relative to the rest of the treatments in the network. Parameter estimation relies on the maximum likelihood theory, with standard errors derived asymptotically from the Hessian matrix. To facilitate the use of our methods, we launched the R package mtrank. We applied our method to two clinical datasets: one comparing 18 antidepressants for major depression and another comparing 6 antihypertensives for the incidence of diabetes. Our approach provided robust, interpretable treatment hierarchies that account for a concrete TCC. We further examined the agreement between the proposed method and existing ranking metrics in 153 published networks, concluding that the degree of agreement depends on the precision of the NMA estimates. Our framework offers a valuable alternative for NMA treatment ranking, mitigating overinterpretation of minor differences. This enables more reliable and clinically meaningful treatment hierarchies.
BACKGROUND:Anxiety disorders often remain undetected and can cause substantial burden. Amongst the many anxiety screening tools, the 7-item Generalized Anxiety Disorder (GAD-7) scale and its short version, the 2-item Generalized Anxiety Disorder (GAD-2) scale, are the most frequently used instruments. OBJECTIVES:Primary: to determine the diagnostic accuracy of GAD-7 and GAD-2 to detect generalised anxiety disorder (GAD) and any anxiety disorder (AAD) in adults. Secondary: to investigate whether their diagnostic accuracy varies by setting, anxiety disorder prevalence, reference standard, and risk of bias; to compare the diagnostic accuracy of GAD-7 and GAD-2; to investigate how diagnostic performance changes with the test threshold. SEARCH METHODS:We searched MEDLINE, Embase, PubMed-not-MEDLINE subset, and PsycINFO from 1990 to 18 January 2024. We checked reference lists of included studies and review articles. SELECTION CRITERIA:We included cross-sectional studies conducted in adults, containing diagnostic accuracy information on GAD-7 and/or GAD-2 questionnaires for the target conditions generalised anxiety disorder and/or any anxiety disorder, and allowing the generation of 2x2 tables. The target conditions must have been diagnosed using a structured or semi-structured clinical interview. We excluded case-control studies and studies in which the time elapsed between the index tests and reference standards exceeded four weeks. We excluded studies involving people (1) seeking help in mental health settings or (2) recruited specifically due to mental health symptoms in other settings. DATA COLLECTION AND ANALYSIS:At least two review authors independently decided on study eligibility, extracted data, and assessed the risk of bias and applicability of included studies. For each questionnaire and each target condition, we present sensitivity and specificity with 95% confidence intervals (95% CI) in forest plots. We used the bivariate model to obtain summary estimates based on cut-offs closest to the recommended values (i.e. within a core range). In secondary analyses, we used the bivariate model and the multiple thresholds model to obtain summary estimates for all available cut-off points. Using the multiple thresholds model, we also calculated the area under the receiver operating characteristic curve to obtain a general indicator of the diagnostic accuracy of GAD-7 and GAD-2. MAIN RESULTS:We included 48 studies with 19,228 participants from 27 different countries, evaluating the GAD-7 and the GAD-2 in 24 different languages. Seven studies were performed in non-clinical settings, nine in clinical settings recruiting participants across conditions, and 32 in clinical settings with participants having specific conditions. Even after categorisation into three settings, the study populations were substantially different. The most frequently studied populations were people: with epilepsy (nine studies); with cancer (five studies); with cardiovascular disease (five studies); and in primary care regardless of their condition (five studies). We considered the risk of bias low in eight studies, and we had low concerns about the applicability of findings in three studies. Thirty-five studies contributed to the primary analyses of GAD-7 for detecting generalised anxiety disorder (median prevalence 12%); 22 studies to analyses of GAD-7 for any anxiety disorder (median prevalence 19%); 24 studies to analyses of GAD-2 for generalised anxiety disorder (median prevalence 9%); and 19 studies to analyses of GAD-2 for any anxiety disorder (median prevalence 19%). At the recommended cut-off of 10 or higher (or the closest available cut-off), the GAD-7 questionnaire yielded a summary sensitivity of 0.64 (95% CI 0.56 to 0.72) and a summary specificity of 0.91 (95% CI 0.87 to 0.93) in detecting generalised anxiety disorder. For detecting any anxiety disorder, summary sensitivity was 0.48 (95% CI 0.40 to 0.57) and summary specificity 0.91 (95% CI 0.89 to 0.93). At the recommended cut-off of 3 or higher (or the closest available cut-off), the GAD-2 yielded a summary sensitivity of 0.68 (95% CI 0.59 to 0.75) and a summary specificity of 0.86 (95% CI 0.82 to 0.89) for detecting generalised anxiety disorder. For detecting any anxiety disorder, the summary sensitivity was 0.53 (95% CI 0.44 to 0.62) and the summary specificity was 0.89 (95% CI 0.86 to 0.91). The 95% prediction region of GAD-7 for detecting generalised anxiety disorder was larger (indicating pronounced statistical heterogeneity) than for the three other analyses. Specificity varied by setting in the analysis of GAD-7 and GAD-2 for detecting any anxiety disorder, and by reference standard in the analysis of GAD-2 for detecting generalised anxiety disorder. Sensitivity varied with prevalence in the analysis of GAD-7 for generalised anxiety disorder. Other investigations of potential sources of heterogeneity did not show statistically significant associations with test accuracy. In all analyses, sensitivity tended to be higher and specificity lower in participants with specific conditions compared to the other two settings. Overall, the heterogeneity in the subgroup analyses remained high. The area under the receiver operating characteristic curve in the multiple thresholds model was 0.86 (95% CI 0.84 to 0.88) for the GAD-7 scale in detecting generalised anxiety disorder, and 0.80 (95% CI 0.78 to 0.82) in detecting any anxiety disorders. For the GAD-2 scale, the value was 0.82 (95% CI 0.81 to 0.86) for detecting generalised anxiety disorder, and 0.77 (95% CI 0.76 to 0.82) for detecting any anxiety disorders. Comparative bivariate analyses revealed no statistically significant differences between the diagnostic test accuracy of GAD-7 and GAD-2. AUTHORS' CONCLUSIONS:The GAD-7 and the GAD-2 scales have been tested in numerous languages and different populations. Overall, the GAD-7 and the GAD-2 seem to have acceptable or good diagnostic accuracy for both generalised anxiety disorder and any anxiety disorder. The GAD-2 scale seems to have similar diagnostic accuracy as the GAD-7 scale. However, due to the diversity of the included studies and the heterogeneity of our findings, our summary estimates of sensitivity and specificity should be interpreted as rough averages. The performance of GAD-7 and GAD-2 may deviate substantially from these values in specific situations.
Network Meta-Analysis (NMA) plays a pivotal role in synthesizing evidence from various sources and comparing multiple interventions. At its core, NMA relies on integrating both direct evidence from head-to-head comparisons and indirect evidence from different paths that link treatments through common comparators. A key aspect is evaluating consistency between direct and indirect sources. Existing methods to detect inconsistency, although widely used, have limitations. For example, they do not account for differences within indirect sources or cannot estimate inconsistency when direct evidence is absent. In this paper, we introduce a path-based approach that explores all sources of evidence without separating direct and indirect. We introduce a measure based on the square of differences to quantitatively capture inconsistency, and propose a Netpath plot to visualize inconsistencies between various paths. We provide an implementation of our path-based method within the netmeta R package. Via application to fictional and real-world examples, we show that our method is able to detect and visualize inconsistency between multiple paths of evidence that would otherwise be masked by considering all indirect sources together. The path-based approach therefore provides a more comprehensive evaluation of inconsistency within a network of treatments.
For network meta-analysis (NMA), we usually assume that the treatment arms are independent within each included trial. This assumption is justified for parallel design trials and leads to a property we call consistency of variances for both multi-arm trials and NMA estimates. However, the assumption is violated for trials with correlated arms, for example, split-body trials. For multi-arm trials with correlated arms, the variance of a contrast is not the sum of the arm-based variances, but comes with a correlation term. This may lead to violations of variance consistency, and the inconsistency of variances may even propagate to the NMA estimates. We explain this using a geometric analogy where three-arm trials correspond to triangles and four-arm trials correspond to tetrahedrons. We also investigate which information has to be extracted for a multi-arm trial with correlated arms and provide an algorithm to analyze NMAs including such trials.
Network meta-analysis (NMA) is an extension of pairwise meta-analysis that facilitates the estimation of relative effects for multiple competing treatments. A hierarchy of treatments is a useful output of an NMA. Treatment hierarchies are produced using ranking metrics. Common ranking metrics include the Surface Under the Cumulative RAnking curve (SUCRA) and P-scores, which are the frequentist analogue to SUCRAs. Both metrics consider the size and uncertainty of the estimated treatment effects, with larger values indicating a more preferred treatment. Although SUCRAs and P-scores themselves consider uncertainty, treatment hierarchies produced by these ranking metrics are typically reported without a measure of certainty, which might be misleading to practitioners. We propose a new metric, Precision of Treatment Hierarchy (POTH), which quantifies the certainty in producing a treatment hierarchy from SUCRAs or P-scores. The metric connects three statistical quantities: The variance of the SUCRA values, the variance of the mean rank of each treatment, and the average variance of the distribution of individual ranks for each treatment. POTH provides a single, interpretable value that quantifies the extent of certainty in producing a treatment hierarchy. We show how the metric can be adapted to apply to subsets of treatments in a network, for example, to quantify the certainty in the hierarchy of the top three treatments. We calculate POTH for a database of NMAs to investigate its empirical properties, and we demonstrate its use on two published networks.
The LFK index has been promoted as an improved method to detect bias in meta-analysis. Putatively, its performance does not depend on the number of studies in the meta-analysis. We conducted a simulation study, comparing the LFK index test to three standard tests for funnel plot asymmetry in settings with smaller or larger group sample sizes. In general, false positive rates of the LFK index test markedly depended on the number and size of studies as well as the between-study heterogeneity with values between 0% and almost 30%. Egger's test adhered well to the pre-specified significance level of 5% under homogeneity, but was too liberal (smaller groups) or conservative (larger groups) under heterogeneity. The rank test was too conservative for most simulation scenarios. The Thompson-Sharp test was too conservative under homogeneity, but adhered well to the significance level in case of heterogeneity. The true positive rate of the LFK index test was only larger compared with classic tests if the false positive rate was inflated. The power of classic tests was similar or larger than the LFK index test if the false positive rate of the LFK index test was used as significance level for the classic tests. Under ideal conditions, the false positive rate of the LFK index test markedly and unpredictably depends on the number and sample size of studies as well as the extent of between-study heterogeneity. The LFK index test in its current implementation should not be used to assess funnel plot asymmetry in meta-analysis.
Quantifying the contributions, or weights, of comparisons or single studies to the estimates in a network meta-analysis (NMA) is an active area of research. We extend this to the contributions of paths to NMA estimates. We present a general framework, based on the path-design matrix, that describes the problem of finding path contributions as a linear equation. The resulting solutions may have negative coefficients. We show that two known approaches, called shortestpath and randomwalk, are special solutions of this equation, and both meet an optimization criterion, as they minimize the sum of absolute path contributions. In general, there is an infinite space of solutions, which can be identified using the generalized inverse (Moore-Penrose pseudoinverse). We consider two further special approaches. For complex networks we find that shortestpath is superior with respect to run time and variability, compared to the other approaches, and is thus recommended in practice. The path-weights framework also has the potential to answer more general research questions in network meta-analysis.
We conducted a systematic review investigating the efficacy and tolerability of adrenocorticotropic hormone (ACTH) and corticosteroids in children with epilepsies other than infantile epileptic spasm syndrome (IESS) that are resistant to anti-seizure medication (ASM). We included retrospective and prospective studies reporting on more than five patients and with clear case definitions and descriptions of treatment and outcome measures. We searched multiple databases and registries, and we assessed the risk of bias in the selected studies using a questionnaire based on published templates. Results were summarized with meta-analyses that pooled logit-transformed proportions or rates. Subgroup analyses and univariable and multivariable meta-regressions were performed to examine the influence of covariates. We included 38 studies (2 controlled and 5 uncontrolled prospective; 31 retrospective) involving 1152 patients. Meta-analysis of aggregate data for the primary outcomes of seizure response and reduction of electroencephalography (EEG) spikes at the end of treatment yielded pooled proportions (PPs) of 0.60 (95% confidence interval [CI] 0.52-0.67) and 0.56 (95% CI 0.43-0.68). The relapse rate was high (PP 0.33, 95% CI 0.27-0.40). Group analyses and meta-regression showed a small benefit of ACTH and no difference between all other corticosteroids, a slightly better effect in electric status epilepticus in slow sleep (ESES) and a weaker effect in patients with cognitive impairment and "symptomatic" etiology. Obesity and Cushing's syndrome were the most common adverse effects, occurring more frequently in trials addressing continuous ACTH (PP 0.73, 95% CI 0.48-0.89) or corticosteroids (PP 0.72, 95% CI 0.54-0.85) than intermittent intravenous or oral corticosteroid administration (PP 0.05, 95% CI 0.02-0.10). The validity of these results is limited by the high risk of bias in most included studies and large heterogeneity among study results. This report was registered under International Prospective Register of Systematic Reviews (PROSPERO) number CRD42022313846. We received no financial support.
The development of methods for the meta-analysis of diagnostic test accuracy (DTA) studies is still an active area of research. While methods for the standard case where each study reports a single pair of sensitivity and specificity are nearly routinely applied nowadays, methods to meta-analyze receiver operating characteristic (ROC) curves are not widely used. This situation is more complex, as each primary DTA study may report on several pairs of sensitivity and specificity, each corresponding to a different threshold. In a case study published earlier, we applied a number of methods for meta-analyzing DTA studies with multiple thresholds to a real-world data example (Zapf et al., Biometrical Journal. 2021; 63(4): 699-711). To date, no simulation study exists that systematically compares different approaches with respect to their performance in various scenarios when the truth is known. In this article, we aim to fill this gap and present the results of a simulation study that compares three frequentist approaches for the meta-analysis of ROC curves. We performed a systematic simulation study, motivated by an example from medical research. In the simulations, all three approaches worked partially well. The approach by Hoyer and colleagues was slightly superior in most scenarios and is recommended in practice.
Zusammenfassung In diesem Bericht fassen wir die Ergebnisse eines systematischen Reviews (SR) zusammen, in dem Daten zur Wirksamkeit und Verträglichkeit von ACTH (adrenocorticotropes Hormon) und Kortikosteroiden (KST) bei Kindern mit anderen Epilepsien als dem infantilen epileptischen Spasmussyndrom (IESS) ausgewertet wurden, die auf Anfallssuppressiva (ASM) nicht angesprochen hatten. Der SR umfasste retrospektive und prospektive Studien, die über mehr als 5 Patienten berichteten und klare Falldefinitionen sowie Beschreibungen der Behandlung und der Ergebnisse enthielten. Achtunddreißig (2 kontrollierte und 5 unkontrollierte prospektive, 31 retrospektive) Studien mit 1152 Patienten wurden eingeschlossen. Die Metaanalyse der aggregierten Daten zur Anfallsreduktion > 50 % und zur Verringerung der EEG(Elektroenzephalogramm)-Spikes am Ende der Initialbehandlung ergab gepoolte Raten (PR) von 0,60 (95 %-KI [Konfidenzintervall] 0,52–0,67) und 0,56 (95 %-KI 0,43–0,68). Die Rückfallquote war hoch (PR 0,33, 95 %-KI 0,27–0,40). Subgruppenanalysen und eine Metaregression zeigten keinen signifikanten Unterschied zwischen den eingesetzten Substanzen, eine etwas bessere Wirkung bei entwicklungsbedingter und/oder epileptischer Enzephalopathie mit Spike-and-Wave-Aktivierung im Schlaf (DEE-SWAS) und eine schwächere Wirkung bei Patienten mit kognitiver Beeinträchtigung und „symptomatischer“ Ätiologie. Die Höhe der kumulativen Dosis der initialen Behandlungsphase hatte keinen Einfluss auf die Behandlungsergebnisse. Adipositas und Cushing-Syndrom waren die häufigsten unerwünschten Wirkungen, die oft in Studien mit kontinuierlicher ACTH- (PR 0,73, 95 %-KI 0,48–0,89) oder KST-Gabe (PR 0,72, 95 %-KI 0,54–0,85), aber selten bei intermittierender intravenöser oder oraler KST-Gabe (PR 0,05, 95 %-KI 0,02–0,10) auftraten. Die Aussagekraft dieser Ergebnisse wird durch ein hohes Verzerrungsrisiko der meisten eingeschlossenen Studien und eine große Heterogenität zwischen den Studiendaten eingeschränkt. Der volle SR wurde unter https://doi.org/10.1111/epi.17918 publiziert.
The placebo effect is the 'effect of the simulation of treatment that occurs due to a participant's belief or expectation that a treatment is effective'. Although the effect might be of little importance for some conditions, it can have a great role in others, mostly when the evaluated symptoms are subjective. Several characteristics that include informed consent, number of arms in a study, the occurrence of adverse events and quality of blinding may influence response to placebo and possibly bias the results of randomised controlled trials. Such a bias is inherited in systematic reviews of evidence and their quantitative components, pairwise meta-analysis (when two treatments are compared) and network meta-analysis (when more than two treatments are compared). In this paper, we aim to provide red flags as to when a placebo effect is likely to bias pairwise and network meta-analysis treatment effects. The classic paradigm has been that placebo-controlled randomised trials are focused on estimating the treatment effect. However, the magnitude of placebo effect itself may also in some instances be of interest and has also lately received attention. We use component network meta-analysis to estimate placebo effects. We apply these methods to a published network meta-analysis, examining the relative effectiveness of four psychotherapies and four control treatments for depression in 123 studies.
We aimed to assess the performance of Ag-RDT and RT-qPCR with regard to detecting infectious SARS-CoV-2 in cell cultures, as their diagnostic test accuracy (DTA) compared to virus isolation remains largely unknown. We searched three databases up to 15 December 2021 for DTA studies. The bivariate model was used to synthesise the estimates. Risk of bias was assessed using QUADAS-2/C. Twenty studies (2605 respiratory samples) using cell culture and at least one molecular test were identified. All studies were at high or unclear risk of bias in at least one domain. Three comparative DTA studies reported results on Ag-RDT and RT-qPCR against cell culture. Two studies evaluated RT-qPCR against cell culture only. Fifteen studies evaluated Ag-RDT against cell culture as reference standard in RT-qPCR-positive samples. For Ag-RDT, summary sensitivity was 93% (95% CI 78; 98%) and specificity 87% (95% CI 70; 95%). For RT-qPCR, summary sensitivity (continuity-corrected) was 98% (95% CI 95; 99%) and specificity 45% (95% CI 28; 63%). In studies relying on RT-qPCR-positive subsamples (n = 15), the summary sensitivity of Ag-RDT was 93% (95% CI 92; 93%) and specificity 63% (95% CI 63; 63%). Ag-RDT show moderately high sensitivity, detecting most but not all samples demonstrated to be infectious based on virus isolation. Although RT-qPCR exhibits high sensitivity across studies, its low specificity to indicate infectivity raises the question of its general superiority in all clinical settings. Study findings should be interpreted with caution due to the risk of bias, heterogeneity and the imperfect reference standard for infectivity.
Network meta-analysis compares different interventions for the same condition, by combining direct and indirect evidence derived from all eligible studies. Network metaanalysis has been increasingly used by applied scientists and it is a major research topic for methodologists. This article describes the R package netmeta, which adopts frequentist methods to fit network meta-analysis models. We provide a roadmap to perform network meta-analysis, along with an overview of the main functions of the package. We present three worked examples considering different types of outcomes and different data formats to facilitate researchers aiming to conduct network meta-analysis with netmeta.
INTRODUCTION: Shivering is a common side effect after general anesthesia. Risk factors are hypothermia, young age and postoperative pain. Severe complications of shivering are rare but can occur due to increased oxygen consumption. Previous systematic reviews are outdated and have summarized the evidence on the topic using only pairwise comparisons. The objective of this manuscript was a quantitative synthesis of evidence on pharmacological interventions to treat postanesthetic shivering. EVIDENCE ACQUSITION: Systematic review and frequentist network meta-analysis using the R package netmeta. Endpoints were the risk ratio (RR) of persistent shivering at one, five and 10 minutes after treatment with saline/placebo as the comparator. Data were retrieved from Medline, Embase, Central and Web of Science up to January 2022. Eligibility criteria were: randomized, controlled, and blinded trials comparing pharmacological interventions to treat shivering after general anesthesia. Studies on shivering during or after any type of regional anesthesia were excluded as well as sedated patients after cardiac surgery.EVIDENCE SYNTHESIS: Thirty-two trials were eligible for data synthesis, including 28 pharmacological interventions. The largest network included 1431 patients. The network geometry was two-centered with most comparisons linked to saline/placebo or pethidine. The best interventions were after one minute: doxapram 2 mg/kg, tramadol 2 mg/kg and nefopam 10 mg, after 5 minutes: tramadol 2 mg/kg, nefopam 10 mg and clonidine 150 mu g and after 10 minutes: nefopam 10 mg, methylphenidate 20 mg and tramadol 1 mg/kg, all reaching statistical significance. Pethidine 25 mg and clonidine 75 mu g also performed well and with statistical significance in all networks.CONCLUSIONS: Nefopam, tramadol, pethidine and clonidine are the most effective treatments to stop postanesthetic shivering. The efficacy of doxapram is uncertain since different doses showed contradictory effects and the evidence for methylphenidate is based on a single comparison in only one network. Furthermore, both lack data on side effects. Further studies are needed to clarify the efficacy of dexmedetomidine to treat postanesthetic shivering.(Cite this article as: Dinges HC, Al-Dahna T, Rucker G, Wulf H, Eberhart L, Wiesmann T, et al. Pharmacologic interventions for the therapy of postanesthetic shivering in adults: a systematic review and network meta-analysis. Minerva Anestesiol 2023;89:923-35. DOI: 10.23736/S0375-9393.23.17410-4)
Hierarchical models are recommended for meta-analysis of test accuracy studies. Hierarchical models such as the bivariate and hierarchical summary receiver operating characteristic models are recommended for test accuracy meta-analysis. The models can be fitted in a frequentist or Bayesian statistics framework. This chapter presents analyses performed using both frequentist and Bayesian approaches wherever possible. It provides an overview of how to fit the hierarchical models within a frequentist framework using three software packages (SAS, Stata and R) and within a Bayesian framework using rjags. The chapter also provides suggestions on meta-analyses of problematic or atypical data sets, including simplifying hierarchical models for meta-analysis of sparse data. Using a Cochrane Review of diagnostic test accuracy as an example, it then illustrates latent class meta-analysis and concludes with a summary and information on additional resources.