Network meta‐analysis compares multiple treatments from studies that form a connected network of evidence. However, for complex networks, it is not easy to see if the network is connected. We use simple techniques from graph theory to test the connectedness of evidence networks in network meta‐analysis. The method is to build the adjacency matrix for a network, with rows and columns corresponding to the treatments in the network and entries being one or zero depending on whether the treatments have been compared or not, and with zeros along the diagonal. Manipulation of this matrix gives the indirect connection matrix. The entries of this matrix determine whether two treatments can be compared, directly or indirectly. We also describe the distance matrix, which gives the minimum number of steps in the network required to compare a pair of treatments. This is a useful assessment of an indirect comparison as each additional step requires further assumptions of homogeneity in, for example, design and target populations of included trials. If there are no loops in the network, the distance is a measure of the degree of assumptions needed; it is approximately this with loops. We illustrate our methods using several constructed examples and giving R code for computation. We have also implemented the techniques in the Stata package “network.” The methods provide a fast way to ensure comparisons are only made between connected treatments and to assess the degree of indirectness of a comparison.
OBJECTIVES:A new method is presented for both synthesizing treatment effects on multiple outcomes subject to measurement error and estimating coherent mapping coefficients between all outcomes. It can be applied to sets of trials reporting different combinations of patient- or clinician-reported outcomes, including both disease-specific measures and generic health-related quality-of-life measures. It is underpinned by a structural equation model that includes measurement error and latent common treatment effect factor. Treatment effects can be expressed on any of the test instruments that have been used.METHODS:This is illustrated in a synthesis of eight placebo-controlled trials of TNF-α inhibitors in ankylosing spondylitis, each reporting treatment effects on between two and five of a total six test instruments.RESULTS:The method has advantages over other methods for synthesis of multiple outcome data, including standardization and multivariate normal synthesis. Unlike standardization, it allows synthesis of treatment effect information from test instruments sensitive to different underlying constructs. It represents a special case of previously proposed multivariate normal models for evidence synthesis, but unlike the former, it also estimates mappings. Combining synthesis and mapping as a single operation makes more efficient use of available data than do current mapping methods and generates treatment effects that are consistent with the mappings. A limitation, however, is that it can only generate mappings to and from those instruments on which some trial data exist.CONCLUSIONS:The method should be assessed in a wide range of data sets on different clinical conditions, before it can be used routinely in health technology assessment.
Develop a method to quickly test whether a network meta-analysis evidence network is connected. BACKGROUND Network meta-analysis, or mixed treatment comparisons, is a method to combine evidence on multiple treatments that have been compared in randomised controlled trials that form a connected network of treatment comparisons. Evidence networks consist of nodes, representing treatments, and edges, representing clinical trials comparing two treatments. If nodes corresponding to treatments are not connected, they cannot be compared. Connectedness is typically tested by visual inspection, however this is time consuming when there are many separate networks representing different outcomes, subgroups, and scenarios, and also prone to error, especially in large networks . Path finding algorithms can be used to automate testing for connectedness, but these are slow and inefficient. We present a fast and simple approach to test connectedness. Our method constructs a symmetric square matrix, called the direct connection matrix, with the number of rows and columns equal to the number of treatments in the network. We fill this matrix with ones where treatments of the corresponding row and column have been compared in a trial, and zeros otherwise. The diagonal is filled with ones. Exponentiation of the matrix to the number of treatments, minus one, gives the indirect connection matrix. Non-zero entries of this final matrix represent treatment combinations that can be compared using available evidence, and vice versa. This test is easy to implement in software and can be conducted rapidly. We prove the validity of the method mathematically and illustrate with application to a network of anticoagulants for the prevention of stroke in atrial fibrillation. We have developed a simple and rapid test of connectedness of networks that is easy to automate and can be applied to any network meta-analysis.
The primary outcomes in trials are usually disease-specific measures (DSMs) designed to be responsive to changes in the condition caused by treatment. For purposes of cost-effectiveness analysis, treatment effects on the DSM are often "mapped" into treatment effects on a generic health-related quality-of-life (QOL) scale, such as EuroQol five-dimensional questionnaire. Trialists have the option of including generic QOL measures as trial outcomes. We consider the relative efficiency (estimate divided by its standard error) of treatment effects derived from the DSM, the generic QOL, the generic QOL indirectly estimated from the mapped DSM, and a pooled estimate combining the direct and indirect information on the generic QOL. By using a "common factor" theory of the relationship between the DSM and the generic QOL, we define the circumstances under which indirectly estimated generic QOL is more efficient than the direct one and when a pooled QOL estimate is more efficient than the DSM estimate. As long as the DSM is more responsive, there is always a threshold sample size above which the indirect estimate has better precision than the direct estimate. This threshold, however, increases as the (1) relative responsiveness ratio of the DSM to the generic QOL increases, (2) precision of the estimated mapping coefficient increases, and (3) true effect becomes smaller. The pooled estimate on the generic QOL may be more efficient than the DSM itself unless the reliability of the DSM is particularly high. Trials powered on DSMs are likely to have sufficient power to detect treatment effect on the generic QOL if a pooled estimate is used. We conclude that generic QOL instruments should be routinely included in randomized controlled trials. Information on mapping coefficients and on relative responsiveness should be collected more systematically to facilitate both evidence synthesis and trial design.
Inconsistency can be thought of as a conflict between "direct" evidence on a comparison between treatments B and C and "indirect" evidence gained from AC and AB trials. Like heterogeneity, inconsistency is caused by effect modifiers and specifically by an imbalance in the distribution of effect modifiers in the direct and indirect evidence. Defining inconsistency as a property of loops of evidence, the relation between inconsistency and heterogeneity and the difficulties created by multiarm trials are described. We set out an approach to assessing consistency in 3-treatment triangular networks and in larger circuit structures, its extension to certain special structures in which independent tests for inconsistencies can be created, and describe methods suitable for more complex networks. Sample WinBUGS code is given in an appendix. Steps that can be taken to minimize the risk of drawing incorrect conclusions from indirect comparisons and network meta-analysis are the same steps that will minimize heterogeneity in pairwise meta-analysis. Empirical indicators that can provide reassurance and the question of how to respond to inconsistency are also discussed.
Inconsistency can be thought of as a conflict between “direct” evidence on a comparison between treatments B and C and “indirect” evidence gained from AC and AB trials. Like heterogeneity, inconsistency is caused by effect modifiers and specifically by an imbalance in the distribution of effect modifiers in the direct and indirect evidence. Defining inconsistency as a property of loops of evidence, the relation between inconsistency and heterogeneity and the difficulties created by multiarm trials are described. We set out an approach to assessing consistency in 3-treatment triangular networks and in larger circuit structures, its extension to certain special structures in which independent tests for inconsistencies can be created, and describe methods suitable for more complex networks. Sample WinBUGS code is given in an appendix. Steps that can be taken to minimize the risk of drawing incorrect conclusions from indirect comparisons and network meta-analysis are the same steps that will minimize heterogeneity in pairwise meta-analysis. Empirical indicators that can provide reassurance and the question of how to respond to inconsistency are also discussed. Keywords Network meta-analysis , inconsistency , indirect evidence , Bayesian
OBJECTIVES:To develop a coherent method for estimating mappings between treatment effects on disease-specific measurement (DSM) instruments and generic health-related quality-of-life (QOL) measures, when both are subject to measurement errors.METHODS:We identified three properties that must be satisfied for mappings to be logically coherent: invertability, transitivity, and invariance to linear transformation. Of the common regressions, ordinary least squares (OLS), geometric mean (GM), and orthogonal regression, only GM has all these properties, and then only in special cases. We developed a common factor model of how DSM and generic QOL scales are related, and derived expressions for coherent mapping coefficients. We showed that these are equivalent to adjusted forms of OLS or GM regressions. Where cohort data are available on just one DSM and one QOL measure, external data on the reproducibility of the DSM are required. In some circumstances, the mappings can be estimated without external data. We illustrated the estimation of mapping coefficients by using data on EuroQol five-dimensional (EQ-5D) questionnaire, 12-item short form health survey (SF-12) Mental Component Summary, and the Beck Depression Inventory (BDI), from a trial of treatments for depression.RESULTS:OLS underestimates and GM overestimates mappings from DSMs to generic QOL measures. Mappings estimated by using external data on reliability were similar to those estimated by using internal data, suggesting approximate adequacy of the common factor model.CONCLUSIONS:Neither OLS nor GM regression, unless corrected, is suitable for estimating mappings between disease-specific and generic QOL scales. OLS systematically underestimates mappings, but it can be adjusted by using external information on test-retest reliability.
Baseline risk is a proxy for unmeasured but important patient‐level characteristics, which may be modifiers of treatment effect, and is a potential source of heterogeneity in meta‐analysis. Models adjusting for baseline risk have been developed for pairwise meta‐analysis using the observed event rate in the placebo arm and taking into account the measurement error in the covariate to ensure that an unbiased estimate of the relationship is obtained. Our objective is to extend these methods to network meta‐analysis where it is of interest to adjust for baseline imbalances in the non‐intervention group event rate to reduce both heterogeneity and possibly inconsistency. This objective is complicated in network meta‐analysis by this covariate being sometimes missing, because of the fact that not all studies in a network may have a non‐active intervention arm. A random‐effects meta‐regression model allowing for inclusion of multi‐arm trials and trials without a ‘non‐intervention’ arm is developed. Analyses are conducted within a Bayesian framework using the WinBUGS software. The method is illustrated using two examples: (i) interventions to promote functional smoke alarm ownership by households with children and (ii) analgesics to reduce post‐operative morphine consumption following a major surgery. The results showed no evidence of baseline effect in the smoke alarm example, but the analgesics example shows that the adjustment can greatly reduce heterogeneity and improve overall model fit. Copyright © 2012 John Wiley & Sons, Ltd.
Mixed treatment comparison (MTC) (also called network meta‐analysis) is an extension of traditional meta‐analysis to allow the simultaneous pooling of data from clinical trials comparing more than two treatment options. Typically, MTCs are performed using general‐purpose Markov chain Monte Carlo software such as WinBUGS, requiring a model and data to be specified using a specific syntax. It would be preferable if, for the most common cases, both could be derived from a well‐structured data file that can be easily checked for errors. Automation is particularly valuable for simulation studies in which the large number of MTCs that have to be estimated may preclude manual model specification and analysis. Moreover, automated model generation raises issues that provide additional insight into the nature of MTC. We present a method for the automated generation of Bayesian homogeneous variance random effects consistency models, including the choice of basic parameters and trial baselines, priors, and starting values for the Markov chain(s). We validate our method against the results of five published MTCs. The method is implemented in freely available open source software. This means that performing an MTC no longer requires manually writing a statistical model. This reduces time and effort, and facilitates error checking of the dataset. Copyright © 2012 John Wiley & Sons, Ltd.
Meta‐analyses that simultaneously compare multiple treatments (usually referred to as network meta‐analyses or mixed treatment comparisons) are becoming increasingly common. An important component of a network meta‐analysis is an assessment of the extent to which different sources of evidence are compatible, both substantively and statistically. A simple indirect comparison may be confounded if the studies involving one of the treatments of interest are fundamentally different from the studies involving the other treatment of interest. Here, we discuss methods for addressing inconsistency of evidence from comparative studies of different treatments. We define and review basic concepts of heterogeneity and inconsistency, and attempt to introduce a distinction between ‘loop inconsistency’ and ‘design inconsistency’. We then propose that the notion of design‐by‐treatment interaction provides a useful general framework for investigating inconsistency. In particular, using design‐by‐treatment interactions successfully addresses complications that arise from the presence of multi‐arm trials in an evidence network. We show how the inconsistency model proposed by Lu and Ades is a restricted version of our full design‐by‐treatment interaction model and that there may be several distinct Lu–Ades models for any particular data set. We introduce novel graphical methods for depicting networks of evidence, clearly depicting multi‐arm trials and illustrating where there is potential for inconsistency to arise. We apply various inconsistency models to data from trials of different comparisons among four smoking cessation interventions and show that models seeking to address loop inconsistency alone can run into problems. Copyright © 2012 John Wiley & Sons, Ltd.
To the Editor – In our article, “Mapping from disease-specific to generic health-related quality of life scales: a common factor model” [1Lu G. Ades A.E. Brazier J. Mapping from disease-specific to generic health-related quality of life scales: a common factor model.Value Health. 2012; : 15Abstract Full Text Full Text PDF Google Scholar], we propose a method for mapping mean treatment effects reported in trials on one scale into mean treatment effects on another scale. Our approach is based on a structural equation model that, in its simplest form, partitions the variances of responses to all scales into two components. One component is that part of the test that responds to treatment, and the other component is the remaining variance. We show that the true mapping coefficient is the signed square root of the ratio of the variances of the first component. We also point out that the ordinary least squares (OLS) regression coefficient will invariably underestimate the true mapping, because of the measurement error inherent in the test instruments. One of the motivations of our approach, which is not mentioned by Palta, is that mappings between mean treatment effects should be both invertible and transitive. In other words, suppose we have estimated a mapping βˆY→Q from disease-specific instrument Y to generic scale Q, then when we observe an estimated treatment effect δˆY in a trial, we would predict an effect δˆQ=βˆY→QδˆY on the generic scale Q. Note that δˆY is usually an unbiased and consistent estimate for the true treatment effect. On the other hand, by the same token, if we observe δˆQ (also unbiased and consistent) in a trial, we would predict δˆY=δˆQ/βˆY→Q on scale Y. In other words, βˆQ→Y=1/βˆY→Q. We show that mappings defined our way have this property. They must also be transitive, which means that a mapping from X to Z must be the product of a mapping from X to Y and a mapping from Y to Z. In her commentary on our article, Palta [2Palta M. Some comments on mapping from disease-specific to generic quality of life scales.Value Health. 2012; : 15Google Scholar] takes issue with our approach on two grounds. First, she claims that the mean treatment effect is still “prone to measurement error,” and therefore that the “correct conversion is [still] the true mapping multiplied by the reliability of the DSM.” She goes on to assert that the correct mapping coefficient to map from X to Y, for either an individual score or an estimated mean score, is the OLS regression, which is the “true” mapping that would be obtained if there were no measurement error in X, multiplied by the reliability of X. In our notation, her proposed mapping coefficient isβX→YOLS=βX→YρX It is easy to see that this will end in a contradiction. The mean treatment effects on scales X and Y have variances that depend on sample size, n in each arm, the variance of the true scores σX2, and its reliability:Var(δˆX)=2τX2n=2σX2nρXandVar(δˆY)=2τY2n=2σY2nρY Under Palta’s proposed mapping, δˆY=βX→YOLSδˆX=βX→YρXδˆX, we would then obtainVar(δˆY)=(βX→YOLS)2Var(δˆX)=(βX→Y)2ρX22σX2n=2σY2nρY This gives us that βX→Y2=σY2/σX2ρX2ρY. But by parity of argument, we can also obtain that βY→X2=σX2σY2ρXρY2 Given that βX→Y=1/βY→X, we end up withσY2σX2ρX2ρY=σY2ρXρY2σX2orσX2σY2=σX2σY2ρX3ρY3which is true only when both reliabilities are 1 (clearly σX2,σY2 are not equal to zero). The same argument can be made in a less technical and perhaps more intuitive way. Imagine a trial of infinite size in which the treatment effect is examined on three test instruments X, Y, and Z. The true mappings are—by definition—the ratios of the treatment effects δX,δY,δZ on each scale. These ratios obviously do have, and must have, the properties of transitivity and invertability. If we imagine, instead, a trial of finite size, the ratios of the estimates must still represent estimates of the mappings, which are still transitive and invertible. But the OLS mappings proposed by Palta, even if they were all estimated from a single (and infinite sized) cohort study, can never have these properties and must always underestimate the correct mappings. Indeed, with her method, if one mapped from X to Y, then back to X, one would not end up where one started. Similarly, one could map from X to Y, and then from Y to Z, but this would end up with a different estimate than mapping from X to Z, even if one had used data from a single cohort study with observations on all three instruments. The second criticism is that the assumptions made by the common factor model may not necessarily be correct for every test, and Palta gives an example where this appears to be the case. We would accept entirely that the common factor model—at least as we have presented it—may be inadequate for some data sets. This is noted in Lu and Brazier [1Lu G. Ades A.E. Brazier J. Mapping from disease-specific to generic health-related quality of life scales: a common factor model.Value Health. 2012; : 15Abstract Full Text Full Text PDF Google Scholar] where possible extensions are suggested. This, however, does not change the fundamental point that when observations are made under measurement error, mappings between mean treatment effects based on OLS regression are incorrect and invariably underestimate the correct mapping. As a result, estimates of the quality of life gain due to treatment based on such methods are invariably underestimates. Mapping from Disease-Specific to Generic Health-Related Quality-of-Life Scales: A Common Factor ModelValue in HealthVol. 16Issue 1PreviewTo develop a coherent method for estimating mappings between treatment effects on disease-specific measurement (DSM) instruments and generic health-related quality-of-life (QOL) measures, when both are subject to measurement errors. Full-Text PDF Open Archive
Mixed treatment comparisons (MTC) meta-analysis synthesises comparative evidence on multiple treatments or other interventions from a collection of randomised controlled trials (RCT) available in a research area, while still respecting the randomisation structure in RCTs. This paper sets out to examine the properties of MTC estimates and elucidate the concept of consistency between direct and indirect evidence in MTC networks. We decompose MTC synthesis into two stages. At the first stage, ordinary meta-analysis is performed in each group of trials that have the same treatment comparators—this provides the ‘direct’ estimates of relative effect parameters. At the second stage, the optimal consistent estimates that minimise the distance between the direct estimates and the consistency hyper-plane can be deduced as the weighted least squares solution to a linear regression model with a specific design matrix that represents the consistency conditions. The consistent MTC estimates can then be represented explicitly as linear combinations of direct estimates, and under normality assumptions the overall evidence consistency can be tested with a likelihood-ratio statistic. This two-stage framework further allows us to use the leverage statistics to diagnose influence of the first-stage evidence and use the regression residuals to assess local inconsistency. The method is illustrated with two examples from medical research. Copyright © 2011 John Wiley & Sons, Ltd.
In this document we describe methods to detect inconsistency in a network meta-analysis. Inconsistency can be thought of as a conflict between “direct” evidence on a comparison between treatments B and C, and “indirect” evidence gained from AC and AB trials. Like heterogeneity, inconsistency is caused by effect-modifiers, and specifically by an imbalance in the distribution of effect modifiers in the direct and indirect evidence. Checking for inconsistency therefore logically comes alongside a consideration of the extent of heterogeneity and its sources, and the possibility of adjustment by meta-regression or bias adjustment (see TSD3). We emphasise that while tests for inconsistency must be carried out, they are inherently underpowered, and will often fail to detect it. Investigators must therefore also ask whether, if inconsistency is not detected, conclusions from combining direct and indirect evidence can be relied upon.
A range of procedures in both robustness and diagnostics require optimisation of a target functional over all subsamples of given size. Whereas such combinatorial problems are extremely difficult to solve exactly, something less than the global optimum can be ‘good enough’ for many practical purposes, as shown by example. Again, a relaxation strategy embeds these discrete, high-dimensional problems in continuous, low-dimensional ones. Overall, nonlinear optimisation methods can be exploited to provide a single, reasonably fast algorithm to handle a wide variety of problems of this kind, thereby providing a certain unity. Four running examples illustrate the approach. On the robustness side, algorithmic approximations to minimum covariance determinant (MCD) and least trimmed squares (LTS) estimation. And, on the diagnostic side, detection of multiple multivariate outliers and global diagnostic use of the likelihood displacement function. This last is developed here as a global complement to Cook’s (in J. R. Stat. Soc. 48:133–169, 1986) local analysis. Appropriate convergence of each branch of the algorithm is guaranteed for any target functional whose relaxed form is—in a natural generalisation of concavity, introduced here—‘gravitational’. Again, its descent strategy can downweight to zero contaminating cases in the starting position. A simulation study shows that, although not optimised for the LTS problem, our general algorithm holds its own with algorithms that are so optimised. An adapted algorithm relaxes the gravitational condition itself.
In mixed treatment comparison (MTC) meta-analysis, modeling the heterogeneity in between-trial variances across studies is a difficult problem because of the constraints on the variances inherited from the MTC structure. Starting from a consistent Bayesian hierarchical model for the mean treatment effects, we represent the variance configuration by a set of triangle inequalities on the standard deviations. We take the separation strategy (Barnard and others, 2000) to specify prior distributions for standard deviations and correlations separately. The covariance matrix of the latent treatment arm effects can be employed as a vehicle to load the triangular constraints, which in addition allows incorporation of prior beliefs about the correlations between treatment effects. The spherical parameterization based on Cholesky decomposition (Pinheiro and Bates, 1996) is used to generate a positive-definite matrix for the prior correlations in Markov chain Monte Carlo (MCMC). Elicited prior information on correlations between treatment arms is introduced in the form of its equivalent data likelihood. The procedure is implemented in a MCMC framework and illustrated with example data sets from medical research practice.
We present a mixed treatment meta-analysis of antivirals for treatment of influenza, where some trials report summary measures on at least one of the two outcomes: time to alleviation of fever and time to alleviation of symptoms. The synthesis is further complicated by the variety of summary measures reported: mean time, median time and proportion symptom free at the end of follow-up. We compare several models using the deviance information criteria and the contribution of different evidence sources to the residual deviance to aid model selection. A Weibull model with exchangeable treatment effects that are independent for each outcome but have a common random effect mean for the two outcomes gives the best fit according to these criteria. This model allows us to summarize treatment effect on two outcomes in a single summary measure and draw conclusions as to the most effective treatment. Amantadine and Oseltamivir were the most effective treatments, with the probability of being most effective of 0.56 and 0.37, respectively. Amantadine reduces the duration of symptoms by an estimated 2.8 days, and Oseltamivir 2.6 days, compared with placebo. The models provide flexible methods for synthesis of evidence on multiple treatments in the absence of head-to-head trial data, when different summary measures are used and either different clinical outcomes are reported or where the same outcomes are reported at different or multiple time points.
Meta-analysis has been well-established for many years, but has been largely confined to pooling evidence on pair-wise contrasts. Broader forms of synthesis have also been described, apparently re-invented in disparate fields, each time taking different computational approaches. The potential value of Bayesian estimation of a joint posterior parameter distribution and simultaneously sampling from it for decision analysis has also been appreciated. However, applications have been relatively few in number, sometimes stylized, and presented mainly to a statistical methods audience. As a result, the potential for multiparameter evidence synthesis in both epidemiology and health technology assessment has remained largely unrecognized. The advent of flexible software for Bayesian Markov chain Monte Carlo in the shape of WinBUGS has the made these earlier strands of work more widely available. Researchers can now carry out synthesis at a realistic level of complexity. The Bristol programme has not only contributed to a growing body of literature on how to synthesize different evidence structures, but also on how to check the consistency of multiple information sources and how to use the resulting models to prioritize future research.
Mixed treatment comparisons (MTC) meta-analysis is a methodology for making inferences on relative treatment effects based on a synthesis of both direct and indirect evidence on multiple treatment contrasts. This is particularly useful in the context of cost-effectiveness analysis and medical decision making. Here, we extend these methods to a more complex situation where trials report results at one or more, different yet fixed, follow-up times. These methods are applied to an illustrative data set combining evidence on healing rates under six different treatments for gastro-esophageal reflux disease (GERD). A series of Bayesian hierarchical models based on piece-wise exponential hazards is developed that borrow strength across the MTC networks and also across time points. These include models for absolute and relative treatment effects, models with fixed or random effects over time, random walk models, and models with homogeneous or heterogeneous between-trials variation. The deviance information criterion (DIC) is used to guide model development and selection. Models for absolute treatment effects generate materially different rankings of the treatments than models that separate the trial-specific baselines from the relative treatment effects. The extent of between-trials heterogeneity in treatment effects depends on treatment contrast. In discussion we note that models of this type have a very wide potential application.
Recently, health systems internationally have begun to use cost-effectiveness research as formal inputs into decisions about which interventions and programmes should be funded from collective resources. This process has raised some important methodological questions for this area of research. This paper considers one set of issues related to the synthesis of effectiveness evidence for use in decision-analytic cost-effectiveness (CE) models, namely the need for the synthesis of all sources of available evidence, although these may not 'fit neatly' into a CE model.Commonly encountered problems include the absence of head-to-head trial evidence comparing all options under comparison, the presence of multiple endpoints from trials and different follow-up periods. Full evidence synthesis for CE analysis also needs to consider treatment effects between patient subpopulations and the use of nonrandomised evidence.Bayesian statistical methods represent a valuable set of analytical tools to utilise indirect evidence and can make a powerful contribution to the decision-analytic approach to CE analysis. This paper provides a worked example and a general overview of these methods with particular emphasis on their use in economic evaluation.
Randomized comparisons among several treatments give rise to an incomplete-blocks structure known as mixed treatment comparisons (MTCs). To analyze such data structures, it is crucial to assess whether the disparate evidence sources provide consistent information about the treatment contrasts. In this article we propose a general method for assessing evidence inconsistency in the framework of Bayesian hierarchical models. We begin with the distinction between basic parameters, which have prior distributions, and functional parameters, which are defined in terms of basic parameters. Based on a graphical analysis of MTC structures, evidence inconsistency is defined as a relation between a functional parameter and at least two basic parameters, supported by at least three evidence sources. The inconsistency degrees of freedom (ICDF) is the number of such inconsistencies. We represent evidence consistency as a set of linear relations between effect parameters on the log odds ratio scale, then relax these relations to allow for inconsistency by adding to the model random inconsistency factors (ICFs). The number of ICFs is determined by the ICDF. The overall consistency between evidence sources can be assessed by comparing models with and without ICFs, whereas their posterior distribution reflects the extent of inconsistency in particular evidence cycles. The methods are elucidated using two published datasets, implemented with standard Markov chain Monte Carlo software.