Aims: The study is aimed to verify Aperio AT2 scanner for reporting on the digital pathology platform (DP) and to validate the cohort of pathologists in the interpretation of DP for routine diagnostic histopathological services in Wales, United Kingdom. Materials, Methods and Results: This was a large multicenter study involving seven hospitals across Wales and unique with 22 (largest number) pathologists participating. 7491 slides from 3001 cases were scanned on Leica Aperio AT2 scanner and reported on digital workstations with Leica software of e-slide manager. A senior pathology fellow compared DP reports with authorized reports on glass slide (GS). A panel of expert pathologists reviewed the discrepant cases under multiheader microscope to establish ground truth. 2745 out of 3001 (91%) cases showed complete concordance between DP and GS reports. Two hundred and fifty-six cases showed discrepancies in diagnosis, of which 170 (5.6%) were deemed of no clinical significance by the review panel. There were 86 (2.9%) clinically significant discrepancies in the diagnosis between DP and GS. The concordance was raised to 97.1% after discounting clinically insignificant discrepancies. Ground truth lay with DP in 28 out of 86 clinically significant discrepancies and with GS in 58 cases. Sensitivity of DP was 98.07% (confidence interval [CI] 97.57–98.56%); for GS was 99.07% (CI 98.72–99.41%). Conclusions: We concluded that Leica Aperio AT2 scanner produces adequate quality of images for routine histopathologic diagnosis. Pathologists were able to diagnose in DP with good concordance as with GS. Strengths and Limitations of this Study: Strengths of this study – This was a prospective blind study. Different pathologists reported digital and glass arms at different times giving an ambience of real-time reporting. There was standardized use of software and hardware across Wales. A strong managerial support from efficiency through the technology group was a key factor for the implementation of the study. Limitations: This study did not include Cytopathology and in situ hybridization slides. Difficulty in achieving surgical pathology practise standardization across the whole country contributed to intra-observer variations.
The adoption of agro-ecological practices in agricultural systems worldwide can contribute to increased food production without compromising future food security, especially under the current biodiversity loss and climate change scenarios. Despite the increase in publications on agro-ecological research and practices during the last 35 years, a weak link between that knowledge and changed farmer practices has led to few examples of agroecological protocols and effective delivery systems to agriculturalists. In an attempt to reduce this gap, we synthesised the main concepts related to biodiversity and its functions by creating a web-based interactive spiral (www. biodiversity function.com). This tool explains and describes a pathway for achieving agro-ecological outcomes, starting from the basic principle of biodiversity and its functions to enhanced biodiversity on farms. Within this pathway, 11 key steps are identified and sequentially presented on a web platform through which key players (farmers, farmer networks, policy makers, scientists and other stakeholders) can navigate and learn. Because in many areas of the world the necessary knowledge needed for achieving the adoption of particular agro-ecological techniques is not available, the spiral approach can provide the necessary conceptual steps needed for obtaining and understanding such knowledge by navigating through the interactive pathway. This novel approach aims to improve our understanding of the sequence from the concept of biodiversity to harnessing its power to improve prospects for 'sustainable intensification' of agricultural systems worldwide.
Key points•Traditional statistical methods were designed to demonstrate differences and cannot easily show that a new treatment is similar to an older one.•Non-inferiority can be shown if the difference between two treatments does not cross a predefined inferiority margin.•Non-inferiority studies need to be carefully planned; failings in the design of the study may make accepting an inferior treatment more likely.Learning objectivesBy reading this article, you should be able to:•Describe why traditional superiority trials cannot demonstrate equivalence.•Describe how a non-inferiority trial works.•Calculate the number of participants required for a simple non-inferiority trial, and analyse the trial result.•Understand and critically appraise a non-inferiority trial. •Traditional statistical methods were designed to demonstrate differences and cannot easily show that a new treatment is similar to an older one.•Non-inferiority can be shown if the difference between two treatments does not cross a predefined inferiority margin.•Non-inferiority studies need to be carefully planned; failings in the design of the study may make accepting an inferior treatment more likely. By reading this article, you should be able to:•Describe why traditional superiority trials cannot demonstrate equivalence.•Describe how a non-inferiority trial works.•Calculate the number of participants required for a simple non-inferiority trial, and analyse the trial result.•Understand and critically appraise a non-inferiority trial. Traditionally, much of medical research has involved finding differences between newer treatments and the previous standard of care, with the hope of showing that the newer treatment is better. Medical statistics has reflected this, with a focus on demonstrating a difference by rejecting the ‘null hypothesis’ that both treatments are similar (so-called ‘superiority studies’). Increasingly, however, new treatments are being developed that may not be better in clinical practice, but which offer advantages in terms of cost, ease of use, or adverse effects. In such a situation it would not be ethical to perform a trial comparing the new treatment to a placebo; instead, the newer drug is usually compared with the working treatment, which is referred to as an ‘active control’. This has led to a new type of statistical test that can demonstrate that two treatments are ‘similar’ to each other in terms of their clinical effectiveness. Although the statistics are straightforward and use familiar concepts, there are important differences in the way that these tests are designed and reported. Such tests can be divided into non-inferiority tests, which try to demonstrate that the new treatment is not worse than the old treatment, and equivalence tests, which attempt to demonstrate that the new treatment is neither better nor worse. As the usual objective is to identify that the new treatment is no worse than the current treatment (and investigators are usually very happy if it should turn out to be better than current treatment), equivalence studies are rare; however, for completeness, they are discussed at the end of this article. When a ‘traditional’ test does not demonstrate a difference between treatments, this is often presented (erroneously) as evidence of similarity, in spite of the fact that we have all been taught that ‘absence of evidence is not evidence of absence’. It may be that no difference exists, or it may be that the study was not of sufficient power to detect the difference between groups. This can be understood more readily by referring to Fig 1, in which a new treatment is compared with an older treatment (often referred to as an ‘active control’). Here we have represented three clinical trials as confidence intervals (CI), bearing in mind that mathematically, testing with CIs is the same as using P-values. In situation A (blue), the 95% CI does not cross the line of no difference and hence we can reject the null hypothesis. In situation B (green) there is no statistically significant difference: the 95% CI includes ‘no difference’, so we cannot reject the null hypothesis. However, this does not mean that we can assume that the two treatments are the same; the null hypothesis is merely one of the range of values that the difference could take. This is further compounded in situation C (red), where an underpowered study has led to a very wide CI, and hence a very large range of possible values for the difference. Ultimately, random samples cannot be used to show that two populations are identical; there will always be a CI representing a range of values, of which ‘identical’ is merely one possibility. The solution is to construct an ‘inferiority margin’ (often represented by the symbol d or dNI, although some authors use Δ). This margin represents the maximum reduction in effectiveness that you would be willing to accept while still considering the treatments to be equal. To illustrate this, imagine that you are considering changing your motor car, and you are worried about fuel efficiency. The newer model has many attractive ‘extras’ that are very tempting, and the salesman assures you that it has the same fuel consumption as your current model. If your current model achieves 40 miles per gallon (mpg), and the new model achieves 39.4 mpg, you may consider this to be close enough to make no difference. However, if the newer model only had a fuel consumption of 30 mpg, you would feel that the salesman had misled you. You might decide to place an inferiority margin at 38.5 mpg; any value less than 38.5 mpg will be considered inferior. With our defined inferiority margin in place we can go on to test for non-inferiority, using a similar process to a traditional superiority trial, but with the inferiority margin taking the place of ‘no difference’ as the null hypothesis. A P-value approach is possible; however, most trials choose to report CIs on the grounds that they are easier to interpret and do not risk confusion with a superiority trial. Possible outcomes are demonstrated in Fig 2. In situation A (blue) the 95% CI is entirely within the zone of non-inferiority, and we can therefore conclude that the new treatment is not inferior; in situation B (green) the 95% CI crosses the inferiority margin, and hence we cannot conclude that it is non-inferior. In case C (red) the entire CI is outside the non-inferiority zone, and in this case we can conclude that the new treatment is inferior. As the 95% CI also excludes the ‘no difference’ line, this trial would also demonstrate a statistically significant difference on traditional superiority testing (with the old treatment demonstrating superiority). As our inferiority margin has replaced ‘no difference’ as the null hypothesis, there are important differences in the error types and error rates when compared with a traditional superiority test. In a superiority test, a type 1 error means finding a significant difference, when no difference exists. In a non-inferiority trial a type 1 error means concluding non-inferiority when the new treatment is in fact inferior. A type 2 error (traditionally, failing to find a difference when a difference exists) now means that inferiority has been concluded in a treatment which is non-inferior. However, in both cases the error rates remain as in superiority studies: type 1 error rate is α (the significance level); type 2 error rate is β (1–power) (see Table 1).Table 1Status of the null hypothesis, error types, and error rates in superiority and non-inferiority studiesSuperiority studyNon-inferiority studyNull hypothesis (H0)Treatment = controlTreatment ≤ inferiority marginAlternate hypothesis (H1)Treatment ≠ controlTreatment > inferiority marginType 1 errorDeciding treatment ≠ control when no difference existsDeciding treatment non-inferior when it is inferiorType 2 errorDeciding treatment = control when a difference existsDeciding treatment inferior when it is non-inferiorType 1 error rateα (significance cut-off)α (significance cut-off)Type 2 error rateβ (1–power)β (1–power) Open table in a new tab A variety of methods exist to help choose the inferiority margin.1Schumi J. Wittes J.T. Through the looking glass: understanding non-inferiority.Trials. 2011; 12: 106Crossref PubMed Scopus (255) Google Scholar However, the figure chosen should be appropriate to clinical practice, and must be set and documented before the trial begins—both to ensure that the trial is fair, but also because the margin is required to calculate the sample size. It is important to ensure that the inferiority margin is set high enough to be better than placebo. For example, if we have an anti-emetic (‘drug A’) that we know works in 50% of patients, a manufacturer might suggest that when testing a new agent (‘drug B’) we would want it to work in at least 30% of patients to be considered non-inferior. This might not seem unreasonable until we discover that in the original trials of drug A, nausea was successfully treated by placebo in 34% of cases. Thus in allowing an inferiority margin of 30%, we would be willing to accept a new drug which may work less well than a placebo. In determining the margin, it may be necessary to perform a meta-analysis of the previous placebo trials. The full details are beyond the scope of this article, and the interested reader is directed to the article by Schumi and Wittes.1Schumi J. Wittes J.T. Through the looking glass: understanding non-inferiority.Trials. 2011; 12: 106Crossref PubMed Scopus (255) Google Scholar The process of calculating a sample size it not dissimilar to the methods used for a superiority study.2Julious S.A. Tutorial in biostatistics: sample sizes for clinical trials with Normal data.Statist Med. 2004; 23: 1921-1986Crossref PubMed Scopus (416) Google Scholar The researcher selects the significance (α) and power levels (1–β), and these need to be converted into their relevant z-values using a table (Table 2); we also require the standard deviation (σ), and the inferiority margin dNI. For normally distributed data, the number needed per arm will be:Numberperarm=2(Z1−β+Z1−α)2σ2((μnew−μcontrol)−dNI)2(1) Table 2z-Values for commonly used percentilesxz1–x0.2000.8420.1001.2820.0501.6450.0251.9600.0102.3260.0013.090 Open table in a new tab The expression (μnew–μcontrol) represents the expected difference between the two treatments. For a type 1 error rate of 2.5% and a power of 0.9, and where we have no evidence of a difference between treatments, this formula reduces to:Numberperarm=21×(σdNI)2,(2) which can easily be calculated by hand. As the test will be one-sided, the type 1 error rate is set at half what of is normally used for a two-sided test.3Ahn S. Park S.H. Lee K.H. How to demonstrate similarity by using noninferiority and equivalence testing in radiology research.Radiology. 2013; 267: 328-338Crossref PubMed Scopus (114) Google Scholar As mentioned above, it is possible to use a P-value approach to analyse the data, but CIs are far more intuitive. The CI we require is the interval around the mean difference between the outcome measures in the treatments, usually at the 95% (1–2α) confidence level (the lower level of which is the one-sided 97.5% CI). The formula is:(μnew−μcontrol)±1.96σnew2nnew+σcontrol2ncontrol,(3) where 1.96 is the appropriate z-value from Table 2. This can now be compared with our inferiority limit. If the interval remains above the inferiority limit, then non-inferiority has been demonstrated according to the standards we have set. Box 1 contains a worked example.Box 1Worked exampleA manufacturer develops a new, short-acting neuromuscular blocking agent which we will call ‘drug X’. The manufacturer believes that its profile makes it a suitable alternative to suxamethonium for rapid-sequence induction of anaesthesia. The researcher undertakes a meta-analysis of the literature and concludes that the mean onset time for suxamethonium is 52 (standard deviation, 8) s, and the manufacturer sets the inferiority margin at 5 s—indicating that we would accept the drug X taking 5 s longer without considering this to be a major disadvantage. They decide that a two-sided 95% confidence limit will be appropriate to compare the outcomes, and wish to have a power of 0.9. Using the briefer Formula (2), and taking dNI=5 and σ=8, the formula becomesNumberneededperarm=21×(85)2=53.76Thus, a suitable number per arm is 54 participants (not including dropouts). If the researcher had evidence to suggest that drug X was better, then Formula (1) could be used; depending on the difference in means this may well lead to a smaller required sample size, essentially because of the greater difference between drug X and the inferiority margin.The researcher decides to proceed, and obtains the following results: mean time to acceptable conditions for tracheal intubation is 56.7 (6.3) s for suxamethonium, and 58.8 (7.3) s for drug X, with 55 patients per arm. Putting these values into Formula (3) gives(58.8−56.7)±1.967.3255+6.3255=−0.448to2.648sAs this is well below the inferiority margin of 5 s, we can conclude that drug X is not inferior to succinylcholine with respect to time to tracheal intubation. A manufacturer develops a new, short-acting neuromuscular blocking agent which we will call ‘drug X’. The manufacturer believes that its profile makes it a suitable alternative to suxamethonium for rapid-sequence induction of anaesthesia. The researcher undertakes a meta-analysis of the literature and concludes that the mean onset time for suxamethonium is 52 (standard deviation, 8) s, and the manufacturer sets the inferiority margin at 5 s—indicating that we would accept the drug X taking 5 s longer without considering this to be a major disadvantage. They decide that a two-sided 95% confidence limit will be appropriate to compare the outcomes, and wish to have a power of 0.9. Using the briefer Formula (2), and taking dNI=5 and σ=8, the formula becomesNumberneededperarm=21×(85)2=53.76 Thus, a suitable number per arm is 54 participants (not including dropouts). If the researcher had evidence to suggest that drug X was better, then Formula (1) could be used; depending on the difference in means this may well lead to a smaller required sample size, essentially because of the greater difference between drug X and the inferiority margin. The researcher decides to proceed, and obtains the following results: mean time to acceptable conditions for tracheal intubation is 56.7 (6.3) s for suxamethonium, and 58.8 (7.3) s for drug X, with 55 patients per arm. Putting these values into Formula (3) gives(58.8−56.7)±1.967.3255+6.3255=−0.448to2.648s As this is well below the inferiority margin of 5 s, we can conclude that drug X is not inferior to succinylcholine with respect to time to tracheal intubation. Because of the way these trials are set up, a trial that is poorly designed or performed will be more likely to find non-inferiority and thus be ‘successful’. In this way, these trials reward poor research practice, and so it is important that the researcher demonstrates that all steps have been taken to ensure a fair trial—even more so than would be expected for a superiority trial. It is particularly important to be aware of a number of factors that have been referred to as the ‘ABC’ of non-inferiority trials: Assay sensitivity, Bias, and the assumption of Constancy.4Flight L. Julious S.A. Practical guide to sample size calculations: non-inferiority and equivalence trials.Pharm Stat. 2016; 15: 80-89Crossref PubMed Scopus (46) Google Scholar Assay sensitivity is the ability of the trial to detect a difference if it exists. As a trivial example, a broken blood pressure machine that always reads the same value would find no difference between patients taking an antihypertensive or a placebo; it is therefore not surprising that this equipment would find any new drug to be non-inferior to the antihypertensive. Many trials rely on surrogate outcome measures, which may appear reasonable but would not be able to detect a difference in any trial. Researchers often try to ensure assay sensitivity by using the same methodology which demonstrated a drug–placebo difference in earlier studies. However, this does rely on the constancy assumption (see below). Bias can be defined as a systematic tendency in a trial that will adversely influence the result. In most clinical trials we can avoid it by randomising and blinding, but in non-inferiority trials we need to be aware that bias can exist between the current trial and any previous placebo-controlled studies. If drug A is better than placebo, and we show that drug B is non-inferior to drug A, we are in effect comparing drug B with the placebo. Is that a fair comparison? We need to be satisfied that the non-inferiority trial has used a similar group of patients and drug dose to the original trials. Underdosing the older, established drug or running the trial in a group of patients where the condition might be easier to treat will give an unfair advantage to the new treatment. The constancy assumption is the requirement that the active control has the same effect now as it always had, and the methods of measuring the outcomes still work. Although this is usually the case, it is worth questioning whether situations have changed. Medicine is complex, and many conditions are now treated with lifestyle advice and medication, meaning that the effect that a particular drug may have had over placebo 40 yrs ago may not be the same as it does now. Of particular importance is the avoidance of ‘biocreep’.5Everson-Stewart S. Emerson S. Bio-creep in non-inferiority clinical trials.Statist Med. 2010; 29: 2769-2780Crossref PubMed Scopus (45) Google Scholar This occurs when non-inferiority comparisons are made on successive drugs, potentially leading to the acceptance of a drug which has no superiority over placebo. For example, if drug A is better than placebo, drug B is non-inferior to drug A by a small margin so drug B becomes standard treatment. Drug C is now compared with drug B, and this process continues, with the inferiority margin ‘creeping’ closer and closer to the placebo effect each time. The fear is that we eventually end up accepting a drug which does not work. RCTs are often analysed using ‘intention to treat’ (ITT) analysis, in which participants are counted in the group they were originally allocated to, even if they discontinued the treatment. There are multiple reasons for this, but one of the main ones is that it gives you a practical ‘real world’ view of a treatment. If a cholesterol-lowering medication treatment works in 95% of cases, but if its adverse-effect profile were so unacceptable that the majority of patients stop taking it, then the actual effect that this drug would have if prescribed would be minimal. The alternative to ITT is per protocol analysis, in which participants are compared on the basis of the treatment they actually received. A useful way to summarise the difference would be to say that ITT shows what happens when a treatment is prescribed, whereas per protocol shows what the treatment actually does when taken. The effect of ITT on superiority studies is to bring the two study arms closer together; hence, we can be more confident in any difference found—in effect, saying ‘we found a difference in spite of the participants who switched groups’. However, in a non-inferiority trial we wish to minimise factors that would make the two study arms seem artificially similar; hence, per protocol analysis is a more correct way to proceed.6Piaggio G. Elbourne D.R. Altman D.G. Pocock S.J. Evans S.J.W. Reporting of noninferiority and equivalence randomized trials.JAMA. 2006; 295: 1152-1160Crossref PubMed Scopus (1039) Google Scholar The most convincing results are those in which non-inferiority is found using both ITT and per protocol analyses.7D’Agostino R.B. Massaro J.M. Sullivan L.M. Non-inferiority trials: design concepts and issues – the encounters of academic consultants in statistics.Statist Med. 2003; 22: 169-186Crossref PubMed Scopus (527) Google Scholar As discussed at the beginning of this article, a superiority trial that fails to find a difference should not be used as evidence of equivalence. However, when a non-inferiority trial finds two treatments to be similar, it may be because they are the same, or it may be because the newer treatment is better. Indeed, once non-inferiority has been shown it is possible to perform a superiority study on the same data without the need for any statistical penalty to preserve the type 1 error rate. This is referred to as an ‘As good as, or better than’ trial. As with all statistical analyses, the statistical procedures and significance levels should be clearly decided and documented before the trial begins. Also the analysis should be per protocol for the non-inferiority analysis, but ITT for the subsequent superiority analysis. It is unusual to use equivalence trials in medicine—normally a treatment which has more than its required effect is considered a useful property. Equivalence can be shown using a ‘two one-sided test’ (TOST) procedure. This is a simple extension of the non-inferiority test described above, with a ‘superiority’ margin (+dNI) and the inferiority margin (–dNI). To conclude equivalence, the appropriate CI must lie completely within the range (–dNI, +dNI) (see Fig. 3). The author declares that they have no conflict of interest. The associated MCQs (to support CME/CPD activity) will be accessible at www.bjaed.org/cme/home by subscribers to BJA Education. Jason Walker FRCA FRSS BSc (Hons) Math Stat is a consultant anaesthetist at Ysbyty Gwynedd Hospital, Bangor, and an honorary senior lecturer at Bangor University. He is vice chair of his local research ethics committee and an examiner for the Primary FRCA examination.
Key points•Hypothesis tests are used to assess whether a difference between two samples represents a real difference between the populations from which the samples were taken.•A null hypothesis of ‘no difference’ is taken as a starting point, and we calculate the probability that both sets of data came from the same population. This probability is expressed as a p-value.•When the null hypothesis is false, p-values tend to be small. When the null hypothesis is true, any p-value is equally likely. •Hypothesis tests are used to assess whether a difference between two samples represents a real difference between the populations from which the samples were taken.•A null hypothesis of ‘no difference’ is taken as a starting point, and we calculate the probability that both sets of data came from the same population. This probability is expressed as a p-value.•When the null hypothesis is false, p-values tend to be small. When the null hypothesis is true, any p-value is equally likely. Learning objectivesBy reading this article, you should be able to:•Explain why hypothesis testing is used.•Use a table to determine which hypothesis test should be used for a particular situation.•Interpret a p-value. By reading this article, you should be able to:•Explain why hypothesis testing is used.•Use a table to determine which hypothesis test should be used for a particular situation.•Interpret a p-value. A hypothesis test is a procedure used in statistics to assess whether a particular viewpoint is likely to be true. They follow a strict protocol, and they generate a ‘p-value’, on the basis of which a decision is made about the truth of the hypothesis under investigation. All of the routine statistical ‘tests’ used in research—t-tests, χ2 tests, Mann–Whitney tests, etc.—are all hypothesis tests, and in spite of their differences they are all used in essentially the same way. But why do we use them at all? Comparing the heights of two individuals is easy: we can measure their height in a standardised way and compare them. When we want to compare the heights of two small well-defined groups (for example two groups of children), we need to use a summary statistic that we can calculate for each group. Such summaries (means, medians, etc.) form the basis of descriptive statistics, and are well described elsewhere.1McCluskey A. Lalkhen A.G. Statistics II: central tendency and spread of data.CEACCP. 2007; 7: 127-130Google Scholar However, a problem arises when we try to compare very large groups or populations: it may be impractical or even impossible to take a measurement from everyone in the population, and by the time you do so, the population itself will have changed. A similar problem arises when we try to describe the effects of drugs—for example by how much on average does a particular vasopressor increase MAP? To solve this problem, we use random samples to estimate values for populations. By convention, the values we calculate from samples are referred to as statistics and denoted by Latin letters (x¯ for sample mean; SD for sample standard deviation) while the unknown population values are called parameters, and denoted by Greek letters (μ for population mean, σ for population standard deviation). Inferential statistics describes the methods we use to estimate population parameters from random samples; how we can quantify the level of inaccuracy in a sample statistic; and how we can go on to use these estimates to compare populations. There are many reasons why a sample may give an inaccurate picture of the population it represents: it may be biased, it may not be big enough, and it may not be truly random. However, even if we have been careful to avoid these pitfalls, there is an inherent difference between the sample and the population at large. To illustrate this, let us imagine that the actual average height of males in London is 174 cm. If I were to sample 100 male Londoners and take a mean of their heights, I would be very unlikely to get exactly 174 cm. Furthermore, if somebody else were to perform the same exercise, it would be unlikely that they would get the same answer as I did. The sample mean is different each time it is taken, and the way it differs from the actual mean of the population is described by the standard error of the mean (standard error, or SEM). The standard error is larger if there is a lot of variation in the population, and becomes smaller as the sample size increases. It is calculated thus:SEM=SDnwhere SD is the sample standard deviation, and n is the sample size. As errors are normally distributed, we can use this to estimate a 95% confidence interval on our sample mean as follows:95%CI=x¯±(1.96×SEM) We can interpret this as meaning ‘We are 95% confident that the actual mean is within this range.’ Some confusion arises at this point between the SD and the standard error. The SD is a measure of variation in the sample. The range x¯±(1.96×SD) will normally contain 95% of all your data. It can be used to illustrate the spread of the data and shows what values are likely. In contrast, standard error tells you about the precision of the mean and is used to calculate confidence intervals. One straightforward way to compare two samples is to use confidence intervals. If we calculate the mean height of two groups and find that the 95% confidence intervals do not overlap, this can be taken as evidence of a difference between the two means. This method of statistical inference is reasonably intuitive and can be used in many situations.2Altman D.G. Machin D. Bryant T.N. Gardner M.J. Statistics with confidence.2nd Edn. BMJ Books, London2000Google Scholar Many journals, however, prefer to report inferential statistics using p-values. In 1925, the British statistician R.A. Fisher described a technique for comparing groups using a null hypothesis, a method which has dominated statistical comparison ever since. The technique itself is rather straightforward, but often gets lost in the mechanics of how it is done. To illustrate, imagine we want to compare the HR of two different groups of people. We take a random sample from each group, which we call our data. Then:(i)Assume that both samples came from the same group. This is our ‘null hypothesis’.(ii)Calculate the probability that an experiment would give us these data, assuming that the null hypothesis is true. We express this probability as a p-value, a number between 0 and 1, where 0 is ‘impossible’ and 1 is ‘certain’.(iii)If the probability of the data is low, we reject the null hypothesis and conclude that there must be a difference between the two groups. Formally, we can define a p-value as ‘the probability of finding the observed result or a more extreme result, if the null hypothesis were true.’ Standard practice is to set a cut-off at p <0.05 (this cut-off is termed the alpha value). If the null hypothesis were true, a result such as this would only occur 5% of the time or less; this in turn would indicate that the null hypothesis itself is unlikely. Fisher described the process as follows: ‘Set a low standard of significance at the 5 per cent point, and ignore entirely all results which fail to reach this level. A scientific fact should be regarded as experimentally established only if a properly designed experiment rarely fails to give this level of significance.’3Fisher R.A. The arrangement of field experiments.J Min Agric Gr Br. 1926; 33: 503-513Google Scholar This probably remains the most succinct description of the procedure. A question which often arises at this point is ‘Why do we use a null hypothesis?’ The simple answer is that it is easy: we can readily describe what we would expect of our data under a null hypothesis, we know how data would behave, and we can readily work out the probability of getting the result that we did. It therefore makes a very simple starting point for our probability assessment. All probabilities require a set of starting conditions, in much the same way that measuring the distance to London needs a starting point. The null hypothesis can be thought of as an easy place to put the start of your ruler. If a null hypothesis is rejected, an alternate hypothesis must be adopted in its place. The null and alternate hypotheses must be mutually exclusive, but must also between them describe all situations. If a null hypothesis is ‘no difference exists’ then the alternate should be simply ‘a difference exists’. The components of a hypothesis test can be readily described using the acronym GOST: identify the Groups you wish to compare; define the Outcome to be measured; collect and Summarise the data; then evaluate the likelihood of the null hypothesis, using a Test statistic. When considering groups, think first about how many. Is there just one group being compared against an audit standard, or are you comparing one group with another? Some studies may wish to compare more than two groups. Another situation may involve a single group measured at different points in time, for example before or after a particular treatment. In this situation each participant is compared with themselves, and this is often referred to as a ‘paired’ or a ‘repeated measures’ design. It is possible to combine these types of groups—for example a researcher may measure arterial BP on a number of different occasions in five different groups of patients. Such studies can be difficult, both to analyse and interpret. In other studies we may want to see how a continuous variable (such as age or height) affects the outcomes. These techniques involve regression analysis, and are beyond the scope of this article. The outcome measures are the data being collected. This may be a continuous measure, such as temperature or BMI, or it may be a categorical measure, such as ASA status or surgical specialty. Often, inexperienced researchers will strive to collect lots of outcome measures in an attempt to find something that differs between the groups of interest; if this is done, a ‘primary outcome measure’ should be identified before the research begins. In addition, the results of any hypothesis tests will need to be corrected for multiple measures. The summary and the test statistic will be defined by the type of data that have been collected. The test statistic is calculated then transformed into a p-value using tables or software. It is worth looking at two common tests in a little more detail: the χ2 test, and the t-test. The χ2 test of independence is a test for comparing categorical outcomes in two or more groups. For example, a number of trials have compared surgical site infections in patients who have been given different concentrations of oxygen perioperatively. In the PROXI trial,4Meyhoff C.S. Wetterslev J. Jorgensen L.N. et al.Effect of high perioperative oxygen fraction on surgical site infection and pulmonary complications after abdominal surgery: the PROXI randomized clinical trial.JAMA. 2009; 302: 1543-1550Crossref PubMed Scopus (305) Google Scholar 685 patients received oxygen 80%, and 701 patients received oxygen 30%. In the 80% group there were 131 infections, while in the 30% group there were 141 infections. In this study, the groups were oxygen 80% and oxygen 30%, and the outcome measure was the presence of a surgical site infection. The summary is a table (Table 1), and the hypothesis test compares this table (the ‘observed’ table) with the table that would be expected if the proportion of infections in each group was the same (the ‘expected’ table). The test statistic is χ2, from which a p-value is calculated. In this instance the p-value is 0.64, which means that results like this would occur 64% of the time if the null hypothesis were true. We thus have no evidence to reject the null hypothesis; the observed difference probably results from sampling variation rather than from an inherent difference between the two groups.Table 1Summary of the results of the PROXI trial. Figures are numbers of patients.GroupOxygen 80%Oxygen 30%OutcomeInfection131141No infection554560Total685701 Open table in a new tab The t-test is a statistical method for comparing means, and is one of the most widely used hypothesis tests. Imagine a study where we try to see if there is a difference in the onset time of a new neuromuscular blocking agent compared with suxamethonium. We could enlist 100 volunteers, give them a general anaesthetic, and randomise 50 of them to receive the new drug and 50 of them to receive suxamethonium. We then time how long it takes (in seconds) to have ideal intubation conditions, as measured by a quantitative nerve stimulator. Our data are therefore a list of times. In this case, the groups are ‘new drug’ and suxamethonium, and the outcome is time, measured in seconds. This can be summarised by using means; the hypothesis test will compare the means of the two groups, using a p-value calculated from a ‘t statistic’. Hopefully it is becoming obvious at this point that the test statistic is usually identified by a letter, and this letter is often cited in the name of the test. The t-test comes in a number of guises, depending on the comparison being made. A single sample can be compared with a standard (Is the BMI of school leavers in this town different from the national average?); two samples can be compared with each other, as in the example above; or the same study subjects can be measured at two different times. The latter case is referred to as a paired t-test, because each participant provides a pair of measurements—such as in a pre- or postintervention study. A large number of methods for testing hypotheses exist; the commonest ones and their uses are described in Table 2. In each case, the test can be described by detailing the groups being compared (Table 2, columns) the outcome measures (rows), the summary, and the test statistic. The decision to use a particular test or method should be made during the planning stages of a trial or experiment. At this stage, an estimate needs to be made of how many test subjects will be needed. Such calculations are described in detail elsewhere.5Columb M.O. Atkinson M.S. Statistical analysis: sample size and power estimations.BJA Educ. 2016; 16: 159-161Abstract Full Text Full Text PDF Scopus (45) Google ScholarTable 2The principle types of hypothesis test. Tests comparing more than two samples can indicate that one group differs from the others, but will not identify which. Subsequent ‘post hoc’ testing is required if a difference is found.Type of dataNumber of groups1 (comparison with a standard)1 (before and after)2More than 2Measured over a continuous rangeCategoricalBinomial testMcNemar's testχ2 test, or Fisher's exact (2×2 tables), or comparison of proportionsχ2 testLogistic regressionContinuous (normal)One-sample t-testPaired t-testIndependent samples t-testAnalysis of variance (ANOVA)Regression analysis, correlationContinuous (non-parametric)Sign test (for median)Sign test, or Wilcoxon matched-pairs testMann–Whitney U testKruskal–Wallis testSpearman's rank correlation Open table in a new tab Although hypothesis tests have been the basis of modern science since the middle of the 20th century, they have been plagued by misconceptions from the outset; this has led to what has been described as a crisis in science in the last few years: some journals have gone so far as to ban p-values outright.6Trafimow D. Marks M. Editorial.Basic Appl Soc Psych. 2015; 37: 1-2Crossref Scopus (442) Google Scholar This is not because of any flaw in the concept of a p-value, but because of a lack of understanding of what they mean. Possibly the most pervasive misunderstanding is the belief that the p-value is the chance that the null hypothesis is true, or that the p-value represents the frequency with which you will be wrong if you reject the null hypothesis (i.e. claim to have found a difference). This interpretation has frequently made it into the literature, and is a very easy trap to fall into when discussing hypothesis tests. To avoid this, it is important to remember that the p-value is telling us something about our sample, not about the null hypothesis. Put in simple terms, we would like to know the probability that the null hypothesis is true, given our data. The p-value tells us the probability of getting these data if the null hypothesis were true, which is not the same thing. This fallacy is referred to as ‘flipping the conditional’; the probability of an outcome under certain conditions is not the same as the probability of those conditions given that the outcome has happened. A useful example is to imagine a magic trick in which you select a card from a normal deck of 52 cards, and the performer reveals your chosen card in a surprising manner. If the performer were relying purely on chance, this would only happen on average once in every 52 attempts. On the basis of this, we conclude that it is unlikely that the magician is simply relying on chance. Although simple, we have just performed an entire hypothesis test. We have declared a null hypothesis (the performer was relying on chance); we have even calculated a p-value (1 in 52, ≈0.02); and on the basis of this low p-value we have rejected our null hypothesis. We would, however, be wrong to suggest that there is a probability of 0.02 that the performer is relying on chance—that is not what our figure of 0.02 is telling us. To explore this further we can create two populations, and watch what happens when we use simulation to take repeated samples to compare these populations. Computers allow us to do this repeatedly, and to see what p-values are generated (see Supplementary online material).7Colquhoun D. An investigation of the false discovery rate and the misinterpretation of p-values.R Soc Open Sci. 2014; 1: 140216Crossref PubMed Scopus (440) Google Scholar Fig 1 illustrates the results of 100,000 simulated t-tests, generated in two set of circumstances. In Fig 1a, we have a situation in which there is a difference between the two populations. The p-values cluster below the 0.05 cut-off, although there is a small proportion with p >0.05. Interestingly, the proportion of comparisons where p <0.05 is 0.8 or 80%, which is the power of the study (the sample size was specifically calculated to give a power of 80%). Figure 1b depicts the situation where repeated samples are taken from the same parent population (i.e. the null hypothesis is true). Somewhat surprisingly, all p-values occur with equal frequency, with p<0.05 occurring exactly 5% of the time. Thus, when the null hypothesis is true, a type I error will occur with a frequency equal to the alpha significance cut-off. Figure 1 highlights the underlying problem: when presented with a p-value <0.05, is it possible with no further information, to determine whether you are looking at something from Fig 1a or Fig 1b? Finally, it cannot be stressed enough that although hypothesis testing identifies whether or not a difference is likely, it is up to us as clinicians to decide whether or not a statistically significant difference is also significant clinically. As mentioned above, some have suggested moving away from p-values, but it is not entirely clear what we should use instead. Some sources have advocated focussing more on effect size; however, without a measure of significance we have merely returned to our original problem: how do we know that our difference is not just a result of sampling variation? One solution is to use Bayesian statistics. Up until very recently, these techniques have been considered both too difficult and not sufficiently rigorous. However, recent advances in computing have led to the development of Bayesian equivalents of a number of standard hypothesis tests.8Ly A. Verhagen J. Wagenmakers E. Harold Jeffreys’s default Bayes factor hypothesis tests: explanation, extension, and application in psychology.J Math Psychol. 2016; 72: 19-32Crossref Scopus (188) Google Scholar These generate a ‘Bayes Factor’ (BF), which tells us how more (or less) likely the alternative hypothesis is after our experiment. A BF of 1.0 indicates that the likelihood of the alternate hypothesis has not changed. A BF of 10 indicates that the alternate hypothesis is 10 times more likely than we originally thought. A number of classifications for BF exist; greater than 10 can be considered ‘strong evidence’, while BF greater than 100 can be classed as ‘decisive’. Figures such as the BF can be quoted in conjunction with the traditional p-value, but it remains to be seen whether they will become mainstream. The author declares that they have no conflict of interest. The associated MCQs (to support CME/CPD activity) will be accessible atwww.bjaed.org/cme/home by subscribers to BJA Education. The following is the Supplementary data to this article: Download .txt (.0 MB) Help with txt files Multimedia component 1 Jason Walker FRCA FRSS BSc (Hons) Math Stat is a consultant anaesthetist at Ysbyty Gwynedd Hospital, Bangor, Wales, and an honorary senior lecturer at Bangor University. He is vice chair of his local research ethics committee, and an examiner for the Primary FRCA.
Morris et al. (1) estimate divergence times for land plants (embryophytes), concluding that they originated in the early Phanerozoic (515 to 473 Ma; midpoint, 494 Ma). In contrast, other molecular clock studies have placed that event 40% earlier, in the Precambrian (707 to 670 Ma) (2⇓–4). Knowing the correct time bears on understanding how land plants have impacted the biosphere (2). Morris et al. (1) conclude that the tree topology and size of the dataset had little impact on their results. They also suggest that their results were robust to “dating strategies,” which included removing a single maximum calibration while keeping all other maximum and minimum calibrations. The 37 minimum … [↵][1]1To whom correspondence should be addressed. Email: sbh{at}temple.edu. [1]: #xref-corresp-1-1
Editor—We agree with Zundert and colleagues1Van Zundert AAJ Gatt SP Kumar CM Van Zundert TCRV Pandit JJ ‘Failed supraglottic airway’: an algorithm for suboptimally placed supraglottic airway devices based on videolaryngoscopy.Br J Anaesth. 2017; 118: 645-649Abstract Full Text Full Text PDF PubMed Scopus (35) Google Scholar that anaesthetists sometimes accept lower standards for supraglottic airway device (SAD) placement than for tracheal tube placement. They advocate the use of videolaryngoscopy to correct a suboptimal SAD position using an algorithm, however they do not make any specific recommendation about which type of videolaryngoscope to use. A variety of classifications of videolaryngoscopes exist,2Mihai R Blair E Kay H Cook TM A quantitative review and meta-analysis of performance of non-standard laryngoscopes and rigid fibreoptic intubation aids.Anaesthesia. 2008; 63: 745-760Crossref PubMed Scopus (144) Google Scholar, 3Pott LM Murray WB Review of video laryngoscopy and rigid fiberoptic laryngoscopy.Curr Opin Anesthesiol. 2008; 21: 750-758Crossref PubMed Scopus (70) Google Scholar, 4Kelly FE Cook TM Seeing is believing: getting the best out of videolaryngoscopy.Br J Anaesth. 2016; 117: i9-13Abstract Full Text Full Text PDF PubMed Scopus (68) Google Scholar, 5Healy DW Maties O Hovord D Kheterpal S A systematic review of the role of videolaryngoscopy in successful orotracheal intubation.BMC Anesthesiol. 2012; 12: 32Crossref PubMed Scopus (109) Google Scholar and some classes (for example, channelled videolaryngoscopes) might prove to be very difficult to use in this situation; some may even cause trauma in the reduced space available when an SAD is present. Zundert and colleagues 6Van Zundert AAJ Gatt SP Kumar CM van Zundert TCRV Vision-guided placement of supraglottic airway device (SAD) prevents airway obstruction—a prospective audit.Br J Anaesth. 2017; 118: 462-463Abstract Full Text Full Text PDF PubMed Scopus (14) Google Scholar have previously referenced the C-MAC videolaryngoscope (Karl Storz, Tuttlingen, Germany) in this situation, but presumably other Macintosh-type bladed scopes would be acceptable. We would be keen to learn whether they have experience of other videolaryngoscopes, and we also wondered whether they had considered the place of the optical stylet. Finally, we were surprised to see a failed intubation algorithm that didn't explicitly recommend when to abandon efforts to place a supraglottic device. We worry that less experienced staff may make continued attempts to optimise placement when it would make more sense simply to intubate the patient. None declared.
Forty anaesthetists calculated maximum permissible doses of eight local anaesthetic formulations for simulated patients three times with three methods: an electronic calculator; nomogram; and pen and paper. Correct dose calculations with the nomogram (85/120) were more frequent than with the calculator (71/120) or pen and paper (57/120), Bayes Factor 4 and 287, p = 0.01 and p = 0.0003, respectively. The rates of calculations at least 120% the recommended dose with each method were different, Bayes Factor 7.9, p = 0.0007: 14/120 with the calculator; 5/120 with the nomogram; 13/120 with pen and paper. The median (IQR [range]) speed of calculation with pen and paper, 38.0 (25.0-56.3 [5-142]) s, was slower than with the calculator, 24.5 (17.8-37.5 [6-204]) s, p = 0.0001, or nomogram, 23.0 (18.0-29.0 [4-100]) s, p = 1 × 10-7 . Local anaesthetic dose calculations with the nomogram were more accurate than with an electronic calculator or pen and paper and were faster than with pen and paper.
Correct dosing of anaesthetic and peri-operative drugs in obese patients is a frequent and difficult problem. The UK guideline developed by Nightingale et al. 1 provides a pragmatic, simplified solution by recommending the use of total body weight (TBW), ideal body weight (IBW), lean body weight (LBW) and adjusted body weight (ABW) to guide dosing of drugs in common use. They provide equations to calculate these parameters 1, but with the exception of IBW, these are complex. This risks leaving the anaesthetist with a choice between making the calculations and making a ‘best guess’ in the interests of saving time. Given the established risks of over- and under-dosing drugs in obese individuals 2, 3, the authors believe that minimising the temptation of guesswork would be beneficial. Two nomograms for LBW and ABW were constructed using methods previously described 4, based on the relevant equations 1. Each nomogram includes a list of relevant perioperative drugs (Fig. 3). The nomograms were evaluated using a randomly generated dataset of 100 simulated patient values for height and weight. Two investigators performed calculations of LBW and ABW using printed A4 copies of each nomogram and a straight edged ruler. Bland-Altman analysis was used to compare the values for LBW and ABW calculated using the nomogram with corresponding values calculated using a spreadsheet. The mean (SD) LBW nomogram bias was 0.09 (0.10) kg, with limits of agreement 0.10 to 0.28 and mean (SD) percentage error 0.15 (0.17) %. The ABW nomogram bias was −0.06 (0.09) kg, with limits of agreement −0.23 to 0.11 and percentage error −0.06 (0.11) %. This demonstrates a high degree of agreement between the nomograms and the calculated values. Nomograms have been validated and used in the healthcare setting for some time and their advantages include accuracy, low cost and convenience. Similar nomograms for LBW calculations in paediatric obesity have been shown to be quicker and less likely to result in mistakes than equations performed with the aid of a calculator 4. Each nomogram is designed to be reproduced as an A4 sheet and can be laminated and placed within the anaesthetic room or in the departmental ‘obesity pack’. Alternatively, the nomograms may be printed and filed in the patient's records for future reference.
The sterile hybrid grass Miscanthus x giganteus (Mxg) can produce more than 30 t dry matter/ha/year. This biomass has a range of uses, including animal bedding and a source of heating fuel. The grass provides a wide range of other ecosystem services (ES), including shelter for crops and livestock, a refuge for beneficial arthropods, reptiles and earthworms and is an ideal cellulosic feedstock for liquid biofuels such as renewable (drop-in) diesel. In this study, the effects of different strains of the beneficial fungus Trichoderma on above-and below-ground biomass of Mxg were evaluated in glasshouse and field experiments, the latter on a commercial dairy farm over two years. Other ES benefits of Trichoderma measured in this study included enhanced leaf chlorophyll content as well as increased digestibility of the dried material for livestock. This study shows, for the first time for a biofuel feedstock plant, how Trichoderma can enhance productivity of such plants and complements other recent work on the wide-ranging provision of ES by this plant species.
We read with fascination and a certain amount of unbridled excitement Dr Goodman's correspondence, in which he correctly points out that the phrase ‘We read with interest’ is much overused. A quick search in Google Books 1 would indicate that the phrase has been around for at least 200 years, reaching a peak use around the 1830s. Its use in Anaesthesia is certainly on the increase. We compared Dr Goodman's reported results with a randomly chosen issue from the shelves (July 1999) and discovered that of 24 items of correspondence (not including replies), only two of them declared any interest in the original article. However, one of the authors was ‘disturbed to read’, another was ‘surprised’, while yet another ‘read with a growing sense of unease’ – which suggests that we were all far easier to upset in the closing days of the 20th Century. ‘We read with interest’ has provided authors with a quick and easy answer to the question ‘how do I start’? The alternative, for example lunging in with something like ‘Dr Seuss's publication on feline headwear raises some important questions’ can seem abrupt, but we're sure readers will adjust with practice
We agree with Chillingworth et al. that it is important to focus on a safe process of tracheal extubation before addressing concerns about theatre utilisation 1, 2. Chillingworth et al. reported 5/835 (0.6%) patients administered general anaesthesia had residual neuromuscular blockade in the postoperative recovery unit. We found a similar prevalence in a three-month prospective audit at our hospital (5/954 (0.5%) 3). The incidence of awareness under general anaesthesia using neuromuscular blockade is less prevalent than these (1:8200 4), but given that 18% of cases of unintentional awareness occur during emergence from general anaesthesia, that 88% of these are preventable and that 51% of patients are distressed by these 5, we think that safety remains paramount. It remains a challenge to reduce the prevalence of residual block to zero 6, but repeat audits after staff education at our hospital have found that it is possible (unpublished data). The All Wales Airway Group (AWAG) are currently undertaking an audit of residual neuromuscular blockade, which should collect further data about the extent of this problem in Wales, and provide a template for monitoring this problem throughout the UK.
Since its original publication, the revised Baux score for mortality prediction in burns patients has been widely adopted. It uses readily available measures, and it is based on regression analysis from actual data rather than a theoretical model. However, the necessary calculations are too complex to perform with anything other than a scientific calculator or dedicated software, which may create issues in a clinical setting where access to electronic devices may be limited. We designed a nomogram capable of performing the calculation to a high degree of accuracy, and evaluated its performance on a set of randomly generated patient data to ensure that the nomogram gives accurate and repeatable results. The nomogram has a bias of -0.003 percentage points, with limits of agreement -0.3619 to 0.3550 and a repeatability coefficient of 0.29 percentage points. We feel that the nomogram's accuracy, low cost, speed and ease of use would make it a very useful adjunct during the initial assessment of burns patients. It could also realistically be used to crosscheck calculations made by other methods.
While local anaesthetic agents are usually safe and are used ubiquitously, inadvertent overdoses may have potentially fatal consequences. Errors in the dosing of local anaesthetics frequently occur due to inherent difficulties in remembering the toxic dosage limits, difficulties in performing the appropriate calculations correctly, and errors in estimating patient weight. We have developed a simple graphical calculation aid (nomogram) to overcome these problems and facilitate rapid cross-checking of the maximum safe dose for a variety of local anaesthetic agents in common use. Standard mathematical techniques were used to draft the nomogram. A randomised blinded study using simulated patient data and Bland-Altman analysis was used to assess the accuracy and precision of the nomogram. The nomogram was found to have a bias of 0.0 ml, with limits of agreement -0.05-0.04 ml. It was found to be easy to use and suitably accurate for clinical use.
Toxic dose limits (mg.kg(-1)) for local anaesthetics based on body weight are well-established, but calculation of the maximum safe volume (ml) of a given agent and formulation is complex, and frequently results in errors. We therefore developed a nomogram to perform this calculation. We compared the performance of the nomogram with a spreadsheet and a general purpose calculator using simulated clinical data. Bland-Altman analysis showed close agreement between the nomogram and spreadsheet, with bias of -0.07 ml and limits of agreement of -0.38 to +0.24 ml (correlation coefficient r(2) = 0.9980; p < 0.001). The nomogram produced fewer and smaller errors compared with the calculator. Our nomogram calculates the maximum safe volume (ml) of local anaesthetic to a clinically acceptable degree of accuracy. It facilitates rapid cross-checking of dosage calculations performed by electronic or other means at negligible cost, and can potentially reduce the incidence of local anaesthetic toxicity.
We were interested to read the interim and follow-up reports into the fire in Bath Intensive Care Unit in 2011 1, 2. We would like to report a near miss that could have resulted in a fire in one of our anaesthetic rooms, the root cause analysis of which uncovered contributory latent and human errors relating to the design of the anaesthetic room, the refrigerator and the daily theatre equipment checklist. Rolls of gamgee tissue that were stored loosely on top of a wall-mounted refrigerator had obstructed airflow through a grille covering the rear mounted condenser, resulting in superheating of the condenser and scorching of the gamgee tissue. The interior temperature of the refrigerator was measured as 22 °C, and drug ampoules stored within it were warm to touch. As part of the daily equipment checklist, the refrigerator temperature had been recorded daily between 20–23 °C for the previous three weeks, without this problem's having been acknowledged. There had been several recent incidents in which thermometers had failed, which may have contributed to this fixation error. The design and wording of the checklist encouraged box-ticking, rather than proper equipment assessment. No harm resulted from this incident, which was investigated according to the hospital's Adverse Incident Monitoring System, with immediate action taken to remove the gamgee tissue. Redesign of the equipment checklist has included the instruction “Confirm fridge temp < 5 degrees”. We have also recommended that either the refrigerator is replaced by a model with the condenser placed underneath, or a shelf is placed the minimum safe height above the refrigerator to allow adequate airflow. We would urge readers to take similar precautions.