Statistical inference is based on the laws of probability. However, frequentists and Bayesians interpret probability differently. Frequentists interpret probability as a long-run frequency over repeated sampling. Consequently, frequentist probability statements are sampling probabilities from a sampling distribution. A sampling distribution is the hypothetical long-run distribution of a statistic we would expect to observe, assuming the population value is fixed. The frequentist interpretation is confusing and leads to the widespread misinterpretation of P-values and confidence intervals. Bayesians interpret probability as a strength of belief. Consequently, Bayesian probability statements are inferential probabilities. An inferential probability is a direct statement about the quantity of interest, that is, the truth of the hypothesis and the size of the treatment effect, given the data observed. As clinicians and researchers, we seek inferential probabilities ('What is the probability the treatment works given the study data?') not sampling probabilities ('If the treatment does not work, how surprising are the study data?'). I explore the frequentist and Bayesian perspectives on probability, address the barriers to adopting Bayesian methods, and make the case for Bayesian inference in anaesthesia research.
Abstract Background Multicentre trials in anaesthesia and critical care report low rates of statistically significant differences. This finding may partly reflect conventional sample size methods, which assume a fixed treatment effect. Assurance methods use a design prior to represent uncertainty in the expected treatment effect, which may provide a more realistic way of estimating sample sizes. Methods We calculated power curves across a range of effect sizes, design priors, and sample sizes using frequentist and Bayesian assurance methods and compared the sample sizes required to achieve 80% and 90% power to the conventional method. We standardised the design priors across effect sizes using the coefficient of variation (CV). We derived a theoretical limit for achievable power. We validated a normal approximation to the Bayesian posterior distribution. Results Frequentist and Bayesian assurance methods produced similar power curves across all scenarios. At a CV of 0.5 – reflecting realistic prior uncertainty in the expected effect size – both methods required sample sizes that were approximately 1.5 to 3.5 times larger than the conventional method. The theoretical power limit depends only on the CV of the design prior and holds true across all effect sizes. The normal approximation to the Bayesian posterior distribution matched the results obtained from Markov chain Monte Carlo sampling. Conclusions Incorporating clinical uncertainty in the expected effect size substantially increases the sample size required to achieve adequate power, which has important implications for the feasibility of randomised trials in anaesthesia and critical care.
AIM:Our aim was to determine the incidence of delirium in a tertiary intensive care unit (ICU) in Auckland, New Zealand compared to other Australasian ICUs. To determine the incidence of delirium among different ethnicities and identify risk factors and outcomes of patients experiencing delirium. METHODS:The design was a retrospective observational study. The setting was a single-centre, 24 bed, tertiary ICU in Auckland, New Zealand. The participants were two hundred and twenty-two patients admitted to the ICU over 10 months in 2019. The main outcome measures were incidence of delirium, identified using the Confusion Assessment Method - ICU (CAM-ICU) screening, antipsychotic prescription, 12-month mortality, and ICU discharge disposition. RESULTS:Fifty of the 222 (23%) patients had delirium. There was no association between the incidence of delirium and ethnicity (p=0.39). The risk of delirium increased with ICU duration of stay (odds ratio [OR]: 1.003, 95% CI, 1.001-1.005, p=0.004), days on vasopressors (p<0.001) and days on mechanical ventilation (p<0.001). Thirty-three of the 50 (66%) patients received at least one antipsychotic medication. Twelve-month mortality was not associated with delirium (OR: 0.97, 95% CI 0.73-1.22, p=0.81). Delirium was not associated with ICU discharge disposition (p=0.20). CONCLUSIONS:The incidence of delirium in this single-centre, tertiary Auckland ICU was comparable to other Australasian ICUs. There was no difference in the incidence of delirium between different ethnicities. Positive associations to delirium included length of stay in ICU, number of days on vasopressors and duration of mechanical ventilation. Delirium was not associated with an increased risk of 12-month mortality and was not associated with ICU discharge disposition.
Days alive out of hospital (DAOH) is defined as the number of days a patient spends alive and out of hospital in a predefined period. Current understanding of the performance of this outcome measure in the intensive care (ICU) population is limited thus we aimed to investigate the relationship between DAOH and prognostic indicators in patients admitted to ICU and to compare the performance of DAOH at 30-days (DAOH30), 90-days (DAOH90) and 180-days (DAOH180). In a retrospective cohort study, all patients aged over 18 years admitted to ICU in New Zealand between 1st January 2017 and 31st December 2023 were eligible for inclusion. DAOH was calculated using merged data from the Australian and New Zealand Intensive Care Society (ANZICS) Adult Patient Database (APD) and the National Minimum Dataset (NMDS) data. 72,546 admissions recorded in the ANZICS APD were linked to NMDS. Median (interquartile range [IQR]) DAOH30 was 15 days (0–22), median (IQR) DAOH90 was 72 days (40–81) and median (IQR) DAOH180 was 160 days (117–171). DAOH90 was negatively correlated with age (ρ= -0.19, p < 0.001) and Acute Physiology and Chronic Health Evaluation (APACHE) II score (ρ = -0.42 p < 0.001). DAOH90 was lower in patients receiving invasive ventilation, vasoactive support and renal replacement therapy and those admitted with severe traumatic brain injury (TBI). Lower median DAOH90 was associated with increased variation in outcome across multiple subgroups. A floor effect was observed for DAOH30 (with a median DAOH30 value of zero days) in patients with high APACHE II score, severe TBI and patients undergoing renal replacement therapy or tracheostomy insertion. The time for the excess hazard of hospitalisation and mortality to reach baseline were 37.3 and 18.6 days respectively but were prolonged in certain patient subgroups. Our study provides evidence to support the use of DAOH for patients admitted to ICU, however DAOH30 is unlikely to be a suitable measure.
Threshold statistical testing produces errors: false positives (type I errors) and false negatives (type II errors). False positives can be quantified as the familywise error rate (FWER) or the false discovery rate (FDR). At the usual significance threshold of 0.05, the chance of a false positive is 5%. However, across a family of independent tests, the chance of at least one false positive (the FWER) is much higher, rising to >60% for 20 tests. A recent simulation study by Chen and Dexter compared different methods of controlling for false positives, which highlights the perils of multiple statistical tests. Here, we discuss their findings and examine the pros and cons of different approaches to multiple testing.
By reading this article you should be able to: •Explain metabolic alkalosis using a traditional (bicarbonate) and physiochemical (Stewart) approach to acid–base disturbances. •Describe the main causes of metabolic alkalosis encountered in anaesthesia and ICU practice and discuss their treatment. •Identify the constituents of mixed acid–base disturbances. •Describe the acid–base effects of commonly used i.v. fluids. David Sidebotham FANZCA is a cardiac anaesthetist and intensivist in Auckland, New Zealand. He is an editorial board member and editor for BJA Education. Michael Park FRCP Edin FCICM is an intensive care physician in Hastings, New Zealand. He is chair of the Central region critical care leadership group and a certified instructor for the BASIC courses.
Randomized controlled trials are one of the best ways of quantifying the effectiveness of medical interventions. Therefore, when the authors of a randomized superiority trial report that differences in the primary outcome between the intervention group and the control group are “significant” (i.e., P ≤ 0.05), we might assume that the intervention has an effect on the outcome. Similarly, when differences between the groups are “not significant,” we might assume that the intervention does not have an effect on the outcome. Nevertheless, both assumptions are frequently incorrect. In this article, we explore the relationship that exists between real treatment effects and declarations of statistical significance based on P values and confidence intervals. We explain why, in some circumstances, the chance an intervention is ineffective when P ≤ 0.05 exceeds 25
OBJECTIVE:Intervention for repair of secondary mitral valve disease is frequently associated with recurrent regurgitation. We sought to determine if there was sufficient evidence to support inclusion of anatomic indices of leaflet dysfunction in the management of secondary mitral valve disease.METHODS:We performed a systematic review and meta-analysis of published reports comparing anatomic indices of leaflet dysfunction with the complexity of valve repair and the outcome from intervention. Patients were stratified by the severity of leaflet dysfunction. A secondary analysis was performed comparing outcomes when procedural complexity was optimally matched to severity of leaflet dysfunction and when intervention was not matched to dysfunction.RESULTS:We identified 6864 publications, of which 65 met inclusion criteria. An association between the severity of leaflet dysfunction and the procedural complexity was highly predictive of satisfactory freedom from recurrent regurgitation. Patients were categorized into 4 groups based on stratification of leaflet dysfunction. Satisfactory results were achieved in 93.7% of patients in whom repair complexity was appropriately matched to severity of leaflet dysfunction and in 68.8% in whom repair was not matched to dysfunction (odds ratio, 0.148; 95% confidence interval, 0.119-0.184; P < .0001).CONCLUSIONS:For patients with secondary mitral valve disease, satisfactory outcome from valve repair improves when procedural complexity is matched to anatomic indices of leaflet dysfunction. Anatomic indices of leaflet dysfunction should be considered when planning interventions for secondary mitral regurgitation. Routine inclusion of anatomic indices in trial design and reporting should facilitate comparison of results and strengthen guidelines. There are sufficient data to support anatomic staging of secondary mitral valve disease.
Objectives: Acute respiratory distress syndrome (ARDS) is associated with high ventilation-perfusion heterogeneity and dead-space ventilation. However, whether the degree of dead-space ventilation is associated with outcomes is uncertain. In this systematic review and meta-analysis, we evaluated the ability of dead-space ventilation measures to predict mortality in patients with ARDS. Data Sources: MEDLINE, CENTRAL, and Google Scholar from inception to November 2022. Study Selection: Studies including adults with ARDS reporting a dead-space ventilation index and mortality. Data Extraction: Two reviewers independently identified eligible studies and extracted data. We calculated pooled effect estimates using a random effects model for both adjusted and unadjusted results. The quality and strength of evidence were assessed using the Quality in Prognostic Studies and Grading of Recommendations, Assessment, Development, and Evaluation, respectively. Data Synthesis: We included 28 studies in our review, 21 of which were included in our meta-analysis. All studies had a low risk of bias. A high pulmonary dead-space fraction was associated with increased mortality (odds ratio [OR], 3.52; 95% CI, 2.22–5.58; p < 0.001; I 2 = 84%). After adjusting for other confounding variables, every 0.05 increase in pulmonary-dead space fraction was associated with an increased odds of death (OR, 1.23; 95% CI, 1.13–1.34; p < 0.001; I 2 = 57%). A high ventilatory ratio was also associated with increased mortality (OR, 1.55; 95% CI, 1.33–1.80; p < 0.001; I 2 = 48%). This association was independent of common confounding variables (OR, 1.33; 95% CI, 1.12–1.58; p = 0.001; I 2 = 66%). Conclusions: Dead-space ventilation indices were independently associated with mortality in adults with ARDS. These indices could be incorporated into clinical trials and used to identify patients who could benefit from early institution of adjunctive therapies. The cut-offs identified in this study should be prospectively validated.
BACKGROUND:The American Statistical Association has highlighted problems with null hypothesis significance testing and outlined alternative approaches that may 'supplement or even replace P-values'. One alternative is to report the false positive risk (FPR), which quantifies the chance the null hypothesis is true when the result is statistically significant. METHODS:We reviewed single-centre, randomised trials in 10 anaesthesia journals over 6 yr where differences in a primary binary outcome were statistically significant. We calculated a Bayes factor by two methods (Gunel, Kass). From the Bayes factor we calculated the FPR for different prior beliefs for a real treatment effect. Prior beliefs were quantified by assigning pretest probabilities to the null and alternative hypotheses. RESULTS:For equal pretest probabilities of 0.5, the median (inter-quartile range [IQR]) FPR was 6% (1-22%) by the Gunel method and 6% (1-19%) by the Kass method. One in five trials had an FPR ≥20%. For trials reporting P-values 0.01-0.05, the median (IQR) FPR was 25% (16-30%) by the Gunel method and 20% (16-25%) by the Kass method. More than 90% of trials reporting P-values 0.01-0.05 required a pretest probability >0.5 to achieve an FPR of 5%. The median (IQR) difference in the FPR calculated by the two methods was 0% (0-2%). CONCLUSIONS:Our findings suggest that a substantial proportion of single-centre trials in anaesthesia reporting statistically significant differences provide limited evidence of real treatment effects, or, alternatively, required an implausibly high prior belief in a real treatment effect. CLINICAL TRIAL REGISTRATION:PROSPERO (CRD42023350783).
Background: The sample size calculation is an important step in designing randomised controlled trials. For a trial comparing a control and an intervention group, where the outcome is binary, the sample size calculation requires choosing values for the anticipated event rates in both the control and intervention groups (the effect size), and the error rates. The Difference ELicitation in TriAls guidance recommends that the effect size should be both realistic, and clinically important to stakeholder groups. Overestimating the effect size leads to sample sizes that are too small to reliably detect the true population effect size, which in turn results in low achieved power. In this study, we use the Delphi approach to gain consensus on what the minimum clinically important effect size is for Balanced-2, a randomised controlled trial comparing processed electroencephalogram-guided ‘light’ to ‘deep’ general anaesthesia on the incidence of postoperative delirium in older adults undergoing major surgery. Methods: Delphi rounds were conducted using electronic surveys. Surveys were administered to two stakeholder groups: specialist anaesthetists from a general adult department in Auckland City Hospital, New Zealand (Group 1), and specialist anaesthetists with expertise in clinical research, identified from the Australian and New Zealand College of Anaesthetist’s Clinical Trials Network (Group 2). A total of 187 anaesthetists were invited to participate (81 from Group 1 and 106 from Group 2). Results from each Delphi round were summarised and presented in subsequent rounds until consensus was reached (>70% agreement). Results: The overall response rate for the first Delphi survey was 47% (88/187). The median minimum clinically important effect size was 5.0% (interquartile range: 5.0–10.0) for both stakeholder groups. The overall response rate for the second Delphi survey was 51% (95/187). Consensus was reached after the second round, as 74% of respondents in Group 1 and 82% of respondents in Group 2 agreed with the median effect size. The combined minimum clinically important effect size across both groups was 5.0% (interquartile range: 3.0–6.5). Conclusions: This study demonstrates that surveying stakeholder groups using a Delphi process is a simple way of defining a minimum clinically important effect size, which aids the sample size calculation and determines whether a randomised study is feasible.
................................................................................................................................................................. Correspondence to: D. Sidebotham Email: dsidebotham@adhb.govt.nz Accepted: 13December 2022
Are the results of randomised trials reliable and are p values and confidence intervals the best way of quantifying efficacy? Low power is common in medical research, which reduces the probability of obtaining a 'significant result' and declaring the intervention had an effect. Metrics derived from Bayesian methods may provide an insight into trial data unavailable from p values and confidence intervals. We did a structured review of multicentre trials in anaesthesia that were published in the New England Journal of Medicine, The Lancet, Journal of the American Medical Association, British Journal of Anaesthesia and Anesthesiology between February 2011 and November 2021. We documented whether trials declared a non-zero effect by an intervention on the primary outcome. We documented the expected and observed effect sizes. We calculated a Bayes factor from the published trial data indicating the probability of the data under the null hypothesis of zero effect relative to the alternative hypothesis of a non-zero effect. We used the Bayes factor to calculate the post-test probability of zero effect for the intervention (having assumed 50% belief in zero effect before the trial). We contacted all authors to estimate the costs of running the trials. The median (IQR [range]) hypothesised and observed absolute effect sizes were 7% (3-13% [0-25%]) vs. 2% (1-7% [0-24%]), respectively. Non-zero effects were declared for 12/56 outcomes (21%). The Bayes factor favouring a zero effect relative to a non-zero effect for these 12 trials was 0.000001-1.9, with post-test zero effect probabilities for the intervention of 0.0001-65%. The other 44 trials did not declare non-zero effects, with Bayes factors favouring zero effect of 1-688, and post-test probabilities of zero effect of 53-99%. The median (IQR [range]) study costs reported by 20 corresponding authors in US$ were $1,425,669 ($514,766-$2,526,807 [$120,758-$24,763,921]). We think that inadequate power and mortality as an outcome are why few trials declared non-zero effects. Bayes factors and post-test probabilities provide a useful insight into trial results, particularly when p values approximate the significance threshold.
In this article, I discuss the potential pitfalls of interpreting p values, confidence intervals, and declarations of statistical significance. To illustrate the issues, I discuss the LOVIT trial, which compared high-dose vitamin C with placebo in mechanically ventilated patients with sepsis. The primary outcome - the proportion of patients who died or had persisting organ dysfunction at day 28 - was significantly higher in patients who received vitamin C (p = .01). The authors had hypothesized that vitamin C would have a beneficial effect, although the prior evidence for benefit was weak. There was no prior evidence for a harmful effect of high-dose vitamin C. Consequently, the pretest probability for harm was low. The sample size was calculated assuming a 10% absolute risk difference, which was optimistic. Overestimating the effect size when calculating the sample size leads to low power. For these reasons, we should be skeptical that vitamin C causes harm in septic patients, despite the significant result. p-values and confidence intervals are probabilities concerning the chance of obtaining the observed data. However, we are more interested in the chance the intervention has a real effect on the outcome. That is to say, we are more interested in whether the hypothesis is true. A Bayesian approach allows us to estimate the false positive risk, which is the post-test probability there is no effect of the intervention. The false positive risk for the LOVIT trial (calculated from the published summary data using uniform priors for the parameter values) is 70%. Most likely, high-dose vitamin C does not cause harm in septic patients. Most likely it has no effect at all. If there is an effect, it is probably small and most likely beneficial.
Echocardiography is an invaluable tool in the management of both veno-arterial (V-A) and veno-venous (V-V) extracorporeal membrane oxygenation (ECMO). Prior to ECMO establishment, echocardiography aids decision-making regarding the appropriate mode of support (V-A or V-V) and provides real-time guidance of guidewire insertion and cannula placement during institution of ECMO. During ECMO support, echocardiography is used to assess the effectiveness of left ventricular (LV) support, to diagnose complications, and to assess readiness to wean from ECMO. In the post-ECMO period, echocardiography is used to assess ongoing cardiac recovery. Consensus guidelines and further studies will enable complications unique to the management of ECMO to be more readily identified.
Background:In medical research, null hypothesis significance testing (NHST) is the dominant framework for statistical inference. NHST involves calculating P-values and confidence intervals to quantify the evidence against the null hypothesis of no effect. However, P-values and confidence intervals cannot tell us the probability that the hypothesis is true. In contrast, false-positive risk (FPR) and false-negative risk (FNR) are post-test probabilities concerning the truth of the hypothesis, that is to say, the probability a real effect exists.Methods:We calculated the FPR or FNR for 53 individual multicentre trials in critical care based on a pretest probability of 0.5 that the hypothesis was true.Results:For trials reporting statistical significance, the FPR varied between 0.1% and 57.6%. For trials reporting non-significance, the FNR varied between 1.7% and 36.9%. Twenty-six of 47 trials (55.3%) reporting non-significance provided strong or very strong evidence in favour of the null hypothesis; the remaining trials provided limited evidence. There was no obvious relationship between the P-value and the FNR.Conclusions:The FPR and FNR showed marked variability, indicating that the probability of a real or absent treatment effect differed substantially between trials. Only one trial reporting statistical significance provided convincing evidence of a real treatment effect, and nearly half of all trials reporting non-significance provided limited evidence for the absence of a treatment effect. Our findings suggest that the quality of evidence from multicentre trials in critical care is highly variable.