This is a protocol for a Cochrane Review (intervention). The objectives are as follows: To assess the effectiveness of social network and social support interventions to support cardiac rehabilitation and secondary prevention in the management of people with heart disease. As a secondary output of this review, and to assist in conceptualising future research focused on social network and social support interventions, we aim to develop a logic model theorising the relationship between social networks or social support and heart disease outcomes. We will draw on existing models of social support for health (e.g. Berkman 2000), as well as established approaches to theorising and implementing behaviour change (e.g. Michie 2011).
There is increased awareness of palliative care needs in people with COPD or interstitial lung disease (ILD). This European Respiratory Society (ERS) task force aimed to provide recommendations for initiation and integration of palliative care into the respiratory care of adult people with COPD or ILD.The ERS task force consisted of 20 members, including representatives of people with COPD or ILD and informal caregivers. Eight questions were formulated, four in the Population, Intervention, Comparison, Outcome format. These were addressed with full systematic reviews and application of Grading of Recommendations Assessment, Development and Evaluation for assessing the evidence. Four additional questions were addressed narratively. An “evidence-to-decision” framework was used to formulate recommendations.The following definition of palliative care for people with COPD or ILD was agreed. A holistic and multidisciplinary person-centred approach aiming to control symptoms and improve quality of life of people with serious health-related suffering because of COPD or ILD, and to support their informal caregivers. Recommendations were made regarding people with COPD or ILD and their informal caregivers: to consider palliative care when physical, psychological, social or existential needs are identified through holistic needs assessment; to offer palliative care interventions, including support for informal caregivers, in accordance with such needs; to offer advance care planning in accordance with preferences; and to integrate palliative care into routine COPD and ILD care. Recommendations should be reconsidered as new evidence becomes available.
Background Summer learning loss has been the subject of longstanding concern among researchers, the public and policy makers. The aim of the current research was to investigate inequality changes in children’s mental health and cognitive ability across the summer holidays. Methods We conducted linear and logistic regression analysis of mental health (borderline-abnormal total difficulty and prosocial scores on the strengths and difficulties questionnaire (SDQ)) and verbal cognitive ability (reading, verbal reasoning or vocabulary) at ages 7, 11 and 14, comparing UK Millennium Cohort Study members who were interviewed before and after the school summer holidays. Inequalities were assessed by including interaction terms in the outcome models between a discrete binary variable with values representing time periods and maternal academic qualifications. Coefficients of the interaction terms were interpreted as changes from the pre- to post-holiday period in the extent of inequality in the outcome between participants whose mothers had high or low educational qualifications. Separate models were fitted for each age group and outcome. We used inverse probability weights to allow for differences in the characteristics of cohort members assessed before and after the summer holidays. Results Mental health (borderline/abnormal SDQ total and prosocial scores) at ages 7 and 14 worsened and verbal cognitive ability scores at age 7 were lower among those surveyed after the summer holidays. Mental health inequalities were larger after the holidays at age 7 ([OR = 1.4; 95%CI (0.6, 3.2) and 14: [OR = 1.5; 95%CI (0.7, 3.2)], but changed little at age 11 (OR = 0.9; 95%CI (0.4, 2.6)]. There were differences in pro-social behaviours among those surveyed before/after the school holidays at age 14 [OR = 1.2; 95%CI (0.5, 3.5)] but not at age 7 or 11. There was little change in inequalities in verbal cognitive ability scores over the school holidays [Age 7: b = 1.3; 95%CI (− 3.3, 6.0); Age 11: b = − 0.7; 95%CI (− 4.3, 2.8); Age 14: b = − 0.3; 95%CI (− 1.0, 0.4)]. Conclusion We found inequalities in mental health and cognitive ability according to maternal education, and some evidence or worsening mental health and mental health inequalities across school summer holidays. We found little evidence of widening inequalities in verbal cognitive ability. Widespread school closures during the COVID-19 restrictions have prompted concerns that prolonged closures may widen health and educational inequalities. Management of school closures should focus on preventing or mitigating inequalities that may arise from differences in the support for mental health and learning provided during closures by schools serving more or less disadvantaged children.
ObjectivesSchool closures have been used as a core non-pharmaceutical intervention (NPI) during the COVID-19 pandemic. This review aims at identifying SARS-CoV-2 transmission in educational settings during the first waves of the pandemic.MethodsThis literature review assessed studies published between December 2019 and 1 April 2021 in Medline and Embase, which included studies that assessed educational settings from approximately January 2020 to January 2021. The inclusion criteria were based on the PCC framework (P-Population, C-Concept, C-Context). The studyPopulationwas restricted to people 1–17 years old (excluding neonatal transmission), theConceptwas to assess child-to-child and child-to-adult transmission, while theContextwas to assess specifically educational setting transmission.ResultsFifteen studies met inclusion criteria, ranging from daycare centres to high schools and summer camps, while eight studies assessed the re-opening of schools in the 2020–2021 school year. In principle, although there is sufficient evidence that children can both be infected by and transmit SARS-CoV-2 in school settings, the SAR remain relatively low—when NPI measures are implemented in parallel. Moreover, although the evidence was limited, there was an indication that younger children may have a lower SAR than adolescents.ConclusionsTransmission in educational settings in 2020 was minimal—when NPI measures were implemented in parallel. However, with an upsurge of cases related to variants of concern, continuous surveillance and assessment of the evidence is warranted to ensure the maximum protection of the health of students and the educational workforce, while also minimising the numerous negative impacts that school closures may have on children.
BACKGROUND AND OBJECTIVE:This article explores the need for conceptual advances and practical guidance in the application of the GRADE approach within public health contexts. METHODS:We convened an expert workshop and conducted a scoping review to identify challenges experienced by GRADE users in public health contexts. We developed this concept article through thematic analysis and an iterative process of consultation and discussion conducted with members electronically and at three GRADE Working Group meetings. RESULTS:Five priority issues can pose challenges for public health guideline developers and systematic reviewers when applying GRADE: (1) incorporating the perspectives of diverse stakeholders; (2) selecting and prioritizing health and "nonhealth" outcomes; (3) interpreting outcomes and identifying a threshold for decision-making; (4) assessing certainty of evidence from diverse sources, including nonrandomized studies; and (5) addressing implications for decision makers, including concerns about conditional recommendations. We illustrate these challenges with examples from public health guidelines and systematic reviews, identifying gaps where conceptual advances may facilitate the consistent application or further development of the methodology and provide solutions. CONCLUSION:The GRADE Public Health Group will respond to these challenges with solutions that are coherent with existing guidance and can be consistently implemented across public health decision-making contexts.
Background: School closures have been used as a core Non pharmaceutical intervention during the COVID-19 pandemic, however the role of educational settings in COVID-19 transmission is still unclear. Methods: This systematic literature review assessed studies published between December 2019 and April 1, 2021 in Medline and Embase, which included studies that assessed educational settings from approximately January 2020 to January 2021. The inclusion criteria were based on the PCC framework (P-Population, C-Concept, C-Context). The study Population was restricted to people 1-17 years old (excluding neonatal transmission), the Concept was to assess child-to-child and child-to-adult transmission, while the Context was to assess specifically educational setting transmission clusters. Results: Fifteen studies met inclusion criteria, ranging from daycare centers to high schools and summer camps, while eight studies assessed the re-opening of schools in the 2020-2021 school year. In principle although there is sufficient evidence that children can both be infected by and transmit SARS-CoV-2 in school settings, the SAR remain relatively low -when NPI measures are implemented in parallel. Moreover, although the evidence was limited there was an indication that younger children may have a lower SAR than adolescents. Conclusions: Transmission in educational settings in 2020 was minimal -when NPI measures were implemented in parallel. However, with an upsurge of cases related to variants of concern, continuous surveillance and assessment of the evidence is warranted to ensure the maximum protection of the health of students and the educational workforce, while also minimising the numerous negative impacts that school closures may have on children.
Effect direction (evidence to indicate improvement, deterioration, or no change in an outcome) can be used as a standardized metric which enables the synthesis of diverse effect measures in systematic reviews. The effect direction (ED) plot was developed to support the synthesis and visualization of effect direction data. Methods for the ED plot require updating in light of new Cochrane guidance on alternative synthesis methods. To update the ED plot, statistical significance was removed from the algorithm for within-study synthesis and use of a sign test was considered to examine whether patterns of ED across studies could be due to chance alone. The revised methods were applied to an existing Cochrane review of the health impacts of housing improvements. The revised ED plot provides a method of data visualization in synthesis without meta-analysis that incorporates information about study characteristics and study quality, using ED as a common metric, without relying on statistical significance to combine outcomes of single studies. The results of sign tests, when appropriate, suggest caution in over-interpreting apparent patterns in effect direction, especially when the number of included studies is small. The revised ED plot meets the need for alternative methods of synthesis and data visualization when meta-analysis is not possible, enabling a transparent link between the data and conclusions of a systematic review. ED plots may be particularly useful in reviews that incorporate nonrandomized studies, complex systems approaches, and diverse sources of evidence, due to the variety of study designs and outcomes in such reviews.
Objectives: The objective of the study is to present the Grading of Recommendations Assessment, Development, and Evaluation (GRADE) conceptual approach to the assessment of certainty of evidence from modeling studies (i.e., certainty associated with model outputs). Study Design and Setting: Expert consultations and an international multidisciplinary workshop informed development of a conceptual approach to assessing the certainty of evidence from models within the context of systematic reviews, health technology assessments, and health care decisions. The discussions also clarified selected concepts and terminology used in the GRADE approach and by the modeling community. Feedback from experts in a broad range of modeling and health care disciplines addressed the content validity of the approach. Results: Workshop participants agreed that the domains determining the certainty of evidence previously identified in the GRADE approach (risk of bias, indirectness, inconsistency, imprecision, reporting bias, magnitude of an effect, dose-response relation, and the direction of residual confounding) also apply when assessing the certainty of evidence from models. The assessment depends on the nature of model inputs and the model itself and on whether one is evaluating evidence from a single model or multiple models. We propose a framework for selecting the best available evidence from models: 1) developing de novo, a model specific to the situation of interest, 2) identifying an existing model, the outputs of which provide the highest certainty evidence for the situation of interest, either "off-the-shelf'' or after adaptation, and 3) using outputs from multiple models. We also present a summary of preferred terminology to facilitate communication among modeling and health care disciplines. Conclusion: This conceptual GRADE approach provides a framework for using evidence from models in health decision-making and the assessment of certainty of evidence from a model or models. The GRADE Working Group and the modeling community are currently developing the detailed methods and related guidance for assessing specific domains determining the certainty of evidence from models across health care-related disciplines (e.g., therapeutic decision-making, toxicology, environmental health, and health economics). (C) 2020 Published by Elsevier Inc.
Background: Regression discontinuity designs are non-randomized study designs that permit strong causal inference with relatively weak assumptions. Interest in these designs is growing but there is limited knowledge of the extent of their application in health. We aimed to conduct a comprehensive systematic review of the use of regression discontinuity designs in health research. Methods: We included studies that used regression discontinuity designs to investigate the physical or mental health outcomes of any interventions or exposures in any populations. We searched 32 health, social science, and gray literature databases (1 January 1960 to 1 January 2019). We critically appraised studies using eight criteria adapted from the What Works Clearinghouse Standards for regression discontinuity designs. We conducted a narrative synthesis, analyzing the forcing variables and threshold rules used in each study. Results: The literature search retrieved 7658 records, producing 325 studies that met the inclusion criteria. A broad range of health topics was represented. The forcing variables used to implement the design were age, socioeconomic measures, date or time of exposure or implementation, environmental measures such as air quality, geographic location, and clinical measures that act as a threshold for treatment. Twelve percent of the studies fully met the eight quality appraisal criteria. Fifteen percent of studies reported a prespecified primary outcome or study protocol. Conclusions: This systematic review demonstrates that regression discontinuity designs have been widely applied in health research and could be used more widely still. Shortcomings in study quality and reporting suggest that the potential benefits of this method have not yet been fully realized.
These guidelines incorporate the recent advances in chronic cough pathophysiology, diagnosis and treatment. The concept of cough hypersensitivity has allowed an umbrella term that explains the exquisite sensitivity of patients to external stimuli such a cold air, perfumes, smoke and bleach. Thus, adults with chronic cough now have a firm physical explanation for their symptoms based on vagal afferent hypersensitivity. Different treatable traits exist with cough variant asthma (CVA)/eosinophilic bronchitis responding to anti-inflammatory treatment and non-acid reflux being treated with promotility agents rather the anti-acid drugs. An alternative antitussive strategy is to reduce hypersensitivity by neuromodulation. Low-dose morphine is highly effective in a subset of patients with cough resistant to other treatments. Gabapentin and pregabalin are also advocated, but in clinical experience they are limited by adverse events. Perhaps the most promising future developments in pharmacotherapy are drugs which tackle neuronal hypersensitivity by blocking excitability of afferent nerves by inhibiting targets such as the ATP receptor (P2X3). Finally, cough suppression therapy when performed by competent practitioners can be highly effective. Children are not small adults and a pursuit of an underlying cause for cough is advocated. Thus, in toddlers, inhalation of a foreign body is common. Persistent bacterial bronchitis is a common and previously unrecognised cause of wet cough in children. Antibiotics (drug, dose and duration need to be determined) can be curative. A paediatric-specific algorithm should be used.
These guidelines incorporate the recent advances in chronic cough pathophysiology, diagnosis and treatment. The concept of cough hypersensitivity has allowed an umbrella term that explains the exquisite sensitivity of patients to external stimuli such a cold air, perfumes, smoke and bleach. Thus, adults with chronic cough now have a firm physical explanation for their symptoms based on vagal afferent hypersensitivity. Different treatable traits exist with cough variant asthma (CVA)/eosinophilic bronchitis responding to anti-inflammatory treatment and non-acid reflux being treated with promotility agents rather the anti-acid drugs. An alternative antitussive strategy is to reduce hypersensitivity by neuromodulation. Low-dose morphine is highly effective in a subset of patients with cough resistant to other treatments. Gabapentin and pregabalin are also advocated, but in clinical experience they are limited by adverse events. Perhaps the most promising future developments in pharmacotherapy are drugs which tackle neuronal hypersensitivity by blocking excitability of afferent nerves by inhibiting targets such as the ATP receptor (P2X3). Finally, cough suppression therapy when performed by competent practitioners can be highly effective. Children are not small adults and a pursuit of an underlying cause for cough is advocated. Thus, in toddlers, inhalation of a foreign body is common. Persistent bacterial bronchitis is a common and previously unrecognised cause of wet cough in children. Antibiotics (drug, dose and duration need to be determined) can be curative. A paediatric-specific algorithm should be used.
Abstract Background Education is widely associated with better physical and mental health, but isolating its causal effect is difficult because education is linked with many socioeconomic advantages. One way to isolate education’s effect is to consider environments where similar students are assigned to different educational experiences based on objective criteria. Here we measure the health effects of assignment to selective schooling based on test score, a widely debated educational policy. Methods In 1960s Britain, children were assigned to secondary schools via a test taken at age 11. We used regression discontinuity analysis to measure health differences in 5039 people who were separated into selective and non-selective schools this way. We measured selective schooling’s effect on six outcomes: mid-life self-reports of health, mental health, and life limitation due to health, as well as chronic disease burden derived from hospital records in mid-life and later life, and the likelihood of dying prematurely. The analysis plan was accepted as a registered report while we were blind to the health outcome data. Results Effect estimates for selective schooling were as follows: self-reported health, 0.1 worse on a 4-point scale (95%CI − 0.2 to 0); mental health, 0.2 worse on a 16-point scale (− 0.5 to 0.1); likelihood of life limitation due to health, 5 percentage points higher (− 1 to 10); mid-life chronic disease diagnoses, 3 fewer/100 people (− 9 to + 4); late-life chronic disease diagnoses, 9 more/100 people (− 3 to + 20); and risk of dying before age 60, no difference (− 2 to 3 percentage points). Extensive sensitivity analyses gave estimates consistent with these results. In summary, effects ranged from 0.10–0.15 standard deviations worse for self-reported health, and from 0.02 standard deviations better to 0.07 worse for records-derived health. However, they were too imprecise to allow the conclusion that selective schooling was detrimental. Conclusions We found that people who attended selective secondary school had more advantaged economic backgrounds, higher IQs, higher likelihood of getting a university degree, and better health. However, we did not find that selective schooling itself improved health. This lack of a positive influence of selective secondary schooling on health was consistent despite varying a wide range of model assumptions.
BACKGROUND:Global prevalence of overweight and obesity are alarming. For tackling this public health problem, preventive public health and policy actions are urgently needed. Some countries implemented food taxes in the past and some were subsequently abolished. Some countries, such as Norway, Hungary, Denmark, Bermuda, Dominica, St. Vincent and the Grenadines, and the Navajo Nation (USA), specifically implemented taxes on unprocessed sugar and sugar-added foods. These taxes on unprocessed sugar and sugar-added foods are fiscal policy interventions, implemented to decrease their consumption and in turn reduce adverse health-related, economic and social effects associated with these food products. OBJECTIVES:To assess the effects of taxation of unprocessed sugar or sugar-added foods in the general population on the consumption of unprocessed sugar or sugar-added foods, the prevalence and incidence of overweight and obesity, and the prevalence and incidence of other diet-related health outcomes. SEARCH METHODS:We searched CENTRAL, Cochrane Database of Systematic Reviews, MEDLINE, Embase and 15 other databases and trials registers on 12 September 2019. We handsearched the reference list of all records of included studies, searched websites of international organisations and institutions, and contacted review advisory group members to identify planned, ongoing or unpublished studies. SELECTION CRITERIA:We included studies with the following populations: children (0 to 17 years) and adults (18 years or older) from any country and setting. Exclusion applied to studies with specific subgroups, such as people with any disease who were overweight or obese as a side-effect of the disease. The review included studies with taxes on or artificial increases of selling prices for unprocessed sugar or food products that contain added sugar (e.g. sweets, ice cream, confectionery, and bakery products), or both, as intervention, regardless of the taxation level or price increase. In line with Cochrane Effective Practice and Organisation of Care (EPOC) criteria, we included randomised controlled trials (RCTs), cluster-randomised controlled trials (cRCTs), non-randomised controlled trials (nRCTs), controlled before-after (CBA) studies, and interrupted time series (ITS) studies. We included controlled studies with more than one intervention or control site and ITS studies with a clearly defined intervention time and at least three data points before and three after the intervention. Our primary outcomes were consumption of unprocessed sugar or sugar-added foods, energy intake, overweight, and obesity. Our secondary outcomes were substitution and diet, expenditure, demand, and other health outcomes. DATA COLLECTION AND ANALYSIS:Two review authors independently screened all eligible records for inclusion, assessed the risk of bias, and performed data extraction.Two review authors independently assessed the certainty of the evidence using the GRADE approach. MAIN RESULTS:We retrieved a total of 24,454 records. After deduplicating records, 18,767 records remained for title and abstract screening. Of 11 potentially relevant studies, we included one ITS study with 40,210 household-level observations from the Hungarian Household Budget and Living Conditions Survey. The baseline ranged from January 2008 to August 2011, the intervention was implemented on September 2011, and follow-up was until December 2012 (16 months). The intervention was a tax - the so-called 'Hungarian public health product tax' - on sugar-added foods, including selected foods exceeding a specific sugar threshold value. The intervention includes co-interventions: the taxation of sugar-sweetened beverages (SSBs) and of foods high in salt or caffeine. The study provides evidence on the effect of taxing foods exceeding a specific sugar threshold value on the consumption of sugar-added foods. After implementation of the Hungarian public health product tax, the mean consumption of taxed sugar-added foods (measured in units of kg) decreased by 4.0% (standardised mean difference (SMD) -0.040, 95% confidence interval (CI) -0.07 to -0.01; very low-certainty evidence). The study was at low risk of bias in terms of performance bias, detection bias and reporting bias, with the shape of effect pre-specified and the intervention unlikely to have any effect on data collection. The study was at unclear risk of attrition bias and at high risk in terms of other bias and the independence of the intervention. We rated the certainty of the evidence as very low for the primary and secondary outcomes. The Hungarian public health product tax included a tax on sugar-added foods but did not include a tax on unprocessed sugar. We did not find eligible studies reporting on the taxation of unprocessed sugar. No studies reported on the primary outcomes of consumption of unprocessed sugar, energy intake, overweight, and obesity. No studies reported on the secondary outcomes of substitution and diet, demand, and other health outcomes. No studies reported on differential effects across population subgroups. We could not perform meta-analyses or pool study results. AUTHORS' CONCLUSIONS:There was very limited evidence and the certainty of the evidence was very low. Despite the reported reduction in consumption of taxed sugar-added foods, we are uncertain whether taxing unprocessed sugar or sugar-added foods has an effect on reducing their consumption and preventing obesity or other adverse health outcomes. Further robustly conducted studies are required to draw concrete conclusions on the effectiveness of taxing unprocessed sugar or sugar-added foods for reducing their consumption and preventing obesity or other adverse health outcomes.
Background: Natural experiments and related study designs such as regression discontinuity (RD) are of increasing interest to researchers and decision makers because of their potential to address confounding and selection effects better than other observational study designs, with potentially greater applicability than controlled experiments. Research methods in health have been relatively slow to incorporate natural experiments compared to other fields such as economics and political science, but interest in these methods is growing rapidly. Objectives: This thesis aimed to (1) investigate the contribution of natural experimental designs to public health research, specifically the evaluation of public health interventions and environmental causes of disease and (2) explore how systematic review methods might be applied to make better use of natural experiments to inform public health and policy. Methods: The thesis comprises four case studies, including a systematic review of RD studies of health outcomes, a systematic review of RD studies of minimum legal drinking age (MLDA) legislation, development of a critical appraisal tool for RD studies, and a meta-review of endocrine-disrupting chemicals (EDCs) and breast cancer risk. Review protocols were registered in the PROSPERO database. Results: The first systematic review identified 181 RD studies of health outcomes which spanned a wide range of public health and policy questions, showing that this natural experimental design has been more widely applied than previously appreciated. Thematic analysis of the forcing variables and threshold rules used in these studies will aid in future applications of the design. The MLDA review of 17 econometric analyses identified challenges in the synthesis of natural experimental studies. The review identified evidence that MLDA has a causal effect on mortality and on alcohol-related hospital admissions. A ten-item checklist specific to the methodological requirements of RD designs was developed based on standards for RD produced by the What Works Clearinghouse; only 5% of the 181 studies met all ten criteria. The meta-review included 15 systematic reviews of EDCs and breast cancer risk; no primary studies in the review were identified as natural experiments. Conclusions: Natural experiments have the potential to support stronger causal inference through designs that address selection effects and confounding. For these designs to be translated into better evidence to inform decision-making, systematic reviews need to be able to identify and represent in detail the differences among non-randomised study designs. To do this requires further development of systematic review methods in order to synthesise results from econometric models and assess the quality of natural experimental studies.
Background Individualised breast cancer risk prediction models may be key for planning risk-based screening approaches. Our aim was to conduct a systematic review and quality assessment of these models addressed to women in the general population. Methods We followed the Cochrane Collaboration methods searching in Medline, EMBASE and The Cochrane Library databases up to February 2018. We included studies reporting a model to estimate the individualised risk of breast cancer in women in the general population. Study quality was assessed by two independent reviewers. Results are narratively summarised. Results We included 24 studies out of the 2976 citations initially retrieved. Twenty studies were based on four models, the Breast Cancer Risk Assessment Tool (BCRAT), the Breast Cancer Surveillance Consortium (BCSC), the Rosner & Colditz model, and the International Breast Cancer Intervention Study (IBIS), whereas four studies addressed other original models. Four of the studies included genetic information. The quality of the studies was moderate with some limitations in the discriminative power and data inputs. A maximum AUROC value of 0.71 was reported in the study conducted in a screening context. Conclusion Individualised risk prediction models are promising tools for implementing risk-based screening policies. However, it is a challenge to recommend any of them since they need further improvement in their quality and discriminatory capacity.
Individualised breast cancer risk prediction models may be key for planning risk-based screening approaches. Our aim was to conduct a systematic review and quality assessment of these models addressed to women in the general population. We followed the Cochrane Collaboration methods searching in Medline, EMBASE and The Cochrane Library databases up to February 2018. We included studies reporting a model to estimate the individualised risk of breast cancer in women in the general population. Study quality was assessed by two independent reviewers. Results are narratively summarised. We included 24 studies out of the 2976 citations initially retrieved. Twenty studies were based on four models, the Breast Cancer Risk Assessment Tool (BCRAT), the Breast Cancer Surveillance Consortium (BCSC), the Rosner & Colditz model, and the International Breast Cancer Intervention Study (IBIS), whereas four studies addressed other original models. Four of the studies included genetic information. The quality of the studies was moderate with some limitations in the discriminative power and data inputs. A maximum AUROC value of 0.71 was reported in the study conducted in a screening context. Individualised risk prediction models are promising tools for implementing risk-based screening policies. However, it is a challenge to recommend any of them since they need further improvement in their quality and discriminatory capacity.
Background: A new tool to assess Risk of Bias In Non-randomised Studies of Interventions (ROBINS-I) was published in Autumn 2016. ROBINS-I uses the Cochrane-approved risk of bias (RoB) approach and focusses on internal validity. As such, ROBINS-I represents an important development for those conducting systematic reviews which include non-randomised studies (NRS), including public health researchers. We aimed to establish the applicability of ROBINS-I using a group of NRS which have evaluated non-clinical public health natural experiments. Methods: Five researchers, all experienced in critical appraisal of non-randomised studies, used ROBINS-I to independently assess risk of bias in five studies which had assessed the health impacts of a domestic energy efficiency intervention. ROBINS-I assessments for each study were entered into a database and checked for consensus across the group. Group discussions were used to identify reasons underpinning lack of consensus for specific questions and bias domains. Results: ROBINS-I helped to systematically articulate sources of bias in NRS. However, the lack of consensus in assessments for all seven bias domains raised questions about ROBINS-I's reliability and applicability for natural experiment studies. The two RoB domains with least consensus were selection (Domain 2) and performance (Domain 4). Underlying the lack of consensus were difficulties in applying an intention to treat or per protocol effect of interest to the studies. This was linked to difficulties in determining whether the intervention status was classified retrospectively at follow-up, i.e. post hoc. The overall risk of bias ranged from moderate to critical; this was most closely linked to the assessment of confounders. Conclusion: The ROBINS-I tool is a conceptually rigorous tool which focusses on risk of bias due to the counterfactual. Difficulties in applying ROBINS-I may be due to poor design and reporting of evaluations of natural experiments. While the quality of reporting may improve in the future, improved guidance on applying ROBINS-I is needed to enable existing evidence from natural experiments to be assessed appropriately and consistently. We hope future refinements to ROBINS-I will address some of the issues raised here to allow wider use of the tool.