Background:Multiple sclerosis is an immune-mediated inflammatory disease, causing long-term disability in young adults. Most cases begin as relapsing-remitting multiple sclerosis. Some people have a form of relapsing-remitting multiple sclerosis known as highly active relapsing-remitting multiple sclerosis, defined as multiple sclerosis with unchanged or increased disease activity despite prior treatment with at least one disease-modifying therapy. Objectives:To appraise the clinical and cost-effectiveness of natalizumab [Tysabri® (Biogen, Cambridge, MA, USA)] and natalizumab biosimilar [Tyruko® (Sandoz)] for treating highly active relapsing-remitting multiple sclerosis compared to other disease-modifying therapy. Design:Systematic review with network meta-analysis and economic model. Searches last updated in April 2024. Results:We included 42 studies (22,409 participants): 40 in people with relapsing-remitting multiple sclerosis and 2 in highly active relapsing-remitting multiple sclerosis. Six studies also reported data separately for highly active relapsing-remitting multiple sclerosis. Only four studies evaluated natalizumab or natalizumab biosimilar; none provided data on those with highly active relapsing-remitting multiple sclerosis. Follow-up ranged from 4 to 36 (median 24) months. Most interventions reduced relapses (39 studies, 17 interventions) and magnetic resonance imaging lesions (19 studies, 11 interventions for gadolinium enhancing lesions and 17 studies, 12 interventions for T2-weighted lesions) compared to placebo. Alemtuzumab, ocrelizumab, cladribine, natalizumab, fingolimod and peginterferon beta-1a reduced disease progression compared to placebo (15 studies, 12 interventions). There were no differences in any adverse events (24 studies, 16 interventions), serious adverse events (31 studies, 15 interventions) or treatment-related adverse events (8 studies, no network meta-analysis) for any intervention compared to placebo. Fingolimod, glatiramer acetate, interferon beta-1a, interferon beta-1b and peginterferon beta-1a were associated with an increased treatment discontinuation (29 studies, 13 interventions). There was little evidence for a difference in quality of life. There was no evidence of a difference between natalizumab and natalizumab biosimilar for relapse rates [rate ratio 0.65 (95% credible interval 0.33 to 1.23], gadolinium enhancing lesions [hazard ratio 1.29 (0.69 to 2.37)], T2-weighted lesions [hazard ratio 1.07 (0.73 to 1.57)], any adverse events [hazard ratio 1.06 (0.77 to 1.46)] or treatment discontinuation [hazard ratio 0.48 (0.13 to 1.76)]. Data in highly active relapsing-remitting multiple sclerosis were available for fingolimod, ocrelizumab, alemtuzumab, cladribine, interferon beta, autologous haematopoietic stem cell treatment and placebo. We also included one study on natalizumab conducted in a population that was close to our definition of highly active relapsing-remitting multiple sclerosis. All interventions except interferon beta-1a were associated with reduced relapse risk compared to placebo (six studies; seven interventions). Compared with natalizumab-intravenous, natalizumab biosimilar-intravenous and natalizumab subcutaneous, all treatments had greater net benefit at £20,000-30,000/quality-adjusted life-year, with the only exception being ocrelizumab, which had lower net benefits. Costs were generally higher on natalizumab than other treatments, though there was no difference in quality-adjusted life-years with 95% credible interval completely overlapping. The results and conclusions were unchanged under all sensitivities. VOI analysis found that the greatest contributor to decision uncertainty was the effectiveness of treatments. Conclusions:There is no direct evidence on the effectiveness of natalizumab or its biosimilar in patients with highly active relapsing-remitting multiple sclerosis. Limited data suggest similar effectiveness in patients with relapsing-remitting multiple sclerosis. The economic model found that natalizumab and natalizumab biosimilar were not cost-effective compared to any of the included comparators in highly active relapsing-remitting multiple sclerosis, with similar quality-adjusted life-years but higher costs, with the only exception being ocrelizumab. Future work:There is need for studies of natalizumab and natalizumab biosimilar in people with highly active relapsing-remitting multiple sclerosis. Study registration:The study is registered as PROSPERO CRD42024556838. Funding:This award was funded by the National Institute for Health and Care Research (NIHR) Evidence Synthesis programme (NIHR award ref: NIHR165943) and is published in full in Health Technology Assessment; Vol. 30, No. 60. See the NIHR Funding and Awards website for further award information.
Background Substantial variation in testing rates in adults with hypertension across UK primary care suggest that patients are not receiving optimal monitoring. Aim To develop a minimal set of evidence-based blood tests for adults with hypertension. Design and setting This was a rapid review, routine data analyses, and consensus study. It was set in primary care. Method This study examined the rationale and evidence for tests recommended by guidelines or used commonly in adults with hypertension using stepwise rapid evidence reviews. A consensus group, including clinicians and patients, voted to include or exclude each test in the testing panel based on the evidence. If there was no consensus (>80%), additional evidence was sought through rapid reviews or analyses of primary care records, which was subject to further voting. Results Sixteen routinely ordered tests were identified. The study found consistent, good evidence that estimated glomerular filtration rate (eGFR) to detect chronic kidney disease and glycosylated haemoglobin (HbA1c) to detect diabetes is beneficial for patients. The study found no or inconsistent evidence of the benefit of routinely measuring lipids, electrolytes, haemoglobin, thyroid function, clotting biomarkers, calcium, ferritin, folate acid, or vitamin B12. Good evidence was found that there is no benefit in routinely monitoring liver function, inflammation markers, or brain natriuretic peptide. Conclusion A minimal set of evidence-based blood tests to monitor adults with hypertension was identified. This panel includes eGFR, HbA1c, potassium, and sodium. Implementing these recommendations could reduce harms associated with unwarranted variation in care. Further research is needed to clarify the role of tests with inconsistent evidence and determine the optimal frequency of testing.
Background:National Institute for Health and Care Excellence technology appraisals assess the effectiveness and cost-effectiveness of medicines at a single point in the treatment pathway. However, for some disease areas, such as non-small cell lung cancer, there are many recommendations, making it difficult to use National Institute for Health and Care Excellence guidance. The treatment pathway for metastatic stage 4 non-small cell lung cancer can be divided into decision points (nodes) based on histology (squamous or non-squamous), programmed death-ligand 1 expression, presence of tumour mutations and line of therapy. The National Institute for Health and Care Excellence commissioned this pilot to assess the potential of taking a 'pathways approach' to technology appraisals. The aim was to build a single disease-specific cost-effectiveness model for metastatic stage 4 non-small cell lung cancer patients not eligible for targeted therapies at first line, that can be updated with economic and clinical data as required. Methods:We conducted a systematic review (searches last updated 11 July 2025) and network meta-analysis of treatment efficacy and safety at each decision node in the pathway. We used flexible fractional polynomial models for primary outcome progression-free survival, required for a model of treatment sequences. We built a novel cost-effectiveness model that compared sequences of treatments, and was populated using network meta-analyses for progression-free survival, data on overall survival after last-line therapy, evidence on treatment sequences from an analysis of systemic anticancer therapy data, and quality of life, cost and resource use estimates from previous technology appraisals. Drug list prices were used, but confidential discounts are available. Results:We included 15 randomised controlled trials and 1 single-arm study in the review, judged as some concerns or low risk of bias. Immunotherapies, in combination with doublet platinum chemotherapy, were most effective first-line treatments, although with higher adverse event rates. Immunotherapy monotherapies were most effective at second line, unless patients were suitable for targeted therapies. Sequences starting with atezolizumab + bevacizumab + doublet platinum chemotherapy had similar costs and quality-adjusted life-years to sequences starting with pembrolizumab + doublet platinum chemotherapy. Sequences starting with pemetrexed + platinum chemotherapy had the lowest costs but also the lowest total quality-adjusted life-years. Sequences for non-squamous non-small cell lung cancer, programmed death-ligand 1 ≥ 50%:Sequences starting with pembrolizumab + doublet platinum chemotherapy had highest quality-adjusted life-years, but higher costs compared to other sequences. Sequences starting with pemetrexed + platinum chemotherapy had the lowest cost and the lowest number of quality-adjusted life-years. Sequences for squamous non-small cell lung cancer, programmed death-ligand 1 < 50%:Sequences starting with pembrolizumab + doublet platinum chemotherapy had higher quality-adjusted life-years and higher costs than sequences starting with platinum chemotherapy. Sequences for squamous non-small cell lung cancer, programmed death-ligand 1 ≥ 50%:Sequences starting with atezolizumab had the highest predicted quality-adjusted life-year gains, slightly higher than for sequences starting with pembrolizumab. Sequences starting with platinum chemotherapy had the lowest quality-adjusted life-years, but the lowest total costs. Patients in the systemic anticancer therapy analysis had a shorter time on treatment than those in trials, resulting in lower treatment costs and causing the immunotherapy sequences to appear more cost effective. Conclusions:Our model can be used to estimate either the most cost-effective sequence of treatments or the most cost-effective treatment at a given point in the pathway, although our results are based on drug list prices and these would need to be updated to draw conclusions about the relative cost-effectiveness of different treatment sequences. We were able to use real-world data from systemic anticancer therapy to estimate model parameters and sequences that reflect clinical practice. Our model can readily be updated and used as a reference model for metastatic non-small cell lung cancer. Study registration:The study is registered as PROSPERO CRD42023470119. Funding:This award was funded by the National Institute for Health and Care Research (NIHR) Evidence Synthesis programme (NIHR award ref: NIHR136097) and is published in full in Health Technology Assessment; Vol. 30, No. 46. See the NIHR Funding and Awards website for further award information.
OBJECTIVE:For surveillance for hepatocellular carcinoma (HCC) to be effective, tests-including imaging, serological biomarkers (conventional and genomic) and algorithms combining multiple tests-must identify early-stage tumours. We aimed to identify, appraise and synthesise studies reporting the accuracy of all such tests in people with cirrhosis. DESIGN:Systematic review and network meta-analysis of diagnostic test accuracy (NMA-DTA) data. DATA SOURCES:MEDLINE and Embase (2005 to September 2025) and a published Cochrane review. ELIGIBILITY CRITERIA:English-language, post-2005, one-gate or two-gate studies quantifying diagnostic accuracy of tests to detect HCC in populations wholly comprising people with cirrhosis, excluding those with pre-existing signs and symptoms of HCC. DATA EXTRACTION AND SYNTHESIS:Data extracted by one reviewer, checked by a second and made available in an open-access database. We assessed risk of bias using QUADAS-2. We synthesised data using Bayesian NMA-DTA, accounting for tumour stage and incorporating continuous tests across all possible thresholds. RESULTS:We included 170 studies (62 643 participants). Of 115 index tests, 97 were amenable to NMA-DTA. Ultrasound appears no better than alpha-fetoprotein at detecting very-early-stage HCC (sensitivity 0.34 (95% CrI 0.21 to 0.55) vs 0.39 (95% CrI 0.28 to 0.47)), only becoming superior as stage advances. Least affected by stage are contrast-enhanced MRI (sensitivity 0.70 (95% CrI 0.50 to 0.84) very-early; 0.86 (95% CrI 0.73 to 0.94) early; 0.90 (95% CrI 0.67 to 0.98) advanced) and CT (0.68 (95% CrI 0.26 to 0.93) very-early; 0.77 (95% CrI 0.29 to 0.95) early; 0.92 (95% CrI 0.49 to 0.99) advanced). No genomic biomarkers show convincing improvements over combinations of conventional blood-markers. Most studies are at high risk of bias, but conclusions are robust when restricting to studies with favourable methodological characteristics. CONCLUSIONS:Using advanced synthesis methods, we found that tests have low sensitivity for detecting early-stage HCC. Given rising prevalence of cirrhosis and HCC, we need better tests and a stronger evidence-base to inform optimal surveillance strategies. PROSPERO REGISTRATION NUMBER:CRD42022357163.
There is substantial variation in routine monitoring blood tests for chronic kidney disease (CKD) stage G3, suggesting that many individuals are not receiving optimal care. This Review aimed to develop evidence-based testing panels for CKD stage G3. Commonly used and guideline-recommended tests were assessed through filtering questions and rapid evidence reviews to determine their clinical utility. A consensus group of general practitioners, primary care nurses, a renal consultant, and patient representatives voted on test inclusion. Wherever evidence was insufficient, additional rapid reviews and routine data analysis were undertaken. The Review found evidence to support routine testing of estimated glomerular filtration rate (eGFR), haemoglobin, and glycated hemoglobin (HbA1c). No evidence of benefit was found for other commonly ordered tests, including urea, lipids, vitamin B12, ferritin, folate, liver function, electrolytes, vitamin D, calcium, thyroid function, clotting, C-reactive protein, erythrocyte sedimentation rate, and B-type natriuretic peptide. These findings offer a simplified approach to CKD stage G3 monitoring, reducing unwarranted variation and focusing resources on evidence-based, sustainable primary care.
Introduction: Diagnostic test accuracy (DTA) systematic reviews bring together findings from DTA studies to summarise the accuracy of a diagnostic test. Studies included in a DTA review should be assessed for risk of bias and applicability concerns, because a review of biased studies, or studies that do not apply directly to the review question, could result in misleading conclusions. Studies are most commonly assessed with the QUADAS-2 tool. Anecdotal evidence has suggested that researchers sometimes struggle to differentiate between risk of bias and applicability. Here, we investigate this distinction for the patient selection and index test domains. Objectives: To develop a framework for assessing the applicability of the study target population and the study index test to the review defined target populations and index test(s). We aimed to explore review authors’ applicability assessments for the QUADAS-2 patient selection and index test domains, to inform the framework. Methods: DTA reviews were eligible for inclusion if they were published in the Cochrane Library, had used QUADAS-2, and had at least one study rated as “high concerns” for applicability of the patient selection or index test domain. Review selection was checked by a second reviewer. From each review, we extracted article identifiers such as title, authors and publication date, extracted the primary objectives an elements of review questions: population, index test(s), target condition, reference standard. For each review, we extracted the rationale provided by the authors for “high concerns” applicability judgements for the respective domains. We also extracted author’s rules how to assess applicability concern with QUADAS-2, whenever it was tailored to the review topic. One reviewer assessed these rationales and rules, which were verified by another. We also recorded any other issues that arose as part of the applicability assessment, such as suboptimal reporting or erroneous applicability assessment. Two reviewers categorized the rationales inductively into themes, which will be discussed in the QUADAS steering group. The final framework will be informed by the identified themes and thorough group discussions. Results: This review is in progress. Of the 186 available Cochrane DTA reviews, 123 met our inclusion criteria: 110 for the patient selection domain, 75 the index test domain. The data extraction process is ongoing and final results will be presented at the conference. The majority of the more recent Cochrane DTA reviews include guidance tailored to the review’s topic. Review authors typically have a broad review question, whereas signaling questions tailored to the review’s topic are more restrictive. Several themes emerged in the patient selection domain. These include: study setting not matching the review question, study unit not matching the review study unit (lesion level versus patient level); study’s target condition not meeting the review’s target condition (in-vitro versus in-vivo); study’s disease spectrum not covering the review’s targeted disease spectrum (i.e. due to sampling methods, choice of selection criteria; inappropriate exclusions) and study’s indication for testing not sufficiently matching the reviews indication. In the index test domain, themes identified so far include a mismatch in the study and the review with respect to: index test technology; test protocol; thresholds to define the target condition; clinical background of the examiner; experience of the examiner; use of consensus test interpretation rather than use of the interpretation of a single examiner, use of clinical information. In both domains, a number of reviews misunderstood QUADAS-2 applicability or provided insufficient information on the rationale for high concerns. Once all themes are identified, the framework will be developed and presented as SISMEC. Conclusions: Clear sources of applicability concerns are identifiable, but several Cochrane review authors struggle to adequately identify and report them. At SISMEC, we will present the applicability framework to guide review authors in their assessment of applicability concerns for the QUADAS patient selection and index test domains
OBJECTIVES:To explore patient, carer and clinician experiences of the QbTest and its impact on patient outcomes for attention deficit hyperactivity disorder (ADHD) diagnosis and medication management. DESIGN:Mixed-methods systematic review. DATA SOURCES:MEDLINE, EMBASE, PsycINFO, CINAHL, ClinicalTrials.gov and WHO ICTRP (from inception to September 2024). STUDY SELECTION:Primary studies, of any design, that evaluated any version of the QbTest (QbMini <5 years, QbTest 6-12 or 12-60 years, QbCheck for remote assessment via webcam or QbMT smartphone version), for ADHD diagnosis and/or medication management and provided data on any of the following outcomes, were eligible: time to assessment/diagnostic decision, use of services, impact on clinical decision-making, healthcare professionals' confidence in assessment, intervention use, morbidity, mortality, health-related quality of life, cost, ease of use, experience and acceptability of the test to patients, carers and clinicians. DATA EXTRACTION AND SYNTHESIS:Two reviewers independently screened titles and abstracts and assessed potentially relevant reports for inclusion. One reviewer conducted data extraction and risk of bias (RoB) assessment, checked by a second reviewer. Mixed-methods synthesis followed the convergent-integrated approach. RESULTS:We identified 10 eligible studies (9 QbTest; 1 QbCheck), including 1 randomised controlled trial (RCT), 2 feasibility RCTs, 5 before-and-after studies, 1 mixed-methods study and 1 diagnostic study. Most studies enrolled children in the UK and included surveys or interviews with patients, carers or clinicians. The RCT and before-and-after studies were judged at high/serious RoB. Six survey components and two qualitative interview components were judged at some concerns of RoB. We identified one ongoing study of the QbMT and no studies for QbMini. We organised themes emerging from the qualitative synthesis into two broad conceptual categories: views around the helpfulness of the QbTest (contribution to ADHD diagnosis, treatment decision-making, communication with caregivers) and barriers to QbTest implementation (practical barriers and acceptability of the test to patients and caregivers). Findings suggested that the addition of the QbTest may reduce time to diagnosis, improve clinician confidence in the diagnostic decision, increase the proportion of patients with a diagnostic decision and reduce cost and number of clinic appointments. The QbTest appeared to be generally well received by clinicians, patients and carers. However, barriers to test implementation were reported. Clinicians cited staffing, room requirements and issues with technology, and patients highlighted the test length and repetitive nature. Little data exist on the use of the QbTest for medication management. CONCLUSIONS:The available evidence suggests the QbTest may be a useful addition to ADHD assessment in children and young people. Further well-designed RCTs with qualitative substudies are required to assess the impact of the QbTest on patient outcomes, user experience and cost, particularly for medication management and in adults, where evidence is scarce. Such RCTs should include economic analyses, direct comparisons to other continuous performance tests with motion trackers and subgroup analyses including age, sex, ethnicity and comorbidities. PROSPERO REGISTRATION NUMBER:CRD42023482963.
BACKGROUND:After testing, ensuring test results are communicated and actioned is important for patient safety, with failure or delay in diagnosis the most common cause of malpractice claims in primary care worldwide. Identifying interventions to improve test communication from the decision to test through to sharing of results has important implications for patient safety, GP workload, and patient engagement.AIM:To assess the factors around communication of blood test results between primary care providers (for example GPs, nurses, reception staff) and their patients and carers.DESIGN & SETTING:A mixed methods systematic review including primary studies involving communication of blood test results in primary care.METHOD:The review will use a segregated convergent synthesis method. Qualitative information will be synthesised using a meta-aggregative approach, and quantitative data will be meta-analysed or synthesised if pooling of studies is appropriate and data are available. If not, data will be presented in tabular and descriptive summary form.CONCLUSION:This review has the potential to provide conclusions about blood test result communication interventions and factors important to stakeholders, including barriers and facilitators to improved communication.
Introduction: Diagnostic test accuracy (DTA) systematic reviews bring together findings from DTA studies to summarise the accuracy of a diagnostic test. Studies included in a DTA review should be assessed for risk of bias and applicability concerns, because a review of biased studies, or studies that do not apply directly to the review question, could result in misleading conclusions. Studies are most commonly assessed with the QUADAS-2 tool. Anecdotal evidence has suggested that researchers sometimes struggle to differentiate between risk of bias and applicability. Here, we investigate this distinction for the patient selection and index test domains. Objectives: To develop a framework for assessing the applicability of the study target population and the study index test to the review defined target populations and index test(s). We aimed to explore review authors’ applicability assessments for the QUADAS-2 patient selection and index test domains, to inform the framework. Methods: DTA reviews were eligible for inclusion if they were published in the Cochrane Library, had used QUADAS-2, and had at least one study rated as “high concerns” for applicability of the patient selection or index test domain. Review selection was checked by a second reviewer. From each review, we extracted article identifiers such as title, authors and publication date, extracted the primary objectives an elements of review questions: population, index test(s), target condition, reference standard. For each review, we extracted the rationale provided by the authors for “high concerns” applicability judgements for the respective domains. We also extracted author’s rules how to assess applicability concern with QUADAS-2, whenever it was tailored to the review topic. One reviewer assessed these rationales and rules, which were verified by another. We also recorded any other issues that arose as part of the applicability assessment, such as suboptimal reporting or erroneous applicability assessment. Two reviewers categorized the rationales inductively into themes, which will be discussed in the QUADAS steering group. The final framework will be informed by the identified themes and thorough group discussions. Results: This review is in progress. Of the 186 available Cochrane DTA reviews, 123 met our inclusion criteria: 110 for the patient selection domain, 75 the index test domain. The data extraction process is ongoing and final results will be presented at the conference. The majority of the more recent Cochrane DTA reviews include guidance tailored to the review’s topic. Review authors typically have a broad review question, whereas signaling questions tailored to the review’s topic are more restrictive. Several themes emerged in the patient selection domain. These include: study setting not matching the review question, study unit not matching the review study unit (lesion level versus patient level); study’s target condition not meeting the review’s target condition (in-vitro versus in-vivo); study’s disease spectrum not covering the review’s targeted disease spectrum (i.e. due to sampling methods, choice of selection criteria; inappropriate exclusions) and study’s indication for testing not sufficiently matching the reviews indication. In the index test domain, themes identified so far include a mismatch in the study and the review with respect to: index test technology; test protocol; thresholds to define the target condition; clinical background of the examiner; experience of the examiner; use of consensus test interpretation rather than use of the interpretation of a single examiner, use of clinical information. In both domains, a number of reviews misunderstood QUADAS-2 applicability or provided insufficient information on the rationale for high concerns. Once all themes are identified, the framework will be developed and presented as SISMEC. Conclusions: Clear sources of applicability concerns are identifiable, but several Cochrane review authors struggle to adequately identify and report them. At SISMEC, we will present the applicability framework to guide review authors in their assessment of applicability concerns for the QUADAS patient selection and index test domains
Objectives: This scoping review was undertaken to identify risk prediction models and pre-operative predictors of surgical site infection (SSI) in adult cardiac surgery. A particular focus was on the identification of novel predictors that could underpin the future development of a risk prediction model to identify individuals at high risk of SSI, and therefore guide a national SSI prevention strategy. Methods: A scoping review to systematically identify and map out existing research evidence on pre-operative predictors of SSI was conducted in two stages. Stage 1 reviewed prediction modelling studies of SSI in cardiac surgery. Stage 2 identified primary studies and systematic reviews of novel cardiac SSI predictors. Results: The search identified 7887 unique reports; 7154 were excluded at abstract screening and 733 were selected for full-text assessment. Twenty-nine studies (across 30 reports) were included in Stage 1 and reported the development (N1/414), validation (N1/413), or both development and validation (N1/42) of 52 SSI risk prediction models including 67 different pre-operative predictors. The remaining 703 reports were reassessed in Stage 2; 49 studies met the inclusion criteria, and 56 novel pre-operative predictors that have not been assessed previously in models were identified.
OBJECTIVES:Assessment of the applicability of primary studies is an essential but often a challenging aspect of systematic reviews of diagnostic test accuracy studies (DTA reviews). We explored review authors' applicability assessments for the QUADAS-2 reference standard domain within Cochrane DTA reviews. We highlight applicability concerns, identify potential issues with assessment, and develop a framework for assessing the applicability of the target condition as defined by the reference standard. STUDY DESIGN AND SETTING:Methodological review. DTA reviews in the Cochrane Library that used QUADAS-2 and judged applicability for the reference standard domain as "high concern" for at least one study were eligible. One reviewer extracted the rationale for the "high concern" and this was checked by a second reviewer. Two reviewers categorized the rationale inductively into themes, and a third reviewer verified these. Discussions regarding the extracted information informed framework development. RESULTS:We identified 50 eligible reviews. Five themes emerged: study uses different reference standard threshold to define the target condition (six reviews), misclassification by the reference standard in the study such that the target condition in the study does not match the review question (11 reviews), reference standard could not be applied to all participants resulting in a different target condition (five reviews), misunderstanding QUADAS-2 applicability (seven reviews), and insufficient information (21 reviews). Our framework for researchers outlines four potential applicability concerns for the assessment of the target condition as defined by the reference standard: different sub-categories of the target condition, different threshold used to define the target condition, reference standard not applied to full study group, and misclassification of the target condition by the reference standard. CONCLUSION:Clear sources of applicability concerns are identifiable, but several Cochrane review authors struggle to adequately identify and report them. We have developed an applicability framework to guide review authors in their assessment of applicability concerns for the QUADAS reference standard domain. PLAIN LANGUAGE SUMMARY:What is the problem? Doctors use tests to help to decide if a person has a certain condition. They want to know how accurate the test is before they use it. This means how well it can tell people who have the condition from people who do not have it. This information can be found in "diagnostic systematic reviews". Diagnostic systematic reviews start with a research question. They bring together findings from studies that have already been done to try to answer this question. It is important for researchers to check that the studies match the review question. This is called an "applicability assessment". For example, if the review looks at children, it is important to check whether the studies look at children too. There is a tool named "QUADAS-2" that can be used to check how well the studies match the review question. This can be hard to do, and there are not many examples available to help people. What did we do? We wanted to know more about how people judge applicability concerns using the QUADAS-2 tool. We also wanted to write guidance to support judgments about applicability. What did we find? We found examples of how people are doing applicability assessment. Many reviews have not done it right, and we explain why this is. We also made guidance to help people with applicability assessment.
Objective Recent evidence supports diagnosing coeliac disease without biopsy in patients with significantly elevated tissue transglutaminase (IgA-tTG) antibodies. However, the implementation of this no-biopsy approach relies on accurate and consistent serological testing across laboratories. In this nationwide survey, we aimed to evaluate the availability and variability of coeliac disease testing across the UK.Methods We conducted a cross-sectional telephone survey of biomedical scientists and laboratory managers from National Health Service trusts and health boards across England, Wales, Scotland, and Northern Ireland. Data collected included assay types, reporting methods, upper limit of normal (ULN) thresholds, turnaround times, total IgA testing, and anti-endomysial antibodies (EMAs) availability.Results A total of 356 sites were approached, with a 96% response rate (n=342). Of responding sites, 177 performed coeliac serology tests in-house, while 165 transferred samples externally. Among sites performing tests, 12 different IgA-tTG assays were identified, with considerable variability in ULN thresholds ranging from 3 to 30 IU/mL, even within laboratories using the same assays. The median turnaround time for IgA-tTG results was 7 days (range 1–21 days). Only 43% of laboratories routinely measured total IgA when IgA-tTG was requested. EMA testing was available in 83% of laboratories.Conclusion Significant variability exists in coeliac serology testing across UK laboratories which poses a challenge for the implementation of the no-biopsy approach in clinical practice. Efforts to standardise serological testing are urgently needed. Until such standardisation is achieved, local assay validation remains critical.
Background:Attention deficit hyperactivity disorder is characterised by inattention, impulsivity and hyperactivity. Diagnosis is complex and time-consuming. Medication requires careful selection and dose titration. Technologies for objective measures of attention deficit hyperactivity disorder that use motion sensors to measure hyperactivity ('sensor continuous performance tests') may help improve the diagnostic process and medication management when used in addition to clinical assessment. Objective:To determine whether sensor continuous performance tests are clinically effective and cost-effective to the National Health Service. Specific objectives were to determine the effectiveness of sensor continuous performance tests for: diagnosis of attention deficit hyperactivity disorder in people referred with suspected attention deficit hyperactivity disorder diagnosis of attention deficit hyperactivity disorder in people referred with suspected attention deficit hyperactivity disorder for whom current assessment cannot reach a diagnosis during initial dose titration and treatment decisions for people with attention deficit hyperactivity disorder evaluating treatment effectiveness during long-term treatment monitoring for people with attention deficit hyperactivity disorder. Design:Systematic review and economic model (searches completed 17 November 2023). Results:Objective 1 [29 studies - 25 QbTest (QbTech Ltd., Stockholm, Sweden), 2 EF Sim (Peili Vision, Oulu, Finland) and 2 Nesplora Kids (Giunti Psychometrics, Florence, Italy)]: most evidence was in children. The AQUA trial was the only study to evaluate the QbTest in combination with clinical assessment and included a comparison with clinical assessment alone. Accuracy was similar and there was no statistical evidence of a difference between groups (p = 0.14), but the study was at high risk of bias. The AQUA trial reported that adding QbTest to the diagnostic process resulted in fewer appointments to reach a diagnosis, reduced consultation time, greater clinician confidence and exclusion of the diagnosis in a more children. Findings were supported by limited data from uncontrolled before-after studies. Qualitative and survey data reported increased clinician confidence in clinical decision-making, reduced time to diagnostic decision and improved communication. Barriers to implementation included staffing, training, technology requirements and length and repetitive content of the test. We found that using QbTest in addition to clinical assessment was likely cost-effective due to the reduced time waiting for assessment, reduced appointments until diagnosis and a higher proportion receiving treatment benefits. Objective 3 (six studies): All evaluated QbTest and most had concerns with risk of bias. Qualitative and survey data suggested that healthcare staff and families valued the QbTest for dose titration, checking medication utility and improving medication adherence. Some data suggested that results may not increase patient understanding and some clinicians highlighted logistical challenges. No studies were identified for objectives 2 and 4. Conclusions:Our results suggest that QbTesting as part of the diagnostic workup for attention deficit hyperactivity disorder in children (age < 18 years), when used in combination with clinical assessment, may be cost-effective. This finding was robust to nearly all assumptions made in the model. There are insufficient data on other sensor continuous performance tests in adults or on medication management. Future work:Diagnostic accuracy study evaluating comparing each of the sensor continuous performance tests plus clinical assessment. This should consider accuracy across different patient subgroups. Trial comparing patient outcomes and process measures in adults and children tested with and without sensor continuous performance tests with separate analyses for difficult-to-diagnose patients. Trial evaluating the role of sensor continuous performance tests in medication management, including long-term follow-up. Limitations:Lack of good-quality data on all tests, both for diagnosis and medication management, particularly when evaluated in combination with clinical information. Study registration:This study is registered as PROSPERO CRD42023482963. Funding:This award was funded by the National Institute for Health and Care Research (NIHR) Evidence Synthesis programme (NIHR award ref: NIHR136009) and is published in full in Health Technology Assessment; Vol. 29, No. 58. See the NIHR Funding and Awards website for further award information.
BACKGROUND:When monitoring long-term conditions, both over- and undertesting can risk patient harm and increase healthcare costs. AIM:To evaluate the evidence base for type 2 diabetes mellitus (T2DM) monitoring tests and develop methods for creating evidence-based testing strategies. DESIGN AND SETTING:Rapid reviews were conducted and a consensus process then used to evaluate the evidence base within primary care settings. METHOD:The authors identified tests that are recommended or used commonly to monitor T2DM. Filtering questions were created to examine the rationale for use of each test, which were answered by stepwise rapid reviews of evidence cited by guidelines, systematic reviews, and individual studies. A consensus group of patient representatives and clinicians voted whether tests should be included or excluded based on the evidence or whether further evidence was needed. RESULTS:Of 15 tests, only haemoglobin A1c, to monitor disease progression and treatment response, and estimated glomerular filtration rate, to detect chronic kidney disease, have a strong evidence base. Based on available evidence and consensus group feedback, routinely testing for fructosamine to monitor disease progression; thyroid function, vitamin B12, ferritin, folate, clotting, bone profile, C-reactive protein, erythrocyte sedimentation rate, and B-type natriuretic peptide; and liver function for adverse treatment effects of metformin was deemed unnecessary. The study found insufficient evidence for lipids and haemoglobin to screen for secondary conditions, and for vitamin B12 to screen for adverse effects in those taking metformin. CONCLUSION:The study found that the evidence base for most T2DM monitoring tests is weak or absent. Clinicians should avoid non-evidence-based tests unless there are additional clinical indications for testing. Standardised evidence-based testing panels for T2DM and other long-term conditions could reduce unnecessary testing.
Introduction: The QUADAS‐2 tool, published in 2011, was designed to evaluate the risk of bias and applicability of diagnostic test accuracy (DTA) studies. The publication reporting QUADAS-2 has been cited over 12, 000 times and it is the recommended tool to assess risk of bias and applicability of studies for major HTA organizations. Although feedback on QUADAS-2 has generally been positive, some signaling questions have been identified as problematic and the tool could be improved based on features included in more recently developed tools. Objectives: To update QUADAS-2 to develop the new QUADAS-3 tool. Methods: We established a core-group of methodological experts to lead the development of the QUADAS-3 tool supported by a wider steering group. We followed the following steps: • Summarised modifications made to QUADAS-2 for the Cochrane Handbook • Web-based survey of reviewers that have used QUADAS-2 • Considered developments from more recent tools in terms of tool structure and implementation • Undertook a review of methodological studies that had evaluated QUADAS-2 • Undertook a review of 50 Cochrane DTA reviews to highlight challenges with the assessment of applicability We have produced a draft tool which has undergone piloting. The results of the piloting, which also included a comparison of the use of signalling questions with signalling statements, was used to inform the final version of the tool. Results The new tool follows a similar structure to the QUADAS-2 tool but with some major updates. Key changes include: • An option to define separate synthesis questions rather than just a single review question • A new section on defining the ideal test accuracy trial for each synthesis question • Assessment of risk of bias and applicability at the accuracy estimate level rather than the study level • A change in answers to signaling questions to include options of “probably yes” and “probably no” and to replace “unclear” with “no information” • Replacement of “Flow and Timing domain” with new “Analysis” domain • Changes to some signaling questions • Inclusion of a section for judging overall risk of bias and applicability (across domains) Conclusions: The QUADAS-3 tool incorporates several changes compared to the previous version (QUADAS-2) which we hope will improve its validity, usability, and usefulness. QUADAS-3 will be introduced at the conference and the results of piloting discussed.
IntroductionAssurance that supporting evidence is based on valid and unbiased assessments, evaluated using rigorously developed risk of bias (validity assessment) tools, is fundamental to good decision-making. Among those available, selecting and correctly using the best tool that is fit-for-purpose is challenging. Collaboration across the global evidence synthesis and health technology assessment (HTA) communities promotes best practices and harmonizes tool use across jurisdictions.MethodsWe have established the LATITUDES Network (https://www.latitudes-network.org/), a publicly available website library of validity assessment tools and resources to guide decision-makers in selecting and applying tools appropriate for particular contexts, including informing HTA reimbursement decisions, clinical guideline development, and stand-alone evidence synthesis projects. The internationally representative leadership team comprises five evidence synthesis experts who have been supported by a competitively awarded academic innovations grant. The 23-member advisory panel representing five continents provided expertise to finalizing criteria for tool inclusion, identifying key tools, and suggesting inclusion of additional tools. Formal launch took place at the 2023 Cochrane Colloquium.ResultsOfficially launched in September 2023, the LATITUDES Network indexes validity assessment tools developed for healthcare studies in an online library. To date, 10 key tools are featured to help reviewers identify the optimal tool for their use. Nineteen additional tools have met all screening criteria and are also recommended. Information characterizing each tool (e.g., citation and training materials) is provided. Seven tools are currently under development. A mechanism for users to suggest new tools is provided. Additional tools and information on toolkits and online training materials, as well as links to courses and events, will be added over time.ConclusionsLATITUDES aims to be the primary resource that provides key information to reviewers conducting validity assessments for evidence synthesis, clinical guideline development, and HTA decision-making. It is intended to increase the robustness of evidence synthesis by improving the process of validity assessment, helping scientists use tools more effectively and efficiently, promoting best practices, and harmonizing validity assessment across the globe.
Background: Point of care tests (POCTs) have the potential to improve the urinary tract infection (UTI) diagnostic pathway, as they can provide a diagnosis quickly in near-patient settings, and some also identify causative pathogens/antimicrobial sensitivity. Objectives: To assess the clinical impact, accuracy, and technical characteristics of POCT for diagnosing UTI. Methods of data synthesis: Narrative summary and bivariate random effects meta-analyses to estimate summary sensitivity and specificity. Data sources: Five electronic databases, two clinical trial registries, study reports and review reference lists, and websites. Study eligibility criteria: Randomized controlled trials/non-randomized studies and diagnostic test accuracy studies published since 2000. Participants: People with suspected UTI. Tests: Rapid tests (results <40 minutes): Astrego PA-10 0 system, Lodestar DX, Uriscreen, UTRiPLEX. Culture tests (results <24 hours): Flexicult Human, ID Flexicult, Diaslide, Dipstreak, Chromostreak, Uricult, Uricult Trio, Uricult Plus. Reference standard: Any. Assessment of risk of bias: Risk of Bias-2, Quality Assessment of Diagnostic Accuracy Studies-2, Quality Assessment of Diagnostic Accuracy Studies-C. Results: Two randomized controlled trials evaluated Flexicult Human (one against standard care; one against ID Flexicult). No difference was reported in antibiotic use concordant with culture results (OR 0.84 95% CI 0.58-1.20) or appropriate antibiotic prescribing (OR 1.44 95% CI 1.03-1.99). Initial antibiotic prescribing was lower with Flexicult than standard care (OR 0.56 95% CI 0.35-0.88). No difference for other measures of antibiotic use, symptom duration, patient enablement, or resource use. Fifteen studies reported accuracy data. Limited data were available, with most POCT evaluated in single studies or not evaluated at all. Uriscreen (four studies), Uricult Trio (three studies), Flexicult Human (four studies), and ID Flexicult (two studies) had modest sensitivity and specificity. POCTs were easier to use and interpret than standard culture. Conclusions: There is currently insufficient evidence to support the use of POCTs in UTI diagnosis. Due to the rapid development of POCT, this review should be updated regularly. Eve Tomlinson, Clin Microbiol Infect 2024;30:197 (c) 2023 European Society of Clinical Microbiology and Infectious Diseases. Published by Elsevier Ltd. All rights reserved.
Objectives: To review the findings of studies that have evaluated the design and/or usability of key risk of bias (RoB) tools for the assessment of RoB in primary studies, as categorized by the Library of Assessment Tools and InsTruments Used to assess Data validity in Evidence Synthesis Network (a searchable library of RoB tools for evidence synthesis): Prediction model Risk Of Bias ASessment Tool (PROBAST) , Risk of Bias-2 (RoB2), Risk Of Bias In Non-randomised Studies of Interventions (ROBINS-I), Quality Assessment of Diagnostic Accuracy Studies-2 (QUADAS-2), Quality Assessment of Diagnostic Accuracy Studies-Comparative (QUADAS-C), Quality Assessment of Prognostic Accuracy Studies (QUAPAS), Risk Of Bias in Non-randomised Studies of Exposures (ROBINS-E), and the COnsensusbased Standards for the selection of health Measurement INstruments (COSMIN) RoB checklist. Study Design and Setting: Systematic review of methodological studies. We conducted a forward citation search from the primary report of each tool, to identify primary studies that aimed to evaluate the design and/or usability of the tool. Two reviewers assessed studies for inclusion. We extracted tool features into Microsoft Word and used NVivo for document analysis, comprising a mix of deductive and inductive approaches. We summarized findings within each tool and explored common findings across tools. Results: We identified 13 tool evaluations meeting our inclusion criteria: PROBAST (3), RoB2 (3), ROBINS-I (4), and QUADAS-2 (3). We identified no evaluations for the other tools. Evaluations varied in clinical topic area, methodology, approach to bias assessment, and tool user background. Some had limitations affecting generalizability. We identified common findings across tools for 6/14 themes: (1) challenging items (eg, RoB2/ROBINS-I "deviations from intended interventions"domain), (2) overall RoB judgment (concerns with overall risk calculation in PROBAST/ROBINS-I), (3) tool usability (concerns about complexity), (4) time to complete tool (varying demands on time, eg, depending on number of outcomes assessed), (5) user agreement (varied across tools), and (6) recommendations for future use (eg, piloting) and development (add intermediate domain answer to QUADAS-2/PROBAST; provide clearer guidance for all tools). Of the other eight themes, seven only had findings for the QUADAS-2 tool, limiting comparison across tools, and one ("reorganization of questions") had no findings. Conclusion: Evaluations of key RoB tools have posited common challenges and recommendations for tool use and development. These findings may be helpful to people who use or develop RoB tools. Guidance is necessary to support the design and implementation of future RoB tool evaluations.