Estimating the number of the number of people from hidden and/or marginalised populations - such as people dependent on opioids or cocaine - is important to guide policy decisions and provision of harm reduction services. Methods such as capture-recapture are widely used, but rely on assumptions that are often violated and not feasible in specific applications. We describe a Bayesian modelling approach called Multi-Parameter Estimation of Prevalence (MPEP). The MPEP approach leverages routinely collected administrative data, starting from a large baseline cohort of individuals from the population of interest and linked events, to estimate the full size of the target population. When multiple event types are included, the approach enables checking of the consistency of evidence about prevalence from different event types. Additional evidence can be incorporated where inconsistencies are identified. In this article, we summarize the general framework of MPEP, with focus on the most recent version, with improved computational efficiency (implemented in STAN). We also explore several extensions to the model that help us understand the sensitivity of the results to modelling assumptions or identify potential sources of bias. We demonstrate the MPEP approach through a case study estimating the prevalence of opioid dependence in Scotland each year from 2014 to 2022.
Meta-analyses of the accuracy of two diagnostic tests typically assume tests are independent conditional on true disease status. This assumption is often unrealistic and violation leads to biased estimates of the accuracy of tests used in combination. Existing models accounting for conditional dependence require `joint classification' data (results for both tests and the `gold standard' on all participants) from all studies and/or suffer from computational instability. We propose a Bayesian hierarchical model for joint meta-analysis of the accuracy of two binary tests, modelling conditional dependence through study-specific log-odds ratios. The model accommodates studies that do not report joint classification data. We show how the model extends to accommodate data from varied study designs, including studies without a gold standard and studies with partial verification, without assuming imperfect reference standards are error-free. We demonstrate the framework with two example meta-analyses. Our modelling framework retains key features of standard diagnostic test accuracy meta-analysis methods, while allowing for conditional dependence. Ignoring conditional dependence yields biased joint accuracy estimates when conditional dependence is substantial. Our parametrisation maintains computational stability and accommodates data from varied study designs, without requiring an initial data imputation step or assuming error-free reference standards in all studies.
Network meta-analysis of diagnostic test accuracy (NMA-DTA) is a relatively new field, involving combining evidence across studies to evaluate and compare the accuracy of different tests for a given condition. However, the methods proposed to date cannot always capture complex aspects of the data. In fact, many commonly used diagnostic tests are continuous biomarkers, whose accuracy is evaluated at multiple thresholds within a study. Using current NMA-DTA methods we are feasibly able to include in our analysis only a few thresholds per study, discarding this way a big amount of data which could have provided us with useful information. We introduce an approach that can efficiently encompass all available data. This is a hierarchical model that incorporates multinomial likelihoods for studies reporting results across multiple thresholds and a parametric structure for the relationship between the probability of testing positive and threshold within each disease class. This approach enables us to obtain accuracy estimates of tests across the whole range of observed thresholds, while it retains all the useful properties of standard NMA-DTA methods. We explore different variations of this model based on different covariance structures, the inclusion of study-level random effects, and the addition of a further hierarchical structure on the test-level variance components. This framework is applied to data from two systematic reviews, allowing the inclusion of a larger number of tests (compared to alternative approaches) and estimation of sensitivity and specificity at different thresholds with increased precision.
OBJECTIVE:For surveillance for hepatocellular carcinoma (HCC) to be effective, tests-including imaging, serological biomarkers (conventional and genomic) and algorithms combining multiple tests-must identify early-stage tumours. We aimed to identify, appraise and synthesise studies reporting the accuracy of all such tests in people with cirrhosis. DESIGN:Systematic review and network meta-analysis of diagnostic test accuracy (NMA-DTA) data. DATA SOURCES:MEDLINE and Embase (2005 to September 2025) and a published Cochrane review. ELIGIBILITY CRITERIA:English-language, post-2005, one-gate or two-gate studies quantifying diagnostic accuracy of tests to detect HCC in populations wholly comprising people with cirrhosis, excluding those with pre-existing signs and symptoms of HCC. DATA EXTRACTION AND SYNTHESIS:Data extracted by one reviewer, checked by a second and made available in an open-access database. We assessed risk of bias using QUADAS-2. We synthesised data using Bayesian NMA-DTA, accounting for tumour stage and incorporating continuous tests across all possible thresholds. RESULTS:We included 170 studies (62 643 participants). Of 115 index tests, 97 were amenable to NMA-DTA. Ultrasound appears no better than alpha-fetoprotein at detecting very-early-stage HCC (sensitivity 0.34 (95% CrI 0.21 to 0.55) vs 0.39 (95% CrI 0.28 to 0.47)), only becoming superior as stage advances. Least affected by stage are contrast-enhanced MRI (sensitivity 0.70 (95% CrI 0.50 to 0.84) very-early; 0.86 (95% CrI 0.73 to 0.94) early; 0.90 (95% CrI 0.67 to 0.98) advanced) and CT (0.68 (95% CrI 0.26 to 0.93) very-early; 0.77 (95% CrI 0.29 to 0.95) early; 0.92 (95% CrI 0.49 to 0.99) advanced). No genomic biomarkers show convincing improvements over combinations of conventional blood-markers. Most studies are at high risk of bias, but conclusions are robust when restricting to studies with favourable methodological characteristics. CONCLUSIONS:Using advanced synthesis methods, we found that tests have low sensitivity for detecting early-stage HCC. Given rising prevalence of cirrhosis and HCC, we need better tests and a stronger evidence-base to inform optimal surveillance strategies. PROSPERO REGISTRATION NUMBER:CRD42022357163.
Background In the absence of a gold standard, latent class models can be used to estimate test accuracy from a study comparing results on multiple tests. Fixed-effect and latent trait models are often used to account for conditional dependencies between tests, mitigating against biased accuracy estimates, but are difficult to implement. Latent class log-linear models are an under-evaluated alternative.Objectives We evaluate the performance of Bayesian two-class log-linear models in estimating sensitivity, specificity, and prevalence under a variety of real-world conditional dependence structures, and the use of shrinkage priors on interaction terms when dependence structures are unknown.Methods Data were simulated from (i) latent trait, (ii) fixed-effect, and (iii) log-linear models, with dependence structures motivated by four real data sets and three sample sizes. We fitted conditional independence models, log-linear models incorporating the "correct" pairwise interactions, and log-linear models incorporating all interactions with shrinkage priors. We report bias, coverage, residual deviance, and DIC.Results Log-linear models incorporating the correct pairwise dependencies exhibited promising but variable performance (likely due to insufficient sample sizes) across data-generating mechanisms. Improvements over conditional independence models were substantial. Shrinkage priors achieved reasonable performance when dependencies existed within a single disease class but showed poor convergence under complex dependence structures.Conclusions Latent class log-linear models offer a relatively robust alternative for estimating accuracy when the conditional dependence structure is known. When this is not the case, shrinkage priors show promise when dependencies are only within one disease state, but there are challenges related to convergence.
Background: In diagnostic test accuracy studies, it is often infeasible to verify the true disease status of all participants with a gold standard. Several alternative study designs with incomplete verification exist. In a ‘check the negatives’ or 'check the positives' design, all participants first receive an imperfect reference standard, and only those testing negative or positive, respectively, undergo verification with the gold standard. We aimed to quantify bias arising from these designs and assess the extent to which this bias can be corrected using a model-based approach. Methods: We algebraically quantify the bias in standard crude estimates of sensitivity and specificity from a 'check the negatives' design. To address this bias, we propose a Bayesian model incorporating informative priors for the imperfect reference standard’s specificity and, if applicable, a conditional dependence parameter. Under conditional independence, we assessed model performance using simulations, including scenarios where the prior for specificity was correctly centred around the true value and scenarios where this parameter was underestimated. Algebraic results and a model-based correction for a 'check the positives' design follow by symmetry -- in this case, requiring prior information on the imperfect reference standard's sensitivity. Results: In a 'check the negatives' study, crude specificity is unbiased and sensitivity is underestimated under conditional independence. Under positive conditional dependence, specificity is overestimated, while bias in sensitivity can be in either direction. Simulations show that the model-based correction reduced or eliminated bias when the prior for specificity was correctly centred. Overly pessimistic priors sometimes over-corrected bias, although this risk was mitigated with lower prior precision. In two applied examples, where data were insufficient to inform a prior or conditional dependence, the model-based adjustment increased estimated sensitivity under an assumption of conditional independence. Conclusions: Researchers can use our algebraic results to assess potential bias before choosing a 'check the negatives' or 'check the positives' design. These designs are efficient, with minimal bias in some scenarios, but substantial bias in others. Bias can be adjusted for if accurate information is available on reference test specificity or sensitivity, respectively, though misspecified priors can lead to incorrect results.
OBJECTIVES:To explore patient, carer and clinician experiences of the QbTest and its impact on patient outcomes for attention deficit hyperactivity disorder (ADHD) diagnosis and medication management. DESIGN:Mixed-methods systematic review. DATA SOURCES:MEDLINE, EMBASE, PsycINFO, CINAHL, ClinicalTrials.gov and WHO ICTRP (from inception to September 2024). STUDY SELECTION:Primary studies, of any design, that evaluated any version of the QbTest (QbMini <5 years, QbTest 6-12 or 12-60 years, QbCheck for remote assessment via webcam or QbMT smartphone version), for ADHD diagnosis and/or medication management and provided data on any of the following outcomes, were eligible: time to assessment/diagnostic decision, use of services, impact on clinical decision-making, healthcare professionals' confidence in assessment, intervention use, morbidity, mortality, health-related quality of life, cost, ease of use, experience and acceptability of the test to patients, carers and clinicians. DATA EXTRACTION AND SYNTHESIS:Two reviewers independently screened titles and abstracts and assessed potentially relevant reports for inclusion. One reviewer conducted data extraction and risk of bias (RoB) assessment, checked by a second reviewer. Mixed-methods synthesis followed the convergent-integrated approach. RESULTS:We identified 10 eligible studies (9 QbTest; 1 QbCheck), including 1 randomised controlled trial (RCT), 2 feasibility RCTs, 5 before-and-after studies, 1 mixed-methods study and 1 diagnostic study. Most studies enrolled children in the UK and included surveys or interviews with patients, carers or clinicians. The RCT and before-and-after studies were judged at high/serious RoB. Six survey components and two qualitative interview components were judged at some concerns of RoB. We identified one ongoing study of the QbMT and no studies for QbMini. We organised themes emerging from the qualitative synthesis into two broad conceptual categories: views around the helpfulness of the QbTest (contribution to ADHD diagnosis, treatment decision-making, communication with caregivers) and barriers to QbTest implementation (practical barriers and acceptability of the test to patients and caregivers). Findings suggested that the addition of the QbTest may reduce time to diagnosis, improve clinician confidence in the diagnostic decision, increase the proportion of patients with a diagnostic decision and reduce cost and number of clinic appointments. The QbTest appeared to be generally well received by clinicians, patients and carers. However, barriers to test implementation were reported. Clinicians cited staffing, room requirements and issues with technology, and patients highlighted the test length and repetitive nature. Little data exist on the use of the QbTest for medication management. CONCLUSIONS:The available evidence suggests the QbTest may be a useful addition to ADHD assessment in children and young people. Further well-designed RCTs with qualitative substudies are required to assess the impact of the QbTest on patient outcomes, user experience and cost, particularly for medication management and in adults, where evidence is scarce. Such RCTs should include economic analyses, direct comparisons to other continuous performance tests with motion trackers and subgroup analyses including age, sex, ethnicity and comorbidities. PROSPERO REGISTRATION NUMBER:CRD42023482963.
OBJECTIVES:We described and compared infectious-cause hospitalisation outcomes among children born without HIV in the Western Cape (WC), South Africa, during the WHO Option B+ (2013-2015) and universal ART (2016-2018) eras by exposure to maternal HIV and ART. DESIGN:Retrospective cohort. METHODS:Using data from the WC Provincial Health Data Centre, we described rates, causes and risk factors of infectious-cause hospitalisations, up to age 3 years, among children born at a public WC health facility. We compared rates of and risk factors for admission, in children exposed to maternal HIV and uninfected (HEU) and children HIV unexposed and uninfected (HUU), in the neonatal, postneonatal (age >28 days to ≤12 months), and age >12-36 month periods using mixed-effects Poisson regression. Regression models were adjusted for maternal age and suburb of residence. RESULTS:We included 398 334 mother-child pairs, 17.2% children HEU and 82.8% HUU. Infectious-cause hospitalisation, between birth and age 3 years, occurred in 11.5% vs. 10.9% of children HEU and HUU, respectively. Children HEU experienced higher rates of hospitalisation than children HUU, irrespective of maternal ART history, during the neonatal period (adjusted incidence rate ratios, aIRRs: 1.34-1.66) and postneonatal period (aIRRs: 1.13-1.42), but not during the >12-36 month period. Among children HEU, maternal viral load (VL) ≥1000/ml vs. <1000/ml during pregnancy was associated with higher admission rates during the postneonatal period (aIRR = 1.15; 95% CI: 1.06-1.25). CONCLUSIONS:Irrespective of timing of maternal ART start, children HEU vs. HUU had higher rates of infectious-cause hospitalisation during the first year of life, but not thereafter.
For many conditions, it is of clinical importance to know not just the ability of a test to distinguish between those with and without the disease, but also the sensitivity to detect disease at different stages: in particular, the test's ability to detect disease at a stage most amenable to treatment. In a systematic review of test accuracy, pooled stage-specific estimates can be produced using subgroup analysis or meta-regression. However, this requires stage-specific data from each study, which is often not reported. Studies may however report test sensitivity for merged stage categories (e.g. stages I-II) or merged across all stages, together with information on the proportion of patients with disease at each stage. We demonstrate how to incorporate studies reporting merged stage data alongside studies reporting stage-specific data, to allow the inclusion of more studies in the meta-analysis. We consider both meta-analysis of tests with binary results, and meta-analysis of tests with continuous results, where the sensitivity to detect disease of each stage across the whole range of observed thresholds is estimated. The methods are demonstrated using a series of simulated datasets and applied to data from a systematic review of the accuracy of tests used to screen for hepatocellular carcinoma in people with liver cirrhosis. We show that incorporating studies with merged stage data can lead to more precise estimates and, in some cases, corrects biologically implausible results that can arise when the availability of stage-specific data is limited.
In the context of an imperfect gold standard, latent class modelling can be used to estimate accuracy of multiple medical tests. However, the conditional independence (CI) assumption is rarely thought to be clinically valid. Two models accommodating conditional dependence are the latent class multivariate probit (LC-MVP) and latent trait models. Despite LC-MVP's greater flexibility - modelling full correlation matrices versus the latent trait's restricted structure - the latent trait has been more widely used. No simulation studies have directly compared these two models. We conducted a comprehensive simulation study comparing both models across five data generating mechanisms: CI, low-heterogeneity (latent trait-generated), and high-heterogeneity (LC-MVP-generated) correlation structures. We evaluated multiple priors, including novel constrained correlation priors using Pinkney's method that preserves prior interpretability. Models were fit using our BayesMVP R package, which achieves GPU-like speed-ups on these inherently serial models. The LC-MVP model demonstrated superior overall performance. Whilst the latent trait model performed acceptably on its own generated data, it failed for high-heterogeneity structures, sometimes performing worse than the CI model. The CI model did badly for most dependent structures. We also found ceiling effects: high sensitivities reduced the importance of correlation recovery, explaining paradoxes where models achieved good performance despite poor correlation recovery. Our results strongly favour LC-MVP for practical applications. The latent trait model's severe consequences under realistic correlation structures make it a more risky choice. However, LC-MVP with custom correlation constraints and priors provides a safer, more flexible framework for test accuracy evaluation without a perfect gold standard.
Background:Attention deficit hyperactivity disorder is characterised by inattention, impulsivity and hyperactivity. Diagnosis is complex and time-consuming. Medication requires careful selection and dose titration. Technologies for objective measures of attention deficit hyperactivity disorder that use motion sensors to measure hyperactivity ('sensor continuous performance tests') may help improve the diagnostic process and medication management when used in addition to clinical assessment. Objective:To determine whether sensor continuous performance tests are clinically effective and cost-effective to the National Health Service. Specific objectives were to determine the effectiveness of sensor continuous performance tests for: diagnosis of attention deficit hyperactivity disorder in people referred with suspected attention deficit hyperactivity disorder diagnosis of attention deficit hyperactivity disorder in people referred with suspected attention deficit hyperactivity disorder for whom current assessment cannot reach a diagnosis during initial dose titration and treatment decisions for people with attention deficit hyperactivity disorder evaluating treatment effectiveness during long-term treatment monitoring for people with attention deficit hyperactivity disorder. Design:Systematic review and economic model (searches completed 17 November 2023). Results:Objective 1 [29 studies - 25 QbTest (QbTech Ltd., Stockholm, Sweden), 2 EF Sim (Peili Vision, Oulu, Finland) and 2 Nesplora Kids (Giunti Psychometrics, Florence, Italy)]: most evidence was in children. The AQUA trial was the only study to evaluate the QbTest in combination with clinical assessment and included a comparison with clinical assessment alone. Accuracy was similar and there was no statistical evidence of a difference between groups (p = 0.14), but the study was at high risk of bias. The AQUA trial reported that adding QbTest to the diagnostic process resulted in fewer appointments to reach a diagnosis, reduced consultation time, greater clinician confidence and exclusion of the diagnosis in a more children. Findings were supported by limited data from uncontrolled before-after studies. Qualitative and survey data reported increased clinician confidence in clinical decision-making, reduced time to diagnostic decision and improved communication. Barriers to implementation included staffing, training, technology requirements and length and repetitive content of the test. We found that using QbTest in addition to clinical assessment was likely cost-effective due to the reduced time waiting for assessment, reduced appointments until diagnosis and a higher proportion receiving treatment benefits. Objective 3 (six studies): All evaluated QbTest and most had concerns with risk of bias. Qualitative and survey data suggested that healthcare staff and families valued the QbTest for dose titration, checking medication utility and improving medication adherence. Some data suggested that results may not increase patient understanding and some clinicians highlighted logistical challenges. No studies were identified for objectives 2 and 4. Conclusions:Our results suggest that QbTesting as part of the diagnostic workup for attention deficit hyperactivity disorder in children (age < 18 years), when used in combination with clinical assessment, may be cost-effective. This finding was robust to nearly all assumptions made in the model. There are insufficient data on other sensor continuous performance tests in adults or on medication management. Future work:Diagnostic accuracy study evaluating comparing each of the sensor continuous performance tests plus clinical assessment. This should consider accuracy across different patient subgroups. Trial comparing patient outcomes and process measures in adults and children tested with and without sensor continuous performance tests with separate analyses for difficult-to-diagnose patients. Trial evaluating the role of sensor continuous performance tests in medication management, including long-term follow-up. Limitations:Lack of good-quality data on all tests, both for diagnosis and medication management, particularly when evaluated in combination with clinical information. Study registration:This study is registered as PROSPERO CRD42023482963. Funding:This award was funded by the National Institute for Health and Care Research (NIHR) Evidence Synthesis programme (NIHR award ref: NIHR136009) and is published in full in Health Technology Assessment; Vol. 29, No. 58. See the NIHR Funding and Awards website for further award information.
Background. The COVID-19 pandemic resulted in the implementation of strict public health and social measures (PHSMs) (including mobility restrictions, social distancing, mask -wearing and hand hygiene), limitations on non -essential healthcare services, and public fear of COVID-19 infection, all of which potentially affected transmission and healthcare use for other diseases such as lower respiratory tract infections (LRTIs). Objective. To determine changes in LRTI hospital admissions and in -facility mortality in children aged <5 years in the Western Cape Province during the pandemic. Methods. We conducted a retrospective analysis of LRTI admissions and in -facility deaths from January 2019 to November 2021. We estimated changes in rates and trends of LRTI admissions during the pandemic compared with pre -pandemic period using interrupted time series analysis, adjusting for key characteristics. Results. There were 36 277 children admitted for LRTIs during the study period, of whom 58% were male and 51% were aged 28 days1 year. COVID-19 restrictions were associated with a 13% step reduction in LRTI admissions compared with the pre-COVID-19 period (incidence rate ratio (IRR) 0.87, 95% confidence interval (CI)) 0.80 - 0.94). The average LRTI admission trend increased on average by 2% per month during the pandemic (IRR 1.02, 95% CI 1.02 - 1.04). Conclusions. The COVID-19 surges and their associated measures were linked to declining LRTI admissions and in -facility deaths, likely driven by a combination of reduced infectious disease transmission and reduced use of healthcare services, with effects diminishing over time. These findings may inform future pandemic response policies.
Background: Point of care tests (POCTs) have the potential to improve the urinary tract infection (UTI) diagnostic pathway, as they can provide a diagnosis quickly in near-patient settings, and some also identify causative pathogens/antimicrobial sensitivity. Objectives: To assess the clinical impact, accuracy, and technical characteristics of POCT for diagnosing UTI. Methods of data synthesis: Narrative summary and bivariate random effects meta-analyses to estimate summary sensitivity and specificity. Data sources: Five electronic databases, two clinical trial registries, study reports and review reference lists, and websites. Study eligibility criteria: Randomized controlled trials/non-randomized studies and diagnostic test accuracy studies published since 2000. Participants: People with suspected UTI. Tests: Rapid tests (results <40 minutes): Astrego PA-10 0 system, Lodestar DX, Uriscreen, UTRiPLEX. Culture tests (results <24 hours): Flexicult Human, ID Flexicult, Diaslide, Dipstreak, Chromostreak, Uricult, Uricult Trio, Uricult Plus. Reference standard: Any. Assessment of risk of bias: Risk of Bias-2, Quality Assessment of Diagnostic Accuracy Studies-2, Quality Assessment of Diagnostic Accuracy Studies-C. Results: Two randomized controlled trials evaluated Flexicult Human (one against standard care; one against ID Flexicult). No difference was reported in antibiotic use concordant with culture results (OR 0.84 95% CI 0.58-1.20) or appropriate antibiotic prescribing (OR 1.44 95% CI 1.03-1.99). Initial antibiotic prescribing was lower with Flexicult than standard care (OR 0.56 95% CI 0.35-0.88). No difference for other measures of antibiotic use, symptom duration, patient enablement, or resource use. Fifteen studies reported accuracy data. Limited data were available, with most POCT evaluated in single studies or not evaluated at all. Uriscreen (four studies), Uricult Trio (three studies), Flexicult Human (four studies), and ID Flexicult (two studies) had modest sensitivity and specificity. POCTs were easier to use and interpret than standard culture. Conclusions: There is currently insufficient evidence to support the use of POCTs in UTI diagnosis. Due to the rapid development of POCT, this review should be updated regularly. Eve Tomlinson, Clin Microbiol Infect 2024;30:197 (c) 2023 European Society of Clinical Microbiology and Infectious Diseases. Published by Elsevier Ltd. All rights reserved.
Abstract Background Latent class models can be used to estimate diagnostic accuracy without a gold standard test. Early studies often assumed independence between tests given the true disease state, however this can lead to biased estimates when there are inter-test dependencies. Residual correlation plots and chi-squared statistics have been commonly utilized to assess the validity of the conditional independence assumption and, when it does not hold, identify which test pairs are conditionally dependent. We aimed to assess the performance of these tools with a simulation study covering a wide range of scenarios. Methods We generated data sets from a model with four tests and a dependence between tests 1 and 2 within the diseased group. We varied sample size, prevalence, covariance, sensitivity and specificity, with 504 combinations of these in total, and 1000 data sets for each combination. We fitted the conditional independence model in a Bayesian framework, and reported absolute bias, coverage, and how often the residual correlation plots, $${G}^{2}$$ G 2 and $${\chi }^{2}$$ χ 2 statistics indicated lack-of-fit globally or for each test pair. Results Across all settings, residual correlation plots, pairwise $${G}^{2}$$ G 2 and $${\chi }^{2}$$ χ 2 detected the correct correlated pair of tests only 12.1%, 10.3%, and 10.3% of the time, respectively, but incorrectly suggested dependence between tests 3 and 4 64.9%, 49.7%, and 49.5% of the time. We observed some variation in this across parameter settings, with these tools appearing to perform more as intended when tests 3 and 4 were both much more accurate than tests 1 and 2. Residual correlation plots, $${G}^{2}$$ G 2 and $${\chi }^{2}$$ χ 2 statistics identified a lack of overall fit in 74.3%, 64.5% and 67.5% of models, respectively. The conditional independence model tended to overestimate the sensitivities of the correlated tests (median bias across all scenarios 0.094, 2.5th and 97.5th percentiles -0.003, 0.397) and underestimate prevalence and the specificities of the uncorrelated tests. Conclusions Residual correlation plots and chi-squared statistics cannot be relied upon to identify which tests are conditionally dependent, and also have relatively low power to detect lack of overall fit. This is important since failure to account for conditional dependence can lead to highly biased parameter estimates.
Using the Data Evaluation and Preparation for HIV-Exposed Uninfected Child Cohorts project’s standardized child HIV exposure definitions, 64%, 64% and 90% of children exposed to HIV in utero could be classified as HIV-uninfected with moderate or high certainty at the ages of 1 and 3 years and at the time of first infectious disease hospitalization, respectively. These definitions can be applied retrospectively to routine datasets with linked mother-child data.
Background:Liver cirrhosis is the largest risk factor for developing hepatocellular carcinoma (HCC), and surveillance is therefore recommended among this population. Current guidance recommends surveillance with ultrasound, with or without alpha-fetoprotein (AFP). This review is part of a larger project looking at benefits, harms and costs of surveillance for HCC in people with cirrhosis. It aims to synthesise the evidence on the diagnostic accuracy of imaging or biomarker tests, alone or in combination, to identify HCC in adults with liver cirrhosis in a surveillance programme. Methods:We will identify studies through a 2021 Cochrane review with similar eligibility criteria, and a database search of MEDLINE, Embase and the Cochrane Database of Systematic Reviews. We will include diagnostic test accuracy studies with adult cirrhosis patients of any aetiology. Studies must assess at least one of the following index tests: ultrasound (US), magnetic resonance imaging (MRI), computerised tomography (CT), alpha-fetoprotein (AFP), des-gamma-carboxyprothrombin (DCP), lens culinaris agglutinin-reactive fraction of AFP (AFP-L3), a genomic biomarker, or a diagnostic prediction model incorporating at least one of the above-mentioned tests. We will assess studies for risk of bias using QUADAS-2 and QUADAS-C. We will combine data using bivariate random effects meta-analyses. For tests evaluated across varying diagnostic thresholds, we will produce pooled estimates of sensitivity and specificity across the full range of numerical thresholds, where possible. Where sufficient studies compare two or more index tests, we will perform additional analyses to compare the accuracy of different tests. Where feasible, we will stratify all meta-analyses by tumour size and patient characteristics, including cirrhosis aetiology and liver disease severity. Discussion:This review will synthesise evidence across the full range of possible surveillance tests, using advanced statistical methods to summarise accuracy across all thresholds and to compare the accuracy of different tests. PROSPERO registration:CRD42022357163.
Background:Acute respiratory infections are a common reason for consultation with primary and emergency healthcare services. Identifying individuals with a bacterial infection is crucial to ensure appropriate treatment. However, it is also important to avoid overprescription of antibiotics, to prevent unnecessary side effects and antimicrobial resistance. We conducted a systematic review to summarise evidence on the diagnostic accuracy of symptoms, signs and point-of-care tests to diagnose bacterial respiratory tract infection in adults, and to diagnose two common respiratory viruses, influenza and respiratory syncytial virus. Methods:The primary approach was an overview of existing systematic reviews. We conducted literature searches (22 May 2023) to identify systematic reviews of the diagnostic accuracy of point-of-care tests. Where multiple reviews were identified, we selected the most recent and comprehensive review, with the greatest overlap in scope with our review question. Methodological quality was assessed using the Risk of Bias in Systematic Reviews tool. Summary estimates of diagnostic accuracy (sensitivity, specificity or area under the curve) were extracted. Where no systematic review was identified, we searched for primary studies. We extracted sufficient data to construct a 2 × 2 table of diagnostic accuracy, to calculate sensitivity and specificity. Methodological quality was assessed using the Quality Assessment of Diagnostic Accuracy Studies version 2 tool. Where possible, meta-analyses were conducted. We used GRADE to assess the certainty of the evidence from existing reviews and new analyses. Results:We identified 23 reviews which addressed our review question; 6 were selected as the most comprehensive and similar in scope to our review protocol. These systematic reviews considered the following tests for bacterial respiratory infection: individual symptoms and signs; combinations of symptoms and signs (in clinical prediction models); clinical prediction models incorporating C-reactive protein; and biological markers related to infection (including C-reactive protein, procalcitonin and others). We also identified systematic reviews that reported the accuracy of specific tests for influenza and respiratory syncytial virus. No reviews were found that assessed the diagnostic accuracy of white cell count for bacterial respiratory infection, or multiplex tests for influenza and respiratory syncytial virus. We therefore conducted searches for primary studies, and carried out meta-analyses for these index tests. Overall, we found that symptoms and signs have poor diagnostic accuracy for bacterial respiratory infection (sensitivity ranging from 9.6% to 89.1%; specificity ranging from 13.4% to 95%). Accuracy of biomarkers was slightly better, particularly when combinations of biomarkers were used (sensitivity 80-90%, specificity 82-93%). The sensitivity and specificity for influenza or respiratory syncytial virus varied considerably across the different types of tests. Tests involving nucleic acid amplification techniques (either single pathogen or multiplex tests) had the highest diagnostic accuracy for influenza (sensitivity 91-99.8%, specificity 96.8-99.4%). Limitations:Most of the evidence was considered low or very low certainty when assessed with GRADE, due to imprecision in effect estimates, the potential for bias and the inclusion of participants outside the scope of this review (children, or people in hospital). Future work:Currently evidence is insufficient to support routine use of point-of-care tests in primary and emergency care. Further work must establish whether the introduction of point-of-care tests adds value, or simply increases healthcare costs. Funding:This article presents independent research funded by the National Institute for Health and Care Research (NIHR) Health Technology Assessment programme as award number NIHR159948.
AIM:Extending faecal immunochemical tests for haemoglobin (FIT) to all primary care patients with symptoms suggestive of colorectal cancer (CRC) could identify people who are likely to benefit from colonoscopy and facilitate earlier treatment. The aim of this work was to investigate the diagnostic accuracy of FIT across different analysers at different thresholds, as a single test or in duplicate (dual FIT). METHOD:This systematic review and meta-analysis searched 10 sources (December 2022). Diagnostic accuracy studies of HM-JACKarc, OC-Sensor, FOB Gold, QuikRead go, NS-Prime and four Immunodiagnostik (IDK) tests in primary care patients were included. Risk of bias was assessed (QUADAS-2). Statistical syntheses produced summary estimates of sensitivity and specificity at any chosen threshold for CRC, inflammatory bowel disease and advanced adenomas separately. Sensitivity analyses investigated reference standard and population type (high, low or all-risk). Subgroup analyses investigated patient characteristics (e.g. anaemia, age, sex, ethnicity). RESULTS:Thirty-seven studies were included. At a threshold of 10 μg/g, pooled results for sensitivity and specificity (95% credible intervals) for CRC, respectively, were: HM-JACKarc (n = 16 studies) 89.5% (84.6%-93.4%) and 82.8% (75.2%-89.6%); OC-Sensor (n = 11 studies) 89.8% (85.9%-93.3%) and 77.6% (64.3%-88.6%); FOB Gold (n = 3 studies), 87.0% (67.3%-98.3%) and 88.4% (81.7%-94.2%). There were limited or no data on the other tests, dual FIT and relating to patient characteristics. CONCLUSION:Test sensitivity at a threshold of 10 μg/g highlights a requirement for adequate safeguards in test-negative patients with ongoing symptoms. Further research is needed into the impact of patient characteristics and dual FIT.