Adaptive clinical dose-finding trials aim to identify an optimal drug dose for use in subsequent phase II and III trials. In adaptive dose-finding trials, dose levels for newly included patients are informed by outcomes of patients that received the drug earlier in the trial. The body of methodological research on adaptive dose-finding trials is extensive, but a clear overview is lacking. The goal of this paper is to provide a knowledge base of designs and statistical methods for adaptive clinical dose-finding trials by means of a literature review. We identified 315 adaptive dose-finding trial methodology articles of which the majority was inspired by oncology. Recent methods focused on identification of an optimal dose considering both toxicity and efficacy endpoints, and addressing challenges related to subgroup-specific dose-finding, dose-finding for combination therapies, and incorporation of delayed outcomes in dose-finding. These developments are driven by the emergence of newer classes of cancer drugs, such as targeted therapies and immunotherapies, and by initiatives like Project Optimus. Most articles focused on model-based designs, like the Continual Reassessment Method (CRM), but recent years have seen a strong increase in model-assisted or interval-based designs, including Toxicity Probability Interval (TPI) and Bayesian Optimal Interval (BOIN) designs and expansions thereof. Considering the increasing availability and large variety of adaptive dose-finding trial designs, it is challenging for researchers to find relevant designs tailored to their needs. Therefore, we provide an interactive map summarizing the classification of our results to facilitate the identification of relevant designs/methods, which we will update regularly.
Anti-arrhythmic drugs (AAD) and catheter ablation are common treatments for rhythm control in patients with atrial fibrillation (AF), but a comprehensive comparison of their benefits and harms is lacking. A systematic review performed by National Institute for Health and Care Excellence (2021) was updated using a database search in MEDLINE, Embase and Cochrane CENTRAL in October 2024. Selection was conducted independently by 2 reviewers. Study characteristics and results were extracted by 1 reviewer and verified by a second. Study quality was assessed using the Cochrane Risk of Bias 2 tool and the certainty of evidence was determined using the Grading of Recommendations, Assessment, Development and Evaluation (GRADE) approach. Twenty-two studies (31 references) were included in the review. Evidence showed that there may be more recurrence of AF in patients with treatment-naïve paroxysmal AF (relative risk [RR], 1.76; 95% confidence interval [CI], 1.34-2.32) and AAD-exposed persistent AF (RR, 1.62; 95% CI, 1.40-1.87), and probably more recurrence in patients with paroxysmal AF previously exposed to catheter ablation (RR, 2.12; 95% CI, 1.61-2.80) when treated with AAD compared with catheter ablation. In both paroxysmal- and persistent AF patients, there may also be more all-cause mortality when treated with AAD vs catheter ablation (RR, 1.79; 95% CI, 0.80-4.02 and RR, 1.45; 95% CI, 0.56-3.75, respectively). Incidence of medication side effects, including gastrointestinal symptoms and syncope, ranged between 0.8 and 3%. Incidence of catheter ablation complications, such as pericardial tamponade and phrenic nerve paralysis, between 0.3 and 2.5%. This review found that, in some subgroups of patients, catheter ablation was better than AAD treatment when considering AF recurrence and mortality. However, catheter ablation has rare but serious complications that should be considered when making informed treatment decisions.
BACKGROUND:A comprehensive overview of the cost-effectiveness of pharmacologic treatments for overweight or obesity is lacking. PURPOSE:To evaluate cost-effectiveness of pharmacologic treatments in adults with overweight or obesity in a U.S. setting. DATA SOURCES:MEDLINE, Embase, and economic databases, searched on 13 October 2025. STUDY SELECTION:Non-industry-sponsored U.S. trial-based and model-based cost-effectiveness evaluations of pharmacologic treatments in adults with overweight or obesity. DATA EXTRACTION:Data on clinical characteristics, economic characteristics (for example, model type), and study outcomes were extracted by one reviewer and verified by a second reviewer. Study quality was assessed using the CHEQUE (Criteria for Health Economic Quality Evaluation) tool; value was assessed using incremental cost-effectiveness ratios (ICERs), with thresholds for high value (<$100 000 per quality-adjusted life-year [QALY], or dominant), intermediate value ($100 000 to $200 000 per QALY), low value (>$200 000 per QALY), and no value (strict or extended dominance, or less costly and less effective); and certainty of evidence was assessed using the GRADE (Grading of Recommendations Assessment, Development and Evaluation) approach. DATA SYNTHESIS:Four out of 9 included studies were at low risk of bias. None of the 42 pairwise comparisons that were reported had high certainty. In the 6 studies with moderate certainty, liraglutide had low value and phentermine-topiramate and tirzepatide had high value when each was compared with lifestyle modification. Semaglutide had low value compared with naltrexone-bupropion and phentermine-topiramate and high value compared with liraglutide. LIMITATIONS:All studies were model-based. ICERs were not reported for all potential treatment comparisons. Most studies had incomplete reporting or were at high risk of bias. CONCLUSION:Current evidence on cost-effectiveness of pharmacologic treatment of overweight or obesity is hampered by poor-quality studies, limiting the ability to draw conclusions. PRIMARY FUNDING SOURCE:American College of Physicians. (PROSPERO: CRD42023491646).
OBJECTIVES:This systematic review and meta-analysis summarized documented breakthrough hepatitis A (HepA) infections and assessed whether they occur more frequently in immunocompromised populations (ICPs). METHODS:We searched Medline, Embase, and Global Index Medicus for records on HepA breakthrough infections, defined as symptomatic HepA among previously vaccinated individuals, published 1991-2024 (PROSPERO: CRD42023450205). Risk of bias was assessed with an adapted Newcastle-Ottawa Scale. Moderate and good-quality studies were analyzed using random-effects and Bayesian hierarchical meta-analyses. The primary outcome compared the proportion of ICPs among fully vaccinated (≥50 days before symptoms) confirmed breakthrough cases to that among all virologically confirmed HepA cases. Meta-analyses assessed the proportion of breakthrough infections among confirmed cases. RESULTS:Of 3345 reports screened, 90 were included, reporting 43,861 HepA cases, including 832 (partially) vaccinated. We identified six confirmed breakthrough infections with proof of complete vaccination (6.4%; 95% confidence interval 2.6-14.3%), including four ICPs (67%). Post-vaccination serology was available in 3/6 cases: two HIV-positive patients were seronegative, and one neutropenic leukemia patient was seropositive but pauci-symptomatic. The pooled proportion of ICPs among confirmed HepA cases was 17.8% (95% confidence interval 7.7-35.9%). CONCLUSIONS:HepA breakthrough infections are rare, but appear more common among ICPs who fail to seroconvert post-vaccination.
Various open science practices have been proposed to improve the reproducibility and replicability of scientific research, but not for all practices, there may be evidence they are indeed effective. Therefore, we conducted a scoping review of the literature on interventions to improve reproducibility. We systematically searched Medline, Embase, Web of Science, PsycINFO, Scopus and Eric, on 18 August 2023. Any study empirically evaluating the effectiveness of interventions aimed at improving the reproducibility or replicability of scientific methods and findings was included. We summarized the retrieved evidence narratively and in evidence gap maps. Of the 105 distinct studies we included, 15 directly measured the effect of an intervention on reproducibility or replicability, while the remainder addressed a proxy outcome that might be expected to increase reproducibility or replicability, such as data sharing, methods transparency or pre-registration. Thirty studies were non-comparative and 27 were comparative but cross-sectional observational designs, precluding any causal inference. Despite studies investigating a range of interventions and addressing various outcomes, our findings indicate that in general the evidence base for which various interventions to improve reproducibility of research remains remarkably limited in many respects.
BACKGROUND:Decision making regarding pharmacologic treatments for the prevention of episodic migraine may depend on the importance that patients place on outcomes and specific treatment preferences. PURPOSE:To assess patients' values and preferences regarding pharmacologic treatments for the prevention of episodic migraine. DATA SOURCES:MEDLINE and CINAHL from inception to April 2024. STUDY SELECTION:Quantitative studies reporting on values and preferences regarding pharmacologic treatments for the prevention of migraine in adults were eligible. DATA EXTRACTION:We extracted data on study design, participants, and findings and assessed risk of bias using a tool developed by the GRADE (Grading of Recommendations Assessment, Development and Evaluation) working group. DATA SYNTHESIS:We included 6 studies (5 discrete choice experiments and 1 survey) comprising a total of 2307 participants and assessing the importance of a total of 32 attributes. Risk of bias was moderate in 3 and low in 3 studies; all but 1 were industry spon sored. Migraine severity, frequency, and duration were found to be the most important among several outcomes including acute medication need and side effects (low certainty). Other outcomes of importance were migraine attack on day 1 postdosing (moderate certainty) and duration of daily activity limitations (high certainty). Patients preferred oral tablets to injections and infusions (moderate certainty). LIMITATIONS:Importance of attributes relative to each other could not be assessed due to the scarcity of direct comparisons. Participant sampling was unclearly reported in most studies. CONCLUSION:Patients may have the strongest preference for oral treatments that reduce the severity, frequency, and duration of their migraine attacks, which reduce the duration of daily activity limitations and reduce risk of migraine on day 1. Side effects may be less important. PRIMARY FUNDING SOURCE:American College of Physicians. (PROSPERO: CRD42023414305).
BACKGROUND:Obstructive sleep apnoea (OSA) is a common cause of sleep disturbance, characterised by the presence of repetitive upper airway obstruction during sleep. OSA is associated with sleepiness during the day, reduced quality of life and an increased risk of cardiovascular disease. OSA can be diagnosed using several different strategies. The current reference test is fully supervised polysomnography, which is expensive and time-consuming. Other diagnostic tests, referred to as limited channel sleep studies because they include fewer parameters than polysomnography, are less resource-intensive but may also have different diagnostic performances, resulting in a difference in clinical outcomes. OBJECTIVES:To assess the clinical impact (outcome on a participant level) of a strategy where treatment follows diagnostic testing (test-treatment combination) using limited channel sleep studies compared to polysomnography in people with suspected obstructive sleep apnoea (OSA). SEARCH METHODS:We searched two databases (CENTRAL, MEDLINE) up to 11 May 2023 using search terms related to OSA and polysomnography developed by our information specialist. SELECTION CRITERIA:We included randomised controlled trials that compared any limited channel sleep studies with Level I fully supervised polysomnography in adults (aged 18 years and older) with suspected OSA. Our primary outcome was sleepiness, and our secondary outcomes were quality of life, all-cause mortality, cardiovascular events and correlating risk factors, continuous positive airway pressure (CPAP) usage, serious adverse events, and cost-effectiveness. DATA COLLECTION AND ANALYSIS:Four review authors extracted data from the included trials and assessed the risk of bias. We summarised treatment effects using random-effects meta-analyses and expressed as mean difference (MD) or standardised mean difference (SMD) with corresponding 95% confidence intervals (CI) where possible. We used GRADE to assess the certainty of the evidence. MAIN RESULTS:We included three trials with 1143 participants. One trial compared Level III sleep studies to a Level I fully supervised polysomnography, one trial compared Level IV sleep studies to Level I sleep studies, and one trial compared Level IV sleep studies versus Level III sleep studies versus Level I sleep studies. The follow-up of these trials ranged from four to six months. Level III sleep studies versus Level I sleep studies There is high-certainty evidence that Level III sleep studies result in little to no difference in sleepiness (MD 0.47, 95% CI -0.23 to 1.18; P = 0.19, I2 = 0%; 2 trials, 701 participants) or quality of life (SMD 0.01, 95% CI -0.14 to 0.16; P = 0.93, I2 = 0%; 2 trials, 701 participants) compared to Level I sleep studies. Level III sleep studies are also probably slightly more cost-effective (moderate-certainty evidence). There is low-certainty evidence that they may result in little to no difference in cardiovascular events and correlating risk factors, CPAP adherence (MD -0.18 hours per day, 95% CI -0.56 to 0.20; P = 0.36, I2 = 0%; 2 trials, 360 participants) or serious adverse events. Level IV sleep studies versus Level I sleep studies There is low-certainty evidence that Level IV sleep studies may not increase sleepiness compared to Level I sleep studies (MD 0.66, 95% CI -0.41 to 1.72; P = 0.23, I2 = 39%; 2 trials, 573 participants). Additionally, there is low-certainty evidence that they may result in little to no difference in cardiovascular events and correlating risk factors. For quality of life, CPAP adherence, serious adverse events and cost-effectiveness, the evidence is very uncertain. None of the included trials reported on all-cause mortality. AUTHORS' CONCLUSIONS:Level III sleep studies may result in little to no difference in clinical outcomes when compared to Level 1 sleep studies in people with suspected OSA. Level IV sleep studies may not increase sleepiness and may result in little to no difference in cardiovascular events and correlating risk factors compared to Level I sleep studies; the evidence was too uncertain to make statements for other outcomes. Overall, the body of evidence was limited, therefore more trials making this comparison are necessary, as are trials with a longer follow-up duration.
BACKGROUND:Evidence on the clinical impact of seasonal respiratory viruses in long-term care facilities (LTCFs) is limited. OBJECTIVES:To provide an overview of the evidence available on the clinical impact of seasonal respiratory viruses on older adults in LTCFs worldwide. DATA SOURCES:Medline (OVID Medline ALL) and Embase (Embase.com) until 13 May 2025. STUDY ELIGIBILITY CRITERIA:Original research articles involving LTCF residents with at least one laboratory-confirmed viral respiratory tract infection (RTI), excluding SARS-CoV-2 and pandemic influenza, reporting any RTI-associated clinical outcome. PARTICIPANTS:LTCF residents (mean or median age >60 years). METHODS OF DATA SYNTHESIS:An evidence gap map was created to visualise the distribution of evidence across viruses and outcomes. Where possible, outcome proportions (defined as the number of cases with the outcome divided by the total number of cases) were extracted from the included studies. These were summarised and visualised in dot plots with medians and interquartile ranges (IQRs) per virus. RESULTS:117 studies were included. The majority of the studies focused on influenza viruses and conventional outcomes including attack rate, lower respiratory tract infection (LRTI), hospitalisation, and mortality. Evidence was limited for human rhinovirus, parainfluenza viruses, enterovirus, adenovirus, and endemic human coronaviruses, as well as for outcomes such as aggravation of underlying diseases and patient-centred outcomes, including quality of life and functional status. Human metapneumovirus (hMPV) was associated with the highest rate of LRTI (median 0.50; IQR 0.45-0.75) as well as the highest mortality rate (median 0.17; IQR 0.10-0.37) found in this study, although based on small numbers of confirmed. CONCLUSIONS:Substantial knowledge gaps remain regarding the impact of seasonal respiratory viruses on older adults in LTCFs. In order to inform decision-making regarding viral RTI management in this population, future studies should prioritise underrepresented viruses including hMPV and incorporate patient-centred outcomes.
BACKGROUND:Various treatments for preventing episodic migraine are available. PURPOSE:To evaluate the comparative effectiveness and harms of pharmacologic prevention of episodic migraine, focusing on treatments already determined to be superior to placebo. DATA SOURCES:MEDLINE, EMBASE, and the Cochrane Central Register of Controlled Trials from inception until April 2024. STUDY SELECTION:Randomized trials evaluating selected efficacious pharmacologic treatments in adults with episodic migraine. Selection was done independently by 2 reviewers. DATA EXTRACTION:Data were extracted by 1 reviewer and checked by a second. Risk of bias and certainty of the evidence were assessed using the Cochrane Risk of Bias tool and the GRADE (Grading of Recommendations Assessment, Development and Evaluation) approach, respectively. DATA SYNTHESIS:Sixty-one studies (20 680 patients) evaluating 16 treatments were included. Nineteen studies had low risk of bias. All selected treatments were deemed efficacious against placebo on the basis of previous systematic reviews. In network meta-analyses, calcitonin gene-related peptide antagonist monoclonal antibodies (CGRP-mAbs) probably resulted in fewer discontinuations due to adverse events than topiramate (risk difference, -16.2% [95% CI, -18.4% to -12.8%]; moderate-certainty evidence), and CGRP-mAbs may result in less migraine-related disability and improved quality of life compared with gepants (mean differences, -4.12 [CI, -9.30 to 1.05] and 2.25 [CI, -0.85 to 5.34], respectively; low-certainty evidence). For other outcomes and comparisons, there was moderate- or low-certainty evidence of no clinically important differences, uncertain evidence, or no evidence. LIMITATIONS:Limited literature was available to determine the minimal important differences. The number of head-to-head comparisons of treatments was limited. CONCLUSION:No high-certainty evidence favored one pharmacologic treatment for prevention of episodic migraine over another. Evidence was mostly insufficient or of low certainty. PRIMARY FUNDING SOURCE:American College of Physicians. (PROSPERO: CRD42023414305).
BACKGROUND:Accurate rapid diagnostic tests for SARS-CoV-2 infection could help manage the COVID-19 pandemic by potentially increasing access to testing and speed detection of infection, as well as informing clinical and public health management decisions to reduce transmission. Previous iterations of this review provided clear and conclusive evidence of superior test performance in those experiencing possible signs and symptoms of Covid-19. However, test performance in asymptomatic individuals and sensitivity by setting and indication for testing remains unclear. This is the fourth iteration of this review, first published in 2020. OBJECTIVES:To assess the diagnostic accuracy of rapid, point-of-care antigen tests (Ag-RDTs) for diagnosis of SARS-CoV-2 infection in asymptomatic population groups. SEARCH METHODS:We searched the COVID-19 Open Access Project living evidence database from the University of Bern (which includes daily updates from MEDLINE and Embase and preprints from medRxiv and bioRxiv) on 17 February 2022. We included independent evaluations from national reference laboratories, FIND and the Diagnostics Global Health website. We did not apply language restrictions. SELECTION CRITERIA:We included test accuracy studies of any design that evaluated commercially produced, rapid antigen tests in asymptomatic people tested because of known or suspected contact with SARS-CoV-2 infection, known SARS-CoV-2 infection or known absence of infection, or those who were being screened for infection. We included evaluations of single applications of a test (one test result reported per person). Reference standards for presence or absence of infection were any laboratory-based molecular test (primarily reverse transcription polymerase chain reaction (RT-PCR)). DATA COLLECTION AND ANALYSIS:We used standard screening procedures with three reviewers. Two reviewers independently carried out quality assessment (using the QUADAS-2 tool) and extracted study results. Other study characteristics were extracted by one review author and checked by a second. We present sensitivity and specificity with 95% confidence intervals (CIs) for each test, and pooled data using the bivariate model. We investigated heterogeneity by including indicator variables in the random-effects logistic regression models. We tabulated results by test manufacturer and compliance with manufacturer instructions for use and according to symptom status. MAIN RESULTS:We included 146 study cohorts (described in 130 study reports). The main results relate to 164 evaluations of single test applications including 144,250 unique samples (7104 with confirmed SARS-CoV-2) obtained from asymptomatic or mainly asymptomatic populations. Studies were mainly conducted in Europe (85/146, 58%), and evaluated 41 different commercial antigen assays (test kit). Only six studies compared two or more brands of test. Nearly all studies (96%) used RT-PCR alone to define presence or absence of infection. Risk of bias was high because of participant selection (13, 9%); interpretation of the index test (3, 2%); weaknesses in the reference standard for absence of infection (3, 2%); and participant flow and timing (46, 32%). Characteristics of participants (11, 8%) and index test delivery (117, 80%) differed from the way in which and in whom the test was intended to be used. Estimates of sensitivity varied considerably between studies, with consistently high specificities. Average sensitivity was 55.0% (95% CI 50.9%, 59.0%) and average specificity was 99.5% (95% CI 99.5%, 99.6%) across the 147 evaluations of Ag-RDTs reporting both sensitivity and specificity (149,251 samples, 7636 cases). Average sensitivity was higher when epidemiological exposure to SARS-CoV-2 was suspected (58.6%, 95% CI 51.4% to 65.5%; 43 evaluations; 15,516 samples, 1483 cases) compared to where COVID-19 testing was reported to be widely available to anyone on presentation for testing (53.0%, 95% CI 48.4% to 57.5%; 103 evaluations; 129,032 samples, 5660 cases); however CIs overlapped, limiting the inference that can be drawn from these data. Average specificity was similarly high for both groups (99.4% and 99.6%). Sensitivity was generally lower when used in a screening context (summary values from 40.6% to 42.1% for three of four screening settings) compared to testing asymptomatic individuals at Covid-19 test centres (56.7%) or emergency departments (54.7%). We observed a decline in summary sensitivities as measures of sample viral load decreased. Sensitivity varied between brands. When tests were used according to manufacturer instructions, average sensitivities by brand ranged from 36.3% to 78.8% in asymptomatic participants (14 assays with sufficient data for pooling). None of the assays met the WHO acceptable performance standard for sensitivity (of 80%) based on meta-analysis; however, sensitivities from individual studies (where meta-analysis was not possible) exceeded 80% for three assays. The WHO acceptable performance criterion of 97% specificity was met by all but four assays (based on individual studies or meta-analysis) when tests were used according to manufacturer instructions. At 0.5% prevalence using summary data for asymptomatic people, where testing was widely available and where epidemiological exposure to COVID-19 was suspected, resulting PPVs would be 40% and 33%, meaning that 3 in 5 or 2 in 3 positive results will be false positives, and between 1 in 2 and 2 in 5 cases will be missed. AUTHORS' CONCLUSIONS:Evidence for antigen testing in asymptomatic cohorts has increased considerably since the publication of the previous update of this review. Average sensitivities remain lower for testing of asymptomatic when compared to symptomatic individuals; however, there is an indication that sensitivities may be higher where epidemiological exposure to SARS-CoV-2 is suspected compared to testing any asymptomatic individual regardless of indication. Sensitivities were particularly low when antigen tests were used in screening settings. Assays from different manufacturers also vary in sensitivity, indicating the need for appropriate clinical validation of a particular antigen test in a given intended use setting prior to more widespread deployment. Further research is needed to evaluate the effectiveness of screening programmes at reducing transmission of infection, whether mass screening or targeted approaches, including schools, healthcare setting and traveller screening. FUNDING:This paper presents independent research supported by the NIHR Birmingham Biomedical Research Centre, University Hospitals Birmingham NHS Foundation Trust, and the University of Birmingham. The views expressed are those of the author(s) and not necessarily those of the NHS, the NIHR or the Department of Health and Social Care. REGISTRATION:Protocol (2020) doi: 10.1002/14651858.CD013596.
The eighth meeting of the International Collaboration for the Automation of Systematic Reviews (ICASR) was held on September 7 and 8, 2023, at the University College London, London, England. ICASR is an interdisciplinary group whose goal is to maximize the use of technology for conducting rapid, accurate, and efficient evidence synthesis, e.g., systematic reviews, evidence maps, and scoping reviews of scientific evidence. In 2023, the major themes discussed were understanding the benefits and harms of automation tools that have become available in recent years, the advantages and disadvantages of large language models in evidence synthesis, and approaches to ensuring the validity of tools for the proposed task.
Background:Many interventions, especially those linked to open science, have been proposed to improve reproducibility in science. To what extent these propositions are based on scientific evidence from empirical evaluations is not clear. Aims:The primary objective is to identify Open Science interventions that have been formally investigated regarding their influence on reproducibility and replicability. A secondary objective is to list any facilitators or barriers reported and to identify gaps in the evidence. Methods:We will search broadly by using electronic bibliographic databases, broad internet search, and contacting experts in the field of reproducibility, replicability, and open science. Any study investigating interventions for their influence on the reproducibility and replicability of research will be selected, including those studies additionally investigating drivers and barriers to the implementation and effectiveness of interventions. Studies will first be selected by title and abstract (if available) and then by reading the full text by at least two independent reviewers. We will analyze existing scientific evidence using scoping review and evidence gap mapping methodologies. Results:The results will be presented in interactive evidence maps, summarized in a narrative synthesis, and serve as input for subsequent research. Review registration:This protocol has been pre-registered on OSF under doi https://doi.org/10.17605/OSF.IO/D65YS.
Background Many interventions, especially those linked to open science, have been proposed to improve reproducibility in science. To what extent these propositions are based on scientific evidence from empirical evaluations is not clear. Aims The primary objective is to identify Open Science interventions that have been formally investigated regarding their influence on reproducibility and replicability. A secondary objective is to list any facilitators or barriers reported and to identify gaps in the evidence. Methods We will search broadly by using electronic bibliographic databases, broad internet search, and contacting experts in the field of reproducibility, replicability, and open science. Any study investigating interventions for their influence on the reproducibility and replicability of research will be selected, including those studies additionally investigating drivers and barriers to the implementation and effectiveness of interventions. Studies will first be selected by title and abstract (if available) and then by reading the full text by at least two independent reviewers. We will analyze existing scientific evidence using scoping review and evidence gap mapping methodologies. Results The results will be presented in interactive evidence maps, summarized in a narrative synthesis, and serve as input for subsequent research. Review registration This protocol has been pre-registered on OSF under doi https://doi.org/10.17605/OSF.IO/D65YS
OBJECTIVES:To give an overview of methods for updating artificial intelligence (AI)-based clinical prediction models based on new data. STUDY DESIGN AND SETTING:We comprehensively searched Scopus and Embase up to August 2022 for articles that addressed developments, descriptions, or evaluations of prediction model updating methods. We specifically focused on articles in the medical domain involving AI-based prediction models that were updated based on new data, excluding regression-based updating methods as these have been extensively discussed elsewhere. We categorized and described the identified methods used to update the AI-based prediction model as well as the use cases in which they were used. RESULTS:We included 78 articles. The majority of the included articles discussed updating for neural network methods (93.6%) with medical images as input data (65.4%). In many articles (51.3%) existing, pretrained models for broad tasks were updated to perform specialized clinical tasks. Other common reasons for model updating were to address changes in the data over time and cross-center differences; however, more unique use cases were also identified, such as updating a model from a broad population to a specific individual. We categorized the identified model updating methods into four categories: neural network-specific methods (described in 92.3% of the articles), ensemble-specific methods (2.5%), model-agnostic methods (9.0%), and other (1.3%). Variations of neural network-specific methods are further categorized based on the following: (1) the part of the original neural network that is kept, (2) whether and how the original neural network is extended with new parameters, and (3) to what extent the original neural network parameters are adjusted to the new data. The most frequently occurring method (n = 30) involved selecting the first layer(s) of an existing neural network, appending new, randomly initialized layers, and then optimizing the entire neural network. CONCLUSION:We identified many ways to adjust or update AI-based prediction models based on new data, within a large variety of use cases. Updating methods for AI-based prediction models other than neural networks (eg, random forest) appear to be underexplored in clinical prediction research. PLAIN LANGUAGE SUMMARY:AI-based prediction models are increasingly used in health care, helping clinicians with diagnosing diseases, guiding treatment decisions, and informing patients. However, these prediction models do not always work well when applied to hospitals, patient populations, or times different from those used to develop the models. Developing new models for every situation is neither practical nor desired, as it wastes resources, time, and existing knowledge. A more efficient approach is to adjust existing models to new contexts ('updating'), but there is limited guidance on how to do this for AI-based clinical prediction models. To address this, we reviewed 78 studies in detail to understand how researchers are currently updating AI-based clinical prediction models, and the types of situations in which these updating methods are used. Our findings provide a comprehensive overview of the available methods to update existing models. This is intended to serve as guidance and inspiration for researchers. Ultimately, this can lead to better reuse of existing models and improve the quality and efficiency of AI-based prediction models in health care.
Childhood stunting is associated with impaired cognitive development and increased risk of infections, morbidity, and mortality. The composition of the enteric microbiota may contribute to the pathogenesis of stunting. We systematically reviewed and synthesized data from studies using high-throughput genomic sequencing methods to characterize the gut microbiome in stunted versus non-stunted children under 5 years in LMICs. We included 14 studies from Asia, Africa, and South America. Most studies did not report any significant differences in the alpha diversity, while a significantly higher beta diversity was observed in stunted children in four out of seven studies that reported beta diversity. At the phylum level, inconsistent associations with stunting were observed for Bacillota, Pseudomonadota, and Bacteroidota phyla. No single genus was associated with stunted children across all 14 studies, and some associations were incongruent by specific genera. Nonetheless, stunting was associated with an abundance of pathobionts that could drive inflammation, such as Escherichia/Shigella and Campylobacter, and a reduction of butyrate producers, including Faecalibacterium, Megasphera, Blautia, and increased Ruminoccoccus. An abundance of taxa thought to originate in the oropharynx was also reported in duodenal and fecal samples of stunted children, while metabolic pathways, including purine and pyrimidine biosynthesis, vitamin B biosynthesis, and carbohydrate and amino acid degradation pathways, predicted linear growth. Current studies show that stunted children can have distinct microbial patterns compared to non-stunted children, which could contribute to the pathogenesis of stunting.
Systematic reviews (SRs) are time-consuming and labor-intensive to perform. With the growing number of scientific publications, the SR development process becomes even more laborious. This is problematic because timely SR evidence is essential for decision-making in evidence-based healthcare and policymaking. Numerous methods and tools that accelerate SR development have recently emerged. To date, no scoping review has been conducted to provide a comprehensive summary of methods and ready-to-use tools to improve efficiency in SR production. To present an overview of primary studies that evaluated the use of ready-to-use applications of tools or review methods to improve efficiency in the review process. We conducted a scoping review. An information specialist performed a systematic literature search in four databases, supplemented with citation-based and grey literature searching. We included studies reporting the performance of methods and ready-to-use tools for improving efficiency when producing or updating a SR in the health field. We performed dual, independent title and abstract screening, full-text selection, and data extraction. The results were analyzed descriptively and presented narratively. We included 103 studies: 51 studies reported on methods, 54 studies on tools, and 2 studies reported on both methods and tools to make SR production more efficient. A total of 72 studies evaluated the validity (n = 69) or usability (n = 3) of one method (n = 33) or tool (n = 39), and 31 studies performed comparative analyses of different methods (n = 15) or tools (n = 16). 20 studies conducted prospective evaluations in real-time workflows. Most studies evaluated methods or tools that aimed at screening titles and abstracts (n = 42) and literature searching (n = 24), while for other steps of the SR process, only a few studies were found. Regarding the outcomes included, most studies reported on validity outcomes (n = 84), while outcomes such as impact on results (n = 23), time-saving (n = 24), usability (n = 13), and cost-saving (n = 3) were less often evaluated. For title and abstract screening and literature searching, various evaluated methods and tools are available that aim at improving the efficiency of SR production. However, only few studies have addressed the influence of these methods and tools in real-world workflows. Few studies exist that evaluate methods or tools supporting the remaining tasks. Additionally, while validity outcomes are frequently reported, there is a lack of evaluation regarding other outcomes.
BACKGROUND:Sample collection is a key driver of accuracy in the diagnosis of SARS-CoV-2 infection. Viral load may vary at different anatomical sampling sites and accuracy may be compromised by difficulties obtaining specimens and the expertise of the person taking the sample. It is important to optimise sampling accuracy within cost, safety and accessibility constraints. OBJECTIVES:To compare the sensitivity of different sampling collection sites and methods for the detection of current SARS-CoV-2 infection with any molecular or antigen-based test. SEARCH METHODS:Electronic searches of the Cochrane COVID-19 Study Register and the COVID-19 Living Evidence Database from the University of Bern (which includes daily updates from PubMed and Embase and preprints from medRxiv and bioRxiv) were undertaken on 22 February 2022. We included independent evaluations from national reference laboratories, FIND and the Diagnostics Global Health website. We did not apply language restrictions. SELECTION CRITERIA:We included studies of symptomatic or asymptomatic people with suspected SARS-CoV-2 infection undergoing testing. We included studies of any design that compared results from different sample types (anatomical location, operator, collection device) collected from the same participant within a 24-hour period. DATA COLLECTION AND ANALYSIS:Within a sample pair, we defined a reference sample and an index sample collected from the same participant within the same clinical encounter (within 24 hours). Where the sample comparison was different anatomical sites, the reference standard was defined as a nasopharyngeal or combined naso/oropharyngeal sample collected into the same sample container and the index sample as the alternative anatomical site. Where the sample comparison was concerned with differences in the sample collection method from the same site, we defined the reference sample as that closest to standard practice for that sample type. Where the sample pair comparison was concerned with differences in personnel collecting the sample, the more skilled or experienced operator was considered the reference sample. Two review authors independently assessed the risk of bias and applicability concerns using the QUADAS-2 and QUADAS-C checklists, tailored to this review. We present estimates of the difference in the sensitivity (reference sample (%) minus index sample sensitivity (%)) in a pair and as an average across studies for each index sampling method using forest plots and tables. We examined heterogeneity between studies according to population (age, symptom status) and index sample (time post-symptom onset, operator expertise, use of transport medium) characteristics. MAIN RESULTS:This review includes 106 studies reporting 154 evaluations and 60,523 sample pair comparisons, of which 11,045 had SARS-CoV-2 infection. Ninety evaluations were of saliva samples, 37 nasal, seven oropharyngeal, six gargle, six oral and four combined nasal/oropharyngeal samples. Four evaluations were of the effect of operator expertise on the accuracy of three different sample types. The majority of included evaluations (146) used molecular tests, of which 140 used RT-PCR (reverse transcription polymerase chain reaction). Eight evaluations were of nasal samples used with Ag-RDTs (rapid antigen tests). The majority of studies were conducted in Europe (35/106, 33%) or the USA (27%) and conducted in dedicated COVID-19 testing clinics or in ambulatory hospital settings (53%). Targeted screening or contact tracing accounted for only 4% of evaluations. Where reported, the majority of evaluations were of adults (91/154, 59%), 28 (18%) were in mixed populations with only seven (4%) in children. The median prevalence of confirmed SARS-CoV-2 was 23% (interquartile (IQR) 13%-40%). Risk of bias and applicability assessment were hampered by poor reporting in 77% and 65% of included studies, respectively. Risk of bias was low across all domains in only 3% of evaluations due to inappropriate inclusion or exclusion criteria, unclear recruitment, lack of blinding, nonrandomised sampling order or differences in testing kit within a sample pair. Sixty-eight percent of evaluation cohorts were judged as being at high or unclear applicability concern either due to inflation of the prevalence of SARS-CoV-2 infection in study populations by selectively including individuals with confirmed PCR-positive samples or because there was insufficient detail to allow replication of sample collection. When used with RT-PCR • There was no evidence of a difference in sensitivity between gargle and nasopharyngeal samples (on average -1 percentage points, 95% CI -5 to +2, based on 6 evaluations, 2138 sample pairs, of which 389 had SARS-CoV-2). • There was no evidence of a difference in sensitivity between saliva collection from the deep throat and nasopharyngeal samples (on average +10 percentage points, 95% CI -1 to +21, based on 2192 sample pairs, of which 730 had SARS-CoV-2). • There was evidence that saliva collection using spitting, drooling or salivating was on average -12 percentage points less sensitive (95% CI -16 to -8, based on 27,253 sample pairs, of which 4636 had SARS-CoV-2) compared to nasopharyngeal samples. We did not find any evidence of a difference in the sensitivity of saliva collected using spitting, drooling or salivating (sensitivity difference: range from -13 percentage points (spit) to -21 percentage points (salivate)). • Nasal samples (anterior and mid-turbinate collection combined) were, on average, 12 percentage points less sensitive compared to nasopharyngeal samples (95% CI -17 to -7), based on 9291 sample pairs, of which 1485 had SARS-CoV-2. We did not find any evidence of a difference in sensitivity between nasal samples collected from the mid-turbinates (3942 sample pairs) or from the anterior nares (8272 sample pairs). • There was evidence that oropharyngeal samples were, on average, 17 percentage points less sensitive than nasopharyngeal samples (95% CI -29 to -5), based on seven evaluations, 2522 sample pairs, of which 511 had SARS-CoV-2. A much smaller volume of evidence was available for combined nasal/oropharyngeal samples and oral samples. Age, symptom status and use of transport media do not appear to affect the sensitivity of saliva samples and nasal samples. When used with Ag-RDTs • There was no evidence of a difference in sensitivity between nasal samples compared to nasopharyngeal samples (sensitivity, on average, 0 percentage points -0.2 to +0.2, based on 3688 sample pairs, of which 535 had SARS-CoV-2). AUTHORS' CONCLUSIONS:When used with RT-PCR, there is no evidence for a difference in sensitivity of self-collected gargle or deep-throat saliva samples compared to nasopharyngeal samples collected by healthcare workers when used with RT-PCR. Use of these alternative, self-collected sample types has the potential to reduce cost and discomfort and improve the safety of sampling by reducing risk of transmission from aerosol spread which occurs as a result of coughing and gagging during the nasopharyngeal or oropharyngeal sample collection procedure. This may, in turn, improve access to and uptake of testing. Other types of saliva, nasal, oral and oropharyngeal samples are, on average, less sensitive compared to healthcare worker-collected nasopharyngeal samples, and it is unlikely that sensitivities of this magnitude would be acceptable for confirmation of SARS-CoV-2 infection with RT-PCR. When used with Ag-RDTs, there is no evidence of a difference in sensitivity between nasal samples and healthcare worker-collected nasopharyngeal samples for detecting SARS-CoV-2. The implications of this for self-testing are unclear as evaluations did not report whether nasal samples were self-collected or collected by healthcare workers. Further research is needed in asymptomatic individuals, children and in Ag-RDTs, and to investigate the effect of operator expertise on accuracy. Quality assessment of the evidence base underpinning these conclusions was restricted by poor reporting. There is a need for further high-quality studies, adhering to reporting standards for test accuracy studies.
Functional constipation is common in children and accurate diagnostic methods are essential for early diagnosis and effective management. The diagnostic accuracy of transabdominal ultrasound to diagnose functional constipation is unclear. To evaluate the diagnostic accuracy of transverse rectal diameter measurement via transabdominal ultrasound in diagnosing children with functional constipation and in identifying fecal impaction. Electronic databases were searched from inception to March 2023. Original studies investigating the diagnostic accuracy of measuring transverse rectal diameter via transabdominal ultrasound, including children with and without functional constipation, or with and without fecal impaction were included. Data extraction and quality assessment were performed independently by two reviewers. Sixteen studies were included (n = 1,801 children, 0–17 years). Thirteen studies investigated the diagnostic accuracy for functional constipation, and five for fecal impaction. High risk of bias was found across the majority of studies mainly due to un-blinded case–control designs. Cut-off transverse rectal diameter values to diagnose functional constipation ranged from 2.4 cm to 3.8 cm. Meta-analysis (seven studies, n = 509 children) estimated mean sensitivity and specificity to diagnose functional constipation were 0.68 (95