Background and Aims:Immunotherapies for cancers are tested in large numbers of clinical trials. It is difficult for clinicians and researchers to stay current with the evidence, and traditional systematic reviews and clinical guidelines are not suited to ensure a continued overview of all trials and their results. To address this problem, we designed the Cancer Immunotherapy Evidence Living (CIEL) Library. Methods:We included planned, ongoing, and completed interventional trials of immunotherapies for cancer, regardless of trial design (e.g., randomization, blinding, and type of comparator). We systematically searched PubMed (for published reports) and ClinicalTrials.gov (for registered clinical trials). PubMed-retrieved records were screened using the AI-assisted software ASReview and manually extracted and curated. We imported data from ClinicalTrials.gov using the Clinical Trials Transformation Initiative database, which then requires further curation. The CIEL-Library was implemented as a web application. It also contained the "Match My Patient" feature, a patient-centered clinical decision support system, aiming to filter planned, ongoing, or completed trials based on four patient characteristics (disease staging, previous treatments, performance status, and location). We piloted our database with one type of cancer immunotherapy, the tumor-infiltrating lymphocytes (TIL) transfer. The CIEL-Library was a prototype and no further developments are planned. Conclusions:The CIEL-Library offers a blueprint for a dynamic evidence synthesis infrastructure by providing a collection of clinical trials with curated trial characteristics and results. This blueprint may be applied across fields, specialties, and topics. The main challenges to making a database of clinical trials are the time and resources needed to populate it with curated and updated data. The CIEL-Library project highlights the potential and the main limitations of designing trial databases intended to be used in routine care.
INTRODUCTION:Conducting systematic reviews of clinical trials is time-consuming and resource-intensive. One potential solution is to design databases that are continuously and automatically populated with clinical trial data from harmonised and structured datasets. This scoping review aimed to identify and map publicly available, continuously updated, topic-specific databases of clinical trials. METHODS:We systematically searched PubMed, Embase, the preprint servers medRxiv, arXiv, Open Science Framework, and Google. We characterised each database using seven predefined features (access model, database type, data input sources, retrieval methods, data-extraction methods, trial presentation, and export options) and narratively summarised the results. RESULTS:We identified 14 continuously updated databases of clinical trials, seven related to COVID-19 (initiated in 2020) and seven non-COVID-19 databases (initiated as early as in 2009). All databases, except one, were publicly funded and accessible without restrictions. Most relied on traditional methods used in static article-based systematic reviews sourcing data from journal publications and trial registries. The COVID-19 databases and some non-COVID-19 databases implemented semi-automated features of data import, which combined automated and manual data curation, whereas the non-COVID-19 databases mainly relied on manual workflows. Most reported information was metadata, such as author names, years of publication, and link to publication or trial registry. Only two databases included trial appraisal information (such as risk of bias assessments). Six databases reported aggregate group-level results, but only one database provided individual participant data on request. DISCUSSION:Continuously updated topic-specific databases of clinical trials remain limited in number, and existing initiatives mainly employ traditional static systematic review methodologies. A key barrier to developing truly living platforms is the lack of accessible, machine-readable, and standardised clinical trial data.
Introduction There is an urgent need to better understand how information from circulating tumour DNA (ctDNA) can be integrated into routine care for patients with advanced solid cancer.Methods and analysis The implementation of liquid biopsies in routine care of patients with advanced solid cancer trial (LIQPLAT) is a single-centre, single-arm trial investigating the implementation of ctDNA in the routine care of patients with advanced solid cancer. We present a mixed-methods process evaluation embedded in the LIQPLAT trial, following Medical Research Council guidance and the Reach, Effectiveness, Adoption, Implementation, Maintenance framework. We show a logic model, which details the causal chain and related assumptions from recruiting patients into the trial to the goal of improving quality of life and survival. Data collection is longitudinal and includes: semistructured interviews with healthcare professionals (pathologists, biologists, oncologists; planned n=20) and patients (planned n=15) to identify implementation barriers and facilitators; recordings of molecular tumour board meetings to analyse clinical decision-making; the 23-item Normalisation MeAsure Development survey for healthcare professionals (planned n=20) at four time points. Quantitative data from hospital records will be used to assess implementation outcomes like patient acceptance rates and ctDNA workflow success. Qualitative data will undergo thematic and content analysis, and quantitative data will be analysed using a Bayesian framework.Ethics and dissemination The LIQPLAT trial was approved by the regional ethics committee of Northwestern and Central Switzerland (BASEC 2024-00358). The qualitative aspects of the process evaluation were exempted from ethics review according to the Swiss Human Research Act. We follow guidelines for data security, confidentiality and information governance. Results will be submitted for publication in peer-reviewed journals and discussed at conferences.Trial registration number NCT06367751, SNCTP000005844.
OBJECTIVE To compare treatment effects from trials designed as decentralised trials or as traditional non-decentralised trials. DESIGN Meta-epidemiological study. DATA SOURCE PubMed database. ELIGIBILITY CRITERIA Trials identified as decentralised and included in meta-analyses alongside corresponding non-decentralised trials, restricted to drug interventions across all populations and outcome types. RESULTS 51 decentralised trials and 86 non-decentralised trials regarding 11 clinical questions were compared. Decentralised trials were larger than non-decentralised trials (median sample size 1175 v 236 participants) and more recent than non-decentralised trials (median year of publication 2012 v 2007). The level of significance agreed between decentralised and non-decentralised trials in nine of 11 clinical questions (82%). No systematic difference between effects from decentralised and non-decentralised trials was found (summary ratio of odds ratio 1.01; 95% confidence interval 0.93 to 1.09; heterogeneity I2=8.6%) with a median absolute deviation of 1.3-fold (absolute ratio of odds ratio 1.30; interquartile range 1.13-1.51). CONCLUSION Decentralised trials provide similar answers to clinical questions as traditional, non-decentralised randomised trials. They do not systematically overestimate or underestimate effect estimates, but modest absolute deviations occur that may lead to different conclusions. Beyond direct benefits to participants, decentralisation opens new avenues for randomised trials, including the facilitation of larger trials.
ASReview is a software that can potentially reduce the workload of literature screening in systematic reviews by ranking the retrieved records. We assessed the tool's feasibility, advantages, and limitations, to populate a database of cancer immunotherapy trials. ASReview is easy to use, and it efficiently identified relevant records. It may save resources compared to traditional systematic reviews using two human reviewers. Predefined procedures are necessary to maintain a transparent and reproducible workflow. Limitations include that adding references to existing projects is difficult and that the algorithm learns from every decision, even when this may not be appropriate.
This paper reviews the scientific evidence on new anti-amyloid monoclonal antibodies for treating Alzheimer's disease as a case study for improving scientific evidence communication. We introduce five guidelines condensed from the biomedical evidence literature but adapted to the short format of science communication in e.g. journal opinion pieces and newspaper articles. Given the major importance and recent confusion regarding the discussed drugs, with certain disagreements seen e.g. between FDA and EMA, the suggested guidelines may be useful to clinicians discussing with their patients and to scientists communicating the evidence in balance. More generally, we hope that the guidelines may help us to improve communication of scientific evidence on complex topics in opinion pieces in the scientific literature, in advocacy, and in media appearances.
Single-arm trials can be used to explore the feasibility, implementation, and effects of treatment. They typically use opportunistic convenience sampling to find potential participants. Their main limitations for health care decision making are lack of generalizability and the poor quality of the comparative evidence they produce. The authors propose a single-arm trial design that can provide greater generalizability and higher quality of comparative evidence than traditional single-arm trials, called a random invitation single-arm trial or RISAT. A RISAT has 4 essential components. First, it has a data infrastructure for routinely collected real-world data (RWD) where participants have provided consent to have their data used for research (for example, a registry or electronic medical record database). Second, a subset of those participants are randomly invited to take part in the RISAT. Those not invited are not contacted. Third, invitees are offered the specific intervention (such as a novel treatment), to which they consent or not. Fourth, all invitees are followed prospectively regardless of their acceptance of the intervention. For an optional randomized comparison, RISATs can use the RWD infrastructure to measure outcomes from invitees and noninvitees. The authors describe the advantages and challenges of this approach, including inferential issues and biases and comparison with other designs. They show how RISATs allow fairer access to participation, improve applicability and generalizability of results compared with traditional single-arm trials, and provide a form of randomized real-world evidence. At negligible added cost beyond already existing infrastructure, this approach may catalyze the early generation of evidence of higher value than that produced with traditional single-arm trials, increasing the credibility and validity of accelerated drug approval processes and enabling better health care decisions.
Background: Advanced solid cancers present significant treatment challenges due to their genomic heterogeneity and resistance. Liquid biopsies, specifically circulating tumour DNA (ctDNA), have emerged as promising tools to support treatment decision making. However, evidence regarding their implementation in routine care remains limited. Methods: LIQPLAT is a single-arm trial (SAT), assessing the feasibility of implementing ctDNA measurements in the usual care of patients with advanced solid cancers excepting primary brain tumours, at the University Hospital Basel. Patients are randomly invited from an ongoing research registry to take part in the SAT. We aim to include 150 randomly invited patients to receive ctDNA measurements alongside standard care. CtDNA samples are collected at baseline, between the second and third months, fifth and seventh months after cancer treatment start, and during serious clinical events such as disease progression or treatment changes. Results are evaluated by a molecular tumor board to guide clinical management. Feasibility outcomes include detectability of ctDNA, identification of actionable alterations, and analysis turnaround time. Other outcomes include patient-reported quality of life, progression-free survival and overall survival, time to next treatment line, and unscheduled hospital and emergency visits, all obtained from routine healthcare data. Discussion: LIQPLAT will examine the implementation of ctDNA measurements in routine care. The random selection for invitation to this SAT within an existing registry embedded in routine care creates a representative sample and allows for better assessment of implementation and generalization of findings. Trial registration: The trial is registered at clinicaltrials.gov (2024-04-11, [NCT06367751][1]) and kofam.ch (2024-03-15, SNCTP000005844). ### Competing Interest Statement BK: Consulting fees: Roche, Dayton Therapeutics, Pharma&. Payment or honoraria for lectures/presentations: Incyte, Roche. Support for attending meetings/travel: Pfizer. HL: Royalties or licenses: Sig9 mAb, Ono Pharmaceuticals. Consulting fees: Servier, Novocure, contributions made to his institution. Stock or stock options: Roche, Novartis. JMS: Grants: work on LIQPLAT is funded by an SNSF grant (Grant No. 323530\_229969). MB: Consulting fees: Jazz Pharmaceuticals. Support for attending meetings/travel: MSD. MSM: Grants: Swiss National Science Foundation (SNSF; Grant No. 320030\_189275). Consulting fees: Thermo Fisher, Merck, GlaxoSmithKline, Janssen-Cilag, Roche, Novartis, Engimmune. Payment or honoraria for lectures: Incyte Biosciences, Astellas Pharma GmbH. All contributions made to his institution. PJ, LGH: The Research Center for Clinical Neuroimmunology and Neuroscience Basel (RC2NB) is supported by the Foundation Clinical Neuroimmunology and Neuroscience Basel. RC2NB has a contract with Roche for a steering committee participation of LGH. All other authors have declared no competing interests. ### Clinical Trial NCT06367751 ### Funding Statement This trial is currently being funded by the Krebsliga beider Basel (KLbB-5845-02-2023) and the Division of Medical Oncology. Krebsliga beider Basel has no role in any aspect of the study and writing of the manuscript. JMS is funded by the Swiss National Science Foundation (grant number: 323530_229969). ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: The Ethics committee of northwestern and central Switzerland (EKNZ) reviewed and gave ethical approval for this work (BASEC 2024-00358). I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes All data produced in the present study will be available upon reasonable request to the authors. [1]: /lookup/external-ref?link_type=CLINTRIALGOV&access_num=NCT06367751&atom=%2Fmedrxiv%2Fearly%2F2025%2F02%2F13%2F2025.02.13.25322206.atom
BACKGROUND:Serum neurofilament light (sNfL) chain levels, a sensitive measure of disease activity in multiple sclerosis (MS), are increasingly considered for individual therapy optimization yet without consensus on their use for clinical application. OBJECTIVE:We here propose treatment decision algorithms incorporating sNfL levels to adapt disease-modifying therapies (DMTs). METHODS:We conducted a modified Delphi study to reach consensus on algorithms using sNfL within typical clinical scenarios. sNfL levels were defined as "high" (>90th percentile) vs "normal" (<80th percentile), based on normative values of control persons. In three rounds, 10 international and 18 Swiss MS experts, and 3 patient consultants rated their agreement on treatment algorithms. Consensus thresholds were defined as moderate (50%-79%), broad (80%-94%), strong (≥95%), and full (100%). RESULTS:The Delphi provided 9 escalation algorithms (e.g. initiating treatment based on high sNfL), 11 horizontal switch (e.g. switching natalizumab to another high-efficacy DMT based on high sNfL), and 3 de-escalation (e.g. stopping DMT or extending intervals in B-cell depleting therapies). CONCLUSION:The consensus reached on typical clinical scenarios provides the basis for using sNfL to inform treatment decisions in a randomized pragmatic trial, an important step to gather robust evidence for using sNfL to inform personalized treatment decisions in clinical practice.
Background and Objective The Pragmatic-Explanatory Continuum Indicator Summary-2 (PRECIS-2) tool has been widely used to help investigators design randomized trials, facilitating the task of aligning design choices with an explanatory or pragmatic primary trial intention. PRECIS-2 is increasingly being used to retrospectively assess the degree of pragmatism or explanatoriness among published trials within reviews. There is little information on the interrater reliability of the tool and no consensus on the preferred method of achieving an accurate and reliable judgment of trial “pragmatism” when using PRECIS-2 retrospectively. The aims of this study were to assess the level of pragmatism or explanatoriness of trials that cite PRECIS-2 and to assess interrater reliability of PRECIS-2 using different scoring approaches. We compared agreement between two independent ratings within a single pair with agreement between consensus scores reached by two independent pairs of reviewers and whether widening the agreement criteria increased interrater reliability. Methods Thirty randomized controlled trials (RCTs) were randomly selected from trials citing the PRECIS-2 tool. Two pairs of reviewers, a clinician paired with a methodologist in each case, were trained and independently scored each trial and reached a consensus score within pairs. Agreement between reviewers within pairs and between consensus scores across pairs was assessed using kappa statistics for each of the nine PRECIS-2 domains. Results RCTs citing PRECIS-2 had predominantly pragmatic design features. Interrater reliability within pairs was low across all domains, with the highest levels found in the two domains of analysis (0.32) and follow-up (0.33). Agreement across pairs on the consensus scores was similarly low. Agreement between reviewers and reviewer pairs was above 70% when agreement was reclassified as “within 1-point difference on the scoring scale” for eight domains, but no improvement was obtained for the remaining domain. Conclusion Trials citing PRECIS-2 tend to have predominantly pragmatic design features. When using PRECIS-2 to retrospectively score trial publications, agreement between consensus scores across pairs of reviewers was no better than agreement within pairs. Reconfiguring the PRECIS scoring scale and improving scoring guidance may provide a more meaningful, easily interpreted measure of “pragmatism” for trialists wishing to use PRECIS-2 as a review tool. Plain Language Summary The Pragmatic-Explanatory Continuum Indicator Summary-2 (PRECIS-2) tool is designed to help researchers match their design decisions to the intended purpose of their trial. The intention of a trial can be “explanatory,” which improves our understanding of how an intervention works, or “pragmatic,” which supports decision-making in health care. Increasingly, the tool has been used for a secondary purpose: in systematic reviews. Here the tool is used to judge the level of “pragmatism” or “explanatoriness” of trials included in the review to aid the understanding of trial results. However, there is debate on the most reliable means of making this judgment. Sometimes judgements are made using one reviewer; other times, multiple reviewers. Our study evaluated interrater reliability of two methods of scoring trial publications using PRECIS-2: individual reviewer scores and pairs of reviewers agreeing on a consensus score. We also found that neither method we tested produced a reliable judgment using PRECIS-2, and the scores from two reviewers agreeing on a consensus were no more reliable than scores from a single reviewer. We performed an additional analysis that showed that simplifying the scoring from the original five-point scale to a three-point scale may give a more reliable judgment of the “pragmatism” or “explanatioriness” of published trials. This simpler method of scoring should be encouraged for retrospective use of PRECIS-2 in systematic reviews.
RATIONALE:People with cardiovascular disease are at risk of recurrent major adverse cardiovascular events, and chronic low-grade inflammation may be a major underlying factor. Treatment with low-dose colchicine has been proposed for the secondary prevention of cardiovascular events in individuals at high cardiovascular risk. A previous Cochrane review showed considerable uncertainty regarding the benefits and harms of this approach. OBJECTIVES:To evaluate the benefits and harms of low-dose colchicine in the prevention of cardiovascular events in adults with a history of stable CVD or following myocardial infarction or stroke. SEARCH METHODS:We conducted a comprehensive search of the literature until February 2025 using Cochrane Central Register of Controlled Trials (CENTRAL), MEDLINE, EMBASE, the drugs@FDA database, references of key papers, and references of included studies. ELIGIBILITY CRITERIA:Randomised controlled trials (RCTs) comparing the use of low-dose colchicine for a minimum of six months versus any control intervention in patients of any age with cardiovascular disease (i.e. history of stable cardiovascular disease, previous myocardial infarction or stroke). OUTCOMES:Our critical outcomes were all-cause mortality, myocardial infarction, and serious adverse events. Our important outcomes were cardiovascular mortality, stroke, all-cause hospitalisations, coronary revascularisation (percutaneous coronary intervention (PCI)/angioplasty or coronary artery bypass graft (CABG)), quality of life, and gastrointestinal adverse events (i.e. diarrhoea, nausea, abdominal pain, or vomiting). RISK OF BIAS:Two authors independently assessed the risk of bias using the Cochrane RoB2 tool. SYNTHESIS METHODS:We conducted meta-analyses using the random-effects model. We generated forest plots to facilitate visualisation of the data. We did not perform any subgroup analysis. We used GRADE to assess the certainty of evidence for all critical outcomes and for cardiovascular mortality, stroke, and coronary revascularisation. This was carried out by two review authors working independently. INCLUDED STUDIES:We included 12 studies involving 22,983 randomised participants. The follow-up in the studies ranged from 6 to 80 months. Overall, 11,524 participants were assigned to low-dose colchicine treatment and 11,459 were assigned to a control intervention, which constituted either usual care plus placebo or usual care only. The doses of colchicine used were 0.5 mg once or twice daily. At baseline, the mean age of participants ranged from 57 to 74 years. Most participants (79.4%) were male. SYNTHESIS OF RESULTS:There is high-certainty evidence that low-dose colchicine treatment reduces the risk of myocardial infarction, with a risk ratio (RR) of 0.74 (95% confidence interval (CI) 0.57 to 0.96; 22,153 participants, 8 studies; I2 = 51%), yielding an absolute risk reduction of 9 fewer events (95% CI 16 fewer to 2 fewer) per 1000 patients, when the myocardial infarction rate is about 4% (36 events per 1000 patients) in the control group. There is also high-certainty evidence that low-dose colchicine reduces the risk of stroke with a RR of 0.67 (95% CI 0.47 to 0.95; 22,483 participants, 10 studies; I2 = 40%), yielding an absolute risk reduction of 8 fewer events (95% CI 12 fewer to 1 fewer) per 1000 patients, when the stroke rate is about 2% (22 events per 1000 patients) in the control group. There is high-certainty evidence that the use of low-dose colchicine does not increase the rate of serious adverse events (RR 0.98, 95% CI 0.94 to 1.02; 15,677 participants, 4 studies; I2 = 0%). However, gastrointestinal adverse events were more common under treatment with colchicine (RR 1.68, 95% CI 1.11 to 2.57; 22,185 participants, 10 studies; I2 = 91%). For all other outcomes assessed, the evidence is of moderate certainty. Colchicine probably results in little to no difference in all-cause mortality (RR 1.01, 95% CI 0.84 to 1.21; 22,747 participants, 10 studies; I2 = 1%; moderate-certainty evidence), in cardiovascular mortality (RR 0.94, 95% CI 0.73 to 1.22; 22,271 participants; 8 studies; I2 = 13%; moderate-certainty evidence), and coronary revascularisation (RR 0.83, 95% CI 0.64 to 1.08; 13,705 participants, 5 studies; I2 = 40%; moderate-certainty evidence). There is no evidence about the benefits or harms of colchicine on quality-of-life or on the risk of all-cause hospitalisation. AUTHORS' CONCLUSIONS:People with cardiovascular disease using low-dose colchicine as secondary prevention for at least six months benefit from reduced rates of myocardial infarction and stroke, without an increase in serious adverse events. Moderate-certainty evidence did not show a benefit from low-dose colchicine for the risk of mortality (i.e. all-cause and cardiovascular mortality) or coronary revascularisation rates. Colchicine use was associated with an increased risk of gastrointestinal adverse events, which were typically described as mild and transient in nature. Additional studies are warranted to investigate the benefits and harms of low-dose colchicine in relevant subgroups and in specific indications, such as long-term use in individuals with stable coronary artery disease versus limited-time use following acute coronary syndrome. FUNDING:Review author FE was supported by the Margot und Erich Goldschmidt & Peter René Jacobson Foundation. Review author CMS was supported by the Janggen Pöhn Foundation and the Swiss National Science Foundation (MD-PhD grant Number: 323530_221860). REGISTRATION:This review is based on its protocol, which is available via DOI 10.1002/14651858.CD014808, and a previous review, which is available via DOI 10.1002/14651858.CD011047.pub2.
Background and Objectives Pragmatic trials are increasingly gaining recognition. However, what pragmatic trials are is frequently misunderstood. They are frequently described superficially by their manifestation and surface only, as studies conducted in “real world” settings, having wide inclusion criteria, and less complicated study procedures. However, these features are neither necessary nor defining characteristics. They also do not guarantee that trials sharing them are useful to inform medical practice. There is a danger of losing sight of the essence of the powerful pragmatic approach. Methods, Results, and Conclusion Here we describe the key elements of the pragmatic approach and the close relationship with the original nature of randomized trials. Our aim is to refocus teaching, research and interpretation of evidence, not as a novel approach but as a return towards the essence of pragmatic evidence and the nature of randomized trials. We first go back to the origin of pragmatism in philosophy and its introduction in medicine and revisit the nature of randomized trials in their pure form. We highlight the critical distinction between assessing treatment decisions and understanding the mechanisms of these decisions. We show why the current view on randomized trials in medicine has lost a pragmatic focus, with the explanatory design features blinding and adherence control often seen as defining characteristics or quality criteria of randomized trials. We then highlight common misunderstandings of pragmatic trials and conclude with an overview of their key features to provide pragmatic evidence.
OBJECTIVES:To systematically evaluate timely reporting of clinical trial results at medical universities and university hospitals in the Nordic countries. STUDY DESIGN AND SETTING:In this cross-sectional study, we included trials (regardless of intervention) registered in the European Union (EU) Clinical Trials Registry and/or ClinicalTrials.gov, completed 2016-2019 and led by a university with medical faculty or university hospital in Denmark, Finland, Iceland, Norway, or Sweden. We identified summary results posted at the trial registries and conducted systematic manual searches for results publications (eg, journal articles, preprints). We present proportions with 95% confidence intervals (CI) and medians with interquartile range (IQR). PROTOCOL:https://osf.io/wua3r. RESULTS:Among 2112 included clinical trials, 1650 (78.1%, 95% CI 76.3%-79.8%) reported any results during our follow-up; 1097 (51.9%, 95% CI 49.8%-54.1%) reported any results within 2 years of the global completion date; and 48 (2.3%, 95% CI 1.7%-3.0%) posted summary results in the registry within 1 year. The median time from global completion date to results reporting was 690 days (IQR 1103). 856/1681 (50.9%) of ClinicalTrials.gov registrations were prospective. Denmark contributed approximately half of all trials. Reporting performance varied widely between institutions. CONCLUSION:Missing and delayed results reporting of academically led clinical trials are a pervasive problem in the Nordic countries. We relied on trial registry information, which can be incomplete. Institutions, funders, and policymakers need to support trial teams, ensure regulation adherence, and secure trial reporting before results are permanently lost. PLAIN LANGUAGE SUMMARY:Reporting of results from clinical trials is necessary for evidence-based clinical decision-making. We followed up reporting of clinical trials in the Nordic countries sponsored by medical universities and university hospitals. Of 2112 studies completed 2016-2019 in two major trials registries, about half reported results in any form within 24 months, and more than one in five did not report results at all. These results show that there is a need for improvement in the reporting of Nordic clinical trials.
BACKGROUND AND OBJECTIVE:It is unknown whether large language models (LLMs) may facilitate time- and resource-intensive text-related processes in evidence appraisal. The objective was to quantify the agreement of LLMs with human consensus in appraisal of scientific reporting (Preferred Reporting Items for Systematic reviews and Meta-Analyses [PRISMA]) and methodological rigor (A MeaSurement Tool to Assess systematic Reviews [AMSTAR]) of systematic reviews and design of clinical trials (PRagmatic Explanatory Continuum Indicator Summary 2 [PRECIS-2]) and to identify areas where collaboration between humans and artificial intelligence (AI) would outperform the traditional consensus process of human raters in efficiency. STUDY DESIGN AND SETTING:Five LLMs (Claude-3-Opus, Claude-2, GPT-4, GPT-3.5, Mixtral-8x22B) assessed 112 systematic reviews applying the PRISMA and AMSTAR criteria and 56 randomized controlled trials applying PRECIS-2. We quantified the agreement between human consensus and (1) individual human raters; (2) individual LLMs; (3) combined LLMs approach; (4) human-AI collaboration. Ratings were marked as deferred (undecided) in case of inconsistency between combined LLMs or between the human rater and the LLM. RESULTS:Individual human rater accuracy was 89% for PRISMA and AMSTAR, and 75% for PRECIS-2. Individual LLM accuracy was ranging from 63% (GPT-3.5) to 70% (Claude-3-Opus) for PRISMA, 53% (GPT-3.5) to 74% (Claude-3-Opus) for AMSTAR, and 38% (GPT-4) to 55% (GPT-3.5) for PRECIS-2. Combined LLM ratings led to accuracies of 75%-88% for PRISMA (4%-74% deferred), 74%-89% for AMSTAR (6%-84% deferred), and 64%-79% for PRECIS-2 (29%-88% deferred). Human-AI collaboration resulted in the best accuracies from 89% to 96% for PRISMA (25/35% deferred), 91%-95% for AMSTAR (27/30% deferred), and 80%-86% for PRECIS-2 (76/71% deferred). CONCLUSION:Current LLMs alone appraised evidence worse than humans. Human-AI collaboration may reduce workload for the second human rater for the assessment of reporting (PRISMA) and methodological rigor (AMSTAR) but not for complex tasks such as PRECIS-2.
In medical research as a whole, frequent inaccurate or biased findings are of international concern. One measure against reporting biases is study registration before the start of data collection (preregistration), preferably together with the statistical analysis plan. This meta-research study systematically evaluated registration of Swedish observational research based on national health registries. In a random sample of registry-based observational studies published 2010-2022, very few were preregistered with a publicly available analysis plan (<1 procent). Ideas from the meta-research literature can be leveraged to strengthen the brand of Swedish registry-based observational studies and counteract reporting bias.
Abstract Background Treatment decisions for persons with relapsing–remitting multiple sclerosis (RRMS) rely on clinical and radiological disease activity, the benefit-harm profile of drug therapy, and preferences of patients and physicians. However, there is limited evidence to support evidence-based personalized decision-making on how to adapt disease-modifying therapy treatments targeting no evidence of disease activity, while achieving better patient-relevant outcomes, fewer adverse events, and improved care. Serum neurofilament light chain (sNfL) is a sensitive measure of disease activity that captures and prognosticates disease worsening in RRMS. sNfL might therefore be instrumental for a patient-tailored treatment adaptation. We aim to assess whether 6-monthly sNfL monitoring in addition to usual care improves patient-relevant outcomes compared to usual care alone. Methods Pragmatic multicenter, 1:1 randomized, platform trial embedded in the Swiss Multiple Sclerosis Cohort (SMSC). All patients with RRMS in the SMSC for ≥ 1 year are eligible. We plan to include 915 patients with RRMS, randomly allocated to two groups with different care strategies, one of them new (group A) and one of them usual care (group B). In group A, 6-monthly monitoring of sNfL will together with information on relapses, disability, and magnetic resonance imaging (MRI) inform personalized treatment decisions (e.g., escalation or de-escalation) supported by pre-specified algorithms. In group B, patients will receive usual care with their usual 6- or 12-monthly visits. Two primary outcomes will be used: (1) evidence of disease activity (EDA3: occurrence of relapses, disability worsening, or MRI activity) and (2) quality of life (MQoL-54) using 24-month follow-up. The new treatment strategy with sNfL will be considered superior to usual care if either more patients have no EDA3, or their health-related quality of life increases. Data collection will be embedded within the SMSC using established trial-level quality procedures. Discussion MultiSCRIPT aims to be a platform where research and care are optimally combined to generate evidence to inform personalized decision-making in usual care. This approach aims to foster better personalized treatment and care strategies, at low cost and with rapid translation to clinical practice. Trial registration ClinicalTrials.gov NCT06095271. Registered on October 23, 2023
Conducting systematic reviews of clinical trials is arduous and resource consuming. One potential solution is to design databases that are continuously and automatically populated with clinical trial data from harmonised and structured datasets. We aimed to map publicly available, continuously updated, topic-specific databases of randomised clinical trials (RCTs). We systematically searched PubMed, Embase, the preprint servers medRxiv, ArXiv, and Open Science Framework, and Google. We described seven features (access model, database architecture, data input sources, retrieval methods, data extraction methods, trial presentation, and export options) and narratively summarised the results. We did not register a protocol for this review. We identified 14 continuously updated clinical trial databases, seven related to COVID-19 (first active in 2020) and seven non-COVID databases (first active in 2009). All databases, except one, were publicly funded and accessible without restrictions. They mainly employed methods similar to those from static article-based systematic reviews and retrieved data from journal publications and trial registries. The COVID-19 databases and some non-COVID databases implemented semi-automated features of data import, which combined automated and manual data curation, whereas the non-COVID databases mainly relied on manual workflows. Most reported information was metadata, such as author names, years of publication, and link to publication or trial registry. Two databases included trial appraisal information (risk of bias assessments). Six databases reported aggregate group level results, but only one database provided individual participant data on request. We identified few continuously updated trial databases, and existing initiatives mainly employ methods known from static article -based reviews. The main limitation to create truly live evidence synthesis is the access and import of machine-readable and harmonised clinical trial data.
Background Technological devices such as smartphones, wearables and virtual assistants enable health data collection, serving as digital alternatives to conventional biomarkers. We aimed to provide a systematic overview of emerging literature on ‘digital biomarkers,’ covering definitions, features and citations in biomedical research.Methods We analysed all articles in PubMed that used ‘digital biomarker(s)’ in title or abstract, considering any study involving humans and any review, editorial, perspective or opinion-based articles up to 8 March 2023. We systematically extracted characteristics of publications and research studies, and any definitions and features of ‘digital biomarkers’ mentioned. We described the most influential literature on digital biomarkers and their definitions using thematic categorisations of definitions considering the Food and Drug Administration Biomarkers, EndpointS and other Tools framework (ie, data type, data collection method, purpose of biomarker), analysing structural similarity of definitions by performing text and citation analyses.Results We identified 415 articles using ‘digital biomarker’ between 2014 and 2023 (median 2021). The majority (283 articles; 68%) were primary research. Notably, 287 articles (69%) did not provide a definition of digital biomarkers. Among the 128 articles with definitions, there were 127 different ones. Of these, 78 considered data collection, 56 data type, 50 purpose and 23 included all three components. Those 128 articles with a definition had a median of 6 citations, with the top 10 each presenting distinct definitions.Conclusions The definitions of digital biomarkers vary significantly, indicating a lack of consensus in this emerging field. Our overview highlights key defining characteristics, which could guide the development of a more harmonised accepted definition.
Background: Pragmatic trials are increasingly recognized for providing real-world evidence on treatment choices. Objective: The objective of this study is to investigate the use and characteristics of pragmatic trials in multiple sclerosis (MS). Methods: Systematic literature search and analysis of pragmatic trials on any intervention published up to 2022. The assessment of pragmatism with PRECIS-2 (PRagmatic Explanatory Continuum Indicator Summary-2) is performed. Results: We identified 48 pragmatic trials published 1967–2022 that included a median of 82 participants (interquartile range (IQR) = 42–160) to assess typically supportive care interventions ( n = 41; 85%). Only seven trials assessed drugs (15%). Only three trials (6%) included >500 participants. Trials were mostly from the United Kingdom ( n = 18; 38%), Italy ( n = 6; 13%), the United States and Denmark (each n = 5; 10%). Primary outcomes were diverse, for example, quality-of-life, physical functioning, or disease activity. Only 1 trial (2%) used routinely collected data for outcome ascertainment. No trial was very pragmatic in all design aspects, but 14 trials (29%) were widely pragmatic (i.e. PRECIS-2 score ⩾ 4/5 in all domains). Conclusion: Only few and mostly small pragmatic trials exist in MS which rarely assess drugs. Despite the widely available routine data infrastructures, very few trials utilize them. There is an urgent need to leverage the potential of this pioneering study design to provide useful randomized real-world evidence.