Background: Subgroup analyses and meta-regression are commonly used to investigate heterogeneity in diagnostic test accuracy (DTA) meta-analyses (MA), but adherence to methodological guidance is unclear. This methodological review summarizes investigations of heterogeneity (IoH) in DTA-MAs, examining their frequency, characteristics, and alignment with recommendations. Methods: We included DTA-MAs published in 2024 reporting at least one pair of summary sensitivity and specificity. Non-DTA reviews, narrative syntheses, studies reporting only alternative measures, and overviews of systematic reviews were excluded. MEDLINE (via Ovid) was searched for English-language publications, with the final search in January 2025. Results: From 403 records, the most recent 100 DTA-MAs were included, each contributing one index test. IoH were reported in 61 analyses. The number of primary studies was positively associated with conducting an investigation (OR 1.66; p = 0.008). Subgroup analyses were used in 35/61 (57
Most original articles published in the medical literature report the results of multiple statistical tests. In a few simple cases, there is agreement on whether to adjust for the number of performed tests. For many cases encountered in practice, however, this is less clear, and the recommendations in the literature are contradictory, along different dimensions, or otherwise confusing. This lack of clear guidance may impair the conduct and interpretation of analyses, and encourage questionable research practices, ultimately jeopardizing the credibility of medical research. In this article, we refine, illustrate, and discuss a unifying guiding principle to assist both statisticians and applied researchers in deciding whether to adjust for multiple testing and, if so, over which set of tests. The principle is that multiple testing should be adjusted for if and only if authors, when reporting and interpreting their findings, put more emphasis on results of one or several of the tests because of their small p-value(s). We relate this principle to previously proposed rules and show how it can guide and clarify the choice of adjustment strategies in three complex multiple testing settings.
Background: In clinical research, the Bland-Altman analysis is commonly used to assess agreement of metric measurements made by two or more techniques, devices or methods. The approach can also deal with repeated measurements per subject or observational unit. However, a strong and implicit assumption is that agreement of methods is homogeneous across subjects. Objective: To extend the previously introduced multivariable modeling of conditional method agreement with single measurements per subject to the frequent case of repeated measurements. Methods: Appropriate regression trees, called conditional method agreement trees (COAT), are generalized to capture the dependence of the parameters of the Bland-Altman analysis on covariates. These parameters, the expectation and variance of the differences between the methods, are decomposed into subject-specific components to account for repeated measurements. Whilst the theoretical, asymptotic properties of tree models are known, a simulation study was carried out to assess the performance of COAT in finite samples. A comparison of devices measuring cardiac output serves as an application example. Results: COAT is applicable to the two relevant cases of paired and unpaired repeated measurements. In the simulation study, it controlled the type-I error at the nominal level and could detect covariate-dependent method agreement with increasing sample size. The Adjusted Rand Index, a measure of concordance between the estimated and true subgroups, reached very high values close to the maximum of 1. The analysis of cardiac output showed that patients' characteristics may influence the agreement between measuring devices, with implications for use in patient care. Conclusion: COAT can explicitly define subgroups of heterogeneous method agreement in dependence of covariates with appropriate statistical testing in case of repeated measurements.
Background: Trastuzumab deruxtecan (T-DXd) is active in HER2-expressing solid tumours, but trials excluded HER2 immunohistochemistry (IHC) 1+ disease, and data in pretreated ovarian cancer are lacking. We evaluated real-world T-DXd activity and genomic correlates in pretreated ovarian cancer, predominantly high-grade serous (HGSOC). Methods: HER2 expression was assessed in an unselected ovarian cancer cohort (N=74). Fifteen patients receiving off-label T-DXd (14 HGSOC, 1 clear cell; IHC 1+ to 3+) had HER2 status centrally confirmed using gastric-type criteria. Activity was assessed by intra-patient growth modulation index (GMI; progression-free survival [PFS] on T-DXd divided by PFS on the prior line; ≥ 1.33 considered meaningful). Patients on treatment at data cut-off were censored. Objective response (RECIST 1.1) was assessed centrally where imaging was available (n=8). Results: Of the 40 HER2-expressing tumours, 15 received T-DXd, limited mainly by reimbursement. Among 14 evaluable patients (median 5 prior lines), 9 reached a GMI ≥ 1.33 (median 1.69); 8 remained on treatment at cut-off, making durability preliminary. Confirmed partial responses occurred across the HER2 spectrum. Benefit was independent of homologous-recombination (HR) status: one HR-proficient, CCNE1-wild-type patient achieved prolonged control and was rendered disease-free after radiotherapy to an oligoprogressive lesion. Exploratory analysis showed all four evaluable CCNE1-amplified tumours had reduced or non-durable benefit. Conclusions: T-DXd shows preliminary, clinically meaningful activity in HER2 IHC 1+ ovarian cancer independent of HR status. CCNE1 amplification may attenuate benefit, a candidate biomarker for WEE1-inhibitor combinations. Approval restricted to IHC 3+ disease would exclude most responders in this cohort. Prospective validation is required. ### Competing Interest Statement M.Ki. has received fees from Springer Press, Biermann Press, Celgene, AstraZeneca, Myriad Genetics, TEVA, Eli Lilly, GSK, Seagen, AllergoSan, NutriImmun, FomF, Jenapharm, Roche, BESINS, Bayer AG, Novartis, Pierre Fabre, Abbvie. M.Ki. has received honoraria for consulting and advisory board participation by Myriad Genetics, Bavarian KVB, DKMS Life, BLAEK, TEVA, Exeltis, Roche, BESINS, Bayer AG, Eurobio-Scientific, Magna Med Group. M.Ki. holds shares in AIM GmbH, in-manas GmbH, Therawis Diagnostic GmbH. N.P. has received honoraria from AstraZeneca, GSK, Johnson & Johnson, Illumina, Novartis, Menarini Stemline, MSD, Merck, Bristol Myers Squibb, PGDX/Labcorp, Eli Lilly, QuIP GmbH. K.I. has received honoraria from Amgen, GSK, MSD und HederaDx. J.L. received honoraria from the Forum for Continuing Medical Education (FomF), AstraZeneca GmbH (Germany), and Novartis Pharma GmbH (Germany). The authors have no additional financial or non-financial conflicts of interest to disclose. ### Clinical Protocols ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: Local ethics approval from TUM University Hospital was granted (No. 2025-362-S-CB), and the study was conducted in accordance with the Declaration of Helsinki. All patients provided written informed consent to participate in the Molecular Tumor Board and for the subsequent use of their clinical and genomic data for research purposes. I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes All data needed to evaluate the conclusions of this study are present in the article and its Supplementary Information. Individual patient-level clinical and raw sequencing data are not publicly available because they could compromise patient privacy and fall outside the consent and ethics terms governing this retrospective study (DRKS00031622; ethics No. 2025-362-S-CB), but de-identified data are available from the corresponding author on reasonable request for non-commercial academic research. Source data for the figures are provided with this paper. Bavarian State Ministry of Health, Care and Prevention, DGP-2024-01 (GO-TWIN)
BACKGROUNDThe Biobeat wrist monitor (BB-613W; Biobeat Technologies, Petah-Tikva, Israel) and the Biobeat chest monitor (BB-613P; Biobeat Technologies) are wearable solutions for continuous noninvasive blood pressure monitoring.OBJECTIVE(S)We aimed to investigate the blood pressure measurement performance of the Biobeat wrist monitor and chest monitor after external calibration.DESIGNA prospective method comparison study.SETTINGUniversity Medical Center Hamburg-Eppendorf, Hamburg, Germany.PATIENTSFifty high-risk patients recovering from noncardiac surgery in an advanced postanaesthesia care unit.MAIN OUTCOME MEASURESWe compared blood pressure measurements from the Biobeat wrist monitor (BPWRIST-ART) and the Biobeat chest monitor (BPCHEST-ART) with intra-arterial blood pressure measurements (BPART). In addition, we aimed to compare blood pressure measurements from the Biobeat wrist monitor (BPWRIST-OSCI) with those from an oscillometric upper-arm cuff (BPOSCI). We used Bland-Altman analysis, four-quadrant plot and error grid analysis for statistical analysis.RESULTSThe mean of the differences +/- standard deviation (95%-limits of agreement) between BPWRIST-ART and BPART was 3 +/- 11 mmHg (-19 to 25 mmHg) for mean blood pressure with a concordance rate to track 15-min blood pressure changes of 51%. The mean of the differences between BPCHEST-ART and BPART was 3 +/- 11 mmHg (-17 to 24 mmHg) for mean blood pressure with a concordance rate to track 15-min blood pressure changes of 61%. The mean of the differences between BPWRIST-OSCI and BPOSCI was 6 +/- 11 mmHg (-16 to 27 mmHg) for mean blood pressure with a concordance rate to track 15-min blood pressure changes of 49%.CONCLUSIONSBlood pressure measurements from the Biobeat wrist monitor and the Biobeat chest monitor did not show clinically acceptable agreement either with intra-arterial blood pressure measurements or with blood pressure measurements from an oscillometric upper-arm cuff in high-risk patients recovering from noncardiac surgery in an advanced postanaesthesia care unit.
BACKGROUND:Cardiac output monitoring is recommended for high-risk surgical patients and critically ill patients with circulatory shock. We performed a systematic review and meta-analysis of clinical studies published since 2010 that compared minimally invasive pulse wave analysis-derived cardiac output or cardiac index measurements with reference measurements by pulmonary artery thermodilution or transpulmonary thermodilution in adult surgical or critically ill patients. METHODS:In a random-effects meta-analysis, we calculated pooled estimates of the percentage error, mean difference and standard deviation, and 95% limits of agreement separately for studies reporting cardiac output or cardiac index. Subgroup analyses were performed by patient population and test device. RESULTS:We included 92 studies divided into 113 data sets with a total of 3111 patients. For 71 data sets reporting cardiac output, the pooled percentage error (95% confidence interval [95% CI]) was 44.0% (38.2%-49.8%) with a mean difference (standard deviation) of -0.1 (1.3) L min-1 with 95% limits of agreement of -2.6 to 2.4 L min-1 (I2=9.8%). For 42 data sets reporting cardiac index, the pooled percentage error was 49.1% (40.6%-57.6%) with a mean difference of -0.1 (0.9) L min-1 m-2 with 95% limits of agreement of -1.8 to 1.6 L min-1 m-2 (I2=7.9%). The percentage error varied substantially across patient populations and devices. Overall risk of bias was low. CONCLUSIONS:The pooled percentage error between minimally invasive pulse wave analysis-derived cardiac output measurements of 44.0% (cardiac output) and 49.1% (cardiac index) in adult surgical or critically ill patients exceeds the 30% threshold for clinically acceptable agreement. However, the percentage error varied depending on the patient population and device used. TRIAL REGISTRATION:PROSPERO (CRD420251090806; submitted July 17, 2025).
Meta-analysis of diagnostic test accuracy studies aggregates information from multiple studies on sensitivity and specificity. Classical approaches select a single pair of sensitivity and specificity per study (single threshold methods, STM), ignoring additional information if studies report results on multiple diagnostic thresholds. Recently, models have been proposed that consider all available information and enable inference on all diagnostic thresholds (multiple threshold methods, MTM). We compare five STM and six MTM to each other in a simulation study, evaluating their performance in various situations. Covering a broad range of real-life settings, we vary eight parameter dimensions in the data-generation mechanisms, including continuous or ordinal outcome type of an index test, and different numbers of diagnostic thresholds available per study. While model performances are comparable regarding bias, empirical coverage, and convergence, we observe a logit GLMM of the MTM type to perform best in many situations. Model performances depend strongest on the outcome type, while the number of thresholds only has a minor impact. We thus find the main advantage of using MTM by getting threshold-dependent estimates of sensitivity and specificity. Additionally, we illustrate differences between model estimates in two real-data examples on diagnosing type 2 diabetes using the continuous biomarker HbA1c and screening for any anxiety disorder using the ordinal questionnaire HADS-A. The applications reveal variations in model estimates within and between STM and MTM, which can be reduced by adjusting for the bias in the simulation settings resembling the real-data situation most closely.
Background International guidelines recommend risk-adapted depression screening in primary care. However, empirical evidence on the diagnostic accuracy of depression screening questionnaires in patients at risk for depression remains limited. Objective To evaluate the diagnostic accuracy of the Patient Health Questionnaire-9 (PHQ-9) for detecting major depressive disorder (MDD) in primary care, stratified by the presence of single depression-related risk factors and the amount of risk factors. Methods This secondary analysis used data from 985 primary care patients participating in the GET.FEEDBACK.GP trial who completed the PHQ-9, a depression-related risk factor assessment, and underwent evaluation for MDD using the Mini-International Neuropsychiatric Interview (MINI). Accounting for partial verification bias, this study applied an inverse probability weighting normalized for a sample of 985. Sensitivity, specificity, positive predictive value (PPV), negative predictive value (NPV), and the area under the curve (AUC) were calculated for different PHQ-9 cut-off scores across single risk factors and the amount of risk factors. The analysis was pre-registered (https://osf.io/wzctq). Results Of 985 participants, 89 (9.1%) had a MDD diagnosis. The best-performing PHQ-9 cut-offs, stratified by the amount of risk factors, varied, ranging from 7 to 13, with a higher number of risk factors being associated with a higher best-performing PHQ-9 cut-off score. Sensitivity ranged from 0.78 to 0.99; specificity from 0.80 to 0.91. PPV ranged from 0.40 to 0.88 and NPV from 0.80 to 0.99. AUCs ranged from 0.92 to 0.97, indicating excellent diagnostic accuracy. Similar results were found when stratifying by single risk factors. Conclusions The PHQ-9 demonstrated high diagnostic accuracy for detecting MDD in patients at risk for depression. Although optimal cut-offs vary slightly according to the number and type of risk factors present, the findings support the validity of risk-adapted depression screening using the PHQ-9 in primary care.
Background This article aims to provide an overview of the risk of bias, concerns regarding applicability, and the reporting of findings from studies examining the diagnostic test accuracy (DTA) of self-report questionnaires used to screen for anxiety disorders. Methods We merged data from 103 studies included in six systematic reviews on the diagnostic accuracy of eight anxiety questionnaires (HADS-A, GAD-7, GAD-2, BAI, STAI-S, STAI-T, PROMIS-A-SF8a, and OASIS) in adults. To be included, primary studies had to report the sensitivity and specificity for at least one cut-off when tested against a (semi-)structured clinical interview used as a reference standard. Risk of bias and applicability were rated with the QUADAS-2 tool (Quality Assessment of Diagnostic Accuracy Studies). Results The 103 studies involved a total of 34,833 participants (range 48 to 3215). The most commonly used questionnaires were the HADS-A (54 studies), the GAD-7 (46), and the GAD-2 (31). Overall, 80% of the studies were rated as having unclear or high risk of bias, and for 88%, concerns regarding applicability were rated as unclear or high. Uncertainties or concerns were particularly common regarding participant recruitment and whether the samples included reflected realistic screening populations. Publications often reported sensitivity and specificity in a very selective manner for a single or a few cut-offs only. Conclusions Most DTA studies on questionnaires used to screen for anxiety have significant shortcomings in the reporting of methods and results. Our findings provide a basis for recommendations for future studies.
BACKGROUND:The Biobeat wrist monitor (BB-613W; Biobeat Technologies, Petah-Tikva, Israel) and the Biobeat chest monitor (BB-613P; Biobeat Technologies) are wearable solutions for continuous noninvasive blood pressure monitoring. OBJECTIVES:We aimed to investigate the blood pressure measurement performance of the Biobeat wrist monitor and chest monitor after external calibration. DESIGN:A prospective method comparison study. SETTING:University Medical Center Hamburg-Eppendorf, Hamburg, Germany. PATIENTS:Fifty high-risk patients recovering from noncardiac surgery in an advanced postanaesthesia care unit. MAIN OUTCOME MEASURES:We compared blood pressure measurements from the Biobeat wrist monitor (BP WRIST-ART ) and the Biobeat chest monitor (BP CHEST-ART ) with intra-arterial blood pressure measurements (BP ART ). In addition, we aimed to compare blood pressure measurements from the Biobeat wrist monitor (BP WRIST-OSCI ) with those from an oscillometric upper-arm cuff (BP OSCI ). We used Bland-Altman analysis, four-quadrant plot and error grid analysis for statistical analysis. RESULTS:The mean of the differences ± standard deviation (95%-limits of agreement) between BP WRIST-ART and BP ART was 3 ± 11 mmHg (-19 to 25 mmHg) for mean blood pressure with a concordance rate to track 15-min blood pressure changes of 51%. The mean of the differences between BP CHEST-ART and BP ART was 3 ± 11 mmHg (-17 to 24 mmHg) for mean blood pressure with a concordance rate to track 15-min blood pressure changes of 61%. The mean of the differences between BP WRIST-OSCI and BP OSCI was 6 ± 11 mmHg (-16 to 27 mmHg) for mean blood pressure with a concordance rate to track 15-min blood pressure changes of 49%. CONCLUSIONS:Blood pressure measurements from the Biobeat wrist monitor and the Biobeat chest monitor did not show clinically acceptable agreement either with intra-arterial blood pressure measurements or with blood pressure measurements from an oscillometric upper-arm cuff in high-risk patients recovering from noncardiac surgery in an advanced postanaesthesia care unit.
BACKGROUND:Anxiety disorders often remain undetected and can cause substantial burden. Amongst the many anxiety screening tools, the 7-item Generalized Anxiety Disorder (GAD-7) scale and its short version, the 2-item Generalized Anxiety Disorder (GAD-2) scale, are the most frequently used instruments. OBJECTIVES:Primary: to determine the diagnostic accuracy of GAD-7 and GAD-2 to detect generalised anxiety disorder (GAD) and any anxiety disorder (AAD) in adults. Secondary: to investigate whether their diagnostic accuracy varies by setting, anxiety disorder prevalence, reference standard, and risk of bias; to compare the diagnostic accuracy of GAD-7 and GAD-2; to investigate how diagnostic performance changes with the test threshold. SEARCH METHODS:We searched MEDLINE, Embase, PubMed-not-MEDLINE subset, and PsycINFO from 1990 to 18 January 2024. We checked reference lists of included studies and review articles. SELECTION CRITERIA:We included cross-sectional studies conducted in adults, containing diagnostic accuracy information on GAD-7 and/or GAD-2 questionnaires for the target conditions generalised anxiety disorder and/or any anxiety disorder, and allowing the generation of 2x2 tables. The target conditions must have been diagnosed using a structured or semi-structured clinical interview. We excluded case-control studies and studies in which the time elapsed between the index tests and reference standards exceeded four weeks. We excluded studies involving people (1) seeking help in mental health settings or (2) recruited specifically due to mental health symptoms in other settings. DATA COLLECTION AND ANALYSIS:At least two review authors independently decided on study eligibility, extracted data, and assessed the risk of bias and applicability of included studies. For each questionnaire and each target condition, we present sensitivity and specificity with 95% confidence intervals (95% CI) in forest plots. We used the bivariate model to obtain summary estimates based on cut-offs closest to the recommended values (i.e. within a core range). In secondary analyses, we used the bivariate model and the multiple thresholds model to obtain summary estimates for all available cut-off points. Using the multiple thresholds model, we also calculated the area under the receiver operating characteristic curve to obtain a general indicator of the diagnostic accuracy of GAD-7 and GAD-2. MAIN RESULTS:We included 48 studies with 19,228 participants from 27 different countries, evaluating the GAD-7 and the GAD-2 in 24 different languages. Seven studies were performed in non-clinical settings, nine in clinical settings recruiting participants across conditions, and 32 in clinical settings with participants having specific conditions. Even after categorisation into three settings, the study populations were substantially different. The most frequently studied populations were people: with epilepsy (nine studies); with cancer (five studies); with cardiovascular disease (five studies); and in primary care regardless of their condition (five studies). We considered the risk of bias low in eight studies, and we had low concerns about the applicability of findings in three studies. Thirty-five studies contributed to the primary analyses of GAD-7 for detecting generalised anxiety disorder (median prevalence 12%); 22 studies to analyses of GAD-7 for any anxiety disorder (median prevalence 19%); 24 studies to analyses of GAD-2 for generalised anxiety disorder (median prevalence 9%); and 19 studies to analyses of GAD-2 for any anxiety disorder (median prevalence 19%). At the recommended cut-off of 10 or higher (or the closest available cut-off), the GAD-7 questionnaire yielded a summary sensitivity of 0.64 (95% CI 0.56 to 0.72) and a summary specificity of 0.91 (95% CI 0.87 to 0.93) in detecting generalised anxiety disorder. For detecting any anxiety disorder, summary sensitivity was 0.48 (95% CI 0.40 to 0.57) and summary specificity 0.91 (95% CI 0.89 to 0.93). At the recommended cut-off of 3 or higher (or the closest available cut-off), the GAD-2 yielded a summary sensitivity of 0.68 (95% CI 0.59 to 0.75) and a summary specificity of 0.86 (95% CI 0.82 to 0.89) for detecting generalised anxiety disorder. For detecting any anxiety disorder, the summary sensitivity was 0.53 (95% CI 0.44 to 0.62) and the summary specificity was 0.89 (95% CI 0.86 to 0.91). The 95% prediction region of GAD-7 for detecting generalised anxiety disorder was larger (indicating pronounced statistical heterogeneity) than for the three other analyses. Specificity varied by setting in the analysis of GAD-7 and GAD-2 for detecting any anxiety disorder, and by reference standard in the analysis of GAD-2 for detecting generalised anxiety disorder. Sensitivity varied with prevalence in the analysis of GAD-7 for generalised anxiety disorder. Other investigations of potential sources of heterogeneity did not show statistically significant associations with test accuracy. In all analyses, sensitivity tended to be higher and specificity lower in participants with specific conditions compared to the other two settings. Overall, the heterogeneity in the subgroup analyses remained high. The area under the receiver operating characteristic curve in the multiple thresholds model was 0.86 (95% CI 0.84 to 0.88) for the GAD-7 scale in detecting generalised anxiety disorder, and 0.80 (95% CI 0.78 to 0.82) in detecting any anxiety disorders. For the GAD-2 scale, the value was 0.82 (95% CI 0.81 to 0.86) for detecting generalised anxiety disorder, and 0.77 (95% CI 0.76 to 0.82) for detecting any anxiety disorders. Comparative bivariate analyses revealed no statistically significant differences between the diagnostic test accuracy of GAD-7 and GAD-2. AUTHORS' CONCLUSIONS:The GAD-7 and the GAD-2 scales have been tested in numerous languages and different populations. Overall, the GAD-7 and the GAD-2 seem to have acceptable or good diagnostic accuracy for both generalised anxiety disorder and any anxiety disorder. The GAD-2 scale seems to have similar diagnostic accuracy as the GAD-7 scale. However, due to the diversity of the included studies and the heterogeneity of our findings, our summary estimates of sensitivity and specificity should be interpreted as rough averages. The performance of GAD-7 and GAD-2 may deviate substantially from these values in specific situations.
Prognostic outcome models might help to minimise the discard rates of renal transplants from deceased donors. To this end, we published 2-Step Scores for both delayed graft function and transplant loss with optional histology. With conventional paraffin PAS histology taking at least 3 h excluding transport time, we tested whether fast, mobile confocal histology with portable VivaScope® 2500 systems might offer a viable and more rapid alternative. After omitting 17 biopsies with less than 12 glomeruli and 1 artery, we collected 14 0-h and 16 renal transplant indication (Tx) biopsies for a combined cohort. All biopsies were scanned in less than 10 min with a VivaScope® 2500, rendering pseudo-HE images, and then underwent our regular paraffin work-up. Banff Lesion Scores ct and cv were assessed on the granular ordinal scale and binary as (ct ≤ 1 vs. ct ≥ 2 and cv ≤ 2 vs. cv3) together with the number of glomeruli as used in the previously published 2-Step Scores in a blinded fashion by an expert nephropathologist on the paraffin sections (P) and VivaScope® 2500 scans (V). Additionally, we examined the ratio of globally sclerotic glomeruli. Correlation and mixed effects linear regressions, as well as Fleiss’ kappa statistics, comparing P and V on the combined cohort were applied, supplemented with Bland-Altman statistics providing limits of agreement. Between P and V, granular Banff ct correlated with a kappa of 0.513 (p = 8.38e−05) and a Kendall’s W of 0.706 (p = 0.0694) on the combined cohort; binary Banff ct correlated with a kappa of 0.869 (p = 1.92e−06). We had to exclude two more biopsies in which no artery was found in the scanning plane in V. Granular Banff cv correlated with a kappa of 0.109 (p = 0.345) and a W of 0.677 (p = 0.103) between P and V on the combined cohort and binary Banff cv with a kappa of 0.24 (p = 0.204). The total number of glomeruli correlated between P and V with an R of 0.75 (p = 2.3e−06), the number of globally sclerotic glomeruli with an R of 0.82 (p = 4.1e−08), and the ratio thereof with an R of 0.86 (p = 8.6e−10). Instant, decentralised VivaScope® 2500 histology might deliver Banff ct and the number of glomeruli with sufficient accuracy for the 2-Step Scores to predict the risk of delayed graft function and 1-year death-censored transplant loss in deceased heart-beating donors. In contrast, VivaScope® 2500 assessment of Banff cv might not be accurate enough for use in 2-Step Scores.
ABSTRACT Background While automated methods for differential diagnosis of parkinsonian syndromes based on MRI imaging have been introduced, their implementation in clinical practice still underlies considerable challenges. Objective To assess whether the performance of classifiers based on imaging derived biomarkers is improved with the addition of basic clinical information and to provide a practical solution to address the insecurity of classification results due to the uncertain clinical diagnosis they are based on. Methods Retro‐ and prospectively collected data from multimodal MRI and standardized clinical datasets of 229 patients with PD (n = 167), PSP (n = 44), or MSA (n = 18) underwent multinomial classification in a benchmark study comparing the performance of nine machine learning methods. A predictor space of imaging variables, either with or without clinical information, was investigated. Classification results were assessed using multiclass AUCs. Individual predicted probabilities were visualized to address diagnostic uncertainty. Results Clinical diagnosis was accurately confirmed using machine learning models with only small differences when using imaging and clinical signs versus imaging variables only (expected multiclass AUC of 0.95 vs. 0.92). Still, multinomial classification is hampered by imbalanced class frequencies. The most discriminatory variables were responsiveness to levodopa, vertical gaze palsy, and the volumes of subcortical structures, including the red nucleus. Conclusion Machine‐learning‐assisted classification of MR‐imaging biomarkers gathered in routine care can assist in the diagnosis of parkinsonian syndromes as part of the diagnostic workup. We provide a visual method that aids the interpretation of neuroimaging‐based classification results of the three main parkinsonian syndromes, improving clinical interpretability.
Background:The course of relapsing-remitting multiple sclerosis (RRMS), frequently preceded by the clinically isolated syndrome (CIS), is variable and challenging to predict. Given many treatment options available, prognostic algorithms are gaining importance in informing initial treatment decisions. However, to date, only a few externally validated exists. External validation, which involves the application of a model to independent data, is essential. Privacy-preserving federated analyses of individual-level data facilitate external validation using clinical datasets that are typically difficult to access. Objectives:Using data from the ProVal-MS study to externally validate the multiple sclerosis treatment decision score (MS-TDS), a predictive algorithm for early RRMS and CIS. The MS-TDS predicts the probability of the occurrence of at least one new or enlarging T2 lesion within 6-24 months following the onset of the disease and supports choosing between initiating platform treatment or a 'wait-and-see' approach. A secondary objective is to demonstrate the feasibility of privacy-preserving federated concepts within the Data Integration for Future Medicine (DIFUTURE) consortium. Design:Prospective, multicentric, non-interventional cohort study (ProVal-MS) within DIFUTURE. Methods:The calibrated MS-TDS was evaluated using the area under the receiver operating characteristic curve (AUROC) and the Brier score in both pooled and distributed settings. A decision curve analysis (DCA) was used to evaluate the net benefit of treatment decisions made by the MS-TDS in comparison to those made by treating neurologists. Results:Of the 271 individuals diagnosed with CIS or early RRMS, 202 (78.2%) received platform treatment, while 59 (21.8%) did not receive treatment. The AUROC was 0.561 (95% CI: 0.492-0.630) in the pooled analysis and 0.567 (95% CI: 0.496-0.634) in the distributed analysis. DCA demonstrated a net benefit that was commensurate with that achieved by decisions made by experienced neurologists. Conclusion:The external validation of the MS-TDS demonstrated low, non-significant predictive performance; however, it may serve as a useful complement, particularly for less-experienced neurologists. The distributed validation was found to be both feasible and compliant with data protection regulations.
RationaleDetection of atrial fibrillation (AFib) and subsequent anticoagulation therapy reduce the risk of recurrent stroke, while prolonged rhythm monitoring significantly increases AFib detection. Thus, prolonged smartwatch-based ECG monitoring after cryptogenic ischemic stroke or transient ischemic attack (TIA) could lead to a reduction of recurrent stroke by prompting adequate anticoagulation therapy.AimWATCH AFib investigates the accuracy of smartwatches for AFib detection in patients with cryptogenic TIA or ischemic stroke compared to an implantable event recorder.Sample size40 cases of AFib are required to estimate the sensitivity for AFib detection per patient with a precision of about 10%. As AFib is observed in 9%−16% of cryptogenic strokes, we intend to enroll 400 patients.MethodsWATCH AFib is a prospective, intraindividual-controlled, multicentre clinical study in patients with cryptogenic ischemic stroke or TIA. ECG-data from smartwatches and event recorders is continuously monitored by two independent cardiologists for a follow-up period of 6 months. If AFib is detected, therapeutic options are discussed at the including center.Primary outcomeTo compare smartwatch- and event recorder- based sensitivity and specificity of AFib detection per patient after 6 months.DiscussionProlonged AFib screening after stroke is currently suboptimal. Smartwatches might be a non-invasive, cost-effective, widely available alternative for prolonged rhythm monitoring. Usability in severely affected patients and patients with persisting neurological deficits might be limited.Trial registrationThe study is registered on clinicaltrials.gov. Registration number: 20230726.
Cardiac output is a key cardiovascular variable quantifying global blood flow. The measurement performance of cardiac output monitoring methods is investigated in validation studies, which are method comparison studies determining the agreement between cardiac output values measured with a test method and those measured with a reference method. The StatistiCal analysis and repOrting of cardiac output Method comPARison studiEs (COMPARE) statement provides a framework for designing, performing, and reporting cardiac output method comparison studies and includes a checklist of 29 items that are essential for reporting of those studies. Considering and reporting the items specified in the COMPARE checklist will help standardize cardiac output method comparison studies and increase the external validity of the results.
Bei vielen Frauen mit hormonsensitivem Brustkrebs lässt sich das Risiko eines Rezidivs durch eine 5‑ bis 10-jährige endokrine Therapie halbieren. Allerdings liegen die Abbruchraten einer adjuvanten endokrinen Therapie bei 31–72
INTRODUCTION:Uncomplicated urinary tract infections (uUTIs) in women are common infections encountered in primary care. Evidence suggests that rapid point-of-care tests (POCTs) to detect bacteria and erythrocytes in urine at presentation may help primary care clinicians to identify women with uUTIs in whom antibiotics can be withheld without influencing clinical outcomes. This pilot study aims to provide preliminary evidence on whether a POCT informed management of uUTI in women can safely reduce antibiotic use.METHODS AND ANALYSIS:This is an open-label two-arm parallel cluster-randomised controlled pilot trial. 20 general practices affiliated with the Bavarian Practice-Based Research Network (BayFoNet) in Germany were randomly assigned to deliver patient management based on POCTs or to provide usual care. POCTs consist of phase-contrast microscopy to detect bacteria and urinary dipsticks to detect erythrocytes in urine samples. In both arms, urine samples will be obtained at presentation for POCTs (intervention arm only) and microbiological analysis. Women will be followed-up for 28 days from enrolment using self-reported symptom diaries, telephone follow-up and a review of the electronic medical record. Primary outcomes are feasibility of patient enrolment and retention rates per site, which will be summarised by means and SDs, with corresponding confidence and prediction intervals. Secondary outcomes include antibiotic use for UTI at day 28, time to symptom resolution, symptom burden, number of recurrent and upper UTIs and re-consultations and diagnostic accuracy of POCTs versus urine culture as the reference standard. These outcomes will be explored at cluster-levels and individual-levels using descriptive statistics, two-sample hypothesis tests and mixed effects models or generalised estimation equations.ETHICS AND DISSEMINATION:The University of Würzburg institutional review board approved MicUTI on 16 December 2022 (protocol n. 109/22-sc). Study findings will be disseminated through peer-reviewed publications, conferences, reports addressed to clinicians and the local citizen's forums.TRIAL REGISTRATION NUMBER:ClinicalTrials.gov NCT05667207.