OBJECTIVES:Large effect sizes (ESs), especially when prominently presented in trial abstracts, draw large attention, but it is important to understand whether they are trustworthy. We aimed to assess indicators of transparency and trustworthiness in randomized controlled trials (RCTs) reporting some large ES for continuous outcomes in their abstract, in comparison with RCTs presenting only nonlarge ESs in their abstract. STUDY DESIGN AND SETTING:We included RCTs indexed in MEDLINE between January 1, 2024 and March 18, 2025, presenting at least one standardized mean differences of absolute value 0.8 or higher (large ES) vs those presenting only smaller absolute standardized mean differences in their abstract. Trial characteristics and methodological features were extracted systematically in large ES and nonlarge ES trials. Primary outcome was prespecified protocol registration, secondary outcomes were any protocol registration and public availability or repository placement of raw data (preregistered protocol: https://osf.io/8xasw). RESULTS:We evaluated 152 trials with large ESs in their abstract and 175 trials with only nonlarge ESs in their abstract. Large ES trials had suggestively lower rates of preregistered protocols (45% vs 61%, P = .0054) and significantly lower rates of any protocol registration (74% vs 87%, P = .0028) than nonlarge ES trials. There was no difference in raw data public availability or repository placement (6% vs 7%). Large ES trials were also less likely to be multicenter (P = .0042), to have high-income country of corresponding author (P = .0001), to be conducted in high-income country site(s) (P = .0003), to have a published statistical analysis plan (P = .0216), to result from between-group comparisons (P < .0001), and tended to report less frequently allocation concealment (P = .0351). Large effects were significantly more likely to involve nondrug/nonpsychological interventions (P = .0001). CONCLUSION:RCTs presenting large ESs in their abstracts are more likely to lack transparency and trustworthiness features and may operate with higher risk of lack of credibility. PLAIN LANGUAGE SUMMARY:When a medical study reports that a treatment has a big effect that naturally grabs attention. However, are impressive-sounding results actually reliable? This study investigated whether clinical trials claiming large treatment effects in their abstracts are as trustworthy as those reporting only more modest results. We compared two groups of recently published randomized controlled trials: 152 trials that reported large effects in their abstracts, and 175 trials that reported only smaller effects. We looked at several markers of good scientific practice, particularly whether researchers had registered their study plans in advance and whether they made their data openly available. We found that trials reporting large effects were less likely to have preregistered their methods and more likely to have no registered protocol at all. The two groups also differed in other aspects of trustworthiness. Large-effect trials were less likely to involve multiple research centers, less likely to have researchers from high-income countries, and less likely to have a published statistical analysis plan. They also tended to test different types of treatments-large effects appeared more often in studies of interventions other than drugs or psychological therapies. Both groups were equally poor at making their raw data publicly available (only about 6%-7%). Overall, when a trial claims to have found impressive treatment effects, some extra scrutiny may be warranted. Before trusting eye-catching results, readers should assess whether trials followed proper scientific checks and transparency safeguards.
Historically, glycomics has lagged behind other omics fields, owing to analytical challenges. However, recent technological advances are rapidly transforming the field, enabling a growing number of large-scale studies to harness the wealth of information encoded by glycans. In this Review, we provide an overview of the current state of large-scale glycomics and its role in understanding the functional importance of glycans. We discuss the main challenges in high-throughput glycoanalytical methods and highlight recent progress, with a particular emphasis on standardization and reproducibility. Special attention is given to the necessity of large, well-designed and ideally longitudinal and multicohort studies that use rigorous statistical approaches to uncover robust biological associations. Finally, we conclude with selected applications and perspectives, emphasizing the potential of large-scale glycomics to drive biomarker discovery and enhance the understanding of glycan-mediated biology.
OBJECTIVE:To estimate the minimal detectable change (MDC) for the Patient Health Questionnaire-9 (PHQ-9) and its eight item (PHQ-8) and two item (PHQ-2) versions including differences by participant and study characteristics. DESIGN:Individual participant data meta-analysis. DATA SOURCES:Medline, Medline In-Process and other non-indexed citations, PsycInfo, and Web of Science, 1 January 2000 to 9 May 2018. ELIGIBILITY CRITERIA FOR SELECTING STUDIES:Datasets from articles in any language if participants were aged ≥18 years, were recruited from any non-psychiatric setting, and were not recruited because they were seeking mental healthcare. Eligible datasets had a classification for major depressive disorder or major depressive episode based on a validated semi-structured or fully structured interview conducted within two weeks of administering the PHQ-9, PHQ-8, or PHQ-2. RESULTS:Pooled MDCs across studies were estimated for the PHQ-9, PHQ-8, and PHQ-2 with random effects meta-analysis for 95% (MDC95), 90% (MDC90), and 67% (MDC67) confidence that change beyond measurement error occurred. PHQ-9, PHQ-8, and PHQ-2 analyses included 42 548 participants (94 studies), 42 592 participants (94 studies), and 44 085 participants (98 studies), respectively. Mean participant age was 49 years (standard deviation 17), and 60% of participants were women. Overall, 10% of participants had major depression (range 1-57% across studies). MDC95 was 5.72 points (95% confidence interval (CI) 5.54 to 5.90, 95% prediction interval (PI) 4.00 to 7.44) for the PHQ-9, 5.51 points (95% CI 5.33 to 5.68, 95% PI 3.87 to 7.15) for the PHQ-8, and 2.26 points (95% CI 2.15 to 2.37, 95% PI 1.20 to 3.32) for the PHQ-2. For the PHQ-9, MDC95 was highest in inpatient healthcare settings at 6.48 (95% CI 6.05 to 6.92) points. MDC95 for the PHQ-9 increased by 0.40 (95% CI 0.25 to 0.55) points for each 10% increase in the proportion of participants with major depression. Sex and age had minimal or no association. Subgroup and meta-regression findings were similar for the PHQ-8 and PHQ-2. CONCLUSIONS:Based on the pooled estimate, a six point difference on the PHQ-9, the PHQ version most used in clinical practice, could be an appropriate MDC threshold in general practice. A higher threshold may be preferred in specialty mental healthcare. MDC67 or MDC90 thresholds would provide less certainty that change has occurred. Alternative strategies, such as using the upper end of a prediction interval, would provide more certainty but a greater likelihood of not recognising change. STUDY REGISTRATION:PROSPERO CRD42014010673.
PurposeThe purpose of this study is to develop a genetic test to aid in diagnosing chronic kidney disease (CKD). One challenge in treating CKD is that 80%–90% of people with it are undiagnosed and thus do not access healthcare promptly. The problem arises because early-stage CKD has no overt symptoms, and the current policy is to perform diagnostic tests only when accompanied by risk factors such as old age, hypertension, and diabetes.MethodsThis study describes the development of the RICK (RIsk for Chronic Kidney disease) algorithm that employs a polygenic risk score for CKD plus clinical risk factors to identify people at risk.ResultsIn data from the United Kingdom biobank, those in the top decile of RICK have a ten-fold increased risk of CKD, and approximately 49% of all those with CKD are included in this decile. Furthermore, targeted creatinine testing for those in the highest RICK decile would potentially increase the number of individuals diagnosed with CKD by 7.4%. However, the RICK algorithm adds little value for detecting CKD defined by elevated uACR (albuminuria).ConclusionUsing RICK to selectively test those in the general population with highest risk may help in the early identification of CKD and facilitate early access to renal healthcare. The effectiveness and cost-effectiveness of such testing require further study.
OBJECTIVES:To build on existing evidence regarding single-item measurement instruments of patient-reported bother or trouble from medical side effects in individuals with rheumatic and musculoskeletal diseases (RMDs). Further, to collect input from the OMERACT community through a structured survey that rated and ranked available options and to seek agreement to advance one or more of these measures for use as exploratory outcomes in future clinical trials. METHODS:At OMERACT 2025 we presented and discussed survey results for domain match, feasibility and ranking of six candidate instruments of bother or trouble from side effects. Collaborator feedback - including comments from patients, clinicians, and researchers - was synthesized with a large-language-model (LLM) to identify key concerns and guide refinement of the instrument's relevance, clarity, and acceptability. The LLM-assisted synthesis of participant comments resulted in a new, single-item instrument designed to improve patient safety reporting from the patient's perspective. RESULTS:The merged and modified version of the instrument was presented at the OMERACT 2025 meeting, where 33 participants approved it as a reasonable approach to incorporate collaborator input. The proposed instrument is feasible (32 [97%]) and voting supported advancing its further assessment (30 [91%]) as an exploratory outcome measurement instrument in coming RMD trials. CONCLUSION:We developed a novel single-item instrument. This is the first known application of LLMs in refining a patient-reported outcome instrument for clinical trials. It is designed to capture the patient perspective on symptomatic treatment-related side effects in RMDs and is supported for exploratory use in trials.
Objectives To assess the construct validity of a modified single-item measure of bother due to side effects (the GP5 item) from the Functional Assessment of Chronic Illness Therapy (FACIT) system by comparing it to current symptomatic side effects from the Patient-Reported Outcomes of the Common Terminology Criteria for Adverse Events (PRO-CTCAE) reported by patients with rheumatoid arthritis (RA). Methods Through a cross-sectional, web-based survey we collected information on the frequency of symptomatic side effects and bother from side effects related to RA medications. We applied multiple correspondence analysis (MCA) to reduce 80 symptomatic side effects into key dimensions (≥5% of the total variance each). We then examined associations among key dimensions, individual items, the sum of current side effects, and the single-item bother measure using Spearman rho. Results A total of 560 patients participated in the online survey. Our scree plot showed a clear elbow point after the first dimension, indicating that keeping just one dimension captured the most meaningful information. This overall side effect burden score appeared to reflect a broad concept influenced by a variety of symptomatic side effects, each having only a negligible to weak impact. Conclusions Our results may indicate that individuals have diverse experiences of side effects, allowing the global index to capture these variations, even when they differ across patients. Thus, a single-item burden measure to side effects can potentially serve as a useful summary indicator, shedding light on the impact of symptomatic side effects experienced by RA patients.
BACKGROUND:Latent factor scoring may provide more precise score estimates than sum scores, but this has not been evaluated for the Hospital Anxiety and Depression Scale (HADS). We investigated whether latent factor scores could improve HADS depression screening accuracy. METHODS:We used a HADS screening accuracy individual participant data meta-analysis (IPDMA) database. We included 42 studies (7982 participants; 12 to 1143 per study) with a semi-structured interview reference standard. We randomly split the database into calibration and validation datasets. In calibration, we estimated latent scores using one-factor models (14-item HADS total scale [HADS-T], 7-item depression subscale [HADS-D]) plus HADS-T two-factor and bi-factor (general factor and two specific factors) models. We estimated cut-offs that maximized combined sensitivity and specificity for each method. In validation, we compared screening accuracy between latent variable approaches and the HADS-D sum score. The process was repeated 1000 times to estimate 95% confidence intervals for parameters. RESULTS:After removing iterations with failed models in confirmatory factor analysis (N = 304) or IPDMA (N = 31), aggregated results showed that confidence intervals for sensitivity, specificity, and combined sensitivity and specificity included 0 for all comparisons between factor scores and sum scores. Statistically significant but minimal advantages appeared in the receiver operating characteristic curve for the two-factor and bi-factor models (0.01, 95% CI [0.01, 0.02]; 0.02, 95% CI [0.01, 0.02]). Sensitivity analysis confirmed findings. CONCLUSIONS:Latent factor scoring did not meaningfully improve HADS screening accuracy compared with sum scores. Sum scores may be preferred in applied settings for their simplicity and feasibility.
Importance:Multiple retractions from the same author often uncover issues affecting their entire work, such as having systematically altered or fabricated data. Objectives:To evaluate the contribution of authors with the most retractions (ie, superretractors) and top-cited scientists with multiple retractions to the retracted randomized clinical trial (RCT) literature. Design, Setting, and Participants:This retrospective cohort study linked an openly available cohort of retracted RCTs (VITALITY) to 3 lists of scientists: (1) superretractors, totaling most retractions in the Retraction Watch Leaderboard; (2) scientists in the top 100 000 or 2% of their subfield in terms of citations (ie, top-cited scientists) over their entire careers who accumulated 10 or more retractions not due to editor or publisher errors; and (3) top-cited scientists in the most recent year (ie, 2024) who accumulated 10 or more retractions not due to editor or publisher errors. The VITALITY cohort was updated up to November 2024. The 3 author lists were updated in August 2025. Main Outcomes and Measures:The main outcomes were authorship and the characteristics of retracted RCTs (publication and retraction year, time between publication and retraction, number of citations). Results:A total of 30 superretractors, 163 career-long top-cited scientists with 10 or more retractions, and 174 recent-year top-cited scientists with 10 or more retractions were included; 1330 retracted RCTs were included. Overall, 6 superretractors (20%), representing anesthesiology as well as endocrinology and metabolism, coauthored 290 retracted RCTs (22%); 18 career-long top-cited scientists with at least 10 retractions, representing 10 fields, coauthored 327 trials (25%), 275 (84%) of which were also coauthored by a superretractor; 7 single-year top-cited scientists with at least 10 retractions coauthored 50 retracted RCTs (4%), all of which were also included in the list of articles authored by career-long top-cited scientists with at least 10 retractions. Articles with superretractor authors vs not were published earlier (median [IQR], 2000 [1997-2005] vs 2020 [2014-2022]); retracted earlier (median [IQR], 2013 [2012-2019] vs 2023 [2018.5-2023]); had a longer lag between publication and retraction (median [IQR], 5111 [3560-6820] days vs 482 [330-1119] days); and accrued more citations (median [IQR], 21 [12-42] vs 5 [1-19]). In multivariable regression models, only time to retraction (β = 0.02; P < .001) was significantly and positively associated with total citations. Results were similar when comparing retracted articles from top-cited scientists with at least 10 retractions vs other articles. Conclusions and Relevance:In this cohort study of 1330 retracted RCTs, a small number of influential authors, often coauthors and concentrated across few fields of medicine, accounted for a significant proportion of retracted clinical trials.
Retractions attract substantial attention and have become more frequent over time. Retractions reflect the self-correcting nature of science but also wasted resources. The Retraction Watch Database (RWD) includes over 60,000 records. We integrated RWD with NIH funding metadata (RePORTER system). As of July 2026, of the 6,081 U.S. affiliated retracted articles, 1,725 (28.4%) were linked to at least one NIH grant. With a mean attributed cost per retracted NIH-funded article of $255,087 in 2026 dollars, the attributed total cost of NIH-funded retracted research is $440 million in 2026 dollars. Grants associated with retracted papers for which the first or last author of the paper was the principal investigator were awarded $4.03 billion in 2026 dollars. NIH-funded articles take longer to be retracted (mean = 6.3 years) than other US-based articles, which may entail greater downstream implications. NIH funding of retracted authors decreased over the 3 years following retraction, particularly among authors with multiple retractions. We have developed a continuously updated dynamic dashboard (https://sandovallentisco.shinyapps.io/nih-retractions/) for the cost of retractions reflecting NIH-funded work.
fMRI research is highly prolific but raises multiple concerns. Many competing statistical methods and respective packages are available using different assumptions, none of which applies equally well to all settings. However, the most fundamental concerns are not about the statistical machinery, but about issues of reproducibility, utility, and even construct validity. One can probe how much the field would benefit by statistical refinements, the conduct of larger studies and/or improved reproducibility practices. Alternatively, maybe fMRI research should largely be abandoned with focus shifting toward developing imaging methods with construct validity for granular neuronal activity and higher potential for clinical utility.
Socioeconomic, demographic, and health system structures may have shaped COVID-19 pandemic impact across populations, but past analyses typically examined few factors. We systematically examined correlates of COVID-era excess mortality, considering 2,745 county-level variables of demography, race/ethnicity, income, insurance, education, employment, housing, and health system. Pearson correlation coefficients (CCs) were obtained for the most recent available pre-pandemic value against age-standardized county excess-death for each year during 2020-2024. Counties were population-weighted. Variables were grouped by meaning into 11 semantic super-clusters. Overall, 17.3% of variables reached at least a moderate correlation level (|CC| > 0.30) and 2.8% reached strong correlations (|CC| > 0.45). Strongest correlations were seen for college attainment (CC -0.54), uninsurance among adults 40-64 (+0.53), and high income (-0.53). At least moderate correlations were seen for 9.1% of variables in 2020 and 8.5% in 2021, but only 1.8%, 0%, and 1.3% in 2022, 2023, and 2024, respectively. Similar patterns of concentration of moderate correlations in the first two pandemic years appeared in both elderly and non-elderly populations. Of 472 variables with |CC| > 0.30, 362/395 moderate-band and 77/77 strong-band variables belonged to demography and socioeconomic super-clusters. Only 7% of health system variables reached |CC| > 0.30, versus 31% of socioeconomic and demographic variables. Using the most recent available value until 2023 or 2015, different population weighting, and Spearman correlations yielded similar results. Overall, these ecological analyses suggest strong relationships of socioeconomic structure and demographics rather than health-care resources/supply with excess mortality across US counties especially during 2020-2021. SIGNIFICANCE STATEMENT:COVID-19 mortality has been linked to poverty, race, and care access, and other diverse socioeconomic and population factors, but typically only a few factors have been assessed and reported each time. Screening the entire Area Health Resources File - 2,745 unselected county-level variables - we found excess death was ecologically associated with many variables (one out of six had absolute correlations > 0.30). Correlations reflected more strongly variables pertaining to demographics and socioeconomic structure rather than baseline hospital capacity or physician supply. The substantive correlation signals were seen almost exclusively in 2020 and 2021, but not in subsequent years. The patterns were robust in sensitivity analyses considering different years of measurement of the county-level variables, different population weighting, and different correlation metrics.
Abstract Objectives To quantify the frequency of baseline control-group use in published long COVID prevalence studies and assess their key methodological features. Methods We performed a meta-epidemiological assessment of 440 post-acute COVID-19 prevalence publications from an existing systematic review. To evaluate study design and methodological transparency, we extracted data on the inclusion and classification of comparator groups, the exclusive use of self-reported outcome measures, and whether uncontrolled investigations explicitly recognized the omission of a control group as a limitation. In addition, we surveyed by email the corresponding authors of these articles to determine if any supplementary comparative data existed. The protocol was prospectively registered (DOI: 10.17605/OSF.IO/T2UP9). Results Among 440 studies, 372 (84.5%) reported no control group. Healthy or uninfected comparators were reported in 55 studies (12.5%) and other comparator types in 14 (3.2%); 1 study included both categories. Solely self-reported outcomes were used in 279 studies (63.4%). Among 372 uncontrolled studies, 244 (65.6%) did not explicitly acknowledge the absence of a baseline comparator as a limitation. Corresponding authors of 140 studies (31.8%) responded to the survey; 126 (90.0%) reported no additional comparative data, while 14 (10.0%) mentioned some available comparative datasets (19 additional datasets). Almost all that information (10/14, 17/19) had been already published in other articles not captured by the index systematic review. Studies with controls had modestly higher citation impact (median 7 versus 4 per year, p=0.002). Conclusions Most published long COVID prevalence studies lacked comparator groups and relied exclusively on self-reported outcomes without acknowledging this limitation. Direct author contact identified little additional comparator information. Much of the long COVID prevalence literature may therefore be poorly suited to estimating burden attributable specifically to SARS-CoV-2. Key Points Question What is the frequency of baseline control group inclusion and the reliance on subjective outcomes in published long COVID prevalence studies? Findings This meta-epidemiological analysis demonstrates that over 80% of published long COVID prevalence studies lacked a baseline non-COVID control group. Most of these investigations also relied exclusively on subjective patient-reported outcomes without explicitly acknowledging the absence of a comparator as a limitation. Meaning These findings suggest that the majority of the long COVID prevalence literature is poorly suited to accurately estimate the symptom burden specifically attributable to SARS-CoV-2.
Death certificates record causes of death as reported by certifiers (Entity Axis) and as standardized by mortality coding rules (Record Axis). Conventional statistics reduce these to a single underlying cause, ignoring other contributing conditions; weighting schemes can instead distribute the burden across all listed causes. We evaluated reclassification from Entity to Record axis and weighting across all 56,986,831 US death certificates from 2003 to 2023, mapping International Classification of Diseases (ICD)-10 codes to 14 broad disease categories and testing three weighting schemes: W1 (50% to the underlying cause, 50% shared equally among contributing causes), W2 (equal weighting across all causes), and W2A (equal weighting at the ICD-10 code level). Entity and Record Axes agreed on underlying cause in 48,313,403 deaths (84.8%) by broad category and 39,260,709 deaths (68.9%) by ICD-10 code (SI Appendix, Table S5 A and B); concordance reached 70.4% using 3-character. Reclassification markedly increased COVID-19 (+92%) and Transport deaths (+43%) while decreasing Other External Causes (-54%). Weighting substantially altered burden attribution: COVID-19 (-44 to -63%) and Falls (-46 to -66%) decreased, Other External causes more than tripled (+204 to 254%), and deviations from unweighted counts were more pronounced with W2 and W2A than W1. Weighting also brought disease-category counts closer to Entity Axis values and restored Respiratory seasonality suppressed during the pandemic. Systematic differences between reported and reclassified causes of death, and the choice of weighting scheme, profoundly alter disease burden estimates for several causes, with major implications for resource allocation and public health priorities.
This Viewpoint examines adversarial collaboration—bringing investigators with opposing views to design, analyze, and publish studies together, often with a neutral arbiter—as a strategy for credible biomedical science.
Scale-up penalty, a common phenomenon in which the promising effects found in early preliminary studies are substantially reduced when evaluated in a subsequent larger trial, can stall the advancement of health behavior interventions. In obesity-related behavioral interventions, changes to key features between a preliminary study and subsequent larger trials inflate scale-up penalty. The purpose of this study is to examine whether changes in intervention features occur in other behavioral disciplines that utilize a similar developmental continuum wherein smaller-scale preliminary studies inform larger-scale trials, and whether changes in key features inflate scale-up penalty. We conducted a systematic review identifying preliminary studies followed by a larger trial conducted by the same author(s) (i.e., a study pair) in four areas—tobacco/smoking cessation, alcohol use, interpersonal violence, and sexually transmitted diseases. We coded intervention features in the preliminary study and larger trial to capture changes in key study features (e.g., who delivered the intervention). Multi-level meta-regressions estimated the association between the changes to key study features and change in standardized mean difference for health outcomes and calculated scale-up penalty. We identified 222 effects across 69 study pairs of preliminary studies with subsequent larger trials. Fifty-eight study pairs (84
Importance The Global Burden of Disease (GBD) reports widely used estimates of mortality and disability-adjusted life-years (DALYs) and related risk factors. However, the overall reliability of these estimates between GBD iterations has not been assessed. Objective To evaluate the instability and inconsistency of GBD risk factor estimates for mortality and DALYs across GBD iterations. Data Sources GBD risk factor collaboration estimates extracted from the published tables of GBD iterations and the Institute for Health Metrics and Evaluation repository. Study Selection GBD risk factor collaboration publications published for 2010 through 2023. Data Extraction and Synthesis Death and DALY estimates were manually extracted by 1 reviewer with independent validation of a random sample of 100 by another with no discrepancies. Risk factor naming was harmonized across iterations to ensure comparability; those with inconsistent definitions were excluded. Main Outcomes and Measures Fluctuations were calculated for numbers of deaths and DALYs for each risk factor across GBD iterations during the study period (2010-2023) and between the original and subsequently revised estimates for each year (1990-2021). Differences were expressed as a ratio of the minimum to maximum range to the mean (R:M) and coefficient of variation. Detail analyses assessed diet and low physical activity. Point estimates were compared to the previous iterations' estimates 95% uncertainty intervals (95% UI) for GBD 2019, 2021, and 2023. Results Across GBD iterations from 2010 to 2023, the median (range) R:M was 0.8 (0-3.8) for deaths, and 0.7 (0.1-3.3) for DALYs. Level 2 dietary and child and maternal malnutrition death estimates showed high instability (R:M >1 for 9 of 16 and 4 of 8 risks, respectively). When comparing original estimates with GBD 2019, 2021, and 2023 estimates for the same years, the median R:M was 0.4 (0-2.9) for both deaths and DALYs. The coefficient of variation was greater than 0.2 for 336 of 675 death estimates (50%). Specifically, 70% to 96% of point estimates for red meat, sugar-sweetened beverages, fruits, vegetables, and seafood omega-3 fatty acids in GBD 2021 fell outside the GBD 2019 95% UI. In GBD 2023, only diet high in trans fats had more than half of point estimates outside the GBD 2021 95% UI. Conclusions and Relevance This meta-epidemiological assessment indicates that GBD estimates are substantially unstable, particularly for behavioral risks, making them unlikely to simply reflect genuine changes over time, and warranting caution in interpretation.