
PURPOSE:Adjustment for a sufficient set of confounders removes bias. When a confounder is unmeasured, approaches exist to estimate bias. Confounders are often correlated, so analyses that ignore correlations overstate bias. METHODS:Using NHANES III, we examined the association between Healthy Eating Index (HEI) and all-cause mortality (n = 2417). A fully adjusted model included tobacco use, sex, age, hypertension, BMI, education, and physical activity. Hazard ratios (HR) that would have been observed-had one of hypertension, BMI, education, or physical activity been "unmeasured"-were estimated by leaving them out. We then performed bias analysis for the unmeasured confounders using uncorrelated bias parameter estimates. RESULTS:The fully adjusted HR comparing HEI Quintile 1 vs. 5 was 1.72 (95% CI 1.24-2.40). After treating variables as "unmeasured" confounders, HRs changed little (range: 1.73-1.98). Bias-adjusted hazard ratios ranged from 1.72 to 1.98 suggesting that substantial unmeasured confounding would be required to explain the associations. CONCLUSIONS:Due to correlations between covariates, the additional bias attributable to unmeasured variables was minimal. QBA produced estimates that often overestimated the impact of the unmeasured confounders. Although QBA is useful for evaluating unmeasured confounding, it may not precisely quantify the strength of bias.
PURPOSE:Cannabis is the most-used substance during pregnancy and has known associations with psychiatric disorders (e.g., psychosis, depression). No prospective studies, however, have investigated associations between prenatal cannabis use disorder (CUD) and psychiatric outcomes in perinatal populations. METHODS:Annual cohorts of women with a live infant delivery in California from 2016 to 2021 were constructed using statewide linked hospital and emergency department (ED) discharge data. We examined temporal trends in prenatal CUD prevalence and compared risk of postpartum psychiatric ED visits between delivering women with prenatal CUD, prenatal alcohol use disorder (AUD; included as a benchmark exposure), and neither disorder. RESULTS:The study sample included 2,036,016 delivering women. Between 2016 and 2021, prenatal CUD prevalence increased 1.0-1.4%. Postpartum psychiatric ED visits occurred in 10.4% of women diagnosed with CUD, 15.6% with AUD, and 2.7% with neither diagnosis. Relative to women with neither, those with CUD had ∼40% higher risk of any postpartum psychiatric ED visit (RRadj = 1.39, 95% CI: 1.32-1.47), while AUD was associated with 80% increased risk (RRadj = 1.80, 95% CI: 1.56-2.09). CONCLUSION:Elevated risk of adverse psychiatric outcomes associated with CUD underscore the need for consistent screening, diagnosis, and treatment among perinatal women with CUD diagnoses.
OBJECTIVE:Community resilience during COVID-19 has been linked to social conditions, public-health capacity, and acute-care resources, but the contribution of pre-pandemic realized preventive-care engagement remains incompletely characterized. We examined whether county-level preventive-care engagement was associated with all-cause excess mortality during the COVID-19 period in the United States. METHODS:In this ecological study, county-specific linear trends fitted to the age-adjusted all-cause premature-death/YPLL-75 indicator for 2012-2018 were used to predict expected mortality in 2020-2022. Excess mortality was the percentage difference between observed and expected values; resilience was its negative. We estimated adjusted associations for four preventive-care indicators using five-fold cross-fitted double/debiased machine learning (DML), with learner, fold, state-clustered, external-demographic, and spatial lag/error sensitivities. RESULTS:Across 3079 counties, mean excess mortality in 2020-2022 was 15.2%. In the primary DML sample (n = 2921), one-standard-deviation higher influenza-vaccination and mammography-screening rates were associated with 0.106 SD (p < 0.001) and 0.118 SD (p < 0.001) higher resilience, respectively. Preventable hospitalization was negatively associated; primary-care physician density was null. Estimates were stable across all sensitivities, although attenuated after broader demographic adjustment, and influenza vaccination predicted a pre-pandemic placebo outcome. CONCLUSIONS:Pre-pandemic realized preventive-care engagement was associated with lower county-level all-cause excess mortality. These ecological associations are not individual-level causal effects and warrant cautious interpretation given demographic attenuation and the positive placebo. Resilience policy should combine rapid screening and vaccination outreach with sustained, equity-oriented primary-prevention infrastructure.
PURPOSE:This study quantified how dementia ascertainment source shapes population-level place of death (POD) estimates among older adults dying with dementia. METHODS:This cross-sectional study (2014-2023) in Taiwan linked death certificates and NHI claims for decedents aged ≥ 65 years identified with dementia. POD was categorized as hospital, home, long-term care (LTC) facility, or inpatient palliative care unit. Ascertainment-related POD differences were assessed using multinomial logistic regression with generalized estimating equations and average marginal effects. RESULTS:Among 291,652 decedents, NHI claims identified 279,074 individuals (underlying cause of death [UCOD], 24,911; multiple causes of death [MCOD], 43,487). Only 5.2% of decedents were identified by all three sources; 83.4% of NHI-identified cases had no dementia recorded on death certificates. NHI ascertainment was associated with significantly lower (vs. UCOD) adjusted odds of palliative unit (aOR 0.32; 95% CI 0.30-0.33), home (aOR 0.68; 0.66-0.70), and LTC deaths (aOR 0.61; 0.58-0.64). NHI-based identification showed a 13.6-percentage-point higher (vs. MCOD) hospital death probability (95% CI 13.2-14.1). CONCLUSIONS:Dementia ascertainment source defines the dementia decedent population monitored and shapes the POD distribution among dementia decedents. Death-certificate-only ascertainment may underestimate hospital deaths. Multisource linkage reveals source-dependent exclusions relevant to end-of-life assessment and policy planning.
BACKGROUND:Loneliness and social isolation are related but distinct dimensions of social disconnection, but their longitudinal associations with chronic condition accumulation remain unclear in non-Western populations. METHODS:We analyzed 8721 adults aged 18 years or older who completed the 2021 and 2025 waves of the Japan "COVID-19 and Society" Internet Survey. Loneliness and social isolation were assessed in 2021, during the COVID-19 pandemic, using the UCLA 3-item Loneliness Scale and the 6-item Lubben Social Network Scale. The outcome was the newly reported presence in 2025 of at least one of 14 physician-diagnosed chronic conditions not reported as currently present in 2021. Standardized inverse probability-weighted logistic regression estimated adjusted odds ratios (ORs) per 1-SD higher loneliness score and per 1-SD lower social network score, adjusting for both exposures, sociodemographic factors, and baseline condition count. RESULTS:Among 8721 participants, 2399 (27.5%) reported at least one new condition. Loneliness was associated with the outcome (OR, 1.24; 95% CI, 1.13-1.36), whereas social isolation was not (OR, 1.06; 95% CI, 0.97-1.17). Additional adjustment for baseline psychological distress attenuated the loneliness association (OR, 1.10; 95% CI, 0.98-1.23). CONCLUSIONS:Loneliness, but not social isolation, was associated with subsequent reporting of new chronic conditions. Psychological distress may partly account for this association.
OBJECTIVE:Traumatic cardiac arrest differs from non-traumatic regarding epidemiology. This study evaluated five machine learning classifiers' ability to identify traumatic cardiac arrest in the Danish Cardiac Arrest Registry, which currently relies on manual review. METHODS:This retrospective study employed split sampling to train and test models using medical records of cardiac arrest patients from 2016 to 2021. Data were preprocessed, resampled, and classified using five classification models. Shapley Additive Explanations values were used to explain the models' predictions except for the BERT model. RESULTS:30,171 medical records included, with 985 involving traumatic cardiac arrests. The histogram gradient boosting model achieved the highest F1-score of 0.62, with precision (positive predictive value) of 0.61 and recall (sensitivity) of 0.62. The BERT model demonstrated a recall of 0.96, a precision of 0.21, and an F1-score of 0.35. For histogram gradient boosting and random forest, the increasing presence of "car", "thorax", and "head" was linked to trauma. Further error analysis revealed that these models exhibited age-related bias by implicitly learning that traumatic cardiac arrest is more common in younger populations and less frequent in older populations. CONCLUSION:This study represents the first steps towards using machine learning algorithms to support manual validation of the Danish Cardiac Arrest Registry. Histogram gradient boosting yielded the best overall results; however, the BERT model demonstrated significantly improved recall in detecting trauma cases. Hence, applying the BERT model to the validation would substantially reduce the present workload.
BACKGROUND:Menstrual cycle length is an important marker of reproductive and overall health; however, the validity of self-reported measures remains debated. PURPOSE:This study estimated agreement between self-reported and calculated menstrual cycle length among premenopausal participants in the Cancer Prevention Study-3. METHODS:Usual cycle length was reported using six predefined categories and calculated using prospectively reported menstrual start dates. Agreement was assessed using descriptive comparisons, Spearman's correlation, and weighted kappa statistics. Analyses incorporated a ± 7-day buffer to account for expected biological variability. RESULTS:Of 55,458 participants, age range: 19-54 years old, 44.9% had exact agreement between self-reported and calculated cycle length. When applying the ±7-day buffer, 82.2% were in agreement. Correlation was weak using exact categories (ρ=0.17) but improved with the ±7-day buffer (ρ=0.53). Agreement was poor using exact categories (κ=0.15) and increased to fair when accounting for biological variation (κ=0.33). Agreement varied across subgroups, with generally higher concordance among participants aged 30-40 years, White or Asian, graduate degree, higher income, or nulliparous and lower concordance among separated/divorced/widowed participants. CONCLUSION:Self-reported menstrual cycle length demonstrates fair agreement when biological variability is considered, supporting its use in epidemiologic research and providing empirical justification for incorporating in menstrual-related analyses.
PURPOSE:Given limited recent nationally representative evidence on depression and metabolic syndrome (MetS), we compared this association in South Korea and the United States (US). METHODS:We analyzed data from the Korea National Health and Nutrition Examination Survey (KNHANES, 2014-2022) and the National Health and Nutrition Examination Survey (NHANES, 2013-2020). Depressive symptoms were assessed using the Patient Health Questionnaire-9 (PHQ-9; categorized as low [0-9], moderate [10-14], or high [15-27]). MetS was defined by the National Cholesterol Education Program Adult Treatment Panel III. Weighted logistic regression analyses were performed to calculate weighted odds ratios (wORs) and 95% confidence intervals (CIs) by PHQ-9 category. RESULTS:We included 25,207 KNHANES participants and 16,447 NHANES participants. MetS prevalence was higher in the high PHQ-9 group (South Korea, 38.19% [95% CI, 33.05-43.33]; the US, 43.20% [36.73-49.68]). Estimated wORs for MetS generally increased at higher PHQ-9 scores. High versus low PHQ-9 scores were associated with higher MetS odds in both countries (South Korea: wOR, 1.34 [95% CI, 1.07-1.66]; the US: 1.44 [1.09-1.91]), with a numerically higher estimate in the US. Subgroup analyses suggested that differences in MetS prevalence across PHQ-9 categories were more evident among females, adults aged 40-64 years, and participants with lower household income. CONCLUSIONS:These findings underscore the need to incorporate depressive symptom assessment into MetS prevention and management, with strategies tailored to national contexts.
PURPOSE:Clear guidelines on using dual-energy X-ray absorptiometry (DXA) are lacking. DXA-derived phenotypes based on whether the person was above or below the median fat- and muscle-mass compared to a reference population were previously constructed. Whether this cutoff can be improved with unsupervised clustering techniques was the study objective. METHODS:Data were from the National Health and Nutrition Examination Survey (1999-2006 cycles, n = 5566; split into 70/30% training and test datasets), a representative U.S. SAMPLE:Phenotypes based on partitioning deciles of fat- and muscle-mass adjusted for age and sex by k-means, and hierarchical clustering were identified. Model fit was assessed using the silhouette and elbow method. Performance of logistic regression models to identify unfavorable cardiometabolic risks was assessed with the area under the receiver operating characteristic curves (ROC-AUC), stratified by sex and incorporating weighting and the complex sampling design. RESULTS:Optimal models were 2-means k-clusters, 4-means k-clusters, and 5 hierarchical clustering phenotypes. ROC-AUCs from 2-means k-clusters (0.52-0.63) were the lowest. Performance of the hierarchical clustering and the 4-means k-cluster phenotypes was higher, but not statistically significantly different from the median-split. CONCLUSIONS:While unsupervised clustering methods improved performance, ROC-AUCs were moderate. Future work investigating other health outcomes is needed.
Aim To estimate the effect of time to treatment initiation on 3- and 5-year survival among women with cervical cancer in the Brazilian Western Amazon. Methods We conducted a retrospective cohort study of women diagnosed with cervical cancer and treated between 2012 and 2017 in Rio Branco, Acre, Brazil. Data were obtained from medical records, and vital status was ascertained using national health information systems, mortality databases, and official registries. Time to treatment initiation was defined as the number of days from diagnosis to first treatment and categorized as ≤60, 61–90, and >90 days. Overall survival at 36 and 60 months was estimated using the Kaplan–Meier method. Crude and adjusted hazard ratios (HRs) with 95% confidence intervals (95% CIs) were estimated using Cox proportional hazards models. Results A total of 388 women were included, of whom 172 (44.3%) died within 5 years. Overall survival at 3 and 5 years was 65.7% and 55.7%, respectively. At 60 months, higher mortality risk was significantly associated with not having a partner (HR = 1.61; 95% CI: 1.17–2.20), advanced stage at diagnosis (IIB–IVA) (HR = 3.93; 95% CI: 2.31–6.68), smoking (HR = 1.49; 95% CI: 1.07–2.08), and receiving radiotherapy (HR = 3.48; 95% CI: 1.85–6.55) or chemoradiation (HR = 7.00; 95% CI: 3.73–13.16). Compared with initiation within ≤60 days, treatment initiation after >90 days was associated with a lower risk of death (HR = 0.44; 95% CI: 0.27–0.70). Conclusion In the Brazilian Western Amazon, advanced stage at diagnosis, receipt of radiotherapy or chemoradiation, and absence of a partner were independently associated with increased 5-year mortality. Unexpectedly, longer time to treatment initiation (>90 days) was associated with lower mortality, warranting further investigation into potential confounding and health system factors.
BACKGROUND:Nationality, ethnicity, and geographic background are frequently required in medical research but are often unavailable in administrative or registry-based datasets. Name-based inference tools have demonstrated good performance for predicting country of origin, yet their ability to approximate legal nationality remains unclear. OBJECTIVE:To evaluate the performance of NamSor in predicting nationality from personal names in a large multinational cohort and to assess whether aggregation into broader geographic or onomastic regions improves classification accuracy. METHODS:This cross-sectional study included 11,989 marathon participants representing 135 nationalities. Self-reported nationality, as recorded in the official race results, served as the reference standard. NamSor predictions were evaluated at the country level, fine/coarse United Nations (UN) regional levels, and predefined onomastic macro-regions. Performance was assessed using classification accuracy (proportion of correct predictions among classified observations) across probability thresholds. RESULTS:Country-level accuracy was 60.2%. Aggregation improved performance to 69.7% for fine UN regions and 75.2% for coarse UN regions. Coarse onomastic macro-regions achieved the highest accuracy (88.3%). Increasing probability thresholds improved accuracy among classified observations (e.g., 92.5% at ≥0.9 at the country level) but substantially reduced the proportion of observations retained for analysis, with similar trade-offs observed for regional and onomastic classifications. CONCLUSIONS:Name-based inference aligns more closely with linguistic-cultural groupings than with exact legal nationality. While country-level prediction showed substantial misclassification in a highly multinational setting, aggregation into broader regional or onomastic categories markedly improved performance. Broader regional or onomastic classifications may therefore represent a pragmatic alternative when direct nationality data are unavailable.