
Participants who continue across the full duration of longitudinal studies often differ from those lost to attrition, defined as discontinuing participation for any reason, and some participants may sporadically miss assessments. Complex participation patterns, such as in multicenter chronic disease cohorts, have not been studied. We aimed to implement an approach to identify participation subgroups and factors associated with theses. We applied our approach in the SPIN Cohort, a multicenter, longitudinal systemic sclerosis (SSc) cohort, because it is an example of an open-ended cohort with no planned end date and complex participation patterns. We used group-based trajectory modeling to identify participation subgroups and multinomial logistic regression to identify predictors of subgroup membership. Data were obtained for 2883 participants from 54 sites in 7 countries. We identified 5 participation subgroups: Ongoing Participation (29%), Immediate Attrition (29%), Mid-Term Attrition (19%), Long-Term Attrition (12%), and Ongoing Sporadic (11%) participation. Older age and sites outside the United States were associated with lower odds of belonging to all subgroups other than Ongoing Participation. Compared to Ongoing Participation, several sociodemographic variables were associated with higher odds of being in the Immediate Attrition group. Our method can be applied to other cohorts with complex participation patterns.
As one of the most disruptive life events, layoffs are pervasive in the US labor market, while little is known about cumulative layoff experience across adulthood in relation to subsequent cognitive function. We investigated the association between cumulative layoffs spanning over 41 years of adulthood and subsequent domain-specific cognition among US working-age adults. Data were from 6421 adults in National Longitudinal Survey of Youth 1979. Cumulative layoff experience was defined as the number of jobs ending in a layoff from 1979-2020: zero (never, n = 4503), one (n = 1402), and ≥ two (n = 516). Domain-specific latent cognitive factor scores were assessed in 2020 using confirmatory factor analysis. Multivariable-adjusted linear regression models were fit to investigate the association of interest. Results suggest that cumulative layoff experience was associated with lower latent factor scores in episodic memory (β = -0.045, 95% CI: -0.100 to 0.010, one layoff vs. never; β = -0.086, 95% CI: -0.171 to -0.001, ≥two layoffs vs. never; P trend 0.037), but not in attention, orientation, or language domains. Given its ubiquity in the US labor market, layoffs may represent an important and understudied risk factor for episodic memory function among US adults.
We investigated combinations of preterm birth, preeclampsia and offspring birthweight by gestational age in a first pregnancy and cross-over risk of gestational diabetes mellitus (GDM) in second pregnancy. The Medical Birth Registry of Norway provided data for 561 873 mothers with first and second births without diabetes prior to second pregnancy, 1985-2020. We combined first pregnancy preterm birth (<37 weeks), preeclampsia and birthweight by gestational age quartiles (Q1-4) into one exposure (16 categories). Relative risks with 95% confidence intervals for GDM in second pregnancy were estimated, keeping mothers with first term birth without preeclampsia and offspring in Q1 as reference. Women with first offsprings in Q4 had high risk of subsequent GDM, however, in combinations with preterm birth and preeclampsia the risk of GDM progressively increased: Compared to 0.9% GDM risk in the reference group, the risk increased to 1.8% for Q4 term births without preeclampsia (aRR 2.1 [95% confidence interval 2.0-2.3]), 2.5% for Q4 preterm births without preeclampsia (aRR 3.3 [2.8-3.8]), 4.2% for Q4 term birth with preeclampsia (aRR 5.6 [4.8-6.4]) and 7.8% for Q4 preterm birth with preeclampsia (aRR 10.1 [7.4-13.7]). Combinations of first pregnancy exposures were associated with a progressive cross-over risk of GDM in second pregnancy.
This study examined changes in COVID-19-related death certification practices (ie, reporting COVID-19 in Part I of the death certificate) and the quality of cause-of-death (COD) reporting (acceptable and rejected causal sequences) in the United States from 2020 to 2023. In 2020, 2021, 2022, and 2023, 385 293, 463 273, 246 196, and 76 024 death certificates, respectively, mentioned COVID-19. COVID-19 was selected as the underlying COD in 91%, 90%, 76%, and 65% of these certificates, corresponding to reporting COVID-19 in Part I in 93%, 91%, 78%, and 67% of certificates, respectively. Acceptable causal sequences were reported in 51%, 51%, 45%, and 41% of certificates, whereas rejected causal sequences were reported in 21%, 22%, 23%, and 23%, respectively. Using 2020 as the reference year, the adjusted odds ratios (95% confidence intervals) for 2023 were 0.18 (0.18, 0.18) for reporting COVID-19 in Part I, 0.77 (0.76, 0.78) for reporting acceptable causal sequences, and 1.14 (1.12, 1.16) for reporting rejected causal sequences. In conclusion, COVID-19-related death certification practices changed substantially over the study period, paralleling the decline in the virulence of SARS-CoV-2. The quality of COVID-19-related COD reporting declined modestly from 2020 to 2023.
The COVID-19 pandemic led to stay-at-home orders, resulting in sudden changes in movement patterns and social restrictions that impacted mental health. This study examined changes in individual behaviors during the pandemic using detailed smartphone-based activity data and determined whether behaviors and green space exposure buffered against mental health concerns. A total of 224 twins from the Washington State Twin Registry provided smartphone location data and completed a baseline survey and up to seven waves of follow-up surveys on mental health outcomes (ie, anxiety, depression, and stress). Objective measures of activities, locations, and time spent in green space were collected using Google Location History. Green space exposures were assessed using all location data. Associations between activities, locations, green space, and each mental health outcome were assessed using linear mixed models. There were changes in activities and locations from baseline to wave 1; by wave 7, the direction and magnitude of changes varied across measures, with some remaining above and others below their baseline levels. Mental health outcomes were associated with a subset of locations, activities, and green space measures although associations were not observed consistently across all exposure-outcome combinations examined.
BACKGROUND:Despite the association between lower perceived neighborhood social cohesion (PNSC) and higher risk of hypertension, there is limited research on this relationship. This population-based study examined the association between PNSC and incident hypertension among a nationally representative population of older adults across racial and/or ethnic groups and sex. METHODS:Data were obtained from 2,998 US older adult participants (mean age=62.5 years, SD=8.12) from the Health and Retirement Study. Normotensive participants at baseline (2006 or 2008) were assessed for hypertension incidence at 4- and 8-year follow-up visits. PNSC was classified into tertiles. Weighted Cox proportional hazards regression was used to estimate the association between PNSC and incident hypertension. Interaction terms (PNSC and racial and/or ethnic group, PNSC and sex) were included in models adjusted for appropriate covariates (sex, racial and/or ethnic group, education, body mass index, alcohol use, cigarette smoking status), and models were stratified by racial and/or ethnic group and sex. RESULTS:Overall, 40.5% of the study population developed hypertension. Low perceived neighborhood social cohesion (versus high) was associated with increased hypertension risk (HR=1.22, 95% CI: 1.08-1.37). This association was the strongest among Black adults (HR=2.28, 95% CI: 1.23-4.25) and Black females (HR=3.40, 95% CI: 1.34-8.62), though the interaction terms were not statistically significant (p-values for interaction >0.05). CONCLUSIONS:Individuals perceiving low neighborhood social cohesion may have increased hypertension risk, particularly among Black females. Future studies could test potential underlying mechanisms that explain this association and implement community-based interventions promoting cohesive communities as a potential strategy for reducing disparities in hypertension risk.
Depression analyses in UK Biobank may be affected by selection bias due to incomplete participation in the optional Mental Health Questionnaire. We used one-sample Mendelian randomization (MR) to assess the causal effects of BMI, educational attainment (EA), and CRP on depression, applying inverse probability of selection weighting (IPW) and the instrumental variable for selection (IVsel) to adjust for selection into the MHQ sample. Sensitivity analyses included the use of psychiatrist consultation as a proxy outcome. MR analyses indicated that higher BMI increased the risk of depression (OR=1.09 [1.05, 1.13]), EA was protective (OR=0.81 [0.75, 0.88]), and there was limited evidence that CRP affected depression (OR=1.01 [0.92, 1.10]). Estimates obtained using IPW and IVsel were broadly consistent, although with IVsel estimates closer to the null, within BMI (ORIPW = 1.10 [1.05, 1.15], ORIVsel = 1.04 [1.02, 1.06]), EA (ORIPW = 0.81 [0.74, 0.88], ORIVsel = 0.93 [0.90, 0.96]) and CRP (ORIPW = 1.03 [0.92, 1.15], ORIVsel = 1.00 [0.97, 1.04]). Proxy outcome analyses suggested similar patterns of selection bias, supporting their utility for assessing potential bias in MHQ-restricted analyses. Bias-adjustment methods and sensitivity analyses provide a framework to assess the impact of selection bias in large cohort studies with optional components.
A recent news article suggested observational studies confuse correlation and causation. Such confusion often arises from unclear research aims rather than inherent limitations of study design. We use this recent public critique of epidemiologic research to illustrate how ambiguity in study questions contributes to misinterpretation of findings. We argue that epidemiologic studies generally fall into three distinct categories (descriptive, predictive, and causal) and that each has different goals, assumptions, analytic approaches, and criteria for evaluation. Descriptive studies characterize the distribution of health outcomes, predictive studies aim to identify who will experience those outcomes, and causal studies seek to estimate the effects of interventions or exposures under well-defined counterfactual contrasts. Failure to clearly state which of these aims is being pursued makes it difficult to evaluate methods, assess validity, and interpret results. It also creates opportunities for critiques that may mischaracterize study intent or overstate limitations. We emphasize that observational studies can contribute to causal inference when aligned with explicit causal questions and supported by appropriate assumptions and design choices. We conclude that explicitly stating study aims and estimands would improve scientific communication, facilitate more appropriate critique, and strengthen the contribution of epidemiologic evidence to public health decision-making.
Health insurance claims-based analyses inform many large epidemiologic studies but may have unmeasured confounding. Electronic health records (EHRs) have more detailed health information, but data may be missing on some individuals. Informed by recent work comparing approaches to address missing data, we selected and applied generalized raking (GR) and multiple imputation (MI) to illustrate how to efficiently combine data from these two sources. We considered a previously conducted claims-based study comparing 90-day arterial thromboembolism (ATE) risk among patients hospitalized with COVID-19 versus influenza, for which body mass index (BMI), a potential confounder, was unavailable. Using linked claims-EHR data, we adjusted for the same covariates as the original study and used GR and MI to additionally control for EHR-ascertained BMI, available for a subset. We included 912 hospitalized Kaiser Permanente Washington patients, 449 with COVID-19 and 463 with influenza (31.0% and 38.5% had EHR-ascertained BMI in the prior 90 days, respectively). Adjusted hazard ratios of ATE were similar with and without adjustment for BMI, in GR and MI analyses. In the setting of a claims-EHR-based study with missing data on a potential confounder, we describe our selection of the missing data approach, and its implementation, to demonstrate a broadly applicable process.
Should original research articles routinely contain prominent policy claims or broad calls to action? Growing emphasis on research impact might be welcome yet have unintended consequences (eg, incentivising overextrapolation, and undermining the perceived objectivity of scientists). We examined 45 807 abstracts from ten leading Epidemiology and Public Health journals (1990-2024). Using a large language model with human validation, we classified policy claims and mapped trends. Claims markedly increased from 17.6% to 35.8%, with wide variation across countries and journals (>60% vs < 4%). Keywords linked to higher claim rates differed by topic and time: some corresponded to topics with clear causal evidence of harm as well as topics with notable advocacy. Claims were most common in qualitative or cross-sectional studies, and less common in cohort, quasi-experimental, or experimental studies. We argue that these patterns reflect a research culture increasingly oriented toward claiming policy relevance-and incentives that encourage attaching claims to single studies. Our findings raise questions about how scientists and journals balance evidence, advocacy, and scientific credibility. Ensuring that policy claims and calls to action remain commensurate with evidence will be central to building trust as policy impact continues to be incentivised.
Prior studies suggest fluoride exposure in drinking water above 1500 μg/L is associated with lower child cognition, but evidence is limited at lower exposure levels and in U.S. populations. We evaluated whether prenatal exposure to fluoride in regulated public drinking water was associated with cognition in a pooled U.S. cohort. We analyzed observational data from the Environmental influences on Child Health Outcomes (ECHO) Cohort, including 2514 children born 2006-2019 across 17 sites in 23 states. Individual prenatal time-weighted average public water fluoride concentrations were estimated by linking census tract-level concentrations to residential addresses across pregnancy. Fluid and crystallized cognition were assessed using NIH Toolbox scores. Generalized estimating equation models estimated adjusted mean differences using restricted cubic spline and linear change-point models. Individual prenatal time-weighted average water fluoride concentrations ranged from < 1.0-1940.0 μg/L (mean = 396.9 μg/L). Cubic spline models showed significant inverse associations for fluid cognition above 1107.0 μg/L. Linear change-point models identified 675 μg/L as the best-fitting change-point for fluid cognition; above this value, fluid scores were 0.67 points lower (95% CI, -0.92, -0.42) per 100 μg/L higher fluoride. These findings indicate that prenatal fluoride exposure in regulated public water is nonlinearly associated with lower fluid cognition scores in U.S. children at concentrations below current WHO and U.S. EPA thresholds.
The potential for trace amounts of selenium supplements to prevent prostate cancer and other neoplasms was studied in SELECT, a randomized trial conducted in North America for 7-12 years. 34 887 eligible men age 55 or older were assigned to take either 200 μg selenium from L-selenomethionine, 400 IU vitamin E, both selenium and vitamin E, or placebo. The trial was stopped for futility and concerns about adverse effects. Subsequently, we used incidence and mortality data from Medicare through 2019 and the National Death Index through 2024 to investigate the extent to which selenium administration was associated with increased long-term risk of cancer and neurodegenerative disease, during 341 697 person-years of follow-up. Lung cancer showed little increased risk. Associations were near null for leukemia, non-Hodgkin lymphoma and Parkinson's disease, and there was a lower risk for melanoma and colorectal cancer. In contrast, we found a higher risk for Hodgkin lymphoma, multiple myeloma and amyotrophic lateral sclerosis, which has been associated with selenium overexposure in some nonexperimental studies. Though these latter associations were based on few cases, they were consistent with findings from a range of other studies, possibly indicating long-term adverse effects of this organic selenium compound. (Registration no. NCT00006392).
Target trial emulation prompts investigators to frame their analysis question in terms of a hypothetical clinical trial. Although this does not solve the problem of confounding, the framework can protect against other sources of bias. A natural question is this: what kinds of trials can be emulated? In a late-phase trial (Phase 3 or 4), the goal is to obtain a well-defined causal estimate that closely approximates the impact of a proposed intervention. In an early-phase trial (Phase 2 or earlier), the estimate is a means to an end rather than an end in itself. An early-phase trial provides proof-of-concept evidence on the impact of an intervention in the exposure in terms of efficacy and safety, but the estimand may not correspond to the intervention to be implemented in practice. In a natural experiment where causal inferences rely on a plausibly random (or quasi-random) comparison, the estimand may not be directly translatable to applied practice. In this case, the analysis may be conceptualized as an early-phase target trial. This provides less specific evidence than a late-stage target trial, but in many cases, a more valid but less applicable comparison is preferable to a more applicable comparison that is more susceptible to bias.