BackgroundThe Pandemic Electronic Benefit Transfer (P-EBT) program replaced free and reduced-price school meals for low-income children during school closures and summers from 2020 to 2023. Evidence indicates that P-EBT disbursements were associated with a 40% reduction in food hardship and better maternal mental health among low-income households with children. As states had discretion in implementing the program, approaches varied across states. This study characterized patterns of P-EBT implementation and disbursement across all 50 states and Washington, D.C., in relation to state characteristics.MethodsWe examined three outcomes, namely timing of first P-EBT disbursement in 2020, inconsistency of summer P-EBT participation through 2023, and annual per capita P-EBT disbursement amounts from 2020 to 2023, using ordered logistic, logistic, and linear regression models. Independent variables included measures of state economic resources, sociopolitical characteristics, and indicators of need for food assistance. To account for regional cost-of-living differences, we conducted a sensitivity analysis using a real state minimum wage adjusted for Regional Price Parity (RPP).ResultsIn unadjusted analyses, Republican gubernatorial affiliation was associated with later first disbursement, lower per capita disbursement, and inconsistent summer participation. In adjusted models, no state characteristics were significantly associated with the timing of the first disbursement. Each additional month of prior-year Supplemental Nutrition Assistance Program (SNAP) emergency allotments was associated with lower odds of inconsistent Summer P-EBT participation [odds ratio (OR) = 0.72, 95% confidence interval (CI): 0.57–0.90]. Higher Medicaid participation rates [regression coefficient (B) = $1,932.40, p < 0.005] and more months of SNAP emergency allotments [B = $23.70, p < 0.05] were associated with higher per capita disbursements. Sensitivity analyses using the RPP-adjusted minimum wage yielded materially unchanged results.ConclusionP-EBT was implemented more comprehensively in states with Democratic governors and in states with more comprehensive SNAP emergency allotments, Medicaid, and other safety-net programs. Indicators of economic conditions, such as poverty and unemployment rates, were not associated with implementation patterns. Findings underscore the role of political context in shaping food assistance implementation, with direct implications for the SUN Bucks program, which uses practices similar to P-EBT but is not yet implemented by all states.
Motivated by the study of state opioid policies, we propose a novel approach that uses autoregressive models for causal effect estimation in settings with panel data and staggered treatment adoption. Specifically, we seek to estimate the impact of key opioid-related policies by quantifying the effects of must access prescription drug monitoring programs (PDMPs), naloxone access laws (NALs), and medical marijuana laws on opioid prescribing. Existing methods, such as differences-in-differences and synthetic controls, are challenging to apply in these types of dynamic policy landscapes where multiple policies are implemented over time and sample sizes are small. Autoregressive models are an alternative strategy that have been used to estimate policy effects in similar settings, but until this paper have lacked formal justification. We outline a set of assumptions that tie these models to causal effects, and we study biases of estimates based on this approach when key causal assumptions are violated. In a set of simulation studies that mirror the structure of our application, we show that our proposed estimators frequently outperform existing estimators. In short, we justify the use of autoregressive models to evaluate the effectiveness of four state policies in combating the opioid crisis.
Background:Real-world psychiatric care is marked by wide heterogeneity in clinical presentations and outcomes, underscoring the need for systematic approaches to outcome measurement. The Clinical Global Impression-Severity (CGI-S) scale is a brief, clinician-rated measure of overall illness severity that is widely used in psychiatric research, but rarely documented in routine care. Large language models (LLMs) may enable automated extraction of CGI-S scores from narrative clinical notes, thereby providing scalable outcome measures for real-world clinical care and research. Objective:The study aimed to evaluate whether LLMs can estimate CGI-S scores from psychiatric clinical notes for patients with major depressive disorder (MDD) and to compare performance across prompting strategies and model architectures. Methods:We extracted psychiatrist-authored notes from the Johns Hopkins electronic health record. Three board-certified psychiatrists independently rated 77 clinical notes using a validated depression-specific Clinical Global Impression (CGI) rubric. Weighted Cohen kappa coefficients were calculated to assess inter-rater reliability and model-human agreement. We evaluated GPT-4o under zero-shot and few-shot prompting conditions and Llama-4 under zero-shot prompting. Model performance was assessed by comparing LLM-generated scores to individual rater scores and consensus ratings. Exploratory analyses evaluated whether agreement varied by patient demographics, care setting, note length, or the percentage of copy-forwarded text within each note. Results:Interrater reliability among psychiatrists was high (κ=0.77-0.78). GPT-4o with zero-shot prompting demonstrated the highest agreement with average human ratings (κ=0.85, 95% CI 0.78-0.90), and few-shot prompting did not improve performance. In contrast, Llama-4 with zero-shot prompting demonstrated lower agreement with average human ratings (κ=0.70, 95% CI 0.55-0.80). Model agreement did not significantly differ across age, sex, race, treatment location, or the percentage of copy-forwarded text, but it was significantly lower for notes below the median note length than for notes at or above the median length (κ=0.72 vs 0.92; P=.003). Conclusions:LLMs can estimate clinician-rated CGI-S scores from psychiatric clinical notes for patients with MDD at a level of agreement comparable to that of expert interrater reliability. Performance varied by model architecture, with GPT-4o outperforming an open-source alternative. If further validated, this approach may support scalable outcome measurement in research settings and inform future efforts to implement measurement-based care in real-world psychiatric practice.
Causal inference in medicine and public health almost always depends on untestable assumptions. Estimating valid causal effects thus often requires substantive knowledge about the study context that quantitative methods alone cannot provide. This paper introduces CACE-MM, a mixed methods framework that integrates qualitative approaches with complier average causal effect (CACE) estimation to strengthen causal decision-making and assess the plausibility of key underlying assumptions. CACE-MM is the first framework to systematically integrate qualitative inquiry with causal effect estimation in the presence of noncompliance. We present a proof-of-concept application of CACE-MM using data from The Youth Empowerment Study (YES), a randomized trial of a trauma-informed intervention for youth involved in the juvenile legal system (n = 630). Following CACE analyses using an instrumental variable approach (invoking the exclusion restriction) and principal score approach (invoking the assumption of principal ignorability), we conducted 10 semi-structured interviews with key informants. Qualitative data were analyzed using inductive open coding followed by deductive mapping to the Capability, Opportunity, Motivation-Behavior (COM-B) model and the Theoretical Domains Framework (TDF). The study team assessed the plausibility of the key assumptions underlying each quantitative approach before and after the qualitative inquiry to generate integrated metainference about assumptions. Qualitative findings identified predictors of participation and outcomes that were not captured in baseline quantitative measures, raising concerns about the plausibility of principal ignorability. Interviews also clarified how meaningful exposure to intervention components was understood by implementers, informing the defensibility of participation thresholds used to invoke the exclusion restriction. More broadly, the findings demonstrate how qualitative inquiry can inform key analytic decisions that shape causal estimates, including how participation is defined, which covariates should be prioritized for measurement, and whether particular identification strategies are appropriate for specific outcomes. Building on these insights, we propose the full CACE-MM framework, incorporating both exploratory and explanatory phases, and outline decision points to guide application in applied health research. CACE-MM offers a rigorous and systematic approach for integrating qualitative evidence into causal analyses and ultimately strengthening the transparency and interpretability of the CACE in applied health research. ClinicalTrials.gov NCT03242447.
BACKGROUND AND AIMS:Dual use of cigarettes and e-cigarettes is often employed by smokers as a harm reduction strategy. This study evaluated the harm reduction potential of e-cigarettes by investigating changes in urinary biomarker of exposure (BOE) to tobacco-specific nitrosamines (TSNAs) and nicotine among smokers transitioning to dual use. METHODS:Data from 8688 adult smokers in waves 1 to 5 of the Population Assessment of Tobacco and Health Study (United States) were analyzed. 'Transitioned dual users' were defined as those changing from exclusive cigarette smoking to dual use across two consecutive waves. BOE levels were assessed before and after this transition and compared with those who continued smoking cigarettes exclusively ('remained smokers') across the same time frame. Propensity score weighting was used to adjust for baseline confounders. Further analyses were stratified by baseline cigarette consumption, dichotomized at the cohort median into heavy and light smoking groups. RESULTS:Observational analysis revealed that transitioned dual users experienced a statistically significant reduction in TSNAs without a statistically significant change in nicotine. The regression analysis comparing transitioned dual users with 'remained smokers' showed that transitioning to dual use was statistically significantly associated with a 13% decrease in 4-(methylnitrosamino)-1-(3-pyridyl)-1-butanol (NNAL) [95% confidence interval (CI) = -20% to -6%] and a 10% decrease in N'-nitrosonornicotine (NNNT) (95% CI = -15% to -4%). In contrast, dual use was statistically significantly associated with a 17% increase in total nicotine equivalents-2 (95% CI = 12%-23%) and 8% increase in total nicotine equivalents-6 (95% CI = 5%-12%). Stratified analysis further showed that the association between dual use and TSNAs reduction was attenuated among the baseline heavy smoking group (>13 cigarettes/day), with one-fourth of the effect estimates and no statistical significance. CONCLUSIONS:Smokers transitioning to dual use alongside e-cigarettes appear to show reductions in tobacco-specific nitrosamines exposure, particularly among lighter smokers, but do not appear to show reductions in nicotine exposure.
Using state-level opioid overdose mortality data (1999-2016), we evaluated the performance of panel data estimators for capturing time-varying impacts of state-level policies. Most health policy evaluations assume static treatment effects, yet many interventions exhibit dynamic impacts, raising methodological questions about optimal estimation strategies. We simulated four time-varying treatment scenarios reflecting common policy dynamics (gradual increase, gradual decline, temporary effects, and inconsistent trajectories) and compared seven methods: two-way fixed effects event study, debiased autoregressive model, augmented synthetic control, difference in differences with staggered adoption, event study with heterogeneous treatment, two-stage differences in differences, and differences-in-differences imputation. Performance was assessed using bias, standard errors, coverage probability, and root mean squared error. Estimator performance varied substantially across scenarios. Augmented synthetic controls showed lower bias but higher variance when policy effectiveness diminished over time. Difference-in-difference approaches provided reasonable coverage in some scenarios but struggled with non-monotonic effects, while autoregressive methods exhibited lower variability but underestimated uncertainty. Overall, no single estimator performed best across settings. For epidemiological policy evaluations, particularly time-sensitive interventions like opioid-related policies, researchers should weigh bias-variance tradeoffs and align methodological choices with expected effect trajectories. Careful selection of analytic approaches is critical to avoid misattribution of policy effects and ensure valid conclusions about population health outcomes.
Randomized clinical trials are considered the gold standard for informing treatment guidelines, but results may not generalize to real-world populations. Generalizability is hindered by distributional differences in baseline covariates and treatment-outcome mediators. Approaches to address differences in covariates are well established, but approaches to address differences in mediators are more limited. Here, we consider the setting where trial activities that differ from usual-care settings (e.g., monetary compensation and follow-up visits frequency) affect treatment adherence. When treatment and adherence data are unavailable for the real-world target population, we cannot identify the mean outcome under a specific treatment assignment (i.e., mean potential outcome) in the target population. Therefore, we propose a sensitivity analysis in which a parameter for the relative difference in adherence to a specific treatment between the trial and the target, possibly conditional on covariates, must be specified. We discuss options for specification of the sensitivity analysis parameter based on external knowledge, including setting a range or specifying a probability distribution from which to repeatedly draw parameter values (i.e., use Monte Carlo sampling). We introduce two estimators for the mean counterfactual outcome in the target, which incorporate this sensitivity parameter, a plug-in estimator, and a one-step estimator that is double robust and supports the use of machine learning for estimating nuisance models. Finally, we apply the proposed approach to the motivating application where we transport the risk of relapse under two different medications for the treatment of opioid use disorder from a trial to a real-world population.
Background:While job loss is associated with adverse mental health, it is unclear if the relationship between job loss and mental health differs across pre-existing financial circumstances (objective financial status or subjective financial strain). Methods:We analyzed data from the nationally representative, longitudinal Cumulative Life Stressors Impact on Mental Health and Well-being study (2023-2024) conducted across the United States. We assessed associations between past-year job loss and probable depression (Patient Health Questionnaire-9, PHQ-9 ≥ 10), balancing on pre-job loss characteristics including objective financial status (USD, household savings ≥$5000) and subjective financial strain (difficulty meeting monthly bills). Propensity score weights were generated using a survey-weighted gradient boosted model. Generalized linear models estimated the odds ratio of probable depression by job loss group, overall and stratified by objective financial status and subjective financial strain. Findings:Among 1023 working-age adults, 79 (8.6%) experienced past-year job loss during the study period. Job loss was associated with higher odds of probable depression among adults with pre-existing subjective financial strain [Conditional Odds Ratio, cOR = 1.62 (95% Confidence Interval (CI): 1.01, 2.59)] but not in the group without subjective financial strain, although the formal test comparing estimates between strain groups was not statistically significant [3.72 (95% CI: 0.39, 35.43)]. Estimates were non-significant across objective financial status groups [<$5000 cOR = 1.56 (95% CI: 0.81, 3.02); ≥$5000 cOR = 1.43 (95% CI: 0.88, 2.32)]. Interpretation:Among working-age adults with pre-existing financial strain, job loss was associated with later probable depression. The presence or absence of $5000 or more in objective savings was not associated with differences in the association between job loss and probable depression. These findings suggest that subjective financial strain prior to job loss may be associated with differential mental health following job loss. Funding:The CLIMB study was supported in part by the de Beaumont Foundation, JHU Nexus Award, InHealth, Hopkins Business of Health Initiative Pilot Award, and Hopkins Center for Health Disparities Solutions Pilot Project Award.
Modified treatment policies (MTPs) are interventions based on each individual's natural treatment value. We study mean outcomes under MTPs for continuous treatments, including exposure mixtures. Positivity is the standard sufficient condition for identifying these mean outcomes without extrapolation: policy-generated values remain supported given covariates. With multivariate treatments or continuous covariates, treatment–covariate combinations can be sparse or unsupported. Retaining the policy, our partial-identification framework decomposes its mean outcome into a point-identified contribution inside a positivity region and one outside. We bound the latter by imposing Lipschitz continuity on conditional mean potential outcomes rather than a parametric extrapolation model. The restriction compares each outside mean with the mean at an anchor inside the region. Metric projection minimizes width among one-anchor intervals but concentrates anchors on a lower-dimensional boundary, making the endpoints not pathwise differentiable. Our novel interior-displaced projection moves anchors inward, restoring pathwise differentiability. With a known region, we derive influence functions, characterize when they are efficient, and obtain asymptotically normal estimators and confidence intervals. In simulations, our intervals attain at least nominal coverage where those assuming positivity undercover. In a pesticide-mixture application, protective associations suggested by methods assuming positivity are not robust to modest outcome variation beyond the estimated region.
Gun violence is a critical public health and safety concern in the United States. There is considerable variability in policy proposals meant to curb gun violence, ranging from increasing gun availability to deter potential assailants (e.g., concealed carry laws or arming school teachers) to restricting access to firearms (e.g., universal background checks or banning assault weapons). Many studies use state-level variation in the enactment of these policies in order to quantify their effect on gun violence. In this paper, we discuss the policy trial emulation framework for evaluating the impact of these policies, and show how to apply this framework to estimating impacts via difference-in-differences and synthetic controls when there is staggered adoption of policies across jurisdictions, estimating the impacts of right-to-carry laws on violent crime as a case study.
Abstract Background Huntington disease (HD) is an inherited neurodegenerative disorder that impairs motor, cognitive, and psychiatric function. Offspring of individuals with HD may experience early caregiving responsibilities, potentially disrupting their educational outcomes. We evaluated the associations between parental age at HD symptom onset, genetic-expansion, sociodemographic, and regional factors with offspring educational attainment outcomes in adulthood. Methods We estimated odds ratios using logistic regression to evaluate associations between higher educational attainment in offspring and parental age at symptom onset, genetic-expansion, race, and region among adults ( $$\:\ge\:$$ 18 years) in the Enroll-HD study. To assess the relative importance of exposures in predicting educational outcomes, we fit a random forest model and ranked these based on mean decrease in accuracy. Results In our explorative analysis, participants whose parents had an earlier age at HD symptom onset were associated with lower odds of attaining a higher education. We also identified a nonlinear, inverted-U association between genetic-expansion and the probability of higher educational attainment–a pattern that has been observed in prior studies of neurocognitive function in children. Marked differences were also observed by race and region: Black, Hispanic/Latino, and Native American participants were associated with lower odds of higher education compared with White participants, and those residing outside Northern America were associated with lower odds of higher educational attainment. Discussion Earlier parental HD onset was associated with lower educational attainment in offspring and disparities were observed across genetic-expansion, sociodemographic, and regional groups. Our exploratory findings may inform future studies aimed at better understanding educational inequities among families affected by HD and related neurodegenerative disorders.
OBJECTIVE:Systematic reviews remain labor-intensive, particularly when extracting methodological details from full texts. Using mediation analysis as a case study, we evaluated whether large language models (LLMs) can match human-expert-level full-text methodological review on key causal assumptions (eg, no unmeasured confounding, temporal ordering) and best practices (eg, sensitivity analyses, interaction assessments, covariate adjustment) for psychiatry and psychology studies. MATERIALS AND METHODS:We evaluated 6 LLMs from 3 major families (ChatGPT-4o-mini/4o/o3/5, Claude Sonnet 4, Gemini 2.5 Flash) on 180 full-text mediation analysis articles from 2013 to 2018 previously reviewed by expert methodologists. LLMs assessed 14 binary methodological criteria ranging from straightforward checks (eg, whether the exposure was randomized) to nuanced assessments (eg, whether the temporal ordering between mediator and outcome was established). Performance was benchmarked against expert consensus labels and individual reviewers using accuracy, precision, recall, F1, AUC, and PR-AUC. RESULTS:LLM performance strongly correlated with human reviewers across methodological criteria (accuracy correlation 0.71; F1 correlation 0.95), indicating tasks difficult for humans were likewise challenging for models. Advanced LLMs achieved near-human accuracy on explicit methodological features but lagged behind top reviewers by up to 15% on inference-intensive tasks. Longer documents reduced model accuracy. Common model errors include overinterpreting on linguistic cues and colloquial use of technical terms. DISCUSSION AND CONCLUSION:Our findings support a criterion-specific human-AI collaboration strategy for full-text methodological assessment and provide a reproducible framework for future testing in other evidence-synthesis settings.
Depression is a prevalent mental health condition, and many patients are treated in primary care, yet optimal antidepressant selection remains uncertain. Real-world data can help address gaps in understanding comparative effectiveness. We examined comparative effectiveness of duloxetine and vortioxetine in patients with depression and evaluated heterogeneity in patient characteristics, comorbidities, and healthcare utilization using electronic health records (EHR) from two large health systems. We conducted a retrospective cohort study using EHR data from Duke University Health System (DUHS) and Johns Hopkins Health System (JHHS) from 2014 to 2021. Eligible participants were adult patients prescribed duloxetine or vortioxetine. We defined a Narrow Cohort requiring a concurrent depression diagnosis and a Broad Cohort based on prescription records. Comparative effectiveness analyses employed propensity score weighting methods. The primary outcome was all-cause hospitalization within one year of cohort entry. Secondary outcomes included emergency department (ED) visits, a composite of hospitalization or ED visit, suicidal encounters, medication switching (either between the study drugs or from the study drugs to other psychiatric medications), and change in PHQ-9 score. The Narrow Cohort included 3,337 (2,937 prescribed duloxetine and 400 vortioxetine) patients in DUHS and 3,237 (2,921 prescribed duloxetine and 316 vortioxetine) in the JHHS. Across both sites, vortioxetine users were more likely to be white, privately insured, and to have prior psychiatric medication use, while duloxetine users had higher comorbidity burden. Hospitalization rates were similar between treatments in both systems (ORDUHS: 0.84 [95
To examine the relationship between family immigration status and adolescent suicidal thoughts and behaviors (STB) in a population-based sample of Latino youth in California, and, to examine whether gender moderates the association between family immigration status and youth STB. Using linked parent and adolescent data from the 2011–2019 California Health Interview Survey (CHIS) (N = 3,155 adolescents self-identifying as Hispanic/Latino), we built a series of logistic regression models, calculating adjusted odds of STB and including a gender-by-family citizenship status interaction term. . Overall, adolescents in mixed status dyads (in which the child was a citizen and the parent child was not a citizen or Lawful Permanent Resident (LPR)), and non-citizen adolescents were less likely to report STB compared to adolescents in dyads where both parent and child were citizens (prevalence of past year suicidal ideation 3.5
Objectives. To examine the association of abortion bans with changes in maternal, pregnancy-related, and pregnancy-associated mortality. Methods. Using national vital statistics data (2016-2023), we used a Bayesian panel model to examine maternal, pregnancy-related, and pregnancy-associated mortality in 14 US states that implemented abortion bans by the end of 2022. Models accounted for temporal trends and state-specific factors. Results. Among the 14 states with abortion bans, there was some evidence of a potential 9.2% (95% credible interval [CI] = -1.6, 20.7) increase in the number of pregnancy-associated deaths above expectation, equivalent to 68 (95% CI = -13, 147) excess deaths; the rate was also higher than expected, with 3.3 (95% CI = -2.6, 9.0) additional deaths per 100 000 live births. Relative changes in pregnancy-related mortality were similar in magnitude but had greater uncertainty. There was no detectable increase in maternal mortality. Conclusions. Abortion bans may be associated with an increase in pregnancy-associated and pregnancy-related mortality, although data limitations and chance variation in these rare outcomes constrain the certainty of these and other findings. (Am J Public Health. 2026;116(6):819-828. https://doi.org/10.2105/AJPH.2026.308465).
This cross-sectional study examines if intersectionality based on race and low-income status is related to food insecurity rates and the role that the Supplemental Nutrition Assistance Program plays.
State-level policy studies often conduct heterogeneity analyses that quantify how treatment effects vary across state characteristics. These analyses may be used to inform state-specific policy decisions, or to infer how the effect of a policy changes in combination with other state characteristics. However, in state-level settings with varied contexts and policy landscapes, multiple versions of similar policies, and differential policy implementation, the causal quantities targeted by these analyses may not align with the inferential goals. This paper clarifies these issues by distinguishing several causal estimands relevant to heterogeneity analyses in state-policy settings, including state-specific treatment effects (ITE), conditional average treatment effects (CATE), and controlled direct effects (CDE). We argue that the CATE is often the easiest to identify and estimate, but may not be the most policy relevant target of inference. Moreover, the widespread practice of coarsening distinct policies or implementations into a single indicator further complicates the interpretation of these analyses. Motivated by these limitations, we propose bounding ITEs as an alternative inferential goal, yielding ranges for each state's policy effect under explicit assumptions that quantify deviations from the ideal identifying conditions. These bounds target a well-defined and policy-relevant quantity, the effect for specific states. We develop this approach within a difference-in-differences framework and discuss how sensitivity parameters may be informed using pre-treatment data. Through simulations we demonstrate that bounding state-specific effects can more reliably determine the sign of the ITEs than CATE estimates. We then illustrate this method to examine the effect of the Affordable Care Act Medicaid expansion on high-volume buprenorphine prescribing.
Dramatic changes in the US abortion policy landscape have led to growing interest in studying the health and social impacts of abortion bans. Many studies of population-level impacts necessarily rely on panel designs using aggregate state-level data to strengthen causal inference, yet such analyses risk pitfalls if they apply generic evaluation frameworks that overlook the complexity of the US abortion context and relevant outcomes. This commentary provides practical guidance for researchers engaged in panel studies of abortion policy, as well as for peer reviewers who may be less familiar with the methodological and substantive considerations in this area. Drawing from recent work, we highlight abortion-specific challenges that require attention, including time-varying confounding and violation of parallel trends, COVID-era disruptions, data suppression, spillover effects, and subgroup heterogeneity. We further recommend assessing sensitivity to including Texas, given its earlier implementation of abortion restrictions and potential outsized influence on results. Ultimately, we emphasize that rigorous evaluation of abortion policies requires thoughtful study design, context-specific considerations, and collaboration between methodologists and subject-matter experts.
Objective Intimate partner violence (IPV) affects an estimated 47% of women living in the USA in their lifetime and is associated with increased risk of physical and mental health concerns. Current prevention efforts focus on individual and family-level interventions rather than macrosystem-level policies. Thus, we sought to test the effects of Medicaid expansion on the rates of IPV and violence more broadly.Methods Present analyses use retrospective longitudinal data from the National Crime Victimization Survey (NCVS). State level rates of total violence and IPV were measured per 1000 population from the NCVS for years 2008-2018 as 3-year averages for each state. A two-way fixed-effects difference-in-differences model was fit to evaluate differences in the change in violence outcomes pre-2014/post-2014 in Medicaid expansion states versus non-expansion states.Results Comparison states had a significantly higher proportion of residents who were black, living below the federal poverty level and with lower educational attainment. Before Medicaid expansion, comparison states had a significantly lower mean rate of total violence and IPV per 1000 population. In two-way fixed effects difference-in-differences models, there was no statistically significant association between Medicaid expansion and IPV or total violence.Discussion Despite null findings, our study adds to the evidence base evaluating the impacts of macro-level policies on different forms of violence. The pathways by which Medicaid expansion could contribute to violence reduction are multifaceted with numerous mediators and those pathways may not be sufficiently strong to generate impacts. Additional work is warranted to further probe Medicaid expansion's impact on violence prevention.