BACKGROUND:The association between wildfire smoke (WFS) exposure and pregnancy loss has been understudied. Here, we examined the association between prenatal wildfire-specific particulate matter ≤2.5 µm (PM2.5) exposure and pregnancy loss in Colorado, USA. METHODS:We retrieved all birth records from the 17 'Front Range' counties (just east of the Rocky Mountains) of Colorado from 2007 to 2018 (n = 614 321). We considered two PM2.5 exposures-wildfire-specific PM2.5 from a novel machine learning model and non-wildfire PM2.5 constructed using the Community Multiscale Air Quality model. We fitted quasi-Poisson distributed lag models to estimate the associations between the two weekly-resolved PM2.5 exposures during pregnancy and live birth-identified conceptions (LBICs) in each county. That is, we used the predicted change in the LBICs to directly infer the change in the number of pregnancy losses due to the exposure. RESULTS:Average weekly non-wildfire PM2.5 was 6.2 µg/m3 (SD 2.3). In weeks with non-zero WFS (27% of all county-weeks), the average wildfire-specific PM2.5 was 0.92 µg/m3 (SD: 1.55). Wildfire-specific PM2.5 appeared important in gestational weeks 6-13-a 1-µg/m3 higher exposure sustained in these gestational weeks was associated with 20 [95% confidence interval (CI): 4-34] losses/year. In contrast, the cumulative association with non-wildfire PM2.5 was stronger-a 1-µg/m3 higher exposure sustained in every week of pregnancy was associated with 84 (95% CI: 46-129) losses/year. CONCLUSION:Our findings suggest that both wildfire-specific and non-wildfire PM2.5 exposures were associated with more pregnancy loss and add to the growing literature on the harmful effects of wildfires and, more broadly, air pollution.
The infant microbiome undergoes rapid changes in composition over time and is associated with long-term risks of conditions such as immune strength, allergy, asthma, and other health outcomes. Modeling the associations between exposures or treatments and microbial composition over time is essential for understanding the factors that drive these changes. Estimating these temporal dynamics has several challenges including repeated measures, overdispersion, compositionality, high-dimensional parameter spaces, and zero-inflation. Many longitudinal regression models used in human microbiome research assume constant effects over time that cannot capture time-varying or functional effects of exposures, ignore the compositional structure of the data by modeling each taxon separately, and are not equipped to handle potential zero-inflation. Dirichlet-multinomial (DM) regression models inherently accommodate overdispersion and the compositional structure of the data and have been extended to account for excess zeros. However, existing DM-based regression models are unable to additionally handle repeated measures designs. To fill this gap, we propose a functional concurrent zero-inflated Dirichlet-multinomial regression model which is designed to model time-varying relations between observed covariates and microbial taxa while accounting for zero-inflation, compositionality, and repeated measures. Through simulation, we demonstrate that the model can accurately estimate the underlying functional relations and scale to large compositional spaces. We apply our model to investigate time-varying associations between infant microbiome composition and observed covariates during the 11-wk postnatal period. We found that $ \boldsymbol{\alpha} $-diversity (ie the diversity of the microbiome within an individual) is positively associated with a higher gestational age and percentage of breast milk in the diet. We provide an accompanying R package and shiny app to implement the method and generate plots.
BACKGROUND:Distributed lag models (DLMs) are widely used in perinatal epidemiology to identify critical windows during pregnancy in which environmental exposures influence pregnancy/birth outcomes. A well-known complication of fitting DLMs in this context is that they require the same length of exposure history for every individual even though not all pregnancies are of the same duration. This misalignment often leads researchers to artificially extend exposure histories to a fixed length and fill post-birth weeks with zeroes (i.e. "zero-filling"). Despite its widespread use, the implications of zero-filling have not been formally evaluated. METHODS:We demonstrate conceptually that zero-filling induces a spurious association between gestational age and late-pregnancy exposures, thus introducing confounding that was otherwise not present in the observed data. We then conducted a simulation study and a real data application using air pollution and birth weight data from a Colorado-based cohort to compare zero-filling with alternative approaches to handle the misalignment between exposure window and gestation. RESULTS:In our simulations, we found that zero-filling produced the largest bias, poorest coverage, and highest root mean squared error. In the analysis of the Colorado birth data, zero-filling produced implausibly strong associations. Adjusting for gestational age attenuated this bias. Alternative approaches of carrying forward the last pre-birth exposure value, using observed post-birth exposures, and, in some situations, truncation at 37 weeks eliminate this bias. CONCLUSION:Zero-filling can cause bias in the estimated associations when using distributed lag models.
BACKGROUND:Exposure to particulate matter, hazardous gases, and noise among workers causes millions of deaths, injuries, and disabilities annually. These exposures are not measured frequently, and most workers are never monitored because resources for monitoring are limited globally. Even in high-income countries, just two or three samples per workplace inspection are collected on average. Considerably larger sample sizes are required to estimate measures of central tendency and upper percentiles of common exposure distributions. Therefore, interventions and epidemiological studies often rely on poor estimates of those parameters. OBJECTIVE:The objective of this study was to demonstrate that a small team can cost-effectively measure all workers' exposures to multiple hazards within a facility (~100 workers) in a single day while minimizing participant burden and loss of productivity. METHODS:We deployed novel, compact personal monitors-called AirPens-to measure exposures to total particulates, formaldehyde, and A-weighted noise among workers at a furniture manufacturing facility during one full work shift. AirPens were preloaded with sampling media and preprogrammed to minimize field deployment time. On-board sensor data were used to identify sampling anomalies, confirm that AirPens were worn, and validate sampling results. Valid samples were then used to estimate exposure distribution parameters. RESULTS:A team of five people deployed 83 AirPens at the facility. After quality assurance screening, 67 PM, 67 noise, and 22 formaldehyde samples remained valid (or 72, 67, and 31, respectively if samples below the limit of detection are counted as valid). PM and noise exposure distribution parameters (e.g., arithmetic and geometric mean, 95th percentile) were estimated with high certainty. Uncertainty (i.e., confidence intervals) grew several fold when smaller sample sizes were analyzed. SIGNIFICANCE:Use of streamlined multi-hazard monitoring technology enables dramatically higher sample throughput for workplace exposure assessment. This approach reduces the uncertainty of workplace risk assessment and facilitates more in-depth analyses of personal exposure. IMPACT STATEMENT:Comprehensive monitoring of personal exposure to air pollution is lacking due to technological and logistical limitations inherent to established monitoring methods. This lack of monitoring limits occupational health practice and epidemiologic research, which rely on sufficient sample sizes to support expert judgments or statistical inferences. This work demonstrates a new wearable sampler for particle, gas, and noise hazards, designed to overcome typical barriers to personal exposure assessment (e.g., cost, time, participant burden). Results show that this new sampling approach can produce measurements of similar quality, but on a much larger scale, compared to established technology.
Residents of agricultural communities may experience higher exposures to pesticides due to their proximity to agricultural operations. We applied a novel measurement approach, using Ultrasonic Personal Air Samplers (UPAS), to quantify particulate matter and organophosphate pesticides in air in California's Central Valley. We collected 124 personal, 126 in-home, and 32 outdoor air samples with 66 adults from 37 rural households in 2023 and 2024. We detected chlorpyrifos, acephate, malathion, diazinon, and naled in air samples. We detected gas-phase chlorpyrifos in 63% of personal samples and 86% of homeseven though use of chlorpyrifos has been banned in California (with few exceptions) since January 2021at 24 h average concentrations ranging up to 13 ng m-3 (personal) and 5.8 ng m-3 (in-home). We did not detect chlorpyrifos in outdoor air samples. Using linear mixed models, we found that higher indoor air temperatures and having more carpets/rugs were associated with higher indoor chlorpyrifos concentrations. The concentrations we measured were well below the California Department of Pesticide Regulation's health screening level of 510 ng m-3 for chronic exposure to chlorpyrifos in air; nevertheless, our results suggest that persistent chlorpyrifos in home environments continues to contribute to nondietary exposure among California residents.
Abstract Precision environmental health seeks to estimate how the effects of the environment vary across the population to inform targeted interventions and public health policy. However, there is a lack of statistical methods to estimate the heterogeneous effects of environmental exposures, particularly mixture exposures that are assessed longitudinally. We examine the heterogeneous exposure effect of weekly average fine particulate matter (PM2.5) and maximal daily temperature during gestation on birth weight using birth registry data in Colorado. We develop a Bayesian additive model represented by an ensemble of tree triplets where a tree triplet consists of two types of binary trees, interacting to model heterogeneous time-structured exposure effects. Our framework provides a tool to estimate individualized and subgroup-specific distributed lag effects of longitudinally assessed mixture exposures. Our method can accommodate a high-dimensional set of candidate modifiers with modifier selection and allows for mixture exposures with time-sensitive interactions. Through simulation, we demonstrate that our model can estimate individualized exposure effects and identify important mixture components and modifying factors. From the Colorado birth registry data, we find evidence of an association between PM2.5 and birth weight with larger effects among mothers younger than 25 years with low income.
There is substantial interest in estimating the health effects of exposure to environmental mixtures. Bayesian kernel machine regression (BKMR) has emerged as a popular tool for mixture analyses. The health effects of environmental exposures, including mixture exposures, often differ among subpopulations. However, there is little guidance on how to assess such heterogeneity for mixture effects. We provide tools and guidance to conduct BKMR analyses with effect modification, including estimating group-specific effects and between-group differences in effects. We propose a new group-separable BKMR variant for mixture analyses with effect modification by a categorical variable. We compare this new method to a stratified analysis and to a model that includes the categorical modifier directly in the BKMR kernel function in both a simulation study and the analysis of a metals mixture on children’s neurodevelopment with child sex as a binary modifier in a rural Bangladesh cohort. Both stratified BKMR and the new group-separable BKMR have the flexibility to capture interactions and estimate between-group differences. The group-separable BKMR has lower variance compared to stratified BKMR, particularly when there are small subgroup sizes. We provide code and data to implement the methods and reproduce simulations and analyses.
An important goal of environmental epidemiology is to quantify the complex health effects posed by a wide array of environmental exposures. In studies of a small number of exposures, flexible models like Bayesian kernel machine regression (BKMR) are appealing because they allow for non-linear and non-additive associations among exposures. However, this flexibility comes at the cost of low power and difficult interpretation, particularly in exposomic analyses when the number of exposures is large. We propose a flexible framework that allows for the separate selection of additive and non-additive effects, unifying additive models and kernel machine regression. The proposed approach yields increased power and simpler interpretation when there is little evidence of interaction. Further, it allows users to specify separate priors for additive and non-additive effect s, and allows for statistical inference on non-additive interactions. We extend the approach to a class of multiple index models, in which the special case of kernel machine-distributed lag models is nested. We apply the method to motivating data from a subcohort of the Human Early Life Exposome (HELIX) study containing 65 mixture components grouped into 13 distinct exposure classes.
Studies document independent effects of prenatal air pollution exposure and social environmental factors, including neighborhood safety, on childhood asthma development with documented sex-specific effects. Further research examining these factors jointly is needed. We examined associations between prenatal residence-level daily fine particulate matter (PM2.5) exposure and child asthma, considering effect modification by a validated Neighborhood Sentiment and Safety Index (NSSI) and child sex. Participants were mothers and full-term (> 37 weeks gestation) singleton-born children from two Boston-area pregnancy cohorts. The Asthma Coalition on Community, Environment and Social Stress (ACCESS) project enrolled 955 pregnant women between August 2002 and July 2009.The Programming of Intergenerational Stress Mechanisms (PRISM) study recruited 390 pregnant women from March 2011 to December 2013. Bayesian distributed lag interaction models (BDLIMs) were implemented to estimate associations between child asthma incidence and daily average maternal PM2.5 exposure across gestation. Effect modification by NSSI and child sex was examined using a BDLIM comparing models with and without effect modification. Women were primarily minorities (29% black, 47% Hispanic) reporting less than a high school education (54%). Children were followed 15.1 ± 3 years; 204 (17%) developed asthma. In the overall sample (n = 1,178), increased PM2.5 exposure between 21 and 27 weeks gestation was associated with increased odds of asthma in children born to women in the high NSSI group (NSSI ≥ 75th percentile representing safer neighborhoods). Both boys and girls were at higher risk of asthma when considering joint effects. These data add to a growing literature highlighting the need to consider both chemical toxins and psychosocial factors operating in communities to better elucidate which factors are driving respiratory health effects. A singular focus on changes to mitigate air pollution may not have high impact on improving respiratory health, particularly in historically under-resourced areas where effects may be more highly driven by social determinants.
When examining the relationship between an exposure and an outcome, there is often a time lag between exposure and the observed effect on the outcome. A common statistical approach for estimating the relationship between the outcome and lagged measurements of exposure is a distributed lag model (DLM). Because repeated measurements are often autocorrelated, the lagged effects are typically constrained to vary smoothly over time. A recent statistical development on the smoothing constraint is a tree structured DLM framework. We present an R package dlmtree, available on CRAN, that integrates tree structured DLM and extensions into a comprehensive software package with user-friendly implementation. A conceptual background on tree structured DLMs and demonstration of the fitting process of each model using simulated data are provided. We also demonstrate inference and interpretation using the fitted models, including summary and visualization. Additionally, a built-in shiny app for heterogeneity analysis is included.
Pregnancy is a critical window for long-term metabolic programming of fetal effects stemming from airborne particulate matter <= 2.5 mu m (PM2.5) exposure. Yet, little is known about long-term metabolic effects of PM2.5 exposure during and surrounding pregnancy in mothers. We assessed potential critical windows of PM2.5 exposure during and surrounding pregnancy with maternal adiposity and lipid measures later in life. We included 517 pregnant women from the PROGRESS cohort with adiposity [body mass index (BMI), waist circumference (WC), % body fat] and lipids [total cholesterol, high-density-lipoprotein (HDL), low-density- lipoprotein (LDL)] measured repeatedly at 4, 6 and 8 years post-delivery. Monthly average PM2.5 exposure was estimated at each participant's address using a validated spatiotemporal model. We employed distributed lag interaction models (DLIMs) adjusting for socio-demographics and clinical covariates. We found that a 1 mu g/m3 increase in PM2.5 exposure throughout mid-/late-pregnancy was associated with higher WC at 6-years post- delivery, peaking at 6 months of gestation: 0.04 cm (95%CI: 0.01, 0.06). We also identified critical windows of PM2.5 exposure during and surrounding pregnancy associated with higher LDL and lower HDL both measured at 4 years post-delivery with peaks at pre-conception for LDL [0.17 mg/dL (95%CI: 0.00, 0.34)] and at the 11th month after conception for HDL [-0.07 mg/dL (95%CI:-0.11,-0.02)]. Stratified analyses by fetal sex indicated stronger associations with adiposity measures in mothers carrying a male, while with lipids in mothers carrying a female fetus. Stratified analyses also indicated potential stronger deleterious lagged effects in women with folic acid intake lower than 600mcg/day during pregnancy.
In environmental health research there is often interest in the effect of an exposure on a health outcome assessed on the same day and several subsequent days or lags. Distributed lag nonlinear models (DLNM) are a well-established statistical framework for estimating an exposure-lag-response function. We propose methods to allow for prior information to be incorporated into DLNMs. First, we impose a monotonicity constraint in the exposure-response at lagged time periods which matches with knowledge on how biological mechanisms respond to increased levels of exposures. Second, we introduce variable selection into the DLNM to identify lagged periods of susceptibility with respect to the outcome of interest. The variable selection approach allows for direct application of informative priors on which lags have nonzero association with the outcome. We propose a tree-of-trees model that uses two layers of trees: one for splitting the exposure time frame and one for fitting exposure-response functions over different time periods. We introduce a zero-inflated alternative to the tree splitting prior in Bayesian additive regression trees to allow for lag selection and the addition of informative priors. We develop a computational approach for efficient posterior sampling and perform a comprehensive simulation study to compare our method to existing DLNM approaches. We apply our method to estimate time-lagged extreme temperature relationships with mortality during summer or winter in Chicago, IL.
Importance:Prior studies report negative associations between prenatal exposure to fine particulate matter (ie, aerodynamic diameter <2.5 µg; PM2.5) and birth weight, but have typically averaged exposure across pregnancy, which may not reveal windows of susceptibility. Objective:To identify windows of prenatal susceptibility to PM2.5. Design, Setting, and Participants:This was a retrospective analysis of a prospectively enrolled cohort study. Participants were enrolled at 1 of 50 sites participating in the US Environmental Influences on Child Health Outcomes Cohort. The study included full-term, singleton births occurring between September 2003 and December 2021. Statistical analyses were conducted from March 2024 to February 2025. Exposures:Daily residential PM2.5 exposure was estimated using a machine-learning model covering the contiguous US and mean exposure estimates were calculated for each week of pregnancy. Main Outcomes and Measures:Bayesian distributed lag interaction models were used to examine cumulative and week-specific associations between PM2.5 exposure and birth weight for gestational age (BWGA) z scores. Interactions with sex, race and ethnicity, and region were also examined. Results:The sample of 16 868 mother-newborn pairs (maternal mean [SD] age, 30.4 [5.5] years; 605 [3.6%] Asian, 2197 [13.0%] Black or Black-Hispanic, 3407 [20.2%] Hispanic, 9251 [54.8%] non-Hispanic White, and 1408 [8.4%] other) included 15 806 unique mothers and 1062 mothers with 2 or more children in the study. Mean (SD) weekly PM2.5 exposure during pregnancy was relatively low, at 8.03 (2.3) µg/m3, and overall mean (SD) birth weight was 3410.7 (464.5) g. In the sample overall, there was a negative association between PM2.5 exposure and BWGA z score (β = -0.06; 95% credible interval [CrI], -0.10 to -0.03), with a critical window in early gestation (weeks 1-5) that persisted only among males (β = -0.06; 95% CrI, -0.10 to -0.02). When examining differences by region, there were negative associations in the Northeast (β = -0.09; 95% CrI, -0.15 to -0.03), Midwest (β = -0.11; 95% CrI, -0.17 to -0.05; critical window, 12-18 weeks), and South (β = -0.18; 95% CrI, -0.17 to -0.05; critical window, 3-9 weeks). Conclusions and Relevance:In this cohort study, higher PM2.5 exposure was associated with lower BWGA z score, with critical windows identified during early pregnancy to midpregnancy; however, findings varied by sex and region. Understanding windows of susceptibility to environmental exposures can help guide research on underlying biological processes and can inform strategies for limiting exposure during certain periods of pregnancy.
Epidemiological evidence supports an association between exposure to air pollution during pregnancy and birth and child health outcomes. Typically, such associations are estimated by regressing an outcome on daily or weekly measures of exposure during pregnancy using a distributed lag model. However, these associations may be modified by multiple factors. We propose a distributed lag interaction model with index modification that allows for effect modification of a functional predictor by a weighted average of multiple modifiers. Our model allows for simultaneous estimation of modifier index weights and the exposure-time-response function via a spline cross-basis in a Bayesian hierarchical framework. Through simulations, we showed that our model out-performs competing methods when there are multiple modifiers of unknown importance. We applied our proposed method to a Colorado birth cohort to estimate the association between birth weight and air pollution modified by a neighborhood-vulnerability index and to a Mexican birth cohort to estimate the association between birthing-parent cardio-metabolic endpoints and air pollution modified by a birthing-parent lifetime stress index.
Exposure to environmental pollutants during the gestational period can significantly impact infant health outcomes, such as birth weight and neurological development. Identifying critical windows of susceptibility, which are specific periods during pregnancy when exposure has the most profound effects, is essential for developing targeted interventions. Distributed lag models (DLMs) are widely used in environmental epidemiology to analyze the temporal patterns of exposure and their impact on health outcomes. However, traditional DLMs focus on modeling the conditional mean, which may fail to capture heterogeneity in the relationship between predictors and the outcome. Moreover, when modeling the distribution of health outcomes like gestational birth weight, it is the extreme quantiles that are of most clinical relevance. We introduce 2 new quantile distributed lag model (QDLM) estimators designed to address the limitations of existing methods by leveraging smoothness and shape constraints, such as unimodality and concavity, to enhance interpretability and efficiency. We apply our QDLM estimators to the Colorado birth cohort data, demonstrating their effectiveness in identifying critical windows of susceptibility and informing public health interventions.
Identifying the determinants of pregnancy loss (PL) is a critical public health concern. However, PL is often not noticed, and even when it is, it is inconsistently recorded. Thus, past studies have been limited to medically identified losses or small, highly selected cohorts, which can lead to biased or nongeneralizable results. We show mathematically and through simulations a novel approach that overcomes this measurement challenge to infer effects about PL by using more available data: the number of conceptions that led to live births (ie, live-birth-identified conceptions [LBICs]). We simulated 10 years of conceptions, pregnancies, losses, and births under several confounding patterns, and 2 nitrogen dioxide (NO2)-PL relationships (no effect, mid-gestation effect). We fitted distributed lag models adjusted for season, year, and temperature, and assessed model performance through bias and coverage. Our simulations showed that our models, across all scenarios, identified the 2 NO2-PL relationships with appropriate coverage (>90% of CIs captured the true effect) and low bias (never exceeded ±2%). In an applied example using NO2-a traffic emissions tracer-and live-birth data from a large tertiary-care hospital in Massachusetts, we found that higher prenatal NO2 exposure was associated with more PLs. Our proposed approach based on LBICs provides an alternative way to study causes of PL.
The COVID-19 infection fatality rate (IFR) is the proportion of individuals infected with SARS-CoV-2 who subsequently die. As COVID-19 disproportionately affects older individuals, age-specific IFR estimates are imperative to facilitate comparisons of the impact of COVID-19 between locations and prioritize distribution of scare resources. However, there lacks a coherent method to synthesize available data to create estimates of IFR and seroprevalence that vary continuously with age and adequately reflect uncertainties inherent in the underlying data. In this paper we introduce a novel Bayesian hierarchical model to estimate IFR as a continuous function of age that acknowledges heterogeneity in population age structure across locations and accounts for uncertainty in the estimates due to seroprevalence sampling variability and the imperfect serology test assays. Our approach simultaneously models test assay characteristic, serology, and death data, where the serology and death data are often available only for binned age groups. Information is shared across locations through hierarchical modeling to improve estimation of the parameters with limited data. Modeling data from 26 developing country locations during the first year of the COVID-19 pandemic, we found seroprevalence did not change dramatically with age, and the IFR at age 60 was above the high-income country benchmark for most locations.
Introduction:Neurotoxicity resulting from air pollution is of increasing concern. Considering exposure timing effects on neurodevelopmental impairments may be as important as the exposure dose. We used distributed lag regression to determine the sensitive windows of prenatal exposure to fine particulate matter (PM2.5) on children's cognition in a birth cohort in Mexico.Methods:Analysis included 553 full-term (>= 37 weeks gestation) children. Prenatal daily PM2.5 exposure was estimated using a validated satellite-based spatiotemporal model. McCarthy Scales of Children's Abilities (MSCA) were used to assess children's cognitive function at 4-5 years old (lower scores indicate poorer performance). To identify susceptibility windows, we used Bayesian distributed lag interaction models to examine associations between prenatal PM2.5 levels and MSCA. This allowed us to estimate vulnerable windows while testing for effect modification.Results:After adjusting for maternal age, socioeconomic status, child age, and sex, Bayesian distributed lag interaction models showed significant associations between increased PM2.5 levels and decreased general cognitive index scores at 31-35 gestation weeks, decreased quantitative scale scores at 30-36 weeks, decreased motor scale scores at 30-36 weeks, and decreased verbal scale scores at 37-38 weeks. Estimated cumulative effects (CE) of PM2.5 across pregnancy showed significant associations with general cognitive index (CE<^> = -0.35, 95% confidence interval [CI] = -0.68, -0.01), quantitative scale (CE<^> = -0.27, 95% CI = -0.74, -0.02), motor scale (CE<^> = -0.25, 95% CI = -0.44, -0.05), and verbal scale (CE<^> = -0.2, 95% CI = -0.43, -0.02). No significant sex interactions were observed.Conclusions:Prenatal exposure to PM2.5, particularly late pregnancy, was inversely associated with subscales of MSCA. Using data-driven methods to identify sensitive window may provide insight into the mechanisms of neurodevelopmental impairment due to pollution.