BACKGROUND:Drinking water can be an important source of exposure to nitrate and disinfection by-products, including trihalomethanes (THMs) and haloacetic acids (HAAs). N-nitroso compounds formed endogenously after nitrate ingestion are animal carcinogens, and THM and HAA exposures increase the risk of some cancers. Our objectives were to evaluate associations of drinking water nitrate and disinfection byproducts with total and aggressive (distant stage, poorly differentiated grade, fatal, or Gleason score ≥7) prostate cancer in the Agricultural Health Study cohort. METHODS:Male participants who were cancer free and used private wells or public water supplies (PWS) for drinking water at enrollment (1993-1997, n = 40 403) were followed through 2021 (mean = 21.9 years). Average nitrate-nitrogen (nitrate-N) concentrations were estimated for private well users based on state-specific geologic and meteorologic factors. We used monitoring data to compute average nitrate-N, THMs, and HAAs for PWS users. We estimated hazard ratios (HRs, 95% CIs) per doubling and categories of exposure for total (n = 3625) and aggressive (n = 2200) prostate cancer using Cox proportional hazards regression. RESULTS:Median (interquartile range) average water nitrate-N was 1.49 (0.76-3.01) mg L-1; 6% >10 mg L-1 (PWS maximum contaminant level). Compared to nitrate-N ≤ 1 mg L-1, exposures >10 mg L-1 were significantly positively associated with total (1.16, 1.01-1.35; P = .10 for trend) and aggressive disease (1.22, 1.02-1.47; P = .03 for trend). We observed weak associations between higher nitrate-N (Q4 vs Q1) and total (1.05, 0.95-1.16) and aggressive (1.13, 0.99-1.27) disease. We did not observe associations with total THMs or HAAs. CONCLUSIONS:These findings suggest that drinking water nitrate-N exposure, at average levels > 10 mg L-1, is a risk factor for prostate cancer, particularly aggressive disease.
BACKGROUND:Epidemiologic studies of the health impact of alcohol consumption have mostly been based on self-reported measures of intake. Objective markers may provide better measures of alcohol intake and its biologic effects, potentially elucidating mechanisms of disease. However, there are currently no validated biomarkers to assess low to moderate drinking. OBJECTIVES:To apply semi-targeted metabolomics to identify biomarkers of low to moderate alcohol consumption using plasma samples collected in the Postmenopausal Women's Alcohol Study (WAS), a randomized controlled crossover feeding study. METHODS:In the WAS, postmenopausal women (n = 51) were randomly assigned to consume 0, 15, or 30 g of alcohol/d (equivalent to 0, 1, or 2 drinks/d) for 8 wk each as part of a controlled diet, with washout periods between treatments. Metabolites were measured in baseline, washout, and post-treatment fasting plasma samples. Linear mixed-effects models were used to identify metabolites significantly altered by alcohol intake. RESULTS:A total of 1422 metabolites were measured, of which 150 were previously reported to be correlated with alcohol in observational studies. Alcohol intake significantly altered plasma levels of 46 metabolites, including xenobiotics directly related to alcohol, ethyl glucuronide and ethyl α-glucopyranoside; α-hydroxyisovalerate; 2-aminobutyrate; androgenic steroids; and multiple phosphatidylcholines. Top-ranking metabolites displayed clear dose-response relationships with alcohol dose-most notably, ethyl α-glucopyranoside, which showed a strong positive relationship (461% and 900% change for 15 and 30 g of alcohol/d, respectively, compared with no alcohol). We replicated associations for 14 alcohol-related metabolites identified in previous studies and discovered a number of new potential biomarkers. CONCLUSIONS:In a tightly controlled feeding study, consumption of 1 or 2 alcoholic drinks/d changed plasma levels of 46 metabolites, suggesting their utility as biomarkers of low to moderate alcohol consumption, with opportunities to conduct etiologic research in cohorts with metabolomics data.
Esophageal squamous cell carcinoma (ESCC) remains a leading cause of cancer-related mortality worldwide, and there are limited molecular screening options. Following a prior discovery and validation studies in subjects from the United States, China, and Iran, we aimed to validate methylated DNA markers (MDM) for ESCC detection in Malawi, a country with one of the highest ESCC incidence. Fourteen MDMs were tested on 177 tissue samples. Ten MDMs showed excellent performance (AUCs ≥ 0.85); a random forest model achieved an AUC of 0.97. The consistent cross-population performance suggests potential for universal application of these MDMs and further justifies exploring nonendoscopic sampling methods, such as swallowed cell collection devices in combination with MDMs, for efficient and cost-effective screening in high-incidence populations. PREVENTION RELEVANCE:This study validates MDMs for early detection of ESCC in Malawi, supporting their potential role in noninvasive, population-based screening strategies in Africa to enable early detection and timely intervention for cancer interception, potentially reducing cancer burden in high-incidence, resource-limited regions.
Abstract Characterizing the variation in clonal trajectory of age-related mosaic chromosomal alterations (mCAs) in circulating leukocytes can identify etiologic factors influencing clonal expansion and hematologic malignancy risk. Here we scan whole blood-derived DNA of 56,324 participants from the Prostate, Lung, Colorectal, Ovarian Cancer Screening Study for mCAs and characterize 1746 longitudinal mCAs. Cross-sectional mCA clonal fraction only moderately correlates with clonal expansion rate, highlighting the utility of serial sampling for evaluating clonal trajectory. The strongest contributor to clonal expansion is genomic location of an mCA with participant age, smoking status, and co-occurring mCA types also modifying expansion rates. Germline susceptibility to mosaic Y or X loss is not associated with expansion rate, supporting independent mechanisms governing generation and expansion. Autosomal mCAs previously associated with hematologic malignancies have the highest rates of expansion, especially myeloid malignancy-associated mCAs, underscoring the importance of tracing mCA clonal dynamics for evaluating hematologic cancer risk.
Abstract: Modeling the effect of smoking on the risk of lung cancer in cohort studies requires specification of time-dependent exposure variables and appropriate handling of confounding. This can be challenging as data on both exposure and confounders are typically incompletely observed. We discuss population exposure and disease processes and review modeling challenges when exposures are dynamic and challenging to summarize. Biases arising from failure to address time-dependent confounding are highlighted, with particular reference to a paradox related to the effect of quitting on risk of lung cancer. We stress the utility of comprehensive joint models.
Pooling genome-wide association studies of multiple related traits can substantially increase power for detecting genetic variants with pleiotropic effects. ASSET, which exhaustively searches all subsets of studies for association signals, has been widely used to detect modest effects and improve interpretability. Under a normality assumption, ASSET computes p-values via an analytic approximation that accounts for multiple testing. However, this approximation has been evaluated only in limited scenarios and for p-values no smaller than 10^-3. A systematic assessment in the extreme tail is therefore needed, yet naïve Monte Carlo methods would require prohibitively many simulations. We develop a computationally efficient importance-sampling (IS) algorithm that provides accurate ASSET p-value estimates for both independent and overlapping studies, achieving substantial efficiency gains over naïve Monte Carlo, particularly for very small p-values. Using IS, we show that ASSET's analytic approximation is highly accurate across nearly the entire p-value range when normality holds. In contrast, when normality is violated (due to small sample sizes, low-frequency variants, or non-normal traits), ASSET p-values can be inflated or deflated by orders of magnitude, whereas our IS approach remains accurate. We illustrate the method through applications to single-cell eQTL mapping using peripheral blood mononuclear cells from the OneK1K cohort and lung cells from a Korean population.
Longitudinal cancer biomarker studies aim to identify markers useful for long-term associations with cancer risk and early detection of cancer diagnosis. Early detection biomarkers show acute associations, meaning longitudinal trajectory changes sharply just before diagnosis, while long-term risk prediction biomarkers are often represented as changes in slopes or levels. We compare three common approaches to identify longitudinal biomarkers associated with survival outcomes: joint models, conditional models, and Cox models with time-varying covariates. Each of the three methods uses a different modeling framework for the joint density of the biomarkers and survival time. Thus, they have distinct advantages and disadvantages for detecting acute and long-term associations. We investigate the power and Type I error rates for the three approaches under different data-generating settings with regular yearly visits and moderate measurement error to evaluate robustness to assumptions and the power gain from using test statistics that match the data's association structure. The Cox model controlled Type I errors rates and maintained high power to detect associations across the scenarios considered. The conditional and joint models had inflated Type I errors under longitudinal model misspecification and only outperformed the Cox model for power when the correct model is used. However, only the conditional model can effectively disentangle the acute and long-term effects. The Cox model's convex likelihood also provides the fastest convergence. We apply all three approaches to a study of the association between CA-125 and ovarian cancer in the National Cancer Institute Prostate, Lung, Colorectal, and Ovarian (PLCO) Cancer Screening Trial. The results follow similar patterns to those of our simulations.
BACKGROUND/AIMS:Children are exposed to persistent organic pollutants (POPs) including polychlorinated biphenyls (PCBs) and organochlorine pesticides through in utero transfer from maternal serum. Few studies have evaluated the relationship between POPs and childhood leukemia risk. METHODS:The Finnish Maternity Cohort is a population-based pregnancy cohort with banked serum and linkages to cancer registry and other databases. We measured first-trimester levels of 25 POPs in maternal sera collected in 1986-2010 for 388 childhood acute lymphoblastic leukemia (ALL) cases (<15 years) and 388 controls matched on sample date, gestational age (≥37 weeks; ± 1 week), birth order (first/second), mother's age, and child's sex. We analyzed lipid-adjusted POPs as continuous (log2-transformed) and categorical exposures and estimated odds ratios (OR) and 95% confidence intervals (CI) using conditional logistic regression with adjustment for correlated POPs (ρ = 0.3-<0.90) in multivariable models. We evaluated joint effects of POPs with >80% quantifications using quantile-based g-computation. RESULTS:POPs detected in >80% of controls included trans-nonachlor (93%), p',p'-dichlorodiphenylethylene (p',p'-DDE; 100%), hexachlorobenzene (99%), β-hexachlorocyclohexane (β-HCH; 84%), and 10 PCBs (range: 84-100%). Increasing trans-nonachlor concentrations were associated with ALL (adjusted ORperlog2 = 1.48,CI = 1.11-1.97); risk was elevated among those in the 90th percentile (adjusted ORQ90vsQ1 = 2.32,CI = 0.98-5.45; p-trend = 0.05). PCB170, 180, 183, and total PCBs were associated with a 34-37% increased risk (e.g., PCB180 adjusted ORperlog2 1.37, CI = 1.05-1.78). β-HCH and p',p'-DDE were inversely associated with ALL. The mixture (hexachlorobenzene, β-HCH, trans-nonachlor, p,p'-DDE, total PCBs) was inversely associated with childhood ALL (ORperlog2 = 0.63, 95% CI 0.43-0.94). CONCLUSION:Our findings suggest that trans-nonachlor and some PCBs may increase childhood leukemia risk with unexplained inverse associations for β-HCH and p',p'-DDE.
Zero-inflated nonnegative continuous longitudinal data frequently arise in biomedical studies where outcomes consist of a mixture of excess zeros and positive continuous measurements. Two widely used approaches for analyzing such data are mixed-model versions of Tobit and the two-part hurdle models. The Tobit assumes a latent regression model that is censored below a specified threshold, while the hurdle separately models the continuous positive outcomes and the binary indicator of being positive. The choice between these models has rarely been systematically discussed, and inappropriate model choice may lead to biased estimation and misleading scientific interpretations. In this paper, we derive rigorous mathematical conditions under which the two models are equivalent and show that the Tobit can be viewed as a special case of the hurdle model when the link function for the binary process is probit. Based on simulation studies, we found that the hurdle is more flexible and robust than the Tobit model. On the other hand, if the assumptions of the Tobit are met, this model is easier to interpret since it does not require distinct inferences on both the continuous and binary processes. We therefore recommend that the Tobit model only be used when these assumptions are scientifically plausible and empirically supported; otherwise, the hurdle model is preferable. We applied both models to study the dynamics of somatic mosaicism using longitudinal clonal fraction measurements from the Prostate, Lung, Colorectal, and Ovarian study data while accounting for excess zero values. Estimates obtained from the Tobit and hurdle models were broadly consistent with those from a standard linear model that ignored zero inflation. These findings provide additional support for previously reported associations in studies of clonal hematopoiesis across different types of mosaic chromosomal alterations.
Serum prostate-specific antigen (PSA) is widely used for prostate cancer screening. While the genetics of PSA levels has been studied to enhance screening accuracy, the genetic basis of PSA velocity, the rate of PSA change over time, remains unclear. The Prostate, Lung, Colorectal, and Ovarian (PLCO) Cancer Screening Trial, a large, randomized study with longitudinal PSA data (15,260 cancer-free males, averaging 5.34 samples per subject) and genome-wide genotype data, provides a unique opportunity to estimate PSA velocity heritability. We developed a mixed model to jointly estimate heritability of PSA levels at age 54 and PSA velocity. To accommodate the large dataset, we implemented two efficient computational approaches: a partitioning and meta-analysis strategy using average information restricted maximum likelihood (AI-REML), and a fast restricted Haseman-Elston (REHE) regression method. Simulations showed that both methods yield unbiased estimates of both heritability metrics, with AI-REML providing smaller variability in the estimation of velocity heritability than REHE. Applying AI-REML to PLCO data, we estimated heritability at 0.32 (s.e. = 0.07) for baseline PSA and 0.45 (s.e. = 0.18) for PSA velocity. These findings reveal a substantial genetic contribution to PSA velocity, supporting future genome-wide studies to identify variants affecting PSA dynamics and improve PSA-based screening.
Supplementary Figure 2 shows associations of relative abundance of a priori-selected bacteria with adenomas
Supplemental Table 1 shows the stability of microbiome metrics across three study timepoints in the four-year Polyp Prevention Trial, 1991-1998
Reaching the national goal of reducing cancer mortality by 50% within 25 years will require improvements in cancer early detection and prevention in addition to treatment. Modern prospective cohorts that capture new and emerging exposures to research cancer etiology are critical to achieve these goals. Profound societal and technological changes in the last decade present opportunities and challenges for the recruitment, engagement, and retention of participants in new cohorts. The Connect for Cancer Prevention Study is a modern cohort of adults, 30-70 years old without a prior cancer diagnosis, recruited from 10 U.S. integrated health care systems. Participants provide information and biospecimens at enrollment and at regular intervals during at least 20 years of follow up. Exposure and outcome information will be captured through online surveys, electronic medical records, medical imaging, geospatial linkages of 20-year residential histories, wearable sensors, and linkages with the National Death Index and the state cancer registries. Collected biospecimens include blood, urine, saliva, and fecal samples, as well as precursor and tumor tissue specimens. Data systems follow F.A.I.R. (findability, accessibility, interoperability, and reuse of digital assets) principles to maximize data and tool re-usability. In 2021, recruitment commenced among an eligible catchment population of more than 3.2 million patients. As of November 2024, over 53, 000 of the expected 200, 000 participants consented; 87% completed at least some baseline activities to date. Baseline recruitment is expected to be complete in 2027. The current study population has a median age of 53 (IQR: 42-61) years, 67% are female, 63% completed at least a bachelor’s degree, and 88% report good to excellent overall health. Among the cohort, 61% of participants self-report as White, 14% self-report as multi-racial, 11% as Black, African American or African, 5% as Asian, and 4% as Hispanic, Latino, or Spanish among other categories. The portion of the latter category is expected to increase as recruitment ramps up at a site with a high Hispanic catchment population, which started enrollment in 2023. The proportion of individuals who have ever smoked cigarettes is 30%, whereas 9% have ever used marijuana, 3% have ever smoked cigars, and 2% have ever vaped nicotine containing electronic cigarettes. The median body mass index of participants is 30 kg/m2 (IQR: 24-36 kg/m2). Approximately, 30, 000 incident precursor and 8, 000 cancer diagnoses are expected to occur during the first 10 years of follow up. Connect for Cancer Prevention Study combines novel approaches in epidemiology and data science to provide a valuable resource for the scientific community to study cancer etiology, natural history, risk prediction, early detection, and survivorship. Individual-level data is expected to be released to the scientific community in 2026. Mia M. Gaudet, Amy Berrington de Gonzalez, Christian C. Abnet, Paul Albert, Jonas S. Almeida, Stephanie Weinstein, Amanda Black, Hannah P. Yang, Michelle Brotzman, Laura Beane-Freeman, Laura Beane-Freeman, Jonine Figueroa, Neal D. Freedman, Nicole M. Gerlanc, Gretchen L. Gierach, Rena R. Jones, Peter Kraft, Charles Matthews, Habibul Ashan, Briseis Aschebrook-Kilfoy, Chun-Hung Chan, Robert T. Greenlee, Stacey A. Honda, Benjamin Rybicki, Katherine Sanchez, Kevin Skyes, Mark A. Schmidt, Larissa L. White, Jeanette Y. Ziegenfuss, Stephen J. Chanock, Montse Garcia-Closas, Nicolas A. Wentzensen. Connect for Cancer Prevention Study: a modern prospective cohort [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2025; Part 1 (Regular Abstracts); 2025 Apr 25-30; Chicago, IL. Philadelphia (PA): AACR; Cancer Res 2025;85(8_Suppl_1):Abstract nr 7362.
Supplemental Table 2 shows the cross-sectional and prospective multivariablea associations of alpha diversity at three timepoints with adenoma recurrence in the four-year Polyp Prevention Trial, 1991-1998, excluding those with hyperplastic polyps from the control group
Carbaryl is a common carbamate insecticide in the United States (USA). Previous epidemiologic investigations, including within the Agricultural Health Study (AHS), have suggested potential associations between carbaryl use and cancer risk. The AHS is a prospective cohort study of licensed pesticide applicators in North Carolina (NC) and Iowa (IA), USA. Information on lifetime pesticide use was reported at enrollment (1993-1997) and follow-up (1999-2005). We evaluated cancer risks associated with ever- and intensity-weighted lifetime days (IWLD) of carbaryl use. Among 52,625 applicators, 8713 incident cancer cases were identified from linkages with state cancer registries through 2014 (NC) or 2017 (IA). We used Poisson regression to estimate rate ratios (RR) and 95 % confidence intervals (CI), controlling for confounders, and evaluated lagged exposures. Approximately 51 % of applicators reported using carbaryl. Increasing IWLD of carbaryl use was associated with increased stomach cancer risk (third tertile vs. never use; RRT3 = 2.07, 95 % CI: 1.05-4.07, p-trend = 0.02), persisting when exposure was lagged by 5-years (RRT3 = 2.20, 95 % CI: 1.12-4.33). We noted elevated risks of esophageal (RR = 1.52, 95 % CI: 1.01-2.27) and tongue (RR = 1.91, 95 % CI: 0.95-3.81) cancers with ever-use. There was an increased risk of aggressive prostate cancer when carbaryl exposure was lagged by 30 years (RRlag30Q4 = 1.56, 95 % CI: 1.18-2.08, p-trend = 0.002). This is the largest and most comprehensive prospective evaluation of carbaryl and cancer risk to date. We provide novel evidence of associations between carbaryl exposure and specific cancers. There is a need for additional studies to confirm these findings and to elucidate the biological mechanisms underlying the observed associations.
Supplemental Table 6 shows the exploratory analysis of cross-sectional and prospective multivariablea associations of select bacteria at three timepoints with adenoma recurrence in the four-year Polyp Prevention Trial, 1991-1998
Supplemental Table 3 shows the cross-sectional and prospective multivariablea associations of alpha diversity at three timepoints with recurrent adenoma/polyp characteristics in the four-year Polyp Prevention Trial, 1991-1998
Supplemental Table 4 shows the cross-sectional multivariablea associations of baseline and year-1 alpha diversity with adenoma/polyp characteristics in the four-year Polyp Prevention Trial, 1991-1998