Importance:Screening by low-dose computed tomography can reduce lung cancer mortality among high-risk individuals, but many lung cancers occur among individuals with a smoking history who are not eligible for screening. Objective:To develop and validate the protein-based Integrative Analysis of Lung Cancer Risk and Etiology (INTEGRAL)-Risk model in individuals with a smoking history from the general population. Design, Setting, and Participants:Cohorts in the Lung Cancer Cohort Consortium recruited research participants in the US, Europe, Asia, and Australia between 1985 and 2009, who were followed up for lung cancer and other health outcomes until 2021. Fourteen case cohorts of 3695 participants with a smoking history within the Lung Cancer Cohort Consortium, including 2305 randomly sampled participants and 1390 patients diagnosed with lung cancer within 3 years after blood sample collection, were designed. Plasma or serum samples from each participant were assayed using the INTEGRAL protein panel in 2022. The INTEGRAL-Risk model was trained using 7 predefined case cohorts (training set; n = 1951) to estimate absolute risk of being diagnosed with lung cancer based on age, smoking history, and 13 proteins. The validity of the INTEGRAL-Risk model was assessed in 7 independent case cohorts (testing set; n = 1744) at 1, 2, and 3 years after blood collection. Exposure:Absolute risk estimates from the protein-based INTEGRAL-Risk model. Main Outcomes and Measures:The primary outcome was the validity of the INTEGRAL-Risk model in the testing set with respect to discrimination (area under the curve [AUC]) and calibration (ratio of expected-to-observed cases [E/O]). Results:A total of 3695 participants were included, with 1951 participants (including 807 with lung cancer) in the training set and 1744 participants (including 583 with lung cancer) in the testing set. In the combined 14 training and testing sets, after application of statistical weights, 323 570 participants were represented (185 016 [57%] female; median [IQR] age, 60 [51-67] years). In the independent testing set, discrimination of the INTEGRAL-Risk model was highest at 1 year of follow-up and exceeded that of the questionnaire-based PLCOm2012 (Prostate, Lung, Colorectal and Ovarian Cancer Screening Trial) model (INTEGRAL-Risk AUC of 0.88 [95% CI, 0.85-0.91] vs PLCOm2012 AUC of 0.79 [95% CI, 0.75-0.83]; P value for difference <.001). Using a risk threshold to achieve the same specificity as US Preventive Services Task Force (USPSTF) 2021 criteria, the INTEGRAL-Risk model captured 85% of lung cancer cases compared with 63% by USPSTF 2021 and 70% by PLCOm2012. Discrimination of the INTEGRAL-Risk model decreased with longer prediction horizons, with a 2-year AUC of 0.84 (95% CI, 0.81-0.86) and 3-year AUC of 0.81 (95% CI, 0.79-0.83). The model was well calibrated (E/O over 3 years, 0.87 [95% CI, 0.69-1.14]). Conclusions and Relevance:Compared with questionnaire-based approaches, the protein-based INTEGRAL-Risk model improved short-term prediction of lung cancer in people with a smoking history. This model has potential to improve selection of high-risk individuals who are most likely to benefit from lung cancer screening.
Importance Screening by low-dose computed tomography can reduce lung cancer mortality among high-risk individuals, but many lung cancers occur among individuals with a smoking history who are not eligible for screening. Objective To develop and validate the protein-based Integrative Analysis of Lung Cancer Risk and Etiology (INTEGRAL)–Risk model in individuals with a smoking history from the general population. Design, Setting, and Participants Cohorts in the Lung Cancer Cohort Consortium recruited research participants in the US, Europe, Asia, and Australia between 1985 and 2009, who were followed up for lung cancer and other health outcomes until 2021. Fourteen case cohorts of 3695 participants with a smoking history within the Lung Cancer Cohort Consortium, including 2305 randomly sampled participants and 1390 patients diagnosed with lung cancer within 3 years after blood sample collection, were designed. Plasma or serum samples from each participant were assayed using the INTEGRAL protein panel in 2022. The INTEGRAL-Risk model was trained using 7 predefined case cohorts (training set; n = 1951) to estimate absolute risk of being diagnosed with lung cancer based on age, smoking history, and 13 proteins. The validity of the INTEGRAL-Risk model was assessed in 7 independent case cohorts (testing set; n = 1744) at 1, 2, and 3 years after blood collection. Exposure Absolute risk estimates from the protein-based INTEGRAL-Risk model. Main Outcomes and Measures The primary outcome was the validity of the INTEGRAL-Risk model in the testing set with respect to discrimination (area under the curve [AUC]) and calibration (ratio of expected-to-observed cases [E/O]). Results A total of 3695 participants were included, with 1951 participants (including 807 with lung cancer) in the training set and 1744 participants (including 583 with lung cancer) in the testing set. In the combined 14 training and testing sets, after application of statistical weights, 323 570 participants were represented (185 016 [57%] female; median [IQR] age, 60 [51-67] years). In the independent testing set, discrimination of the INTEGRAL-Risk model was highest at 1 year of follow-up and exceeded that of the questionnaire-based PLCOm2012 (Prostate, Lung, Colorectal and Ovarian Cancer Screening Trial) model (INTEGRAL-Risk AUC of 0.88 [95% CI, 0.85-0.91] vs PLCOm2012 AUC of 0.79 [95% CI, 0.75-0.83]; P value for difference <.001). Using a risk threshold to achieve the same specificity as US Preventive Services Task Force (USPSTF) 2021 criteria, the INTEGRAL-Risk model captured 85% of lung cancer cases compared with 63% by USPSTF 2021 and 70% by PLCOm2012. Discrimination of the INTEGRAL-Risk model decreased with longer prediction horizons, with a 2-year AUC of 0.84 (95% CI, 0.81-0.86) and 3-year AUC of 0.81 (95% CI, 0.79-0.83). The model was well calibrated (E/O over 3 years, 0.87 [95% CI, 0.69-1.14]). Conclusions and Relevance Compared with questionnaire-based approaches, the protein-based INTEGRAL-Risk model improved short-term prediction of lung cancer in people with a smoking history. This model has potential to improve selection of high-risk individuals who are most likely to benefit from lung cancer screening.
Previous research suggests higher risk of prostate cancer diagnosis in men with high vitamin D serology, but there are limited data related to prostate cancer survival. In this study, we examine the associations between prediagnostic circulating levels of 25-hydroxyvitamin D [25(OH)D] with risk of prostate cancer-specific mortality, as well as mortality from other causes, in participants diagnosed with prostate cancer. We conducted a prospective multi-national collaborative analysis of pre-diagnostic circulating 25(OH)D, the major biochemical indicator of vitamin D status, with data for 12,635 incident prostate cancer cases from 13 cohorts. Multivariable-adjusted Cox proportional hazards regression models estimated hazard ratios (HRs) and 95
The ABO locus is associated with pancreatic ductal adenocarcinoma (PDAC). Potential metabolic mechanisms underlying these associations have not been investigated. We examined associations between genotype-derived ABO blood group (rs505922 and rs8176746) and 1478 prediagnostic serum metabolites in 4042 participants from 8 nested case-control studies within the Prostate, Lung, Colorectal, and Ovarian Cancer Screening Trial and Alpha-Tocopherol, Beta-Carotene Cancer Prevention Study using linear regression and fixed-effect meta-analysis. We then examined associations between the identified ABO-associated metabolites and PDAC in 2 nested case-control studies (493 cases, 640 controls) using logistic regression and evaluated metabolite mediation of the ABO-PDAC association. Non-O and A (vs O) blood groups were associated with 13 and 20 metabolites, respectively, at false discovery rate <0.20, with 9 in common. The ABO-associated metabolites, sphingosine (non-O: β = 0.15), aspartate (A: β = -0.11), and aspartylphenylalanine (A: β = -0.16) were positively, and fibrinopeptide B (1-13) (non-O: β = 0.13; A: β = 0.21) was inversely associated with PDAC (OR = 0.96-1.07 per SD change log10-metabolite, P <.05) after adjustment for blood group. Non-O (OR = 1.50 [95% CI, 1.16-1.94]) and A (OR = 1.46 [95% CI, 1.10-1.92]) (vs O) blood groups were associated with PDAC, however none of the ABO associated metabolites significantly mediated the association between ABO blood group and PDAC. Our results suggest the ABO-associated metabolites are independent risk factors for PDAC. Trial registration: NCT00342992 and NCT00339495.
BACKGROUND:Evidence supports a modest positive association between alcohol intake and pancreatic ductal adenocarcinoma (PDAC); however, knowledge regarding mechanisms underlying the association is scarce. Investigation of lipidomic metabolites may provide mechanistic insights into this association. METHODS:We measured 611 lipid species across 14 lipid classes in serum samples collected up to 24 years before PDAC diagnosis in 2 nested case-control studies (706 matched sets) within American and Finnish cohorts. We conducted cross-sectional analyses using multivariable linear regressions to examine associations between log-transformed self-reported alcohol intake and log-transformed lipid concentrations among controls within each cohort. The identified alcohol-associated lipids in both cohorts were then evaluated for PDAC risk using multivariable conditional logistic regressions and fixed-effects meta-analyses to estimate overall odds ratios across the 2 cohorts. RESULTS:Alcohol intake was associated with 21 lipid species, 11 class-specific fatty acids (FA), 3 total FA, and 1 lipid class at Bonferroni significance thresholds with similar directions of associations in both cohorts. Among them, total pentadecanoic acid (FA15:0) and 7 lipid species-TAG(49:3-FA18:2), TAG(51:3-FA18:2), TAG(49:2-FA18:2), TAG(51:3-FA15:0), TAG(51:2-FA18:2), TAG(51:2-FA15:0), and PC(15:0-18:2)-were inversely associated with alcohol intake and with PDAC risk at false discovery rate <0.10, with overall odds ratios ranging from 0.82 to 0.86, without evidence of heterogeneity by smoking habits. CONCLUSION:Findings from 2 prospective cohorts identified 7 lipid species and 1 FA inversely associated with both alcohol intake and PDAC risk. These results suggest that alcohol intake may be positively associated with PDAC through downregulation of circulating lipids years before PDAC diagnosis.
Abstract Introduction The Connect for Cancer Prevention study is a new prospective cohort with repeated exposure assessment and long-term follow-up with the goals to study cancer initiation, multi-step carcinogenesis, early detection, and outcomes in a US study population. Over 85,000 participants have been recruited so far at 10 U.S. integrated healthcare systems. Biospecimens are a critical component to achieve Connect’s goals. Designing biospecimen collections in prospective cohort studies needs to balance the desire for large, repeated collections of various biospecimens with participant burden and cost. Methods The biospecimen collection protocol was informed by literature review, expert consultations, and pilot studies evaluating the effect of pre-analytical factors on commonly measured biomarkers. Baseline biospecimen collection includes a blood draw with serum, plasma, cell-free DNA collection tubes, a urine collection, and a mouthwash sample. Biospecimen collection is performed at over 50 collection locations, including within the clinical phlebotomy infrastructure and dedicated research laboratories. All biospecimens are shipped to a NCI central laboratory for processing and long-term storage. Process metrics include sample completeness, sample deviations, temperature logging, and needle-to-processing time, among others. Repeated biospecimen collections to study different exposure windows and biomarker changes within individuals are planned every three years, with more frequent collections among participants age 50 and older to pursue cancer early detection aims. Results As of October 2025, 56,445 participants of 81,030 enrolled (70%) donated blood and urine samples, with additional collections underway. We observed higher proportions of biospecimen donations in older age groups, ranging from 56% among participants age 30-34 to 83% among 66-70 year olds. Biospecimen collection participation was similar by sex and race/ethnicity. Among the collections, 60% were from clinical sites, and 40% from research laboratories. 82% of biospecimens collected at research laboratories were received at NCI within one day, and over 95% of all biospecimens were received within 4 days. The return of home-collected mouthwash samples was 77% among those sent a kit. Among 422,000 biospecimen tubes collected, 94% were complete with no deviations recorded. Over 90% of participants submitted a short survey at the time or shortly after biospecimen collection. Conclusions The Connect Cohort for Cancer Prevention combines electronic health record data, state-of-the-art surveys, and repeated biospecimen collections to address critical questions on cancer etiology and prevention. We successfully implemented a robust and efficient biospecimen collection approach at 10 recruitment sites across the U.S. Citation Format: Nicolas A. Wentzensen, Stephanie J. Weinstein, Amanda Black, Erin Schwartz, Hannah P. Yang, Michelle Brotzman, Paul Albert, Laura E. Beane Freeman, Amy Berrington de Gonzalez, Jonas S. Almeida, Jonine Figueroa, Montserrat García-Closas, Nicole Gerlanc, Gretchen L. Gierach, Rena Jones, Peter Kraft, Autumn Hullings, Charles E. Matthews, Habibul Ahsan, Brisa Aschebrook-Kilfoy, Chun-Hung Chan, Robert Greenlee, Stacey Honda, Benjamin A. Rybicki, Blythe Ryerson, Katherine Sanchez, Mark A. Schmidt, Kevin Skyes, Larissa L. White, Jeanette Ziegenfuss, Stephen J. Chanock, Christian C. Abnet, Mia M. Gaudet. Biospecimen collections for cancer etiology and prevention research in the Connect for Cancer Prevention Study: Guiding principles, approach, and key metrics [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2026; Part 1 (Regular Abstracts); 2026 Apr 17-22; San Diego, CA. Philadelphia (PA): AACR; Cancer Res 2026;86(7 Suppl):Abstract nr 5098.
This table includes the overall sample description stratified by colorectal cancer (CRC) status and smoking status.
Reaching the national goal of reducing cancer mortality by 50% within 25 years will require improvements in cancer early detection and prevention in addition to treatment. Modern prospective cohorts that capture new and emerging exposures to research cancer etiology are critical to achieve these goals. Profound societal and technological changes in the last decade present opportunities and challenges for the recruitment, engagement, and retention of participants in new cohorts. The Connect for Cancer Prevention Study is a modern cohort of adults, 30-70 years old without a prior cancer diagnosis, recruited from 10 U.S. integrated health care systems. Participants provide information and biospecimens at enrollment and at regular intervals during at least 20 years of follow up. Exposure and outcome information will be captured through online surveys, electronic medical records, medical imaging, geospatial linkages of 20-year residential histories, wearable sensors, and linkages with the National Death Index and the state cancer registries. Collected biospecimens include blood, urine, saliva, and fecal samples, as well as precursor and tumor tissue specimens. Data systems follow F.A.I.R. (findability, accessibility, interoperability, and reuse of digital assets) principles to maximize data and tool re-usability. In 2021, recruitment commenced among an eligible catchment population of more than 3.2 million patients. As of November 2024, over 53, 000 of the expected 200, 000 participants consented; 87% completed at least some baseline activities to date. Baseline recruitment is expected to be complete in 2027. The current study population has a median age of 53 (IQR: 42-61) years, 67% are female, 63% completed at least a bachelor’s degree, and 88% report good to excellent overall health. Among the cohort, 61% of participants self-report as White, 14% self-report as multi-racial, 11% as Black, African American or African, 5% as Asian, and 4% as Hispanic, Latino, or Spanish among other categories. The portion of the latter category is expected to increase as recruitment ramps up at a site with a high Hispanic catchment population, which started enrollment in 2023. The proportion of individuals who have ever smoked cigarettes is 30%, whereas 9% have ever used marijuana, 3% have ever smoked cigars, and 2% have ever vaped nicotine containing electronic cigarettes. The median body mass index of participants is 30 kg/m2 (IQR: 24-36 kg/m2). Approximately, 30, 000 incident precursor and 8, 000 cancer diagnoses are expected to occur during the first 10 years of follow up. Connect for Cancer Prevention Study combines novel approaches in epidemiology and data science to provide a valuable resource for the scientific community to study cancer etiology, natural history, risk prediction, early detection, and survivorship. Individual-level data is expected to be released to the scientific community in 2026. Mia M. Gaudet, Amy Berrington de Gonzalez, Christian C. Abnet, Paul Albert, Jonas S. Almeida, Stephanie Weinstein, Amanda Black, Hannah P. Yang, Michelle Brotzman, Laura Beane-Freeman, Laura Beane-Freeman, Jonine Figueroa, Neal D. Freedman, Nicole M. Gerlanc, Gretchen L. Gierach, Rena R. Jones, Peter Kraft, Charles Matthews, Habibul Ashan, Briseis Aschebrook-Kilfoy, Chun-Hung Chan, Robert T. Greenlee, Stacey A. Honda, Benjamin Rybicki, Katherine Sanchez, Kevin Skyes, Mark A. Schmidt, Larissa L. White, Jeanette Y. Ziegenfuss, Stephen J. Chanock, Montse Garcia-Closas, Nicolas A. Wentzensen. Connect for Cancer Prevention Study: a modern prospective cohort [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2025; Part 1 (Regular Abstracts); 2025 Apr 25-30; Chicago, IL. Philadelphia (PA): AACR; Cancer Res 2025;85(8_Suppl_1):Abstract nr 7362.
This file details the two-step interaction tests, and the gene-based aggregate test.
This file includes the expression imputation statistics and included SNPs from the elastin net models.
Background:High intake of red and/or processed meat are established colorectal cancer (CRC) risk factors. Genome-wide association studies (GWAS) have reported 204 variants (G) associated with CRC risk. We used functional annotation data to identify subsets of variants within known pathways and constructed pathway-based Polygenic Risk Scores (pPRS) to model pPRS x environment (E) interactions. Methods:A pooled sample of 30,812 cases and 40,504 CRC controls of European ancestry from 27 studies were analyzed. Quantiles for red and processed meat intake were constructed. The 204 GWAS variants were annotated to genes with AnnoQ and assessed for overrepresentation in PANTHER-reported pathways. pPRS's were constructed from significantly overrepresented pathways. Covariate-adjusted logistic regression models evaluated pPRSxE interactions with red or processed meat intake in relation to CRC risk. Results:A total of 30 variants were overrepresented in four pathways: Alzheimer disease-presenilin, Cadherin/WNT-signaling, Gonadotropin-releasing hormone receptor, and TGF-β signaling. We found a significant interaction between TGF-β-pPRS and red meat intake (p = 0.003). When variants in the TGF-β pathway were assessed, significant interactions with red meat for rs2337113 (intron SMAD7 gene, Chr18), and rs2208603 (intergenic region BMP5, Chr6) (p = 0.013 & 0.011, respectively) were observed. We did not find evidence of pPRS x red meat interactions for other pathways or with processed meat. Conclusions:This pathway-based interaction analysis revealed a significant interaction between variants in the TGF-β pathway and red meat consumption that impacts CRC risk. Impact:These findings shed light into the possible mechanistic link between CRC risk and red meat consumption.
This file includes the association parameters (OR [95% CI]) for the identified SNPs with significant interaction term for smoking intensity by genotypes.
Pancreatic ductal adenocarcinoma (PDAC) is highly fatal, with incidence rising worldwide. Metabolomics may provide insight into etiology and mechanisms contributing to pancreatic carcinogenesis. We examined associations between 1483 prediagnostic (up to 24 years) serum metabolites and PDAC in nested case–control studies within a cohort of male Finnish smokers and another of American men and women ( n = 732 matched pairs). We used conditional logistic regression to calculate odds ratios (OR) and 95% confidence intervals per standard deviation increase in log‐metabolite level within each cohort and combined using fixed‐effect meta‐analyses. We performed elastic net regression (EN) to select metabolites and calculated area under the curve (AUC) for established PDAC risk factors (smoking, diabetes, and overweight/obesity), selected metabolites, and their combination. Sixty‐six metabolites were associated with PDAC at false discovery rate <0.05, with 26 below Bonferroni threshold ( p < 3.4 × 10 −5 ) and 38 not reported previously. Notable findings include fibrinopeptide B (1–9); 13 modified, di‐ or poly‐peptides; 11 tobacco‐chemical related xenobiotics; glycolysis–gluconeogenesis–tricarboxylic acid (TCA) cycle metabolites (aspartate, glutamate, lactate, α‐ketoglutarate, and pyruvate); and four secondary and two primary bile acids that were positively (OR = 1.18–1.58) and five fibrinogen cleavage peptides that were inversely (OR = 0.70–0.84) associated with PDAC. AUCs for combined metabolites‐risk factors outperformed known risk factors ( p ≤ .01) but not metabolites ( p ≥ .31) alone. Systemic metabolism is prospectively associated with PDAC. New metabolite associations include those related to immune response, tobacco, microbiome, glycolysis–gluconeogenesis and TCA cycle, and adiposity or diabetes. The EN selected metabolites were more sensitive indicators of prediagnostic metabolic processes and exposures associated with PDAC than established risk factors.
This table includes the association parameters of smoking habits for colorectal cancer risk stratified by study type.
Background:Studies have reported higher lung cancer incidence among groups with lower socioeconomic position (SEP). However, it is not known how this difference in lung cancer incidence between SEP groups varies across different geographical settings. Furthermore, most prior studies that assessed the association between SEP and lung cancer incidence were conducted without detailed adjustment for smoking. Therefore, we aimed to assess this relationship across world regions. Methods:In this international prospective cohort consortium study, we used data from the Lung Cancer Cohort Consortium (LC3), which includes 20 prospective population cohorts from 16 countries in North America, Europe, Asia, and Australia. Participants were enrolled between 1985 and 2010 and followed for cancer outcomes using registry linkages and/or active follow-up. We estimated hazard ratios (HRs) for the association between educational level (our primary measure of SEP, in 4 categories) and incident lung cancer using Cox proportional hazards models separately for participants with and without a smoking history. The models were adjusted for age, sex, cohort (when multiple cohorts were included), smoking duration, cigarettes per day, and time since cessation. Findings:Among 2,487,511 participants, 53,830 developed lung cancer during a 13.5-year median follow-up (IQR = 6.5-15.0 years). Among participants with a smoking history, higher education was associated with decreased lung cancer incidence in nearly every cohort after detailed smoking adjustment. By world region, this association was observed in North America (HR per one-category increase in education [HRtrend] = 0.88, 95% CI = 0.87-0.89), Europe (HRtrend = 0.89, 95% CI = 0.88-0.91), and Asia (HRtrend = 0.91, 95% CI = 0.86-0.96), but not in the Australian study (HRtrend = 1.02, 95% CI = 0.95-1.09). By histological subtype, education associated most strongly with squamous cell carcinoma and more weakly with adenocarcinoma (p-heterogeneity < 0.0001). Among participants who never smoked, there was no association between education and lung cancer incidence in any cohort (all p-trend > 0.05), except the USA Southern Community Cohort Study (HRtrend = 0.75, 95% CI = 0.62-0.90). Interpretation:Based on longitudinal data from 2.5 million participants from 16 countries, our findings suggest that higher educational attainment was associated with lower lung cancer risk among participants with a smoking history, but not among participants who never smoked. Limitations of our study include that cohort participants cannot fully represent the general populations of the geographical regions included, and education was the only measure of SEP consistently available across our consortium. Funding:This study was supported in part by the National Cancer Institute (NCI), the Lung Cancer Research Foundation (LCRF), and the World Cancer Research Fund (WCRF).
This figure depicts the LocusZoom plots for SNPs interacting with smoking intensity for colorectal cancer risk.
This file includes the association parameters of the interaction component (OR [95% CI]) for the identified SNPs for smoking habits stratified by tumor molecular markers.