Emerging evidence suggests a critical role of the tumor microenvironment (TME) in breast cancer (BC) development and outcomes, yet factors that modify the TME are poorly understood. We investigated the relationship between BC etiological factors and the tumor and TME using 110 histological features of the epithelium, stroma, and immune infiltration, computationally quantified in 3724 H&E slides from three prospective cohort studies. Age, race, hormonal, and lifestyle factors were associated with features of the breast TME. Menopausal hormone therapy was associated with epithelial and stromal features found in less aggressive tumors, while higher body mass index (BMI) was associated with two histological features not captured by grade, and both were associated with poor prognosis. These two features mediated the BMI and BC-specific mortality association by 18.1%. Our findings provide novel insights into the role of etiological factors on the TME including modifiable factors that have implications for prevention and outcomes.
statistics fine-mapping methods offer advantages over classical methods, including avoiding data-sharing constraints and improved modelling of correlated variables and sparse effects. However, its performance has not been comprehensively evaluated in breast cancer using real-world data. Previous multinomial stepwise regression (MNR) fine-mapping analyses for breast cancer identified 196 credible sets. Here, we apply summary statistics fine-mapping, compare methods, and assess parameters influencing performance. Using summary statistics from the Breast Cancer Association Consortium, we compared finiMOM, SuSiE, and FINEMAP to published MNR results across 129 regions. Performance was assessed by recall using in-sample and out-of-sample LD. Discordant credible sets were examined for technical factors, and target genes were defined using the INQUISIT pipeline. SuSiE showed the closest agreement with MNR. Results varied across regions depending on the assumed number of causal variants (L), with higher values reducing recall and no single L maximising performance. At optimal L per region, SuSiE identified 8,192 CCVs in 244 credible sets, with recall of 88%, 86%, and 72% for overall, ER-positive, and ER-negative breast cancer. Thirty MNR sets were missed. Discordance was partially explained by allele flips, imputation quality, and array heterogeneity. Fifty-two MNR-identified genes, including BRCA2, WNT7B and CREBBP were not recovered, while additional candidate genes were identified. Using out-of-sample LD reduced recall by 3% but identified novel variants. Fine-mapping results vary across methods, and no single approach is sufficient. The choice of L strongly influences results, and combining analytical approaches with functional validation can improve causal variant identification.
Abstract Background: The tumor immune microenvironment may provide key information on breast cancer prognosis. Several modifiable risk factors influence the immune response, including weight, physical activity, and diet; however, whether these factors alter the breast tumor immune microenvironment remains unknown. Methods: Participants enrolled in the Nurses’ Health Studies diagnosed with invasive breast cancer and available tumor and/or normal adjacent tissue were included (N=945). Immune cell abundance was deconvoluted using CIBERSORTx. Gene expression signatures were derived for components of the immune profile such as immune checkpoint markers, co-regulatory signal and antigen presentation, and cytokine signaling. Participant weight, BMI, alcohol use, and smoking status were collected via bi-annual questionnaires. Overall diet was summarized by the alternative healthy eating index (AHEI) and the empirical dietary inflammatory pattern (EDIP), derived from food frequency questionnaires completed every 4 years. Linear regression was used to test the association between pre-diagnostic exposures (from questionnaires closest to diagnosis) and immune cell abundance and gene expression. Tumor and normal-adjacent tissue were analyzed separately. Models were adjusted for year and age of diagnosis, menopausal status, and estrogen-receptor (ER) status. Multiple testing of immune components was controlled via the false discovery rate. Models were also stratified by ER and menopausal status. Results: Among 875 women with available tumor tissue, mean age at diagnosis was 59 years (SD=11.4). The majority (73%) were post-menopausal, had ER-positive tumors (77%), and were diagnosed at stage 1-2 (91%). In ER-positive tumor tissue of postmenopausal women, weight gain since age 18 was positively associated with interferon signaling, MHCII, and PD1 expression (all padj<0.05). Higher physical activity was associated with enriched CD8:CD68 ratio in this group. In ER-negative tumor tissue of pre- and postmenopausal women, consuming more drinks per week was associated with higher PDL1 and lower CSR expression. In normal-adjacent tissue, current smoking was associated with enriched cytokine signaling in premenopausal women and total pack-years was associated with higher B cell plasma in postmenopausal women. Higher EDIP was associated with heightened MHCII and IL12 expression in ER-negative tumor tissue while higher AHEI was associated with lower lymphocyte infiltration expression in normal adjacent tissue, regardless of menopausal status. Conclusions: Weight gain, smoking, alcohol use, and diet (AHEI and EDIP) are associated with differences in the breast cancer immune microenvironment, with distinct associations by ER and menopausal status. Changes in lifestyle behaviors may influence the tumor immune response and impact patient prognosis, though further studies are needed. Citation Format: Kristen D. Brantley, Cheng Peng, Clara Bodelon, Deborah A. Tadesse, Peter Kraft, Rulla M. Tamimi. Modifiable lifestyle factors and immune gene expression in breast tumor and normal-adjacent tissue [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2026; Part 1 (Regular Abstracts); 2026 Apr 17-22; San Diego, CA. Philadelphia (PA): AACR; Cancer Res 2026;86(7 Suppl):Abstract nr 2090.
Abstract Artificial intelligence (AI)-scores estimated from digital mammograms predict future breast cancer (BC) risk. MIRAI is a deep learning BC risk model that provides a continuous 5-year risk of overall BC from four screening full field digital mammograms. Common germline genetic variation in the form of a polygenic risk score (BC-PRS) is associated with increased BC risk and may improve MIRAI’s 5-year risk prediction. Our goal was to develop an updated MIRAI 5-year risk model that incorporates the BC-PRS (MIRAI+PRS) and to evaluate its discriminatory accuracy and calibration for both overall and invasive BC compared to MIRAI alone. We developed the MIRAI+PRS model by multiplying each woman’s 5-year MIRAI risk estimate by their relative risk based on their BC-PRS relative to the population mean. We evaluated the models within the Mayo Clinic Biobank mammography cohort, comprised of 12,307 women without a prior history of BC; 176 invasive and 250 overall BC were diagnosed within 5 years. MIRAI was estimated on screening mammograms closest to enrollment but at least 6 months prior to BC. Discriminatory accuracy, assessed by C-statistic, was high and similar for MIRAI+PRS vs. MIRAI models, for overall BC and invasive BC (Table). Calibration assessed by observed to expected (O/E) ratios was also similar for MIRAI+PRS compared to MIRAI predictions for BC outcomes (Table), although there was improvement in decile-specific O/E ratios across the lowest risk deciles (<1.67%) for MIRAI+PRS. Calibration for invasive cancer was poor for MIRAI with or without PRS. In summary, the MIRAI+PRS risk model did not result in significant difference of discriminatory accuracy or overall calibration compared to the MIRAI model, but there was evidence for improved calibration for women with 5-year risk below 1.67%. For invasive BC, the model had poor calibration regardless of whether PRS was included, underscoring the importance of training AI models for BC outcomes that are associated with a clinical intervention. Citation Format: Christopher G. Scott, Peter Kraft, Imon Banerjee, Ramon Correa Medero, Aaron D. Norman, Fergus J. Couch, Karla Kerlikowske, Stacey J. Winham, Celine M. Vachon. Development and evaluation of a MIRAI 5-year risk model with a breast cancer polygenic risk score [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2026; Part 1 (Regular Abstracts); 2026 Apr 17-22; San Diego, CA. Philadelphia (PA): AACR; Cancer Res 2026;86(7 Suppl):Abstract nr 2779.
Childhood cancer radiation therapy (RT) increases subsequent neoplasm risk. Radiation dose may modulate DNA damage responses, but the small sample sizes of prior human studies of homologous recombination repair hampered dose-specific investigations. We pooled data for 12 180 survivors (8339 from the Childhood Cancer Survivor Study and 3841 from the St Jude Lifetime Cohort) to estimate associations between deleterious homologous recombination repair variants and RT-related subsequent neoplasms (most commonly breast cancer, meningioma, thyroid cancer, and sarcoma) using conditional logistic regression with matched controls. In all, 1253 (10.3%) survivors were homologous recombination repair variant carriers, and 1301 (10.7%) developed at least 1 RT-related subsequent neoplasms. Variants increased the risk of out-of-field RT-related subsequent neoplasms (cases, 40/190 [21.1%]; control individuals, 9.7%; odds ratio [OR] = 2.5, 95% confidence interval [CI] = 1.7 to 3.6; P = 4.80 ×10-6), with consistent results across cohorts (Childhood Cancer Survivor Study, OR = 2.5, 95% CI = 1.6 to 3.7, P = 3.77 ×10-5; St Jude Lifetime Cohort, OR = 2.5, 95% CI = 1.0 to 6.4, P = 3.07 ×10-2). No association was observed for in-field or near-field subsequent neoplasms or individuals not undergoing RT. Findings emphasize homologous recombination repair variant-conferred susceptibility to RT-related subsequent neoplasms and dose-dependent DNA damage repair.
Abstract Polygenic risk scores (PRSs) may enhance risk stratification for pancreatic ductal adenocarcinoma (PDAC), but existing models vary widely in design, predictive performance, and cross-ancestry transferability. We developed genome-wide PRSs using Bayesian methods (LDpred2 and PRS-CS) and p value thresholding (PRSice-2) and systematically evaluated these alongside 13 published PRSs to identify models with robust predictive performance across ancestries. Using GWAS summary statistics from 7531 cases and 10,631 controls, we derived the PRSs and tested associations in an independent sample of 4508 PDAC cases and 46,189 controls, with adjustment for well-established PDAC risk factors. Among all models, the genome-wide LDpred2-based PRS showed the strongest association with PDAC (OR = 1.57 per standard deviation increase; 95% CI: 1.51–1.62) and significantly improved discrimination beyond established risk factors alone (AUC = 0.74–0.76; p < 0.0001). Importantly, the genome-wide LDpred2 PRS demonstrated consistent associations across African, Admixed American, and European ancestry groups, whereas the best-performing published PRS was associated with PDAC risk only in individuals of European ancestry. These findings support genome-wide PRSs as a promising framework for multi-ancestry risk stratification for PDAC and to inform targeted early detection strategies.
Abstract Background: Prior genome-wide association studies (GWAS) of acute myeloid leukemia (AML) have been limited by sample size and ancestry representation. We conducted the largest multi-ancestry GWAS to-date to identify germline susceptibility loci associated with AML. Methods: The study is part of the NCI-CIBMTR® collaborative Genomic Studies in Blood and Marrow Transplantation (GS-BMT) project. Patients with AML were allogeneic hematopoietic cell transplantation (HCT) recipients with blood samples collected before HCT (82% were in complete morphologic remission). AML-free controls included HCT donors from GS-BMT and participants from the Prostate, Lung, Colorectal and Ovarian (PLCO) Cancer Screening Trial. Samples were genotyped on Illumina Global Screening Array platforms, imputed to the TOPMed v3 reference panel, and analyzed using REGENIE v4.1 under an additive model with Firth-approximate likelihood ratio tests. Models were adjusted for sex, age, and genetic principal components. Variants with imputation quality score > 0.2 and minor allele count ≥ 30 were tested. Results: The primary analysis included 10,937 cases and 100,705 controls spanning multiple genetically inferred ancestry groups, including 1,776 cases of non-European ancestry. Test statistic inflation was limited (λ1,000=1.001; LDSCR intercept=1.014), indicating minimal residual confounding. We observed eight genomic regions with genome-wide significant associations. Among these, we highlight three signals on chromosome 5, including rs141601766 at 5q35 (OR = 22.79, 95% CI = 14.34-36.20, P = 1.05×10-40), a rare non-synonymous variant in DDX41 classified as pathogenic or likely pathogenic in ClinVar and previously implicated in familial AML predisposition; rs552806293 at 5q35 (OR = 22.97, 95% CI = 13.26-39.80, P = 1.03×10-28), a rare intronic variant within UIMC1, a gene with established roles in DNA damage response; and rs7705526 at 5p15 (TERT locus; OR = 1.18, 95% CI = 1.15-1.22, P = 6.65×10-25), a common variant previously associated with leukocyte telomere length, clonal hematopoiesis, myeloproliferative neoplasms, hematologic quantitative traits, and ovarian serous carcinoma in large GWAS. Conditional analyses suggested the two 5q35 signals were statistically independent. Conclusions: This multi-ancestry GWAS identifies multiple germline susceptibility loci for AML risk. Ongoing work includes ancestry-specific, molecular-stratification, and outcome analyses. Citation Format: Xueyao Wu, Filip Pirsl, Gabrielle Schmidt, Maryam Rafati, Aurélie Vogt, Herbert Higson, Jia Liu, Jiahui Wang, Shilpa Gaddam, Shengchao Li, Wael Saber, Yung-Tsi Bolon, Steven Moore, Sharon A. Savage, Stephen Chanock, Stephen Spellman, Peter Kraft, Shahinaz M. Gadalla. Genome-wide association study identifies germline susceptibility loci for acute myeloid leukemia [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2026; Part 2 (Late-Breaking, Clinical Trial, and Invited Abstracts); 2026 Apr 17-22; San Diego, CA. Philadelphia (PA): AACR; Cancer Res 2026;86(8_Suppl):Abstract nr LB387.
Background: Mammograms contain imaging biomarkers that can predict future breast cancer risk using deep learning (DL) models. We evaluated whether adding a polygenic risk score (PRS) improves performance of the image-only DL breast cancer risk model Mirai. Methods: This nested case-control study within the Nurses' Health Study 2 included 902 women (270 cases, 632 controls) who underwent bilateral 2D digital screening mammography between 2001-2017. Risk was assessed using Mirai and, for clinical comparison, the Gail 5-year model. A PRS was calculated using 313 breast cancer-associated single-nucleotide polymorphisms. The primary outcome was incident breast cancer within five years of the index mammogram. Discrimination was evaluated using area under the receiver operating characteristic curve (AUC), with comparisons using the DeLong test. Results: Mean age was 55.5 years(SD 5.3). Among cases, median time from index mammogram to diagnosis was 2.0 years (IQR0.5-4.0). Mirai alone achieved an AUC of 0.66 (95% CI: 0.62-0.70), increasing to 0.73 (95% CI 0.67-0.78; P = 0.05) with PRS. The Gail model improved from 0.52 (95% CI: 0.47-0.57) to 0.69 (95% CI: 0.62-0.76; P < 0.001) with PRS. Mirai+PRS significantly outperformed Gail+PRS (P < 0.001). Conclusions: Integrating PRS with DL-based mammographic models modestly improves risk discrimination and may enhance personalized screening.
The increasing availability of diverse biobanks has enabled multi-ancestry genome-wide association studies (GWASs) to enhance the discovery of genetic variants across traits and diseases. However, the choice of an optimal method remains debated, due to challenges in statistical power differences across ancestral groups and approaches to account for population structure. Two primary strategies exist: (1) pooled analysis, which combines individuals from all genetic backgrounds into a single dataset while adjusting for population stratification using principal components, increasing the sample size and statistical power but requiring careful control of population stratification; and (2) meta-analysis, which performs ancestry-group-specific GWASs and subsequently combines summary statistics, potentially capturing fine-scale population structure but facing limitations in handling admixed individuals. Using large-scale simulations with varying sample sizes and ancestry compositions, we compare these methods alongside real data analyses of eight continuous and five binary traits from the UK Biobank (N ≈ 324,000) and the All of Us Research Program (N ≈ 207,000). Our results demonstrate that pooled analysis generally exhibits better statistical power while effectively adjusting for population stratification. We further present a theoretical framework linking power differences to allele-frequency variations across populations. These findings, validated across both biobanks, highlight pooled analysis as a powerful and scalable strategy for multi-ancestry GWASs, improving genetic discovery while maintaining rigorous population structure control.
Polycystic ovary syndrome (PCOS) and its underlying features remain poorly understood. In this genetic study (n = 544,513), we expand the number of genetic loci from 16 to 29, and additionally identify 31 associated plasma proteins. Many risk-increasing loci were associated with later age at menopause, underscoring the reproductive longevity related to an increased oocyte number and/or availability across the lifespan. Hormonal regulation in the etiology of this condition, through metabolic and reproductive features, was emphasized. The proteomic analysis highlighted metabolic biology known to be related to PCOS. A polygenic risk score (PRS) was associated with adverse cardiometabolic outcomes, with differing relevance of testosterone and body mass index in women and men. Finally, while oligo-anovulation and anovulatory infertility are features of PCOS, we observed no impact of PCOS susceptibility on childlessness. We suggest that PCOS susceptibility confers balanced pleiotropic influences on fertility in women, and life-long adverse metabolic consequences in both sexes.
Conventional prediction models incorporating genetic and clinical factors including breast density underperform in non- European populations. We investigated an artificial intelligence-derived mammogram risk score (MRS), a summary of texture features that captures intrinsic breast- tissue characteristics-the substrate for cancer devel-opment. This study leveraged data from two North American screening cohorts totaling >226,000 women, includ-ing non- Hispanic white, non- Hispanic Black, East Asian, South Asian, and Indigenous women. MRS distributions showed nonsignificant shifts (and similar SDs) between cohorts and across race and ethnic subgroups. MRS in-creased with age and was significantly associated with breast cancer risk, with hazard ratios per SD ranging from 2.24 [95% confidence interval (CI), 2.03 to 2.46] to 2.32 (95% CI, 2.25 to 2.39) after age adjustment. Associations remained significant within all subgroups. Calibration was excellent across the racial and ethnic groups and across full- field digital mammograms and tomosynthesis. These findings establish MRS as a strong predictor that is inde-pendent of race or ethnicity, demonstrating its potential for broader clinical utility.
PRSs predict complex traits by aggregating genetic effects across the genome, yet most models focus on common variants, overlooking rare variants that may contribute to hidden heritability. Here, we develop RICE, a PRS framework integrating both common and rare variants to improve genetic risk prediction across diverse ancestries. RICE constructs separate PRSs: for common variants, it integrates methods using ensemble learning; for rare variants, it uses gene-level testing with functional annotations and penalized regression. We evaluate RICE using simulated datasets and sequencing data from UK Biobank and All of Us, involving up to 740 million genetic variants from 361,939 individuals across diverse ancestries and 11 complex traits. In real data analysis, RICE improves predictive accuracy compared to leading common variant methods for traits with distinct rare variant architectures, particularly lipids and height. For lipid traits, incorporating rare variants increased R2 by up to ~11.2% in Europeans and ~60.7% in African ancestry compared to common variant PRS alone. Notably, for lipid traits, RICE captures substantial predictive signal beyond established high-penetrance genes, validating its ability to leverage the broader polygenic architecture of rare variation.
Abstract Background: Renal cell carcinoma (RCC), the predominant form of kidney cancer, is influenced by several risk factors (RFs) including obesity, hypertension, and smoking. However, the molecular mechanisms linking these RFs to RCC remain unclear. Methods: We investigated plasma proteins (PP) as potential intermediates of markers of the effects of RFs on RCC using two-stage Mendelian randomization (TSMR) approach. In stage 1, we identified PPs associated with each of the 19 RFs evaluated (e.g., anthropometric traits, blood pressure, smoking behavior, blood cell counts, and kidney function), leveraging summary-level proteogenetic data on PPs from the UK Biobank Pharma Proteomics Project (N=34,557). In stage 2, we evaluated the effects of these RF-associated PPs on RCC, using the largest-to-date RCC GWAS (N = 864,690; cases=29,020). Results: Among 2,940 PPs, 2,339 were significantly associated (P<1.7E-05) with at least one of the 19 RFs. Of these, 33 showed a significant effect on RCC (FDR<5%) with 28 mapping outside RCC GWAS loci. Using multivariable MR, we estimated mediation effects of associated PPs, finding that proteins such as CDA and PILRB mediated up to 17.41% of BMI’s effect on RCC, and APOL1 mediated 2.76% of white blood cell’s effect. Convergent evidence from multiple in silico analyses with cis-MR, colocalization, and TCGA differential expression further prioritized TYMP, UMOD and USP28 as key protein intermediaries of the RF effects on RCC. TYMP and USP28, inversely associated with RCC risk, showed immune-related and tumor-suppressive effects, while UMOD was positively associated, potentially linking renal dysfunction to carcinogenesis. Functional annotation revealed enhancer activity and transcription factor (HIF) binding near these proteins. Conclusion: Our approach and results identify molecular intermediates that may link epidemiologic risk factors to RCC and highlight actionable candidates for laboratory investigation. Citation Format: Ibrahim Hossain Sajal, Andrew J. Song, Kevin M. Brown, Mitchell J. Machiela, Peter Kraft, Stephen J. Chanock, Mark P. Purdue, Diptavo Dutta. Integrative analysis identifies potential proteomic intermediates associated with renal cell carcinoma and its risk factors [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2026; Part 1 (Regular Abstracts); 2026 Apr 17-22; San Diego, CA. Philadelphia (PA): AACR; Cancer Res 2026;86(7 Suppl):Abstract nr 1481.
ABSTRACT Background Several breast cancer (BC) risk prediction models have been developed to provide personal risk assessments. Though individually validated, their performance has not been systematically evaluated across a wide range of populations or ages. Methods We harmonized individual-level baseline questionnaire data and incident BC diagnoses from 21 cohorts from North America, Europe, and Australia participating in the Breast Cancer Risk Prediction Project. Five-year absolute risk of invasive BC was estimated for five established risk prediction models using classical risk factors only. Discrimination was evaluated by area under the curve (AUC). Calibration was assessed using average and risk-decile specific expected to observed (E/O) ratios. Performance metrics were meta-analyzed across cohorts and models. Metaregression tested associations between cohort characteristics and performance metrics. Results This analysis included 1,595,977 women aged 20-75 years, enrolled in studies between 1976-2015, with 19,062 (1.2%) invasive BC cases ascertained within 5 years from exposure assessment. Age-adjusted AUCs were similar across models and cohorts (pooled AUCs by model: 0.57-0.58), while E/O ratios varied substantially (pooled E/O ratios by model: 0.83-1.25). Overestimation was common among predicted high-risk individuals (>3%). No appreciable differences in model performance by cohort age, birth year, race, and variable missingness emerged. Calibration improved after assigning race-specific incidence rates. Conclusion Existing BC risk prediction models provided similar risk discrimination across multiple cohorts, although there was overestimation of risk for high-risk individuals. Performance variation across cohorts was not driven by specific characteristics, which supports development of a unified risk model for diverse populations that leverages appropriate incidence rates. Key messages When using classical risk factor components of existing risk prediction models, we found similar discriminatory ability of models across diverse cohorts. Aside from underlying cancer incidence rate, which heavily influenced calibration, no cohort-specific characteristics were consistently associated with model performance. Risk was underestimated at lower predicted risk deciles and overestimated at higher predicted risk deciles, indicating a need to improve model fit by integrating more complex risk-factor relationships.
BACKGROUND:The associations between different types of diabetes, characterized by distinct pathophysiology and genetic architecture, and pancreatic ductal adenocarcinoma (PDAC) risk are not understood. METHODS:We investigated associations of genetic susceptibility to type 2 diabetes (T2D), 8 T2D mechanistic clusters, type 1 diabetes (T1D), and maturity-onset diabetes of the young (MODY) with PDAC risk. We used genome-wide association study (GWAS) summary-level statistics for T2D (242 283 cases, 1 569 734 controls), T1D (18 942 cases, 501 638 controls), and PDAC (10 244 cases and 360 535 controls) in individuals of European ancestry. RESULTS:Two-sample Mendelian randomization (MR) using the Robust Adjusted Profile Score (MR-RAPS) method indicated that genetically predicted T2D was associated with PDAC risk (OR = 1.10; 95% CI = 1.05 to 1.15), particularly the T2D obesity (OR = 1.28; 95% CI = 1.15 to 1.42) and lipodystrophy (OR = 1.25; 95% CI = 1.03 to 1.51) clusters. No association was observed for T1D with PDAC risk (OR = 1.01; 95% CI = 0.99 to 1.02). Pathway/gene-set analysis using the summary-based Adaptive Rank Truncated Product (sARTP) method revealed a significant association between the MODY gene-sets and PDAC risk (P = 1.5 × 10-8), which remained after excluding 20 known PDAC GWAS loci (P = 7.6 × 10-4). HNF1A, FOXA3, and HNF4A were the top contributing genes after excluding the previously identified GWAS loci regions. CONCLUSIONS:Our results from this genetic association study support that T2D, particularly the obesity and lipodystrophy mechanistic clusters, and MODY genomic susceptibility regions play a role in the etiology of PDAC.
Abstract Triple negative breast cancer (TNBC) exhibits distinct evolutionary pathways reflected in heterogeneous mutational and immune profiles. To better understand the relationships between prediagnostic exposures, inherited genetic variation, mutational and immune profiles in TNBCs, the PRediagnostic Exposures, Mutations, Immune SignaturEs-Triple Negative (PREMISE-TN) project performed whole exome sequencing (WES) of matched formalin-fixed paraffin embedded tumor tissue and germline DNA samples from 322 TNBC patients from four prospective cohort studies, the Nurses’ Health Study (NHS, NHS II) and the Cancer Prevention Study (CPS II, CPS3). After excluding 66 mis-matched tumor normal-pairs and 32 pairs where either the tumor or blood sample did not reach the target coverage (70% of bases covered at 20x), 224 pairs were available for analysis. The median sequencing depth for tumor samples in these pairs was 111.2x (range=7.4x-481x); the median for blood samples was 283.4x (range=156.2x-652.1x). Mutational calling and sequencing quality assessment and control is underway. Patients’ age at diagnosis ranged from 34-86 years (median=58), and year of diagnosis ranged from 1976-2018 (median=2004). 56 (28.7%) of the patients were premenopausal at diagnosis. Most tumors (n=162, 88.5%) were stage I-II; 19 (10.3%) were stage III and 1 (0.5%) was stage IV. PREMISE-TN integrated somatic mutational profiles with other data from these cohorts, including: germline genome-wide association study data, breast cancer risk factors, tumor immune signatures, and radiologic and pathologic phenotypes derived from digital images. Preliminary studies identified associations between germline polygenic risk scores (PRS) and breast tumor immune features, including inverse associations between PRS for immune-mediated conditions and interferon signaling in both breast tumors and adjacent normal tissue. Other studies examined the influence of reproductive factors on the breast tumor microenvironment. Future work will assess associations of these features with tumor mutational profiles. The data resource generated by PREMISE-TN will enable the investigation of how genetic and nongenetic risk factors influence breast tumor mutational signatures and immune response. Citation Format: Deborah A. Tadesse, Clara Bodelon, Cheng Peng, Kristen D. Brantley, Margaux Delporte, Yujing J. Heng, Lauren Teras, Rulla M. Tamimi, Peter Kraft. Pre-diagnostic exposures, mutational signatures, and immune profiles in triple-negative breast cancer: An overview of the PREMISE-TN project [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2026; Part 2 (Late-Breaking, Clinical Trial, and Invited Abstracts); 2026 Apr 17-22; San Diego, CA. Philadelphia (PA): AACR; Cancer Res 2026;86(8_Suppl):Abstract nr LB391.
Renal cell carcinoma (RCC), the predominant form of kidney cancer, is influenced by several risk factors (RFs) including obesity, hypertension, and smoking. However, the molecular mechanisms linking these RFs to RCC remain unclear. We investigated plasma proteins (PP) as potential intermediates of the effects of RFs on RCC using two-step Mendelian randomization (TSMR). In step 1, we identified PPs influenced by each RF, leveraging GWAS summary-statistics from the UK Biobank Proteomics (N = 34,557) for PPs and 19 RFs encompassing anthropometric traits, blood pressure, smoking behavior, blood cell, and kidney function. In step 2, we evaluated the effects of these RF-associated PPs on RCC, using the largest-to-date RCC GWAS (cases=29,020). Proteins significantly associated with both RFs and RCC were further prioritized through convergent evidence from multiple downstream analyses. Among 2,940 PPs, 2,339 were significantly associated (P < 1.7E-05) with at least one of the 19 RFs. Of these, 33 showed a significant effect on RCC (FDR<5%) with 28 mapping outside RCC GWAS loci. Using multivariable MR, we estimated mediation effects of associated PPs, finding that proteins such as APOL1 mediated 2.76% of white blood cell's effect. Convergent evidence from multiple downstream analyses prioritized TYMP, UMOD, and USP28 as key protein intermediaries of the RF effects on RCC. TYMP and USP28, both inversely associated with RCC risk, showed immune-related and tumor-suppressive effects, while UMOD was positively associated, potentially linking renal dysfunction to carcinogenesis. Our analysis highlights PPs as intermediates of the effect of distinct RFs on RCC, thereby nominating molecular targets for further investigation.
Polygenic risk score (PRS) models effectively predict breast cancer (BC) risk in European-ancestry women but have limited accuracy for African-ancestry women, particularly for aggressive subtypes. We developed PRS models for overall BC, estrogen receptor (ER)-positive, ER-negative and triple-negative BC (TNBC) in African-ancestry women using data from the African Ancestry Breast Cancer Genetics consortium (17,391 cases and 18,800 controls). We applied several PRS methods and integrated information across ancestries and BC subtypes. The best models for overall, ER-positive, ER-negative and TNBC showed an area under the receiving operating curve of 0.612, 0.621, 0.611 and 0.639, respectively, and maintained predictive accuracy in external validation studies with area under the receiving operating curves of 0.612, 0.640, 0.605 and 0.652. We further introduce a parsimonious 162-variant PRS for TNBC with comparable accuracy (0.626). These findings demonstrate markedly improved PRS accuracy for BC risk prediction in African-ancestry women. Using these PRS models for screening will help promote more equitable cancer prevention efforts.
Abstract Thyroid cancer is the most common endocrine malignancy, yet its biological underpinnings remain incompletely understood. Here we show that common risk alleles for thyroid cancer point to distinct biological pathways underlying disease susceptibility. We perform a multi-ancestry genome-wide association meta-analysis of thyroid cancer (16,167 cases and 2,430,374 controls), identifying 51 independent loci, including 21 not previously reported. By integrating these loci with genetic associations for 151 thyroid-cancer-related traits, we identify pleiotropic mechanistic clusters linked to thyroid function, oncogenic pathways, and mixed physiological function. Two thyroid-specific clusters, associated with thyroid stimulating hormone or thyroid growth and function, are enriched in thyroid tissues. Oncogenic clusters include DNA repair ( ATM , CHEK2 , TP53 ) and telomere maintenance ( TERT ) genes, implicating shared cancer mechanisms. Cluster-specific polygenic scores are associated with thyroid disease, cancer, and metabolic traits across ancestry groups, suggesting distinct genetic subtypes of thyroid cancer risk and supporting pleiotropy-based approaches to genetic risk stratification.