Supplementary Table 1: Study population and exclusion criteria by cohorts as established by the ACC reproductive factor working group
Abstract Background & Aims Dietary pattern analysis is essential in nutritional epidemiology, yet traditional clustering approaches may be limited by their inability to capture latent dietary structures. This study compared three dimensionality reduction techniques—Principal Component Analysis (PCA), Uniform Manifold Approximation and Projection (UMAP), and Autoencoders (AE)—for dietary pattern development, and further examined associations of AE-derived dietary patterns with cancer incidence in a large prospective cohort study. Methods Data were obtained from 130,472 participants enrolled in the Health Examinees-Gem (HEXA-G) study (2004-2013), who completed a validated food frequency questionnaire. PCA, UMAP, and AE were each applied prior to k-means clustering. Cluster quality was assessed using silhouette coefficients, and variable contributions were evaluated using SHAP values. External validation was conducted by applying the HEXA-trained encoder to the Korean National Health and Nutrition Examination Survey (KNHANES). Cancer incidence was ascertained through linkage with the Korea Central Cancer Registry up to December 31, 2018. Multivariable Cox proportional hazards models estimated hazard ratios (HRs) and 95% confidence intervals (CIs) for total and site-specific cancers, focusing on the seven most common cancers in Korea. Results Without dimensionality reduction, the silhouette coefficient was 0.05; PCA rarely exceeded 0.2, UMAP reached ∼0.4, and AE achieved >0.35, providing competitive cluster quality with the most balanced variable contributions. Ten dietary patterns were identified: Balanced, Selective, Rice, Bread, Vegetables, Dairy, Meat, Processed meat, Noodles, and Salty. External validation using KNHANES produced similar silhouette values (∼0.36) and preserved centroid positions, confirming transferability. Over a median follow-up of 9.4 years, 7,390 cancer cases occurred. No significant associations were observed for total cancer; however, site-specific analyses revealed that the Processed meat pattern in men was associated with higher colorectal cancer risk (HR = 1.98, 95% CI: 1.12-3.49), and the Selective pattern with higher gastric cancer risk (HR = 1.32, 95% CI: 1.03-1.70) compared to the Balanced pattern. In women, the Bread pattern was associated with lower gastric cancer risk (HR = 0.53, 95% CI: 0.32-0.89). Conclusion Among the dimensionality reduction techniques, AE achieved the most favorable balance of cluster quality and variable contribution balance, supporting its utility for developing dietary patterns. These findings demonstrate that machine learning-based dimensionality reduction methods, particularly AEs, can strengthen dietary pattern development and capture meaningful associations with cancer risk. Citation Format: Hyobin Lee, Dongseok Heo, Sukhong Min, Sinyoung Cho, So-Yoon Lee, Ji-Yeob Choi, Bongwon Suh, Daehee Kang. Comparison of machine learning-based dimensionality reduction methods for dietary patterns and their predictability of cancer risk in a large cohort study [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2026; Part 1 (Regular Abstracts); 2026 Apr 17-22; San Diego, CA. Philadelphia (PA): AACR; Cancer Res 2026;86(7 Suppl):Abstract nr 5053.
Supplementary Table 3: Pooled relative risks for recategorized age at menarche and age at menopause & incident thyroid cancer risk, Overall and papillary type
Supplementary Table 2: Distribution of total cases according to histology according to participating cohorts
Polygenic risk scores (PRSs), which quantify inherited susceptibility to complex traits and diseases, have emerged as valuable tools for risk stratification and precision medicine. Despite their promise, PRS developed on European cohorts often demonstrate substantially reduced predictive accuracy in non-European populations, due to differences in genetic architecture. The disproportionate representation of European ancestry cohorts in genome-wide association studies (GWAS) leads to inequitable deployment of PRS technologies across diverse populations. Here, we introduce PRANA (Polygenic Risk Adaptation via Neural-network Architecture), a deep learning framework that adapts an existing PRS developed on one population to other ancestries. Unlike methods that require large-scale GWAS in the target population, PRANA leverages pre-trained PRS models derived from European cohorts and adapts them using modestly sized cohorts from the target population. We evaluated PRANA on seven complex traits in South Asian, East Asian and Ashkenazi Jewish populations, as well as in selected smaller East Asian subpopulations where the scarcity of training data poses a particular challenge. PRANA mostly improved predictive performance of the baseline PRS models by 5%-20% in terms of effect size (β) and Nagelkerke's R2, and, in most cases, outperformed existing cross-ancestry multi-PRS approaches. These results highlight PRANA as a scalable and practical strategy to reduce disparities in genomic risk prediction and advance the equitable application of PRS in diverse populations.
Supplementary Figure 3: Forest plots of the pooled hazard ratios (HRs) and 95% confidence intervals (CIs) generated by combining cohort-specific HRs for the association between reproductive factors and the overall risk of thyroid cancer in the Asia Cohort Consortium. A - Forest plot for the pooled HRs and CIs for breastfeeding status and thyroid cancer risk, overall B - Forest plot for the pooled HRs and CIs for postmenopausal status and thyroid cancer risk, overall C - Forest plot for the pooled HRs and CIs for age at menopause and thyroid cancer risk, overall
Supplementary Methods 1: Details on the development of the Asia Cohort Consortium reproductive factor working group protocol
Abstract Background: Colorectal cancer (CRC) remains one of the most common cancers in Korea, underscoring the importance of identifying modifiable risk factors. Although insulin resistance has been implicated in CRC development, existing evidence remains inconsistent, and direct measures of insulin resistance are not routinely collected in clinical practice. This limitation highlights the potential utility of simple proxy markers of insulin resistance in epidemiologic and clinical settings. In this study, we investigated the associations between several proxy insulin resistance markers and CRC risk among Korean adults. Methods: Using data from the Korean Genome and Epidemiology Study Health Examinee cohort, we evaluated the associations between several proxy markers of insulin resistance, including the TG/HDL ratio, triglyceride-glucose index (TyG), TyG-BMI index, TyG-waist circumference index (TyG-WC), TyG-waist-to-height ratio index (TyG-WHTR), and the metabolic score for insulin resistance (METS-IR), and CRC risk using Cox regression models. Subgroup analyses were stratified by sex, age, diabetes status, and prior screening experience, and sensitivity analyses were conducted based on varying follow-up durations. Results: During a median follow-up period of 9.3 years, 795 new CRC cases were observed among 106,965 Koreans aged 40-69 years (36,899 men and 70,066 women). For the TG/HDL ratio, individuals in the highest quartile had a significantly elevated CRC risk compared with the lowest quartile (HR: 1.23, 95% CI: 1.00-1.53). A similar pattern was observed for the TyG and TyG-WC indices, where quartile 4 was associated with increased CRC incidence (TyG Q4 HR: 1.30, 95% CI: 1.04-1.63; TyG-WC Q4 HR: 1.36, 95% CI: 1.01-1.83). Finally, the METS-IR index showed a graded association, and the highest quartile was significantly associated with CRC incidence (Q4 HR: 1.26, 95% CI: 1.02-1.57). Conclusions: In this large population-based cohort, multiple surrogate markers of insulin resistance were independently associated with an increased risk of colorectal cancer, with the strongest effects observed in the highest quartiles. These findings indicate that metabolic dysregulation related to insulin resistance contributes to colorectal carcinogenesis and that simple lipid-glucose-anthropometric markers may help identify individuals at elevated risk. Citation Format: Sukhong Min, Hyobin Lee, Sinyoung Cho, So-Yoon Lee, Jeongheon Kim, Ji-Yeob Choi, Daehee Kang. Association between proxy markers of insulin resistance and colorectal cancer risk: results from a large scale prospective cohort of Korean adults [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2026; Part 1 (Regular Abstracts); 2026 Apr 17-22; San Diego, CA. Philadelphia (PA): AACR; Cancer Res 2026;86(7 Suppl):Abstract nr 924.
Supplementary Figure 4: Forest plots of the pooled hazard ratios (HRs) and 95% confidence intervals (CIs) generated by combining cohort-specific HRs for the association between reproductive factors and the overall risk of thyroid cancer in the Asia Cohort Consortium. A - Forest plot for the pooled HRs and CIs for oral contraceptive use and thyroid cancer risk, overall B - Forest plot for the pooled HRs and CIs for hormone replace therapy use and thyroid cancer risk, overall
Supplementary Figure 5: Forest plot of stratified analysis between number of children/deliveries and thyroid cancer risk by birth years in the Asia Cohort Consortium. The Pooled Hazard Ratios (HRs) with 95% Confidence intervals (CIs) were generated by combining cohort-specific HRs. Models were adjusted for smoking status, alcohol drinking status and Body mass index. a Significant (p-value <0.05) trend across categories of the reproductive factor. b Significant (p-value <0.05) for interaction indicating a modifying effect on the association between the reproductive factor and thyroid cancer risk.
Supplementary Figure 6: Forest plot of pooled hazard ratios (HRs) and 95% confidence intervals (CIs) for the association between reproductive factors and thyroid cancer risk, by age of diagnosis in the Asia Cohort Consortium. The Pooled Hazard Ratios (HRs) with 95% Confidence intervals (CIs) were generated by combining cohort-specific HRs. Models were adjusted for smoking status, alcohol drinking status and Body mass index. a Significant (p-value <0.05). b The model included all 9 cohorts. c The model for Breastfeeding included 6 cohorts, that for Oral contraceptive use included 5 cohorts and that for hormone replacement therapy included 6 cohorts.
Genome-wide association studies (GWAS) have identified numerous genetic variants linked to breast cancer risk, but most discoveries come from European populations, limiting their applicability to other populations. Here, we show that the choice of genotype imputation reference panel, an essential step for GWAS, affects variant detection in Asian populations. Using two large breast cancer datasets from the Breast Cancer Association Consortium (n = 38 954 Asian samples), we compared the 1000 Genomes (1KG) reference panel with SG10K_Health (SG10K), an Asian-specific panel. SG10K imputed more rare variants and achieved higher accuracy for rare alleles (MAF < 0.001), while 1KG performed better for common variants in some contexts. Differences in panel performance influenced association signals, including breast cancer candidate loci such as FGFR2, TOX3, and ESR1. Together, these findings support the use of population-specific imputation panels as a means to improve variant discovery in underrepresented populations.
This study aimed to determine the association between trajectories of obesity status and prediabetes reversion to normoglycemia or progression to diabetes. The study included 14,452 participants from the National Health Insurance Service-National Health Screening (NHIS-HEALS) cohort who continuously had prediabetes glycemic status during the index period (2002-2008), defined by their fasting plasma glucose. The exposure of the study was the trajectories of obesity (defined by body mass index ≥ 25 kg/m2) generated using latent class growth analysis. The outcomes were reversion to normoglycemia or progression to diabetes, whichever occurred first during the follow-up period (2009-2016). The association between trajectories and changes in prediabetes status were examined using cause-specific hazard regression by obtaining the hazard ratio (HR) with a 95% CI. We identified three distinct trajectories which were "Stable obese", "Stable non-obese" and "Obese to non-obese". After a median follow-up of 2 years, 51.99% of participants had their glycemic status back to normoglycemia and 32.17% developed diabetes. Compared with participants in the "Stable obese" group, those in "Stable non-obese" and "Obese to non-obese" groups were more likely to have reversion to normoglycemia (HR with a 95% CI = 1.30 [1.23-1.37] and 1.15 [1.07-1.24], respectively) and lower risk of developing diabetes (0.78 [0.73-0.84] and 0.90 [0.82-0.98], respectively). The findings suggest that maintaining or achieving a non-obese status is linked to higher reversion to normoglycemia as well as lower risks of developing diabetes.
Supplementary Figure 1: Forest plots of the pooled hazard ratios (HRs) and 95% confidence intervals (CIs) generated by combining cohort-specific HRs for the association between reproductive factors and the overall risk of thyroid cancer in the Asia Cohort Consortium. A - Forest plot for the pooled HRs and CIs for age at menarche and thyroid cancer risk, overall B - Forest plot for the pooled HRs and CIs for age at first delivery and thyroid cancer risk, overall
Supplementary Figure 2: Forest plots of the pooled hazard ratios (HRs) and 95% confidence intervals (CIs) generated by combining cohort-specific HRs for the association between reproductive factors and the overall risk of thyroid cancer in the Asia Cohort Consortium. A - Forest plot for the pooled HRs and CIs for parity status and thyroid cancer risk, overall B - Forest plot for the pooled HRs and CIs for number of children/deliveries and thyroid cancer risk, overall C - Forest plot for the pooled HRs and CIs for recategorized number of children/deliveries and thyroid cancer risk, overall
Genome-wide association studies (GWAS) have identified over 200 genetic risk loci for breast cancer, yet the target genes in these loci remain largely unknown. To address this knowledge gap, we conducted a series of multi-ancestry transcriptome-wide association studies (TWAS) to discover potential breast cancer susceptibility genes. We developed and validated ancestry-specific genetic models to predict levels of gene expression, alternative splicing, and 3' UTR alternative polyadenylation, using genomic and transcriptomic data from normal breast tissue samples of 652 females of African, Asian, or European ancestry. These models were then applied to GWAS data of 178,534 breast cancer cases and 248,300 controls from these ancestry groups for association analyses. We identified 290 genes associated with breast cancer risk, including 103 previously unreported in TWAS and 46 located at least 500Kb away from any previously identified risk variants. Among them, 39 genes exhibited distinct associations with breast cancer risk by estrogen receptor status. The identified genes were enriched in pathways related to homologous recombination, apoptosis, p53, PI3K/AKT/mTOR, estrogen, and IL-2/STAT5 signaling. Single-cell RNA sequencing and in vitro experiment data provided additional functional evidence for 169 genes. Our study uncovered large numbers of candidate breast cancer susceptibility genes and contributed valuable insights into the genetics and biology of this common cancer.
Abstract Background: The prognostic utility of the fatty liver index (FLI, a steatosis index derived from BMI, waist circumference, triglycerides, and GGT) and AST/ALT (De Ritis) ratio for hepatocellular carcinoma (HCC) risk in low-risk Asian populations is not well defined. We evaluated their independent and incremental predictive value in a large Korean cohort. Methods: We analyzed 43,981 Korean men (376 HCC cases) from the Health Examinees-Gem (HEXA-G) cohort (2004-2013). A low-risk subcohort (n = 39,033) was defined by excluding individuals with diabetes or chronic hepatitis. Multivariable Cox models assessed associations with log-transformed FLI and the AST/ALT ratio after adjusting for demographic, lifestyle, socioeconomic, and metabolic factors. Incremental predictive value was evaluated using likelihood ratio tests (LRTs), changes in C-statistics, and Bayesian Information Criterion (BIC). Sensitivity analyses using penalized regression produced similar results. Results: In univariable analyses, the AST/ALT ratio showed crude associations with HCC, whereas FLI did not. After adjustment, this pattern reversed: the AST/ALT ratio became non-predictive—consistent with confounding by age, alcohol consumption, smoking, and metabolic factors—while log(FLI) emerged as a modest independent predictor in the low-risk subcohort (adjusted HR 1.08; 95% CI, 1.01-1.16; p = 0.04). Discrimination improved minimally (C-statistic 0.715 to 0.719), and conventional FLI categories failed to stratify risk. In the full cohort, log(FLI) remained independently associated with HCC (adjusted HR 1.11; 95% CI, 1.03-1.20; p = 0.009) and modestly improved model fit (LRT p = 0.011) without meaningfully improving discrimination (C-statistic 0.795 to 0.797). The De Ritis ratio added no independent or incremental value in any model. Conclusions: Across low-risk and mixed-risk Korean men, FLI remained an independent predictor of HCC after adjustment, although the effect size was modest. In contrast, the AST/ALT ratio lost all prognostic value after accounting for demographic, lifestyle, and metabolic confounding. The reversal of crude versus adjusted associations underscores substantial confounding for AST/ALT and a small, metabolically related signal for FLI. Overall, these findings highlight that FLI provides some independent information in low-risk settings, but substantial improvement in HCC risk prediction will require more robust biomarkers. Citation Format: So Yoon Lee, Hyobin Lee, Sukhong Min, Sinyoung Cho, Jeongheon Kim, Ji-Yeob Choi, Daehee Kang. Fatty liver index and AST/ALT ratio for hepatocellular carcinoma prediction in low-risk Korean men: Results from the HEXA-G cohort [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2026; Part 1 (Regular Abstracts); 2026 Apr 17-22; San Diego, CA. Philadelphia (PA): AACR; Cancer Res 2026;86(7 Suppl):Abstract nr 7585.
Supplementary Figure 7: Forest plot of stratified analyses between reproductive factors and thyroid cancer risk by body mass index (BMI) and smoking status in the Asia Cohort Consortium. The Pooled Hazard Ratios (HRs) with 95% Confidence intervals (CIs) were generated by combining cohort-specific HRs. Models for stratified analyses by smoking status were adjusted for alcohol drinking status and BMI and those for BMI were adjusted for smoking status and alcohol drinking status. a Significant (p-value <0.05). b Significant (p-value <0.05) for interaction indicating a modifying effect on the association between the reproductive factor and thyroid cancer risk. c The model included all 9 cohorts. d The model for Breastfeeding included 6 cohorts, that for Oral contraceptive use included 5 cohorts, and that for hormone replacement therapy included 6 cohorts