Supplementary Figure from Beyond GWAS of Colorectal Cancer: Evidence of Interaction with Alcohol Consumption and Putative Causal Variant for the 10q24.2 Region
Post-Acute Sequelae of SARS-CoV-2 infection (PASC or “Long COVID”), includes numerous chronic conditions associated with widespread morbidity and rising healthcare costs. PASC has highly variable clinical presentations, and likely includes multiple molecular subtypes, but it remains poorly understood from a molecular and mechanistic standpoint. This hampers the development of rationally targeted therapeutic strategies. The NIH-sponsored “Researching COVID to Enhance Recovery” (RECOVER) initiative includes several retrospective/prospective observational cohort studies enrolling adult, pregnant adult and pediatric patients respectively. RECOVER formed an “OMICS” multidisciplinary task force, including clinicians, pathologists, laboratory scientists and data scientists, charged with developing recommendations to apply cutting-edge system biology technologies to achieve the goals of RECOVER. The task force met biweekly over 14 months, to evaluate published evidence, examine the possible contribution of each “omics” technique to the study of PASC and develop study design recommendations. The OMICS task force recommended an integrated, longitudinal, simultaneous systems biology study of participant biospecimens on the entire RECOVER cohorts through centralized laboratories, as opposed to multiple smaller studies using one or few analytical techniques. The resulting multi-dimensional molecular dataset should be correlated with the deep clinical phenotyping performed through RECOVER, as well as with information on demographics, comorbidities, social determinants of health, the exposome and lifestyle factors that may contribute to the clinical presentations of PASC. This approach will minimize lab-to-lab technical variability, maximize sample size for class discovery, and enable the incorporation of as many relevant variables as possible into statistical models. Many of our recommendations have already been considered by the NIH through the peer-review process, resulting in the creation of a systems biology panel that is currently designing the studies we proposed. This system biology strategy, coupled with modern data science approaches, will dramatically improve our prospects for accurate disease subtype identification, biomarker discovery and therapeutic target identification for precision treatment. The resulting dataset should be made available to the scientific community for secondary analyses. Analogous system biology approaches should be built into the study designs of large observational studies whenever possible.
BACKGROUND:Obesity has been positively associated with most molecular subtypes of colorectal cancer (CRC); however, the magnitude and the causality of these associations is uncertain. METHODS:We used Mendelian randomization (MR) to examine potential causal relationships between body size traits (body mass index [BMI], waist circumference, and body fat percentage) with risks of Jass classification types and individual subtypes of CRC (microsatellite instability [MSI] status, CpG island methylator phenotype [CIMP] status, BRAF and KRAS mutations). Summary data on tumour markers were obtained from two genetic consortia (CCFR, GECCO). FINDINGS:A 1-standard deviation (SD:5.1 kg/m2) increment in BMI levels was found to increase risks of Jass type 1MSI-high,CIMP-high,BRAF-mutated,KRAS-wildtype (odds ratio [OR]: 2.14, 95% confidence interval [CI]: 1.46, 3.13; p-value = 9 × 10-5) and Jass type 2non-MSI-high,CIMP-high,BRAF-mutated,KRAS-wildtype CRC (OR: 2.20, 95% CI: 1.26, 3.86; p-value = 0.005). The magnitude of these associations was stronger compared with Jass type 4non-MSI-high,CIMP-low/negative,BRAF-wildtype,KRAS-wildtype CRC (p-differences: 0.03 and 0.04, respectively). A 1-SD (SD:13.4 cm) increment in waist circumference increased risk of Jass type 3non-MSI-high,CIMP-low/negative,BRAF-wildtype,KRAS-mutated (OR 1.73, 95% CI: 1.34, 2.25; p-value = 9 × 10-5) that was stronger compared with Jass type 4 CRC (p-difference: 0.03). A higher body fat percentage (SD:8.5%) increased risk of Jass type 1 CRC (OR: 2.59, 95% CI: 1.49, 4.48; p-value = 0.001), which was greater than Jass type 4 CRC (p-difference: 0.03). INTERPRETATION:Body size was more strongly linked to the serrated (Jass types 1 and 2) and alternate (Jass type 3) pathways of colorectal carcinogenesis in comparison to the traditional pathway (Jass type 4). FUNDING:Cancer Research UK, National Institute for Health Research, Medical Research Council, National Institutes of Health, National Cancer Institute, American Institute for Cancer Research, Brigham and Women's Hospital, Prevent Cancer Foundation, Victorian Cancer Agency, Swedish Research Council, Swedish Cancer Society, Region Västerbotten, Knut and Alice Wallenberg Foundation, Lion's Cancer Research Foundation, Insamlingsstiftelsen, Umeå University. Full funding details are provided in acknowledgements.
Abstract We used deep learning (DL) to predict multiple genetic biomarkers of colorectal cancer (CRC) from routine histopathology images, aiming to identify morphological changes associated with the mutational profile of tumors. A new Transformer-based prediction model was developed which can simultaneously predict multiple biomarkers from a single histopathology image. We compared the performance of the multi-target approach to conventional single-target models. We analyzed five patient cohorts from the Genetics and Epidemiology of CRC Consortium (GECCO) with colorectal cancer (N=1,385 patients). Our model was trained on an internal data set (N=739) to predict the presence of >200 genetic alterations, focusing on genes with non-silent mutations and mutational signatures, which were assessed using targeted sequencing. We used a 7-fold cross validation and validated the model on an independent, external data set (N=646). Morphological features of microsatellite instability (MSI) are detectable by DL models in histopathology images, making MSI a potential confounding factor. Thus, we also assessed whether our model can predict genetic alterations regardless of MSI status, or whether the model identifies the MSI phenotype. For the majority of biomarkers, the multi-target transformer reached a higher prediction performance as measured by the mean Area Under the Receiver Operating Characteristic (AUROC) Curve and lower standard deviation than single-target DL models. For the external validation set, the mean AUROC of the multi-target versus the single-target model was 0.78 (+/- 0.01) versus 0.72 (+/- 0.06) for BRAF mutations, 0.88 (+/- 0.01) versus 0.86 (+/- 0.03) for hypermutation status, 0.94 (+/- 0.01) versus 0.91 (+/- 0.02) for MSI, 0.86 (+/- 0.01) versus 0.80 (+/- 0.05) for RNF43 mutations and 0.72 (+/- 0.02) versus 0.69 (+/- 0.05) for TP53 mutations, respectively. However, our analysis of the individual prediction scores suggests that the model does not identify the phenotype of the alterations themselves. Instead, the model seems to determine the status of specific genetic alterations based on their correlation with MSI. Consequently, biomarkers which are associated with microsatellite stability have prediction scores negatively correlated with predicted MSI scores. Conversely, biomarkers which are associated with MSI have prediction scores positively correlated with predicted MSI scores. Our results demonstrate potential benefits of multi-target DL models in improving the predictive power and efficacy of histology-based biomarker extraction compared to single-target DL models. However, our results crucially highlight the importance of considering potential confounders, such as MSI status in CRC, and emphasize the need for extensive analysis beyond AUROC in studies focusing on DL-based biomarker detection. Citation Format: Marco Gustav, Marko van Treeck, Zunamys I. Carrero, Chiara M. Loeffler, Nic G. Reitsam, Bruno Märkl, Asier Rabasco Meneghetti, Lisa A. Boardman, Amy J. French, Ellen L. Goode, Andrea Gsur, Stefanie Brezina, Marc J. Gunter, Neil Murphy, Paul Limburg, Stephen Thibodeau, Sebastian Foersch, Robert Steinfelder, Tabitha Harrison, Ulrike Peters, Amanda Phipps, Jakob N. Kather. Assessing microsatellite instability dominance in colorectal cancer phenotype: A multi-study initiative using multi-target transformers for genomic biomarker prediction [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2024; Part 1 (Regular Abstracts); 2024 Apr 5-10; San Diego, CA. Philadelphia (PA): AACR; Cancer Res 2024;84(6_Suppl):Abstract nr 4924.
BACKGROUND:Menopausal hormone therapy (MHT), a common treatment to relieve symptoms of menopause, is associated with a lower risk of colorectal cancer (CRC). To inform CRC risk prediction and MHT risk-benefit assessment, we aimed to evaluate the joint association of a polygenic risk score (PRS) for CRC and MHT on CRC risk. METHODS:We used data from 28,486 postmenopausal women (11,519 cases and 16,967 controls) of European descent. A PRS based on 141 CRC-associated genetic variants was modeled as a categorical variable in quartiles. Multiplicative interaction between PRS and MHT use was evaluated using logistic regression. Additive interaction was measured using the relative excess risk due to interaction (RERI). 30-year cumulative risks of CRC for 50-year-old women according to MHT use and PRS were calculated. RESULTS:The reduction in odds ratios by MHT use was larger in women within the highest quartile of PRS compared to that in women within the lowest quartile of PRS (p-value = 2.7 × 10-8). At the highest quartile of PRS, the 30-year CRC risk was statistically significantly lower for women taking any MHT than for women not taking any MHT, 3.7% (3.3%-4.0%) vs 6.1% (5.7%-6.5%) (difference 2.4%, P-value = 1.83 × 10-14); these differences were also statistically significant but smaller in magnitude in the lowest PRS quartile, 1.6% (1.4%-1.8%) vs 2.2% (1.9%-2.4%) (difference 0.6%, P-value = 1.01 × 10-3), indicating 4 times greater reduction in absolute risk associated with any MHT use in the highest compared to the lowest quartile of genetic CRC risk. CONCLUSIONS:MHT use has a greater impact on the reduction of CRC risk for women at higher genetic risk. These findings have implications for the development of risk prediction models for CRC and potentially for the consideration of genetic information in the risk-benefit assessment of MHT use.
Single nucleotide polymorphism (SNP) interactions are the key to improving polygenic risk scores. Previous studies reported several significant SNP–SNP interaction pairs that shared a common SNP to form a cluster, but some identified pairs might be false positives. This study aims to identify factors associated with the cluster effect of false positivity and develop strategies to enhance the accuracy of SNP–SNP interactions. The results showed the cluster effect is a major cause of false-positive findings of SNP–SNP interactions. This cluster effect is due to high correlations between a causal pair and null pairs in a cluster. The clusters with a hub SNP with a significant main effect and a large minor allele frequency (MAF) tended to have a higher false-positive rate. In addition, peripheral null SNPs in a cluster with a small MAF tended to enhance false positivity. We also demonstrated that using the modified significance criterion based on the 3 p-value rules and the bootstrap approach (3pRule + bootstrap) can reduce false positivity and maintain high true positivity. In addition, our results also showed that a pair without a significant main effect tends to have weak or no interaction. This study identified the cluster effect and suggested using the 3pRule + bootstrap approach to enhance SNP–SNP interaction detection accuracy.
Supplementary Table 1 from Origins and Prevalence of the American Founder Mutation of <i>MSH2</i>
Supplementary Tables S1-3. Supplementary Table S1 Description of the characteristics by study. Supplementary Table S2 Associations between confounders and body mass index and the weighted genetic risk score among controls. Supplementary Table S3 Odds ratios (OR) and 95% confidence intervals (CI) for the association between a 1-unit increase in the weighted genetic risk score (the instrumental variable [IV]) and risk of colorectal cancer.
PDF - 112KB, Supplemental Table 1: Cancer-Telomere length association studies. Supplemental Table 2: Validation Set: Differences in Peripheral Blood RTL by DNA Extraction Method (Southern Blots). Supplemental Table 3: Fixed Effects and Variance Estimates of relative telomere length (log T/S ratio) by Extraction Method (Southern Blots).
Documents sent to hospitals to collect tumor samples for the Cancer Prevention Study-II Nutrition Cohort Colorectal Tissue Repository.
Supplementary Figure S1. LocusZoom regional association plots for the seven new cross-cancer loci that were > 1 Mb from known index SNPs. Supplementary Figure S2A-B. Box plots showing eQTL associations between (A) rs9375701 and L3MBTL3 in normal breast and prostate tissues and (B) rs8037137 and RCCD1 in normal breast and ovarian tissues. Supplementary Figure S3. Interactions between BCL2L11 and the 32 Biocarta "Death Pathway" genes. Interactions were identified using the GeneMania server. Circles contain gene names, lines represent interactions, and the color of the line indicates a specific type of interaction as listed in the legend.
Supplementary Figure 2 from Frequent Truncating Mutation of TFAM Induces Mitochondrial DNA Depletion and Apoptotic Resistance in Microsatellite-Unstable Colorectal Cancer
PDF file - 108KB, Predictive value of specific KRAS mutations in patients with BRAF-wild type resected stage III colon cancer treated with adjuvant FOLFOX chemotherapy with vs without cetuximab. Unadjusted hazard ratios (HR) for disease-free survival by KRAS mutation strata are shown for patients who were randomized during the period when KRAS-mutated and -wild type tumors were eligible. HR < 1 indicates benefit from cetuximab. No estimates produced due to insufficient sample size in one subgroup.
Supplementary Table S1. Descriptive factors for KP study population, broken down by study. Supplementary Table S2. Genome-wide significant SNPs found in our cohort. Supplementary Table S3. Cis-eQTL expression of rs4646284. Supplementary Table S4. Results at the 105 loci previously found to be associated with prostate cancer. Supplementary Table S5. Risk score of and variance explained by the 105 previously reported hits. Supplementary Figure S1. Manhattan and Q-Q plots of each race/ethnicity and meta-analysis. Supplementary Figure S2. Local plots of novel replicated rs4646284 and suggestive rs2659124. Supplementary Figure S3. Cis-eQTL of SLC22A1 and SLC22A3. Supplementary Figure S4. Comparison of ORs of KP to previous reports by race/ethnicity. Supplementary Figure S5. KP AUC estimates.
Description of supplementary files - PDF file 44K, File contains description of supplementary files.
Supplementary Figure 1. Forest Plot (Fixed effects model) of the 22 prostate cancer risk associated miRSNPs.