Metabolic dysfunction-associated steatotic liver disease (MASLD), formerly known as non-alcoholic fatty liver disease (NAFLD), exhibits marked racial disparities in incidence, progression, and clinical outcomes. While diet, lifestyle, and socioeconomic factors have been shown to influence these disparities, biological mechanisms underlying racial differences in MASLD risk remain poorly understood. We hypothesized that race-associated variation in hepatic gene expression may contribute to differential susceptibility and progression of MASLD. To test this hypothesis, we analyzed publicly available gene expression data from liver biopsies obtained from over 300 Black and White individuals undergoing bariatric surgery. Gene expression profiles were compared across four histological stages: normal liver, MASLD, metabolic dysfunction-associated steatohepatitis (MASH) without fibrosis, and MASH with fibrosis. We identified more than 200 genes that were significantly differentially expressed between Black and White individuals. Genes associated with MASLD progression were significantly enriched among race-specific genes, supporting the hypothesis that racial differences in hepatic gene expression contribute to disease risk and progression. Using histologically normal liver as a reference, we identified race-specific candidate genes potentially driving MASLD progression. These included UCN3 and PRSS3 in Black individuals, and MMP15, LAMB2, LEPR, ELOVL2, CD48, COL5A2, and ICAM1 in White individuals. Notably, divergence in gene expression profiles between racial groups became more pronounced with advancing disease stages, suggesting that race may play an increasingly important role in later phases of MASLD progression. Our findings indicate that differential modulation of hepatic gene expression represents a potential biological mechanism contributing to racial disparities in MASLD. These results highlight the importance of considering race-specific molecular signatures in understanding MASLD pathogenesis and in developing targeted prevention and therapeutic strategies.
Loss-of-function (LoF) mutations in genes essential for cancer cell survival and proliferation are expected to be selectively depleted in tumors; therefore, a deficit of LoF mutations can serve as a marker of cancer-specific essentiality. Using somatic mutation data from the Catalogue Of Somatic Mutations In Cancer (COSMIC), we analyzed lung tumors and modeled gene-level LoF mutation counts (nonsense, frameshift, and canonical splice-site variants) using Poisson regression. The number of LoF mutations per gene was modeled as a function of the number of potential LoF sites and gene-specific mutability, estimated from synonymous mutation rates. Deviance residuals were used to identify genes with significant deficits (negative outliers) or excesses (positive outliers) of LoF mutations relative to expectation. We observed a strong association between deviance residuals and prior evidence of cancer relevance, with a U-shaped relationship indicating that both LoF-depleted and LoF-enriched genes are functionally important in tumorigenesis. We estimated the degree of evolutionary constraint on these outlier genes using allele frequencies from the Genome Aggregation Database (gnomAD) consortium data. Focusing on genes with significant LoF depletion and low evolutionary constraint enabled the identification of potential candidate lung cancer–essential genes as potential therapeutic targets for lung cancer treatment.
INTRODUCTION:Rare, deleterious germline variants are key contributors to inherited lung cancer (LC) risk. The Genetic Epidemiology of LC Consortium (GELCC) has curated valuable high-risk LC families and is uniquely positioned to uncover rare, high-penetrance variants underlying familial LC (FLC). METHODS:We performed whole-genome and exome sequencing on germline DNA from 120 high-risk LC families (177 FLC cases, 309 unaffected relatives). We prioritized rare (allele frequency <1% in the genome aggregation database), potentially deleterious variants present in two or more FLC cases. These variants were then validated in 10,085 sporadic LC (SLC) cases and 612,970 controls. RESULTS:We identified 118 candidate variants, 28 of which were validated in SLC with strong statistical support. We discovered a novel pathogenic axis of three truncating variants in GALNT6, MUC4, and ERBB3 genes, which are critical regulators of mucin-type O-glycosylation. Nine top hits were mapped to the known 6q23-25 linkage region (ROS1, LAMA2, PRKN, SYNE1). Other candidates were clustered in DNA repair (ATM, BRCA2, MLH1), oncogenic signaling (ERBB3, JAK1, PIM1), and extracellular matrix genes (COL6A3, FLG). Carriers of two or more variant alleles had a strong dose-dependent risk. Furthermore, gene-based burden tests revealed strong associations between RARB, MGMT, and EBF1 with FLC susceptibility. CONCLUSION:Our findings underscore the important role of rare, high-penetrance genetic variants in FLC susceptibility, particularly in mucin glycosylation and DNA repair genes. These findings offer promising targets for early detection and personalized therapies.
BACKGROUND:Lung adenocarcinoma (LUAD) in never-smokers is a major public health burden, especially among East Asian women. Polygenic risk scores (PRSs) are promising for risk stratification but are primarily developed in European-ancestry populations. We aimed to develop and validate single- and multi-ancestry PRSs for East Asian never-smokers to improve LUAD risk prediction. METHODS:PRSs were developed using genome-wide association study summary statistics from East Asian (8,002 cases; 20,782 controls) and European (2,058 cases; 5,575 controls) populations. Single-ancestry models included PRS-25, PRS-CT, and LDpred2; multi-ancestry models included LDpred2+PRS-EUR128, PRS-CSx, and CT-SLEB. Performance was evaluated in independent East Asian data from the Female Lung Cancer Consortium (FLCCA) and externally validated in the Nanjing Lung Cancer Cohort (NJLCC). We assessed predictive accuracy via AUC, with 10-year and (age 30-80) absolute risks estimates. RESULTS:The best multi-ancestry PRS, using East Asian and European data via CT-SLEB (clumping and thresholding, super learning, empirical Bayes), outperformed the best East Asian-only PRS (LDpred2; AUC = 0.629, 95% CI:0.618,0.641), achieving an AUC of 0.640 (95% CI : 0.629,0.653) and odds ratio of 1.71 (95% CI : 1.61,1.82) per SD increase. NJLCC Validation confirmed robust performance (AUC =0.649, 95% CI: 0.623, 0.676). The top 20% PRS group had a 3.92-fold higher LUAD risk than the bottom 20%. Further, the top 5% PRS group reached a 6.69% lifetime absolute risk. Notably, this group reached the average population 10-year LUAD risk at age 50 (0.42%) by age 41, nine years earlier. CONCLUSIONS:Multi-ancestry PRS approaches enhance LUAD risk stratification in East Asian never-smokers, with consistent external validation, suggesting future clinical utility.
PURPOSE:Patients with stage II and resected stage III melanomas have variable clinical outcomes, providing evidence of underlying biological differences in tumors and/or the patients themselves, beyond stage. The approval of adjuvant immunotherapy for stage IIB/C and resected stage III/IV disease (and adjuvant targeted therapy for resected stage III disease) has created a pressing need to develop biomarkers to accurately distinguish patients at low risk versus high risk for recurrence and death from melanoma. miRNAs are promising biomarkers because of their stability in tissues and fluids and their demonstrated functional and prognostic roles in melanoma. We hypothesized that miRNA expression could be integrated into prognostic models that would classify 5-year survival outcomes more accurately than clinical factors alone. EXPERIMENTAL DESIGN:Using a NanoString miRNA Expression Assay, we analyzed 715 primary melanomas from patients with stage II or stage III disease within the InterMEL consortium and examined associations between miRNA expression and melanoma-specific death. RESULTS:When integrated into clinical prognostic models for 5-year melanoma-specific survival, miRNA signatures improved the area under the receiver operating characteristic curve for patients in stage II from 0.71 for a "clinical factors-only" model to 0.81 for a "clinical plus miRNA" model in an independent test set, an improvement of 0.10 with a 95% confidence interval (0.03-0.19). The improvement was more modest for patients in stage III who were included in the analysis. CONCLUSIONS:Incorporating miRNA expression in primary melanomas may enhance the accuracy of clinical prognostic models and potentially aid in the selection of patients with melanoma for adjuvant treatment and clinical trials.
Lung adenocarcinoma (LUAD) has several histologically distinct subtypes that differ by a number of clinical features including patient survival. Molecular mechanisms underlying histological and clinical differences between subtypes remain poorly understood. We conducted a comparative analyses of gene expression in acinar, lepidic, papillary and solid subtypes, as well as mucinous adenocarcinoma. We used a novel, more efficient approach to identify subtype-specific genes. We compared the mean gene expression level separately for tumors and adjacent normal tissue with pure or a highly represented (≥ 75
Missing gene expression values are a common issue in RNAseq-based analyses of gene expression. However, an analysis of genetic and environmental factors contributing to data missingness in RNAseq-based assessment of gene expression has never been conducted. In this study we tried to identify factors in RNAseq data missingness. We used RNAseq data from 66 lung adenocarcinoma tumors and corresponding adjacent normal lung tissues. We found a strong negative association between the gene expression level and missingness, supporting the idea that the borderline expression level is a key contributor to missingness. In a more detailed analysis, the relationship between gene expression and missingness was more complex: while the expected negative association between missingness and the expression level was observed for genes with low missingness, mean expression spiked at the right end of the distribution which included genes with very high missingness. We hypothesized that genes with a high missing rate include not only genes with borderline expression but also genes with high expression in some individuals but no expression in others (true biological missingness, TBM). The results of the comparative analysis of missingness in smokers and nonsmokers, an examination of the proportion of known tobacco smoke-sensitive genes by missing rate, and gene enrichment analysis support the hypothesis. We argue that it would be beneficial first to check data for the presence of genes with true biological missingness. The presence of highly expressed genes with missingness is an indication of TBM related to inter-individual variation in gene expression level. The results of our analysis call for caution in indiscriminatory imputation of missing values. When true biological missingness is present, it is advisable to identify genes with true biological missingness and analyze them separately because including such genes in imputation will lead to a bias: expression values will be assigned to a subset of the genes that are not expressed.
Epigenome-wide association studies (EWAS) profile DNA methylation across the human genome to identify associations with diseases and exposures. Most employ Illumina methylation arrays; this platform, however, under-samples interindividual epigenetic variation. The systemic and stable nature of epigenetic variation at correlated regions of systemic interindividual variation (CoRSIVs) should be advantageous to EWAS. Here, we analyze 2,203 published EWAS to determine whether Illumina probes within CoRSIVs are over-represented in the literature. Enrichment of CoRSIV-overlapping probes was observed for most classes of disease, particularly for neurodevelopmental disorders and type 2 diabetes, indicating an opportunity to improve the power of EWAS by over 200- and over 100-fold, respectively. EWAS targeting all known CoRSIVs should accelerate discovery of associations between individual epigenetic variation and risk of disease.
Supplementary Table S7 shows the signals from interaction with smoking status for the identified variants in lung cancer.
Ideally, detection of somatic mutations in a tumor is accomplished using a patient-matched sample of normal cells as the benchmark. In this way somatic mutations can be distinguished from rare germline mutations. In large retrospective studies, archival tissue collection can pose challenges in obtaining samples of normal DNA. In this article we propose a protocol that improves somatic mutation analysis in the absence of a matched normal sample. The method was motivated by the InterMEL study, a large-scale epidemiologic investigation involving multiomic, multi-institutional genomic profiling of 1000 primary melanoma samples. The key insight for accomplishing improved mutation calling is the fact that germline mutations should produce a variant allele frequency (VAF) of around 50%. While a similar VAF of 50% would also be expected for somatic mutations in pure tumor samples, typically the tumor purity is much less than 50%, resulting in a considerably lower VAF. Making use of a technique that can simultaneously estimate both tumor purity and VAF from tumor-only samples we have developed a method for better distinguishing somatic versus germline variants. Based on 137 melanomas from the InterMEL Study with matched normal tissue to provide a gold standard we show that the conventional pipeline using a panel of (unmatched) normal samples has a false positive rate of 15.6% and a false negative rate of 3.5%. Our new technique improves these error rates to 6.4% and 2.1%, respectively.
Supplementary Table S1 displays the number of ever- and never-smokers from each of the ten studies.
Single nucleotide substitutions are the most common type of somatic mutations in cancer genome. The goal of this study was to use publicly available somatic mutation data to quantify negative and positive selection in individual lung tumors and test how strength of directional and absolute selection is associated with clinical features. The analysis found a significant variation in strength of selection (both negative and positive) among tumors, with median selection tending to be negative even though tumors with strong positive selection also exist. Strength of selection estimated as the density of missense mutations relative to the density of silent mutations showed only a weak correlation with tumor mutation burden. In the "all histology together" analysis we found that absolute strength of selection was strongly correlated with all clinically relevant features analyzed. In histology-stratified analysis selection was strongest in small cell lung cancer. Selection in adenocarcinoma was somewhat higher compared to squamous cell carcinoma. The study suggests that somatic mutation- based quantifying of directional and absolute selection in individual tumors can be a useful biomarker of tumor aggressiveness.
Supplementary Figure S4 shows the results of eQTL analysis of rs968516 in GTEx.
Supplementary Table S4 shows the imputation quality score of identified variants in different studies.
PURPOSE Patients with stage II and III cutaneous primary melanoma vary considerably in their risk of melanoma-related death. We explore the ability of methylation profiling to distinguish primary melanoma methylation classes and their associations with clinicopathologic characteristics and survival. MATERIALS AND METHODS InterMEL is a retrospective case-control study that assembled primary cutaneous melanomas from American Joint Committee on Cancer (AJCC) 8th edition stage II and III patients diagnosed between 1998 and 2015 in the United States and Australia. Cases are patients who died of melanoma within 5 years from original diagnosis. Controls survived longer than 5 years without evidence of melanoma recurrence or relapse. Methylation classes, distinguished by consensus clustering of 850K methylation data, were evaluated for their clinicopathologic characteristics, 5-year survival status, and differentially methylated gene sets. RESULTS Among 422 InterMEL melanomas, consensus clustering revealed three primary melanoma methylation classes (MethylClasses): a CpG island methylator phenotype (CIMP) class, an intermediate methylation (IM) class, and a low methylation (LM) class. CIMP and IM were associated with higher AJCC stage (both P = .002), Breslow thickness (CIMP P = .002; IM P = .006), and mitotic index (both P < .001) compared with LM, while IM had higher N stage than CIMP ( P = .01) and LM ( P = .007). CIMP and IM had a 2-fold higher likelihood of 5-year death from melanoma than LM (CIMP odds ratio [OR], 2.16 [95% CI, 1.18 to 3.96]; IM OR, 2.00 [95% CI, 1.12 to 3.58]) in a multivariable model adjusted for age, sex, log Breslow thickness, ulceration, mitotic index, and N stage. Despite more extensive CpG island hypermethylation in CIMP, CIMP and IM shared similar patterns of differential methylation and gene set enrichment compared with LM. CONCLUSION Melanoma MethylClasses may provide clinical value in predicting 5-year death from melanoma among patients with primary melanoma independent of other clinicopathologic factors.
Supplementary Table S5 shows the signals for the variants in stratified analysis by smoking behavior
Supplementary Figure S3 displays the Validation of the association effect of rs6757077 at IKZF2 using data from six independent study sites in China.
BACKGROUND:Although the associations between genetic variations and lung cancer risk have been explored, the epigenetic consequences of DNA methylation in lung cancer development are largely unknown. Here, the genetically predicted DNA methylation markers associated with non-small cell lung cancer (NSCLC) risk by a two-stage case-control design were investigated. METHODS:The genetic prediction models for methylation levels based on genetic and methylation data of 1595 subjects from the Framingham Heart Study were established. The prediction models were applied to a fixed-effect meta-analysis of screening data sets with 27,120 NSCLC cases and 27,355 controls to identify the methylation markers, which were then replicated in independent data sets with 7844 lung cancer cases and 421,224 controls. Also performed was a multi-omics functional annotation for the identified CpGs by integrating genomics, epigenomics, and transcriptomics and investigation of the potential regulation pathways. RESULTS:Of the 29,894 CpG sites passing the quality control, 39 CpGs associated with NSCLC risk (Bonferroni-corrected p ≤ 1.67 × 10-6 ) were originally identified. Of these, 16 CpGs remained significant in the validation stage (Bonferroni-corrected p ≤ 1.28 × 10-3 ), including four novel CpGs. Multi-omics functional annotation showed nine of 16 CpGs were potentially functional biomarkers for NSCLC risk. Thirty-five genes within a 1-Mb window of 12 CpGs that might be involved in regulatory pathways of NSCLC risk were identified. CONCLUSIONS:Sixteen promising DNA methylation markers associated with NSCLC were identified. Changes of the methylation level at these CpGs might influence the development of NSCLC by regulating the expression of genes nearby. PLAIN LANGUAGE SUMMARY:The epigenetic consequences of DNA methylation in lung cancer development are still largely unknown. This study used summary data of large-scale genome-wide association studies to investigate the associations between genetically predicted levels of methylation biomarkers and non-small cell lung cancer risk at the first time. This study looked at how well larotrectinib worked in adult patients with sarcomas caused by TRK fusion proteins. These findings will provide a unique insight into the epigenetic susceptibility mechanisms of lung cancer.
Supplementary Table S2 shows the signals at known susceptibility loci from ever- and never-smoking lung cancer.