Supplementary Figure S17 shows Kaplan-Meier (KM) survival curves for relapse-free survival stratified by NS-LUAD expression subtypes in (A) patients with stage IA tumors, (B) patients with stage I tumors harboring no EGFR mutations or ALK fusions, or with no clinical records indicating treatment with tyrosine kinase inhibitor (non-TKI subset). P-values and hazard ratios (HR) were calculated using Cox proportional hazards models adjusted for age and sex. The number of samples in each subgroup is indicated on the KM curves. Unless otherwise specified, p-values are based on 10-year overall survival; for comparisons involving the chaotic subtype, 5-year overall survival rates were used owing to its association with advanced tumor stage. For the non-TKI subset, comparisons between the chaotic and steady subtypes were based on 10-year overall survival due to the absence of death events within the 5-year time window.
Despite promising results in using deep learning to infer genetic features from histological whole-slide images (WSIs), no prior studies have specifically applied these methods to lung adenocarcinomas from subjects who have never smoked tobacco (NS-LUAD) – a molecularly and histologically distinct subset of lung cancer. Existing models have focused on LUAD from predominantly smoker populations, with limited molecular scope and variable performance. Here, we propose a customized deep convolutional neural network based on ResNet50 architecture, optimized for multilabel classification for NS-LUAD, enabling simultaneous prediction of 16 molecular alterations from a single H&E-stained WSI. Key architectural modifications included a simplified two-layer residual block without bottleneck layers, selective shortcut connections, and a sigmoid-based classification head for independent prediction of each alteration, designed to reduce computational complexity while maintaining predictive accuracy. The model was trained and evaluated on 495 WSIs from the Sherlock-Lung study (70% training with 10% internal test set for 10-fold cross-validation, and 30% held-out validation set for final evaluation). For the held-out validation data, our model achieved high areas under the receiver operating characteristic curve [AUROC] values =0.84-0.93 for detecting 11 features: EGFR, KRAS, TP53, RBM10 mutations, MDM2 amplification, kataegis, CDKN2A deletion, ALK fusion, whole-genome doubling, and EGFR hotspot mutations (p.L858R and p.E746_A750del). Performance was low to moderate for tumor mutational burden (AUROC=0.67), APOBEC mutational signature (AUROC=0.57), and KRAS hotspot mutations (p.G12C: AUROC=0.74, p.G12V: AUROC=0.55, p.G12D: AUROC=0.43). Compared to results from established architectures such as Inception-v3 on the same WSIs, our model demonstrated significantly improved performance for most features. With further optimization, our model could support triaging for molecular testing and inform precision treatment strategies for NS-LUAD patients.
OBJECTIVE:The association between occupational exposure to chlorinated solvents and lung cancer remains inconclusive. This study investigated this relationship using data from the internationally pooled SYNERGY study. METHODS:Data from 14 case-control studies conducted in 13 European countries and Canada were pooled, including 28 048 participants (12 329 cases and 15 719 controls). Lifetime occupational exposure to chlorinated solvents was assessed using the ALOHA+job-exposure matrix. ORs and 95% CIs were estimated using unconditional logistic regression, adjusted for study centre, age, sex, smoking (pack-years and cessation), cumulative exposure to five occupational lung carcinogens (asbestos, hexavalent chromium, polycyclic aromatic hydrocarbons, respirable crystalline silica and diesel engine exhaust), cumulative benzene exposure and employment in high-risk occupations ('List A' jobs). Associations were estimated across categories of exposure levels, durations and analyses stratified by smoking status and lung cancer subtypes. RESULTS:We found no evidence of an association between ever exposure to chlorinated solvents and lung cancer risk (OR 1.03; 95% CI 0.96 to 1.10). Among exposed individuals, a positive trend with cumulative exposure was observed (p=0.031), but not when non-exposed individuals were included (p=0.173). Positive trends were found with exposure duration (p=0.005 for exposed; p=0.048 overall); risks were modestly elevated (OR 1.11) in those exposed for 20 or more years. No increased risk was observed across smoking strata or lung cancer subtypes. CONCLUSIONS:This pooled analysis provides limited evidence of an association between occupational exposure to chlorinated solvents and lung cancer, though exposure-response trends were noted among exposed individuals.
BACKGROUND:Transcriptome-wide association studies (TWAS) integrate gene expression and genome-wide association studies (GWAS) to identify disease susceptibility genes. Because gene expression varies substantially across cell types within tissues, cell type-specific prediction models may enhance the power of TWAS. METHODS:We conducted cell type-specific TWAS leveraging single-cell RNA sequencing data from the OneK1K cohort (14 immune cell types, 1.27 million cells) and GWAS summary statistics for 7 cancers (>290 000 cases in total). To improve prediction accuracy, we developed a modeling framework that incorporates shared gene expression effects across cell types. RESULTS:At a false discovery rate of 5%, we identified 106 (Bonferroni 5%: 13) previously unreported loci for breast cancer, 51 (4) loci for prostate cancer, 11 (4) loci for lung cancer, 39 (5) loci for melanoma, 9 (1) loci for ovarian cancer, and 2 (1) loci for diffuse large B-cell lymphoma, with most genes exhibiting cell type specificity. Gene set analyses confirmed joint associations of unreported genes with breast and prostate cancer risk in UK Biobank data. Additional lung tissue single-cell RNA sequencing data with 113 individuals validated 18 of 32 (56.3%) statistically significant genes for lung cancer. Across cancers, 139 statistically significant genes were shared by at least 2 cancer types and were primarily enriched in specific immune cell types. CONCLUSION:Cell type-specific TWAS improve the identification of novel cancer susceptibility loci and provide insights into the immune landscape of cancer etiology.
Abstract Background: Lung cancer in never smokers (LCINS) most often presents as non-mucinous adenocarcinoma. The International Association for the Study of Lung Cancer (IASLC) system is the current standard for histologic grading, but its use is limited by interobserver variability, time-intensive manual assessment, and limited scalability. We evaluated whether a deep learning model applied to routine hematoxylin and eosin (H&E) slides could predict overall survival and improve prognostic stratification beyond conventional histologic grading. Methods: We analyzed 595 stage I-III whole-slide images from the Sherlock-Lung study, each from a unique patient. For the full cohort, median follow-up was 37 months (range 1-120 months); 55 deaths occurred by 5 years and 102 deaths by 10 years from diagnosis, out of 190 events overall. Data were split into training (n=409), internal cross-validation (n=45), and held-out validation (n=141) sets. A convolutional neural network generated continuous patient-level risk scores from H&E images, which were dichotomized into high-risk (n=34) and low-risk (n=107) groups in the validation cohort using the Youden index. Prognostic discrimination for overall survival in the validation set was assessed using time-dependent AUCs and Cox models. We compared AUCs for 5- and 10-year survival probabilities estimated from IASLC grade and deep-learning risks in univariate analyses and based on the following Cox models: (1) baseline (age, sex, ancestry, tumor stage) + IASLC grade; (2) baseline + deep learning; (3) baseline + IASLC grade + deep learning. Results: Deep learning yielded higher AUCs than IASLC grade for 5 year (0.84 [0.76-0.92] vs 0.70 [0.62-0.78]; p=0.01) and 10 years overall survival (0.75 [0.61-0.90] vs 0.64 [0.48-0.79]; p=0.59). In multivariable analyses, at 5 years the AUCs were 0.79 for baseline + IASLC, 0.86 for baseline + deep learning (p<0.01 vs baseline + IASLC), and 0.87 for the full model (better than simpler models, p=0.04); at 10 years the corresponding AUCs were 0.82, 0.86, and 0.87, with no statistically significant differences overall (p=0.63), possibly reflecting fewer late events. In a Cox model for 5 years of follow-up the deep-learning low-risk group had improved overall survival versus the high-risk group (HR 0.31, 95% CI 0.13-0.76). Conclusions: Deep learning on routine H&E slides provides prognostic information independent of IASLC grade in LCINS. Integrating deep learning with grade and clinical factors improves overall survival prediction, supporting AI-augmented pathology for precision risk stratification and personalized treatment/surveillance in non-mucinous lung adenocarcinoma among never-smokers. Citation Format: Monjoy Saha, Thi-Van-Trinh Tran, Huu Phuc Hoang, Praphulla MS Bhawsar, Robert Homer, Marina K. Baine, Lynette M. Sholl, Philippe Joubert, Charles Leduc, William D. Travis, Ruth M. Pfeiffer, Jonas S. Almeida, Soo-Ryum Yang, Maria Teresa Landi. Deep learning of H&E slides adds prognostic value beyond IASLC grading in non-mucinous lung adenocarcinoma among never-smokers [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2026; Part 1 (Regular Abstracts); 2026 Apr 17-22; San Diego, CA. Philadelphia (PA): AACR; Cancer Res 2026;86(7 Suppl):Abstract nr 1458.
Supplementary Figure S16 shows proportions of (A) interchromosome and intrachromosome fusions and (B) fusions supported by structural variants (SV) across NS-LUAD expression subtypes. (C) Summary of types of genes involved in fusions. Circle sizes indicate the numbers of genes. D-E, (D) Numbers of protein-coding gene fusions per sample and (E) numbers of in-frame gene fusions per sample across NS-LUAD expression subtypes. Mean values are indicated by the yellow lines. P-values from two-sided Mann-Whitney U-test are shown.
Abstract Approximately 10% of cutaneous malignant melanoma cases are familial. Variants in CDKN2A account for ∼40% of melanoma-prone families, with 10% more explained by other established genes. Strikingly, many unexplained families nonetheless show genetic linkage to chromosomal band 9p21, which harbors CDKN2A, suggesting that high-penetrance non-coding variants in the region may contribute to familial risk.As a part of a large-scale study of high-risk melanoma families from the Mediterranean region, we conducted germline whole-genome sequencing (WGS) on a melanoma case from a four-case family from Genoa, Italy, that is negative for variation in known susceptibility genes. We identified a novel 100kb deletion mapping to 9p21 205kb from CDKN2A and verified cosegregation in the two other available cases. The deletion is not protein-coding but contains annotated melanocyte enhancer elements and open chromatin that we find to interact with the promoters of the p16 and p14 transcripts of CDKN2A in melanocytes using H3K27Ac Hi-ChIP.To identify additional deletion carriers in whole-exome sequencing (WES) data, we searched for nearby exonic variants that cosegregate in the discovery family to serve as a proxy of the deletion, identifying a rare cosegregating missense variant in MTAP (rs755147810; chr9: 21,802,755:T:G; p.Ser3Ala; gnomADv4.1 Non-Finish European MAF: 3.4x10-6). We assessed this variant in WES data from 3,574 high-risk Mediterranean melanoma patients and 2,673 controls, identifying 17 additional Italian melanoma cases (mostly from Genoa) and a single control, and subsequently confirmed all carriers of rs755147810 harbored the deletion. We verified cosegregation of the deletion in all available cases from four additional families (3/3, 2/2, 2/2 and 2/2 cases). The association of the MTAP variant with melanoma risk was highly significant when considering unrelated individuals from Italy, Spain, and Greece (3,219 cases and 2,266 controls; P = 4.34 x 10-4; OR = 13.88) as well as Italian samples alone (2,170 cases, 1,888 controls; P = 4.14 x 10-4; OR = 14.01). Analysis of a shared haplotype among carriers suggests a founder haplotype with the most recent common ancestor (MRCA) dating approximately 26 generations, i.e. the 17th century, when Italy was repeatedly struck by devastating outbreaks of plague that profoundly affected major urban and commercial centers. The coincidence of the estimated MRCA timeframe with this major population collapse suggests that the demographic effects of the plague epidemic may have contributed to the persistence of this genetic variant in modern populations.In conclusion, our study provides evidence of a distant high-penetrance non-coding variant conferring melanoma susceptibility, with potential translational impact on facilitating screening and early-detection efforts in high-risk individuals. Citation Format: Linh T. Bui-Raborn, Lorenza Pastorino, Huu Phuc Hoang, Bruna Dalmasso, Jessica Scales, Sophie Papiernik, Xiaoyu Wang, Rohit Thakur, Mai Xu, Joshuah Yon, William Bruno, Wen Lou, Aurélie L. Vogt, Jia Liu, Kristine Jones, The Melanostrum Consortium, Jianxin Shi, Paola Ghiorzo, Kevin M. Brown, Maria Teresa Landi. A large-scale sequencing study identifies a novel non-coding structural variant as a high-penetrance susceptibility variant for melanoma [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2026; Part 1 (Regular Abstracts); 2026 Apr 17-22; San Diego, CA. Philadelphia (PA): AACR; Cancer Res 2026;86(7 Suppl):Abstract nr 5939.
Supplementary Table S5 shows gene signatures of hallmark genes and lung developmental pathways.
Abstract We performed a new melanoma GWAS meta-analysis, roughly doubling the sample size to ∼70,000 cases compared to the most recent study by Landi and colleagues. In the new meta-analysis, we identify 116 independent risk loci, replicating 52/54 loci previously reported. Post-GWAS variant-to-gene mapping remains critical to functionally interpret new loci discovered by this analysis. Most loci contain many candidate causal variants in linkage disequilibrium, with very few altering protein-coding sequence, suggesting cis-regulatory function underlying the majority of loci. Quantitative trait locus (QTL) colocalization methods can effectively prioritize target genes at susceptibility loci but many loci remain without assigned targets, perhaps reflecting context-specific variant effects not well-reflected in QTL datasets. To complement QTLs and more comprehensively identify potential causal genes, including those beyond 1 Mb, we applied H3K27ac-HiChIP in human primary melanocytes, the cell type of origin for melanoma. H3K27ac-HiChIP combines a Hi-C approach with a H3K27ac ChIP step to detect enhancer-promoter interactions at the risk loci. Statistically significant chromatin interactions were identified using the FitHiChIP pipeline. We subsequently performed variant-to-gene (V2G) mapping at all genome-wide significant melanoma loci, nominating target genes where fine-mapped variants overlapped or physically interacted with their respective promoters. HiChIP-based V2G mapping approach identified target genes at 94% loci (779 genes nominated at 109/116 loci), outperforming other gene nomination approaches (QTL colocalization, protein-coding variants), which nominated 102 candidate genes at 57% of risk loci (67/116 loci). Among genes identified by melanocyte or melanoma eQTL colocalization, a majority (19 of 32) were also nominated by V2G mapping. Likewise, 3 of 5 splice QTL genes and 21 of 34 meQTL genes overlapped with the V2G gene set. Next, we compared gene sets including the V2G mapping-identified candidates to those identified solely by QTLs and protein-coding variants. The gene set incorporating V2G candidates showed significant enrichment for oncogenic signaling pathways, including WNT/β-catenin signaling (FDR: Pwith V2G gene set= 6.3 × 10-05 vs Pwithout V2G gene set= 0.2), aryl hydrocarbon receptor signaling (P= 1.6 × 10-4 vs 0.2), and signaling by NOTCH1 (P=0.002 vs0.5). Notably, V2G mapping linked risk-associated variants to known cancer drivers (e.g. PIK3CA, NOTCH2, MDM4etc.) at 46% of loci (53/116; nominated 76 cancer drivers), with several located >1Mb away. Strikingly, we detected a highly significant long-range interaction (∼2 Mb) connecting fine-mapped variants near the locus 8q24.21 to cancer driver MYC. Overall, H3K27ac-HiChIP-based V2G mapping greatly improves the interpretation of melanoma susceptibility loci by identifying distant susceptibility genes, highlighting known cancer drivers as potential targets, and revealing that these loci converge on key oncogenic signaling pathways. Citation Format: Rohit Thakur, G J M Shanika R Jayasinghe, Mai Xu, Linh Bui-Raborn, Jianxin Shi, Diptavo Dutta, Phuc H. Hoang, Mathias Seviiri, Christopher I. Amos, Andrew Bakshi, Anne E. Cust, Florence Demenais, David L. Duffy, Lars G. Fritsche, Jiali Han, Nicholas K. Hayward, Kiarash Khosrotehrani, Rajiv Kumar, John F. Thompson, Stuart MacGregor, Miguel Renteria, Diane T. Smelser, Sarah V. Ward, Maria Concetta Fargnoli, Paola Ghiorzo, Alisa M. Goldstein, Chiara Menin, David Millan-Esteban, Eduardo Nagore, Cristina Pellegrini, Susana Puig, Alex Stratigos, David C. Whiteman, Melanoma Meta-Analysis Consortium, Lee E. Whelees, Rebecca I. Hartman, Maria Teresa Landi, Matthew H. Law, Kevin M. Brown. H3K27ac-HiChIP variant-to-gene mapping confirms the importance of cancer drivers and oncogenic signaling pathways in melanoma risk [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2026; Part 2 (Late-Breaking, Clinical Trial, and Invited Abstracts); 2026 Apr 17-22; San Diego, CA. Philadelphia (PA): AACR; Cancer Res 2026;86(8_Suppl):Abstract nr LB384.
Understanding tumor cell dynamics can improve prognosis and treatment but remains limited for lung adenocarcinoma in people who have never smoked (NS-LUAD). With RNA sequencing data from 684 NS-LUAD cases and validation in an independent dataset, we identified three subtypes with distinct phenotypic traits and cell compositions. Additional genomic and histologic data further characterized the subtypes. "Steady," marked by low proliferation, high alveolar cell fraction, moderate-to-well differentiation, and fewer driver gene alterations, is linked to prolonged survival and low immune evasion. "Proliferative" shows high proliferation markers, TP53 mutations, and gene fusions. "Chaotic," with high epithelial-to-mesenchymal transition markers, has the worst prognosis, even within stage I tumors. Lacking known molecular or histologic characteristics, this aggressive subtype is solely identified by transcriptomic data. A 60-gene signature recapitulates the classification and predicts survival even within subgroups based on tumor stage or known genomic features, emphasizing its potential for improving early-stage NS-LUAD prognostication in clinical settings. SIGNIFICANCE:The transcriptome of 684 NS-LUAD identifies three subtypes with different cellular dynamics and genomic and morphologic features. A 60-gene signature accurately stratifies subjects for mortality risk, even in stage I, offering a potential clinically applicable tool for treatment decision-making in patients with NS-LUAD. See related commentary by Azizi et al., p. 423.
Genome-wide association studies identified a melanoma- and nevus count-associated locus on chromosome band 9q34.13. Fine-mapping and melanocyte expression data collectively suggest two potential causal genes with opposite association with risk: higher levels of Rap guanine nucleotide exchange factor 1 ( RAPGEF1 ) and lower levels of uridine-cytidine kinase 1 ( UCK1 ). Colocalization analyses and conditional TWAS suggest multiple causal cis -regulatory sequence variants in partial linkage disequilibrium (LD) to each other. Melanocyte capture-HiC and CRISPR-inhibition demonstrated regulatory interactions between fine-mapped variants and the RAPGEF1 and UCK1 promoters. Focusing on RAPGEF1 , we demonstrate RAPGEF1 expression promotes melanocyte growth and drives malignant transformation of human immortalized melanocytes. Following treatment with human EGF, RAPGEF1 overexpression activated both RAP1 and RAS. Further, we show RAPGEF1 expression is significantly enriched in melanomas lacking strongly activating RAS-MAPK mutations, suggesting that RAPGEF1 may promote oncogenic RAS-MAPK signaling in melanomas. Furthermore, in these tumors, we provide preliminary evidence to support the prognostic relevance of RAPGEF1 expression in patients lacking RAS or BRAF mutations. Together with other recent studies, these data suggest that germline variation influencing RAS activation may play a key role in nevus development and melanoma risk.
Supplementary Table S10 shows NS-LUAD expression subtypes predicted by the 60-gene signature in the GIS cohort.
Supplementary Table S8 shows centroids of the 60-gene signature for the classification of NS-LUAD expression subtype.
Abstract Background We investigated whether markers, genes or terms of the Human Phenotype Ontology associated with genetic or rare diseases (GARDs) that affect airway or lung function are associated with lung cancer. Methods Genes of interest were extracted from GARD (Genetic and Rare Diseases Information Center), OMIM (Online Mendelian Inheritance in Man®), ORPHANET and Monarch Initiative. Individual SNP, gene level and gene-set analyses were performed for 52,207 SNPs, 1677 genes or for 620 terms of the Human Phenotype Ontology. The analysis included 14,068 lung cancer cases and 12,390 cancer-free control subjects of European descent from the International Lung Cancer Consortium ILCCO. Results The marker rs56113850 (OR=0.893, 95%CI: 0.862-0.924) was associated with lung cancer (p=1.2x10-10). This marker is located in CYP2A6 as well as in an enhancer region of LTBP4, which is associated with cutis laxa. A suggestive significant association was observed for two markers associated with the DMD gene, which is linked to Duchenne muscular dystrophy. The gene sets "Abnormal circulating adrenocorticotropin concentration" and "Central nervous system neoplasm" were found to be significantly enriched with GARD genes, and can therefore be considered to be associated with lung cancer. Conclusions Genes associated with genetic and rare lung diseases do not generally appear to carry risk factors for lung cancer. However, genes associated with the hypothalamic-pituitary-adrenal axis show some, but rather weak or complex, associations with lung cancer. Tests at the gene level provide extremely inhomogeneous results, even when applied to the same data.
Approximately 10% of cutaneous malignant melanoma cases are familial. Variants in CDKN2A account for up to 40% of melanoma-prone families, with an additional ~10% explained by other genes. Many CDKN2A mutation-negative families show linkage to chromosome-band 9p21, which harbors CDKN2A, suggesting non-coding variants may contribute to familial risk. Here, whole-genome sequencing revealed a novel 100 kb deletion mapping to 9p21 in a gene-desert region, 205 kb from CDKN2A, cosegregating in a four-case family from Genoa, Italy. The deletion overlaps melanocyte enhancers that interact with the promoters of CDKN2A p16 and p14 transcripts and is predicted to reduce p16 expression. Using a nearby rare exonic variant in MTAP (rs755147810) on the deletion haplotype, we searched for deletion carriers in WES data from high-risk melanoma patients and controls and identified 22 cases and a single control carrying rs755147810 and the deletion. The association with melanoma in case-control analysis was highly significant (3,319 cases and 5,680 controls; P=1.27x10-6; OR=27.50). We observe loss-of-heterozygosity of the wild-type allele in a carrier’s tumor sample. The founder haplotype with the most recent common ancestor dates approximately 26 generations, broadly overlapping the period when Italy was struck by devastating outbreaks of plague that decimated the population creating a genetic bottleneck. Our results provide evidence of a high-penetrance intergenic variant conferring melanoma susceptibility, with potential for genetic screening of high-risk individuals.
Osteosarcoma, the most common childhood bone tumor, can occur in rare cancer predisposition syndromes; however, most are sporadic with no known predisposing factors. We investigated the frequency of SMARCAL1 putative pathogenic variants in our large ongoing study of 2119 osteosarcoma patients, their relation to patient characteristics, and the population prevalence. Our analysis uncovered a higher frequency of SMARCAL1 pathogenic variants across 3 osteosarcoma patient sets (1.8%, n = 2119) than in 2625 comparably sequenced cancer-free individuals (0.3%; P < .001). Patients with SMARCAL1 pathogenic variants had statistically significantly improved overall survival compared with patients without these variants (hazard ratio [HR] = 0.36, 95% confidence interval [CI] = 0.14 to 0.96; P = .034). In the UK Biobank (469 557 exomes), there was a 33-fold increased risk of osteosarcoma in individuals with SMARCAL1 pathogenic variants. These results identify SMARCAL1 as a new osteosarcoma predisposition gene and thus warrant follow-up to identify the mechanisms by which SMARCAL1 contributes to the etiology of osteosarcoma.
Abstract Lung cancer is the leading cause of cancer death worldwide in both men and women. It is a highly heterogeneous disease primarily driven by tobacco smoking. About 20% of lung cancers occur among never smokers. There are major differences between lung cancers in patients who have never smoked (LCINS) and who have smoked in patient ancestry, sex, tumor histology and clinical features. LCINS occur more frequently in Asians and females and are predominantly adenocarcinomas. EGFR mutations are enriched in LCINS, whereas KRAS and TP53 mutations are more common in tumors from patients who have smoked. Our understanding of the etiology of LCINS is still limited. Here, we perform a comprehensive study in 1,209 whole-genome sequenced lung cancers collected from 30 different locations across four continents through Sherlock-Lung Project as well as published studies. Among these tumors, 1024 of them are adenocarcinomas and 864 are LCINS. This cohort represents the largest genomics cohort in lung cancer to date. Our study focuses on somatic structural variations (SVs) which are large-scale chromosomal rearrangements. In total, we detect 182,429 somatic SVs in 1209 tumor samples with an average of 151 SVs per sample. Complex SVs, such as chromothripsis, are studied separately from simple SVs because they arise through one-time catastrophic chromosome shattering and rejoining. About two-thirds of the SVs in our cohort are part of complex events. We deconvolute a total of 8 complex SV signatures based on event topology and 8 simple SV signatures using non-negative matrix factorization. They likely represent divergent molecular mechanisms. Among these, chromatin bridge signature, driven by dicentric chromosomes, is the most abundant complex SVs; and median size deletion is the most common simple SVs. SVs are more abundant in tumors from patients who have smoked; however, they are more complex and play more important roles in tumorigenesis in LCINS. The SV breakpoints have distinct distributions across the genome depending on the signatures due to combined effects of mutagenesis and positive selection. Many established cancer-driving genes are recurrently rearranged by multiple SV signatures suggesting functional convergence of these genome instability mechanisms. EGFR mutations and KRAS mutations profoundly and independently shape the SV landscape. EGFR mutant tumors have higher SV burden and more cancer-driving SVs. In contrast, KRAS mutations are associated with lower SV burden and fewer driver SVs. Our study has deepened our understanding of genomic landscape and etiology of lung cancer. Citation Format: YANG YANG, Tongwu Zhang, Lixing Yang, Maria Teresa Landi. Panorama of chromosomal instability in 1209 lung cancer genomes [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2026; Part 1 (Regular Abstracts); 2026 Apr 17-22; San Diego, CA. Philadelphia (PA): AACR; Cancer Res 2026;86(7 Suppl):Abstract nr 5937.
Genetic regulation of splicing uniquely contributes to trait-associated genome-wide association studies (GWAS) signals. However, quantitative trait loci (QTL) analysis using short-read sequencing of bulk tissues fails to capture full-length and cell-type-specific isoforms. Here, we present an isoform-level lung cell atlas from 129 never-smoking Korean women using single-cell long-read RNA-sequencing, identifying abundant unannotated and cell-type-specific isoforms. Isoform-level signatures of 37 lung cell types display a larger difference and therefore improve cell-type classification compared to gene-level expression. Notably, isoform-QTLs (isoQTLs) detect unannotated and/or cell-type-specific isoforms with independent genetic regulation from expression-QTL (eQTL), supported by enriched splicing functional elements. IsoQTLs nominate susceptibility isoforms from previously unexplained lung function and cancer GWAS loci, via eQTL-independent signals. We highlight a potentially functional novel variant of PPIL6 in multiciliated cells underlying lung cancer risk through alternative splicing. This isoform-level resource advances our understanding of cell-type-specific isoform regulation and its contribution to lung traits and diseases.