Rationale: Lung cancer is the leading cause of cancer-related death globally, and rising incidence among traditionally low-risk individuals intensifies the need for improved early-detection methods that the lower-airway microbiome may inform. Objectives: To evaluate how respiratory tract sampling method shapes inferred microbiome structure, and whether bronchial brushing recovers a microbial community ecologically distinct from BAL and oral rinse. Methods: Prospective cohort of 33 participants (8 lung cancer, 25 non-cancer controls) underwent oral rinse, bilateral bronchoalveolar lavage (BAL), and bronchial brushing. Microbiome structure was characterized by small subunit ribosomal rRNA gene amplicon sequence variant (ASV) profiling, with indicator-species analysis and SPIEC-EASI correlation-network mapping used to identify ASVs associated with sample type or cancer status. Measurements and Main Results: Sampling method was the dominant axis of variation. BAL communities closely resembled oral rinse (65 jointly indicative ASVs; none shared with brushing), whereas bronchial brushing yielded a distinct but low-biomass signal. After excluding host-derived sequences, two brush-specific indicator ASVs affiliated with Sphingomonadaceae and an uncultured Steroidobacteraceae were identified that co-localized within a single co-occurrence module. Cancer-status effects were not detectable in this pilot cohort, consistent with limited statistical power. Conclusions: Sampling method is the primary determinant of inferred respiratory microbiome structure. Bronchial brushing recovers a distinct but low-biomass signal that is obscured when BAL is used in isolation. However, overlap between this low-biomass signal with host and contaminant sequences indicates that confidently resolving a discrete lower-airway community will require deeper sequencing and dedicated contamination controls. These methodological findings directly inform the design of future cancer- and disease-association studies.
In the past two decades, lung cancer screening (LCS) with low-dose computed tomography (LDCT) has emerged as one of the most effective strategies for reducing lung cancer mortality. Landmark trials, including NLST and NELSON, demonstrated mortality reductions exceeding 20%, establishing LDCT as the standard of care for early detection in high-risk populations. Currently, 13 countries have implemented national or regional LCS programs, with additional nations preparing for rollout. Advances in risk-prediction models, volumetric nodule assessment, and structured management protocols have improved precision and efficiency. Integration of artificial intelligence is enhancing nodule detection, prediction of malignancy risk, individualized screening intervals, and workflow optimization. Real-world evidence confirms improved stage distribution and suggests reduction in lung cancer mortality. Initiatives such as promoting community engagement, equitable access through geospatial mapping, and mobile screening will improve screening uptake and retention. Embedding tobacco dependence treatment within LCS further augments life-years gained. Complementary incidental pulmonary-nodule programs and expanding studies in people who have never smoked are extending the reach of early detection, whereas biomarker research is progressing toward integration with imaging-based screening. The potential to use LDCT scans to detect coronary heart disease and chronic obstructive pulmonary disease may have a major impact on future health care benefits. Ongoing efforts to harmonize data collection standards, establish quality indicators, and strengthen workforce training are essential to sustain high-quality implementation. As LCS evolves into a cornerstone of lung cancer control, continued innovation in risk stratification, imaging technologies, and biomarker integration will be key to maximizing global benefit and equity.
Abstract Background We investigated whether markers, genes or terms of the Human Phenotype Ontology associated with genetic or rare diseases (GARDs) that affect airway or lung function are associated with lung cancer. Methods Genes of interest were extracted from GARD (Genetic and Rare Diseases Information Center), OMIM (Online Mendelian Inheritance in Man®), ORPHANET and Monarch Initiative. Individual SNP, gene level and gene-set analyses were performed for 52,207 SNPs, 1677 genes or for 620 terms of the Human Phenotype Ontology. The analysis included 14,068 lung cancer cases and 12,390 cancer-free control subjects of European descent from the International Lung Cancer Consortium ILCCO. Results The marker rs56113850 (OR=0.893, 95%CI: 0.862-0.924) was associated with lung cancer (p=1.2x10-10). This marker is located in CYP2A6 as well as in an enhancer region of LTBP4, which is associated with cutis laxa. A suggestive significant association was observed for two markers associated with the DMD gene, which is linked to Duchenne muscular dystrophy. The gene sets "Abnormal circulating adrenocorticotropin concentration" and "Central nervous system neoplasm" were found to be significantly enriched with GARD genes, and can therefore be considered to be associated with lung cancer. Conclusions Genes associated with genetic and rare lung diseases do not generally appear to carry risk factors for lung cancer. However, genes associated with the hypothalamic-pituitary-adrenal axis show some, but rather weak or complex, associations with lung cancer. Tests at the gene level provide extremely inhomogeneous results, even when applied to the same data.
Background:Genome-wide association studies (GWAS) have identified numerous lung cancer susceptibility loci based on single nucleotide polymorphisms (SNPs), yet a substantial proportion of heritability remains unexplained. We therefore evaluated germline copy number variants (CNVs) as an underexplored source of genetic susceptibility and potential contributors to genomic instability in lung cancer. Methods:We conducted a genome-wide analysis of germline CNVs using 19,342 cases and 15,917 controls from the Transdisciplinary Research in Cancer of the Lung (TRICL) consortium, with replication in two independent cohorts. High-confidence CNVs were identified by integrating two CNV callers including PennCNV and modSaRa2. Association analyses were performed using both gene-based and CNV region-based approaches. Polygenic risk scores (PRS) were constructed from top loci, and functional validation was conducted using siRNA-mediated knockdown in lung fibroblast cells. Results:We identified CNVs in four genomic regions (1p36.22, 2q31.2, 6p21.32, and 19q13.32) significantly associated with lung cancer risk. Two loci (1p36.22 and 2q31.2) were consistently supported across both analytical strategies. A CNV-based PRS constructed from key genes (CLCN6, NFE2L2, OPA3, and PSMB8) was significantly associated with lung cancer risk and replicated across independent datasets. Functional assays demonstrated that knockdown of NFE2L2 and OPA3 increased endogenous DNA damage, supporting a role in genomic stability. Conclusions:Germline CNVs contribute to lung cancer susceptibility and may influence carcinogenesis through mechanisms related to genomic instability. Impact:These findings expand the genetic architecture of lung cancer and highlight CNVs as potential biomarkers for improving risk stratification and informing precision prevention strategies.
Purpose To compare the performance of an artificial intelligence (AI) system with that of radiologists for estimating malignancy risk of indeterminate-size nodules (5-15 mm) at low-dose CT (LDCT) within a standardized and transparent evaluation framework. Materials and Methods Teams participating in the AI study had access to a public dataset of 555 malignant and 5608 benign nodules on 4069 baseline LDCT scans from the National Lung Screening Trial to develop AI systems. External testing was performed on 156 malignant and 312 benign size-matched nodules, all of indeterminate size, from 463 baseline scans collected from three large European lung cancer screening trials, and the best-performing AI system (based on area under the receiver operating characteristic curve [AUC]) was selected. An observer study was conducted in which radiologists assessed 300 randomly selected nodules (100 malignant, 200 benign) from the external test set. Radiologists categorized nodules as low, intermediate, or high risk, and the threshold of intermediate or greater risk (intermediate or high-risk) was used to define a positive test result. The selected AI system was compared with radiologists on this subset using the AUC. Results The selected AI system demonstrated superior performance to the 65 radiologists' mean (AUC, 0.78 [95% CI: 0.73, 0.84] vs 0.70 [95% CI: 0.65, 0.74]; P = .001). With use of the intermediate risk or greater threshold, the AI system correctly classified 12% more malignant nodules at matched specificity and yielded 20% fewer false-positive results at matched sensitivity. Conclusion The selected AI system was superior to radiologists in estimating malignancy risk of indeterminate lung nodules at LDCT. Keywords: CT, Thorax, Lung, Observer Performance, Screening, Supervised Learning, Lung Cancer Screening, Radiologists, Artificial Intelligence, Benchmarking, Pulmonary Nodule Malignancy Risk, Deep Learning Supplemental material is available for this article. © RSNA, 2026 See also commentary by Júdice de Mattos Farina and Szarf in this issue.
Air pollution, such as particulate matter with a diameter of ≤2.5 μm (PM2.5) is a major contributor to lung cancer in the never-smoking population. Anthracotic pigments (black deposits in the lungs) are physical evidence of environmental exposures. While links between anthracosis and lung disease have been established, anthracosis has not been routinely used as a measure of environmental exposure in research, because there is neither a standard quantitative measure of anthracosis nor an efficient method for quantifying it. We developed 'Slide-based methods for High-throughput Anthracosis Detection and Estimation' (SHADE). SHADE is an automated workflow that quantifies anthracotic pigments on scanned images of whole H&E slides. SHADE was optimized to identify anthracosis while avoiding detection of artefacts in images. SHADE scores and manual pathologist rankings of 10-image patches demonstrated high concordance (R = 0.864). Application of SHADE to background lung sections from 140 never-smoking lung cancer patients demonstrated significant associations between high anthracosis and older age, male sex, 30-year residential PM2.5 estimates, and birthplace in Asia (p < 0.05). Low anthracosis was associated with less aggressive adenocarcinomas (p = 0.052). No significant relationship was identified between anthracosis and 3-year residential PM2.5 estimates, lobe sampled, ethnicity, EGFR mutation status or pathologic tumor stage. Anthracosis levels in the never-smoking cohort were similar to that of background lung sections from 59 ever-smoking patients. Exploratory analysis of background lung sections of 15 individuals who had bronchoalveolar lavage specimen gene expression data available, SHADE scores were associated with the expression of several immunoregulatory genes (CEACAM6, CXCL6, IL5RA, LCN2, SPA17), the activation of inflammatory response and cytokine gene sets, and suppression of the antigen processing and presentation gene set (false discovery rate < 0.1). SHADE provides an automated workflow for anthracosis quantification and has demonstrated utility in uncovering insight into anthracosis-associated molecular changes in lung cancer.
Predicting lung cancer risk would enhance prevention trials. Although the Canakinumab Anti-inflammatory Thrombosis Outcome Study (CANTOS) trial demonstrated reduced lung cancer incidence with interleukin (IL)-1β inhibition, the high number needed to treat (NNT) to prevent lung cancer limits its use in unselected populations. Using machine learning, we identified a 14-protein plasma signature predicting lung cancer more than 5 years before diagnosis. The signature, validated across eight cohorts, was elevated in current smokers and individuals exposed to particulate matter (PM) and linked to lung myeloid and alveolar cells. In epidermal growth factor receptor (EGFR)-driven lung adenocarcinoma, diverse epithelial lineages converged on a keratin8+/claudin4+ alveolar transitional state (KAC), whose transcriptional programs correlated with signature emergence. Components of the signature were induced by PM, oncogenic EGFR, or IL-1β, whereas IL-1β inhibition restrained PM-driven KAC expansion and early tumorigenesis. In CANTOS, the signature identified individuals who seemed to benefit more from anti-IL-1β therapy, lowering the NNT threshold and nominating circulating signals of tumor promotion for prevention.
Exhaled breath (EB) testing holds potential for non-invasive screening and early detection of lung cancer. Development of such a test requires knowledge of volatile organic compounds (VOCs) originating in the lung microenvironment, rather than exogenous sources or non-lung endogenous sources such as the GI tract and oral cavity. Direct evidence linking peripheral lung-derived VOCs to EB is lacking. The peripheral lung air (⩾4th generation airways) of thirty-five participants (9 lung cancers; 26 controls) was sampled during bronchoscopy using micro-thermal desorption-gas chromatography-ion mobility spectrometry (µTD-GC-IMS) for direct on-site analysis and was compared to EB obtained immediately prior. Data processing included signal-normalized background subtraction to evaluate which VOCs originate in the peripheral lung and comparison to EB. Forty-three IMS clusters (features) were determined to originate from the lung microenvironment. All forty-three lung-originating features were detected in EB. Twenty-four out of forty-three features had higher median signal in either the EB or peripheral lung air compared to the environmental background. Some features had higher signal in EB compared to peripheral lung air, while others showed the reverse trend. Direct analysis of peripheral lung air, proximal to the tumour, usingµTD-GC-IMS was clinically feasible. These findings help address a key translational barrier in breath-based respiratory biomarker development and support the feasibility of non-invasive approaches for lung cancer screening grounded in lung-specific VOC biology, and will inform larger clinical trials directed at breath biomarker discovery.
BACKGROUND:Low-dose CT (LDCT) imaging screening reduces lung cancer mortality, the leading cause of cancer deaths globally. Segmentation-free deep learning (DL) models such as Sybil can improve screening efficiency but require extensive validation and possible improvement. RESEARCH QUESTION:Can the integration of DL based on LDCT scans and clinical data improve lung cancer risk prediction? STUDY DESIGN AND METHODS:Retrospective cohort data from 4 different screening programs, 1 used for model training and 3 used for external validation. Data were collected between 2002 and 2021. The median follow-up period was 7 years. All participants had a history of either current or former smoking, with at least 10 pack-years of smoking or who smoked over 20 years. The area under the receiver operating characteristic curve (AUC) was calculated for lung cancer risk within 1 to 6 years, stratified by pulmonary nodule presence and size. Key clinical and epidemiologic factors were evaluated for their added predictive value. RESULTS:This analysis used 52,482 LDCT scan series from 22,469 participants. Sybil's AUC ranged from 0.93 in year 1 and reduced to 0.79 in year 6 in the independent cohorts. The predictive performance was suboptimal in the absence of documented nodules (AUC, 0.64) and for small nodules (AUC, 0.61) in year 6. Our new model, Sybil-Epi, trained with baseline scans, achieved higher predictive performance (AUC, 0.83; 95% CI, 0.81-0.85) compared with Sybil (AUC, 0.80; 95% CI, 0.78-0.82) in year 6. The difference is most notable when nodules are absent. Sybil-Epi's AUC was 0.76 (95% CI, 0.70-0.82) and Sybil's AUC was 0.64 (95% CI, 0.57-0.70). INTERPRETATION:Our results show that Sybil performs better for short-term lung cancer risk, but the predictive accuracy was suboptimal when nodules were absent. Our integrated Sybil-Epi model with DL and clinical and epidemiologic factors significantly improved model predictive performance.
Heterozygosity at human leukocyte antigen (HLA) loci may improve lung cancer immunosurveillance by increasing recognition of the tumor by the immune system. Previous studies utilizing data from population-level biobanks, such as the United Kingdom Biobank and FinnGen, have identified an association between germline HLA class II (HLA-II) heterozygosity and reduced lung cancer risk in smokers. In the present study, we evaluate the association between HLA heterozygosity and lung cancer in a large case-control study (15,302 cases and 14,580 controls) with imputed HLA allele-type information, comparing differences in HLA heterozygosity between smokers and non-smokers, among lung cancer subtypes, and at 2- and 4-digit HLA allele resolution. We identify a strong protective association of HLA-II heterozygosity in smokers compared to non-smokers, particularly at the HLA-DPB1 and HLA-DPA1 loci, and provide subtype-specific resolution. Finally, analysis of the additive effects of HLA allele heterozygosity in smokers identified significant associations with several 4-digit HLA alleles, including HLA-B∗08:01, HLA-A∗01:01, HLA-C∗07:01, HLA-DQA1∗05:01, HLA-DRB1∗03:01, and HLA-C∗03:04. Our study provides additional evidence, with added histologic subtype information, that germline HLA-II heterozygosity is inversely associated with lung cancer risk.
BACKGROUND:Lung cancer (LC) remains the deadliest cancer, often diagnosed at advanced stages. Screening reduces mortality in high-risk individuals. Eligibility criteria in European and US screening guidelines have recently expanded. Therefore, we conducted an updated systematic review of risk-based models for identifying candidates for low-dose computed tomography screening and post-screening nodule classification. METHODS:We systematically searched Embase and Medline (January 2020-January 2026), identifying studies proposing new risk models in the context of LC screening. We separated models by pre- and post-screening risk stratification. Data extraction included study design, population, model type, risk horizon and model performance metrics. We performed an exploratory meta-regression of areas under the curve (AUCs) to assess whether sample size, model type, validation type and inclusion of biomarkers were associated with performance. RESULTS:Of 2462 records, 91 were included. 56 models were for screening selection (30 included biomarkers) and 35 for post-screening nodule classification. Regression-based models predominated, though machine-learning approaches were increasingly common. Discrimination ranged from moderate (AUC∼0.70) to excellent (>0.90), with biomarker and imaging-enhanced models often outperforming models without. Calibration was inconsistently reported and fewer than half underwent external validation. CONCLUSION:We identified 91 risk prediction models for LC, developed after 2020. Although many demonstrated promising discrimination across both screening selection and post-screening management, most remain insufficiently mature for clinical adoption, as their performance and practical value outside the original study setting are uncertain. Future work should prioritise external validation, updating and comparative evaluation of existing models, and prospective implementation studies rather than continued development of additional models.
BACKGROUND:In chronic obstructive pulmonary disease (COPD), patients with peripheral airways infiltrated by eosinophils and neutrophils ("mixed granulocytic COPD") show worse outcomes than those without. OBJECTIVES:We sought to examine the expression of immune response and lung tissue remodeling pathways in patients with mixed granulocytic COPD. METHODS:In this post hoc study of the DISARM (A Study to Investigate the Differential Effects of Inhaled Symbicort and Advair on Lung Microbiota) randomized controlled trial, mixed granulocytic COPD was defined by eosinophils >1% and neutrophils >3% of the total leukocyte count in bronchoalveolar lavage (BAL). We compared clinical outcomes of mixed granulocytic COPD with 2 other phenotypes (neutrophilic, pauci-granulocytic). We then examined canonical pathways using gene set enrichment analysis and expression of immune cell and tissue remodeling gene signatures in BAL across phenotypes. RESULTS:Among 54 patients, 33% had mixed granulocytic COPD. Patients with mixed granulocytic COPD had the lowest FEV1, the most radiographic emphysema, and the highest annualized exacerbation rates. Cell pellets of BAL in mixed granulocytic COPD showed upregulated gene expression of type 1 (TNFA, IL6, and IFNG signaling) and type 2 (IL4/13 signaling) immune responses relative to the neutrophilic and pauci-granulocytic phenotypes. Mixed granulocytic COPD was marked by increased expression of gene signatures for natural killer cells, B cells, and CD4 and CD8 naive and memory/effector T cells. Mixed granulocytic COPD showed the highest expression of tissue remodeling processes, which significantly associated with lower FEV1. CONCLUSIONS:Mixed granulocytic COPD is marked by complex immune responses in the peripheral airways and is associated with increased tissue remodeling that associates with more severe airflow limitation.
Objective: Pembrolizumab monotherapy is an anti-PD-1 immunotherapy that is approved as a first-line treatment for non-small cell lung cancer (NSCLC) patients with high PD-L1 expression (≥50%). However, approximately 55% of these patients do not respond. Early identification of likely non-responders is critical to enable timely transition to alternative treatments. Materials: This study analyzed a retrospective cohort of NSCLC patients treated with first-line PD-L1 monotherapy, divided into a discovery training set (n: 97; 27 non-responders) and a preliminary test set (n: 17; 9 non-responders). Treatment response was assessed using baseline and follow-up CT scans in accordance with the response evaluation criteria in solid tumors (RECIST v1.1). Methods: Our objective was to extract deep learning (DL) features from the two groups of patients and apply transfer learning techniques to identify patients at risk of progression on pembrolizumab monotherapy. A nonparametric statistical test (Mann-Whitney U) was employed to rank the discriminative power of the 128 features from these training groups. Two types of support vector machine (SVM-RBF and SVM-Polynomial) classifiers were employed to investigate the discriminating power of the highest-ranked features as measured by F1 score and AUC values over ROC curves at the three levels of the data (slice, lesion, and patient) with and without clinical descriptors. Results: SVM-RBF performed best when trained on the 10 highest-ranked DL features and five clinical descriptors, achieving AUC of 0.742 (CI 95% 0.47-1.00), SN of 88.9%, SP of 75% and F1 score of 84.2% on preliminary test set patients, whereas an AUC of 0.902 ± 0.031, SN of 81.5%, SP of 81.4% and F1 score of 71% were observed for the discovery training set. Conclusions: Integrating CT-based DL features with clinical descriptors demonstrated balanced performance, offering a promising tool to identify patients at risk of progression on pembrolizumab monotherapy to support first-line treatment decisions in PD-L1-high NSCLC.
Abstract Microbial communities inhabiting the respiratory tract contribute to health status through interactions with host physiology, immune function, and local environmental conditions. Advances in small subunit ribosomal RNA (SSU or 16S rRNA) gene amplicon sequencing enable culture-independent profiling of microbial communities as amplicon sequence variants (ASVs), revealing links between microbial dysbiosis and respiratory diseases, and the use of mass spectrometry to measure volatile organic compounds (VOCs) in exhaled breath shows emerging promise for biomarker discovery. Here we present ASPIRE, the Amplicon Sequencing Profiler for Investigating Respiratory Ecosystems, an accessible Nextflow workflow for processing, analyzing, and interpreting linked ASV-VOC data from respiratory microbiome studies. ASPIRE is designed to support scalable comparative analysis across respiratory sample types while preserving intermediate file outputs for inspection and reuse within a standardized file structure. Availability and implementation The ASPIRE workflow, quick start guide, and usage are freely available on the GitHub platform https://github.com/hallamlab/ASPIRE . ASPIRE is implemented in Nextflow with a modular design to allow for reproducibility, extensibility and scalability.
Rationale: A 4-protein biomarker panel (4MP) can improve estimation of lung cancer risk and identify individuals who may benefit most from lung cancer screening. In the current study, we evaluate the performance of the 4MP for risk determination of lung cancer in individuals undergoing lung cancer screening. Methods: We obtained two cohorts of samples from the National Lung Screening Trial (NLST). The first cohort consisted of 675 samples from individuals with CT scans without suspicious findings, including 135 eventually diagnosed with lung cancer and 540 matched control samples. The second cohort consisted of 715 samples from individuals with screen-detected pulmonary nodules, including 143 diagnosed with cancer and 572 matched controls. We additionally obtained plasma from 27 cases lung cancer cases and 1308 non-cases with screen-detected pulmonary nodules that were drawn from the University of British Columbia site of the International Lung Screening Trial (ILST). The 4MP was measured using a multiplex bead-based immunoassay using coefficients fixed from a previously developed logistic regression model. Performance was evaluated using receiver operating characteristic analysis, including computing the area under the curve (AUCs) as well as a net reclassification index (NRI) to estimate how well the 4MP improved existing risk prediction models such as the Brock nodule risk model. Results: In the NLST cohort with negative CTs, for those individuals eventually diagnosed with stage II or higher lung cancer, the 4MP showed an AUC of 0.67 (95% CI 0.59-0.74). For all those with advanced (stage III and higher), AUC was 0.71 (95% CI 0.61-0.81). For individuals diagnosed within 1 year of blood draw, the AUC was 0.71 (95% CI 0.51-0.91). In those with indeterminate nodules, the 4MP showed similar performance in both the NLST and ILST samples, with an AUC of 0.64 (95% CI 0.59-0.70) in those eventually diagnosed with stage II or higher lung cancer from the NLST. We further evaluated the performance of the 4MP to reclassify pulmonary nodules, roughly based on Lung-RADS risk category. This model was developed in the ILST samples and validated in the NLST samples. The 4MP, added to the Brock model, showed an NRI of 0.26 compared to the Brock model alone in the validation set. Conclusions: The 4MP may be a useful adjunct to screening, especially in identifying those who will develop more advanced stage disease in the following year. This could help to identify individuals who may benefit from closer clinical follow-up.