Artificial intelligence applications in biomedicine face major challenges from data privacy requirements. To address this issue for clinically annotated tissue proteomic data, we developed a federated deep learning approach (ProCanFDL), training local models on simulated sites containing data from a pan-cancer cohort (n = 1,260) and 29 cohorts held behind private firewalls (n = 6,265), representing 19,930 replicate data-independent acquisition mass spectrometry runs. Local parameter updates were aggregated to build the global model, achieving a 43% performance gain on the hold-out test set (n = 625) in 14 cancer subtyping tasks compared with local models and matching centralized model performance. The approach's generalizability was demonstrated by retraining the global model with data from two external, data-independent acquisition mass spectrometry cohorts (n = 55) and eight acquired by tandem mass tag proteomics (n = 832). ProCanFDL presents a solution for internationally collaborative machine learning initiatives using proteomic data, for example, for discovering predictive biomarkers or treatment targets while maintaining data privacy. SIGNIFICANCE:A federated deep learning approach applied to human proteomic data, acquired using two distinct proteomic technologies from 40 tumor cohorts across eight countries, enabled accurate cancer histopathologic subtyping while preserving data privacy. This approach will enable the privacy-compliant development of large-scale proteomic artificial intelligence models, including foundation models, across institutions globally.
Feature importance for selected proteins with utility at distinguishing cancer subtypes
Abstract Purpose: Nonsmokers account for 10% to 13% of all lung cancer cases in the United States. Etiology is attributed to multiple risk factors including exposure to secondhand smoking, asbestos, environmental pollution, and radon, but these exposures are not within the current eligibility criteria for early lung cancer screening by low-dose CT (LDCT). Experimental Design: Urine samples were collected from two independent cohorts comprising 846 participants (exploratory cohort) and 505 participants (validation cohort). The cancer urinary biomarkers, creatine riboside (CR) and N-acetylneuraminic acid (NANA), were analyzed and quantified using liquid chromatography–mass spectrometry to determine if nonsmoker cases can be distinguished from sex and age-matched controls in comparison with tobacco smoker cases and controls, potentially leading to more precise eligibility criteria for LDCT screening. Results: Urinary levels of CR and NANA were significantly higher and comparable in nonsmokers and tobacco smoker cases than population controls in both cohorts. Receiver operating characteristic analysis for combined CR and NANA levels in nonsmokers of the exploratory cohort resulted in better predictive performance with the AUC of 0.94, whereas the validation cohort nonsmokers had an AUC of 0.80. Kaplan–Meier survival curves showed that high levels of CR and NANA were associated with increased cancer-specific death in nonsmokers as well as tobacco smoker cases in both cohorts. Conclusions: Measuring CR and NANA in urine liquid biopsies could identify nonsmokers at high risk for lung cancer as candidates for LDCT screening and warrant prospective studies of these biomarkers.
Association between rs4809294 and rs2292975 and lung cancer risk stratified by exposure to secondhand smoke during adulthood in the NCI-MD Cohort (European Americans).
NCI-MD Case-Control Study Genomic DNA was isolated from buffy coat samples containing white blood cells using Flexigene DNA Kit (Qiagen) according to the manufacturer's instructions.
Association between 3'UTR SNPs and lung cancer in the NCI-MD Cohort (European Americans).