Abstract Leading cellular foundation models have been trained on hundreds of millions of single-cell transcriptomes, with progress increasingly driven by larger datasets and model scaling. Here, we asked whether adding a proteomics modality can improve gene-level and cell-level representations beyond scaling RNA-only models. We introduce cross-modal continued pretraining, fine-tuning a published single-cell model (Tahoe-x1) on a large corpus of proteomic profiles. Training a 70M-parameter Tahoe-x1 model for a single epoch on 48843 proteomic samples from 440 diverse mass-spectrometry studies matched or exceeded 1B- and 3B-parameter RNA-only models across most of the original Tahoe-x1 evaluation benchmarks. This shows that with the right training recipe, heterogeneous proteomics data can improve the learned representations of single-cell RNAseq samples, demonstrating strong out-of-distribution generalization. Cross-modal pretraining also improves transfer to a held-out protein perturbation benchmark, where scaling the RNA-only model does not provide comparable benefits. These results demonstrate that careful targeted curation of proteomics data can provide larger benefits than increasing the model size alone and suggest that multimodal pretraining is a promising path toward more informative biological foundation models.
Abstract Background Blood-based biomarkers are transforming Alzheimer’s Disease (AD) diagnosis and staging. Recent multi-protein plasma panels accurately identify individuals with advanced Braak pathology, but whether these circulating biomarkers truly reflect the molecular remodeling underlying AD neuropathology remains unclear. Pre-existing datasets could answer such questions, but their reuse requires harmonized protein expression data, standardized sample annotations, and consistent study metadata. Methods A recently reported seven-protein plasma staging panel was evaluated across multiple post-mortem proteomics brain cohorts using pre-structured datasets on the Tesorai platform. Five datasets with Braak stage data were identified and re-analyzed. A linear model was used to distinguish late (V-VI) from early Braak stage (0-IV). Because phosphorylated tau 217 (p-tau217) and amyloid beta 1-40 (Aβ40) were not reported for any studies, raw spectra were reprocessed to quantify these proteoforms. Results Across the cohorts, the plasma-derived biomarker panel consistently discriminated early from late disease despite heterogeneous protein coverage. Reprocessing of raw spectra recovered tau phosphopeptides absent from the original protein summaries, enabling inclusion of p-tau217, while Aβ40 remained undetected. Together, these findings show that proteins comprising a recently proposed blood-based staging panel are associated with proteomic remodeling in the AD brain across independent cohorts. Conclusions These findings provide biological validation for a recently proposed blood-based seven-protein staging panel by demonstrating that its constituent biomarkers are associated with disease-stage proteomic changes in AD brain tissue. More broadly, re-mining legacy mass spectrometry data for disease-relevant proteoforms, combined with pre-structured datasets, can accelerate evaluation of emerging blood-based biomarkers against neuropathological and molecular features of disease.
The original mass spectrometry search engines used simple algorithms for peptide identification. Recent tools improved accuracy by adding several extra components such as fragment ion intensities or retention times prediction and training target-decoy classifiers on-the-fly, leading to sometimes inconsistent results. Our study explores the impact of replacing those extra components with a deep-learning pretrained model that directly learns the complex relationship between the full spectra and associated peptide sequence, without using decoys. This simplified workflow has fewer parameters to tweak, making it easier to use and perform robustly on data from instruments and use-cases never seen during training. Surprisingly, our approach consistently identifies more peptides than FragPipe, PEAKS, and Proteome Discoverer (12%, 9%, and 21% more, respectively, across a range of datasets). Tesorai Search is also fast - 250 immunopeptidomics searches in 45 min - and free for academics, available as a webserver at console.tesorai.com.
pIHC and segmentation algorithms were developed for mIF model development and evaluation. A, Examples of mIF (top) and the corresponding pIHC (bottom) for PanCK, PD-L1, CD3, and CD8. As per conventional practice, the residual unmixed AF is not shown in the mIF to maintain appropriate visual quality. B, Example of the segmentations of tumor regions based on the PanCK stain (dashed outline, green) and negative (solid outline, blue) and positive (solid outline, respective color) cells for PD-L1, CD3, and CD8.
T-cell-based immunotherapies have revolutionized cancer treatment, yet only a minority of patients are eligible for these approaches, significantly constrained by the limited knowledge of tumor-specific antigens. Here we present ImmunoVerse, a comprehensive map of T cell targets across 21 cancer types, revealing actionable tumor-specific targets in 89% of tumors analyzed. To define the repertoire of actionable T cell targets, we conducted an exhaustive pan-cancer analysis, integrating data from 7,188 RNA-Seq, 1,771 immunopeptidomes from 512 biological samples and 208 single-cell cancer datasets using novel AI methods, and compared these against 17,384 normal samples covering 51 tissues. Our analysis uncovered 62 viable surface protein targets and 28,446 tumor-specific HLA-presented antigens, deriving from 11 distinct molecular events, across 21 tumor types. Among these, we identified 5,928 previously uncharacterized neoantigens, new tumor self-antigens, peptides derived from tumor-specific cryptic ORFs, tumor-associated microbial targets and a novel splicing-derived PMEL peptide (sPMEL) with enhanced abundance and safety compared to the canonical clinical targets. We successfully expanded sPMEL-specific T cells, validating the therapeutic potential of these targets in functional assays. We highlight 153 promising new tumor targets and experimentally validate 19 targets representing six antigen classes. In addition to being the most comprehensive atlas of targets in scope, ImmunoVerse offers the most extensively annotated resource with key parameters for target selection, providing critical insights for therapeutic prioritization and clinical translation. To catalyze therapeutic development, we released our pan-cancer target atlas through an interactive web portal (https://www.immuno-verse.com) and made the accompanying toolkits available to the scientific community. This work redefines the landscape of therapeutic T cell targets and provides a foundational resource to unlock immunotherapy development across multiple cancers previously considered intractable.
BACKGROUND:Understanding diabetes at the molecular level can help refine diagnostic approaches and personalized treatment efforts. METHODS:We generated proteomic data from plasma collected from participants enrolled in the longitudinal observational cohort study Project Baseline Health Study (PBHS) (evaluated cohort, n = 738, 27.9% of the total PBHS cohort), and integrated those data with information from their medical history and laboratory tests to determine diabetes status. We then identified biomarker proteins associated with diabetes status. RESULTS:Here we identify 87 differentially expressed proteins in people with diabetes compared to those without diabetes, 71 of which show higher expression. This proteomic profile, integrated with clinical data into a logistic regression model, can discriminate diabetes status with over 85% balanced accuracy. CONCLUSIONS:Our approach indicates that proteomic data can enhance diabetes phenotyping, showing potential for marker-based stratification of diabetes diagnosis. These results suggest that a holistic molecular-clinical approach to diagnosis might help personalize treatments or interventions for people with diabetes.
Pearson’s correlation (95% confidence interval) between measurements on real and virtual stains obtained from the cell segmentation–based analysis in Visiopharm software for PanCK, DAPI, PD-L1, CD3, and CD8 on testing slides. Analysis was performed according to three different definitions of the region of interest
Abstract Virtual staining for digital pathology has great potential to enable spatial biology research, improve efficiency and reliability in the clinical workflow, as well as conserve tissue samples in a nondestructive manner. In this study, we demonstrate the feasibility of generating virtual stains for hematoxylin and eosin (H&E) and a multiplex immunofluorescence (mIF) immuno-oncology panel (DAPI, PanCK, PD-L1, CD3, and CD8) from autofluorescence (AF) images of unstained non–small cell lung cancer tissue by combining high-throughput hyperspectral fluorescence microscopy and machine learning. Using domain-specific computational methods, we evaluated the accuracy of virtual H&E staining for histologic subtyping and virtual mIF for cell segmentation–based measurements, including clinically relevant measurements such as tumor area, T-cell density, and PD-L1 expression (tumor proportion score and combined positive score). The virtual stains reproduce key morphologic features and protein biomarker expressions at both tissue and cell levels compared with real stains, enable the identification of key immune phenotypes important for immuno-oncology, and show moderate to good performance across various evaluation metrics. This study extends our previous work on virtual staining from AF in liver disease and prostate cancer, further demonstrating the generalizability of this deep learning technique to a different disease (lung cancer) and stain modality (mIF). Significance: We extend the capabilities of virtual staining from AF to a different disease and stain modality. Our work includes newly developed virtual stains for H&E and a multiplex immunofluorescence panel (DAPI, PanCK, PD-L1, CD3, and CD8) for non–small cell lung cancer, which reproduce the key features of real stains.
Coleoid cephalopods exhibit the highest levels of ADAR-mediated RNA editing of any known animal, yet the functional consequences of most recoding events remain largely unknown. We integrate proteomics with biochemical and cellular assays to characterize thousands of recoding events across the Doryteuthis pealeii proteome. Using quantitative and functional mass spectrometry, we show that RNA edit-driven recoding reshapes the cellular proteome to alter protein stability, subcellular localization, post-translational modifications, and enzymatic activity. -Recoding can regulate post-translational modifications through their creation or ablation, and this has direct effects on protein function and protein-protein interactions. Recoding of the E3 ligase MARCHF5 drives widespread changes in substrate ubiquitylation and perturbs mitochondrial homeostasis, illustrating how RNA editing can influence organelle function. These data provide the first proteome-scale view of how extensive RNA recoding diversifies protein function in coleoid cephalopods and offers a new framework for understanding how RNA-level plasticity shapes protein function and cellular physiology.
Average Dice score (mean ± SD, median) between segmentations on real and virtual stains of combined tumor, leukocyte aggregates, necrosis, and “other” categories on testing slides
Current database search algorithms for mass spectrometry proteomics typically fail to identify up to 75% of spectra. To mitigate the issue, tools such as the widely used Percolator and PeptideProphet leverage machine learning for rescoring. Though this can boost peptide and protein identification rates, the fact that they are trained to separate targets from decoys can lead to inaccurate false-discovery rate (FDR) estimates. In this paper, we propose a novel approach to peptide-spectrum matching, based on a pre-trained large deep learning model, which does not utilize decoys during training and does not require training a new model for every new sample. We trained the Tesorai model on over 100M real peptide-spectrum pairs and demonstrated that the approach performs robustly across a wide range of use cases including standard trypsin-digested human samples, immunopeptidomics, metaproteomics, single-cell and isobaric-labeled samples. In addition to providing robust FDR control, our method increases identifications by up to 170%. Furthermore, we incorporated it into a cloud-native, end-to-end system that can process hundreds of samples in just a few hours. To facilitate new discoveries, we made the model publicly available at www.tesorai.com. ### Competing Interest Statement MB, DS, ST, MM and PC are employees and stockholders of Tesorai, Inc. Other authors declare no competing interest.
The tissue diagnosis of adenocarcinoma and intraductal carcinoma of the prostate includes Gleason grading of tumor morphology on the hematoxylin and eosin stain and immunohistochemistry markers on the prostatic intraepithelial neoplasia-4 stain (CK5/6, P63, and AMACR). In this work, we create an automated system for producing both virtual hematoxylin and eosin and prostatic intraepithelial neoplasia-4 immunohistochemistry stains from unstained prostate tissue using a high-throughput hyperspectral fluorescence microscope and artificial intelligence and machine learning. We demonstrate that the virtual stainer models can produce high-quality images suitable for diagnosis by genitourinary pathologists. Specifically, we validate our system through extensive human review and computational analysis, using a previously validated Gleason scoring model, and an expert panel, on a large data set of test slides. This study extends our previous work on virtual staining from autofluorescence, demonstrates the clinical utility of this technology for prostate cancer, and exemplifies a rigorous standard of qualitative and quantitative evaluation for digital pathology.
Abstract Checkpoint blockade immunotherapy is a cornerstone of lung cancer treatment, but there is a need to improve the identification of patients who will respond favorably. Here, we explored a deep learning approach to predict immunotherapy outcomes from hematoxylin and eosin (H&E) images in non-small cell lung cancer (NSCLC). We included 150 unique cases with metastatic NSCLC (113 adenocarcinoma, 29 squamous cell, 8 other) treated with anti-PD-1/PD-L1 immunotherapy (56 nivolumab, 49 atezolizumab, 44 pembrolizumab, 1 durvalumab) as mono or combination (14 with chemotherapy, 1 with ipilimumab) therapy in a single institution. Each case consisted of a representative H&E whole slide image (53 biopsies, 50 needle core biopsies, 47 resections) obtained prior to immunotherapy, and the outcome reported as the 1-year overall survival (OS). PD-L1 status (tumor proportion score ≥ 1%) was known for 70 cases. We preprocessed the H&E images using two deep learning models previously developed using The Cancer Genome Atlas dataset. First, we used a classification model to identify tumor regions and randomly sampled a fixed number of tumor patches for each case. Then, we used a self-supervised pathology foundation model to obtain a compressed visual representation of each patch, known as an embedding. Next, using our dataset, we trained a deep multiple instance learning (DeepMIL) model with a gated attention mechanism to predict the binary 1-year OS status (0=deceased, 1=alive) for each case. As a baseline, we also trained a linear-probe (logistic regression) model using the averaged embeddings. Given the small dataset size, 5-fold cross-validation was used to train and evaluate both the DeepMIL and linear-probe models, with cases randomly split across folds. For evaluation, we used survival analysis to compare the 0/1 case groups. Overall, across all 150 cases, univariable Cox regression showed that 1-year OS was more strongly associated with the DeepMIL status (46/104, HR=0.55, p=0.03) than the linear-probe status (55/95, HR=0.81, p=0.44). Results were consistent on the subset of 70 cases with known PD-L1 status, whereby OS was most strongly associated with the DeepMIL status (19/51, HR=0.40, p=0.04) compared to the linear-probe status (31/39, HR=0.46, p=0.09) and PD-L1 status (30/40, HR=0.65, p=0.32). In multivariable Cox regression adjusting for age group and smoking status, OS remained more strongly associated with the DeepMIL status (HR=0.45, p=0.08) than PD-L1 status (HR=0.72, p=0.47). In conclusion, the DeepMIL status predicted from H&E images showed a stronger association with outcomes compared to PD-L1 status, a standard biomarker for immunotherapy in NSCLC. These exploratory results demonstrate the potential of deep learning using pathology foundation models to improve immunotherapy outcomes prediction, even with small datasets. Such approaches may even enable the discovery of novel biomarkers from H&E images to advance precision medicine. Citation Format: Jessica Loo, Yang Wang, Pok Fai Wong, Ellery Wulczyn, Jeremy Lai, Peter Cimermancic, David F. Steiner, Shamira S. Weaver. Predicting immunotherapy outcomes from H&E images in lung cancer [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2024; Part 1 (Regular Abstracts); 2024 Apr 5-10; San Diego, CA. Philadelphia (PA): AACR; Cancer Res 2024;84(6_Suppl):Abstract nr 7380.
Conventional histopathology involves expensive and labor-intensive processes that often consume tissue samples, rendering them unavailable for other analyses. We present a novel end-to-end workflow for pathology powered by hyperspectral microscopy and deep learning. First, we developed a custom hyperspectral microscope to nondestructively image the autofluorescence of unstained tissue sections. We then trained a deep learning model to use autofluorescence to generate virtual histologic stains, which avoids the cost and variability of chemical staining procedures and conserves tissue samples. We showed that the virtual images reproduce the histologic features present in the real-stained images using a randomized nonalcoholic steatohepatitis (NASH) scoring comparison study, where both real and virtual stains are scored by pathologists (D.T., A.D.B., R.K.P.). The test showed moderate-to-good concordance between pathologists' scoring on corresponding real and virtual stains. Finally, we developed deep learning-based models for automated NASH Clinical Research Network score prediction. We showed that the end-to-end automated pathology platform is comparable with an independent panel of pathologists for NASH Clinical Research Network scoring when evaluated against the expert pathologist consensus scores. This study provides proof of concept for this virtual staining strategy, which could improve cost, efficiency, and reliability in pathology and enable novel approaches to spatial biology research.
Both histologic subtypes and tumor mutation burden (TMB) represent important biomarkers in lung cancer, with implications for patient prognosis and treatment decisions. Typically, TMB is evaluated by comprehensive genomic profiling but this requires use of finite tissue specimens and costly, time-consuming laboratory processes. Histologic subtype classification represents an established component of lung adenocarcinoma histopathology, but can be challenging and is associated with substantial inter-pathologist variability. Here we developed a deep learning system to both classify histologic patterns in lung adenocarcinoma and predict TMB status using de-identified Hematoxylin and Eosin (H&E) stained whole slide images. We first trained a convolutional neural network to map histologic features across whole slide images of lung cancer resection specimens. On evaluation using an external data source, this model achieved patch-level area under the receiver operating characteristic curve (AUC) of 0.78–0.98 across nine histologic features. We then integrated the output of this model with clinico-demographic data to develop an interpretable model for TMB classification. The resulting end-to-end system was evaluated on 172 held out cases from TCGA, achieving an AUC of 0.71 (95% CI 0.63–0.80). The benefit of using histologic features in predicting TMB is highlighted by the significant improvement this approach offers over using the clinical features alone (AUC of 0.63 [95% CI 0.53–0.72], p = 0.002). Furthermore, we found that our histologic subtype-based approach achieved performance similar to that of a weakly supervised approach (AUC of 0.72 [95% CI 0.64–0.80]). Together these results underscore that incorporating histologic patterns in biomarker prediction for lung cancer provides informative signals, and that interpretable approaches utilizing these patterns perform comparably with less interpretable, weakly supervised approaches.
3122 Background: The current standard work-up for both diagnosis and predictive biomarker testing in metastatic non-small cell lung cancer (NSCLC), can exhaust an entire tumor specimen. Notably, gene mutation panels or tumor mutation burden (TMB) testing currently requires 10 tissue slides and ranges from 10 days to 3 weeks from sample acquisition to test result. As more companion diagnostic (CDx)-restricted drugs are developed for NSCLC, rapid, tissue-sparing tests are sorely needed. We investigated whether TMB, T-effector (TEFF) gene signatures and PD-L1 status can be inferred from H&E images alone using a machine learning approach. Methods: Algorithm development included two steps: First, a neural network was trained to segment hand-annotated, pathologist-confirmed biological features from H&E images, such as tumor architecture and cell types. Second, these feature maps were fed into a classification model to predict the biomarker status. Ground truth biomarker status of the H&E-associated tumor samples came from whole exome sequencing (WES) for TMB, RNAseq for the TEFF gene signatures or reverse-phase protein array for PD-L1. Digital H&E images of NSCLC adenocarcinoma for model development were obtained from the cancer genome atlas (TCGA) and commercial sources. Results: This approach achieves > 75% accuracy in predicting TMB, TEFF and PD-L1 status, offers a way to interpret the model, and provides biological insights into the tumor-host microenvironment. Conclusions: These findings suggest that biomarker inference from H&E images is feasible, and may be sufficiently accurate to supplement or replace current tissue-based tests in a clinical setting. Our approach utilizes biological features for inference, and is thus robust, interpretable, and readily verifiable by pathologists. Finally, biomarker status inference from a single H&E image may enable testing in patients whose tumor tissue has been exhausted, spare further tissue use, and return test results within hours to enable rapid treatment decision-making to maximize patient benefit.
Most clinical drugs are based on microbial natural products, with compound classes including polyketides (PKS), non-ribosomal peptides (NRPS), fluoroquinones and ribosomally synthesized and post-translationally modified peptides (RiPPs). While variants of biosynthetic gene clusters (BGCs) for known classes of natural products are easy to identify in genome sequences, BGCs for new compound classes escape attention. In particular, evidence is accumulating that for RiPPs, subclasses known thus far may only represent the tip of an iceberg. Here, we present decRiPPter (Data-driven Exploratory Class-independent RiPP TrackER), a RiPP genome mining algorithm aimed at the discovery of novel RiPP classes. DecRiPPter combines a Support Vector Machine (SVM) that identifies candidate RiPP precursors with pan-genomic analyses to identify which of these are encoded within operon-like structures that are part of the accessory genome of a genus. Subsequently, it prioritizes such regions based on the presence of new enzymology and based on patterns of gene cluster and precursor peptide conservation across species. We then applied decRiPPter to mine 1,295 Streptomyces genomes, which led to the identification of 42 new candidate RiPP families that could not be found by existing programs. One of these was studied further and elucidated as a novel subfamily of lanthipeptides, designated Class V. Two previously unidentified modifying enzymes are proposed to create the hallmark lanthionine bridges. Taken together, our work highlights how novel natural product families can be discovered by methods going beyond sequence similarity searches to integrate multiple pathway discovery criteria. Code and data availability The source code of DecRiPPter is freely available online at https://github.com/Alexamk/decRiPPter . Results of the data analysis are available online at http://www.bioinformatics.nl/~medem005/decRiPPter_strict/index.html and http://www.bioinformatics.nl/~medem005/decRiPPter_mild/index.html (for the strict and mild filters, respectively). All training data and code used to generate these, as well as outputs of the data analyses, are available on Zenodo at doi:10.5281/zenodo.3834818 .