Artificial intelligence (AI) is transforming scientific research, including proteomics. In this Perspective, we highlight key mass spectrometry (MS)-based proteomics areas where AI is driving innovation, ranging from protein identification to building AI virtual cells. These include improving peptide and protein identification and quantification; characterizing protein-protein interactions and protein complexes; advancing spatial and perturbation proteomics; integrating multi-omics data; and, ultimately, enabling AI virtual cells. Finally, we call for global collaboration among data producers, data consumers and other stakeholders to establish an AI-friendly ecosystem for MS-based proteomics, laying the foundation for transformative advancements in proteomics driven by AI.
Abstract Understanding drug effects and drug–drug interactions is essential for developing combination therapies. We present Drug-Prot, a computational framework that leverages large-scale perturbation proteomics to quantify causal drug effects, drug–drug interactions, and dynamic protein relationships. Using data from 63 single drugs and 59 drug combinations applied to 18 breast cancer cell lines at 6, 24, and 48 hours, Drug-Prot estimates drug effects on protein expression and reconstructs directed temporal protein dependency networks. The publicly available software enables targeted analyses of user-defined protein sets, substantially reducing the multiple-testing burden. Through an interactive web application, users obtain corrected p-values for single-drug and combination effects, directed temporal dependency networks, and downloadable results without requiring access to the underlying proteomic dataset. As a use case, we apply invariance-regularized Random Forests to triple-negative breast cancer cell lines to identify proteins associated with drug response. Querying these proteins in Drug-Prot reveals drug-specific and interaction effects at the protein-network level, illustrating how the framework links candidate causal protein features to actionable drug combinations.
The Mycobacterium tuberculosis complex (MTBC) includes ten human-adapted lineages with varying geography and pathogenicity. Lineage 1 (L1) shows low virulence while Lineage 2 (L2) is hyper-virulent, more transmissible, and associated with drug-resistance. We performed comparative analyses integrating whole-genome sequencing with transcriptomic and proteomic profiling of L1 and L2 clinical strains under two in vitro growth conditions. Transcript-protein correlations varied by strain and gene category, suggesting lineage-specific post-translational regulation. Expression differences scaled with phylogenetic distance, one in three SNPs affected gene expression. A new transcriptional regulatory model identified master transcription factors, linked to the sigma factor network, whose targets were differentially expressed between L1 and L2. For instance, DosR regulon proteins had higher basal levels and exhibited a stronger nitric oxide response in L2. Time-course experiments involving LysG (Rv1985c) induction and wild-type H37Rv under hypoxia and subsequent reaeration confirmed that LysG contributes to reduced metabolic activity, thereby promoting increased tolerance to the novel tuberculosis drug bedaquiline in L2 strains relative to L1. Overall, our findings show how limited genetic variation in the MTBC can yield major phenotypic differences through differential regulation of key transcriptional networks.
A detailed, spatially resolved quantitative map of the human proteome is essential for a deeper understanding of human biology and disease1-4. Here we present a comprehensive human proteomic landscape, generated by profiling more than 13,000 proteins across 2,856 samples using data-independent acquisition mass spectrometry. The dataset spans 58 major tissue types, 251 specific tissue subtypes and 25 distinct carcinomas. This resource enables the depiction of spatially resolved proteome trajectories across tissue types and physiological states, including fetal, tumour, adjacent non-tumour and healthy adult tissue, thereby providing insight into both developmental processes and oncogenic progression. Furthermore, quantitative proteomics comparisons across diverse tissue types and states facilitate the indication of organ-specific toxicity, the identification of repurposable anticancer drug candidates and the prioritization of therapeutic targets for cancers. This study establishes a quantitative resource for navigating the proteome in the human body and in common cancers.
Abstract Even though our meta-analysis ranks Mycobacterium tuberculosis genomes among the bacterial pathogens that are most straightforward to assemble, most available assemblies relied on short-read sequencing and contain genomic blind spots that miss functionally important genes. Complete genomes are essential for functional genomics, particularly for identifying small ORF-encoded proteins (SEPs; ≤100 amino acids), which can play critical biological roles yet are frequently missed by standard annotations. Here, we generated complete long-read assemblies for six clinical reference strains representing lineage 1 and the more pathogenic lineage 2, followed by comparative genomic and proteogenomic analyses. We additionally provide software to predict comprehensive sets of mycobacteria-specific proline-glutamic acid (PE) and PPE family genes, including lineage-specific variants. Using parallel accumulation–serial fragmentation mass spectrometry, we detected approximately two-thirds of each strain’s annotated proteome from unfractionated cell extracts. Extending our proteogenomic framework across related strains, and adding rigorous control of proteogenomic discovery rates using entrapment strategies, we revealed 12–24 previously unannotated proteins per strain, predominantly SEPs, 56–60 alternative translation start sites, and 9–17 expressed pseudogenes. Newly identified proteins included conserved and lineage-specific SEPs, an antitoxin, candidate antimicrobial peptides and novel proteins under purifying selection. Overall, applying this improved proteogenomics method to phylogenomically selected clinical reference strains provides a valuable approach for discovering candidate diagnostics or therapeutics, as illustrated here for a WHO-listed critical bacterial pathogen.
The HUPO Human Proteome Project (HPP) aims to complete the human protein parts list by detecting evidence of expression and of function for all proteins in the human proteome, and make proteomics an integral part of multiomics studies of health and disease. Here we describe the state of the 2025 HPP reference proteome of 19,435 proteins, based on GENCODE v48, UniProtKB 2025_03, Human Protein Atlas 24, MassIVE-KB 2023, and PeptideAtlas 2025-01. We evaluate the progress in the past year, with 93.6% of the proteome detected, and examine the proteins that have not yet been detected to determine where further progress can be made. We also evaluate the progress in determining at least one function for every protein in the HPP target list, finding an increase of 288 proteins in the highest category (FE1) to 5562. Finally, we provide highlights from 12 Biology/Disease-based HPP initiatives, HPP resource pillars, and π-HuB.
Protein-protein interactions are central to virtually all biological processes, forming intricate networks that operate in a highly regulated manner. These interactions are not permanent but rather continuously adapt to environmental changes, developmental cues, or disease-related stress. Understanding which protein interactions are present in a specific cellular state and how they adapt to specific stimuli is one of the long-standing goals of modern systems biology. Mass spectrometry (MS)-based proteomics has emerged as the primary tool for charting these networks. Over the past two decades, continuous advances in instrumentation, sample preparation, and data analysis have enabled researchers to explore the protein interaction landscape with increasing depth and accuracy. This has led to important discoveries in areas ranging from fundamental cell signalling to the identification of new therapeutic targets. We present the current state of MS-based protein interaction analysis, focusing on the three most widely utilized approaches: affinity purification, proximity labelling, and co-fractionation MS. For each, we discuss the fundamental approach, technical considerations, limitations, and highlight the potential integration with future technologies and datasets. Recent innovations such as short-gradient chromatography and faster data acquisition have further improved sensitivity and throughput. Together, these developments are bringing researchers closer to mapping the dynamic, context-dependent architecture of protein networks in unprecedented detail.
The state of a cell depends not only on protein abundance, but also on the biochemical and cellular activities of proteins, which are largely invisible to abundance profiling alone. Here, we introduce a multi-omics framework that infers context-specific protein activities from transcriptomic, phosphoproteomic, and protein correlation-based protein-protein interaction data, integrating modality-specific algorithms via network diffusion. Applying it to a panel of phenotypically diverse HeLa cell lines, whose genetic drift provides a natural perturbation system, we make three findings. First, physical separation of monomeric and assembled protein fractions by protein correlation profiling provides direct evidence that complex assembly buffers variation in gene copy number and transcription, a mechanism previously only inferred from bulk measurements. Second, using Let7 perturbation data, CRISPR gene dependency scores, and subcellular localization, we orthogonally validate that inferred protein activities capture functional regulation linked to cellular phenotypes inaccessible from abundance data alone. Third, differential analysis of context-specific activity profiles identifies molecular mechanisms underlying phenotypic divergence, including a WIPF1/WIPF2–Arp2/3 axis governing invadopodium formation and infection susceptibility, and an immunoproteasome switch linked to immune adaptation. A multi-omics framework infers protein activities from co-fractionation, phosphoproteomic, and transcriptomic data. Applied to a panel of HeLa cell lines, it reveals how protein complex assembly buffers genetic variation and drives phenotypic divergence. A multi-omics framework infers protein activities from co-fractionation, phosphoproteomic, and transcriptomic data. Applied to a panel of HeLa cell lines, it reveals how protein complex assembly buffers genetic variation and drives phenotypic divergence.
This study demonstrates the ability of a protein-based signature to risk-stratify prostate cancer patients with Gleason grade groups 2 and 3 and predict the risk of biochemical recurrence.
Protein-RNA interactions underpin many critical biological processes, demanding the development of technologies to precisely characterize their nature and functions. Many such technologies depend upon cross-linking under mild irradiation conditions to stabilize contacts between amino acids and nucleobases; for example, the cross-linking of stable isotope labelled RNA coupled to mass spectrometry (CLIR-MS) method. A deeper understanding of the CLIR-MS workflow is required to maximize its impact for structural biology, particularly addressing the low abundance of cross-linking products and the information content of spatial/geometric restraints reflected by a cross-link. Here, we present a vastly improved CLIR-MS pipeline that features enhanced sample preparation, data acquisition and interpretation. These advances significantly increase the number of detected cross-link products per sample. We demonstrate that the procedure is robust against variation of key experimental parameters, including irradiation energy and temperature. Using this improved protocol on four protein-RNA complexes representing canonical and non-canonical RNA-binding domains, we propose for the first time the distances encoded by protein-RNA cross-links, enabling their use as structural restraints. We also compared the cross-linking of canonical RNA with 4-thiouracil-labeled counterparts, showing slight, but noticeable differences. The improved understanding of protein-RNA cross-links refines their structural interpretation and facilitates the adoption of the method in integrative/hybrid structural biology.
Artificial intelligence (AI) is transforming scientific research, including proteomics. Advances in mass spectrometry (MS)-based proteomics data quality, diversity, and scale, combined with groundbreaking AI techniques, are unlocking new challenges and opportunities in biological discovery. Here, we highlight key areas where AI is driving innovation, from data analysis to new biological insights. These include developing an AI-friendly ecosystem for proteomics data generation, sharing, and analysis; improving peptide and protein identification and quantification; characterizing protein-protein interactions and protein complexes; advancing spatial and perturbation proteomics; integrating multi-omics data; and ultimately enabling AI-empowered virtual cells.
Despite progress in mapping protein-protein interactions, their tissue specificity is understudied. Here, given that protein coabundance is predictive of functional association, we compiled and analyzed protein abundance data of 7,811 proteomic samples from 11 human tissues to produce an atlas of tissue-specific protein associations. We find that this method recapitulates known protein complexes and the larger structural organization of the cell. Interactions of stable protein complexes are well preserved across tissues, while cell-type-specific cellular structures, such as synaptic components, are found to represent a substantial driver of differences between tissues. Over 25% of associations are tissue specific, of which <7% are because of differences in gene expression. We validate protein associations for the brain through cofractionation experiments in synaptosomes, curation of brain-derived pulldown data and AlphaFold2 modeling. We also construct a network of brain interactions for schizophrenia-related genes, indicating that our approach can functionally prioritize candidate disease genes in loci linked to brain disorders.
As Molecular Systems Biology marks its 20th anniversary, we take this moment to reflect on two decades of discovery and innovation. Since its launch, the journal has stood at the forefront of integrating quantitative biology, computational modeling, and systems science—helping to shape how we understand complex biological systems.
Cell function studies primarily focus on measuring overall molecular abundances while often overlooking critical clues-including protein modifications and molecular interaction networks-that critically determine the functional properties of the cell. In prior work, we introduced a suite of methods to reveal context-specific transcription factor-gene regulatory networks, kinase-substrate networks, and protein interaction networks and leveraged them to gain deeper insights into transcriptional regulation and signal transduction. However, the complex interdependencies between these networks are still elusive. To address this challenge, we introduce a multi-omics framework, aimed at harnessing measured or inferred protein activity in context-specific networks, which yields deeper functional insights into mechanisms underlying molecular phenotypes, compared to protein abundance alone. As proof of concept, we utilized progressively differentiated instances of HeLa CCL2 and Kyoto cell lines to explore the role of protein complexes and interactions in cell doubling time and susceptibility to Salmonella Typhimurium infection. Notably, this analysis underscores the pivotal role of protein interaction networks in linking molecular profiles to phenotypic outcomes, thus providing a highly generalizable framework for multi-omics dataset analysis.
Apolipoprotein E (apoE) polymorphism is associated with different pathologies such as atherosclerosis and Alzheimer's disease. Knowledge of the three-dimensional structure of apoE and isoform-specific structural differences are prerequisites for the rational design of small molecule structure modulators that correct the detrimental effects of pathological isoforms. In this study, cross-linking mass spectrometry (XL-MS) targeting Asp, Glu and Lys residues was used to explore the intramolecular interactions in the E2, E3 and E4 isoforms of apoE. The resulting quantitative XL-MS data combined with molecular modeling revealed isoform-specific characteristics of the N- and C-terminal domain interfaces as well as the isoform-dependent dynamic equilibrium of these interfaces. Finally, the data identified a network of salt bridges formed by R61-R112-E109 residues in the N-terminal helical bundle as a modulator of the interaction with the C-terminal domain making this network a potential drug target.
A comprehensive spatial distribution of the proteome in human body and cancers is fundamental for understanding human biology and diseases including cancers. Here, we present an anatomically resolved human proteome derived from 1781 benign and malignant samples from 58 major tissue types encompassing 251 specific tissues and 25 carcinomas. Based on a spectral library covering over 75% of the human protein-coding genes with 208 understudied and 82 missing proteins characterized, we quantified over 13,000 proteins in these samples using data-independent acquisition proteomics. This data resource presents the so far most comprehensive quantitative proteomic landscape of human tissues and common carcinomas. It allows systematic evaluation of tissue-specific drug responses, identification of drug candidates that may be repurposed as antineoplastics, and discovery of novel targets for anticancer therapy. This resource, available as an online knowledgebase, refines our knowledge of spatial distribution of the human proteome and tumor-specific protein modulation. ### Competing Interest Statement T.Guo and Y. Z. are shareholders of Westlake Omics Inc. The other co-authors declare no competing interests. A.I.N. and F.Y. receive royalties from the University of Michigan for the sale of MSFragger and IonQuant software licenses to commercial entities. All license transactions are managed by the University of Michigan Innovation Partnerships office, and all proceeds are subject to university technology transfer policy.
Systematic inference of enzyme activity in human tumors is key to understanding cancer progression and resistance to therapy. However, standard protein or transcript abundances are blind to the activity status of the measured enzymes, regulated, for example, by active-site amino acid mutations or post-translational protein modifications. Current methods for activity-based proteome profiling (ABPP), which combine mass spectrometry (MS) with chemical probes, quantify the fraction of enzymes that are catalytically active. Here, we describe depletion-dependent ABPP (dd-ABPP) combined with automated SWATH/DIA-MS, which simultaneously determines three molecular layers of studied enzymes: i) catalytically active enzyme fractions, ii) enzyme and background protein abundances, and iii) context-dependent enzyme-protein interactions. We demonstrate the utility of the method in advanced lung adenocarcinoma (LUAD) by monitoring nearly 4000 protein groups and 200 serine hydrolases (SHs) in tumor and adjacent tissue sections routinely collected for patient histopathology. The activity profiles of 23 SHs and the abundance of 59 proteins associated with these enzymes retrospectively classified aggressive LUAD. The molecular signature revealed accelerated lipoprotein depalmitoylation via palmitoyl(protein)hydrolase activities, further confirmed by excess palmitate and its metabolites. The approach is universal and applicable to other enzyme families with available chemical probes, providing clinicians with a biochemical rationale for tumor sample classification.
The Human Proteome Project (HPP), the flagship initiative of the Human Proteome Organization (HUPO), has pursued two goals: (1) to credibly identify at least one isoform of every protein-coding gene and (2) to make proteomics an integral part of multiomics studies of human health and disease. The past year has seen major transitions for the HPP. neXtProt was retired as the official HPP knowledge base, UniProtKB became the reference proteome knowledge base, and Ensembl-GENCODE provides the reference protein target list. A function evidence FE1-5 scoring system has been developed for functional annotation of proteins, parallel to the PE1-5 UniProtKB/neXtProt scheme for evidence of protein expression. This report includes updates from neXtProt (version 2023-09) and UniProtKB release 2024_04, with protein expression detected (PE1) for 18138 of the 19411 GENCODE protein-coding genes (93%). The number of non-PE1 proteins ("missing proteins") is now 1273. The transition to GENCODE is a net reduction of 367 proteins (19,411 PE1-5 instead of 19,778 PE1-4 last year in neXtProt). We include reports from the Biology and Disease-driven HPP, the Human Protein Atlas, and the HPP Grand Challenge Project. We expect the new Functional Evidence FE1-5 scheme to energize the Grand Challenge Project for functional annotation of human proteins throughout the global proteomics community, including π-HuB in China.