Escherichia coli is a widely used model organism in molecular biology. Despite its pivotal role, a comprehensive proteome resource covering the E. coli pan-proteome and its post-translational modifications (PTMs) has been lacking. Here we present the E. coli PeptideAtlas build, the first comprehensive pan-proteome analysis of E. coli, generated from 40 high-quality public and in-house data sets spanning a broad diversity of strains, sample types, and experimental conditions, and comprising over 73 million MS/MS spectra. All data sets were reprocessed using both a closed search (Trans-Proteomic Pipeline using MSFragger) and an open search (ionbot). The E. coli PeptideAtlas build provides evidence for 4755 proteins, including 1376 previously lacking protein-level support in UniProtKB. The resource offers protein coverage, modification sites, raw spectra with matched peptides, and manually annotated metadata for the E. coli pan-proteome. PTM profiling identified over 10,000 modification sites, including phosphorylation (3806), acetylation (754), methylation (730), glutathionylation (352) and phosphoribosylation (226). Analysis of the glutathionylation sites revealed potential links to metal binding regulation. We also detected proteins likely stemming from phages, underscoring the value of pan-proteomic approaches for studying host-phage interactions. All identifications are publicly accessible and traceable through the PeptideAtlas interface. We expect that the E. coli PeptideAtlas build will provide a useful resource for the community, which supports, for example, targeted MS experiment design, PTM enrichment method development, and strain typing. It allows straightforward lookups of protein and peptide identifications and facilitates comparative proteomic analyses by enabling the assessment of protein presence and variability across different E. coli strains. The build is available at https://peptideatlas.org/builds/ecoli/.
Non-canonical (i.e., unannotated) open reading frames (ncORFs) have until recently been omitted from reference genome annotations, despite evidence of their translation, limiting their incorporation into biomedical research. To address this, in 2022, we initiated the TransCODE consortium and built the first community-driven consensus catalog of human ncORFs, which was openly distributed to the research community via Ensembl-GENCODE. While this catalog represented a starting point for reference ncORF annotation, major technical and scientific issues remained. In particular, this initial catalogue had no standardized framework to judge the evidence of translation for individual ncORFs. Here, we present an expanded and refined catalog of the human reference annotation of ncORFs. By incorporating more datasets and by lifting constraints on ORF length and start-codon, we define a comprehensive set of 28,359 ncORFs that is nearly four times the size of the previous catalog. Furthermore, to aid users who wish to work with ncORFs with the strongest and most reproducible signals of translation, we utilized a data-driven framework (i.e. translation signature scores) to assess the accumulated evidence for any individual ncORF. Using this approach, we derive a subset of 7,888 ncORFs with translation evidence on par with canonical protein-coding genes, which we refer to as the Primary set. This set can serve as a reliable reference for downstream analyses and validation, with a particular emphasis on high quality. Overall, this update reflects continual community-driven efforts to make ncORFs accessible and actionable to the broader research public and further iterations of the catalog will continue to expand and refine this resource.
Thousands of short open reading frames (sORFs) are translated outside of annotated coding sequences. Recent studies have pioneered searching for sORF-encoded microproteins in mass spectrometry (MS)-based proteomics and peptidomics datasets. Here, we assessed literature-reported MS-based identifications of unannotated human proteins. We find that studies vary by three orders of magnitude in the number of unannotated proteins they report. Of nearly 10,000 reported sORF-encoded peptides, 96% were unique to a single study, and 12% mapped to annotated proteins or proteoforms. Manual curation of a benchmark dataset of 406 manually evaluated spectra from 204 sORF-encoded proteins revealed large variation in peptide-spectrum match (PSM) quality between studies, with immunopeptidomics studies generally reporting higher quality PSMs than conventional enzymatic digests of whole cell lysates. We estimate that 65% of predicted sORF-encoded protein detections in immunopeptidomics studies were supported by high-quality PSMs versus 7.8% in non-immunopeptidomics datasets. Our work stresses the need for standardized protocols and analysis workflows to guide future advancements in microprotein detection by MS towards uncovering how many human microproteins exist.
The rice genome underpins fundamental research and breeding, but the Nipponbare (japonica) reference does not fully encompass the genetic diversity of Asian rice. To address this gap, the Rice Population Reference Panel (RPRP) was developed, comprising high-quality assemblies of 16 rice cultivars to represent japonica, indica, aus, and aromatic varietal groups. The RPRP has been consistently annotated, supported by extensive experimental data and here we report the computational assignment, characterization and dissemination of stably identified pan-genes. We identified 25,178 core pan-genes shared across all cultivars, alongside cultivar-specific and family-enriched genes. Core genes exhibit higher gene expression and proteomic evidence, higher confidence protein domains and AlphaFold structures, while cultivar-specific genes were enriched for domains under selective breeding pressure, such as for disease resistance. This resource, integrated into public databases, enables researchers to explore genetic and functional diversity via a population-aware "reference guide" across rice genomes, advancing both basic and applied research.
Knowledge graphs are increasingly being used to integrate heterogeneous biomedical knowledge and data. General-purpose graph database management systems such as Neo4j are often used to host and search knowledge graphs, but such tools come with overhead and leave biomedical-specific standards compliance and reasoning to the user. Interoperability across biomedical knowledge bases and reasoning systems necessitates the use of standards such as those adopted by the Biomedical Data Translator consortium. We present PloverDB, a comprehensive software platform for hosting and efficiently serving biomedical knowledge graphs as standards-compliant web application programming interfaces. In addition to fundamental back-end knowledge reasoning tasks, PloverDB automatically handles entity resolution, exposure of standardized metadata and test data, and multiplexing of knowledge graphs, all in a single platform designed specifically for efficient query answering and ease of deployment. PloverDB increases data accessibility and utility by allowing data providers to quickly serve their biomedical knowledge graphs as standards-compliant web services. Availability and Implementation:PloverDB's source code and technical documentation are publicly available under an MIT License at github:RTXteam/PloverDB, archived on Zenodo at doi:10.5281/zenodo.15454600.
Proteomics data-dependent acquisition data sets collected with high-resolution mass-spectrometry (MS) can achieve very high-quality results, but nearly every analysis yields results that are thresholded at some accepted false discovery rate, meaning that a substantial number of results are incorrect. For study conclusions that rely on a small number of peptide-spectrum matches being correct, it is thus important to examine at least some crucial spectra to ensure that they are not one of the incorrect identifications. We present Quetzal, a peptide fragment ion spectrum annotation tool to assist researchers in annotating and examining such spectra to ensure that they correctly support study conclusions. We describe how Quetzal annotates spectra using the new Human Proteome Organization (HUPO) Proteomics Standards Initiative (PSI) mzPAF standard for fragment ion peak annotation, including the Python-based code, a web-service end point that provides annotation services, and a web-based application for annotating spectra and producing publication-quality figures. We illustrate its functionality with several annotated spectra of varying complexity. Quetzal provides easily accessible functionality that can assist in the effort to ensure and demonstrate that crucial spectra support study conclusions. Quetzal is publicly available at https://proteomecentral.proteomexchange.org/quetzal/.
Protein phosphorylation, a key post-translational modification, is central to cellular signaling and disease pathogenesis. The development of high-throughput proteomics pipelines has led to the discovery of large numbers of phosphorylated protein motifs and sites (phosphosites) across many eukaryotic species. However, the majority of phosphosites are reported from human samples, with most species having a few experimentally confirmed or computationally predicted phosphosites. Furthermore, only a small fraction of the characterized human phosphoproteome has an annotated functional role. A common way of predicting functional phosphosites is through conservation-based sequence analysis, but large-scale evolutionary studies are scarce. In this study, we explore the conservation of 20,751 confident human phosphosites across 100 eukaryotic species and investigate the evolution of associated protein domains and kinases. We categorize protein functions based on phosphosite conservation patterns and demonstrate the importance of conservation analysis in identifying organisms suitable as biological models for studying conserved signaling pathways relevant to human biology and disease. Finally, we use human protein sequences as a reference for propagating over 1,000,000 potential phosphosites to other eukaryotes. Our results can improve proteome annotations of several species and help direct research aimed at exploring the evolution and functional relevance of phosphorylation.
A major scientific drive is to characterize the protein-coding genome as it provides the primary basis for the study of human health. But the fundamental question remains: what has been missed in prior genomic analyses? Over the past decade, the translation of non-canonical open reading frames (ncORFs) has been observed across human cell types and disease states, with major implications for proteomics, genomics, and clinical science. However, the impact of ncORFs has been limited by the absence of a large-scale understanding of their contribution to the human proteome. Here, we report the collaborative efforts of stakeholders in proteomics, immunopeptidomics, Ribo-seq ORF discovery, and gene annotation, to produce a consensus landscape of protein-level evidence for ncORFs. We show that at least 25% of a set of 7,264 ncORFs give rise to translated gene products, yielding over 3,000 peptides in a pan-proteome analysis encompassing 3.8 billion mass spectra from 95,520 experiments. With these data, we developed an annotation framework for ncORFs and created public tools for researchers through GENCODE and PeptideAtlas. This work will provide a platform to advance ncORF-derived proteins in biomedical discovery and, beyond humans, diverse animals and plants where ncORFs are similarly observed.
Artificial intelligence (AI) is transforming scientific research, including proteomics. Advances in mass spectrometry (MS)-based proteomics data quality, diversity, and scale, combined with groundbreaking AI techniques, are unlocking new challenges and opportunities in biological discovery. Here, we highlight key areas where AI is driving innovation, from data analysis to new biological insights. These include developing an AI-friendly ecosystem for proteomics data generation, sharing, and analysis; improving peptide and protein identification and quantification; characterizing protein-protein interactions and protein complexes; advancing spatial and perturbation proteomics; integrating multi-omics data; and ultimately enabling AI-empowered virtual cells.
We describe a new release of the Candida albicans PeptideAtlas proteomics spectral resource (build 2024-03), providing a sequence coverage of 79.5% at the canonical protein level, matched mass spectrometry spectra, and experimental evidence identifying 3382 and 536 phosphorylated serine and threonine sites with false localization rates of 1% and 5.3%, respectively. We provide a tutorial on how to use the PeptideAtlas and associated tools to access this information. The C. albicans PeptideAtlas summary web page provides “Build overview”, “PTM coverage”, “Experiment contribution”, and “Data set contribution” information. The protein and peptide information can also be accessed via the Candida Genome Database via hyperlinks on each protein page. This allows users to peruse identified peptides, protein coverage, post-translational modifications (PTMs), and experiments that identify each protein. Given the value of understanding the PTM landscape in the sequence of each protein, a more detailed explanation of how to interpret and analyze PTM results is provided in the PeptideAtlas of this important pathogen. Candida albicans PeptideAtlas web page: https://db.systemsbiology.net/sbeams/cgi/PeptideAtlas/buildDetails?atlas_build_id=578.
Advances in mass spectrometry (MS) instrumentation, including higher resolution, faster scan speeds, and improved sensitivity, have dramatically increased the data volume and complexity. The adoption of imaging and ion mobility further amplifies these challenges in proteomics, metabolomics, and lipidomics. Current open formats such as mzML and imzML struggle to keep pace due to large file sizes, slow data access, and limited metadata support. Vendor-specific formats offer faster access but lack interoperability and long-term archival guarantees. We here lay the groundwork for mzPeak, a next-generation community data format designed to address these challenges and support high-throughput, multidimensional MS workflows. By adopting a hybrid model that combines efficient binary storage for numerical data and both human- and machine-readable metadata storage, mzPeak will reduce file sizes, accelerate data access, and offer a scalable, adaptable solution for evolving MS technologies. For researchers, mzPeak will support complex workflows and regulatory compliance through faster access, improved metadata, and interoperability. For vendors, it offers a streamlined, open alternative to proprietary formats. mzPeak aims to become a cornerstone of MS data management, enabling sustainable, high-performance solutions for future data types and fostering collaboration across the mass spectrometry community.
Dataset acquisition and curation are often the hardest and most time-consuming parts of a machine learning endeavor. This is especially true for proteomics-based LC-IM-MS datasets, due to the high-throughput data structure with high levels of noise and complexity between raw and machine learning-ready formats. While predictive proteomics is a field on the rise, when predicting peptide behavior in LC-IM-MS setups, each lab often uses unique and complex data processing pipelines in order to maximize performance, at the cost of accessibility and reproducibility. For this reason we introduce ProteomicsML, an online resource for proteomics-based datasets and tutorials across most of the currently explored physicochemical peptide properties. This community-driven resource makes it simple to access data in easy-to-process formats, and contains easy-to-follow tutorials that allow new users to interact with even the most advanced algorithms in the field. ProteomicsML provides datasets that are useful for comparing state-of-the-art (SOTA) machine learning algorithms, as well as providing introductory material for teachers and newcomers to the field alike. The platform is freely available on https://www.proteomicsml.org/ and we welcome the entire proteomics community to contribute to the project at https://github.com/proteomicsml/.
Protein sulfation can be crucial in regulating protein–protein interactions but remains largely underexplored. Sulfation is nearly isobaric to phosphorylation, making it particularly challenging to investigate using mass spectrometry. The degree to which tyrosine sulfation (sY) is misidentified as phosphorylation (pY) is, thus, an unresolved concern. This study explores the extent of sY misidentification within the human phosphoproteome by distinguishing between sulfation and phosphorylation based on their mass difference. Using Gaussian mixture models (GMMs), we screened ∼45 M peptide-spectrum matches (PSMs) from the PeptideAtlas human phosphoproteome build for peptidoforms with mass error shifts indicative of sulfation. This analysis pinpointed 104 candidate sulfated peptidoforms, backed up by Gene Ontology (GO) terms and custom terms linked to sulfation. False positive filtering by manual annotation resulted in 31 convincing peptidoforms spanning 7 known and 7 novel sY sites. Y47 in calumenin was particularly intriguing since mass error shifts, acidic motif conservation, and MS2 neutral loss patterns characteristic of sulfation provided strong evidence that this site is sulfated rather than phosphorylated. Overall, although misidentification of sulfation in phosphoproteomics data sets derived from cell and tissue intracellular extracts can occur, it appears relatively rare and should not be considered a substantive confounding factor in high-quality phosphoproteomics data sets.
One aim of the international Human Proteome Organization (HUPO) Human Proteome Project (HPP) is to obtain high-confidence translation evidence for every human protein-coding gene established in its target list of 19,433 entries based on the protein-coding genes from Ensembl-GENCODE. However, 76 are annotated in UniProtKB (as of release 2024_06) with PE5, indicating skepticism in the protein's existence from a manual curator, so it is unclear if these entries belong in the HPP target list. Here, we review these 76 entries by assembling evidence from the literature, reference databases, and genome alignments with other species to conclude whether these entries should be freed from their PE5 status to become annotated with PE1-4 in UniProtKB. We find that 17 of these have credible translation evidence and therefore should be upgraded to PE1. Another 15 lack translation evidence but have transcription evidence, the evolutionary hallmarks of protein-coding genes, and are presumed to produce functional proteins. 41 have no translational or transcriptional evidence, although they still bear the evolutionary hallmarks of protein-coding genes; currently, it remains unclear if these are protein-coding, so their representation becomes a matter of policy. Only 3 entries still seem best categorized as PE5 and excluded from the HUPO-HPP target list.
The growing availability of biomedical data offers vast potential to improve human health, but the complexity and lack of integration of these datasets often limit their utility. To address this, the Biomedical Data Translator Consortium has developed an open-source knowledge graph-based system-Translator-designed to integrate, harmonize, and make inferences over diverse biomedical data sources. We announce here Translator's initial public release and provide an overview of its architecture, standards, user interface, and core features. Translator employs a scalable, federated, knowledge graph framework for the integration of clinical, genomic, pharmacological, and other biomedical knowledge sources, enabling query retrieval, inference, and hypothesis generation. Translator's user interface is designed to support the exploration of knowledge relationships and the generation of insights, without requiring deep technical expertise and gradually revealing more detailed evidence, provenance, and confidence information, as needed by a given user. To demonstrate Translator's application and impact, we highlight features of the user interface in the context of three real-world use cases: suggesting potential therapeutics for patients with rare disease; explaining the mechanism of action of a pipeline drug; and screening and validating drug candidates in a model organism. We discuss strengths and limitations of reasoning within a largely federated system and the need for rich concept modeling and deep provenance tracking. Finally, we outline future directions for enhancing Translator's functionality and expanding its data sources. Translator represents a significant step forward in making complex biomedical knowledge more accessible and actionable, aiming to accelerate translational research and improve patient care.
The Human Proteome Project (HPP), the flagship initiative of the Human Proteome Organization (HUPO), has pursued two goals: (1) to credibly identify at least one isoform of every protein-coding gene and (2) to make proteomics an integral part of multiomics studies of human health and disease. The past year has seen major transitions for the HPP. neXtProt was retired as the official HPP knowledge base, UniProtKB became the reference proteome knowledge base, and Ensembl-GENCODE provides the reference protein target list. A function evidence FE1-5 scoring system has been developed for functional annotation of proteins, parallel to the PE1-5 UniProtKB/neXtProt scheme for evidence of protein expression. This report includes updates from neXtProt (version 2023-09) and UniProtKB release 2024_04, with protein expression detected (PE1) for 18138 of the 19411 GENCODE protein-coding genes (93%). The number of non-PE1 proteins ("missing proteins") is now 1273. The transition to GENCODE is a net reduction of 367 proteins (19,411 PE1-5 instead of 19,778 PE1-4 last year in neXtProt). We include reports from the Biology and Disease-driven HPP, the Human Protein Atlas, and the HPP Grand Challenge Project. We expect the new Functional Evidence FE1-5 scheme to energize the Grand Challenge Project for functional annotation of human proteins throughout the global proteomics community, including π-HuB in China.
Mass spectral libraries are collections of reference spectra, usually associated with specific analytes from which the spectra were generated, that are used for further downstream analysis of new spectra. There are many different formats used for encoding spectral libraries, but none have undergone a standardization process to ensure broad applicability to many applications. As part of the Human Proteome Organization Proteomics Standards Initiative (PSI), we have developed a standardized format for encoding spectral libraries, called mzSpecLib (https://psidev.info/mzSpecLib). It is primarily a data model that flexibly encodes metadata about the library entries using the extensible PSI-MS controlled vocabulary, and can be encoded in and converted between different serialization formats. We have also developed a standardized data model and serialization for fragment ion peak annotations, called mzPAF (https://psidev.info/mzPAF). It is defined as a separate standard since it may be used for other applications besides spectral libraries. The mzSpecLib and mzPAF standards are compatible with existing PSI standards such as ProForma 2.0 and the Universal Spectrum Identifier. The mzSpecLib and mzPAF standards have been primarily defined for peptides in proteomics applications, with basic small molecule support. They could be extended in the future to other fields that need to encode spectral libraries for non-peptidic analytes.
Arabidopsis (Arabidopsis thaliana) ecotype Col-0 has plastid and mitochondrial genomes encoding over 100 proteins. Public databases (e.g. Araport11) have redundancy and discrepancies in gene identifiers for these organelle-encoded proteins. RNA editing results in changes to specific amino acid residues or creation of start and stop codons for many of these proteins, but the impact of RNA editing at the protein level is largely unexplored due to the complexities of detection. Here, we assembled the nonredundant set of identifiers, their correct protein sequences, and 452 predicted nonsynonymous editing sites of which 56 are edited at lower frequency. We then determined accumulation of edited and/or unedited proteoforms by searching ∼259 million raw tandem MS spectra from ProteomeXchange, which is part of PeptideAtlas (www.peptideatlas.org/builds/arabidopsis/). We identified all mitochondrial proteins and all except 3 plastid-encoded proteins (NdhG/Ndh6, PsbM, and Rps16), but no proteins predicted from the 4 ORFs were identified. We suggest that Rps16 and 3 of the ORFs are pseudogenes. Detection frequencies for each edit site and type of edit (e.g. S to L/F) were determined at the protein level, cross-referenced against the metadata (e.g. tissue), and evaluated for technical detection challenges. We detected 167 predicted edit sites at the proteome level. Minor frequency sites were edited at low frequency at the protein level except for cytochrome C biogenesis 382 at residue 124 (Ccb382-124). Major frequency sites (>50% editing of RNA) only accumulated in edited form (>98% to 100% edited) at the protein level, with the exception of Rpl5-22. We conclude that RNA editing for major editing sites is required for stable protein accumulation.
Recent improvements in proteomics technologies have fundamentally altered our capacities to characterize human biology. There is an ever-growing interest in using these novel methods for studying the circulating proteome, as blood offers an accessible window into human health. However, every methodological innovation and analytical progress calls for reassessing our existing approaches and routines to ensure that the new data will add value to the greater biomedical research community and avoid previous errors. As representatives of HUPO’s Human Plasma Proteome Project (HPPP), we present our 2024 survey of the current progress in our community, including the latest build of the Human Plasma Proteome PeptideAtlas that now comprises 4608 proteins detected in 113 data sets. We then discuss the updates of established proteomics methods, emerging technologies, and investigations of proteoforms, protein networks, extracellualr vesicles, circulating antibodies and microsamples. Finally, we provide a prospective view of using the current and emerging proteomics tools in studies of circulating proteins.