The Scythians, often described as mounted horse-back warriors of the Iron Age steppe with lavish burial goods, have attracted increasing scientific interest over the past years. Recent genetic and multi-isotopic studies have uncovered that the 'Scythians' were neither a homogenous political nor a cultural group, but rather diverse populations of heterogeneous origins with intricate socio-political systems. Although populational differences in agro-pastoral subsistence regimes of Northern Black Sea Region groups have previously been identified through stable isotope analysis, it remains unclear which animal products were consumed. Here we investigate the dietary systems of two Scythian-era populations in present-day Ukraine using protein analysis of ancient dental calculus. Various dietary proteins and their taxonomic origin were identified revealing the consumption of milk from ruminant and equine species. This study supplements previous findings that Scythians engaged in complex, agro-pastoralist subsistence strategies in forest-steppe and steppe environments.
Recent advances in liquid chromatography mass spectrometry (LCMS) have accelerated the adoption of high-throughput workflows that deliver deep proteome coverage using minimal sample amounts. This trend is largely driven by clinical and single-cell proteomics, where sensitivity and reproducibility are essential. Here, we extend our previous benchmark dataset (PXD028735) using next-generation LC-MS platforms optimized for rapid proteome analysis. We generated an extensive DDA/DIA dataset using a human-yeast-E. coli hybrid proteome. The proteome sample was distributed across multiple laboratories together with standardized analytical protocols specifying two short LC gradients (5 and 15 min) and low sample input amounts. This dataset includes data acquired on four different platforms, and features new scanning quadrupole-based implementations, extending coverage across different instruments and acquisition strategies. Our comprehensive evaluation highlights how technological advances and reduced LC gradients may affect proteome depth, quantitative precision, and cross-instrument consistency. The release of this benchmark dataset via ProteomeXchange (PXD070049 and PXD071205), allows for the acceleration of cross-platform algorithm development, enhance data mining strategies, and supports standardization of short-gradient, high-throughput LC-MS-based proteomics. ### Competing Interest Statement Frederic Fontaine is employed by Thermo Fisher Scientific. Ihor Batruch, Patrick Pribil and Jean-Baptiste Vincendet are employed by SCIEX. Bart Van Puyvelde joined SCIEX after the completion of this work. Research Foundation - Flanders, https://ror.org/03qtxy027, 1278023N, 1SH9O24N, G010023N, 12A6L24N Ghent University Special Research Fund, BOF/PDO/2025/049, BOF21/GOA/033 Horizon Europe, 101080544, 101191739 CHIST-ERA, G0GDV23N European Molecular Biology Laboratory, 208391/Z/17/Z, 223745/Z/21/Z BBSRC, BB/X001911/1 Agence Nationale de la Recherche - French Proteomic Infrastructure (ProFI), ProFI UAR2048, ANR-10-INBS08-03, ANR-24-INBS-0015 Region Grand-Est, SC-Proteomics project ITMO Cancer of Aviesan the Interdisciplinary Thematic Institute IMS IdEx Unistra, ANR-10-IDEX-0002 SFRI-STRATUS, ANR-20-SFRI-0012
Quality control procedures play a pivotal role in ensuring the reliability and consistency of data generated in mass spectrometry-based proteomics laboratories. However, the lack of standardized quality control practices across laboratories poses challenges for data comparability and reproducibility. In response, we conducted a harmonization study within proteomics laboratories of the Core for Life alliance with the aim of establishing a common quality control framework, which facilitates comprehensive quality assessment and identification of potential sources of performance drift. Through collaborative efforts, we developed a consensus quality control standard for longitudinal assessment and adopted common processing software. We generated a 4-year longitudinal data set from multiple instruments and laboratories, which enabled us to assess intra- and interlaboratory variability, to identify causes of performance drift, and to establish community reference values for several quality control parameters. Our study enhances data comparability and reliability and fosters a culture of collaboration and continuous improvement within the proteomics community to ensure the integrity of proteomics data.
AbstractStaphylococcus aureusandPseudomonas aeruginosafrequently co-occur in infections, and there is evidence that their interactions can negatively affect disease outcomes.P. aeruginosais known to be dominant, often compromisingS. aureusthrough the secretion of inhibitory compounds. We previously demonstrated thatS. aureuscan become resistant to growth-inhibitory compounds during experimental evolution. While resistance arose rapidly, the underlying mechanisms were not obvious as only a few genetic mutations were associated with resistance, while ample phenotypic changes occurred. We thus hypothesize that resistance may result from a combination of phenotypic responses and genetic adaptation. Here, we tested this hypothesis using proteomics. We first focused on an evolved strain that acquired a single mutation intcyA(encoding a transmembrane transporter unit) upon exposure toP. aeruginosasupernatant. We show that this mutation leads to a complete abolishment of transporter synthesis, which confers moderate protection against PQS and selenocystine, two toxic compounds produced byP. aeruginosa. However, this genetic effect was minor compared to the fundamental phenotypic changes observed at the proteome level when both ancestral and evolvedS. aureusstrains were exposed toP. aeruginosasupernatant. Major changes involved the downregulation of virulence factors, metabolic pathways, membrane transporters, and the upregulation of ROS scavengers and an efflux pump. Our results suggest that the observed multi-variate phenotypic response is a powerful adaptive strategy, offering instant protection against competitors in fluctuating environments and reducing the need for hard-wired genetic adaptations.ImportanceDifferent bacterial pathogens can co-occur in infections, where they interact with one another and influence disease severity. Previous research showed that pathogens can evolve and adapt to co-infecting species. Here, we show that evolution through genetic mutations and selection are not necessarily required to change pathogen behavior. Instead, we found that the human pathogenStaphylococcus aureusis able to plastically respond to the presence ofPseudomonas aeruginosa, a competing pathogen. Through proteomics and metabolomics, we demonstrate thatS. aureusundergoes substantial proteomic alterations in response toP. aeruginosaby down-regulating virulence factor expression, changing metabolism, and mounting protective measures against toxic compounds. Our work highlights that pathogens possess sophisticated mechanisms to respond to competitors to secure growth and survival in polymicrobial infections. We predict such plastic responses to have significant impacts on infection outcomes.
Recent developments in machine learning (ML) and deep learning have immense potential for applications in proteomics, such as generating spectral libraries, improving peptide identification, and optimizing targeted acquisition modes. Although new ML models are regularly published, the rate at which the community adopts these models is slow. This is in part due to a lack of findability and accessibility of these models as well as the technical challenges involved in incorporating these models into data analysis pipelines and demonstrating their reusability for end-users. Here we show Koina, an open-source decentralized and online-accessible model repository to facilitate publication of ML models. Koina enables ML model usage via an easy-to-use online interface, facilitating the integration of ML models in data analysis pipelines. Using the widely used FragPipe computational platform as an example, we demonstrate how Koina can be integrated with existing proteomics software tools and how these integrations improve data analysis.
Mass spectrometry (MS)-based proteomics is a well-established strategy for analyzing complex biological mixtures. Many MS instruments and data acquisition strategies are available, and the data they acquire differ substantially, thus requiring tailored analysis algorithms. Hence, many dedicated bioinformatics workflows are developed. These are in constant evolution, and the community lacks a centralized platform for comparing their performance. Here, we propose ProteoBench, a single platform that brings together software developers and software users to provide an ever-evolving comparison of state-of-the-art proteomics data processing tools. ProteoBench is an open-source resource that enables the community to evaluate data analysis workflows, develop benchmarking modules dedicated to specific comparisons, and discuss the best methods to compare software tools. The platform ensures that the benchmark evolves alongside advances in proteomics data analysis workflows. ProteoBench guides researchers towards the best-suited tool and parameters for their specific project and data according to their needs, and developers can test their newly developed tools or workflows privately, before adding them as public references. This community-driven effort will increase transparency and reproducibility between MS data analysis workflows, as well as facilitate the development and publication of software workflows in the field. ### Competing Interest Statement The authors have declared no competing interest.
Mass spectrometry is a cornerstone of quantitative proteomics, enabling relative protein quantification and differential expression analysis (DEA) of proteins. As experiments grow in complexity, involving more samples, groups, and identified proteins, interactive differential expression analysis tools become impractical. The prolfquapp addresses this challenge by providing a command-line interface that simplifies DEA, making it accessible to nonprogrammers and seamlessly integrating it into workflow management systems. Prolfquapp streamlines data processing and result visualization by generating dynamic HTML reports that facilitate the exploration of differential expression results. These reports allow for investigating complex experiments, such as those involving repeated measurements or multiple explanatory variables. Additionally, prolfquapp supports various output formats, including XLSX files, SummarizedExperiment objects and rank files, for further interactive analysis using spreadsheet software, the exploreDE Shiny application, or gene set enrichment analysis software, respectively. By leveraging advanced statistical models from the prolfqua R package, prolfquapp offers a user-friendly, integrated solution for large-scale quantitative proteomics studies, combining efficient data processing with insightful, publication-ready outputs.
Staphylococcus aureus and Pseudomonas aeruginosa frequently co-occur in infections, and there is evidence that their interactions can negatively affect disease outcomes. P. aeruginosa is known to be dominant, often compromising S. aureus through the secretion of inhibitory compounds. We previously demonstrated that S. aureus can become resistant to growth-inhibitory compounds during experimental evolution. While resistance arose rapidly, the underlying mechanisms were not obvious as only a few genetic mutations were associated with resistance, while ample phenotypic changes occurred. We thus hypothesize that resistance may result from a combination of phenotypic responses and genetic adaptation. Here, we tested this hypothesis using proteomics. We first focused on an evolved strain that acquired a single mutation in tcyA (encoding a transmembrane transporter unit) upon exposure to P. aeruginosa supernatant. We show that this mutation leads to a complete abolishment of transporter synthesis, which confers moderate protection against Pseudomonas quinolone signal and selenocystine, two toxic compounds produced by P. aeruginosa. However, this genetic effect was minor compared to the fundamental phenotypic changes observed at the proteome level when both ancestral and evolved S. aureus strains were exposed to P. aeruginosa supernatant. Major changes involved the downregulation of virulence factors, metabolic pathways, and membrane transporters, and the upregulation of reactive oxygen species scavengers and an efflux pump. Our results suggest that the observed multivariate phenotypic response is a powerful adaptive strategy, offering instant protection against competitors in fluctuating environments and reducing the need for hard-wired genetic adaptations.IMPORTANCEDifferent bacterial pathogens can co-occur in infections, where they interact with one another and influence disease severity. Previous research showed that pathogens can evolve and adapt to co-infecting species. Here, we show that evolution through genetic mutations and selection is not necessarily required to change pathogen behavior. Instead, we found that the human pathogen Staphylococcus aureus is able to plastically respond to the presence of Pseudomonas aeruginosa, a competing pathogen. Through proteomics and metabolomics, we demonstrate that S. aureus undergoes substantial proteomic alterations in response to P. aeruginosa by downregulating virulence factor expression, changing metabolism, and mounting protective measures against toxic compounds. Our work highlights that pathogens possess sophisticated mechanisms to respond to competitors to secure growth and survival in polymicrobial infections. We predict such plastic responses to have significant impacts on infection outcomes.
Recent developments in machine-learning (ML) and deep-learning (DL) have immense potential for applications in proteomics, such as generating spectral libraries, improving peptide identification, and optimizing targeted acquisition modes. Although new ML/DL models for various applications and peptide properties are frequently published, the rate at which these models are adopted by the community is slow, which is mostly due to technical challenges. We believe that, for the community to make better use of state-of-the-art models, more attention should be spent on making models easy to use and accessible by the community. To facilitate this, we developed Koina, an open-source containerized, decentralized and online-accessible high-performance prediction service that enables ML/DL model usage in any pipeline. Using the widely used FragPipe computational platform as example, we show how Koina can be easily integrated with existing proteomics software tools and how these integrations improve data analysis.
The spotted pod borer, Maruca vitrata (Lepidoptera: Crambidae) is a destructive insect pest that inflicts signifi-cant productivity losses on important leguminous crops. Unravelling insect proteomes is vital to comprehend their fundamental molecular mechanisms. This research delved into the proteome profiles of four distinct stages-three larval and pupa of M. vitrata, utilizing LC-MS/MS label-free quantification-based methods. Employing comprehensive proteome analysis with fractionated datasets, we mapped 75 % of 3459 Drosophila protein orthologues out of which 2695 were identified across all developmental stages while, 137 and 94 were exclusive to larval and pupal stages respectively. Cluster analysis of 2248 protein orthologues derived from MaxQuant quantitative dataset depicted six clusters based on expression pattern similarity across stages. Consequently, gene ontology and protein-protein interaction network analyses using STRING database identified cluster 1 (58 proteins) and cluster 6 (25 proteins) associated with insect immune system and lipid metabolism. Furthermore, qRT-PCR-based expression analyses of ten selected proteins-coding genes authenticated the proteome data. Subsequently, functional validation of these chosen genes through gene silencing reduced their transcript abundance accompanied by a marked increase in mortality among dsRNA-injected larvae. Overall, this is a pioneering study to effectively develop a proteome atlas of M. vitrata as a potential resource for crop protection programs.
Mass spectrometry is widely used for quantitative proteomics studies, relative protein quantification, and differential expression analysis of proteins. There is a large variety of quantification software and analysis tools. Nevertheless, there is a need for a modular, easy-to-use application programming interface in R that transparently supports a variety of well principled statistical procedures to make applying them to proteomics data, comparing and understanding their differences easy. The prolfqua package integrates essential steps of the mass spectrometry-based differential expression analysis workflow: quality control, data normalization, protein aggregation, statistical modeling, hypothesis testing, and sample size estimation. The package makes integrating new data formats easy. It can be used to model simple experimental designs with a single explanatory variable and complex experiments with multiple factors and hypothesis testing. The implemented methods allow sensitive and specific differential expression analysis. Furthermore, the package implements benchmark functionality that can help to compare data acquisition, data preprocessing, or data modeling methods using a gold standard data set. The application programmer interface of prolfqua strives to be clear, predictable, discoverable, and consistent to make proteomics data analysis application development easy and exciting. Finally, the prolfqua R-package is available on GitHub https://github.com/fgcz/ prolfqua, distributed under the MIT license. It runs on all platforms supported by the R free software environment for statistical computing and graphics.
Core facilities have to offer technologies that best serve the needs of their users and provide them a competitive advantage in research. They have to set up and maintain instruments in the range of ten to a hundred, which produce large amounts of data and serve thousands of active projects and customers. Particular emphasis has to be given to the reproducibility of the results. More and more, the entire process from building the research hypothesis, conducting the experiments, doing the measurements, through the data explorations and analysis is solely driven by very few experts in various scientific fields. Still, the ability to perform the entire data exploration in real-time on a personal computer is often hampered by the heterogeneity of software, the data structure formats of the output, and the enormous data sizes. These impact the design and architecture of the implemented software stack. At the Functional Genomics Center Zurich (FGCZ), a joint state-of-the-art research and training facility of ETH Zurich and the University of Zurich, we have developed the B-Fabric system, which has served for more than a decade, an entire life sciences community with fundamental data science support. In this paper, we sketch how such a system can be used to glue together data (including metadata), computing infrastructures (clusters and clouds), and visualization software to support instant data exploration and visual analysis. We illustrate our in-daily life implemented approach using visualization applications of mass spectrometry data.
Here, we present the Universal Spectrum Explorer (USE), a web-based tool based on IPSA for cross-resource (peptide) spectrum visualization and comparison (https://www.proteomicsdb.org/use/). Mass spectra under investigation can be either provided manually by the user (table format) or automatically retrieved from online repositories supporting access to spectral data via the universal spectrum identifier (USI), or requested from other resources and services implementing a newly designed REST interface. As a proof of principle, we implemented such an interface in ProteomicsDB thereby allowing the retrieval of spectra acquired within the ProteomeTools project or real-time prediction of tandem mass spectra from the deep learning framework Prosit. Annotated mirror spectrum plots can be exported from the USE as editable scalable high-quality vector graphics. The USE was designed and implemented with minimal external dependencies allowing local usage and integration into other web sites (https://github.com/kusterlab/universal_spectrum_explorer).
We use `prolfqua` to develop highly customizable, visually appealing, and interactive data analysis reports in pdf or HTML format for quantification experiments. We use `prolfqua` to visualize and model simple experimental designs with a single explanatory variable and complex experiments with multiple factors. The `prolfqua` package integrates essential steps of the data analysis workflow: quality control, data normalization, protein aggregation, sample size estimation, modeling, and hypothesis testing. We further use `prolfqua` to benchmark data acquisition, data preprocessing or data modeling methods. We developed and improved the package by applying the "Eating your own dog food" principle, making it easy to use. We store all the data needed for analysis in a single data frame in a tidy table, i.e., every column is a variable, every row is an observation, every cell is a single value. Using an R6 configuration object, we specify what variable is in which column, making it easy to integrate new inputs in prolfqua if provided in tidy tables. For example, to visualize tidy Spectronaut, or Skyline outputs, or data in MSStats format, only a few lines of code to update the prolfqua configuration are needed. For popular software like MaxQuant or MSFragger, which stores the same variable (e.g., intensity) in multiple columns, one for each sample, we implemented methods that transform the data into tidy tables. Relying on the tidy data table enabled us to easily interface with many data manipulation, visualization, and modeling methods implemented in base R and the tidyverse. We use R's linear model and mixed model formula in prolfqua. R linear model and linear mixed effect models allow modeling parallel designs, repeated measurements, factorial designs, and many more. R's formula interface for linear models is flexible, widely used, and well documented. This approach makes it easy to reproduce an analysis performed with prolfqua in any other statistical programming language. We implemented features specific to high throughput experiments, such as the experimental Bayes variance and p-value moderation, which utilizes the parallel structure of the protein measurements and the analysis [limma]. We also compute probabilities of differential protein regulation based on peptide level models [ropeca]. Contrasts to test hypothesis can intuitively be specified in prolfqua using descriptive variable names, e.g., “Treatment_drug - Treatment_placebo”. The Benchmark functionality of prolfqua includes ROC curves and computes partial areas under those curves (pAUC) and other scores. We use it to study how well linear, mixed effect models or p-value moderation models quantitative mass spectrometric high throughput experiments. The benchmarking results enable us to choose the best method depending on the number of missing values present or the research hypothesis. Based on these results, we discuss how much each of these Mass Spectrometric, computational and statistical methods can improve the sensitivity and specificity of the analysis. How to install and use prolfqua to analyze your data, or benchmark your MS method or software, is shown in various tutorials available at https://wolski.github.io/prolfqua/ . Prolfqua is an easy-to-use R package to analyze quantitative mass spectrometric data and to report results. We used it here to benchmark MS software and statistical methods.
The Bioconductor project (Nat. Methods 2015, 12 (2), 115-121) has shown that the R statistical environment is a highly valuable tool for genomics data analysis, but with respect to proteomics, we are still missing low-level infrastructure to enable performant and robust analysis workflows in R. Fundamentally important are libraries that provide raw data access. Our R package rawDiag (J. Proteome Res. 2018, 17 (8), 2908-2914) has provided the proof-of-principle how access to mass spectrometry raw files can be realized by wrapping a vendor-provided advanced programming interface (API) for the purpose of metadata analysis and visualization. Our novel package rawrr now provides complete, OS-independent access to all spectral data logged in Thermo Fisher Scientific raw files. In this technical note, we present implementation details and describe the main functionalities provided by the rawrr package. In addition, we report two use cases inspired by real-world research tasks that demonstrate the application of the package. The raw data used for demonstration purposes was deposited as MassIVE data set MSV000086542. Availability: https://github.com/fgcz/rawrr.
Proteomics research infrastructures and core facilities within the Core for Life alliance advocate for community policies for quality control to ensure high standards in proteomics services.
The package for pro teomics l abel - f ree qua ntification `prolfqua` (read: prolevka) evolved from functions and code snippets used to visualize and analyze label-free quantification data. To compute protein fold changes among treatment conditions, we first used t-test or linear models and then used functions implemented in the package limma to obtain moderated p-values. We evaluated MSStats , ROPECA , or MSqRob , all implemented in R, to integrate the various approaches. Although all these packages are written in R, model specification, input, and output formats differ widely and wildly, making our first attempt to use the original implementations challenging. Therefore, and also to better understand the algorithms used, we attempted to reimplement those methods, where possible. The R-package prolfqua is the outcome of this venture. When developing prolfqua , we draw inspiration from packages such as sf , which uses data in a long tidy table format, dplyr for data transformation ggplot2 for visualization. In the long table format, each column stores a different attribute, e.g., there is only a single column with the intensities. In prolfqua , the data needed for analysis is represented using a single data-frame in a long, tidy format and an R6 configuration object. The configuration annotates the data-frame, i.e., specifies what information is in which column. The use of an annotated table makes integrating new data if provided in long formatted tables simple. Therefore, all that is needed to incorporate Spectronaut, Skyline text output, or MSStats inputs is to update the configuration object. For software like MaxQuant writing the data in a wide table format, with several intensity columns, one for each sample, we implemented methods that transform the data into a long format. Relying on the long tidy data table format enabled us to easily access various useful data manipulation and visualization methods implemented in the R packages dplyr and ggplot2 . A further design decision was to embraces R's linear model formula interface, including the lme4 mixed effect models formula interface. R's formula interface for linear models is flexible, widely used, and well documented. These interfaces allow specifying a wide range of essential models, including parallel designs, factorial designs, repeated measurements, and many more. Since `prolfqua` uses R modeling infrastructure directly, we can fit all these models to proteomics data. This is not easily possible with other package dedicated to proteomics data analysis. For instance, MSStats, although using the same modeling infrastructure, supports only a subset of possible models. Limma supports the R formula interface but not for linear mixed models. Since ROPECA relies on limma , it is limited to the same set of models. MSqRob allows specifying fixed and random effects but using their own model specification, and it is unclear how interactions among factors can be specified, estimated, or tested. R's formula interface does not limit prolfqua to the output provided by the R modeling infrastructure. prolfqua implements p-value moderations and computes probabilities of differential regulation, as suggested by ROPECA . Last but not least, ANOVA analysis or model selection using the likelihood ratio test for thousand of proteins can also be performed. To use prolfqua, knowledge of the R regression model infrastructure is of advantage. Acknowledging the formula interface's complexity, we plan to provide an MSstats emulator, which derives the model formula from an MSstats formatted input file. We benchmarked all the methods implemented in prolfqua : linear models, mixed effect models, p-value moderation, ROPECA, and Bayesian regression models implemented in brms using a benchmark dataset, enabling us to evaluate the practical relevance of these methods. Finally, prolfqua supports the LFQ data analysis workflow elements, e.g., computing coefficients of Variations (CV) for peptide and proteins, sample size estimation, visualization and summarization of missing data, intensities, multivariate analysis, etc. It also implements various protein intensity summarization and inference methods, e.g., top 3, or Tukeys median polish. Our package makes it relatively easy to perform proteomics data analysis to generate visualizations and reports using Rmarkdown. We will continue extending the package's functionality. The package can be installed from www.github.com/wolski/prolfqua .
Mike Sips合作论文数Stanford University
Gates Computer Science
Graphics Lab13