Quantitative proteomics relies on accurate selection of fragment ions for quantification, yet most current algorithms apply simple strategies such as median intensity or single quality filters. Modern data-independent acquisition (DIA) searches generate rich features such as fragment ion correlations, retention time and many others that could be leveraged to assess fragment quality. We introduce QuantSelect, a novel strategy to select optimal fragments by systematically integrating these features via self-supervised deep learning. QuantSelect uses a regularized, weighted-variance loss on intensity traces normalized via our directLFQ algorithm. This allows learning a fragment quality score without ground truth labels, enabling on-the-fly training on label-free DIA datasets. Integrated within our alphaDIA pipeline, QuantSelect significantly improves quantitative accuracy and in some cases substantially corrects protein intensity estimation. Sensitivity in differential expression improved by 68% in a mixed-species benchmarking dataset and by 18% in single-cell data. QuantSelect provides a practical framework for data-driven fragment selection that improves accuracy, precision and downstream inference in DIA proteomics. ### Competing Interest Statement Matthias Mann is an indirect investor in Evosep Biosciences. The other authors declare no relevant competing interests.
Scientific discovery relies on innovative software as much as experimental methods, especially in proteomics, where computational tools are essential for mass spectrometer setup, data analysis, and interpretation. Since the introduction of SEQUEST, proteomics software has grown into a complex ecosystem of algorithms, predictive models, and workflows, but the field faces challenges, including the increasing complexity of mass spectrometry data, limited reproducibility due to proprietary software, and difficulties integrating with other omics disciplines. Closed-source, platform-specific tools exacerbate these issues by restricting innovation, creating inefficiencies, and imposing hidden costs on the community. Open-source software (OSS), aligned with the FAIR Principles (Findable, Accessible, Interoperable, Reusable), offers a solution by promoting transparency, reproducibility, and community-driven development, which fosters collaboration and continuous improvement. In this manuscript, we explore the role of OSS in computational proteomics, its alignment with FAIR principles, and its potential to address challenges related to licensing, distribution, and standardization. Drawing on lessons from other omics fields, we present a vision for a future where OSS and FAIR principles underpin a transparent, accessible, and innovative proteomics community.
Deep and accurate proteome analysis is crucial for understanding cellular processes and disease mechanisms; however, it is challenging to implement in routine settings. In this protocol, we combine a robust chromatographic platform with a high-performance mass spectrometric setup to enable routine yet in-depth proteome coverage for a broad community. This entails tip-based sample preparation and pre-formed gradients (Evosep One) combined with a trapped ion mobility time-of-flight mass spectrometer (timsTOF, Bruker). The timsTOF enables parallel accumulation–serial fragmentation (PASEF), in which ions are accumulated and separated by their ion mobility, maximizing ion usage and simplifying spectra. Combined with data-independent acquisition (DIA), it offers high peak sampling rates and near-complete ion coverage. Here, we explain how to balance quantitative accuracy, specificity, proteome coverage and sensitivity by choosing the best PASEF and DIA method parameters. The protocol describes how to set up the liquid chromatography–mass spectrometry system and enables PASEF method generation and evaluation for varied samples by using the py_diAID tool to optimally position isolation windows in the mass-to-charge and ion mobility space. Biological projects (e.g., triplicate proteome analysis in two conditions) can be performed in 3 d with ~3 h of hands-on time and minimal marginal cost. This results in reproducible quantification of 7,000 proteins in a human cancer cell line in quadruplicate 21-min injections and 29,000 phosphosites for phospho-enriched quadruplicates. Synchro-PASEF, a highly efficient, specific and novel scan mode, can be analyzed by Spectronaut or AlphaDIA, resulting in superior quantitative reproducibility because of its high sampling efficiency. Aligning trapped ion mobility with a mass-selective quadrupole and time-of-flight mass spectrometry enables parallel accumulation–serial fragmentation, which improves proteomic analysis. This protocol comprehensively describes PASEF workflows.
The scale of data generated for mass-spectrometry-based proteomics and modern acquisition strategies poses a challenge to bioinformatic analysis. Search engines need to make optimal use of the data for biological discoveries while remaining statistically rigorous, transparent and performant. Here we present alphaDIA, a modular open-source search framework for data-independent acquisition (DIA) proteomics. We developed a feature-free identification algorithm that performs machine learning directly on the raw signal and is particularly suited for detecting patterns in data produced by time-of-flight instruments. Benchmarking demonstrates competitive identification and quantification performance. While the method supports empirical spectral libraries, we propose a search strategy named DIA transfer learning that uses fully predicted libraries. This entails continuously optimizing a deep neural network for predicting machine-specific and experiment-specific properties, enabling the generic DIA analysis of any post-translational modification. AlphaDIA provides a high performance and accessible framework running locally or in the cloud, opening DIA analysis to the community.
Machine learning increasingly uncovers rules of biology directly from data, enabled by large, standardized datasets. Microscopy images provide rich information on cellular architecture and are accessible at scale across biological systems, making them an ideal foundation for modeling cell behavior. However, a standardized image format does not exist at the single-cell level. Here we present scPortrait, an scverse software package for generation, storage, and application of single-cell image datasets. scPortrait reads, stitches and segments raw fields of view with out-of-core computation scaling to larger-than-memory datasets. Parallelization enables rapid extraction of individual cells into a standardized single-cell image format with fast access to accelerate machine learning. scPortrait enables analysis across modalities including images, proteomics and transcriptomics, identifying cancer-associated macrophage subpopulations by morphology and embedding single-cell images into transcriptome atlases. scPortrait turns microscopy images into a reusable resource for integrative cell modeling, establishing single-cell images as a core modality in systems biology. ### Competing Interest Statement S.C.M. consulted for Lamin Labs GmbH and is a future employee of Ensocell Therapeutics. N.A.S. consulted for Lamin Labs GmbH. A.N. and L.H. are employees of Lamin Labs GmbH. G.W. is a founder of Aplusia GmbH. F.J.T. consults for Immunai, Singularity Bio, CytoReason, Cellarity, and Omniscope and has ownership interest in Dermagnostix and Cellarity. M.M. is an indirect investor in Evosep Biosystems and OmicVision Biosciences. European Research Council, https://ror.org/0472cxd90, ERC-2020-ADG–101018672 ENGINES, DeepCell - 101054957 Helmholtz Association’s Initiative and Networking Fund, Interlabs-0029 Max-Planck Society for the Advancement of Science Joachim Herz Stiftung, https://ror.org/01j8qgg43, Add-on Fellowship Boehringer Ingelheim Fonds, https://ror.org/00dkye506, PhD fellowship Konrad Zuse School of Excellence in Learning and Intelligent Systems, DAAD program Konrad Zuse Schools of Excellence in Artificial Intelligence
Erythroderma is an acute and potentially life-threatening inflammatory condition characterized by redness and scaling of > 90% of the skin. Its treatment is challenging because various underlying skin diseases can cause erythroderma and are difficult to distinguish. Here, we performed in-depth proteomics and transcriptomics analyses of skin from 96 patients with erythroderma caused by five different diseases, including pityriasis rubra pilaris, psoriasis, atopic dermatitis, cutaneous T-cell lymphoma, and drug-induced maculopapular rash. High-throughput workflows enabled in-depth molecular profiling, identifying over 9,300 proteins and 17,200 protein coding transcripts, revealing distinct molecular signatures for each disease. The proteome showed elevated expression of type 2 immunity associated Charcot-Leyden crystal in skin of atopic dermatitis, potentially contributing to NLRP3-driven chronic inflammation in this disease. Complementary transcriptomic analysis demonstrated selective upregulation of IL17C in pityriasis rubra pilaris, strongly correlating with increased IL1 family cytokine expression. Interestingly, only a subset of these patients expressed this IL17C-IL1 signature, suggesting treatment-relevant disease endotypes. Through multi-omics integration, we uncovered disease-specific molecular signatures consistently altered at both protein and transcript levels. In particular, we identified elevated expression of T-cell regulator RASAL3 in cutaneous T-cell lymphoma, which has not been explored in its pathogenesis so far. To translate these molecular profiles into clinical utility, we expanded our adaptive machine-learning algorithm (ADAPT-Mx) for tissue based-disease classification. This achieved 76.6% diagnostic accuracy, substantially outperforming combined conventional clinical and histopathological methods (59.5%). This study provides a template for precision diagnostics in erythroderma and demonstrates the clinical potential of multi-omic profiling in severe inflammatory skin diseases.
The response to proteotoxic stresses such as heat shock allows organisms to maintain protein homeostasis under changing environmental conditions. We asked what happens if an organism can no longer react to cytosolic proteotoxic stress. To test this, we deleted or depleted, either individually or in combination, the stress-responsive transcription factors Msn2, Msn4, and Hsf1 in Saccharomyces cerevisiae. Our study reveals a combination of survival strategies, which together protect essential proteins. Msn2 and 4 broadly reprogram transcription, triggering the response to oxidative stress, as well as biosynthesis of the protective sugar trehalose and glycolytic enzymes, while Hsf1 mainly induces the synthesis of molecular chaperones and reverses the transcriptional response upon prolonged mild heat stress (adaptation).
Mass spectrometry (MS)-based proteomics continues to evolve rapidly, opening more and more application areas. The scale of data generated on novel instrumentation and acquisition strategies pose a challenge to bioinformatic analysis. Search engines need to make optimal use of the data for biological discoveries while remaining statistically rigorous, transparent and performant. Here we present alphaDIA, a modular open-source search framework for data independent acquisition (DIA) proteomics. We developed a feature-free identification algorithm particularly suited for detecting patterns in data produced by sensitive time-of-flight instruments. It naturally adapts to novel, more eTicient scan modes that are not yet accessible to previous algorithms. Rigorous benchmarking demonstrates competitive identification and quantification performance. While supporting empirical spectral libraries, we propose a new search strategy named end-to-end transfer learning using fully predicted libraries. This entails continuously optimizing a deep neural network for predicting machine and experiment specific properties, enabling the generic DIA analysis of any post-translational modification (PTM). AlphaDIA provides a high performance and accessible framework running locally or in the cloud, opening DIA analysis to the community. ### Competing Interest Statement MM is an indirect investor in Evosep.
Mass spectrometry (MS) enables specific and accurate quantification of proteins with ever-increasing throughput and sensitivity. Maximizing this potential of MS requires optimizing data acquisition parameters and performing efficient quality control for large datasets. To facilitate these objectives for data-independent acquisition (DIA), we developed a second version of our framework for data-driven optimization of MS methods (DO-MS). The DO-MS app v2.0 (do-ms.slavovlab.net) allows one to optimize and evaluate results from both label-free and multiplexed DIA (plexDIA) and supports optimizations particularly relevant to single-cell proteomics. We demonstrate multiple use cases, including optimization of duty cycle methods, peptide separation, number of survey scans per duty cycle, and quality control of single-cell plexDIA data. DO-MS allows for interactive data display and generation of extensive reports, including publication of quality figures that can be easily shared. The source code is available at github.com/SlavovLab/DO-MS.
Forward genetic screening associates phenotypes with genotypes by randomly inducing mutations and then identifying those that result in phenotypic changes of interest. Here we present spa tially r esolved C RISPR s creening (SPARCS), a platform for microscopy-based genetic screening for spatial cellular phenotypes. SPARCS uses automated high-speed laser microdissection to physically isolate phenotypic variants in situ from virtually unlimited library sizes. We demonstrate the potential of SPARCS in a genome-wide CRISPR-KO screen on autophagosome formation in 40 million cells. Coupled to deep learning image analysis, SPARCS recovered almost all known macroautophagy genes in a single experiment and discovered a role for the ER-resident protein EI24 in autophagosome biogenesis. Harnessing the full power of advanced imaging technologies, SPARCS enables genome-wide forward genetic screening for diverse spatial phenotypes in situ .
Data-independent acquisition (DIA) methods have become increasingly popular in mass spectrometry-based proteomics because they enable continuous acquisition of fragment spectra for all precursors simultaneously. However, these advantages come with the challenge of correctly reconstructing the precursor-fragment relationships in these highly convoluted spectra for reliable identification and quantification. Here, we introduce a scan mode for the combination of trapped ion mobility spectrometry with parallel accumulation-serial fragmentation (PASEF) that seamlessly and continuously follows the natural shape of the ion cloud in ion mobility and peptide precursor mass dimensions. Termed synchro-PASEF, it increases the detected fragment ion current several-fold at sub-second cycle times. Consecutive quadrupole selection windows move synchronously through the mass and ion mobility range. In this process, the quadrupole slices through the peptide precursors, which separates fragment ion signals of each precursor into adjacent synchro-PASEF scans. This precisely defines precursor-fragment relationships in ion mobility and mass dimensions and effectively deconvolutes the DIA fragment space. Importantly, the partitioned parts of the fragment ion transitions provide a further dimension of specificity via a lock-and-key mechanism. This is also advantageous for quantification, where signals from interfering precursors in the DIA selection window do not affect all partitions of the fragment ion, allowing to retain only the specific parts for quantification. Overall, we establish the defining features of synchro-PASEF and explore its potential for proteomic analyses.
Current mass spectrometry methods enable high-throughput proteomics of large sample amounts, but proteomics of low sample amounts remains limited in depth and throughput. To increase the throughput of sensitive proteomics, we developed an experimental and computational framework, called plexDIA, for simultaneously multiplexing the analysis of peptides and samples. Multiplexed analysis with plexDIA increases throughput multiplicatively with the number of labels without reducing proteome coverage or quantitative accuracy. By using three-plex non-isobaric mass tags, plexDIA enables quantification of threefold more protein ratios among nanogram-level samples. Using 1-hour active gradients, plexDIA quantified ~8,000 proteins in each sample of labeled three-plex sets and increased data completeness, reducing missing data more than twofold across samples. Applied to single human cells, plexDIA quantified ~1,000 proteins per cell and achieved 98% data completeness within a plexDIA set while using ~5 minutes of active chromatography per cell. These results establish a general framework for increasing the throughput of sensitive and quantitative protein analysis. Proteomics of small sample sizes using data-independent acquisition methods achieves higher throughput with multiplexing.
Glycan–protein interactions are highly specific yet transient, rendering glycans ideal recognition signals in a variety of biological processes. In human norovirus (HuNoV) infection, histo-blood group antigens (HBGAs) play an essential but poorly understood role. For murine norovirus infection (MNV), sialylated glycolipids or glycoproteins appear to be important. It has also been suggested that HuNoV capsid proteins bind to sialylated ganglioside head groups. Here, we study the binding of HBGAs and sialoglycans to HuNoV and MNV capsid proteins using NMR experiments. Surprisingly, the experiments show that none of the norovirus P-domains bind to sialoglycans. Notably, MNV P-domains do not bind to any of the glycans studied, and MNV-1 infection of cells deficient in surface sialoglycans shows no significant difference compared to cells expressing respective glycans. These findings redefine glycan recognition by noroviruses, challenging present models of infection.
3000 x 3000 px test images for the VIPER pipeline
Bile acids have been reported as important cofactors promoting human and murine norovirus (NoV) infections in cell culture. The underlying mechanisms are not resolved. Through the use of chemical shift perturbation (CSP) NMR experiments, we identified a low-affinity bile acid binding site of a human GII.4 NoV strain. Long-timescale MD simulations reveal the formation of a ligand-accessible binding pocket of flexible shape, allowing the formation of stable viral coat protein-bile acid complexes in agreement with experimental CSP data. CSP NMR experiments also show that this mode of bile acid binding has a minor influence on the binding of histo-blood group antigens and vice versa. STD NMR experiments probing the binding of bile acids to virus-like particles of seven different strains suggest that low-affinity bile acid binding is a common feature of human NoV and should therefore be important for understanding the role of bile acids as cofactors in NoV infection.