Abstract Cross-linking mass spectrometry is a powerful method for structural analysis, but choosing between cleavable and non-cleavable cross-linkers remains challenging. We rigorously compared non-cleavable DSS with cleavable DSSO and found that DSS consistently yields more cross-link identifications from isolated protein complexes to bacterial lysates. The advantage of DSS diminishes as sample complexity increases. At the highest complexity tested—human cell lysate—the trend reverses, with DSSO outperforming DSS. The superior performance of DSS in less complex samples is likely explained by its longer and more flexible spacer arm, which interrogates a spatial volume >40% larger than that of DSSO. For both cross-linkers, the number of identified cross-links decreases as the search space expands, but more steeply for DSS. This sharper decline arises from DSS cross-links producing slightly lower fragment ion coverage, not from the absence of signature ions that could reduce search space. Fragment ion coverage is key to interactome mapping: when coverage reaches 85% or above, identification sensitivity hardly decreases as the search space expands, regardless of the cross-linker used. In summary, we recommend DSS for samples no more complex than bacterial lysates. For interactome mapping of mammalian cells, although DSSO outperforms DSS, neither achieves deep interactome coverage.
Chemical cross-linking of proteins coupled with mass spectrometry provides structural insights by identifying cross-linked peptide pairs, abbreviated as cross-links. Presently, cross-link identification suffers from ambiguity and poor sensitivity because they are typically of lower abundance and consequently of lower MS2 quality than linear peptides present in the same sample. Here, we present target-enhanced accurate inclusion mass screening (TAIMS), a meticulously optimized targeted mass spectrometry method. TAIMS significantly improved the quality of MS2, as indicated by fragment ion coverage (FIC) and other metrics. From data-dependent acquisition (DDA) to TAIMS, high-FIC cross-links increased by 359 or 678 on yeast ribosome or Escherichia coli lysate, respectively, or from about 40% to around 90%. As a result, TAIMS substantially enhanced the accuracy of cross-link site localization and mitigated sensitivity loss in cross-link identification from large database searches. The enhanced identification sensitivity of TAIMS is further evidenced by its capacity to recover genuine cross-link identifications from data that would typically be discarded. On cross-linked E. coli lysate, 10.5% (284/2711) of the inclusion-list entries generated from unidentified cross-link-spectrum matches gained identity through TAIMS. Of these, 230 were linear peptides and 54 were cross-links, including 10 intermolecular cross-links missed entirely by DDA. Additionally, we demonstrate that TAIMS is a general method for the identification of low-abundance, post-translationally modified peptides. On a mouse brain sample, TAIMS increased the number of phosphopeptides identified with an accurate phosphosite assignment by 67%. These findings indicate that TAIMS has broad applicability in proteomics.
Ubiquitin-like proteins (UBLs) constitute a family of evolutionarily conserved proteins that share similarities with ubiquitin in 3D structures and modification mechanisms. For most UBLs including small-ubiquitin-like modifiers (SUMO), their modification sites on substrate proteins cannot be identified using the mass spectrometry-based method that has been successful for identifying ubiquitination sites, unless a UBL protein is mutated accordingly. To identify UBL modification sites without having to mutate UBL, we have developed a dedicated search engine pLink-UBL on the basis of pLink, a software tool for the identification of cross-linked peptide pairs. pLink-UBL exhibited superior precision, sensitivity, and speed than "make-do" search engines such as MaxQuant, pFind, and pLink. For example, compared to MaxQuant, pLink-UBL increased the number of identified SUMOylation sites by 50 ∼ 300% from the same datasets. Additionally, we present a method for identifying small-molecule modifications of UBLs. This method involves antibody enrichment of a UBL C-terminal peptide following enrichment of a UBL protein, followed by LC-MS/MS analysis and a pFind 3 blind search to identify unexpected modifications. Using this method, we have discovered nonprotein substrates of SUMO, of which spermidine is the major one for fission yeast SUMO Pmt3. Spermidine can be conjugated to the C-terminal carboxylate group of Pmt3 through its N1 or also likely, N8 amino group in the presence of SUMO E1, E2, and ATP. Pmt3-spermidine conjugation does not require E3 and can be reversed by SUMO isopeptidase Ulp1. SUMO-spermidine conjugation is present in mice and humans. Also, spermidine can be conjugated to ubiquitin in vitro by E1 and E2 in the presence of ATP. The above observations suggest that spermidine may be a common small molecule substrate of SUMO and possibly ubiquitin across eukaryotic species.
Mass spectrometry-based proteomics aims to identify peptides and proteins to give direct proofs of gene expressions, analyze structures and functions of proteins, study the relationship between proteins and diseases, and provide targeted treatment options. All these studies are based on the credibility of identified peptides and proteins. However, it is impossible to manually check all identified peptides because a large number of identifications can be collected from one mass spectrometry experiment. Thus, target-decoy approach (TDA) is proposed and always used to control the quality of identified peptides and proteins, and has been expanded to subclasses of peptides (including ordinary subclasses of peptides, variant peptides, and modified peptides) and cross-linking peptides. However, TDA still has two limitations:(1) the estimation of false discovery rate (FDR) is inaccurate and (2) validation of single identification cannot be supported. Thus, the identification results that passed the TDA-based FDR control need to be further validated and other validation methods which are used after TDA-FDR filtration (referred to as Beyond-TDA methods) have been developed to enhance peptide validation.This paper reviews TDA and its extensions as well as Beyond-TDA methods and discusses the advantages and disadvantages of each method. In the first part of this paper, we introduce the goal of proteomics, the process of mass spectrometry acquisition and analysis, the validation problem, and the early statistical methods to evaluate the identification credibility. Then, in the second part of this paper, we describe in detail the ordinary TDA-FDR method, including the assumption that random matches are equally likely to appear in target and decoy databases,the construction methods to generate the decoy database, and the computational formula of TDA-FDR. We also introduce the extensions of TDA-FDR on ordinary subclasses of peptides, variant peptides, modified peptides,proteogenomics peptides, cross-linking peptides, and glycopeptides. However, TDA cannot model the homologous incorrect peptides, thus TDA-FDR underestimates the actual false rate. So, after TDA-FDR filtration,it is necessary to use more strict validation methods,i.e., Beyond-TDA methods, which are reviewed in detail in the third part of this paper, to control validation credibility. In this part, four kinds of methods are introduced,including validation methods based on search space (trap database validation and open search validation), spectra similarity (synthetic peptide validation and theoretical spectra prediction), chemical information (retention time prediction and stable isotopic labeling validation) and machine learning technology (Percolator, pValid, and DeepRescore). Lastly, we summarize the content of this paper and discuss the future improvement directions of validation methods
The remarkable advancement of top-down proteomics in the past decade is driven by the technological development in separation, mass spectrometry (MS) instrumentation, novel fragmentation, and bioinformatics. However, the accurate identification and quantification of proteoforms, all clearly-defined molecular forms of protein products from a single gene, remain a challenging computational task. This is in part due to the complicated mass spectra from intact proteoforms when compared to those from the digested peptides. Herein, pTop 2.0 is developed to fill in the gap between the large-scale complex top-down MS data and the shortage of high-accuracy bioinformatic tools. Compared with pTop 1.0, the first version, pTop 2.0 concentrates mainly on the identification of the proteoforms with unexpected modifications or a terminal truncation. The quantitation based on isotopic labeling is also a new function, which can be carried out by the convenient and user-friendly "one-key operation," integrated together with the qualitative identifications. The accuracy and running speed of pTop 2.0 is significantly improved on the test data sets. This chapter will introduce the main features, step-by-step running operations, and algorithmic developments of pTop 2.0 in order to push the identification and quantitation of intact proteoforms to a higher-accuracy level in top-down proteomics.
Chemical cross-linking coupled with mass spectrometry (CXMS) is an important tool to analyze protein structures and protein-protein interactions. In the last five years, CXMS has made great progress in both methods and applications. In terms of methods, on the one hand, cleavable cross-linkers and new separation and enrichment methods have shown good prospects, and on the other hand, more efficient cross-linked peptide search engines and quality control methods provide powerful tools for CXMS data analysis. In terms of applications, on the one hand, CXMS combined with cryo-electron microscopy has determined a large number of protein structures, and on the other hand, CXMS has shown the potential to analyze protein-protein interaction networks at a proteome scale. The intensive research on CXMS in methods and applications reflect the important role of this technology. Here we review the various aspects of CXMS, including cross-linker selection, cross-linking reaction, protein digestion, separation and enrichment, data acquisition, cross-linked peptide identification, quality control, and application, and mainly focus on progress in the last five years. Lastly, we discuss the challenges and opportunities of CXMS in the future. In section 1, we review cross-linkers from the aspects of reactive group and spacer arm. In section 2, we give tips for the cross-linking reaction. In section 3, we describe the sequential digestion strategy for cross-linked proteins. In section 4, we elaborate enrichment methods for cross-linked peptides, including affinity purification, chromatographic separation, and ion mobility mass spectrometry. In section 5, we elaborate data acquisition methods for cross-linked peptides, and compare three methods for MS-cleavable cross-linked peptides. In section 6, we elaborate search engines for cross-linked peptide identification. In section 7, we describe quality control methods for cross-linked peptide identification. In section 8, we review applications of CXMS and list some proteome-wide CXMS studies. In section 9, we conclude the paper and discuss the challenges and opportunities of CXMS in the future.
Tandem mass spectrometry has been the principal method in shotgun proteomics for peptide and protein identification. However, incorrect identifications reported by proteome search engines are still unknown, and further validation methods are needed. We have proposed a validation method pValid before, but its scope of application is limited because two features used in pValid are related to open database search and sub-optimal peptide candidates for tandem mass spectra, and the performance on complex datasets still has room for improvement. In this study, we developed a more comprehensive validation method, pValid 2, to break these limitations by removing the two features and bringing in a new feature related to the retention time predicted by a deep learning-based method pPredRT. pValid 2 yielded an average false positive rate of 0.03% and an average false negative rate of 1.37% on three testing datasets, better than those of pValid, and flagged 8.47% to 11.31% more incorrect identifications than pValid on two complex datasets. Moreover, pValid 2 flagged almost all decoy identifications in validating the open-search datasets. In addition, the function of validating identifications given by MaxQuant and MS-GF+ was implemented in pValid 2, and the validation results showed that pValid 2 performed dramatically better than three metabolic labeling validation methods. Further considering its cost-effectiveness as a pure computational approach, pValid 2 has the potential to be a widely used validation tool for peptide identifications of any proteome search engines in shotgun proteomics. SIGNIFICANCE: Identification results given by shotgun proteomics are vital to life science research. The correctness of identifications deeply affects the precision of the subsequent studies about protein structures and functions, protein-protein interactions, pathogenic mechanism, and targeted drugs. Thus, validating the correctness of identifications is crucial and urgent. In 2019, we developed an identification credibility validation method named pValid, whose false positive rate (FPR) is 0.03% and false negative rate (FNR) is 1.79%, comparable to those of the gold standard, i.e., the Synthetic-peptide validation method. However, pValid can only be used for validating the results from pFind, and its validation performance on a few complex datasets still has room for improvement. So, in this submission, we proposed pValid 2, a more comprehensive computational validation method that can validate identifications from any proteome search engines with increased discriminating power.
Great advances have been made in mass spectrometric data interpretation for intact glycopeptide analysis. However, accurate identification of intact glycopeptides and modified saccharide units at the site-specific level and with fast speed remains challenging. Here, we present a glycan-first glycopeptide search engine, pGlyco3, to comprehensively analyze intact N- and O-glycopeptides, including glycopeptides with modified saccharide units. A glycan ion-indexing algorithm developed for glycan-first search makes pGlyco3 5-40 times faster than other glycoproteomic search engines without decreasing accuracy or sensitivity. By combining electron-based dissociation spectra, pGlyco3 integrates a dynamic programming-based algorithm termed pGlycoSite for site-specific glycan localization. Our evaluation shows that the site-specific glycan localization probabilities estimated by pGlycoSite are suitable to localize site-specific glycans. With pGlyco3, we confidently identified N-glycopeptides and O-mannose glycopeptides that were extensively modified by ammonia adducts in yeast samples. The freely available pGlyco3 is an accurate and flexible tool that can be used to identify glycopeptides and modified saccharide units.
We present a glycan-first glycopeptide search engine, pGlyco3, to comprehensively analyze intact N- and O-glycopeptides, including glycopeptides with modified saccharide units. A novel glycan ion-indexing algorithm developed in this work for glycan-first search makes pGlyco3 5-40 times faster than other glycoproteomic search engines without decreasing the accuracies and sensitivities. By combining electron-based dissociation spectra, pGlyco3 integrates a fast, dynamic programming-based algorithm termed pGlycoSite for site-specific glycan localization (SSGL). Our evaluation based on synthetic and natural glycopeptides showed that the SSGL probabilities estimated by pGlycoSite were proved to be appropriate to localize site-specific glycans. With pGlyco3, we found that N-glycopeptides and O-mannose glycopeptides in yeast samples were extensively modified by ammonia adducts on Hex (aH) and verified the aH-glycopeptide identifications based on released N-glycans and 15N/13C-labeled data. Thus pGlyco3, which is freely available on https://github.com/pFindStudio/pGlyco3/releases, is an accurate and flexible tool to identify glycopeptides and modified saccharide units.
In cross-linking mass spectrometry, the identification of cross-linked peptide pairs heavily relies on the ability of a database search engine to measure the similarities between experimental and theoretical MS/MS spectra. However, the lack of accurate ion intensities in theoretical spectra impairs the performance of search engines, in particular, on proteome scales. Here we introduce pDeepXL, a deep neural network to predict MS/MS spectra of cross-linked peptide pairs. To train pDeepXL, we used the transfer-learning technique because it facilitated the training with limited benchmark data of cross-linked peptide pairs. Test results on more than ten data sets showed that pDeepXL accurately predicted the spectra of both noncleavable DSS/BS3/Leiker crosslinked peptide pairs (>80% of predicted spectra have Pearson's r values higher than 0.9) and cleavable DSSO/DSBU cross-linked peptide pairs (>75% of predicted spectra have Pearson's r values higher than 0.9). pDeepXL also achieved the accurate prediction on unseen data sets using an online fine-tuning technique. Lastly, integrating pDeepXL into a database search engine increased the number of identified cross-link spectra by 18% on average.
We describe pLink 2, a search engine with higher speed and reliability for proteome-scale identification of cross-linked peptides. With a two-stage open search strategy facilitated by fragment indexing, pLink 2 is ~40 times faster than pLink 1 and 3~10 times faster than Kojak. Furthermore, using simulated datasets, synthetic datasets, 15 N metabolically labeled datasets, and entrapment databases, four analysis methods were designed to evaluate the credibility of ten state-of-the-art search engines. This systematic evaluation shows that pLink 2 outperforms these methods in precision and sensitivity, especially at proteome scales. Lastly, re-analysis of four published proteome-scale cross-linking datasets with pLink 2 required only a fraction of the time used by pLink 1, with up to 27% more cross-linked residue pairs identified. pLink 2 is therefore an efficient and reliable tool for cross-linking mass spectrometry analysis, and the systematic evaluation methods described here will be useful for future software development.
The usage of design techniques in design processes is an important driver for the success of digital services. However, before using design techniques, suitable techniques need to be selected. With the continuous growth of the number of design techniques, the selection of appropriate ones becomes more difficult, especially for design novices with limited knowledge and expertise. In order to support the selection process, we propose design principles for the development of an advisory platform that interacts with design novices to suggest design techniques for different design situations using artificial intelligence (AI) techniques. Specifically, we leverage conversational agents, recommender techniques, and taxonomic background knowledge to conceptualize and implement an AI-based advisory platform. Following a design science research methodology, we contribute design knowledge for the class of advanced advisory platforms. Furthermore, from a practical point of view, we help design novices with our implemented advisory platform in the contextualized selection process of design techniques.
In the past decade, tandem mass spectrometry (MS/MS)-based bottom-up proteomics has become the method of choice for analyzing post-translational modifications (PTMs) in complex mixtures. The key to the identification of the PTM-containing peptides and localization of the PTM-modified residues is to measure the similarities between the theoretical spectra and the experimental ones. An accurate prediction of the theoretical MS/MS spectra of the modified peptides will improve the similarity measurement. Here, we proposed the deep-learning-based pDeep2 model for PTMs. We used the transfer learning technique to train pDeep2, facilitating the training with a limited scale of benchmark PTM data. Using the public synthetic PTM data sets, including the synthetic phosphopeptides and 21 synthetic PTMs from ProteomeTools, we showed that the model trained by transfer learning was accurate (>80% Pearson correlation coefficients were higher than 0.9), and was significantly better than the models trained without transfer learning. We also showed that accurate prediction of the fragment ion intensities of the PTM neutral loss, for example, the phosphoric acid loss (-98 Da) of the phosphopeptide, will improve the discriminating power to distinguish the true phosphorylated residue from its adjacent candidate sites. pDeep2 is available at https://github.com/pFindStudio/pDeep/tree/master/pDeep2 .
Mass spectrometry (MS) analysis of peptides has traditionally been conducted in the positive ion mode on protonated peptides, although negative ion-mode analysis of deprotonated peptides can provide complementary and sometimes critical structural information. This is partly due to insufficient understanding of the fragmentation behaviors of deprotonated peptides. Here, using ion-trap collision-induced dissociation (CID) and higher-energy collisional dissociation (HCD, a beam-type CID), we characterized the fragmentation patterns of 36 deprotonated peptides of 4–16 amino acids on a high-resolution, high-mass accuracy instrument. Our study finds that among the backbone cleavage products, y-, c- and z-type ions (using a nomenclature similar to that of positive ions) are the most dominant species in both CID and HCD spectra of deprotonated peptides, accompanied by abundant neutral loss (NL) peaks. Similar to the charge-dependency of collisional energy of protonated peptides, we find that singly charged deprotonated peptides require significantly higher collisional energy than their doubly or multiply charged counterparts to reach 50% fragmentation. HCD is generally better than CID for peptide sequencing in the negative ion mode, since HCD generates more backbone cleavage products whereas CID produces predominately NL peaks of precursors. For disulfide-bonded peptides and C-terminally amidated peptides, unusual fragmentation is observed in the negative ion mode. The fragmentation behaviors of deprotonated peptides reported in this study will promote further investigation of the fundamental mechanisms and facilitate algorithm development for peptide sequencing in the negative mode.
We study theoretically and experimentally the extent to which communication can solve coordination problems when there is some conflict of interest. We investigate various communication protocols, including one in which players chat sequentially and free-format. We develop a model based on the ‘feigned-ignorance principle’, according to which players ignore any communication unless they reach an agreement in which both players are (weakly) better off. With standard preferences, the model predicts that communication is effective in Battle-of-the-Sexes but futile in Chicken. A remarkable implication is that increasing players' payoffs can make them worse off, by making communication futile. Our experimental findings provide strong support for these and some other predictions.
As the de facto validation method in mass spectrometry-based proteomics, the target-decoy approach determines a threshold to estimate the false discovery rate and then filters those identifications beyond the threshold. However, the incorrect identifications within the threshold are still unknown and further validation methods are needed. In this study, we characterized a framework of validation and investigated a number of common and novel validation methods. We first defined the accuracy of a validation method by its false-positive rate (FPR) and false-negative rate (FNR) and, further, proved that a validation method with lower FPR and FNR led to identifications with higher sensitivity and precision. Then we proposed a validation method named pValid that incorporated an open database search and a theoretical spectrum prediction strategy via a machine-learning technology. pValid was compared with four common validation methods as well as a synthetic peptide validation method. Tests on three benchmark data sets indicated that pValid had an FPR of 0.03% and an FNR of 1.79% on average, both superior to the other four common validation methods. Tests on a synthetic peptide data set also indicated that the FPR and FNR of pValid were better than those of the synthetic peptide validation method. Tests on a large-scale human proteome data set indicated that pValid successfully flagged the highest number of incorrect identifications among all five methods. Further considering its cost-effectiveness, pValid has the potential to be a feasible validation tool for peptide identification.
Motivation De novo peptide sequencing based on tandem mass spectrometry data is the key technology of shotgun proteomics for identifying peptides without any database and assembling unknown proteins. However, owing to the low ion coverage in tandem mass spectra, the order of certain consecutive amino acids cannot be determined if all of their supporting fragment ions are missing, which results in the low precision of de novo sequencing. Results In order to solve this problem, we developed pNovo 3, which used a learning-to-rank framework to distinguish similar peptide candidates for each spectrum. Three metrics for measuring the similarity between each experimental spectrum and its corresponding theoretical spectrum were used as important features, in which the theoretical spectra can be precisely predicted by the pDeep algorithm using deep learning. On seven benchmark datasets from six diverse species, pNovo 3 recalled 29-102% more correct spectra, and the precision was 11-89% higher than three other state-of-the-art de novo sequencing algorithms. Furthermore, compared with the newly developed DeepNovo, which also used the deep learning approach, pNovo 3 still identified 21-50% more spectra on the nine datasets used in the study of DeepNovo. In summary, the deep learning and learning-to-rank techniques implemented in pNovo 3 significantly improve the precision of de novo sequencing, and such machine learning framework is worth extending to other related research fields to distinguish the similar sequences. Availability and implementation pNovo 3 can be freely downloaded from http://pfind.ict.ac.cn/software/pNovo/index.html. Supplementary information Supplementary data are available at Bioinformatics online.
Chemical cross-linking of proteins coupled with mass spectrometry analysis (CXMS) is widely used to study protein-protein interactions (PPI), protein structures, and even protein dynamics. However, structural information provided by CXMS is still limited, partly because most CXMS experiments use lysine-lysine (K-K) cross-linkers. Although superb in selectivity and reactivity, they are ineffective for lysine deficient regions. Herein, we develop aromatic glyoxal cross-linkers (ArGOs) for arginine-arginine (R-R) cross-linking and the lysine-arginine (K-R) cross-linker KArGO. The R-R or K-R cross-links generated by ArGO or KArGO fit well with protein crystal structures and provide information not attainable by K-K cross-links. KArGO, in particular, is highly valuable for CXMS, with robust performance on a variety of samples including a kinase and two multi-protein complexes. In the case of the CNGP complex, KArGO cross-links covered as much of the PPI interface as R-R and K-K cross-links combined and improved the accuracy of Rosetta docking substantially.