Knowledge graphs have emerged as a powerful paradigm for structuring, organizing and reasoning over complex scientific knowledge, and are increasingly recognized as catalysts for accelerating AI for science. This study provides a comprehensive survey of scientific knowledge graphs (SciKGs), covering their construction methodologies and diverse applications across biology, chemistry and materials science. We examine how SciKGs support tasks such as drug development, omics analysis, reaction prediction and materials design, and highlight how the synergistic integration of SciKGs and large language models (LLMs) forms a knowledge- and language-driven framework for scientific discovery, in which SciKGs serve as the foundational knowledge infrastructure and LLMs act as dynamic semantic engines. We further identify key challenges and outline emerging opportunities for building auditable, interoperable and self-evolving SciKGs. Looking forward, we envision a new generation of SciKG-centered ecosystems where self-updating graphs, co-evolving with LLMs and embodied within AI scientists, become core infrastructures that autonomously drive, verify and accelerate scientific discovery.
Simultaneous characterization of proteomic and metabolomic profiles at the single-cell level is crucial for deciphering cellular heterogeneity and elucidating disease mechanisms. However, it is still a great challenge to achieve high-depth dual-omics analysis in the same single cell. Here, we propose a unified strategy called one-shot hybrid-mode single-cell proteome and metabolome analysis (hybrid-scPMA), in which the mass spectrometry (MS) detection mode of data-independent acquisition (DIA) is utilized for the analysis of peptides from protein digestion, and the data-dependent acquisition (DDA) mode is used for metabolite analysis in a single liquid chromatography-mass spectrometry (LC-MS) analysis run, enabling deep analysis of the proteome and metabolome in single-cell samples. Building upon the strategy, we established an improved single-cell multiomics analysis workflow that integrated automated single-cell capture, simplified sample pretreatment, LC injection and separation, and DIA-DDA hybrid-mode MS detection. With this approach, we achieved an average identification of 3510 protein groups and 255 metabolites from single HepG2 cells, representing a substantial increase in identification depth over previous approaches. We also performed time-resolved proteomic and metabolomic profiling of HepG2 single cells undergoing sorafenib drug intervention, resolving drug response characteristics at the single-cell level and providing multiomics insights into drug mechanisms from a proteo-metabolomic perspective.
Retrieving molecular structures from tandem mass spectra is a crucial step in rapid compound identification. Existing retrieval methods, such as traditional mass spectral library matching, suffer from limited spectral library coverage, while recent cross-modal representation learning frameworks often encounter modality misalignment, resulting in suboptimal retrieval accuracy and generalization. To address these limitations, we propose GLMR, a Generative Language Model-based Retrieval framework that mitigates the cross-modal misalignment through a two-stage process. In the pre-retrieval stage, a contrastive learning-based model identifies top candidate molecules as contextual priors for the input mass spectrum. In the generative retrieval stage, these candidate molecules are integrated with the input mass spectrum to guide a generative model in producing refined molecular structures, which are then used to re-rank the candidates based on molecular similarity. Experiments on both MassSpecGym and the proposed MassRET-20k dataset demonstrate that GLMR significantly outperforms existing methods, achieving over 40% improvement in top-1 accuracy and exhibiting strong generalizability.
Recent years have seen a rise of single-cell proteomics by data-independent acquisition mass spectrometry (DIA MS). While diverse data analysis strategies have been reported in literature, their impact on the outcome of single-cell proteomic experiments has been rarely investigated. Here, we present a framework for benchmarking data analysis strategies for DIA-based single-cell proteomics. This framework provides a comprehensive comparison of popular DIA data analysis software tools and searching strategies, as well as a systematic evaluation of method combinations in subsequent informatic workflow, including sparsity reduction, missing value imputation, normalization, batch effect correction, and differential expression analysis. Benchmarking on simulated single-cell samples consisting of mixed proteomes and real single-cell samples with a spike-in scheme, recommendations are provided for the data analysis for DIA-based single-cell proteomics.
The efficacy of cancer immunotherapy is significantly influenced by the heterogeneity of individual tumors and immune responses. To investigate this phenomenon, a microfluidic platform is constructed for profiling immune-cancer cell interactions at the single-cell proteomics level for the first time. Based on the platform, a comprehensive workflow is proposed for achieving accurate single-cell pairing of an immune cell and a cancer cell with low cell damage and high success rate up to 95%, cell pair co-culture, and real-time microscopic monitoring of the cell-pair interactions, cell pair retrieval, mass spectrometry-based proteomic analysis of singe cell pairs, and decoupling of the proteomic information for each cell within the cell pair with the stable-isotope labeling method. With the workflow, the interactions of single natural killer (NK) cells and single K562 tumor cells are investigated based on real-time images and single cell-pair proteomics. Notably, an identification depth of over 1000 protein groups in a single cell-pair is achieved, leading to the discovery of sub-clusters of NK cells with different functions and the identification of important biomarkers for cancer treatments. This demonstrates the unique capability of the present platform in providing substantial and comprehensive datasets for profiling immune-cancer cell interactions, discovering heterogeneous immune responses, and predicting biomarkers in the study of cancer immunotherapy.
Investigating the heterogeneous responses of individual cancer cells to chemotherapeutic drugs is crucial for deciphering the mechanisms of cancer drug resistance. In recent years, single-cell proteomics has demonstrated its significant capability in exploring drug response in cancer cells. Meanwhile, there are increasing reports suggesting that the cellular morphology is potentially associated with drug resistance. However, integrating the single-cell proteomic results with morphological information remains challenging. Here, we present a morphology-aware single-cell proteomic analysis (Morp-SCP) platform to precisely capture cells of interest with real-time and high-resolution imaging, and to conduct deep proteomic analysis at the single-cell level, providing multidimensional information on the target single cells. The Morp-SCP platform was applied for exploring the time-dependent proteomic alterations of human nonsmall cell lung cancer cells (A549) upon cisplatin exposure. Subpopulations of drug-resistant A549 cells were identified, which exhibited distinct proteomic and morphological patterns when resisting cell death induced by cisplatin exposure. By revealing the proteomics-morphology relationship, the Morp-SCP platform offers an effective strategy to provide insights into the heterogeneity of drug resistance at the single-cell level.
Matrix-assisted laser desorption/ionization time-of-flight mass spectrometry (MALDI-TOF MS) has been widely used for identification of microorganisms. In a typical MALDI-TOF MS analysis of microorganisms, spectra of unknown samples are compared to reference libraries of spectra of known microorganisms by spectral pattern matching. This chapter provides an overview of the data analysis workflow for MALDI-TOF MS-based identification of microorganisms, including spectrum preprocessing, spectral matching, and result interpretation. The existing computational methods for the three steps of data analysis and available software solutions are summarized. In addition, bioinformatic methods that do not require a reference spectral library are introduced as alternatives to typical spectral matching approaches. Finally, the current challenges and outlook of MALDI-TOF MS data analysis for microorganism identification are discussed.
Proteome analysis currently heavily relies on tandem mass spectrometry (MS/MS), which does not fully utilize MS1 features, as many precursors remain unselected for MS/MS fragmentation, especially in the cases of low abundance samples and wide abundance dynamic range samples. Therefore, leveraging MS1 features as a complement to MS/MS has become an attractive option to improve the coverage of feature identification. Herein, we propose MonoMS1, an approach combining deep learning-based retention time, ion mobility, detectability prediction, and logistic regression-based scoring for MS1 feature identification. The approach achieved a significant increase in MS1 feature identification based on an E. coli data set. Application of MonoMS1 to data sets with wide dynamic range, such as human serum proteome samples, and with low sample abundance, such as single-cell proteome samples, enabled substantial complementation of MS/MS-based peptide and protein identification. This method opens a new avenue for proteomic analysis and can boost proteomic research on complex samples.
Single-cell proteomics allows revealing precisely the differences of proteins between individual cells,which has become a research hotspot showing indispensable application value in many important fields.Its difficulties lie in the fact that the proteins in a single cell are of extremely low abundance,which calls for ingenious solutions to the problems of sample loss during preparation,low sensitivity of chromatography-mass spectrometry detection,and insufficient analysis of spectral data with low signal intensities.This review summarizes the current research progress of mass spectrometry-based single-cell proteomic analysis,including single-cell sorting,sample preparation,chromatography-mass spectrometry acquisition,and data analysis,as well as its applications in biomedical fields.Its potential future development is also discussed.
Intact glycopeptide characterization by mass spectrometry has proven to be a versatile tool for site-specific glycoproteomics analysis and biomarker screening. Here, we present a method using a new model of a Q-TOF instrument equipped with a Zeno trap for intact glycopeptide identification and demonstrate its ability to analyze large-cohort glycoproteomes. From 124 clinical serum samples of breast cancer, noncancerous diseases, and nondisease controls, a total of 6901 unique site-specific glycans on 807 glycosites of proteins were detected. Much more differences of glycoproteome were observed in breast diseases than the proteome. By employing machine learning, 15 site-specific glycans were determined as potential glyco-signatures in detecting breast cancer. The results demonstrate that our method provides a powerful tool in glycoproteomic studies.
The shotgun proteomic analysis is currently the most promising single-cell protein sequencing technology, however its identification level of ~1000 proteins per cell is still insufficient for practical applications. Here, we develop a pick-up single-cell proteomic analysis (PiSPA) workflow to achieve a deep identification capable of quantifying up to 3000 protein groups in a mammalian cell using the label-free quantitative method. The PiSPA workflow is specially established for single-cell samples mainly based on a nanoliter-scale microfluidic liquid handling robot, capable of achieving single-cell capture, pretreatment and injection under the pick-up operation strategy. Using this customized workflow with remarkable improvement in protein identification, 2449–3500, 2278–3257 and 1621–2904 protein groups are quantified in single A549 cells ( n = 37), HeLa cells ( n = 44) and U2OS cells ( n = 27) under the DIA (MBR) mode, respectively. Benefiting from the flexible cell picking-up ability, we study HeLa cell migration at the single cell proteome level, demonstrating the potential in practical biological research from single-cell insight.
Deep learning has achieved a notable success in mass spectrometry-based proteomics and is now emerging in glycoproteomics. While various deep learning models can predict fragment mass spectra of peptides with good accuracy, they cannot cope with the non-linear glycan structure in an intact glycopeptide. Herein, we present DeepGlyco, a deep learning-based approach for the prediction of fragment spectra of intact glycopeptides. Our model adopts tree-structured long-short term memory networks to process the glycan moiety and a graph neural network architecture to incorporate potential fragmentation pathways of a specific glycan structure. This feature is beneficial to model explainability and differentiation ability of glycan structural isomers. We further demonstrate that predicted spectral libraries can be used for data-independent acquisition glycoproteomics as a supplement for library completeness. We expect that this work will provide a valuable deep learning resource for glycoproteomics.
Metaproteomics offers a direct avenue to identify microbial proteins in microbiota, enabling the compositional and functional characterization of microbiota. Due to the complexity and heterogeneity of microbial communities, in-depth and accurate metaproteomics faces tremendous limitations. One challenge in metaproteomics is the construction of a suitable protein sequence database to interpret the highly complex metaproteomic data, especially in the absence of metagenomic sequencing data. Herein, we present a high-abundance protein-guided hybrid spectral library strategy for in-depth data independent acquisition (DIA) metaproteomic analysis (HAPs-hyblibDIA). A dedicated high-abundance protein database of gut microbial species is constructed and used to mine the taxonomic information on microbiota samples. Then, a sample-specific protein sequence database is built based on the taxonomic information using Uniprot protein sequence for subsequent analysis of the DIA data using hybrid spectral library-based DIA analysis. We evaluated the accuracy and sensitivity of the method using synthetic microbial community samples and human gut microbiome samples. It was demonstrated that the strategy can successfully identify taxonomic compositions of microbiota samples and that the peptides identified by HAPs-hyblibDIA overlapped greatly with the peptides identified using a metagenomic sequencing-derived database. At the peptide and species level, our results can serve as a complement to the results obtained using a metagenomic sequencing-derived database. Furthermore, we validated the applicability of the HAPs-hyblibDIA strategy in a cohort of human gut microbiota samples of colorectal cancer patients and controls, highlighting its usability in biomedical research.
Microbiota are closely associated to human health and disease. Metaproteomics can provide a direct means to identify microbial proteins in microbiota for compositional and functional characterization. However, in-depth and accurate metaproteomics is still limited due to the extreme complexity and high diversity of microbiota samples. One of the main challenges is constructing a protein sequence database that best fits the microbiota sample. Herein, we proposed an accurate taxonomic annotation pipeline from metagenomic data for deep metaproteomic coverage, namely contigs directed gene annotation (ConDiGA). We mixed 12 known bacterial species to derive a synthetic microbial community to benchmark metagenomic and metaproteomic pipelines. With the optimized taxonomic annotation strategy by ConDiGA, we built a protein sequence database from the metagenomic data for metaproteomic analysis and identified about 12,000 protein groups, which was very close to the result obtained with the reference proteome protein sequence database of the 12 species. We also demonstrated the practicability of the method in real fecal samples, achieved deep proteome coverage of human gut microbiome, and compared the function and taxonomy of gut microbiota at metagenomic level and metaproteomic level. Our study can tackle the current taxonomic annotation reliability problem in metagenomics-derived protein sequence database for metaproteomics. The unique dataset of metagenomic and the metaproteomic data of the 12 bacterial species is publicly available as a standard benchmarking sample for evaluating various analysis pipelines. The code of ConDiGA is open access at GitHub for the analysis of real microbiota samples.
IntroductionThe special flavor and fragrance of Chinese liquor are closely related to microorganisms in the fermentation starter Daqu. The changes of microbial community can affect the stability of liquor yield and quality.MethodsIn this study, we used data-independent acquisition mass spectrometry (DIA-MS) for cohort study of the microbial communities of a total of 42 Daqu samples in six production cycles at different times of a year. The DIA MS data were searched against a protein database constructed by metagenomic sequencing.ResultsThe microbial composition and its changes across production cycles were revealed. Functional analysis of the differential proteins was carried out and the metabolic pathways related to the differential proteins were explored. These metabolic pathways were related to the saccharification process in liquor fermentation and the synthesis of secondary metabolites to form the unique flavor and aroma in the Chinese liquor.DiscussionWe expect that the metaproteome profiling of Daqu from different production cycles will serve as a guide for the control of fermentation process of Chinese liquor in the future.
Metaproteomics can provide valuable insights into the functions of human gut microbiota (GM), but is challenging due to the extreme complexity and heterogeneity of GM. Data-independent acquisition (DIA) mass spectrometry (MS) has been an emerging quantitative technique in conventional proteomics, but is still at the early stage of development in the field of metaproteomics. Herein, we applied library-free DIA (directDIA)-based metaproteomics and compared the directDIA with other MS-based quantification techniques for metaproteomics on simulated microbial communities and feces samples spiked with bacteria with known ratios, demonstrating the superior performance of directDIA by a comprehensive consideration of proteome coverage in identification as well as accuracy and precision in quantification. We characterized human GM in two cohorts of clinical fecal samples of pancreatic cancer (PC) and mild cognitive impairment (MCI). About 70,000 microbial proteins were quantified in each cohort and annotated to profile the taxonomic and functional characteristics of GM in different diseases. Our work demonstrated the utility of directDIA in quantitative metaproteomics for investigating intestinal microbiota and its related disease pathogenesis.
IntroductionDaqu, the Chinese liquor fermentation starter, contains complex microbial communities that are important for the yield, quality, and unique flavor of produced liquor. However, the composition and metabolism of microbial communities in the different types of high-temperature Daqu (i.e., white, yellow, and black Daqu) have not been well understood. MethodsHerein, we used quantitative metaproteomics based on data-independent acquisition (DIA) mass spectrometry to analyze a total of 90 samples of white, yellow, and black Daqu collected in spring, summer, and autumn, revealing the taxonomic and metabolic profiles of different types of Daqu across seasons. ResultsTaxonomic composition differences were explored across types of Daqu and seasons, where the under-fermented white Daqu showed the higher microbial diversity and seasonal stability. It was demonstrated that yellow Daqu had higher abundance of saccharifying enzymes for raw material degradation. In addition, considerable seasonal variation of microbial protein abundance was discovered in the over-fermented black Daqu, suggesting elevated carbohydrate and amino acid metabolism in autumn black Daqu. DiscussionWe expect that this study will facilitate the understanding of the key microbes and their metabolism in the traditional fermentation process of Chinese liquor production.
Large-scale profiling of intact glycopeptides is critical but challenging in glycoproteomics. Data-independent acquisition (DIA) mass spectrometry is an emerging technology with deep proteome coverage as well as accurate quantitative capability for large-scale proteomics studies and has also been applied to the field of glycoproteomics. In this protocol, we describe how to analyze data from a DIA experiment for profiling serum intact N-glycopeptides. We present a comprehensive data analysis workflow using GproDIA, including glycopeptide spectral library building, chromatographic feature extraction from the DIA data, and feature scoring with appropriate statistical control of error rates. We anticipate that this method could provide a powerful tool to explore the serum glycoproteome.
Protein phosphorylation is a post-translational modification crucial for many cellular processes and protein functions. Accurate identification and quantification of protein phosphosites at the proteome-wide level are challenging, not least because efficient tools for protein phosphosite false localization rate (FLR) control are lacking. Here, we propose DeepFLR, a deep learning-based framework for controlling the FLR in phosphoproteomics. DeepFLR includes a phosphopeptide tandem mass spectrum (MS/MS) prediction module based on deep learning and an FLR assessment module based on a target-decoy approach. DeepFLR improves the accuracy of phosphopeptide MS/MS prediction compared to existing tools. Furthermore, DeepFLR estimates FLR accurately for both synthetic and biological datasets, and localizes more phosphosites than probability-based methods. DeepFLR is compatible with data from different organisms, instruments types, and both data-dependent and data-independent acquisition approaches, thus enabling FLR estimation for a broad range of phosphoproteomics experiments.
Protein post-translational modifications (PTMs) increase the functional diversity of the cellular proteome. Accurate and high throughput identification and quantification of protein PTMs is a key task in proteomics research. Recent advancements in data-independent acquisition (DIA) mass spectrometry (MS) technology have achieved deep coverage and accurate quantification of proteins and PTMs. This review provides an overview of DIA data processing methods that cover three aspects of PTMs analysis, that is, detection of PTMs, site localization, and characterization of complex modification moieties, such as glycosylation. In addition, a survey of deep learning methods that boost DIA-based PTMs analysis is presented, including in silico spectral library generation, as well as feature scoring and error rate control. The limitations and future directions of DIA methods for PTMs analysis are also discussed. Novel data analysis methods will take advantage of advanced MS instrumentation techniques to empower DIA MS for in-depth and accurate PTMs measurements.