Abstract Accurate recognition and deidentification of sensitive health information (SHI) in spoken dialogues requires multimodal algorithms that can understand medical language and contextual nuance. However, the recognition and deidentification risks expose sensitive health information (SHI). Additionally, the variability and complexity of medical terminology, along with the inherent biases in medical datasets, further complicate this task. This study introduces the SREDH/AI-Cup 2025 Medical Speech Sensitive Information Recognition Challenge, which focuses on two tasks: Task-1: Speech transcription systems must accurately transcribe speech into text; and Task-2: Medical speech de-identification to detect and appropriately classify mentions of SHI. The competition attracted 246 teams; top-performing systems achieved a mixed error rate (MER) of 0.1147 and a macro F1-score of 0.7103, with average MER and macro F1-score of 0.3539 and 0.2696, respectively. Results were presented at the IW-DMRN workshop in 2025. Notably, the results reveal that LLMs were prevalent across both tasks: 97.5% of teams adopted LLMs for Task 1 and 100% for Task 2. Highlighting their growing role in healthcare. Furthermore, we finetuned six models, demonstrating strong precision (∼0.885–0.889) with slightly lower recall (∼0.830–0.847), resulting in F1-scores of 0.857–0.867.
ABSTRACT Objectives Due to privacy constraints, the sensitivity of medical speech, and the complexity of speech-level annotation, publicly available datasets for clinical speech de-identification remain scarce. To address this gap, we constructed the SREDH-AI Cup Sensitive Health Information (SHI) speech corpus, a time-aligned clinical speech dataset annotated for 26 SHI categories. Methods We compiled the corpus by integrating two English medical speech resources with Mandarin Chinese utterances collected from Taiwanese television dramas. In addition to using an existing open dataset for automated medical transcription, we used the OpenDeID v2 corpus into spoken-style clinical scripts and generated corresponding audio records with 25 voice contributors. All audio data were manually annotated with millisecond-level, time-aligned SHI segments. The annotation schema covered 26 SHI subcategories. Results The final corpus comprises approximately 20 hours of annotated audio and supports evaluation of both automatic speech recognition (ASR) and SHI recognition. It is divided into training (10 hours, 1,539 files), validation (5 hours, 775 files), and test (5 hours, 710 files) subsets. The corpus contains 7,830 SHI entities. Its language distribution reflects the composition of the selected source materials, with 19.36 hours of English and 0.89 hours of Chinese speech. Inter-annotator agreement, measured using Fleiss’s kappa, reached 0.907. In the SREDH/AICUP competition, the top-performing team achieved a mixed error rate of 0.1147 for ASR and a macro F1-score of 0.7103 for SHI recognition. Discussion The corpus poses several important challenges for clinical speech de-identification. In particular, SHI categories are distributed in a long-tailed manner, and the amount of English and Mandarin Chinese data is highly imbalanced. The Mandarin Chinese subset should be understood as a supplementary multilingual extension rather than as a parallel monolingual benchmark. Our experience constructing this subset highlights both the scarcity of usable Chinese clinical speech resources and the practical difficulty of producing time-aligned SHI annotations for low-resource clinical speech settings. Conclusion The SREDH-AI Cup SHI speech corpus provides a time-aligned speech dataset that approximates dialogue-based clinical scenarios. It supports research on automated medical speech de-identification, ASR, and SHI recognition, while also offering an initial reference resource for future work on English and low-resource clinical speech privacy protection. Key message What is already known on this topic - There is a scarcity of publicly available clinical speech datasets containing time-aligned sensitive health information (SHI) annotations for clinical speech de-identification research. What this study adds - The SREDH-AI Cup SHI speech corpus introduces a clinically-oriented speech dataset with millisecond-level, time-aligned SHI annotations across 26 categories. - The corpus integrates structured annotation protocols and standardized data processing to support reproducible benchmarking of speech-based de-identification models. How this study might affect research, practice or policy - The availability of time-aligned SHI annotations may facilitate research on real-time or streaming de-identification systems beyond conventional transcription-focused approaches. - The dataset may support the development of English and low-resource bilingual conditions privacy-preserving technologies in clinical speech environments.
BACKGROUND:Peptides have emerged as promising therapeutic agents for drug development against cancer, immune disorders, hypertension, and microbial infections. Peptide drugs have the advantage of high selectivity, low production cost, and fewer side effects compared to traditional small molecule-based drugs. However, one main challenge that hinders the adoption of peptide therapeutics is that some peptides are prone to be hemolytic, leading to the disruption of erythrocyte membranes and decreasing the life span of red blood cells. A computational model for hemolytic peptide identification would be a valuable tool for peptide drug discovery. RESULTS:In this study, we present HEPAD, a machine learning predictor to identify hemolytic peptides based on adaptive feature engineering and diverse sequence descriptors. Sequence descriptors were applied for feature encoding, generating a feature vector of nearly 4000 numeric values for each peptide. Next, an adaptive feature engineering method was proposed to produce a customized feature subset for a given dataset. The four datasets considered in this study were associated with 250, 350, 90, and 130 selected features. Five machine learning methods of different rationale were employed to perform cross validation and independent tests. HEPAD yields Matthew's correlation coefficients (MCCs) of 0.973, 0.643, and 0.609, respectively, for three independent datasets. The improvements in MCC compared to existing approaches range from 1.9 to 13.3% for three independent tests. Moreover, data visualization reveals that the customized feature subsets can effectively separate hemolytic peptides from random peptides. CONCLUSIONS:HEPAD offers efficient identification of potential hemolytic peptides, thereby expediting experimental procedures in drug discovery. The source code, datasets, and machine learning models are available at https://github.com/csh07/HEPAD .
This research explores the utilization of artificial intelligence (AI) language generation models for the de-identification of medical case narratives, effectively anonymizing patient identifiers to prevent unauthorized disclosure of sensitive health information. This ensures that patient data remains unrecognizable even when accessed, offering two significant advantages: compliance with legal standards forbidding the indiscriminate release of medical records, and bolstered trust in healthcare providers which encourages patients to share necessary details for therapeutic and investigative purposes. Historically, de-identification processes necessitated the manual examination of records by specialists to identify and sanitize personal identifiers, it is a time-consuming task with potential for error. The digitalization of medical records now offers avenues for automated de-identification, simultaneously enhancing data accuracy and facilitating the ethical utilization of extensive datasets in healthcare research, public health strategy, and policymaking. The emergence of sophisticated generative AI technologies has augmented the capacity for comprehensive de-identification. Platforms such as ChatGPT and Bing have evolved to address complex human inquiries. Yet, employing these tools without proper safeguards potentially compromises de-identification objectives by risking data exposure. To advance the integrity of privacy measures and maintain data governance, this study fine-tunes a pre-established vast language AI model, Pythia, for localized training with state-of-the-art technology with rigorous privacy assurance.
Mass spectrometry‑based proteomics using isobaric labeling technology has become popular for proteomic quantitation. Existing approaches rely on the mechanism of target-decoy search and false discovery rate control to examine whether a peptide-spectrum match (PSM) is utilized for quantitation. However, some PSMs passing the examination may still exhibit high quantitation errors, which can deteriorate the overall quantitation accuracy. We present IQUP, a machine learning-based method to identify quantitatively unreliable PSMs, termed QUPs. PSMs were characterized by 16 spectral and distance-based features for machine learning. Independent test results reveal that the best-performing models for the three datasets achieve accuracies of 0.883-0.966, AUCs of 0.924-0.963, and MCCs of 0.596-0.691. Notably, the distributions of relative errors for QUPs and quantitatively reliable PSMs (QRPs) exhibit significant differences. By using only the predicted QRPs for peptide-level quantitation, the proportions of peptides with larger relative errors decrease significantly, with a range between 15.3 and 83.3% for the three datasets; in the meantime, the proportions of peptides with smaller relative errors increase by 3.1-25.5%. Our experimental results demonstrate that IQUP provides robust performance and strong generalizability across multiple datasets and has great potential in improving proteomic quantitation accuracy at PSM and peptide levels for isobaric labeling experiments.
Secondary use of electronic health record notes enhances clinical outcomes and personalized medicine, but risks sensitive health information (SHI) exposure. Inconsistent time formats hinder interpretation, necessitating deidentification and temporal normalization. The SREDH/AI CUP 2023 competition explored large language models (LLMs) for these tasks using 3,244 pathology reports with surrogated SHIs and normalized dates. The competition drew 291 teams; the top teams achieved macro-F1 scores >0.8. Results were presented at the IW-DMRN workshop in 2024. Notably, 77.2% used LLMs, highlighting their growing role in healthcare. This study compares competition results with in-context learning and fine-tuned LLMs. Findings show that fine-tuning, especially with lower-rank adaptation, boosts performance but plateaus or degrades in models over 6 B parameters due to overfitting. Our findings highlight the value of data augmentation, training strategies, and hybrid approaches. Effective LLM-based deidentification requires balancing performance with legal and ethical demands, ensuring privacy and interpretability in regulated healthcare settings.
Aging is a natural phenomenon characterized by the loss of normal morphology and physiological functioning of the body, causing wrinkles on the skin, loss of hair, and compromised immune systems. Peptide therapies have emerged as a promising approach in aging studies because of their excellent tolerability, low immunogenicity, and high specificity. Computational methods can significantly expedite wet lab-based anti-aging peptide discovery by predicting potential candidates with high specificity and efficacy. We propose AAGP, an anti-aging peptide predictor based on diverse physicochemical and compositional features. Two datasets were constructed, both shared anti-aging peptides as positives, with the first using antimicrobial peptides as negatives and the second using random peptides as negatives. Peptides were encoded using 4,305 features, followed by adaptive feature selection with a heuristic algorithm on both datasets. Nine machine learning models were used for cross-validation and independent tests. AAGP achieves reasonably accurate prediction performance, with MCCs of 0.692 and 0.580 and AUCs of 0.963 and 0.808 on the two independent test datasets, respectively. Our feature importance analysis shows that physicochemical features are more crucial for the first dataset, whereas compositional features hold greater importance for the second. The source code of AAGP is available at https://github.com/saptawtf/AAGP .
Electronic Medical Records (EMR) implementation benefits the medical industry with streamlined data analysis, increased patient medication safety, reduced expenses for pathology report storage, and improved medical care efficiency. The EMR text notes hold a patient's clinical record, comprising notes written by the medical staff and further analyzed by the doctor following diagnosis. However, utilizing the EMR text note in its raw form can expose sensitive personal data belonging to patients and medical personnel. Therefore, safeguarding this private information is of utmost importance. Furthermore, the expression of time information in EMR text notes varies across institutions, which can significantly impact the accuracy and reliability of temporal information analysis. Therefore, normalizing temporal information is also a critical issue. The study presents a competition titled Privacy Protection and Standardization of Electronic Medical Record Competition that addresses recognizing Sensitive health information (SHI) recorded in EMR text notes and normalizing temporal information that poses a risk of identity theft. The competition released a corpus containing synthesized SHIs and normalized temporal information. The highest performance for the SHI recognition (subtask 1) and temporal information normalization (subtask 2) are micro-/macro-F of 0.949/0.912 and 0.844/0.869, respectively. Overall, the average micro/macro score for subtasks 1 and 2 were 0.666/0.496 and 0.6/0.394.
Angiogenesis is a key process for the proliferation and metastatic spread of cancer cells. Anti-angiogenic peptides (AAPs), with the capability of inhibiting angiogenesis, are promising candidates in cancer treatment. We propose AAPL, a sequence-based predictor to identify AAPs with machine learning models of improved prediction accuracy. Each peptide sequence was transformed to a vector of 4335 numeric values according to 58 different feature types, followed by a heuristic algorithm for feature selection. Next, the hyperparameters of six machine learning models were optimized with respect to the feature subset. We considered two datasets, one with entire peptide sequences and the other with 15 amino acids from peptide N-termini. AAPL achieved Matthew's correlation coefficients of 0.671 and 0.756 for independent tests based on the two datasets, respectively, outperforming existing predictors by a range of 5.3% to 24.6%. Further analyses show that AAPL yields higher prediction accuracy for peptides with more hydrophobic residues, and fewer hydrophilic and charged residues. The source code of AAPL is available at https://github.com/yunzheng2002/Anti-angiogenic .
BackgroundThe widespread use of electronic health records in the clinical and biomedical fields makes the removal of protected health information (PHI) essential to maintain privacy. However, a significant portion of information is recorded in unstructured textual forms, posing a challenge for deidentification. In multilingual countries, medical records could be written in a mixture of more than one language, referred to as code mixing. Most current clinical natural language processing techniques are designed for monolingual text, and there is a need to address the deidentification of code-mixed text. ObjectiveThe aim of this study was to investigate the effectiveness and underlying mechanism of fine-tuned pretrained language models (PLMs) in identifying PHI in the code-mixed context. Additionally, we aimed to evaluate the potential of prompting large language models (LLMs) for recognizing PHI in a zero-shot manner. MethodsWe compiled the first clinical code-mixed deidentification data set consisting of text written in Chinese and English. We explored the effectiveness of fine-tuned PLMs for recognizing PHI in code-mixed content, with a focus on whether PLMs exploit naming regularity and mention coverage to achieve superior performance, by probing the developed models’ outputs to examine their decision-making process. Furthermore, we investigated the potential of prompt-based in-context learning of LLMs for recognizing PHI in code-mixed text. ResultsThe developed methods were evaluated on a code-mixed deidentification corpus of 1700 discharge summaries. We observed that different PHI types had preferences in their occurrences within the different types of language-mixed sentences, and PLMs could effectively recognize PHI by exploiting the learned name regularity. However, the models may exhibit suboptimal results when regularity is weak or mentions contain unknown words that the representations cannot generate well. We also found that the availability of code-mixed training instances is essential for the model’s performance. Furthermore, the LLM-based deidentification method was a feasible and appealing approach that can be controlled and enhanced through natural language prompts. ConclusionsThe study contributes to understanding the underlying mechanism of PLMs in addressing the deidentification process in the code-mixed context and highlights the significance of incorporating code-mixed training instances into the model training phase. To support the advancement of research, we created a manipulated subset of the resynthesized data set available for research purposes. Based on the compiled data set, we found that the LLM-based deidentification method is a feasible approach, but carefully crafted prompts are essential to avoid unwanted output. However, the use of such methods in the hospital setting requires careful consideration of data security and privacy concerns. Further research could explore the augmentation of PLMs and LLMs with external knowledge to improve their strength in recognizing rare PHI.
Cancer immunotherapy enhances the body’s natural immune system to combat cancer, offering the advantage of lowered side effects compared to traditional treatments because of its high selectivity and efficacy. Utilizing computational methods to identify tumor T cell antigens (TTCAs) is valuable in unraveling the biological mechanisms and enhancing the effectiveness of immunotherapy. In this study, we present ENCAP, a predictor for TTCA based on ensemble classifiers and diverse sequence features. Sequences were encoded as a feature vector of 4349 entries based on 57 different feature types, followed by feature engineering and hyperparameter optimization for machine learning models, respectively. The selected feature subsets of ENCAP are primarily composed of physicochemical properties, with several features specifically related to hydrophobicity and amphiphilicity. Two publicly available datasets were used for performance evaluation. ENCAP yields an AUC (Area Under the ROC Curve) of 0.768 and an MCC (Matthew’s Correlation Coefficient) of 0.522 on the first independent test set. On the second test set, it achieves an AUC of 0.960 and an MCC of 0.789. Performance evaluations show that ENCAP generates 4.8% and 13.5% improvements in MCC over the state-of-the-art methods on two popular TTCA datasets, respectively. For the third test dataset of 71 experimentally validated TTCAs from the literature, ENCAP yields prediction accuracy of 0.873, achieving improvements ranging from 12% to 25.7% compared to three state-of-the-art methods. In general, the prediction accuracy is higher for sequences of fewer hydrophobic residues, and more hydrophilic and charged residues. The source code of ENCAP is freely available at https://github.com/YnnJ456/ENCAP.
Isobaric labeling relative quantitation is one of the dominating proteomic quantitation technologies. Traditional quantitation pipelines for isobaric-labeled mass spectrometry data are based on sequence database searching. In this study, we present a novel quantitation pipeline that integrates sequence database searching, spectral library searching, and a feature-based peptide-spectrum-match (PSM) filter using various spectral features for filtering. The combined database and spectral library searching results in larger quantitation coverage, and the filter removes PSMs with larger quantitation errors, retaining those with higher quantitation accuracy. Quantitation results show that the proposed pipeline can improve the overall quantitation accuracy at the PSM and protein levels. To our knowledge, this is the first study that utilizes spectral library searching to improve isobaric labeling-based quantitation. For users to conveniently perform the proposed pipeline, we have implemented the feature-based filter being executable on both Windows and Linux platforms; its executable files, user manual, and sample data sets are freely available at https://ms.iis.sinica.edu.tw/comics/Software_FPF.html. Furthermore, with the developed filter, the proposed pipeline is fully compatible with the Trans-Proteomic Pipeline.
Antimicrobial resistance is one of the most serious issue for human health. Compared to existing antibiotics, antimicrobial peptides have the advantage of efficient killing microbes and other pathogens without inducing drug resistance. Large-scale experimental methods to characterize AMPs require wet-lab resources and longer time. In silico prediction of AMP, on the other hand, is an attractive strategy to lower the cost and time in the discovery of new AMPs. In this study, we proposed a CatBoost model for AMP prediction. We included various features for numerical representation of peptides, and then employed a systematic approach to select 130 important features for our machine learning models. The CatBoost model achieves an accuracy, F1-score, MCC, and AUC of 0.758, 0.750, 0.518, and 0.831, respectively, for cross validation. For an independent test based on 188 peptide sequences, the proposed model achieves an accuracy, MCC, and AUC of 0.814, 0.632, and 0.884, respectively, all of which are the best compared to five state-of-art methods. Our model improves the MCC of five existing methods by 2.6% to 21.1%, and improves the AUC of them by 1.3% to 13.3%, respectively. The results demonstrate that our CatBoost model is capable of yielding reliable results, and can be of great help in discovering novel AMPs.
Identifying peptides and proteins from mass spectrometry (MS) data, spectral library searching has emerged as a complementary approach to the conventional database searching. However, for the spectrum-centric analysis of data-independent acquisition (DIA) data, spectral library searching has not been widely exploited because existing spectral library search tools are mainly designed and optimized for the analysis of data-dependent acquisition (DDA) data. We present Calibr, a spectral library search tool for spectrum-centric DIA data analysis. Calibr optimizes spectrum preprocessing for pseudo MS2 spectra, generating an 8.11% increase in spectrum–spectrum match (SSM) number and a 7.49% increase in peptide number over the traditional preprocessing approach. When searching against the DDA-based spectral library, Calibr improves SSM number by 17.6–26.65% and peptide number by 18.45–37.31% over two state-of-the-art tools on three different data sets. Searching against the public spectral library from MassIVE, Calibr improves state-of-the-art tools in SSM and peptide numbers by more than 31.49% and 25.24%, respectively, for two data sets. Our analyses indicate higher sensitivity of Calibr results from the use of various spectral similarity measures and statistical scores, coupled with machine learning-based statistical validation for FDR control. Calibr executable files including a graphical user-interface application are available at https://ms.iis.sinica.edu.tw/COmics/Software_CalibrWizard.html and https://sourceforge.net/projects/comics-calibr .
Mass spectrometry-based proteomics using isobaric labeling for multiplex quantitation has become a popular approach for proteomic studies. We present Multi-Q 2, an isobaric-labeling quantitation tool which can yield the largest quantitation coverage and improved quantitation accuracy compared to three state-of-the-art methods. Multi-Q 2 supports identification results from several popular proteomic data analysis platforms for quantitation, offering up to 12% improvement in quantitation coverage for accepting identification results from multiple search engines when compared with MaxQuant and PatternLab. It is equipped with various quantitation algorithms, including a ratio compression correction algorithm, and results in up to 336 algorithmic combinations. Systematic evaluation shows different algorithmic combinations have different strengths and are suitable for different situations. We also demonstrate that the flexibility of Multi-Q 2 in customizing algorithmic combination can lead to improved quantitation accuracy over existing tools. Moreover, the use of complementary algorithmic combinations can be an effective strategy to enhance sensitivity when searching for biomarkers from differentially expressed proteins in proteomic experiments. Multi-Q 2 provides interactive graphical interfaces to process quantitation and to display ratios at protein, peptide, and spectrum levels. It also supports a heatmap module, enabling users to cluster proteins based on their abundance ratios and to visualize the clustering results. Multi-Q 2 executable files, sample data sets, and user manual are freely available at http://ms.iis.sinica.edu.tw/COmics/Software_Multi-Q2.html .
Lung cancer in East Asia is characterized by a high percentage of never-smokers, early onset and predominant EGFR mutations. To illuminate the molecular phenotype of this demographically distinct disease, we performed a deep comprehensive proteogenomic study on a prospectively collected cohort in Taiwan, representing early stage, predominantly female, non-smoking lung adenocarcinoma. Integrated genomic, proteomic, and phosphoproteomic analysis delineated the demographically distinct molecular attributes and hallmarks of tumor progression. Mutational signature analysis revealed age- and gender-related mutagenesis mechanisms, characterized by high prevalence of APOBEC mutational signature in younger females and over-representation of environmental carcinogen-like mutational signatures in older females. A proteomics-informed classification distinguished the clinical characteristics of early stage patients with EGFR mutations. Furthermore, integrated protein network analysis revealed the cellular remodeling underpinning clinical trajectories and nominated candidate biomarkers for patient stratification and therapeutic intervention. This multi-omic molecular architecture may help develop strategies for management of early stage never-smoker lung adenocarcinoma.
N-linked glycosylation is one of the predominant post-translational modifications involved in a number of biological functions. Since experimental characterization of glycosites is challenging, glycosite prediction is crucial. Several predictors have been made available and report high performance. Most of them evaluate their performance at every asparagine in protein sequences, not confined to asparagine in the N-X-S/T sequon. In this paper, we present N-GlyDE, a two-stage prediction tool trained on rigorously-constructed non-redundant datasets to predict N-linked glycosites in the human proteome. The first stage uses a protein similarity voting algorithm trained on both glycoproteins and non-glycoproteins to predict a score for a protein to improve glycosite prediction. The second stage uses a support vector machine to predict N-linked glycosites by utilizing features of gapped dipeptides, pattern-based predicted surface accessibility, and predicted secondary structure. N-GlyDE’s final predictions are derived from a weight adjustment of the second-stage prediction results based on the first-stage prediction score. Evaluated on N-X-S/T sequons of an independent dataset comprised of 53 glycoproteins and 33 non-glycoproteins, N-GlyDE achieves an accuracy and MCC of 0.740 and 0.499, respectively, outperforming the compared tools. The N-GlyDE web server is available at http://bioapp.iis.sinica.edu.tw/N-GlyDE/ .
Protein and peptide identification and quantitation are essential tasks in proteomics research and involve a series of steps in analyzing mass spectrometry data. Trans-Proteomic Pipeline (TPP) provides a wide range of useful tools through its web interfaces for analyses such as sequence database search, statistical validation, and quantitation. To utilize the powerful functionality of TPP without the need for manual intervention to launch each step, we developed a software tool, called WinProphet, to create and automatically execute a pipeline for proteomic analyses. It seamlessly integrates with TPP and other external command-line programs, supporting various functionalities, including database search for protein and peptide identification, spectral library construction and search, data-independent acquisition (DIA) data analysis, and isobaric labeling and label-free quantitation. WinProphet is a standalone, installation-free tool with graphical interfaces for users to configure, manage, and automatically execute pipelines. The constructed pipelines can be exported as XML files with all of the parameter settings for reusability and portability. The executable files, user manual, and sample data sets of WinProphet are freely available at http://ms.iis.sinica.edu.tw/COmics/Software_WinProphet.html.
BACKGROUND:Tandem mass spectrometry allows biologists to identify and quantify protein samples in the form of digested peptide sequences. When performing peptide identification, spectral library search is more sensitive than traditional database search but is limited to peptides that have been previously identified. An accurate tandem mass spectrum prediction tool is thus crucial in expanding the peptide space and increasing the coverage of spectral library search.RESULTS:We propose MS2CNN, a non-linear regression model based on deep convolutional neural networks, a deep learning algorithm. The features for our model are amino acid composition, predicted secondary structure, and physical-chemical features such as isoelectric point, aromaticity, helicity, hydrophobicity, and basicity. MS2CNN was trained with five-fold cross validation on a three-way data split on the large-scale human HCD MS2 dataset of Orbitrap LC-MS/MS downloaded from the National Institute of Standards and Technology. It was then evaluated on a publicly available independent test dataset of human HeLa cell lysate from LC-MS experiments. On average, our model shows better cosine similarity and Pearson correlation coefficient (0.690 and 0.632) than MS2PIP (0.647 and 0.601) and is comparable with pDeep (0.692 and 0.642). Notably, for the more complex MS2 spectra of 3+ peptides, MS2PIP is significantly better than both MS2PIP and pDeep.CONCLUSIONS:We showed that MS2CNN outperforms MS2PIP for 2+ and 3+ peptides and pDeep for 3+ peptides. This implies that MS2CNN, the proposed convolutional neural network model, generates highly accurate MS2 spectra for LC-MS/MS experiments using Orbitrap machines, which can be of great help in protein and peptide identifications. The results suggest that incorporating more data for deep learning model may improve performance.
Ting-Yi Sung合作论文数Institute of Information Science
Academia Sinica, Taiwan14