Peptide therapeutics occupy a unique chemical space between small molecules and biologics, combining high target specificity with structural programmability and favorable safety profiles. Recent regulatory approvals and expanding clinical pipelines underscore the growing therapeutic and commercial relevance of peptide-based drugs. This review outlines chemical modification approaches and contemporary design strategies, and evaluates their impact on proteolytic stability, pharmacokinetics, membrane permeability, and target engagement. We then highlight recent advances in artificial intelligence (AI)-guided peptide drug design, including machine learning models, protein language models, and generative architectures that enable high-throughput activity prediction, property optimization, and de novo sequence generation. These approaches collectively accelerate the traditional discovery-design-validation cycle while reducing experimental attrition through data-driven, structure-informed modeling frameworks. Among these applications, AI also enables the rational design of cell-penetrating peptides (CPPs) to enhance intracellular delivery and biological activity. Building on these methodological advances, we further examine their application to peptide therapeutics, with particular emphasis on AI-based predictive models for CPPs as well as on therapeutic applications within the central nervous and pulmonary systems. We conclude by outlining future perspectives and emphasize that the systematic integration of AI-enabled sequence design with rational chemical engineering and advanced delivery technologies, supported by rigorous experimental validation, will be critical for developing robust and clinically durable peptide-based medicines.
Cell-penetrating peptides (CPPs) facilitate the intracellular delivery of therapeutic molecules. However, their accurate identification and design remain challenging because of the complexity of their structural and physicochemical characteristics. This study aimed to develop an interpretable predictive model that enables reliable CPP discovery and provides interpretable descriptors suggestive of the molecular properties underlying their activity. Peptide samples were represented using two-dimensional descriptors generated by Mordred. A novel two-stage feature selection method was introduced, combining a correlation-based filter with Shapley Additive exPlanations (SHAP). The model was trained on the CPP1708 dataset and built using an ensemble learning strategy integrating multiple machine learning algorithms. The ensemble framework, combining Extreme Gradient Boosting and Light Gradient Boosting Machine, identified five Mordred descriptors-5-ordered bonding information content (BIC5), Extended Topochemical Atom epsilon 5 (ETA_epsilon_5), averaged and centered Moreau-Broto autocorrelation of lag 0 weighted by ionization potential (AATSC0i), centered Moreau-Broto autocorrelation of lag 2 weighted by mass (ATSC2m), and first highest eigenvalue of Burden matrix weighted by gasteiger charge (BCUTc-1h)-as critical features for CPP prediction. The model achieved an accuracy of 82.0% and an area under the curve of 87.5% on the CPP1708 test set, outperforming existing predictors. This interpretable, high-performing prediction model supports the rational design of CPPs and advances peptide-based drug development. The SHAP-guided feature selection framework improves both efficiency and interpretability, with potential applications across diverse peptide classes. Furthermore, the identification of five mechanistic descriptors offers deeper insight into the structural, electronic, and physicochemical properties underpinning CPP activity.
G-protein-coupled receptors (GPCRs) represent a compelling intersection of structural biology, systems pharmacology, and clinical medicine. They function as essential molecular regulators of cardiovascular physiology, governing processes from the initiation of cardiac rhythm in the sinoatrial node to the regulation of vascular tone. Their widespread expression in cardiac and vascular tissues, combined with their ability to integrate diverse extracellular signals, positions GPCRs as key regulators of heart rate, contractility, vascular tone, inflammation, and metabolic homeostasis. Recent advances in GPCR structural biology, coupled with mechanistic insights into G-protein- and β-arrestin-mediated signaling mechanisms, underscore their central role in both cardiovascular health and disease. Notably, dysregulation of GPCR signaling has emerged as a unifying mechanism across major cardiovascular diseases (CVDs), contributing to pathological remodeling, impaired contractile function, and maladaptive vascular responses. This review uniquely integrates GPCR structural biology, signaling dynamics, and mechanosensitive and biased signaling mechanisms within the cardiovascular system. We discuss the contributions of class A, class B, and other GPCR families to cardiovascular physiology and pathology emphasizing their relevance to the development of targeted interventions for hypertension, heart failure, arrhythmias, atherosclerosis, and related CVDs. Together, these insights establish a contemporary framework for advancing precision GPCR-directed therapies in CVDs. By framing these developments within a mechanistic and translational context, the review offers a timely and clinically significant resource for both basic researchers and clinicians.
Interleukin-6 (IL-6) is a key immunomodulatory cytokine implicated in diverse physiological processes and pathological conditions, including autoimmune diseases, cancers, and cytokine storms. Immunogenic peptides capable of inducing IL-6 expression are key modulators of host immune responses and represent promising candidates for therapeutic design and epitope-based vaccine development. However, experimental identification of IL-6-inducing peptides remains laborious and unsuitable for large-scale screening. Although existing computational approaches show promise, many often struggle to capture both global contextual semantics and local motif-level features essential for peptide immunogenicity. To address these limitations, we present CONTRA-IL6, a novel deep learning framework that integrates Transformer fusion and convolutional localization modules with stacked pretrained protein language model embeddings to predict IL-6-inducing peptides. Comprehensive benchmarking on an independent dataset demonstrates that CONTRA-IL6 achieves superior predictive performance over six state-of-the-art predictors. Notably, it achieves the highest Matthews correlation coefficient (MCC, 0.504) and F1 (0.549) and improves over the best-performing existing method by 3.2% in MCC and 4.3% in F1, demonstrating balanced and robust performance. Feature space visualizations (uniform manifold approximation and projection, kernel density estimation) showed clear class separation, while 1D gradient-weighted class activation mapping++ highlighted strong attention to specific C-terminal regions. Crucially, we moved beyond these attribution methods by employing in silico mutagenesis, which causally confirmed the functional importance and physicochemical constraints. Ablation studies further confirmed the synergistic contribution of global and local modules to model performance. CONTRA-IL6 offers a robust, scalable, and interpretable solution for immunoinformatics research. The standalone package is freely available at https://pypi.org/project/contra-il6/ to facilitate broader community use.
Anticancer peptide (ACP) has emerged as potent therapeutic agents owing to its ability to selectively target cancer cells while minimising toxicity to healthy cells. However, the accurate computational prediction of ACP remains challenging because of the complex molecular mechanisms underlying cancer. In this study, we introduce EnsemPred-ACP, an innovative ensemble framework that combines machine learning (ML) and deep learning (DL) approaches to enhance ACP prediction. Our primary innovation is the introduction of binary profile features (BPF) to augment pre-trained protein embeddings, thereby capturing position-specific patterns crucial for ACP identification. The framework used a dual-pipeline architecture; ML models processed handcrafted sequence features and embeddings, whereas DL models handled BPF-enhanced embeddings. Upon evaluation with independent datasets, EnsemPred-ACP achieved an accuracy of 0.863, sensitivity of 0.897, and specificity of 0.830, notably outperforming existing methods. The model demonstrated a strong generalisation performance, achieving an area under the receiver operating characteristic curve of 0.93. Ablation studies on independent datasets further highlighted the substantial impact of BPF, enhancing the prediction accuracy by 2.5 % and 11.1 % when integrated with ESM2 and ProtT5 embeddings, respectively. These results demonstrate the effectiveness of our integrated approach in accurately identifying potential therapeutic peptides, thereby contributing to the advancement of peptide-based cancer therapeutics.
Cell-penetrating peptides (CPPs) have gained significant attention for biomedical applications, including drug delivery and therapeutic development, due to their ability to penetrate cell membranes. The accurate prediction of CPPs is critical for accelerating the design and development of novel peptide-based therapies. Approaches for CPP prediction primarily depend on either peptide characteristic-based conventional features or one or two protein language models (PLMs), but these methods often fail to fully leverage the potential of combining diverse features. To address this limitation, we propose CPPpred-En, a prediction model that evaluates multiple conventional and PLM-based features across various machine learning classifiers, selects high-performing feature-classifier combinations, and integrates them through ensemble learning. The CPPpred-En model, which was trained on both the CPP924 and MLCPP 2.0 datasets, outperformed existing state-of-the-art predictors, achieving an accuracy (Acc) of 97.27 % and a matthews correlation coefficient (MCC) of 0.964 on the CPP924 dataset and an Acc of 96.10 % and an MCC of 0.707 on the MLCPP 2.0 dataset. The ensemble-based strategy demonstrated robustness across different datasets, highlighting the strong ability of the model to generalise. The combination of conventional and PLM features in an ensemble framework is promising approach for improving peptide-based therapeutics. The CPPpred-En model is a highly accurate and reliable tool for the identification of CPPs and their application in drug delivery and targeted therapy.
Pancreatic α-amylase breaks down starch into isomaltose and maltose, which are further hydrolyzed by α-glucosidase in the intestine into monosaccharides, rapidly raising blood sugar levels and contributing to type 2 diabetes mellitus (T2DM). Synthetic inhibitors of carbohydrate-digesting enzymes are used to manage T2DM but may harm organ function over time. Bioactive peptides offer a safer alternative, avoiding such adverse effects. Computational methods for predicting antidiabetic peptides (ADPs) can significantly reduce the time and cost of experimental testing. While machine learning (ML) has been applied to identify ADPs, advancements in data analysis and algorithms continue to drive progress in the field. To address this, we developed AntiT2DMP-Pred, the first ML-based tool specifically designed for predicting type 2 antidiabetic peptides (T2ADPs). This tool employs a feature fusion strategy, combining ten highly discriminative feature descriptors chosen from a pool of 32 descriptors and eight ML algorithms, tested across a range of baseline models. AntiT2DMP-Pred demonstrated excellent performance, surpassing both baseline and feature-optimized models, with an accuracy (ACC) and Matthews’ correlation coefficient (MCC) of 0.976 and 0.953 on the training dataset, and an ACC and MCC of 0.957 and 0.851 on the independent dataset. The web server (https://balalab-skku.org/AntiT2DMP-Pred) is freely accessible, enabling researchers worldwide to utilize it in their experimental workflows and contribute to the discovery and understanding of T2ADPs, ultimately supporting peptide-based therapeutic development for diabetes management.
The rationale for using ADMET prediction tools in the early drug discovery paradigm is to guide the design of new compounds with favorable ADMET properties and ultimately minimize the attrition rates of drug failures. Artificial intelligence (AI) in in silico ADMET modeling has gained momentum due to its high-throughput and low-cost attributes. In this study, we developed a machine learning model capable of predicting 11 ADMET properties of chemical compounds. Each model was constructed by combining one of 40 classification algorithms including random forest (RF), extreme gradient boosting (XGB), support vector machine (SVM), and gradient boosting (GB) with one of three predefined hyperparameter configurations. This process can be efficiently performed using automated machine learning (AutoML) methods, which automatically search for the best combination of model algorithms and optimized hyperparameters. We developed optimal predictive models for 11 different ADMET properties using the Hyperopt-sklearn AutoML method. All of the developed models depicted an area under the ROC curve (AUC) >0.8. Furthermore, our developed models outperformed most of the ADMET properties and showed comparable performance in other properties when evaluated on external data sets and compared with published predictive models. Our results support the applicability of AutoML in ADMET prediction and will be helpful for ADMET prediction in early-stage drug discovery.
BACKGROUND:Amyotrophic lateral sclerosis (ALS) is a serious neurodegenerative disorder affecting nerve cells in the brain and spinal cord that is caused by mutations in the superoxide dismutase 1 (SOD1) enzyme. ALS-related mutations cause misfolding, dimerisation instability, and increased formation of aggregates. The underlying allosteric mechanisms, however, remain obscure as far as details of their fundamental atomistic structure are concerned. Hence, this gap in knowledge limits the development of novel SOD1 inhibitors and the understanding of how disease-associated mutations in distal sites affect enzyme activity.METHODS:We combined microsecond-scale based unbiased molecular dynamics (MD) simulation with network analysis to elucidate the local and global conformational changes and allosteric communications in SOD1 Apo (unmetallated form), Holo, Apo_CallA (mutant and unmetallated form), and Holo_CallA (mutant form) systems. To identify hotspot residues involved in SOD1 signalling and allosteric communications, we performed network centrality, community network, and path analyses.RESULTS:Structural analyses showed that unmetallated SOD1 systems and cysteine mutations displayed large structural variations in the catalytic sites, affecting structural stability. Inter- and intra H-bond analyses identified several important residues crucial for maintaining interfacial stability, structural stability, and enzyme catalysis. Dynamic motion analysis demonstrated more balanced atomic displacement and highly correlated motions in the Holo system. The rationale for structural disparity observed in the disulfide bond formation and R143 configuration in Apo and Holo systems were elucidated using distance and dihedral probability distribution analyses.CONCLUSION:Our study highlights the efficiency of combining extensive MD simulations with network analyses to unravel the features of protein allostery.
Peptide hormones were first used in medicine in the early 20th century, with the pivotal event being the isolation and purification of insulin in 1921. These hormones are integral to a sophisticated system that emerged early in evolution to regulate growth, development, and homeostasis. They serve as targeted signaling molecules that transfer specific information between cells and organs, ensuring coordinated and precise physiological responses. While experimental methods for identifying peptide hormones present challenges such as low abundance, stability issues, and complexity, computational methods offer promising alternatives. Advances in machine learning and bioinformatics have facilitated the prediction of peptide hormones, further enhancing their therapeutic potential. In this study, we explored three different computational frameworks for peptide hormone identification and determined that the meta-approach was the most suitable. Firstly, we evaluated the discriminative power of 26 feature descriptors using a series of baseline models and identified seven feature descriptors with high predictive potential. Through a systematic approach, we then selected the top 20 performing baseline models and integrated their predicted probabilities to train a meta-model, leveraging the strengths of multiple prediction strategies. Our final light gradient boosting-based meta-model, mHPpred, significantly outperformed the existing method, HOPPred, on both benchmarking and independent datasets. Notably, mHPpred also demonstrated superior performance compared to the hybrid and integrative framework approaches employed in this study. This superiority demonstrates the effectiveness of our multi-view feature learning strategy in capturing discriminative features and providing a more accurate prediction model for peptide hormones. mHPpred is publicly accessible at: https://balalab-skku.org/mHPpred.
Allergy is a hypersensitive condition in which individuals develop objective symptoms when exposed to harmless substances at a dose that would cause no harm to a “normal” person. Most current computational methods for allergen identification rely on homology or conventional machine learning using limited set of feature descriptors or validation on specific datasets, making them inefficient and inaccurate. Here, we propose SEP-AlgPro for the accurate identification of allergen protein from sequence information. We analyzed 10 conventional protein-based features and 14 different features derived from protein language models to gauge their effectiveness in differentiating allergens from non-allergens using 15 different classifiers. However, the final optimized model employs top 10 feature descriptors with top seven machine learning classifiers. Results show that the features derived from protein language models exhibit superior discriminative capabilities compared to traditional feature sets. This enabled us to select the most discriminatory baseline models, whose predicted outputs were aggregated and used as input to a deep neural network for the final allergen prediction. Extensive case studies showed that SEP-AlgPro outperforms state-of-the-art predictors in accurately identifying allergens. A user-friendly web server was developed and made freely available at https://balalab-skku.org/SEP-AlgPro/, making it a powerful tool for identifying potential allergens.
RNA polymers undergo extensive modifications following transcription by enzymes called "RNA modification enzymes." More than 170 distinct post-transcriptional modifications have been identified and this number has been constantly increasing. Diverse types of RNA including ribosomal (rRNA), transfer (tRNA), messenger (mRNA), and long non-coding RNA undergo this type of post-transcriptional modifications. The enthusiasm for RNA modification research has been rekindled as more evidence demonstrates that it plays a significant role in the regulation of gene expression.1Brégeon D. Pecqueur L. Toubdji S. Sudol C. Lombard M. Fontecave M. de Crécy-Lagard V. Motorin Y. Helm M. Hamdane D. Dihydrouridine in the transcriptome: new life for this ancient RNA chemical modification.ACS Chem. Biol. 2022; 17: 1638-1657Crossref PubMed Scopus (4) Google Scholar One of the most abundant modified bases in tRNA, dihydrouridine (D), has just entered the world of mRNA modifications, demonstrating critical physiological functions in cell growth.2Finet O. Yague-Sanz C. Krüger L.K. Tran P. Migeot V. Louski M. Nevers A. Rougemaille M. Sun J. Ernst F.G.M. et al.Transcription-wide mapping of dihydrouridine reveals that mRNA dihydrouridylation is required for meiotic chromosome segregation.Mol. Cell. 2022; 82: 404-419.e9Abstract Full Text Full Text PDF PubMed Scopus (20) Google Scholar D is a modified uridine nucleotide that is catalyzed by D synthase (DUS) enzyme and the second most prevalent modification in tRNAs. Generally, epitranscriptome profiling approaches are expensive, time-consuming, and labor-intensive. However, computational methods could offer a rapid, efficient, and affordable alternative to conventional experimental approaches. Several computational methods have been developed to predict D sites, but these tools are all trained on tRNAs, and their generalization on mRNAs is obscure. In this issue, Wang et al.,3Wang Y. Wang X. Cui X. Meng J. Rong R. Self-attention enabled deep learning of dihydrouridine (D) modification on mRNAs unveiled a distinct sequence signature from tRNAs.Mol. Ther. Nucleic Acids. 2023; 31: 411-420Abstract Full Text Full Text PDF PubMed Scopus (2) Google Scholar develop the first computational tool, DPred, for predicting D modification sites on mRNAs in yeast using local self-attention and a convolutional framework. Their paper highlights the differences in the D site sequence motifs among mRNAs and tRNAs, suggesting putative variations observed in the formation mechanisms of D on diverse RNA types. The authors emphasize that mixed predictions based on tRNA and mRNA datasets containing D are ineffective, and that their predictions should be clearly differentiated. A major limitation of their study is that the predictions for multiple species were not conducted due to insufficient data. The authors constructed D-containing tRNA and mRNA sequences from the literature and RMBase 2.0 (Figure 1). Due to the absence of experimentally verified unmodified sequences, they randomly selected the positive D site transcripts and considered them as negative samples. A balanced dataset was constructed, where 80% of the samples were selected for developing the prediction model, while the remaining samples were selected to test the model transferability. Using the training dataset, they assessed four feature encodings, including one hot encoding (OH), nucleotide chemical properties (NCP), nucleotide density (ND), and electron-ion interaction potential (EIIP). Instead of individual encodings, they generated four hybrid feature sets, including OH_ND, OH_EIIP, NCP_ND, and NCP_EIIP. These were trained using a deep neural network that comprises an additive local self-attention and a convolutional neural network (CNN) layer. In contrast to global attention, local attention considers only a subset of states when computing attention weights.4Soydaner D. Attention mechanism in neural networks: where it comes and where it goes.Neural Comput. Appl. 2022; 34: 13371-13385Crossref Scopus (7) Google Scholar As a result of these constraints, it is easier to implement and train, particularly when dealing with smaller datasets. Among these features, NCP_ND achieved the highest performance with an area under the receiver operating curve of 0.917 and 0.903 during training and independent evaluations, respectively. In the same study, DPred outperformed four conventional machine learning classifiers, including random forest, logistic regression, extreme gradient boosting, and support vector machine. Furthermore, the authors demonstrated that the proposed DPred could accurately predict D-related tRNA and mRNA across a vast range of species, including humans and mice. The researchers investigated whether a D-related tRNA-trained model can be applied to mRNA datasets or vice versa. Considering the results, it appears that data-driven D-prediction tools based on mRNA and tRNA cannot be used to predict each other, but should be clearly distinguished due to highly distinct sequence signatures around D. Nevertheless, this also raises concerns regarding the limitations of sequence-based computational models.5Li Z. Gao E. Zhou J. Han W. Xu X. Gao X. Applications of deep learning in understanding gene regulation.Cell Rep. Methods. 2023; 3: 100384Abstract Full Text Full Text PDF PubMed Scopus (3) Google Scholar The data-driven model generally needs to be retrained using experimental profiles based on the new condition type, which may not always be available. To improve the generalization ability of the model and further increase our understanding of D sites, other features, such as secondary structures and cell types, could be incorporated to the predictive model in the future. Another major limitation of their study to be addressed is the lack of multi-species prediction due to the limitation of high-quality D epitranscriptome data. Overall, their study provides putative insights into the study of D in the transcriptome from two different perspectives: First, CNN and local self-attention algorithms were utilized to provide the first data-driven predictive model that can predict D modification on mRNAs. Next, it revealed that mRNAs and tRNAs have distinct nucleotide preferences around D sites, implying that their mechanisms of formation and functions may differ. In the coming years, the pace of diverse RNA epigenetic modification site characterization by experimental approaches is likely to increase exponentially. Consequently, sequence-based computational approaches, as presented by Wang et al., will be crucial for comprehending biological functions. Their article has an intriguing finding that the sequence signature around D sites on mRNAs differs significantly from those on tRNAs. However, future studies are essential to examine the detailed mechanisms behind the formation of modified mRNAs and their biological implications. In addition, researchers should not overlook rRNAs, as DUS enzymes do not appear to be involved in D synthesis, suggesting that another enzyme system might be responsible for its biosynthesis.1Brégeon D. Pecqueur L. Toubdji S. Sudol C. Lombard M. Fontecave M. de Crécy-Lagard V. Motorin Y. Helm M. Hamdane D. Dihydrouridine in the transcriptome: new life for this ancient RNA chemical modification.ACS Chem. Biol. 2022; 17: 1638-1657Crossref PubMed Scopus (4) Google Scholar The authors declare no competing interests.
Diabetes mellitus has become a major public health concern associated with high mortality and reduced life expectancy and can cause blindness, heart attacks, kidney failure, lower limb amputations, and strokes. A new generation of antidiabetic peptides (ADPs) that act on β-cells or T-cells to regulate insulin production is being developed to alleviate the effects of diabetes. However, the lack of effective peptide-mining tools has hampered the discovery of these promising drugs. Hence, novel computational tools need to be developed urgently. In this study, we present ADP-Fuse, a novel two-layer prediction framework capable of accurately identifying ADPs or non-ADPs and categorizing them into type 1 and type 2 ADPs. First, we comprehensively evaluated 22 peptide sequence-derived features coupled with eight notable machine learning algorithms. Subsequently, the most suitable feature descriptors and classifiers for both layers were identified. The output of these single-feature models, embedded with multiview information, was trained with an appropriate classifier to provide the final prediction. Comprehensive cross-validation and independent tests substantiate that ADP-Fuse surpasses single-feature models and the feature fusion approach for the prediction of ADPs and their types. In addition, the SHapley Additive exPlanation method was used to elucidate the contributions of individual features to the prediction of ADPs and their types. Finally, a user-friendly web server for ADP-Fuse was developed and made publicly accessible (https://balalab-skku.org/ADP-Fuse), enabling the swift screening and identification of novel ADPs and their types. This framework is expected to contribute significantly to antidiabetic peptide identification.
Nanoparticles have garnered significant interest in neurological research in recent years owing to their efficient penetration of the blood–brain barrier (BBB). However, significant concerns are associated with their harmful effects, including those related to the immune response mediated by microglia, the resident immune cells in the brain, which are exposed to nanoparticles. We analysed the cytotoxic effects of silica-coated magnetic nanoparticles containing rhodamine B isothiocyanate dye [MNPs@SiO2(RITC)] in a BV2 microglial cell line using systems toxicological analysis. We performed the invasion assay and the exocytosis assay and transcriptomics, proteomics, metabolomics, and integrated triple-omics analysis, generating a single network using a machine learning algorithm. The results highlight alteration in the mechanisms of the nanotoxic effects of nanoparticles using integrated omics analysis.
Air pollution exerts several deleterious effects on the cardiovascular system, with cardiovascular disease (CVD) accounting for 80% of all premature deaths caused by air pollution. Short-term exposure to particulate matter 2.5 (PM2.5) leads to acute CVD-associated deaths and nonfatal events, whereas long-term exposure increases CVD-associated risk of death and reduces longevity. Here, we summarize published data illustrating how PM2.5 may impact the cardiovascular system to provide information on the mechanisms by which it may contribute to CVDs. We provide an overview of PM2.5, its associated health risks, global statistics, mechanistic underpinnings related to mitochondria, and hazardous biological effects. We elaborate on the association between PM2.5 exposure and CVD development and examine preventive PM2.5 exposure measures and future strategies for combating PM2.5-related adverse health effects. The insights gained can provide critical guidelines for preventing pollution-related CVDs through governmental, societal, and personal measures, thereby benefitting humanity and slowing climate change.
Acetylation on lysine residues is considered one of the most potent protein post-translational modifications, owing to its crucial role in cellular metabolism and regulatory processes. Recent advances in experimental techniques have unraveled several lysine acetylation substrates and sites. However, owing to its cost-ineffectiveness, cumbersome process, time-consumption, and labor-intensiveness, several efforts have been geared towards the development of computational tools. In particular, machine learning (ML)-based approaches hold great promise in the rapid discovery of lysine acetylation modification sites, which could be witnessed by the growing number of prediction tools. Recently, several ML methods have been developed for the prediction of lysine acetylation sites, owing to their time- and cost-effectiveness. In this review, we present a complete survey of the state-of-the-art ML predictors for lysine acetylation. We discuss a variety of key aspects for developing a successful predictor, including operating ML algorithms, feature selection methods, validation techniques, and software utility. Initially, we review lysine acetylation site databases, current ML approaches, working principles, and their performances. Lastly, we discuss the shortcomings and future directions of ML approaches in the prediction of lysine acetylation sites. This review may act as a useful guide for the experimentalists in choosing the right ML tool for their research. Moreover, it may help bioinformaticians in the development of more accurate and advanced MLbased predictors in protein research.
N-7-methylguanosine (m7G) is an essential, ubiquitous, and positively charged modification at the 50 cap of eukaryotic mRNA, modulating its export, translation, and splicing processes. Although several machine learning (ML)-based computational predictors for m7G have been developed, all utilized specific computational framework. This study is the first instance we explored four different computational frameworks and identified the best approach. Based on that we developed a novel predictor, THRONE (A three layer ensemble predictor for identifying human RNA N-7-methylguanosine sites) to accurately identify m7G sites from the human genome. THRONE employs a wide range of sequence-based features inputted to several ML classifiers and combines these models through ensemble learning. The three step ensemble learning is as follows: 54 baseline models were constructed in the first layer and the predicted probability of m7G was considered as a new feature vector for the sequential step. Subsequently, six meta-models were created using the new feature vector and their predicted probability was yet again considered as novel features. Finally, random forest was deemed as the best super classifier learner for the final prediction using a systematic approach incorporated with novel features. Interestingly, THRONE outperformed other existing methods in the prediction of m7G sites on both cross-validation analysis and independent evaluation. The proposed method is publicly accessible at: http://thegleelab.org/THRONE/ and expects to help the scientific community identify the putative m7G sites and formulate a novel testable biological hypothesis. (C) 2022 Elsevier Ltd. All rights reserved.
More than 150 genes are involved in amyotrophic lateral sclerosis (ALS), with superoxide dismutase 1 (SOD1) being one of the most studied. Mutations in SOD1 gene, which encodes the enzyme SOD1 is the second most prevalent and studied cause of familial ALS. SOD1 is a ubiquitous, homodimeric metalloenzyme that forms a critical component of the cellular defense against reactive oxygen species. Several mutations in the SOD1 enzyme cause misfolding, dimerization instability, and increased aggregate formation in ALS. However, there is a lack of information on the dimerization of SOD1 monomers and the mechanistic underpinnings on how the pathogenic mutations disrupt the dimerization mechanism. Here, we presented microsecond-scale molecular dynamics (MD) simulations to unravel how interface-based mutations compromise SOD1 dimerization and provide mechanistic understanding into the corresponding process using WT and three interface-based mutant systems (A4V, T54R, and I113T). Structural stability analysis showed that the mutant systems displayed disparate variations in the catalytic sites which may directly alter the stability and activity of the SOD1 enzyme. Based on the dynamic network analysis and principal component analysis, it has been identified that the mutations weakened the correlated motions along the dimer interface and altered the protein conformational behavior, thus weakening the stability of dimer formation. Moreover, the simulation results identified crucial residues such as G51, D52, G114, I151, and Q153 in establishing the dimerization interaction network, which were weakened or absent in the presence of interfacial mutants. Surface potential analysis on mutant systems also displayed changes in the dimerization potential, thus showing the unfavorable dimer formation. Furthermore, network analysis identified the hotspot residues necessary for SOD1 signal transduction which were surprisingly found in the catalytic sites rather than the anticipated dimerization interface.
Protein post-translational modification (PTM) is an important regulatory mechanism that plays a key role in both normal and disease states. Acetylation on lysine residues is one of the most potent PTMs owing to its critical role in cellular metabolism and regulatory processes. Identifying protein lysine acetylation (Kace) sites is a challenging task in bioinformatics. To date, several machine learning-based methods for the in silico identification of Kace sites have been developed. Of those, a few are prokaryotic species-specific. Despite their attractive advantages and performances, these methods have certain limitations. Therefore, this study proposes a novel predictor STALLION (STacking-based Predictor for ProkAryotic Lysine AcetyLatION), containing six prokaryotic species-specific models to identify Kace sites accurately. To extract crucial patterns around Kace sites, we employed 11 different encodings representing three different characteristics. Subsequently, a systematic and rigorous feature selection approach was employed to identify the optimal feature set independently for five tree-based ensemble algorithms and built their respective baseline model for each species. Finally, the predicted values from baseline models were utilized and trained with an appropriate classifier using the stacking strategy to develop STALLION. Comparative benchmarking experiments showed that STALLION significantly outperformed existing predictor on independent tests. To expedite direct accessibility to the STALLION models, a user-friendly online predictor was implemented, which is available at: http://thegleelab.org/STALLION.
Background Nanoparticles have been utilized in brain research and therapeutics, including imaging, diagnosis, and drug delivery, owing to their versatile properties compared to bulk materials. However, exposure to nanoparticles leads to their accumulation in the brain, but drug development to counteract this nanotoxicity remains challenging. To date, concerns have risen about the potential toxicity to the brain associated with nanoparticles exposure via penetration of the brain blood barrier to address this issue. Methods Here the effect of silica-coated-magnetic nanoparticles containing the rhodamine B isothiocyanate dye [MNPs@SiO 2 (RITC)] were assessed on microglia through toxicological investigation, including biological analysis and integration of transcriptomics, proteomics, and metabolomics. MNPs@SiO 2 (RITC)-induced biological changes, such as morphology, generation of reactive oxygen species, intracellular accumulation of MNPs@SiO 2 (RITC) using transmission electron microscopy, and glucose uptake efficiency, were analyzed in BV2 murine microglial cells. Each omics data was collected via RNA-sequencing-based transcriptome analysis, liquid chromatography-tandem mass spectrometry-based proteome analysis, and gas chromatography- tandem mass spectrometry-based metabolome analysis. The three omics datasets were integrated and generated as a single network using a machine learning algorithm. Nineteen compounds were screened and predicted their effects on nanotoxicity within the triple-omics network. Results Intracellular reactive oxygen species production, an inflammatory response, and morphological activation of cells were greater, but glucose uptake was lower in MNPs@SiO 2 (RITC)-treated BV2 microglia and primary rat microglia in a dose-dependent manner. Expression of 121 genes (from 41,214 identified genes), and levels of 45 proteins (from 5918 identified proteins) and 17 metabolites (from 47 identified metabolites) related to the above phenomena changed in MNPs@SiO 2 (RITC)-treated microglia. A combination of glutathione and citrate attenuated nanotoxicity induced by MNPs@SiO 2 (RITC) and ten other nanoparticles in vitro and in the murine brain, protecting mostly the hippocampus and thalamus. Conclusions Combination of glutathione and citrate can be one of the candidates for nanotoxicity alleviating drug against MNPs@SiO 2 (RITC) induced detrimental effect, including elevation of intracellular reactive oxygen species level, activation of microglia, and reduction in glucose uptake efficiency. In addition, our findings indicate that an integrated triple omics approach provides useful and sensitive toxicological assessment for nanoparticles and screening of drug for nanotoxicity. Graphical Abstract