Hepatocellular carcinoma (HCC) progression involves disruption of oncogenic and tumor-suppressive signaling networks. DEPDC7 (DEP domain-containing protein 7), a liver-specific gene associated with intercellular communication, is highly expressed in normal hepatocytes but markedly downregulated in HCC. Here, we investigated the tumor-suppressive mechanisms of DEPDC7 using Huh-7 cells. Structural analysis revealed conserved DEP and RhoGAP domains, with multiple predicted post-translational modification sites suggesting regulatory potential. DEPDC7 expression was significantly reduced in HCC cells and localized to both cytoplasm and nucleus. Functionally, DEPDC7 overexpression inhibited cell proliferation and migration. RNA-seq analysis identified the JAK1/STAT3 pathway was the most suppressed upon DEPDC7 overexpression, with downregulation of JAK1 and STAT3. Molecular docking and co-immunoprecipitation confirmed direct interaction between DEPDC7 and the JAK1 kinase domain, indicating regulation through physical binding. Moreover, DEPDC7 overexpression suppressed epithelial-mesenchymal transition (EMT), increasing E-cadherin while reducing N-cadherin and vimentin. Morphological changes observed by scanning electron microscopy supported reduced migratory capacity. Collectively, DEPDC7 exerts tumor-suppressive effects by (1) promoting cell cycle arrest and apoptosis, (2) inhibiting JAK1/STAT3 signaling, and (3) attenuating EMT. These findings provide mechanistic evidence that DEPDC7 functions as a tumor suppressor in HCC, highlighting its potential as a therapeutic target.
Biomolecule sensing for recognition is exhibited as the fundamental upstream step concerning target identification during the metabolism of individual life. Nevertheless, it is always a complicated work that leverages both in vitro and in vivo experiments to discriminate the corresponding interaction, affinity, structure, activity, and toxicity concerning target biomolecules. Simultaneously, biological investigation with intelligent computing has extended to bio-sequence analysis and biomedical image processing, especially biomolecule identification in multi-view and multimodal. This review presents a panorama of contemporary development among biomolecular omics and computing biological sensing, machine learning scenarios, and heterogeneous information with multi-view, multi-modal, structured, and unstructured text and biomedical images. After being given the background, the concept and database of biomolecule interaction, affinity, and structure are introduced. Then, the machine learning paradigms in bioinformatics and biomedical engineering are demonstrated according to epigenetics-centered or pharmacogenomics. Next, the multi-view or multi- modal learning algorithms and optimization strategies with structured and unstructured data formats, including texts and biomedical images are listed in detail. By comparing and analyzing the state-of-the-art works, this study has summarized the advantages of existing methods in target biomolecule identification and the challenges. Finally, future developments are prospected, including the trend of research in robustness, data augmentation, generalized model delineated, and acceleration.
The complexes of long non-coding RNAs bound to proteins can be involved in regulating life activities at various stages of organisms. However, in the face of the growing number of lncRNAs and proteins, verifying LncRNA-Protein Interactions (LPI) based on traditional biological experiments is time-consuming and laborious. Therefore, with the improvement of computing power, predicting LPI has met new development opportunity. In virtue of the state-of-the-art works, a framework called LncRNA-Protein Interactions based on Kernel Combinations and Graph Convolutional Networks (LPI-KCGCN) has been proposed in this article. We first construct kernel matrices by taking advantage of extracting both the lncRNAs and protein concerning the sequence features, sequence similarity features, expression features, and gene ontology. Then reconstruct the existent kernel matrices as the input of the next step. Combined with known LPI interactions, the reconstructed similarity matrices, which can be used as features of the topology map of the LPI network, are exploited in extracting potential representations in the lncRNA and protein space using a two-layer Graph Convolutional Network. The predicted matrix can be finally obtained by training the network to produce scoring matrices w.r.t. lncRNAs and proteins. Different LPI-KCGCN variants are ensemble to derive the final prediction results and testify on balanced and unbalanced datasets. The 5-fold cross-validation shows that the optimal feature information combination on a dataset with 15.5% positive samples has an AUC value of 0.9714 and an AUPR value of 0.9216. On another highly unbalanced dataset with only 5% positive samples, LPI-KCGCN also has outperformed the state-of-the-art works, which achieved an AUC value of 0.9907 and an AUPR value of 0.9267.
Introduction Accurately filling out death certificates is essential for death surveillance. However, manually determining the underlying cause of death is often imprecise. In this study, we investigate the Wide and Deep framework as a method to improve the accuracy and reliability of inferring the underlying cause of death. Methods Death report data from national-level cause of death surveillance sites in Fujian Province from 2016 to 2022, involving 403,547 deaths, were analyzed. The Wide and Deep embedded with Convolutional Neural Networks (CNN) was developed. Model performance was assessed using weighted accuracy, weighted precision, weighted recall, and weighted area under the curve (AUC). A comparison was made with XGBoost, CNN, Gated Recurrent Unit (GRU), Transformer, and GRU with Attention. Results The Wide and Deep achieved strong performance metrics on the test set: precision of 95.75%, recall of 92.08%, F1 Score of 93.78%, and an AUC of 95.99%. The model also displayed specific F1 Scores for different cause-of-death chain lengths: 97.13% for single causes, 95.08% for double causes, 91.24% for triple causes, and 79.50% for quadruple causes. Conclusions The Wide and Deep significantly enhances the ability to determine the root causes of death, providing a valuable tool for improving cause-of-death surveillance quality. Integrating artificial intelligence (AI) in this field is anticipated to streamline death registration and reporting procedures, thereby boosting the precision of public health data.
Triple-negative breast cancer (TNBC) is known for its aggressive nature, lack of effective diagnostic tools and treatments, and generally poor prognosis. The objective of this study was to investigate metabolic changes in TNBC using metabolomics approaches and explore the underlying mechanisms through integrated analysis with transcriptomics. In this study, serum untargeted metabolic profiles were first examined between 18 TNBC patients and 21 healthy control (HC) subjects using liquid chromatography-mass spectrometry (LC-MS), identifying a total of 22 significantly differential metabolites (DMs). Subsequently, receiver operating characteristic analysis revealed that 7-methylguanine could serve as a potential biomarker for TNBC in both the discovery and validation sets. Additionally, transcriptomic datasets were retrieved from the GEO database to identify differentially expressed genes (DEGs) between TNBC and normal tissues. An integrative analysis of the DMs and DEGs was conducted, uncovering potential molecular mechanisms underlying TNBC. Notably, three pathways—tyrosine metabolism, phenylalanine metabolism, and glycolysis/gluconeogenesis—were enriched, providing insight into the energy metabolism disorders in TNBC. Within these pathways, two DMs (4-hydroxyphenylacetaldehyde and oxaloacetic acid) and six DEGs (MAOA, ADH1B, ADH1C, AOC3, TAT, and PCK1) were identified as key components. In summary, this study highlights metabolic biomarkers that could potentially be used for the diagnosis and screening of TNBC. The comprehensive analysis of metabolomics and transcriptomics data offers a validated and in-depth understanding of TNBC metabolism.
Objective: We aim to accurately distinguish ubiquitin-specific proteases (USPs) from other members within the deubiquitinating enzyme families based on protein sequences. Additionally, we seek to elucidate the specific regulatory mechanisms through which USP26 modulates Kruppel-like factor 6 (KLF6) and assess the subsequent effects of this regulation on both the proliferation and migration of cervical cancer cells.Methods: All the deubiquitinase (DUB) sequences were classified into USPs and non-USPs. Feature vectors, including 188D, n-gram, and 400D dimensions, were extracted from these sequences and subjected to binary classification via the Weka software. Next, thirty human USPs were also analyzed to identify conserved motifs and ascertained evolutionary relationships. Experimentally, more than 90 unique DUB-encoding plasmids were transfected into HeLa cell lines to assess alterations in KLF6 protein levels and to isolate a specific DUB involved in KLF6 regulation. Subsequent experiments utilized both wild-type (WT) USP26 overexpression and shRNAmediated USP26 knockdown to examine changes in KLF6 protein levels. The half-life experiment was performed to assess the influence of USP26 on KLF6 protein stability. Immunoprecipitation was applied to confirm the USP26-KLF6 interaction, and ubiquitination assays to explore the role of USP26 in KLF6 deubiquitination. Additional cellular assays were conducted to evaluate the effects of USP26 on HeLa cell proliferation and migration.Results: 1. Among the extracted feature vectors of 188D, 400D, and n-gram, all 12 classifiers demonstrated excellent performance. The RandomForest classifier demonstrated superior performance in this assessment. Phylogenetic analysis of 30 human USPs revealed the presence of nine unique motifs, comprising zinc finger and ubiquitin-specific protease domains. 2. Through a systematic screening of the deubiquitinase library, USP26 was identified as the sole DUB associated with KLF6. 3. USP26 positively regulated the protein level of KLF6, as evidenced by the decrease in KLF6 protein expression upon shUSP26 knockdown in both 293T and Hela cell lines. Additionally, half-life experiments demonstrated that USP26 prolonged the stability of KLF6. 4. Immunoprecipitation experiments revealed a strong interaction between USP26 and KLF6. Notably, the functional interaction domain was mapped to amino acids 285-913 of USP26, as opposed to the 1-295 region. 5. WT USP26 was found to attenuate the ubiquitination levels of KLF6. However, the mutant USP26 abrogated its deubiquitination activity. 6. Functional biological assays demonstrated that overexpression of USP26 inhibited both proliferation and migration of HeLa cells. Conversely, knockdown of USP26 was shown to promote these oncogenic properties.Conclusions: 1. At the protein sequence level, members of the USP family can be effectively differentiated from non-USP proteins. Furthermore, specific functional motifs have been identified within the sequences of human USPs. 2. The deubiquitinating enzyme USP26 has been shown to target KLF6 for deubiquitination, thereby modulating its stability. Importantly, USP26 plays a pivotal role in the modulation of proliferation and migration in cervical cancer cells.
BACKGROUND:Breast cancer (BC), the most common form of malignant cancer affecting women worldwide, was characterized by heterogeneous metabolic disorder and lack of effective biomarkers for diagnosis. The purpose of this study is to search for reliable metabolite biomarkers of BC as well as triple-negative breast cancer (TNBC) using serum metabolomics approach. METHODS:In this study, an untargeted metabolomics technique based on ultra-high performance liquid chromatography combined with mass spectrometry (UHPLC-MS) was utilized to investigate the differences in serum metabolic profile between the BC group (n = 53) and non-BC group (n = 57), as well as between TNBC patients (n = 23) and non-TNBC subjects (n = 30). The multivariate data analysis, determination of the fold change and the Mann-Whitney U test were used to screen out the differential metabolites. Additionally, machine learning methods including receiver operating curve analysis and logistic regression analysis were conducted to establish diagnostic biomarker panels. RESULTS:There were 36 metabolites found to be significantly different between BC and non-BC groups, and 12 metabolites discovered to be significantly different between TNBC and non-TNBC patients. Results also showed that four metabolites, including N-acetyl-D-tryptophan, 2-arachidonoylglycerol, pipecolic acid and oxoglutaric acid, were considered as vital biomarkers for the diagnosis of BC and non-BC with an area under the curve (AUC) of 0.995. Another two-metabolite panel of N-acetyl-D-tryptophan and 2-arachidonoylglycerol was discovered to discriminate TNBC from non-TNBC and produced an AUC of 0.965. CONCLUSION:This study demonstrated that serum metabolomics can be used to identify BC specifically and identified promising serum metabolic markers for TNBC diagnosis.
Mitochondrial ribosomal protein L27 (MRPL27) is a member of mitochondrial ribosomal proteins (MRPs). However, the biological function of MRPL27 in hepatocellular carcinoma (LIHC) is still unclear. Use UALCAN, TIMER, TISIDB, Kaplan Meier, and GEPIA database systems to analyze the expression, prognostic value, and relationship between MRPL27 and immune infiltration in LIHC. The expression and clinical significance of MRPL27 in LIHC patients were validated using tissue microarray. Conduct cell function experiments to detect the effects of overexpression and knockdown of MRPL27 on the proliferation, migration, and invasion of Huh7 cells. The tissue microarray results confirmed that MRPL27 expression is upregulated in LIHC, and high MRPL27 expression is associated with poorer prognosis. The expression of MRPL27 is significantly correlated with the infiltration levels of immune modulators, chemokines, and various immune cells. In addition, MRPL27 affects the proliferation, migration, and invasion of LIHC cells. These data indicate that MRPL27 is a new prognostic biomarker for LIHC and is associated with immune infiltration in LIHC.
With the increasing amount of recognized lncRNAs, people are paying much more attention than before mining their potential function of them, which performs biological functions by interacting with proteins. However, facing the huge amount of biological data, it is obvious that biological experiments only in vitro and in vivo are time-consuming and insufficient. To this end, this study proposed a framework for LncRNA-Protein Interactions prediction based on Multi-kernel fusion and Graph Auto-Encoders (LPI-MGAE). First, three feature kernels should be constructed in lncRNA and protein space, respectively. Secondly, three feature kernels are fused separately using the average weighted strategy. Then, the embedding of the feature kernels can be extracted with a graph auto-encoder consisting of two-layer graph convolutional networks. Finally, a regularized least squares classifier can be used to derive the final prediction results. 5-fold cross-validation shows that LPI-MGAE obtains successful outcomes on both Dataset1 and Dataset2, with an AUPR of 19.37
The Src Homology 2 (SH2) domain plays an important role in the signal transmission mechanism in organisms. It mediates the protein-protein interactions based on the combination between phosphotyrosine and motifs in SH2 domain. In this study, we designed a method to identify SH2 domain-containing proteins and non-SH2 domain-containing proteins through deep learning technology. Firstly, we collected SH2 and non-SH2 domain-containing protein sequences including multiple species. We built six deep learning models through DeepBIO after data preprocessing and compared their performance. Secondly, we selected the model with the strongest comprehensive ability to conduct training and test separately again, and analyze the results visually. It was found that 288-dimensional (288D) feature could effectively identify two types of proteins. Finally, motifs analysis discovered the specific motif YKIR and revealed its function in signal transduction. In summary, we successfully identified SH2 domain and non-SH2 domain proteins through deep learning method, and obtained 288D features that perform best. In addition, we found a new motif YKIR in SH2 domain, and analyzed its function which helps to further understand the signaling mechanisms within the organism.
Skin-lesions segmentation plays a prominent role in computer-aided diagnosis systems for skin cancer, especially the remarkable success of the convolutional neural network (CNN) approaches in skin-lesions segmentation. However, it faces intractable challenges such as variable shape and blurred skin lesions boundaries. To this end, past research has employed cutting-edge mechanisms, including diverse attention modules. Inspired by state-of-the-art works, this study proposed a Dual Encoder framework with a Text-Guided Attention Network (DETA-Net) which can accurately and efficiently segment various and blurred lesions. Firstly, we designed a multi-scale joint encoder that took the advantage of both the CNNs and Transformer to extract features under the blurred lesion background condition. In addition, we introduced text-guided attention to propel classification in the manner of text-based embedding in the DETA-Net so that the variation in the size and number of the lesion can be efficiently accommodated. Experimental results demonstrated that DETA-Net provided better performance across multiple datasets compared with state-of-the-art on variable-sized skin lesion datasets in Skin-Cancer detection. We also evaluated the effectiveness of DETA-Net through extensive ablation studies on three different datasets, including ISIC 2016, ISIC 2018, and PH2 datasets. The baseline achieved 0.8838 Dice on ISIC 2016, 0.8864 Dice on ISIC 2018, and 0.8695 Dice on PH2.
Background: RNA Secondary Structure (RSS) has drawn growing concern, both for their pivotal roles in RNA tertiary structures prediction and critical effect in penetrating the mechanism of functional non-coding RNA. Computational techniques that can reduce the in vitro and in vivo experimental costs have become popular in RSS prediction. However, as an NP-hard problem, there is room for improvement that the validity of the prediction RSS with pseudoknots in traditional machine learning predictors.Results: In this essay, by integrating the bidirectional GRU (Gated Recurrent Unit) with the attention, we propose a multilayered neural network called BAT-Net to predict RSS. Different from the state-of-the-art works, BAT-Net can not only make full use of the information about the direct predecessor and direct successor of the predicted base in the RNA sequence but also dynamically adjust the corresponding loss function. The experimental results on five representative datasets extracted from the RNA STRAND database show that the sensitivity, precision, accuracy, and MCC (Matthews Correlation Coefficient) of the BAT-Net have improved by 8.52%, 8.28%, 5.66% and 9.82%, respectively, compared with the benchmark approaches on the best averages.Conclusions: BAT-Net can provide users with more credible RSS results since it has further utilized the source information of the dataset. Comparative results show that the proposed BAT-Net is superior to the other existing methods on the relevant indicators.
Background: The identification of DNA binding proteins (DBP) is an important research field. Experiment-based methods are time-consuming and labor-intensive for detecting DBP. Objective: To solve the problem of large-scale DBP identification, some machine learning methods are proposed. However, these methods have insufficient predictive accuracy. Our aim is to develop a sequence- based machine learning model to predict DBP. Methods: In our study, we extracted six types of features (including NMBAC, GE, MCD, PSSM-AB, PSSM-DWT, and PsePSSM) from protein sequences. We used Multiple Kernel Learning based on Hilbert- Schmidt Independence Criterion (MKL-HSIC) to estimate the optimal kernel. Then, we constructed a hypergraph model to describe the relationship between labeled and unlabeled samples. Finally, Laplacian Support Vector Machines (LapSVM) is employed to train the predictive model. Our method is tested on PDB186, PDB1075, PDB2272 and PDB14189 data sets. Result: Compared with other methods, our model achieved best results on benchmark data sets. Conclusion: The accuracy of 87.1% and 74.2% are achieved on PDB186 (Independent test of PDB1075) and PDB2272 (Independent test of PDB14189), respectively.
Drawing support from an effective Medical Image Segmentation (MIS) is conducive to a substantial diagnostic basis for the physicians to identify the focus lesion in the patient body and give the subsequent clinical assessment of the patient status. Although various works have tried the challenging quantitative analysis problem, it is still difficult to conduct precise automatic segmentation, especially the soft tissue organs. In this decade, with the increased amount of available datasets, deep learning-based networks have achieved remarkable performance in image processing. Inspired by the state-of-the-art deep learning works, in this paper, we propose an end-to-end multi-layer network named RCGA-Net. It consists of an encoder-decoder backbone that integrates a coordinate attention mechanism based on space and channel and a global context extraction module to highlight more valuable information. To evaluate the performance of RCGA-Net, we apply it to different kinds of clinical and experimental MIS tasks to testify its generalization ability. Extensive experiments represent that our schema has taken the outperform or compatible results among the comparison methods group. Specifically, the numeric result of RCGA-Net on the pulmonary dataset has achieved a 99.12% optimum F1-score.
To distinguish Methicillin-Resistant Staphylococcus aureus (MRSA) from Methicillin-Sensitive Staphylococcus aureus (MSSA) in the protein sequences level, test the susceptibility to antibiotic of all Staphylococcus aureus isolates from Quanzhou hospitals, define the virulence factor and molecular characteristics of the MRSA isolates. MRSA and MSSA Pfam protein sequences were used to extract feature vectors of 188D, n-gram and 400D. Weka software was applied to classify the two Staphylococcus aureus and performance effect was evaluated. Antibiotic susceptibility testing of the 81 Staphylococcus aureus was performed by the Mérieux Microbial Analysis Instrument. The 65 MRSA isolates were characterized by Panton-Valentine leukocidin ( PVL ), X polymorphic region of Protein A ( spa ), multilocus sequence typing test ( MLST ), staphylococcus chromosomal cassette mec ( SCC mec) typing. After comparing the results of Weka six classifiers, the highest correctly classified rates were 91.94, 70.16, and 62.90% from 188D, n-gram and 400D, respectively. Antimicrobial susceptibility test of the 81 Staphylococcus aureus: Penicillin-resistant rate was 100%. No resistance to teicoplanin, linezolid, and vancomycin. The resistance rate of the MRSA isolates to clindamycin, erythromycin and tetracycline was higher than that of the MSSAs. Among the 65 MRSA isolates, the positive rate of PVL gene was 47.7% (31/65). Seventeen sequence types ( STs ) were identified among the 65 isolates, and ST59 was the most prevalent. SCCmec type III and IV were observed at 24.6 and 72.3%, respectively. Two isolates did not be typed. Twenty-one spa types were identified, spa t 437 (34/65, 52.3%) was the most predominant type. MRSA major clone type of molecular typing was CC59-ST59-spa t437-IV (28/65, 43.1%). Overall, 188D feature vectors can be applied to successfully distinguish MRSA from MSSA. In Quanzhou, the detection rate of PVL virulence factor was high, suggesting a high pathogenic risk of MRSA infection. The cross-infection of CA-MRSA and HA-MRSA was presented, the molecular characteristics were increasingly blurred, HA-MRSA with typical CA-MRSA molecular characteristics has become an important cause of healthcare-related infections. CC59-ST59-spa t437-IV was the main clone type in Quanzhou, which was rare in other parts of mainland China.
深度学习教学模式的构建能有效促进本科课程教学水平的提升.本文以医学院校本科生生物化学与分子生物学课程为例,通过分析当前该本科课程教与学过程中存在问题的根源、探讨深度学习教学模式的特征及构建方案,从课程教学平台完善、资源库建设、教学手段及方法改进、效果评价等方面入手构建此模式,并在本科生课堂中进行了实践.结果 表明,实验班更能把握课程核心、创建知识关联和迁移、有效进行反思性学习,且期末成绩优于对照班.
目的 研究内皮型一氧化氮合酶(eNOS)基因敲除小鼠脑内细胞色素P450-1B1(CYP1B1)蛋白的表达变化及其意义.方法 按照体重将小鼠分为4组:野生型组(n=4)、eNOS基因敲除组(n=4)、实验组(2种小鼠各4只)和对照组(2种小鼠各4只).实验组腹腔注射2,3′,4,5′-tetramethoxystilbene(TMS,CYP1B1抑制剂)300μg·kg-1,对照组腹腔注射等剂量溶剂二甲基亚砜30μL,每天1次,连续1周.以免疫荧光染色和免疫印迹法测定小鼠的CYP1B1与半胱氨酸天冬氨酸蛋白酶-3(Caspase-3)p17表达.结果 野生型组和基因敲除组小鼠的前脑皮质区CYP1 B1免疫荧光阳性细胞数分别是(106±21)个/200×视野、(249±17)个/200×视野;这2组在纹状体区阳性细胞数分别是(85±16)个/200×、(211±23)个/200×视野,2组比较差异均有统计学意义(均P<0.001).野生型组和基因敲除组小鼠的脑CYP1 B1蛋白表达灰度值分别是0.45±0.04,1.15±0.15,2组比较差异均有统计学意义(均P<0.001).TMS干预后,实验组中野生型小鼠和eNOS基因敲除小鼠脑Caspase-3 p17蛋白灰度值分别是1.24±0.21,2.21±0.17,均显著高于对照组中的野生型小鼠和eNOS基因敲除小鼠,Caspase-3 p17蛋白灰度值分别是0.23±0.03,0.76±0.08,组间比较差异均有统计学意义(均P<0.001).结论 eNOS基因敲除小鼠脑内CYP1B1蛋白表达升高并且起到抗凋亡作用,但其作用机制目前尚不明确.
目的 探索将微信平台应用于护理学本科“生物化学”课程中的混合式教学方式的应用效果.方法 本文将手机媒体结合到传统课堂中进行混合式教学,从微信平台增加了课前预习引导和课后练习巩固环节,采用教师主导和学生主体的学习方式,初步应用于大班护理学本科“生物化学”教学中,并借助“互联网+”重新进行教学设计、实施和分析.结果 实验组学生期末成绩优于对照组,尤其是高分段的优良生占比显著增多.结论 这种混合式教学模式值得继续完善并推广应用于本课程的教学实践中.
MOTIVATION:Accurate identification of N4-methylcytosine (4mC) modifications in a genome wide can provide insights into their biological functions and mechanisms. Machine learning recently have become effective approaches for computational identification of 4mC sites in genome. Unfortunately, existing methods cannot achieve satisfactory performance, owing to the lack of effective DNA feature representations that are capable to capture the characteristics of 4mC modifications. RESULTS:In this work, we developed a new predictor named 4mcPred-IFL, aiming to identify 4mC sites. To represent and capture discriminative features, we proposed an iterative feature representation algorithm that enables to learn informative features from several sequential models in a supervised iterative mode. Our analysis results showed that the feature representations learnt by our algorithm can capture the discriminative distribution characteristics between 4mC sites and non-4mC sites, enlarging the decision margin between the positives and negatives in feature space. Additionally, by evaluating and comparing our predictor with the state-of-the-art predictors on benchmark datasets, we demonstrate that our predictor can identify 4mC sites more accurately. AVAILABILITY AND IMPLEMENTATION:The user-friendly webserver that implements the proposed 4mcPred-IFL is well established, and is freely accessible at http://server.malab.cn/4mcPred-IFL. SUPPLEMENTARY INFORMATION:Supplementary data are available at Bioinformatics online.
Objective: Small GTPase is an important molecular switch that plays an important role in numerous signaling transduction pathways, the aim is to explore its binary classification features with machine learning algorithms. Methods: The sequences including small GTPases and non small GTPases were clustered to remove similar entries, respectively. Then, they were divided into 10 datasets, each containing equal entries of small GTPases and non small GTPases. These datasets extracted three feature vectors that included188-dimensional(188D), 400D, and motif-based features (608D). The next step was classification based on easy-classify.py software in scikit-learn, which integrated 12 classifiers and finally discovered the conserved motifs by MEME suite. Results: The three best performed classifiers were logistic regression (LR), gradient boosting decision tree (GBDT), and bagging for 400D features, LibSVM, GBDT, and bagging for 188D features, and GBDT, bagging, and AdaBoost for 608D features, respectively. The top four classifiers were GBDT, bagging, LR, and AdaBoost according to commonly evaluated indices as a whole. GBDT obtained the highest area under the curve (AUC) value at 88.61%. The 400D features performed better than the 188D and 608D ones. Five conserved G-box motifs were discovered in the sequences of human small GTPases. Conclusion: This study provides the first description of GBDT algorithm performed best for small GTPases classification.