DNA N4-methylcytosine (4mC), a key epigenetic modification regulating DNA repair and replication, requires efficient computational detection methods due to experimental limitations. Although machine learning predictors have been proposed, their performance could be enhanced through systematic optimization of feature encoding schemes. Here, we propose EnDeep4mC, a dual-adaptive framework integrating species-specific modeling with ensemble deep learning architectures to systematically optimize feature encoding schemes. Evaluated across six species, EnDeep4mC demonstrates commendable prediction performance and significantly outperforms current state-of-the-art predictors. Cross-species validation confirms its robust transferability from animal to microbe groups. Evolutionary analysis further uncovers the functional differentiation of 4mC sequences in biological evolution: Prokaryotic 4mC relies on stable patterns, whereas eukaryotes achieve regulatory plasticity through dynamic sequence combinations, which provides experimental evidence for species-adaptive encoding strategies.
Bitter peptides are a practical barrier in food-grade protein hydrolysates, fermented products, and peptide-based supplements because they can compromise flavor before nutritional or functional value is realized. Sensory panels and mass-spectrometry-based identification remain reliable, but their throughput is limited for early screening of large peptide pools. Existing predictors usually emphasize either interpretable hand-crafted descriptors or deep sequence representations, whereas these two information sources may be complementary for food-oriented bitter peptide screening. Here, we propose iBitter-HF, a hybrid feature embedding method that integrates seven classes of hand-crafted descriptors with Unified Representation (UniRep) features. Light Gradient Boosting Machine (LGBM)-based feature-importance ranking was used to organize the candidate embeddings, and eXtreme Gradient Boosting (XGB) was used for classification of the selected feature subset. On the public BTP640 benchmark, the finalized 135-feature model achieved 96.9% accuracy on the independent test set. Literature-based comparison indicated competitive performance relative to eight reported bitter peptide predictors, and dimensionality reduction visualization suggested clearer local organization of bitter and non-bitter peptides after feature optimization. These results support iBitter-HF as a computational aid for sequence-level bitter peptide screening and debittering-oriented design of protein hydrolysates.
Anticancer peptides (ACPs) are a promising focus in clinical oncology due to their ability to inhibit tumor cell proliferation with minimal side effects. Nevertheless, large-scale, expeditious and efficacious identification of ACPs is hindered by the high cost and time demands of conventional wet-lab experiments. Therefore, we introduced a new method called iACP-SEI to identify ACPs using sequence evolution information. iACP-SEI method utilizes the ESM2 protein language model, based on Transformer architecture, to extract feature vectors that encapsulate evolutionary information from peptide sequences. These vectors underwent feature selection via the light gradient boosting machine and used in an ensemble learning approach. Using the AntiCP2.0 main and alternate datasets, iACP-SEI model achieved independent test accuracies of 77.78% and 94.82%, respectively. Furthermore, it outperformed current methods on an unbalanced dataset, achieving a cross-validation accuracy of 90.39%, demonstrating improved robustness in handling imbalanced class samples. Although iACP-SEI demonstrated higher predictive performance and robustness than other methods, some limitations of it are also discussed. Received: 18 November 2024 | Revised: 8 January 2025 | Accepted: 23 January 2025 Conflicts of Interest The authors declare that they have no conflicts of interest to this work. Data Availability Statement Data available on request from the corresponding author upon reasonable request. Author Contribution Statement Bowen Zheng: Validation, Formal analysis, Investigation, Resources, Writing – original draft, Writing – review & editing, Supervision, Project administration. Rujun Li: Investigation, Data curation, Visualization. Haotian Wang: Validation, Investigation, Visualization. Sheng Wang: Investigation. Shiyu Peng: Data curation. Mingxin Li: Formal analysis. Liangzhen Jiang: Writing – review & editing. Zhibin Lv: Conceptualization, Methodology, Software, Supervision, Project administration, Funding acquisition.
RNA methylation, particularly through m6A modification, represents a crucial epigenetic mechanism that governs gene expression and influences a range of biological functions. Accurate identification of methylation sites is crucial for understanding their biological functions. Traditional experimental methods, however, are often costly and can be influenced by experimental conditions, making machine learning, especially deep learning techniques, a vital tool for m6A site identification. Despite their utility, current machine learning models struggle with unbalanced datasets, a common issue in bioinformatics. This study addresses the RNA methylation site data imbalance problem from three key perspectives: feature encoding representation, deep learning models, and data resampling strategies. Using the K-mer one-hot encoding strategy, we effectively extracted RNA sequence features and developed classification prediction models utilizing long short-term memory networks (LSTM) and its variant, Multiplicative LSTM (mLSTM). We further enhanced model performance by ensemble and weighted strategy models. Additionally, we utilized the sequence generative adversarial network (SeqGAN) and the synthetic minority resampling technique (SMOTE) to construct balanced datasets for RNA methylation sites. The prediction results were rigorously analyzed using the Wilcoxon test and multivariate linear regression to explore the effects of different K-mer values, model architectures, and sampling methods on classification outcomes. The analysis underscored the significant impact of feature selection, model architecture, and sampling techniques in addressing data imbalance. Notably, the optimal prediction performance was achieved with a K value of 5 using the mLSTM-ensemble model. These findings not only offer new insights and methodologies for RNA methylation site identification but also provide valuable guidance for addressing similar challenges in bioinformatics.
Purpose:To retrospectively analyse the different imaging manifestations of acquired immunodeficiency syndrome-associated hepatic Kaposi's sarcoma (AIDS-HKS) on CT, MRI, and Ultrasound.Patients and Methods:Eight patients were enrolled in the study. Laboratory tests of liver function were performed. The CT, MRI, and Ultrasound manifestations were reviewed by two radiologists and two sonographers, respectively. The distribution and imaging signs of AIDS-HKS were evaluated.Results:AIDS-HKS patients commonly presented multiple lesions, mainly distributed around the portal vein on CT, MRI, and Ultrasound. AIDS-HKS presented as ring enhancement in the arterial phase on contrast-enhanced CT and MRI scanning, and nodules gradually strengthen in the portal venous phase and the delayed phase. AIDS-HKS presented as intrahepatic bile duct dilatation and bile duct wall thickening around the lesion. Five patients (62.5%, 5/8) were followed up. After chemotherapy, the lesions were completely relieved (60.0%), or decreased (40.0%).Conclusion:AIDS-HKS presented as multiple nodular lesions with different imaging features. The combination of different imaging methods was helpful for the imaging diagnosis of AIDS-HKS.
Objective To determine the high-efficiency ancillary features (AFs) screened from LR-3/4 lesions and the HCC/non-HCC group and the diagnostic performance of LR3/4 observations. Materials and methods We retrospectively analyzed a total of 460 patients (with 473 nodules) classified into LR-3-LR-5 categories, including 311 cases of hepatocellular carcinoma (HCC), 6 cases of non-HCC malignant tumors, and 156 cases of benign lesions. Two faculty abdominal radiologists with experience in hepatic imaging reviewed and recorded the major features (MFs) and AFs of the Liver Imaging Reporting and Data System (LI-RADS). The frequency of the features and diagnostic performance were calculated with a logistic regression model. After applying the above AFs to LR-3/LR-4 observations, the sensitivity and specificity for HCC were compared. Results The average age of all patients was 54.24 ± 11.32 years, and the biochemical indicators ALT ( P = 0.044), TBIL ( P = 0.000), PLT ( P = 0.004), AFP ( P = 0.000) and Child‒Pugh class were significantly higher in the HCC group. MFs, mild-moderate T2 hyperintensity, restricted diffusion and AFs favoring HCC in addition to nodule-in-nodule appearance were common in the HCC group and LR-5 category. AFs screened from the HCC/non-HCC group (AF-HCC) were mild–moderate T2 hyperintensity, restricted diffusion, TP hypointensity, marked T2 hyperintensity and HBP isointensity ( P = 0.005, < 0.001, = 0. 032, p < 0.001, = 0.013), and the AFs screened from LR-3/4 lesions (AF-LR) were restricted diffusion, mosaic architecture, fat in mass, marked T2 hyperintensity and HBP isointensity ( P < 0.001, = 0.020, = 0.036, < 0.001, = 0.016), which were not exactly the same. After applying AF-HCC and AF-LR to LR-3 and LR-4 observations in HCC group and Non-HCC group, After the above grades changed, the diagnostic sensitivity for HCC were 84.96% using AF-HCC and 85.71% using AF-LR, the specificity were 89.26% using AF-HCC and 90.60% using AF-LR, which made a significant difference ( P = 0.000). And the kappa value for the two methods of AF-HCC and AF–LR were 0.695, reaching a substantial agreement. Conclusion When adjusting for LR-3/LR-4 lesions, the screened AFs with high diagnostic ability can be used to optimize LI-RADS v 2018; among them, AF-LR is recommended for better diagnostic capabilities.
Neuropeptides are biomolecules with crucial physiological functions. Accurate identification of neuropeptides is essential for understanding nervous system regulatory mechanisms. However, traditional analysis methods are expensive and laborious, and the development of effective machine learning models continues to be a subject of current research. Hence, in this research, we constructed an SVM-based machine learning neuropeptide predictor, iNP_ESM, by integrating protein language models Evolutionary Scale Modeling (ESM) and Unified Representation (UniRep) for the first time. Our model utilized feature fusion and feature selection strategies to improve prediction accuracy during optimization. In addition, we validated the effectiveness of the optimization strategy with UMAP (Uniform Manifold Approximation and Projection) visualization. iNP_ESM outperforms existing models on a variety of machine learning evaluation metrics, with an accuracy of up to 0.937 in cross-validation and 0.928 in independent testing, demonstrating optimal neuropeptide recognition capabilities. We anticipate improved neuropeptide data in the future, and we believe that the iNP_ESM model will have broader applications in the research and clinical treatment of neurological diseases.
Thermophilic proteins, mesophiles proteins and psychrophilic proteins have wide industrial applications, as enzymes with different optimal temperatures are often needed for different purposes. Convenient methods are needed to determine the optimal temperatures for proteins; however, laboratory methods for this purpose are time-consuming and laborious, and existing machine learning methods can only perform binary classification of thermophilic and non-thermophilic proteins, or psychrophilic and non-psychrophilic proteins. Here, we developed a deep learning model, PSTP-BERT, based on protein sequences that can directly perform Three classes identification of thermophilic, mesophilic, and psychrophilic proteins. By comparing BERT-bfd with other deep learning models using five-fold cross-validation, we found that BERT-bfd-extracted features achieved the highest accuracy under six classifiers. Furthermore, to improve the model's accuracy, we used SMOTE (synthetic minority oversampling technique) to balance the dataset and light gradient-boosting machine to rank BERT-bfd-extracted features according to their weights. We obtained the best-performing model with five-fold cross-validation accuracy of 89.59 % and independent test accuracy of 85.42 %. The performance of the PSTP-BERT is significantly better than that of existing models in Three classes identification task. In order to compare with previous binary classification models, we used PSTP-BERT to perform binary classification tasks of thermophilic and non-thermophilic protein, and psychrophilic and non-psychrophilic protein on an independent test set. PSTP-BERT achieved the highest accuracy on both binary classification tasks, with an accuracy of 93.33 % for thermophilic protein binary classification and 88.33 % for psychrophilic protein binary classification. The accuracy of the independent test of the model can reach between 89.8 % and 92.9 % after training and optimization of the training set with different sequence similarities, and the prediction accuracy of the new data can exceed 97 %. For the convenience of future researchers to use and reference, we have uploaded source code of PSTP-BERT to GitHub.
Anti-coronavirus peptides (ACVPs) represent a relatively novel approach of inhibiting the adsorption and fusion of the virus with human cells. Several peptide-based inhibitors showed promise as potential therapeutic drug candidates. However, identifying such peptides in laboratory experiments is both costly and time consuming. Therefore, there is growing interest in using computational methods to predict ACVPs. Here, we describe a model for the prediction of ACVPs that is based on the combination of feature engineering (FE) optimization and deep representation learning. FEOpti-ACVP was pre-trained using two feature extraction frameworks. At the next step, several machine learning approaches were tested in to construct the final algorithm. The final version of FEOpti-ACVP outperformed existing methods used for ACVPs prediction and it has the potential to become a valuable tool in ACVP drug design. A user-friendly webserver of FEOpti-ACVP can be accessed at http://servers.aibiochem.net/soft/FEOpti-ACVP/.
Background:: Chronic liver disease (CLD) will affect the enhancement of hepatic parenchyma and portal vein on abdominal-enhanced MRI. Objective:: To investigate the difference in liver parenchyma and portal vein enhancement in patients with CLD of different liver function grades between Gd- EOB-DTPA and Gd-DPTA in the portal venous phase (PVP). Methods:: This retrospective study included 218 patients with CLD who had undergone abdominal enhanced MRI from January 2019 to June 2020. Patients with various degrees of liver dysfunction were identified with Child-Turcotte-Pugh and albumin-bilirubin grade. Two readers measured the precontrast and PVP signal intensities of liver parenchyma, portal vein, spleen, and psoas muscle. Relative liver enhancement, liver-to-spleen contrast index, portal vein image contrast, and portal vein-to-liver contrast were calculated. Results:: The relative enhancement of liver parenchyma was significantly lower for the Gd-EOB-DTPA group in any degree of liver function than the Gd- DTPA group in the PVP. The Gd-EOB-DTPA group showed significantly lower portal vein-to-liver contrast in the overall study population, CTP class B, and ALBI grade 2 patients compared to the group of Gd-DTPA at PVP. No significant difference was noted in the portal vein image contrast between the two contrast agents, regardless of CTP and ALBI grading. Conclusion:: In CLD patients, Gd-EOB-DTPA yielded lower liver parenchymal enhancement and similar portal vein image contrast compared to Gd-DTPA in the PVP. Portal vein-to-liver contrast in the Gd-EOB-DTPA group was lower in the CTP class B and ALBI grade 2 subgroups compared to the Gd- DTPA group.
Umami peptides enhance the umami taste of food and have good food processing properties, nutritional value, and numerous potential applications. Wet testing for the identification of umami peptides is a time-consuming and expensive process. Here, we report the iUmami-DRLF that uses a logistic regression (LR) method solely based on the deep learning pre-trained neural network feature extraction method, unified representation (UniRep based on multiplicative LSTM), for feature extraction from the peptide sequences. The findings demonstrate that deep learning representation learning significantly enhanced the capability of models in identifying umami peptides and predictive precision solely based on peptide sequence information. The newly validated taste sequences were also used to test the iUmami-DRLF and other predictors, and the result indicates that the iUmami-DRLF has better robustness and accuracy and remains valid at higher probability thresholds. The iUmami-DRLF method can aid further studies on enhancing the umami flavor of food for satisfying the need for an umami-flavored diet.
Anticancer peptides (ACPs) represent a promising new therapeutic approach in cancer treatment. They can target cancer cells without affecting healthy tissues or altering normal physiological functions. Machine learning algorithms have increasingly been utilized for predicting peptide sequences with potential ACP effects. This study analyzed four benchmark datasets based on a well-established random forest (RF) algorithm. The peptide sequences were converted into 566 physicochemical features extracted from the amino acid index (AAindex) library, which were then subjected to feature selection using four methods: light gradient-boosting machine (LGBM), analysis of variance (ANOVA), chi-squared test (Chi2), and mutual information (MI). Presenting and merging the identified features using Venn diagrams, 19 key amino acid physicochemical properties were identified that can be used to predict the likelihood of a peptide sequence functioning as an ACP. The results were quantified by performance evaluation metrics to determine the accuracy of predictions. This study aims to enhance the efficiency of designing peptide sequences for cancer treatment.
T cell proliferation regulators (Tcprs), which are positive regulators that promote T cell function, have made great contributions to the development of therapies to improve T cell function. CAR (chimeric antigen receptor) -T cell therapy, a type of adoptive cell transfer therapy that targets tumor cells and enhances immune lethality, has led to significant progress in the treatment of hematologic tumors. However, the applications of CAR-T in solid tumor treatment remain limited. Therefore, in this review, we focus on the development of Tcprs for solid tumor therapy and prognostic prediction. We summarize potential strategies for targeting different Tcprs to enhance T cell proliferation and activation and inhibition of cancer progression, thereby improving the antitumor activity and persistence of CAR-T. In summary, we propose means of enhancing CAR-T cells by expressing different Tcprs, which may lead to the development of a new generation of cell therapies.
Interleukin-10 (IL-10) has anti-inflammatory properties and is a crucial cytokine in regulating immunity. The identification of IL-10 through wet laboratory experiments is costly and time-intensive. Therefore, a new IL-10-induced peptide recognition method, IL10-Stack, was introduced in this research, which was based on unified deep representation learning and a stacking algorithm. Two approaches were employed to extract features from peptide sequences: Amino Acid Index (AAindex) and sequence-based unified representation (UniRep). After feature fusion and optimized feature selection, we selected a 1900-dimensional UniRep feature vector and constructed the IL10-Stack model using stacking. IL10-Stack exhibited excellent performance in IL-10-induced peptide recognition (accuracy (ACC) = 0.910, Matthews correlation coefficient (MCC) = 0.820). Relative to the existing methods, IL-10Pred and ILeukin10Pred, the approach increased in ACC by 12.1% and 2.4%, respectively. The IL10-Stack method can identify IL-10-induced peptides, which aids in the development of immunosuppressive drugs.
EDITORIAL article Front. Genet., 03 February 2023Sec. Computational Genomics Volume 14 - 2023 | https://doi.org/10.3389/fgene.2023.1150688
Objective: To investigate the chest computed tomography (CT) findings and dynamic changes in patients infected with SARS-CoV-2 Omicron variants. Methods: 200 patients infected with SARS-CoV-2 Omicron variants were collected in Beijing Ditan Hospital, Capital Medical University from November 2022 to January 2023. These patients were divided into mild group, moderate group and severe/critical group according to the clinical classification. All patients’ clinical, laboratory and chest CT data were retrospectively analyzed. Results: Among 200 cases infected with SARS-CoV-2 omicron variant, the main clinical manifestations were fever, cough, sore throat and fatigue. There was a statistically significant difference in white blood cell count between the mild group and the medium group, and between the mild group and the severe/critical group. There were significant differences in erythrocyte sedimentation rate between mild and moderate groups, and between mild and severe/critical groups. Most of the lesions in mild group were subpleural (53.6%), while most of the lesions in moderate group (77.9%) and severe/critical group (88.9%) were mixed. The crazy-paving sign was statistically significant between the mild and severe/critical groups, and between the mild and moderate groups. There were significant differences in air bronchogram sign between the mild and the severe/critical groups, the mild and severe/critical groups , and the mild and moderate groups. The frequency of Ground Glass Opacity (GGO) was the highest at different intervals between the onset and the first chest CT. The proportion of GGO with consolidation/consolidation and air bronchogram sign gradually increased when the interval was more than 4 days. The proportion of GGO with consolidation/consolidation and air bronchogram sign gradually increased when the interval was more than 4 days. The highest proportion (95.4%) of crazy-paving sign appeared within the interval of 5-9 days, after which the proportion decreased. The frequency of irregular linear opacities, the proportion of pleural thickening and pleural effusion increased in patients with an interval more than 14 days. The median times to occurrence of GGO, GGO with crazy-paving sign, GGO with consolidation or consolidation and irregular linear opacities respectively were 4 days (2 days, 7 days), 9 days (7 days, 11 days), 13 days (10 days, 16 days) and 16 days (13 days, 19 days). Conclusions: Chest CT can reflect the distribution, morphology, dynamic imaging development and outcome of lesions in patients infected with SARS-CoV-2 Omicron variant, which is helpful for clinical treatment decision-making and efficacy evaluation.
Thermophilic proteins have great potential to be utilized as biocatalysts in biotechnology. Machine learning algorithms are gaining increasing use in identifying such enzymes, reducing or even eliminating the need for experimental studies. While most previously used machine learning methods were based on manually designed features, we developed BertThermo, a model using Bidirectional Encoder Representations from Transformers (BERT), as an automatic feature extraction tool. This method combines a variety of machine learning algorithms and feature engineering methods, while relying on single-feature encoding based on the protein sequence alone for model input. BertThermo achieved an accuracy of 96.97% and 97.51% in 5-fold cross-validation and in independent testing, respectively, identifying thermophilic proteins more reliably than any previously described predictive algorithm. Additionally, BertThermo was tested by a balanced dataset, an imbalanced dataset and a dataset with homology sequences, and the results show that BertThermo was with the best robustness as comparied with state-of-the-art methods. The source code of BertThermo is available.
Abstract We performed a retrospectively study in a tertiary infectious diseases hospital in Beijing to explore the prevalence and risk factors of NTM among individuals with symptoms suggestive of pulmonary TB. This was a retrospective study of characteristics of patients with suggestive of active TB at Beijing Ditan Hospital. TB accounted for 93.3% of the burden of disease in Beijing cohort of HIV-infected patients with mycobacterial infections, whereas the other 6.7% were due to NTM infections. The receiver operating characteristic curve (ROC) of Albumin combined with CD4/CD8 value for diagnosing active TB from NTM cases was 0.638, and the optimal cut-off values for Albumin and CD4/CD8 were determined as 36.15 g/L and 0.17, respectively. Overall, the most prevalent NTM species associated with pulmonary infections in HIV-infected individuals was M. intracellulare. CD4/CD8 ratio and albumin level indicating their potential as surrogate marker to differentiate TB and NTM infection in HIV-infected population.
Background Acute-on-chronic liver failure (ACLF) is a syndrome with high 28- and 90-day mortality rates. Magnetic resonance imaging (MRI) has been widely used to diagnose and evaluate liver disease. Our purpose is to determine the value of the imaging features derived from Gd-DTPA-enhanced MRI for predicting the poor outcome of patients with ACLF and develop a clinically practical radiological score. Methods This retrospective study comprised 175 ACLF patients who underwent Gd-DTPA-enhanced abdominal MRI from January 2017 to December 2021. The primary end-point was 90-day mortality. Imaging parameters, such as diffuse hyperintense of the liver on T2WI, patchy enhancement of the liver at the arterial phase, uneven enhancement of the liver at the portal vein phase, gallbladder wall edema, periportal edema, ascites, esophageal and gastric varix, umbilical vein patefac, portal vein thrombosis, and splenomegaly were screened. Cox proportional hazard regression models were used to evaluate prognostic factors and develop a prediction model. The accuracy of the model was evaluated by receiver operating characteristic (ROC) curves. Results During the follow-up period, 31 of the 175 ACLF patients died within 90 days. In the multivariate analysis, three imaging parameters were independently associated with survival: diffuse hyperintense on T2WI (p = 0.007; HR = 3.53 [1.40–8.89]), patchy enhancement at the arterial phase (p = 0.037; HR = 2.45 [1.06–5.69]), moderate ascites (vs. mild) (p = 0.006; HR = 4.12 [1.49–11.36]), and severe ascites (vs. mild) (p = 0.005; HR = 4.29 [1.57–11.71]). A practical radiological score was proposed, based on the presence of diffuse hyperintense (7 points), patchy enhancement (5 points), and ascites (6, 8, and 8 points for mild, moderate, and severe, respectively). Further analysis showed that a cut-off at 14 points was optimum to distinguish high-risk (score > 14) from the low-risk group (score ≤ 14) for 90-day survival and demonstrated a mean area under the ROC curve of 0.774 in ACLF patients. Conclusions Gd-DTPA-enhanced MR imaging features can predict poor outcomes in patients with ACLF, based on which we proposed a clinically practical radiological score allowing stratification of the 90-day survival.
Among many machine learning models for analyzing the relationship between miRNAs and diseases, the prediction results are optimized by establishing different machine learning models, and less attention is paid to the feature information contained in the miRNA sequence itself. This study focused on the impact of the different feature information of miRNA sequences on the relationship between miRNA and disease. It was found that when the graph neural network used was the same and the miRNA features based on the K-spacer nucleic acid pair composition (CKSNAP) feature were adopted, a better graph neural network prediction model of miRNA–disease relationship could be built (AUC = 93.71%), which was 0.15% greater than the best model in the literature based on the same benchmark dataset. The optimized model was also used to predict miRNAs related to lung tumors, esophageal tumors, and kidney tumors, and 47, 47, and 37 of the top 50 miRNAs related to three diseases predicted separately by the model were consistent with descriptions in the wet experiment validation database (dbDEMC).