With the frequent outbreak of viral pandemics, the search for efficient antiviral drugs has become an urgent task. Antiviral peptides (AVPs) have been proven to prevent the infection of host cells by viruses. This study proposes a novel tool called Datt-AVP for AVP prediction based on the peptide sequences. A dual-channel deep learning model of Long Short-Term Memory network and convolutional neural network, along with a pre-trained protein language model as a feature extractor, was applied to capture features in the antiviral peptide sequences simultaneously. Furthermore, self-attention module was added to promote the performance of our prediction results. It showed good recognition ability for antiviral peptides with different lengths. Our tool achieved an accuracy rate of 96.1% on the benchmark dataset, which out-performed state-of-the-art tools.
BACKGROUND:Tuberculosis (TB) represents a major global health challenge. Drug resistance in Mycobacterium tuberculosis (MTB) poses a substantial obstacle to effective TB treatment. Identifying genomic mutations in MTB isolates holds promise for unraveling the underlying mechanisms of drug resistance in this bacterium.METHODS:In this study, we investigated the roles of single nucleotide variants (SNVs) in MTB isolates resistant to four antibiotics (moxifloxacin, ofloxacin, amikacin, and capreomycin) through whole-genome analysis. We identified the drug-resistance-associated SNVs by comparing the genomes of MTB isolates with reference genomes using the MuMmer4 tool.RESULTS:We observed a strikingly high proportion (94.2%) of MTB isolates resistant to ofloxacin, underscoring the current prevalence of drug resistance in MTB. An average of 3529 SNVs were detected in a single ofloxacin-resistant isolate, indicating a mutation rate of approximately 0.08% under the selective pressure of ofloxacin exposure. We identified a set of 60 SNVs associated with extensively drug-resistant tuberculosis (XDR-TB), among which 42 SNVs were non-synonymous mutations located in the coding regions of nine key genes (ctpI, desA3, mce1R, moeB1, ndhA, PE_PGRS4, PPE18, rpsA, secF). Protein structure modeling revealed that SNVs of three genes (PE_PGRS4, desA3, secF) are close to the critical catalytic active sites in the three-dimensional structure of the coding proteins.CONCLUSION:This comprehensive study elucidates novel resistance mechanisms in MTB against antibiotics, paving the way for future design and development of anti-tuberculosis drugs.
Klebsiella pneumoniae is a type of Gram-negative bacterium which can cause a range of infections in human. In recent years, an increasing number of strains of K. pneumoniae resistant to multiple antibiotics have emerged, posing a significant threat to public health. The protein function of this bacterium is not well known, thus a systematic investigation of K. pneumoniae proteome is in urgent need. In this study, the protein functions of this bacteria were re-annotated, and their function groups were analyzed. Moreover, three machine learning models were built to identify novel virulence factors. Results showed that the functions of 16 uncharacterized proteins were first annotated by sequence alignment. In addition, K. pneumoniae proteins share a high proportion of homology with Haemophilus influenzae and a low homology proportion with Chlamydia pneumoniae. By sequence analysis, 10 proteins were identified as potential drug targets for this bacterium. Our model achieved a high accuracy of 0.901 in the benchmark dataset. By applying our models to K. pneumoniae, we identified 39 virulence factors in this pathogen. Our findings could provide novel clues for the treatment of K. pneumoniae infection.
Introduction. The resistance rate of Klebsiella pneumoniae (K. pneumoniae) to imipenem is increasing year by year, and the imipenem resistance mechanism of K. pneumoniae is complex. Therefore, it is urgent to develop new strategies to explore the resistance mechanism of imipenem for its effective and accurate use in clinical practice.Hypothesis/Gap sStatement. Machine learning could identify resistance features and biological process that influence microbial resistance from whole-genome sequencing (WGS) data.Aims. This work aimed to predict imipenem resistance genetic features in K. pneumoniae from whole-genome k-mer features, and analyse their function for understanding its resistance mechanism.Methods. This study analysed WGS data of K. pneumoniae combined with resistance phenotype for imipenem, and established K. pneumoniae to imipenem genotype-phenotype model to predict resistance features using chi-squared test and random forest. An external clinical dataset was used to verify prediction power of resistance features. The potential genes were identified through alignment the resistance features with the K. pneumoniae reference genome using blastn, the functions of potential genes were further analysed to explore its resistance-related signalling pathways with GO and KEGG analysis, the resistance sequence patterns were screened using streme software. Finally, the resistance features were combined and modelled through four machine-learning algorithms (logistic regression, SVM, GBDT and XGBoost) to evaluate their phenotype prediction ability.Results. A total of 16 670 imipenem resistance features were predicted from genotype-phenotype model. The 30 potential genes were identified by annotating the resistance features and corresponded to known antibiotic-related genes (mdtM, dedA, rne, etc.). GO and KEGG pathway analyses indicated the possible association of imipenem resistance with metabolism process and cell membrane. CRYCAGCDN and CGRDAAAN were found from the imipenem resistance features, which were widely presented in the reported β-lactam resistance genes (bla SHV, bla CTX-M, bla TEM, etc.), and YCYAGCMCAST with metabolic functions (organic substance metabolic process, nitrogen compound metabolic process and cellular metabolic process) was identified from the top 50 resistance features. The 25 resistance genes in the training dataset included 19 genes in the external dataset, which verified the accuracy of prediction. The area under curve values of logistics regression, SVM, GBDT and XGBoost were 0.965, 0.966, 0.969 and 0.969, respectively, indicating that the imipenem resistance features have a strong prediction power.Conclusion. Machine-learning methods could effectively predict the imipenem resistance feature in K. pneumoniae, and provide resistance sequence profiles for predicting resistance phenotype and exploring potential resistance mechanisms. It provides an important insight into the potential therapeutic strategies of K. pneumoniae resistance to imipenem, and speed up the application of machine learning in routine diagnosis.
Modules consisting of antibiotic resistance genes (ARGs) flanked by inverted repeat Xer-specific recombination sites were thought to be mobile genetic elements that promote horizontal transmission. Less frequently, the presence of mobile modules in plasmids, which facilitate a pdif-mediated ARGs transfer, has been reported. Here, numerous ARGs and toxin-antitoxin genes have been found in pdif site pairs. However, the mechanisms underlying this apparent genetic mobility is currently not understood, and the studies relating to pdif-mediated ARGs transfer onto most bacterial genera are lacking. We developed the web server pdifFinder based on an algorithm called PdifSM that allows the prediction of diverse pdif-ARGs modules in bacterial genomes. Using test set consisting of almost 32 thousand plasmids from 717 species, PdifSM identified 481 plasmids from various bacteria containing pdif sites with ARGs. We found 28-bp-long elements from different genera with clear base preferences. The data we obtained indicate that XerCD-dif site-specific recombination mechanism may have evolutionary adapted to facilitate the pdif-mediated ARGs transfer. Through multiple sequence alignment and evolutionary analyses of duplicated pdif-ARGs modules, we discovered that pdif sites allow an interspecies transfer of ARGs but also across different genera. Mutations in pdif sites generate diverse arrays of modules which mediate multidrug-resistance, as these contain variable numbers of diverse ARGs, insertion sequences and other functional genes. The identification of pdif-ARGs modules and studies focused on the mechanism of ARGs co-transfer will help us to understand and possibly allow controlling the spread of MDR bacteria in clinical settings. The pdifFinder code, standalone software package and description with tutorials are available at https://github.com/mjshao06/pdifFinder.
Probiotics are widely used in the fields of food and medicine. Some of these beneficial bacteria carry various antibiotic resistance genes and virulence factors, which have adverse effects on human health. Safety evaluation of probiotics is a unique challenge since the potential risks are related to multiple factors. This study constructed a Potential Risk Comprehensive Evaluation model of Probiotic Species (PRCE-PS), which could provide a quantitative risk evaluation of probiotics based on complete genome sequences. The model employed the analytic hierarchy process (AHP) to determine the weight of risk factors and the fuzzy comprehensive evaluation (FCE) method to form the comprehensive risk-decision process at the species level. The PRCE-PS model yielded four ranks (I, high-risk; II, medium-risk; III, low-risk; IV, safe) to represent the risk status of different probiotic species. By applying the PRCE-PS model to evaluate 32 popular beneficial species, most of them were relatively safe, but some species showed unsafe risks. Lactobacillus plantarum and B. animalis were evaluated as medium-risk, showing the greatest potential risk of antibiotic resistance genes transfer onto all the Lactobacillus and Bifidobacterium. Bacillus subtilis, E. faecalis, and E. faecium were assessed as high-risk species with great risks both on antibiotic resistance genes transferring and exotoxin-related virulence genes carrying. The PRCE-PS model provides a clear and comprehensive potential risk evaluation result to probiotic species, and helps to improve the understanding of the safety of these beneficial species.