The pharmacokinetic profile of a potential drug is largely determined by its metabolic stability, which reflects its susceptibility to biotransformation. Metabolic stability data allow one to assess the therapeutic value of a compound and its toxicological risk. This assesment relies primarily on pharmacokinetic parameters, particularly half-life (t1/2) and clearance (CL), which are typically determined using in vitro systems including hepatocytes and liver microsomal fractions. Using the publicly available ChEMBL v. 35 and PubChem databases, we collected over 8000 chemical compounds with experimental intrinsic CL and/or half-life data from liver microsome assays obtained in mice, rats, and humans. Different thresholds were applied to differentiate the stable and unstable molecules. The Naive Bayesian classifier with MNA (Multilevel Neighborhoods of Atoms) descriptors and Self-Consistent Extreme Classifier (SCEC) with QNA (Quantitative Neighborhoods of Atoms) descriptors were used for creating classification models. The accuracy (AUC) of most classification models exceeded 0.85. Self-Consistent Regression was used to create quantitative models. The coefficient of determination of the regression models varied from 0.35 (rat, t1/2) to 0.7 (human, CLint). These models were integrated into the freely available web application MetaStab-Analyzer, which provides a unique combination of qualitative (stable/unstable/moderate) and quantitative predictions for three species. A key feature of the application is the providing of numerical metrics for each prediction, which increases its interpretability. This combination of innovative algorithms (SCR and SCEC), dual qualitative-quantitative assessment, and a user-friendly interface is not available in any existing tool. MetaStab-Analyzer is freely available at https://www.way2drug.com/metastab/.
Background and Objectives The development of viral resistance can significantly reduce the effectiveness of therapy. Human immunodeficiency virus type 1 is the cause of chronic immune dysfunction, leading to the development of co-infections and serious complications. Despite worldwide progress and consolidated efforts to overcome HIV drug resistance, the development of novel approaches for rational drug therapy of HIV infection is still needed for building models with high accuracy of prediction and that can be applied for evaluation of resistance against wide variety of inhibitors. Our study is dedicated to the development of a novel computational ML-driven approach for the ternary classification of HIV protease, reverse transcriptase, and integrase sequences. Binary classification approaches naturally are not applicable to capture clinically important intermediate resistance levels, motivating the use of a ternary classification model. Methods For the model development we used the Self-Consistent Extreme Classifier. One-versus-rest and one-versus-one ternary approaches were applied to sequences related resistance data from Stanford University HIV Drug Resistance Database (StDB). Results For the final classifiers we selected the most appropriate models with 0.913 sensitivity, 0.894 specificity, 0.741 precision and 0.953 area under ROC, all values provided in average. We tested our approach in a clinical task and performed prospective validation for eight sequences of HIV protease and reverse transcriptase obtained from treatment-naive HIV-positive male patients. We performed a prediction and compared the results with the therapeutic outcome, in particular, with the viral load decline at 24 weeks. Conclusions The results of the prospective validation are generally consistent with the results of the therapeutic outcome and confirm the possibility of using the developed approach for the selection of the most appropriate therapeutic regimens.
Understanding the biotransformation of xenobiotics in the human body is critical for a comprehensive assessment of drug effects since pharmacologically active drug metabolites may exhibit a range of biological effects that often differ from those of the original pharmaceutical agent. Studies of the biotransformation mechanisms of xenobiotics have resulted in numerous publications. Extracting information about the parent compounds (substrates) and their metabolites from the texts allows retrieval of information on their biological activities, molecular mechanisms of action, and toxicity. Manual curation of the names of xenobiotics, their metabolites, and biotransformation reactions in the text is a challenging task due to the large number of publications related to studies of pharmaceutical agents metabolism. Our aim is to create an annotated corpus of texts that can be used for automated extraction of the names of xenobiotics, including pharmaceutical agents that undergo biotransformation and their metabolites. Prior to manual annotation of the corpus, semiautomatic annotation was carried out based on the earlier developed rule-based method for parent compounds and their metabolites extraction. To create XenoMet, we automatically extracted relevant texts from PubMed using a query based on MeSH terms. The names of biotransformation reactions were recognized by using an in-house-developed dictionary. Then, we manually verified the extracted data by correcting errors in the named entity annotation and identified the associations between substrates and metabolites. We tested the applicability of XenoMet for the reconstruction of a metabolic tree and for the automated extraction of the chemical names of substrates, metabolites, and reactions of biotransformation. Classification of the named entities of metabolites, substrates, and biotransformation reactions by a conditional random fields approach using XenoMet as the training set provides an F1-score of 0.79.
In the human body, pharmacological substances undergo biotransformation, therefore, during drugs development, it is necessary to take into account the biological activity spectra of their metabolites. Previously, we created the MetaPASS web application to analyze the probable spectra of biological activity of drug-like organic compounds taking into account their metabolism. Here we describe a new version of MetaPASS 2024 (https://www.way2drug.com/metapass), containing increased number of known metabolic pathways, and added procedures for searching structural similarity based on MNA and QNA descriptors and searching for compounds with the highest probability estimate for target biological activity; we have also implemented representation of the spectrum of biological activity in the form of treemaps.
This study presents an approach for the in silico assessment of potential geroprotectors that target the multifaceted mechanisms of aging, implemented in the PASS GERO web application. This work is timely given the societal impact of aging—the primary risk factor for major chronic diseases. The urgent need to extend healthspan—the period of life spent in good health—motivates the search for compounds that modulate fundamental aging mechanisms. The model estimates the probabilities of 117 aging-related biological activities with high predictive accuracy, achieving an average Invariant Accuracy of Prediction (IAP) of 0.967 under cross-validation. Validation using known geroprotectors (rapamycin, metformin, and resveratrol) demonstrated strong concordance between predicted activities and documented molecular mechanisms of action. For instance, the model correctly predicted rapamycin’s inhibition of mTOR and metformin’s activation of AMPK. The PASS GERO web application provides a systematic strategy to prioritize novel compound candidates for experimental evaluation in anti-aging research. We discuss challenges including the chemical diversity of the training data, the need for validated biomarkers, and the limitations of translating computational predictions into clinical outcomes, positioning the tool as robust application for activity profiling in discovery workflows.
Drug resistance of pathogens, including viruses, is one of the reasons for decreased efficacy of therapy. Considering the impact of HIV type 1 (HIV-1) on the development of progressive immune dysfunction and the rapid development of drug resistance, the analysis of HIV-1 resistance is of high significance. Currently, a substantial amount of data has been accumulated on HIV-1 drug resistance that can be used to build both qualitative and quantitative models of HIV-1 drug resistance. Quantitative models of drug resistance can enrich the information about the efficacy of a particular drug in the scheme of antiretroviral therapy. In our study, we investigated the possibility of developing models for quantitative prediction of HIV-1 resistance to eight protease inhibitors based on the analysis of amino acid sequences of HIV-1 protease for 900 virus variants. We developed random forest regression (RFR), support vector regression (SVR), and self-consistent regression (SCR) models using binary vectors containing values from 0 or 1, depending on the presence of a specific peptide fragment in each amino acid sequence as independent variables, while fold ratio, reflecting the level of resistance, was the predicted variable. The SVR and SCR models showed the highest predictive performances. The models built demonstrate reasonable performances for eight out of nine (R2 varied from 0.828 to 0.909) protease inhibitors, while R2 for predicting tipranavir fold ratio was lower (R2 was 0.642). We believe that the developed approach can be applied to evaluate drug resistance of molecular targets of other viruses where appropriate experimental data are available.
Being widely accepted tools in computational drug search, the (Q)SAR methods have limitations related to data incompleteness. The proteochemometrics (PCM) approach expands the applicability area by using description for both protein and ligand structures. The PCM algorithms are urgently required for the development of new antiviral agents. We suggest the PCM method using the TLMNA descriptors, combining the MNA descriptors of ligands and protein sequence N-grams. Our method was validated on the viral chymotrypsin-like proteases and their ligands. We have developed an original protocol allowing us to collect a comprehensive set of 15 protein sequences and more than 9000 ligands from the ChEMBL database. The N-grams were derived from the 3D-based alignment, accurately superposing ligand-binding regions. In testing the ligand set in SAR mode with MNA descriptors, an accuracy above 0.95 was determined that shows the perspective of the antiviral drug search in virtual chemical libraries. The effective PCM models were built with the TLMNA descriptor. The strong validation procedure with pair exclusion simulated the prediction of interactions between the new ligands and new targets, resulting in accuracy estimation up to 0.89. The PCM approach shows slightly lower accuracy caused by more uncertainty compared with SAR, but it overcomes the problem of data incompleteness.
The analysis of drug-induced gene expression profiles (DIGEP) is widely used to estimate the potential therapeutic and adverse drug effects as well as the molecular mechanisms of drug action. However, the corresponding experimental data is absent for many existing drugs and drug-like compounds. To solve this problem, we created the DIGEP-Pred 2.0 web application, which allows predicting DIGEP and potential drug targets by structural formula of drug-like compounds. It is based on the combined use of structure-activity relationships (SARs) and network analysis. SAR models were created using PASS (Prediction of Activity Spectra for Substances) technology for data from the Comparative Toxicogenomics Database (CTD), the Connectivity Map (CMap) for the prediction of DIGEP, and PubChem and ChEMBL for the prediction of molecular mechanisms of action (MoA). Using only the structural formula of a compound, the user can obtain information on potential gene expression changes in several cell lines and drug targets, which are potential master regulators responsible for the observed DIGEP. The mean accuracy of prediction calculated by leave-one-out cross validation was 86.5 % for 13377 genes and 94.8 % for 2932 proteins (CTD data), and it was 97.9 % for 2170 MoAs. SAR models (mean accuracy-87.5 %) were also created for CMap data given on MCF7, PC3, and HL60 cell lines with different threshold values for the logarithm of fold changes: 0.5, 0.7, 1, 1.5, and 2. Additionally, the data on pathways (KEGG, Reactome), biological processes of Gene Ontology, and diseases (DisGeNet) enriched by the predicted genes, together with the estimation of target-master regulators based on OmniPath data, is also provided. DIGEP-Pred 2.0 web application is freely available at https://www.way2drug.com/digep-pred.
Novel antimycobacterial compounds are needed to expand the existing toolbox of therapeutic agents, which sometimes fail to be effective. In our study we extracted, filtered, and aggregated the diverse data on antimycobacterial activity of chemical compounds from the ChEMBL database version 24.1. These training sets were used to create the classification and regression models with PASS and GUSAR software. The IOC chemical library consisting of approximately 200,000 chemical compounds was screened using these (Q)SAR models to select novel compounds potentially having antimycobacterial activity. The QikProp tool (Schrodinger) was used to predict ADME properties and find compounds with acceptable ADME profiles. As a result, 20 chemical compounds were selected for further biological evaluation, of which 13 were the Schiff bases of isoniazid. To diversify the set of selected compounds we applied substructure filtering and selected an additional 10 compounds, none of which were Schiff bases of isoniazid. Thirty compounds selected using virtual screening were biologically evaluated in a REMA assay against the M. tuberculosis strain H37Rv. Twelve compounds demonstrated MIC below 20 mu M (ranging from 2.17 to 16.67 mu M) and 18 compounds demonstrated substantially higher MIC values. The discovered antimycobacterial agents represent different chemical classes.
The accurate prediction of secondary structures of proteins (SSPs) is a critical challenge in molecular biology and structural bioinformatics. Despite recent advancements, this task remains complex and demands further exploration. This study presents a novel approach to SSP prediction using atom-centric substructural multilevel neighborhoods of atoms (MNA) descriptors for protein molecular fragments. A dataset comprising over 335,000 SSPs, annotated by the Dictionary of Secondary Structure in Proteins (DSSP) software from 37,000 proteins, was constructed from Protein Data Bank (PDB) records with a resolution of 2 Å or better. Protein fragments were converted into structural formulae using the RDKit Python package and stored in SD files using the MOL V3000 format. Classification sequence–structure–property relationships (SSPR) models were developed with varying levels of MNA descriptors and a Bayesian algorithm implemented in MultiPASS software. The average prediction accuracy (AUC) for eight SSP types, calculated via leave-one-out cross-validation, was 0.902. For independent test sets (ASTRAL and CB513 datasets), the best SSPR models achieved AUC, Q3, and Q8 values of 0.860, 77.32%, 70.92% and 0.889, 78.78%, 74.74%, respectively. Based on the created models, a freely available web application MNA-PSS-Pred was developed.
Potently affecting human and animal brain and behavior, hallucinogenic drugs have recently emerged as potentially promising agents in psychopharmacotherapy. Complementing laboratory rodents, the zebrafish (Danio rerio) is a powerful model organism for screening neuroactive drugs, including hallucinogens. Here, we tested four novel N-benzyl-2-phenylethylamine (NBPEA) derivatives with 2,4- and 3,4-dimethoxy substitutions in the phenethylamine moiety and the -F, -Cl, and -OCF3 substitutions in the ortho position of the phenyl ring of the N-benzyl moiety (34H-NBF, 34H-NBCl, 24H-NBOMe(F), and 34H-NBOMe(F)), assessing their behavioral and neurochemical effects following chronic 14 day treatment in adult zebrafish. While the novel tank test behavioral data indicate anxiolytic-like effects of 24H-NBOMe(F) and 34H-NBOMe(F), neurochemical analyses reveal reduced brain norepinephrine by all four drugs, and (except 34H-NBCl) - reduced dopamine and serotonin levels. We also found reduced turnover rates for all three brain monoamines but unaltered levels of their respective metabolites. Collectively, these findings further our understanding of complex central behavioral and neurochemical effects of chronically administered novel NBPEAs and highlight the potential of zebrafish as a model for preclinical screening of small psychoactive molecules.
In silico prediction of cell line cytotoxicity considerably decreases time and financial costs during drug development of new antineoplastic agents. (Q)SAR models for the prediction of drug-like compound cytotoxicity in relation to nine breast cancer cell lines (T47D, ZR-75-1, MX1, Hs-578T, MCF7-DOX, MCF7, Bcap37, MCF7R, BT-20) were created by GUSAR software based on the data from ChEMBL database (v. 30). The separate datasets related with IC50 and IG50 values were used for the creation of (Q)SAR models for each cell line. Based on leave-one-out and 5F CV procedures, 24 reasonable (Q)SAR models were selected for the creation of a freely available web-application (BC CLC-Pred: https://www.way2drug.com/bc/) to predict substance cytotoxicity in relation to human breast cancer cell lines. The mean accuracies of prediction r2, RMSE, Balance Accuracy for the selected (Q)SAR models calculated by 5F CV were 0.599, 0.679 and 0.875, respectively. As a result, BC CLC-Pred provides simultaneous quantitative and qualitative predictions of IC50 and IG50 values for most of the nine breast cancer cell lines, which may be helpful in selecting promising compounds and optimizing lead compounds during the development of new antineoplastic agents against breast cancer.
Predicting viral drug resistance is a significant medical concern. The importance of this problem stimulates the continuous development of experimental and new computational approaches. The use of computational approaches allows researchers to increase therapy effectiveness and reduce the time and expenses involved when the prescribed antiretroviral therapy is ineffective in the treatment of infection caused by the human immunodeficiency virus type 1 (HIV-1). We propose two machine learning methods and the appropriate models for predicting HIV drug resistance related to amino acid substitutions in HIV targets: (i) k-mers utilizing the random forest and the support vector machine algorithms of the scikit-learn library, and (ii) multi-n-grams using the Bayesian approach implemented in MultiPASSR software. Both multi-n-grams and k-mers were computed based on the amino acid sequences of HIV enzymes: reverse transcriptase and protease. The performance of the models was estimated by five-fold cross-validation. The resulting classification models have a relatively high reliability (minimum accuracy for the drugs is 0.82, maximum: 0.94) and were used to create a web application, HVR (HIV drug Resistance), for the prediction of HIV drug resistance to protease inhibitors and nucleoside and non-nucleoside reverse transcriptase inhibitors based on the analysis of the amino acid sequences of the appropriate HIV proteins from clinical samples.
The search for the relationships between CDR3 TCR sequences and epitopes or MHC types is a challenging task in modern immunology. We propose a new approach to develop the classification models of structure-activity relationships (SAR) using molecular fragment descriptors MNA (Multilevel Neighbourhoods of Atoms) to represent CDR3 TCR sequences and the naïve Bayes classifier algorithm. We have created the freely available TCR-Pred web application (http://way2drug.com/TCR-pred/) to predict the interactions between α chain CDR3 TCR sequences and 116 epitopes or 25 MHC types, as well as the interactions between β chain CDR3 TCR sequences and 202 epitopes or 28 MHC types. The TCR-Pred web application is based on the data (more 250 000 unique CDR3 TCR sequences) from VDJdb, McPAS-TCR, and IEDB databases and the proposed approach. The average AUC values of the prediction accuracy calculated using a 20-fold cross-validation procedure varies from 0.857 to 0.884. The created web application may be useful in studies related with T-cell profiling based on CDR3 TCR sequences.
The metagenome of bacteria colonizing the human intestine is a set of genes that is almost 150 times greater than the set of host genes. Some of these genes encode enzymes whose functioning significantly expands the number of potential pathways for xenobiotic metabolism. The resulting metabolites can exhibit activity different from that of the parent compound. This can decrease the efficacy of pharmacotherapy as well as induce undesirable and potentially life-threatening side effects. Thus, analysis of the biotransformation of small drug-like compounds mediated by the gut microbiota is an important step in the development of new pharmaceutical agents and repurposing of the approved drugs. In vitro research, the interaction of drug-like compounds with the gut microbiota is a multistep and time-consuming process. Systematic testing of large sets of chemical structures is associated with a number of challenges, including the lack of standardized techniques and significant financial costs to identify the structure of the final metabolites. Estimation of the compounds' ability to be biotransformed by the gut microbiota and prediction of the structures of their metabolites are possible in silico. However, the development of computational approaches is limited by the lack of information about chemical structures metabolized by microbiota enzymes. The aim of this study is to create a database containing information on the metabolism of drug-like compounds by the gut microbiota. We created the data set containing information about 368 structures metabolized and 310 structures not metabolized by the human gut microbiota. The HGMMX database is freely available at https://www.way2drug.com/hgmmx. The information presented will be useful in the development of computational approaches for analyzing the impact of the human microbiota on metabolism of drug-like molecules.
The human gut microbiota (HGM) comprises a complex population of microorganisms that significantly affect human health, including their influence on xenobiotics metabolism. Many pharmaceuticals are taken orally and thus come into contact with HGM, which can metabolize them. Therefore, it is necessary to evaluate the effect of HGM on the fate of pharmaceuticals in the organism. We have collected information about over 600 compounds from more than eighty publications. At least half of them (329 compounds) are known to be metabolized by HGM. We have used PASS (Prediction of Activity Spectra for Substances) software to build three classification SAR models for HGM-mediated drug metabolism prediction. The first model with an accuracy of prediction 0.85 estimates whether compounds will be metabolized by HGM. The second model with an average accuracy of prediction 0.92 estimates which bacterial genera are responsible for the drug metabolism. The third model with an average accuracy of prediction 0.92 estimates the biotransformation reactions during HGM-mediated drug metabolism. The created models were used to develop the freely available web application MDM-Pred (http://www.way2drug.com/mdm-pred/).
Many human diseases including cancer, degenerative and autoimmune disorders, diabetes and others are multifactorial. Pharmaceutical agents acting on a single target do not provide their efficient curation. Multitargeted drugs exhibiting pleiotropic pharmacological effects have certain advantages due to the normalization of the complex pathological processes of different etiology. Extracts of medicinal plants (EMP) containing multiple phytocomponents are widely used in traditional medicines for multifactorial disorders' treatment. Experimental studies of pharmacological potential for multicomponent compositions are quite expensive and time-consuming. In silico evaluation of EMP the pharmacological potential may provide the basis for selecting the most promising directions of testing and for identifying potential additive/synergistic effects. Multiphytoadaptogen (MPhA) containing 70 major phytocomponents of different chemical classes from 40 medicinal plant extracts has been studied in vitro, in vivo and in clinical researches. Antiproliferative and anti-tumor activities have been shown against some tumors as well as evidence-based therapeutic effects against age-related pathologies. In addition, the neuroprotective, antioxidant, antimutagenic, radioprotective, and immunomodulatory effects of MPhA were confirmed. Analysis of the PASS profiles of the biological activity of MPhA phytocomponents showed that most of the predicted anti-tumor and anti-metastatic effects were consistent with the results of laboratory and clinical studies. Antimutagenic, immunomodulatory, radioprotective, neuroprotective and anti-Parkinsonian effects were also predicted for most of the phytocomponents. Effects associated with positive effects on the male and female reproductive systems have been identified too. Thus, PASS and PharmaExpert can be used to evaluate the pharmacological potential of complex pharmaceutical compositions containing natural products
Next Generation Sequencing (NGS) technologies are rapidly entering clinical practice. A promising area for their use lies in the field of newborn screening. The mass screening of newborns using NGS technology leads to the discovery of a large number of new missense variants that need to be assessed for association with the development of hereditary diseases. Currently, the primary analysis and identification of pathogenic variations is carried out using bioinformatic tools. Although extensive efforts have been made in the computational approach to variant interpretation, there is currently no generally accepted pathogenicity predictor. In this study, we used the sequence–structure–property relationships (SSPR) approach, based on the representation of protein fragments by molecular structural formula. The approach predicts the pathogenic effect of single amino acid substitutions in proteins related with twenty-five monogenic heritable diseases from the Uniform Screening Panel for Major Conditions recommended by the Advisory Committee on Hereditary Disorders in Newborns and Children. In order to create SSPR models of classification, we modified a piece of cheminformatics software, MultiPASS, that was originally developed for the prediction of activity spectra for drug-like substances. The created SSPR models were compared with traditional bioinformatic tools (SIFT 4G, Polyphen-2 HDIV, MutationAssessor, PROVEAN and FATHMM). The average AUC of our approach was 0.804 ± 0.040. Better quality scores were achieved for 15 from 25 proteins with a significantly higher accuracy for some proteins (IVD, HADHB, HBB). The best SSPR models of classification are freely available in the online resource SAV-Pred (Single Amino acid Variants Predictor).
Biotransformation of drug-like compounds in the human body may lead to the adverse or toxic effects caused by their metabolites. During early pharmaceutical drug R&D, experimental metabolite structure is often not yet available. To increase the safety profile of novel pharmaceutical agents, a computer-aided assessment of toxicity should be performed based on the structural formulae of both parent compounds and their metabolites. In this chapter, we survey current approaches to the drug metabolite prediction, with subsequent estimation of their action potentially leading to undesirable biological effects. Herein, we propose the concept of integral toxicity that concomitantly reflects the overall biological activity of a pharmaceutical substance and its metabolites. The current possibilities and limitations of the multifaceted computational assessment of xenobiotics toxicity are discussed.