
Cyclin-dependent kinase 8 (CDK8) has emerged as a promising therapeutic target for acute myeloid leukaemia (AML). We developed an integrated computational pipeline combining neural network-based potency prediction with molecular dynamics (MD) simulations for CDK8 inhibitor discovery. A curated dataset of 1,200 unique CDK8 inhibitors was assembled from ChEMBL. The optimal neural network architecture (two hidden layers, 512→128 units) with dropout regularization achieved a test set r2 of 0.47 and RMSE of 0.87. Y-randomization testing confirmed genuine structure-activity relationships. Using an extrapolation-focused training strategy combined with a SELFIES-based genetic recombination algorithm, we generated novel molecular structures beyond the training distribution. SHAP analysis revealed fingerprint bits as critical determinants of potency. Applicability‑domain analysis confirmed that the novel hit falls within validated chemical space. The identified candidate exhibited a predicted pIC50 of 11.97, substantially exceeding the most potent training compound (pIC50 = 10.09) and compound 12 (pIC50 = 7.47). Molecular docking revealed Moldock scores of -143.9 kcal/mol for the novel hit versus -113.9 kcal/mol for compound 12. MD simulations demonstrated stable binding with the novel hit forming a highly stable hydrogen bond with Asp98. MM-GBSA calculations showed superior binding free energy for the novel hit (-99.10 vs. -47.77 kcal/mol).
Per- and polyfluoroalkyl substances (PFAS) are persistent environmental contaminants associated with adverse human health outcomes. However, experimental toxicity data remain unavailable for the vast majority of PFAS, limiting comprehensive risk assessment. In this study, ML-based classification Quantitative Structure Toxicity Relationship (QSTR) models were developed to predict PFAS toxicity across three biologically relevant high-throughput screening endpoints. These included AID-1030 (ALDH1A1 inhibition associated with reproductive toxicity), AID-504444 (Nrf2 pathway inhibition responsible for vascular disruption, hepatic steatosis, lung carcinogenesis, and infertility), and AID-588855 (inhibition of TGF-β/Smad3 signalling linked to developmental toxicity and tumour progression). To address pronounced class imbalance in these datasets, multiple data balancing techniques (ADASYN, SMOTE, Borderline-SMOTE, SVMSMOTE, and random oversampling) were applied. Fourteen ML classifiers were trained for each balanced dataset, yielding 70 models per endpoint. Sum-of-Ranking-Differences (SRD) analysis identified the most robust models with Gradient Boosting, Random Forest, and Support Vector Classifier models emerging as optimal for AID-1030, AID-504444, and AID-588855, respectively. SHAP and substructure analyses provided mechanistic interpretability, by linking PFAS structural features and AOP progression. The optimized models were further applied to an independent external dataset of 2,361 PFAS, and a Python-based screening tool, PERSIST, was developed to screen PFAS.
Leishmaniasis remains a major public health burden due to limited vaccines, toxic treatments, and drug resistance. The M32 metallocarboxypeptidase of Leishmania donovani (LdMCP1) represents a promising target due to its presence in trypanosomatids and absence in mammals. In this study, an in silico strategy was applied to predict the three-dimensional structure of LdMCP1, revealing a canonical α/β topology and the characteristic catalytic motifs of the M32 family. Dimeric interface analysis indicated that the N-terminal subdomain plays a critical role in inter-subunit interactions and overall dimer stability. Molecular docking delineated the binding mode of a peptide substrate within the active site, highlighting key interactions with catalytic residues. Virtual screening of GSK anti-kinetoplastid libraries identified three compounds as lead molecules based on docking scores, substrate interaction patterns, and previously reported in vitro inhibitory activity against TcMCP1. These compounds exhibited favourable physicochemical properties, ADMET and pharmacokinetic profiles. All protein-ligand complexes demonstrated structural integrity, compactness, and flexibility during the molecular dynamics simulations. MM-PBSA binding free energy calculations indicated TCMDC-143515 and TCMDC-143620 bind more strongly than the substrate. Collectively, these findings suggest that LdMCP1 is a druggable enzyme and identified lead molecules provide a rational starting point for the design and optimization of specific antileishmanial agents.
Alzheimer's disease (AD) is a progressive neurodegenerative disorder associated with amyloid-beta accumulation, neuroinflammation, and synaptic dysfunction, making beta-secretase 1 (BACE1) a promising therapeutic target. In the current investigation, we used an integrated computational approach, including virtual screening, machine learning, quantum-chemical analysis, molecular dynamics simulations, ADMET profiling, and network pharmacology, to identify marine fungal metabolites with potential BACE1-binding activity. A virtual chemical library consisting of 4,683 marine-derived compounds was virtually screened against the BACE1 catalytic subdomain to select 100 top compounds with docking energy ranging from -11.9 to -10.2 kcal/mol. The pharmacokinetic study showed promising CNS characteristics. CatBoost was identified as the most reliable model (r2 = 0.243 ± 0.049), predicting pIC50 values of 7.51-7.95 for shortlisted marine fungal metabolites, suggesting potential BACE1 inhibitory activity. Density functional theory computations revealed good ligand reactivity and stability. Molecular dynamics simulation validated the stable ligand-protein interactions, and MM/GBSA analysis selected CMNPD7259 (ΔG_bind = -68.45 kcal/mol) as the most promising molecule. Network pharmacology further associated the metabolites with pathways involved in inflammation, apoptosis, immunological control, and neuronal function. Overall, the results indicate that metabolites from marine fungi are intriguing candidates for developing safer, more efficacious treatments for AD.
Colorectal cancer (CRC) significantly contributes to global cancer-related morbidity and mortality. As systemic toxicity, poor selectivity, and drug resistance restrict the usefulness of available chemotherapy and radiotherapy, phytochemicals present a promising substitute strategy with their generally low toxicity and potential for multitarget action. In this study, key proteins involved in CRC signalling were selected as therapeutic targets, and baicalin, berberine, luteolin, quercetin, and licorice glycoside D1 (licorice) were evaluated as potential multitarget therapeutics. Molecular docking revealed key stabilizing interactions, such as hydrogen bonds, π-π stacking, and π-cation contacts. Additionally, eight complexes investigated via 200 ns all-atom molecular dynamics simulations consistently showed stability in the protein RMSD profiles. MM/PBSA binding free energy calculations identified the BCL-2-baicalin, IL-1β-baicalin, IL-1β-licorice, and TNF-α-licorice complexes (ΔGbind of -27.66, -26.47, -29.69, and -29.87 kcal/mol, respectively) as exhibiting the strongest protein-ligand affinities, where strong van der Waals interactions effectively offset opposing polar contributions. Free-energy surface analyses revealed prominent protein conformations as well as ligand conformations that varied depending on the bound protein. Overall, insights from our computational analyses identified baicalin and licorice as promising multitarget inhibitors while also providing energetic and mechanistic understanding that can be leveraged in future rational drug design.
The dopamine D2 receptor (DRD2) is a key therapeutic target for several neuropsychiatric disorders, driving the need for new ligands with improved safety and efficacy. To find possible DRD2 inhibitors, we developed an integrated in silico workflow in this study that combines drug-likeness filtering, machine learning-based quantitative structure-activity relationship (ML-QSAR) modelling, and structure-based virtual screening. A standardized dataset of 1,128 DRD2 ligands with experimental inhibition constants was assembled, using pKi50 as the activity metric. Regression-based ML-QSAR models were constructed using the PubChem database and Substructure fingerprints. Random Forest techniques demonstrated the best prediction performance and robustness among these models. Crucial DRD2-binding motifs included aromatic systems, heterocycles, alkyl-aryl ethers, and halogenated groups. Strong agreement between predicted and experimental pKi50 values for FDA-approved antipsychotic medications further confirmed the validity of the model. A CNS-targeted chemical library was subjected to virtual screening, and the lead compounds were evaluated using molecular docking against the crystal structure of the dopamine D2 receptor (DRD2). VS012-7128 demonstrated strong binding affinities and formed essential interactions inside the receptor binding pocket in the molecular dynamics simulation. The study documented the effectiveness and reliability of the employed computational approach for identifying potential DRD2 ligands.
Ovarian cancer remains a major global health concern and leading cause of mortality among women due to late diagnosis, therapeutic resistance, and limited predictive biomarkers for treatment response. There is an urgent need for integrative approaches to improve early detection and treatment outcomes. In this study, we integrated machine learning and pharmacogenomics to identify drug-sensitive biomarkers and prioritize therapeutic candidates in ovarian cancer. Pharmacogenomic data were obtained from the Cancer Cell Line Encyclopaedia (CCLE) and Genomics of Drug Sensitivity in Cancer (GDSC-v2), including gene expression and drug response profiles across ovarian cancer cell lines. Predictive models were developed for six FDA-approved drugs using Elastic Net, Ridge, and Lasso regression, demonstrating robust predictive performance achieving Pearson correlation (r) to 0.65 and Spearman correlation (ρ) to 0.63 on validation sets. Biomarker analysis identified key genes associated with drug response, including CCR10 and PLEKHH2, WNT9B, ITPRID1, CHRNG, and DIRAS3. Ligand-based similarity screening against the COCONUT database followed by molecular docking and MD simulation identified three promising compounds (662142, 733302, and 883576) with improved binding affinity and conserved interactions with Topoisomerase-1. This integrative framework highlights the potential of combining machine learning, pharmacogenomics, and molecular modelling for biomarker discovery and drug prioritization in ovarian cancer.
In this study, the multi-target QSAR (mt-QSAR) models were constructed which can predict the inhibitory activity of compounds against various class I HDACs isoforms under different experimental conditions. Models based on mt-QSAR classification (a linear model and six non-linear models) were constructed using 1215 compounds obtained from the ChEMBL database by the Box - Jenkins moving average method using 13 deviation descriptors. The high predictive performance was found in the non-linear models based on Support Vector Classification, Random Forest and Gradient Boosting with accuracies exceeding 90% for the sub-training, test and validation sets. Additionally, virtual screening was performed using the ZINC library to identify a potential hit compound. The SwissADME was used for in silico predictions to assess drug-likeness of the identified virtual hit. Finally, docking and molecular dynamics simulations were performed to study the interactions of target proteins with the hit compound. It was found that coordination of the ligand with the catalytic zinc ion is essential for inhibitory activity. The results offer important insights for the search for new inhibitors of class I HDACs as potential therapeutic agents.
Aedes aegypti is the principal vector responsible for the transmission of several arboviral diseases, including dengue, Zika, chikungunya, and yellow fever, posing a significant threat to global public health. Controlling its population is crucial to limiting the spread of these infections. Among the most effective strategies is the use of larvicides derived from plant sources, which interfere with the mosquito's development and reduce its proliferation. This study presents a QSAR analysis of 255 plant-based larvicidal compounds against Aedes aegypti (Zika vector), using pLC50 values as the measure of bioactivity. QSAR models were developed using the CORAL software, which applies a Monte Carlo optimization algorithm to determine the correlation weights of molecular descriptors. The predictive power and reliability of the models were extensively validated using several statistical indicators. Results demonstrated that the models were robust, simple, and predictive, with validation metrics including r2, Q2, IIC, and CII ranging from 0.6877 to 0.7820, 0.6622 to 0.7643, 0.5638 to 0.8324, and 0.8248 to 0.8548, respectively. Based on these models, important molecular descriptors influencing the enhancement or reduction of larvicidal activity were determined. Furthermore, molecular docking studies were performed to investigate the binding interactions and conformational poses of the selected compounds.
Malignant diseases are considered the most prominent and widespread causes of death affecting populations globally. Synergistic drug combinations have shown beneficial therapeutic results in the treatment of malignant diseases. Although techniques such as clinical trials and high-throughput drug screening are commonly used to discover promising synergistic drug pairs, they are time-consuming and expensive. Over the past years, various AI-based drug synergy techniques including machine learning and deep learning have been utilized in finding synergistic drug combinations. Individually using these methods for synergy prediction has the disadvantages of overfitting and lack of interpretability. Combining different AI methods through ensemble learning provides better predictions by more closely representing the underlying distribution of data. This study utilized the heterogeneous stacking ensemble approach (HTeSyn) by aggregating four machine learning methods as base learners and one neural method as meta-learner.This multi-faceted approach helps in correcting classification results and provides more reliable synergy predictions, which is crucial for identifying effective drug combinations. For the bliss independence synergy task, HTeSyn outperforms the state-of-the-art synergy prediction method with an accuracy of 94%, RMSE of 12.5, and r2 of 0.8.
The histamine H3 receptor (H3R) is a GPCR that regulates the release of multiple neurotransmitters and has emerged as an attractive target for CNS disorders. An integrated computational workflow was applied to identify H3R ligands from natural products by combining pharmacophore modelling, QSAR, docking and molecular dynamics (MD) algorithms. Structure-based and ligand-based pharmacophore models were developed and validated using ROC analysis against a set of active H3R inhibitors and DUD-E decoys. The resulting models (PHARM-1 and PHARM-2) achieved AUC values of 0.716 and 0.959, respectively, supporting their ability to discriminate active from inactive compounds. Both models were used to screen AnalytiCon Discovery database of natural products, yielding 23 hits that were subsequently prioritized by docking and QSAR model generated using genetic function approximation. Six candidates docked successfully and preserved the expected anchoring interaction with the conserved Asp114; four compounds showed favourable QSAR-predicted affinities (pKi ≈ 5.01-7.06) and high consensus docking scores. Finally, MD simulations of H3R complexed with the two promising hits, physostigmine and catharanthine, indicate stable complexes, and receptor fluctuations reflecting ligand-dependent conformational sampling. These results highlight several natural products as promising H3R ligand candidates and provide a practical multi-filter pipeline for prioritizing hits for experimental validation.
The growing energy demand has accelerated the search for renewable energy sources, with dye-sensitized solar cells (DSSCs) emerging as promising candidates. To streamline the experimental process and reduce associated costs, we developed predictive models based on linear regression to predict the efficiency of DSSCs. These models, use quantum molecular descriptors (QMDs) derived from molecular electronics structure and excited states of the dyes. Our study focused on evaluating organic dyes based in imidazole, BODIPY, and squaraine for DSSC applications. The resulting linear models were simple, robust, and predictive, satisfying all standard validation metrics. Furthermore, each model's descriptors reveal key electronic/spectroscopic characteristics for efficiency enhancement, including: (i) increased molecular mass through branching, (ii) HOMO-LUMO gap control, and (iii) planar π-bridge optimization.
Water solubility is an important factor in environmental and toxicological science because it determines the mobility, bioavailability, and potential for absorption by living organisms. Higher solubility often correlates with greater environmental mobility, and toxicity is often inversely related to solubility. In this work, the solubility of 4453 compounds in water was examined using quantitative structure-property relationships (QSPR). The effectiveness of the Monte Carlo method, the correlation intensity index (CII), and the Las Vegas algorithm for developing organic compound solubility models is assessed using CORAL software. Both factors (CII and the Las Vegas algorithm) have been shown to improve the statistical quality of the model set for the calibration and validation sets. The average coefficient of determination for the validation set is 0.94 ± 0.01.
Accurate uncertainty quantification is a prerequisite for reliable toxicity assessments in drug discovery. Traditional QSAR models provide point estimates but fail to communicate prediction reliability, particularly for structurally complex compounds. We propose Conformalized Two-Stage Heteroscedastic BART (C2S-HBART), a novel framework addressing homoscedastic modelling assumptions. A first-stage BART model estimates the mean toxicity response; a second stage explicitly estimates local predictive variance from SMILES-derived structural descriptors. Split Conformal Prediction is then integrated to provide distribution-free validity guarantees. Evaluated on the Tox21 dataset, C2S-HBART achieves near-nominal coverage (0.952 vs. 0.880 for the uncalibrated baseline) and reduces the rate of silent failures - confident but incorrect predictions - from 12.02% to 1.8%. Compared to a standard conformalized baseline requiring wide intervals (Avg Width: 1.63), the heteroscedastic approach achieves equivalent safety with sharper predictions (Avg Width: 1.20), representing a 26% improvement in information efficiency. Variable importance analysis further reveals that molecular size and topological complexity are primary drivers of predictive uncertainty. C2S-HBART provides a statistically rigorous, transparent decision-support tool for preclinical screening, enabling toxicologists to prioritize safer compounds and flag structurally complex molecules for experimental validation.
Bisphenol A (BPA) substitutes are increasingly prevalent in consumer products, yet their potential to induce oligoasthenospermia (OAS) remains poorly understood. This study combined network toxicology, molecular docking, and 500-ns molecular dynamics (MD) simulations to systematically compare the toxicity mechanisms of BPA, BPS, BPF, and BPAF in the context of OAS. Intersection analysis identified 25 shared targets, with topological clustering highlighting ESR1, AR, CYP19A1, TNF, and IL6 as central hubs linking endocrine disruption with inflammatory pathways. Molecular docking revealed broad receptor engagement (-5.4 to -8.7 kcal/mol), with BPAF exhibiting the strongest affinities, particularly for ESR1 (-8.7 kcal/mol) and CYP19A1 (-8.4 kcal/mol). Molecular Dynamics (MD) simulations confirmed the dynamic stability of the ESR1-BPAF complex, demonstrating a compact fold (SASA ~124 nm2), minimal structural drift (ligand RMSD ~0.16 nm; protein RMSD ~0.20 nm), and persistent hydrogen bonding anchored by residues GLU353 and LEU428. These computational findings predict that BPA alternatives - especially BPAF - may significantly perturb steroidogenic and inflammatory axes to drive spermatogenic impairment. This study provides a predictive hazard prioritization framework that challenges the safety assumption of 'BPA-free' substitutes and warrants urgent experimental validation.
The rapid advancement of ML algorithms and the increasing availability of large datasets have significantly transformed the landscape of predictive modelling in scientific research. In this context, we introduce Optima (OPTimized Interpretable Model Building & Analysis Toolkit), a user-friendly, Python-based GUI designed to simplify and accelerate the development of interpretable ML-based classification QSAR/QSPR/QSTR models. This toolkit offers an intuitive graphical user interface (GUI), enabling users with domain knowledge but limited coding experience to efficiently optimize, construct, and interpret various ML-based classification models. A most highlighting feature of this GUI is its fully customizable settings panel, allowing users to modify colour schemes, font sizes, axis labels, and plot dimensions to suit publication or presentation needs. By combining robust optimization with an explainable approach, the Optima toolkit improves classification QSAR model performance while ensuring transparency and reproducibility. This platform addresses a critical need by providing an intuitive GUI for rational dataset splitting, efficient feature selection, and the development of seven different ML-based classification QSAR models, covering the entire workflow from optimization to interpretation. The toolkit is presently available for Windows and can be downloaded from the provided link (https://github.com/Rahul-Roy-21/OPTIMA).
Accurate prediction of acetylcholinesterase (AChE) inhibitory activity is important in drug discovery and environmental toxicology because AChE inhibition represents a key mechanism underlying neurotoxicity associated with pharmaceuticals and environmental contaminants. In this study, machine learning approaches were used to develop predictive models for AChE inhibitory activity using experimentally measured bioactivity data for small molecules targeting human AChE. A curated dataset containing 5795 molecules was compiled from BindingDB to support reliable model development. Fifteen predictive models were evaluated, including twelve individual machine learning and deep learning models and three hybrid fusion models, using multiple molecular representations such as physicochemical descriptors derived from RDKit and PaDEL and graph-based molecular structures. Among the individual models, tree-based ensemble methods demonstrated strong baseline performance, indicating that physicochemical descriptors capture important chemical features associated with AChE inhibition. Graph neural networks, particularly Graph Isomorphism Network effectively learn structural patterns related to inhibitory activity. To integrate complementary molecular information, a late-fusion hybrid framework combining descriptor-based predictions and graph-based representations was implemented using leakage-safe stacking with a Ridge regression meta-learner. Across ten independent train-test splits, the best-performing hybrid model integrating PaDEL-based XGBoost and GIN achieved r2 = 0.7400 ± 0.0138, demonstrating improved and stable predictive performance over individual models.
Predicting acetylcholinesterase (AChE) inhibitory activity is important in drug discovery. This study evaluates molecular descriptor - based machine learning models to predict AChE activity as pIC50 values. The primary objective was to comparatively investigate the impact of different data preprocessing strategies on prediction performance and model selection under challenging chemical datasets exhibiting low correlation structures. Tree based gradient boosting algorithms, namely CatBoost and XGBoost, together with sensitive regression models including Support Vector Regression and Multilayer Perceptron, were examined, and model specific data preparation pipelines were applied according to their structural assumptions. The target variable was stabilized through logarithmic transformation and winsorization of IC50 values. Model performance was assessed using both a 70-15-15 train-validation-test split and a 10-fold cross validation protocol. Furthermore, stacking based ensemble learning strategies were explored to enhance generalization capability. The results demonstrate that predictive performance is predominantly constrained by intrinsic dataset characteristics rather than algorithmic selection. Optimized tree-based models achieved the highest accuracy, while stacking provided only marginal improvements over the best individual learners. To improve interpretability, SHAP based explainable artificial intelligence analysis was conducted, highlighting the contributions of biologically meaningful molecular descriptors, and offers guidance for future studies addressing comparable biochemical modelling challenges.
Epilepsy is a chronic neurological disorder characterized by recurrent seizures resulting from abnormal neuronal excitability and ion channel dysfunction. This study explored the antiepileptic potential of Jatropha integerrima using integrated in vivo and in silico approaches to identify safer alternatives to conventional therapies. Ethanolic extracts of the plant were fractionated into petroleum ether (PE) and ethyl acetate (EA) fractions and analyzed via LC-MS for tentative phytoconstituent identification. Anticonvulsant activity of both fractions was evaluated in Swiss albino mice using pentylenetetrazol (PTZ) and maximal electroshock seizure (MES) models. The PE fraction significantly reduced seizure frequency, duration, and mortality (p < 0.05), demonstrating effects comparable to diazepam. Molecular docking and MM/GBSA analyses revealed strong binding affinities of major compounds toward GABAA and Nav1.2 receptors. Notably, (-)-jatrointelignan A (C₃₁H₃₆O₁₁) exhibited the highest docking scores (-10.20 kcal/mol for GABAA and -11.47 kcal/mol for Nav1.2) and favorable binding free energies (-25.38 and -70.63 kcal/mol, respectively). ADME and toxicity predictions suggested good pharmacokinetic properties with low toxicity, while molecular dynamics simulations confirmed stable receptor ligand interactions. Overall, J. integerrima and its lead compound demonstrate promising antiepileptic potential via dual modulation of GABAergic and sodium channel pathways.
Tuberculosis affects 10 million people around the world, for this reason, the development of new drugs against M. tuberculosis (Mtb) is an urgent issue. Enoyl ACP reductase (InhA) is a relevant biological target because it's the role in the synthesis of mycolic acid, the building block of the cell wall. Here, we propose new anti-tuberculosis candidates based on a family of tested InhA pyrrolidine-carboxamides. The strategy starts building QSAR models using, for the first time, a set of topological-descriptors from QTAIM (LDMtrace, Ndeloc, Ntotal, and Max_eigenv), hydrogen bond energy (HB) of Docking simulations, and standard QSAR molecular-descriptors (MATS4m, VE1_Dzp, and RDF70m). The model with the best performance displays robust external and internal validation parameters. Finally, based on the most reliable QSAR model, a set of new molecules were proposed; six of them (p69, p71, p73, p75, p77, and p78) show an efficient calculated IC50 and adequate ADMET properties, suggesting an enhanced biological activity. Molecular Dynamics and MM/PBSA results show that p71 is the most reliable candidate for the inhibition of InhA, displaying a stable complex p71-InhA along 200 ns of simulation and a negative free binding energy, ∆Gbinding < 0.