
Thermoplastic starch (TPS) has emerged as a biodegradable alternative to conventional plastics. However, the literature indicates that its functional performance varies with its composition and processing conditions, hindering efficient material design with traditional trial-and-error approaches. To address this need, this study develops machine learning-based predictive models to estimate the mechanical, thermal, permeability, barrier, and biodegradability properties of TPS as a function of formulation and processing variables. A dataset of 98 variables and 962 observations, obtained from previous studies, was constructed for this purpose. The CRISP-ML(Q) methodology was followed, and three supervised algorithms were implemented: Random Forest (RF), Extreme Gradient Boosting (XGBoost), and Artificial Neural Networks (ANN). Model performance was evaluated using hold-out validation (70:20:10) and standard regression metrics. The results showed that XGBoost is the algorithm with the best overall predictive ability. The highest levels of accuracy were achieved for WVP (R2 = 0.981) and Td (R2 = 0.939). The main predictive limitations were observed in some ANN models, especially with the variable’s Tensile strength, Young's modulus, and T50, which reported negative or near-zero R2 values. The developed models constitute a practical decision-support tool, allowing researchers and non-specialist users to virtually evaluate different combinations of formulation and processing conditions before conducting laboratory tests, thus reducing the number of experiments during the initial stages of new material development.
Recent advances in large language models (LLMs) have opened new avenues for automating scientific research workflows. Systems like the AI Scientist [1] demonstrate fully autonomous endto-end research pipelines for machine learning, but their applicability to domains requiring expensive high-performance computing (HPC) resources, specialized quantum chemistry software, and rigorous thermochemical accuracy remains unexplored. Here we present VirtualLab_CC, a hybrid LLM-augmented virtual laboratory for computational chemistry that integrates Claude (Anthropic) as an orchestration layer with a multi-agent architecture, gated lifecycle management, and bidirectional HPC cluster synchronization. Unlike fully autonomous systems, VirtualLab_CC enforces human-in-the-loop quality gates at every critical decision point while automating routine operations: conformer searches, input generation, job submission and monitoring, output parsing, failure diagnosis, and manuscript preparation. The system manages 21 specialized skills spanning literature review, cluster operations, result extraction, novelty verification, AI-assisted peer review, and citation validation. We demonstrate VirtualLab_CC on a prototype case study of 1,3-dipolar cycloadditions between stabilized Criegee intermediates and biogenic terpene dipolarophiles, involving 17 molecular species, 34 conformers, and over 100 SLURM jobs across multiple calculation stages. The system successfully identified and autocorrected cluster failures, detected a methodological anomaly (a suspiciously low dioxirane barrier requiring IRC verification), and coordinated multi-method validation across three DFT functionals. VirtualLab_CC bridges the gap between fully autonomous AI research and the rigorous quality control demands of computational chemistry, providing a practical framework for LLM-augmented scientific workflows in resource-constrained HPC environments.
Accurate prediction and generation of reaction temperature remain open challenges in data-driven chemical synthesis, owing to the continuous nature heterogeneous experimental protocols, and the need for strict physical plausibility. To address this, we introduce the Coherent Guided Generative Model (CGGM), a hybrid deep-learning architecture that integrates variational inference, adversarial learning, and smooth quadratic physical constraints to map reaction temperatures directly from molecular descriptors. Benchmarked on a rigorously curated dataset of 554,974 reactions from the Open Reaction Database (ORD), the CGGM exposes a fundamental precision-generalization trade-off in chemical AI. While a standalone Variational Autoencoder (VAE) minimizes pointwise error (MAE = 0.5105) through conservative mean-regression, the hybrid CGGM achieves superior global density estimation, yielding the highest log-likelihood (-5.2621) and accurately capturing rare cryogenic and high-temperature regimes. Crucially, the CGGM enforces thermodynamic boundaries, restricting out-of-range predictions to just 0.05%, nearly identical to the curated empirical ground truth (0.04%) and significantly outperforming a standalone GAN (0.16%). Latent space topology and manifold traversal analyses confirm a smooth, continuous embedding space with monotonic temperature encoding (rho > 0.99, p < 0.001), allowing seamless interpolation between cryogenic (approximate to-78 degrees C and reflux (>100 degrees C) domains without structural degradation. Finally, stratified evaluations across Suzuki, Buchwald-Hartwig, and esterification reactions highlight that while classical regressors are optimal for static pointwise lookups, the primary merit of the CGGM lies in its exceptional distributional fidelity, tail-risk modeling, and physically bounded generation. This hybrid framework establishes a robust, uncertainty-aware computational engine for continuous chemical space exploration and closed-loop autonomous synthesis planning.
Quantitative simulation of trivalent f-block chelates in water remains challenging because bonded and non-bonded force-field models make different approximations for coordination structure, exchange dynamics, and ion–ligand interactions in highly charged systems. Here, we develop a hybrid machine-learning/molecular-mechanics (ML/MM) framework for Ac3+–DOTA in explicit solvent by training an E(3)-equivariant neural network potential (MACELES) on mechanically embedded QM/MM data for Ac aquo and Ac–DOTA species and coupling it to NAMD 2.14 with particle-mesh Ewald electrostatics. Nanosecond ML/MM trajectories remain numerically stable and preserve chelate integrity, yielding a compact DOTA inner shell with an inner-sphere water coordination number of CNAc,Ow≈1.7 arising from a dynamic equilibrium between one- and two-water states (37.5% and 59.9% of frames; three waters 2.5%). A 5 ns potential of mean force shows two low-lying basins at CNAc,Ow≈1 and CNAc,Ow≈2. DFT end-state free energies are consistent with the ML/MM profile, and DFT minimum-energy paths provide a qualitative electronic-structure reference for the observed basin connectivity. State-resolved kinetics reveal picosecond water-exchange pathways that couple hydration changes to transient DOTA arm fluctuations, and training-set comparisons show that temperature-matched Ac–DOTA data optimize energy/force accuracy while more diverse solvated data improve charge prediction. Overall, the present hybrid ML/MM model provides a practical description of Ac3+–DOTA hydration thermodynamics and short-time exchange behavior in explicit water at MD-like cost.
Identifying explosive molecules is crucial for safety and materials research. In this work, we propose a machine learning framework for classifying 119 molecules using images of their electrostatic potential maps (MEPs), computed from quantum-chemical methods and depicting the molecules' positive and negative electric charge regions. The training dataset was augmented by rotating each image 30 times, and a Convolutional Neural Network (CNN) was subsequently trained to classify molecules. The CNN classification model was combined with advanced gradient-based Explainable Artificial Intelligence (XAI) techniques and visualization methods. The CNN achieved 100% accuracy on the test set, demonstrating its effectiveness. In a separate experiment, it achieved a 81.25% accuracy using molecules chemically distinct from the training set: the CNN model classified eight as explosive, including two known simulants and two non-explosive compounds containing strong electron-withdrawing groups, thus indicating sensitivity to chemically relevant features not included in the training distribution. To gain a deeper understanding of the model’s decision-making process, we used the Saliency Map technique with the reduce_max function to identify the most critical regions in the MEP images for the classification task. Neutral and homogeneous MEP regions primarily characterize non-explosive molecules. In contrast, explosive compounds are associated with pronounced electrostatic potential gradients and localized charge features, particularly in regions related to nitro groups and trigger bonds. This work highlights the potential of combining MEPs computed from quantum-chemical methods with machine learning to analyze and classify molecules with targeted properties.
We present an automated pipeline for extracting scalar coupling constants J and chemical shifts from strongly coupled 1H NMR spectra of ABC and ABCD spin systems. It operates on an isolated peak list of the target spin system without requiring initial parameters or manual transition assignment. It combines a neural-network (NN) stage with a quantum-mechanical refinement. The two spin-system classes are handled by two architectures tailored to their respective inverse problems: an ABC network built on a Set Transformer backbone and an ABCD network designed to be topology-agnostic through a coupling-conditioned shift-prediction scheme. On synthetic and experimentally acquired 1H spectra at 60, 80 and 500 MHz, the ABC pipeline achieves mean sub-0.1 Hz accuracy on both J and chemical shifts across the experimental benchmark. The ABCD network, trained on a topology-diverse synthetic data set spanning four structurally distinct molecular families, reaches MAE(J) ti 0.13 Hz and MAE(nu) ti 0.33 Hz on a held-out synthetic evaluation subset comprising the validation and test partitions. On an ABCD validation set comprising experimental and DFT-based spectra, the full pipeline further achieves MAE(J) ti 0.08 Hz and MAE(nu) ti 0.30 Hz without requiring topology information. These accuracies are below, or well within, typical 1H NMR linewidths at all three field strengths considered, and the entire pipeline runs in tens of milliseconds per spectrum, combining sub-linewidth precision with throughput suitable for routine automated processing of large experimental datasets.
The vast compositional space of ABX3 perovskites presents a dual bottleneck for materials discovery: prohibitive screening costs and inefficient hyperparameter optimization (HPO) required to develop accurate predictive models. To address this computational challenge, we introduce a systematic machine learning framework that rigorously benchmarks six meta-heuristic algorithms (MHAs) for HPO, systematically quantifying the critical trade-off between prediction accuracy and computational cost. The optimized models achieve competitive predictive performance, with test-set R2 values exceeding 0.9653 for formation energy, 0.9196 for energy above hull, and 0.8669 for bandgap, alongside a thermodynamic stability classification accuracy of 0.8646. Our analysis reveals that the Sooty Tern Optimization Algorithm (STOA) and Particle Swarm Optimization (PSO) show a favorable balance, confirming that the ideal MHA choice is task-specific—a key finding for efficient model development. Beyond prediction, the interpretable framework uses Shapley analysis to uncover core physicochemical drivers, revealing, for instance, that stability is strongly influenced by range electronegativity and space group, while band gaps are governed by formation energy and transition metal fraction. This accelerated, insight-driven framework culminated in the high-throughput screening of 3864 compounds, identifying 54 stable, lead-free candidates for photovoltaics (PV) and 51 for photoelectrochemical (PEC) applications.
This study investigates the atomic-scale structure of a glassy system using first-principles molecular dynamics (FPMD) and machine-learned interatomic potentials (MLIP). The structural models generated using MLIP demonstrate excellent agreement with FPMD results while offering substantial computational efficiency, enabling the exploration of much larger system sizes that are prohibitive for FPMD while retaining comparable accuracy. Validation against experimental X-ray total structure factors and pair distribution functions indicates the reliability of both the FPMD- and MLIP-generated structures. The glass network is found to be composed of a mix of VO4, VO5 and VO6 polyhedra, TeO3, TeO4 and TeO5 units. Local environments are studied through maximally localized Wannier function (MLWF)-based bonding analysis, coordination number analysis, and bond angle distributions. The MLIP model is further used to study glass compositions with varying amounts of V5+ and V4+ to investigate oxidation-state-dependent structural changes. Our results demonstrate that melt-quench glass models produced through MLIP-based molecular dynamics and subsequently relaxed via a single-shot DFT calculation can capture changes in local topology, coordination, and V–O bonding motifs with high accuracy, revealing the reliability of the MLIP model in describing mixed-valence oxide glass systems. This finding can be explained by charge self-regulation in transition-metal oxides, where changes in formal oxidation state are accommodated by local metal–O electronic redistribution and accompanying structural rearrangements. This work highlights the capability of MLIPs to accurately describe complex oxide glasses with multicomponent compositions and variable oxidation states, extending their applicability to chemically diverse disordered systems.
Achieving true transferability remains the central challenge for Machine Learning Interatomic Potentials (ML-IAPs) in modeling complex bimetallic nanoclusters across their vast potential energy surfaces. We systematically investigate data selection strategies to optimize the Chebyshev Interaction Model for Efficient Simulation (ChIMES) potential for the Bi-Pt nanoclusters by comparing three innovative sampling methods: Principal Component Analysis (PCA)/k-means (structural diversity), t-distributed Stochastic Neighbor Embedding (t-SNE)/k-means (force-space diversity), and hierarchical clustering. Quantitatively, the PCA/k-means strategy proved most effective for global accuracy, yielding the lowest force errors and achieving energy root mean square errors (RMSE) values competitive with Density Functional Theory (DFT), demonstrating excellent accuracy (19.16 meV/atom). Structural validation on 34 unique DFT-optimized isomers further confirmed the potential's high fidelity, with the best model PCA/k-means reproducing structures with an average root mean square deviation (RMSD) of 0.10 & Aring;. However, the t-SNE methods, by maximizing diversity in the force space, demonstrated superior extrapolative power, leading to the more precise prediction of a novel stellated octadecagon Bi18Pt24 cage structure, demonstrating the potential for exploring previously unseen morphologies. Our results establish a clear methodology for strategic data sampling that successfully maximizes ML-IAP transferability, providing an accurate and computationally efficient tool that accelerates the theoretical discovery of complex bimetallic architectures.
Machine learning force fields trained on the fly on ab initio molecular dynamics simulations using density functional theory are developed to explore the evolution of five selected morphologies of Pd147 nanoparticles (approximate to 1.6 nm) at finite temperature. At 373 K, the simulations performed from the trained force fields predict no morphological change, with a preference for the strongly irregular truncated octahedral and the defective Marks-decahedral nanoclusters, in agreement with geometry optimizations. At 600 K, three isomers among five undergo morphological changes which could not have been guessed without machine learning force field simulations of one nanosecond, implying millions of structures, with the ab initio accuracy. A geometric descriptor based on the average metal-metal distance along the trajectories captures these changes quite relevantly and illustrates the rapid transformation of fcc and hcp nanoclusters, while the fcc stacking-faulted isomer undergoes a very progressive transformation into a defective icosahedral structure. The robustness of these trained machine learning force fields at a target temperature of 485 K was tested by either performing simulations at a higher temperature of 600 K (extrapolation) or crossing them between selected morphologies (transferability). The extrapolation with respect to temperature led to plausible results whereas the transferability between morphologies raised the question of the predicted relative stability order, although the obtained structures were correct. This study opens promising perspectives concerning multiple crossed training of machine learning force fields and molecular dynamics simulations beyond one nanosecond with the ab initio precision for palladium nanoparticles.
Antimicrobial peptides (AMPs) emerge as a type of promising therapeutic compounds that exhibit broad spectrum antimicrobial activity with high specificity and good tolerability. However, current AI-based AMP design strategies, which primarily rely on learning the distribution of natural AMPs, fail to overcome the inherent trade-off between antimicrobial activity and toxicity, thereby hindering their clinical translation. In this work, we propose PepGen-FB, a novel multi-objective optimization method for optimizing desired properties for AMPs iteratively. It employs a curriculum learning-guided feedback mechanism to iteratively guide the generative model to smoothly optimize AMPs with improving antibacterial activity and decreasing toxicity. Comprehensive experiments demonstrate that the AMPs designed by PepGen-FB substantially outperform natural prototypes in achieving an optimal balance between high antimicrobial activity and low cytotoxicity, improved the generation success rate from 7.1% to 96.2% compared to the ProGen2 model. Further motif analyses provide interpretative support for the optimization process. PepGen-FB enables the seamless integration of arbitrary black-box predictors while ensuring optimization stability through curriculum-guided feedback, which establishes a novel data-driven optimization paradigm.
Montmorillonite plays a central role in environmental barrier systems because its hydrated interlayers, high cation-exchange capacity (CEC), and structural flexibility strongly influence heavy-metal retention. However, predictive evaluation across varying humidity histories, cation identities, and solution chemistries remains limited when based solely on empirical trends. In this study, experimentally established X-ray diffraction (XRD) observations from the literature-describing discrete hydration states (0 W/1 W/2 W), mixed-layer interstratification, and the cation-dependent classification of swelling behaviour-are incorporated as structural constraints within a physics-informed machine-learning (ML) framework. A curated dataset (N = 600), augmented with descriptors reflecting hydration energy, ionic radius, layer-charge proxies, CEC, pH, and ionic strength, is used to train gradient-boosting and ensemble models. Model performance on held-out data reaches R2 approximate to 0.68, which is consistent with nonlinear interactions between hydration-controlled descriptors. Molecular-dynamics (MD) simulations are employed to validate predicted basal-spacing trends and to capture the & Aring;-scale interlayer reorganizations-such as heterogeneous water-sheet distributions and cation-specific coordination environments-that underlie continuous swelling beyond discrete hydration states. By aligning ML attributions with established structural behavior from XRD studies and validating them against MD-resolved microscopic configurations, the resulting framework provides a rapid, interpretable, and physically consistent approach. Physical knowledge is incorporated through structure-aware feature design and class-consistent priors rather than through explicit governing equations, making the approach best described as physics-guided machine learning for predicting heavy-metal uptake by montmorillonite and guiding the design of robust clay-based barrier materials.
Heavy metal contamination in water is a major environmental challenge due to its toxicity, persistence, and potential for bioaccumulation in living organisms. Among the various treatment technologies, adsorption using activated carbon remains one of the most effective and economical methods for removing metal ions such as Cu(II), Zn(II), Ni(II), Pb(II), Cd(II), Cr(VI), and As(V). However, exploring and predicting adsorption performance through conventional experimentation is often time-consuming and resource-intensive. In this study, a machine learning based predictive framework was developed to estimate the amount of metal adsorbed at equilibrium (Qe, in mg g-1) on activated carbon. A large dataset comprising 1,528 experimental records was assembled, incorporating key adsorbent textural parameters (BET surface area, pore diameter, pore volume), operational variables (pH, temperature, initial concentration, dosage, contact time), and intrinsic ionic properties (hydrated radius, van der Waals radius, molar mass, electronegativity). Four supervised learning algorithms were implemented and evaluated: Random Forest Regressor (RF), Extra Trees regressor (ET), Gradient Boosting Regressor (GBR), and Extreme Gradient Boosting Regressor (XGBR). Among them, the Gradient Boosting regressor showed the best predictive performance on the test set (RMSE = 4.035, MAE = 2.205, R2 = 0.9648). Beyond prediction, the proposed machine learning approach enables the identification of complex, nonlinear relationships governing adsorption behavior and highlights the relative importance of key physicochemical parameters. It therefore represents a relevant complementary tool to experimental studies for improving the design of water treatment systems.
Quantum-Aided Drug Design (QuADD) is a platform that utilizes quantum computing to formulate molecular design as a multi-objective optimization problem, enabling the generation of drug-like molecules optimized for interactions within a defined binding pocket. Generative artificial intelligence has similarly emerged as a strategy for exploring chemical space and designing novel molecular structures. The Bond and Interaction generating Diffusion model (BInD), an AI-based application, applies reverse diffusion techniques to produce structurally diverse candidate molecules. In this work, we present a controlled comparison between these two structure-based molecular generation systems. Both approaches produced novel molecules compatible with the evaluated binding site. While BInD generated candidates exhibiting greater structural diversity, QuADD-generated molecules more consistently satisfied prioritization criteria related to predicted binding affinity, drug-likeness, and preservation of key protein–ligand interactions. These findings suggest that constraint-driven optimization can provide advantages in generating synthetically plausible, pocket-aware candidate molecules suitable for early-stage lead discovery, while probabilistic generative strategies may promote broader structural exploration.
High-entropy alloys (HEAs) present significant challenges for property prediction due to their vast compositional freedom and limited reliable data. Without tractable governing equations, elastic properties have been rationalized using physically motivated descriptors such as elastic bounds and empirical parameters. While incorporating such constraints improves robustness, predictive performance remains sensitive to how multiple constraints are weighted.Here, we propose a physics-guided neural network integrating four complementary constraints: Voigt and Reuss elastic bounds from classical elasticity theory, valence electron concentration (VEC) for phase stability, and atomic size mismatch (δ) reflecting lattice distortion. These are incorporated as soft regularization terms, automatically balanced via Lagrangian-dual adaptive weighting.The framework was evaluated on 1,117 HEAs with bulk moduli ranging from 12.5 to 428.7GPa. Four-fold cross-validation demonstrated robust convergence with coefficients of variation below 7% for all constraint weights. Compared with an unconstrained neural network (MAE: 5.583GPa, R²: 0.8657) and a fixed-weight physics-constrained model (MAE: 5.240GPa, R²: 0.8526), the proposed framework achieved best performance (MAE: 4.573GPa, R²: 0.9002), corresponding to MAE reductions of 18.1% and 12.7%, respectively.These results demonstrate that combining physically motivated constraints with automatic λ tuning enhances extrapolation performance while maintaining consistency with fundamental mechanical bounds. The approach provides a physically interpretable and data-efficient framework for accelerating materials design in high-dimensional alloy systems under limited data conditions.
Viral infections remain a major global health challenge, with current antiviral therapies often limited by drug resistance and high development costs. Antiviral peptides (AVPs) are promising alternatives to traditional antiviral drugs owing to their safety and broad-spectrum activity. Most existing AVP prediction classifiers rely solely on sequence-derived features while neglecting three-dimensional structural information, which limits their generalization ability under highly imbalanced virtual screening conditions. To overcome these limitations, we propose a graph-based deep learning framework that explicitly integrates residue-level three-dimensional structural information with sequence semantics for AVPs identification. Residue-level graphs are constructed from ESMFold-predicted structures and enriched with embeddings from the ESMC protein language model, enabling the model to capture both spatially proximal and sequentially distant interactions that are inaccessible to sequence-only approaches. These graphs are processed using a graph attention network with multiscale pooling to learn structure-aware representations. Evaluated on a large, imbalanced, and independent test set, our model demonstrates substantially improved robustness to class imbalance and structural variability, outperforming state-of-the-art sequence-based predictors. Notably, the proposed framework reduces false positives by 54% relative to Stack-AVP, improves the Matthews correlation coefficient by 29%, and achieves an accuracy of 84.1% and specificity of 84.8%. Furthermore, Grad-CAM-based interpretability analysis provides residue-level mechanistic insights, highlighting structurally and functionally relevant amino acids driving antiviral activity. By unifying sequence semantics with explicit structural constraints, this work advances AVPs prediction beyond sequence-only paradigms and provides a practical, interpretable tool for antiviral peptide discovery under realistic, imbalanced conditions.
Carbon capture, utilization, and storage (CCUS) technologies are now being developed to meet the global net-zero emissions target. Due to their high surface area and high porosity, metal-organic frameworks (MOFs) are promising solid sorbent candidates for post-combustion carbon capture. Recent studies now use machine learning (ML) to accelerate the high-throughput screening of MOFs. However, most studies only rely on supervised learning to do structure-to-property predictions and offer little understanding about the MOF chemical space. In this paper, we aim to provide a more interpretable ML workflow for finding best-performing MOFs for carbon capture using manifold learning and Bayesian optimization. We posed an optimization problem whose objective is to find MOFs from the CoRE MOF 2019 database that minimize the required energy when simulated in a pressure-swing adsorber, subject to having high CO2 purity and high CO2/N2 selectivity. Different from existing literature, we introduce uniform manifold approximation and projection (UMAP) to first embed the pore geometry descriptors data in a 2-D manifold subspace, hence visualizing the MOF chemical space. We then explored the possibility of performing Bayesian optimization in the UMAP subspace to search MOFs more efficiently. Our results include a 2-D mapping of MOFs, a ranking of best MOFs, and an analysis of regions in the geometric descriptor subspace that led to best process performance. We hope that these contributions can help increase our understanding of the relationships between the material-level and process-level metrics of MOFs for carbon capture.
Developing high-performance Energetic Materials (EMs) with small environmental footprint remains a critical challenge for both civilian and military applications. The conventional EMs such as Research Department Explosive (RDX) and trinitrotoluene (TNT) exhibit strong performance, but they tend to release toxic byproducts, posing significant environmental risk. In this work, we introduce a quantum-chemistry informed Bayesian Optimization (BO) for accelerated discovery of novel pyrazole-based EMs. The employed BO and Quantum chemistry calculations to identify pyrazole derivatives with high potential for EMs applications. Using a library of 350 pyrazole-based EMs generated through systematic enumeration of pyrazole scaffold with various explosophoric groups, and their computed density values. BO was employed to rapidly identify pyrazole derivatives with optimal density values in the chemical space within only few iterations. The energetic potential of the BO-selected pyrazole derivatives was further ascertained using density functional theory (DFT) calculations at the B3LYP/6-311G (d, p) level of theory. The generalizability of the BO framework to rapidly identify pyrazole derivatives with high energetic potential was further validated using an expanded library of 1500 pyrazole derivatives. The top five candidates identified by the BO algorithm demonstrates impressive energetic potential with DFT computed detonation parameters comparable or exceeding those of benchmarked explosives. The DFT computed detonation parameters of the BO-selected pyrazole derivatives including crystalline density (1.93-2.25gcm-3), heats of formation (712.4-876.2kJmol-1), detonation velocities (8.78-9.22kms-1), and detonation pressure (18.60-25.11GPa) were found to be favorable for EMs applications. This study establishes a data-driven workflow for rapid EMs discovery and highlights pyrazole scaffolds as promising platforms for safer, greener EMs, with direct relevance for civilian and military applications.
The computational repurposing of existing drugs has proven to be a fast-track strategy in the development of new cancer therapy. Nonetheless, integrated mechanistically grounded approaches are required for a more reliable in silico identification of new drug-target pairs. In this study, we introduce an integrated workflow combining machine learning with molecular docking, kinase selectivity profiling, molecular dynamics simulations, and ADMET assessment to systematically repurpose FDA-approved drugs as inhibitors of the oncogenic kinase PI3K alpha. A robust quantitative structure-activity relationship (QSAR) model (test set R2 = 0.825) prioritized candidates from a library of 2458 approved drugs. Through this pipeline, three promising candidates (Vemurafenib, Fedratinib, and Zafirlukast) were identified, each exhibiting stable interactions with PI3K alpha, favorable binding free energies, and targeted polypharmacology rather than promiscuous inhibition. Molecular dynamics simulations confirmed that ligand binding reduces protein flexibility and confines the conformational landscape. These results provide a multi-layered computational rationale for repurposing these FDA-approved drugs as PI3K alpha inhibitors. Overall, this hybrid approach illustrates how combining data-driven and physics-based methods can enhance the precision of computational drug repurposing, effectively transforming the large library of approved drugs into a tractable source for novel targeted cancer therapies.
Generative AI and deep learning improve molecular simulations and drug development. Traditional computational methods like MD, MC, and QM/MM have been crucial in investigating biomolecular interactions and thermodynamics. However, processing power and speed restrict their scalability. This article provides a comprehensive review and comparative analysis of how advanced neural network architectures and generative AI models address these computational limitations. This review analyses how advanced neural network architectures and generative AI models satisfy these restrictions. Neural network potentials trained on high-quality quantum datasets achieve ab initio precision at low processing cost. We tested convolutional (CNNs), recurrent (RNNs), graph neural networks (GNNs), and transformers to evaluate how well they could describe molecular changes over time and predict structural changes. Researchers have investigated generative frameworks including variational autoencoders (VAEs), generative adversarial networks (GANs), and diffusion models to develop medications with superior binding affinity and pharmacokinetic characteristics. The findings reveal that AI-driven modelling and physics-based simulations create a closed-loop system where MD or QM/MM simulations enhance AI-generated molecules repeatedly. This feedback loop speeds up hit-to-lead optimisation, increases ADMET prediction, and enhances protein folding and shape information. This paradigm shift from descriptive to predictive and generative frameworks using AI and molecular modelling improves computational drug discovery's scalability, interpretability, and creativity. AI is used as a computational tool and a collaborator to speed up molecular discovery. Overall, this manuscript serves as a critical review summarizing state-of-the-art progress, challenges, and future prospects at the interface of AI and molecular simulation research.