Deep learning models are data-hungry, and synthetic (artificial) data has been shown to be invaluable when data availability is low. While this has been demonstrated in certain technology areas, adopting such an approach is new in machine learning (ML) applications in chemistry, except for some pre-training tasks. In drug discovery, predicting binding energy between proteins and ligands is crucial. Many ML-based studies have been proposed to predict protein-ligand binding affinity using existing experimental data. However, these models suffer from inherent biases. Recent efforts have produced PLAS-20k, a synthetic dataset of multiple protein-ligand complex (PLC) conformations generated using molecular dynamics (MD) simulations as a viable option to complement existing experimental data and improve binding affinity prediction. For the binding affinity prediction task, we employ Pafnucy, a deep convolutional neural network, and propose using multiple structures for each PLC from PLAS-20k for training. We compare four different statistical and ML-based result-aggregation techniques. This work demonstrates the utility of dynamic datasets in enhancing binding affinity predictions, laying the foundation for future improvements in predicting similar protein properties using synthetic datasets and more sophisticated models and methods. We propose that physics-based synthetic datasets can significantly help develop more accurate data-driven methods. Scientific contribution This study shows that incorporating synthetic molecular dynamics data improves deep learning models for protein–ligand binding affinity prediction beyond static experimental structures. By systematically evaluating frame selection and prediction aggregation strategies, we demonstrate that training on diverse conformational snapshots significantly enhances generalization and accuracy. Our results highlight that dynamic synthetic datasets can enable deep learning models to outperform conventional methods such as MM-PBSA while remaining computationally efficient.
Antimicrobial peptides (AMPs) are emerging as potent alternatives to conventional antibiotics, yet their diverse nature due to divergent mechanisms of action hinders rational design. Here, we present an electrostatics-stratified computational framework that uncovers key physicochemical principles governing AMP activity. Experimentally validated peptides were grouped by average charge per residue (i.e., the charge/length of the peptide) and analyzed through integrated sequence-, structure-, and chemistry-based descriptors. Distinct molecular signatures emerged across electrostatic regimes: low-charge/length peptides rely on amphipathic organization via structural compactness, whereas the intermediate-charge/length peptides exhibit balanced hydrophobicity and electrostatics. The high-charge peptides couple strong cationic attraction with lipophilicity and tryptophan anchoring to mainly disrupt membranes. Interestingly, hydrophobic moment, which is a measure of the amphipathicity, is found to be important in all three classes of AMPs. This study identifies distinguishing features of AMP sub-groups and suggests design guidelines for developing selective and potent next-generation AMPs.
Gold nanoclusters have been found to catalyze chemically significant reactions like oxidation of greenhouse gases and selective oxidation of alcohols. Further incorporation of Cu atoms in Au alloys enhances their catalytic activity and selectivity. Experimental results have demonstrated different rates of benzyl alcohol conversion to benzaldehyde relative to different ratios of Au and Cu in the catalyst nanoparticles. The rates of conversion demonstrate higher catalytic activity of pure Au clusters compared to pure Cu clusters. In this work, the catalytic roles of Au and Cu in the aerobic oxidation of benzyl alcohol to benzaldehyde were analyzed using density functional theory (DFT) methods. Au-rich and Cu-rich clusters were considered as catalysts, namely Au13, Au12Cu and Cu13, Cu12Au clusters, respectively. The mechanism of the minimum energy path focuses on the oxidation of two molecules of benzyl alcohol by a single molecule of O2 to produce two molecules of benzaldehyde and H2O each. The calculations reveal that the reaction rate is predominantly governed by the adsorption of reactants and the desorption of products, consistent with experimental observations. In Cu-rich clusters, the adsorption of O2 is more favorable due to effective charge transfer between the cluster and O2 resulting in high adsorption energy. Additionally, the desorption of the products is significantly more endothermic on Cu-rich catalysts than on Au-rich catalysts.
Drug-Drug Interactions (DDIs) often cause significant adverse events, making early identification crucial, especially with rising polypharmacy. While most methods focus on binary prediction, multiclass DDI prediction is more advantageous as it identifies specific interaction types. However, this is challenging due to class imbalance, with many rare interaction types compared to dominant ones. We propose a multiclass DDI prediction framework comprising three components: Stage 1 involves feature extraction, where the molecular structures of two drugs are modeled as graphs using a graph attention network (GAT) as an encoder, then combined with Drug Repurposing Knowledge Graph (DRKG) embeddings to produce enhanced drug features. Stage 2 involves expert modelling, where the paired drug representations are processed by two Mamba State Space Model (SSM)- based experts: a standard Mamba expert and a class-reweighted Mamba expert, which efficiently capture sequential dependencies and interaction patterns between the drug representations. Stage 3 involves gated fusion, where a confidence-based gating strategy integrates the two experts to generate final predictions. The performance on DrugBank and DDIMDL benchmark datasets achieved macro F1 scores of 90.68 https://github.com/devalab/MambaDDI .
Exploring promiscuous catalytic activity of enzymes to stereoselective transformations uncovers ecofriendly asymmetric catalysts and provides scope to execute new-to-nature chemistry. Biocatalytic promiscuous diastereoselective Henry reaction (DHR) is restricted to a few hydroxynitrile lyases (HNLs), often engineered, with narrow synthetic scope and anti-selectivity. Here, we report a single native enzyme exhibiting complementary diastereoselectivity in an asymmetric Henry reaction. Baliospermum montanum HNL (BmHNL)-catalyzed DHR enabled the production of thirty-two anti-(1S,2R)-beta-nitroalcohols with up to >99% conversion, >99% ee, and >99% de using longer nitroalkanes. It afforded syn diastereomers, i.e., (1S,2S)-beta-nitroalcohols as thermodynamically controlled products by manipulating the reaction conditions. Reuse of BmHNL in DHR for 20 cycles has empowered the synthesis of (1S,2R)-2-nitro-1-phenylpropan-1-ol (NPP), where it retained 82% of its initial activity, maintained >99% ee on each cycle, and provided total turnover number >3.7 x 10(4), which is 140-fold higher compared to the only existing biocatalytic DHR. Our study demonstrates a gram-scale synthesis of (1S,2R)-NPP and a preparative-scale synthesis of (1S,2S)-NPP, a precursor to d-norpseudoephedrine used for the treatment of obesity. The origin of the inverse diastereoselectivity was probed using isotope labeling studies. Molecular modeling of the reaction using density functional theory methods reveals the underlying mechanism of the diastereoselectivity and the competing kinetic vs thermodynamic nature of the reaction. Despite a promiscuous catalytic activity, the outstanding catalytic efficiency, broad synthetic scope, and complementary diastereoselectivity of the native enzyme of BmHNL on DHR are remarkable, which illustrates it as a specialized biocatalyst for greener technology and stereocontrolled production of diverse beta-nitroalcohols.
Spectroscopy is the study of how matter interacts with electromagnetic radiations of specific frequencies that has led to several monumental discoveries in science. The spectra of any particular molecule is highly information-rich, yet the inverse relation from the spectra to the molecular structure is still an unsolved problem. Nuclear Magnetic Resonance (NMR) spectroscopy is one such critical tool in the tool-set for scientists to characterise any chemical sample. In this work, a novel framework is proposed that attempts to solve this inverse problem by navigating the chemical space to find the correct structure that resulted in the target spectra. The proposed framework uses a combination of online Monte- Carlo-Tree-Search (MCTS) and a set of offline trained Graph Convolution Networks to build a molecule iteratively from scratch. Our method is able to predict the correct structure of the molecule ∼80% of the time in its top 3 guesses. We believe that the proposed framework is a significant step in solving the inverse design problem of NMR spectra to molecule.
Self-driving or autonomous labs are platforms that integrate artificial intelligence (AI)/machine learning (ML), robotics and high-throughput experiments, and are capable of designing, synthesizing, testing and optimizing materials for specific purposes with minimal human intervention. Crucial components are algorithms and methods that are capable of generating new materials with desired properties. A recent study by Zeni et al. has reported a material generative model, MatterGen that is fine-tuned to propose new materials conditioned on user-specific mechanical, electronic and magnetic properties. Fully autonomous chemistry laboratories bring together advanced generative AI algorithms, robotic systems and high-throughput experiments that can computationally design, and experimentally synthesize and characterize molecules/materials. In this news story, we discuss the importance of methods such as recently proposed MatterGen in making self-driving laboratories a reality.
AuCu nanoclusters have widespread application in reactions like activation of , selective oxidation, and cross‐coupling reactions. In this study, we investigate the stepwise doping of copper atoms in a pure 13‐atom gold cluster, denoted as ( m + n = 13). The genetic algorithm based on the artificial bee colony algorithm has been utilized to model various isomers of each composition. The potential energy landscape of these clusters was analyzed by means of the density functional theory method with pure Perdew−Burke−Ernzerhof (PBE) functional. We identify the minimum energy isomer for each cluster composition to evaluate molecular properties like HOMO–LUMO gap, binding energy/atom, second order difference in energy, vertical ionization energy, and vertical electron affinity. Notably, the introduction of copper atoms in these clusters enhances their stability and reactivity. Distinct odd–even oscillations due to close shell electronic configurations are absent, as all cluster compositions have an overall open shell configuration. To assess the catalytic activity of the clusters, we study the adsorption energies of small molecules like and on all available sites on the cluster. This study thereby comprehensively explores the range of copper‐doped 13‐atom gold cluster compositions and their implications on their structure–property relationships vital for catalysis and nanomaterial applications.
Chemical reaction prediction, encompassing forward synthesis and retrosynthesis, stands as a fundamental challenge in organic synthesis. A widely adopted computational approach frames synthesis prediction as a sequence-to-sequence translation task, using the commonly used SMILES representation for molecules. The current evaluation of machine learning methods for retrosynthesis assumes perfect training data, overlooking imperfections in reaction equations in popular datasets, such as missing reactants, products, other physical and practical constraints such as temperature and cost, primarily due to a focus on the target molecule. This limitation leads to an incomplete representation of viable synthetic routes, especially when multiple sets of reactants can yield a given desired product. In response to these shortcomings, this study examines the prevailing evaluation methods and introduces comprehensive metrics designed to address imperfections in the dataset. Our novel metrics not only assess absolute accuracy by comparing predicted outputs with ground truth but also introduce a nuanced evaluation approach. We provide scores for partial correctness and compute adjusted accuracy through graph matching, acknowledging the inherent complexities of retrosynthetic pathways. Additionally, we explore the impact of small molecular augmentations while preserving chemical properties and employ similarity matching to enhance the assessment of prediction quality. We introduce SynFormer, a sequence-to-sequence model tailored for SMILES representation. It incorporates architectural enhancements to the original transformer, effectively tackling the challenges of chemical reaction prediction. SynFormer achieves a Top-1 accuracy of 53.2% on the USPTO-50k dataset, matching the performance of widely accepted models like Chemformer, but with greater efficiency by eliminating the need for pre-training.
Determining spin state energy gaps (SSE) of 3d transition metal complexes (TMCs) is a major challenge in theoretical chemistry, as high-level quantum methods, though reliable, are computationally impractical for large-scale studies. This work explores a machine learning (ML)-based approach to predict DFT adiabatic SSE gaps using descriptors derived from a single high-spin DFT calculation. This approach is adopted to eliminate the differential treatment of electronic correlation between high-spin and low-spin structures. Our descriptors aim to incorporate the knowledge of crystal field theory into the ML model. They include atomic energy levels of bare metal ions, natural charges of ligating atoms, d-orbital molecular orbital eigenvalues derived from an high spin calculation, HOMO-LUMO gaps of free ligands, and simple identity-based features. We train ML models on 1434 SSE values spanning 934 complexes and demonstrate their transferability to more challenging complexes having bidentate π-bonding ligands despite being trained on simpler Werner-type monodentate complexes. We achieved a minimum MAE of 4.0 kcal mol-1 on the monodentate test set, and maintained a comparable MAE of 6.6 kcal mol-1 in the transferability assessment. This approach bypasses the need for multi-reference low-spin optimizations while retaining predictive accuracy, offering a cost-effective strategy for SSE estimation in transition metal chemistry. We hope the insights covered in this study will contribute to the development of additional electronic structure-based descriptors for SSE predictions.
Recent progress and development of artificial intelligence and machine learning (AI/ML) techniques have enabled addressing complex biomolecular problems. AI/ML models learn the underlying distribution of data they are trained on and when exposed to new inputs, they make predictions based on patterns and relationships previously observed in the training set. Further, generative artificial intelligence (GenAI) can be used to accurately generate protein structure or sequence from specific selected properties. This review specifically focuses on the applications of AI/ML in predicting important functional properties of proteins, and the potential prospects of reverseengineering in depicting the sequence and structure, from available protein-property information.
Intrinsically disordered proteins (IDPs), unlike globular proteins, lack stable secondary structure and exist as dynamic ensembles of conformations in physiological conditions. These conformations allow them to adapt to many roles while interacting with other proteins, forming partially folded soluble oligomers or insoluble plaques rich in β -sheet, and are responsible for different pathological diseases. The dynamic nature of IDPs and the lack of well-defined binding sites make drug discovery and development both challenging. Although computational protocols, predominantly involving molecular dynamics (MD) simulations studies, have been complementing in understanding the structure and dynamics of IDPs, the choice of force fields become important in reproducing experimental observations. We herewith provide a systematic study to investigate the structural propensity of four IDPs (A β 42, Tau43, amylin and α S), using 13 combinations of force field-water models (FF-wm). In addition to validating with NMR observables, we also examine the water dynamics surrounding the chosen IDPs to underscore the generalizability of chosen FF-wms. C36-IDPSFF with amber99sb-dispersion water model (A99SB-disp) is found to perform optimal in simulating IDPs in comparison with other FF-wm combination based on structural propensity and comparison with experimental NMR 3 J HN−Hα -Coupling data as well as dynamical observables. ### Competing Interest Statement The authors have declared no competing interest. DST-SERB, CRG/2021/008036 Ihub-Data, NA
Molecular Property Diagnostic Suite (MPDS) was conceived and developed as an open-source disease-specific web portal based on Galaxy. MPDSCOVID-19 was developed for COVID-19 as a one-stop solution for drug discovery research. Galaxy platforms enable the creation of customized workflows connecting various modules in the web server. The architecture of MPDSCOVID-19 effectively employs Galaxy v22.04 features, which are ported on CentOS 7.8 and Python 3.7. MPDSCOVID-19 provides significant updates and the addition of several new tools updated after six years. Tools developed by our group in Perl/Python and open-source tools are collated and integrated into MPDSCOVID-19 using XML scripts. Our MPDS suite aims to facilitate transparent and open innovation. This approach significantly helps bring inclusiveness in the community while promoting free access and participation in software development. Availability & Implementation The MPDSCOVID-19 portal can be accessed at https://mpds.neist.res.in:8085/.
Inferring complete molecular structure from infrared (IR) spectra is a challenging task. In this work, we propose SMEN (Spectra and Molecule Encoder Network), a framework for scoring molecules against given IR spectra. The proposed framework uses contrastive optimization to obtain similar embedding for a molecule and its spectra. For this study, we consider the QM9 dataset with molecules consisting of less than 9 heavy atoms and obtain simulated spectra. Using the proposed method, we can rank the molecules using embedding similarity and obtain a Top 1 accuracy of similar to 81%, Top 3 accuracy of similar to 96%, and Top 10 accuracy of similar to 99% on the evaluation set. We extend SMEN to build a generative transformer for a direct molecule prediction from IR spectra. The proposed method can significantly help molecule library ranking tasks and aid the problem of inferring molecular structures from spectra.
In recent years, the rapid advancement of generative artificial intelligence (GenAI) has revolutionized the landscape of drug design, offering innovative solutions to potentially expedite the discovery of novel therapeutics. GenAI encompasses algorithms and models that autonomously create new data, including text, images, and molecules, often mirroring characteristics of existing datasets. This comprehensive review delves into the realm of GenAI for drug design, emphasizing recent advancements and methodologies that have propelled the field forward. Specifically, we focus on three prominent paradigms: transformers, diffusion models, and reinforcement learning algorithms, which have been exceptionally impactful in the last few years. By synthesizing insights from a myriad of studies and developments, we elucidate the potential of these approaches in accelerating the drug discovery process. Through a detailed analysis, we explore the current state and future directions of GenAI in the context of drug design, highlighting its transformative impact on pharmaceutical research and development.
Drug-Drug Interactions (DDI) can trigger unexpected pharmacological consequences, including adverse drug events (ADE). The rise in polypharmacy underscores the importance of understanding how different drug molecules influence each other's pharmacological activities and necessitates the investigation of potential interactions between newly developed drugs and existing medications. Traditional laboratory methods for DDI detection are time-consuming, making the development of computational prediction methods crucial. This study introduces GraphDDI, a machine learning method utilizing Graph Neural Networks (GNN) to predict DDI accurately. The proposed methodology, trained end-to-end, comprises three stages. (1) Featurization stage: A GNN extracts atomic features from two drugs separately. (2) Interaction stage: an interaction map is calculated between all atom pairs of the drugs. (3) Prediction stage: the model combines the interaction map and drug features to create a unified representation of the drug molecules. Subsequently, the model concatenates these representations and employs a feed-forward neural network to predict the DDI. We demonstrate the efficacy of our proposed model in predicting both the presence of DDI and the specific types of interactions (DDI events). Comparative analysis reveals that our framework surpasses existing models, achieving an F1 score of 0.98 in predicting the existence of drug-drug interactions and 0.90 in categorizing DDI event types. The code is available in our GitHub repository (https://github.com/devalab/GraphDDI).
Pure and doped gold clusters have been of immense use in catalyzing reactions and assembling nano functional materials for various applications. In this work, we focus on stepwise doping of copper atoms in pure 13 atom gold clusters, thereby the cluster composition investigated is Aum Cun (m+n = 13). We employ the genetic algorithm ABCluster which uses the artificial bee colony (ABC) algorithm to model the various isomers of each cluster composition. We applied DFT functional PBE and LANL2DZ basis functions to model the potential energy surface of the clusters. The minimum energy isomer of each composition was then used to study various molecular properties like binding energy, second order difference in energy, vertical ionization energy, vertical electron affinity, HOMO-LUMO gap, second order difference in energy. Odd-even oscillations in the molecular properties reveal the competing shell closing stabilization between Cu and Au atoms. To compare the activity of the clusters in catalysis, adsorption studies of small molecules O2 and C2H2 were done. This work aims to study the entire range of Cu doped 13 atom Au cluster compositions.