Despite continuing hype about the role of AI in drug discovery, no "AI-discovered drugs" have so far received regulatory approval. Here we assess one of the latest AI-based tools in this domain. Boltz-2, a recently developed biomolecular foundation model, aims to bridge the gap between AI efficiency and physics-based precision through a joint "cofolding" approach. In this study, we provide an extensive evaluation of Boltz-2 using two large-scale data sets: 16780 compounds for 3CLPro and 21702 compounds for TNKS2. We compare Boltz-2 predicted structures with traditional docking and binding affinities with binding free energies derived from the physics-based ESMACS protocol. Structural analysis reveals significant global RMSD variations, indicating that Boltz-2 predicts multiple protein conformations and ligand binding positions rather than a single converged pose. Energetic evaluations exhibit only weak to moderate correlations across the global data sets. Furthermore, a focused analysis of the top 100 compounds yields no significant correlation between the Boltz-2 predictions and the binding free energies from fine-grained ESMACS, alongside frequently observed saturation-state errors in Boltz-2 predicted ligand structures. Our results show that Boltz-2 lacks the energetic resolution required for lead identification. These findings highlight the necessity of employing physics-based methods for the reliability and refinement of AI-derived models.
The work of Martin Karplus, who passed away on Dec 28 2024, was at the forefront of computational chemistry and molecular biophysics for a period of more than sixty years. His career started with a PhD in theoretical chemistry at Caltech with Linus Pauling in 1953 . After performing leading research in molecular quantum chemistry for two decades, in the 1970s he began to incorporate work on biological systems, and his work was instrumental in creating modern computational molecular biophysics. In 2013, he was awarded the Nobel Prize together with Arieh Warshel and Michael Levitt “for the development of multiscale models for complex chemical systems”. This article aims at briefly reviewing the main achievements of the research performed in the Karplus lab from the point of view of the people working under his unique mentorship.
We present a generative active learning (GAL) framework for molecular design that integrates the generative AI platform REINVENT with physics-based free-energy estimation via ESMACS, specifically addressing synthetic tractability. In a previous study we have shown that iterative generation and optimization of molecules for protein binding affinity can be effectively achieved. For drug discovery, however, molecules need to be synthesized for experimental validation. Molecules with unclear synthetic routes will be de-prioritized regardless of their predicted binding affinity. Here, we address this by incorporating additional synthesizability constraints into the computational design process, framing the problem as a multi-objective optimization task. We show that it is possible to simultaneously optimize candidate molecules for both binding affinity and synthetic accessibility. Our results show that integrating synthesizability prediction into physics-based GAL workflows enables the efficient design of compounds that are chemically diverse and predicted to be strong binders and synthetically tractable at the same time, demonstrating efficient and practical computational drug design.
The drug discovery process remains very slow and unreliable, requiring about 10 years and $2-3 billion for a single new drug at a success rate of around 5%. To address this, we have recently created a complex and highly heterogeneous workflow called IMPECCABLE which is designed to transform much of the drug discovery process into a far more reliable in silico environment. It runs on Frontier, the world’s first exascale machine and has recently been implemented on the second such machine, Aurora. A variety of in silico discovery methods exist, including physics-based and artificial intelligence (AI) approaches. Despite their limitations, these methods are complementary, and we combine them here within IMPECCABLE, an integrated framework which efficiently and reliably explores the vast chemical space of small molecule binders. IMPECCABLE involves screening, generating and evaluating candidate compounds. The environment is comprised of several interlocking and interacting stages. In the first of these, an exceptionally fast machine learning model acts as a docking surrogate for rapidly screening billions of compounds and provides a ranked series of binding poses for the best ones. In a generative multi-objective active learning loop, the accurate binding free energies from physics-based free energy methods, including ESMACS and TIES, guide REINVENT to predict compounds not only with higher binding affinities, but which are also synthesizable while ensuring diversity and novelty. The full iterative workflow spans both hit identification and hit-to-lead optimization and includes medicinal chemists in the process. Applying IMPECCABLE to two disease-related protein targets, WDR91 and TNKS2, we demonstrate its capability to iteratively produce experimentally makeable potent binders. The potential hits discovered were further optimized and chiral purified via an iterative REINVENT-TIES cycle. We identified a potent hit for WDR91, which was synthesised in five steps, exhibiting nanomolar binding affinity with a KD value of 352 ± 32 nM.
This chapter presents the probabilistic formulation of molecular dynamics, offering a comprehensive overview of the underlying theories and their implications. It begins with an introduction to probabilistic dynamical systems, setting the stage for a detailed discussion on ergodic theory and its relevance to molecular dynamics. The chapter establishes the connection with kinetic theory, followed by an exploration of unstable periodic orbits and the Ruelle zeta function, highlighting their significance in understanding molecular behaviour. The implications of these probabilistic approaches for classical molecular dynamics are thoroughly examined, providing insights into how they enhance the accuracy and precision of simulations. The chapter concludes with a summary of key findings and their impact on the field of molecular dynamics.
Retinoids, such as all-trans retinoic acid (ATRA), are the active metabolite forms of endogenous Vitamin A and function as key signaling molecules involved in the regulation of a variety of cellular processes. Due to their highly diverse biological roles, retinoids have been implicated in a wide range of diseases such as neurological disorders and some cancers. However, their therapeutic potential is limited due to their chemical and metabolic instability and adverse side effects. Synthetic retinoid analogues with increased stability and specificity have therefore attracted significant attention. In this study, we developed a scalable synthetic platform to generate a library of novel synthetic retinoids. Twenty-three new compounds were synthesized, and their receptor binding was assessed by an in vitro fluorescence competition binding assay, complemented by molecular docking and molecular dynamics (MD) simulations. We show that while computational studies are extremely useful for predicting binding modes and hence can guide synthetic efforts, the binding assays demonstrated that these novel retinoids exhibit strong binding albeit with limited selectivity for the different retinoic acid receptors (RARs). Therefore, their biological activity was measured by assessing their genomic and nongenomic activities in neuroblastoma cells with the goal of correlating binding properties and pathway activation to neuro-regenerative potential measured by neurite outgrowth. Importantly, four of the novel retinoids are shown to bind tightly to RARs and exhibit dual action in the relevant cellular models, with an ability to induce both genomic and nongenomic responses as well as significant neurite outgrowth. The compound with the highest biological activity possesses significant potential to be used as therapeutics for treating a wide range of neurological disorders like Alzheimer's disease and motor neuron disease.
Active learning (AL) is a specific instance of sequential experimental design and uses machine learning to intelligently choose the next data point or batch of molecular structures to be evaluated. In this sense it closely mimics the iterative design-make-test-analysis cycle of laboratory experiments to find optimized compounds for a given design task. Here we describe an AL protocol which combines generative molecular AI, using REINVENT, and physics-based absolute binding free energy molecular dynamics simulation, using ESMACS, to discover new ligands for two different target proteins, 3CLpro and TNKS2. We have deployed our generative active learning (GAL) protocol on Frontier, the world’s only exa-scale machine. We show that the protocol can find better binders compared to baseline, a surrogate ML docking model for 3CLpro and compounds with experimentally determined binding affinities for TNKS2. The ligands found are also chemically diverse and occupy a different chemical space than the baseline. We vary the batch sizes that are put forward for free energy assessment in each GAL cycle to assess the impact on their efficiency on the GAL protocol and recommend their optimal values in different scenarios. Overall, we demonstrate a powerful capability of the combination of physics-based and AI methods which yields effective chemical space sampling at an unprecedented scale, of immediate and direct relevance for modern, data-driven drug discovery.
The advent of hybrid computing platforms consisting of quantum processing units integrated with conventional high-performance computing brings new opportunities for algorithms design. By strategically offloading select portions of the workload to classical hardware where tractable, we may broaden the applicability of quantum computation in the near term. In this perspective, we review techniques that facilitate the study of subdomains of chemical systems with quantum computers and present a proof-of-concept demonstration of quantum-selected configuration interaction deployed within a multiscale/multiphysics simulation workflow leveraging classical molecular dynamics, projection-based embedding and qubit subspace tools. This allows the technology to be utilised for simulating systems of real scientific and industrial interest, which not only brings true quantum utility closer to realisation but is also relevant as we look forward to the fault-tolerant regime.
As one of the deadliest infectious diseases in the world, tuberculosis is responsible for millions of new cases and deaths reported annually. The rise of drug-resistant tuberculosis, particularly resistance to first-line treatments like rifampicin, presents a critical challenge for global health, which complicates the treatment strategies and calls for effective diagnostic and predictive tools. In this study, we apply an ensemble-based molecular dynamics computer simulation method, TIES_PM, to estimate the binding affinity through free energy calculations and predict rifampicin resistance in RNA polymerase. By analyzing 61 mutations, including those in the rifampicin resistance-determining region, TIES_PM produces reliable results in good agreement with clinical reference and identifies abnormal data points indicating alternative mechanisms of resistance. In the future, TIES_PM is capable of identifying and selecting leads with a lower risk of resistance evolution and, for smaller proteins, it may systematically predict antibiotic resistance by analyzing all possible codon permutations. Moreover, its flexibility allows for extending predictions to other first-line drugs and drug-resistant diseases. TIES_PM provides a rapid, accurate, low-cost, and scalable supplement to current diagnostic pipelines, particularly for drug resistance screening in both research and clinical domains.IMPORTANCEAntimicrobial resistance (AMR), a global threat, challenges early diagnosis and treatment of tuberculosis (TB). This study employs TIES_PM, a free-energy calculation method, to efficiently predict AMR by quantifying how mutations in bacterial RNA polymerase (RNAP) affect rifampicin (RIF) binding. On simulating 61 clinically observed mutations, the results align with WHO classifications and reveal ambiguous cases, suggesting alternative resistance mechanisms. Each mutation requires ~5 h, offering rapid, cost-effective predictions. An ensemble approach ensures statistical robustness. TIES_PM can be extended to smaller proteins for systematic codon permutation analysis, enabling comprehensive antibiotic resistance prediction, or adapted to identify low-resistance-risk drug leads. It also applies to other TB drugs and resistant pathogens, supporting personalized therapy and global AMR surveillance. This work provides novel tools to refine resistance mutation databases and phenotypic classification standards, enhancing early diagnosis while advancing translational research and infectious disease control.
Free energy calculations for protein-ligand complexes have become widespread in recent years owing to several conceptual, methodological and technological advances. Central among these is the use of ensemble methods which permits accurate, precise and reproducible predictions and are necessary for uncertainty quantification. Absolute binding free energies (ABFEs) are challenging to predict using alchemical methods and their routine application in drug discovery has remained out of reach until now. Here, we apply ensemble alchemical ABFE methods to a large dataset comprising 219 ligand-protein complexes and obtain statistically robust results with high accuracy (< 1 kcal/mol). We compare equilibrium and non-equilibrium methods for ABFE predictions at large scale and provide a systematic critical assessment of each method. The equilibrium method is more accurate, precise, faster, computationally more cost-effective and requires a much simpler protocol, making it preferable for large scale and blind applications. We find that the calculated free energy distributions are non-normal and discuss the consequences. We recommend a definitive protocol to perform ABFE calculations optimally. Using this protocol, it is possible to perform thousands of ABFE calculations within a few hours on modern exascale machines.
Active learning (AL) is a specific instance of sequential experimental design and uses machine learning to intelligently choose the next data point or batch of molecular structures to be evaluated. In this sense, it closely mimics the iterative design-make-test-analysis cycle of laboratory experiments to find optimized compounds for a given design task. Here, we describe an AL protocol which combines generative molecular AI, using REINVENT, and physics-based absolute binding free energy molecular dynamics simulation, using ESMACS, to discover new ligands for two different target proteins, 3CL(pro) and TNKS2. We have deployed our generative active learning (GAL) protocol on Frontier, the world's only exa-scale machine. We show that the protocol can find higher-scoring molecules compared to the baseline, a surrogate ML docking model for 3CL(pro) and compounds with experimentally determined binding affinities for TNKS2. The ligands found are also chemically diverse and occupy a different chemical space than the baseline. We vary the batch sizes that are put forward for free energy assessment in each GAL cycle to assess the impact on their efficiency on the GAL protocol and recommend their optimal values in different scenarios. Overall, we demonstrate a powerful capability of the combination of physics-based and AI methods which yields effective chemical space sampling at an unprecedented scale and is of immediate and direct relevance to modern, data-driven drug discovery.
The domain of computational biomedicine is a new and burgeoning one. Its areas of concern cover all scales of human biology, physiology, and pathology, commonly referred to as medicine, from the genomic to the whole human and beyond, including epidemiology and population health. Computational biomedicine aims to provide high-fidelity descriptions and predictions of the behavior of biomedical systems of both fundamental scientific and clinical importance. Digital twins and virtual humans aim to reproduce the extremely accurate duplicate of real-world human beings in cyberspace, which can be used to make highly accurate predictions that take complicated conditions into account. When that can be done reliably enough for the predictions to be actionable, such an approach will make an impact in the pharmaceutical industry by reducing or even replacing the extremely laboratory-intensive preclinical process of making and testing compounds in laboratories, and in clinical applications by assisting clinicians to make diagnostic and treatment decisions.
Uncertainty quantification (UQ) is rapidly becoming a sine qua non for all forms of computational science out of which actionable outcomes are anticipated. Much of the microscopic world of atoms and molecules has remained immune to these developments but due to the fundamental problems of reproducibility and reliability, it is essential that practitioners pay attention to the issues concerned. Here a UQ study is undertaken of classical molecular dynamics with a particular focus on uncertainties in the high-dimensional force-field parameters, which affect key quantities of interest, including material properties and binding free energy predictions in drug discovery and personalized medicine. Using scalable UQ methods based on active subspaces that invoke machine learning and Gaussian processes, the sensitivity of the input parameters is ranked. Our analyses reveal that the prediction uncertainty is dominated by a small number of the hundreds of interaction potential parameters within the force fields employed. This ranking highlights what forms of interaction control the prediction uncertainty and enables systematic improvements to be made in future optimizations of such parameters.
PDF file - 98K, The binding positions of gefitinib and ATP in the binding site of EGFR
Relative binding free energy (RBFE) calculations are widely used to aid the process of drug discovery. TIES, Thermodynamic Integration with Enhanced Sampling, is a dual-topology approach to RBFE calculations with support for NAMD and OpenMM molecular dynamics engines. The software has been thoroughly validated on publicly available datasets. Here we describe the open source software along with a web portal (https://ccs-ties.org) that enables users to perform such calculations correctly and rapidly.
Significantly more ‘outliers’ can be produced from a non-Gaussian distribution than one would anticipate were the statistics to conform to a normal distribution. Using ensemble simulations consisting of 25 replicas, we have previously identified a considerable percentage of ligand-protein systems which present non-Gaussian distributions in calculated binding free energies. Here we report on the statistics of much larger ensemble and find that the free energy distributions are definitively non-Gaussian for these systems.
It is increasingly widely recognized that ensemble-based approaches are required to achieve reliability, accuracy, and precision in molecular dynamics calculations. The purpose of the present article is to address a frequently raised question: what is the optimal way to perform ensemble simulation to calculate quantities of interest?