Protein flexibility is central to function—governing allosteric regulation, catalysis, signal transduction, and drug binding. Yet, this flexibility is often underrepresented in structural models derived from X-ray or Cryo-EM experiments, where flexible loops, alternate protonation states, and mobile side chains are frequently unresolved or omitted. These omissions limit the accuracy of molecular dynamics (MD), free energy calculations, and mechanistic insights. In this talk, we present a density-driven, quantum-informed protocol for reconstructing disordered and flexible regions of macromolecular structures, enabling more complete and chemically accurate models suitable for dynamic simulation and predictive design.Implemented in the DivCon Discovery Suite, our approach rebuilds missing backbone loops using φ/ψ sampling with geometry and density-aware filtering. Real-space refinement is applied at each stage—backbone closure, rotamer placement, and protonation state selection—using QM, MM, or hybrid QM/MM Hamiltonians combined with X-ray or Cryo-EM electron density gradients. Z-score of the difference density (ZDD) is used throughout to evaluate and select the most experimentally consistent conformations. The workflow also includes tautomer/protomer determination and hydrogen-bond network-aware protonation state sampling.We demonstrate how this method restores conformational diversity essential for understanding loop gating, ligand-induced rearrangements, and proton-coupled mechanisms. Case studies reveal improved structural stability during simulations, stronger correlation between ligand and loop motion in MD, and enhanced accuracy in binding free energy predictions. These refinements not only support more faithful molecular “movies” of protein function but also improve the quality of AI/ML training datasets used for structure and property prediction.This work highlights the importance of recovering conformationally dynamic features from density data to enable flexible, functionally relevant, and therapeutically actionable protein models.
Conventional structural refinement relies on fixed stereochemical restraints for all components within the protein-ligand complex, using pre- determined parameters such as bond lengths, angles, and torsions in the unbound ligand conformation. This approach can be challenging for non-standard modalities like covalent modifiers and macrocycles, where conformational flexibility is inherent to their mechanism of action, and the bound vs. unbound states differ significantly. To address this limitation and achieve a more complete model, we employ an automated molecular refinement approach that replaces stereochemical restraint gradients with quantum-mechanic/molecular mechanic (QM/MM) gradients in standard refinement packages, Phenix and Buster. In this context, we present user test cases in QM/MM structure refinement using QuantumBio’s DivCon plugin with XModeScore to explore potential ligand states and evaluate binding modes based on ligand strain energy and the z-score of difference density.
With the steady, decades-long advance of faster, smarter computers, structure-based drug discovery (SBDD) and computer-aided drug design (CADD) have become indispensable tools for pharmaceutical research. Today, most pharmaceutical companies and research laboratories (academic as well as industrial) employ target-ligand structures to help inform the drug discovery effort. At the outset of any drug discovery campaign, if one or more experimental models are available, they are generally reviewed and scrutinized to determine what protein–ligand interactions are critical and what interactions could be better optimized in to build the next generation of compounds that will be safer and more effective than the last. At the same time, the tools that have become integral to pharmaceutical research – docking, scoring, structure prediction, dynamics, free energy modeling, etc. – are constantly evolving to increase speed and improve accuracy. Finally, recent advances in machine learning and artificial intelligence have led to advances in structure prediction based on the availability of pertinent models in the protein data bank (PDB). At the core of all this research and development is the availability of high-quality and highly accurate experimental models for these research efforts: garbage-in/garbage-out is a real and persistent problem as we look to next-generation cures and tools. X-ray crystallography and, to a lesser extent, nuclear magnetic resonance (NMR) have been the workhorses of structure solution. And with recent advances in cryogenic electron microscopy (Cryo-EM), we can now solve structures that previously would have been difficult to solve. But these experimental protocols all require significant computation themselves in the form of structure optimization (refinement) in the presence of experimental restraints (density). In the present work, we summarize recent advances in the field through the introduction of higher, more accurate levels of theory, and we discuss how these methods impact our understanding of the structure and how we can use this better understanding to inform both the SBDD process and the CADD tools at our disposal.
X -ray crystallography and cryogenic electron microscopy (cryo-EM) are the primary experimental techniques used to determine the three-dimensional (3D) structure of protein:ligand and protein:protein complexes, and these methods play a central role in Structure Based Drug Design (SBDD). To overcome typical weaknesses of the conventional X - ray refinement due to using stereochemical - restraints we incorporated advanced QM or QM/MM functionals into the crystal structure refinement [1, 2]. Later we also demonstrated the critical role of the QM/MM refinement for the affinity predictions in SBDD [3]. During the model building stage of X -ray and Cryo-EM refinement, a significant challenge in ligand model placement is determination of the correct ligand orientation especially when the experimental density is weak or incomplete. We have addressed this ligand placement and refinement problem with the use of our MovableType fast free energy based d ocking algorithm (MTDock) [4] coupled with a built-in realspace (RS) refinement engine in which geometry energy and gradients - derived using a QM, MM or mixed QM/MM Hamiltonian - are combined with electron density values/gradients derived from X -ray data / Cryo-EM maps [1, 2]. In this work, we will present this integrated RS-MTDock protocol which involves automated electron density blob selection followed by MT - driven ligand placement and RS refinement. Each refined pose is then scored using MT Score to determine which pose is correct. For the maximum performance, these steps have been fully integrated within QuantumBio's Discovery Suite to provide user - friendly, fully automated and fast software to directly use in any SBDD protocol. We have validated the RS-MTDock protocol using 480 protein:ligand structures taken from the Iridium and PDBBind sets. We demonstrated top scored ligand poses have RMSD to the published structures below 0.3 Å
Fast and accurate biomolecular free energy estimation has been a significant interest for decades, and with recent advances in computer hardware, interest in new method development in this field has even grown. Thorough configurational state sampling using molecular dynamics (MD) simulations has long been applied to the estimation of the free energy change corresponding to the receptor-ligand complexing process. However, performing large-scale simulation is still a computational burden for the high-throughput hit screening. Among molecular modeling tools, docking and scoring methods are widely used during the early stages of the drug discovery process in that they can rapidly generate discrete receptor-ligand binding modes and their individual binding affinities. Unfortunately, the lack of thorough conformational sampling in docking and scoring protocols leads to difficulty discovering global minimum binding modes on a complicated energy landscape. The Movable Type (MT) method is a novel absolute binding free energy approach which has demonstrated itself to be robust across a wide range of targets and ligands. Traditionally, the MT method is used with protein-ligand binding modes generated with rigid-receptor or flexible-receptor (induced fit) docking protocols; however, these protocols are by their nature less likely to be effective with more highly flexible targets or with those situations in which binding involves multiple step pathways. In these situations, more thorough samplings are required to better explain the free energy of binding. Therefore, to explore the prediction capability and computational efficiency of the MT method when using more thorough protein-ligand conformational sampling protocols, in the present work, we introduced a series of binding mode modeling protocols ranging from conventional docking routines to single-trajectory conventional molecular dynamics (cMD) and parallel Monte Carlo molecular dynamics (MCMD). Through validation against several structurally and mechanistically diverse protein-ligand test sets, we explore the performance of the MT method as a virtual screening tool to work with the docking protocols and as an MD simulation-based binding free energy tool.
Conventional protein:ligand crystallographic refinement uses stereochemistry restraints coupled with a rudimentary energy functional to ensure the correct geometry of the model of the macromolecule—along with any bound ligand(s)—within the context of the experimental, X-ray density. These methods generally lack explicit terms for electrostatics, polarization, dispersion, hydrogen bonds, and other key interactions, and instead they use pre-determined parameters (e.g. bond lengths, angles, and torsions) to drive structural refinement. In order to address this deficiency and obtain a more complete and ultimately more accurate structure, we have developed an automated approach for macromolecular refinement based on a two layer, QM/MM (ONIOM) scheme as implemented within our DivCon Discovery Suite and "plugged in" to two mainstream crystallographic packages: PHENIX and BUSTER. This implementation is able to use one or more region layer(s), which is(are) characterized using linear-scaling, semi-empirical quantum mechanics, followed by a system layer which includes the balance of the model and which is described using a molecular mechanics functional. In this work, we applied our Phenix/DivCon refinement method—coupled with our XModeScore method for experimental tautomer/protomer state determination—to the characterization of structure sets relevant to structure-based drug design (SBDD). We then use these newly refined structures to show the impact of QM/MM X-ray refined structure on our understanding of function by exploring the influence of these improved structures on protein:ligand binding affinity prediction (and we likewise show how we use post-refinement scoring outliers to inform subsequent X-ray crystallographic efforts). Through this endeavor, we demonstrate a computational chemistry ↔ structural biology (X-ray crystallography) "feedback loop" which has utility in industrial and academic pharmaceutical research as well as other allied fields.
For decades, the complicated energy surfaces found in macromolecular protein:ligand structures, which require large amounts of computational time and resources for energy state sampling, have been an inherent obstacle to fast, routine free energy estimation in industrial drug discovery efforts. Beginning in 2013, the Merz research group addressed this cost with the introduction of a novel sampling methodology termed “Movable Type” (MT). Using numerical integration methods, the MT method reduces the computational expense for energy state sampling by independently calculating each atomic partition function from an initial molecular conformation in order to estimate the molecular free energy using ensembles of the atomic partition functions. In this work, we report a software package, the DivCon Discovery Suite with the MovableType module from QuantumBio Inc., that performs this MT free energy estimation protocol in a fast, fully encapsulated manner. We discuss the computational procedures and improvements to the original work, and we detail the corresponding settings for this software package. Finally, we introduce two validation benchmarks to evaluate the overall robustness of the method against a broad range of protein:ligand structural cases. With these publicly available benchmarks, we show that the method can use a variety of input types and parameters and exhibits comparable predictability whether the method is presented with “expensive” X-ray structures or “inexpensively docked” theoretical models. We also explore some next steps for the method. The MovableType software is available at http://www.quantumbioinc.com/
Active site protonation states play a critical role in enzyme mechanisms of action and understanding these states -in ligands, cofactors, and pocket residues -can be integral to successful structure based drug design (SBDD) campaigns.Neutron diffraction data are typically required for an accurate assessment of hydrogen positions in macromolecular crystallography.However, despite recent methodological advances, routing neutron diffraction experiments still remain very challenging.Recently we developed the XModeScore[1] method, which employs linear scaling, QM-based X-ray macromolecular refinement[2] coupled with rigorous experimental density statistical analysis in order to determine the correct protomer and tautomer states of protein residues and/or bound ligands even when challenged with data determined at routine resolutions.We will discuss the influence of protonation states on reactivity, mechanisms of action, and ligand binding within the active through the combinatorial application of the XModeScore methodology as applied to an active site involving a mixture of multiple His, Glu and Asp residues.In particular, the impact of cooperative effects of protonation of key catalytic residues, as well as the influence of coordinated metals and bound water molecules, will be discussed.
X-ray crystallography is the primarily technique used to determine the three-dimensional (3D) structure of protein:ligand and protein:protein complexes, and it plays a central role in Structure Based Drug Design (SBDD).Recently, we integrated our quantum mechanics toolkit with the Phenix crystallographic package to replace the conventional stereochemical restraints with a more accurate QM/MM-based energy functional in "real time" during the refinement and we expanded this tool to include density driven solvation and tautomer/protomer/rotamer determination.While these methods are powerful, they suffer from the same limited radius of convergence as conventional methods, and their success is ultimately dependent upon the initial atom placement of the model.We have addressed this ligand placement problem with the use of a cutting edge free energy based algorithm (i.e.MovableType or MT).Unlike conventional free energy methods, the MT method -along with the directly associated MT Dock tool -is an entirely novel way to assemble partition functions and, hence, free energies.Importantly, this method is based on fundamental statistical mechanics and does not rely on the use of expensive molecular dynamics, and it also does not suffer from many of the issues associated with conventional, energy-based found in typical docking and scoring methodologies.In this study, the MT Dock algorithm -coupled with the Phenix/DivCon package -was validated against the highly diverse PDBBind set.We showed that we were able to routinely place ligands with a high degree of accuracy as measured by the crystallographic Z-score of the difference density metric (ZDD).
Conventional macromolecular crystallographic refinement relies on often dubious stereochemical restraints, the preparation of which often requires human validation for unusual species, and on rudimentary energy functionals that are devoid of nonbonding effects owing to electrostatics, polarization, charge transfer or even hydrogen bonding. While this approach has served the crystallographic community for decades, as structure-based drug design/discovery (SBDD) has grown in prominence it has become clear that these conventional methods are less rigorous than they need to be in order to produce properly predictive protein–ligand models, and that the human intervention that is required to successfully treat ligands and other unusual chemistries found in SBDD often precludes high-throughput, automated refinement. Recently, plugins to thePython-based Hierarchical ENvironment for Integrated Xtallography(PHENIX) crystallographic platform have been developed to augment conventional methods with thein situuse of quantum mechanics (QM) applied to ligand(s) along with the surrounding active site(s) at each step of refinement [Borbulevychet al.(2014),Acta CrystD70, 1233–1247]. This method (Region-QM) significantly increases the accuracy of the X-ray refinement process, and this approach is now used, coupled with experimental density, to accurately determine protonation states, binding modes, ring-flip states, water positions and so on. In the present work, this approach is expanded to include a more rigorous treatment of the entire structure, including the ligand(s), the associated active site(s) and the entire protein, using a fully automated, mixed quantum-mechanics/molecular-mechanics (QM/MM) Hamiltonian recently implemented in theDivConpackage. This approach was validated through the automatic treatment of a population of 80 protein–ligand structures chosen from the Astex Diverse Set. Across the entire population, this method results in an average 3.5-fold reduction in ligand strain and a 4.5-fold improvement inMolProbityclashscore, as well as improvements in Ramachandran and rotamer outlier analyses. Overall, these results demonstrate that the use of a structure-wide QM/MM Hamiltonian exhibits improvements in the local structural chemistry of the ligand similar to Region-QM refinement but with significant improvements in the overall structure beyond the active site.
Conventional macromolecular crystallographic refinement relies on stereochemistry restraints and rudimentary energy functionals to ensure the correct geometry of the model of the macromolecule, along with any bound ligand(s), within the experimental, X-ray density.Traditionally, these highly approximate functionals lack explicit, chemically rigorous terms for electrostatics, polarization, dispersion, hydrogen bonds, and other interactions, and they often rely on pre-determined parameters to capture the a priori understanding of the structure.Previously, we addressed this problem, especially for active sites and ligands, through the integration of our DivCon, linear scaling, semiempricial quantum mechanics software with the PHENIX package.However, this implementation was limited in that it still treated the remainder of the system using conventional, stereochemical restraints without taking into account, longrange interactions between this part and the QM region.
This presentation will describe how structural biology, molecular pharmacology, and medicinal chemistry studies can be combined with molecular modeling and chemoinformatics analyses for a more accurate description and prediction of structural determinants of protein-ligand binding, functional activity, and selectivity.The challenges and possibilities of structural chemogenomics studies will be discussed, including the integration of large volumes of heterogeneous pharmacological and chemical data for different protein targets and the development of structure-based virtual screening and computer-aided drug design approaches to discover novel small molecule ligands with well defined functional activity and protein selectivity profiles.The potential of molecular dynamics simulation methods to complement hybrid structural biology studies will be demonstrated for the investigation the mechanisms of conformational selection and protein-ligand binding kinetics.In the final part of the presentation structural protein-ligand interaction databases will be described that link structure-based protein-ligand interaction maps to protein ligand topology and can be used as structural chemogenomics tools to navigate medicinal chemistry space.