Shadow molecular dynamics provide an efficient and stable atomistic simulation framework for flexible charge models with long-range electrostatic interactions. Shadow molecular dynamics simulations are driven by approximate "shadow" Born-Oppenheimer potentials for which the exact charges and forces are directly accessible without relying on costly (and approximate) iterative solvers. While previous implementations have been limited to atomic monopole charge distributions, we extend this approach to flexible multipole models. We derive detailed expressions for the shadow energy functions, potentials, and force terms, explicitly incorporating monopole-monopole, dipole-monopole, and dipole-dipole interactions. In our formulation, both atomic monopoles and atomic dipoles are treated as extended dynamical variables alongside the propagation of the nuclear degrees of freedom. We demonstrate that introducing the additional dipole degrees of freedom preserves the stability and accuracy previously seen in monopole-only shadow molecular dynamics simulations. In addition, we present a shadow molecular dynamics scheme where the monopole charges are held fixed while the dipoles remain flexible. Our extended shadow dynamics provide a framework for stable, computationally efficient, and versatile molecular dynamics simulations involving long-range interactions between flexible multipoles. This is of particular current interest in combination with machine-learned interatomic potentials, including long-range electrostatic interactions.
Graph-based linear-scaling electronic-structure theory provides a scalable framework for parallel quantum-mechanical molecular dynamics (QMD) simulations by exploiting the nearsightedness of the non-local electronic connectivity in non-metallic systems. When combined with recent shadow molecular dynamics in an extended-Lagrangian formulation, it enables stable long-time simulations of large, chemically active systems. This article introduces the Scalable Ecosystem, Driver, and Analyzer for Complex Chemistry Simulations (SEDACS), which integrates all these advances within a modular, Python-based software package for large-scale QMD simulations driven by external electronic-structure codes. SEDACS provides a tunable, adaptive graph construction in which edges encode the non-local electronic overlap between atoms. This graph is then decomposed into a set of smaller, overlapping subgraphs, where the electronic structure of each of these subgraphs is solved for independently and in parallel using an external electronic-structure code. SEDACS can be coupled to a variety of external electronic-structure solvers with minimal modifications to their software, enabling rapid adoption of the graph-based QMD approach. In this way, SEDACS can greatly extend the capability of existing electronic-structure packages by enabling stable QMD simulations of systems that were previously computationally inaccessible. We demonstrate highly efficient and stable QMD simulations for chemically active systems with tens of thousands of atoms by interfacing SEDACS with an external Fortran-based electronic-structure code based on self-consistent-charge density functional tight-binding theory.
A complete understanding of enzyme mechanisms requires atomistic details of chemical reactions. Quantum-based molecular dynamics simulations (QMD) are a potential source of this information, but tradeoffs between accuracy and computational cost have limited their use. Here, we develop a reactive QMD approach to investigate mechanisms of isocyanide hydratase (ICH) catalysis. In QMD simulations, molecular analogs of ICH active site residues reacted with para-nitrophenol isocyanide, forming a thioimidate intermediate. Analysis of simulated atomic configurational and charge dynamics revealed a pathway where protonation of the isocyanide carbon occurs prior to thioimidate formation. X-ray crystallography and functional assays of ICH mutants suggest this order of events might occur during enzyme catalysis. Mobile protons play essential roles in many enzymes, yet they are difficult to observe experimentally, making the ordering of proton-dependent steps ambiguous in many enzyme mechanisms. The ability to directly simulate reactions relevant to enzyme catalysis involving mobile protons demonstrates the significance of our reactive QMD approach and motivates further biological applications.
A complete understanding of enzyme mechanisms requires atomistic details of chemical reactions. Quantum-based molecular dynamics simulations (QMD) are a potential source of this information, but trade-offs between accuracy and computational cost have limited their use. We previously developed extended Lagrangian Born-Oppenheimer molecular dynamics (XL-BOMD) methods that leverage a negligible compromise in accuracy to substantially decrease the cost of QMD simulations. Here, we develop a reactive QMD approach using the latest XL-BOMD formulation, which enables efficient simulations of highly reactive systems, and use it to investigate mechanisms of intermediate formation in isocyanide hydratase (ICH) catalysis. In QMD simulations, molecular analogs of ICH active site residues reacted with para-nitrophenyl isocyanide, forming a thioimidate. Analysis of simulated atomic configurational and charge dynamics revealed a pathway where protonation of the isocyanide carbon occurs prior to thioimidate formation and suggested a possible role of Asp17 as a proton donor in the early phase of ICH catalysis. To test whether the pathway seen using the reactive QMD approach might be relevant to ICH catalysis, we performed X-ray crystallography and pre-steady-state enzyme kinetics studies of wild-type and D17N mutant ICH. Both the structure and kinetics are sensitive to the D17N mutation in a manner that is consistent with the order of the reaction steps seen in the simulations. Mobile protons play essential roles in many enzymes, yet they are difficult to observe experimentally, making the ordering of proton-dependent steps ambiguous in many enzyme mechanisms. The ability to directly simulate model reactions for the design of experiments that provide information about enzyme mechanisms involving mobile protons demonstrates the significance of our reactive QMD approach and motivates further biological applications.
Graph-based electronic structure theory offers a scalable approach to study large, complex atomistic systems using distributed and hybrid computational platforms. We demonstrate the coupling of graph-based linear scaling electronic structure theory, as implemented in the Scalable Ecosystem, Driver, and Analyzer for Complex Chemistry Simulations (SEDACS), with semiempirical quantum chemistry methods as implemented in the PySEQM code, with Graphics Processing Unit (GPU) acceleration. This powerful combination enables efficient, scalable electronic structure calculations over many nodes, significantly reducing computational cost while naturally harnessing parallelism. Detailed analyses of parallelization efficiency, computational accuracy, and communication overheads are provided, highlighting an order-of-magnitude speedup for systems of up to 10,000 atoms.
We present a framework for atomistic simulations of surface catalysis under electrochemical bias. The framework makes use of extended Lagrangian Born-Oppenheimer quantum-based molecular dynamics (XL-BOMD) simulations, which provide the speed and accuracy required for explicit atomistic treatment of both electrode and electrolyte. Simulations of solvated O_2 near nitrogen-doped graphene (NG) were performed to gain insight into the oxygen reduction reaction (ORR). Different mechanisms were observed, depending on the applied bias. Under high bias ORR occurred by an outer sphere mechanism, without adsorption of O_2 to NG. In this mechanism, electron transfer between the catalyst and the O_2 was mediated by the solvent. Under low bias ORR occurred by an inner sphere mechanism involving adsorption of O_2 to NG, leading to direct electron transfer. Combining quantum accuracy with explicit solvation and bias, XL-BOMD opens a route to predictive, atomistic insight into electrocatalytic processes beyond the reach of traditional methods.
This review article provides an overview of structurally oriented experimental datasets that can be used to benchmark protein force fields, focusing on data generated by nuclear magnetic resonance (NMR) spectroscopy and room temperature (RT) protein crystallography. We discuss what the observables are, what they tell us about structure and dynamics, what makes them useful for assessing force field accuracy, and how they can be connected to molecular dynamics simulations carried out using the force field one wishes to benchmark. We also touch on statistical issues that arise when comparing simulations with experiment. We hope this article will be particularly useful to computational researchers and trainees who develop, benchmark, or use protein force fields for molecular simulations.
This review article provides an overview of structurally oriented experimental datasets that can be used to benchmark protein force fields, focusing on data generated by nuclear magnetic resonance (NMR) spectroscopy and room temperature (RT) protein crystallography. We discuss what the observables are, what they tell us about structure and dynamics, what makes them useful for assessing force field accuracy, and how they can be connected to molecular dynamics simulations carried out using the force field one wishes to benchmark. We also touch on best practices for setup and analysis of benchmark simulations. We hope this article will be particularly useful to computational researchers and trainees who develop, benchmark, or use protein force fields or machine learning models that generate protein ensembles.
Co-design across the Exascale Computing Project has been critical for both enabling science applications and bringing disparate communities together. Developing and porting applications to the various high-performance computing architectures on pre-exascale and exascale computers has been quite challenging due to the diversity of hardware features and software stacks. The Co-design Center for Particle Applications (CoPA) has developed and enhanced the Cabana and Parallel, Rapid O(N), and Graph-Based Recursive Electronic Structure Solver (PROGRESS)/Basic Matrix Library (BML) libraries to facilitate the creation of new particle applications, make existing particle applications exascale capable, and allow teams to explore new capabilities. Particle methods from the atomistic, mesoscale, and continuum through cosmological scales have been built with Cabana, along with new possibilities for application coupling. Similarly, the PROGRESS/BML library has enabled quantum particle applications with linear algebra solvers to use advanced hardware. Across these CoPA-developed libraries, the co-design abstraction layer combines performance portability with math library support to facilitate the separation of concerns and directly support science runs.
SummaryThe upcoming exascale computing systems Frontier and Aurora will draw much of their computing power from GPU accelerators. The hardware for these systems will be provided by AMD and Intel, respectively, each supporting their own GPU programming model. The challenge for applications that harness one of these exascale systems will be to avoid lock‐in and to preserve performance portability. We report here on our results of using Kokkos to accelerate a real‐world application on NERSC's Perlmutter Phase 1 (using NVIDIA A100 accelerators) and Crusher, the testbed system for OLCF's Frontier (using AMD MI250X). By porting to Kokkos, we successfully ran the same X‐ray tracing code on both systems and achieved speed‐ups between 13 % and 66 % compared to the original CUDA code. These results are a highly encouraging demonstration of using Kokkos to accelerate production science code.
ExaFEL is an HPC-capable X-ray Free Electron Laser (XFEL) data analysis software suite for both Serial Femtosecond Crystallography (SFX) and Single Particle Imaging (SPI) developed in collaboration with the Linac Coherent Lightsource (LCLS), Lawrence Berkeley National Laboratory (LBNL) and Los Alamos National Laboratory. ExaFEL supports real-time data analysis via a cross-facility workflow spanning LCLS and HPC centers such as NERSC and OLCF. Our work therefore constitutes initial path-finding for the US Department of Energy's (DOE) Integrated Research Infrastructure (IRI) program. We present the ExaFEL team's 7 years of experience in developing real-time XFEL data analysis software for the DOE's exascale supercomputers. We present our experiences and lessons learned with the Perlmutter and Frontier supercomputers. Furthermore we outline essential data center services (and the implications for institutional policy) required for real-time data analysis. Finally we summarize our software and performance engineering approaches and our experiences with NERSC's Perlmutter and OLCF's Frontier systems. This work is intended to be a practical blueprint for similar efforts in integrating exascale compute resources into other cross-facility workflows.
To address the challenge of performance portability and facilitate the implementation of electronic structure solvers, we developed the basic matrix library (BML) and Parallel, Rapid O(N), and Graph-based Recursive Electronic Structure Solver (PROGRESS) library. The BML implements linear algebra operations necessary for electronic structure kernels using a unified user interface for various matrix formats (dense and sparse) and architectures (CPUs and GPUs). Focusing on density functional theory and tight-binding models, PROGRESS implements several solvers for computing the single-particle density matrix and relies on BML. In this paper, we describe the general strategies used for these implementations on various computer architectures, using OpenMP target functionalities on GPUs, in conjunction with third-party libraries to handle performance critical numerical kernels. We demonstrate the portability of this approach and its performance in benchmark problems.
Enzymes populate ensembles of structures necessary for catalysis that are difficult to experimentally characterize. We use time-resolved mix-and-inject serial crystallography at an x-ray free electron laser to observe catalysis in a designed mutant isocyanide hydratase (ICH) enzyme that enhances sampling of important minor conformations. The active site exists in a mixture of conformations, and formation of the thioimidate intermediate selects for catalytically competent substates. The influence of cysteine ionization on the ICH ensemble is validated by determining structures of the enzyme at multiple pH values. Large molecular dynamics simulations in crystallo and time-resolved electron density maps show that Asp 17 ionizes during catalysis and causes conformational changes that propagate across the dimer, permitting water to enter the active site for intermediate hydrolysis. ICH exhibits a tight coupling between ionization of active site residues and catalysis-activated protein motions, exemplifying a mechanism of electrostatic control of enzyme dynamics.
This chapter discusses the use of diffraction simulators to improve experimental outcomes in macromolecular crystallography, in particular for future experiments aimed at diffuse scattering. Consequential decisions for upcoming data collection include the selection of either a synchrotron or free electron laser X-ray source, rotation geometry or serial crystallography, and fiber-coupled area detector technology vs. pixel-array detectors. The hope is that simulators will provide insights to make these choices with greater confidence. Simulation software, especially those packages focused on physics-based calculation of the diffraction, can help to predict the location, size, shape, and profile of Bragg spots and diffuse patterns in terms of an underlying physical model, including assumptions about the crystal's mosaic structure, and therefore can point to potential issues with data analysis in the early planning stages. Also, once the data are collected, simulation may offer a pathway to improve the measurement of diffraction, especially with weak data, and might help to treat problematic cases such as overlapping patterns.
Graph-based linear scaling electronic structure theory for quantum-mechanical molecular dynamics simulations [A. M. N. Niklasson et al., J. Chem. Phys. 144, 234101 (2016)] is adapted to the most recent shadow potential formulations of extended Lagrangian Born-Oppenheimer molecular dynamics, including fractional molecular-orbital occupation numbers [A. M. N. Niklasson, J. Chem. Phys. 152, 104103 (2020) and A. M. N. Niklasson, Eur. Phys. J. B 94, 164 (2021)], which enables stable simulations of sensitive complex chemical systems with unsteady charge solutions. The proposed formulation includes a preconditioned Krylov subspace approximation for the integration of the extended electronic degrees of freedom, which requires quantum response calculations for electronic states with fractional occupation numbers. For the response calculations, we introduce a graph-based canonical quantum perturbation theory that can be performed with the same natural parallelism and linear scaling complexity as the graph-based electronic structure calculations for the unperturbed ground state. The proposed techniques are particularly well-suited for semi-empirical electronic structure theory, and the methods are demonstrated using self-consistent charge density-functional tight-binding theory both for the acceleration of self-consistent field calculations and for quantum-mechanical molecular dynamics simulations. Graph-based techniques combined with the semi-empirical theory enable stable simulations of large, complex chemical systems, including tens-of-thousands of atoms.
It is investigated whether molecular-dynamics (MD) simulations can be used to enhance macromolecular crystallography (MX) studies. Historically, protein crystal structures have been described using a single set of atomic coordinates. Because conformational variation is important for protein function, researchers now often build models that contain multiple structures. Methods for building such models can fail, however, in regions where the crystallographic density is difficult to interpret, for example at the protein-solvent interface. To address this limitation, a set of MD-MX methods that combine MD simulations of protein crystals with conventional modeling and refinement tools have been developed. In an application to a cyclic adenosine monophosphate-dependent protein kinase at room temperature, the procedure improved the interpretation of ambiguous density, yielding an alternative water model and a revised protein model including multiple conformations. The revised model provides mechanistic insights into the catalytic and regulatory interactions of the enzyme. The same methods may be used in other MX studies to seek mechanistic insights.
Structural dynamics of water and its hydrogen-bondingnetworksplay an important role in enzyme function via the transport of protons,ions, and substrates. To gain insights into these mechanisms in thewater oxidation reaction in Photosystem II (PS II), we have performedcrystalline molecular dynamics (MD) simulations of the dark-stableS(1) state. Our MD model consists of a full unit cell with8 PS II monomers in explicit solvent (861 894 atoms), enablingus to compute the simulated crystalline electron density and to compareit directly with the experimental density from serial femtosecondX-ray crystallography under physiological temperature collected atX-ray free electron lasers (XFELs). The MD density reproduced theexperimental density and water positions with high fidelity. The detaileddynamics in the simulations provided insights into the mobility ofwater molecules in the channels beyond what can be interpreted fromexperimental B-factors and electron densities alone.In particular, the simulations revealed fast, coordinated exchangeof waters at sites where the density is strong, and water transportacross the bottleneck region of the channels where the density isweak. By computing MD hydrogen and oxygen maps separately, we developeda novel Map-based Acceptor-Donor Identification (MADI) techniquethat yields information which helps to infer hydrogen-bond directionalityand strength. The MADI analysis revealed a series of hydrogen-bondwires emanating from the Mn cluster through the Cl1 and O4 channels;such wires might provide pathways for proton transfer during the reactioncycle of PS II. Our simulations provide an atomistic picture of thedynamics of water and hydrogen-bonding networks in PS II, with implicationsfor the specific role of each channel in the water oxidation reaction.
Molecular-dynamics (MD) simulations of protein crystals enable the prediction of structural and dynamical features of both the protein and the solvent components of macromolecular crystals, which can be validated against diffraction data from X-ray crystallographic experiments. The simulations have been useful for studying and predicting both Bragg and diffuse scattering in protein crystallography; however, the preparation is not yet automated and includes choices and tradeoffs that can impact the results. Here we examine some of the intricacies and consequences of the choices involved in setting up MD simulations of protein crystals for the study of diffraction data, and provide a recipe for preparing the simulations, packaged in an accompanying Jupyter notebook. This article and the accompanying notebook are intended to serve as practical resources for researchers wishing to put these models to work.
algebra operations for distributed architectures. A significant reduction in the wallclock time for an MD timestep is necessary to achieve practical large scale QMD. Equally challenging is to develop scalable distributed implementations. This requires the exploration of efficient data formats (dense, sparse, etc.), data decomposition and distribution across compute nodes using block- and graph-based methods. Traditional programming models such as MPI/OpenMP, hybrid MPI/OpenMP/Accelerator, as well as asynchronous task-based models will be explored to achieve maximum parallelism. Parallelism at the level of a matrix operation, solver, and/or application needs to be explored to gain the most benefit on new and emerging architectures. This report serves as a “living document” of the additions and enhancements throughout the ECP CoPA project for $$\mathcal{O}(N)$$ QMD. The following sections address the applications we are working with, the quantum chemistry libraries (BML and PROGRESS), and our co-design exploration through the use of a proxy application.
Water often plays a key role in protein structure, molecular recognition, and mediating protein-ligand interactions. Thus, free energy calculations must adequately sample water motions, which often proves challenging in typical MD simulation time scales. Thus, the accuracy of methods relying on MD simulations ends up limited by slow water sampling. Particularly, as a ligand is removed or modified, bulk water may not have time to fill or rearrange in the binding site. In this work, we focus on several molecular dynamics (MD) simulation-based methods attempting to help rehydrate buried water sites: BLUES, using nonequilibrium candidate Monte Carlo (NCMC); grand, using grand canonical Monte Carlo (GCMC); and normal MD. We assess the accuracy and efficiency of these methods in rehydrating target water sites. We selected a range of systems with varying numbers of waters in the binding site, as well as those where water occupancy is coupled to the identity or binding mode of the ligand. We analyzed the rehydration of buried water sites in binding pockets using both clustering of trajectories and direct analysis of electron density maps. Our results suggest both BLUES and grand enhance water sampling relative to normal MD and grand is more robust than BLUES, but also that water sampling remains a major challenge for all of the methods tested. The lessons we learned for these methods and systems are discussed.
William S. Hlavacek合作论文数Theoretical Biology and Biophysics Group, Theoretical Division, Los Alamos National Laboratory, Los Alamos, NM 87545, USA11