The critical importance of water in sustaining life highlights the need for accurate water models in computer simulations, aiming to mimic biochemical processes experimentally. The polarizable Gaussian multipole (pGM) model, recently introduced for biomolecular simulations, improves the handling of complex biomolecular interactions. As an integral part of our initial exploration, we examined a minimalist fixed geometry three-center pGM water model using ab initio quantum mechanical calculations of water oligomers. However, our final model development was based on liquid-phase water properties, leveraging automated machine learning (AutoML) techniques for optimization. This allows the development of a framework to refine both van der Waals and electrostatic parameters of the pGM model, aiming to accurately reproduce specific properties such as the oxygen-oxygen radial distribution function, density, and dipole moment, all at 298 K and 1.0 bar pressure. The efficacy of the optimized three-center pGM water model, pGM3P-25, was assessed through simulations of a water box of 512 water molecules, showcasing marked enhancements in both accuracy and practical utility. Notably, the model accurately reproduces thermodynamic properties not explicitly included in training while significantly reducing the time and human effort required for optimization. It was found that pGM3P-25 can reproduce temperature-dependent properties such as density, self-diffusion constants, heat capacity, second virial coefficient, and dielectric constant, which are important in biomolecular simulations. This study underscores the potential of AutoML-driven frameworks to streamline parameter refinement for molecular dynamics simulations, paving the way for broader applications in computational chemistry and beyond.
The van der Waals potential is an integral component in molecular mechanics force fields. In this work, we document our effort to develop the parameters for the Lennard-Jones (LJ) 12-6 potential as part of our effort to develop a polarizable Gaussian multipole force field. Starting from the GAFF2 van der Waals parameter set, atom types were assigned using the Antechamber program in the AMBER24 simulation package. Atomic charges and permanent dipoles were calculated via PCMRESP, based on electrostatic potentials computed at the MP2/aug-cc-pVTZ//MP2/6-311++G(d,p) level in four solvents with dielectric constants ranging from 4.24 to 78.36. The optimization of the LJ parameters was conducted in two stages. In the first stage, we employed CCSD(T)/CBS interaction energies of about 4.6 million dimer conformations, including those from Shaw and co-workers, and applied an iterative least-squares fitting procedure. This reduced the Boltzmann factor-weighted root-mean-square error (RMSE) from 2.44 to 0.97 kcal/mol. When tested on 161 neat liquids, the simulated densities yielded an average unsigned percent error of 5.3% from the experimental values. In the second stage, the parameters were further refined by comparing simulated and experimental densities of the same 161-liquid training set using a gradient-guided iterative search based on short molecular dynamics simulations. After 24 iterations, the average unsigned percent error was reduced to 1.71%. Extended 10 ns simulations yielded an average unsigned percent error of 1.69% for the training set densities, which compares favorably with the 2.45% error obtained from GAFF2. The RMSE for heats of vaporization (H vap) was 1.39 kcal/mol. This result is encouraging, even though it is slightly higher than the 1.23 kcal/mol RMSE from GAFF2 because H vap was not used in the training, where energetic terms, including H vap, were part of the training data in GAFF2. Evaluations on a test set further demonstrated that the optimized parameters reproduce experimental densities and H vap with errors of 1.25% and 1.50 kcal/mol, respectively; both are better than GAFF2, which yielded 1.62% and 2.43 kcal/mol errors, respectively. Given GAFF2's broad adoption and independent validation by the broad community, we consider the accuracy of the optimized pGM LJ parameters to be acceptable.
Proteins' flexibility is a feature in communicating changes in cell signaling instigated by binding with secondary messengers, such as calcium ions, associated with the coordination of muscle contraction, neurotransmitter release, and gene expression. When binding with the disordered parts of a protein, calcium ions must balance their charge states with the shape of calcium-binding proteins and their versatile pool of partners depending on the circumstances they transmit. Accurately determining the ionic charges of those ions is essential for understanding their role in such processes. However, it is unclear whether the limited experimental data available can be effectively used to train models to accurately predict the charges of calcium-binding protein variants. Here, we developed a chemistry-informed, machine-learning algorithm that implements a game theoretic approach to explain the output of a machine-learning model without the prerequisite of an excessively large database for high-performance prediction of atomic charges. We used the ab initio electronic structure data representing calcium ions and the structures of the disordered segments of calcium-binding peptides with surrounding water molecules to train several explainable models. Network theory was used to extract the topological features of atomic interactions in the structurally complex data dictated by the coordination chemistry of a calcium ion, a potent indicator of its charge state in protein. Our design created a computational tool of CaXML, which provided a framework of explainable machine learning model to annotate ionic charges of calcium ions in calcium-binding proteins in response to the chemical changes in an environment. Our framework will provide new insights into protein design for engineering functionality based on the limited size of scientific data in a genome space.
Atomic polarizabilities are considered to be fundamental parameters in polarizable molecular mechanical force fields that play pivotal roles in determining model transferability across different electrostatic environments. In an earlier work, the atomic polarizabilities were obtained by fitting them to the B3LYP/aug-cc-pvtz molecular polarizability tensors of mainly small molecules. Taking advantage of the recent PCMRESPPOL method, we refine the atomic polarizabilities for condensed-phase simulations using a polarizable Gaussian Multipole (pGM) force field. Departing from earlier works, in this work, we incorporated polarizability tensors of a large number of dimers and electrostatic potentials (ESPs) in multiple solvents. We calculated 1565 × 4 ESPs of small molecule monomers and dimers of noble gas and small molecules and 4742 × 4 ESPs of small molecule dimers in four solvents (diethyl ether, ε = 4.24, dichloroethane, ε = 10.13, acetone, ε = 20.49, and water, ε = 78.36). For the gas-phase polarizability tensors, we supplemented the molecule set that was used in our earlier work by adding both the 4252 monomer and dimer sets studied by Shaw and co-workers and the 7211 small molecule monomers listed in the QM7b database to a combined total of 13,523 molecular polarizability tensors of monomers and dimers. The QM7b polarizability set was obtained from quantum-machine.org and was calculated at the LR-CCSD/d-aug-cc-pVDZ level of theory. All other polarizability tensors and all ESPs were calculated at the ωB97X-D/aug-cc-pVTZ level of theory. The atomic polarizabilities were developed using all polarizability tensors and the 1565 × 4 ESPs of small molecule monomers and were then assessed by comparing them to the 4742 × 4 ab initio ESPs of small molecule dimers. The predicted dimer ESPs had an average relative root-mean-square error (RRMSE) of 9.30%, which was only slightly larger than the average fitting RRMSE of 9.15% of the monomer ESPs. The transferability of the polarizability set was further evaluated by comparing the ESPs calculated using parameters developed in another dielectric environment for both tetrapeptide and DES monomer data sets. It was observed that the polarizabilities of this work retained or slightly improved the transferability over the one discussed in earlier work even though the number of parameters in the present set is about half of that in the earlier set. Excluding the gas-phase data, for the DES monomer set, the average transfer RRMSEs were 16.25% and 10.83% for pGM-ind and pGM-perm methods, respectively, comparable to the average fitting RRMSEs of 16.03% and 10.54%; for tetrapeptides, the average transfer RRMSEs were 5.62% and 3.95% for pGM-ind and pGM-perm methods, respectively, slightly larger than 5.41% and 3.61% of the fitting RRMSEs. Therefore, we conclude that the pGM methods with updated polarizabilities achieved remarkable transferability from monomer to dimer and from one solvent to another.
Accurate parametrization of amino acids is pivotal for the development of reliable force fields for molecular modeling of biomolecules such as proteins. This study aims to assess amino acid electrostatic parametrizations with the polarizable Gaussian Multipole (pGM) model by evaluating the performance of the pGM-perm (with atomic permanent dipoles) and pGM-ind (without atomic permanent dipoles) variants compared to the traditional RESP model. The 100-conf-combterm fitting strategy on tetrapeptides was adopted, in which (1) all peptide bond atoms (-CO-NH-) share identical set of parameters and (2) the total charges of the two terminal N-acetyl (ACE) and N-methylamide (NME) groups were set to neutral. The accuracy and transferability of electrostatic parameters across peptides with varying lengths and real-world examples were examined. The results demonstrate the enhanced performance of the pGM-perm model in accurately representing the electrostatic properties of amino acids. This insight underscores the potential of the pGM-perm model and the 100-conf-combterm strategy for the future development of the pGM force field.
The transferability of force field parameters is a crucial aspect of high-quality force fields. Previous investigations have affirmed the transferability of electrostatic parameters derived from polarizable Gaussian multipole models (pGMs) when applied to water oligomer clusters, polypeptides across various conformations, and different sequences. In this study, we introduce PCMRESP, a novel method for electrostatic parametrization in solution, intended for the development of polarizable force fields. We utilized this method to assess the transferability of three models: a fixed charge model and two variants of pGM models. Our analysis involved testing these models on 377 small molecules and 100 tetra-peptides in five representative dielectric environments: gas, diethyl ether, dichloroethane, acetone, and water. Our findings reveal that the inclusion of atomic polarization significantly enhances transferability and the incorporation of permanent atomic dipoles, in the form of covalent bond dipoles, leads to further improvements. Moreover, our tests on dual-solvent strategies demonstrate consistent transferability for all three models, underscoring the robustness of the dual-solvent approach. In contrast, an evaluation of the traditional HF/6-31G* method indicates poor transferability for the pGM-ind and pGM-perm models, suggesting the limitations of this conventional approach.
Proteins' fuzziness are features for communicating changes in cell signaling instigated by binding with secondary messengers, such as calcium ions, associated with the coordination of muscle contraction, neurotransmitter release, and gene expression. Binding with the disordered parts of a protein, calcium ions must balance their charge states with the shape of calcium-binding proteins and their versatile pool of partners depending on the circumstances they transmit, but it is unclear whether the limited experimental data available can be used to train models to accurately predict the charges of calcium-binding protein variants. Here, we developed a chemistry-informed, machine-learning algorithm that implements a game theoretic approach to explain the output of a machine-learning model without the prerequisite of an excessively large database for high-performance prediction of atomic charges. We used the ab initio electronic structure data representing calcium ions and the structures of the disordered segments of calcium-binding peptides with surrounding water molecules to train several explainable models. Network theory was used to extract the topological features of atomic interactions in the structurally complex data dictated by the coordination chemistry of a calcium ion, a potent indicator of its charge state in protein. With our designs, we provided a framework of explainable machine learning model to annotate atomic charges of calcium ions in calcium-binding proteins with domain knowledge in response to the chemical changes in an environment based on the limited size of scientific data in a genome space.
Accurate characterization of electrostatic interactions is crucial in molecular simulation. Various methods and programs have been developed to obtain electrostatic parameters for additive or polarizable models to replicate electrostatic properties obtained from experimental measurements or theoretical calculations. Electrostatic potentials (ESPs), a set of physically well-defined observables from quantum mechanical (QM) calculations, are well suited for optimization efforts due to the ease of collecting a large amount of conformation-dependent data. However, a reliable set of QM ESP computed at an appropriate level of theory and atomic basis set is necessary. In addition, despite the recent development of the PyRESP program for electrostatic parameterizations of induced dipole-polarizable models, the time-consuming and error-prone input file preparation process has limited the widespread use of these protocols. This work aims to comprehensively evaluate the quality of QM ESPs derived by eight methods, including wave function methods such as Hartree-Fock (HF), second-order Møller-Plesset (MP2), and coupled cluster-singles and doubles (CCSD), as well as five hybrid density functional theory (DFT) methods, used in conjunction with 13 different basis sets. The highest theory levels CCSD/aug-cc-pV5Z (a5z) and MP2/aug-cc-pV5Z (a5z) were selected as benchmark data over two homemade data sets. The results show that the hybrid DFT method, ωB97X-D, combined with the aug-cc-pVTZ (a3z) basis set, performs well in reproducing ESPs while taking both accuracy and efficiency into consideration. Moreover, a flexible and user-friendly program called PyRESP_GEN was developed to streamline input file preparation. The restraining strengths, along with strategies for polarizable Gaussian multipole (pGM) model parameterizations, were also optimized. These findings and the program presented in this work facilitate the development and application of induced dipole-polarizable models, such as pGM models, for molecular simulations of both chemical and biological significance.
Zika virus (ZIKV) serine protease, indispensable for viral polyprotein processing and replication, is composed of the membrane-anchored NS2B polypeptide and the N-terminal domain of the NS3 polypeptide (NS3pro). The C-terminal domain of the NS3 polypeptide (NS3hel) is necessary for helicase activity and contains an ATP-binding site. We discovered that ZIKV NS2B-NS3pro binds single-stranded RNA with a Kd of ~0.3 μM, suggesting a novel function. We tested various structural modifications of NS2B-NS3pro and observed that constructs stabilized in the recently discovered "super-open" conformation do not bind RNA. Likewise, stabilizing NS2B-NS3pro in the "closed" (proteolytically active) conformation using substrate inhibitors abolished RNA binding. We posit that RNA binding occurs when ZIKV NS2B-NS3pro adopts the "open" conformation, which we modeled using highly homologous dengue NS2B-NS3pro crystallized in the open conformation. We identified two positively charged fork-like structures present only in the open conformation of NS3pro. These forks are conserved across Flaviviridae family and could be aligned with the positively charged grove on NS3hel, providing a contiguous binding surface for the negative RNA strand exiting helicase. We propose a "reverse inchworm" model for a tightly intertwined NS2B-NS3 helicase-protease machinery, which suggests that NS2B-NS3pro cycles between open and super-open conformations to bind and release RNA enabling long-range NS3hel processivity. The transition to the closed conformation, likely induced by the substrate, enables the classical protease activity of NS2B-NS3pro.
Accuracy and transferability are the two highly desirable properties of molecular mechanical force fields. Compared with the extensively used point-charge additive force fields that apply fixed atom-centered point partial charges to model electrostatic interactions, polarizable force fields are thought to have the advantage of modeling the atomic polarization effects. Previous works have demonstrated the accuracy of the recently developed polarizable Gaussian multipole (pGM) models. In this work, we assessed the transferability of the electrostatic parameters of the pGM models with (pGM-perm) and without (pGM-ind) atomic permanent dipoles in terms of reproducing the electrostatic potentials surrounding molecules/oligomers absent from electrostatic parameterizations. Encouragingly, both the pGM-perm and pGM-ind models show significantly improved transferability than the additive model in the tests (1) from water monomer to water oligomer clusters; (2) across different conformations of amino acid dipeptides and tetrapeptides; (3) from amino acid tetrapeptides to longer polypeptides; and (4) from nucleobase monomers to Watson-Crick base pair dimers and tetramers. Furthermore, we demonstrated that the double-conformation fittings using amino acid tetrapeptides in the αR and β conformations can result in good transferability not only across different tetrapeptide conformations but also from tetrapeptides to polypeptides with lengths ranging from 1 to 20 repetitive residues for both the pGM-ind and pGM-perm models. In addition, the observation that the pGM-ind model has significantly better accuracy and transferability than the point-charge additive model, even though they have an identical number of parameters, strongly suggest the importance of intramolecular polarization effects. In summary, this and previous works together show that the pGM models possess both accuracy and transferability, which are expected to serve as foundations for the development of next-generation polarizable force fields for modeling various polarization-sensitive biological systems and processes.
Induced dipole models have proven to be effective tools for simulating electronic polarization effects in biochemical processes, yet their potential has been constrained by energy conservation issue, particularly when historical data is utilized for dipole prediction. This study identifies error outliers as the primary factor causing this failure of energy conservation and proposes a comprehensive scheme to overcome this limitation. Leveraging maximum relative errors as a convergence metric, our data demonstrates that energy conservation can be upheld even when using historical information for dipole predictions. Our study introduces the multi-order extrapolation method to quicken induction iteration and optimize the use of historical data, while also developing the preconditioned conjugate gradient with local iterations to refine the iteration process and effectively remove error outliers. This scheme further incorporates a "peek" step via Jacobi under-relaxation for optimal performance. Simulation evidence suggests that our proposed scheme can achieve energy convergence akin to that of point-charge models within a limited number of iterations, thus promising significant improvements in efficiency and accuracy.
Dysregulation of autophagic pathways leads to accumulation of abnormal proteins and damaged organelles in many neurodegenerative disorders, including Parkinson's disease (PD) and Lewy body dementia (LBD). Autophagy-related dysfunction may also trigger secretion and spread of misfolded proteins, such as α-synuclein (α-syn), the major misfolded protein found in PD/LBD. However, the mechanism underlying these phenomena remains largely unknown. Here, we used cell-based models, including human induced pluripotent stem cell-derived neurons, CRISPR/Cas9 technology, and male transgenic PD/LBD mice, plus vetting in human postmortem brains (both male and female). We provide mechanistic insight into this pathologic pathway. We find that aberrant S-nitrosylation of the autophagic adaptor protein p62 causes inhibition of autophagic flux and intracellular buildup of misfolded proteins, with consequent secretion resulting in cell-to-cell spread. Thus, our data show that pathologic protein S-nitrosylation of p62 represents a critical factor not only for autophagic inhibition and demise of individual neurons, but also for α-syn release and spread of disease throughout the nervous system.SIGNIFICANCE STATEMENTIn Parkinson's disease and Lewy body dementia, dysfunctional autophagy contributes to accumulation and spread of aggregated α-synuclein. Here, we provide evidence that protein S-nitrosylation of p62 inhibits autophagic flux, contributing to α-synuclein aggregation and spread.
Amyloid-β (Aβ) peptides are involved in Alzheimer’s disease (AD) development. The interactions of these peptides with copper and zinc ions also seem to be crucial for this pathology. Although Cu(II) and Zn(II) ions binding by Aβ peptides has been scrupulously investigated, surprisingly, this phenomenon has not been so thoroughly elucidated for N-truncated Aβ4−x—probably the most common version of this biomolecule. This negligence also applies to mixed Cu–Zn complexes. From the structural in silico analysis presented in this work, it appears that there are two possible mixed Cu–Zn(Aβ4−x) complexes with different stoichiometries and, consequently, distinct properties. The Cu–Zn(Aβ4−x) complex with 1:1:1 stoichiometry may have a neuroprotective superoxide dismutase-like activity. On the other hand, another mixed 2:1:2 Cu–Zn(Aβ4−x) complex is perhaps a seed for toxic oligomers. Hence, this work proposes a novel research direction for our better understanding of AD development.
Electrostatic interactions are of fundamental importance for the structures and functions of biomolecules. Their accurate modeling is crucial in the design and development of physical models in computational studies of biomolecules. It is known that widely used point-charge models cannot capture subtle multibody effects and molecular polarization anisotropicity. Recently emerged point multipole models, in the meantime, face a so-called “polarization catastrophe” difficulty that requires artificial truncation and omission of important short-range interactions.
Molecular modeling on atomic level have been applied in a wide range of biological systems. The widely adopted additive force fields typically use fixed atom-centered partial charges to model electrostatic interactions. However, the additive force fields cannot accurately model polarization effects, leading to unrealistic simulations in polarization sensitive processes. Numerous attempts have been directed to developing induced dipole based polarizable models, including the Applequist point dipole model, the Thole damped dipole model, and the recently proposed polarizable Gaussian multipole (pGM) model.
Our previous article has established the theory of molecular dynamics (MD) simulations for systems modeled with the polarizable Gaussian multipole (pGM) electrostatics [Wei et al., J. Chem. Phys. 153(11), 114116 (2020)]. Specifically, we proposed the covalent basis vector framework to define the permanent multipoles and derived closed-form energy and force expressions to facilitate an efficient implementation of pGM electrostatics. In this study, we move forward to derive the pGM internal stress tensor for constant pressure MD simulations with the pGM electrostatics. Three different formulations are presented for the flexible, rigid, and short-range screened systems, respectively. The analytical formulations were implemented in the SANDER program in the Amber package and were first validated with the finite-difference method for two different boxes of pGM water molecules. This is followed by a constant temperature and constant pressure MD simulation for a box of 512 pGM water molecules. Our results show that the simulation system stabilized at a physically reasonable state and maintained the balance with the externally applied pressure. In addition, several fundamental differences were observed between the pGM and classic point charge models in terms of the simulation behaviors, indicating more extensive parameterization is necessary to utilize the pGM electrostatics.
Molecular modeling at the atomic level has been applied in a wide range of biological systems. The widely adopted additive force fields typically use fixed atom-centered partial charges to model electrostatic interactions. However, the additive force fields cannot accurately model polarization effects, leading to unrealistic simulations in polarization-sensitive processes. Numerous efforts have been invested in developing induced dipole-based polarizable force fields. Whether additive atomic charge models or polarizable induced dipole models are used, proper parameterization of the electrostatic term plays a key role in the force field developments. In this work, we present a Python program called PyRESP for performing atomic multipole parameterizations by reproducing ab initio electrostatic potential (ESP) around molecules. PyRESP provides parameterization schemes for several electrostatic models, including the RESP model with atomic charges for the additive force fields and the RESP-ind and RESP-perm models with additional induced and permanent dipole moments for the polarizable force fields. PyRESP is a flexible and user-friendly program that can accommodate various needs during force field parameterizations for molecular modeling of any organic molecules.
A key advantage of polarizable force fields is their ability to model the atomic polarization effects that play key roles in the atomic many-body interactions. In this work, we assessed the accuracy of the recently developed polarizable Gaussian Multipole (pGM) models in reproducing quantum mechanical (QM) interaction energies, many-body interaction energies, as well as the nonadditive and additive contributions to the many-body interactions for peptide main-chain hydrogen-bonding conformers, using glycine dipeptide oligomers as the model systems. Two types of pGM models were considered, including that with (pGM-perm) and without (pGM-ind) permanent atomic dipoles. The performances of the pGM models were compared with several widely used force fields, including two polarizable (Amoeba13 and ff12pol) and three additive (ff19SB, ff15ipq, and ff03) force fields. Encouragingly, the pGM models outperform all other force fields in terms of reproducing QM interaction energies, many-body interaction energies, as well as the nonadditive and additive contributions to the many-body interactions, as measured by the root-mean-square errors (RMSEs) and mean absolute errors (MAEs). Furthermore, we tested the robustness of the pGM models against polarizability parameterization errors by employing alternative polarizabilities that are either scaled or obtained from other force fields. The results show that the pGM models with alternative polarizabilities exhibit improved accuracy in reproducing QM many-body interaction energies as well as the nonadditive and additive contributions compared with other polarizable force fields, suggesting that the pGM models are robust against the errors in polarizability parameterizations. This work shows that the pGM models are capable of accurately modeling polarization effects and have the potential to serve as templates for developing next-generation polarizable force fields for modeling various biological systems.
Proteases comprise an important class of enzymes, whose activity is central to many physiologic and pathologic processes. Detailed knowledge of protease specificity is key to understanding their function. Although many methodologies have been developed to profile specificities of proteases, few have the diversity and quantitative grasp necessary to fully define specificity of a protease, both in terms of substrate numbers and their catalytic efficiencies. We have developed a concept of “selectome”, which defines the set of substrates that uniquely represents specificity of a protease. We applied it to two closely related members of the Matrixin family – MMP-2 and MMP-9 by using substrate phage display coupled with Next Generation Sequencing and information theory-based data analysis. We have also derived a quantitative measure of substrate specificity, which accounts for both the numbers and relative catalytic efficiencies of substrates. Using these advances greatly facilitates uncovering selectivity between closely related members of protease families and provides insight into to the degree of contribution of catalytic cleft specificity to protein substrate recognition, thus providing basis to overcoming two of the major challenges in the field of proteolysis: 1) development of highly selective activity probes and inhibitors for studying proteases with overlapping specificities, and 2) distinguishing targeted proteolysis from bystander proteolytic events.
Yong Duan (段勇)合作论文数University of California,Davis26
Ray Luo合作论文数School of Biological Sciences, University of California22