Force Field X (FFX) is an open-source software package for atomic resolution modeling of genetic variants and organic crystals that leverages advanced potential energy functions and experimental data. FFX currently consists of nine modular packages with novel algorithms that include global optimization via a many-body expansion, acid-base chemistry using polarizable constant-pH molecular dynamics, estimation of free energy differences, generalized Kirkwood implicit solvent models, and many more. Applications of FFX focus on the use and development of a crystal structure prediction pipeline, biomolecular structure refinement against experimental datasets, and estimation of the thermodynamic effects of genetic variants on both proteins and nucleic acids. The use of Parallel Java and OpenMM combines to offer shared memory, message passing, and graphics processing unit parallelization for high performance simulations. Overall, the FFX platform serves as a computational microscope to study systems ranging from organic crystals to solvated biomolecular systems.
Computational methods developed to find the global free energy minimum of amino acid sequences are increasingly successful, but limitations in both accuracy and efficiency remain. Optimization algorithms are typically focused on proteins of modest size (i.e. of approximately 100 residues) and utilize potential energy functions based on fixed charged force fields, statistical or knowledge based potentials, and/or potentials incorporating experimental data. Although the aforementioned methods are widely used, known limitations include 1) search protocols that are inefficient or not deterministic due to rough energy landscapes characterized by large energy barriers between multiple minima and 2) use of a target function whose global minimum does not correspond to the actual free energy minimum. To overcome the first limitation, this work describes a global optimization approach based on metadynamics to drive the search of conformational space toward unexplored regions by adding a time-dependent bias to the objective function. To overcome the second limitation, a hybrid objective function is defined as the sum of the polarizable AMOEBA polarizable force field and an experimental X-ray crystallography target. As metadynamics drives the search, periodic quenching via local minimization is used to access structure quality via evaluation of Rwork. Thus, the overall method is called AMOEBA Metadynamics with Minimization (AMM), and is suitable for optimization of side-chains, ligand binding poses, protein loops or even protein complexes. Here we focus on characterizing the ability of AMM to elucidate the structural details of missing protein loops, which are often excluded from experimental X-ray crystallography structures due to conformational heterogeneity and/or limitations in the resolution of the data. We first show that the correlation between experimental data and AMOEBA structural minima is stronger than that for OPLS-AA/L (i.e. a fixed charge force field). Next, missing protein loops are optimized using 5 nsec of sampling for both AMM and simulated annealing with OPLS-AA/L. The AMM procedure provides more accurate structures in terms of both experimental (i.e. lower Rfree values) and structural metrics (i.e. MolProbity). In addition to providing more accurate loop conformations, AMM converged faster than the simulated annealing protocol. Overall, this work suggests that AMM is well-suited to refine or predict the coordinates of missing amino acid residues and/or protein loops due to both the increased accuracy of the target function relative to OPLS-AA/L and more rapid convergence of the metadynamics driven search compared to simulation annealing.
A balance of van der Waals, electrostatic, and hydrophobic forces drive the folding and packing of protein side chains. Although such interactions between residues are often approximated as being pairwise additive, in reality, higher-order many-body contributions that depend on environment drive hydrophobic collapse and cooperative electrostatics. Beginning from dead-end elimination, we derive the first algorithm, to our knowledge, capable of deterministic global repacking of side chains compatible with many-body energy functions. The approach is applied to seven PCNA x-ray crystallographic data sets with resolutions 2.5-3.8 angstrom (mean 3.0 angstrom) using an open-source software. While PDB_REDO models average an R-free value of 29.5% and MOLPROBITY score of 2.71 angstrom (77th percentile), dead-end elimination with the polarizable AMOEBA force field lowered R-free by 2.8-26.7% and improved mean MOLPROBITY score to atomic resolution at 1.25 angstrom (100th percentile). For structural biology applications that depend on side-chain repacking, including x-ray refinement, homology modeling, and protein design, the accuracy limitations of pairwise additivity can now be eliminated via polarizable or quantum mechanical potentials.
Global optimization of protein side-chains using discrete rotamer libraries is a challenging problem due to the large number of permutations that exist. For example, a small protein with 100 amino acid residues, each with 3 energetically favorable conformations, gives 3100 (or 1047) permutations to be searched. A systematic search over all permutations is computationally intractable, even in this simple case. To reduce the search space, rigorous inequalities have been described that eliminate high-energy rotamers, rotamer pairs, and so on (i.e. dead-end elimination), based on the assumption of a pairwise decomposable energy function. However, important biomolecular driving forces, including the hydrophobic effect and electronic polarization, are fundamentally many-body in nature. Neglect of higher order many-body forces has unduly limited the accuracy protein side-chain optimization and its applications to protein structure refinement, protein design and drug discovery. Here we present new rotamer elimination criteria that make possible the provable global optimization of protein side-chains using a polarizable (many-body) force field. Our many-body elimination criteria are only modestly more complex than the earlier pairwise versions when truncated at 3- or 4-body interactions. This opens the door to the rigorous use of not only polarizable force fields, which are now widely available for protein simulations, but also quantum mechanical potentials and/or implicit solvents such as Poisson-Boltzmann. Application to the protein "Proliferating Cell Nuclear Antigen" (PCNA) will be presented, which is a eukaryotic sliding clamp involved in numerous DNA processing pathways.
We also propose a website to host such a DBMS that can facilitate researchers to upload their data and perform analysis efficiently. Users can analyze their data using efficient functions implemented to access the data. Index structures are generated to store all results of analysis that may be interesting to other users, so that the results of analysis are readily available without the need to duplicate the analysis. The data upload feature can be made available through application program interfaces (APIs). Users can upload their data using the APIs. The DBMS takes care of generating indexes, storing results of analysis and retrieving efficiently, whenever users run analysis query or request information.
SNARE proteins promote membrane fusion by forming a four-stranded parallel helical bundle that brings the membranes into close proximity. Post-fusion, the complex is disassembled by an AAA+ ATPase called N-ethylmaleimide-sensitive factor (NSF). We present evidence that NSF uses a processive unwinding mechanism to disassemble SNARE proteins. Using a real-time disassembly assay based on fluorescence dequenching, we correlate NSF-driven disassembly rates with the SNARE-activated ATPase activity of NSF. Neuronal SNAREs activate the ATPase rate of NSF by similar to 26-fold. One SNARE complex takes an average of similar to 5 s to disassemble in a process that consumes similar to 50 ATP. Investigations of substrate requirements show that NSF is capable of disassembling a truncated SNARE substrate consisting of only the core SNARE domain, but not an unrelated four-stranded coiled-coil. NSF can also disassemble an engineered double-length SNARE complex, suggesting a processive unwinding mechanism. We further investigated processivity using single-turnover experiments, which show that SNAREs can be unwound in a single encounter with NSF. We propose a processive helicase-like mechanism for NSF in which similar to 1 residue is unwound for every hydrolyzed ATP molecule.
Enzymes stabilize transition states of reactions while limiting binding to ground states, as is generally required for any catalyst. Alkaline Phosphatase (AP) and other nonspecific phosphatases are some of Nature's most impressive catalysts, achieving preferential transition state over ground state stabilization of more than 10²²-fold while utilizing interactions with only the five atoms attached to the transferred phosphorus. We tested a model that AP achieves a portion of this preference by destabilizing ground state binding via charge repulsion between the anionic active site nucleophile, Ser102, and the negatively charged phosphate monoester substrate. Removal of the Ser102 alkoxide by mutation to glycine or alanine increases the observed Pi affinity by orders of magnitude at pH 8.0. To allow precise and quantitative comparisons, the ionic form of bound P(i) was determined from pH dependencies of the binding of Pi and tungstate, a P(i) analog lacking titratable protons over the pH range of 5-11, and from the ³¹P chemical shift of bound P(i). The results show that the Pi trianion binds with an exceptionally strong femtomolar affinity in the absence of Ser102, show that its binding is destabilized by ≥10⁸-fold by the Ser102 alkoxide, and provide direct evidence for ground state destabilization. Comparisons of X-ray crystal structures of AP with and without Ser102 reveal the same active site and P(i) binding geometry upon removal of Ser102, suggesting that the destabilization does not result from a major structural rearrangement upon mutation of Ser102. Analogous Pi binding measurements with a protein tyrosine phosphatase suggest the generality of this ground state destabilization mechanism. Our results have uncovered an important contribution of anionic nucleophiles to phosphoryl transfer catalysis via ground state electrostatic destabilization and an enormous capacity of the AP active site for specific and strong recognition of the phosphoryl group in the transition state.
Hydrogen bond networks are key elements of protein structure and function but have been challenging to study within the complex protein environment. We have carried out in-depth interrogations of the proton transfer equilibrium within a hydrogen bond network formed to bound phenols in the active site of ketosteroid isomerase. We systematically varied the proton affinity of the phenol using differing electron-withdrawing substituents and incorporated site-specific NMR and IR probes to quantitatively map the proton and charge rearrangements within the network that accompany incremental increases in phenol proton affinity. The observed ionization changes were accurately described by a simple equilibrium proton transfer model that strongly suggests the intrinsic proton affinity of one of the Tyr residues in the network, Tyr16, does not remain constant but rather systematically increases due to weakening of the phenol-Tyr16 anion hydrogen bond with increasing phenol proton affinity. Using vibrational Stark spectroscopy, we quantified the electrostatic field changes within the surrounding active site that accompany these rearrangements within the network. We were able to model these changes accurately using continuum electrostatic calculations, suggesting a high degree of conformational restriction within the protein matrix. Our study affords direct insight into the physical and energetic properties of a hydrogen bond network within a protein interior and provides an example of a highly controlled system with minimal conformational rearrangements in which the observed physical changes can be accurately modeled by theoretical calculations.
Central to most forms of autophagy are two ubiquitin-like proteins (UBLs), Atg8 and Atg12, which play important roles in autophagosome biogenesis, substrate recruitment to autophagosomes, and other aspects of autophagy. Typically, UBLs are activated by an E1 enzyme that (1) catalyzes adenylation of the UBL C terminus, (2) transiently covalently captures the UBL through a reactive thioester bond between the E1 active site cysteine and the UBL C terminus, and (3) promotes transfer of the UBL C terminus to the catalytic cysteine of an E2 conjugating enzyme. The E2, and often an E3 ligase enzyme, catalyzes attachment of the UBL C terminus to a primary amine group on a substrate. Here, we summarize our recent work reporting the structural and mechanistic basis for E1-E2 protein interactions in autophagy.
Vesicle trafficking in eukaryotic cells is facilitated by SNARE-mediated membrane fusion. The ATPase NSF (N-ethylmaleimide- sensitive factor) and the adaptor protein alpha-SNAP (soluble NSF attachment protein) disassemble all SNARE complexes formed throughout different pathways, but the effect of SNARE sequence and domain variation on the poorly understood disassembly mechanism is unknown. By measuring SNARE-stimulated ATP hydrolysis rates, Michaelis-Menten constants for disassembly, and SNAP-SNARE binding constants for four different ternary SNARE complexes and one binary complex, we found a conserved mechanism, not influenced by N-terminal SNARE domains. alpha-SNAP and the ternary SNARE complex form a 1:1 complex as revealed by multiangle light scattering. We propose a model of NSF-mediated disassembly in which the reaction is initiated by a 1:1 interaction between alpha-SNAP and the ternary SNARE complex, followed by NSF binding. Subsequent additional alpha-SNAP binding events may occur as part of a processive disassembly mechanism.
Macromolecular crystallography is typified by a "static" crystal structure that results from the experimental refinement. However, we posit that the single atomic model is a myopic representation of the statistical ensemble, and derive the Bayesian logic to illustrate a rigorous, probabilistic representation of the data using a model ensemble. Using ubiquitin and a complex of a viral entry protein bound to its cellular receptor as examples, we will show the utility of this method in describing biological function and protein dynamics in the context of X-ray crystallography that was previously not possible. Finally, a brief description of the software with which this is being accomplished will be presented.
Comparisons among evolutionarily related enzymes offer opportunities to reveal how structural differences produce different catalytic activities. Two structurally related enzymes, Escherichia coli alkaline phosphatase (AP) and Xanthomonas axonopodis nucleotide pyrophosphatase/phosphodiesterase (NPP), have nearly identical binuclear Zn2+ catalytic centers but show tremendous differential specificity for hydrolysis of phosphate monoesters or phosphate diesters. To determine if there are differences in Zn2+ coordination in the two enzymes that might contribute to catalytic specificity, we analyzed both x-ray absorption spectroscopic and x-ray crystallographic data. We report a 1.29-Å crystal structure of AP with bound phosphate, allowing evaluation of interactions at the AP metal site with high resolution. To make systematic comparisons between AP and NPP, we measured zinc extended x-ray absorption fine structure for AP and NPP in the free-enzyme forms, with AMP and inorganic phosphate ground-state analogs and with vanadate transition-state analogs. These studies yielded average zinc–ligand distances in AP and NPP free-enzyme forms and ground-state analog forms that were identical within error, suggesting little difference in metal ion coordination among these forms. Upon binding of vanadate to both enzymes, small increases in average metal–ligand distances were observed, consistent with an increased coordination number. Slightly longer increases were observed in NPP relative to AP, which could arise from subtle rearrangements of the active site or differences in the geometry of the bound vanadyl species. Overall, the results suggest that the binuclear Zn2+ catalytic site remains very similar between AP and NPP during the course of a reaction cycle.
Understanding the electrostatic forces and features within highly heterogeneous, anisotropic, and chemically complex enzyme active sites and their connection to biological catalysis remains a longstanding challenge, in part due to the paucity of incisive experimental probes of electrostatic properties within proteins. To quantitatively assess the landscape of electrostatic fields at discrete locations and orientations within an enzyme active site, we have incorporated site-specific thiocyanate vibrational probes into multiple positions within bacterial ketosteroid isomerase. A battery of X-ray crystallographic, vibrational Stark spectroscopy, and NMR studies revealed electrostatic field heterogeneity of 8 MV/cm between active site probe locations and widely differing sensitivities of discrete probes to common electrostatic perturbations from mutation, ligand binding, and pH changes. Electrostatic calculations based on active site ionization states assigned by literature precedent and computational pKa prediction were unable to quantitatively account for the observed vibrational band shifts. However, electrostatic models of the D40N mutant gave qualitative agreement with the observed vibrational effects when an unusual ionization of an active site tyrosine with a pKa near 7 was included. UV-absorbance and 13C NMR experiments confirmed the presence of a tyrosinate in the active site, in agreement with electrostatic models. This work provides the most direct measure of the heterogeneous and anisotropic nature of the electrostatic environment within an enzyme active site, and these measurements provide incisive benchmarks for further developing accurate computational models and a foundation for future tests of electrostatics in enzymatic catalysis.