Steroid hormones like estradiol and progesterone bind to nuclear hormone receptors and regulate gene transcription. They compete with pollutants and Endocrine Distrupting Chemicals (EDCs) like bisphenol A. When administered as medication, estradiol and progesterone are themselves considered EDCs. To allow modeling studies, parameters for the polarizable AMOEBA force field were derived here for these three molecules, along with the host molecule cyclodextrin and two phosphorylated forms of p-cresol and tyrosine. Indeed, phosphorylation of an estrogen receptor tyrosine regulates estradiol action. AMOEBA-optimized molecular structures were in good agreement with quantum mechanics, and interaction energies with water or ammonium were well-reproduced. Molecular dynamics simulations of crystal structures showed improved agreement with experiment over an additive force field; the mean relative error for unit cell volumes was reduced by more than two. A short MD simulation of the estradiol-estrogen receptor complex was done as a proof of principle and was structurally stable. Although additional simulations are desirable, the developed parameters should already be useful for studying hormone and EDC interactions with their hosts.
Our approach to computational protein design is physics-based. We develop a software called Proteus, allowing both physics-based energy evaluation and sequence-conformation exploration. Unlike knowledge-based models, a physics-based energy function facilitates the consideration of unusual chemical entities, such as novel ligands or transition states. Additionally, the adaptive landscape flattening method allows direct sampling on free energy difference between two states. We show here how these ingredients combined can benefit enzyme design. As an illustration, we revisit the stereospecificity inversion of tyrosyl-tRNA synthetase. Following a tutorial presentation, we explore various design criteria that can be related to enzyme experimental parameters. Our model is able to retrieve the native sequence when targeting L-tyrosine. Mutations predicted to favor D-tyrosine are proposed.
Protein kinases are key actors of signaling networks and important drug targets. They cycle between active and inactive conformations, distinguished by a few elements within the catalytic domain. One is the activation loop, whose conserved DFG motif can occupy DFG-in, DFG-out, and some rarer conformations. Annotation and classification of the structural kinome are important, as different conformations can be targeted by different inhibitors and activators. Valuable resources exist; however, large-scale applications will benefit from increased automation and interpretability of structural annotation. Interpretable machine learning models are described for this purpose, based on ensembles of decision trees. To train them, a set of catalytic domain sequences and structures was collected, somewhat larger and more diverse than existing resources. The structures were clustered based on the DFG conformation and manually annotated. They were then used as training input. Two main models were constructed, which distinguished active/inactive and in/out/other DFG conformations. They considered initially 1692 structural variables, spanning the whole catalytic domain, then identified ("learned") a small subset that sufficed for accurate classification. The first model correctly labeled all but 3 of 3289 structures as active or inactive, while the second assigned the correct DFG label to all but 17 of 8826 structures. The most potent classifying variables were all related to well-known structural elements in or near the activation loop and their ranking gives insights into the conformational preferences. The models were used to automatically annotate 3850 kinase structures predicted recently with the Alphafold2 tool, showing that Alphafold2 reproduced the active/inactive but not the DFG-in proportions seen in the Protein Data Bank. We expect the models will be useful for understanding and engineering kinases.
Specific amino acid (AA) binding by aminoacyl-tRNA synthetases (aaRSs) is necessary for correct translation of the genetic code. Sequence and structure analyses have revealed the main specificity determinants and allowed a partitioning of aaRSs into two classes and several subclasses. However, the information contributed by each determinant has not been precisely quantified, and other, minor determinants may still be unidentified. Growth of genomic data and development of machine learning classification methods allow us to revisit these questions. This work considered the subclass IIb, formed by the three enzymes aspartyl-, asparaginyl-, and lysyl-tRNA synthetase (LysRS). Over 35,000 sequences from the Pfam database were considered, and used to train a machine-learning model based on ensembles of decision trees. The model was trained to reproduce the existing classification of each sequence as AspRS, AsnRS, or LysRS, and to identify which sequence positions were most important for the classification. A few positions (5-8 depending on the AA substrate) sufficed for accurate classification. Most but not all of them were well-known specificity determinants. The machine learning models thus identified sets of mutations that distinguish the three subclass members, which might be targeted in engineering efforts to alter or swap the AA specificities for biotechnology applications.
Enzyme design is an important application of computational protein design (CPD). It can benefit enormously from the additional chemistries provided by noncanonical amino acids (ncAAs). These can be incorporated into an 'expanded' genetic code, and introduced in vivo into target proteins. The key step for genetic code expansion is to engineer an aminoacyl-transfer RNA (tRNA) synthetase (aaRS) and an associated tRNA that handles the ncAA. Experimental directed evolution has been successfully used to engineer aaRSs and incorporate over 200 ncAAs into expanded codes. But directed evolution has severe limits, and is not yet applicable to noncanonical AA backbones. CPD can help address several of its limitations, and has begun to be applied to this problem. We review efforts to redesign aaRSs, studies that designed new proteins and functionalities with the help of ncAAs, and some of the method developments that have been used, such as adaptive landscape flattening Monte Carlo, which allows an enzyme to be redesigned with substrate or transition state binding as the design target.
Amino acids (AAs) with a noncanonical backbone would be a valuable tool for protein engineering, enabling new structural motifs and building blocks. To incorporate them into an expanded genetic code, the first, key step is to obtain an appropriate aminoacyl-tRNA synthetase. Currently, directed evolution is not available to optimize AAs with noncanonical backbones, since an appropriate selective pressure has not been discovered. Computational protein design (CPD) is an alternative. We used a new CPD method to redesign MetRS and increase its activity towards β-Met, which has an extra backbone methylene. The new method considered a few active site positions for design and used a Monte Carlo exploration of the corresponding sequence space. During the exploration, a bias energy was adaptively learned, such that the free energy landscape of the apo enzyme was flattened. Enzyme variants could then be sampled, in the presence of the ligand and the bias energy, according to their β-Met binding affinities. Eighteen predicted variants were chosen for experimental testing; 10 exhibited detectable activity for β-Met adenylation. Top predicted hits were characterized experimentally in detail. Dissociation constants, catalytic rates, and Michaelis constants for both α-Met and β-Met were measured. The best mutant retained a preference for α-Met over β-Met; however, the preference was reduced, compared to the wildtype, by a factor of 29. For this mutant, high resolution crystal structures were obtained in complex with both α-Met and β-Met, indicating that the predicted, active conformation of β-Met in the active site was retained.
Methionine γ-lyase (MGL) breaks down methionine, with the help of its cofactor pyridoxal-5'-phosphate (PLP), or vitamin B6. Methionine depletion is damaging for cancer cells but not normal cells, so MGL is of interest as a therapeutic protein. To increase our understanding and help engineer improved activity, we focused on the reactive, Michaelis complex M between MGL, covalently bound PLP, and substrate Met. M is not amenable to crystallography, as it proceeds to products. Experimental activity measurements helped exclude a mechanism that would bypass M. We then used molecular dynamics and alchemical free energy simulations to elucidate its structure and dynamics. We showed that the PLP phosphate has a pKa strongly downshifted by the protein, whether Met is present or not. Met binding affects the structure surrounding the reactive atoms. With Met, the Schiff base linkage between PLP and a nearby lysine shifts from a zwitterionic, keto form to a neutral, enol form that makes it easier for Met to approach its labile, target atom. The Met ligand also stabilizes the correct orientation of the Schiff base, more strongly than in simulations without Met, and in agreement with structures in the Protein Data Bank, where the Schiff base orientation correlates with the presence or absence of a co-bound anion or substrate analogue in the active site. Overall, the Met ligand helps organize the active site for the enzyme reaction by reducing fluctuations and shifting protonation states and conformational populations.
In response to antibiotics that inhibit a bacterial enzyme, resistance mutations inevitably arise. Predicting them ahead of time would aid target selection and drug design. The simplest resistance mechanism would be to reduce antibiotic binding without sacrificing too much substrate binding. The property that reflects this is the enzyme “vitality”, defined here as the difference between the inhibitor and substrate binding free energies. To predict such mutations, we borrow methodology from computational protein design. We use a Monte Carlo exploration of mutation space and vitality changes, allowing us to rank thousands of mutations and identify ones that might provide resistance through the simple mechanism considered. As an illustration, we chose dihydrofolate reductase, an essential enzyme targeted by several antibiotics. We simulated its complexes with the inhibitor trimethoprim and the substrate dihydrofolate. 20 active site positions were mutated, or “redesigned” individually, then in pairs or quartets. We computed the resulting binding free energy and vitality changes. Out of seven known resistance mutations involving active site positions, five were correctly recovered. Ten positions exhibited mutations with significant predicted vitality gains. Direct couplings between designed positions were predicted to be small, which reduces the combinatorial complexity of the mutation space to be explored. It also suggests that over the course of evolution, resistance mutations involving several positions do not need the underlying point mutations to arise all at once: they can appear and become fixed one after the other.
Pyridoxal-5′-phosphate (PLP) is a cofactor in the reactions of over 160 enzymes, several of which are implicated in diseases. Methionine γ-lyase (MGL) is of interest as a therapeutic protein for cancer treatment. It binds PLP covalently through a Schiff base linkage and digests methionine, whose depletion is damaging for cancer cells but not normal cells. To improve MGL activity, it is important to understand and engineer its PLP binding. We develop a simulation model for MGL, starting with force field parameters for PLP in four main states: two phosphate protonation states and two tautomeric states, keto or enol for the Schiff base moiety. We used the force field to simulate MGL complexes with each form, and showed that those with a fully-deprotonated PLP phosphate, especially keto, led to the best agreement with MGL structures in the PDB. We then confirmed this result through alchemical free energy simulations that compared the keto and enol forms, confirming a moderate keto preference, and the fully-deprotonated and singly-protonated phosphate forms. Extensive simulations were needed to adequately sample conformational space, and care was needed to extrapolate the protonation free energy to the thermodynamic limit of a macroscopic, dilute protein solution. The computed phosphate pK a was 5.7, confirming that the deprotonated, −2 form is predominant. The PLP force field and the simulation methods can be applied to all PLP enzymes and used, as here, to reveal fine details of structure and dynamics in the active site.
Physics and physical chemistry are an important thread in computational protein design, complementary to knowledgebased tools. They provide molecular mechanics scoring functions that need little or no ad hoc parameter readjustment, methods to thoroughly sample equilibrium ensembles, and different levels of approximation for conformational flexibility. They led recently to the successful redesign of a small protein using a physics-based folded state energy. Adaptive Monte Carlo or molecular dynamics schemes were discovered where protein variants are populated as per their ligand-binding free energy or catalytic efficiency. Molecular dynamics have been used for backbone flexibility. Implicit solvent models have been refined, polarizable force fields applied, and many physical insights obtained.
The design of proteins and miniproteins is an important challenge. Designed variants should be stable, meaning the folded/unfolded free energy difference should be large enough. Thus, the unfolded state plays a central role. An extended peptide model is often used, where side chains interact with solvent and nearby backbone, but not each other. The unfolded energy is then a function of sequence composition only and can be empirically parametrized. If the space of sequences is explored with a Monte Carlo procedure, protein variants will be sampled according to a well-defined Boltzmann probability distribution. We can then choose unfolded model parameters to maximize the probability of sampling native-like sequences. This leads to a well-defined maximum likelihood framework. We present an iterative algorithm that follows the likelihood gradient. The method is presented in the context of our Proteus software, as a detailed downloadable tutorial. The unfolded model is combined with a folded model that uses molecular mechanics and a Generalized Born solvent. It was optimized for three PDZ domains and then used to redesign them. The sequences sampled are native-like and similar to a recent PDZ design study that was experimentally validated.
This chapter describes two computational methods for PDZ-peptide binding: high-throughput computational protein design (CPD) and a medium-throughput approach combining molecular dynamics for conformational sampling with a Poisson-Boltzmann (PB) Linear Interaction Energy for scoring. A new CPD method is outlined, which uses adaptive Monte Carlo simulations to efficiently sample peptide variants that tightly bind a PDZ domain, and provides at the same time precise estimates of their relative binding free energies. A detailed protocol is described based on the Proteus CPD software. The mediumthroughput approach can be performed with standard MD and PB software, such as NAMD and Charmm. For 40 complexes between Tiam1 and peptide ligands, it gave high a2ccuracy, with mean errors of around 0.5 kcal/mol for relative binding free energies and no large errors. It requires a moderate amount of parameter fitting before it can be applied, and its transferability to other protein families is still untested.
In this review, applications of various molecular modelling methods in the study of estrogens and xenoestrogens are summarized. Selected biomolecules that are the most commonly chosen as molecular modelling objects in this field are presented. In most of the reviewed works, ligand docking using solely force field methods was performed, employing various molecular targets involved in metabolism and action of estrogens. Other molecular modelling methods such as molecular dynamics and combined quantum mechanics with molecular mechanics have also been successfully used to predict the properties of estrogens and xenoestrogens. Among published works, a great number also focused on the application of different types of quantitative structure–activity relationship (QSAR) analyses to examine estrogen’s structures and activities. Although the interactions between estrogens and xenoestrogens with various proteins are the most commonly studied, other aspects such as penetration of estrogens through lipid bilayers or their ability to adsorb on different materials are also explored using theoretical calculations. Apart from molecular mechanics and statistical methods, quantum mechanics calculations are also employed in the studies of estrogens and xenoestrogens. Their applications include computation of spectroscopic properties, both vibrational and Nuclear Magnetic Resonance (NMR), and also in quantum molecular dynamics simulations and crystal structure prediction. The main aim of this review is to present the great potential and versatility of various molecular modelling methods in the studies on estrogens and xenoestrogens.
We describe methods and software for physics-based protein design. The folded state energy combines molecular mechanics with Generalized Born solvent. Sequence and conformation space are sampled with Replica Exchange Monte Carlo, assuming one or a few fixed protein backbone structures and discrete side chain rotamers. Whole protein design and enzyme design are presented as illustrations. Full redesign of three PDZ domains was done using a simple, empirical, unfolded state model. Designed sequences were very similar to natural ones. Enzyme redesign exploited a powerful, adaptive, importance sampling approach that allows the design to directly target substrate binding, reaction rate, catalytic efficiency, or the specificity of these properties. Redesign of tyrosyl-tRNA synthetase stereospecificity is reported as an example.
Designed enzymes are of fundamental and technological interest. Experimental directed evolution still has significant limitations, and computational approaches are a complementary route. A designed enzyme should satisfy multiple criteria: stability, substrate binding, transition state binding. Such multi-objective design is computationally challenging. Two recent studies used adaptive importance sampling Monte Carlo to redesign proteins for ligand binding. By first flattening the energy landscape of the apo protein, they obtained positive design for the bound state and negative design for the unbound. We have now extended the method to design an enzyme for specific transition state binding, i.e., for its catalytic power. We considered methionyl-tRNA synthetase (MetRS), which attaches methionine (Met) to its cognate tRNA, establishing codon identity. Previously, MetRS and other synthetases have been redesigned by experimental directed evolution to accept noncanonical amino acids as substrates, leading to genetic code expansion. Here, we have redesigned MetRS computationally to bind several ligands: the Met analog azidonorleucine, methionyl-adenylate (MetAMP), and the activated ligands that form the transition state for MetAMP production. Enzyme mutants known to have azidonorleucine activity were recovered by the design calculations, and 17 mutants predicted to bind MetAMP were characterized experimentally and all found to be active. Mutants predicted to have low activation free energies for MetAMP production were found to be active and the predicted reaction rates agreed well with the experimental values. We suggest the present method should become the paradigm for computational enzyme design.
Computational protein design relies on simulations of a protein structure, where selected amino acids can mutate randomly, and mutations are selected to enhance a target property, such as stability. Often, the protein backbone is held fixed and its degrees of freedom are modeled implicitly to reduce the complexity of the conformational space. We present a hybrid method where short molecular dynamics (MD) segments are used to explore conformations and alternate with Monte Carlo (MC) moves that apply mutations to side chains. The backbone is fully flexible during MD. As a test, we computed side chain acid/base constants or pKa's in five proteins. This problem can be considered a special case of protein design, with protonation/deprotonation playing the role of mutations. The solvent was modeled as a dielectric continuum. Due to cost, in each protein we allowed just one side chain position to change its protonation state and the other position to change its type or mutate. The pKa's were computed with a standard method that scans a range of pH values and with a new method that uses adaptive landscape flattening (ALF) to sample all protonation states in a single simulation. The hybrid method gave notably better accuracy than standard, fixed-backbone MC. ALF decreased the computational cost a factor of 13.
Computational protein design (CPD) can address the inverse folding problem, exploring a large space of sequences and selecting ones predicted to fold. CPD was used previously to redesign several proteins, employing a knowledge-based energy function for both the folded and unfolded states. We show that a PDZ domain can be entirely redesigned using a "physics-based" energy for the folded state and a knowledge-based energy for the unfolded state. Thousands of sequences were generated by Monte Carlo simulation. Three were chosen for experimental testing, based on their low energies and several empirical criteria. All three could be overexpressed and had native-like circular dichroism spectra and 1D-NMR spectra typical of folded structures. Two had upshifted thermal denaturation curves when a peptide ligand was present, indicating binding and suggesting folding to a correct, PDZ structure. Evidently, the physical principles that govern folded proteins, with a dash of empirical post-filtering, can allow successful whole-protein redesign.
We describe methods for physics-based protein design and some recent applications from our work. We present the physical interpretation of a MC simulation in sequence space and show that sequences and conformations form a well-defined statistical ensemble, explored with Monte Carlo and Boltzmann sampling. The folded state energy combines molecular mechanics for solutes with continuum electrostatics for solvent. We usually assume one or a few fixed protein backbone structures and discrete side chain rotamers. Methods based on molecular dynamics, which introduce additional backbone and side chain flexibility, are under development. The redesign of a PDZ domain and an aminoacyl-tRNA synthetase enzyme were successful. We describe a versatile, adaptive, Wang-Landau MC method that can be used to design for substrate affinity, catalytic rate, catalytic efficiency, or the specificity of these properties. The methods are transferable to all biomolecules, can be systematically improved, and give physical insights.
PDZ (PSD-95/Dlg/Zo-1) domains are small, structurally conserved protein-protein interaction modules (∼90 amino acids), which consist of six beta-strands and two alpha-helices. They are highly represented in the human proteome and regulate the spatial and temporal function of proteins or signaling complexes, such as synaptic transmembrane and cell adhesion proteins. PDZ domains selectively interact with linear peptide motifs known as PDZ-binding-motifs of their binding partner proteins and are associated with several human disorders. Here, we describe two approaches towards designing non-native, artificial PDZ domains. The first approach uses molecular dynamics (MD) methods to identify diverse amino acid sequences that fold into a desired PDZ domain structure. Two template structures were selected and candidate design sequences that remained stable in MD simulations were experimentally probed for their thermal stability and overall structure using DSF, CD and NMR. These experiments indicate that the candidate designs ranged from molten globular intermediates to native-like folded proteins. The second approach relies on using the consensus amino acid sequence of PDZ domains to design artificial proteins. Three candidate PDZ designs were obtained representing the full alignment or a subset representing their typical classification (i.e. Full, Class I and Class II). The three artificial PDZ domains were soluble and stable towards thermal and chemical denaturation. NMR and CD data indicate that they contain regular secondary structure and are folded. The crystal structure of the Class II PDZ domain design was determined and it was found to be a domain swapped dimer. Finally, the designed PDZ domains had a broad range of affinities (Kd ∼25 to >250 μM) for a small set of native ligands. These studies show promise for the design of artificial PDZ domains.
The association of Mg(2+) and H2 PO4 (-) in water can give insights into Mg:phosphate interactions in general, which are very widespread, but for which experimental data is surprisingly sparse. It is studied through molecular dynamics simulations (>100 ns) by using the polarizable AMOEBA force field, and the association free energy is computed for the first time. Explicit consideration of outer-sphere and two types of inner-sphere association provides considerable insight into the dynamics and thermodynamics of ion pairing. After careful assessment of the computational approximations, the agreement with experimental values indicates that the methodology can be extended to other inorganic and biological Mg:phosphate interactions in solution.