Cyclic peptides are an important class of pharmaceutical drugs. We used replica-exchange molecular dynamics (REMD) and simulated tempering (ST) simulations to explore the conformational landscape of a set of nine cyclic peptides. The N-ter to C-ter backbone-cyclized peptides of 7-10 residues were previously designed for high conformational stability with a mixture of l- and d-amino acids. Their experimental NMR structures are available in the protein data bank (PDB). For each peptide, we tested several force fields, namely, Amber96, Amber14, RSFF2C, and Charmm36m in implicit and explicit solvents. We find that the variability of the free energy maps obtained from several protocols is larger than the variability obtained by just repeating the same protocol. Running multiple protocols is therefore important for the convergence assessment of REMD or ST simulations. The majority of the free energy maps showed clusters with a high RMSD compared to the native structures, revealing the residual flexibility of this set of cyclic peptides. The high RMSD clusters had in some cases the lowest free energy, rendering the prediction of the native structure more difficult with a single protocol. Fortunately, the combination of four implicit solvent REMD and ST simulations, mixing the Amber96 and Amber14 force fields, predicted robustly the native structure. As implicit solvent simulations in the REMD or ST setup are up to one hundred times faster than explicit solvent simulations, running four implicit solvent simulations is a good practical choice. We checked that the use of an explicit solvent REMD or ST simulation, taken alone or combined with implicit solvent simulations, did not significantly improve our results. It results in our combination of four implicit solvent simulations being tied in terms of success rate with much more expensive combinations that include explicit solvent simulations. This may be used as a guideline for further studies of cyclic peptide conformations.
Evaluation of the structural perturbations introduced by a single amino acid mutation is the main issue for protein structural biology. We propose here to present some recent advances in methods, allowing the splitting of distortion between the actual substitution effect and the contribution of the local flexibility of the position where the mutation occurs. Its main drawback is the need of many structures with a single mutation in each of them. To bypass this difficulty, we propose to use molecular modeling tools, with several software enabling us to build a model from a template, given the sequence. As a proof of concept, we rely on a gold standard, the human lysozyme. Both wild-type and three mutant structures are available in the PDB. Two of these mutations result in amyloid fibril formation, and the last one is neutral. As a conclusion, irrespective of the algorithm used for modeling, side chain conformations at the site of mutation are reliable, although long-range effects are out of reach of these tools.
Folds are the architecture and topology of a protein domain. Categories of folds are very few compared to the astronomical number of sequences. Eukaryotes have more protein folds than Archaea and Bacteria. These folds are of two types: shared with Archaea and/or Bacteria on one hand and specific to eukaryotic clades on the other hand. The first kind of folds is inherited from the first endosymbiosis and confirms the mixed origin of eukaryotes. In a dataset of 1073 folds whose presence or absence has been evidenced among 210 species equally distributed in the three super-kingdoms, we have identified 28 eukaryotic folds unambiguously inherited from Bacteria and 40 eukaryotic folds unambiguously inherited from Archaea. Compared to previous studies, the repartition of informational function is higher than expected for folds originated from Bacteria and as high as expected for folds inherited from Archaea. The second type of folds is specifically eukaryotic and associated with an increase of new folds within eukaryotes distributed in particular clades. Reconstructed ancestral states coupled with dating of each node on the tree of life provided fold appearance rates. The rate is on average twice higher within Eukaryota than within Bacteria or Archaea. The highest rates are found in the origins of eukaryotes, holozoans, metazoans, metazoans stricto sensu, and vertebrates: the roots of these clades correspond to bursts of fold evolution. We could correlate the functions of some of the fold synapomorphies within eukaryotes with significant evolutionary events. Among them, we find evidence for the rise of multicellularity, adaptive immune system, or virus folds which could be linked to an ecological shift made by tetrapods.
GTPases constitute a superclass of proteins with a common fold. Five specific G motifs located in loops are signatures of this superclass. Nevertheless, some proteins may share the fold of the small GTPases, although their functions are totally unrelated. To retrieve them, we specifically searched in the BLAST output listings for non-GTPases with available 3D structure, starting from a canonical GTPase sequence as a query. We then performed both a sequence analysis by means of HCA and a structural comparison with an established GTPase. It results that, although sequence identity is in the twilight zone, i.e. below 25%, one can evidence some conservations of the catalytic motifs. Nevertheless, mutations have occurred that produced a new function while the global fold is maintained. We discuss whether non-GTPases presumably originated from a common ancestor with an ancient G domain. The evolutionary mechanisms relating non-GTPases to GTPases that we can advance are sequence divergence, convergence, and DNA recombination. We conclude that the most probable evolutionary pathway leading to such structural similarities is that all the studied proteins must have evolved by sequence divergence from a primordial GTP-binding domain.
Several studies showed that folds (topology of protein secondary structures) distribution in proteomes may be a global proxy to build phylogeny. Then, some folds should be synapomorphies (derived characters exclusively shared among taxa). However, previous studies used methods that did not allow synapomorphy identification, which requires congruence analysis of folds as individual characters. Here, we map SCOP folds onto a sample of 210 species across the tree of life (TOL). Congruence is assessed using retention index of each fold for the TOL, and principal component analysis for deeper branches. Using a bicluster mapping approach, we define synapomorphic blocks of folds (SBF) sharing similar presence/absence patterns. Among the 1232 folds, 20% are universally present in our TOL, whereas 54% are reliable synapomorphies. These results are similar with CATH and ECOD databases. Eukaryotes are characterized by a large number of them, and several SBFs clearly support nested eukaryotic clades (divergence times from 1100 to 380 mya). Although clearly separated, the three superkingdoms reveal a strong mosaic pattern. This pattern is consistent with the dual origin of eukaryotes and witness secondary endosymbiosis in their phothosynthetic clades. Our study unveils direct analysis of folds synapomorphies as key characters to unravel evolutionary history of species.
Chapter 2 Impact of a Point Mutation in a Protein Structure Mathilde Carpentier, Mathilde CarpentierSearch for more papers by this authorJacques Chomilier, Jacques ChomilierSearch for more papers by this author Mathilde Carpentier, Mathilde CarpentierSearch for more papers by this authorJacques Chomilier, Jacques ChomilierSearch for more papers by this author Book Editor(s):Philippe Grandcolas, Philippe GrandcolasSearch for more papers by this authorMarie-Christine Maurel, Marie-Christine MaurelSearch for more papers by this author First published: 02 April 2021 https://doi.org/10.1002/9781119476870.ch2 AboutPDFPDF ToolsRequest permissionExport citationAdd to favoritesTrack citation ShareShareShare a linkShare onEmailFacebookTwitterLinkedInRedditWechat Summary Proteins are involved in most cellular functions at all levels, from DNA duplication to chemical metabolism, cell structuring and signal transmission. "Water-soluble" proteins fold into a compact globular form (unlike fibrous, membrane and "disordered" proteins). The hydrophobic nature of certain amino acids makes this compact folding necessary. Proteins are said to be "marginally stable": typically, there is a difference of 3–7 kcal/mol in free folding energy between folded and unfolded conformations. The substitution of one amino acid by another (a mutation) is one of the fundamental events of molecular evolution, with variable consequences in proteins. In order to determine whether Root-Mean-Square Deviations tend to be larger in terms of mutations, it is necessary to be able to compare several structures that only differ from each other by a single point mutation. 2.8. References Alberts, B., Bray, D., Lewis, J., Raff, M., Toberts, K., and Watson, J. (1994). Molecular Biology of the Cell. Garland Publishing, New York. Anfinsen, C.B. (1973). Principles that govern the folding of protein chains. Science, 181, 223–230. Berman, H.M., Westbrook, J., Feng, Z., Gilliland, G., Bhat, T.N., Weissig, H., Shindyalov, I.N., and Bourne, P.E. (2000). The Protein Data Bank. Nucleic Acids Research, 28, 235–242. Bloom, J.D., Raval, A., and Wilke, C.O. (2007). Thermodynamics of neutral protein evolution [Online]. Genetics, 175, 255–266. Available: 10.1534/genetics.106.061754. Bordner, A.J. and Abagyan, R.A. (2004). Large-scale prediction of protein geometry and stability changes for arbitrary single point mutations [Online]. Proteins, 57, 400–413. Available: 10.1002/prot.20185. Bressler, S. and Talmud, D. (1944). On the nature of globular proteins. Comptes rendus de l'Académie des sciences de l'URSS, 43, 310–314. Davis, I.W., Arendall, W.B., Richardson, D.C., and Richardson, J.S. (2006). The backrub motion: How protein backbone shrugs when a sidechain dances [Online]. Structure, 14, 265–274. Available: 10.1016/j.str.2005.10.007. DePristo, M.A., Weinreich, D.M., and Hartl, D.L. (2005). Missense meanderings in sequence space: A biophysical view of protein evolution [Online]. Nature Reviews Genetics, 6, 678–687. Available: 10.1038/nrg1672. Dunbrack, R.L. (2002). Rotamer libraries in the 21st century [Online]. Current Opinion in Structural Biology, 12, 431–440. Available: 10.1016/S0959-440X(02)00344-5. Gong, S., Worth, C.L., Bickerton, G.R.J., Lee, S., Tanramluk, D., and Blundell, T.L. (2009). Structural and functional restraints in the evolution of protein families and superfamilies [Online]. Biochemical Society Transactions, 37, 727–733. Available: 10.1042/BST0370727. Gromiha, M.M. and Sarai, A. (2010). Thermodynamic database for proteins: Features and applications [Online]. Methods in Molecular Biology, 609, 97–112. Available: 10.1007/978-1-60327-241-4_6. Guerois, R., Nielsen, J.E., and Serrano, L. (2002). Predicting changes in the stability of proteins and protein complexes: A study of more than 1000 mutations [Online]. Journal of Molecular Biology, 320, 369–387. Available: 10.1016/S0022-2836(02)00442-4. Guo, H.H., Choe, J., and Loeb, L.A. (2004). Protein tolerance to random amino acid change [Online]. Proceedings of the National Academy of Sciences, 101, 9205–9210. Available: 10.1073/pnas.0403255101. Kellogg, E.H., Leaver-Fay, A., and Baker, D. (2011). Role of conformational sampling in computing mutation-induced changes in protein structure and stability [Online]. Proteins, 79, 830–838. Available: 10.1002/prot.22921. Kosloff, M. and Kolodny, R. (2008). Sequence-similar, structure-dissimilar protein pairs in the PDB [Online]. Proteins, 71, 891–902. Available: 10.1002/prot.21770. Lauck, F., Smith, C.A., Friedland, G.F., Humphris, E.L., and Kortemme, T. (2010). RosettaBackrub – A web server for flexible backbone protein structure modeling and design [Online]. Nucleic Acids Research, 38, W569–W575. Available: 10.1093/nar/gkq369. Lonquety, M., Lacroix, Z., Papandreou, N., and Chomilier, J. (2009). SPROUTS: A database for the evaluation of protein stability upon point mutation [Online]. Nucleic Acids Research, 37, D374-9. Available: 10.1093/nar/gkn704. Luzzati, V. (1952). Traitement statistique des erreurs dans la determination des structures cristallines [Online]. Acta Crystallographica, 5, 802–810. Available: 10.1107/S0365110X52002161. Religa, T.L., Markson, J.S., Mayor, U., Freund, S.M.V., and Fersht, A.R. (2005). Solution structure of a protein denatured state and folding intermediate [Online]. Nature, 437, 1053–1056. Available: 10.1038/nature04054. Sander, C. and Schneider, R. (1991). Database of homology-derived protein structures and the structural meaning of sequence alignment [Online]. Proteins, 9, 56–68. Available: 10.1002/prot.340090107. Schaefer, C. and Rost, B. (2012). Predict impact of single amino acid change upon protein structure [Online]. BMC Genomics, 13, S4. Available: 10.1186/1471-2164-13-S4-S4. Shakhnovich, E.I. and Gutin, A.M. (1991). Influence of point mutations on protein structure: Probability of a neutral mutation. J. Theor. Biol., 149, 537–546. Shanthirabalan, S., Chomilier, J., and Carpentier, M. (2018). Structural effects of point mutations in proteins [Online]. Proteins: Structure, Function, and Bioinformatics, 86, 853–867. Available: 10.1002/prot.25499. Stryer, L. (1994). Biochemistry. W.H. Freeman, New York. Studer, R.A., Dessailly, B.H., and Orengo, C.A. (2013). Residue mutations and their impact on protein structure and function: Detecting beneficial and pathogenic changes [Online]. Biochemical Journal, 449, 581–594. Available: 10.1042/BJ20121221. Taverna, D.M. and Goldstein, D. (2002). Why are proteins marginally stable? [Online]. Proteins, 46(1), 105–109. Available: 10.1002/prot.10016. Zeldovich, K.B., Chen, P., and Shakhnovich, E.I. (2007). Protein stability imposes limits on organism complexity and speed of molecular evolution [Online]. Proceedings of the National Academy of Sciences, 104, 16152–16157. Available: 10.1073/pnas.0705366104. Zhou, R., Eleftheriou, M., Royyuru, A.K., and Berne, B.J. (2007). Destruction of long-range interactions by a single mutation in lysozyme [Online]. Proceedings of the National Academy of Sciences, 104, 5824–5829. Available: 10.1073/pnas.0701249104. Systematics and the Exploration of Life ReferencesRelatedInformation
Ferredoxin I and II are proteins carrying a specific ligand—an iron-sulfur cluster—which allows transport of electrons. These two classes of ferredoxin in their monomeric and dimeric forms are the object of this work. Characteristic of hydrophobic core in both molecules is analyzed via fuzzy oil drop model (FOD) to show the specificity of their structure enabling the binding of a relatively large ligand and formation of the complex. Structures of FdI and FdII are a promising example for the discussion of influence of hydrophobicity on biological activity but also for an explanation how FOD model can be used as an initial stage adviser (or a scoring function) in the search for locations of ligand binding pockets and protein–protein interaction areas. It is shown that observation of peculiarities in the hydrophobicity distribution present in the molecule (in this case—of a ferredoxin) may provide a promising starting location for computer simulations aimed at the prediction of quaternary structure of proteins.
The effects of a single residue substitution on the protein backbone are frequently quite small and there are many other potential sources of structural variation for protein. We present here a methodology considering different sources of distortions in order to isolate the very effect of the mutation. To validate our methodology, we consider a well-studied family with many single mutants: the human lysozyme. Most of the perturbations are expected to be at the very localisation of the mutation, but in many cases the effects are propagated at long range. We show that the distances between the mutated residue and the 5% most disturbed residues exponentially decreases. One third of the affected residues are in direct contact with the mutated position; the remaining two thirds are potential allosteric effects. We confirm the reliability of the residues identified as significantly perturbed by comparing our results to experimental studies. We confirm with the present method all the previously identified perturbations. This study shows that mutations have long-range impact on protein backbone that can be detected, although the displacement of the affected atoms is small.
ABSTRACT Facing the huge increase of information about proteins, classification has reached the level of a compulsory task, essential for assigning a function to a given sequence, by means of comparison to existing data. Multiple sequence alignment programs have been proven to be very useful and they have already been evaluated. In this paper we wished to evaluate the added value provided by taking into account structures. We compared the multiple alignments resulting from 24 programs, either based on sequence, structure, or both, to reference alignments deposited in five databases. Reference databases, on their side, can be split in two: more automatic ones, and more manually ones. Scores have been attributed to each program. As a global rule of thumb, five groups of methods emerge, with the lead to two of the structure-based programs. This advantage is increased at low levels of sequence identity among aligned proteins, or for residues in regular secondary structures or buried. Concerning gap management, sequence-based programs place less gaps than structure-based programs. Concerning the databases, the alignments from the manually built databases are the more challenging for the programs.
Evaluating the model quality of protein structures that evolve in environments with particular physicochemical properties requires scoring functions that are adapted to their specific residue compositions and/or structural characteristics. Thus, computational methods developed for structures from the cytosol cannot work properly on membrane or secreted proteins. Here, we present MyPMFs, an easy-to-use tool that allows users to train statistical potentials of mean force (PMFs) on the protein structures of their choice, with all parameters being adjustable. We demonstrate its use by creating an accurate statistical potential for transmembrane protein domains. We also show its usefulness to study the influence of the physical environment on residue interactions within protein structures. Our open-source software is freely available for download at https://github.com/bibip-impmc/mypmfs.
Small cyclic peptides represent a promising class of therapeutic molecules with unique chemical properties. However, the poor knowledge of their structural characteristics makes their computational design and structure prediction a real challenge. In order to better describe their conformational space, we developed a method, named EGSCyP, for the exhaustive exploration of the energy landscape of small head-to-tail cyclic peptides. The method can be summarized by (i) a global exploration of the conformational space based on a mechanistic representation of the peptide and the use of robotics-based algorithms to deal with the closure constraint and (ii) an all-atom refinement of the obtained conformations. EGSCyP can handle D-form residues and N-methylations. Two strategies for the side-chains placement were implemented and compared. To validate our approach, we applied it to a set of three variants of cyclic RGDFV pentapeptides, including the drug candidate Cilengitide. A comparative analysis was made with respect to replica exchange molecular dynamics simulations in implicit solvent. Its results show that the EGSCyP method provides a very complete characterization of the conformational space of small cyclic pentapeptides.
Understanding the folding, function and evolution of a protein requires separating its whole macromolecular structure into simpler constituent parts, which can be studied independently. Thus, proteins are most often decomposed into structural domains, considered for long as evolutionary building blocks, and secondary structures. However, because defining a proper level of analysis is a crucial challenge, authors have regularly proposed other descriptors of the protein architecture, with intermediate sizes between domains and secondary structures, such as the intrinsically stable ‘supersecondary structures’ of Rossmann. The underlying assumption, in part of the literature, is that contemporary proteins might result from the accretion of short peptide ancestors (Alva, S€ oding, & Lupas, 2015). Thus, a thorough analysis of sequence and structure comparison among folds should permit to decipher these ancestral fragments. Other types of structural subdomain units are defined either as connected regular secondary structure elements (from a static analysis of domains) or by using criteria based on folding, evolution (Nepomnyachiy, Ben-Tal, & Kolodny, 2017) or inter-residue contacts. The compactness of common globular protein folds requires that the amino acid polymer chain returns back to itself, forming loop-like trajectories, with ‘closures’ (defined as the Ca–Ca distance between the ends of the loop) typically below 10Å. This has led to the definition of another class of structural subdomains, which are named either ‘closed loops’ (Berezovsky, Grosberg, & Trifonov, 2000; Ittah & Haas, 1995) or ‘tightened end fragments’ (TEFs) (Lamarine, Mornon, Berezovsky, & Chomilier, 2001). In this article, we will use the two names interchangeably. Closed loops are not usual loops linking two successive secondary structure elements, but larger subdomain fragments, between 10 and 100 residues in size, which may contain partial or entire secondary structure elements. The relevance of closed loops as descriptors of protein architecture is supported by the fact that highly conserved hydrophobic positions in protein sequence alignments statistically match the terminal residues of TEFs (Lamarine et al., 2001). This finding has been later confirmed by Reynolds and coworkers, who define closed loops based on insertions and deletions in homologous proteins (Chintapalli et al., 2010). Interestingly, despite a different strategy for delimiting closed loops, their results are very similar to those of Berezovsky and co-workers, which indicates how robust the concept of closed loop is. Finally, Reynolds and co-workers have demonstrated the usefulness of TEFs as a means of structural analysis, by using their decomposition method for interpreting nuclear magnetic resonance (NMR) data of the unfolding of three proteins (cytochrome c, cytochrome b562 and triosephosphate isomerase). Although the pertinence of fragmenting protein structures into closed loops is established, there remains a problem regarding the power of the methods developed so far, in the sense that they tend to leave large parts of the protein chain unassigned (Figure 1). Such a limited sequence coverage not only leads to a biased and incomplete analysis of protein structures but also incompatible with the current hypothesis of closed loops as ubiquitous building blocks of globular proteins. To address this critical issue, we propose here TEF 2.0, a new computational method for decomposing protein structures into closed loops. Unlike its previous version (TEF 1.0) (Lamarine et al., 2001), which only maximized the proximity between the two terminal positions of the fragment, TEF 2.0 additionally maximizes the sequence coverage. To optimize these two criteria, we have introduced another novelty in TEF 2.0, by using a graph representation of the problem. Thanks to this approach, TEF 2.0 can rapidly find the optimal decomposition into closed loops (i.e. with the tightest closure and the largest coverage) among the numerous partitioning solutions. To validate our method, we compare the results produced by TEF 2.0 with those obtained by two other algorithms aimed at delineating TEFs. Moreover, we explore the scope of the method by studying the dependence of the sequence coverage on protein size and shape. Finally, we
Abstract The relation between distribution of hydrophobic amino acids along with protein chains and their structure is far from being completely understood. No reliable method allows ab initio prediction of the folded structure from this distribution of physicochemical properties, even when they are highly degenerated by considering only two classes: hydrophobic and polar. Establishment of long-range hydrophobic three dimension (3D) contacts is essential for the formation of the nucleus, a key process in the early steps of protein folding. Thus, a large number of 3D simulation studies were developed to challenge this issue. They are nowadays evaluated in a specific chapter of the molecular modeling competition, Critical Assessment of Protein Structure Prediction. We present here a simulation of the early steps of the folding process for 850 proteins, performed in a discrete 3D space, which results in peaks in the predicted distribution of intra-chain noncovalent contacts. The residues located at these peak positions tend to be buried in the core of the protein and are expected to correspond to critical positions in the sequence, important both for folding and structural (or similarly, energetic in the thermodynamic hypothesis) stability. The degree of stabilization or destabilization due to a point mutation at the critical positions involved in numerous contacts is estimated from the calculated folding free energy difference between mutated and native structures. The results show that these critical positions are not tolerant towards mutation. This simulation of the noncovalent contacts only needs a sequence as input, and this paper proposes a validation of the method by comparison with the prediction of stability by well-established programs.
A structural database of 11 families of chains differing by a single amino acid substitution has been built. Another structural dataset of 5 families with identical sequences has been used for comparison. The RMSD computed after a global superimposition of the mutated protein on each native one is smaller than the RMSD calculated among proteins of identical sequences. The effect of the perturbation is very local, and not necessarily the highest at the position of the mutation. A RMSD between mutated and native proteins is computed over a 3-residue or a 7-residue window at each position. To separate the effects of structural fluctuations due to point mutations from other sources, pair RMSD have been translated into P values which themselves are included in a score called P-RANK. This score allows highlighting small backbone distortions by comparing these RMSD between mutated and native positions to the RMSD at the same positions in the absence of a mutation. It results from the P-RANK that 38% of all mutations produce a significant effect on the displacement. When compared with a random distribution of RMSD at un-mutated positions, we show that, even if the RMSD is greater when the mutation is in loops than in regular secondary structure, the relative effect is more important for regular secondary structures and for buried positions. We confirm the absence of correlation between RMSD and the predicted variation of free energy of folding but we found a small correlation between high RMSD and the error in the prediction of ΔΔG.