
The nature of the nucleation-collapse mechanism in protein folding is probed using 27-mer and 36-mer lattice models. Three different forms for the interaction potentials are used. Three of the four 27-mer sequences have maximally compact and identical native state while the other has a non-compact native conformation. All the sequences fold thermodynamically and kinetically by a two-state process. Analysis of individual trajectories for each sequence using a self-organizing neural net algorithm shows that upon formation of a critical set of contacts the polypeptide chain rapidly reaches the native conformation which is consistent with a nucleation-collapse mechanism. The algorithm, which reduces the identification of the folding nucleus for each trajectory to one of pattern recognition, is used to show that there are multiple folding nuclei. There is a distribution of nucleation contacts in the transition states with some of them occurring with more probability (when averaged over the denatured ensemble) than others. We also show that there is a distribution in the size of the nuclei with the average number of residues in the folding nuclei being less than about one-third of the chain size. The fluctuations in the sizes of the nuclei are large, suggestive of a broad transition region. The folding nuclei, the structures of each are the corresponding transition states, have varying degree of overlap with the native conformation. The distribution of the radius of gyration of the transition states shows that these structures are an expanded form (by about 25% in the radius of gyration) of the native conformation. Local contacts are most dominant in the folding nuclei while a certain fraction of non-local contacts is necessary to stabilize the transition states. The search for the critical nuclei initially involves the formation of local contacts, while non-local contacts are formed later. The fractional values of ΦF for the two 27-mer mutants found by using the protein engineering protocol are consistent with the microscopic picture of partial formation of structures involving these residues in the transition state. These observations lead to a multiple folding nuclei (MFN) model for nucleation-collapse mechanism in protein folding. The major implication of the MFN model is that, even if the residues whose tertiary interactions are formed nearly completely in the transition state are mutated, it does not disrupt the nature of the nucleation-collapse mechanism. We analyze the experiments on chymotrypsin inhibitor 2 and α-spectrin SH3 domain and two circular permutants in light of the MFN model. It is shown that the ΦF-value analysis for these proteins gives considerable support to the MFN model. The theoretical and experimental studies give a coherent picture of the nucleation-collapse mechanism in which there is a distribution of folding nuclei with some more probable than others. The formation of any specific nucleus is not necessary for efficient two-state folding.
Background: NMR experiments show that even water molecules that are well ordered in a crystal structure exchange with the external solvent. Despite crucial progress on the understanding of the exchange of crystal-buried water molecules, the detailed pathways followed by a water molecule to escape from or penetrate into the protein interior are unknown. Results: The exchange of a crystal water molecule buried in the low-density lipoprotein receptor-binding domain of human apolipoprotein E with a water molecule from the external solvent was detected and monitored in a molecular dynamics simulation. This simulation shows that the escape of the crystal water molecule from the protein interior and the penetration of the water molecule from the bulk occur by a single-pathway mechanism involving conformational fluctuations of arginine and tryptophan sidechains. Along the pathway the exchanging water molecule interacts specifically with protein atoms by way of a varying pattern of hydrogen bonds. Conclusions:The exchange pathway revealed by the molecular dynamics trajectory suggests a mechanism by which hydrogen bonds work in relay to permit either the penetration or the expulsion of a water molecule. This result may have important implications not only on the process of water exchange but also to probe ligand binding to proteins.
Background: Peptides have ubiquitous roles in all biological systems and are thus of interest in both basic and applied research. The rational design of bioactive peptides could be greatly enhanced by an efficient method for accurately predicting the conformations that these molecules can adopt in solution, As a design process inevitably requires testing numerous molecules, an efficient method would require the calculations to be pet-formed in vacuum.Results: Attempts to predict the conformations of cyclic peptides using a simulated annealing protocol with the Amber/OPLS potential in vacuum resulted, not unexpectedly, in overly packed, non-native conformations We therefore empirically modified the potential by several cycles of structure prediction and function refinement until a good fit between experimental and predicted conformations was obtained, Three major modifications to the potential were required in order to reproduce the solution structures of cyclic peptides: explicit torsional energies for the peptide backbone torsional angles; explicit hydrogen-bonding energies for backbone hydrogen bonds; and a penalty for close approaches between uncharged and charged atoms.Conclusions: Using the modified potential, we predicted the solution conformations of cyclic peptides in the size range of 5-10 residues with reasonable accuracy.
In this commentary, I compare and discuss different views on the nucleation mechanisms in protein folding
Background: Transmissible spongiform encephalopathies are a group of neurodegenerative disorders of man and animals that are believed to be caused by an α-helical to β-sheet conformational change in the prion protein, PrP. Recently determined NMR structures of recombinant PrP (residues 121–231 and 90–231) have identified a short two-stranded anti-parallel β sheet in the normal cellular form of the protein (PrPC). This β sheet has been suggested to be involved in seeding the conformational transition to the disease-associated form (PrPSc) via a partially unfolded intermediate state.Results: We describe CD and NMR studies of three peptides (125–170, 142–170 and 156–170) that span the β-sheet and helix 1 region of PrP, forming a large part of the putative PrPSc–PrPC binding site that has been proposed to be important for self-seeding replication of PrPSc. The data suggest that all three peptides in water have predominantly helical propensities, which are enhanced in aqueous methanol (as judged by deviations from random-coil Hα chemical shifts and 3JHα–NH values). Although the helical propensity is most marked in the region corresponding to helix 1 (144–154), it is also apparent for residues spanning the two β-strand sequences.Conclusions:We have attempted to model the conformational properties of a partially unfolded state of PrP using peptide fragments spanning the region 125–170. We find no evidence in the sequence for any intrinsic conformational preference for the formation of extended β-like structure that might be involved in promoting the PrPC–PrPSc conformational transition.
Background: Many attempts have been made to resolve in time the folding of model proteins in computer simulations. Different computational approaches have emerged. Some of these approaches suffer from the insensitivity to the geometrical properties of the proteins (lattice models), while others are computationally heavy (traditional MD). Results: We use a recently-proposed approach of Zhou and Karplus to study the folding of the protein model based on the discrete time molecular dynamics algorithm. We show that this algorithm resolves with respect to time the folding --- unfolding transition. In addition, we demonstrate the ability to study the coreof the model protein. Conclusion: The algorithm along with the model of inter-residue interactions can serve as a tool to study the thermodynamics and kinetics of protein models.
The authors of the two commentaries that follow have done a service to the folding community by presenting their contrasting views of the concept of nucleation in protein folding. The realization that folding might resemble the process of crystallization is an old one that can be traced back to Pauling. The seductive term ‘nucleation’ is familiar to most scientists despite the complexity of defining that word precisely for many situations. Workers in physical chemistry are quite aware of the difficulties in even formulating a quantitative nucleation theory. Although the idea was introduced by Becker, Doering, and Frenkel earlier in this century, there have been quantitative arguments about the nature of nucleation in simple systems such as crystallizing liquids that rage up until the present day. This is why metallurgy is by no means a dead subject. Thus, it is no surprise that in the context of protein folding many different ideas about nucleation would be bandied about.Many of the qualitative notions of nucleation theory are still unconnected from precise quantifications of the kinetic process. The status of the issue is made clear by noting that neither commentary mentions how one could compute a folding rate de novo using either nucleation picture. A precise definition requires such an algorithmic recipe. In this situation, it seems clear that Thirumalai is using the term ‘nucleus’ in the sense of a critical nucleus that determines rates through an activation free energy in a Kramers-like description, whereas Shakhnovich introduces another specific term, ‘postcritical nucleus’, which has some additional criteria on the transmission factor involved in its formulation. The latter may prove to be an interesting concept for visualizing the folding process.We can see that the term ‘nucleus’ in protein folding has clearly evolved since the earliest times when it seemed to mean a pre-existing structure in the denatured ensemble. It is quite useful, therefore, that the authors of these commentaries quote directly the definitions that were used at the time the various papers were published. In fact, most readers will be able to see that real progress has been made over the years of discussion of this concept. The study of simple models has now prepared the way for quantitatively addressing the factors that determine the folding speeds of real proteins in the laboratory.
Background: In vitro selection has been shown previously to be a powerful method for isolating nucleic acids with specific ligand-binding functions (‘aptamers’). Given this capacity, we have sought to isolate RNA motifs that can confer fluorescent labeling to tagged RNA transcripts, potentially allowing in vivo detection and in vitro spectroscopic analysis of RNAs.Results: Two aptamers that recognize the fluorophore sulforhodamine B were isolated by the in vitro selection process. An unusually large motif of approximately 60 nucleotides is responsible for binding in one RNA (SRB-2). This motif consists of a three-way helical junction with two large, highly conserved unpaired regions. Phosphorothioate mapping with an iodoacetamide-tagged form of the ligand shows that these two regions make close contacts with the fluorophore, suggesting that the two loops combine to form separate halves of a binding pocket. The aptamer binds the fluorophore with high affinity, recognizing both the planar aromatic ring system and a negatively charged sulfonate, a rare example of anion recognition by RNA. An aptamer (FB-1) that specifically binds fluorescein has also been isolated by mutagenesis of a sulforhodamine aptamer followed by re-selection. In a simple in vitro test, SRB-2 and FB-1 have been shown to discriminate between sulforhodamine and fluorescein, specifically localizing each fluorophore to beads tagged with the corresponding aptamer.Conclusions:In addition to serving as a model system for understanding the basis of RNA folding and function, these experiments demonstrate potential applications for the aptamers in transcript double labeling or fluorescence resonance energy transfer studies.
Background: Protein glycosylation, the covalent attachment of carbohydrates, is very common, but in many cases the biological function of glycosylation is not well understood. Recently, fluorescence energy transfer experiments have shown that glycosylation can strongly change the global conformational distributions of peptides. We intend to show the physical mechanism behind this structural effect using a theoretical model.Results: The framework of the hp model of Dill and coworkers is used to describe peptides and their glycosylated counterparts. Conformations are completely enumerated and exact results are obtained for the effect of glycosylation. On glycosylation, the model peptides experience conformational changes similar to those seen in experiments. This effect is highly specific for the sequence of amino acids and also depends on the size of the glycan. Experimentally testable predictions are made for related peptides.Conclusions:Glycans can, by means of entropic contributions, modulate the free energy landscape of polypeptides and thereby specifically stabilize polypeptide conformations. With respect to glycoproteins, the results suggest that the loss of chain entropy during protein folding is partly balanced by an increase in carbohydrate entropy.
During folding of RNase A an initial global hydrophobicity is not observed, which contradicts the view that this is a general requirement for protein folding. This folding behavior could be typical of similar, moderately hydrophobic proteins.
Background: The most conspicuous feature of a right-handed alpha helix is the presence of hydrogen bonds between the backbone carbonyl oxygen and NH groups along the chain. A simple off-lattice model that includes hydrogen bond interactions using virtual atoms is used to examine the stability, cooperativity and kinetics of the helix-coil transition.Results: We have studied the thermodynamics (using multiple histogram method) and kinetics (by Brownian dynamics simulations) of 16-mer minimal off-lattice models of four-turn alpha-helix sequences. The carbonyl and NH groups are represented as virtual moieties located between two alpha-carbon atoms along the polypeptide chain. The characteristics of the native conformations of the model helices, such as the helical pitch and angular correlations, coincide with those found in real proteins. The transition from coil to helix is quite broad, which is typical of these finite-sized systems. The cooperativity, as measured by a dimensionless parameter, Omega(c), that takes into account the width and the slope of the transition curves, is enhanced when hydrogen bonds are taken into account. The value of Omega(c) for our model is consistent with that inferred from experiment for an alanine-based helix-forming peptide. The folding time tau(F) ranges from 6 to 1000 ns in the temperature range 0.7-1.9 T-F where T-F is the helix-coil transition temperature. These values are in excellent agreement with the results from recent fast folding experiments. The temperature dependence of tau(F) exhibits a nearly Arrhenius behavior. Thermally induced unfolding occurs on a time scale that is less than 40-170 ps depending on the final temperature. Our calculations also predict that, although tau(F) can be altered by changes in the sequence, the dynamic range over which such changes take place is not as large as that predicted for beta-turn formation.Conclusions: Hydrogen bonds not only affect the stability of alpha-helix formation but also have profound influence on the kinetics. The excellent agreement between our calculations and experiments suggests that these models can be used to investigate the effects of sequence, temperature and viscosity on the helix-coil transition.
Background: A detailed knowledge of three-dimensional conformations is necessary in order to understand the close relationship between protein structure and function. Among current methodologies, homology modeling is an important tool for obtaining reliable geometries and it provides a direct alternative to X-ray or NMR techniques. In contrast, predictive methods with no three-dimensional template (non-homology) still require further validation and systematization.Results: Here, we present a non-homology knowledge-based strategy for the structural prediction of the proregion of a cysteine proteinase zymogen. This method analyzes individual sequences and multiple alignments of homologous sequences, making use of different published algorithms and incorporating all available structure-related information to obtain improved predictions. Our strategy yielded acceptable secondary structure and general three-dimensional assignments when compared with crystallographic data from homologous proteins.Conclusions:We discuss our successes and failures as a contribution to non-homology prediction development. In addition, based on the information analyzed and generated in this work, we propose plausible folding and activation mechanisms for thiol-proteinase precursors that attempt to shed light on the molecular basis of prosegment functions.
Background: The partial specific volume of a protein is an experimental quantity containing information about solute-solvent interactions and protein hydration, We use a hydration-shell model to partition the partial specific volume into an intrinsic volume occupied by the protein and a change in the volume occupied by the solvent resulting from the solvent interactions with the protein. We seek to extract microscopic information about protein hydration and unfolding from experimental volume measurements without using computer simulations. We employ the idea that the protein-solvent interaction will be proportional to the surface area of the protein.Results: A linear relationship is obtained when the difference between the experimental protein partial specific volume and its intrinsic volume is plotted as a function of the protein solvent-accessible surface area. The effect of using different protein volume definitions on the analysis of protein volumetric properties is discussed. Volumetric data are used to test a model for the unfolded state of proteins and to make predictions about the denatured state.Conclusions: The linear relationship between hydration-shell volume change and accessible surface area reflects the similar surface properties (fractional composition of nonpolar, polar and charged surface) among a diverse set of proteins. This linear relationship is found to be independent of how the solution is partitioned into solute and solvent components. The interpretation of hydration shell versus bulk water properties is found to be very model dependent, however, The maximally exposed unfolded protein model is found to be inconsistent with experimental volume changes of unfolding.
Background: There have been many attempts to approximate realistic protein interaction energies by coarse graining (i.e. considering interactions between amino acids rather than those between atoms). In particular, many 20-letter models have been derived (corresponding to the 20 naturally occurring amino acids). Because such models remain computationally infeasible, many two-letter models have been proposed as further simplifications. The choice of which model to use remains arbitrary, however. In this work, we formulate the framework within which the quality of approximate interaction potentials with respect to folding can be defined explicitly.Results: Using a recently proposed criterion for comparing interaction matrices, we compare various 20 × 20 interaction matrices and obtain the two-letter model that most closely approximates each 20 × 20 matrix. We find that there are considerable differences among the 20 × 20 matrices. In particular, some matrices are much more similar to the hydrophobic model than others. Furthermore, we find that although the best two-letter approximation of a 20-letter model is a significantly better approximation than a random two-letter model, it is still a poor approximation of realistic protein interactions.Conclusions: The determination of the best two-letter approximations of various 20-letter models of protein interaction energies reveals the degree to which hydrophobic interactions dominate in each of the models and hence in proteins.
Background: Proteins seem to have their native structure in a global minimum of free energy. No mechanism is known, however, for ensuring this property. Furthermore, computational complexity studies suggest that such a mechanism is not feasible. These seemingly contradictory observations can be reconciled by the suggestion that evolutionary selection can yield proteins whose native conformation is in the global minimum of free energy. The aim of this study is to investigate such evolutionary processes in a simple model of protein folding.Results: Three possible evolutionary processes are explored. First, if the free energy of the chain is kept below a fixed threshold there is no improvement towards the global minimum. Second, if free energy is minimized directly, sequences emerge whose native conformation is in the global minimum of free energy. Third, when evolutionary pressure is applied within a small set of close homologs, sequences emerge whose functional conformation is in the global minimum of free energy.Conclusions: Although minimizing free energy does select for sequences whose functional conformation is in the global free energy minimum, we argue that for most proteins, which typically have free energy values of only 5-15 kcal/mol, such evolutionary pressure cannot be considered biologically plausible. In contrast, by repeatedly forcing sequences to avoid drifting towards competing 'non-native' conformations, sequences emerge whose native conformation becomes very close to the global minimum of free energy. We argue that such a mechanism is both efficient and biologically plausible.
Background: Slow refolding of 3-phosphoglycerate kinase is supposed to be caused mainly by its domain structure: folding of the C-terminal domain and/or domain pairing has been suggested to be the rate-limiting step. A slow isomerization has been observed during refolding of the isolated C-terminal proteolytic fragment (larger than the C-domain of about 22 kDa by 5 kDa) of the pig muscle enzyme. Here, the role of this step in the reformation of the active enzyme species is investigated.Results: The time course of reactivation during refolding of 3-phosphoglycerate kinase or its complementary proteolytic fragments (residues 1-155 and 156-416) exhibits a pronounced lag-phase indicating the formation of an inactive folding intermediate. The whole process, which leads to a high (60-85%) recovery of the enzyme activity, can be described by two consecutive first-order steps (with rate constants 0.012 +/- 0.0035 and 0.007 +/- 0.0020 s(-1)). A prior renaturation of the C-fragment restores MgATP binding by the C-domain and abolishes the faster step, allowing the separate observation of the slower step. In accordance with this, refolding of the C-domain as monitored by a change in Trp fluorescence occurs at a rate similar to that of the faster step.Conclusions: In addition to the previously observed slow refolding step (0.012 s(-1)) within the C-domain, the occurrence of another slow step (0.007 s(-1)), probably within the N-domain, is detected. The independence of the folding of the C-domain is demonstrated whereas, from the comparative kinetic analysis, independent folding of the N-domain looks less probable. Our data are more compatible with a sequential, rather than random, mechanism and suggest that folding of the C-domain, leading to an inactive intermediate, occurs first, followed by folding of the N-domain.
Background: Uncharacterized proteins from newly sequenced genomes provide perfect targets for fold and function prediction.Results: For 38% of the entire genome of Mycoplasma genitalium, sequence similarity to a protein with a known structure can be recognized using a new sequence alignment algorithm. When comparing genomes of M. genitalium and Escherichia coli, > 80% of M. genitalium proteins have a significant sequence similarity to a protein in E. coli and there are > 40 examples that have not been recognized before. For all cases of proteins with significant profile similarities, there are strong analogies in their functions, if the functions of both proteins are known. The results presented here and other recent results strongly support the argument that such proteins are actually homologous. Assuming this homology allows one to make tentative functional assignments for > 50 previously uncharacterized proteins, including such intriguing cases as the putative β-lactam antibiotic resistance protein in M. genitalium.Conclusions:Using a new profile-to-profile alignment algorithm, the three-dimensional fold can be predicted for almost 40% of proteins from a genome of the small bacterium M. genitalium, and tentative function can be assigned to almost 80% of the entire genome. Some predictions lead to new insights about known functions or point to hitherto unexpected features of M.genitalium
Background: The influenza M2 protein is a simple membrane protein, containing a single transmembrane helix. It is representative of a very large family of single-transmembrane helix proteins. The functional protein is a tetramer, with the four transmembrane helices forming a proton-permeable channel across the bilayer. Two independently derived models of the M2 channel domain are compared, in order to assess the success of applying molecular modelling approaches to simple membrane proteins.Results: The Cα RSMD between the two models is 1.7 å. Both models are composed of a left-handed bundle of helices, with the helices tilted roughly 15° relative to the (presumed) bilayer normal. The two models have similar pore radius profiles, with a pore cavity lined by the Ser31 and Gly34 residues and a pore constriction formed by the ring of His37 residues.Conclusions:Independent studies of M2 have converged on the same structural model for the channel domain. This model is in agreement with solid state NMR data. In particular, both model and NMR data indicate that the M2 helices are tilted relative to the bilayer normal and form a left-handed bundle. Such convergence suggests that, at least for simple membrane proteins, restraints-directed modelling might yield plausible models worthy of further computational and experimental investigation.