Proteins have evolved to exploit long-range structural and dynamic effects as a means of regulating function. Understanding communication between sites in proteins is therefore vital to our comprehension of such phenomena as allostery, catalysis, and ligand binding/ejection. Double mutant cycle analysis has long been used to determine the existence of communication between pairs of sites, proximal or distal, in proteins. Typically, nonadditivity (or "Thermodynamic coupling") is measured from global transitions in concert with a single probe. Here, we have applied the atomic resolution of NMR in tandem with native-state hydrogen exchange (HX) to probe the structure/energy landscape for information transduction between a large number of distal sites in a protein. Considering the event of amide proton exchange as an energetically quantifiable structural perturbation, m n-dimensional cycles can be constructed from mutation of n - 1 residues, where m is the number of residues for which HX data is available. Thus, efficient mapping of a large number of couplings is made possible. We have applied this technique to one additive and two nonadditive double mutant cycles in a model system, eglin c. We find heterogeneity of HX-monitored couplings for each cycle, yet averaging results in strong agreement with traditionally measured values. Furthermore, long-range couplings observed at locally exchanging residues indicate that the basis for communication can occur within the native state ensemble, a conclusion not apparent from traditional measurements. We propose that higher-order couplings can be obtained and show that such couplings provide a mechanistic basis for understanding lower-order couplings via "spheres of perturbation". The method is presented as an additional tool for identifying a large number of couplings with greater coverage of the protein of interest.
CheY is a response regulator protein involved in bacterial chemotaxis. Much is known about its active and inactive conformations, but little is known about the mechanisms underlying long-range interactions or correlated motions. To investigate these events, molecular dynamics simulations were performed on the unphosphorylated, inactive structure from Salmonella typhimurium and the CheY−BeF3− active mimic structure (with BeF3− removed) from Escherichia coli. Simulations utilized both sequences in each conformation to discriminate sequence- and structure-specific behavior. The previously identified conformational differences between the inactive and active conformations of the strand-4-helix-4 loop, which are present in these simulations, arise from the structural, and not the sequence, differences. The simulations identify previously unreported structure-specific flexibility features in this loop and sequence-specific flexibility features in other regions of the protein. Both structure- and sequence-specific long-range interactions are observed in the active and inactive ensembles. In the inactive ensemble, two distinct mechanisms based on Thr-87 or Ile-95 rotameric forms, are observed for the previously identified g+ and g− rotamer sampling by Tyr-106. These molecular dynamics simulations have thus identified both sequence- and structure-specific differences in flexibility, long-range interactions, and rotameric form of key residues. Potential biological consequences of differential flexibility and long-range correlated motion are discussed.
Eglin c is a small protease inhibitor whose structural and thermodynamic properties have been well studied. Previous thermodynamic measurements on mutants at solvent-accessible positions in the protein's helix have shown the unexpected result that the data could be best fit by the inclusion of residue- and position-specific parameters to the model. To explore the origins of this surprising result, long molecular dynamics simulations in explicit solvent have been performed. These simulations indicate specific long-range interactions between the solvent-exposed residues in the eglin c alpha-helix and binding loop, an unexpected observation for such a small protein. The residues involved in the interaction are on opposite sides of the protein, about 25 A apart. Simulations of alanine substitutions at the solvent-exposed helix positions, arginine 22, glutamic acid 23, threonine 26, and leucine 27, show both small and large perturbations of eglin c dynamics. Two mutations exhibit large impacts on the long-range helix-loop interactions. Previous stability measurements (Yi et al., Biochemistry 2003;42:7594-7603) had indicated that an alanine substitution at position 27 was less stabilizing than at other solvent-exposed positions in the helix. The L27A mutation effects observed in these simulations suggest that the position-dependent loss of stability measured in wet bench experiments is derived from changes in dynamics that involve long-range interactions; thus, these simulations support the hypothesis that solvent-exposed positions in helices are not always equivalent.
Protein stability is a standard metric to study the relationship between protein sequence and stability. Here we show that laboratory automation can be used to increase both the precision and the throughput of stability measurements. A laboratory automation workstation was used to purify his-tagged eglin c proteins in 96-well format and to prepare protein solutions for solvent denaturation on a semiautomated titrating fluorometer. Using this setup, we have attained a throughput of about 20 stability measurements per day. The robotics facilitate precise stability measurements in which the standard deviation of values from multiple protein preparations (±0.05 kcal/mol) differs little from multiple measurements from a single preparation (±0.04 kcal/mol), while the measurement errors in the literature ranges from 0.10-0.60 kcal/mol. Our approach can be applied to evaluate the consequences of mutation in any fast-folding protein. Large numbers of high-precision stability values derived from mutation are useful for parameterizing models about protein stability determinants and testing models for predicting protein stability and structure based on protein sequence information.
To characterize the long-range structure that persists in the unfolded form of the 70-residue protein eglin C, residual dipolar couplings (RDCs) for HN-N and HA-CA bond vectors were measured by NMR spectroscopy for both its low pH, urea denatured state and its native state. When the data sets for the two different structural states were compared, a statistically significant correlation was found, with both sets of dipolar couplings yielding a correlation coefficient of r = 0.47 to 0.51. This finding directly demonstrates that the denatured state of eglin C has a nativelike global structure, a conclusion reached indirectly for staphylococcal nuclease by combining two different types of NMR data. A simple computer simulation showed that the degree of variation in phi and psi angles that yields the RDC correlation of r = 0.5 was inversely dependent on the statistical segment length, ranging from +/-6 to +/-30 degrees at the upper limit. Stable nativelike topologies that persist on unfolding would explain the rapid refolding kinetics displayed by many proteins and might provide a natural barrier against amyloid fibril formation.
The use of statistical modeling to test hypotheses concerning the determinants of protein structure requires stability data (e.g., the free energy of denaturation in H(2)O, DeltaG(HOH)) from hundreds of protein mutants. Fluorescence-monitored chemical denaturation provides a convenient method for high-precision, high-throughput DeltaG(HOH) determination. For eglin c we find that a throughput of about 20 min per protein can be attained in a two-channel semiautomated titrating fluorometer. We find also that the use of robotics for protein purification and preparation of the solutions for chemical denaturation gives highly precise DeltaG(HOH) values in which the standard deviation of values from multiple preparations (+/-0.051 kcal/mol) differs very little from multiple measurements from a single preparation (+/-0.040 kcal/mol). Since the variance introduced into model fitting by DeltaG(HOH) increases as the square of measurement error, there is a premium on precision. In fact, the fraction of stability behavior explicable by otherwise perfect models goes from 98% to only 50% over the error range commonly reported for chemical denaturation measurements (0.1-0.6 kcal/mol). We have found that the precision of chemical denaturation DeltaG(HOH) measurements depends most heavily on the precision of the instrument used, followed by protein purity and the capacity to precisely prepare the solutions used for titrations.
Statistical modeling provides the mathematics to use data from large numbers of mutant proteins to generate information about hypotheses concerning protein structure not easily obtained from anecdotal studies on small numbers of mutants. Here we use the unfolding free energies of 303 unique eglin c mutant proteins obtained from high-precision, high-throughput chemical denaturation measurements to assess models concerning helix stability. A model with helix propensity as the sole determinant of stability accounts for 83% of the mutant-to-mutant variation in stability for 99% of the mutant proteins (three outliers). When position effects and side chain-side chain interactions are added to the model, the fraction of variation explained increases to 92%. The propensity parameters in this model are identical to helix propensity values derived from other approaches. Measurement error accounts for another 1% of the mutant-to-mutant variation in stability. While the data support terms for several of the expected stabilizing/destabilizing effects, it does not support terms for several others, including i, i + 3 effects in the center of the helix and helix-dipole effects. In addition, the model does better with terms for several stabilizing/destabilizing effects for which we cannot identify the physical basis. The precision of our unfolding stability measurements (+/-0.087 kcal/mol) allows us to conclude that the 7% of variation in stabilities of the mutant proteins not accounted for by the model or by measurement variation is both real and large with respect to the nonpropensity terms in the model. The analysis also shows that the common practice of using C(m)m(av) instead of C(m)m(mut) to calculate DeltaG(HOH,N-D) values for each mutant protein results in a loss of information. We see no correlation between the residuals derived from the full model and m(mut) - m(wt), and hence it is unlikely our m(mut) values reflect mutant-to-mutant differences in the denatured state.
We welcome the ‘Research News’ article about our recent paper [ 1 Carter C.W. et al. Four-body potentials reveal protein-specific correlations to stability changes caused by hydrophobic core mutations. J. Mol. Biol. 2001; 311: 625-638 Crossref PubMed Scopus (103) Google Scholar ] establishing a connection between four-body database-derived likelihood potentials and experimental mutational free energy changes. The Delaunay tetrahedron is the simplest three-dimensional packing motif; we noted in our article, and presented more fully elsewhere [ 2 Cammer, S. et al. Identification of sequence-specific tertiary packing motifs in protein structures using Delaunay tessellation. In Lecture Notes in Computational Science and Engineering (Schlick, T., ed). Springer-Verlag, New York (in press) Google Scholar ] that, unlike potentials of smaller dimensionality, joint consideration of four interacting side chains achieves excellent proportionality between statistical and thermodynamic four-body potentials (calculated by summing the transfer free energies, or hydrophobicities, of the four interacting side chains). Decomposing tertiary structure into elementary three-dimensional simplices therefore appears to afford a decisive new coherence to side-chain packing analysis.
Mutational experiments show how changes in the hydrophobic cores of proteins affect their stabilities. Here, we estimate these effects computationally, using four-body likelihood potentials obtained by simplicial neighborhood analysis of protein packing (SNAPP). In this procedure, the volume of a known protein structure is tiled with tetrahedra having the center of mass of one amino acid side-chain at each vertex. Log-likelihoods are computed for the 8855 possible tetrahedra with equivalent compositions from structural databases and amino acid frequencies. The sum of these four-body potentials for tetrahedra present in a given protein yields the SNAPP score. Mutations change this sum by changing the compositions of tetrahedra containing the mutated residue and their related potentials. Linear correlation coefficients between experimental mutational stability changes, Δ(ΔGunfold), and those based on SNAPP scoring range from 0.70 to 0.94 for hydrophobic core mutations in five different proteins. Accurate predictions for the effects of hydrophobic core mutations can therefore be obtained by virtual mutagenesis, based on changes to the total SNAPP likelihood potential. Significantly, slopes of the relation between Δ(ΔGunfold) and ΔSNAPP for different proteins are statistically distinct, and we show that these protein-specific effects can be estimated using the average SNAPP score per residue, which is readily derived from the analysis itself. This result enhances the predictive value of statistical potentials and supports previous suggestions that “comparable” mutations in different proteins may lead to different Δ(ΔGunfold) values because of differences in their flexibility and/or conformational entropy.
Recent progress has been remarkable in identifying mutations which cause diseases (mostly uncommon) that are inherited simply. Unfortunately, the common diseases of humankind with a strong genetic component, such as those affecting cardiovascular function, have proved less tractable. Their etiology is complex with substantial environmental components and strong indications that multiple genes are implicated. In this article, we consider the genetic etiology of essential hypertension. After presenting the distribution of blood pressures in the population, we propose the hypothesis that essential hypertension is the consequence of different combinations of genetic variations that are individually of little consequence. The candidate gene approach to finding relevant genes is exemplified by studies that identified potentially causative variations associated with quantitative differences in the expression of the angiotensinogen gene (AGT). Experiments to test causation directly are possible in mice, and we describe their use to establish that blood pressures are indeed altered by genetic changes in AGT expression. Tests of differences in expression of the genes coding for the angiotensin-converting enzyme (ACE) and for the natriuretic peptide receptor A are also considered, and we provide a tabulation of all comparable experiments in mice. Computer simulations are presented that resolve the paradoxical finding that while ACE inhibitors are effective, genetic variations in the expression of the ACE gene do not affect blood pressure. We emphasize the usefulness of studying animals heterozygous for an inactivating mutation and a wild-type allele, and briefly discuss a way of establishing causative links between complex phenotypes and single nucleotide polymorphisms.
We studied the thermal denaturation of eglin c by using CD spectropolarimetry and differential scanning calorimetry (DSC). At low protein concentrations, denaturation is consistent with the classical two-state model. At concentrations greater than several hundred μM, however, the calorimetric enthalpy and the midpoint transition temperature increase with increasing protein concentration. These observations suggested the presence of intermediates and/or native state aggregation. However, the transitions are symmetric, suggesting that intermediates are absent, the DSC data do not fit models that include aggregation, and analytical ultracentrifugation (AUC) data show that native eglin c is monomeric. Instead, the AUC data show that eglin c solutions are nonideal. Analysis of the AUC data gives a second virial coefficient that is close to values calculated from theory and the DSC data are consistent with the behavior expected for nonideal solutions. We conclude that the concentration dependence is caused by differential nonideality of the native and denatured states. The nondeality arises from the high charge of the protein at acid pH and is exacerbated by low buffer concentrations. Our conclusion may explain differences between van't Hoff and calorimetric denaturation enthalpies observed for other proteins whose behavior is otherwise consistent with the classical two-state model. © 1999 John Wiley & Sons, Inc. Biopoly 49: 471–479, 1999
Site-directed mutagenesis and combinatorial libraries are powerful tools for providing information about the relationship between protein sequence and structure. Here we report two extensions that expand the utility of combinatorial mutagenesis for the quantitative assessment of hypotheses about the determinants of protein structure, First, we show that resin-splitting technology, which allows the construction of arbitrarily complex libraries of degenerate oligonucleotides, can be used to construct more complex protein libraries for hypothesis testing than can be constructed from oligonucleotides limited to degenerate codons. Second, using eglin c as a model protein, we show that regression analysis of activity scores from library data can be used to assess the relative contributions to the specific activity of the amino acids that were varied in the library, The regression parameters derived from the analysis of a 455-member sample from a library wherein four solvent-exposed sites in an alpha-helix can contain any of nine different amino acids are highly correlated (P < 0.0001, R-2 = 0.97) to the relative helix propensities for those amino acids, as estimated by a variety of biophysical and computational techniques.
The feasibility of using zinc fingers as the affinity tag for the purification of recombinant proteins with immobilized metal affinity chromatography (IMAC) was investigated. It was shown that upon the release of pre-existing metal ions from the model zinc-finger fusion protein by dialysis with buffer containing 500 mM NaCl, the recovery of the zinc-finger fusion protein was increased from 47% to 87%. The purity of the zinc-finger fusion protein was further improved with a pre-precipitation step at pH 5.0. The adsorption capacities for the zinc-finger fusion protein were in the order of Ni(II) > Cu(II) > Zn (II) > Co(II). Based on the low affinity of Zn(II)-loaded IMAC adsorbent for the model protein, a dual column process with Zn(II)-IMAC as the negative column for the removal of unwanted proteins and Ni(II)-IMAC as the positive column for the adsorption of the zinc-finger fusion protein was proposed, giving a purity of 98%. Purification of the zinc-finger fusion protein under denaturing conditions was also demonstrated. The results of this study suggest that it is possible to purify zinc finger proteins with IMAC without the need for external affinity tags, which is particularly useful for the purification of zinc finger proteins for structural analysis.
All mammalian genomes contain approximately 100,000 copies of the transposable element LINES-1 (L1). Phylogenetic analysis indicates that the L1 progenitor predates the mammalian radiation; since that time, the open reading frames encoded in L1 have evolved under selection. The least conserved regions within L1 are the 5'-terminal transcriptional regulatory sequences. In rodents, four types of L1 elements (A, F, and V from mouse and R from rat) have been defined according to the type of apparently nonhomologous promoter sequence present at the 5' end. In this study, we investigate the relationships between these four types of promoters. DNA sequence was determined from approximately 1.5-kb regions from the 5' ends of seven F- and three V-type L1 elements. These sequences were aligned with 29 previously reported L1 elements. Phylogenetic analysis was then performed on the homologous regions of the alignment. The results indicate that in mouse all of the A-, F-, and V-type elements belong to a single dominant lineage but were inserted into the genome during different time periods; V-type elements are the oldest, while A-type elements are the most recently inserted. V-type elements also appear ancestral to the R-type elements found in rat and therefore were replicatively competent prior to the divergence of rat and mouse. Analysis of sequence identity indicates that the different 5' promoters did not derive from a common ancestor. Therefore, the dominant L1 lineage appears to have acquired novel promoter sequences from non-L1 sources. Transposable elements from a wide range of species show similar structural rearrangements, suggesting that acquisition of new sequences may be a common theme in their evolution.
The F-type subfamily of LINE-1 or L1 retroposons [for long interspersed (repetitive) element 1] was dispersed in the mouse genome several million years ago. This subfamily appears to be both transcriptionally and transpositionally inactive today and therefore may be considered evolutionarily extinct. We hypothesized that these F-type L1s are inactive because of the accumulation of mutations. To test this idea we used phylogenetic analysis to deduce the sequence of a transpositionally active ancestral F-type promoter, resurrected it by chemical synthesis, and showed that it has promoter activity. In contrast, F-type sequences isolated from the modern genome are inactive. This approach, in which the automated DNA synthesizer is used as a "time machine," should have broad application in testing models derived from evolutionary studies.
It is generally thought that only a few Alu elements are capable of retrotransposition and that these ‘master’ sources produce inactive copies. Here, we use a network phylogenetic approach to demonstrate that recently integrated human-specific Alu subfamilies typically contain 10–20% of secondary source elements that contributed 20–40% of all subfamily members. This multiplicity of source elements provides new insight into the remarkably successful amplification strategy of the Alu family.
L1 retroposons are represented in mice by subfamilies of interspersed sequences of varied abundance. Previous analyses have indicated that subfamilies are generated by duplicative transposition of a small number of members of the L1 family, the progeny of which then become a major component of the murine L1 population, and are not due to any active processes generating homology within preexisting groups of elements in a particular species. In mice, more than a third of the L1 elements belong to a clade that became active approximately 5 Mya and whose elements are > or = 95% identical. We have collected sequence information from 13 L1 elements isolated from two species of voles (Rodentia: Microtinae: Microtus and Arvicola) and have found that divergence within the vole L1 population is quite different from that in mice, in that there is no abundant subfamily of homologous elements. Individual L1 elements from voles are very divergent from one another and belong to a clade that began a period of elevated duplicative transposition approximately 13 Mya. Sequence analyses of portions of these divergent L1 elements (approximately 250 bp each) gave no evidence for concerted evolution having acted on the vole L1 elements since the split of the two vole lineages approximately 3.5 Mya; that is, the observed interspecific divergence (6.7%-24.7%) is not larger than the intraspecific divergence (7.9%-27.2%), and phylogenetic analyses showed no clustering into Arvicola and Microtus clades.