The flexibility of biological macromolecules is an important structural determinant of function. Unfortunately, the correlations between different motional modes are poorly captured by discrete ensemble representations. Here, we present new ways to both represent and visualize correlated interdomain motions. Interdomain motions are determined directly from residual dipolar couplings, represented as a continuous conformational distribution, and visualized using the disk-on-sphere representation. Using the disk-on-sphere representation, features of interdomain motions, including correlations, are intuitively visualized. The representation works especially well for multidomain systems with broad conformational distributions.This analysis also can be extended to multiple probability density modes, using a Bingham mixture model. We use this new paradigm to study the interdomain motions of staphylococcal protein A, which is a key virulence factor contributing to the pathogenicity of Staphylococcus aureus. We capture the smooth transitions between important states and demonstrate the utility of continuous distribution functions for computing the reorientational components of binding thermodynamics. Such insights allow for the dissection of the dynamic structural components of functionally important intermolecular interactions.
The flexibility of biological macromolecules is an important structural determinant of function. Unfortunately, the correlations between different motional modes are poorly captured by discrete ensemble representations. Here, we present new ways to both represent and visualize correlated interdomain motions. Interdomain motions are determined directly from residual dipolar couplings, represented as a continuous conformational distribution, and visualized using the disk-on-sphere representation. Using the disk-on-sphere representation, features of interdomain motions, including correlations, are intuitively visualized. The representation works especially well for multidomain systems with broad conformational distributions.This analysis also can be extended to multiple probability density modes, using a Binghammixture model. We use this new paradigm to study the interdomain motions of staphylococcal protein A, which is a key virulence factor contributing to the pathogenicity of Staphylococcus aureus. We capture the smooth transitions between important states and demonstrate the utility of continuous distribution functions for computing the reorientational components of binding thermodynamics. Such insights allow for the dissection of the dynamic structural components of functionally important intermolecular interactions. © 2018 Elsevier Ltd. All rights reserved.
SpA-N is the N-terminal half of Staphylococcal protein A and it is composed of five protein binding domains. The five domains could bind to antibody, TNFR1 and von Willebrand factor and facilitate the evasion of Staphylococcus aureus into the human immune system. The functional plasticity is suggested to be the result of structural flexibility. Heteronuclear spin relaxation experiments demonstrated that the five domains of SpA-N are connected by four flexible linkers. In order to obtain additional dynamic information, we constructed a di-domain mimic of SpA-N with a lanthanide binding tag and measured residual dipolar couplings (RDCs). The di-domain construct can bind lanthanide ions and consequently be aligned in the presence of a strong magnetic field. By using different combinations of lanthanide ions and protein constructs, we obtained two orthogonal alignments, which have enough information content to determine a model with maximally 10 parameters. In addition, we designed a de novo method to extract dynamic information from RDCs. Instead of determining conventional structure ensembles, our method determines an inter-domain orientation distribution to describe the structure of a flexible protein using continuous probability distributions. By using continuous models, orientation distributions can be parameterized by a small number of variables and still remain general enough to describe a broad spectrum of inter-domain motions. As a result, the method can avoid the over-fitting problem even with sparse data. Because no force field is applied, the generated distribution is purely determined by RDC data and least biased. Using the method, we determined the inter-domain orientation distribution of the di-domain construct with only two orthogonal alignments. A strong correlation was observed in the distribution and conformations were well populated in a limited region.
Residual dipolar coupling (RDC) and residual chemical shift anisotropy (RCSA) provide orientational restraints on internuclear vectors and the principal axes of chemical shift anisotropy (CSA) tensors, respectively. Mathematically, while an RDC represents a single sphero-conic, an RCSA can be interpreted as a linear combination of two sphero-conics. Since RDCs and RCSAs are described by a molecular alignment tensor, they contain inherent structural ambiguity due to the symmetry of the alignment tensor and the symmetry of the molecular fragment, which often leads to more than one orientation and conformation for the fragment consistent with the measured RDCs and RCSAs. While the orientational multiplicities have been long studied for RDCs, structural ambiguities arising from RCSAs have not been investigated. In this paper, we give exact and tight bounds on the number of peptide plane orientations consistent with multiple RDCs and/or RCSAs measured in one alignment medium. We prove that at most 16 orientations are possible for a peptide plane, which can be computed in closed form by solving a merely quadratic equation, and applying symmetry operations. Furthermore, we show that RCSAs can be used in the initial stages of structure determination to obtain highly accurate protein backbone global folds. We exploit the mathematical interplay between sphero-conics derived from RCSA and RDC, and protein kinematics, to derive quartic equations, which can be solved in closed-form to compute the protein backbone dihedral angles (φ, ψ). Building upon this, we designed a novel, sparse-data, polynomial-time divide-and-conquer algorithm to compute protein backbone conformations. Results on experimental NMR data for the protein human ubiquitin demonstrate that our algorithm computes backbone conformations with high accuracy from 13C′ -RCSA or 15N-RCSA, and N-HN RDC data. We show that the structural information present in 13C′ -RCSA and 15N-RCSA can be extracted analytically, and used in a rigorous algorithmic framework to compute a high-quality protein backbone global fold, from a limited amount of NMR data. This will benefit automated NOE assignment and high-resolution protein backbone structure determination from sparse NMR data.
In structural studies of large proteins by NMR, global fold determination plays an increasingly important role in providing a first look at a target's topology and reducing assignment ambiguity in NOESY spectra of fully protonated samples. In this work, we demonstrate the use of ultrasparse sampling, a new data processing algorithm, and a 4-D time-shared NOESY experiment (1) to collect all NOEs in (2)H/(13)C/(15)N-labeled protein samples with selectively protonated amide and ILV methyl groups at high resolution in only four days, and (2) to calculate global folds from this data using fully automated resonance assignment. The new algorithm, SCRUB, incorporates the CLEAN method for iterative artifact removal but applies an additional level of iteration, permitting real signals to be distinguished from noise and allowing nearly all artifacts generated by real signals to be eliminated. In simulations with 1.2% of the data required by Nyquist sampling, SCRUB achieves a dynamic range over 10000:1 (250× better artifact suppression than CLEAN) and completely quantitative reproduction of signal intensities, volumes, and line shapes. Applied to 4-D time-shared NOESY data, SCRUB processing dramatically reduces aliasing noise from strong diagonal signals, enabling the identification of weak NOE crosspeaks with intensities 100× less than those of diagonal signals. Nearly all of the expected peaks for interproton distances under 5 Å were observed. The practical benefit of this method is demonstrated with structure calculations for 23 kDa and 29 kDa test proteins using the automated assignment protocol of CYANA, in which unassigned 4-D time-shared NOESY peak lists produce accurate and well-converged global fold ensembles, whereas 3-D peak lists either fail to converge or produce significantly less accurate folds. The approach presented here succeeds with an order of magnitude less sampling than required by alternative methods for processing sparse 4-D data.
Nuclear magnetic resonance (NMR) spectroscopy is a primary tool to perform structural studies of proteins in the physiologically-relevant solution-state. Restraints on distances between pairs of nuclei in the protein, derived from the nuclear Overhauser effect (NOE) for example, provide information about the structure of the protein in its folded state. NMR studies of symmetric protein homo-oligomers present a unique challenge. Current techniques can determine whether an NOE restrains a pair of protons across different subunits or within a single subunit, but are unable to determine in which subunits the restrained protons lie. Consequently, it is difficult to assign NOEs to particular pairs of subunits with certainty, thus hindering the structural analysis of the oligomeric state. Hence, computational approaches are needed to address this subunit ambiguity. We reduce the structure determination of protein homo-oligomers with cyclic symmetry to computing geometric arrangements of unions of annuli in a plane. Our algorithm, disco, runs in expected O(n 2) time, where n is the number of distance restraints, and is guaranteed to report the exact set of oligomer structures consistent with ambiguously-assigned inter-subunit distance restraints and orientational restraints from residual dipolar couplings (RDCs). Since the symmetry axis of an oligomeric complex must be parallel to an eigenvector of the alignment tensor of RDCs, we can represent each distance restraint as a union of annuli in a plane encoding the configuration space of the symmetry axis. Oligomeric protein structures with the best restraint satisfaction correspond to faces of the arrangement contained in the greatest number of unions of annuli. We demonstrate our method using two symmetric protein complexes: the trimeric E. coli Diacylglycerol Kinase (DAGK), whose distance restraints possess at least two possible subunit assignments each; and a dimeric mutant of the immunoglobulin-binding domain B1 of streptococcal protein G (GB1) using ambiguous NOEs. In both cases, disco computes oligomer structures with high accuracy.
High-resolution structure determination of homo-oligomeric protein complexes remains a daunting task for NMR spectroscopists. Although isotope-filtered experiments allow separation of intermolecular NOEs from intramolecular NOEs and determination of the structure of each subunit within the oligomeric state, degenerate chemical shifts of equivalent nuclei from different subunits make it difficult to assign intermolecular NOEs to nuclei from specific pairs of subunits with certainty, hindering structural analysis of the oligomeric state. Here, we introduce a graphical method, DISCO, for the analysis of intermolecular distance restraints and structure determination of symmetric homo-oligomers using residual dipolar couplings. Based on knowledge that the symmetry axis of an oligomeric complex must be parallel to an eigenvector of the alignment tensor of residual dipolar couplings, we can represent distance restraints as annuli in a plane encoding the parameters of the symmetry axis. Oligomeric protein structures with the best restraint satisfaction correspond to regions of this plane with the greatest number of overlapping annuli. This graphical analysis yields a technique to characterize the complete set of oligomeric structures satisfying the distance restraints and to quantitatively evaluate the contribution of each distance restraint. We demonstrate our method for the trimeric E. coli diacylglycerol kinase, addressing the challenges in obtaining subunit assignments for distance restraints. We also demonstrate our method on a dimeric mutant of the immunoglobulin-binding domain B1 of streptococcal protein G to show the resilience of our method to ambiguous atom assignments. In both studies, DISCO computed oligomer structures with high accuracy despite using ambiguously assigned distance restraints.
Symmetric homo-oligomers represent a majority of proteins, and determining their structures helps elucidate important biological processes, including ion transport, signal transduction, and transcriptional regulation. In order to account for the noise and sparsity in the distance restraints used in Nuclear Magnetic Resonance (NMR) structure determination of cyclic (C(n)) symmetric homo-oligomers, and the resulting uncertainty in the determined structures, we develop a Bayesian structural inference approach. In contrast to traditional NMR structure determination methods, which identify a small set of low-energy conformations, the inferential approach characterizes the entire posterior distribution of conformations. Unfortunately, traditional stochastic techniques for inference may under-sample the rugged landscape of the posterior, missing important contributions from high-quality individual conformations and not accounting for the possible aggregate effects on inferred quantities from numerous unsampled conformations. However, by exploiting the geometry of symmetric homo-oligomers, we develop an algorithm that provides provable guarantees for the posterior distribution and the inferred mean atomic coordinates. Using experimental restraints for three proteins, we demonstrate that our approach is able to objectively characterize the structural diversity supported by the data. By simulating spurious and missing restraints, we further demonstrate that our approach is robust, degrading smoothly with noise and sparsity.
Protein complexes play vital roles in the fundamental processes of life. In particular, homo-oligomers are involved in cell signaling, regulation, and transport. To make detailed studies of these symmetric proteins, they need to be discovered, and their structures need to be solved. For nuclear magnetic resonance (NMR) structure determination, a protein of interest must be first selected, then NMR data must be assigned and then applied to a structure determination method. In this thesis we present several algorithms for each of these steps (selection, assignment, and structure determination). Our particular area of focus is on the NMR structure determination of symmetric homo-oligomers because they form a large fraction of all known protein structures. We present algorithms to compute the structure of an entire symmetric homo-oligomer, given the structure of the sub-unit, and a sparse number of additional restraints, such as restraints from the nuclear Overhauser effect (NOE), along with the possible addition of residual dipolar couplings (RDCs). Traditional methods for structure determination are typically incomplete in that they do not consider all possibilities in the configuration-space. As a result, they often have no provable guarantees for returning all valid solutions. For example, stochastic approaches (such as simulated annealing and Monte-Carlo methods) only consider those configurations which they have explicitly sampled. Because their sampling set is discrete and finite, they can fail to discover important solutions that are not in their sample set. Gridsearch approaches suffer from the same disadvantage. Local differential approaches (such as gradient-descent, molecular dynamics (MD)), may become trapped in local-minima and not consider valid solutions in other regions of the configuration space. In contrast, we present structure determination algorithms which are provably complete. By complete we mean that the algorithm returns a super-set of all possible solutions, and therefore is guaranteed never to miss any valid solutions. Our algorithm achieves completeness by using geometric bounds to consider all possibilities in the configuration space, and to provably eliminate large sections of the configuration space which cannot contain a solution. In addition, our geometric bound also allows us to remain complete, even when restraints are ambiguous with many possible interpretations.
We present a novel structure determination approach that exploits the global orientational restraints from RDCs to resolve ambiguous NOE assignments. Unlike traditional approaches that bootstrap the initial fold from ambiguous NOE assignments, we start by using RDCs to compute accurate secondary structure element (SSE) backbones at the beginning of structure calculation. Our structure determination package, called rdc-Panda (RDC-based SSE PAcking with NOEs for Structure Determination and NOE Assignment), consists of three modules: (1) rdc-exact; (2) Packer; and (3) hana (HAusdorff-based NOE Assignment). rdc-exact computes the global optimal solution of backbone dihedral angles for each secondary structure element by exactly solving a system of quartic RDC equations derived by Wang and Donald (Proceedings of the IEEE computational systems bioinformatics conference (CSB), Stanford, CA, 2004a; J Biomol NMR 29(3):223–242, 2004b), and systematically searching over the roots, each of which is a backbone dihedral ϕ- or ψ-angle consistent with the RDC data. Using a small number of unambiguous inter-SSE NOEs extracted using only chemical shift information, Packer performs a systematic search for the core structure, including all SSE backbone conformations. hana uses a Hausdorff-based scoring function to measure the similarity between the experimental spectra and the back-computed NOE pattern for each side-chain from a statistically-diverse rotamer library, and drives the selection of optimal position-specific rotamers for filtering ambiguous NOE assignments. Finally, a local minimization approach is used to compute the loops and refine side-chain conformations by fixing the core structure as a rigid body while allowing movement of loops and side-chains. rdc-Panda was applied to NMR data for the FF Domain 2 of human transcription elongation factor CA150 (RNA polymerase II C-terminal domain interacting protein), human ubiquitin, the ubiquitin-binding zinc finger domain of the human Y-family DNA polymerase Eta (pol η UBZ), and the human Set2-Rpb1 interacting domain (hSRI). These results demonstrated the efficiency and accuracy of our algorithm, and show that rdc-Panda can be successfully applied for high-resolution protein structure determination using only a limited set of NMR data by first computing RDC-defined backbones.
Symmetric homo-oligomers are protein complexes with similar subunits arranged symmetrically [10]. Figure 1 illustrates the structure of a symmetric homo-oligomer called phospholamban. Phospholamban is a membrane protein that helps regulate the calcium level inside the cell and hence aids in muscle contraction and relaxation [7]; ion conductance studies [5] also suggest that phospholamban might have a separate role as an ion channel. A detailed molecular-level understanding of homo-oligomeric structures provides insights into their functions and, in some cases, how to design appropriate drugs. Nuclear Magnetic Resonance (NMR) spectroscopy underlies many structural studies of homo-oligomers, but poses significant computational challenges in inferring threedimensional structures from indirect (and often sparse) measurements of geometry.
Assignment of nuclear Overhauser effect (NOE) data is a key bottleneck in structure determination by NMR. NOE assignment resolves the ambiguity as to which pair of protons generated the observed NOE peaks, and thus should be restrained in structure determination. In the case of intersubunit NOEs in symmetric homo-oligomers, the ambiguity includes both the identities of the protons within a subunit, and the identities of the subunits to which they belong. This paper develops an algorithm for simultanous intersubunit NOE assignment and Cn symmetric homo-oligomeric structure determinations, given the subunit structure. By using a configuration space framework, our algorithm guarantees completeness, in that it identifies structures representing, to within a user-defined similarity level, every structure consistent with the available data (ambiguous or not). However, while our approach is complete in considering all conformations and assignments, it avoids explicit enumeration of the exponential number of combinations of possible assignments. Our algorithm can draw two types of conclusions not possible under previous methods: (1) that different assignments for an NOE would lead to different structural classes, or (2) that it is not necessary to uniquely assign an NOE, since it would have little impact on structural precision. We demonstrate on two test proteins that our method reduces the average number of possible assignments per NOE by a factor of 2.6 for MinE and 4.2 for CCMP. It results in high structural precision, reducing the average variance in atomic positions by factors of 1.5 and 3.6, respectively.
Structural studies of symmetric homo-oligomers provide mechanistic insights into their roles in essential biological processes, including cell signaling and cellular regulation. This paper presents a novel algorithm for homo-oligomeric structure determination, given the subunit structure, that is both complete, in that it evaluates all possible conformations, and data-driven, in that it evaluates conformations separately for consistency with experimental data and for quality of packing. Completeness ensures that the algorithm does not miss the native conformation, and being data-driven enables it to assess the structural precision possible from data alone. Our algorithm performs a branch-and-bound search in the symmetry configuration space, the space of symmetry axis parameters (positions and orientations) defining all possible C-n homo-oligomeric complexes for a given subunit structure. It eliminates those symmetry axes inconsistent with intersubunit nuclear Overhauser effect (NOE) distance restraints and then identifies conformations representing any consistent, well-packed structure to within a user-defined similarity level. For the human phospholamban pentamer in dodecylphosphocholine micelles, using the structure of one subunit determined from a subset of the experimental NMR data, our algorithm identifies a diverse set of complex structures consistent with the nine intersubunit NOE restraints. The distribution of determined structures provides an objective characterization of structural uncertainty: backbone RMSD to the previously determined structure ranges from 1.07 to 8.85 angstrom, and variance in backbone atomic coordinates is an average of 12.32 angstrom(2). Incorporating vdW packing reduces structural diversity to a maximum backbone RMSD of 6.24 angstrom and an average backbone variance of 6.80 angstrom(2). By comparing data consistency and packing quality under different assumptions of oligomeric number, our algorithm identifies the pentamer as the most likely oligomeric state of phospholamban, demonstrating that it is possible to determine the oligomeric number directly from NMR data. Additional tests on a number of homo-oligomers, from dimer to heptamer, similarly demonstrate the power of our method to provide unbiased determination and evaluation of homo-oligomeric complex structures.