The prediction of the quaternary structure of biomolecular macromolecules is of paramount importance for fundamental understanding of cellular processes and drug design. In the era of integrative structural biology, one way of increasing the accuracy of modeling methods used to predict the structure of biomolecular complexes is to include as much experimental or predictive information as possible in the process. This has been at the core of our information-driven docking approach HADDOCK. We present here the updated version 2.2 of the HADDOCK portal, which offers new features such as support for mixed molecule types, additional experimental restraints and improved protocols, all of this in a user-friendly interface. With well over 6000 registered users and 108,000 jobs served, an increasing fraction of which on grid resources, we hope that this timely upgrade will help the community to solve important biological questions and further advance the field. The HADDOCK2.2 Web server is freely accessible to non-profit users at http://haddock.science.uu.nl/services/HADDOCK2.2.
ABSTRACT Information‐driven docking is currently one of the most successful approaches to obtain structural models of protein interactions as demonstrated in the latest round of CAPRI. While various experimental and computational techniques can be used to retrieve information about the binding mode, the availability of three‐dimensional structures of the interacting partners remains a limiting factor. Fortunately, the wealth of structural information gathered by large‐scale initiatives allows for homology‐based modeling of a significant fraction of the protein universe. Defining the limits of information‐driven docking based on such homology models is therefore highly relevant. Here we show, using previous CAPRI targets, that out of a variety of measures, the global sequence identity between template and target is a simple but reliable predictor of the achievable quality of the docking models. This indicates that a well‐defined overall fold is critical for the interaction. Furthermore, the quality of the data at our disposal to characterize the interaction plays a determinant role in the success of the docking. Given reliable interface information we can obtain acceptable predictions even at low global sequence identity. These results, which define the boundaries between trustworthy and unreliable predictions, should guide both experts and nonexperts in defining the limits of what is achievable by docking. This is highly relevant considering that the fraction of the interactome amenable for docking is only bound to grow as the number of experimentally solved structures increases. Proteins 2013; 81:2119–2128. © 2013 Wiley Periodicals, Inc.
ABSTRACTWe report the first assessment of blind predictions of water positions at protein–protein interfaces, performed as part of the critical assessment of predicted interactions (CAPRI) community‐wide experiment. Groups submitting docking predictions for the complex of the DNase domain of colicin E2 and Im2 immunity protein (CAPRI Target 47), were invited to predict the positions of interfacial water molecules using the method of their choice. The predictions—20 groups submitted a total of 195 models—were assessed by measuring the recall fraction of water‐mediated protein contacts. Of the 176 high‐ or medium‐quality docking models—a very good docking performance per se—only 44% had a recall fraction above 0.3, and a mere 6% above 0.5. The actual water positions were in general predicted to an accuracy level no better than 1.5 Å, and even in good models about half of the contacts represented false positives. This notwithstanding, three hotspot interface water positions were quite well predicted, and so was one of the water positions that is believed to stabilize the loop that confers specificity in these complexes. Overall the best interface water predictions was achieved by groups that also produced high‐quality docking models, indicating that accurate modelling of the protein portion is a determinant factor. The use of established molecular mechanics force fields, coupled to sampling and optimization procedures also seemed to confer an advantage. Insights gained from this analysis should help improve the prediction of protein–water interactions and their role in stabilizing protein complexes. Proteins 2014; 82:620–632. © 2013 Wiley Periodicals, Inc.
The WeNMR ( http://www.wenmr.eu ) project is a European Union funded international effort to streamline and automate analysis of Nuclear Magnetic Resonance (NMR) and Small Angle X-Ray scattering (SAXS) imaging data for atomic and near-atomic resolution molecular structures. Conventional calculation of structure requires the use of various software packages, considerable user expertise and ample computational resources. To facilitate the use of NMR spectroscopy and SAXS in life sciences the WeNMR consortium has established standard computational workflows and services through easy-to-use web interfaces, while still retaining sufficient flexibility to handle more specific requests. Thus far, a number of programs often used in structural biology have been made available through application portals. The implementation of these services, in particular the distribution of calculations to a Grid computing infrastructure, involves a novel mechanism for submission and handling of jobs that is independent of the type of job being run. With over 450 registered users (September 2012), WeNMR is currently the largest Virtual Organization (VO) in life sciences. With its large and worldwide user community, WeNMR has become the first Virtual Research Community officially recognized by the European Grid Infrastructure (EGI).
Inaccuracies in computational molecular modeling methods are often counterweighed by brute‐force generation of a plethora of putative solutions. These are then typically sieved via structural clustering based on similarity measures such as the root mean square deviation (RMSD) of atomic positions. Albeit widely used, these measures suffer from several theoretical and technical limitations (e.g., choice of regions for fitting) that impair their application in multicomponent systems ( N > 2), large‐scale studies (e.g., interactomes), and other time‐critical scenarios. We present here a simple similarity measure for structural clustering based on atomic contacts—the fraction of common contacts—and compare it with the most used similarity measure of the protein docking community—interface backbone RMSD. We show that this method produces very compact clusters in remarkably short time when applied to a collection of binary and multicomponent protein–protein and protein–DNA complexes. Furthermore, it allows easy clustering of similar conformations of multicomponent symmetrical assemblies in which chain permutations can occur. Simple contact‐based metrics should be applicable to other structural biology clustering problems, in particular for time‐critical or large‐scale endeavors.Proteins 2012; © 2012 Wiley Periodicals, Inc.
Paramagnetic metal ions generate pseudocontact shifts (PCSs) in nuclear magnetic resonance spectra that are manifested as easily measurable changes in chemical shifts. Metals can be incorporated into proteins through metal binding tags, and PCS data constitute powerful long-range restraints on the positions of nuclear spins relative to the coordinate system of the magnetic susceptibility anisotropy tensor (Δχ-tensor) of the metal ion. We show that three-dimensional structures of proteins can reliably be determined using PCS data from a single metal binding site combined with backbone chemical shifts. The program PCS-ROSETTA automatically determines the Δχ-tensor and metal position from the PCS data during the structure calculations, without any prior knowledge of the protein structure. The program can determine structures accurately for proteins of up to 150 residues, offering a powerful new approach to protein structure determination that relies exclusively on readily measurable backbone chemical shifts and easily discriminates between correctly and incorrectly folded conformations.
In order to enhance the structure determination process of macromolecular assemblies by NMR, we have implemented long-range pseudocontact shift (PCS) restraints into the data-driven protein docking package HADDOCK. We demonstrate the efficiency of the method on a synthetic, yet realistic case based on the lanthanide-labeled N-terminal ε domain of the E. coli DNA polymerase III (ε186) in complex with the HOT domain. Docking from the bound form of the two partners is swiftly executed (interface RMSDs < 1 Å) even with addition of very large amount of noise, while the conformational changes of the free form still present some challenges (interface RMSDs in a 3.1–3.9 Å range for the ten lowest energy complexes). Finally, using exclusively PCS as experimental information, we determine the structure of ε186 in complex with the HOT-homologue θ subunit of the E. coli DNA polymerase III.
Understanding biological phenomena at atomic resolution is one of the keys to modern drug design. In particular, knowledge of 3D structures of proteins and their interactions with other macromolecules are necessary for designing chemical compounds that modify biological processes. Conventional methods for protein structure determinations comprise X-ray crystallography and nuclear magnetic resonance (NMR) spectroscopy. These techniques can also determine the binding mode of chemical compounds. Either technique can be slow and costly, making it highly relevant to explore alternative strategies. Paramagnetic NMR spectroscopy is emerging as such an alternative technique. In order to measure the paramagnetic effects, two NMR spectra are compared that have been measured with and without a bound paramagnetic metal ion. In particular, pseudocontact shifts (PCS) of nuclear spins are easily measured as the difference (in ppm) of the chemical shifts between the two spectra. PCSs provide long range and orientation dependent restraints, allowing positioning of the spin with respect to the magnetic susceptibility tensor anisotropy (Δχ-tensor) of the metal ion. In this thesis, I used the PCS effect to computationally extract information from NMR spectra. I developed (i) a tool (called Possum) to automatically assign diamagnetic and paramagnetic spectra of the methyl groups of amino acid side chains, given structural information of the protein studied and prior knowledge of the Δχ-tensor; (ii) I designed a comprehensive software package (called Numbat) to extract Δχ-tensor parameters from assigned PCS values and the available 3D structure; and (iii) I incorporated PCS-based restraints into the protein structure prediction software CS-ROSETTA and demonstrated that this combination (PCS-ROSETTA) presents a significant improvement for de novo structure determination. The three projects serve different purposes at different stages of protein NMR studies. They could be combined in the following manner: Starting from assigned backbone PCSs, PCS-Rosetta could be used to determine the 3D structure of the protein. Possum can then be used to automatically assign the NMR resonances of the methyl groups using PCSs. Finally, Numbat can be used to fit improved Δχ-tensors to all the PCS data, analyze the quality of the Δχ-tensors and identify possible wrong assignments. Iterative repetition of this protocol would give a 3D structural model of the protein with a minimum of data. Alternatively, the Δχ-tensor parameters and PCSs could be used as input for a traditional software package such as Xplor-NIH to compute a 3D structure of the protein.
A new lanthanide tag was designed for site-specific labeling of proteins with paramagnetic lanthanide ions. The tag, 4-mercaptomethyl-dipicolinic acid, binds lanthanide ions with nanomolar affinity, is readily attached to proteins via a disulfide bond, and avoids the problems of diastereomer formation associated with most of the conventional lanthanide tags. The high lanthanide affinity of the tag opens the possibility to measure residual dipolar couplings in a single sample containing a mixture of paramagnetic and diamagnetic lanthanides. Using the DNA-binding domain of the E. coli arginine repressor as an example, it is demonstrated that the tag allows immobilization of the lanthanide ion in close proximity of the protein by additional coordination of the lanthanide by a carboxyl group of the protein. The close proximity of the lanthanide ion promotes accurate determinations of magnetic susceptibility anisotropy tensors. In addition, the small size of the tag makes it highly suitable for studies of intermolecular interactions.
Pseudocontact shift (PCS) effects induced by a paramagnetic lanthanide bound to a protein have become increasingly popular in NMR spectroscopy as they yield a complementary set of orientational and long-range structural restraints. PCS are a manifestation of the chi-tensor anisotropy, the Deltachi-tensor, which in turn can be determined from the PCS. Once the Deltachi-tensor has been determined, PCS become powerful long-range restraints for the study of protein structure and protein-ligand complexes. Here we present the newly developed package Numbat (New User-friendly Method Built for Automatic Deltachi-Tensor determination). With a Graphical User Interface (GUI) that allows a high degree of interactivity, Numbat is specifically designed for the computation of the complete set of Deltachi-tensor parameters (including shape, location and orientation with respect to the protein) from a set of experimentally measured PCS and the protein structure coordinates. Use of the program for Linux and Windows operating systems is illustrated by building a model of the complex between the E. coli DNA polymerase III subunits epsilon186 and theta using PCS.
Pseudocontact shifts (PCSs) induced by a site-specifically bound paramagnetic lanthanide ion are shown to provide fast access to sequence-specific resonance assignments of methyl groups in proteins of known three-dimensional structure. Stereospecific assignments of Val and Leu methyls are obtained as well as resonance assignments of all other methyls, including Met epsilonCH3 groups. No prior assignments of the diamagnetic protein are required nor are experiments that transfer magnetization between the methyl groups and the protein backbone. Methyl Cz-exchange experiments were designed to provide convenient access to PCS measurements in situations where a paramagnetic lanthanide is in exchange with a diamagnetic lanthanide. In the absence of exchange, simultaneous 13C-HSQC assignments and PCS measurements are delivered by the newly developed program Possum. The approaches are demonstrated with the complex between the N-terminal domain of the subunit epsilon and the subunit theta of the Escherichia coli DNA polymerase III.
Anisotropic magnetic susceptibility tensors χ of paramagnetic metal ions are manifested in pseudocontact shifts, residual dipolar couplings, and other paramagnetic observables that present valuable long-range information for structure determinations of protein-ligand complexes. A program was developed for automatic determination of the χ-tensor anisotropy parameters and amide resonance assignments in proteins labeled with paramagnetic metal ions. The program requires knowledge of the three-dimensional structure of the protein, the backbone resonance assignments of the diamagnetic protein, and a pair of 2D 15N-HSQC or 3D HNCO spectra recorded with and without paramagnetic metal ion. It allows the determination of reliable χ-tensor anisotropy parameters from 2D spectra of uniformly 15N-labeled proteins of fairly high molecular weight. Examples are shown for the 185-residue N-terminal domain of the subunit ε from E. coli DNA polymerase III in complex with the subunit θ and La3+ in its diamagnetic and Dy3+, Tb3+, and Er3+ in its paramagnetic form.