Numerous affinity purification-mass spectrometry (AP-MS) and yeast two-hybrid screens have each defined thousands of pairwise protein-protein interactions (PPIs), most of which are between functionally unrelated proteins. The accuracy of these networks, however, is under debate. Here, we present an AP-MS survey of the bacterium Desulfovibrio vulgaris together with a critical reanalysis of nine published bacterial yeast two-hybrid and AP-MS screens. We have identified 459 high confidence PPIs from D. vulgaris and 391 from Escherichia coli. Compared with the nine published interactomes, our two networks are smaller, are much less highly connected, and have significantly lower false discovery rates. In addition, our interactomes are much more enriched in protein pairs that are encoded in the same operon, have similar functions, and are reproducibly detected in other physical interaction assays than the pairs reported in prior studies. Our work establishes more stringent benchmarks for the properties of protein interactomes and suggests that bona fide PPIs much more frequently involve protein partners that are annotated with similar functions or that can be validated in independent assays than earlier studies suggested.
Identifying protein-protein interactions (PPIs) at an acceptable false discovery rate (FDR) is challenging. Previously we identified several hundred PPIs from affinity purification - mass spectrometry (AP-MS) data for the bacteria Escherichia coli and Desulfovibrio vulgaris. These two interactomes have lower FDRs than any of the nine interactomes proposed previously for bacteria and are more enriched in PPIs validated by other data than the nine earlier interactomes. To more thoroughly determine the accuracy of ours or other interactomes and to discover further PPIs de novo, here we present a quantitative tagless method that employs iTRAQ MS to measure the copurification of endogenous proteins through orthogonal chromatography steps. 5273 fractions from a four-step fractionation of a D. vulgaris protein extract were assayed, resulting in the detection of 1242 proteins. Protein partners from our D. vulgaris and E. coli AP-MS interactomes copurify as frequently as pairs belonging to three benchmark data sets of well-characterized PPIs. In contrast, the protein pairs from the nine other bacterial interactomes copurify two- to 20-fold less often. We also identify 200 high confidence D. vulgaris PPIs based on tagless copurification and colocalization in the genome. These PPIs are as strongly validated by other data as our AP-MS interactomes and overlap with our AP-MS interactome for D.vulgaris within 3% of expectation, once FDRs and false negative rates are taken into account. Finally, we reanalyzed data from two quantitative tagless screens of human cell extracts. We estimate that the novel PPIs reported in these studies have an FDR of at least 85% and find that less than 7% of the novel PPIs identified in each screen overlap. Our results establish that a quantitative tagless method can be used to validate and identify PPIs, but that such data must be analyzed carefully to minimize the FDR.
Tilted electron microscope images are routinely collected for an ab initio structure reconstruction as a part of the Random Conical Tilt (RCT) or Orthogonal Tilt Reconstruction (OTR) methods, as well as for various applications using the "free-hand" procedure. These procedures all require identification of particle pairs in two corresponding images as well as accurate estimation of the tilt-axis used to rotate the electron microscope (EM) grid. Here we present a computational approach, PCT (particle correspondence from tilted pairs), based on tilt-invariant context and projection matching that addresses both problems. The method benefits from treating the two problems as a single optimization task. It automatically finds corresponding particle pairs and accurately computes tilt-axis direction even in the cases when EM grid is not perfectly planar.
Electron tomography of intact cells has the potential to reveal the entire cellular content at a resolution corresponding to individual macromolecular complexes. Characterization of macromolecular complexes in tomograms is nevertheless an extremely challenging task due to the high level of noise, and due to the limited tilt angle that results in missing data in Fourier space. By identifying particles of the same type and averaging their 3D volumes, it is possible to obtain a structure at a more useful resolution for biological interpretation. Currently, classification and averaging of sub-tomograms is limited by the speed of computational methods that optimize alignment between two sub-tomographic volumes. The alignment optimization is hampered by the fact that the missing data in Fourier space has to be taken into account during the rotational search. A similar problem appears in single particle electron microscopy where the random conical tilt procedure may require averaging of volumes with a missing cone in Fourier space. We present a fast implementation of a method guaranteed to find an optimal rotational alignment that maximizes the constrained cross-correlation function (cCCF) computed over the actual overlap of data in Fourier space.
Cell membranes represent the "front line" of cellular defense and the interface between a cell and its environment. To determine the range of proteins and protein complexes that are present in the cell membranes of a target organism, we have utilized a "tagless" process for the system-wide isolation and identification of native membrane protein complexes. As an initial subject for study, we have chosen the Gram-negative sulfate-reducing bacterium Desulfovibrio vulgaris. With this tagless methodology, we have identified about two-thirds of the outer membrane- associated proteins anticipated. Approximately three-fourths of these appear to form homomeric complexes. Statistical and machine-learning methods used to analyze data compiled over multiple experiments revealed networks of additional protein-protein interactions providing insight into heteromeric contacts made between proteins across this region of the cell. Taken together, these results establish a D. vulgaris outer membrane protein data set that will be essential for the detection and characterization of environment-driven changes in the outer membrane proteome and in the modeling of stress response pathways. The workflow utilized here should be effective for the global characterization of membrane protein complexes in a wide range of organisms.
Biological macromolecules can adopt multiple conformational and compositional states due to structural flexibility and alternative subunit assemblies. This structural heterogeneity poses a major challenge in the study of macromolecular structure using single-particle electron microscopy. We propose a fully automated, unsupervised method for the three-dimensional reconstruction of multiple structural models from heterogeneous data. As a starting reference, our method employs an initial structure that does not account for any heterogeneity. Then, a multi-stage clustering is used to create multiple models representative of the heterogeneity within the sample. The multi-stage clustering combines an existing approach based on Multivariate Statistical Analysis to perform clustering within individual Euler angles, and a newly developed approach to sort out class averages from individual Euler angles into homogeneous groups. Structural models are computed from individual clusters. The whole data classification is further refined using an iterative multi-model projection-matching approach. We tested our method on one synthetic and three distinct experimental datasets. The tests include the cases where a macromolecular complex exhibits structural flexibility and cases where a molecule is found in ligand-bound and unbound states. We propose the use of our approach as an efficient way to reconstruct distinct multiple models from heterogeneous data.
An unbiased survey has been made of the stable, most abundant multi-protein complexes in Desulfovibrio vulgaris Hildenborough ( Dv H) that are larger than Mr ≈ 400 k. The quaternary structures for 8 of the 16 complexes purified during this work were determined by single-particle reconstruction of negatively stained specimens, a success rate ≈10 times greater than that of previous “proteomic” screens. In addition, the subunit compositions and stoichiometries of the remaining complexes were determined by biochemical methods. Our data show that the structures of only two of these large complexes, out of the 13 in this set that have recognizable functions, can be modeled with confidence based on the structures of known homologs. These results indicate that there is significantly greater variability in the way that homologous prokaryotic macromolecular complexes are assembled than has generally been appreciated. As a consequence, we suggest that relying solely on previously determined quaternary structures for homologous proteins may not be sufficient to properly understand their role in another cell of interest.
We propose a feature-based image alignment method for single-particle electron microscopy that is able to accommodate various similarity scoring functions while efficiently sampling the two-dimensional transformational space. We use this image alignment method to evaluate the performance of a scoring function that is based on the Mutual Information (MI) of two images rather than one that is based on the cross-correlation function. We show that alignment using MI for the scoring function has far less model-dependent bias than is found with cross-correlation based alignment. We also demonstrate that MI improves the alignment of some types of heterogeneous data, provided that the signal-to-noise ratio is relatively high. These results indicate, therefore, that use of MI as the scoring function is well suited for the alignment of class-averages computed from single-particle images. Our method is tested on data from three model structures and one real dataset.
Primary amino acid content and the geometry of the folded protein 3D structure are major parameters of protein function. During the course of evolution the protein 3D structure is more preserved than its primary sequence. Thus, analysis of protein structures is expected to lead to a deep insight into protein function. Recognition of a structural core common to a set of protein structures serves as a basic tool for the studies of protein evolution and classification, analysis of similar structural motifs and functional binding sites, and for homology modeling and threading. In this chapter, we discuss several biologically related computational aspects of the multiple structure alignment and propose a method that provides solutions to these problems. Finally, we address the problem of structure-based multiple sequence alignment and propose an optimization method that unifies primary sequence and 3D structure information.
Prediction of protein–RNA interactions at the atomic level of detail is crucial for our ability to understand and interfere with processes such as gene expression and regulation. Here, we investigate protein binding pockets that accommodate extruded nucleotides not involved in RNA base pairing. We observed that most of the protein-interacting nucleotides are part of a consecutive fragment of at least two nucleotides whose rings have significant interactions with the protein. Many of these share the same protein binding cavity and more than 30% of such pairs are π-stacked. Since these local geometries cannot be inferred from the nucleotide identities, we present a novel framework for their prediction from the properties of protein binding sites. First, we present a classification of known RNA nucleotide and dinucleotide protein binding sites and identify the common types of shared 3-D physicochemical binding patterns. These are recognized by a new classification methodology that is based on spatial multiple alignment. The shared patterns reveal novel similarities between dinucleotide binding sites of proteins with different overall sequences, folds and functions. Given a protein structure, we use these patterns for the prediction of its RNA dinucleotide binding sites. Based on the binding modes of these nucleotides, we further predict an RNA fragment that interacts with those protein binding sites. With these knowledge-based predictions, we construct an RNA fragment that can have a previously unknown sequence and structure. In addition, we provide a drug design application in which the database of all known small-molecule binding sites is searched for regions similar to nucleotide and dinucleotide binding patterns, suggesting new fragments and scaffolds that can target them.
Analysis of protein-ligand complexes and recognition of spatially conserved physico-chemical properties is important for the prediction of binding and function. Here, we present two webservers for multiple alignment and recognition of binding patterns shared by a set of protein structures. The first webserver, MultiBind (http://bioinfo3d.cs.tau.ac.il/MultiBind), performs multiple alignment of protein binding sites. It recognizes the common spatial chemical binding patterns even in the absence of similarity of the sequences or the folds of the compared proteins. The input to the MultiBind server is a set of protein-binding sites defined by interactions with small molecules. The output is a detailed list of the shared physico-chemical binding site properties. The second webserver, MAPPIS (http://bioinfo3d.cs.tau.ac.il/MAPPIS), aims to analyze protein-protein interactions. It performs multiple alignment of protein-protein interfaces (PPIs), which are regions of interaction between two protein molecules. MAPPIS recognizes the spatially conserved physico-chemical interactions, which often involve energetically important hot-spot residues that are crucial for protein-protein associations. The input to the MAPPIS server is a set of protein-protein complexes. The output is a detailed list of the shared interaction properties of the interfaces.
Background: Conservation of the spatial binding organizations at the level of physico-chemical interactions is important for the formation and stability of protein-protein complexes as well as protein and drug design. Due to the lack of computational tools for recognition of spatial patterns of interactions shared by a set of protein-protein complexes, the conservation of such interactions has not been addressed previously.Results: We performed extensive spatial comparisons of physico-chemical interactions common to different types of protein-protein complexes. We observed that 80% of these interactions correspond to known hot spots. Moreover, we show that spatially conserved interactions allow prediction of hot spots with a success rate higher than obtained by methods based on sequence or backbone similarity. Detection of spatially conserved interaction patterns was performed by our novel MAPPIS algorithm. MAPPIS performs multiple alignments of the physico-chemical interactions and the binding properties in three dimensional space. It is independent of the overall similarity in the protein sequences, folds or amino acid identities. We present examples of interactions shared between complexes of colicins with immunity proteins, serine proteases with inhibitors and T-cell receptors with superantigens. We unravel previously overlooked similarities, such as the interactions shared by the structurally different RNase-inhibitor families.Conclusion: The key contribution of MAPPIS is in discovering the 3D patterns of physicochemical interactions. The detected patterns describe the conserved binding organizations that involve energetically important hot spot residues and are crucial for the protein-protein associations.
Electron microscopy (EM) is one of the key experimental techniques in modern structural biology[4]. Three-dimensional structures of large macro-molecular complexes can be solved by EM which otherwise would be impossible using X-ray crystallography or NMR. A common reconstruction technique, from 2D EM images to 3D structure, assigns 2D images to 2D projections of some initial 3D template[5, 1]. Due to high levels of noise and unknown orientation of imaged particles, the quality of the reconstruction procedure relies on two factors: an accurate alignment between a raw image and a template, and a scoring function that should be able to assign a raw particle to a correct template projection. Most of the methods for particle alignment apply a very efficient Fast Fourier Transform to compute a superposition between two images. This approach implies that only a normalized cross-correlation (NCC) function, or its variants, can be used as a similarity measure between two images[5, 1]. One of the major problems in EM reconstruction is heterogeneous data. In such cases a single template model does not accurately account for the heterogeneous data. Consequently, an application of the cross-correlation function produces inaccurate alignments. Here we present an alignment method that is able to accommodate various similarity scoring functions while efficiently sampling the 2D transformational space. In our preliminary results we apply a scoring function based on Mutual Information of two images. For a heterogeneous sample containing incomplete molecules it allows accurate alignment using a model of the complete structure. In our preliminary results, we successfully tested our approach on a model data set containing a mixed population of 70S and 50S E. coli Ribosomes.
Cryo-EM has become an increasingly powerful technique for elucidating the structure, dynamics, and function of large flexible macromolecule assemblies that cannot be determined at atomic resolution. However, due to the relatively low resolution of cryo-EM data, a major challenge is to identify components of complexes appearing in cryo-EM maps. Here, we describe EMatch, a novel integrated approach for recognizing structural homologues of protein domains present in a 6-10 A resolution cryo-EM map and constructing a quasi-atomic structural model of their assembly. The method is highly efficient and has been successfully validated on various simulated data. The strength of the method is demonstrated by a domain assembly of an experimental cryo-EM map of native GroEL at 6 Aring resolution
In the era of structural genomics, comparing a large number of protein structures can be a dauntingly time-consuming task. Traditional structural alignment methods, although offer accurate comparison, are not fast enough. Therefore, a number of databases storing pre-computed structural similarities are created to handle structural comparison queries efficiently. However, these databases cannot be updated in a timely fashion due to the sheer burden of computational requirements, thus offering only a rigid classification by some predefined parameters. Therefore, there is an increasingly urgent need for algorithms that can rapidly compare a large set of structures. Recently proposed projection methods, e.g., [1,2,3,4,5], show good promise for the development of fast structural database search solutions. Projection methods map a structure into a point in a high dimensional space and compare two structures by measuring distance between their projected points. These methods offer a tremendous increase in speed over residue-level structural alignment methods. However, current projection methods are not practical, partly because they are unable to identify local similarities. We propose a new projection-based approach that can rapidly detect global as well as local structural similarities. Local structural search is enabled by a topology-based writhe decomposition protocol (inspired by [4]) that produces a small number of fragments while ensuring that similar structures are cut in a similar manner. In a benchmark test for local structural similarity detection, we show that our method, Writher, dramatically improves accuracy over current leading projection methods [4,5] in terms of recognizing SCOP domains out of multidomain proteins.
Routinely used multiple‐sequence alignment methods use only sequence information. Consequently, they may produce inaccurate alignments. Multiple‐structure alignment methods, on the other hand, optimize structural alignment by ignoring sequence information. Here, we present an optimization method that unifies sequence and structure information. The alignment score is based on standard amino acid substitution probabilities combined with newly computed three‐dimensional structure alignment probabilities. The advantage of our alignment scheme is in its ability to produce more accurate multiple alignments. We demonstrate the usefulness of the method in three applications: 1) computing more accurate multiple‐sequence alignments, 2) analyzing protein conformational changes, and 3) computation of amino acid structure‐sequence conservation with application to protein–protein docking prediction. The method is available at http://bioinfo3d.cs.tau.ac.il/staccato/ . Proteins 2006. © 2005 Wiley‐Liss, Inc.
Protein sequence and structure are fundamental objects in computational biology. The sequence comparison problem has been widely addressed, resulting in a spectrum of algorithms ranging from the sensitive ones such as profile-HMM to fast ones such as k-mer indexing, arguably culminated in BLAST, where a practical balance of sensitivity and speed is achieved. Current structural comparison methods achieve results generally satisfactory to biologists. However, fast and accurate data base searches, in spirit to BLAST, are not possible due to the nature of the structural comparison methodology. Similarity of protein structures is typically measured at the residue level via structural alignment, whose goal is to find a 3D transformation that brings into correspondence the largest number of atoms. The quality of a 3D superposition is typically measured by the number of matched C-alpha atoms and their RMSD. The exact solution for the pairwise structural alignment is computationally expensive [1]. Therefore, heuristic approaches have been developed to find a good solution efficiently (for a review see [3]). An alternative approach to assess protein structure similarity is based on global topological properties, for example, by means of writhe number [2] and Gauss integrals (GIs) [5], or by means of secondary structure footprints [6]. The advantage of this approach is that each structure is represented by a constant number of features. This concise representation tolerates small structural distortions. More importantly, unlike in the structural alignment approach, global topological features of proteins can be trivially compared in constant time, e.g. by the Euclidean distance of GI vectors [5]. This offers potential for fast database search. The global descriptor approach may suffer from drawbacks. First, it is unable to detect local similarities, i.e., matching of substructures. For example, it cannot detect the similarity between a single domain protein to one of the domains in a multidomain protein. Second, certain, relatively small, structural changes in a protein structure, e.g. loop movement or loop indels, may cause significant changes in a global descriptor. We propose a new scheme that unifies the above two approaches for structural comparison. Instead of using one global descriptor for the entire protein backbone we consider descriptors for all possible fragments [i, j]. The overall similarity between two structures can be defined as the sum of matching scores of a set of sequential, non-overlapping (or not-so-much overlapping) fragment pairs, normalized by their lengths. We designed a dynamic programming algorithm variant to calculate the optimal matching. The similarity between a pair of segments is measured in the same fashion as in [5]. We reduce the running time by exploiting the redundancy in the set of [i, j] descriptors. The running time of the dynamic programming method is Θ(n4) if all Θ(n2) fragments from each protein are considered. However, we notice that the number of fragments whose descriptors are sufficient
John-Marc Chandonia合作论文数ASTRAL5