The study of protein structures’ local conformations has a long history principally based on the analysis of the classical repetitive structures (i.e. α-helix and β-sheet), and also on the characterization of some particular structures in the coil state (e.g. turns). The secondary structures are interesting for describing the global protein fold but miss all the orientations of the connecting regions and so neglect many particularities of the coil state. In order to take these structural features into account, we have identified a local structural alphabet composed of 16 folding patterns of five consecutive residues, called Protein Blocks (PBs). Conversely to the secondary structures, the PBs are able to approximate every part of the protein structures. These PBs have been used both to describe precisely the 3D protein backbones with an average rmsd of 0.42 Å, and to perform a local structure prediction with a rate of correct prediction of 48.7%. In this chapter, we present the interest of the Protein Blocks by comparing the secondary structure assignment with the assignment in terms of PBs. We highlight the discrepancies between different secondary structure assignment methods and show some interesting correspondence between particular local folds and the Protein Blocks. Then, we use the Protein Block prediction to classify proteins into the classical structural classes, namely all α, all β and mixed. The prediction rate of theses different classes is good, i.e. 71.5%, with no confusion between all α and all β classes. Finally, we present a new approach named TopKAPi that stands for “Triangular Kohonen Map for Analyzing Proteins”. It enables to classify and analyze proteins H A L athor m anscript inerm -004564, version 1
The description of protein 3D structures can be performed through a library of 3D fragments, named a structural alphabet. Our structural alphabet is composed of 16 small protein fragments of 5 Cα in length, called protein blocks (PBs). It allows an efficient approximation of the 3D protein structures and a correct prediction of the local structure. The 72 most frequent series of 5 consecutive PBs, called structural words (SWs) are able to cover more than 90% of the 3D structures. PBs are highly conditioned by the presence of a limited number of transitions between them. In this study, we propose a new method called “pinning strategy” that used this specific feature to predict long protein fragments. Its goal is to define highly probable successions of PBs. It starts from the most probable SW and is then extended with overlapping SWs. Starting from an initial prediction rate of 34.4%, the use of the SWs instead of the PBs allows a gain of 4.5%. The pinning strategy simply applied to the SWs increases the prediction accuracy to 39.9%. In a second step, the sequence-structure relationship is optimized, the prediction accuracy reaches 43.6%.
Protein sequence world is considerably larger than structure world. In consequence, numerous non-related sequences may adopt similar 3D folds and different kinds of amino acids may thus be found in similar 3D structures. By grouping together the 20 amino acids into a smaller number of representative residues with similar features, sequence world simplification may be achieved. This clustering hence defines a reduced amino acid alphabet (reduced AAA). Numerous works have shown that protein 3D structures are composed of a limited number of building blocks, defining a structural alphabet. We previously identified such an alphabet composed of 16 representative structural motifs (5-residues length) called Protein Blocks (PBs). This alphabet permits to translate the structure (3D) in sequence of PBs (1D). Based on these two concepts, reduced AAA and PBs, we analyzed the distributions of the different kinds of amino acids and their equivalences in the structural context. Different reduced sets were considered. Recurrent amino acid associations were found in all the local structures while other were specific of some local structures (PBs) (e.g Cysteine, Histidine, Threonine and Serine for the α-helix Ncap). Some similar associations are found in other reduced AAAs, e.g Ile with Val, or hydrophobic aromatic residues Trp with Phe and Tyr. We put into evidence interesting alternative associations. This highlights the dependence on the information considered (sequence or structure). This approach, equivalent to a substitution matrix, could be useful for designing protein sequence with different features (for instance adaptation to environment) while preserving mainly the 3D fold.
Three‐dimensional protein structures can be described with a library of 3D fragments that define a structural alphabet. We have previously proposed such an alphabet, composed of 16 patterns of five consecutive amino acids, called Protein Blocks (PBs). These PBs have been used to describe protein backbones and to predict local structures from protein sequences. The Q 16 prediction rate reaches 40.7% with an optimization procedure. This article examines two aspects of PBs. First, we determine the effect of the enlargement of databanks on their definition. The results show that the geometrical features of the different PBs are preserved (local RMSD value equal to 0.41 Å on average) and sequence–structure specificities reinforced when databanks are enlarged. Second, we improve the methods for optimizing PB predictions from sequences, revisiting the optimization procedure and exploring different local prediction strategies. Use of a statistical optimization procedure for the sequence–local structure relation improves prediction accuracy by 8% ( Q 16 = 48.7%). Better recognition of repetitive structures occurs without losing the prediction efficiency of the other local folds. Adding secondary structure prediction improved the accuracy of Q 16 by only 1%. An entropy index ( N eq ), strongly related to the RMSD value of the difference between predicted PBs and true local structures, is proposed to estimate prediction quality. The N eq is linearly correlated with the Q 16 prediction rate distributions, computed for a large set of proteins. An “expected” prediction rate Q E 16 is deduced with a mean error of 5%. Proteins 2005. © 2005 Wiley‐Liss, Inc.
We developed a novel approach for predicting local protein structure from sequence. It relies on the Hybrid Protein Model (HPM), an unsupervised clustering method we previously developed. This model learns three‐dimensional protein fragments encoded into a structural alphabet of 16 protein blocks (PBs). Here, we focused on 11‐residue fragments encoded as a series of seven PBs and used HPM to cluster them according to their local similarities. We thus built a library of 120 overlapping prototypes (mean fragments from each cluster), with good three‐dimensional local approximation, i.e., a mean accuracy of 1.61 Å Cα root‐mean‐square distance. Our prediction method is intended to optimize the exploitation of the sequence‐structure relations deduced from this library of long protein fragments. This was achieved by setting up a system of 120 experts, each defined by logistic regression to optimize the discrimination from sequence of a given prototype relative to the others. For a target sequence window, the experts computed probabilities of sequence‐structure compatibility for the prototypes and ranked them, proposing the top scorers as structural candidates. Predictions were defined as successful when a prototype <2.5 Å from the true local structure was found among those proposed. Our strategy yielded a prediction rate of 51.2% for an average of 4.2 candidates per sequence window. We also proposed a confidence index to estimate prediction quality. Our approach predicts from sequence alone and will thus provide valuable information for proteins without structural homologs. Candidates will also contribute to global structure prediction by fragment assembly. Proteins 2006. © 2005 Wiley‐Liss, Inc.
Prediction of protein three-dimensional structures constitutes a major scientific stake. However, it remains difficult even though it is established that the information needed to specify the complex 3D structure of a protein is contained in its amino acid sequence. Our goal is to analyze protein structures at a local level since protein local structural information is encoded to an extent in local amino acid sequences. In our study, we used an unsupervised clustering method called “Hybrid Protein Model” (HPM) to perform the compression of a nonredundant protein structure databank into a library of overlapping 3D structural fragments [1]. The library obtained is composed of a limited number of structural classes grouping together fragments sharing similar local structures. These classes are characterized by average structural prototypes and are representative of all local folds observed in the protein structure databank. From this library, we analyzed the relation between amino acid sequence and local structure. We characterized for each structural prototype the amino acid specificities and preferences.
Predicting protein structure from amino acid sequence is one of the main challenges of genomics. Various computational methods have been developed during the last decade to reach this goal. However, the problem of structure prediction remains difficult. Before facing this complex problem, our goal is to focus on the accurate analysis of protein structures at a local level. In our study, we present an approach called "hybrid protein model" (HPM) which uses a training procedure similar to the one of the self-organizing maps. It allows the compression of a non-redundant protein structure databank into a library of overlapping 3D structural fragments. The "hybrid protein model" carries out a multiple alignment of structural fragments. We present in this study an improvement of this strategy by introducing gaps in the local structures, and a sensitivity study of the training according to the control parameters. The library obtained is composed of a finite number of structural classes, each class including fragments sharing similar local structures. These classes are representative of the structural motifs found in the protein structures from the databank. Thus, this library constitutes an efficient tool for determining structural similarities between proteins and especially for predicting the local protein structure from the amino acid sequence.