The third-generation sequencing technology, PacBio, has shown an ability to sequence the HIV virus amplicons in their full length. The long read of PaBio offers a distinct advantage to comprehensively understand the virus evolution complexity at quasispecies level (i.e. maintaining linkage information of variants) comparing to the short reads from Illumina shotgun sequencing. However, due to the highnoise nature of the PacBio reads, it is still a challenge to build accurate contigs at high sensitivity. Most of previously developed NGS assembly tools work with the assumption that the input reads are fairly accurate, which is largely true for the data derived from Sanger or Illumina technologies. When applying these tools on PacBio high-noise reads, they are largely driven by noise rather than true signal eventually leading to poor results in most cases. In this study, we propose the de novo assembly procedure, which comprises a positivefocused strategy, and linkage-frequency noise reduction so that it is more suitable for PacBio high-noise reads. We further tested the unique de novo assembly procedure on HIV PacBio benchmark data and clinical samples, which accurately assembled dominant and minor populations of HIV quasispecies as expected. The improved de novo assembly procedure shows potential ability to promote PacBio technology in the field of HIV drug-resistance clinical detection, as well as in broad HIV phylogenetic studies.
A query protein structure is compared with the VAST program to a database of target structures from the PDB (PDB40, list of protein structures having less than 40% of identical residues: 19 500 structures version 2011). The threshold of the VAST program is lowered in order to find the largest possible number of structures having a local similarity with the query protein. The purpose of the web site is to define structural domains in the query protein using the recurrence of these locally similar substructures. (http://genome.jouy.inra.fr/domire/). The list of matches is subsequently sorted according to two criteria: the number of aligned residues by VAST is at least 40% of the number of residues of the target, and 80% of the target length is aligned including gaps of non aligned residues if less than 40. Besides this list, a residue-residue alignment of the structural neighbour on the amino acid sequence of the query protein is provided together with a 3D view of their superposition. The object of this sorting is to help in detecting remote homologues and isolated protein structures matching the domain structures of a protein.
the long fiber axis while the b-strands run perpendicular (cross-b) to the fiber axis. b-helices are a parallel b-sheet motif that have similar properties to amyloid fibers. In this work, we explore the utility of the b-helix as a building block to capture and test attributes of amyloid fibers relevant to the design of selfassembling nanostructures. We have isolated and characterized a five-rung portion of the triple-stranded, homotrimeric b-helix, GP5, of the bacteriophage T4. In order to construct a self-assembling b-helix, we have developed a small library based on amyloid sequence analysis and applied it to the five-rung helix. Specific changes targeting internal interactions of the helix have been made. The mutations in these positions have been selected to allow for greater control of self-assembly and hold the promise of breaking the homomeric assembly into a heteromeric construction. A split adenylate kinase selection assay has been developed to monitor self-assembly of these alternative structures.
SUMMARY:The DOMIRE web server implements a novel, automatic, protein structural domain assignment procedure based on 3D substructures of the query protein which are also found within structures of a non-redundant protein database. These common 3D substructures are transformed into a co-occurrence matrix that offers a global view of the protein domain organization. Three different algorithms are employed to define structural domain boundaries from this co-occurrence matrix. For each query, a list of structural neighbors and their alignments are provided. DOMIRE, by displaying the protein structural domain organization, can be a useful tool for defining protein common cores and for unravelling the evolutionary relationship between different proteins.AVAILABILITY:http://genome.jouy.inra.fr/domireCONTACT:jean.garnier@jouy.inra.fr.
Domains are basic units of protein structure and essential for exploring protein fold space and structure evolution. With the structural genomics initiative, the number of protein structures in the Protein Databank (PDB) is increasing dramatically and domain assignments need to be done automatically. Most existing structural domain assignment programs define domains using the compactness of the domains and/or the number and strength of intra‐domain versus inter‐domain contacts. Here we present a different approach based on the recurrence of locally similar structural pieces (LSSPs) found by one‐against‐all structure comparisons with a dataset of 6373 protein chains from the PDB. Residues of the query protein are clustered using LSSPs via three different procedures to define domains. This approach gives results that are comparable to several existing programs that use geometrical and other structural information explicitly. Remarkably, most of the proteins that contribute the LSSPs defining a domain do not themselves contain the domain of interest. This study shows that domains can be defined by a collection of relatively small locally similar structural pieces containing, on average, four secondary structure elements. In addition, it indicates that domains are indeed made of recurrent small structural pieces that are used to build protein structures of many different folds as suggested by recent studies. Proteins 2011. Published 2010 Wiley‐Liss, Inc.
Domains are basic units of protein structure and essential for exploring protein fold space and structure evolution. With the NIH Protein Structure Initiative and other structural genomics initiatives worldwide, the number of protein structures in PDB is increasing dramatically and domain parsing needs to be done automatically. Most of the existing structural domain parsing programs consider the compactness of the domains and/or the number and strength of internal (intra-domain) versus external (inter-domain) contacts. Here we present a completely different approach. Taking advantage of the growing number of known structures in the PDB, the chains are parsed solely by using recurrence of similar structures that appear in the structural database. A non-redundant set of 6373 protein chains was selected as the target data set and 128 benchmark chains from pDomains were used as query chains. For each query chain, one against all target structure comparisons were performed using VAST. Then the VAST cliques were collected and the protein residues were clustered using mathematical procedures akin to those used for analyzing the microarray data. These clusters define domains. NDO scores were used to compare the results with SCOP and CATH domain boundaries as well as with those from other parsing programs. Our algorithm gave results that were comparable to those of several existing programs. It handles segmented domains equally well as non-segmented domains. The structures that contribute the cliques that define a domain may contain distant evolutionary information of the domain.
BACKGROUND:Formal classification of a large collection of protein structures aids the understanding of evolutionary relationships among them. Classifications involving manual steps, such as SCOP and CATH, face the challenge of increasing volume of available structures. Automatic methods such as FSSP or Dali Domain Dictionary, yield divergent classifications, for reasons not yet fully investigated. One possible reason is that the pairwise similarity scores used in automatic classification do not adequately reflect the judgments made in manual classification. Another possibility is the difference between manual and automatic classification procedures. We explore the degree to which these two factors might affect the final classification.RESULTS:We use DALI, SHEBA and VAST pairwise scores on the SCOP C class domains, to investigate a variety of hierarchical clustering procedures. The constructed dendrogram is cut in a variety of ways to produce a partition, which is compared to the SCOP fold classification.Ward's method dendrograms led to partitions closest to the SCOP fold classification. Dendrogram- or tree-cutting strategies fell into four categories according to the similarity of resulting partitions to the SCOP fold partition. Two strategies which optimize similarity to SCOP, gave an average of 72% true positives rate (TPR), at a 1% false positive rate. Cutting the largest size cluster at each step gave an average of 61% TPR which was one of the best strategies not making use of prior knowledge of SCOP. Cutting the longest branch at each step produced one of the worst strategies. We also developed a method to detect irreducible differences between the best possible automatic partitions and SCOP, regardless of the cutting strategy. These differences are substantial. Visual examination of hard-to-classify proteins confirms our previous finding, that global structural similarity of domains is not the only criterion used in the SCOP classification.CONCLUSION:Different clustering procedures give rise to different levels of agreement between automatic and manual protein classifications. None of the tested procedures completely eliminates the divergence between automatic and manual protein classifications. Achieving full agreement between these two approaches would apparently require additional information.
BACKGROUND:Current classification of protein folds are based, ultimately, on visual inspection of similarities. Previous attempts to use computerized structure comparison methods show only partial agreement with curated databases, but have failed to provide detailed statistical and structural analysis of the causes of these divergences.RESULTS:We construct a map of similarities/dissimilarities among manually defined protein folds, using a score cutoff value determined by means of the Receiver Operating Characteristics curve. It identifies folds which appear to overlap or to be "confused" with each other by two distinct similarity measures. It also identifies folds which appear inhomogeneous in that they contain apparently dissimilar domains, as measured by both similarity measures. At a low (1%) false positive rate, 25 to 38% of domain pairs in the same SCOP folds do not appear similar. Our results suggest either that some of these folds are defined using criteria other than purely structural consideration or that the similarity measures used do not recognize some relevant aspects of structural similarity in certain cases. Specifically, variations of the "common core" of some folds are severe enough to defeat attempts to automatically detect structural similarity and/or to lead to false detection of similarity between domains in distinct folds. Structures in some folds vary greatly in size because they contain varying numbers of a repeating unit, while similarity scores are quite sensitive to size differences. Structures in different folds may contain similar substructures, which produce false positives. Finally, the common core within a structure may be too small relative to the entire structure, to be recognized as the basis of similarity to another.CONCLUSION:A detailed analysis of the entire available protein fold space by two automated similarity methods reveals the extent and the nature of the divergence between the automatically determined similarity/dissimilarity and the manual fold type classifications. Some of the observed divergences can probably be addressed with better structure comparison methods and better automatic, intelligent classification procedures. Others may be intrinsic to the problem, suggesting a continuous rather than discrete protein fold space.