
A novel and generally applicable method is described for the detection of homology in distantly related proteins using a new domain sequence database that contains over 20,000 protein sequence segments of known function. The use of the method is illustrated on distantly related domains shared by complement components C1S and C1R, calcium-dependent serine proteinase and bone morphogenetic protein 1. New homologies are shown between human adducin and the actin-binding domains of alfa-actinin and dystrophin.
Two inhibitors (SI alpha 4 and SI alpha 5) of the alpha-amylases from insect and mammalian sources were purified from seeds of Sorghum bicolor by saline extraction, precipitation with ammonium sulphate, affinity chromatography on Red Sepharose, and preparative and analytical reverse-phase HPLC on columns of Vydac C18. The complete primary structures of these two inhibitors were determined by automated degradation of the intact, reduced and S-alkylated proteins and by manual 4-N,N-dimethylaminoazobenzene-4-isothiocyanate/phenyl isothiocyanate microsequencing of peptides derived from them following enzyme digests. The amino acid sequences were as follows: SI alpha 4: TVDVTACAPGLAIPAPPLPTCRTFARPRTCGLGGPYGPVDPSPVLKQ- RCCRELAAVPSRCRCAALGFMMDGVDAPLQDFRGCTREMQRIYAVSRLTRAAECNLPTIPGGGCHLSNS PR; and SI alpha 5: ANWCEPGLVIPLNPLPSCRTYMVRRACGVSIGPVVPLPVLKERCCSELEKLV- PYCRCGALRTALDSMMTGYEMRPTCSWGGLLTFAPTIVCYRECNLRTLHGRPFCYALGAEGTTT. Comparisons of these sequences with one another and with those of other proteins in the US National Biomedical Research Foundation Databank indicated that the two Sorghum proteins had significant similarities (21%-42% identity) with the members of the cereal superfamily of enzyme inhibitors.
The acidic ribosomal protein family of eukaryotic cells is thought to form a complex on ribosomes mainly by hydrophobic forces. To investigate the structural basis of how they associate with one another, the primary sequences of the related proteins accumulated from various organisms were analyzed searching for evolutionarily conserved hydrophobic motifs. Initially it is shown that all the P1-type 13-kDa proteins contain a bilateral hydrophobic zipper on a putative alpha-helix, which consists of two periodic arrays of hydrophobic amino acid residues arranged on the opposite sides of an alpha-helix. The P2-type 13-kDa proteins, except for those from the yeast Saccharomyces cerevisiae, are shown to contain two kinds of hydrophobic areas on putative alpha-helices, which can sterically bind to each other in a hook-and-eye fashion. On the other hand, the 38-kDa proteins contain a hydrophobic zipper and a hydrophobic hook in different helical regions. Thus, it is proposed that the 13-kDa proteins associate with the 38-kDa proteins via the hydrophobic zipper or hydrophobic hook-and-eye, and associate with one another with these hydrophobic elements.
The presence of 22-residue repeats, each with a preferential potential to form an amphipathic alpha-helix, is a unique feature of the plasma apolipoproteins. There are 27 such repeats in the three human apolipoproteins A-I, A-IV, and E. The extent of similarities and differences among these repeats have been estimated by computing correlation coefficients, Dayhoff scores, secondary structure difference profiles, and discrete Fourier transforms. The results reveal that there is a high level of similarity among the repeats of apo A-IV, and a low level of similarity in the repeats of apo E. Within each protein, similarity among some specified repeat pairs is distinctively higher than the others. A high order of similarity is also found among certain segments of each protein with those in the other two. The repeats prefer a mostly alpha-helical structure that is amphipathic in nature. Among the repeats of the three proteins, those of apo E show a high level of divergence among themselves. A consensus alignment of the residues of the 27 repeats into a hydrophobic versus hydrophilic pattern brings to focus the possible specific structure-stabilizing factors, such as the leucine zipper and the salt bridge. The recently reported crystal structures of the human apolipoprotein E and locust apolipophorin-III support many of the predictions made in this study.
It is demonstrated that the amino acid sequences of the products of E. coli genes sbcC and prrC, and bacteriophage P2 gene old encompass the four conserved motifs typical of the superfamily of UvrA-related ATPases. A more pronounced statistically significant similarity was revealed between SbcC protein, bacteriophage T4 endonuclease component gp46 and bacteriophage T5 protein D13. It is suggested that the newly identified members of the superfamily might all be ATPase components of the respective nucleases, and that the reactions catalyzed by these enzymes are probably ATP dependent.
Two forms of canine pancreatic kallikrein, designated as canine pancreatic kallikrein A and B, were separately isolated by ion-exchange, affinity and hydrophobic chromatographies. These enzymes had similar apparent molecular masses, substrate specificities and pH optima. However, kallikrein B was inhibited by soybean trypsin inhibitor, while kallikrein A was not. Both kallikrein A and B were shown by sodium dodecyl sulfate-polyacryl amide gel electrophoresis to consist of two polypeptide chains, designated alpha and beta chains, and binding by disulfide bond(s). The N-terminal amino acid sequences of each alpha and beta chains of kallikrein A and B were determined.
The expanding family of cytokines, interleukins and colony-stimulatory factors has made it difficult to readily access their structural and biological properties for comparative purposes. Here their aligned amino acid sequences, biological actions and some structural predictions are presented together for ready comparisons
The amino acid sequences of ferredoxin isoproteins (Fd A and Fd B) from Alocasia macrorrhiza Schott in Papua New Guinea were determined. They consisted of single polypeptide chains of 97 and 98 residues, respectively, and both Fds had a molecular mass of 10,800 Da. There was an 88% identity between the sequences of the isoproteins (Fd A and Fd B). These sequences were compared with those of the closely related plant Fds and their phylogenetic relationships are discussed.
The complete amino acid sequence of equine miniplasminogen (Mr 37,132, 338 residues) was determined with the aid of fragments obtained by cleavage with 2-(2-nitrophenylsulfenyl)-3-methyl-3'-bromoindolenine, cyanogen bromide or clostripain. The fragments were aligned with overlapping sequences. Sequence comparison with other species gave identities in the range of 76% (bovine) and 81% (canine), indicating the presence of the same structural and functional domains as in the other species. Sequence comparison of different miniplasminogens showed that positions 49 (Arg), 83 (Arg) and 161 (Ser) may play a role in the interaction between plasminogen and streptokinase.
Ubiquitin has been isolated and purified from rabbit brain using gel permeation and reverse-phase high-performance liquid chromatography. The 76-residue protein exhibits one difference towards a murine form, is identical to other characterized vertebrate ubiquitins, and confirms an extensive conservation of the ubiquitin structure. No positional microheterogeneities were detectable between two sub-forms.
The presence of two types of kallikrein inhibitor (cationic and anionic inhibitors) was demonstrated in bovine pituitary gland. These kallikrein inhibitors were separated from the homogenate of bovine posterior pituitary by successive CM-Sephadex chromatography. The major cationic inhibitor was further purified to homogeneity by affinity chromatography using porcine pancreatic beta-kallikrein immobilized on Sepharose 4B and gel filtration. The complete amino acid sequence of this inhibitor was first determined, and it was shown to be a peptide of 58 residues with a calculated molecular weight of 6,511. The Ki value against bovine pituitary kallikrein was 6 x 10(-9) M. The cationic inhibitor was found to be identical with basic pancreatic trypsin inhibitor.
The 8-kDa protein in Photosystem I (PS I) reaction center complex was isolated from a thermophilic cyanobacterium, Synechococcus elongatus, by SDS-polyacrylamide gel electrophoresis using TRIS-Tricine buffer system. The complete amino acid sequence of the protein was determined. The 8-kDa protein consisted of 73 amino acid residues giving a calculated molecular weight of 7,472. No significant sequence homology were observed with the known other small subunits in PS I reaction center complex, except for the 6.5-kDa protein in PS I from another thermophilic cyanobacterium, S. vulcanus. The 8-kDa protein was characteristically rich in hydrophobic amino acid residues, especially the content of leucine. These suggest that the 8-kDa subunit is an intrinsic structure component in PS I core complex for stabilization of the reaction center.
Recently, we have developed a sequence-structure database of protein information, NRL_3D, that is extracted from the Protein Data Bank (PDB) of the Brookhaven National Laboratory. NRL_3D provides a vehicle for the retrieval of the three-dimensional coordinates of protein fragments as identified by sequence properties. These data are formulated to allow access by standard sequence analysis programs such as those provided by the Protein Identification Resource (PIR). Because the PDB is updated four times per year, semimanual construction of NRL_3D in coordination with these updates becomes a time-consuming and inefficient task. Hence, we have developed a computer program (PRENRL_3D) in the "C" computer language that automatically extracts NRL_3D from the PDB. Although the program was developed in a VAX/VMS environment, care was taken to ensure its portability to other computer systems. Customized versions of the NRL_3D database can be created from the PDB entry files using various options available in PRENRL_3D, such as selection of entries determined at high resolution and with low R-value. The program has been developed modularly and it contains a number of generalized procedures for manipulating various information in the PDB.
Recently, we have developed a sequence-structure database of protein information, NRL_3D, that is extracted from the Protein Data Bank (PDB) of the Brookhaven National Laboratory. NRL_3D provides a vehicle for the retrieval of the three-dimensional coordinates of protein fragments as identified by sequence properties. These data are formulated to allow access by standard sequence analysis programs such as those provided by the Protein Identification Resource (PIR). Because the PDB is updated four times per year, semimanual construction of NRL_3D in coordination with these updates becomes a time-consuming and inefficient task. Hence, we have developed a computer program (PRENRL_3D) in the "C" computer language that automatically extracts NRL_3D from the PDB. Although the program was developed in a VAX/VMS environment, care was taken to ensure its portability to other computer systems. Customized versions of the NRL_3D database can be created from the PDB entry files using various options available in PRENRL_3D, such as selection of entries determined at high resolution and with low R-value. The program has been developed modularly and it contains a number of generalized procedures for manipulating various information in the PDB.
A new method has been developed for detecting the similarity between distantly related families of proteins. The amino acid sequences of each family of proteins are vertically aligned by a homologous alignment method and the physico-chemical properties of the amino acid residues at the corresponding site are evaluated simultaneously, by method of the principal component analysis. Taking into account the species diversity of each family of proteins, we assign the similar regions between the different families of proteins by the overlapping degree of the standard deviations around the mean values of the first principal component. To investigate the homologous relationship between the electron transport proteins in photosynthetic and O2 respiratory systems, this method has been applied to 70 species of mitochondrial cytochrome c, 4 species of cytochrome c1 and 7 species of cytochrome f. This analysis reveals that both cytochrome f and cytochrome c1 have large regions which are similar to those of cytochrome c. Assuming that these similar regions have the same stereochemical structures as those in cytochrome c, we can predict the outlines of the tertiary structures of cytochrome c1 and cytochrome f, respectively, each able to interact with its electron acceptor, cytochrome c and plastocyanin.
The nucleotide sequences of extragenic regions in the Escherichia coli genome are statistically analyzed. Sequence elements with high occurrence frequencies are identified; these elements are: (1) extragenic palindromic sequences, which are markedly distinguishable from the already identified repetitive extragenic palindromic sequences; (2) promoter sequences of purine biosynthetic genes; and (3) rho-independent terminator sequences. The repetitious occurrence and extensive sequence similarities suggest that these elements share common evolutionary origins. Copies of one sequence element would have become distributed to various positions on the genome during evolution and have been fixed at locations that provide a selective advantage. The extragenic regions of the E. coli genome seem to consist of various regulatory 'building blocks', similar to a protein which consists of modules or domains.
Aconitase has been purified from membranes prepared from both the human gastric carcinoma cell line Okajima and from porcine gastric mucosa by chromatography on concanavalin A-Sepharose and carboxymethyl-Sepharose, and preparative polyacrylamide gel electrophoresis. Automated Edman degradation of the intact proteins yielded no N-terminal amino acid sequence due, presumably, to N-terminal blockage. Sequence analysis of tryptic peptides derived from S-carboxymethyl porcine and human aconitases established the positions of 95 and 64 amino acid residues, respectively. The amino acid sequence data for porcine aconitase was in perfect agreement with the previously reported cDNA-deduced amino acid sequence [Zheng et al. (1990) J Biol Chem 265:2814-2821]. Comparison of the human amino acid sequence data with the cDNA-deduced amino acid sequence of porcine aconitase indicated that these two proteins have 95% amino acid sequence identity within the sequenced region.