The accuracy of multiple sequence alignment program MAFFT has been improved. The new version (5.3) of MAFFT offers new iterative refinement options, H-INS-i, F-INS-i and G-INS-i, in which pairwise alignment information are incorporated into objective function. These new options of MAFFT showed higher accuracy than currently available methods including TCoffee version 2 and CLUSTAL W in benchmark tests consisting of alignments of >50 sequences. Like the previously available options, the new options of MAFFT can handle hundreds of sequences on a standard desktop computer. We also examined the effect of the number of homologues included in an alignment. For a multiple alignment consisting of ∼8 sequences with low similarity, the accuracy was improved (2–10 percentage points) when the sequences were aligned together with dozens of their close homologues (E-value < 10−5–10−20) collected from a database. Such improvement was generally observed for most methods, but remarkably large for the new options of MAFFT proposed here. Thus, we made a Ruby script, mafftE.rb, which aligns the input sequences together with their close homologues collected from SwissProt using NCBI-BLAST.
A multiple sequence alignment program, MAFFT, has been developed. The CPU time is drastically reduced as compared with existing methods. MAFFT includes two novel techniques. (i) Homo logous regions are rapidly identified by the fast Fourier transform (FFT), in which an amino acid sequence is converted to a sequence composed of volume and polarity values of each amino acid residue. (ii) We propose a simplified scoring system that performs well for reducing CPU time and increasing the accuracy of alignments even for sequences having large insertions or extensions as well as distantly related sequences of similar length. Two different heuristics, the progressive method (FFT-NS-2) and the iterative refinement method (FFT-NS-i), are implemented in MAFFT. The performances of FFT-NS-2 and FFT-NS-i were compared with other methods by computer simulations and benchmark tests; the CPU time of FFT-NS-2 is drastically reduced as compared with CLUSTALW with comparable accuracy. FFT-NS-i is over 100 times faster than T-COFFEE, when the number of input sequences exceeds 60, without sacrificing the accuracy.
An important biochemical feature of autotrophs, land plants and algae, is their incorporation of inorganic nitrogen, nitrate and ammonium, into the carbon skeleton. Nitrate and ammonium are converted into glutamine and glutamate to produce organic nitrogen compounds, for example proteins and nucleic acids. Ammonium is not only a preferred nitrogen source but also a key metabolite, situated at the junction between carbon metabolism and nitrogen assimilation, because nitrogen compounds can choose an alternative pathway according to the stages of their growth and environmental conditions. The enzymes involved in the reactions are nitrate reductase (EC 1.6.6.1-2), nitrite reductase (EC 1.7.7.1), glutamine synthetase (EC 6.3.1.2), glutamate synthase (EC 1.4.1.13-14, 1.4.7.1), glutamate dehydrogenase (EC 1.4.1.2-4), aspartate aminotransferase (EC 2.6.1.1), asparagine synthase (EC 6.3.5.4), and phosphoenolpyruvate carboxylase (EC 4.1.1.31). Many of these enzymes exist in multiple forms in different subcellular compartments within different organs and tissues, and play sometimes overlapping and sometimes distinctive roles. Here, we summarize the biochemical characteristics and the physiological roles of these enzymes. We also analyse the molecular evolution of glutamine synthetase, glutamate synthase and glutamate dehydrogenase, and discuss the evolutionary relationships of these three enzymes.
Previously we showed that the evolutionary rates of the Pax proteins are markedly reduced in higher vertebrates, as compared with those in the ancestral lineage of vertebrates, and we suggested that the reduced Pax protein evolution might be explained by increased functional constraints due to gene recruitment for other purposes or repeated expression in different developmental stages. To clarify the problem of whether the evolutionary rate variation found in the Pax proteins is an evolutionary feature generally recognized in most transcription factors, we have cloned and sequenced cDNAs encoding the TATA-box binding protein (TBP), a general transcription factor of eukaryotes, from Oryzias latipes, a Japanese medaka, Lampetra reissneri, a lamprey, and Ephydatia fluviatilis, a freshwater sponge. An evolutionary rate analysis of TBP has revealed that the evolutionary rate of TBP is extremely low in higher vertebrates, but not in the ancestral lineage of vertebrates, as found in the Pax proteins. In contrast, no marked reduction of the evolutionary rate in higher vertebrates is observed in the aldolase C, a house keeping enzyme. It is therefore likely that the increased functional constraint on TBP is responsible for the extremely low evolutionary rate in higher vertebrates. The temporal pattern of the evolutionary rate variation during vertebrate evolution was discussed.
In mammals and the amphibian, Xenopus, isotypes of antibodies have been shown to be changed through class switch recombination within the IgH chain gene locus. Here, we identified switch (S) repetitive sequences in the 5' introns of the Ig C(mu) and C(gamma) genes of the chicken. The S(mu) region is composed of two homologous regions, S(mu)1 and S(mu)2. The S(mu)1 region is an upstream 3.7 kb sequence composed of 37 repeats of a consensus sequence containing tandem repeats of the decamer ACCAGTATGG. The S(mu)2 region is a downstream 1.4 kb sequence consisting of simple tandem repeats of a decamer CCCAGTACAG. The S(gamma) region contains repeats of the decamer TATGGGGCAG. Analysis of chicken IgG-producing hybridomas revealed that the C(mu) gene was deleted from the chromosome by the recombination occurring between the S(mu) and S(gamma) regions. Recombination breakpoints at the C(mu) gene of splenocytes from an immunized chicken were scattered around the S(mu) region and two such breakpoints, the precise position of which were determined, were located within possible hairpin loop structures at the palindromic sequence of S(mu)1. A primordial palindromic sequence from which the prevalent switch repeat motifs of mammals, chickens and amphibians may have diverged is presented.
The animal cyclic nucleotide phosphodiesterases (PDEs) comprise at least seven subtypes, PDE1–7, which differ from each other in domain organization and primary function, and they diverged from an ancestral gene by gene duplication and domain shuffling during animal evolution. To obtain rough estimates for the divergence times of these subtypes, cloning of PDE cDNAs from Ephydatia fluviatilis (freshwater sponge) by RT‐PCR was carried out. We obtained four cDNAs, EFPDE1, EFPDE2, EFPDE3, and EFPDE4, which are possibly homologs of the vertebrate PDE1, PDE2, PDE3, and PDE4, respectively, judging from the sequence similarity, domain organization, and branching pattern in the phylogenetic tree. The phylogenetic tree of the PDE family revealed that most gene duplications and domain shufflings that gave rise to different subtypes had been completed in the early evolution of animals before the separation of sponges and eumetazoans.
A novel human protein with a molecular mass of 55 kD, designated RanBPM, was isolated with the two-hybrid method using Ran as a bait. Mouse and hamster RanBPM possessed a polypeptide identical to the human one. Furthermore, Saccharomyces cerevisiae was found to have a gene, YGL227w, the COOH-terminal half of which is 30% identical to RanBPM. Anti-RanBPM antibodies revealed that RanBPM was localized within the centrosome throughout the cell cycle. Overexpression of RanBPM produced multiple spots which were colocalized with gamma-tubulin and acted as ectopic microtubule nucleation sites, resulting in a reorganization of microtubule network. RanBPM cosedimented with the centrosomal fractions by sucrose- density gradient centrifugation. The formation of microtubule asters was inhibited not only by anti- RanBPM antibodies, but also by nonhydrolyzable GTP-Ran. Indeed, RanBPM specifically interacted with GTP-Ran in two-hybrid assay. The central part of asters stained by anti-RanBPM antibodies or by the mAb to gamma-tubulin was faded by the addition of GTPgammaS-Ran, but not by the addition of anti-RanBPM anti- bodies. These results provide evidence that the Ran-binding protein, RanBPM, is involved in microtubule nucleation, thereby suggesting that Ran regulates the centrosome through RanBPM.
The complete nucleotide sequence of the 957-kb DNA of the human immunoglobulin heavy chain variable (VH) region locus was determined and 43 novel VH segments were identified. The region contains 123 VH segments classifiable into seven different families, of which 79 are pseudogenes. Of the 44 VH segments with an open reading frame, 39 are expressed as heavy chain proteins and 1 as mRNA, while the remaining 4 are not found in immunoglobulin cDNAs. Combinatorial diversity of VH region was calculated to be ∼6,000. Conservation of the promoter and recombination signal sequences was observed to be higher in functional VH segments than in pseudogenes. Phylogenetic analysis of 114 VH segments clearly showed clustering of the VH segments of each family. However, an independent branch in the tree contained a single VH, V4-44.1P, sharing similar levels of homology to human VH families and to those of other vertebrates. Comparison between different copies of homologous units that appear repeatedly across the locus clearly demonstrates that dynamic DNA reorganization of the locus took place at least eight times between 133 and 10 million years ago. One nonimmunoglobulin gene of unknown function was identified in the intergenic region.
The protein tyrosine kinases (PTKs) are a large protein family consisting of many subfamilies with a variety of domain structures. The basic functions are thought to differ for different subfamilies. To know the dates at which the subfamilies diverged by gene duplications, a phylogenetic tree of the PTKs was inferred by comparing sequences from a wide range of species covering diploblasts and triploblasts. The PTK tree revealed that almost all of the gene duplications that gave rise to different subfamilies occurred rapidly before the diploblast-triploblast split, accompanying with rapid amino acid substitutions. This type of gene duplication was, however, rarely observed after that split. Long after the subfamily divergence, another type of gene duplication that gave rise to diverse tissue-specific genes occurred in each subfamily on the chordate lineage since the separation from arthropods. This type of gene duplication occurred frequently before the fish-tetrapod split, accompanying with rapid amino acid substitutions. In contrast, both the frequency of gene duplications and the rate of the amino acid substitutions were considerably reduced after that split. These results strongly suggest that the PTKs diverged intermittently, but not gradually, during animal evolution.
We identified a novel gene,kf-1,highly expressed in the normal cerebellum but not in the cerebral cortex, the expression of which could have been augmented in the cerebral cortex of a sporadic Alzheimer's disease patient. We cloned human and mouse entirekf-1cDNAs encoding conserved 79 kDa proteins containing a zinc-binding RING-H2 finger motif at the carboxy-terminus as found in acetylcholine receptor-associated protein (RAPsyn). The 3′-untranslated regions are highly conserved between human and mouse as to constitute a common mRNA secondary structure.In situhybridization analysis of mouse brain sections revealed strongkf-1expression in the cerebellum and hippocampus. We propose that KF-1 is involved in membranous protein-sorting apparatus similarly to RAPsyn. We mapped the humankf-1gene to 2p11.2.
To determine a possible relationship between organismal and molecular evolution, the divergence patterns of gene families were examined by taking special notice of functional difference, tissue distribution, and intracellular localization of the members. A phylogenetic analysis of 25 different gene families revealed interesting patterns of divergence of these families: Most gene duplications giving rise to different functions antedate the vertebrates-arthropods separation. On the other hand, in a group of members carrying virtually identical function to one another but differing in tissue distribution (tissue-specific isoform), most gene duplications have occurred independently in each of vertebrates and arthropods after the separation of the two animal groups. In family members encoding molecules localizing in cell compartments (compartmentalized isoforms), the gene duplications antedate the animals-fungi separation. In the cases of the Ca2+ pump and rab subfamilies, the compartmentalized isoforms were shown to have diverged during the early evolution of eukaryotes. A phylogenetic analysis of the tissue-specific isoforms from 26 different subfamilies revealed extensive gene duplications and rapid rates of amino acid substitutions in the early evolution of chordates before the separation of fishes and tetrapods. On the contrary, the genetic variations are relatively low in the later period. This pattern of evolution observed at the molecular level is correlated well with that of tissue evolution based on fossil evidence and morphological data, and thus evolution at the two levels may be related.
To estimate approximate times of divergence of animal phyla lacking fossil data, it is important to find a molecule that evolves with an approximately constant rate over a wide evolutionary distance covering the whole animal phyla. For this purpose, the evolutionary rate constancy has been examined for 20 proteins. It was found that four proteins, particularly the aldolase C, involved in the glycolitic pathway, had evolved with rates that are approximately constant not only among different classes of vertebrates, but also between vertebrates and arthropods. The evolutionary rate (= 0.26 x 10(-9)/site/year) of the aldolase C is likely to have remained essentially unchanged even between animals and plants.
We have isolated cDNA clones encoding the rat and human forms of a novel protein kinase, termed TESK1 (testis-specific protein kinase 1). Sequence analysis indicates that rat TESK1 contains 628 amino acid residues, composed of an N-terminal protein kinase consensus sequence followed by a C-terminal proline-rich region. Human TESK1 contains 626 amino acids, sharing 92% amino acid identity with its rat counterpart. The protein kinase domain of TESK1 is structurally similar to those of LIMK (LIM motif-containing protein kinase)-1 and LIMK2, with 49-50% sequence identity. Phylogenetic analysis of the protein kinase domains revealed that TESK1 is most closely related to a LIMK subfamily. Chromosomal localization of human TESK1 gene was assigned to 9p13. Anti-TESK1 antibody raised against the C-terminal peptide of TESK1 recognized two polypeptides of 68 and 80 kDa in cell lysates of COS cells transfected with human TESK1 cDNA expression plasmid. TESK1 protein expressed in COS cells exhibited serine/threonine kinase activity, when myelin basic protein was used as a substrate. Northern blot analysis revealed that TESK1 mRNA was specifically expressed in rat and mouse testicular germ cells. The TESK1 mRNA in the testis was detectable only after the 18th day of postnatal development of mice and was mainly expressed in the round spermatids. These observations suggest that TESK1 has a specific function in spermatogenesis.
In the protein kinase family, the basic function of kinase domain is similar among members. According to the standard view of functional constraint, the molecular evolutionary rate depends on functional and structural features characteristic of individual molecules (local constraint). Thus the evolutionary rate of the kinase domain is expected to be similar for different members. Contrary to this expectation, a comparison of the evolutionary rates revealed a wide difference among members; it amounts to about 100 times difference between the maximum and minimum rates. A similar result was also found in members of the immunoglobulin (Ig) family. In addition, significant correlations in evolutionary rate were observed between the kinase domain and the Ig-like domain in the receptor protein tyrosine kinases and between the kinase domain and the SH domain in the nonreceptor-type kinases. Furthermore, the evolutionary rates of family members that are expressed tissue specifically differ widely, depending on their tissue distribution: members expressed in the brain evolve with significantly slower rates than those expressed in the immune system. These results strongly suggest the presence of an alternative constraint (global constraint) against changes on molecules derived from higher levels like tissues or organs.
In this paper, we reviewed our recent works on a possible link between molecular evolution and tissue evolution. The evolutionary rates of genes that are expressed tissue specifically were shown to differ widely to one another, depending on tissues: Brain specific genes evolve with significantly slower rate than immune specific genes. The tissue dependence of molecular evolutionary rate strongly suggests the presence of functional constraints against molecular changes from tissue level. A molecular phylogenetic analysis of tissue specific isoforms that are identical to one another in function, but differ only in tissue distribution revealed frequent gene duplications and rapid accumulations of amino acid substitutions during the early evolution of chordates, where rapid evolution at the tissue or organ levels is thought to have occurred. On the basis of functional constraints, a possible explanation for the correlation between evolution at the two levels was presented.
By low-stringency screening of a human hepatoma HepG2 cell cDNA library, using the genomic fragment of chick c-sea receptor tyrosine kinase as a probe, we isolated overlapping cDNAs encoding a novel protein kinase, which we termed LIM-kinase (LIMK).* The predicted open reading frame encodes a 647-amino-acid polypeptide containing a putative protein kinase structure in the C-terminal half. In addition, LIMK has two repeats of cysteine-rich LIM/double zinc finger motif at the most N-terminus. To our knowledge, this is the first protein kinase seen to contain the LIM motif(s) in the molecule. Although the protein kinase domain of LIMK has highly conserved sequence elements of protein kinases, phylogenetic analysis revealed that LIMK cannot be classified into any subfamily of known protein kinases. Northern blot analysis revealed that the single species of LIMK mRNA of 3.3 kb was expressed in various human epithelial and hematopoietic cell lines. In rat tissues, LIMK mRNA was expressed in the brain, at the highest level. LIM is suggested to be involved in protein-protein interactions by binding to another LIM motif. As the LIM domain is frequently present in the homeodomain-containing transcriptional regulators and oncogenic nuclear proteins, LIMK may be involved in developmental or oncogenic processes through interactions with these LIM-containing proteins.