The nucleotide sequence of the entire genome of a cyanobacterium Gloeobacter violaceus PCC 7421 was determined. The genome of G. violaceus was a single circular chromosome 4,659,019 bp long with an average GC content of 62%. No plasmid was detected. The chromosome comprises 4430 potential protein-encoding genes, one set of rRNA genes, 45 tRNA genes representing 44 tRNA species and genes for tmRNA, B subunit of RNase P, SRP RNA and 6Sa RNA. Forty-one percent of the potential protein-encoding genes showed sequence similarity to genes of known function, 37% to hypothetical genes, and the remaining 22% had no apparent similarity to reported genes. Comparison of the assigned gene components with those of other cyanobacteria has unveiled distinctive features of the G. violaceus genome. Genes for PsaI, PsaJ, PsaK, and PsaX for Photosystem I and PsbY, PsbZ and Psb27 for Photosystem II were missing, and those for PsaF, PsbO, PsbU, and PsbV were poorly conserved. cpcG for a rod core linker peptide for phycobilisomes and nblA related to the degradation of phycobilisomes were also missing. Potential signal peptides of the presumptive products of petJ and petE for soluble electron transfer catalysts were less conserved than the remaining portions. These observations may be related to the fact that photosynthesis in G. violaceus takes place not in thylakoid membranes but in the cytoplasmic membrane. A large number of genes for sigma factors and transcription factors in the LuxR, LysR, PadR, TetR, and MarR families could be identified, while those for major elements for circadian clock, kaiABC were not found. These differences may reflect the phylogenetic distance between G. violaceus and other cyanobacteria.
The entire genome of a thermophilic unicellular cyanobacterium, Thermosynechococcus elongatus BP-1, was sequenced. The genome consisted of a circular chromosome 2,593,857 bp long, and no plasmid was detected. A total of 2475 potential protein-encoding genes, one set of rRNA genes, 42 tRNA genes representing 42 tRNA species and 4 genes for small structural RNAs were assigned to the chromosome by similarity search and computer prediction. The translated products of 56% of the potential protein-encoding genes showed sequence similarity to experimentally identified and predicted proteins of known function, and the products of 34% of these genes showed sequence similarity to the translated products of hypothetical genes. The remaining 10% lacked significant similarity to genes for predicted proteins in the public DNA databases. Sixty-three percent of the T. elongatus genes showed significant sequence similarity to those of both Synechocystis sp. PCC 6803 and Anabaena sp. PCC 7120, while 22% of the genes were unique to this species, indicating a high degree of divergence of the gene information among cyanobacterial strains. The lack of genes for typical fatty acid desaturases and the presence of more genes for heat-shock proteins in comparison with other mesophilic cyanobacteria may be genomic features of thermophilic strains. A remarkable feature of the genome is the presence of 28 copies of group II introns, 8 of which contained a presumptive gene for maturase/reverse transcriptase. A trace of genome rearrangement mediated by the group II introns was also observed.
The nucleotide sequence of the entire genome of a filamentous cyanobacterium, Anabaena sp. strain PCC 7120, was determined. The genome of Anabaena consisted of a single chromosome (6,413,771 bp) and six plasmids, designated pCC7120alpha (408,101 bp), pCC7120beta (186,614 bp), pCC7120gamma (101,965 bp), pCC7120delta (55,414 bp), pCC7120epsilon (40,340 bp), and pCC7120zeta (5,584 bp). The chromosome bears 5368 potential protein-encoding genes, four sets of rRNA genes, 48 tRNA genes representing 42 tRNA species, and 4 genes for small structural RNAs. The predicted products of 45% of the potential protein-encoding genes showed sequence similarity to known and predicted proteins of known function, and 27% to translated products of hypothetical genes. The remaining 28% lacked significant similarity to genes for known and predicted proteins in the public DNA databases. More than 60 genes involved in various processes of heterocyst formation and nitrogen fixation were assigned to the chromosome based on their similarity to the reported genes. One hundred and ninety-five genes coding for components of two-component signal transduction systems, nearly 2.5 times as many as those in Synechocystis sp. PCC 6803, were identified on the chromosome. Only 37% of the Anabaena genes showed significant sequence similarity to those of Synechocystis, indicating a high degree of divergence of the gene information between the two cyanobacterial strains.
The complete nucleotide sequence of the genome of a symbiotic bacterium Mesorhizobium loti strain MAFF303099 was determined. The genome of M. loti consisted of a single chromosome (7,036, 071 bp) and two plasmids, designated as pMLa (351, 911 bp) and pMLb (208, 315 bp). The chromosome comprises 6752 potential protein-coding genes, two sets of rRNA genes and 50 tRNA genes representing 47 tRNA species. Fifty-four percent of the potential protein genes showed sequence similarity to genes of known function, 21% to hypothetical genes, and the remaining 25% had no apparent similarity to reported genes. A 611-kb DNA segment, a highly probable candidate of a symbiotic island, was identified, and 30 genes for nitrogen fixation and 24 genes for nodulation were assigned in this region. Codon usage analysis suggested that the symbiotic island as well as the plasmids originated and were transmitted from other genetic systems. The genomes of two plasmids, pMLa and pMLb, contained 320 and 209 potential protein-coding genes, respectively, for a variety of biological functions. These include genes for the ABC-transporter system, phosphate assimilation, two-component system, DNA replication and conjugation, but only one gene for nodulation was identified.
Arabidopsis thaliana is an important model system for plant biologists. In 1996 an international collaboration (the Arabidopsis Genome Initiative) was formed to sequence the whole genome of Arabidopsis and in 1999 the sequence of the first two chromosomes was reported. The sequence of the last three chromosomes and an analysis of the whole genome are reported in this issue. Here we present the sequence of chromosome 3, organized into four sequence segments (contigs). The two largest (13.5 and 9.2 Mb) correspond to the top (long) and the bottom (short) arms of chromosome 3, and the two small contigs are located in the genetically defined centromere. This chromosome encodes 5,220 of the roughly 25,500 predicted protein-coding genes in the genome. About 20% of the predicted proteins have significant homology to proteins in eukaryotic genomes for which the complete sequence is available, pointing to important conserved cellular functions among eukaryotes.
The genome of the model plant Arabidopsis thaliana has been sequenced by an international collaboration, The Arabidopsis Genome Initiative. Here we report the complete sequence of chromosome 5. This chromosome is 26 megabases long; it is the second largest Arabidopsis chromosome and represents 21% of the sequenced regions of the genome. The sequence of chromosomes 2 and 4 have been reported previously and that of chromosomes 1 and 3, together with an analysis of the complete genome sequence, are reported in this issue. Analysis of the sequence of chromosome 5 yields further insights into centromere structure and the sequence determinants of heterochromatin condensation. The 5,874 genes encoded on chromosome 5 reveal several new functions in plants, and the patterns of gene organization provide insights into the mechanisms and extent of genome evolution in plants.
A fine physical map of Arabidopsis thaliana chromosome 5 was constructed by ordering the clones from YAC, P1, TAC and BAC libraries of the genome using the sequences of a variety of genetic and EST markers and terminal sequences of clones. The markers used were 88 genetic markers, 13 EST markers, 87 YAC end probes, 100 YAC subclone end probes, and 390 end probes of P1, TAC and BAC clones. The entire genome of chromosome 5, except for the centromeric and telomeric regions, was covered by two large contigs 11.6 Mb and 14.2 Mb long separated by the centromeric region. The minimum tiling path of the chromosome was constituted by a total of 430 P1, TAC and BAC clones. The map information is available at the Web site http://www.kazusa.or.jp/arabi/.
The sequence determination of the entire genome of the Synechocystis sp. strain PCC6803 was completed. The total length of the genome finally confirmed was 3,573,470 bp, including the previously reported sequence of 1,003,450 bp from map position 64% to 92% of the genome. The entire sequence was assembled from the sequences of the physical map-based contigs of cosmid clones and of lambda clones and long PCR products which were used for gap-filling. The accuracy of the sequence was guaranteed by analysis of both strands of DNA through the entire genome. The authenticity of the assembled sequence was supported by restriction analysis of long PCR products, which were directly amplified from the genomic DNA using the assembled sequence data. To predict the potential protein-coding regions, analysis of open reading frames (ORFs), analysis by the GeneMark program and similarity search to databases were performed. As a result, a total of 3,168 potential protein genes were assigned on the genome, in which 145 (4.6%) were identical to reported genes and 1,257 (39.6%) and 340 (10.8%) showed similarity to reported and hypothetical genes, respectively. The remaining 1,426 (45.0%) had no apparent similarity to any genes in databases. Among the potential protein genes assigned, 128 were related to the genes participating in photosynthetic reactions. The sum of the sequences coding for potential protein genes occupies 87% of the genome length. By adding rRNA and tRNA genes, therefore, the genome has a very compact arrangement of protein- and RNA-coding regions. A notable feature on the gene organization of the genome was that 99 ORFs, which showed similarity to transposase genes and could be classified into 6 groups, were found spread all over the genome, and at least 26 of them appeared to remain intact. The result implies that rearrangement of the genome occurred frequently during and after establishment of this species.