We have sequenced and annotated the genome of fission yeast (Schizosaccharomyces pombe), which contains the smallest number of protein-coding genes yet recorded for a eukaryote: 4,824. The centromeres are between 35 and 110 kilobases (kb) and contain related repeats including a highly conserved 1.8-kb element. Regions upstream of genes are longer than in budding yeast (Saccharomyces cerevisiae), possibly reflecting more-extended control regions. Some 43% of the genes contain introns, of which there are 4,730. Fifty genes have significant similarity with human disease genes; half of these are cancer related. We identify highly conserved genes important for eukaryotic cell organization including those required for the cytoskeleton, compartmentation, cell-cycle control, proteolysis, protein phosphorylation and RNA splicing. These genes may have originated with the appearance of eukaryotic life. Few similarly conserved genes that are important for multicellular organization were identified, suggesting that the transition from prokaryotes to eukaryotes required more new genes than did the transition from unicellular to multicellular organization.
Nature 415, 871–880 (2002). In this Article, the author Andreas Düsterhöft was mistakenly omitted: his name and affiliation (footnote 6) should have been inserted between M. Fuchs and C. Fritzc in the author list. In addition, the name of L. Cerutti (in the last line of the author list) was misspelled.
Arabidopsis thaliana is an important model system for plant biologists. In 1996 an international collaboration (the Arabidopsis Genome Initiative) was formed to sequence the whole genome of Arabidopsis and in 1999 the sequence of the first two chromosomes was reported. The sequence of the last three chromosomes and an analysis of the whole genome are reported in this issue. Here we present the sequence of chromosome 3, organized into four sequence segments (contigs). The two largest (13.5 and 9.2 Mb) correspond to the top (long) and the bottom (short) arms of chromosome 3, and the two small contigs are located in the genetically defined centromere. This chromosome encodes 5,220 of the roughly 25,500 predicted protein-coding genes in the genome. About 20% of the predicted proteins have significant homology to proteins in eukaryotic genomes for which the complete sequence is available, pointing to important conserved cellular functions among eukaryotes.
Arabidopsis thaliana has a relatively small genome of approximately 130 Mb containing about 10% repetitive DNA. Genome sequencing studies reveal a gene-rich genome, predicted to contain approximately 25 000 genes spaced on average every 4.5 kb. Between 10 to 20% of the predicted genes occur as clusters of related genes, indicating that local sequence duplication and subsequent divergence generates a significant proportion of gene families. In addition to gene families, repetitive sequences comprise individual and small clusters of two to three retroelements and other classes of smaller repeats. The clustering of highly repetitive elements is a striking feature of the A. thaliana genome emerging from sequence and other analyses.
The higher plant Arabidopsis thaliana (Arabidopsis) is an important model for identifying plant genes and determining their function. To assist biological investigations and to define chromosome structure, a coordinated effort to sequence the Arabidopsis genome was initiated in late 1996. Here we report one of the first milestones of this project, the sequence of chromosome 4. Analysis of 17.38 megabases of unique sequence, representing about 17% of the genome, reveals 3,744 protein coding genes, 81 transfer RNAs and numerous repeat elements. Heterochromatic regions surrounding the putative centromere, which has not yet been completely sequenced, are characterized by an increased frequency of a variety of repeats, new repeats, reduced recombination, lowered gene density and lowered gene expression. Roughly 60% of the predicted protein-coding genes have been functionally characterized on the basis of their homology to known genes. Many genes encode predicted proteins that are homologous to human and Caenorhabditis elegans proteins.
The nucleotide sequences of five major regions from chromosome VII of Saccharomyces cerevisiae have been determined and analysed. These regions represent 203 kilobases corresponding to approximately one-fifth of the complete yeast chromosome VII. Two fragments originate from the left arm of this chromosome. The first one of about 15.8 kb starts approximately 75 kb from the left telomere and is bordered by the SK18 chromosomal marker. The second fragment covers the 72.6 kb region between the chromosomal markers CYH2 and ALG2. On the right chromosomal arm three regions, a 70.6 kb region between the MSB2 and the KSS1 chromosomal markers and two smaller regions dominated by the KRE11 marker and another one in the vicinity of the SER2 marker were sequenced. We found a total of 114 open reading frames (ORFs), 13 of which were completely overlapping with larger ORFs running in the opposite direction. A total of 44 yeast genes, the physiological functions of which are known, could be precisely mapped on this chromosome. Of the remaining 57 ORFs, 26 shared sequence homologies with known genes, among which were 13 other S. cerevisiae genes and five genes from other organisms. No homology with any sequence in the databases could be found for 31 ORFs. Furthermore, five Ty elements were found, one of which may not be functional due to a frame shift in its Ty1B amino acid sequence. The five chromosomal regions harboured five potential ARS elements and one sigma element together with eight tRNA genes and two snRNAs, one of which is encoded by an intron of a protein-coding gene.
The removal of the mRNA poly(A) tail in the yeast Saccharomyces cerevisiae is stimulated by the poly(A)-binding protein (Pab1p). A large scale purification of the Pab1p-stimulated poly(A) ribonuclease (PAN) identifies a 76-kDa and two 135-Da polypeptides as candidate enzyme subunits. Antibodies against the Pan1p protein, which is the minor 135-kDa protein in the preparation, can immunodeplete Pan1p but not PAN activity. The protein sequence of the major 135-kDa protein, Pan2p, reveals a novel protein that was also found in the previously reported PAN purification (Sachs, A. B., and Deardorff, J. A.(1992) Cell 70, 961-973). Deletion of the non-essential PAN2 gene results in an increase of the average length of mRNA poly(A) tails in vivo, and a loss of Pab1p-stimulated PAN activity in crude extracts. These data confirm that Pan2p and not Pan1p is required for PAN activity, and they suggest that ribonucleases other than the Pab1p-stimulated PAN are capable of shortening poly(A) tails in vivo.