We report on the comparative analysis of the sequences of four and five copies of the Ck1 (caseine kinase-like) gene in Arabidopsis thaliana and Brassica oleracea, respectively. These duplicate genes could be separated into two major groups based on their number of exons. The first group, including most of the genes, contained 14 exons, whereas the second one had 13 exons. Based on this separation as well as on the DNA and amino acid sequence identity and phylogenetic analysis, it was possible to assign orthology and paralogy to these genes and to draw conclusions on their possible evolutionary relationships.
In Saccharomyces cerevisiae, the SAC1 gene encodes a polyphosphoinositide phosphatase (PPIPase) that modulates the levels of phosphoinositides, which are key regulators of a number of signal transduction processes. SAC1p has been implicated in multiple cellular functions: actin cytoskeleton organization, secretory functions, inositol metabolism, ATP transport, and multiple-drug sensitivity. Here, we describe the characterization of three genes in Arabidopsis thaliana, AtSAC1a, AtSAC1b, and AtSAC1c, encoding proteins similar to those of yeast SAC1p. We demonstrated that the three AtSAC1 proteins are functional homologs of the yeast SAC1p because they can rescue the cold-sensitive and inositol auxotroph yeast sac1-null mutant strain. The fact that Arabidopsis and yeast SAC1 genes derived from a common ancestor suggests that this plant multigenic family is involved in the phosphoinositide pathway and in a range of cellular functions similar to those in yeast. Using GFP fusion experiments, we demonstrate that the three AtSAC1 proteins are targeted to the endoplasmic reticulum. Their expression patterns are overlapping, with at least two members expressed in each organ. Remarkably, AtSAC1 genes are not expressed during seed development, and therefore additional phosphatases are required to control phosphoinositide levels in seeds.
The region corresponding to the ABI1-Rps2-Ck1 segment on chromosome 4 of Arabidopsis thaliana was sequenced in Brassica oleracea. Similar to A. thaliana, the B. oleracea homolog BoRps2 is present in single copy. The B. oleracea orthologous segment was located on chromosome 4 and can be distinguished by the presence of an N-myristoyl transferase coding gene (N-myr) between the Rps2 and Ck1 (BoCk1a) genes. The N-myr homologs in Arabidopsis are on chromosomes 2 and 5. Additional homologs for Ck1 are located on these two chromosomes. A second Ck1 homolog found on B. oleracea (BoCk1b) chromosome 7 served to define another orthologous segment located in Arabidopsis chromosome 1. The two segments displayed identical gene content and order in both species, namely BoCK1b, a gene encoding a hypothetical protein (BohypothA) and transcription factor eiF4A. High levels of sequence identity were observed for the coding sequences of all genes examined. Although in general larger spacers were found in Brassica than in A. thaliana, this was not always the case. Promoters were poorly conserved, except for several sequence stretches of a few nucleotides. Comparative sequencing revealed microsyntenic changes resulting from chromosomal structural rearrangements, which are often undetectable by genetic mapping.
The region corresponding to the ABI1-Rps2-Ck1 segment on chromosome 4 of Arabidopsis thaliana was sequenced in Brassica oleracea. Similar to A. thaliana, the B. oleracea homolog BoRps2 is present in single copy. The B. oleracea orthologous segment was located on chromosome 4 and can be distinguished by the presence of an N-myristoyl transferase coding gene (N-myr) between the Rps2 and Ck1 (BoCk1a) genes. The N-myr homologs in Arabidopsis are on chromosomes 2 and 5. Additional homologs for Ck1 are located on these two chromosomes. A second Ck1 homolog found on B. oleracea (BoCk1b) chromosome 7 served to define another orthologous segment located in Arabidopsis chromosome 1. The two segments displayed identical gene content and order in both species, namely BoCK1b, a gene encoding a hypothetical protein (BohypothA) and transcription factor eiF4A. High levels of sequence identity were observed for the coding sequences of all genes examined. Although in general larger spacers were found in Brassica than in A. thaliana, this was not always the case. Promoters were poorly conserved, except for several sequence stretches of a few nucleotides. Comparative sequencing revealed microsyntenic changes resulting from chromosomal structural rearrangements, which are often undetectable by genetic mapping.
This paper reviews part of our studies on gene regulation during the Arabidopsis seed maturation phase. Essentially, three complementary strategies have been used. The first one consisted in identifying genes expressed during this period by random sequencing of EST from a dry seed cDNA library and by comparing their frequency with that in an immature cDNA library. The second strategy focused on the detailed analysis of the expression of a specific group of genes coding for the class I LEA proteins, the Em genes, and analysis of their promoter. Finally we evaluated the expression of a number of LEA gene in various regulatory mutants, including abi3, lec1 and abi5.Altogether, our results illustrate the complexity of expression patterns and the interaction of various factors to define several distinct regulatory pathways.
Several cDNA clones encoding three different lipid transfer proteins (LTPs) have been isolated from rice (Oryza sativa L.) in order to analyse the complexity, the evolution and the expression of the LTP gene family. The mature proteins deduced from three clones exhibited a molecular mass of 9 kDa, in agreement with the molecular mass of other LTPs from plants. The clones were shown to be homologous in the coding region, while the 3′ non-coding regions diverged strongly between the clones. The occurrence of at least three small multigene families encoding these proteins in rice was confirmed by Southern blot analysis. When compared with each other and with LTPs from other plants, the cluster including rice LTPs and other cereal LTPs indicated that these genes duplicated rather recently and independently in the different plant phyla. The expression pattern of each gene family was also investigated. Northern blot experiments demonstrated that they are differentially regulated in the different tissues analysed. Components such as salt, salicylic acid and abscisic acid were shown to modulate Ltp gene expression, depending on tissues and gene classes, suggesting a complex regulation of these genes.
The genome of the model plant Arabidopsis thaliana is being analyzed in more and more detail. This paper reviews recent progress over the last 5 years. A first goal was to establish a catalogue of expressed genes using the EST (expressed sequence tag) strategy. Two consortia (French and American) have together released close to 30 000 EST representing approximately 10 000 genes. Such a catalogue has already facilitated a number of biological analyses. The next step, which is sequencing the whole genome, has already started with a European Union pilot project, which has demonstrated the feasibility of the large scale sequencing of this genome. During the last 3 years 2.5 Mbp have been determined and data acquisition is accelerating tremendously Two major questions remain for the future. What is the function of the genes with no known homology? How can this enormous information resource be used for the benefit of other plants? A few current ideas and perspectives are discussed.
© 1997 Federation of European Biochemical Societies.
Arabidopsis is a crucifer weed with a small genome of about 120 Mbp which has been chosen as a model species for plant molecular genetics. Four years ago, a consortium of nine French laboratories, including ours, initiated a project aimed at mapping the transcribed regions of the genome. The strategy employed was to systematically and randomly sequence cDNA clones isolated from libraries made from different tissues and organs of plants grown under various physiological conditions. The consortium released about 7,000 expressed sequenced tags (ESTs) in the dbEST database corresponding to approximately 3,500 unique genes. In the next phase of the programme, a YAC library with average inserts of 500 kbp has been prepared. We have now started to use the EST information to map the cDNA clones on these YACs. The most recent aspect of Arabidopsis sequencing is the ESSA (European Scientists Sequencing Arabidopsis) project, in which the aim is to describe 2.5 Mbp by the end of 1996. Genomic sequencing has revealed a very high gene density. Comparison of present genomic sequencing results with the EST data suggests that up to half of the genes might already be tagged with an EST. In collaboration with Carlos Quiros' group in Davis we have also analysed the conservation of a 30 kbp locus (Em 1, a late embryogenesis abundant protein gene) on chromosome 3 between Arabidopsis and several Brassica species. Progress on these various aspects will be reviewed. We shall also present some sequence comparisons between Arabidopsis and rice ESTs. These results suggest that it should be possible in the very near future to map a pool of common genes onto many different plant genomes. This should provide a common framework to integrate maps from different species and facilitate mapbased cloning of genes of agronomical importance.
Nearly 7000 Arabidopsis thaliana-expressed sequence tags (ESTs) from 10 cDNA libraries have been sequenced, of which almost 5000 non-redundant tags have been submitted to the EMBL data bank. The quality of the cDNA libraries used is analysed. Similarity searches in international protein data banks have allowed the detection of significant similarities to a wide range of proteins from many organisms. Alignment with ESTs from the rice systematic sequencing project has allowed the detection of amino acid motifs which are conserved between the two organisms, thus identifying tags to genes encoding highly conserved proteins. These genes are candidates for a common framework in genome mapping projects in different plants.
During the course of an Arabidopsis thaliana genome sequencing project, we identified a gene, G4, with a derived amino acid sequence showing homology to the product of the Rhodobacter capsulatus bchG locus which is involved in the esterification of bacterio-chlorophyllide with geranylgeraniol. The relationship between this gene and bchG was confirmed by the isolation and analysis of a corresponding full-length cDNA. Comparison of genomic and cDNA sequences indicated that the gene is made up of 14 exons, some of them being very short. Southern and Northern analyses showed that this sequence represents a single-copy gene and its transcript is detected only in green or greening tissues. Both homologies and expression data suggest that this gene encodes a chlorophyll synthetase, one of the last enzymes of chlorophyll biosynthesis, and thus represents a new example of a nuclear gene encoding an enzyme of this pathway in higher plants.
The major simple sequence repeats present in the Arabidopsis genome were identified by Southern hybridizations with 49 oligonucleotide probes matching all the possible combinations of motifs up to 4 nucleotides long. The method used allowed us to perform all the hybridizations under the same temperature conditions. A good correlation was observed with the data obtained from database analysis, indicating that the method can be useful for identifying the major classes of microsatellite loci in species for which few or no sequence data are available. AG/CT, AAG/CTT, ATG/CAT and GTG/CAC are the major motifs present in the Arabidopsis genome that can be used as convenient probes to isolate microsatellite loci by screening libraries. AAG/CTT is the more frequent of these motifs, and its relative frequency in Arabidopsis is much higher than averagely found in the plant kingdom. About 8% of the cDNA clones from an immature silique library contains AG/CT, AAG/CTT or ATG/CAT microsatellite loci. Several microsatellite loci were isolated by screening genomic and cDNA libraries. Twenty-six tri-nucleotide loci were PCR amplified from four different ecotypes, and polymorphism was observed for 12 of them; 10 loci showing two alleles and 2 loci showing three alleles.
The cloning and sequence analysis of a gene that encodes a lipid transfer protein (LTP) from rice is reported. A genomic DNA library from Oryza sativa was screened using a cDNA encoding a maize LTP. One genomic clone containing the gene (Ltp) was partially sequenced and analyzed. The open reading frame is interrupted by an 89-bp intron. From the results of Southern hybridizations, Ltp appears to be a member of a small multigenic family. Transcripts of the corresponding gene were detected in several tissues including coleoptile, leaf, endosperm, scutellum and root. The transcription start point was determined by primer extension. The deduced amino-acid sequence of the Ltp product is shown to be homologous to LTPs from other crops.
The genetic relationships between two Prunus species, involved in rootstock breeding, were examined at the level of the ribosomal RNA genes. Twenty clones of P. cerasifera, a diploid species, and 12 clones of P. spinosa, a tetraploid wild species, were studied. The use of three heterologous ribosomal DNA probes covering different regions of the ribosomal tandem repeats enabled us to construct restriction maps for EcoRI and BamHI. We identified two unit types (unit I and unit II) in P. cerasifera. In P. spinosa, P. cerasifera units were present in addition to a third ribosomal unit type (unit III). These results appeared to confirm previous cytological studies (Salesses 1973) indicating that one of the genomes in P.spinosa has homology with the one from P. cerasifera.
The intergenic spacer of a rice ribosomal RNA gene repeating unit has been completely sequenced. The spacer contains three imperfect, direct repeated regions of 264-253 bp, followed by a related but more highly divergent region. Detailed analysis of the sequence allows the presentation of an evolutionary scenario in which the 264-253-bp repeats are derived from an ancestral 150-bp sequence by deletion and amplification. Comparison of the rice sequence with those of maize, wheat, and rye shows that, despite considerable divergence from the ancestral sequence, several regions have been highly conserved, suggesting that they may play an important role in the structure and/or expression of the ribosomal genes.
Partial and systematic sequencing of an Arabidopsis thaliana immature silique cDNA library has revealed a clone with high amino acid sequence homology with a C12:0-acyl-carrier protein thioesterase from California bay (Umbellularia californica) seeds. Isolation of this clone is the first evidence that a gene for, this enzyme is expressed in Arabidopsis sp.
Mature seeds contain a significant stock of stored mRNA, the life-span of which is as long as the seed-life (Payne, 1976; Delseny et al., 1977). A basic question in seed biology is the role and function of this stored mRNA. During the last ten years, many plant molecular biologists have addressed this question. As a result, seed development has been extensively studied. Most results concern the easiest genes to deal with, those coding for the storage proteins. However this gives only a partial view of seed development (Dure, 1985; Casey et al., 1986). Trying to answer questions concerning mRNA stored in mature seeds we have been led to analyse gene expression at various developmental stages and to realise that during seed development a number of genes are differentially expressed and sequentially switched on and off.