We investigated the potential of Agrobacterium tumefaciens-mediated transformation for producing a genome-wide library of TDNA rice plants individually characterized by their flanking sequence tag (FST). A japonica rice callus transformation procedure—which relies both on high frequency (75% and 96% in Taipei 309 and Nipponbare, respectively) of cocultured calliyielding resistant cell lines and independent generation of multiple (2 to 30) resistant cell lines per cocultured callus—was optimized. Potential efficiencies were as high as 5 and 7 independent transformants per cocultured callus in cultivars Taipei 309 and Nipponbare. We further studied the integration of the T-DNA in more than 200 transgenic plants. Around 35% of the T 0 plants were found to harbor one copy of the T-DNA in both populations. In Taipei 309, 90% of these plants did not have integrated plasmid backbone sequences and 95% segregated the hygromycin resistance trait according to a 3:1 ratio in their T 1 progenies. In multiple-copy plants, 31% of the plants had integrated plasmid backbone sequences and 52% and 10% segregated according to a 3:1 and 15:1 ratio, respectively. Using an efficient polymerase chain reaction walking method, 82 (58%) DNA regions adjacent to the left border of T-DNA inserts (FSTs) have been isolated and sequenced. Results of a homology search indicated that eight known genes had been tagged with the T-DNA.
Single nucleotide polymorphisms (SNP) are the most abundant type of DNA polymorphism found in animal and plant genomes. They provide an important new source of molecular markers that are useful in genetic mapping, map-based positional cloning, quantitative trait locus mapping and the assessment of genetic distances between individuals. Very little is known on the frequency of SNPs in cassava. We have exploited the recently-developed collection of cassava expressed sequence tags (ESTs) to detect SNPs in the five cultivars of cassava used to generate the sequences. The frequency of intra-cultivar and inter-cultivar SNPs after analysis of 111 contigs was one polymorphism per 905 and one per 1,032 bp, respectively; totaling 1 each 509 bp. We have obtained further information on the frequency of SNPs in six cassava cultivars by analysis of 33 amplicons obtained from 3' EST and BAC end sequences. Overall, about 11 kb of DNA sequence was obtained for each cultivar. A total of 186 SNPs (136 and 50 from ESTs and BAC ends, respectively) were identified. Among these, 146 were intra-cultivar polymorphisms, while 80 were inter-cultivar polymorphisms. Thus the total frequency of SNPs was one per 62 bp. This information will help to develop new strategies for EST mapping as well as their association with phenotypic characteristics.
The Cdc25 protein phosphatase is a key enzyme involved in the regulation of the G(2)/M transition in metazoans and yeast. However, no Cdc25 ortholog has so far been identified in plants, although functional studies have shown that an activating dephosphorylation of the CDK-cyclin complex regulates the G(2)/M transition. In this paper, the first green lineage Cdc25 ortholog is described in the unicellular alga Ostreococcus tauri. It encodes a protein which is able to rescue the yeast S. pombe cdc25-22 conditional mutant. Furthermore, microinjection of GST-tagged O. tauri Cdc25 specifically activates prophase-arrested starfish oocytes. In vitro histone H1 kinase assays and anti-phosphotyrosine Western Blotting confirmed the in vivo activating dephosphorylation of starfish CDK1-cyclinB by recombinant O. tauri Cdc25. We propose that there has been coevolution of the regulatory proteins involved in the control of M-phase entry in the metazoan, yeast and green lineages.
Several cDNA libraries were constructed using mRNA isolated from roots, panicles, cell suspensions and leaves of non-stressed Oryza sativa indica (IR64) and japonica (Azucena) plants, from wounded leaves, and from leaves of both cultivars inoculated with Rice Yellow Mottle Virus (RYMV). A total of 5549 cleaned expressed sequence tags (ESTs) were generated from these libraries. They were classified into functional categories on the basis of homology, and analyzed for redundancy within each library. The expression profiles represented by each library revealed great differences between indica and japonica backgrounds. EST frequencies during the early stages of RYMV infection indicated that changes in the expression of genes involved in energy metabolism and photosynthesis are differentially accentuated in susceptible and partially resistant cultivars. Mapping of these ESTs revealed that several co-localize with previously described resistance gene analogs and QTLs (quantitative trait loci).
Plant disease resistance genes (R genes) show significant similarity amongst themselves in terms of both their DNA sequences and structural motifs present in their protein products. Oligonucleotide primers designed from NBS (Nucleotide Binding Site) domains encoded by several R-genes have been used to amplify NBS sequences from the genomic DNA of various plant species, which have been called Resistance Gene Analogues (RGAs) or Resistance Gene Candidates (RGCs). Using specific primers from the NBS and TIR (Toll/Interleukin-1 Receptor) regions, we identified twelve classes of RGCs in cassava (Manihot esculenta Crantz). Two classes were obtained from the PCR-amplification of the TIR domain. The other 10 classes correspond to the NBS sequences and were grouped into two subfamilies. Classes RCa1 to RCa5 are part of the first subfamily and were linked to a TIR domain in the N terminus. Classes RCa6 to RCa10 corresponded to non-TIR NBS-LRR encoding sequences. BAC library screening with the 12 RGC classes as probes allowed the identification of 42 BAC clones that were assembled into 10 contigs and 19 singletons. Members of the two TIR and non-TIR NBS-LRR subfamilies occurred together within individual BAC clones. The BAC screening and Southern hybridization analyses showed that all RGCs were single copy sequences except RCa6 that represented a large and diverse gene family. One BAC contained five NBS sequences and sequence analysis allowed the identification of two complete RGCs encoding two highly similar proteins. This BAC was located on linkage group J with three other RGC-containing BACs. At least one of these genes, RGC2, is expressed constitutively in cassava tissues.
Expressed sequence tags (ESTs) from the Arabidopsis thaliana sequencing project were used to construct a genetic RFLP map for Brassica oleracea. Of the 110 A. thaliana ESTs tested, 95 were found to be informative RFLP probes in map construction. In total, 212 new loci corresponding to the 95 ESTs were added to the existing genetic map of B. oleracea. The enriched map covers all nine basic linkage groups and confirms that the chromosomes of B. oleracea and A. thaliana are similar in linear organization. However, varying levels of sequence conservation between the chromosomes of B. oleracea and A. thaliana were detected in different regions of the genomes. Long conserved regions encompassing entire chromosome arms in both genomes were identified; these are probably shared by descent. On the other hand, extensive rearrangements were observed in numerous chromosome regions, producing a mosaic of A. thaliana -like segments in the genome of Brassica. The presence of extensive chromosome duplication in A. thaliana was taken into consideration in the construction of the comparative maps of B. oleracea and A. thaliana.
We investigated the potential of an improved Agrobacterium tumefaciens-mediated transformation procedure of japonica rice (Oryza sativa L.) for generating large numbers of T-DNA plants that are required for functional analysis of this model genome. Using a T-DNA construct bearing the hygromycin resistance (hpt), green fluorescent protein (gfp) and β-glucuronidase (gusA) genes, each individually driven by a CaMV 35S promoter, we established a highly efficient seed-embryo callus transformation procedure that results both in a high frequency (75–95%) of co-cultured calli yielding resistant cell lines and the generation of multiple (10 to more than 20) resistant cell lines per co-cultured callus. Efficiencies ranged from four to ten independent transformants per co-cultivated callus in various japonica cultivars. We further analysed the T-DNA integration patterns within a population of more than 200 transgenic plants. In the three cultivars studied, 30–40% of the T0 plants were found to have integrated a single T-DNA copy. Analyses of segregation for hygromycin resistance in T1 progenies showed that 30–50% of the lines harbouring multiple T-DNA insertions exhibited hpt gene silencing, whereas only 10% of lines harbouring a single T-DNA insertion was prone to silencing. Most of the lines silenced for hpt also exhibited apparent silencing of the gus and gfp genes borne by the T-DNA. The genomic regions flanking the left border of T-DNA insertion points were recovered in 477 plants and sequenced. Adapter-ligation Polymerase chain reaction analysis proved to be an efficient and reliable method to identify these sequences. By homology search, 77 T-DNA insertion sites were localized on BAC/PAC rice Nipponbare sequences. The influence of the organization of T-DNA integration on subsequent identification of T-DNA insertion sites and gene expression detection systems is discussed.
Eukaryotic ribosomes are made of two components, four ribosomal RNAs, and approximately 80 ribosomal proteins (r-proteins). The exact number of r-proteins and r-protein genes in higher plants is not known. The strong conservation in eukaryotic r-protein primary sequence allowed us to use the well-characterized rat (Rattus norvegicus) r-protein set to identify orthologues on the five haploid chromosomes of Arabidopsis. By use of the numerous expressed sequence tag (EST) accessions and the complete genomic sequence of this species, we identified 249 genes (including some pseudogenes) corresponding to 80 (32 small subunit and 48 large subunit) cytoplasmic r-protein types. None of the r-protein genes are single copy and most are encoded by three or four expressed genes, indicative of the internal duplication of the Arabidopsis genome. The r-proteins are distributed throughout the genome. Inspection of genes in the vicinity of r-protein gene family members confirms extensive duplications of large chromosome fragments and sheds light on the evolutionary history of the Arabidopsis genome. Examination of large duplicated regions indicated that a significant fraction of the r-protein genes have been either lost from one of the duplicated fragments or inserted after the initial duplication event. Only 52 r-protein genes lack a matching EST accession, and 19 of these contain incomplete open reading frames, confirming that most genes are expressed. Assessment of cognate EST numbers suggests that r-protein gene family members are differentially expressed.
P67, a new protein binding to a specific RNA probe, was purified from radish seedlings [Echeverria, M. and Lahmy, S. (1995) Nucleic Acids Res. 23, 4963–4970]. Amino acid sequence information obtained from P67 microsequencing allowed the isolation of genes encoding P67 in radish and Arabidopsis thaliana. Immunolocalisation experiments in transfected protoplasts demonstrated that this protein is addressed to the chloroplast. The RNA‐binding activity of recombinant P67 was found to be similar to that of the native protein. A significant similarity with the maize protein CRP1 [Fisk, D.G., Walker, M.B. and Barkan, A. (1999) EMBO J. 18, 2621–2630] suggests that P67 belongs to the PPR family and could be involved in chloroplast RNA processing.
Arabidopsis thaliana is an important model system for plant biologists. In 1996 an international collaboration (the Arabidopsis Genome Initiative) was formed to sequence the whole genome of Arabidopsis and in 1999 the sequence of the first two chromosomes was reported. The sequence of the last three chromosomes and an analysis of the whole genome are reported in this issue. Here we present the sequence of chromosome 3, organized into four sequence segments (contigs). The two largest (13.5 and 9.2 Mb) correspond to the top (long) and the bottom (short) arms of chromosome 3, and the two small contigs are located in the genetically defined centromere. This chromosome encodes 5,220 of the roughly 25,500 predicted protein-coding genes in the genome. About 20% of the predicted proteins have significant homology to proteins in eukaryotic genomes for which the complete sequence is available, pointing to important conserved cellular functions among eukaryotes.
Almost all the nuclear genes of four Gramineae (maize, wheat, barley, rice) and pea are located in DNA fractions covering only a 1–2% GC range and representing between 10 and 25% of the different genomes. These DNA fractions comprise large gene‐rich regions (collectively called the ‘gene space’) separated by vast gene‐empty, repeated sequences. In contrast, in Arabidopsis thaliana, genes are distributed in DNA fractions covering an 8% GC range and representing 85% of the genome. Here, we investigated the integration of a transferred DNA (T‐DNA) in the genomes of Arabidopsis and rice and found different patterns of integration, which are correlated with the different gene distributions. While T‐DNA integrates essentially everywhere in the Arabidopsis genome, integration was detected only in the gene space, namely in the gene‐rich, transcriptionally active, regions of the rice genome. The implications of these results for the integration of foreign DNA are discussed.
In the past several years our work has concentrated on the construction of a genetic map of Brassica oleracea using known DNA probes from the genome of the model plant, Arabidopsis thaliana. Gene probes from the entire A. thaliana genome were used but with emphasis on those from chromosomes 4 and 5. Mapping was done on a set of 67 Fl lines resulting from a cross of collard (B, oleracea var, acephala) x cauliflower (B, oleracea var, botrytis). One hundred and twenty cDNA clones of A. thaliana were used as probes of which 115 clones had high sequence homology to known genes. Linkages and chromosomal order of loci were evaluated using MapMaker 3.0. As a result of this effort, 229 new loci were added to the existing genetic map of Brassica oleracea. The map covers nine basic linkage groups. Many homoeologous chromosome segments among B. oleracea and A. thaliana were detected. This confirms previous reports that many chromosome segments in B. oleracea and A. thaliana display similar linear organization. Our work also shows extensive rearrangements of numerous chromosome regions leading to a mosaic of A. thaliana-like segments in the genomes of Brassica.
Arabidopsis thaliana L. leafy cotyledon1 (lec1) and fusca3 (fus3) mutants show multiple phenotypic defects during seed development. In this report the effects of these mutations are examined at the molecular level. The patterns of protein accumulation in lec1 and fus3 seeds are severly altered. In lec1 seeds the steady-state mRNA levels of several late embryogenesis genes were reduced. Different patterns of expression were observed, indicating the occurrence of several regulatory pathways. The effect of lec1 mutations on the expression of the late-embryogenesis abundant AtEm1 gene was examined in detail. In lec1-1 seeds, the AtEm1 gene was expressed at a higher level than in the wild type and earlier in development. The activity of an AtEm1 promoter/beta-glucuronidase reporter gene construct in transgenic A. thaliana plants was studied. Changes in promoter activity in lec1-1 with respect to wild-type seeds were correlated with changes in corresponding mRNA steady-state levels. fus3-2 mutation produced similar changes in AtEm1 promoter activity as lec1-1, which is consistent with the hypothesis that LEC1 and FUS3 might act in the same regulatory pathway. Transgenic analysis using 5'-promoter deletions demonstrated that at least two regions of AtEm1 gene promoter interact with the LEC1-dependent transcriptional regulatory pathway. In spite of expression of the AtEm1 promoter and accumulation of AtEm1 mRNA, the corresponding Em1 protein does not accumulate in lec1-1 seeds. The ABA inducibility of the AtEm1 promoter was not affected by the lec1 mutation.
Arabidopsis thaliana has a relatively small genome of approximately 130 Mb containing about 10% repetitive DNA. Genome sequencing studies reveal a gene-rich genome, predicted to contain approximately 25 000 genes spaced on average every 4.5 kb. Between 10 to 20% of the predicted genes occur as clusters of related genes, indicating that local sequence duplication and subsequent divergence generates a significant proportion of gene families. In addition to gene families, repetitive sequences comprise individual and small clusters of two to three retroelements and other classes of smaller repeats. The clustering of highly repetitive elements is a striking feature of the A. thaliana genome emerging from sequence and other analyses.
Seedless fruits are a desirable commodity for consumers, and have been produced using traditional farming and breeding methods for many centuries. Evidence that seedless forms of Vitis vinifera grapes have been prized for many centuries as dried fruit is provided by Greek philosophers such as Hippocrate, Platon and in the writings of ancient Egypt of 3000 bc. However, the use of current agricultural practices to achieve seedlessness has in-built disadvantages. Here we discuss novel approaches that have emerged over the past few years that open up new possibilities for breeding seedless plants. These include quantitative trait loci, manipulating genes that promote parthenocarpy and interfering with seed development using ‘terminator’ technology. Several patents based on recombinant DNA techniques illustrate the current industrial interest in this field, and we discuss the likely positive and negative impacts of these novel strategies on food production.
The higher plant Arabidopsis thaliana (Arabidopsis) is an important model for identifying plant genes and determining their function. To assist biological investigations and to define chromosome structure, a coordinated effort to sequence the Arabidopsis genome was initiated in late 1996. Here we report one of the first milestones of this project, the sequence of chromosome 4. Analysis of 17.38 megabases of unique sequence, representing about 17% of the genome, reveals 3,744 protein coding genes, 81 transfer RNAs and numerous repeat elements. Heterochromatic regions surrounding the putative centromere, which has not yet been completely sequenced, are characterized by an increased frequency of a variety of repeats, new repeats, reduced recombination, lowered gene density and lowered gene expression. Roughly 60% of the predicted protein-coding genes have been functionally characterized on the basis of their homology to known genes. Many genes encode predicted proteins that are homologous to human and Caenorhabditis elegans proteins.
Arabidopsis AtEm1 and AtEm6 proteins were overexpressed in E. coli and used to raise antibodies which were used to analyse Em expression at the protein level. One of the sera is specific for AtEm1 protein whereas the second reacts with both AtEm1 and AtEm6 proteins. These antibodies were used to analyse expression of Em genes at the protein level and to complete previous studies at the mRNA level. During seed maturation, AtEm1 protein accumulates earlier than AtEm6, in parallel with the corresponding mRNA but with a 3 d delay. During germination, AtEm1 protein undergoes two successive cleavages before being degraded. Both proteins are much more stable than the corresponding mRNA. Soaking of dormant seeds indicates that imbibition is sufficient to induce Em protein degradation and that germination per se is not required. AtEm1 and AtEm6 mRNA can be precociously induced by ABA in immature siliques, but protein accumulation could not be observed. A similar observation was made with leaves of transgenic plants ectopically expressing ABI3. These results establish clearly that Em protein accumulation is also tightly controlled by post-transcriptional mechanisms in addition to transcriptional ones.