The genome of tomato (Solanum lycopersicum) is being sequenced by an international consortium of 10 countries (Korea, China, the United Kingdom, India, The Netherlands, France, Japan, Spain, Italy and the United States) as part of a larger initiative called the 'International Solanaceae Genome Project (SOL): Systems Approach to Diversity and Adaptation'. The goal of this grassroots initiative, launched in November 2003, is to establish a network of information, resources and scientists to ultimately tackle two of the most significant questions in plant biology and agriculture: (1) How can a common set of genes/proteins give rise to a wide range of morphologically and ecologically distinct organisms that occupy our planet? (2) How can a deeper understanding of the genetic basis of plant diversity be harnessed to better meet the needs of society in an environmentally friendly and sustainable manner? The Solanaceae and closely related species such as coffee, which are included in the scope of the SOL project, are ideally suited to address both of these questions. The first step of the SOL project is to use an ordered BAC approach to generate a high quality sequence for the euchromatic portions of the tomato as a reference for the Solanaceae. Due to the high level of macro and micro-synteny in the Solanaceae the BAC-by-BAC tomato sequence will form the framework for shotgun sequencing of other species. The starting point for sequencing the genome is BACs anchored to the genetic map by overgo hybridization and AFLP technology. The overgos are derived from approximately 1500 markers from the tomato high density F2-2000 genetic map (http://sgn.cornell.edu/). These seed BACs will be used as anchors from which to radiate the tiling path using BAC end sequence data. Annotation will be performed according to SOL project guidelines. All the information generated under the SOL umbrella will be made available in a comprehensive website. The information will be interlinked with the ultimate goal that the comparative biology of the Solanaceae-and beyond-achieves a context that will facilitate a systems biology approach.
Arabidopsis thaliana has a relatively small genome of approximately 130 Mb containing about 10% repetitive DNA. Genome sequencing studies reveal a gene-rich genome, predicted to contain approximately 25 000 genes spaced on average every 4.5 kb. Between 10 to 20% of the predicted genes occur as clusters of related genes, indicating that local sequence duplication and subsequent divergence generates a significant proportion of gene families. In addition to gene families, repetitive sequences comprise individual and small clusters of two to three retroelements and other classes of smaller repeats. The clustering of highly repetitive elements is a striking feature of the A. thaliana genome emerging from sequence and other analyses.
As part of the European Scientists Sequencing Arabidopsis program, a contiguous region (396 607 bp) located on chromosome 4 around the APETALA2 gene was sequenced. Analysis of the sequence and comparison to public databases predicts 103 genes in this area, which represents a gene density of one gene per 3.85 kb. Almost half of the genes show no significant homology to known database entries. In addition, the first 45 kb of the contig, which covers 11 genes, is similar to a region on chromosome 2, as far as coding sequences are concerned. This observation indicates that ancient duplications of large pieces of DNA have occurred in Arabidopsis.
The higher plant Arabidopsis thaliana (Arabidopsis) is an important model for identifying plant genes and determining their function. To assist biological investigations and to define chromosome structure, a coordinated effort to sequence the Arabidopsis genome was initiated in late 1996. Here we report one of the first milestones of this project, the sequence of chromosome 4. Analysis of 17.38 megabases of unique sequence, representing about 17% of the genome, reveals 3,744 protein coding genes, 81 transfer RNAs and numerous repeat elements. Heterochromatic regions surrounding the putative centromere, which has not yet been completely sequenced, are characterized by an increased frequency of a variety of repeats, new repeats, reduced recombination, lowered gene density and lowered gene expression. Roughly 60% of the predicted protein-coding genes have been functionally characterized on the basis of their homology to known genes. Many genes encode predicted proteins that are homologous to human and Caenorhabditis elegans proteins.
The NIM1 (for noninducible immunity) gene product is involved in the signal transduction cascade leading to both systemic acquired resistance (SAR) and gene-for-gene disease resistance in Arabidopsis. We have isolated and characterized five new alleles of nim1 that show a range of phenotypes from weakly impaired in chemically induced pathogenesis-related protein-1 gene expression and fungal resistance to very strongly blocked. We have isolated the NIM1 gene by using a map-based cloning procedure. Interestingly, the NIM1 protein shows sequence homology to the mammalian signal transduction factor I kappa B subclass alpha. NF-kappa B/I kappa B signaling pathways are implicated in disease resistance responses in a range of organisms from Drosophila to mammals, suggesting that the SAR signaling pathway in plants is representative of an ancient and ubiquitous defense mechanism in higher organisms.