AbstractThe sections in this article areIntroductionGenome SequencingStrategies for Understanding Gene FunctionProspectsAcknowledgements
Pseudomonas fluorescens F113 is a plant growth-promoting rhizobacterium (PGPR) that has biocontrol activity against fungal plant pathogens and is a model for rhizosphere colonization. Here, we present its complete genome sequence, which shows that besides a core genome very similar to those of other strains sequenced within this species, F113 possesses a wide array of genes encoding specialized functions for thriving in the rhizosphere and interacting with eukaryotic organisms.
We studied the physical and genetic organization of chromosome 6 of tomato (Solanum lycopersicum) cv. Heinz 1706 by combining bacterial artificial chromosome (BAC) sequence analysis, high-information-content fingerprinting, genetic analysis, and BAC-fluorescent in situ hybridization (FISH) mapping data. The chromosome positions of 81 anchored seed and extension BACs corresponded in most cases with the linear marker order on the high-density EXPEN 2000 linkage map. We assembled 25 BAC contigs and eight singleton BACs spanning 2.0 Mb of the short-arm euchromatin, 1.8 Mb of the pericentromeric heterochromatin and 6.9 Mb of the long-arm euchromatin. Sequence data were combined with their corresponding genetic and pachytene chromosome positions into an integrated map that covers approximately a third of the chromosome 6 euchromatin and a small part of the pericentromeric heterochromatin. We then compared physical length (Mb), genetic (cM) and chromosome distances (microm) for determining gap sizes between contigs, revealing relative hot and cold spots of recombination. Through sequence annotation we identified several clusters of functionally related genes and an uneven distribution of both gene and repeat sequences between heterochromatin and euchromatin domains. Although a greater number of the non-transposon genes were located in the euchromatin, the highly repetitive (22.4%) pericentromeric heterochromatin displayed an unexpectedly high gene content of one gene per 36.7 kb. Surprisingly, the short-arm euchromatin was relatively rich in repeats as well, with a repeat content of 13.4%, yet the ratio of Ty3/Gypsy and Ty1/Copia retrotransposable elements across the chromosome clearly distinguished euchromatin (2:3) from heterochromatin (3:2).
BACKGROUND:Alternative splicing (AS) is a widespread phenomenon in higher eukaryotes but the extent to which it leads to functional protein isoforms and to proteome expansion at large is still a matter of debate. In contrast to animal species, for which AS has been studied extensively at the protein and functional level, protein-centered studies of AS in plant species are scarce. Here we investigate the functional impact of AS in dicot and monocot plant species using a comparative approach.RESULTS:Detailed comparison of AS events in alternative spliced orthologs from the dicot Arabidopsis thaliana and the monocot Oryza sativa (rice) revealed that the vast majority of AS events in both species do not result from functional conservation. Transcript isoforms that are putative targets for the nonsense-mediated decay (NMD) pathway are as likely to contain conserved AS events as isoforms that are translated into proteins. Similar results were obtained when the same comparison was performed between the two more closely related monocot species rice and Zea mays (maize).Genome-wide computational analysis of functional protein domains encoded in alternatively and constitutively spliced genes revealed that only the RNA recognition motif (RRM) is overrepresented in alternatively spliced genes in all species analyzed. In contrast, three domain types were overrepresented in constitutively spliced genes. AS events were found to be less frequent within than outside predicted protein domains and no domain type was found to be enriched with AS introns. Analysis of AS events that result in the removal of complete protein domains revealed that only a small number of domain types is spliced-out in all species analyzed. Finally, in a substantial fraction of cases where a domain is completely removed, this domain appeared to be a unit of a tandem repeat.CONCLUSION:The results from the ortholog comparisons suggest that the ability of a gene to produce more than one functional protein through AS does not persist during evolution. Cross-species comparison of the results of the protein-domain oriented analyses indicates little correspondence between the analyzed species. Based on the premise that functional genetic features are most likely to be conserved during evolution, we conclude that AS has only a limited role in functional expansion of the proteome in plants.
ABSTRACT Pseudomonas fluorescens is of agricultural and economic importance as a biological control agent largely because of its plant association and production of secondary metabolites, in particular 2,4-diacetylphloroglucinol (2,4-DAPG). This polyketide, which is encoded by the eight-gene phl cluster, has antimicrobial effects on phytopathogens, promotes amino acid exudation from plant roots, and induces systemic resistance in plants. Despite its importance, 2,4-DAPG production is limited to a subset of P. fluorescens strains. Determination of the evolution of the phl cluster and understanding the selective pressures promoting its retention or loss in lineages of P. fluorescens will help in the development of P. fluorescens as a viable and effective inoculant for application in agriculture. In this study, genomic and sequence-based approaches were integrated to reconstruct the phylogeny of P. fluorescens and the phl cluster. It was determined that 2,4-DAPG production is an ancestral trait in the species P. fluorescens but that most lineages have lost this capacity through evolution. Furthermore, intragenomic recombination has relocated the phl cluster within the P. fluorescens genome at least three times, but the integrity of the cluster has always been maintained. The possible evolutionary and functional implications for retention of the phl cluster and 2,4-DAPG production in some lineages of P. fluorescens are discussed.
The genome of tomato (Solanum lycopersicum L.) is being sequenced by an international consortium of 10 countries (Korea, China, the United Kingdom, India, the Netherlands, France, Japan, Spain, Italy, and the United States) as part of the larger “International Solanaceae Genome Project (SOL): Systems Approach to Diversity and Adaptation” initiative. The tomato genome sequencing project uses an ordered bacterial artificial chromosome (BAC) approach to generate a high‐quality tomato euchromatic genome sequence for use as a reference genome for the Solanaceae and euasterids. Sequence is deposited at GenBank and at the SOL Genomics Network (SGN). Currently, there are around 1000 BACs finished or in progress, representing more than a third of the projected euchromatic portion of the genome. An annotation effort is also underway by the International Tomato Annotation Group. The expected number of genes in the euchromatin is ∼40,000, based on an estimate from a preliminary annotation of 11% of finished sequence. Here, we present this first snapshot of the emerging tomato genome and its annotation, a short comparison with potato (Solanum tuberosum L.) sequence data, and the tools available for the researchers to exploit this new resource are also presented. In the future, whole‐genome shotgun techniques will be combined with the BAC‐by‐BAC approach to cover the entire tomato genome. The high‐quality reference euchromatic tomato sequence is expected to be near completion by 2010.
Background Tomato ( Solanum lycopersicon ) and potato ( S. tuberosum ) are two economically important crop species, the genomes of which are currently being sequenced. This study presents a first genome-wide analysis of these two species, based on two large collections of BAC end sequences representing approximately 19% of the tomato genome and 10% of the potato genome. Results The tomato genome has a higher repeat content than the potato genome, primarily due to a higher number of retrotransposon insertions in the tomato genome. On the other hand, simple sequence repeats are more abundant in potato than in tomato. The two genomes also differ in the frequency distribution of SSR motifs. Based on EST and protein alignments, potato appears to contain up to 6,400 more putative coding regions than tomato. Major gene families such as cytochrome P450 mono-oxygenases and serine-threonine protein kinases are significantly overrepresented in potato, compared to tomato. Moreover, the P450 superfamily appears to have expanded spectacularly in both species compared to Arabidopsis thaliana , suggesting an expanded network of secondary metabolic pathways in the Solanaceae . Both tomato and potato appear to have a low level of microsynteny with A. thaliana . A higher degree of synteny was observed with Populus trichocarpa , specifically in the region between 15.2 and 19.4 Mb on P. trichocarpa chromosome 10. Conclusion The findings in this paper present a first glimpse into the evolution of Solanaceous genomes, both within the family and relative to other plant species. When the complete genome sequences of these species become available, whole-genome comparisons and protein- or repeat-family specific studies may shed more light on the observations made here.
Ongoing genomics projects of tomato (Solanum lycopersicum) and potato (S. tuberosum) are providing unique tools for comparative mapping studies in Solanaceae. At the chromosomal level, bacterial artificial chromosomes (BACs) can be positioned on pachytene complements by fluorescence in situ hybridization (FISH) on homeologous chromosomes of related species. Here we present results of such a cross-species multicolor cytogenetic mapping of tomato BACs on potato chromosomes 6 and vice versa. The experiments were performed under low hybridization stringency, while blocking with Cot-100 was essential in suppressing excessive hybridization of repeat signals in both within-species FISH and cross-species FISH of tomato BACs. In the short arm we detected a large paracentric inversion that covers the whole euchromatin part with breakpoints close to the telomeric heterochromatin and at the border of the short arm pericentromere. The long arm BACs revealed no deviation in the colinearity between tomato and potato. Further comparison between tomato cultivars Cherry VFNT and Heinz 1706 revealed colinearity of the tested tomato BACs, whereas one of the six potato clones (RH98-856-18) showed minor putative rearrangements within the inversion. Our results present cross-species multicolor BAC-FISH as a unique tool for comparative genetic studies across Solanum species.
Within the framework of the International Solanaceae Genome Project, the genome of tomato (Solanum lycopersicum) is currently being sequenced. We follow a 'BAC-by-BAC' approach that aims to deliver high-quality sequences of the euchromatin part of the tomato genome. BACs are selected from various libraries of the tomato genome on the basis of markers from the F2.2000 linkage map. Prior to sequencing, we validated the precise physical location of the selected BACs on the chromosomes by five-colour high-resolution fluorescent in situ hybridization (FISH) mapping. This paper describes the strategies and results of cytogenetic mapping for chromosome 6 using 75 seed BACs for FISH on pachytene complements. The cytogenetic map obtained showed discrepancies between the actual chromosomal positions of these BACs and their markers on the linkage group. These discrepancies were most notable in the pericentromere heterochromatin, thus confirming previously described suppression of cross-over recombination in that region. In a so called pooled-BAC FISH, we hybridized all seed BACs simultaneously and found a few large gaps in the euchromatin parts of the long arm that are still devoid of seed BACs and are too large for coverage by expanding BAC contigs. Combining FISH with pooled BACs and newly recruited seed BACs will thus aid in efficient targeting of novel seed BACs into these areas. Finally, we established the occurrence of repetitive DNA in heterochromatin/euchromatin borders by combining BAC FISH with hybridization of a labelled repetitive DNA fraction (Cot-100). This strategy provides an excellent means to establish the borders between euchromatin and heterochromatin in this chromosome.
Pseudomonas putida PCL1445 secretes two cyclic lipopeptides, putisolvin I and putisolvin II, which possess a surface-tension-reducing ability, and are able to inhibit biofilm formation and to break down biofilms of Pseudomonas species including Pseudomonas aeruginosa. The putisolvin synthetase gene cluster (pso) and its surrounding region were isolated, sequenced and characterized. Three genes, termed psoA, psoB and psoC, were identified and shown to be involved in putisolvin biosynthesis. The gene products encode the 12 modules responsible for the binding of the 12 amino acids of the putisolvin peptide moiety. Sequence data indicate that the adenylation domain of the 11th module prioritizes the recognition of Val instead of Leu or Ile and consequently favours putisolvin I production over putisolvin II. Detailed analysis of the thiolation domains suggests that the first nine modules recognize the d form of the amino acid residues while the two following modules recognize the l form and the last module the l or d form, indifferently. The psoR gene, which is located upstream of psoA, shows high similarity to luxR-type regulatory genes and is required for the expression of the pso cluster. In addition, two genes, macA and macB, located downstream of psoC were identified and shown to be involved in putisolvin production or export.
Chromosomal coexpression domains are found in a number of different genomes under various developmental conditions. The size of these domains and the number of genes they contain vary. Here, we define local coexpression domains as adjacent genes where all possible pair-wise correlations of expression data are higher than 0.7. In rice, such local coexpression domains range from predominantly two genes, up to 4, and make up ∼5% of the genomic neighboring genes, when examining different expression platforms from the public domain. The genes in local coexpression domains do not fall in the same ontology category significantly more than neighboring genes that are not coexpressed. Duplication, orientation or the distance between the genes does not solely explain coexpression. The regulation of coexpression is therefore thought to be regulated at the level of chromatin structure. The characteristics of the local coexpression domains in rice are strikingly similar to such domains in the Arabidopsis genome. Yet, no microsynteny between local coexpression domains in Arabidopsis and rice could be identified. Although the rice genome is not yet as extensively annotated as the Arabidopsis genome, the lack of conservation of local coexpression domains may indicate that such domains have not played a major role in the evolution of genome structure or in genome conservation.
Cover: A number of diverse Arabidopsis thaliana accessions overlay a map of the world demonstrating the wide genetic variation that exists in this reference plant. A recent analysis of twenty natural accessions resulted in the identification of over 500,000 unique single-nucleotide polymorphisms (SNPs) making Arabidopsis the organism with the highest SNP density relative to genome size. Studies of natural variation and comparative genomics, current areas of focus for the MASC, will be enabled by the availability of these data. Cover image concept: Philip Benfey, Duke University. Cover image design:
It is believed that CLAVATA3 (CLV3) encodes a peptide ligand that interacts with the CLV1/CLV2 receptor complex to limit the number of stem cells in the shoot apical meristem of Arabidopsis thaliana; however, the exact composition of the functional CLV3 product remains a mystery. A recent study on CLV3 shows that the CLV3/ESR (CLE) motif, together with the adjacent C-terminal sequence, is sufficient to execute CLV3 function when fused behind an N-terminal sequence of ERECTA. Here we show that most of the sequences flanking the CLE motif of CLV3 can be deleted without affecting CLV3 function. Using a liquid culture assay, we demonstrate that CLV3p, a synthetic peptide corresponding to the CLE motif of CLV3, is able to restrict the size of the shoot apical meristem in clv3 seedlings but not in clv1 seedlings. In accordance with this decrease in meristem size, application of CLV3p to in vitro-grown clv3 seedlings restricts the expression of the stem cell-promoting transcription factor WUSCHEL. Thus, we propose that the CLE motif is the functional region of CLV3 and that this region acts independently of its adjacent sequences.
It is believed that CLAVATA3 ( CLV3) encodes a peptide ligand that interacts with the CLV1/CLV2 receptor complex to limit the number of stem cells in the shoot apical meristem of Arabidopsis thaliana; however, the exact composition of the functional CLV3 product remains a mystery. A recent study on CLV3 shows that the CLV3/ESR (CLE) motif, together with the adjacent C-terminal sequence, is sufficient to execute CLV3 function when fused behind an N-terminal sequence of ERECTA. Here we show that most of the sequences flanking the CLE motif of CLV3 can be deleted without affecting CLV3 function. Using a liquid culture assay, we demonstrate that CLV3p, a synthetic peptide corresponding to the CLE motif of CLV3, is able to restrict the size of the shoot apical meristem in clv3 seedlings but not in clv1 seedlings. In accordance with this decrease in meristem size, application of CLV3p to in vitro-grown clv3 seedlings restricts the expression of the stem cell-promoting transcription factor WUSCHEL. Thus, we propose that the CLE motif is the functional region of CLV3 and that this region acts independently of its adjacent sequences.
In both the monocot rice and the dicot Arabidopsis, highly expressed genes have more and longer introns and a larger primary transcript than genes expressed at a low level: higher expressed genes tend to be less compact than lower expressed genes. In animal genomes, it is the other way round. Although the length differences in plant genes are much smaller than in animals, these findings indicate that plant genes are in this respect different from animal genes. Explanations for the relationship between gene configuration and gene expression in animals might be (or might have been) less important in plants. We speculate that selection, if any, on genome onfiguration has taken a different turn after the divergence of plants and animals.
The potential of increased allergenicity of genetically modified (GM) food is a major issue in the biosafety assessment of GM crops. Food allergy is a complex adverse reaction towards food that is difficult to diagnose or treat. Potential unintended effects in food and food processing related to allergenicity increase the complexity of the safety assessments. A detailed decision tree to evaluate the potential allergenicity of GM food starts with the source of the gene. Comparison to known allergens is the only known way to predict the allergenic potential of a protein without laboratory tests. Biochemical characteristics have limited predictive power. Here, an overview is given of bioinformatics approaches towards prediction of allergenicity. These are based on advanced multi-alignment, clustering and pattern recognition tools as well as on three-dimensional structural analysis, using a variety of databases. The criteria what establishes 'sufficient similarity' to identify allergens are still being discussed. Cur-rent criteria may need improvements. Ongoing developments will contribute to the future ability to predict the potential allergenicity of proteins in silico and may result in new approaches to circumvent the issue for future novel crops.
The genome of tomato (Solanum lycopersicum) is being sequenced by an international consortium of 10 countries (Korea, China, the United Kingdom, India, The Netherlands, France, Japan, Spain, Italy and the United States) as part of a larger initiative called the 'International Solanaceae Genome Project (SOL): Systems Approach to Diversity and Adaptation'. The goal of this grassroots initiative, launched in November 2003, is to establish a network of information, resources and scientists to ultimately tackle two of the most significant questions in plant biology and agriculture: (1) How can a common set of genes/proteins give rise to a wide range of morphologically and ecologically distinct organisms that occupy our planet? (2) How can a deeper understanding of the genetic basis of plant diversity be harnessed to better meet the needs of society in an environmentally friendly and sustainable manner? The Solanaceae and closely related species such as coffee, which are included in the scope of the SOL project, are ideally suited to address both of these questions. The first step of the SOL project is to use an ordered BAC approach to generate a high quality sequence for the euchromatic portions of the tomato as a reference for the Solanaceae. Due to the high level of macro and micro-synteny in the Solanaceae the BAC-by-BAC tomato sequence will form the framework for shotgun sequencing of other species. The starting point for sequencing the genome is BACs anchored to the genetic map by overgo hybridization and AFLP technology. The overgos are derived from approximately 1500 markers from the tomato high density F2-2000 genetic map (http://sgn.cornell.edu/). These seed BACs will be used as anchors from which to radiate the tiling path using BAC end sequence data. Annotation will be performed according to SOL project guidelines. All the information generated under the SOL umbrella will be made available in a comprehensive website. The information will be interlinked with the ultimate goal that the comparative biology of the Solanaceae-and beyond-achieves a context that will facilitate a systems biology approach.
CLAVATA3 (CLV3), CLV3/ESR19 (CLE19), and CLE40 belong to a family of 26 genes in Arabidopsis thaliana that encode putative peptide ligands with unknown identity. It has been shown previously that ectopic expression of any of these three genes leads to a consumption of the root meristem. Here, we show that in vitro application of synthetic 14-amino acid peptides, CLV3p, CLE19p, and CLE40p, corresponding to the conserved CLE motif, mimics the overexpression phenotype. The same result was observed when CLE19 protein was applied externally. Interestingly, clv2 failed to respond to the peptide treatment, suggesting that CLV2 is involved in the CLE peptide signaling. Crossing of the CLE19 overexpression line with clv mutants confirms the involvement of CLV2. Analyses using tissue-specific marker lines revealed that the peptide treatments led to a premature differentiation of the ground tissue daughter cells and misspecification of cell identity in the pericycle and endodermis layers. We propose that these 14-amino acid peptides represent the major active domain of the corresponding CLE proteins, which interact with or saturate an unknown cell identity-maintaining CLV2 receptor complex in roots, leading to consumption of the root meristem.