Mobile insertion elements such as transposons and T-DNA generate useful genetic variation and are important tools for functional genomics studies in plants and animals. The spectrum of mutations obtained in different systems can be highly influenced by target site preferences inherent in the mechanism of DNA integration. We investigated the target site preferences of Agrobacterium T-DNA insertions in the chromosomes of the model plant Arabidopsis thaliana . The relative frequencies of insertions in genic and intergenic regions of the genome were calculated and DNA composition features associated with the insertion site flanking sequences were identified. Insertion frequencies across the genome indicate that T-strand integration is suppressed near centromeres and rDNA loci, progressively increases towards telomeres, and is highly correlated with gene density. At the gene level, T-DNA integration events show a statistically significant preference for insertion in the 5′ and 3′ flanking regions of protein coding sequences as well as the promoter region of RNA polymerase I transcribed rRNA gene repeats. The increased insertion frequencies in 5′ upstream regions compared to coding sequences are positively correlated with gene expression activity and DNA sequence composition. Analysis of the relationship between DNA sequence composition and gene activity further demonstrates that DNA sequences with high CG-skew ratios are consistently correlated with T-DNA insertion site preference and high gene expression. The results demonstrate genomic and gene-specific preferences for T-strand integration and suggest that DNA sequences with a pronounced transition in CG- and AT-skew ratios are preferred targets for T-DNA integration.
After the completion of the genomic sequence of Arabidopsis thaliana, it is now a priority to identify all the genes, their patterns of expression and functions. Transcript profiling is playing a substantial role in annotating and determining gene functions, having advanced from one-gene-at-a-time methods to technologies that provide a holistic view of the genome. In this review, comprehensive transcript profiling methodologies are described, including two that are used extensively by the authors, cDNA-AFLP and cDNA microarraying. Both these technologies illustrate the requirement to integrate molecular biology, automation, LIMS and data analysis. With so much uncharted territory in the Arabidopsis genome, and the desire to tackle complex biological traits, such integrated systems will provide a rich source of data for the correlative, functional annotation of genes.
cDNA-AFLP, a technology historically used to identify small numbers of differentially expressed genes, was adapted as a genome-wide transcript profiling method. mRNA levels were assayed in a diverse range of tissues from Arabidopsis thaliana plants grown under a variety of environmental conditions. The resulting cDNA-AFLP fragments were sequenced. By linking cDNA-AFLP fragments to their corresponding mRNAs via these sequences, a database was generated that contained quantitative expression information for up to two-thirds of gene loci in A. thaliana, ecotype Ws. Using this resource, the expression levels of genes, including those with high nucleotide sequence similarity, could be determined in a high-throughput manner merely by comparing cDNA-AFLP profiles with the database. The lengths of cDNA-AFLP fragments inferred from their electrophoretic mobilities correlated well with actual fragment lengths determined by sequencing. In addition, the concentrations of AFLP fragments from single cDNAs were highly correlated, illustrating the validity of cDNA-AFLP as a quantitative, genome-wide, transcript profiling method. cDNA-AFLP profiles were also qualitatively consistent with mRNA profiles obtained from parallel microarray analysis, and with data from previous studies.
A polypeptide of approximately 11 000 daltons (11 kDa protein) encoded by an open reading frame (10.9 ORF) from the virion sense of maize streak virus (MSV) DNA has been detected among the products of in vitro translation reactions programmed with RNA from infected maize plants and also in total protein extracts from infected leaves. The 11 kDa protein has not been detected in virions and is therefore proposed to have a nonstructural role.