The Oryza officinalis complex is the largest species group in Oryza, with more than nine species from four continents, and is a tertiary gene pool that can be exploited in breeding programs for the improvement of cultivated rice. Most diploid and tetraploid members of this group have a C genome. Using a new reference C genome for the diploid species O. officinalis, and draft genomes for two other C genome diploid species Oryza eichingeri and Oryza rhizomatis, we examine the influence of transposable elements on genome structure and provide a detailed phylogeny and evolutionary history of the Oryza C genomes. The O. officinalis genome is 1.6 times larger than the A genome of cultivated Oryza sativa, mostly due to proliferation of Gypsy type long-terminal repeat transposable elements, but overall syntenic relationships are maintained with other Oryza genomes (A, B, and F). Draft genome assemblies of the two other C genome diploid species, Oryza eichingeri and Oryza rhizomatis, and short-read resequencing of a series of other C genome species and accessions reveal that after the divergence of the C genome progenitor, there was still a substantial degree of variation within the C genome species through proliferation and loss of both DNA and long-terminal repeat transposable elements. We provide a detailed phylogeny and evolutionary history of the Oryza C genomes and a genomic resource for the exploitation of the Oryza tertiary gene pool.
51) Int. Cl. ................. A61K 3/495; A61K 31/445; CO7D 295/00; CO7D 211/30 52 U.S. Cl. .................................... 514/255; 514/210; 514/212; 514/315; 514/423; 514/452; 514/548; 540/575; 540/610,544/387: 544/388; 546/247; 548/530; 548/537; 548/950; 549/333; 549/372; 500/155 58) Field of Search ................ 540/575, 610,544/387, 544/388; 546/247; 548/530, 537,950, 514/210, 212, 255,315,423,426, 452,548; 549/333, 372; 560/155
Understanding the processes that regulate plant sink formation and development at the molecular level will contribute to the areas of crop breeding, food production and plant evolutionary studies. We report the annotation and analysis of the draft genome sequence of the radish Raphanus sativus var. hortensis (long and thick root radish) and transcriptome analysis during root development. Based on the hybrid assembly approach of next-generation sequencing, a total of 383 Mb (N50 scaffold: 138.17 kb) of sequences of the radish genome was constructed containing 54,357 genes. Syntenic and phylogenetic analyses indicated that divergence between Raphanus and Brassica coincide with the time of whole genome triplication (WGT), suggesting that WGT triggered diversification of Brassiceae crop plants. Further transcriptome analysis showed that the gene functions and pathways related to carbohydrate metabolism were prominently activated in thickening roots, particularly in cell proliferating tissues. Notably, the expression levels of sucrose synthase 1 (SUS1) were correlated with root thickening rates. We also identified the genes involved in pungency synthesis and their transcription factors.
We elucidated the genome sequence ofGlycine maxcv. Enrei to provide a reference for characterization of Japanese domestic soybean cultivars. The whole genome sequence obtained using a next-generation sequencer was used for reference mapping into the current genome assembly ofG. maxcv. Williams 82 obtained by the Soybean Genome Sequencing Consortium in the USA. After sequencing and assembling the whole genome shotgun reads, we obtained a data set with about 928 Mbs total bases and 60,838 gene models. Phylogenetic analysis provided glimpses into the ancestral relationships of both cultivars and their divergence from the complex that include the wild relatives of soybean. The gene models were analyzed in relation to traits associated with anthocyanin and flavonoid biosynthesis and an overall profile of the proteome. The sequence data are made available in DAIZUbase in order to provide a comprehensive informatics resource for comparative genomics of a wide range of soybean cultivars in Japan and a reference tool for improvement of soybean cultivars worldwide.
Soybean [Glycine max (L) Merrill] is one of the most important leguminous crops and ranks fourth after to rice, wheat and maize in terms of world crop production. Soybean contains abundant protein and oil, which makes it a major source of nutritious food, livestock feed and industrial products. In Japan, soybean is also an important source of traditional staples such as tofu, natto, miso and soy sauce. The soybean genome was determined in 2010. With its enormous size, physical mapping and genome sequencing are the most effective approaches towards understanding the structure and function of the soybean genome. We constructed bacterial artificial chromosome (BAC) libraries from the Japanese soybean cultivar, Enrei. The end-sequences of approximately 100,000 BAC clones were analyzed and used for construction of a BAC-based physical map of the genome. BLAST analysis between Enrei BAC-end sequences and the Williams82 genome was carried out to increase the saturation of the map. This physical map will be used to characterize the genome structure of Japanese soybean cultivars, to develop methods for the isolation of agronomically important genes and to facilitate comparative soybean genome research. The current status of physical mapping of the soybean genome and construction of database are presented.
Similarity of gene expression across a wide range of biological conditions can be efficiently used in characterization of gene function. We have constructed a rice gene coexpression database, RiceFREND (http://ricefrend.dna.affrc.go.jp/), to identify gene modules with similar expression profiles and provide a platform for more accurate prediction of gene functions. Coexpression analysis of 27 201 genes was performed against 815 microarray data derived from expression profiling of various organs and tissues at different developmental stages, mature organs throughout the growth from transplanting until harvesting in the field and plant hormone treatment conditions, using a single microarray platform. The database is provided with two search options, namely, 'single guide gene search' and 'multiple guide gene search' to efficiently retrieve information on coexpressed genes. A user-friendly web interface facilitates visualization and interpretation of gene coexpression networks in HyperTree, Cytoscape Web and Graphviz formats. In addition, analysis tools for identification of enriched Gene Ontology terms and cis-elements provide clue for better prediction of biological functions associated with the coexpressed genes. These features allow users to clarify gene functions and gene regulatory networks that could lead to a more thorough understanding of many complex agronomic traits.
A wide range of resources on gene expression profiling enhance various strategies in plant molecular biology particularly in characterization of gene function. We have updated our gene expression profile database, RiceXPro (http://ricexpro.dna.affrc.go.jp/), to provide more comprehensive information on the transcriptome of rice encompassing the entire growth cycle and various experimental conditions. The gene expression profiles are currently grouped into three categories, namely, ‘field/development’ with 572 data corresponding to 12 data sets, ‘plant hormone’ with 143 data corresponding to 13 data sets and ‘cell- and tissue-type’ comprising of 38 microarray data. In addition to the interface for retrieving expression information of a gene/genes in each data set, we have incorporated an interface for a global approach in searching an overall view of the gene expression profiles from multiple data sets within each category. Furthermore, we have also added a BLAST search function that enables users to explore expression profile of a gene/genes with similarity to a query sequence. Therefore, the updated version of RiceXPro can be used more efficiently to survey the gene expression signature of rice in sufficient depth and may also provide clues on gene function of other cereal crops.
There is controversy as to whether gene expression is silenced in the functional centromere. The complete genomic sequences of the centromeric regions in higher eukaryotes have not been fully elucidated, because the presence of highly repetitive sequences complicates many aspects of genomic sequencing. We performed resequencing, assembly, and sequence finishing of two P1-derived artificial chromosome clones in the centromeric region of rice (Oryza sativa L.) chromosome 5 (Cen5). The pericentromeric region, where meiotic recombination is silenced, is located at the center of chromosome 5 and is 2.14 Mb long; a total of six restriction-fragment-length polymorphism markers (R448, C1388, S20487S, E3103S, C53260S, and R2059) genetically mapped at 54.6 cM were located in this region. In the pericentromeric region, 28 genes were annotated on the short arm and 45 genes on the long arm. To quantify all transcripts in this region, we performed massive parallel sequencing of mRNA. Transcriptional density (total length of transcribed region/length of the genomic region) and expression level (number of uniquely mapped reads/length of transcribed region) were calculated on the basis of the mapped reads on the rice genome. Transcriptional density and expression level were significantly lower in Cen5 than in the average of the other chromosomal regions. Moreover, transcriptional density in Cen5 was significantly lower on the short arm than on the long arm; the distribution of transcriptional density was asymmetric. The genomic sequence of Cen5 has been integrated into the most updated reference rice genome sequence constructed by the International Rice Genome Sequencing Project.
Plants have developed several morphological and physiological strategies to adapt to phosphate stress. We analyzed the inducible transcripts associated with phosphate starvation and over-abundant phosphate supply to characterize the transcriptome in rice seedlings using the mRNA-Seq strategy. Fifty-three million reads obtained from 16 libraries under various phosphate stress and recovery treatments were uniquely mapped to the rice genome. Transcripts identified specifically tagged to 40,574 (root) and 39,748 (shoot) Rice Annotation Project (RAP) transcripts. Additionally, we detected uniquely 10,388 transcripts with no match to any RAP transcript. These transcripts that showed specific response to Pi stress include those without ORFs that may act as non-protein coding transcripts. With an accompanying browser of the transcriptome under Pi stress, a deeper understanding of the structural and functional features of both annotated and unannotated Pi stress-responsive transcripts can provide useful information in improving Pi acquisition and utilization in rice and other cereal crops.
Full-length cDNA (FLcDNA) libraries consisting of 172,000 clones were constructed from a two-row malting barley cultivar (Hordeum vulgare ‘Haruna Nijo’) under normal and stressed conditions. After sequencing the clones from both ends and clustering the sequences, a total of 24,783 complete sequences were produced. By removing duplicates between these and publicly available sequences, 22,651 representative sequences were obtained: 17,773 were novel barley FLcDNAs, and 1,699 were barley specific. Highly conserved genes were found in the barley FLcDNA sequences for 721 of 881 rice (Oryza sativa) trait genes with 50% or greater identity. These FLcDNA resources from our Haruna Nijo cDNA libraries and the full-length sequences of representative clones will improve our understanding of the biological functions of genes in barley, which is the cereal crop with the fourth highest production in the world, and will provide a powerful tool for annotating the barley genome sequences that will become available in the near future.
A vagrant individual Branta canadensis visited Hokkaido on migration in spring 2006.We present details of its eventual identification as Branta canadensis parvipes and our observations of this rare visitor, which has not been formally described from Japan.We describe its staging at two separate locations in southern and northern Hokkaido, its association with other geese, and its length of stay.
P>Here we present the genomic sequence of the African cultivated rice, Oryza glaberrima, and compare these data with the genome sequence of Asian cultivated rice, Oryza sativa. We obtained gene-enriched sequences of O. glaberrima that correspond to about 25% of the gene regions of the O. sativa (japonica) genome by methylation filtration and subtractive hybridization of repetitive sequences. While patterns of amino acid changes did not differ between the two species in terms of the biochemical properties, genes of O. glaberrima generally showed a larger synonymous-nonsynonymous substitution ratio, suggesting that O. glaberrima has undergone a genome-wide relaxation of purifying selection. We further investigated nucleotide substitutions around splice sites and found that eight genes of O. sativa experienced changes at splice sites after the divergence from O. glaberrima. These changes produced novel introns that partially truncated functional domains, suggesting that these newly emerged introns affect gene function. We also identified 2451 simple sequence repeats (SSRs) from the genomes of O. glaberrima and O. sativa. Although tri-nucleotide repeats were most common among the SSRs and were overrepresented in the protein-coding sequences, we found that selection against indels of tri-nucleotide repeats was relatively weak in both African and Asian rice. Our genome-wide sequencing of O. glaberrima and in-depth analyses provide rice researchers not only with useful genomic resources for future breeding but also with new insights into the genomic evolution of the African and Asian rice species.
Gene duplication occurs by either DNA- or RNA-based processes; the latter duplicates single genes via retroposition of messenger RNA. The expression of a retroposed gene copy (retrocopy) is expected to be uncorrelated with its source gene because upstream promoter regions are usually not part of the retroposition process. In contrast, DNA-based duplication often encompasses both the coding and the intergenic (promoter) regions; hence, expression is often correlated, at least initially, between DNA-based duplicates. In this study, we identified 150 retrocopies in rice (Oryza sativa L. ssp japonica), most of which represent ancient retroposition events. We measured their expression from high-throughput RNA sequencing (RNAseq) data generated from seven tissues. At least 66% of the retrocopies were expressed but at lower levels than their source genes. However, the tissue specificity of retrogenes was similar to their source genes, and expression between retrocopies and source genes was correlated across tissues. The level of correlation was similar between RNA-and DNA-based duplicates, and they decreased over time at statistically indistinguishable rates. We extended these observations to previously identified retrocopies in Arabidopsis thaliana, suggesting they may be general features of the process of retention of plant retrogenes.
BACKGROUND:Microarray technology is limited to monitoring the expression of previously annotated genes that have corresponding probes on the array. Computationally annotated genes have not fully been validated, because ESTs and full-length cDNAs cannot cover entire transcribed regions. Here, mRNA-Seq (an Illumina cDNA sequencing application) was used to monitor whole mRNAs of salinity stress-treated rice tissues.RESULTS:Thirty-six-base-pair reads from whole mRNAs were mapped to the rice genomic sequence: 72.0% to 75.2% were mapped uniquely to the genome, and 5.0% to 5.7% bridged exons. From the piling up of short reads mapped on the genome, a series of programs (Bowtie, TopHat, and Cufflinks) comprehensively predicted 51,301 (shoot) and 54,491 (root) transcripts, including 2,795 (shoot) and 3,082 (root) currently unannotated in the Rice Annotation Project database. Of these unannotated transcripts, 995 (shoot) and 1,052 (root) had ORFs similar to those encoding the amino acid sequences of functional proteins in a BLASTX search against UniProt and RefSeq databases. Among the unannotated genes, 213 (shoot) and 436 (root) were differentially expressed in response to salinity stress. Sequence-based and array-based measurements of the expression ratios of previously annotated genes were highly correlated.CONCLUSION:Unannotated transcripts were identified on the basis of the piling up of mapped reads derived from mRNAs in rice. Some of these unannotated transcripts encoding putative functional proteins were expressed differentially in response to salinity stress.