The Wheat@URGI portal has been developed to provide the international community of researchers and breeders with access to the bread wheat reference genome sequence produced by the International Wheat Genome Sequencing Consortium. Genome browsers, BLAST, and InterMine tools have been established for in-depth exploration of the genome sequence together with additional linked datasets including physical maps, sequence variations, gene expression, and genetic and phenomic data from other international collaborative projects already stored in the GnpIS information system. The portal provides enhanced search and browser features that will facilitate the deployment of the latest genomics resources in wheat improvement.
Agronomical characters such as yield, biotic and abiotic resistance are determined by the genetic information carried by the plant genome. Within the framework of the International Wheat Genome Sequencing Initiative (IWGSC) effort for obtaining a reference sequence of the bread wheat genome and to provide the scientific communities dealing with large and complex genomes a versatile, easy-to-use online annotation tool, we have developed the TriAnnot pipeline (Leroy et al. 2012). TriAnnot has already been used to annotate the wheat chromosome 3B (Choulet et al. 2014) and 4D (Helguera et al. 2015). Annotation of chromosome 1B is currently underway (INRA – GDEC), as well as the annotation of chromosome 7A in collaboration with the University of Murdoch (Australia). TriAnnot has the ambition to federate the international community around the annotation of the 21 wheat chromosomes and, in this perspective, the TriAnnot source code has been strongly improved to facilitate its deployment on external computing resources such as: IEB, Olomouc (Czech Republic); The Pawsey Supercomputing Centre, WA (Australia); National Research Council of Canada, Saskatoon (Canada); CRRI, Clermont-Ferrand and ABiMS CNRS bioinformatics platform, Roscoff (France). The development of a virtual machine is also underway in collaboration with ABiMS and the French Institute of bioinformatics (IFB). TriAnnot was also adapted for the annotation of other plant genomes such as barley, maize, rice and oak. A public instance of TriAnnot is currently deployed on the cluster of the URGI bioinformatics platform (Versailles) and usable through a user-friendly web interface or in command line.
An ordered draft sequence of the 17-gigabase hexaploid bread wheat ( Triticum aestivum ) genome has been produced by sequencing isolated chromosome arms. We have annotated 124,201 gene loci distributed nearly evenly across the homeologous chromosomes and subgenomes. Comparative gene analysis of wheat subgenomes and extant diploid and tetraploid wheat relatives showed that high sequence similarity and structural conservation are retained, with limited gene loss, after polyploidization. However, across the genomes there was evidence of dynamic gene gain, loss, and duplication since the divergence of the wheat lineages. A high degree of transcriptional autonomy and no global dominance was found for the subgenomes. These insights into the genome biology of a polyploid crop provide a springboard for faster gene isolation, rapid genetic marker development, and precise breeding to meet the needs of increasing food demand worldwide.
We produced a reference sequence of the 1-gigabase chromosome 3B of hexaploid bread wheat. By sequencing 8452 bacterial artificial chromosomes in pools, we assembled a sequence of 774 megabases carrying 5326 protein-coding genes, 1938 pseudogenes, and 85% of transposable elements. The distribution of structural and functional features along the chromosome revealed partitioning correlated with meiotic recombination. Comparative analyses indicated high wheat-specific inter- and intrachromosomal gene duplication activities that are potential sources of variability for adaption. In addition to providing a better understanding of the organization, function, and evolution of a large and polyploid genome, the availability of a high-quality sequence anchored to genetic maps will accelerate the identification of genes underlying important agronomic traits.
In support of the international effort to obtain a reference sequence of the bread wheat genome and to provide plant communities dealing with large and complex genomes with a versatile, easy-to-use online automated tool for annotation, we have developed the TriAnnot pipeline. Its modular architecture allows for the annotation and masking of transposable elements, the structural, and functional annotation of protein-coding genes with an evidence-based quality indexing, and the identification of conserved non-coding sequences and molecular markers. The TriAnnot pipeline is parallelized on a 712 CPU computing cluster that can run a 1-Gb sequence annotation in less than 5 days. It is accessible through a web interface for small scale analyses or through a server for large scale annotations. The performance of TriAnnot was evaluated in terms of sensitivity, specificity, and general fitness using curated reference sequence sets from rice and wheat. In less than 8 h, TriAnnot was able to predict more than 83% of the 3,748 CDS from rice chromosome 1 with a fitness of 67.4%. On a set of 12 reference Mb-sized contigs from wheat chromosome 3B, TriAnnot predicted and annotated 93.3% of the genes among which 54% were perfectly identified in accordance with the reference annotation. It also allowed the curation of 12 genes based on new biological evidences, increasing the percentage of perfect gene prediction to 63%. TriAnnot systematically showed a higher fitness than other annotation pipelines that are not improved for wheat. As it is easily adaptable to the annotation of other plant genomes,TriAnnot should become a useful resource for the annotation of large and complex genomes in the future.
License Commons Creative . http://creativecommons.org/licenses/by-nc/3.0/ described at as a Creative Commons License (Attribution-NonCommercial 3.0 Unported License), ). After six months, it is available under http://genome.cshlp.org/site/misc/terms.xhtml for the first six months after the full-issue publication date (see This article is distributed exclusively by Cold Spring Harbor Laboratory Press
The comparison of the chromosome numbers of today's species with common reconstructed paleo-ancestors has led to intense speculation of how chromosomes have been rearranged over time in mammals. However, similar studies in plants with respect to genome evolution as well as molecular mechanisms leading to mosaic synteny blocks have been lacking due to relevant examples of evolutionary zooms from genomic sequences. Such studies require genomes of species that belong to the same family but are diverged to fall into different subfamilies. Our most important crops belong to the family of the grasses, where a number of genomes have now been sequenced. Based on detailed paleogenomics, using inference from n = 5-12 grass ancestral karyotypes (AGKs) in terms of gene content and order, we delineated sequence intervals comprising a complete set of junction break points of orthologous regions from rice, maize, sorghum, and Brachypodium genomes, representing three different subfamilies and different polyploidization events. By focusing on these sequence intervals, we could show that the chromosome number variation/reduction from the n = 12 common paleo-ancestor was driven by nonrandom centric double-strand break repair events. It appeared that the centromeric/telomeric illegitimate recombination between nonhomologous chromosomes led to nested chromosome fusions (NCFs) and synteny break points (SBPs). When intervals comprising NCFs were compared in their structure, we concluded that SBPs (1) were meiotic recombination hotspots, (2) corresponded to high sequence turnover loci through repeat invasion, and (3) might be considered as hotspots of evolutionary novelty that could act as a reservoir for producing adaptive phenotypes.
On 12 January, 2010, the International Wheat Genome Sequencing Consortium (IWGSC) organized a workshop to develop and discuss protocols and standards for the physical mapping of the hexaploid wheat genome and develop a consensus. In addition, the workshop surveyed the sequencing efforts undertaken within the consortium to coördinate the studies carried out in member laboratories. The goal was to ensure homogeneity in the procedures used for constructing the wheat physical maps by providing guidelines developed in expert laboratories and distributing these to the groups participating in the physical mapping and sequencing of bread wheat chromosomes under the auspice of the IWGSC.
Paleogenomics seeks to reconstruct ancestral genomes from the genes of today's species. The characterization of paleo-duplications represented by 11,737 orthologs and 4,382 paralogs identified in five species belonging to three of the agronomically most important subfamilies of grasses, that is, Ehrhartoideae (rice) Panicoideae (sorghum, maize), and Pooideae (wheat, barley), permitted us to propose a model for an ancestral genome with a minimal size of 33.6 Mb structured in five proto-chromosomes containing at least 9,138 predicted proto-genes. It appears that only four major evolutionary shuffling events (alpha, beta, gamma, and delta) explain the divergence of these five cereal genomes during their evolution from a common paleo-ancestor. Comparative analysis of ancestral gene function with rice as a reference indicated that five categories of genes were preferentially modified during evolution. Furthermore, alignments between the five grass proto-chromosomes and the recently identified seven eudicot proto-chromosomes indicated that additional very active episodes of genome rearrangements and gene mobility occurred during angiosperm evolution. If one compares the pace of primate evolution of 90 million years (233 species) to 60 million years of the Poaceae (10,000 species), change in chromosome structure through speciation has accelerated significantly in plants.
Recent updates in comparative genomics among cereals have provided the opportunity to identify conserved orthologous set (COS) DNA sequences for cross-genome map-based cloning of candidate genes underpinning quantitative traits. New tools are described that are applicable to any cereal genome of interest, namely, alignment criterion for orthologous couples identification, as well as the Intron Spanning Marker software to automatically select intron-spanning primer pairs. In order to test the software, it was applied to the bread wheat genome, and 695 COS markers were assigned to 1,535 wheat loci (on average one marker/2.6 cM) based on 827 robust rice-wheat orthologs. Furthermore, 31 of the 695 COS markers were selected to fine map a pentosan viscosity quantitative trait loci (QTL) on wheat chromosome 7A. Among the 31 COS markers, 14 (45%) were polymorphic between the parental lines and 12 were mapped within the QTL confidence interval with one marker every 0.6 cM defining candidate genes among the rice orthologous region.
Anchored physical maps represent essential frameworks for map-based cloning, comparative genomics studies, and genome sequencing projects. High throughput anchoring can be achieved by polymerase chain reaction (PCR) screening of bacterial artificial chromosome (BAC) library pools with molecular markers. However, for large genomes such as wheat, the development of high dimension pools and the number of reactions that need to be performed can be extremely large making the screening laborious and costly. To improve the cost efficiency of anchoring in such large genomes, we have developed a new software named Elephant (electronic physical map anchoring tool) that combines BAC contig information generated by FingerPrinted Contig with results of BAC library pools screening to identify BAC addresses with a minimal amount of PCR reactions. Elephant was evaluated during the construction of a physical map of chromosome 3B of hexaploid wheat. Results show that a one dimensional pool screening can be sufficient to anchor a BAC contig while reducing the number of PCR by 384-fold thereby demonstrating that Elephant is an efficient and cost-effective tool to support physical mapping in large genomes.