Wheat is the most widely cultivated crop in the world, with over 215 million hectares grown annually. The 10+ Wheat Genomes Project recently sequenced and assembled to chromosome-level the genomes of nine wheat cultivars, uncovering genetic diversity and selection within the pan-genome of wheat. Here, we provide a wheat pan-transcriptome with de novo annotation and differential expression analysis for these wheat cultivars across multiple tissues. Using the de novo annotations we identify cultivar-specific genes and define the core and dispensable genomes. Expression analysis across cultivars and tissues reveals conservation in expression between a large core set of homeologous genes, in addition to widespread changes in subgenome homeolog expression bias between cultivars and cultivar-specific expression profiles. We utilise both the newly constructed gene-based wheat pan-genome and pan-transcriptome, demonstrating variation in the prolamin superfamily and immune-reactive proteins across cultivars.
Alopecurus aequalis is a winter annual or short-lived perennial bunchgrass which has in recent years emerged as the dominant agricultural weed of barley and wheat in certain regions of China and Japan, causing significant yield losses. Its robust tillering capacity and high fecundity, combined with the development of both target and non-target-site resistance to herbicides means it is a formidable challenge to food security. Here we report on a chromosome-scale assembly of A. aequalis with a genome size of 2.83 Gb. The genome contained 33,758 high-confidence protein-coding genes with functional annotation. Comparative genomics revealed that the genome structure of A. aequalis is more similar to Hordeum vulgare rather than the more closely related Alopecurus myosuroides.
Wheat is the most widely cultivated crop in the world with over 215 million hectares grown annually. However, to meet the demands of a growing global population, breeders face the challenge of increasing wheat production by approximately 60% within the next 40 years. The 10+ Wheat Genomes Project recently sequenced and assembled the genomes of 15 wheat cultivars to develop our understanding of genetic diversity and selection within the pan-genome of wheat. Here, we provide a wheat pan-transcriptome with de novo annotation and differential expression analysis for nine of these wheat cultivars, across multiple different tissues and whole seedlings sampled at dusk/dawn. Analysis of these de novo annotations facilitated the discovery of genes absent from the Chinese Spring reference, identified genes specific to particular cultivars and defined the core and dispensable genomes. Expression analysis across cultivars and tissues revealed conservation in expression between a large core set of homoeologous genes, but also widespread changes in sub-genome homoeolog expression bias between cultivars. Co-expression network analysis revealed the impact of divergence of sub-genome homoeolog expression and identified tissue-associated cultivar-specific expression profiles. In summary, this work provides both a valuable resource for the wider wheat community and reveals diversity in gene content and expression patterns between global wheat cultivars. ### Competing Interest Statement The authors have declared no competing interest.
Fragilariopsis cylindrus CCMP1102 is characterised by a complex genome with significant levels of heterozygosity between haplotypes, > 35% repeats, and an unknown karyotype. This complexity hindered prior assemblies, which show coverage discrepancies indicative of incompleteness. Here, we use a k-mer spectra analysis to reveal the coverage signature for a third haplotype. We applied a novel haplotype-specific assembly method to reconstruct the F. cylindrus CCMP1102 genome, producing 10 fully assembled chromosomes capped by telomeres, and a putative chromosome with a single breakpoint. Our analysis shows triploidy, two cases of aneuploidy, and several truncations. We also present evidence that F. cylindrus reproduces sexually. Taken together, our analytical approach is capable of haplotype-resolved assemblies from structurally complex, poly-ploid genomes, making it suitable for complex genomes of non-model organisms, including those with unknown karyotype.
Barley (Hordeum vulgare) is one of the most important crops worldwide and is also considered a research model for the large-genome small grain temperate cereals. Despite genomic resources improving all the time, they are limited for the cv Golden Promise, the most efficient genotype for genetic transformation. We have developed a barley cv Golden Promise reference assembly integrating Illumina paired-end reads, long mate-pair reads, Dovetail Chicago in vitro proximity ligation libraries and chromosome conformation capture sequencing (Hi-C) libraries into a contiguous reference assembly. The assembled genome of 7 chromosomes and 4.13Gb in size, has a super-scaffold N50 after Chicago libraries of 4.14Mb and contains only 2.2% gaps. Using BUSCO (benchmarking universal single copy orthologous genes) as evaluation the genome assembly contains 95.2% of complete and single copy genes from the plant database. A high-quality Golden Promise reference assembly will be useful and utilized by the whole barley research community but will prove particularly useful for CRISPR-Cas9 experiments.
Advances in genomics have expedited the improvement of several agriculturally important crops but similar efforts in wheat ( Triticum spp.) have been more challenging. This is largely owing to the size and complexity of the wheat genome 1 , and the lack of genome-assembly data for multiple wheat lines 2,3 . Here we generated ten chromosome pseudomolecule and five scaffold assemblies of hexaploid wheat to explore the genomic diversity among wheat lines from global breeding programs. Comparative analysis revealed extensive structural rearrangements, introgressions from wild relatives and differences in gene content resulting from complex breeding histories aimed at improving adaptation to diverse environments, grain yield and quality, and resistance to stresses 4,5 . We provide examples outlining the utility of these genomes, including a detailed multi-genome-derived nucleotide-binding leucine-rich repeat protein repertoire involved in disease resistance and the characterization of Sm1 6 , a gene associated with insect resistance. These genome assemblies will provide a basis for functional gene discovery and breeding to deliver the next generation of modern wheat cultivars.
Background Sequence exchange between homologous chromosomes through crossing over and gene conversion is highly conserved among eukaryotes, contributing to genome stability and genetic diversity. A lack of recombination limits breeding efforts in crops; therefore, increasing recombination rates can reduce linkage drag and generate new genetic combinations. Results We use computational analysis of 13 recombinant inbred mapping populations to assess crossover and gene conversion frequency in the hexaploid genome of wheat ( Triticum aestivum ). We observe that high-frequency crossover sites are shared between populations and that closely related parents lead to populations with more similar crossover patterns. We demonstrate that gene conversion is more prevalent and covers more of the genome in wheat than in other plants, making it a critical process in the generation of new haplotypes, particularly in centromeric regions where crossovers are rare. We identify quantitative trait loci for altered gene conversion and crossover frequency and confirm functionality for a novel RecQ helicase gene that belongs to an ancient clade that is missing in some plant lineages including Arabidopsis. Conclusions This is the first gene to be demonstrated to be involved in gene conversion in wheat. Harnessing the RecQ helicase has the potential to break linkage drag utilizing widespread gene conversions.
The Sequence Distance Graph (SDG) framework works with genome assembly graphs and raw data from paired, linked and long reads. It includes a simple deBruijn graph module, and can import graphs using the graphical fragment assembly (GFA) format. It also maps raw reads onto graphs, and provides a Python application programming interface (API) to navigate the graph, access the mapped and raw data and perform interactive or scripted analyses. Its complete workspace can be dumped to and loaded from disk, decoupling mapping from analysis and supporting multi-stage pipelines. We present the design and implementation of the framework, and example analyses scaffolding a short read graph with long reads, and navigating paths in a heterozygous graph for a simulated parent-offspring trio dataset. SDG is freely available under the MIT license at https://github.com/bioinfologics/sdg
Primula vulgaris (primrose) exhibits heterostyly: plants produce self-incompatible pin- or thrum-form flowers, with anthers and stigma at reciprocal heights. Darwin concluded that this arrangement promotes insect-mediated cross-pollination; later studies revealed control by a cluster of genes, or supergene, known as the S (Style length) locus. The P. vulgaris S locus is absent from pin plants and hemizygous in thrum plants (thrum-specific); mutation of S locus genes produces self-fertile homostyle flowers with anthers and stigma at equal heights. Here, we present a 411 Mb P. vulgaris genome assembly of a homozygous inbred long homostyle, representing ~87% of the genome. We annotate over 24,000 P. vulgaris genes, and reveal more genes up-regulated in thrum than pin flowers. We show reduced genomic read coverage across the S locus in other Primula species, including P. veris, where we define the conserved structure and expression of the S locus genes in thrum. Further analysis reveals the S locus has elevated repeat content (64%) compared to the wider genome (37%). Our studies suggest conservation of S locus genetic architecture in Primula, and provide a platform for identification and evolutionary analysis of the S locus and downstream targets that regulate heterostyly in diverse heterostylous species.
Accelerating international trade and climate change make pathogen spread an increasing concern. Hymenoscyphus fraxineus , the causal agent of ash dieback, is a fungal pathogen that has been moving across continents and hosts from Asian to European ash. Most European common ash trees ( Fraxinus excelsior ) are highly susceptible to H. fraxineus , although a minority (~5%) have partial resistance to dieback. Here, we assemble and annotate a H. fraxineus draft genome, which approaches chromosome scale. Pathogen genetic diversity across Europe and in Japan, reveals a strong bottleneck in Europe, though a signal of adaptive diversity remains in key host interaction genes. We find that the European population was founded by two divergent haploid individuals. Divergence between these haplotypes represents the ancestral polymorphism within a large source population. Subsequent introduction from this source would greatly increase adaptive potential of the pathogen. Thus, further introgression of H. fraxineus into Europe represents a potential threat and Europe-wide biological security measures are needed to manage this disease.
Accelerating international trade and climate change make pathogen spread an increasing concern. Hymenoscyphus fraxineus, the causal agent of ash dieback is one such pathogen, moving across continents and hosts from Asian to European ash. Most European common ash (Fraxinus excelsior) trees are highly susceptible to H. fraxineus although a small minority (~5%) evidently have partial resistance to dieback. We have assembled and annotated a draft of the H. fraxineus genome which approaches chromosome scale. Pathogen genetic diversity across Europe, and in Japan, reveals a tight bottleneck into Europe, though a signal of adaptive diversity remains in key host interaction genes (effectors). We find that the European population was founded by two divergent haploid individuals. Divergence between these haplotypes represents the 'shadow' of a large source population and subsequent introduction would greatly increase adaptive potential and the pathogen's threat. Thus, EU wide biological security measures remain an important part of the strategy to manage this disease.
Bioinformatic analyses and tools make extensive use of k-mers (fixed contiguous strings of k nucleotides) as an informational unit. K-mer analyses are both useful and fast, but are strongly affected by single nucleotide polymorphisms or sequencing errors, effectively hindering direct-analyses of whole regions and decreasing their usability between evolutionary distant samples. Q-grams or spaced seeds, subsequences generated with a pattern of used-and-skipped nucleotides, overcome many of these limitations but introduce larger complexity which hinders their wider adoption. We introduce a concept of skip-mers, a cyclic pattern of used-and-skipped positions of k nucleotides spanning a region of size S ≥ k , and show how analyses are improved by using this simple subset of q-grams as a replacement for k-mers. The entropy of skip-mers increases with the larger span, capturing information from more distant positions and increasing the specificity, and uniqueness, of larger span skip-mers within a genome. In addition, skip-mers constructed in cycles of 1 or 2 nucleotides in every 3 (or a multiple of 3) lead to increased sensitivity in the coding regions of genes, by grouping together the more conserved nucleotides of the protein-coding regions. We implemented a set of tools to count and intersect skip-mers between different datasets, a simple task given that the properties of skip-mers make them a direct substitute for k-mers. We used these tools to show how skip-mers have advantages over k-mers in terms of entropy and increased sensitivity to detect conserved coding sequence, allowing better identification of genic matches between evolutionarily distant species. We then show benefits for multi-genome analyses provided by increased and better correlated coverage of conserved skip-mers across multiple samples. Software availability the skm-tools implementing the methods described in this manuscript are available under MIT license at http://github.com/bioinfologics/skm-tools/
ABSTRACT Motivation De novo assembly of whole genome shotgun (WGS) next-generation sequencing (NGS) data benefits from high-quality input with high coverage. However, in practice, determining the quality and quantity of useful reads quickly and in a reference-free manner is not trivial. Gaining a better understanding of the WGS data, and how that data is utilised by assemblers, provides useful insights that can inform the assembly process and result in better assemblies. Results We present the K-mer Analysis Toolkit (KAT): a multi-purpose software toolkit for reference-free quality control (QC) of WGS reads and de novo genome assemblies, primarily via their k-mer frequencies and GC composition. KAT enables users to assess levels of errors, bias and contamination at various stages of the assembly process. In this paper we highlight KAT’s ability to provide valuable insights into assembly composition and quality of genome assemblies through pairwise comparison of k-mers present in both input reads and the assemblies. Availability KAT is available under the GPLv3 license at: https://github.com/TGAC/KAT . Contact bernardo.clavijo@earlham.ac.uk Supplementary Information Supplementary Information (SI) is available at Bioinformatics online. In addition, the software documentation is available online at: http://kat.readthedocs.io/en/latest/ .
Advances in genome sequencing and assembly technologies are generating many high-quality genome sequences, but assemblies of large, repeat-rich polyploid genomes, such as that of bread wheat, remain fragmented and incomplete. We have generated a new wheat whole-genome shotgun sequence assembly using a combination of optimized data types and an assembly algorithm designed to deal with large and complex genomes. The new assembly represents >78% of the genome with a scaffold N50 of 88.8 kb that has a high fidelity to the input data. Our new annotation combines strand-specific Illumina RNA-seq and Pacific Biosciences (PacBio) full-length cDNAs to identify 104,091 high-confidence protein-coding genes and 10,156 noncoding RNA genes. We confirmed three known and identified one novel genome rearrangements. Our approach enables the rapid and scalable assembly of wheat genomes, the identification of structural variants, and the definition of complete gene models, all powerful resources for trait analysis and breeding of this key global crop.
Producing high-quality whole-genome shotgun de novo assemblies from plant and animal species with large and complex genomes using low-cost short read sequencing technologies remains a challenge. But when the right sequencing data, with appropriate quality control, is assembled using approaches focused on robustness of the process rather than maximization of a single metric such as the usual contiguity estimators, good quality assemblies with informative value for comparative analyses can be produced. Here we present a complete method described from data generation and qc all the way up to scaffold of complex genomes using Illumina short reads and its application to data from plants and human datasets. We show how to use the w2rap pipeline following a metric-guided approach to produce cost-effective assemblies. The assemblies are highly accurate, provide good coverage of the genome and show good short range contiguity. Our pipeline has already enabled the rapid, cost-effective generation of de novo genome assemblies from large, polyploid crop species with a focus on comparative genomics. Availability w2rap is available under MIT license, with some subcomponents under GPL-licenses. A ready-to-run docker with all software pre-requisites and example data is also available. http://github.com/bioinfologics/w2rap http://github.com/bioinfologics/w2rap-contigger
Darwin's studies on heterostyly in Primula described two floral morphs, pin and thrum, with reciprocal anther and stigma heights that promote insect-mediated cross-pollination. This key innovation evolved independently in several angiosperm families. Subsequent studies on heterostyly in Primula contributed to the foundation of modern genetic theory and the neo-Darwinian synthesis. The established genetic model for Primula heterostyly involves a diallelic S locus comprising several genes, with rare recombination events that result in self-fertile homostyle flowers with anthers and stigma at the same height. Here we reveal the S locus supergene as a tightly linked cluster of thrum-specific genes that are absent in pins. We show that thrums are hemizygous not heterozygous for the S locus, which suggests that homostyles do not arise by recombination between S locus haplotypes as previously proposed. Duplication of a floral homeotic gene 51.7 million years (Myr) ago, followed by its neofunctionalization, created the current S locus assemblage which led to floral heteromorphy in Primula. Our findings provide new insights into the structure, function and evolution of this archetypal supergene.
Summary Heteromorphic flower development in Primula is controlled by the S locus. The S locus genes, which control anther position, pistil length and pollen size in pin and thrum flowers, have not yet been characterized. We have integrated S‐linked genes, marker sequences and mutant phenotypes to create a map of the P. vulgaris S locus region that will facilitate the identification of key S locus genes. We have generated, sequenced and annotated BAC sequences spanning the S locus, and identified its chromosomal location. We have employed a combination of classical genetics and three‐point crosses with molecular genetic analysis of recombinants to generate the map. We have characterized this region by Illumina sequencing and bioinformatic analysis, together with chromosome in situ hybridization. We present an integrated genetic and physical map across the P. vulgaris S locus flanked by phenotypic and DNA sequence markers. BAC contigs encompass a 1.5‐Mb genomic region with 1 Mb of sequence containing 82 S‐linked genes anchored to overlapping BACs. The S locus is located close to the centromere of the largest metacentric chromosome pair. These data will facilitate the identification of the genes that orchestrate heterostyly in Primula and enable evolutionary analyses of the S locus.
Duplication of genes is thought to facilitate increasing organismal complexity. Duplicated genes may be retained in the genome if one of the duplicates acquires a novel function, or the functions of the ancestral gene are subdivided between the duplicates. A likely process for the retention of duplicates is the loss or gain of regulatory elements in the promoters of duplicated genes.The objective of this study was to explore the evolution of gene promoters using the multigene family of fatty acid‐binding proteins (fabp) in zebrafish. Previous studies had implicated the peroxisome proliferator‐activated receptors (PPARs) in the regulation of some fabp genes. The promoters of the zebrafish fabp1a, fabp1b.1 and fabp1b.2 genes, duplicated by a whole genome duplication (fabp1a and ancestral fabp1b) followed by a tandem gene duplication (fabp1b.1 and fabp1b.2), were cloned into firefly luciferase reporter plasmids, transfected into HEK cells, and tested in the presence of PPARα‐ and PPARγ‐selective agonists. Expression of these genes was further analyzed by qRT‐PCR in explant cultures of zebrafish liver and intestine treated with PPAR agonists.fabp1a expression was selectively increased by PPARα‐agonism in HEK cells and intestine. fabp1b.1 expression was selectively increased by PPARγ‐agonism in HEK cells and liver. fabp1b.2 promoter activity was not induced by PPAR modulation. The differential PPAR regulation of the duplicated fabp1 genes provide evidence of divergence of regulatory elements in these gene promoters, which may account for the retention of the duplicated fabp1 genes in the zebrafish genome.Funding from NSERC, CIHR, and Dalhousie University.
Summary In Primula vulgaris outcrossing is promoted through reciprocal herkogamy with insect‐mediated cross‐pollination between pin and thrum form flowers. Development of heteromorphic flowers is coordinated by genes at the S locus. To underpin construction of a genetic map facilitating isolation of these S locus genes, we have characterised Oakleaf, a novel S locus‐linked mutant phenotype. We combine phenotypic observation of flower and leaf development, with classical genetic analysis and next‐generation sequencing to address the molecular basis of Oakleaf. Oakleaf is a dominant mutation that affects both leaf and flower development; plants produce distinctive lobed leaves, with occasional ectopic meristems on the veins. This phenotype is reminiscent of overexpression of Class I KNOX‐homeodomain transcription factors. We describe the structure and expression of all eight P. vulgaris PvKNOX genes in both wild‐type and Oakleaf plants, and present comparative transcriptome analysis of leaves and flowers from Oakleaf and wild‐type plants. Oakleaf provides a new phenotypic marker for genetic analysis of the Primula S locus. We show that none of the Class I PvKNOX genes are strongly upregulated in Oakleaf leaves and flowers, and identify cohorts of 507 upregulated and 314 downregulated genes in the Oakleaf mutant.
An ordered draft sequence of the 17-gigabase hexaploid bread wheat ( Triticum aestivum ) genome has been produced by sequencing isolated chromosome arms. We have annotated 124,201 gene loci distributed nearly evenly across the homeologous chromosomes and subgenomes. Comparative gene analysis of wheat subgenomes and extant diploid and tetraploid wheat relatives showed that high sequence similarity and structural conservation are retained, with limited gene loss, after polyploidization. However, across the genomes there was evidence of dynamic gene gain, loss, and duplication since the divergence of the wheat lineages. A high degree of transcriptional autonomy and no global dominance was found for the subgenomes. These insights into the genome biology of a polyploid crop provide a springboard for faster gene isolation, rapid genetic marker development, and precise breeding to meet the needs of increasing food demand worldwide.