Lignocellulose biomass is one of the most abundant resources for sustainable biofuels. However, scaling up the biomass-to-biofuels conversion process for widespread usage is still pending. Bottlenecks during the process of enzymatic hydrolysis are the high cost of enzymes and the labor-intensive need for substrate-dependent enzyme mixtures. Current research efforts are therefore targeted at searching for or engineering lignocellulolytic enzymes of high efficiency. One way is to engineer multi-enzyme complexes that mimic the bacterial cellulosomal system, known to increase degradation efficiency up to 50-fold when compared to freely-secreted enzymes. However, these designer cellulosomes are instable and less efficient than wild type cellulosomes. Fungi cellulosomes discovered in recent years have significant differences from bacterial counterparts and hold great potential for industrial applications, both as designer cellulosomes and as additions to the enzymatic repertoire. Up to date, they are only found in a few anaerobic fungi. In this review, we extensively compared the degradation mechanisms in bacteria and fungi, and highlighted the essential gaps in applying these mechanisms in industrial applications. To better understand cellulosomes in microorganisms, we examined their sequences in 66,252 bacterial species and 823 fungal species and identified several bacterial species that are potentially cellulosome-producing. These findings act as a valuable resource in the biomass community for further proteomic and genetic sequence analysis. We also collated the current strategies of bioengineering lignocellulose degradation to suggest concepts that could be favorable for industrial usage. ### Competing Interest Statement The authors have declared no competing interest.
Tweetable abstract Monitoring changes in methylation heterogeneity can be powerful in detecting disease progression early. This editorial highlights the importance of profiling methylation heterogeneity and identifies existing measures and research gaps.
KEY MESSAGE:Transcriptomic and epigenomic profiling of gene expression and small RNAs during seed and seedling development reveals expression and methylation dominance levels with implications on early stage heterosis in oilseed rape. The enhanced performance of hybrids through heterosis remains a key aspect in plant breeding; however, the underlying mechanisms are still not fully elucidated. To investigate the potential role of transcriptomic and epigenomic patterns in early expression of hybrid vigor, we investigated gene expression, small RNA abundance and genome-wide methylation in hybrids from two distant Brassica napus ecotypes during seed and seedling developmental stages using next-generation sequencing. A total of 31117, 344, 36229 and 7399 differentially expressed genes, microRNAs, small interfering RNAs and differentially methylated regions were identified, respectively. Approximately 70% of the differentially expressed or methylated features displayed parental dominance levels where the hybrid followed the same patterns as the parents. Via gene ontology enrichment and microRNA-target association analyses during seed development, we found copies of reproductive, developmental and meiotic genes with transgressive and paternal dominance patterns. Interestingly, maternal dominance was more prominent in hypermethylated and downregulated features during seed formation, contrasting to the general maternal gamete demethylation reported during gametogenesis in angiosperms. Associations between methylation and gene expression allowed identification of putative epialleles with diverse pivotal biological functions during seed formation. Furthermore, most differentially methylated regions, differentially expressed siRNAs and transposable elements were in regions that flanked genes without differential expression. This suggests that differential expression and methylation of epigenomic features may help maintain expression of pivotal genes in a hybrid context. Differential expression and methylation patterns during seed formation in an F1 hybrid provide novel insights into genes and mechanisms with potential roles in early heterosis.
The mechanism of the addition of a methyl group to cytosine has been identified as one of several heritable epigenetic mechanisms. In plants, DNA methylation is involved in mediating response to stress, plant development, polyploidy, and domestication through regulation of gene expression. The correlation of epigenetic variation to phenotypic traits expands our understanding toward plant evolution, and provides new source for targeted manipulation in crop improvement. To address the increasing interest to map methylation landscape in plant species, this chapter describes methods to analyze bisulfite sequencing data and identify epigenetic variation between samples. We also detailed guidelines to highlight possible optimizations, as well as ways to tailor parameters according to data and biological variability.
In a cross between two homozygous Brassica napus plants of synthetic and natural origin, we demonstrate that novel structural genome variants from the synthetic parent cause immediate genome diversification among F1 offspring. Long read sequencing in twelve F1 sister plants revealed five large-scale structural rearrangements where both parents carried different homozygous alleles but the heterozygous F1 genomes were not identical heterozygotes as expected. Such spontaneous rearrangements were part of homoeologous exchanges or segmental deletions and were identified in different, individual F1 plants. The variants caused deletions, gene copy-number variations, diverging methylation patterns and other structural changes in large numbers of genes and may have been causal for unexpected phenotypic variation between individual F1 sister plants, for example strong divergence of plant height and leaf area. This example supports the hypothesis that spontaneous de novo structural rearrangements after de novo polyploidization can rapidly overcome intense allopolyploidization bottlenecks to re-expand crops genetic diversity for ecogeographical expansion and human selection. The findings imply that natural genome restructuring in allopolyploid plants from interspecific hybridization, a common approach in plant breeding, can have a considerably more drastic impact on genetic diversity in agricultural ecosystems than extremely precise, biotechnological genome modifications.
SummaryGenome structural variation (SV) contributes strongly to trait variation in eukaryotic species and may have an even higher functional significance than single‐nucleotide polymorphism (SNP). In recent years, there have been a number of studies associating large chromosomal scale SV ranging from hundreds of kilobases all the way up to a few megabases to key agronomic traits in plant genomes. However, there have been little or no efforts towards cataloguing small‐ (30–10 000 bp) to mid‐scale (10 000–30 000 bp) SV and their impact on evolution and adaptation‐related traits in plants. This might be attributed to complex and highly duplicated nature of plant genomes, which makes them difficult to assess using high‐throughput genome screening methods. Here, we describe how long‐read sequencing technologies can overcome this problem, revealing a surprisingly high level of widespread, small‐ to mid‐scale SV in a major allopolyploid crop species, Brassica napus. We found that up to 10% of all genes were affected by small‐ to mid‐scale SV events. Nearly half of these SV events ranged between 100 bp and 1000 bp, which makes them challenging to detect using short‐read Illumina sequencing. Examples demonstrating the contribution of such SV towards eco‐geographical adaptation and disease resistance in oilseed rape suggest that revisiting complex plant genomes using medium‐coverage long‐read sequencing might reveal unexpected levels of functional gene variation, with major implications for trait regulation and crop improvement.
Key message A novel structural variant was discovered in the FLOWERING LOCUS T orthologue BnaFT.A02 by long-read sequencing. Nested association mapping in an elite winter oilseed rape population revealed that this 288 bp deletion associates with early flowering, putatively by modification of binding-sites for important flowering regulation genes. Abstract Perfect timing of flowering is crucial for optimal pollination and high seed yield. Extensive previous studies of flowering behavior in Brassica napus (canola, rapeseed) identified mutations in key flowering regulators which differentiate winter, semi-winter and spring ecotypes. However, because these are generally fixed in locally adapted genotypes, they have only limited relevance for fine adjustment of flowering time in elite cultivar gene pools. In crosses between ecotypes, the ecotype-specific major-effect mutations mask minor-effect loci of interest for breeding. Here, we investigated flowering time in a multiparental mapping population derived from seven elite winter oilseed rape cultivars which are fixed for major-effect mutations separating winter-type rapeseed from other ecotypes. Association mapping revealed eight genomic regions on chromosomes A02, C02 and C03 associating with fine modulation of flowering time. Long-read genomic resequencing of the seven parental lines identified seven structural variants coinciding with candidate genes for flowering time within chromosome regions associated with flowering time. Segregation patterns for these variants in the elite multiparental population and a diversity set of winter types using locus-specific assays revealed significant associations with flowering time for three deletions on chromosome A02. One of these was a previously undescribed 288 bp deletion within the second intron of FLOWERING LOCUS T on chromosome A02, emphasizing the advantage of long-read sequencing for detection of structural variants in this size range. Detailed analysis revealed the impact of this specific deletion on flowering-time modulation under extreme environments and varying day lengths in elite, winter-type oilseed rape.
Plant genomes demonstrate significant presence/absence variation (PAV) within a species; however, the factors that lead to this variation have not been studied systematically in Brassica across diploids and polyploids. Here, we developed pangenomes of polyploid Brassica napus and its two diploid progenitor genomes B. rapa and B. oleracea to infer how PAV may differ between diploids and polyploids. Modelling of gene loss suggests that loss propensity is primarily associated with transposable elements in the diploids while in B. napus, gene loss propensity is associated with homoeologous recombination. We use these results to gain insights into the different causes of gene loss, both in diploids and following polyploidization, and pave the way for the application of machine learning methods to understanding the underlying biological and physical causes of gene presence/absence.
Blackleg is one of the major fungal diseases in oilseed rape/canola worldwide. Most commercial cultivars carry R gene-mediated qualitative resistances that confer a high level of race-specific protection against Leptosphaeria maculans, the causal fungus of blackleg disease. However, monogenic resistances of this kind can potentially be rapidly overcome by mutations in the pathogen’s avirulence genes. To counteract pathogen adaptation in this evolutionary arms race, there is a tremendous demand for quantitative background resistance to enhance durability and efficacy of blackleg resistance in oilseed rape. In this study, we characterized genomic regions contributing to quantitative L. maculans resistance by genome-wide association studies in a multiparental mapping population derived from six parental elite varieties exhibiting quantitative resistance, which were all crossed to one common susceptible parental elite variety. Resistance was screened using a fungal isolate with no corresponding avirulence (AvrLm) to major R genes present in the parents of the mapping population. Genome-wide association studies revealed eight significantly associated quantitative trait loci (QTL) on chromosomes A07 and A09, with small effects explaining 3–6% of the phenotypic variance. Unexpectedly, the qualitative blackleg resistance gene Rlm9 was found to be located within a resistance-associated haploblock on chromosome A07. Furthermore, long-range sequence data spanning this haploblock revealed high levels of single-nucleotide and structural variants within the Rlm9 coding sequence among the parents of the mapping population. The results suggest that novel variants of Rlm9 could play a previously unknown role in expression of quantitative disease resistance in oilseed rape.
The cultivated Brassica species include numerous vegetable and oil crops of global importance. Three genomes (designated A, B and C) share mesohexapolyploid ancestry and occur both singly and in each pairwise combination to define the Brassica species. With organizational errors (such as misplaced genome segments) corrected, we showed that the fundamental structure of each of the genomes is the same, irrespective of the species in which it occurs. This enabled us to clarify genome evolutionary pathways, including updating the Ancestral Crucifer Karyotype (ACK) block organization and providing support for the Brassica mesohexaploidy having occurred via a two-step process. We then constructed genus-wide pan-genomes, drawing from genes present in any species in which the respective genome occurs, which enabled us to provide a global gene nomenclature system for the cultivated Brassica species and develop a methodology to cost-effectively elucidate the genomic impacts of alien introgressions. Our advances not only underpin knowledge-based approaches to the more efficient breeding of Brassica crops but also provide an exemplar for the study of other polyploids.
Recent advances in long-read sequencing have the potential to produce more complete genome assemblies using sequence reads which can span repetitive regions. However, overlap based assembly methods routinely used for this data require significant computing time and resources. Here, we have developed RefKA, a reference-based approach for long read genome assembly. This approach relies on breaking up a closely related reference genome into bins, aligning k -mers unique to each bin with PacBio reads, and then assembling each bin in parallel followed by a final bin-stitching step. During benchmarking, we assembled the wheat Chinese Spring (CS) genome using publicly available PacBio reads in parallel in 168 wall hours on a 250 CPU system. The maximum RAM used was 300 Gb and the computing time was 42,000 CPU hours. The approach opens applications for the assembly of other large and complex genomes with much-reduced computing requirements. The RefKA pipeline is available at ### Competing Interest Statement The authors have declared no competing interest.
Rapeseed (Brassica napus), the second most important oilseed crop globally, originated from an interspecific hybridization between B. rapa and B. oleracea. After this genome collision, B. napus underwent extensive genome restructuring, via homoeologous chromosome exchanges, resulting in widespread segmental deletions and duplications. Illicit pairing among genetically similar homoeologous chromosomes during meiosis is common in recent allopolyploids like B. napus, and post-polyploidization restructuring compounds the difficulties of assembling a complex polyploid plant genome. Specifically, genomic rearrangements between highly similar chromosomes are challenging to detect due to the limitation of sequencing read length and ambiguous alignment of reads. Recent advances in long read sequencing technologies provide promising new opportunities to unravel the genome complexities of B. napus by encompassing breakpoints of genomic rearrangements with high specificity. Moreover, recent evidence revealed ongoing genomic exchanges in natural B. napus, highlighting the need for multiple reference genomes to capture structural variants between accessions. Here we report the first long-read genome assembly of a winter B. napus cultivar. We sequenced the German winter oilseed rape accession ‘Express 617’ using 54.5x of long reads. Short reads, linked reads, optical map data and high-density genetic maps were used to further correct and scaffold the assembly to form pseudochromosomes. The assembled Express 617 genome provides another valuable resource for Brassica genomics in understanding the genetic consequences of polyploidization, crop domestication, and breeding of recently-formed crop species.
We report the first annotated chromosome-level reference genome assembly for pea, Gregor Mendel's original genetic model. Phylogenetics and paleogenomics show genomic rearrangements across legumes and suggest a major role for repetitive elements in pea genome evolution. Compared to other sequenced Leguminosae genomes, the pea genome shows intense gene dynamics, most likely associated with genome size expansion when the Fabeae diverged from its sister tribes. During Pisum evolution, translocation and transposition differentially occurred across lineages. This reference sequence will accelerate our understanding of the molecular basis of agronomically important traits and support crop improvement.
Seagrasses are marine angiosperms that live fully submerged in the sea. They evolved from land plant ancestors, with multiple species representing at least three independent return-to-the-sea events. This raises the question of whether these marine angiosperms followed the same adaptation pathway to allow them to live and reproduce under the hostile marine conditions. To compare the basis of marine adaptation between seagrass lineages, we generated genomic data for Halophila ovalis and compared this with recently published genomes for two members of Zosteraceae, as well as genomes of five non-marine plant species (Arabidopsis, Oryza sativa, Phoenix dactylifera, Musa acuminata, and Spirodela polyrhiza). Halophila and Zosteraceae represent two independent seagrass lineages separated by around 30 million years. Genes that were lost or conserved in both lineages were identified. All three species lost genes associated with ethylene and terpenoid biosynthesis, and retained genes related to salinity adaptation, such as those for osmoregulation. In contrast, the loss of the NADH dehydrogenase-like complex is unique to H. ovalis. Through comparison of two independent return-to-the-sea events, this study further describes marine adaptation characteristics common to seagrass families, identifies species-specific gene loss, and provides molecular evidence for convergent evolution in seagrass lineages.
Individual cells in an organism are variable, which strongly impacts cellular processes. Advances in sequencing technologies have enabled single-cell genomic analysis to become widespread, addressing shortcomings of analyses conducted on populations of bulk cells. While the field of single-cell plant genomics is in its infancy, there is great potential to gain insights into cell lineage and functional cell types to help understand complex cellular interactions in plants. In this review, we discuss current approaches for single-cell plant genomic analysis, with a focus on single-cell isolation, DNA amplification, next-generation sequencing, and bioinformatics analysis. We outline the technical challenges of analysing material from a single plant cell, and then examine applications of single-cell genomics and the integration of this approach with genome editing. Finally, we indicate future directions we expect in the rapidly developing field of plant single-cell genomic analysis.
There is an increasing understanding that variation in gene presence-absence plays an important role in the heritability of agronomic traits; however, there have been relatively few studies on variation in gene presence-absence in crop species. Hexaploid wheat is one of the most important food crops in the world and intensive breeding has reduced the genetic diversity of elite cultivars. Major efforts have produced draft genome assemblies for the cultivar Chinese Spring, but it is unknown how well this represents the genome diversity found in current modern elite cultivars. In this study we build an improved reference for Chinese Spring and explore gene diversity across 18 wheat cultivars. We predict a pangenome size of 140 500 ± 102 genes, a core genome of 81 070 ± 1631 genes and an average of 128 656 genes in each cultivar. Functional annotation of the variable gene set suggests that it is enriched for genes that may be associated with important agronomic traits. In addition to variation in gene presence, more than 36 million intervarietal single nucleotide polymorphisms were identified across the pangenome. This study of the wheat pangenome provides insight into genome diversity in elite wheat as a basis for genomics-based improvement of this important crop. A wheat pangenome, GBrowse, is available at http://appliedbioinformatics.com.au/cgi-bin/gb2/gbrowse/WheatPan/, and data are available to download from http://wheatgenome.info/wheat_genome_databases.php.
SummaryAs an increasing number of plant genome sequences become available, it is clear that gene content varies between individuals, and the challenge arises to predict the gene content of a species. However, genome comparison is often confounded by variation in assembly and annotation. Differentiating between true gene absence and variation in assembly or annotation is essential for the accurate identification of conserved and variable genes in a species. Here, we present the de novo assembly of the B. napus cultivar Tapidor and comparison with an improved assembly of the Brassica napus cultivar Darmor‐bzh. Both cultivars were annotated using the same method to allow comparison of gene content. We identified genes unique to each cultivar and differentiate these from artefacts due to variation in the assembly and annotation. We demonstrate that using a common annotation pipeline can result in different gene predictions, even for closely related cultivars, and repeat regions which collapse during assembly impact whole genome comparison. After accounting for differences in assembly and annotation, we demonstrate that the genome of Darmor‐bzh contains a greater number of genes than the genome of Tapidor. Our results are the first step towards comparison of the true differences between B. napus genomes and highlight the potential sources of error in future production of a B. napus pangenome.
We developed runBNG, an open-source software package which wraps BioNano genomic analysis tools into a single script that can be run on the command line. runBNG can complete analyses, including quality control of single molecule maps, optical map de novo assembly, comparisons between different optical maps, super-scaffolding and structural variation detection. Compared to existing software BioNano IrysView and the KSU scripts, the major advantages of runBNG are that the whole pipeline runs on one single platform and it has a high customizability.
Seagrass meadows are disappearing at alarming rates as a result of increasing coastal development and climate change. The emergence of omics and molecular profiling techniques in seagrass research is timely, providing a new opportunity to address such global issues. Whilst these applications have transformed terrestrial plant research, they have only emerged in seagrass research within the past decade; In this time frame we have observed a significant increase in the number of publications in this nascent field, and as of this year the first genome of a seagrass species has been sequenced. In this review, we focus on the development of omics and molecular profiling and the utilization of molecular markers in the field of seagrass biology. We highlight the advances, merits and pitfalls associated with such technology, and importantly we identify and address the knowledge gaps, which to this day prevent us from understanding seagrasses in a holistic manner. By utilizing the powers of omics and molecular profiling technologies in integrated strategies, we will gain a better understanding of how these unique plants function at the molecular level and how they respond to on-going disturbance and climate change events.
Seagrasses are marine angiosperms that evolved from land plants but returned to the sea around 140 million years ago during the early evolution of monocotyledonous plants. They successfully adapted to abiotic stresses associated with growth in the marine environment, and today, seagrasses are distributed in coastal waters worldwide. Seagrass meadows are an important oceanic carbon sink and provide food and breeding grounds for diverse marine species. Here, we report the assembly and characterization of the Zostera muelleri genome, a southern hemisphere temperate species. Multiple genes were lost or modified in Z. muelleri compared with terrestrial or floating aquatic plants that are associated with their adaptation to life in the ocean. These include genes for hormone biosynthesis and signaling and cell wall catabolism. There is evidence of whole-genome duplication in Z. muelleri; however, an ancient pan-commelinid duplication event is absent, highlighting the early divergence of this species from the main monocot lineages.