Abstract Camelina sativa , an oilseed crop of the Brassicaceae family, has close relatives that vary in ploidy levels, providing a unique platform for studying plant genome evolution. Here, we report an improved assembly of the widely used C. sativa reference DH55 and three additional genome assemblies of Camelina microcarpa : one tetraploid, and two hexaploids with divergent chromosome numbers, Type 1 (2n = 40) and Type 2 (2n = 38). We uncover the fourth subgenome of the Camelina genus that of C. microcarpa Type 2, which shows numerous unique chromosomal rearrangements differentiating it from other characterized Camelina subgenomes. In this recently formed species, the second subgenome displays gene expression dominance, contrary to expectations from the two-step evolutionary process invoked in the generation of related Brassicaceae species. The observed gene dominance is negatively correlated with inter-subgenome chromatin interaction frequencies, suggesting chromosome conformation and proximity in the nucleus contribute to this mechanism of genome evolution.
Camelina sativa is an oilseed of the Brassicaceae , whose close relatives vary in ploidy number, providing a novel platform for studying plant genome evolution. The availability of diploid, tetraploid and hexaploid species of Camelina allow the evolutionary trajectory and fate of duplicated genes in the neopolyploid Camelina species to be elucidated. Here we report an improved assembly of the widely used C. sativa reference DH55 and three new genome assemblies of Camelina microcarpa ; one tetraploid CN119243 (2n = 26), and two hexaploids with divergent chromosome numbers, Type 1 - CN119205 (2n=40) and Type 2 - CN120025 (2n=38). The tetraploid represents the first step in the evolutionary path to form C. sativa , while the hexaploids suggest three divergent lineages in the formation of higher ploidy Camelina species. The previously uncharacterized fourth subgenome found in C. microcarpa Type 2, although showing some homology to the C. sativa diploid progenitor genome, C. neglecta , showed numerous unique chromosomal rearrangements differentiating it from other subgenomes present in known Camelina species. Although this species was recently formed, the second subgenome showed gene expression dominance, which was in contrast to both 2n=40 Camelina species where the third subgenome was dominant. The expression dominance in Type 2 C. microcarpa contradicted the accepted two-step evolutionary process which led to the generation of related Brassicaceae species. However, the observed genome dominance in all Camelina species was negatively correlated with inter-subgenome chromatin interaction frequencies, suggesting that chromosome confirmation and proximity in the nucleus contributes to this mechanism of genome evolution. Despite the differences in genome structure, successful inter-specific hybridization provided evidence of chromosomal exchange between the divergent third sub-genomes of C. sativa and C. microcarpa Type 2, opening up a novel avenue to new diversity in the established oilseed. Key points ### Competing Interest Statement The authors have declared no competing interest. All sequence data is deposited at the EBI-ENA under accession: PRJEB96055. The genome assemblies are available at
Genomic prediction (GP) significantly enhances genetic gain by improving selection efficiency and shortening crop breeding cycles. Using a nested association mapping population, a set of diverse scenarios were assessed to evaluate GP for important agronomic traits in Brassica napus, including plant height, days to flowering, 1000-kernel weight, and yield. GP accuracy was examined on each trait by employing eight different models, eight marker sets, varying population sizes and marker densities, and incorporating trait-associated markers identified through genome-wide association study analysis. Eight models, including linear and semi-parametric approaches, were tested. The choice of model minimally impacted GP accuracy across traits. Employing a training population of 1500 lines or more resulted in increased prediction accuracies. Inclusion of single nucleotide absence polymorphism markers with single-nucleotide polymorphism markers significantly improved prediction accuracy, with gains of up to 15%. The study provided estimates of GPs for major agronomic traits through varied prediction scenarios, shedding light on achievable genetic gains. These insights, coupled with marker application, can advance the breeding cycle acceleration in B. napus.
Seed quality traits of oilseed rape, Brassica napus (B. napus), exhibit quantitative inheritance determined by its genetic makeup and the environment via the mediation of a complex genetic architecture of hundreds to thousands of genes. Thus, instead of single gene analysis, network-based systems genomics and genetics approaches that combine genotype, phenotype, and molecular phenotypes offer a promising alternative to uncover this complex genetic architecture. In the current study, systems genetics approaches were used to explore the genetic regulation of lignin traits in B. napus seeds. Four QTL (qLignin_A09_1, qLignin_A09_2, qLignin_A09_3, and qLignin_C08) distributed on two chromosomes were identified for lignin content. The qLignin_A09_2 and qLignin_C08 loci were homologous QTL from the A and C subgenomes, respectively. Genome-wide gene regulatory network analysis identified eighty-three subnetworks (or modules); and three modules with 910 genes in total, were associated with lignin content, which was confirmed by network QTL analysis. eQTL (expression quantitative trait loci) analysis revealed four cis-eQTL genes including lignin and flavonoid pathway genes, cinnamoyl-CoA-reductase (CCR1), and TRANSPARENT TESTA genes TT4, TT6, TT8, as causal genes. The findings validated the power of systems genetics to identify causal regulatory networks and genes underlying complex traits. Moreover, this information may enable the research community to explore new breeding strategies, such as network selection or gene engineering, to rewire networks to develop climate resilience crops with better seed quality.
Camelina neglecta is a diploid species from the genus Camelina, which includes the versatile oilseed Camelina sativa. These species are closely related to Arabidopsis thaliana and the economically important Brassica crop species, making this genus a useful platform to dissect traits of agronomic importance while providing a tool to study the evolution of polyploids. A highly contiguous chromosome-level genome sequence of C. neglecta with an N50 size of 29.1 Mb was generated utilizing Pacific Biosciences (PacBio, Menlo Park, CA) long-read sequencing followed by chromosome conformation phasing. Comparison of the genome with that of C. sativa shows remarkable coincidence with subgenome 1 of the hexaploid, with only one major chromosomal rearrangement separating the two. Synonymous substitution rate analysis of the predicted 34 061 genes suggested subgenome 1 of C. sativa directly descended from C. neglecta around 1.2 mya. Higher functional divergence of genes in the hexaploid as evidenced by the greater number of unique orthogroups, and differential composition of resistant gene analogs, might suggest an immediate adaptation strategy after genome merger. The absence of genome bias in gene fractionation among the subgenomes of C. sativa in comparison with C. neglecta, and the complete lack of fractionation of meiosis-specific genes attests to the neopolyploid status of C. sativa. The assembled genome will provide a tool to further study genome evolution processes in the Camelina genus and potentially allow for the identification and exploitation of novel variation for Camelina crop improvement.
AbstractVernalization requirement is an integral component of flowering in winter-type plants. The availability of winter ecotypes amongCamelinaspecies facilitated the mapping of QTL for vernalization requirement inC. sativa. An inter- and intraspecific crossing scheme between relatedCamelinaspecies, where two different sources of the winter-type habit were used, resulted in the development of two segregating populations. Linkage maps generated with sequence-based markers identified three QTL associated with vernalization requirement inC. sativa; two from the inter-specific (chromosomes 13 and 20) and one from the intra-specific cross (chromosome 8). Notably, the three loci were mapped to different homologous regions of the hexaploidC. sativagenome. All three QTL were found in proximity toFLOWERING LOCUS C(FLC), variants of which have been reported to affect the vernalization requirement in plants. Temporal transcriptome analysis for winter-typeCamelina alyssumdemonstrated reduction in expression ofFLCon chromosomes 13 and 20 during cold treatment, which would trigger flowering, sinceFLCwould be expected to suppress floral initiation.FLCon chromosome 8 also showed reduced expression in theC. sativassp.pilosawinter parent upon cold treatment, but was expressed at very high levels across all time points in the spring-typeC. sativa. The chromosome 8 copy carried a deletion in the spring-type line, which could impact its functionality. Contrary to previous reports, all threeFLCloci can contribute to controlling the vernalization response inC. sativaand provide opportunities for manipulating this requirement in the crop.Significance StatementDeveloping winterC. sativagermplasm is an important breeding goal for this alternative oilseed, with application in the food, fuel and bioproduct industries. Studying the genetic architecture of the vernalization response has shown that contrary to previous reports all threeFLCloci inCamelinaspecies could be exploited to manipulate this important trait.
Vernalization requirement is an integral component of flowering in winter-type plants. The availability of winter ecotypes among Camelina species facilitated the mapping of quantitative trait loci (QTL) for vernalization requirement in Camelina sativa. An inter and intraspecific crossing scheme between related Camelina species, where one spring and two different sources of winter-type habit were used, resulted in the development of two segregating populations. Linkage maps generated with sequence-based markers identified three QTLs associated with vernalization requirement in C. sativa; two from the interspecific (chromosomes 13 and 20) and one from the intraspecific cross (chromosome 8). Notably, the three loci were mapped to different homologous regions of the hexaploid C. sativa genome. All three QTLs were found in proximity to Flowering Locus C (FLC), variants of which have been reported to affect the vernalization requirement in plants. Temporal transcriptome analysis for winter-type Camelina alyssum demonstrated reduction in expression of FLC on chromosomes 13 and 20 during cold treatment, which would trigger flowering, since FLC would be expected to suppress floral initiation. FLC on chromosome 8 also showed reduced expression in the C. sativa ssp. pilosa winter parent upon cold treatment, but was expressed at very high levels across all time points in the spring-type C. sativa. The chromosome 8 copy carried a deletion in the spring-type line, which could impact its functionality. Contrary to previous reports, all three FLC loci can contribute to controlling the vernalization response in C. sativa and provide opportunities for manipulating this requirement in the crop.
Phenotyping is considered a significant bottleneck impeding fast and efficient crop improvement. Similar to many crops, Brassica napus, an internationally important oilseed crop, suffers from low genetic diversity, and will require exploitation of diverse genetic resources to develop locally adapted, high yielding and stress resistant cultivars. A pilot study was completed to assess the feasibility of using indoor high-throughput phenotyping (HTP), semi-automated image processing, and machine learning to capture the phenotypic diversity of agronomically important traits in a diverse B. napus breeding population, SKBnNAM, introduced here for the first time. The experiment comprised 50 spring-type B. napus lines, grown and phenotyped in six replicates under two treatment conditions (control and drought) over 38 days in a LemnaTec Scanalyzer 3D facility. Growth traits including plant height, width, projected leaf area, and estimated biovolume were extracted and derived through processing of RGB and NIR images. Anthesis was automatically and accurately scored (97% accuracy) and the number of flowers per plant and day was approximated alongside relevant canopy traits (width, angle). Further, supervised machine learning was used to predict the total number of raceme branches from flower attributes with 91% accuracy (linear regression and Huber regression algorithms) and to identify mild drought stress, a complex trait which typically has to be empirically scored (0.85 area under the receiver operating characteristic curve, random forest classifier algorithm). The study demonstrates the potential of HTP, image processing and computer vision for effective characterization of agronomic trait diversity in B. napus, although limitations of the platform did create significant variation that limited the utility of the data. However, the results underscore the value of machine learning for phenotyping studies, particularly for complex traits such as drought stress resistance.
Turnip mosaic virus (TuMV) induces disease in susceptible hosts, notably impacting cultivation of important crop species of the Brassica genus. Few effective plant viral disease management strategies exist with the majority of current approaches aiming to mitigate the virus indirectly through control of aphid vector species. Multiple sources of genetic resistance to TuMV have been identified previously, although the majority are strain-specific and have not been exploited commercially. Here, two Brassica juncea lines (TWBJ14 and TWBJ20) with resistance against important TuMV isolates (UK 1, vVIR24, CDN 1, and GBR 6) representing the most prevalent pathotypes of TuMV (1, 3, 4, and 4, respectively) and known to overcome other sources of resistance, have been identified and characterized. Genetic inheritance of both resistances was determined to be based on a recessive two-gene model. Using both single nucleotide polymorphism (SNP) array and genotyping by sequencing (GBS) methods, quantitative trait loci (QTL) analyses were performed using first backcross (BC1) genetic mapping populations segregating for TuMV resistance. Pairs of statistically significant TuMV resistance-associated QTLs with additive interactive effects were identified on chromosomes A03 and A06 for both TWBJ14 and TWBJ20 material. Complementation testing between these B. juncea lines indicated that one resistance-linked locus was shared. Following established resistance gene nomenclature for recessive TuMV resistance genes, these new resistance-associated loci have been termed retr04 (chromosome A06, TWBJ14, and TWBJ20), retr05 (A03, TWBJ14), and retr06 (A03, TWBJ20). Genotyping by sequencing data investigated in parallel to robust SNP array data was highly suboptimal, with informative data not established for key BC1 parental samples. This necessitated careful consideration and the development of new methods for processing compromised data. Using reductive screening of potential markers according to allelic variation and the recombination observed across BC1 samples genotyped, compromised GBS data was rendered functional with near-equivalent QTL outputs to the SNP array data. The reductive screening strategy employed here offers an alternative to methods relying upon imputation or artificial correction of genotypic data and may prove effective for similar biparental QTL mapping studies.
Ethiopian mustard (Brassica carinata A. Braun) is an emerging sustainable source of vegetable oil, in particular for the biofuel industry. The present study exploited genome assemblies of the Brassica diploids, Brassica nigra and Brassica oleracea, to discover over 10,000 genome-wide SNPs using genotype by sequencing of 620 B. carinata lines. The analyses revealed a SNP frequency of one every 91.7 kb, a heterozygosity level of 0.30, nucleotide diversity levels of 1.31 × 10−05, and the first five principal components captured only 13% molecular variation, indicating low levels of genetic diversity among the B. carinata collection. Genome bias was observed, with greater SNP density found on the B subgenome. The 620 lines clustered into two distinct sub-populations (SP1 and SP2) with the majority of accessions (88%) clustered in SP1 with those from Ethiopia, the presumed centre of origin. SP2 was distinguished by a collection of breeding lines, implicating targeted selection in creating population structure. Two selective sweep regions on B3 and B8 were detected, which harbour genes involved in fatty acid and aliphatic glucosinolate biosynthesis, respectively. The assessment of genetic diversity, population structure, and LD in the global B. carinata collection provides critical information to assist future crop improvement.
High-quality nanopore genome assemblies were generated for two Brassica nigra genotypes (Ni100 and CN115125); a member of the agronomically important Brassica species. The N50 contig length for the two assemblies were 17.1 Mb (58 contigs) and 0.29 Mb (963 contigs), respectively, reflecting recent improvements in the technology. Comparison with a de novo short read assembly for Ni100 corroborated genome integrity and quantified sequence related error rates (0.002%). The contiguity and coverage allowed unprecedented access to low complexity regions of the genome. Pericentromeric regions and coincidence of hypo-methylation enabled localization of active centromeres and identified a novel centromere-associated ALE class I element which appears to have proliferated through relatively recent nested transposition events (<1 million years ago). Computational abstraction was used to define a post-triplication Brassica specific ancestral genome and to calculate the extensive rearrangements that define the genomic distance separating B. nigra from its diploid relatives.
It is only recently, with the advent of long-read sequencing technologies, that we are beginning to uncover previously uncharted regions of complex and inherently recursive plant genomes. To comprehensively study and exploit the genome of the neglected oilseed Brassica nigra , we generated two high-quality nanopore de novo genome assemblies. The N50 contig lengths for the two assemblies were 17.1 Mb (12 contigs), one of the best among 324 sequenced plant genomes, and 0.29 Mb (424 contigs), respectively, reflecting recent improvements in the technology. Comparison with a de novo short-read assembly corroborated genome integrity and quantified sequence-related error rates (0.2%). The contiguity and coverage allowed unprecedented access to low-complexity regions of the genome. Pericentromeric regions and coincidence of hypomethylation enabled localization of active centromeres and identified centromere-associated ALE family retro-elements that appear to have proliferated through relatively recent nested transposition events (<1 Ma). Genomic distances calculated based on synteny relationships were used to define a post-triplication Brassica- specific ancestral genome, and to calculate the extensive rearrangements that define the evolutionary distance separating B. nigra from its diploid relatives.
Summary Ensuring faithful homologous recombination in allopolyploids is essential to maintain optimal fertility of the species. Variation in the ability to control aberrant pairing between homoeologous chromosomes in Brassica napus has been identified. The current study exploited the extremes of such variation to identify genetic factors that differentiate newly resynthesised B. napus, which is inherently unstable, and established B. napus, which has adapted to largely control homoeologous recombination. A segregating B. napus mapping population was analysed utilising both cytogenetic observations and high‐throughput genotyping to quantify the levels of homoeologous recombination. Three quantitative trait loci (QTL) were identified that contributed to the control of homoeologous recombination in the important oilseed crop B. napus. One major QTL on BnaA9 contributed between 32 and 58% of the observed variation. This study is the first to assess homoeologous recombination and map associated QTLs resulting from deviations in normal pairing in allotetraploid B. napus. The identified QTL regions suggest candidate meiotic genes that could be manipulated in order to control this important trait and further allow the development of molecular markers to utilise this trait to exploit homoeologous recombination in a crop.
The heavy selection pressure due to intensive breeding of Brassica napus has created a narrow gene pool, limiting the ability to produce improved varieties through crosses between B. napus cultivars. One mechanism that has contributed to the adaptation of important agronomic traits in the allotetraploid B. napus has been chromosomal rearrangements resulting from homoeologous recombination between the constituent A and C diploid genomes. Determining the rate and distribution of such events in natural B. napus will assist efforts to understand and potentially manipulate this phenomenon. The Brassica high-density 60K SNP array, which provides genome-wide coverage for assessment of recombination events, was used to assay 254 individuals derived from 11 diverse cultivated spring type B. napus These analyses identified reciprocal allele gain and loss between the A and C genomes and allowed visualization of de novo homoeologous recombination events across the B. napus genome. The events ranged from loss/gain of 0.09 Mb to entire chromosomes, with almost 5% aneuploidy observed across all gametes. There was a bias toward sub-telomeric exchanges leading to genome homogenization at chromosome termini. The A genome replaced the C genome in 66% of events, and also featured more dominantly in gain of whole chromosomes. These analyses indicate de novo homoeologous recombination is a continuous source of variation in established Brassica napus and the rate of observed events appears to vary with genetic background. The Brassica 60K SNP array will be a useful tool in further study and manipulation of this phenomenon.
AbstractThe heavy selection pressure due to intensive breeding of Brassica napus has created a narrow gene pool, limiting the ability to produce improved varieties through crosses between B. napus cultivars. One mechanism that has contributed to the adaptation of important agronomic traits in the allotetraploid B. napus has been chromosomal rearrangements resulting from homoeologous recombination between the constituent A and C diploid genomes. Determining the rate and distribution of such events in natural B. napus will assist efforts to understand and potentially manipulate this phenomenon. The Brassica high-density 60K SNP array, which provides genome-wide coverage for assessment of recombination events, was used to assay 254 individuals derived from 11 diverse cultivated spring type B. napus. These analyses identified reciprocal allele gain and loss between the A and C genomes and allowed visualization of de novo homoeologous recombination events across the B. napus genome. The events ranged from loss/gain of 0.09 Mb to entire chromosomes, with almost 5% aneuploidy observed across all gametes. There was a bias toward sub-telomeric exchanges leading to genome homogenization at chromosome termini. The A genome replaced the C genome in 66% of events, and also featured more dominantly in gain of whole chromosomes. These analyses indicate de novo homoeologous recombination is a continuous source of variation in established Brassica napus and the rate of observed events appears to vary with genetic background. The Brassica 60K SNP array will be a useful tool in further study and manipulation of this phenomenon.
The Brassica napus 60K Illumina Infinium™ SNP array has had huge international uptake in the rapeseed community due to the revolutionary speed of acquisition and ease of analysis of this high-throughput genotyping data, particularly when coupled with the newly available reference genome sequence. However, further utilization of this valuable resource can be optimized by better understanding the promises and pitfalls of SNP arrays. We outline how best to analyze Brassica SNP marker array data for diverse applications, including linkage and association mapping, genetic diversity and genomic introgression studies. We present data on which SNPs are locus-specific in winter, semi-winter and spring B. napus germplasm pools, rather than amplifying both an A-genome and a C-genome locus or multiple loci. Common issues that arise when analyzing array data will be discussed, particularly those unique to SNP markers and how to deal with these for practical applications in Brassica breeding applications.
Brassica napus seed composition traits (fibre, protein, oil and fatty acid profiles), seed colour and yield-associated traits are regulated by a complex network of genetic factors. Although previous studies have attempted to dissect the underlying genetic basis for these traits, a more complete picture of the available quantitative trait loci (QTL) variation and any interaction between the different traits is required. In this study, QTL mapping for eleven seed composition traits, seed colour and a yield-related trait (TSW) was conducted in a spring-type canola-quality B. napus doubled haploid (DH) population from a cross between black-seeded (DH12075) and yellow-seeded (YN01-429) lines across five environments. A major QTL associated with fibre traits (acid detergent fibre, acid detergent lignin and neutral detergent fibre) and seed colour (whiteness index) was mapped on chromosome N9 across the five environments. Multi-trait analysis identified QTL which had pleiotropic effect for seed colour and other composition traits. Multi-environment analysis revealed genetic (QTL) × environment effects on most QTL. These findings provide a more detailed insight into the complex QTL networks controlling seed composition and yield-associated traits in canola-quality B. napus.
The Brassica napus Illumina array provides genome-wide markers linked to the available genome sequence, a significant tool for genetic analyses of the allotetraploid B. napus and its progenitor diploid genomes. A high-density single nucleotide polymorphism (SNP) Illumina Infinium array, containing 52,157 markers, was developed for the allotetraploid Brassica napus. A stringent selection process employing the short probe sequence for each SNP assay was used to limit the majority of the selected markers to those represented a minimum number of times across the highly replicated genome. As a result approximately 60 % of the SNP assays display genome-specificity, resolving as three clearly separated clusters (AA, AB, and BB) when tested with a diverse range of B. napus material. This genome specificity was supported by the analysis of the diploid ancestors of B. napus, whereby 26,504 and 29,720 markers were scorable in B. oleracea and B. rapa, respectively. Forty-four percent of the assayed loci on the array were genetically mapped in a single doubled-haploid B. napus population allowing alignment of their physical and genetic coordinates. Although strong conservation of the two positions was shown, at least 3 % of the loci were genetically mapped to a homoeologous position compared to their presumed physical position in the respective genome, underlying the importance of genetic corroboration of locus identity. In addition, the alignments identified multiple rearrangements between the diploid and tetraploid Brassica genomes. Although mostly attributed to genome assembly errors, some are likely evidence of rearrangements that occurred since the hybridisation of the progenitor genomes in the B. napus nucleus. Based on estimates for linkage disequilibrium decay, the array is a valuable tool for genetic fine mapping and genome-wide association studies in B. napus and its progenitor genomes.
Camelina sativa, a largely relict crop, has recently returned to interest due to its potential as an industrial oilseed. Molecular markers are key tools that will allow C. sativa to benefit from modern breeding approaches. Two complementary methodologies, capture of 3' cDNA tags and genomic reduced-representation libraries, both of which exploited second generation sequencing platforms, were used to develop a low density (768) Illumina GoldenGate single nucleotide polymorphism (SNP) array. The array allowed 533 SNP loci to be genetically mapped in a recombinant inbred population of C. sativa. Alignment of the SNP loci to the C. sativa genome identified the underlying sequenced regions that would delimit potential candidate genes in any mapping project. In addition, the SNP array was used to assess genetic variation among a collection of 175 accessions of C. sativa, identifying two sub-populations, yet low overall gene diversity. The SNP loci will provide useful tools for future crop improvement of C. sativa.