Flax (Linum usitatissimum L.) has been domesticated for dual end uses as linseed and fiber flax, yet the genomic basis of morphotype divergence remains unclear. Here, we constructed a morphotype-resolved pangenome by integrating three newly generated near telomere-to-telomere genome assemblies with 14 previously published ones. Despite substantial variation in assembly size, driven primarily by DNA transposons, gene content was highly conserved, with little evidence for significant morphotype-specific gene presence-absence variation. Population genomic analyses of 407 accessions revealed that fiber flax had reduced nucleotide diversity, extended linkage disequilibrium, and a more compact population structure relative to linseed, consistent with stronger selection and a narrower genetic base. Genome-wide differentiation was heterogeneous and concentrated in discrete regions. Integration of FST, nucleotide diversity ratios, Tajima's D, and genome-wide association signals identified morphotype-enriched genomic blocks distributed across the genome. Many candidate regions are primarily supported by directional shifts in nucleotide diversity rather than extreme differentiation, indicating selection on standing genetic variation. Genome-wide association analyses identified 1,712 unique quantitative trait nucleotides (QTNs), with predominantly small effect sizes and strong enrichment in gene-proximal regions, consistent with a polygenic architecture. Overall, fiber flax traits tend to be controlled by fewer loci with moderate-to-large effects, whereas linseed traits exhibit a more diffuse genetic architecture. Patterns of Tajima's D further support non-classical selection dynamics, with predominantly positive values in linseed and localized negative values in fiber flax, consistent with selection on standing genetic variation. Together, our results suggest that flax morphotype divergence is driven primarily by selection on pre-existing allelic variation within a conserved gene repertoire. This study provides a comprehensive framework linking genome structure, population genomics, and trait architecture, and highlights the importance of standing genetic variation as a key resource for flax breeding and improvement.
We report complete genome sequences of Rhizobium favelukesii strains T136, T1470, and T1473 isolated from root-nodules of Medicago and Melilotus. Each genome comprises a chromosome (~4.2 Mb) and four to seven plasmids (~0.01–1.99 Mb) and contains predicted nodulation and nitrogen-fixation genes but lacks predicted type III secretion system genes.
ABSTRACT Chromosome number variation and structural reorganization are key drivers of plant evolution, yet their genomic basis remains unclear due to incomplete representation of repetitive regions in existing assemblies. The Linum genus exhibits exceptional karyotypic diversity ( n = 7–43), providing a powerful system to investigate chromosome evolution. Here, we generated near telomere-to-telomere (T2T) genome assemblies for four species, including cultivated flax ( L. usitatissimum cv. CDC Bethune; n = 15), its wild progenitor ( L. bienne ; n = 15), and two related species ( L. decumbens and L. grandiflorum ; n = 8). Together with published genomes of L. lewisii ( n = 9) and L. tenue ( n = 10), these enabled reconstruction of chromosome evolution across six lineages. Phylogenomic analyses revealed a shared ancestral whole-genome duplication (WGD) associated with the n = 9 karyotype, followed by lineage-specific WGDs and divergent diploidization. The transition from n = 8 to the derived n = 15 flax lineage not only occurred without chromosome length expansion, but also with genome size reduction, indicating extensive internal restructuring. Comparative analyses showed that this restructuring was associated with lineage-specific expansion of a single DNA transposon family ( TE_00003234 ; hAT ), which is highly enriched in expansive pericentromeric regions that are characterized by low gene density and nucleotide diversity, suppressed recombination, segregation distortion, and extensive synteny disruption, unlike the LTR retrotransposon-rich pericentromeres typical of most plant genomes. These findings support a model in which lineage-specific DNA transposon expansion is associated with remodeling of pericentromeric architecture and large-scale chromosome restructuring following polyploidization.
We report complete genome sequences of nitrogen-fixing Sinorhizobium medicae strains T2 and T10 isolated from Medicago and Melilotus. Each genome comprises a chromosome (~4.1 to 5.5 Mb) and three plasmids (~0.17 to 1.58 Mb) encoding predicted nodulation and nitrogen-fixation genes but not predicted type III secretion system genes.
Genomic selection (GS) is a core strategy in modern breeding programs, yet the rapid expansion of statistical, machine-learning (ML), and deep-learning (DL) models has made systematic evaluation and practical deployment increasingly challenging. To address these issues, we developed MultiGS, a unified and user-friendly framework that integrates linear, ML, DL, hybrid, and ensemble GS models within a standardized and computationally efficient workflow. MultiGS is implemented through two complementary pipelines: MultiGS-R, a Java/R pipeline implementing 12 statistical and ML models, and MultiGS-P, a Python pipeline integrating 17 models including five linear models, three ML approaches, and nine recently developed DL architectures implemented within the framework. We benchmarked MultiGS using wheat, maize, and flax datasets representing contrasting prediction scenarios. Wheat and maize were evaluated using random training-test splits within the same population, reflecting suitable conditions for assessing model capacity and scalability. Under these scenarios, several DL, hybrid, and ensemble models achieved prediction accuracies comparable to RR-BLUP and consistently exceeded those of GBLUP. In contrast, the flax dataset represented a true across-population prediction scenario with limited training set size and strong population structure. In this challenging context, classical linear models provided stable baselines, while a subset of DL architectures - particularly graph-based models and BLUP-integrated hybrids - demonstrated comparatively improved generalization across populations. Comparisons with previously published DL tools showed that MultiGS models achieved comparable or improved prediction accuracies while requiring lower computational costs, enabling routine retraining and large-scale evaluation. Overall, MultiGS informs, scenario-specific model selection and provides a practical platform for deploying genomic prediction under realistic breeding conditions. The software is freely available on GitHub (https://github.com/AAFC-ORDC-Crop-Bioinfomatics/MultiGS). ### Competing Interest Statement The authors have declared no competing interest. Genome Canada, https://ror.org/029s29983, 4DWheat Canadian Agricultural Partnership AgriScience, SCAP-ACS-05, SCAP-ASC-08
Genomic selection (GS) is a promising strategy to improve breeding efficiency for complex traits such as seed yield by enabling early selection and reducing reliance on extensive field testing. However, practical deployment of GS remains challenging due to limited training populations sizes and reduced predictive ability when models are applied to true breeding germplasm. In this study, we evaluated GS for flax (Linum usitatissimum L.) seed yield under realistic breeding scenarios, with a focus on across-population prediction (APP) and breeding decision support rather than model benchmarking. Using historical germplasm collections and a newly developed breeding-oriented population as training sets, GS performance was assessed across multiple independent test populations representing contemporary breeding lines evaluated in replicated yield trials. APP predictive abilities ranged from r = 0.67 to 0.84 depending on population relatedness when training and test populations were genetically aligned, supporting routine breeding deployment. Training population composition emerged as a key determinant of prediction success, with breeding-oriented populations consistently outperforming broad germplasm collections for predicting true breeding lines. Check-based selection analyses showed that GS reliably reproduced phenotypic advancement decisions while eliminating 61–91
Four field trials were conducted in western Canada in 2020 and 2021 to assess cadmium (Cd) concentration in seed from 166 pure lines of flax originating from 26 countries that were derived by single seed descent from a genetically diverse flax core collection preserved by Plant Gene Resources of Canada. Associations of Cd concentration with morphological and phenological traits, as well as the country of origin were considered. The mean Cd concentration in the seed ranged from 0.31 to 1.50 mg/kg with an overall mean value of 0.93 ± 0.22 mg/kg. The Cd in the soil from the field sites ranged from 0.35 to 2.80 mg/kg and was positively correlated with the Cd concentration in the seed harvested from the respective sites. The consistency of the Cd concentration in the seed of pure lines across the four site-years was low, with correlation coefficients ranging from 0.35 to 0.53. Pure lines with consistently low Cd concentration in the seed were identified and may be useful for breeding linseed cultivars for human consumption while pure lines with consistently high Cd concentration might be used for phytoremediation of Cd-contaminated soils. Cd concentration in the seed was not significantly correlated with thousand seed weight, plant height, petal colour, seed colour, or country of origin. As a tendency, the longer the vegetative period and the later the accessions matured, the more Cd they accumulated in the seed. For linseed production in western Canada, choosing locations with a low Cd content in the soil is very important when growing flax for human consumption or feed use.
We announce the complete genome sequence of Bradyrhizobium diazoefficiens 172S4, a highly efficient nitrogen-fixing symbiont isolated from a soybean root nodule. The chromosome (~9.1 Mb) exhibits a large inversion (~5.3 Mb) and encodes genes for symbiosis, nitrogen fixation, N2O reduction, and hydrogen uptake, highlighting the potential of 172S4 for sustainable agriculture.
Two novel bacterial strains isolated from root nodules of white sweet clover ( Melilotus albus ) plants grown at a Canadian site were previously characterized and placed in the genus Phyllobacterium . Here, we present phylogenomic and phenotypic data to support the description of strain T1293 T as representative of a novel species and present the first complete closed genome sequence of a bacterial strain (T1018) representing the species ‘ Phyllobacterium pellucidum ’. Phylogenetic analysis of genome sequences, as well as analysis of 53 core genes, placed novel strain T1293 T in a highly supported cluster of strains distinct from named Phyllobacterium species with Phyllobacterium myrsinacearum and Phyllobacterium calauticae as closest relatives. The highest average nucleotide identity and digital DNA–DNA hybridization values of genome sequences of T1293 T compared to closest species type strains (84.1 and 26.5%, respectively) are well below the threshold values for bacterial species circumscription. The genome of strain T1293 T has a size of 5,074,034 bp with a DNA G+C content of 55.1 mol% and possesses three plasmids with sizes of 397,619 bp, 476,847 bp and 519,835 bp. Detected in the genome were type III and type VI secretion system genes, implicated in plant–microbe and microbe–microbe interactions, but key nodulation, nitrogen-fixation and photosystem genes were not detected. A novel prophage (size ~41.5 kb) was also detected in the genome of T1293 T . Tests using defined culture media revealed that novel strains T1293 T and T1018 were highly resistant to, and able to metabolize glyphosate, a widely used herbicide that has negative consequences for the environment and human health. Data for multiple morphological, physiological, biochemical and plant tests complemented the sequence-based data. The data presented support the description of a new species, and the name Phyllobacterium meliloti sp. nov. is proposed with T1293 T =LMG 32641 T =HAMBI 3765 T as the species type strain.
Key message Fine-mapping of a locus on chromosome 1 of flax identified an S-lectin receptor-like kinase (SRLK) as the most likely candidate for a major Fusarium wilt resistance gene. Abstract Fusarium wilt, caused by the soil-borne fungal pathogen Fusarium oxysporum f. sp. lini , is a devastating disease in flax. Genetic resistance can counteract this disease and limit its spread. To map major genes for Fusarium wilt resistance, a recombinant inbred line population of more than 700 individuals derived from a cross between resistant cultivar ‘Bison’ and susceptible cultivar ‘Novelty’ was phenotyped in Fusarium wilt nurseries at two sites for two and three years, respectively. The population was genotyped with 4487 single nucleotide polymorphism (SNP) markers. Twenty-four QTLs were identified with IciMapping, 18 quantitative trait nucleotides with 3VmrMLM and 108 linkage disequilibrium blocks with RTM-GWAS. All models identified a major QTL on chromosome 1 that explained 20–48% of the genetic variance for Fusarium wilt resistance. The locus was estimated to span ~ 867 Kb but included a ~ 400 Kb unresolved region. Whole-genome sequencing of ‘CDC Bethune’, ‘Bison’ and ‘Novelty’ produced ~ 450 Kb continuous sequences of the locus. Annotation revealed 110 genes, of which six were considered candidate genes. Fine-mapping with 12 SNPs and 15 Kompetitive allele-specific PCR (KASP) markers narrowed down the interval to ~ 69 Kb, which comprised the candidate genes Lus10025882 and Lus10025891 . The latter, a G-type S-lectin receptor-like kinase (SRLK) is the most likely resistance gene because it is the only polymorphic one. In addition, Fusarium wilt resistance genes previously isolated in tomato and Arabidopsis belonged to the SRLK class. The robust KASP markers can be used in marker-assisted breeding to select for this major Fusarium wilt resistance locus.
Wheat, particularly common wheat (Triticum aestivum L.), is a major crop accounting for 25
Two novel bacterial strains isolated from root-nodules of white sweet clover (Melilotus albus) plants grown at a Canadian site were previously characterized and placed in the genus Phyllobacterium. Here we present phylogenomic and phenotypic data to support the description of strain T1293T as representative of a novel species and present the first complete closed genome sequence of a bacterial strain (T1018) representing the species "P. pellucidum". Phylogenetic analysis of genome sequences as well as analysis of 53 core genes placed novel strain T1293T in a highly supported cluster of strains distinct from named Phyllobacterium species with P. myrsinacearum and P. calauticae as closest relatives. The highest average nucleotide identity (ANI) and digital DNA-DNA hybridization (dDDH) values of genome sequences of T1293T compared to closest species type strains (84.1% and 26.5%, respectively) are well below the threshold values for bacterial species circumscription. The genome of strain T1293T has a size of 5074034 bp with a DNA G+C content of 55 mol% and possesses three plasmids with sizes of 397619 bp, 476847 bp and 519835 bp. Detected in the genome were Type III and Type VI secretion system genes, implicated in plant-microbe and microbe-microbe interactions, but key nodulation, nitrogen-fixation and photosystem genes were not detected. Further analysis revealed that T1293T, like other Phyllobacterium species, possesses key genes encoding an enzyme complex implicated in the degradation of glyphosate, a widely used broad-spectrum herbicide that has negative consequences for many microorganisms including the human gut microbiome. A novel prophage (size ~ 41.5 kb) was also detected in the genome of T1293T. Data for multiple phenotypic tests complemented the sequence-based characterization of strain T1293T. The data presented support the description of a new species and the name Phyllobacterium meliloti sp. nov. is proposed with T1293T = LMG32641T = HAMBI 3765T as the species type strain. ### Competing Interest Statement The authors have declared no competing interest.
Bacterial strain A19 T was previously isolated from a root-nodule of Aeschynomene indica (Indian jointvetch) and assigned to a new lineage in the genus Bradyrhizobium. Here data are presented for the detailed phylogenomic and taxonomic characterisation of strain A19 T . Phylogenetic analysis of whole genome sequences as well as 51 concatenated core gene sequences placed strain A19 T in a highly supported lineage that was distinct from described Bradyrhizobium species; B. oligotrophicum , a symbiont of A. indica, was the most closely related species. The digital DNA-DNA hybridization and average nucleotide identity values for strain A19 T in pair-wise comparisons with close relatives were far lower than the respective threshold values of 70% and ~96% for definition of species boundaries. The complete genome of strain A19 T consists of a single 8.44 Mbp chromosome (DNA G+C content, 64.9 mol%) and contains a photosynthesis gene cluster, nitrogen-fixation genes and genes encoding a complete denitrifying enzyme system including nitrous oxide reductase. Nodulation and type III secretion system genes, needed for nodulation by most rhizobia, were not detected in the genome of A19 T . Data for multiple phenotypic tests complemented the sequence-based analyses. Strain A19 T elicits nitrogen-fixing nodules on stems and roots of A. indica plants but not on soybeans or Macroptilium atropurpureum . Based on the data presented, a new species named Bradyrhizobium ontarionense sp. nov. is proposed with strain A19 T (= LMG 32638 T = HAMBI 3761 T ) as the type strain.
Flax is an important crop producing high-quality fiber from its stem and seeds that are rich in omega-3 alpha-linolenic acid (ALA). During the past decades, studies on flax have produced large amounts of data, including genomic and phenotypic data, such as RNA and DNA sequences, genome assemblies and their annotations, single-nucleotide polymorphisms (SNPs), quantitative trait loci (QTLs), quantitative trait nucleotides (QTNs), and predicted candidate genes associated with important traits. These data are represented in FlaxDB, a substantial body of resources for future genomics-assisted breeding in flax. Most flax sequences are deposited in the National Center for Biotechnology Information (NCBI) databases, whereas annotated genomic information is available primarily through the published literature and their supplementary files or in online databases. A comprehensive database integrating genomic information such as SNPs, QTLs/QTNs and their candidate genes, and trait phenotypes of genetic and breeding populations is warranted for efficient use of the data in flax research and breeding. FlaxDB is such a database. It integrates a wide array of flax data such as genome sequences, markers, SNPs, QTLs, a variety of pedigrees, germplasm, and phenotypes. This chapter briefly reviews the latest progress in flax genomic resources and phenotypic data, and describes a number of online databases available to the flax community.
OBJECTIVE:The 1,000 wheat exome project captured the single nucleotide variants in the coding regions of a diverse set of 890 wheat accessions to analyse the contribution of introgression to adaptation of wheat. However, this highly useful single nucleotide polymorphism (SNP) dataset is based on RefSeq v1.0 of the International Wheat Genome Sequencing Consortium (IWGSC) assembly of the bread wheat genome of Chinese Spring. This reference sequence has recently been updated using optical maps and long-read sequencing to produce the improved RefSeq v2.1. Our objective was to develop a reliable high-density SNP dataset positioned onto RefSeq v2.1 because it is the current standard reference sequence used by wheat researchers.RESULTS:The 3,039,822 SNPs originally positioned on RefSeq v1.0 were projected to v2.1 using Liftoff with four different flanking regions, and 2,946,536 SNPs were consistently lifted to the same location irrespective of the flanking region lengths. Of these, 2,799,166 were located on the '+' ve strand. The distribution of the SNPs across the 21 chromosomes on RefSeq v2.1 was similar to that of RefSeq v1.0. Among the SNPs that were based on unanchored scaffolds in RefSeq v1.0, 11,938 were projected to one of the 21 pseudomolecules in the upgraded assembly. This SNP dataset constitutes a much-needed standardized resource for the wheat research community.
Rust, pasmo (PAS), powdery mildew (PM), and Fusarium wilt (FW) are four major diseases of flax, causing yield losses as well as seed and fiber quality reduction. Flax rust resistance is controlled by few major genes, some of which have been cloned. Currently, all new Canadian flax cultivars are immune to rust, exemplifying the success of genetic improvement for rust resistance. Resistance to PAS, PM, or FW is a heritable quantitative trait, involving multiple small-effect polygenes. In the last decade, significant advances have been made in genotyping core collections and biparental populations along their large-scale phenotyping for disease resistance over multiple years and locations. Hundreds of quantitative trait loci (QTLs) or quantitative trait nucleotides (QTNs) and the linked candidate genes for all three diseases have been identified using the rapidly evolving genome-wide association study (GWAS) statistical models. The effects of these QTLs/QTNs and candidate genes are mostly additive, offering valuable genomic resources and potential to pyramid the resistant alleles into future cultivars or pre-breeding germplasm via genomics-assisted breeding technologies, such as marker-assisted selection, genomic selection, and gene editing.
IntroductionWheat rust diseases are widespread and affect all wheat growing areas around the globe. Breeding strategies focus on incorporating genetic disease resistance. However, pathogens can quickly evolve and overcome the resistance genes deployed in commercial cultivars, creating a constant need for identifying new sources of resistance. MethodsWe have assembled a diverse tetraploid wheat panel comprised of 447 accessions of three Triticum turgidum subspecies and performed a genome-wide association study (GWAS) for resistance to wheat stem, stripe, and leaf rusts. The panel was genotyped with the 90K Wheat iSelect single nucleotide polymorphism (SNP) array and subsequent filtering resulted in a set of 6,410 non-redundant SNP markers with known physical positions. ResultsPopulation structure and phylogenetic analyses revealed that the diversity panel could be divided into three subpopulations based on phylogenetic/geographic relatedness. Marker-trait associations (MTAs) were detected for two stem rust, two stripe rust and one leaf rust resistance loci. Of them, three MTAs coincide with the known rust resistance genes Sr13, Yr15 and Yr67, while the other two may harbor undescribed resistance genes. DiscussionThe tetraploid wheat diversity panel, developed and characterized herein, captures wide geographic origins, genetic diversity, and evolutionary history since domestication making it a useful community resource for mapping of other agronomically important traits and for conducting evolutionary studies.
A bacterial strain, designated T173T, was previously isolated from a root-nodule of a Melilotus albus plant growing in Canada and identified as a novel Ensifer lineage that shared a clade with the non-symbiotic species, Ensifer adhaerens. Strain T173T was also previously found to harbour a symbiosis plasmid and to elicit root-nodules on Medicago and Melilotus species but not fix nitrogen. Here we present data for the genomic and taxonomic description of strain T173T. Phylogenetic analyses including the analysis of whole genome sequences and multiple locus sequence analysis (MLSA) of 53 concatenated ribosome protein subunit (rps) gene sequences confirmed placement of strain T173T in a highly supported lineage distinct from named Ensifer species with E. morelensis Lc04T as the closest relative. The highest digital DNA–DNA hybridization (dDDH) and average nucleotide identity (ANI) values of genome sequences of strain T173T compared with closest relatives (35.7 and 87.9%, respectively) are well below the respective threshold values of 70% and 95–96% for bacterial species circumscription. The genome of strain T173T has a size of 8,094,229 bp with a DNA G + C content of 61.0 mol%. Six replicons were detected: a chromosome (4,051,102 bp) and five plasmids harbouring plasmid replication and segregation (repABC) genes. These plasmids were also found to possess five apparent conjugation systems based on analysis of TraA (relaxase), TrbE/VirB4 (part of the Type IV secretion system (T4SS)) and TraG/VirD4 (coupling protein). Ribosomal RNA operons encoding 16S, 23S, and 5S rRNAs that are usually restricted to bacterial chromosomes were detected on plasmids pT173d and pT173e (946,878 and 1,913,930 bp, respectively) as well as on the chromosome of strain T173T. Moreover, plasmid pT173b (204,278 bp) was found to harbour T4SS and symbiosis genes, including nodulation (nod, noe, nol) and nitrogen fixation (nif, fix) genes that were apparently acquired from E. medicae by horizontal transfer. Data for morphological, physiological and symbiotic characteristics complement the sequence-based characterization of strain T173T. The data presented support the description of a new species for which the name Ensifer canadensis sp. nov. is proposed with strain T173T (= LMG 32374T = HAMBI 3766T) as the species type strain.
Wheat was one of the crops domesticated in the Fertile Crescent region approximately 10,000 years ago. Despite undergoing recent polyploidization, hull-to-free-thresh transition events, and domestication bottlenecks, wheat is now grown in over 130 countries and accounts for a quarter of the world’s cereal production. The main reason for its widespread success is its broad genetic diversity that allows it to thrive in different environments. To trace historical selection and hybridization signatures, genome scans were performed on two datasets: approximately 113K SNPs from 921 predominantly bread wheat accessions and approximately 110K SNPs from about 400 wheat accessions representing all ploidy levels. To identify environmental factors associated with the loci, a genome–environment association (GEA) was also performed. The genome scans on both datasets identified a highly differentiated region on chromosome 4A where accessions in the first dataset were dichotomized into a group (n = 691), comprising nearly all cultivars, wild emmer, and most landraces, and a second group (n = 230), dominated by landraces and spelt accessions. The grouping of cultivars is likely linked to their potential ancestor, bread wheat cv. Norin-10. The 4A region harbored important genes involved in adaptations to environmental conditions. The GEA detected loci associated with latitude and temperature. The genetic signatures detected in this study provide insight into the historical selection and hybridization events in the wheat genome that shaped its current genetic structure and facilitated its success in a wide spectrum of environmental conditions. The genome scans and GEA approaches applied in this study can help in screening the germplasm housed in gene banks for breeding, and for conservation purposes.