Plant stress responses occur within daily cycles of physiology, metabolism, and growth, making timing a critical dimension of acclimation. In Arabidopsis, circadian and diel regulation influence responses to abiotic stress, including cold, but how this temporal regulation is conserved, diversified, or expanded in crop genomes remains unclear. This question is especially challenging in Brassica rapa, which underwent a genome triplication after diverging from Arabidopsis, resulting in multiple retained paralogs that can be grouped by Arabidopsis orthology and ancient homeologous relationships. Here, we generated a B. rapa pangenome spanning six morphotypes and used it to profile diel (24 h) cold acclimation responses across diverse accessions differing in freeze tolerances. Cold altered peak expression time for thousands of genes, which we grouped into distinct phase-change groups. Circadian leaf movement assays revealed accession-specific differences in clock period and temperature compensation under cold, suggesting that altered clock behavior may contribute in part to the diel transcriptome retiming. At the individual gene level, inferred gene regulatory networks (GRNs) were highly accession-specific and lost shared connectivity under cold stress. However, grouping these paralogs by their Arabidopsis orthologs revealed a highly conserved regulatory architecture that was otherwise masked by paralog diversification. Integrating these networks with functional pathways identified key candidate regulators of retimed processes, including modules linked to nighttime phosphorylation and daytime photosynthesis. Finally, analyzing conserved noncoding sequences across the pangenome prioritized specific regulatory targets within cold-retimed groups. Together, these results demonstrate that cold acclimation in B. rapa is shaped by a combination of diel retiming, paralog-specific regulation, and deeply conserved programs.
Persistent drought affects global crop production and is becoming more severe in many parts of the world in recent decades. Deciphering how plants respond to drought will facilitate the development of flexible mitigation strategies. Sorghum bicolor L. Moench (sorghum), a major cereal crop and an emerging bioenergy crop, exhibits remarkable resilience to drought. To better understand the molecular traits that underlie sorghum's remarkable drought tolerance, we undertook a large-scale sorghum gene expression profiling effort, totaling nearly 1500 transcriptome profiles, across a 3-year field study with replicated plots in California's Central Valley. This study included time-resolved gene expression data from roots and leaves of two sorghum genotypes, BTx642 and RTx430, with different pre-flowering and post-flowering drought-tolerance adaptations under control and drought conditions. Quantification of genotype-specific drought tolerance effects was enabled by de novo sequencing, assembly, and annotation of both BTx642 and RTx430 genomes. These reference-quality genomes were used to construct a pangene set for characterizing conserved and genotype-specific expression. By integrating time-resolved transcriptomic responses to drought in the field across three consecutive years, we identified a set of 726 drought-responsive genes that responded similarly in all 3 years of our field study. Functional enrichment analysis identified abiotic stress, secondary cell wall-related processes and metabolism as particularly affected under both types of drought stress. We also found that some glyoxylate cycle pathway genes, including malate synthase and isocitrate lyase, are differentially regulated particularly during post-flowering drought stress, implicating this pathway as potentially important for drought responsiveness. This expansive dataset represents a unique resource for sorghum and drought research communities and provides a methodological framework for the integration of multi-faceted time-resolved transcriptomic datasets.
The legume family originated ca. 60-65 million years ago and soon diversified into at least six lineages (now extant subfamilies). The signal of whole genome duplications (WGD) is apparent in species sampled from all six subfamilies. The early diversification has posed difficulties for resolving the legume backbone structure and the timing of WGDs, especially in Caesalpinioideae where the diversification and WGD signals coincide. In this study, we report the genome sequences and annotations for Cercis canadensis (Cercidoideae) and Chamaecrista fasciculata (Caesalpinioideae) to help resolve the timings of WGDs relative to subfamily origins and the ancestral legume karyotype. Analyses of genome assemblies from four subfamilies within Fabaceae show that the last common ancestor of all legumes likely had seven chromosomes, with a genome structure similar to the extant Cercis genome. The retained karyotype structure, the lack of a WGD in the last 100+ Mya (Cercis and the lineage leading to it following the eudicot γ whole-genome triplication), and the unusually slow rates of nucleotide substitution and structural evolution in the Cercis genome underscore its utility as a genomic proxy for the last common ancestor of all legume species. Our analysis supports an allopolyploid origin of Caesalpinioideae, with progenitors from lineages along the backbone of the legume phylogeny. Rapid diversification and the inferred allopolyploid origin of Caesalpinioideae provide a partial explanation for the difficulty in resolving the backbone of the legume phylogeny and early Caesalpinioideae diversification.
Although the green revolution adapted a handful of crops to homogeneous and high-input industrialized agriculture, much of the global population still relies on the local production of variable crop cultivars by low-input smallholder farms. This diversity of unhomogenized crops1, like that of the grain and bioenergy crop sorghum2-5, offers raw materials for genetic gain and cultivar improvement. However, breeding efforts can be constrained by highly specialized traits and breeding targets6. Here, to bridge this diversity, we constructed a 33-member pangenome reference and a diversity panel across 1,984 cultivars and landraces. We leveraged these resources to explore the complex interplay among historical contingency, ongoing adaptation and previously uncharacterized structural diversity. Specifically, our analyses conclusively demonstrated multiple nested and deeply diverged structural variants in the domestication gene SHATTERING1, which distinguish the previously established multicentric origin of sorghum. We then applied landscape genomics to reveal how gene flow and secondary contact created the complex genetic mosaic in contemporary breeding networks. As proof of concept for pangenome-accelerated trait discovery, we connected biosynthetic gene cluster structural variation to phenotypic leaf concentration of the cyanogenic glucoside dhurrin. Combined, these approaches will accelerate breeding and trait discovery and provide a framework for similar applications in other crops.
While the green revolution adapted a handful of crops to homogenous and high-input industrialized agriculture, much of the global population still relies on local food production from low-input smallholder farms that grow highly variable crop cultivars. The high diversity of the grain and bioenergy crop sorghum [1][1]–[4][2], and many other crops that were not homogenized during the green revolution [5][3], not only provides the raw materials for breeders to make substantial gains in cultivar improvement, but also constrains breeding efforts due to highly specialized locally adapted plant phenotypes [6][4]. Here, we construct a 33-member pangenome and identify trait-associated variants in 1,988 cultivars and landraces. We then apply these resources to explore the complex interplay between historical contingency, ongoing adaptation, and the potential for future gains through climate-aware genome-enabled breeding. Specifically, our analyses conclusively demonstrate that multiple nested, deeply diverged, and previously uncharacterized structural variants in the domestication gene SHATTERING1 distinguish the previously established multicentric origin of sorghum. We then apply landscape genomics tests to reveal how gene flow, adaptation, and secondary contact created the complex genetic mosaic in current global breeding networks. Further analysis of climate-gene associations highlights candidate loci underlying adaptation, including the biosynthetic gene cluster for the cyanogenic glucoside dhurrin. Combined, the pangenome-informed variants developed here will enable both trait discovery and subsequent marker assays to accelerate breeding and provide a framework for similar applications in other diverse and non-model crops. ### Competing Interest Statement The authors have declared no competing interest. United States Department of Energy, DE-AC02-05CH11231 [1]: #ref-1 [2]: #ref-4 [3]: #ref-5 [4]: #ref-6
Polyploidy is ubiquitous across North American prairies, which provide essential ecosystem services and rich soil for agriculture. Yet the mechanism driving polyploid abundance is unclear. Multiple hypotheses have been proposed including polyploid abundance is proportional to the opportunity for whole genome duplication (WGD), and WGD alters phenotypes that may increase fitness. We tested these two hypotheses together in the mixed-ploidy species Andropogon gerardi , a dominant grass species in endangered North American tallgrass prairies. Leveraging a novel, phased allopolyploid reference genome, we found the A. gerardi hexaploid arose after the C 4 grassland expansion in the early Pleistocene, when glacial cycles likely increased secondary contact between the diploid progenitors. We sequenced A. gerardi from 25 popula-tions and examined cytotype performance and morphology in a controlled environment to investigate the consequences of the contemporary mixed-ploidy populations. We found the 9 x A. gerardi cytotype is a neopolyploid and a result of recurrent WGD events. Further, we demonstrate the 9 x neopolyploids have greater growth and a decreased stomatal pore index, which is adaptive in xeric climates where the 9 x cy-totype is most common. Together, our results support both hypotheses for polyploid abundance in North America: WGD is a product of opportunity and can have immediate fitness consequences. Although the changes to fitness may provide an advantage to 9 x A. gerardi , the establishment of 9 x may lower overall population fitness due to the lower reproductive viability of 9 x individuals. Polyploid species are abundant in North American prairies and make up many of the dominant species in the ecosystem. This prominence could be a result of whole genome duplication conferring an advantage that increases the frequency of polyploids or could simply indicate that the opportunity for whole genome duplication is higher in this ecosystem, or both. Through examining three polyploidization events in A. gerardi , the dominant species in endangered tallgrass North American prairies, we found whole genome duplication is both surprisingly common and confers traits that are beneficial in some environments.
Ancient whole-genome duplications are believed to facilitate novelty and adaptation by providing the raw fuel for new genes. However, it is unclear how recent whole-genome duplications may contribute to evolvability within recent polyploids. Hybridization accompanying some whole-genome duplications may combine divergent gene content among diploid species. Some theory and evidence suggest that polyploids have a greater accumulation and tolerance of gene presence-absence and genomic structural variation, but it is unclear to what extent either is true. To test how recent polyploidy may influence pangenomic variation, we sequenced, assembled, and annotated 12 complete, chromosome-scale genomes of Camelina sativa, an allohexaploid biofuel crop with 3 distinct subgenomes. Using pangenomic comparative analyses, we characterized gene presence-absence and genomic structural variation both within and between the subgenomes. We found over 75% of ortholog gene clusters are core in C. sativa and <10% of sequence space was affected by genomic structural rearrangements. In contrast, 19% of gene clusters were unique to one subgenome, and the majority of these were Camelina specific (no ortholog in Arabidopsis). We identified an inversion that may contribute to vernalization requirements in winter-type Camelina and an enrichment of Camelina-specific genes with enzymatic processes related to seed oil quality and Camelina's unique glucosinolate profile. Genes related to these traits exhibited little presence-absence variation. Our results reveal minimal pangenomic variation in this species and instead show how hybridization accompanied by whole-genome duplication may benefit polyploids by merging diverged gene content of different species.
Cultivar 'Williams 82' has served as the reference genome for the soybean research community since 2008, but is known to have areas of genomic heterogeneity among different sub-lines. This work provides an updated assembly (version Wm82.a6) derived from a specific sub-line known as 'Wm82-ISU-01' (seeds available under USDA accession PI 704477). The genome was assembled using Pacific BioSciences HiFi reads and integrated into chromosomes using HiC. The 20 soybean chromosomes assembled into a genome of 1.01Gb, consisting of 36 contigs. The genome annotation identified 48,387 gene models, named in accordance with previous assembly versions Wm82.a2 and Wm82.a4. Comparisons of Wm82.a6 with other near-gapless assemblies of 'Williams 82' reveal large regions of genomic heterogeneity, including regions of differential introgression from the genotype 'Kingwa' within approximately 30 Mb and 25 Mb segments on chromosomes 03 and 07, respectively. Additionally, our analysis revealed a previously unknown large (~20 Mb) heterogeneous region in the pericentromeric region of chromosome 12, where Wm82.a6 matches the 'Williams' haplotype while the other two near-gapless assemblies do not match the haplotype of either parent of 'Williams 82'. In addition to the Wm82.a6 assembly, we also assembled the genome of soybean line 'Fiskeby III', a rich resource for abiotic stress resistance genes. A genome comparison of Wm82.a6 with 'Fiskeby III' revealed the nucleotide and structural polymorphisms between the two genomes within a QTL region for iron deficiency chlorosis resistance. The Wm82.a6 and 'Fiskeby III' genomes described here will enhance comparative and functional genomics capacities and applications in the soybean community.
The grass family (Poaceae, Poales) holds immense economic and ecological significance, exhibiting unique metabolic traits, including dual starch and lignin biosynthetic pathways. To investigate when and how the metabolic innovations known in grasses evolved, we sequenced the genomes of a non-core grass, Pharus latifolius, the non-grass graminids, Joinvillea ascendens and Ecdeiocolea monostachya, representing the sister clade to Poaceae, and Typha latifolia, representing the sister clade to the remaining Poales. The rho whole genome duplication (ρWGD) in the ancestral lineage for all grasses contributed to the gene family expansions underlying cytosolic starch biosynthesis, whereas an earlier tandem duplication of phenylalanine ammonia lyase (PAL) gave rise to phenylalanine/tyrosine ammonia lyase (PTAL) responsible for the dual lignin biosynthesis. Two mutations were sufficient to expand ancestral PAL function into PTAL. The integrated genomic and biochemical analyses of grass relatives in Poales revealed the evolutionary and molecular basis of key metabolic innovations of grasses. ### Competing Interest Statement Y.T-K, B.M., and H.A.M. have a pending patent application related to the two mutations that can convert PAL enzymes into PTAL enzymes. The other authors declare that they have no competing interests related to this work.
Homosporous lycophytes (Lycopodiaceae) are a deeply diverged lineage in the plant tree of life, having split from heterosporous lycophytes (Selaginella and Isoetes) ~400 Mya. Compared to the heterosporous lineage, Lycopodiaceae has markedly larger genome sizes and remains the last major plant clade for which no chromosome-level assembly has been available. Here, we present chromosomal genome assemblies for two homosporous lycophyte species, the allotetraploid Huperzia asiatica and the diploid Diphasiastrum complanatum. Remarkably, despite that the two species diverged ~350 Mya, around 30% of the genes are still in syntenic blocks. Furthermore, both genomes had undergone independent whole genome duplications, and the resulting intragenomic syntenies have likewise been preserved relatively well. Such slow genome evolution over deep time is in stark contrast to heterosporous lycophytes and is correlated with a decelerated rate of nucleotide substitution. Together, the genomes of H. asiatica and D. complanatum not only fill a crucial gap in the plant genomic landscape but also highlight a potentially meaningful genomic contrast between homosporous and heterosporous species.
### Competing Interest Statement The authors have declared no competing interest.
Sex chromosomes have evolved hundreds of times across the flowering plant tree of life; their recent origins in some members of this clade can shed light on the early consequences of suppressed recombination, a crucial step in sex chromosome evolution. Amborella trichopoda, the sole species of a lineage that is sister to all other extant flowering plants, is dioecious with a young ZW sex determination system. Here we present a haplotype-resolved genome assembly, including highly contiguous assemblies of the Z and W chromosomes. We identify a similar to 3-megabase sex-determination region (SDR) captured in two strata that includes a similar to 300-kilobase inversion that is enriched with repetitive sequences and contains a homologue of the Arabidopsis METHYLTHIOADENOSINE NUCLEOSIDASE (MTN1-2) genes, which are known to be involved in fertility. However, the remainder of the SDR does not show patterns typically found in non-recombining SDRs, such as repeat accumulation and gene loss. These findings are consistent with the hypothesis that dioecy is derived in Amborella and the sex chromosome pair has not significantly degenerated.
Cotton (Gossypium hirsutum L.) is the key renewable fibre crop worldwide, yet its yield and fibre quality show high variability due to genotype-specific traits and complex interactions among cultivars, management practices and environmental factors. Modern breeding practices may limit future yield gains due to a narrow founding gene pool. Precision breeding and biotechnological approaches offer potential solutions, contingent on accurate cultivar-specific data. Here we address this need by generating high-quality reference genomes for three modern cotton cultivars (‘UGA230’, ‘UA48’ and ‘CSX8308’) and updating the ‘TM-1’ cotton genetic standard reference. Despite hypothesized genetic uniformity, considerable sequence and structural variation was observed among the four genomes, which overlap with ancient and ongoing genomic introgressions from ‘Pima’ cotton, gene regulatory mechanisms and phenotypic trait divergence. Differentially expressed genes across fibre development correlate with fibre production, potentially contributing to the distinctive fibre quality traits observed in modern cotton cultivars. These genomes and comparative analyses provide a valuable foundation for future genetic endeavours to enhance global cotton yield and sustainability.
Soapwort (Saponaria officinalis) is a flowering plant from the Caryophyllaceae family with a long history of human use as a traditional source of soap. Its detergent properties are because of the production of polar compounds (saponins), of which the oleanane-based triterpenoid saponins, saponariosides A and B, are the major components. Soapwort saponins have anticancer properties and are also of interest as endosomal escape enhancers for targeted tumor therapies. Intriguingly, these saponins share common structural features with the vaccine adjuvant QS-21 and, thus, represent a potential alternative supply of saponin adjuvant precursors. Here, we sequence the S. officinalis genome and, through genome mining and combinatorial expression, identify 14 enzymes that complete the biosynthetic pathway to saponarioside B. These enzymes include a noncanonical cytosolic GH1 (glycoside hydrolase family 1) transglycosidase required for the addition of d-quinovose. Our results open avenues for accessing and engineering natural and new-to-nature pharmaceuticals, drug delivery agents and potential immunostimulants. Soapwort (Saponariaofficinalis) is a rich reservoir of triterpenoid glycosides that often have important pharmaceutical, nutraceutical and agronomical potential. Here, the authors elucidate the complete biosynthetic pathway of saponarioside B, a major saponin constituent in soapwort.
Populus species play a foundational role in diverse ecosystems and are important renewable feedstocks for bioenergy and bioproducts. Hybrid aspen Populus tremula × P. alba INRA 717-1B4 is a widely used transformation model in tree functional genomics and biotechnology research. As an outcrossing interspecific hybrid, its genome is riddled with sequence polymorphisms which present a challenge for sequence-sensitive analyses. Here we report a telomere-to-telomere genome for this hybrid aspen with two chromosome-scale, haplotype-resolved assemblies. We performed a comprehensive analysis of the repetitive landscape and identified both tandem repeat array-based and array-less centromeres. Unexpectedly, the most abundant satellite repeats in both haplotypes lie outside of the centromeres, consist of a 147 bp monomer PtaM147, frequently span >1 megabases, and form heterochromatic knobs. PtaM147 repeats are detected exclusively in aspens (section Populus) but PtaM147-like sequences occur in LTR-retrotransposons of closely related species, suggesting their origin from the retrotransposons. The genomic resource generated for this transformation model genotype has greatly improved the design and analysis of genome editing experiments that are highly sensitive to sequence polymorphisms. The work should motivate future hypothesis-driven research to probe into the function of the abundant and aspen-specific PtaM147 satellite DNA.
Five versions of the Chlamydomonas reinhardtii reference genome have been produced over the last two decades. Here we present version 6, bringing significant advances in assembly quality and structural annotations. PacBio-based chromosome-level assemblies for two laboratory strains, CC-503 and CC-4532, provide resources for the plus and minus mating-type alleles. We corrected major misassemblies in previous versions and validated our assemblies via linkage analyses. Contiguity increased over ten-fold and >80% of filled gaps are within genes. We used Iso-Seq and deep RNA-seq datasets to improve structural annotations, and updated gene symbols and textual annotation of functionally characterized genes via extensive manual curation. We discovered that the cell wall-less classical reference strain CC-503 exhibits genomic instability potentially caused by deletion of the helicase RECQ3, with major structural mutations identified that affect >100 genes. We therefore present the CC-4532 assembly as the primary reference, although this strain also carries unique structural mutations and is experiencing rapid proliferation of a Gypsy retrotransposon. We expect all laboratory strains to harbor gene-disrupting mutations, which should be considered when interpreting and comparing experimental results. Collectively, the resources presented here herald a new era of Chlamydomonas genomics and will provide the foundation for continued research in this important reference organism.
A contiguous assembly of the inbred ‘EL10’ sugar beet (Beta vulgaris ssp. vulgaris) genome was constructed using PacBio long-read sequencing, BioNano optical mapping, Hi-C scaffolding, and Illumina short-read error correction. The EL10.1 assembly was 540 Mb, of which 96.2% was contained in nine chromosome-sized pseudomolecules with lengths from 52 to 65 Mb, and 31 contigs with a median size of 282 kb that remained unassembled. Gene annotation incorporating RNA-seq data and curated sequences via the MAKER annotation pipeline generated 24,255 gene models. Results indicated that the EL10.1 genome assembly is a contiguous genome assembly highly congruent with the published sugar beet reference genome. Gross duplicate gene analyses of EL10.1 revealed little large-scale intra-genome duplication. Reduced gene copy number for well-annotated gene families relative to other core eudicots was observed, especially for transcription factors. Variation in genome size in B. vulgaris was investigated by flow cytometry among 50 individuals producing estimates from 633 to 875 Mb/1C. Read-depth mapping with short-read whole-genome sequences from other sugar beet germplasm suggested that relatively few regions of the sugar beet genome appeared associated with high-copy number variation.
Gene functional descriptions offer a crucial line of evidence for candidate genes underlying trait variation. Conversely, plant responses to environmental cues represent important resources to decipher gene function and subsequently provide molecular targets for plant improvement through gene editing. However, biological roles of large proportions of genes across the plant phylogeny are poorly annotated. Here we describe the Joint Genome Institute (JGI) Plant Gene Atlas, an update able data resource consisting of transcript abundance assays spanning 18 diverse species. To integrate across these diverse genotypes, we anal yzed e xpression pr ofiles, b uilt gene c lusters that exhibited tissue / condition specific expression, and tested for transcriptional response to environmental queues. We discovered extensive phylogenetically constrained and condition-specific expression profiles f or genes without an y previously documented functional annotation. Such conserved expression patterns and tightly co-expressed gene clusters let us assign expression derived additional biological information to 64 495 genes with otherwise unknown functions. The ever-expanding Gene Atlas resource is available at JGI Plant Gene Atlas ( https://plantgeneatlas.jgi.doe.gov ) and Phytozome ( https://phytozome .jgi.doe .gov/), providing bulk access to data and user-specified queries of gene sets. Combined, these web interfaces let users access differentiall y e xpressed genes, track orthologs across the Gene Atlas plants, graphically represent co-expressed genes, and visualize gene ontology and pathway enrichments.
The native, perennial shrub American hazelnut (Corylus americana) is cultivated in the Midwestern United States for its significant ecological benefits, as well as its high-value nut crop. Implementation of modern breeding methods and quantitative genetic analyses of C. americana requires high-quality reference genomes, a resource that is currently lacking. We therefore developed the first chromosome-scale assemblies for this species using the accessions 'Rush' and 'Winkler'. Genomes were assembled using HiFi PacBio reads and Arima Hi-C data, and Oxford Nanopore reads and a high-density genetic map were used to perform error correction. N50 scores are 31.9 Mb and 35.3 Mb, with 90.2% and 97.1% of the total genome assembled into the 11 pseudomolecules, for 'Rush' and 'Winkler', respectively. Gene prediction was performed using custom RNAseq libraries and protein homology data. 'Rush' has a BUSCO score of 99.0 for its assembly and 99.0 for its annotation, while 'Winkler' had corresponding scores of 96.9 and 96.5, indicating high-quality assemblies. These two independent assemblies enable unbiased assessment of structural variation within C. americana, as well as patterns of syntenic relationships across the Corylus genus. Furthermore, we identified high-density SNP marker sets from genotyping-by-sequencing data using 1343 C. americana, C. avellana and C. americana × C. avellana hybrids, in order to assess population structure in natural and breeding populations. Finally, the transcriptomes of these assemblies, as well as several other recently published Corylus genomes, were utilized to perform phylogenetic analysis of sporophytic self-incompatibility (SSI) in hazelnut, providing evidence of unique molecular pathways governing self-incompatibility in Corylus.
Cowpea, Vigna unguiculata L. Walp., is a diploid warm-season legume of critical importance as both food and fodder in sub-Saharan Africa. This species is also grown in Northern Africa, Europe, Latin America, North America, and East to Southeast Asia. To capture the genomic diversity of domesticates of this important legume, de novo genome assemblies were produced for representatives of six subpopulations of cultivated cowpea identified previously from genotyping of several hundred diverse accessions. In the most complete assembly (IT97K-499-35), 26,026 core and 4963 noncore genes were identified, with 35,436 pan genes when considering all seven accessions. GO terms associated with response to stress and defense response were highly enriched among the noncore genes, while core genes were enriched in terms related to transcription factor activity, and transport and metabolic processes. Over 5 million single nucleotide polymorphisms (SNPs) relative to each assembly and over 40 structural variants >1 Mb in size were identified by comparing genomes. Vu10 was the chromosome with the highest frequency of SNPs, and Vu04 had the most structural variants. Noncore genes harbor a larger proportion of potentially disruptive variants than core genes, including missense, stop gain, and frameshift mutations; this suggests that noncore genes substantially contribute to diversity within domesticated cowpea.