Eukaryotic genomes harbor many forms of variation, including nucleotide diversity and structural polymorphisms, which experience natural selection and contribute to genome evolution and biodiversity. Harnessing this variation for agriculture hinges on our ability to detect, quantify, catalog, and deploy genetic diversity. Here, we explore seven complete genomes of the emerging biofuel crop pennycress (Thlaspi arvense) drawn from across the species' current genetic diversity to catalog variation in genome structure and content. Across this new pangenome resource, we find contrasting evolutionary modes in different genomic zones. Gene-poor, repeat-rich pericentromeric regions experience frequent rearrangements, including repeated centromere repositioning. By contrast, conserved gene-dense chromosome arms maintain large-scale synteny across accessions even in fast-evolving NOD-like receptor immune genes, where microsynteny breaks down across species, but gene cluster positioning macrosynteny is maintained. Our findings highlight that multiple elements of the genome experience dynamic evolution that conserves functional content on the chromosome scale but allows repositioning and presence-absence variation on a local scale. This diversity is invisible to classical reference-based strategies and highlights the strength and utility of pangenomic resources. These results provide a valuable case study of rapid genomic structural evolution within a species and powerful resources for crop development in an emerging biofuel crop.
Sorghum bicolor (Sorghum) is a drought and heat tolerant C4 grass crop used to produce grain, forage, biofuels, and other bioproducts. Genetic improvement of sorghum hybrid crops is aided by a large and diverse germplasm, sorghum's diploid inbreeding genetics, and a relatively small genome that has facilitated genomic research. Over the past 20 years, the sorghum research community characterized the cytogenetic and recombinant landscapes of sorghum's 10 chromosomes, sequenced and annotated the sorghum genome, and used that information to identify genes/alleles that modulate flowering time, plant height, seed shattering, and other important traits. More recently, >1000 RNA-seq transcriptome profiles were collected from 15 sorghum genotypes to help understand the genetic basis of variation in growth and development of sorghum stems, tillers, roots, and leaves, and the regulation of biosynthetic pathways that produce epicuticular wax, dhurrin, and RFOs, compounds that contribute to sorghum's resilience. Transcriptome studies were designed to identify differentially expressed genes that are co-expressed during development or in response to a treatment to enable construction of gene regulatory networks. Co-expression and network analysis identified transcription factors and their cognate binding sites in target gene promoters and signaling pathways that modulate gene regulatory networks providing gene editing targets for further trait optimization. RNA-seq data from >20 experiments targeting sorghum organs, tissues, cell types, developmental stages, and responses to environmental conditions (i.e., diel, day-length, shading, water-deficit, temperature) has been compiled in a sorghum transcriptome compendium. The goal of this resource paper is to describe compendium content, accessibility, and a compendium data analysis pipeline and to illustrate the types of information that can be derived from the compendium with a focus on the elucidation of gene regulatory networks useful for guiding the improvement of sorghum traits through gene editing.
Characterizing population structure and admixture events between ancestral groups plays a key role in understanding the evolutionary history of species and crops. Most tools for inferring admixture have been developed for diploids and are not suitable for polyploids, in particular those with high and mixed ploidy such as Saccharum. Here we present AdmixPoly, an R-package designed to infer admixture in polyploid species both at the genome-wide scale and locally along chromosomes. We compare AdmixPoly with state-of-the-art methods using simulations, demonstrating its precision and computational efficiency. Notably, local admixture inference in complex scenarios, such as high ploidy levels, large numbers of ancestral groups and alleles per marker is enabled through efficient approximations of emission and transition probabilities within a hidden Markov model framework. We apply this approach to characterize the contributions of wild Saccharum species to the complex polyploid genome of modern sugarcane cultivars. A panel of wild and cultivated Saccharum accessions is genotyped for 80K genomic regions, each revealing approximately 50 read-scale haplotypes. The results reveal that most of the approximately 12 copies of each basic chromosome in modern cultivars are derived from the domesticated species Saccharum officinarum, with one to four copies typically contributed by distinct subgroups of the wild species Saccharum spontaneum. In addition, contributions from an unknown wild Saccharum group originating from the Pacific were identified in most cultivars. The conserved pattern of these introgressions suggests that they can be traced back to the early stages of sugarcane breeding approximately a century ago.
Background Sorghum ( Sorghum bicolor ) is a versatile C4 crop used for food and feed and as biomass for bioproducts and energy. Improving nitrogen use efficiency (NUE) in sorghum is important because fertilizer is costly and excessive fertilizer use has negative environmental impacts. Leaf senescence mediates nutrient recycling, but its dynamic progression is difficult to quantify at scale. We evaluated whether visible-near-infrared hyperspectral imaging can provide high-throughput measures of N-limitation-induced senescence in sorghum and link these phenotypes to gene expression. Sorghum Tx430 plants were grown under four N treatments (6, 9, 12, and 15 mM), imaged from vegetative growth through grain fill, and destructively sampled for RNA-seq at four developmental stages. Results A supervised support vector machine with a radial basis function kernel classified pixels from a hyperspectral image of sorghum plants grown under different N levels into green leaf, yellow leaf, dry leaf, stalk, panicle, and background classes with 0.93 accuracy. We defined the senescence ratio as the sum of yellow and dry leaf areas divided by the green leaf area and computed it across multiple growth stages and nitrogen levels. The senescence ratio did not differ among N treatments during vegetative growth, but it declined with increasing N during boot, anthesis, and grain fill, indicating earlier senescence under N limitation. Among the genes whose expression positively correlated with senescence ratio were 13 putative transcription factors, including SbiRTX430.02G247100, a WRKY1/ZAP1 homolog and a WRKY4 homolog. Gene regulatory network analysis of the top 1% of genes associated with SbiRTX430.02G247100 showed enrichment for processes associated with leaf senescence and chlorophyll catabolism. In contrast, the network associated with the WRKY4 homolog was enriched for autophagy-related terms. Conclusions Our study shows that automated hyperspectral imaging is highly effective for monitoring dynamic plant phenotypes, such as stress-induced senescence, that are difficult to visually score with the naked eye. Here, nitrogen deficiency served as the stress condition. Still, this approach supports large-scale phenotypic data collection for any such stressor and enables analyses with greater statistical power, yielding more robust conclusions and the potential for new insights that can be applied to engineering and breeding better crops.
Brachypodium is a powerful model system for investigating grass genome evolution, yet genomic resources remain concentrated in the three annual species, whereas perennial species are less sampled. Here, we present chromosome-level assemblies for two perennial species, Brachypodium mexicanum and B. arbuscula, which represent the earliest-diverging lineages of the genus and the earliest-diverging lineage of the core perennial clade, respectively. Synteny-based phylogenomics indicate that B. mexicanum is a meso-allotetraploid composed of two closely related but temporally distinct x=10 subgenomes, here designed as P and U, each carrying subgenome-specific chromosome rearrangements. We further show that the unusually large B. mexicanum genome, in contrast to the reduced genomes of most other Brachypodium species, is primarily due to transposable elements distributed across all chromosomal regions. By contrast, the diploid genome of the earliest-diverging core perennial, B. arbuscula, contains few transposable elements, whereas the most recent diverged diploid perennial B. sylvaticum shows evidence of a secondary TEs proliferation. Comparisons of lineage-specific and functionally enriched orthogroups among B. mexicanum, core perennial species and annual species suggest that ancestral hybridization between annual and perennial lineages may have contributed to the origin of allotetraploid B. mexicanum. These assemblies provide a framework for testing how polyploidy, descending dysploidy, transposable-element turnover, and life-history evolution jointly shaped genome architecture in Brachypodium.
Developing native perennial crops is vital for climate-resilient agriculture, yet their domestication is often hindered by a lack of cost-effective genomic resources. To build a framework for genomics-accelerated domestication of perennials with large, complex genomes, we generated chromosome-scale, haplotype-phased assemblies for Silphium integrifolium Michx. and Silphium perfoliatum L., two deep-rooted native North American prairie species valued for drought tolerance and oil production. The genomes are characterized by the presence of a putative helical structure preserved during interphase with a loop circumference of 43 Mb. Using targeted sequencing of 14 Silphium species, we first refined the phylogeny of the genus, recovering the Composita and Silphium subgenera while identifying taxonomic discrepancies. We used a combination of low-coverage short-read sequencing and target-sequencing to characterize a 258-accession diversity panel to define ancestral populations and perform genome-wide association studies (GWAS) on several traits. We identified 81 loci associated with environmental adaptation and domestication traits; notably, variants in a MATE transporter, Alpha-Beta hydrolase and a ACR4 -related protein explained significant variance in seed number and floral architecture. These findings establish the genomic framework necessary to accelerate the domestication of Silphium and provide a model for unlocking the potential of other complex wild genomes.
Plant stress responses occur within daily cycles of physiology, metabolism, and growth, making timing a critical dimension of acclimation. In Arabidopsis, circadian and diel regulation influence responses to abiotic stress, including cold, but how this temporal regulation is conserved, diversified, or expanded in crop genomes remains unclear. This question is especially challenging in Brassica rapa, which underwent a genome triplication after diverging from Arabidopsis, resulting in multiple retained paralogs that can be grouped by Arabidopsis orthology and ancient homeologous relationships. Here, we generated a B. rapa pangenome spanning six morphotypes and used it to profile diel (24 h) cold acclimation responses across diverse accessions differing in freeze tolerances. Cold altered peak expression time for thousands of genes, which we grouped into distinct phase-change groups. Circadian leaf movement assays revealed accession-specific differences in clock period and temperature compensation under cold, suggesting that altered clock behavior may contribute in part to the diel transcriptome retiming. At the individual gene level, inferred gene regulatory networks (GRNs) were highly accession-specific and lost shared connectivity under cold stress. However, grouping these paralogs by their Arabidopsis orthologs revealed a highly conserved regulatory architecture that was otherwise masked by paralog diversification. Integrating these networks with functional pathways identified key candidate regulators of retimed processes, including modules linked to nighttime phosphorylation and daytime photosynthesis. Finally, analyzing conserved noncoding sequences across the pangenome prioritized specific regulatory targets within cold-retimed groups. Together, these results demonstrate that cold acclimation in B. rapa is shaped by a combination of diel retiming, paralog-specific regulation, and deeply conserved programs.
Wild perennial plants can be domesticated to make agriculture more diverse and resilient, but many have large genomes that have been recalcitrant to analysis. Here, we report phased genome assemblies for Silphium integrifolium Michx. and S. perfoliatum L., two species native to North America under domestication, and demonstrate the utility of trio-binning for genome assembly using an interspecific hybrid. These genomes have chromosomes reaching 1.8 Gb and a helical structure preserved during interphase with a loop circumference of 43 Mb. A genome-informed low coverage and target sequencing strategy enables the refinement of the genus phylogeny, reveals the spatial distribution and structure of natural populations, and identifies 81 loci associated with environmental and domestication traits. Variants in a MATE transporter, α/β hydrolase, and ortholog of Arabidopsis ACT Domain Repeat (ACR4) protein explain significant variance in floral architecture. These advances in genome assembly and genotyping could expand the range of candidates for de novo crop domestication.
One hundred diatom species have been selected for genome and transcriptome sequencing. The 100 Diatom Genomes Project aims to provide a scalable framework for understanding diatom biodiversity, ecology and evolution, and for investigating their use in biotechnology.
Copy-number variants at genomic loci evolve at a high rate, are linked to many different diseases, and play a role in adaptive evolution in humans and other organisms. Here, we show that stickleback fish from freshwater environments have rapidly and repeatedly evolved an expanded number of copies of a gene family involved in muscle development, myosin heavy chain 3 cluster C (MYH3C), compared with marine populations. Differences in copy number between marine and freshwater fish are maintained even in the presence of gene flow, suggesting that MYH3C changes represent adaptive divergence between ecotypes. Copy-number expansion occurs by tandem duplication of MYH3C coding and regulatory regions on the stickleback sex chromosome. We identify a muscle regulatory enhancer within the expanded MYH3C region and show that elevated copy number is associated with developmental and tissue-specific increases in corresponding mRNA expression levels in skeletal muscle. Common MYH3C clusters include 3-, 4-, 5-, and 6-copy variants that likely evolved through a combination of microhomology-mediated break repair and non-allelic homologous recombination. Our results provide a new example of copy-number changes in a wild species and identify copy-number variations (CNVs) as potential “hotspots” of repeated adaptive evolution.
Persistent drought affects global crop production and is becoming more severe in many parts of the world in recent decades. Deciphering how plants respond to drought will facilitate the development of flexible mitigation strategies. Sorghum bicolor L. Moench (sorghum), a major cereal crop and an emerging bioenergy crop, exhibits remarkable resilience to drought. To better understand the molecular traits that underlie sorghum's remarkable drought tolerance, we undertook a large-scale sorghum gene expression profiling effort, totaling nearly 1500 transcriptome profiles, across a 3-year field study with replicated plots in California's Central Valley. This study included time-resolved gene expression data from roots and leaves of two sorghum genotypes, BTx642 and RTx430, with different pre-flowering and post-flowering drought-tolerance adaptations under control and drought conditions. Quantification of genotype-specific drought tolerance effects was enabled by de novo sequencing, assembly, and annotation of both BTx642 and RTx430 genomes. These reference-quality genomes were used to construct a pangene set for characterizing conserved and genotype-specific expression. By integrating time-resolved transcriptomic responses to drought in the field across three consecutive years, we identified a set of 726 drought-responsive genes that responded similarly in all 3 years of our field study. Functional enrichment analysis identified abiotic stress, secondary cell wall-related processes and metabolism as particularly affected under both types of drought stress. We also found that some glyoxylate cycle pathway genes, including malate synthase and isocitrate lyase, are differentially regulated particularly during post-flowering drought stress, implicating this pathway as potentially important for drought responsiveness. This expansive dataset represents a unique resource for sorghum and drought research communities and provides a methodological framework for the integration of multi-faceted time-resolved transcriptomic datasets.
Many hermaphroditic species increase outcrossing rates by partitioning reproduction so that male and female organs mature at different times, a phenomenon known as dichogamy. Previous work has documented that dichogamy in pecan trees is governed by the Mendelian super-gene 'G-locus'; however, its non-recombinant sex chromosome-like architecture has impeded quantitative genetic exploration and candidate gene discovery. Here, we probe the genetic drivers of the G-locus through a pangenome-integrated quantitative genomics experiment. We first provide one of the clearest examples to date of mapping bias, where a linear reference-based GWAS discovered 66 off-target peaks while mapping with a pangenome graph reference resolved the known single Mendelian locus. Across six new genome assemblies, the fully haplotype phased G-locus QTL spanned 223-491kb and included 25 candidate gene families. The strongest candidate gene encoded a MATE efflux protein and had dominant allele-specific action during male flower developmental stages. Combined, these candidates and genomic resources provide a powerful foundation for breeding and optimal dichogamy phenotype engineering for future pecan orchards.
Benzylisoquinoline alkaloids (BIAs) represent a vast group of specialized plant metabolites with diverse pharmaceutical applications, synthesized by a variety of gene families. Among the multiple plant lineages that produce BIAs, the most notable is the poppy family (Papaveraceae), with California poppy (Eschscholzia californica) emerging as a model organism. Here, we report a haplotype-resolved genome assembly, in combination with a high-density expression atlas, for California poppy. Genome analyses reveal recent diversification of BIA biosynthesis genes in poppy through localized duplications. Furthermore, we demonstrate that the degree of phylogenetic relatedness among paralogs within BIA biosynthesis-associated gene families correlates with similarities in gene expression. In contrast, gene families involved in carotenoid biosynthesis, which contributes to the intense orange petal pigmentation, are not phylogenetically clustered, and floral developmental regulators exhibit a high degree of retention of gene duplicates associated with ancient polyploidy events. These findings illustrate alternative roles for gene and genome duplications as drivers of trait evolution. Given the position of California poppy in the angiosperm phylogeny, the high-quality genomic resources generated for this work constitute a valuable resource for comparative genomic and transcriptomic analyses for poppies and flowering plants more generally.
The legume family originated ca. 60-65 million years ago and soon diversified into at least six lineages (now extant subfamilies). The signal of whole genome duplications (WGD) is apparent in species sampled from all six subfamilies. The early diversification has posed difficulties for resolving the legume backbone structure and the timing of WGDs, especially in Caesalpinioideae where the diversification and WGD signals coincide. In this study, we report the genome sequences and annotations for Cercis canadensis (Cercidoideae) and Chamaecrista fasciculata (Caesalpinioideae) to help resolve the timings of WGDs relative to subfamily origins and the ancestral legume karyotype. Analyses of genome assemblies from four subfamilies within Fabaceae show that the last common ancestor of all legumes likely had seven chromosomes, with a genome structure similar to the extant Cercis genome. The retained karyotype structure, the lack of a WGD in the last 100+ Mya (Cercis and the lineage leading to it following the eudicot γ whole-genome triplication), and the unusually slow rates of nucleotide substitution and structural evolution in the Cercis genome underscore its utility as a genomic proxy for the last common ancestor of all legume species. Our analysis supports an allopolyploid origin of Caesalpinioideae, with progenitors from lineages along the backbone of the legume phylogeny. Rapid diversification and the inferred allopolyploid origin of Caesalpinioideae provide a partial explanation for the difficulty in resolving the backbone of the legume phylogeny and early Caesalpinioideae diversification.
More than a century after two introduced pathogens killed billions of American chestnut trees, introgression of resistance alleles from Chinese chestnuts has contributed to the recovery of self-sustaining populations. However, progress has been slow because of the complex genetic architecture of resistance. To better understand blight resistance, we compared reference genomes, gene expression responses, and stem metabolite profiles of the resistant Chinese and susceptible American chestnut species. To accelerate resistance breeding, we conducted large-scale phenotyping and genotyping in hybrids of these species. Simulation and inoculation experiments suggest that significant resistance gains are possible through selectively breeding trees with an average of 70 to 85% American chestnut ancestry. The resources developed in this work are foundational for breeding to create diverse restoration populations with sufficient disease resistance and competitive growth.
The model woody plant Populus trichocarpa displays an atypical alkene-diverse wax cuticle likely driven by copy number variation (CNV) of 3-ketoacyl-CoA synthases (KCS), which has been difficult to confirm with short-read assemblies. Long-read sequencing enables the development of telomere-to-telomere resources to detect cryptic variation, including CNVs, which are currently missed. Integrating this information can improve genomic prediction for breeding and provide insights into the evolutionary basis of important traits. Our analysis of 78 long-read haplotypes from chromosome 10 identified more than twice as many KCS genes as previously reported, and numerous intragenic non-synonymous substitutions. Random Forest predictive models highlighted the importance of Potri.010G079500 in producing very long chain alkenes; however, its absence did not predict previously reported alkene-deficient phenotypes. Instead, alkene levels are best predicted by the combinations of KCS copies. Additionally, amino acid substitutions clustered around ligand and donor binding pockets, suggesting they contribute to differing wax cuticle composition. Finally, each KCS gene and copy was linked to a Helitron transposon. A phylogenetic analysis suggests Helitrons are the evolutionary mechanism for generating KCS tandem arrays. Long-read generated telomereto-telomere assemblies of P. trichocarpa chromosome 10 revealed large-effect loci critical to genetic studies that are unattainable from short-reads. This new resource produced novel insights into genome structure and function, and a novel mechanism for generating tandem gene duplication. Our results highlight that, given current challenges in annotation and assembly, detailed and focused long-read sequences are key to interpreting complex genomic regions that contain tandem copy number variants.
Although the green revolution adapted a handful of crops to homogeneous and high-input industrialized agriculture, much of the global population still relies on the local production of variable crop cultivars by low-input smallholder farms. This diversity of unhomogenized crops1, like that of the grain and bioenergy crop sorghum2-5, offers raw materials for genetic gain and cultivar improvement. However, breeding efforts can be constrained by highly specialized traits and breeding targets6. Here, to bridge this diversity, we constructed a 33-member pangenome reference and a diversity panel across 1,984 cultivars and landraces. We leveraged these resources to explore the complex interplay among historical contingency, ongoing adaptation and previously uncharacterized structural diversity. Specifically, our analyses conclusively demonstrated multiple nested and deeply diverged structural variants in the domestication gene SHATTERING1, which distinguish the previously established multicentric origin of sorghum. We then applied landscape genomics to reveal how gene flow and secondary contact created the complex genetic mosaic in contemporary breeding networks. As proof of concept for pangenome-accelerated trait discovery, we connected biosynthetic gene cluster structural variation to phenotypic leaf concentration of the cyanogenic glucoside dhurrin. Combined, these approaches will accelerate breeding and trait discovery and provide a framework for similar applications in other crops.
Crassulacean acid metabolism (CAM) is a specialized photosynthetic pathway that enhances water-use efficiency by temporally separating nocturnal CO2 uptake from daytime decarboxylation and carbon fixation. To uncover the regulatory mechanisms coordinating these temporal dynamics, we generated high-resolution, 48 h time-course transcriptomes for the CAM model Kalanchoe fedtschenkoi under both 12 h/12 h light/dark (LD) cycles and continuous light (LL). A rhythmicity analysis revealed that diel light cues are the dominant driver of transcript oscillations: 16,810 genes (54.3% of annotated genes) exhibited rhythmic expression only under LD, whereas just 399 genes (1.3%) remained rhythmic under LL. A smaller set of 3009 genes (9.7%) oscillated in both conditions, indicating that the intrinsic circadian clock sustains rhythmicity for a limited subset of the transcriptome. A gene co-expression network analysis revealed extensive integration between circadian clock components, core CAM pathway enzymes, and stomatal regulators, defining regulatory modules that coordinate metabolic and physiological timing. Notably, key hub genes associated with post-translational and post-transcriptional regulation, including the E3 ubiquitin ligase HUB2 and several pentatricopeptide repeat (PPR) proteins, act as central nodes in CAM-associated networks. This discovery implicates epigenetic and organellar regulation as previously unrecognized critical tiers of control in CAM. Together, our results support a regulatory model in which CAM rhythmicity is governed by both external light/dark cues and the endogenous circadian clock through multi-level control spanning transcriptional and protein-level regulation. To support community exploration, we also provide an interactive eFP (electronic Fluorescent Pictograph) browser for visualizing time-resolved gene expression profiles.
Increasing the genomic resources of emerging aquaculture crop targets can expedite breeding processes as seen in molecular breeding advances in agriculture. High quality annotated reference genomes are essential to implement this relatively new molecular breeding scheme and benefit research areas such as population genetics, gene discovery, and gene mechanics by providing a tool for standard comparison. The brown macroalga Saccharina latissima (sugar kelp) is an ecologically and economically important kelp that is found in both the northern Pacific and Atlantic Oceans. Cultivation of Saccharina latissima for human consumption has increased significantly this century in both North America and Europe, and its single blade morphology allows for dense seeding practices used in the cultivation of its Asian sister species, Saccharina japonica. While Saccharina latissima has potential as a human food crop, insufficient information from genetic resources has limited molecular breeding in sugar kelp aquaculture. We present scaffolded and annotated Saccharina latissima nuclear and organelle genomes from a female gametophyte collected from Black Ledge, Groton, Connecticut. This Saccharina latissima genome compares well with other published kelp genomes and contains 218 scaffolds with a scaffold N50 of 1.35 Mb, a GC content of 49.84%, and 25,012 predicted genes. We also validated this genome by comparing the synteny and completeness of this Saccharina latissima genome to other kelp genomes. Our team has successfully performed initial genomic selection trials with sugar kelp using a draft version of this genome. This Saccharina latissima genome expands the genetic toolkit for the economically and ecologically important sugar kelp and will be a fundamental resource for future foundational science, breeding, and conservation efforts.
Phaeocystales, comprising the genus Phaeocystis and an uncharacterized sister lineage, are nanoplanktonic haptophytes widespread in the global ocean. Several species form mucilaginous colonies and influence key biogeochemical cycles, yet their underlying diversity and ecological strategies remain underexplored. Here, we present new genomic data from 13 strains, including three high-quality reference genomes (N50 > 30 kbp), and integrate previous metagenome-assembled genomes to resolve a robust phylogeny. Divergence timing of P. antarctica aligns with Miocene cooling and Southern Ocean isolation. Genomic traits reveal metabolic flexibility, including mixotrophic nitrogen acquisition in temperate waters and gene expansions linked to polar nutrient adaptation. Concordantly, transcriptomic comparisons between temperate and polar Phaeocystis suggest Southern Ocean populations experience iron and B12 limitation. We also identify signatures of horizontal gene transfer and endogenous giant virus/virophage insertions. Together, these findings highlight Phaeocystales as an ecologically versatile and geographically widespread lineage shaped by evolutionary innovation and adaptation to contrasting environmental stressors.