More than a century after two introduced pathogens killed billions of American chestnut trees, introgression of resistance alleles from Chinese chestnuts has contributed to the recovery of self-sustaining populations. However, progress has been slow because of the complex genetic architecture of resistance. To better understand blight resistance, we compared reference genomes, gene expression responses, and stem metabolite profiles of the resistant Chinese and susceptible American chestnut species. To accelerate resistance breeding, we conducted large-scale phenotyping and genotyping in hybrids of these species. Simulation and inoculation experiments suggest that significant resistance gains are possible through selectively breeding trees with an average of 70 to 85% American chestnut ancestry. The resources developed in this work are foundational for breeding to create diverse restoration populations with sufficient disease resistance and competitive growth.
Although the green revolution adapted a handful of crops to homogeneous and high-input industrialized agriculture, much of the global population still relies on the local production of variable crop cultivars by low-input smallholder farms. This diversity of unhomogenized crops1, like that of the grain and bioenergy crop sorghum2-5, offers raw materials for genetic gain and cultivar improvement. However, breeding efforts can be constrained by highly specialized traits and breeding targets6. Here, to bridge this diversity, we constructed a 33-member pangenome reference and a diversity panel across 1,984 cultivars and landraces. We leveraged these resources to explore the complex interplay among historical contingency, ongoing adaptation and previously uncharacterized structural diversity. Specifically, our analyses conclusively demonstrated multiple nested and deeply diverged structural variants in the domestication gene SHATTERING1, which distinguish the previously established multicentric origin of sorghum. We then applied landscape genomics to reveal how gene flow and secondary contact created the complex genetic mosaic in contemporary breeding networks. As proof of concept for pangenome-accelerated trait discovery, we connected biosynthetic gene cluster structural variation to phenotypic leaf concentration of the cyanogenic glucoside dhurrin. Combined, these approaches will accelerate breeding and trait discovery and provide a framework for similar applications in other crops.
While the green revolution adapted a handful of crops to homogenous and high-input industrialized agriculture, much of the global population still relies on local food production from low-input smallholder farms that grow highly variable crop cultivars. The high diversity of the grain and bioenergy crop sorghum [1][1]–[4][2], and many other crops that were not homogenized during the green revolution [5][3], not only provides the raw materials for breeders to make substantial gains in cultivar improvement, but also constrains breeding efforts due to highly specialized locally adapted plant phenotypes [6][4]. Here, we construct a 33-member pangenome and identify trait-associated variants in 1,988 cultivars and landraces. We then apply these resources to explore the complex interplay between historical contingency, ongoing adaptation, and the potential for future gains through climate-aware genome-enabled breeding. Specifically, our analyses conclusively demonstrate that multiple nested, deeply diverged, and previously uncharacterized structural variants in the domestication gene SHATTERING1 distinguish the previously established multicentric origin of sorghum. We then apply landscape genomics tests to reveal how gene flow, adaptation, and secondary contact created the complex genetic mosaic in current global breeding networks. Further analysis of climate-gene associations highlights candidate loci underlying adaptation, including the biosynthetic gene cluster for the cyanogenic glucoside dhurrin. Combined, the pangenome-informed variants developed here will enable both trait discovery and subsequent marker assays to accelerate breeding and provide a framework for similar applications in other diverse and non-model crops. ### Competing Interest Statement The authors have declared no competing interest. United States Department of Energy, DE-AC02-05CH11231 [1]: #ref-1 [2]: #ref-4 [3]: #ref-5 [4]: #ref-6
Over a century after two introduced pathogens decimated American chestnut populations, breeding programs continue to incorporate resistance from Chinese chestnut to recover self-sustaining populations. Due to complex genetics of chestnut blight resistance, it is challenging to obtain trees with sufficient resistance and competitive growth. We developed high quality reference genomes for Chinese and American chestnut and leveraged large disease phenotype and genotype datasets to develop accurate genomic selection. Inoculation and simulation results indicate that resistance may be substantially increased in trees that inherited 70% to 100% of their genome from American chestnut. To facilitate gene editing, we integrated multiple lines of evidence to discover candidate alleles for blight resistance and susceptibility. These genomic resources provide a strong foundation to accelerate restoration of this iconic tree. ### Competing Interest Statement The authors have declared no competing interest.
Yellow monkeyflowers (Mimulus guttatus complex, Phrymaceae) are a powerful system for studying ecological adaptation, reproductive variation, and genome evolution. To initiate pan-genomics in this group, we present four chromosome-scale assemblies and annotations of accessions spanning a broad evolutionary spectrum: two from a single M. guttatus population, one from the closely related selfing species M. nasutus, and one from a more divergent species M. tilingii. All assemblies are highly complete and resolve centromeric and repetitive regions. Comparative analyses reveal such extensive structural variation in repeat-rich, gene-poor regions that large portions of the genome are unalignable across accessions. As a result, this Mimulus pan-genome is primarily informative in genic regions, underscoring limitations of resequencing approaches in such polymorphic taxa. We document gene presence-absence, investigate the recombination landscape using high-resolution linkage data, and quantify nucleotide diversity. Surprisingly, pairwise differences at fourfold synonymous sites are exceptionally high-even in regions of very low recombination-reaching ~3.2% within a single M. guttatus population, ~7% within the interfertile M. guttatus species complex (approximately equal to SNP divergence between great apes and Old World monkeys), and ~7.4% between that complex and the reproductively isolated M. tilingii. Genome-wide patterns of nucleotide variation show little evidence of linked selection, and instead suggest that the concentration of genes (and likely selected sites) in high-recombination regions may buffer diversity loss. These assemblies, annotations, and comparative analyses provide a robust genomic foundation for Mimulus research and offer new insights into the interplay of recombination, structural variation, and molecular evolution in highly diverse plant genomes.
Cultivar 'Williams 82' has served as the reference genome for the soybean research community since 2008, but is known to have areas of genomic heterogeneity among different sub-lines. This work provides an updated assembly (version Wm82.a6) derived from a specific sub-line known as 'Wm82-ISU-01' (seeds available under USDA accession PI 704477). The genome was assembled using Pacific BioSciences HiFi reads and integrated into chromosomes using HiC. The 20 soybean chromosomes assembled into a genome of 1.01Gb, consisting of 36 contigs. The genome annotation identified 48,387 gene models, named in accordance with previous assembly versions Wm82.a2 and Wm82.a4. Comparisons of Wm82.a6 with other near-gapless assemblies of 'Williams 82' reveal large regions of genomic heterogeneity, including regions of differential introgression from the genotype 'Kingwa' within approximately 30 Mb and 25 Mb segments on chromosomes 03 and 07, respectively. Additionally, our analysis revealed a previously unknown large (~20 Mb) heterogeneous region in the pericentromeric region of chromosome 12, where Wm82.a6 matches the 'Williams' haplotype while the other two near-gapless assemblies do not match the haplotype of either parent of 'Williams 82'. In addition to the Wm82.a6 assembly, we also assembled the genome of soybean line 'Fiskeby III', a rich resource for abiotic stress resistance genes. A genome comparison of Wm82.a6 with 'Fiskeby III' revealed the nucleotide and structural polymorphisms between the two genomes within a QTL region for iron deficiency chlorosis resistance. The Wm82.a6 and 'Fiskeby III' genomes described here will enhance comparative and functional genomics capacities and applications in the soybean community.
The grass family (Poaceae, Poales) holds immense economic and ecological significance, exhibiting unique metabolic traits, including dual starch and lignin biosynthetic pathways. To investigate when and how the metabolic innovations known in grasses evolved, we sequenced the genomes of a non-core grass, Pharus latifolius, the non-grass graminids, Joinvillea ascendens and Ecdeiocolea monostachya, representing the sister clade to Poaceae, and Typha latifolia, representing the sister clade to the remaining Poales. The rho whole genome duplication (ρWGD) in the ancestral lineage for all grasses contributed to the gene family expansions underlying cytosolic starch biosynthesis, whereas an earlier tandem duplication of phenylalanine ammonia lyase (PAL) gave rise to phenylalanine/tyrosine ammonia lyase (PTAL) responsible for the dual lignin biosynthesis. Two mutations were sufficient to expand ancestral PAL function into PTAL. The integrated genomic and biochemical analyses of grass relatives in Poales revealed the evolutionary and molecular basis of key metabolic innovations of grasses. ### Competing Interest Statement Y.T-K, B.M., and H.A.M. have a pending patent application related to the two mutations that can convert PAL enzymes into PTAL enzymes. The other authors declare that they have no competing interests related to this work.
Stress-sensitive and stress-adapted plants respond differently to environmental stresses. To explore the cellular-level stress adaptations, we built root single-cell transcriptome atlases for diverse Brassicaceae species: stress-sensitive plants ( Arabidopsis thaliana and Sisymbrium irio ), extremophytes ( Eutrema salsugineum and Schrenkiella parvula ) and a polyploid crop ( Camelina sativa ), under control, NaCl, and abscisic acid treatments. Approximately half of Arabidopsis cell-type markers lacked expression conservation across species. We identified new conserved cell-type markers, along with orthologs showing divergent expressions. We experimentally mapped distinct cortex sub-populations to different cortex layers across species. We found distinct cell-type-specific transcriptomic responses between species and treatments. Lineage-specific losses of stress responses were less prevalent but evolutionarily more favored than gains. In C. sativa , sub-genomes contributed equally to stress responses and homeologs with divergent stress responses typically did not exhibit high coding sequence or expression divergence. Our study provides a foundational root atlas and an analytical framework for multi-species single-cell transcriptomics. ### Competing Interest Statement The authors have declared no competing interest.
### Competing Interest Statement The authors have declared no competing interest.
Cotton (Gossypium hirsutum L.) is the key renewable fibre crop worldwide, yet its yield and fibre quality show high variability due to genotype-specific traits and complex interactions among cultivars, management practices and environmental factors. Modern breeding practices may limit future yield gains due to a narrow founding gene pool. Precision breeding and biotechnological approaches offer potential solutions, contingent on accurate cultivar-specific data. Here we address this need by generating high-quality reference genomes for three modern cotton cultivars (‘UGA230’, ‘UA48’ and ‘CSX8308’) and updating the ‘TM-1’ cotton genetic standard reference. Despite hypothesized genetic uniformity, considerable sequence and structural variation was observed among the four genomes, which overlap with ancient and ongoing genomic introgressions from ‘Pima’ cotton, gene regulatory mechanisms and phenotypic trait divergence. Differentially expressed genes across fibre development correlate with fibre production, potentially contributing to the distinctive fibre quality traits observed in modern cotton cultivars. These genomes and comparative analyses provide a valuable foundation for future genetic endeavours to enhance global cotton yield and sustainability.
Phaeocystis is a genus of nanoplanktonic haptophytes, prevalent in all of the world’s oceans. At least three Phaeocystis species form large mucilaginous colonies that contribute to high-biomass blooms, making them major players in biogeochemical cycles, especially of carbon and sulfur. To investigate the ecological success of these primary producers, we assembled genomic data for thirteen Phaeocystis strains and annotated three with the highest contiguity. Using these data and additional metagenome-assembled genomes, we present a robust phylogeny for Phaeocystis and their sister clade, identifying several undescribed but abundant lineages. We distinguish two major ecological distribution patterns, cosmopolitan and polar, with fine-tuned preferences for nutrients, temperature, and motility. The gene repertoires of three annotated Phaeocystis reference genomes reflect a unique strategy of gene expansion among algae: repetitive elements, horizontal gene transfer, and full-length endogenous virus insertions. Phaeocystales thus emerge as an ecologically versatile group with diverse adaptations to biotic and abiotic stressors.
Five versions of the Chlamydomonas reinhardtii reference genome have been produced over the last two decades. Here we present version 6, bringing significant advances in assembly quality and structural annotations. PacBio-based chromosome-level assemblies for two laboratory strains, CC-503 and CC-4532, provide resources for the plus and minus mating-type alleles. We corrected major misassemblies in previous versions and validated our assemblies via linkage analyses. Contiguity increased over ten-fold and >80% of filled gaps are within genes. We used Iso-Seq and deep RNA-seq datasets to improve structural annotations, and updated gene symbols and textual annotation of functionally characterized genes via extensive manual curation. We discovered that the cell wall-less classical reference strain CC-503 exhibits genomic instability potentially caused by deletion of the helicase RECQ3, with major structural mutations identified that affect >100 genes. We therefore present the CC-4532 assembly as the primary reference, although this strain also carries unique structural mutations and is experiencing rapid proliferation of a Gypsy retrotransposon. We expect all laboratory strains to harbor gene-disrupting mutations, which should be considered when interpreting and comparing experimental results. Collectively, the resources presented here herald a new era of Chlamydomonas genomics and will provide the foundation for continued research in this important reference organism.
Gene functional descriptions offer a crucial line of evidence for candidate genes underlying trait variation. Conversely, plant responses to environmental cues represent important resources to decipher gene function and subsequently provide molecular targets for plant improvement through gene editing. However, biological roles of large proportions of genes across the plant phylogeny are poorly annotated. Here we describe the Joint Genome Institute (JGI) Plant Gene Atlas, an update able data resource consisting of transcript abundance assays spanning 18 diverse species. To integrate across these diverse genotypes, we anal yzed e xpression pr ofiles, b uilt gene c lusters that exhibited tissue / condition specific expression, and tested for transcriptional response to environmental queues. We discovered extensive phylogenetically constrained and condition-specific expression profiles f or genes without an y previously documented functional annotation. Such conserved expression patterns and tightly co-expressed gene clusters let us assign expression derived additional biological information to 64 495 genes with otherwise unknown functions. The ever-expanding Gene Atlas resource is available at JGI Plant Gene Atlas ( https://plantgeneatlas.jgi.doe.gov ) and Phytozome ( https://phytozome .jgi.doe .gov/), providing bulk access to data and user-specified queries of gene sets. Combined, these web interfaces let users access differentiall y e xpressed genes, track orthologs across the Gene Atlas plants, graphically represent co-expressed genes, and visualize gene ontology and pathway enrichments.
Cowpea, Vigna unguiculata L. Walp., is a diploid warm-season legume of critical importance as both food and fodder in sub-Saharan Africa. This species is also grown in Northern Africa, Europe, Latin America, North America, and East to Southeast Asia. To capture the genomic diversity of domesticates of this important legume, de novo genome assemblies were produced for representatives of six subpopulations of cultivated cowpea identified previously from genotyping of several hundred diverse accessions. In the most complete assembly (IT97K-499-35), 26,026 core and 4963 noncore genes were identified, with 35,436 pan genes when considering all seven accessions. GO terms associated with response to stress and defense response were highly enriched among the noncore genes, while core genes were enriched in terms related to transcription factor activity, and transport and metabolic processes. Over 5 million single nucleotide polymorphisms (SNPs) relative to each assembly and over 40 structural variants >1 Mb in size were identified by comparing genomes. Vu10 was the chromosome with the highest frequency of SNPs, and Vu04 had the most structural variants. Noncore genes harbor a larger proportion of potentially disruptive variants than core genes, including missense, stop gain, and frameshift mutations; this suggests that noncore genes substantially contribute to diversity within domesticated cowpea.
Perennial grasses are important forage crops and emerging biomass crops and have the potential to be more sustainable grain crops. However, most perennial grass crops are difficult experimental subjects due to their large size, difficult genetics, and/or their recalcitrance to transformation. Thus, a tractable model perennial grass could be used to rapidly make discoveries that can be translated to perennial grass crops. Brachypodium sylvaticum has the potential to serve as such a model because of its small size, rapid generation time, simple genetics, and transformability. Here, we provide a high-quality genome assembly and annotation for B. sylvaticum, an essential resource for a modern model system. In addition, we conducted transcriptomic studies under 4 abiotic stresses (water, heat, salt, and freezing). Our results indicate that crowns are more responsive to freezing than leaves which may help them overwinter. We observed extensive transcriptional responses with varying temporal dynamics to all abiotic stresses, including classic heat-responsive genes. These results can be used to form testable hypotheses about how perennial grasses respond to these stresses. Taken together, these results will allow B. sylvaticum to serve as a truly tractable perennial model system.
The nutrient-rich tubers of the greater yam, Dioscorea alata L., provide food and income security for millions of people around the world. Despite its global importance, however, greater yam remains an orphan crop. Here, we address this resource gap by presenting a highly contiguous chromosome-scale genome assembly of D. alata combined with a dense genetic map derived from African breeding populations. The genome sequence reveals an ancient allotetraploidization in the Dioscorea lineage, followed by extensive genome-wide reorganization. Using the genomic tools, we find quantitative trait loci for resistance to anthracnose, a damaging fungal pathogen of yam, and several tuber quality traits. Genomic analysis of breeding lines reveals both extensive inbreeding as well as regions of extensive heterozygosity that may represent interspecific introgression during domestication. These tools and insights will enable yam breeders to unlock the potential of this staple crop and take full advantage of its adaptability to varied environments.
The "genomic shock" hypothesis posits that unusual challenges to genome integrity such as whole genome duplication may induce chaotic genome restructuring. Decades of research on polyploid genomes have revealed that this is often, but not always the case. While some polyploids show major chromosomal rearrangements and derepression of transposable elements in the immediate aftermath of whole genome duplication, others do not. Nonetheless, all polyploids show gradual diploidization over evolutionary time. To evaluate these hypotheses, we produced a chromosome-scale reference genome for the natural allotetraploid grass Brachypodium hybridum, accession "Bhyb26." We compared 2 independently derived accessions of B. hybridum and their deeply diverged diploid progenitor species Brachypodium stacei and Brachypodium distachyon. The 2 B. hybridum lineages provide a natural timecourse in genome evolution because one formed 1.4 million years ago, and the other formed 140 thousand years ago. The genome of the older lineage reveals signs of gradual post-whole genome duplication genome evolution including minor gene loss and genome rearrangement that are missing from the younger lineage. In neither B. hybridum lineage do we find signs of homeologous recombination or pronounced transposable element activation, though we find evidence supporting steady post-whole genome duplication transposable element activity in the older lineage. Gene loss in the older lineage was slightly biased toward 1 subgenome, but genome dominance was not observed at the transcriptomic level. We propose that relaxed selection, rather than an abrupt genomic shock, drives evolutionary novelty in B. hybridum, and that the progenitor species' similarity in transposable element load may account for the subtlety of the observed genome dominance.
The development of multiple chromosome-scale reference genome sequences in many taxonomic groups has yielded a high-resolution view of the patterns and processes of molecular evolution. Nonetheless, leveraging information across multiple genomes remains a significant challenge in nearly all eukaryotic systems. These challenges range from studying the evolution of chromosome structure, to finding candidate genes for quantitative trait loci, to testing hypotheses about speciation and adaptation. Here, we present GENESPACE, which addresses these challenges by integrating conserved gene order and orthology to define the expected physical position of all genes across multiple genomes. We demonstrate this utility by dissecting presence–absence, copy-number, and structural variation at three levels of biological organization: spanning 300 million years of vertebrate sex chromosome evolution, across the diversity of the Poaceae (grass) plant family, and among 26 maize cultivars. The methods to build and visualize syntenic orthology in the GENESPACE R package offer a significant addition to existing gene family and synteny programs, especially in polyploid, outbred, and other complex genomes.
ABSTRACTGene functional descriptions, which are typically derived from sequence similarity to experimentally validated genes in a handful of model species, offer a crucial line of evidence when searching for candidate genes that underlie trait variation. Plant responses to environmental cues, including gene expression regulatory variation, represent important resources for understanding gene function and crucial targets for plant improvement through gene editing and other biotechnologies. However, even after years of effort and numerous large-scale functional characterization studies, biological roles of large proportions of protein coding genes across the plant phylogeny are poorly annotated. Here we describe the Joint Genome Institute (JGI) Plant Gene Atlas, a public and updateable data resource consisting of transcript abundance assays from 2,090 samples derived from 604 tissues or conditions across 18 diverse species. We integrated across these diverse conditions and genotypes by analyzing expression profiles, building gene clusters that exhibited tissue/condition specific expression, and testing for transcriptional modulation in response to environmental queues. For example, we discovered extensive phylogenetically constrained and condition-specific expression profiles across many gene families and genes without any functional annotation. Such conserved expression patterns and other tightly co-expressed gene clusters let us assign expression derived functional descriptions to 64,620 genes with otherwise unknown functions. The ever-expanding Gene Atlas resource is available at JGI Plant Gene Atlas (https://plantgeneatlas.jgi.doe.gov) and Phytozome (https://phytozome-next.jgi.doe.gov), providing bulk access to data and user-specified queries of gene sets. Combined, these web interfaces let users access differentially expressed genes, track orthologs across the Gene Atlas plants, graphically represent co-expressed genes, and visualize gene ontology and pathway enrichments.
The development of multiple high-quality reference genome sequences in many taxonomic groups has yielded a high-resolution view of the patterns and processes of molecular evolution. Nonetheless, leveraging information across multiple reference haplotypes remains a significant challenge in nearly all eukaryotic systems. These challenges range from studying the evolution of chromosome structure, to finding candidate genes for quantitative trait loci, to testing hypotheses about speciation and adaptation in nature. Here, we address these challenges through the concept of a pan-genome annotation, where conserved gene order is used to restrict gene families and define the expected physical position of all genes that share a common ancestor among multiple genome annotations. By leveraging pan-genome annotations and exploring the underlying syntenic relationships among genomes, we dissect presence-absence and structural variation at four levels of biological organization: among three tetraploid cotton species, across 300 million years of vertebrate sex chromosome evolution, across the diversity of the Poaceae (grass) plant family, and among 26 maize cultivars. The methods to build and visualize syntenic pan-genome annotations in the GENESPACE R package offer a significant addition to existing gene family and synteny programs, especially in polyploid, outbred and other complex genomes. ### Competing Interest Statement The authors have declared no competing interest.