Abstract Background Pangenomes are a promising new approach to genomics that can reduce reference bias in genotyping, but the reliability of such a data model remains unclear in tracking variation across species. To test the utility of graph-based pangenomes for interspecific breeding, we developed a Minigraph-Cactus super-pangenome representing four Citrus species derived from the founder lines of a citrus breeding program. To benchmark SNP calling accuracy using graph and linear-based approaches, we performed whole genome short read sequencing for two sets of pedigreed progeny: 30 F1 hybrids and 244 advanced hybrids from an F1 crossed with a parent not included in the pangenome. Results The linear approach yielded more SNP calls than the graph-based approach, however, both methods exhibited similar Mendelian Inheritance Error Rates (MIER) in a tool-dependent manner. Reconstruction of parental haplotype blocks in the advanced hybrids revealed a striking improvement in performance in the pangenome graph-based calls, suggesting MIER is vulnerable to error when reference bias influences both parental and progeny genotype calls. Masking of regions diverged from the reference path improved MIER accuracy metrics and haplotype block reconstruction in both the linear and graph-based SNP calls. Conclusions In non-model systems, inheritance patterns observed from pedigreed hybrids provide a framework for benchmarking variant-calling accuracy using pangenomes. SNP miscalls originating from diverged regions can falsely satisfy MIER filters, thus we recommend haplotype blocks. The inherent structure of the pangenome graph has promising applications for removing regions of unreliable mapping quality, which cannot otherwise be reliably removed using traditional filtering metrics.
Advances in agricultural genetic, genomic, and breeding (GGB) technologies generate increasingly large and complex datasets that need to be adequately managed and shared. While several agricultural biological databases maintain and curate GGB data, not all scientists are aware of them and how they can be used to access and share data. In addition, there is the need to increase scientists’ awareness that appropriate data archiving and curation increases data longevity and value and bolsters scientific discoveries’ reproducibility and transparency. The AgBioData Education working group aims to address these unmet needs and developed a modular curriculum for educators teaching the basics of biological databases and the findable, accessible, interoperable, and reusable (FAIR) principles to undergraduate and graduate students (https://www.agbiodata.org/). The present paper provides an overview of the topics covered within the curriculum, called ‘AgBioData Curriculum for Ag FAIR Data,’ its audience and modalities, and how it will positively impact all the different stakeholders of the agricultural database ecosystem. We hope the modular curriculum presented here can help scientists and students understand and support database use in all aspects of improving our global food system. Database URL: https://zenodo.org/records/14278084
White oak (Quercus alba) is an abundant forest tree species across eastern North America that is ecologically, culturally, and economically important. We report the first haplotype-resolved chromosome-scale genome assembly of Q. alba and conduct comparative analyses of genome structure and gene content against other published Fagaceae genomes. We investigate the genetic diversity of this widespread species and the phylogenetic relationships among oaks using whole genome data. Despite strongly conserved chromosome synteny and genome size across Quercus, certain gene families have undergone rapid changes in size, including defense genes. Unbiased annotation of resistance (R) genes across oaks revealed that the overall number of R genes is similar across species - as are the chromosomal locations of R gene clusters - but, gene number within clusters is more labile. We found that Q. alba has high genetic diversity, much of which predates its divergence from other oaks and likely impacts divergence time estimations. Our phylogenetic results highlight widespread phylogenetic discordance across the genus. The white oak genome represents a major new resource for studying genome diversity and evolution in Quercus. Additionally, we show that unbiased gene annotation is key to accurately assessing R gene evolution in Quercus.
Motivation High-quality genome assemblies and pangenomes are increasingly accessible and achievable due to advances in third-generation sequencing and assembly algorithms, but genome annotation remains a critical bottleneck. Existing gene annotation pipelines often require complex installations, multiple steps, long runtimes, and produce variable results, which impede quality and the downstream usage of the annotations. Results We developed RAGNAROK (RApid GeNe Annotation ROcKs), a modular annotation pipeline built on Nextflow and Apptainer that integrates ab initio prediction, transcriptome and protein evidence, repeat annotation, functional assignment, and quality assessment into a single reproducible workflow. Moreover, RAGNAROK utilizes GPU acceleration and parallelization to produce high-quality gene annotations. We benchmarked against BRAKER3 across five diverse, reference-quality plant genomes and demonstrated that RAGNAROK achieved higher sensitivity, precision, and F1 scores at the exon, transcript, and gene levels. Furthermore, we demonstrated RAGNAROK’s improvement in re-annotating a suite of five Rosaceae genomes that were previously annotated using MAKER. Overall, RAGNAROK produced more ideal mono:multi-exonic gene ratios, improved BUSCO completeness scores, and reduced missing gene content compared to other pipelines. Additionally, RAGNAROK consistently outperformed BRAKER3 in runtime, scaling efficiently from small to gigabase-scale genomes. RAGNAROK provides a flexible, rapid, scalable, and accurate solution for de novo and re-annotation of plant genomes. Its modular design and workflow scalability lay the foundation for future extensions to animal, fungal, and other eukaryotic genomes. Availability and Implementation RAGNAROK is available as a GitHub repository at . ### Competing Interest Statement The authors have declared no competing interest.
Huanglongbing (HLB) is a devastating citrus disease that threatens the citrus industry worldwide. HLB is associated with the bacteria Candidatus Liberibacter asiaticus (CLas) and as of today, there are no tools for economically viable disease management. Several wild Australian limes have been identified to be HLB resistant and their resistance is hypothesized to be conferred by resistance genes (R-genes), which mediate pathogen-specific defense responses. The aim of this study was to gain insight into the genomic features of R-genes in Australian limes, in comparison to susceptible citrus cultivars. In this study, we used five citrus genomes, including three Australian limes (Citrus australasica, C. glauca and C. inodora) and two cultivated citrus species (C. clementina and C. sinensis). Our results indicate up to 70% of the R-genes were identified in the unannotated regions in the original genome annotation of each species, owing to the use of a R-gene specific pipeline. Surprisingly, the two cultivated species harbored 15.8 to 104% more R-genes than the Australian limes. In all species, over 75% of the R-genes occurred in clusters and nearly 80% were concentrated in three chromosomes (Chr3, 5 and 7). The syntenic R-gene based phylogenic classification grouped the five species according to their HLB-resistance levels, reflecting the association between these R-genes and their distinct Australian origins. Domain structure analysis revealed substantial similarities in the R-genes between wild Australian limes and cultivated citrus. Investigation of chromosomal sites underlying Australian specific R genes revealed diversifying selection signatures on several chromosomal regions. The findings in this study will aid in the development of tools for genome-assisted breeding for HLB-resistant varieties.
BACKGROUND:Transgenic crops expressing Cry and Vip3Aa insecticidal proteins from the bacterium Bacillus thuringiensis are a primary tool for controlling fall armyworm (Spodoptera frugiperda) populations. The evolution of resistance to Cry proteins in the native range of the fall armyworm has increased reliance and intensified the selection of resistance to Vip3Aa. In this study, we identified mechanisms of resistance to Vip3Aa in the LA-RR strain of S. frugiperda originating from Louisiana (USA). RESULTS:Midgut epithelial damage in susceptible larvae was evidenced by a significant drop in midgut pH after feeding on either Vip3Aa protoxin or activated toxin. In contrast, this midgut pH drop was only detected for activated Vip3Aa toxin in LA-RR larvae. Midgut fluids from LA-RR larvae displayed delayed processing of Vip3Aa protoxin when compared to fluids from susceptible larvae, and this slower processing was associated with reduced activity and expression of trypsin and chymotrypsin enzyme genes in the LA-RR strain. In bioassays, LA-RR larvae were significantly more susceptible to Vip3Aa protoxin pre-processed by midgut fluids from susceptible than from LA-RR larvae. In addition, midgut brush border membrane vesicles from LA-RR larvae exhibited lower specific Vip3Aa toxin binding than vesicles from the susceptible strain. CONCLUSION:The results of this study support that both slower proteolytic processing and reduced specific binding are associated with resistance to Vip3Aa in a S. frugiperda strain from the Western hemisphere, the native range of this pest. This information increases our understanding of resistance to Vip3Aa and advances monitoring and fall armyworm management. © 2025 Society of Chemical Industry.
Chestnut blight (caused by Cryphonectria parasitica), together with Phytophthora root rot (caused by Phytophthora cinnamomi), has nearly extirpated American chestnut (Castanea dentata) from its native range. In contrast to the susceptibility of American chestnut, many Chinese chestnut (C. mollissima) genotypes are resistant to blight. In this research, we performed a series of genome-wide association studies for blight resistance originating from three unrelated Chinese chestnut trees (Mahogany, Nanking and M16) and a Quantitative Trait Locus (QTL) study on a Mahogany-derived inter-species F2 family. We evaluated trees for resistance to blight after artificial inoculation with two fungal strains and scored nine morpho-phenological traits that are the hallmarks of species differentiation between American and Chinese chestnuts. Results support a moderately complex genetic architecture for blight resistance, as 31 QTLs were found on 12 chromosomes across all studies. Additionally, although most morpho-phenological trait QTLs overlap or are adjacent to blight resistance QTLs, they tend to aggregate in a few genomic regions. Finally, comparison between QTL intervals for blight resistance and those previously published for Phytophthora root rot resistance, revealed five common disease resistance regions on chromosomes 1, 5, and 11. Our results suggest that it will be difficult, but still possible to eliminate Chinese chestnut alleles for the morpho-phenological traits while achieving relatively high blight resistance in a backcross hybrid tree. We see potential for a breeding scheme that utilizes marker-assisted selection early for relatively large effect QTLs followed by genome selection in later generations for smaller effect genomic regions.
Quercus alba L., also known as white oak, eastern white oak, or American white oak, is a quintessential North American species within the white oak section (Quercus) of the genus Quercus, subgenus Quercus. This species plays a vital role as a keystone species in eastern North American forests and plays a significant role in local and regional economies. As a long-lived woody perennial covering an extensive natural range, Q. alba’s biology is shaped by a myriad of adaptations accumulated throughout its natural history. Populations of Q. alba are crucial repositories of genetic, genomic, and evolutionary insights, capturing the essence of successful historical adaptations and ongoing responses to contemporary environmental challenges in the Anthropocene. This intersection offers an exceptional opportunity to integrate genomic knowledge with the discovery of climate-relevant traits, advancing tree improvement, forest ecology, and forest management strategies. This review provides a comprehensive examination of the current understanding of Q. alba’s biology, considering past, present, and future research perspectives. It encompasses aspects such as distribution, phylogeny, population structure, key adaptive traits to cyclical environmental conditions (including water use, reproduction, propagation, and growth), as well as the species’ resilience to biotic and abiotic stressors. Additionally, this review highlights the state-of-the-art research resources available for the Quercus genus, including Q. alba, showcasing developments in genetics, genomics, biotechnology, and phenomics tools. This overview lays the groundwork for exploring and elucidating the principles of longevity in plants, positioning Q. alba as an emerging model tree species, ideally suited for investigating the biology of climate-relevant traits.
Stewartia ovata (cav.) Weatherby, commonly known as mountain stewartia, is an understory tree native to the southeastern United States (U.S.). This relatively rare species occurs in isolated populations in Virginia, Kentucky, Tennessee, North Carolina, South Carolina, Georgia, Alabama, and Mississippi. As a species, S. ovata has largely been overlooked, and limited information is available regarding its ecology, which presents obstacles to conservation efforts. Stewartia ovata has vibrant, large white flowers that bloom in summer with a variety of filament colors, suggesting potential horticultural traits prized by ornamental industry. However, S. ovata is relatively slow growing and, due to long seed dormancy, propagation is challenging with limited success rates. This has created a need to assess the present genetic diversity in S. ovata populations to inform potential conservation and restoration of the species. Here, we employ a genotyping-by-sequencing (GBS) approach to characterize the spatial distribution and genetic diversity of S. ovata in the southern Appalachia region of the eastern United States. A total of 4475 single nucleotide polymorphisms (SNPs) were identified across 147 individuals from 11 collection sites. Our results indicate low genetic diversity (He = 0.216), the presence of population structure (K = 2), limited differentiation (F-ST = 0.039), and high gene flow (Nm = 6.16) between our subpopulations. Principal component analysis corroborated the findings of STRUCTURE, confirming the presence of two distinct S. ovata subpopulations. One subpopulation mainly contains genotypes from the Cumberland Plateau, Tennessee, while the other consists of genotypes present in the Great Smoky Mountain ranges in Tennessee, North Carolina, and portions of Nantahala, Chattahoochee-Oconee national forests in Georgia, highlighting that elevation likely plays a major role in its distribution. Our results further suggested low inbreeding coefficient (F-IS = 0.070), which is expected with an outcrossing tree species. This research further provides necessary insight into extant subpopulations and has generated valuable resources needed for conservation efforts of S. ovata.
Huanglongbing (HLB) is a severe citrus disease worldwide. Wild Australian limes like Citrus australasica, C. inodora, and C. glauca possess beneficial HLB resistance traits. Individual trees of the three taxa were extensively used in a breeding program for over a decade to introgress resistance traits into commercial-quality citrus germplasm. We generated high-quality, phased, de novo genome assemblies of the three Australian limes using PacBio long-read sequencing. The genome assembly sizes of the primary and alternate haplotypes were determined for C. australasica (337 Mb/335 Mb), C. inodora (304 Mb/299 Mb), and C. glauca (376 Mb/379 Mb). The nine chromosome-scale scaffolds included 86–91% of the genome sequences generated. The integrity and completeness of the assembled genomes were estimated to be at 97.2–98.8%. Gene annotation studies identified 25,461 genes in C. australasica, 27,665 in C. inodora, and 30,067 in C. glauca. Genes belonging to 118 orthogroups were specific to Australian lime genomes compared to other citrus genomes analyzed. Significantly fewer canonical resistance (R) genes were found in C. inodora and C. glauca (319 and 449, respectively) compared to C. australasica (576), C. clementina (579), and C. sinensis (651). Similar patterns were observed for other gene families associated with potential HLB resistance, including Phloem protein 2 (PP2) and Callose synthase (CalS) genes predicted in the Australian lime genomes. The genomic information on Australian limes developed in the present study will help understand the genetic basis of HLB resistance.
BACKGROUND:The application of reduced metagenomic sequencing approaches holds promise as a middle ground between targeted amplicon sequencing and whole metagenome sequencing approaches but has not been widely adopted as a technique. A major barrier to adoption is the lack of read simulation software built to handle characteristic features of these novel approaches. Reduced metagenomic sequencing (RMS) produces unique patterns of fragmentation per genome that are sensitive to restriction enzyme choice, and the non-uniform size selection of these fragments may introduce novel challenges to taxonomic assignment as well as relative abundance estimates.RESULTS:Through the development and application of simulation software, readsynth, we compare simulated metagenomic sequencing libraries with existing RMS data to assess the influence of multiple library preparation and sequencing steps on downstream analytical results. Based on read depth per position, readsynth achieved 0.79 Pearson's correlation and 0.94 Spearman's correlation to these benchmarks. Application of a novel estimation approach, fixed length taxonomic ratios, improved quantification accuracy of simulated human gut microbial communities when compared to estimates of mean or median coverage.CONCLUSIONS:We investigate the possible strengths and weaknesses of applying the RMS technique to profiling microbial communities via simulations with readsynth. The choice of restriction enzymes and size selection steps in library prep are non-trivial decisions that bias downstream profiling and quantification. The simulations investigated in this study illustrate the possible limits of preparing metagenomic libraries with a reduced representation sequencing approach, but also allow for the development of strategies for producing and handling the sequence data produced by this promising application.
Summary White oak ( Quercus alba ) is an abundant forest tree species across eastern North America that is ecologically, culturally, and economically important. We report the first haplotype-resolved chromosome-scale genome assembly of Q. alba and conduct comparative analyses of genome structure and gene content against other published Fagaceae genomes. In addition, we probe the genetic diversity of this widespread species and investigate its phylogenetic relationships with other oaks using whole-genome data. Our genome assembly comprises two haplotypes each consisting of 12 chromosomes. We found that the species has high genetic diversity, much of which predates the divergence of Q. alba from other oak species and likely impacts divergence time estimation in Quercus . Our phylogenetic results highlight phylogenetic discordance across the genus and suggest different relationships among North American oaks than have been reported previously. Despite a high preservation of chromosome synteny and genome size across the Quercus phylogeny, certain gene families have undergone rapid changes in size including resistance genes (R genes). The white oak genome represents a major new resource for studying genome diversity and evolution in Quercus and forest trees more generally. Future research will continue to reveal the full scope of genomic diversity across the white oak clade.
Northern red oak (Quercus rubra L.) is an ecologically and economically important forest tree native to the northeastern United States. We present a chromosome-scale, haplotype-resolved genome of Q. rubra, a representative red oak species, generated by the combination of PacBio sequences and chromatin conformation capture (Hi-C) scaffolding. This is the first reference genome from the red oak clade (section Lobatae). The Q. rubra assembly spans 739 Megabases (Mb) with 95.27% of the genome sequences scaffolded into 12 chromosomes and 33,333 protein-coding genes. Comparisons to the genomes of Q. lobata and Q. mongolica reveal high collinearity, with intrachromosomal structural variants present. Orthologous gene family analysis with other oak and rosid tree species revealed that gene families associated with defense response were expanding and contracting simultaneously across the Q. rubra genome. Quercus rubra had the most CC-NBS-LRR and TIR-NBS-LRR resistance genes out of the nine species analyzed. Terpene synthase gene family comparisons further reveal tandem gene duplications in TPS-b subfamily, similar to Q. robur. Single major QTL regions were identified for vegetative bud break and marcescence which contain candidate genes for further research, including a putative ortholog of the circadian clock constituent cryptochrome (CRY2) and a family of eight tandemly duplicated genes for serine protease inhibitors, respectively. Genome-environment associations across natural populations identified candidate abiotic stress tolerance genes and predicted performance in a common garden. This high-quality red oak genome represents an essential resource to the oak genomics community which will further supplement the knowledge of Quercus genomics.
BACKGROUND:Muscadine grape (Vitis rotundifolia) is resistant to many of the pathogens that negatively impact the production of common grape (V. vinifera), including the bacterial pathogen Xylella fastidiosa subsp. fastidiosa (Xfsf), which causes Pierce's Disease (PD). Previous studies in common grape have indicated Xfsf delays host immune response with a complex O-chain antigen produced by the wzy gene. Muscadine cultivars range from tolerant to completely resistant to Xfsf, but the mechanism is unknown.RESULTS:We assembled and annotated a new, long-read genome assembly for 'Carlos', a cultivar of muscadine that exhibits tolerance, to build upon the existing genetic resources available for muscadine. We used these resources to construct an initial pan-genome for three cultivars of muscadine and one cultivar of common grape. This pan-genome contains a total of 34,970 synteny-constrained entries containing genes of similar structure. Comparison of resistance gene content between the 'Carlos' and common grape genomes indicates an expansion of resistance (R) genes in 'Carlos.' We further identified genes involved in Xfsf response by transcriptome sequencing 'Carlos' plants inoculated with Xfsf. We observed 234 differentially expressed genes with functions related to lipid catabolism, oxidation-reduction signaling, and abscisic acid (ABA) signaling as well as seven R genes. Leveraging public data from previous experiments of common grape inoculated with Xfsf, we determined that most differentially expressed genes in the muscadine response were not found in common grape, and three of the R genes identified as differentially expressed in muscadine do not have an ortholog in the common grape genome.CONCLUSIONS:Our results support the utility of a pan-genome approach to identify candidate genes for traits of interest, particularly disease resistance to Xfsf, within and between muscadine and common grape.
Inquiline ant social parasites exploit other ant species for their reproductive benefit because they do not possess a worker caste. Due to their relative rarity in nature, the biology and natural history of inquilines are largely unknown. Likewise, not much research exists that details the close relationship between inquilines and their host(s), and how each organism influences the genetic structure of the other. Here, we conducted a comparative population genetics study to assess patterns of genetic structure within and among populations of inquiline Solenopsis daguerrei and its known fire ant hosts, which includes invasive Solenopsis invicta. Using nuclear and mitochondrial markers, we show that four genetically distinct groups of S. daguerrei likely exist, each with different degrees of host association. Consistent with previous inferences of the inquiline lifestyle, we find that inbreeding is common in S. daguerrei, presumably a result of intranidal mating and restricted dispersal. Results from this study, specifically host association patterns, may inform future biological control strategies to mitigate invasive S. invicta populations.
Northern red oak (Quercus rubra L.) is an ecologically and economically important forest tree native to North America. We present a chromosome-scale genome of Q. rubra generated by the combination of PacBio sequences and chromatin conformation capture (Hi-C) scaffolding. This is the first reference genome from the red oak clade (section Lobatae). The Q. rubra assembly spans 739 Mb with 95.27% of the genome in 12 chromosomes and 33,333 protein-coding genes. Comparisons to the genomes of Quercus lobata and Quercus mongolica revealed high collinearity, with intrachromosomal structural variants present. Orthologous gene family analysis with other tree species revealed that gene families associated with defense response were expanding and contracting simultaneously across the Q. rubra genome. Quercus rubra had the most CC-NBS-LRR and TIR-NBS-LRR resistance genes out of the 9 species analyzed. Terpene synthase gene family comparisons further reveal tandem gene duplications in TPS-b subfamily, similar to Quercus robur. Phylogenetic analysis also identified 4 subfamilies of the IGT/LAZY gene family in Q. rubra important for plant structure. Single major QTL regions were identified for vegetative bud break and marcescence, which contain candidate genes for further research, including a putative ortholog of the circadian clock constituent cryptochrome (CRY2) and 8 tandemly duplicated genes for serine protease inhibitors, respectively. Genome-environment associations across natural populations identified candidate abiotic stress tolerance genes and predicted performance in a common garden. This high-quality red oak genome represents an essential resource to the oak genomic community, which will expedite comparative genomics and biological studies in Quercus species.
It is critical to gather biological information about rare and endangered plants to incorporate into conservation efforts. The secondary metabolism of Pityopsis ruthii, an endangered flowering plant that only occurs along limited sections of two rivers (Ocoee and Hiwassee) in Tennessee, USA was studied. Our long-term goal is to understand the mechanisms behind P. ruthii's adaptation to restricted areas in Tennessee. Here, we profiled the secondary metabolites, specifically in flowers, with a focus on terpenes, aiming to uncover the genomic and molecular basis of terpene biosynthesis in P. ruthii flowers using transcriptomic and biochemical approaches. By comparative profiling of the nonpolar portion of metabolites from various tissues, P. ruthii flowers were rich in terpenes, which included 4 monoterpenes and 10 sesquiterpenes. These terpenes were emitted from flowers as volatiles with monoterpenes and sesquiterpenes accounting for almost 68% and 32% of total emission of terpenes, respectively. These findings suggested that floral terpenes play important roles for the biology and adaptation of P. ruthii to its limited range. To investigate the biosynthesis of floral terpenes, transcriptome data for flowers were produced and analyzed. Genes involved in the terpene biosynthetic pathway were identified and their relative expressions determined. Using this approach, 67 putative terpene synthase (TPS) contigs were detected. TPSs in general are critical for terpene biosynthesis. Seven full-length TPS genes encoding putative monoterpene and sesquiterpene synthases were cloned and functionally characterized. Three catalyzed the biosynthesis of sesquiterpenes and four catalyzed the biosynthesis of monoterpenes. In conclusion, P. ruthii plants employ multiple TPS genes for the biosynthesis of a mixture of floral monoterpenes and sesquiterpenes, which probably play roles in chemical defense and attracting insect pollinators alike.