Tropical milkweed (Asclepias curassavica) serves as a host plant for monarch butterflies (Danaus plexippus) and other insect herbivores that can tolerate the abundant cardiac glycosides that are characteristic of this species. Cardiac glycosides, along with additional specialized metabolites, also contribute to the ethnobotanical uses of A. curassavica. To facilitate further research on milkweed metabolism, we assembled the 197-Mbp genome of a fifth-generation inbred line of A. curassavica into 619 contigs, with an N50 of 10 Mbp. Scaffolding resulted in 98% of the assembly being anchored to 11 chromosomes, which are mostly colinear with the previously assembled common milkweed (A. syriaca) genome. Assembly completeness evaluations showed that 98% of the BUSCO gene set is present in the A. curassavica genome assembly. The transcriptomes of six tissue types (young leaves, mature leaves, stems, flowers, buds, and roots), with and without defense elicitation by methyl jasmonate treatment, showed both tissue-specific gene expression and induced expression of genes that may be involved in cardiac glycoside biosynthesis. Expression of a CYP87A gene, the predicted first gene in the cardiac glycoside biosynthesis pathway, was observed only in the stems and roots and was induced by methyl jasmonate. Together, this genome sequence and transcriptome analysis provide important resources for further investigation of the ecological and medicinal uses of A. curassavica.
Basil, Ocimum basilicum L., is a widely cultivated aromatic herb, prized for its culinary and medicinal uses, predominantly owing to its unique aroma, primarily determined by eugenol for Genovese cultivars or methyl chavicol for Thai cultivars. To date, a comprehensive basil reference genome has been lacking, with only a fragmented draft available. To fill this gap, we employed PacBio HiFi and Hi-C sequencing to construct a homeolog-phased chromosome-level genome for basil. The tetraploid basil genome was assembled into 26 pseudomolecules and further categorized into subgenomes. High levels of synteny were observed between the two basil subgenomes but comparisons to Salvia rosmarinus show collinearity quickly breaks down in near relatives. We utilized a bi-parental population derived from a Genovese × Thai cross to map quantitative trait loci (QTL) for the aroma chemotype. We discovered a single QTL governing the eugenol/methyl chavicol ratio, which encompassed a genomic region with 95 genes, including 15 genes encoding a shikimate O-hydroxycinnamoyltransferase (HCT/CST) enzyme. Of them, only ObHCT1 exhibited significantly higher expression in the Genovese cultivar and showed a trichome-specific expression. ObHCT1 was functionally confirmed as a genuine HCT enzyme using an in vitro assay. The high-quality, contiguous basil reference genome is now publicly accessible at BasilBase, a valuable resource for the scientific community. Combined with insights into cell-type-specific gene expression, it promises to elucidate specialized metabolite biosynthesis pathways at the cellular level.
Plants in the genus Erysimum produce both glucosinolates and cardenolides as a defense mechanism against herbivory. Two natural isolates of Erysimum cheiranthoides (wormseed wallflower) differed in their glucosinolate content, cardenolide content, and their resistance to Myzus persicae (green peach aphid), a broad generalist herbivore. Both classes of defensive metabolites were produced constitutively and were not further induced by aphid feeding. To investigate the relative importance of glucosinolates and cardenolides in E. cheiranthoides defense, we generated an improved genome assembly, genetic map, and segregating F2 population. The genotypic and phenotypic analysis of the F2 plants identified quantitative trait loci, which affected glucosinolates and cardenolides, but not the aphid resistance. The abundance of most glucosinolates and cardenolides was positively correlated in the F2 population, indicating that similar processes regulate their biosynthesis and accumulation. Aphid reproduction was positively correlated with glucosinolate content. Although the overall cardenolide content had little effect on aphid growth and survival, there was a negative correlation between aphid reproduction and helveticoside abundance. However, this variation in defensive metabolites could not explain the differences in aphid growth on the two parental lines, suggesting that processes other than the abundance of glucosinolates and cardenolides have a predominant effect on aphid resistance in E. cheiranthoides.
Sunflower (Helianthus annuus L.) is widely utilized for seed oil production. Priming seeds prior to sowing is a technique used to enhance the germination rate and uniformity of seedling growth. Priming times of 6 and 18 h were selected to be the optimal and extended durations, respectively. Three biological replicates per treatment were used for next-generation sequencing via the Illumina platform. In this study, priming conditions decreased the time to 50% radicle emergence, implying earlier and more uniform germination without affecting the final germination percentage. Transcriptome analysis identified 6171 differentially expressed genes (DEGs) for 0-6 h, 7241 DEGs for 0-18 h, and 1527 DEGs for 6-18 h. The number of expressed DEGs during the early hours of priming exceeded those identified during the extended priming time, suggesting more metabolic processes were taking place in the earlier priming period. Noteworthy DEGs are related to stress responses, the biosynthesis of metabolites, and the activation of antioxidants and phytohormones. Moreover, these DEGs were significantly enriched in translation, metabolic processes, and biosynthesis for the priming times of 0-6 h and 0-18 h, whereas anaerobic respiration, detection of hypoxia, and lipid transport were differentially expressed between 6 and 18 h. Furthermore, expansin genes, namely, EXPA1, EXPA9, EXPA15, EXLA3, and EXP/Lol pI 3B GAG-pre-integrase domain, were significantly expressed. These expansins are cell wall-loosening proteins that are involved in the processes of physiological development in higher plants, including seed germination. The data obtained in this study contribute toward deciphering the molecular mechanisms associated with seed priming in sunflower seeds.
Coffee leaf rust (CLR) is one of the most economically important diseases affecting Coffea arabica production, having a significant economic impact. Among the main goals of coffee breeding programs is the development of cultivars resistant to this disease. A source of resistance genes is Híbrido de Timor (HdT), a spontaneous hybrid originated from the cross between C. arabica and C. canephora. Previously, in a transcriptome study, the Ca TDF77 NBS-LRR gene from HdT involved in resistance to CLR was identified. Hence, our aim was to characterize the genomic region surrounding the Ca TDF77 NBS-LRR gene in Coffea spp. Furthermore, we aimed to analyze the transcriptional profile of this gene, in the C. arabica cultivar IAPAR 59, which is originated from HdT introgression and is resistant to CLR race II. The outcome delineated the gene’s localization on chromosome 11 (canephora subgenome) of C. arabica, spotlighting intragenic polymorphisms between HdT and Arabica coffee susceptible to CLR race II. The genomic region surrounding the gene in Coffea spp. revealed a tandem structure and transposable elements. Notably, within IAPAR 59, the gene exhibited significant upregulation at 24 and 72 h post CLR infection, contrasting starkly with the susceptible genotype. This observation validates its role in fortifying the defense mechanism of this particular cultivar. This study enriches our understanding of the evolutionary dynamics of Coffea spp. genomes and also provides genomic resources instrumental in devising biotechnological strategies for resistance to CLR.
In intimate ecological interactions, the interdependency of species may result in correlated demographic histories. For species of conservation concern, understanding the long-term dynamics of such interactions may shed light on the drivers of population decline. Here, we address the demographic history of the monarch butterfly, Danaus plexippus, and its dominant host plant, the common milkweed Asclepias syriaca (A. syriaca), using broad-scale sampling and genomic inference. Because genetic resources for milkweed have lagged behind those for monarchs, we first release a chromosome-level genome assembly and annotation for common milkweed. Next, we show that despite its enormous geographic range across eastern North America, A. syriaca is best characterized as a single, roughly panmictic population. Using approximate Bayesian computation with random forests (ABC-RF), a machine learning method for reconstructing demo-graphic histories, we show that both monarchs and milkweed experienced population expansion during the most recent recession of North American glaciers 10,000-20,000 years ago. Our data also identify concurrent population expansions in both species during the large-scale clearing of eastern forests (-200 years ago). Finally, we find no evidence that either species experienced a reduction in effective population size over the past 75 years. Thus, the well-documented decline of monarch abundance over the past 40 years is not visible in our genomic dataset, reflecting a possible mismatch of the overwintering census population to effective population size in this species.
Water availability influences all aspects of plant growth and development; however, most studies of plant responses to drought have focused on vegetative organs, notably roots and leaves. Far less is known about the molecular bases of drought acclimation responses in fruits, which are complex organs with distinct tissue types. To obtain a more comprehensive picture of the molecular mechanisms governing fruit development under drought, we profiled the transcriptomes of a spectrum of fruit tissues from tomato (Solanum lycopersicum), spanning early growth through ripening and collected from plants grown under varying intensities of water stress. In addition, we compared transcriptional changes in fruit with those in leaves to highlight different and conserved transcriptome signatures in vegetative and reproductive organs. We observed extensive and diverse genetic reprogramming in different fruit tissues and leaves, each associated with a unique response to drought acclimation. These included major transcriptional shifts in the placenta of growing fruit and in the seeds of ripe fruit related to cell growth and epigenetic regulation, respectively. Changes in metabolic and hormonal pathways, such as those related to starch, carotenoids, jasmonic acid, and ethylene metabolism, were associated with distinct fruit tissues and developmental stages. Gene coexpression network analysis provided further insights into the tissue-specific regulation of distinct responses to water stress. Our data highlight the spatiotemporal specificity of drought responses in tomato fruit and indicate known and unrevealed molecular regulatory mechanisms involved in drought acclimation, during both vegetative and reproductive stages of development.
In intimate ecological interactions, the interdependency of species may result in correlated demographic histories. For species of conservation concern, understanding the long-term dynamics of such interactions may shed light on the drivers of population decline. Here we address the demographic history of the monarch butterfly, Danaus plexippus , and its dominant host plant, the common milkweed Asclepias syriaca , using broad-scale sampling and genomic inference. Because genetic resources for milkweed have lagged behind those for monarchs, we first release a chromosome-level genome assembly and annotation for common milkweed. Next, we show that despite its enormous geographic range across eastern North America, A. syriaca is best characterized as a single, roughly panmictic population. Using Approximate Bayesian Computation via Random Forests (ABC-RF), a machine learning method for reconstructing demographic histories, we show that both monarchs and milkweed experienced concurrent range expansion during the most recent recession of North American glaciers ∼12,000 years ago. Our data identify an expansion of milkweed during the large-scale clearing of eastern forests (∼200 years ago) but was inconclusive as to expansion or contraction of the monarch butterfly population during this time. Finally, our results indicate that neither species experienced a population contraction over the past 75 years. Thus, the well-documented decline of monarch abundance over the past 40 years is not visible in our genomic dataset, reflecting a possible mismatch of the overwintering census population to effective population size in this species. ### Competing Interest Statement The authors have declared no competing interest.
SUMMARY Wild relatives of tomato are a valuable source of natural variation in tomato breeding, as many can be hybridized to the cultivated species ( Solanum lycopersicum ). Several, including Solanum lycopersicoides , have been crossed to S. lycopersicum for the development of ordered introgression lines (ILs), facilitating breeding for desirable traits. Despite the utility of these wild relatives and their associated ILs, few finished genome sequences have been produced to aid genetic and genomic studies. Here we report a chromosome‐scale genome assembly for S. lycopersicoides LA2951, which contains 37 938 predicted protein‐coding genes. With the aid of this genome assembly, we have precisely delimited the boundaries of the S. lycopersicoides introgressions in a set of S. lycopersicum cv. VF36 × LA2951 ILs. We demonstrate the usefulness of the LA2951 genome by identifying several quantitative trait loci for phenolics and carotenoids, including underlying candidate genes, and by investigating the genome organization and immunity‐associated function of the clustered Pto gene family. In addition, syntenic analysis of R2R3MYB genes sheds light on the identity of the Aubergine locus underlying anthocyanin production. The genome sequence and IL map provide valuable resources for studying fruit nutrient/quality traits, pathogen resistance, and environmental stress tolerance. We present a new genome resource for the wild species S. lycopersicoides , which we use to shed light on the Aubergine locus responsible for anthocyanin production. We also provide IL boundary mappings, which facilitated identifying novel carotenoid quantitative trait loci of which one was likely driven by an uncharacterized lycopene β‐cyclase whose function we demonstrate.
Modern breeding methods integrate next-generation sequencing and phenomics to identify plants with the best characteristics and greatest genetic merit for use as parents in subsequent breeding cycles to ultimately create improved cultivars able to sustain high adoption rates by farmers. This data-driven approach hinges on strong foundations in data management, quality control, and analytics. Of crucial importance is a central database able to (1) track breeding materials, (2) store experimental evaluations, (3) record phenotypic measurements using consistent ontologies, (4) store genotypic information, and (5) implement algorithms for analysis, prediction, and selection decisions. Because of the complexity of the breeding process, breeding databases also tend to be complex, difficult, and expensive to implement and maintain. Here, we present a breeding database system, Breedbase (https://breedbase.org/, last accessed 4/18/2022). Originally initiated as Cassavabase (https://cassavabase.org/, last accessed 4/18/2022) with the NextGen Cassava project (https://www.nextgencassava.org/, last accessed 4/18/2022), and later developed into a crop-agnostic system, it is presently used by dozens of different crops and projects. The system is web based and is available as open source software. It is available on GitHub (https://github.com/solgenomics/, last accessed 4/18/2022) and packaged in a Docker image for deployment (https://hub.docker.com/u/breedbase, last accessed 4/18/2022). The Breedbase system enables breeding programs to better manage and leverage their data for decision making within a fully integrated digital ecosystem.
The tomato (Solanum lycopersicum L.) family, Solanaceae, is a model clade for a wide range of applied and basic research questions. Currently, reference‐quality genomes are available for over 30 species from seven genera, and these include numerous crops as well as wild species [e.g., Jaltomata sinuosa (Miers) Mione and Nicotiana attenuata Torr. ex S. Watson]. Here we present the genome of the showy‐flowered Andean shrub Iochroma cyaneum (Lindl.) M. L. Green, a woody lineage from the tomatillo (Physalis philadelphica Lam.) subfamily Physalideae. The assembled size of the genome (2.7 Gb) is more similar in size to pepper (Capsicum annuum L.) (2.6 Gb) than to other sequenced diploid members of the berry clade of Solanaceae [e.g., potato (Solanum tuberosum L.), tomato, and Jaltomata]. Our assembly recovers 92% of the conserved orthologous set, suggesting a nearly complete genome for this species. Most of the genomic content is repetitive (69%), with Gypsy elements alone accounting for 52% of the genome. Despite the large amount of repetitive content, most of the 12 I. cyaneum chromosomes are highly syntenic with tomato. Bayesian concordance analysis provides strong support for the berry clade, including I. cyaneum, but reveals extensive discordance along the backbone, with placement of chili pepper and Jaltomata being highly variable across gene trees. The I. cyaneum genome contributes to a growing wealth of genomic resources in Solanaceae and underscores the need for expanded sampling of diverse berry genomes to dissect major morphological transitions.
The maize (Zea mays) genome encodes three indole-3-glycerolphosphate synthase enzymes (IGPS1, 2, and 3) catalyzing the conversion of 1-(2-carboxyphenylamino)-l-deoxyribulose-5-phosphate to indole-3-glycerolphosphate. Three further maize enzymes (BX1, benzoxazinoneless 1; TSA, tryptophan synthase alpha subunit; and IGL, indole glycerolphosphate lyase) convert indole-3-glycerolphosphate to indole, which is released as a volatile defense signaling compound and also serves as a precursor for the biosynthesis of tryptophan and defense-related benzoxazinoids. Phylogenetic analyses showed that IGPS2 is similar to enzymes found in both monocots and dicots, whereas maize IGPS1 and IGPS3 are in monocot-specific clades. Fusions of yellow fluorescent protein with maize IGPS enzymes and indole-3-glycerolphosphate lyases were all localized in chloroplasts. In bimolecular fluorescence complementation assays, IGPS1 interacted strongly with BX1 and IGL, IGPS2 interacted primarily with TSA, and IGPS3 interacted equally with all three indole-3-glycerolphosphate lyases. Whereas IGPS1 and IGPS3 expression was induced by insect feeding, IGPS2 expression was not. Transposon insertions in IGPS1 and IGPS3 reduced the abundance of both benzoxazinoids and free indole. Spodoptera exigua (beet armyworm) larvae show improved growth on igps1 mutant maize plants. Together, these results suggest that IGPS1 and IGPS3 function mainly in the biosynthesis of defensive metabolites, whereas IGPS2 may be involved in the biosynthesis of tryptophan. This metabolic channeling is similar to, though less exclusive than, that proposed for the three maize indole-3-glycerolphosphate lyases.
Cassava, a food security crop in Africa, is grown throughout the tropics and subtropics. Although cassava can provide high productivity in suboptimal conditions, the yield in Africa is substantially lower than in other geographies. The yield gap is attributable to many challenges faced by cassava in Africa, including susceptibility to diseases and poor soil conditions. In this study, we carried out 3’RNA sequencing on 150 accessions from the National Crops Resources Research Institute, Uganda for 5 tissue types, providing population-based transcriptomics resources to the research community in a web-based queryable cassava expression atlas. Differential expression and weighted gene co-expression network analysis were performed to detect 8820 significantly differentially expressed genes (DEGs), revealing similarity in expression patterns between tissue types and the clustering of detected DEGs into 18 gene modules. As a confirmation of data quality, differential expression and pathway analysis targeting cassava mosaic disease (CMD) identified 27 genes observed in the plant–pathogen interaction pathway, several previously identified CMD resistance genes, and two peroxidase family proteins different from the CMD2 gene. Present research work represents a novel resource towards understanding complex traits at expression and molecular levels for the development of resistant and high-yielding cassava varieties, as exemplified with CMD.
Understanding the diversity and genetic relationships among and within crop germplasm is invaluable for genetic improvement. This study assessed genetic diversity in a panel of 173 D. rotundata accessions using joint analysis for 23 morphological traits and 136,429 SNP markers from the whole-genome resequencing platform. Various diversity matrices and clustering methods were evaluated for a comprehensive characterization of genetic diversity in white Guinea yam from West Africa at phenotypic and molecular levels. The translation of the different diversity matrices from the phenotypic and genomic information into distinct groups varied with the hierarchal clustering methods used. Gower distance matrix based on phenotypic data and identity by state (IBS) distance matrix based on SNP data with the UPGMA clustering method found the best fit to dissect the genetic relationship in current set materials. However, the grouping pattern was inconsistent (r = − 0.05) between the morphological and molecular distance matrices due to the non-overlapping information between the two data types. Joint analysis for the phenotypic and molecular information maximized a comprehensive estimate of the actual diversity in the evaluated materials. The results from our study provide valuable insights for measuring quantitative genetic variability for breeding and genetic studies in yam and other root and tuber crops.
Summary Wild relatives of tomato are a valuable source of natural variation in tomato breeding, as many can be hybridized to the cultivated species ( Solanum lycopersicum ). Several, including Solanum lycopersicoides , have been crossed to S. lycopersicum for the development of ordered introgression lines (ILs). Despite the utility of these wild relatives and their associated ILs, limited finished genomes have been produced to aid genetic and genomic studies. We have generated a chromosome-scale genome assembly for Solanum lycopersicoides LA2951 using PacBio sequencing, Illumina, and Hi-C. We identified 37,938 genes based on Illumina and Isoseq and compared gene function to the available cultivated tomato genome resources, in addition to mapping the boundaries of the S. lycopersicoides introgressions in a set of cv. VF36 x LA2951 introgression lines (IL). The genome sequence and IL map will support the development of S. lycopersicoides as a model for studying fruit nutrient/quality, pathogen resistance, and environmental stress tolerance traits that we have identified in the IL population and are known to exist in S. lycopersicoides .
Capsicum annuum is one of the most important horticultural crops worldwide. Anthracnose disease ( Colletotrichum spp.) is a major constraint for chili production, causing substantial losses. Capsidiol is a sesquiterpene phytoalexin present in pepper fruits that can enhance plant resistance. The genetic mechanisms involved in capisidiol biosynthesis are still poorly understood. In this study, a 3′ RNA sequencing approach was used to develop the transcriptional profile dataset of C. annuum genes in unripe (UF) and ripe fruits (RF) in response to C. scovillei infection. Results showed 4,845 upregulated and 4,720 downregulated genes in UF, and 2,560 upregulated and 1,762 downregulated genes in RF under fungus inoculation. Four capsidiol-related genes were selected for RT-qPCR analysis, two 5-epi-aristolochene synthase ( CA12g05030 , CA02g09520) and two 5-epi-aristolochene-1,3-dihydroxylase genes ( CA12g05070 , CA01g05990 ). CA12g05030 and CA01g05990 genes showed an early response to fungus infection in RF (24 h post-inoculation—HPI), being 68-fold and 53-fold more expressed at 96 HPI, respectively. In UF, all genes showed a late response, especially CA12g05030 , which was 700-fold more expressed at 96 HPI compared to control plants. We are proving here the first high-throughput expression dataset of pepper fruits in response to anthracnose disease in order to contribute for future pepper breeding programs.
Phytochemical diversity is thought to result from coevolutionary cycles as specialization in herbivores imposes diversifying selection on plant chemical defenses. Plants in the speciose genus Erysimum (Brassicaceae) produce both ancestral glucosinolates and evolutionarily novel cardenolides as defenses. Here we test macroevolutionary hypotheses on co-expression, co-regulation, and diversification of these potentially redundant defenses across this genus. We sequenced and assembled the genome of E. cheiranthoides and foliar transcriptomes of 47 additional Erysimum species to construct a phylogeny from 9868 orthologous genes, revealing several geographic clades but also high levels of gene discordance. Concentrations, inducibility, and diversity of the two defenses varied independently among species, with no evidence for trade-offs. Closely related, geographically co-occurring species shared similar cardenolide traits, but not glucosinolate traits, likely as a result of specific selective pressures acting on each defense. Ancestral and novel chemical defenses in Erysimum thus appear to provide complementary rather than redundant functions.
Modern breeding programs routinely use genome-wide information for selecting individuals to advance. The large volumes of genotypic information required present a challenge for data storage and query efficiency. Major use cases require genotyping data to be linked with trait phenotyping data. In contrast to phenotyping data that are often stored in relational database schemas, next-generation genotyping data are traditionally stored in non-relational storage systems due to their extremely large scope. This study presents a novel data model implemented in Breedbase () for uniting relational phenotyping data and non-relational genotyping data within the open-source PostgreSQL database engine. Breedbase is an open-source, web-database designed to manage all of a breeder's informatics needs: management of field experiments, phenotypic and genotypic data collection and storage, and statistical analyses. The genotyping data is stored in a PostgreSQL data-type known as binary JavaScript Object Notation (JSONb), where the JSON structures closely follow the Variant Call Format (VCF) data model. The Breedbase genotyping data model can handle different ploidy levels, structural variants, and any genotype encoded in VCF. JSONb is both compressed and indexed, resulting in a space and time efficient system. Furthermore, file caching maximizes data retrieval performance. Integration of all breeding data within the Chado database schema retains referential integrity that may be lost when genotyping and phenotyping data are stored in separate systems. Benchmarking demonstrates that the system is fast enough for computation of a genomic relationship matrix (GRM) and genome wide association study (GWAS) for datasets involving 1,325 diploid Zea mays, 314 triploid Musa acuminata, and 924 diploid Manihot esculenta samples genotyped with 955,690, 142,119, and 287,952 genotype-by-sequencing (GBS) markers, respectively.
Physcomitrella patens is a bryophyte model plant that is often used to study plant evolution and development. Its resources are of great importance for comparative genomics and evo-devo approaches. However, expression data from Physcomitrella patens were so far generated using different gene annotation versions and three different platforms: CombiMatrix and NimbleGen expression microarrays and RNA sequencing. The currently available P. patens expression data are distributed across three tools with different visualization methods to access the data. Here, we introduce an interactive expression atlas, Physcomitrella Expression Atlas Tool (PEATmoss), that unifies publicly available expression data for P. patens and provides multiple visualization methods to query the data in a single web-based tool. Moreover, PEATmoss includes 35 expression experiments not previously available in any other expression atlas. To facilitate gene expression queries across different gene annotation versions, and to access P. patens annotations and related resources, a lookup database and web tool linked to PEATmoss was implemented. PEATmoss can be accessed at https://peatmoss.online.uni-marburg.de.