Solanum microdontum Bitter is a diploid wild Andean relative of potato that has shaped the domestication and adaptation of modern cultivated potato to diverse environments. S. microdontum has the potential to provide a wealth of untapped genetic material for use in addressing current challenges in potato breeding. Here, we report a high-quality 772 Mb reference genome sequence for S. microdontum that is anchored to 12 chromosomes. The resulting genome assembly has 99.0% complete benchmarking universal single copy orthologs and an N50 scaffold length of over 57 Mb, indicating a high level of completeness. Annotation of the assembly resulted in the identification of 37,324 protein-coding genes and 65% repetitive sequence. A total of 1,187 nucleotide-binding leucine-rich repeat genes were predicted from the assembly, of which, 93.1% overlapped an annotated high-confidence gene model. A k-mer-based kinship matrix derived from a 107-member S. microdontum diversity panel revealed an underlying population structure that corresponds to geographic proximity. The S. microdontum dataset enhances publicly available potato genome resources by providing breeders with genetic, molecular, and germplasm resources for newly developed diploid potato breeding programs.
Yellow wood sorrel (Oxalis stricta L.), also known as sourgrass, juicy fruit, or sheep weed, is a member of the Oxalidaceae family. Yellow wood sorrel is commonly considered a weed, and while native to North America, it is distributed across Europe, Asia, and Africa. To date, only 2 other genomes from the Oxalidaceae family have been published, star fruit (Averrhoa carambola L.) and Oxalis articulata Savigny. Here, we present a chromosome-scale assembly for O. stricta, revealing its allotetraploid nature and synteny within its 2 subgenomes as well as synteny with A. carambola and O. articulata. Using Oxford Nanopore Technologies long-read sequences coupled with chromatin capture sequencing, we generated a 436 Mb chromosome-scale assembly of O. stricta with a scaffold N50 length of 36.2 Mb that is anchored to 12 chromosomes across the 2 subgenomes. Assessment of the final genome assembly using the Long Terminal Repeat Assembly Index yielded a score of 13.12, and assessment of Benchmarking Universal Single Copy Orthologs revealed 99.6% complete orthologs; both metrics are suggestive of a high-quality reference genome. Total repetitive sequence content in the O. stricta genome was 39.7% with retroelements being the largest class of transposable elements. Annotation of protein-coding genes yielded 61,550 high-confidence genes encoding 115,089 gene models. Synteny between the 2 O. stricta subgenomes was present in 93 syntenic blocks containing 40,750 genes, of which, 76.47% were present in 1:1 syntenic relationships between the 2 subgenomes. The availability of an annotated chromosome-scale high-quality genome assembly for O. stricta will provide a launching point to understand the high fecundity of this weed and provide further foundation for comparative genomics within the Oxalidaceae.
Teucrium chamaedrys, commonly known as wall germander, is a small woody shrub native to the Mediterranean region. Its name is derived from the Greek words meaning "ground oak," as its tiny leaves resemble those of an oak tree. Teucrium species are prolific producers of diterpenes, endowing them with valuable properties widely utilized in traditional and modern medicine. Sequencing and assembly of the 3-Gbp tetraploid T. chamaedrys genome revealed 74 diterpene synthase genes, with a substantial number of these genes clustered at four synteny genomic loci, each harboring a copy of a large diterpene biosynthetic gene cluster. Comparative genomics revealed that this cluster is conserved in the closely related species Teucrium marum. Along with the presence of several cytochrome p450 sequences, this region is among the largest biosynthetic gene clusters identified. Teucrium is well known for accumulating clerodane-type diterpenoids, which are produced from a kolavenyl diphosphate precursor. To elucidate the complex biosynthetic pathways of these medicinal compounds, we identified and functionally characterized several kolavenyl diphosphate synthases from T. chamaedrys. The remarkable chemical diversity and tetraploid nature of T. chamaedrys make it a valuable model for studying genomic evolution and adaptation in plants.
Iridoids are specialized monoterpenes ancestral to asterid flowering plants1,2 that play key roles in defence and are also essential precursors for pharmacologically important alkaloids3,4. The biosynthesis of all iridoids involves the cyclization of the reactive biosynthetic intermediate 8-oxocitronellyl enol. Here, using a variety of approaches including single-nuclei sequencing, we report the discovery of iridoid cyclases from a phylogenetically broad sample of asterid species that synthesize iridoids. We show that these enzymes catalyse formation of 7S-cis-trans and 7R-cis-cis nepetalactol, the two major iridoid stereoisomers found in plants. Our work uncovers a key missing step in the otherwise well-characterized early iridoid biosynthesis pathway in asterids. This discovery unlocks the possibility to generate previously inaccessible iridoid stereoisomers, which will enable metabolic engineering for the sustainable production of valuable iridoid and iridoid-derived compounds.
Breeding for sweetpotato (Ipomea batatas) resistance requires accelerating our understanding of genomic sources of resistance. Nucleotide-binding domain leucine-rich repeat receptor (NLR) proteins represent a key component of the plant immune system that mediate plant immune responses. We cataloged the NLR diversity in 32 hexaploid sweetpotato genotypes and three diploid wild relatives using resistance gene enrichment sequencing (RenSeq) to capture and sequence full NLRs. A custom-designed NLR bait library enriched NLR genes with an average 97% target capture rate. We employed a curated database of cloned and functionally characterized NLRs to assign sequenced sweetpotato NLRs to canonical phylogenetic clades. We identified between 800 and 1,200 complete NLRs, highlighting the expanded diversity of coiled-coil NLRs (CNLs) across all genotypes. NLRs among sweetpotato genotypes exhibited large conservation across genotypes. The phylogenetic distance between 6× (hexaploid) and 2× (diploid) genotypes revealed that a small repertoire of I. batatas CNLs diverged from the sweetpotato wild relatives. Finally, we obtained chromosome coordinates in hexaploid (Beauregard) and diploid (Ipomoea trifida) genomes and recorded clustering of NLRs on chromosomes arms. Our study provides a catalog of NLR genes that can be used to accelerate breeding and increase our understanding of the evolutionary dynamics of sweetpotato NLRs. [Formula: see text] Copyright © 2025 The Author(s). This is an open access article distributed under the CC BY-NC-ND 4.0 International license.
Monoterpene indole alkaloids (MIAs) are a large, structurally diverse class of bioactive natural products. These compounds are biosynthetically derived from a stereoselective Pictet-Spengler condensation that generates a tetrahydro-β-carboline scaffold characterized by a 3S stereocenter. However, a subset of MIAs contains a noncanonical 3R stereocenter. Here we report the basis for 3R-MIA biosynthesis in Mitragyna speciosa (kratom). We discover the presence of the iminium species (20S)-3-dehydrocorynantheidine, which supports isomerization of 3S to 3R via oxidation and stereoselective reduction downstream of the initial Pictet-Spengler condensation. Isotopologue feeding experiments identify the sites for downstream MIA pathway biosynthesis as well as the oxidase/reductase pair that catalyzes this epimerization. This oxidase/reductase pair has broad substrate specificity, suggesting that this pathway may be responsible for the formation of many 3R-MIAs and downstream spirooxindole alkaloids in kratom. The elucidation of this epimerization mechanism allows biocatalytic access to a range of pharmacologically active spirooxindole alkaloid compounds.
The medicinal tree Camptotheca acuminata produces camptothecin, a monoterpenoid indole al- kaloid (MIA) precursor for several leading chemotherapeutic agents ([Lorence and Nessler 2004][1]). Alt- hough the biosynthesis of camptothecin remains poorly understood, a putative route has been hypothe- sized based on in planta metabolite profiling and feeding studies ([Fig. 1a][2]) ([Sheriha and Rapoport 1976][3]; [Sadre et al. 2016][4]). However, pathways proposed on whole-tissue or organ-level metabolomic and tran- scriptomic data lack resolution on the intricate, cell-specific compartmentalization of natural products biosynthesis. Such gaps could be addressed by single cell technologies, which have recently shown tremendous potential to transform gene discovery in herbaceous plants ([Li et al. 2023][5]; [Zhan et al. 2023][6]; [Vu et al. 2024][7]; [Wu et al. 2024][8]; [McClune et al. 2025][9]). Nevertheless, single cell mass spectrometry (scMS) has not been adapted for woody species, largely due to the challenges associated with their highly lignified tissue and complex cellular architecture. In this study, we developed an scMS pipeline for the woody tree C. acuminata to investigate the intercellular organization of camptothecin biosynthesis. ### Competing Interest Statement The authors have declared no competing interest. NSERC, ALLRP 571673 – 21, ALLRP 579871 – 22 Michael Smith Health BC, SCH-2020-0401 National Science Foundation, MCB-2309665 Georgia Research Alliance, https://ror.org/030689123 Georgia Seed Development Max Planck Gesellschaft [1]: #ref-2 [2]: #F1 [3]: #ref-5 [4]: #ref-4 [5]: #ref-1 [6]: #ref-8 [7]: #ref-6 [8]: #ref-7 [9]: #ref-3
The dry forests of northern Peru are dominated by the legumous tree Neltuma pallida which is adapted to hot arid and semiarid conditions in the tropics. Despite having been successfully introduced in multiple other areas around the world, N. pallida is currently threatened in its native area, where it is invaluable for the dry forest ecosystem and human subsistence. A major tool for enhancing ecosystem conservation and understanding the adaptive properties of N. pallida to dry forest ecosystems is the construction of a reference genome sequence. Here, we report on a high-quality reference genome for N. pallida. The final genome assembly size is 403.7 Mb, consisting of 14 pseudochromosomes and 63 scaffolds with an N50 size of 26.2 Mb and a 34.3% GC content. Use of Benchmarking Universal Single Copy Orthologs revealed 99.2% complete orthologs. Long terminal repeat elements dominated the repetitive sequence content which was 51.2%. Genes were annotated using N. pallida transcripts, plant protein sequences, and ab initio predictions resulting in 22,409 protein-coding genes encoding 24,607 gene models. Comparative genomic analysis showed evidence of rapidly evolving gene families related to disease resistance, transcription factors, and signaling pathways. The chromosome-scale N. pallida reference genome will be a useful resource for understanding plant evolution in extreme and highly variable environments.
Plant specialized metabolism is intricately regulated and often compartmentalized at the cell-type level. Understanding where and when metabolites accumulate is essential for uncovering their function, biosynthesis, and regulation. Historically, studies have inferred metabolite localization based on the expression patterns of genes encoding biosynthetic enzymes, but these approaches fall short due to the complexity of metabolite transport and the discrepancy between transcript, protein, and metabolite abundance. Recent advances in mass spectrometry imaging, single-cell transcriptomics, and multiomics have enabled the direct visualization and quantification of metabolites and gene expression at cellular resolution. These technologies have revealed striking cell type- and organ-specific patterns of metabolite accumulation, as well as the underlying transcriptional and chromatin regulatory networks. In this review, we describe case studies in several model and medicinal plant species that highlight the roles of rare or specialized cell types in specialized metabolite biosynthesis and the importance of spatiotemporal regulation. In addition, we discuss why it is becoming increasingly important to transition from single- to multiomics approaches. As new tools continue to evolve, the regulation of plant metabolism will be uncovered at higher resolution, enabling precise pathway discovery and metabolic engineering for agriculture, biotechnology, and medicine.
Potato is a key food crop with a complex, polyploid genome. Advancements in sequencing technologies coupled with improvements in genome assembly algorithms have enabled generation of phased, chromosome-scale genome assemblies for cultivated tetraploid potato. The SpudDB database houses potato genome sequence and annotation, with the doubled monoploid DM 1-3 516 R44 (hereafter DM) genome serving as the reference genome and haplotype. Diverse annotation data types for DM genes are provided through a suite of Gene Report Pages including gene expression profiles across 438 potato samples. To further annotate potato genes based on expression, 65 gene co-expression modules were constructed that permit the identification of tightly co-regulated genes within DM across development and responses to wounding, abiotic stress, and biotic stress. Genome browser views of DM and 28 other potato genomes are provided along with a download page for genome sequence and annotation. To link syntenic genes within and between haplotypes, syntelogs were identified across 25 cultivated potato genomes. Through access to potato genome sequences and associated annotations, SpudDB can enable potato biologists, geneticists, and breeders to continue to improve this key food crop.
BACKGROUND:Tocopherols are a class of lipid-soluble compounds that have multiple functional roles in plants and exhibit vitamin E activity, an essential nutrient for human and animal health. The tocopherol biosynthetic pathway is conserved across the plant kingdom, but source of the key tocopherol pathway precursor, phytol, is unclear. Two protochlorophyllide reductases (POR1 and POR2) were previously identified as loci controlling the natural variation of total tocopherols in maize grain, a non-photosynthetic tissue. POR1 and POR2 are key genes in chlorophyll biosynthesis yet the contribution of the chlorophyll biosynthetic pathway to tocopherol biosynthesis is still not understood. RESULTS:We took two approaches to alter the activity of these two POR genes within kernel tissue, physiological treatments and CRISPR/Cas9-mediated knockouts, to determine the role of chlorophyll biosynthesis for tocopherol content. Since light is required for POR enzymatic activity, we imposed a dark treatment on developing kernels, which reduced chlorophyll a and tocopherols levels in embryo tissue by 92-99% and 87-90%, respectively, compared to the light treatment. In CRISPR/Cas9-mediated knockouts, the levels of chlorophyll a and tocopherols in embryos of the por1 por2 double homozygous mutant were reduced by 98-100% and 76-83%, respectively, compared to WT. CONCLUSION:These findings demonstrate that tocopherol synthesis in maize grain depends almost entirely on phytol derived from chlorophyll biosynthesis within the embryo. POR1 and POR2 activity play crucial roles in chlorophyll biosynthesis, underscoring the importance of POR alleles and their activity in the biofortification of vitamin E levels in non-photosynthetic grain of maize.
Potato (Solanum tuberosum) is the third-most important food crop in the world. Although the potato genome has been fully sequenced, functional genomics research of potato lags behind that of other major food crops, largely due to the lack of a model experimental potato line. Here, we present a diploid potato line, 'Jan,' which possesses all essential characteristics for facile functional genomics studies. Jan exhibits a high level of homozygosity after seven generations of self-pollination. Jan is vigorous, highly fertile and produces tubers with outstanding traits. Additionally, it demonstrates high regeneration rates and excellent transformation efficiencies. We generated a chromosome-scale genome assembly for Jan, annotated its genes and identified syntelogs relative to the potato reference genome assembly DMv6.1 to facilitate functional genomics. To miniaturize plant architecture, we developed two 'mini-Jan' lines with compact and dwarf plant stature through CRISPR/Cas9-mediated mutagenesis targeting the Dwarf and Erecta genes involved in growth. One mini-Jan mutant, mini-JanE, is fully fertile and will permit higher-throughput studies in limited growth chamber and greenhouse space. Thus, Jan and mini-Jan offer a robust model system that can be leveraged for gene editing and functional genomics research in potato.
The hexaploid sweetpotato (Ipomoea batatas [L.] Lam.) is a globally important stable crop that plays a key role in biofortification. Its high resilience and adaptability provide distinct advantages in addressing food security and climate challenges. Here we report a haplotype-resolved chromosome-level genome assembly of an African cultivar, ‘Tanzania’, revealing mosaic genomic origins along haplotype-phased chromosomes. The wild tetraploid I. aequatoriensis, currently found in coastal Ecuador, contributes to a substantial fraction of the sweetpotato genome. Another large proportion of the genome shows a closer genetic relationship to the wild tetraploid I. batatas 4×, distributed in Central America. The sequences contributed by different wild species are not distributed in typical subgenomes but are intertwined along chromosomes, possibly owing to the known non-preferential recombination among sweetpotato haplotypes. This study improves our understanding of sweetpotato origin and genome architecture and provides valuable genomic resources to accelerate sweetpotato breeding. This study developed a phased chromosome-level genome assembly for the African sweetpotato cultivar ‘Tanzania’, revealing the complex genome architecture of hexaploid sweetpotato and providing valuable insights into its origin.
Tubers are modified underground stems that enable asexual, clonal reproduction and serve as a mechanism for overwintering and avoidance of herbivory. Potato (Solanum tuberosum L.) is cultivated for its tubers, which serve as a major crop. Genes responsible for tuber initiation and disease resistance have been characterized in potato including StSP6A, a homolog of flowering time, that functions as a tuberigen, the equivalent of a florigen. To elucidate additional molecular and genetic mechanisms underlying potato biology including tuber initiation, tuber development, and stress responses, we generated a developmental and abiotic/biotic-stress gene expression atlas from 34 tissues and treatments of the tetraploid potato cultivar, Atlantic. Using the haplotype-phased tetraploid Atlantic genome assembly and expression abundances of 129 218 genes, we constructed gene coexpression modules that represent networks associated with distinct developmental stages as well as stress responses. Functional annotations were given to modules and used to identify genes involved in tuberization and stress resilience. Structural variation from a pan-genomic analysis across four cultivated potato genome assemblies as well as domestication and wild introgression data allowed for deeper insights into the modules to identify key genes involved in tuberization and stress responses. This study underscores the importance of transcriptional regulation in tuberization and provides a comprehensive framework for future research on potato development and improvement.
Solanum verrucosum Schlechtendal (2x = 2n = 24) is unique among the clade 4 Solanum Sect Petota species. In addition to being one of the only fully self-compatible diploid potato species, S. verrucosum is the only clade 4 species that lacks prezygotic interspecific reproductive barriers. This allows S. verrucosum to accept pollen from a broad range of Solanum species and thereby serving as a genetic "bridge" between the cultivated or primary potato gene pool and distantly related wild relatives in the tertiary gene pool. The genetic mechanisms underlying self-compatibility in Solanum often underpin interspecific compatibility interactions, which in S. verrucosum, has been attributed to the lack of S-RNase expression. Using an interspecific F2 mapping population (n = 150), we investigated the genetic mechanisms responsible for the lack of interspecific reproductive barriers in S. verrucosum. This F2 population was evaluated for the ability to accept pollen from two clade 1, 1 EBN species (S. pinnatisectum and S. tarnii); from which two QTL for interspecific compatibility were identified on chromosomes 1 and 11, explaining 56.6% of the phenotypic variation observed. To identify the genetic basis of interspecific compatibility, we generated a chromosome-scale genome assembly of S. verrucosum MSII1813-2 and performed gene expression profiling of reproductive organs. Differential gene expression of S-RNase, located within the chromosome 1 QTL, confirmed the central role of the S-locus and specifically, S-RNase, in interspecific compatibility. Discovery of a non-S-locus QTL is consistent with previous findings that other non-S-locus factors are necessary for interspecific compatibility in S. verrucosum.
BACKGROUND:Photoperiodic changes in diel cycles of gene expression are pervasive in plants. The timing of circadian regulators, together with light signals, regulate multiple photoperiod-dependent responses such as growth, flowering or tuber formation. However, for most genes, the importance of cyclic mRNA levels is less clear. We analyzed the diel transcriptome of modern cultivated potato, a highly heterozygous autotetraploid. Clonal propagation and limited meiosis have led to the accumulation of deleterious alleles, making tetraploid potato an ideal model system to investigate the conservation of cyclic expression and cyclic genes during artificial selection and clonal propagation. RESULTS:Our results indicate that rhythmic alleles of cultivated potato are more highly expressed than non-rhythmic genes and are highly co-expressed not only under diel cycles but also across tissues, developmental stages, and stress conditions. Moreover, the smaller ratio of non-synonymous to synonymous differences within rhythmic versus non-rhythmic allelic groups indicates that cyclic genes, in general, have more conserved core functions than non-cyclic genes. In accordance with this observation, fully rhythmic allelic groups are highly enriched in photosynthesis and ribosome biogenesis genes, which have core functions in plants. Furthermore, we investigated differences in cyclic expression patterns between photoperiods identifying potential regulators for the strong changes in phase of expression of ribosome biogenesis and pathogen response genes. Finally, analyses of genes involved in tuber formation suggests that the regulation of CO gene transcription is not the only factor enabling tuberization under long days in modern cultivated potato. CONCLUSIONS:This study not only provides high quality diel transcriptomic datasets of cultivated potato but also provides important insight on the role of allelic diversity in rhythmic expression in plants.
Camptotheca acuminata Decne is a woody medicinal tree that produces over a hundred bioactive compounds, including camptothecin, which has been used as the starting material to semi-synthesize many leading anticancer drugs ([Lorence and Nessler 2004][1]). Camptothecin and its derivatives are potent inhibitors of DNA topoisomerase I and are widely used for the treatment of lung, cervical, ovarian, and colon cancers. Camptothecin biosynthesis in C. acuminata involves complex catalytic steps, most of which remain undeciphered. In this pathway, tryptamine and secologanic acid are coupled, leading to strictosidinic acid. The formation of strictosidinic acid is catalyzed by strictosidine/strictosidine acid syn-thase enzymes (STR) ([Fig. 1A][2]). While a biosynthetic route for the conversion of the indole ring to the quinoline ring has been proposed, most of the underlying biosynthetic genes have yet to be identified ([Fig. 1A][2]) ([Sadre et al. 2016][3]). In addition, the cell type specificity of this pathway also remains undescribed. Here, we generated a single cell multiome (RNA-seq and Assay for Transposase Accessible Chromatin by sequencing [ATAC-seq] from the same nuclei) to probe the cell type specificity of camptothecin biosyn-thetic genes. ### Competing Interest Statement The authors have declared no competing interest. Georgia Research Alliance, https://ror.org/030689123 Georgia Seed Development National Science Foundation, MCB-2309665 Max Planck Gesellschaft NSERC Alliance Collaboration, ALLRP 571673 – 21 Catalyst, ALLRP 579871 – 22 [1]: #ref-4 [2]: #F1 [3]: #ref-6
CRISPR/Cas9 is the most popular genome editing platform for investigating gene function or improving traits in plants. The specificity of gene editing has yet to be evaluated at a genome-wide scale in seed-propagated Camelina sativa (L.) Crantz (camelina) or clonally propagated Solanum tuberosum L. (potato). In this study, seven potato and nine camelina stable transgenic Cas9-edited plants were evaluated for on and off-target editing outcomes using 55x and 60x coverage whole genome shotgun sequencing data, respectively. For both potato and camelina, a prevalence of mosaic somatic edits from constitutive Cas9 expression was discovered as well as evidence of transgenerational editing in camelina. CRISPR/Cas9 editing provided negligible off-target activity compared to background variation in both species. The results from this study guide deployment and risk assessment of genome editing in commercially relevant traits in food crops.