Vegetables are a critical component of diets, with inadequate intake of this essential food group leading to poor dietary quality and malnutrition. Food system assessments identify insufficient production, comparatively high prices, and sociocultural barriers as key constraints to vegetable consumption. We argue that vegetable biodiversity, spanning vegetable species and their varieties, as well as their wild relative species, is a central yet underutilized lever for enhancing vegetable consumption. Vegetables span a wider phylogenetic range than any other plant-derived food group, offering options for different climatic, cultural, and market niches worldwide. However, vegetable biodiversity is declining due to market homogenization, land-use change, and other threats. Its current conservation is insufficient, restricting in turn access to this diversity for research, breeding, and innovation, and making it more difficult to bridge the gap between current and recommended vegetable intake. Reversing this trend globally requires aligning conservation with dietary goals through four complementary action areas: i) Securing vegetable biodiversity by collecting, regenerating, and conserving local crop varieties and wild relatives of key vegetable species in biodiversity hotspots; ii) Harnessing vegetable biodiversity to deliver new varieties through collaborative research, breeding, and variety testing; iii) Promoting vegetable biodiversity to diversify diets, particularly among children and other vulnerable groups, by including nutrient-dense, climate-resilient vegetables into home and school meals; and iv) Integrating vegetable biodiversity into policy frameworks for long-term impact. Implementing this integrated approach in hotspots where vegetable biodiversity and malnutrition overlap can transform an overlooked opportunity into a cornerstone strategy for healthier diets.
Abstract Rosa, belonging to the family Rosaceae, encompasses more than 150 species widely distributed across the northern hemisphere. Renowned for their beauty, roses are cultivated throughout the world for ornamental purposes and the production of essential oils and perfumes. Despite their cultural and commercial significance, the genomic resources of wild Rosa species have not been studied comprehensively, hampering the understanding of their genetic diversity, evolutionary history, and breeding potential. Here we present a Rosaceae panproteome and a Rosa pangenome, spanning wild, traditional garden, and modern rose lineages, constructed using a De Bruijn graph (DBG)-based approach, and introduce two high-quality de novo genomes for Rosa sericea and Rosa rugosa. A phylogeny of 18 Rosa haplotypes based on 4367 single-copy core homology groups (genes) provided robust evolutionary inference. Our analysis revealed substantial interspecific genomic diversity in core gene repertoires, structural features, and a transposable element (TE) landscape that shaped genome size differences and is potentially linked to phenotypic plasticity. We provide two examples of the types of analyses that become possible with this pangenome. First, the pangenome serves as a quality-aware lens, exposing discrepancies arising from assembly and annotation variability and helping separate technical artifacts from genuine biological signal. Second, the pangenome provides locus-level resolution: analysis of MYB114, a key regulator of anthocyanin accumulation, reveals lineage-specific presence–absence patterns and TE-associated regulatory variation. This pangenomic study deepens our understanding of the genetic diversity and genome evolution of Rosa species and establishes a resource to resolve the genetic bases of key traits, thereby informing and supporting rose breeding.
The shea tree (Vitellaria paradoxa C.F. Gaertn.) is an important source of oil, primarily used in the chocolate, cosmetics, and food industries. However, the absence of elite domesticated varieties limits breeding and optimal utilization. This study explores the genetic diversity, population structure and biochemical variation to identify genotypes with high oilseed potential in Benin. A total of 167 trees were genotyped using DArTseq technology, generating 3,720 high-quality single nucleotide polymorphism (SNP) markers after strict filtering. Data for two years (2023-2024) were collected, including agromorphological and biochemical data (fat content and fatty acid profile). Genetic diversity analysis revealed a moderate genetic diversity, with expected heterozygosity values of He = 0.21. Analysis of Molecular Variance (AMOVA) revealed significant differentiation (p < 0.01), with 7.9% of variance attributed to differences among populations. Population structure analysis distinguished three genetic groups (Pop1, Pop2, Pop3), with strong differentiation between Pop1 and Pop3 (FST > 0.85). Correlation analysis revealed positive relationships between nut traits and fatty acid composition. Five major fatty acids were found in the shea kernels: stearic (C18:0; 141.5 mg/g approximate to 45.6%), oleic (C18:1n-9, 125.49 mg/g approximate to 43.7%), linoleic (C18:2n-6, 17.43 mg/g approximate to 5.98%), palmitic (C16:0, 9.9 mg/g approximate to 3.6%), and arachidic (C20:0, 4.63 mg/g approximate to 1%). Genotypes BG104, BG109, BG114, BG214 and BG258 were elite candidates for fat and fatty acid content. This study offers an overview of the genetic potential of shea trees in Benin, identifying the best genotypes with high concentrations of fat and fatty acid composition for future breeding programs.
The shea tree (Vitellaria paradoxa C.F. Gaertn.) is an important source of oil, primarily used in the chocolate, cosmetics, and food industries. However, the absence of elite domesticated varieties limits breeding and optimal utilization. This study explores the genetic diversity, population structure and biochemical variation to identify genotypes with high oilseed potential in Benin. A total of 167 trees were genotyped using DArTseq technology, generating 3,720 high-quality single nucleotide polymorphism (SNP) markers after strict filtering. Data for two years (2023–2024) were collected, including agromorphological and biochemical data (fat content and fatty acid profile). Genetic diversity analysis revealed a moderate genetic diversity, with expected heterozygosity values of He = 0.21. Analysis of Molecular Variance (AMOVA) revealed significant differentiation (p < 0.01), with 7.9
Quantitative genetic parameters associated with the inheritance of biomass yield and related traits were evaluated in Gynandropsis gynandra for selecting appropriate breeding methods for cultivar development. A total of 331 F1 hybrids generated from a North Carolina mating design II and their 39 parental lines were evaluated in the field and greenhouse for gene action, combining ability effects and heterosis value of the biomass yield and related traits. The evaluation was done across seven environments between 2019 and 2021 using an alpha lattice design with two replicates per environment. Significant differences (P < 0.05) were observed among and between hybrids and parents for all agronomic traits. Overall, hybrids performed better than their parents for stem diameter (21.1
BACKGROUND AND AIMS:The Brassiceae tribe encompasses many economically important crops and exhibits high intra- and interspecific phenotypic variation. After a shared whole-genome triplication (WGT) event (Br-α, ~15.9 Mya), differential lineage diversification and genomic changes contributed to an array of divergence in morphology, biochemistry and physiology underlying photosynthesis-related traits. Here, the C3 species Hirschfeldia incana is studied because it displays high photosynthetic rates in high-light conditions. Our aim was to elucidate the evolution that gave rise to the genome of H. incana and its high-photosynthesis traits. METHODS:We reconstructed a chromosome-level genome assembly for H. incana (Nijmegen, v.2.0) using nanopore and chromosome conformation capture (Hi-C) technologies, with 409 Mb in size and an N50 of 52 Mb (a 10× improvement over the previously published scaffold-level v.1.0 assembly). The updated assembly and annotation were subsequently used to investigate the WGT history of H. incana in a comparative phylogenomic framework from the Brassiceae ancestral genomic blocks and related diploidized crops. KEY RESULTS:Hirschfeldia incana (x = 7) shares extensive genome collinearity with Raphanus sativus (x = 9). These two species share some commonalities with Brassica rapa and Brassica oleracea (A genome, x = 10 and C genome, x = 9, respectively) and other similarities with Brassica nigra (B genome, x = 8). Phylogenetic analysis revealed that H. incana and R. sativus form a monophyletic clade in between the Brassica A/C and B genomes. We postulate that H. incana and R. sativus genomes are results of hybridization or introgression of the Brassica A/C and B genome types. Our results might explain the discrepancy observed in published studies regarding phylogenetic placement of H. incana and R. sativus in relationship to the 'triangle of U' species. Expression analysis of WGT retained gene copies revealed sub-genome expression divergence, probably attributable to neo- or sub-functionalization. Finally, we highlight genes associated with physio-biochemical-anatomical adaptive changes observed in H. incana, which are likely to facilitate its high-photosynthesis traits under high light. CONCLUSIONS:The improved H. incana genome assembly, annotation and results presented in this work will be a valuable resource for future research to unravel the genetic basis of its ability to maintain a high photosynthetic efficiency in high-light conditions and thereby improve photosynthesis for enhanced agricultural production.
Motivation In the era of multiple genome references, researchers often align sequencing reads against distinct assemblies or even multiple references simultaneously. This enables applications such as the detection of introgressed segments or highly variable genomic regions, which are especially prevalent in large-genome crop species such as lettuce or wheat. However, these applications come at the cost of increased computational burden, inconsistencies in mapping methods, and reduced reproducibility across studies. To address these limitations, we developed KCFtools, a Java-based toolkit that identifies the presence and absence of k-mers in nonoverlapping genomic or transcriptomic windows by comparing query and reference genomes. This alignment-free approach enables the efficient computation of an identity score for each window, thereby facilitating robust detection of introgressed or variable regions across genomes.Results We systematically evaluated the performance and accuracy of the k-mer-based method implemented in KCFtools, benchmarking it against conventional single nucleotide variation-based introgression detection pipelines. Our results demonstrate that KCFtools effectively captures introgressed segments and structurally diverse regions, even in species with fragmented or highly divergent reference genomes. In addition, we extended KCFtools to generate genotype matrices from k-mer variation tables. These matrices are compatible with genome-wide association studies software and allow the identification of loci associated with phenotypic traits. We showcase the utility of this approach by detecting known and novel associations for downy mildew resistance in lettuce, underscoring the pipeline's potential for high-resolution, reference-agnostic population genetic analysis.Availability and implementation https://github.com/sivasubramanics/kcftools
Sundaland, a subcontinent situated in the Indo-Malay Archipelago, is characterized by a dynamic climatic, geological, and geographical history and a correspondingly diverse biota. The distribution of this biota follows complex biogeographical patterns, with high levels of local endemism and frequent instances of hidden diversity-greater genetic differentiation than expected based on phenotype. Biogeographical studies on Sundaic insects are scarce, despite high species richness and complex distributions. By applying molecular "museomics" techniques to a species group of Troides birdwing butterflies, we generated inter- and intraspecific datasets comprising 39,485 to 59,905 SNP positions and complete mitogenomes from museum specimens collected between the 1890s and 1990s. These datasets allowed the investigation of taxonomic status, biogeography, and divergence timing within the Sundaland endemic amphrysus group. The origin of the group was estimated between the late Miocene and early Pliocene, and subsequent intraspecific divergence took place during the Pleistocene. Earlier-diverging taxa were found on Borneo, while later-diverging taxa had distributions centered in western Sundaland. Divergence between subspecies of later-diverging species complexes occurred <1 million years ago, overlapping with marine incursions on the Sunda Shelf. Javan populations showed comparatively high genetic differentiation from those on other Sundaic landmasses. We found genetic evidence of possible hidden species-level diversity within two species complexes. Contrastingly, T. amphrysus euthydemus (Fruhstorfer 1913) syn. nov. is synonymized with T. amphrysus ruficollis (Butler 1879), given the low divergence between these two species. Together, these results provide a perspective on Triodes evolutionary dynamics and the role of Pleistocene glaciations in shaping patterns of Sundaland's insect diversity.
Abstract Plants produce the most diverse blends of specialized metabolites on earth. Natural products derived from plants are valuable resources for drug development, food chemistry, and crop resistance breeding. Phenotypes of specialized metabolite profiles can be captured by untargeted mass-spectrometry across species phylogeny, tissues, and genotypes. Here, we collected metabolic fingerprints of 17 Brassicaceae species across three tissues (paired leaf and root; flower) using liquid chromatography-tandem mass spectrometry (LC-MS/MS) in positive and negative ionization mode. Corresponding metadata has been refined for reuse according to ReDU guidelines, and for integration with public genomic and transcriptomic data. Standardization of in vitro growth conditions, and data processing workflows enables integration of acquired raw and processed data across platforms for single- and multi-omics analysis. Further, the inclusion of tissue-specific metabolic profiles across ploidy levels, as well as across crop species and wild relatives, makes this dataset a valuable resource for natural product discovery.
The Asteraceae (Compositae) is the largest flowering plant family, ubiquitous in most terrestrial communities, and morphologically diverse. A two-step, ancient whole genome triplication (paleohexaploidization) occurred at approximately the same time as the evolutionary innovation and adaptive radiation of the family during the middle Eocene. Despite its importance, the consequences of this triplication have yet to be tracked in context of the Asteraceae genome evolution. To do so, we applied a synteny oriented phylogenomic analysis of 23 Asterales genomes. We identified 16 genomic groups that date back to the common diploid ancestor of all Asteraceae. Each group underwent triplication, resulting in 48 genomic blocks (16 × 3) that collectively represent the ancestral Asteraceae genome, excluding the early-diverging lineages which do not share the second step. We then analyzed the evolutionary dynamics of the 48 genomic blocks across the Asteraceae phylogeny. We found that modern Asteraceae genomes are genetic mosaics of three progenitor genomes, shaped by genomic exchanges, chromosomal rearrangements, and gene fractionation. One hundred fifty-seven genes retained three paleohexaploid-derived syntenic paralogs across most Asteraceae species. Transcription factors and auxin-related genes are significantly overrepresented in these triplets, and expression of the paleohexaploidy paralogs is spatiotemporally differentiated. These genes are involved in the development of floral capitula, a remarkable morphological innovation of the family. The discovery of the 157 triplicated genes can direct further study to understand the evolutionary innovation, and the synteny-phylogenomic framework provides a comparative framework to characterize newly sequenced Asteraceae genomes.
Abstract Understanding the genome‐wide diversity of jute mallow (Corchorus olitorius L.) is crucial for unlocking the potential of global genebank collections, enabling the discovery and use of traits that support climate resilience, improve nutrition, and increase productivity. Using 23,471 high‐quality diversity array technology sequencing single‐nucleotide polymorphisms (SNPs), this study assessed the genetic diversity, population structure, and linkage disequilibrium (LD) of 607 accessions. Moderate genetic diversity was detected with a total gene diversity of 0.28, an expected heterozygosity of 0.26, and a Shannon index of 0.42. Four distinct genetic clusters were identified, reflecting geographic patterns, where Cluster 1 (n = 62) and Cluster 4 (n = 354) were predominantly composed of West African accessions. An analysis of molecular variance revealed significant genetic structuring (p < 0.001), with most genetic variation occurring within countries (45.2%), followed by within individuals (32.5%), while differentiation among clusters accounted for 18.2% and variation among regions was minimal (2.9%). LD revealed low genome‐wide r2 values (mean = 0.028; r290 = 0.067) and a very rapid decay (LD50 ≈ 1 bp), with only 4.2% of SNP pairs showing significant LD (r2 > 0.1, p < 0.05), indicating extensive historical recombination. The findings suggest that a significant portion of the existing genetic variation remains untapped in breeding. Strategic conservation of the unique genetic variants through core and mini‐core collections, coupled with targeted crosses among diverse regional accessions, can broaden the genetic base and support the development of resilient, high‐yielding, and nutrient‐rich dual‐purpose varieties (i.e., leafy vegetables and industrial fibers) across diverse environments.
Abstract Not all odors influencing mating behavior evolve as sex pheromones. Female butterflies’ post-mating odors have been considered species-specific anti-aphrodisiac pheromones shaped by sexual selection but may also serve broader ecological roles shaped by natural selection. Males transfer odors to females that repel rivals, yet the widespread use of these compounds across phyla makes them targets for eavesdropping, such as by phoretic egg parasitoids. We show that in cabbage white butterflies ( Pieris spp.), these odors are highly variable and attract parasitoids, deter predators, and influence oviposition. Using gas chromatography and electroantennography, we demonstrate that odor emission and perception lack species-specificity: compounds once thought unique to P. brassicae and P. rapae are shared across Pieridae. In P. napi , odor variation among populations correlates with parasitoid pressure, but not with latitude, genetic distance, or mating frequency, suggesting ecological rather than sexual drivers. In P. brassicae , CRISPR/Cas9 disruption of odor perception alters oviposition and increases susceptibility to parasitism. Moreover, these odors render females unpalatable to birds. Together, our results show that post-mating odors in Pieris butterflies may act as aposematic signals. We provide evidence that these signals evolve under multiple selective pressures, balancing deterrence of mates and predators, parasitoid avoidance, and host-plant interactions. These findings suggest that chemical signals should be viewed as integrating ecological and reproductive pressures, rather than being interpreted solely through the lens of sexual communication.
Urbanisation is rapidly expanding worldwide, posing a significant threat to biodiversity through habitat loss, fragmentation, and degradation. To counter these negative effects, restoring sufficient interconnected green spaces is crucial. One promising solution is an urban green infrastructure network including green roofs. Despite being promoted for their ecosystem services and habitat connectivity, the actual contribution of green roofs to biodiversity, especially for insects, is debatable. This review challenges current thinking in green roof design through an ecological perspective on green roof development, highlighting the importance of insects. Our approach encompasses: (1) an overview of the different roof types and how they influence insect diversity, (2) a summary of the impact of green roof factors on insect diversity, and (3) an application of neutral and niche theory to understand the ecological processes – both stochastic and deterministic – that shape insect communities on green roofs over time. We highlight how these processes drive the development towards increasingly complex green roof communities, which are crucial for resilient, biodiverse and low‐maintenance green roofs. Finally, (4) we discuss the application of network analysis and other techniques to assess patterns in ecosystem development. This review aims to inform landscape planning, design, management, and policy on how to build towards a more biodiverse green roof landscape in cities. This approach aligns with the EU Biodiversity Strategy for 2030, advocating for the strategic use of blue and green infrastructure to enhance urban biodiversity.
Summary Plants produce diverse bouquets of specialized metabolites (SMs), yet only a fraction of the vast phytochemical space has been explored to date. Comparative analysis of SM profiles can reveal hotspots of biochemical novelty, while systematic profiling across taxonomic levels does presently not cover large plant families. To study core and accessory SM profiles in the Brassicaceae plant family, we fingerprinted 14 species by Liquid-Chromatography Mass-Spectrometry (LCMS/MS). We develop standardized experimental and computational workflows integrating in silico annotation tools to study consensus compound class and substructure distributions of SMs. Furthermore, we investigate the congruence of chemotaxonomy and species phylogeny across an extended panel of 17 species. Unique metabolite profiles were outstanding in Camelina sativa, Capsella rubella , and B. vulgaris , with the largest unique terpenoid profile annotated in C. sativa , accounting for 33.5% and 55.6% in positive and negative ionization mode, respectively. Substructure motifs were found to overlap with compound class predictions, highlighted for triterpenoids in Camelinodae. Furthermore, dual-tissue chemotaxonomic clustering resembled relationships of Brassica subgenomes across tissues. We anticipate that our systematic approach can serve as a blueprint for investigating biochemical diversity in other plant lineages and can boost the characterization of plant natural product pathways.
Derived woodiness—the evolution of woody growth from herbaceous ancestors—has arisen hundreds of times across flowering plants, yet the environmental conditions associated with its repeated evolution remain poorly understood. Here, we analyse woodiness evolution in the mustard family (Brassicaceae; ~4,150 species) using a time-calibrated phylogeny of 2,927 species, including 374 of the 385 known woody species, together with global growth-form and climatic niche data. We infer 231 independent origins of woodiness alongside 176 reversals to herbaceousness, indicating that woodiness evolves repeatedly but remains evolutionarily unstable. Although woody species are enriched on islands, most occur on the mainland, where woodiness is consistently associated with drought and reduced frost. Correlated-evolution analyses reveal that drought is associated primarily with the persistence of woodiness, whereas reduced frost is associated with gains of woodiness and increased frost with its loss. These findings identify distinct climatic associations with the gain, persistence, and loss of woody growth forms.
Lettuce (Lactuca sativa L.) is an economically important leafy vegetable within the Asteraceae family, cultivated worldwide across diverse agricultural systems. Recent advances in genomic and transcriptomic resources have positioned lettuce as a promising model system for functional genomics in the Asteraceae. However, currently available gene expression datasets lack comprehensive tissue-specific resolution, primarily focus on a single cultivar and are not visualised in an interpretable manner, limiting their utility for broader genetic and physiological studies. To bridge this gap, we developed the Lettuce Expression Browser (LEB), a publicly available platform providing high-resolution gene expression maps across various organs, tissues and developmental stages in both cultivated and wild lettuce species. The LEB integrates transcriptomic data from finely dissected seedlings, shoot tissues at various developmental stages and seedlings subjected to abiotic stresses (salt and far-red), visualised using the ggPlantmap R package. This platform offers an intuitive interface for exploring gene expression patterns and serves as a valuable resource for those studying lettuce development, stress responses and evolutionary genomics. The LEB is hosted on the LettuceKnow Web Portal (https://lettuce.bioinformatics.nl) and can be expanded to include additional datasets, enhancing its role as a key tool for lettuce research and crop improvement.
With the current speed of sequencing, there is a desire for standardized and automated genome assembly and annotation to produce high-quality genomes as input for comparative (pan)genomics. Therefore, we created a convenience pipeline using existing tools that creates annotated genome assemblies from HiFi (and optionally ultra-long ONT and/or Hi-C) reads for a set of related individuals as well as a related reference genome. Our pipeline is species-agnostic and generates an extensive quality assessment report that can be used for manual filtering and refinement of the assembly and annotation. It includes statistics for individual completeness and contamination assessments as well as a concise pangenome view. The pipeline is implemented in Snakemake and available with a GPLv3 licence at GitHub under github.com/dirkjanvw/MoGAAAP, at Zenodo under doi.org/10.5281/zenodo.14833021, and can be installed through Bioconda.
The Solanum lycopersicum (tomato) genome encodes seven malic enzymes (MEs): three cytosolic NADP-ME isoforms (SlNADP-ME1, -2, and -3), two plastidic NADP-ME proteins (SlNADP-ME4a and -4b), and two mitochondrial NAD-dependent subunits (SlNAD-ME1 and -2) of a heteromeric enzyme. Except for SlNADP-ME3, which is almost exclusively expressed in the root, all the genes are active in the major tissues of flowering plants (leaf, stem, flower, and root), with SlNADP-ME2 and -4b transcripts accumulating at relative high levels. During the cell expansion and ripening phases of fruit development, SlNADP-ME3 and -4b expression increased in the pericarp and seeds while SlNADP-ME2 remained at a high level. In green fruit, SlNADP-ME3 and -4b are down-regulated by ethylene, with SlNAD-ME1 and -2 being up-regulated. A correlation between SlNADP-ME4b accumulation and NADP-ME activity was observed from the early immature green to the mature green fruit stages, linking this plastidic isoform to starch and lipid biosynthesis. SlNADP-ME1 and -4a expression increase with temperature, suggesting involvement in defense mechanisms and supported by cis-elements composition in their promoters. Interesting, SlNADP-ME3 biochemical properties and its accumulation in seeds is part of an inherited and long-conserved genetic lineage in Angiosperms including Arabidopsis NADP-ME1, involved in normal seed germination. By analyzing the phylogeny and synteny of SlNAD(P)-ME genes and determining the biochemical properties of the recombinant proteins, we seek functional similarities to Arabidopsis NAD(P)-ME members. Our findings enhance the understanding of malate metabolism in tomatoes, which could inform strategies to improve fruit quality. Additionally, they prompt a re-evaluation of pre-existing notions regarding the functional orthology of enzymes based on phylogenetic relationships. ### Competing Interest Statement The authors have declared no competing interest.
IntroductionUnderstanding the genome-wide variation pattern in crop germplasm is required in profiling breeding products and defining conservation units. Yet, such knowledge was missing for the large germplasm collection of Corchorus olitorius in Benin at CalaviGen (the University of Abomey-Calavi genebank), the world’s largest holder of the crop germplasm with 1,566 accessions conserved.MethodsUsing 1,114 high-quality SNPs, this study: i) investigated the spatial variation of the genetic structure of 305 accessions sampled along the South-North ecological gradient of Benin, ii) derived a core collection from the batch of accessions and iii) gauged the extent of (mal)adaptation of this core set.Results and discussionOverall, we detected a moderate diversity with a total gene diversity of 0.28 and an expected heterozygosity estimate of 0.27. The spatial variation of the genomic diversity painted an increasing trend following the South-North ecological gradient, giving rise to four optimal genetic groups based on STRUCTURE analysis while the neighbour-joining analysis revealed three clusters. The ShinyCore algorithm application yielded a core set of 54 accessions that echoed a good geographical representativeness and encompassed a level of diversity comparable to that of the whole collection. Nearly 88% of this core set accessions were characterized by a low genomic offset score, which suggests a strong adaptation potential to future climate. This SNP-based core collection represents a unique and viable working asset for accelerated traits-discovery, in the species and should play a pivotal role in international collaborative initiatives dedicated to promoting C. olitorius use and conservation.
Following publication, concerns were raised to the Editorial Office relating to a potential conflict of interest between one of the authors and the Academic Editor that supervised the peer-review of this article [...]