Abstract Rosa, belonging to the family Rosaceae, encompasses more than 150 species widely distributed across the northern hemisphere. Renowned for their beauty, roses are cultivated throughout the world for ornamental purposes and the production of essential oils and perfumes. Despite their cultural and commercial significance, the genomic resources of wild Rosa species have not been studied comprehensively, hampering the understanding of their genetic diversity, evolutionary history, and breeding potential. Here we present a Rosaceae panproteome and a Rosa pangenome, spanning wild, traditional garden, and modern rose lineages, constructed using a De Bruijn graph (DBG)-based approach, and introduce two high-quality de novo genomes for Rosa sericea and Rosa rugosa. A phylogeny of 18 Rosa haplotypes based on 4367 single-copy core homology groups (genes) provided robust evolutionary inference. Our analysis revealed substantial interspecific genomic diversity in core gene repertoires, structural features, and a transposable element (TE) landscape that shaped genome size differences and is potentially linked to phenotypic plasticity. We provide two examples of the types of analyses that become possible with this pangenome. First, the pangenome serves as a quality-aware lens, exposing discrepancies arising from assembly and annotation variability and helping separate technical artifacts from genuine biological signal. Second, the pangenome provides locus-level resolution: analysis of MYB114, a key regulator of anthocyanin accumulation, reveals lineage-specific presence–absence patterns and TE-associated regulatory variation. This pangenomic study deepens our understanding of the genetic diversity and genome evolution of Rosa species and establishes a resource to resolve the genetic bases of key traits, thereby informing and supporting rose breeding.
Motivation In the era of multiple genome references, researchers often align sequencing reads against distinct assemblies or even multiple references simultaneously. This enables applications such as the detection of introgressed segments or highly variable genomic regions, which are especially prevalent in large-genome crop species such as lettuce or wheat. However, these applications come at the cost of increased computational burden, inconsistencies in mapping methods, and reduced reproducibility across studies. To address these limitations, we developed KCFtools, a Java-based toolkit that identifies the presence and absence of k-mers in nonoverlapping genomic or transcriptomic windows by comparing query and reference genomes. This alignment-free approach enables the efficient computation of an identity score for each window, thereby facilitating robust detection of introgressed or variable regions across genomes.Results We systematically evaluated the performance and accuracy of the k-mer-based method implemented in KCFtools, benchmarking it against conventional single nucleotide variation-based introgression detection pipelines. Our results demonstrate that KCFtools effectively captures introgressed segments and structurally diverse regions, even in species with fragmented or highly divergent reference genomes. In addition, we extended KCFtools to generate genotype matrices from k-mer variation tables. These matrices are compatible with genome-wide association studies software and allow the identification of loci associated with phenotypic traits. We showcase the utility of this approach by detecting known and novel associations for downy mildew resistance in lettuce, underscoring the pipeline's potential for high-resolution, reference-agnostic population genetic analysis.Availability and implementation https://github.com/sivasubramanics/kcftools
Lettuce (Lactuca sativa L.) is an economically important leafy vegetable within the Asteraceae family, cultivated worldwide across diverse agricultural systems. Recent advances in genomic and transcriptomic resources have positioned lettuce as a promising model system for functional genomics in the Asteraceae. However, currently available gene expression datasets lack comprehensive tissue-specific resolution, primarily focus on a single cultivar and are not visualised in an interpretable manner, limiting their utility for broader genetic and physiological studies. To bridge this gap, we developed the Lettuce Expression Browser (LEB), a publicly available platform providing high-resolution gene expression maps across various organs, tissues and developmental stages in both cultivated and wild lettuce species. The LEB integrates transcriptomic data from finely dissected seedlings, shoot tissues at various developmental stages and seedlings subjected to abiotic stresses (salt and far-red), visualised using the ggPlantmap R package. This platform offers an intuitive interface for exploring gene expression patterns and serves as a valuable resource for those studying lettuce development, stress responses and evolutionary genomics. The LEB is hosted on the LettuceKnow Web Portal (https://lettuce.bioinformatics.nl) and can be expanded to include additional datasets, enhancing its role as a key tool for lettuce research and crop improvement.
Due to their ability to kill closely related strains, phage tail-like bacteriocins, also called tailocins, play an important role in shaping bacterial communities. One such tailocin, called carotovoricin, is also known to be present in the Pectobacterium genus. However, little is known about its evolutionary dynamics and the scope of impact on species interactions in this genus. To investigate the diversity and evolution of carotovoricin, we performed a genus-wide, phylogenetically-structured pangenome study. This analysis inferred that the gene cluster responsible for carotovoricin biosynthesis is conserved across the genus and is located in the same gene neighborhood in all the species. Within the carotovoricin cluster, the tail fiber genes, which determine the host range specificity, exhibit high variability and discordance with the species phylogeny. We show evidence for an evolutionary mechanism involving recombination-mediated exchange of these tail fiber loci across the entire Pectobacterium genus, which complements the previously known mechanism for DNA sequence inversion to maintain tailocin polymorphism at the population level. In addition, the ability to exchange tail-fiber loci in a highly targeted and genus-wide manner could influence the community dynamics in nutrient rich environments such as infected plant tissues. In conclusion, the strong signal for carotovoricin retention and ability to exchange tail fibers indicates that it significantly contributes to the community interactions of the Pectobacterium phytopathogens.
With the current speed of sequencing, there is a desire for standardized and automated genome assembly and annotation to produce high-quality genomes as input for comparative (pan)genomics. Therefore, we created a convenience pipeline using existing tools that creates annotated genome assemblies from HiFi (and optionally ultra-long ONT and/or Hi-C) reads for a set of related individuals as well as a related reference genome. Our pipeline is species-agnostic and generates an extensive quality assessment report that can be used for manual filtering and refinement of the assembly and annotation. It includes statistics for individual completeness and contamination assessments as well as a concise pangenome view. The pipeline is implemented in Snakemake and available with a GPLv3 licence at GitHub under github.com/dirkjanvw/MoGAAAP, at Zenodo under doi.org/10.5281/zenodo.14833021, and can be installed through Bioconda.
This study presents three genome assemblies within the Capsicum genus, enabling comprehensive comparative analyses for the Annuum and Baccatum complexes within the genus. We produced highly continuous assemblies of the nuclear genomes and complete chloroplast assemblies. Subsequent genome annotation identified 34,580 genes in nonpungent C. annuum cv. ECW, and 32,704 and 33,994 genes in pungent C. chacoense and C. galapagoense, respectively. These assemblies, including the first complete genomes for C. chacoense and C. galapagoense, provide additional genomic resolution within the Capsicum genus. The novel genomes were analyzed within a pangenomic framework, integrating 16 Capsicum genomes across the Annuum, Baccatum, and Pubescens complexes. Homology grouping was used to identify core, accessory and unique genes and showed a wide spectrum of genetic diversity, particularly in homology groups exclusive to C. chacoense and C. galapagoense. Out of 79,267 homology groups identified, 13% were core groups, present in all accessions, corresponding to approximately 30% of core genes per genome. Comparative analyses revealed distinct species and genus-specific genomic characteristics. Additionally, we used the graph pangenome to illustrate locus-level exploration by examining the Pun1 locus associated with capsaicinoid biosynthesis, identifying multiple Pun1-like genes including their genomic position and homology information. The integration of these new resources into a dynamic Capsicum pangenome framework provides a versatile platform for extracting genetic information relevant to both fundamental research and breeding applications.
Comparing gene organization across genomic sequences reveals insights into evolutionary and functional diversity among different organisms and varieties. Performing this task across many sequences, such as from a pangenome, is challenging because of the scale, the density of information, and the inherent variation. Often, analyses are centered on a genomic region of interest-a locus that might be associated with a trait or contain genes within the same family or biological pathway. Within these regions, researchers examine the conservation of gene order and orientation across organisms and assess sequence similarity, along with other gene content features such as gene size, to find biological variations or potential errors in the data. Automated methods in comparative genomics struggle to identify meaningful patterns due to varying and often unknown features of interest, leaving manual, time-intensive, and scalability-challenged visualization as the primary alternative. To address these challenges, we present a multiscale design for studying gene organization within pangenomes, developed in close collaboration with domain experts. Our tool, Multipla, enables users to explore organization at multiple levels of detail in a decluttered manner through layout abstractions, semantic zooming, and layouts with flexible distance definitions and feature selections, combining the advantages of manual and automated methods used in practice. We evaluate the design of Multipla through two pangenomic use cases and conclude with lessons learned from designing multiscale views for pangenomic locus analysis.
Bacterial pathogens of the genus Pectobacterium are responsible for soft-rot and blackleg diseases in a wide range of crops and have a global impact on food production. The emergence of new lineages and their competitive succession is frequently observed in Pectobacterium species, in particular in Pectobacterium brasiliense . With a focus on one such recently emerged P. brasiliense lineage in the Netherlands that causes blackleg in potatoes, we studied genome evolution in this genus using a reference-free graph-based pangenome approach. We clustered 1,977,865 proteins from 454 Pectobacterium spp . genomes into 30,156 homology groups. The Pectobacterium genus pangenome is open, and its growth is mainly contributed by the accessory genome. Bacteriophage genes were enriched in the accessory genome and contributed 16% of the pangenome. Blackleg-causing P. brasiliense isolates had increased genome size with high levels of prophage integration. To study the diversity and dynamics of these prophages across the pangenome, we developed an approach to trace prophages across genomes using pangenome homology group signatures. We identified lineage-specific as well as generalist bacteriophages infecting Pectobacterium species. Our results capture the ongoing dynamics of mobile genetic elements, even in the clonal lineages. The observed lineage-specific prophage dynamics provide mechanistic insights into Pectobacterium pangenome growth and contribution to the radiating lineages of P. brasiliense .
Rosa , belonging to the family Rosaceae, encompasses more than 150 species which are widely distributed in the northern hemisphere. Renowned for their beauty, roses are cultivated throughout the world for ornamental purposes and the production of essential oils and perfumes. Despite their cultural and commercial significance, the genomic resources of wild Rosa species have not been studied comprehensively, hampering the understanding of their genetic diversity, evolutionary history, and breeding potential. Here we report on high-quality de novo genomes for Rosa sericea and Rosa rugosa . By integrating these two de novo genomes with existing public genomic resources, we have built a Rosaceae panproteome and a Rosa pangenome (spanning wild, traditional garden, and modern rose lineages) using a De Bruijn graph (DBG)-based approach. A maximum likelihood (ML) phylogeny of 18 Rosa haplotypes based on 4,367 single-copy core homology groups (genes) provided robust evolutionary inference, confirming the basal position of R. sericea , and enabled a gene-based macrosynteny analysis across the pangenome. Our analyses revealed significant genomic diversity among species, extensive variation in core gene content, and lineage-specific transposable element (TE) expansion patterns that contribute to the variation in Rosa genome size and to species-specific adaptations. The pangenome also revealed biased diversification of homology groups potentially linked to phenotypic plasticity in Rosa . Specifically, our analysis of the rose scent-related gene family, NUDX1 , uncovered its evolutionary trajectory in Rosa , in which TEs insertions provided putative novel regulatory elements that facilitated adaptive evolution in metabolic pathways. This pangenomic study deepens our understanding of the genetic diversity and evolution of traits within the Rosa genus. In addition, the findings lay the foundation for future efforts to understand the genetic mechanisms driving trait evolution, which can support rose breeding. ### Competing Interest Statement The authors have declared no competing interest.
The field of comparative genomics is shifting toward pangenomics, aiming to alleviate the reference bias observed in reference-based approaches. High-quality genomes are needed as input for pangenomics, as explained by the principle of "garbage in, garbage out". Errors in an assembly or annotation will lead to technical variation in a pangenome, while it is meant to reveal genuine biological variation only. Achieving the required assembly and annotation quality remains challenging, particularly in plants, given the complexity in genome size, ploidy level, and repeat content present in the plant kingdom.Nevertheless, the (technical) variation uncovered by pangenomics can, in turn, guide iterative refinements that yield high-quality genomes. The comparative approach is especially powerful to identify "abnormalities" in a set of genomes, which cannot be detected from traditional, stand-alone quality assessment. Two use cases, on roses and chili peppers, illustrate the role of pangenomics in this iterative process leading toward meaningful pangenome analyses. Thus, we argue that high-quality genomes can be achieved through pangenomics, facilitating the convergence on real genetic variation underlying complex traits and diseases, not only in plants but also in other organisms.
With advances in long-read sequencing and assembly techniques, haplotype-resolved (phased) genome assemblies are becoming more common, also in the field of plant genomics. Computational tools to effectively explore these phased genomes, particularly for polyploid genomes, are currently limited. Here we describe a new strategy adopting a pangenome approach. To analyse both intra- and intergenomic variation in phased genome assemblies, we have made the software package PanTools ploidy-aware by updating the pangenome graph representation and adding several novel functionalities to assess synteny and gene retention, profile repeats and calculate synonymous and nonsynonymous mutation rates. Using PanTools, we constructed and analysed a pangenome comprising of one diploid and four tetraploid potato cultivars, and a pangenome of five diploid apple species. Both pangenomes show high intra- and intergenomic allelic diversity in terms of gene absence/presence, SNPs, indels and larger structural variants. Our findings show that the new functionalities and visualizations are useful to discover introgressions and detect likely misassemblies in phased genomes. PanTools is available at https://git.wur.nl/bioinformatics/pantools.
Summary The growing number of sequences and increasing proof that single references create reference bias have driven the development of pangenomes to represent the genomic diversity of species. To leverage this complex diversity information for biological insights, analysis and visualization support are needed to explore the variants in the context of metadata and phylogenies. We developed PanVA, an interactive visual analytics tool for exploring sequence variants in groups of homologous sequences in their biological context. PanVA is a web application that allows users to explore existing instances or create new ones to visualize their own data. Availability and Implementation The PanVA source code is available on GitHub at under the GPLv3 License. Documentation and and public demo instances showcasing examples can be accessed at . ### Competing Interest Statement The authors have declared no competing interest. Netherlands eScience Center, https://ror.org/00rbjv475, ETEC.2019.019 TKI Agri & Food, TU18034
Genomics researchers increasingly use multiple reference genomes to comprehensively explore genetic variants underlying differences in detectable characteristics between organisms. Pangenomes allow for an efficient data representation of multiple related genomes and their associated metadata. However, current visual analysis approaches for exploring these complex genotype-phenotype relationships are often based on single reference approaches or lack adequate support for interpreting the variants in the genomic context with heterogeneous (meta)data. This design study introduces PanVA, a visual analytics design for pangenomic variant analysis developed with the active participation of genomics researchers. The design uniquely combines tailored visual representations with interactions such as sorting, grouping, and aggregation, allowing users to navigate and explore different perspectives on complex genotype-phenotype relations. Through evaluation in the context of plants and pathogen research, we show that PanVA helps researchers explore variants in genes and generate hypotheses about their role in phenotypic variation.
Photosynthesis is the only yield-related trait not yet substantially improved by plant breeding. Previously, we have established H. incana as the model plant for high photosynthetic light-use efficiency (LUE). Now we aim to unravel the genetic basis of this trait in H. incana, potentially contributing to the improvement of photosynthetic LUE in other species. Here, we compare its transcriptomic response to high light with that of Arabidopsis thaliana, Brassica rapa, and Brassica nigra, 3 fellow Brassicaceae members with lower photosynthetic LUE. We built a high-light, high-uniformity growing environment, in which the plants developed normally without signs of stress. We compared gene expression in contrasting light conditions across species, utilizing a panproteome to identify orthologous proteins. In-depth analysis of 3 key photosynthetic pathways showed a general trend of lower gene expression under high-light conditions for all 4 species. However, several photosynthesis-related genes in H. incana break this trend. We observed cases of constitutive higher expression (like antenna protein LHCB8), treatment-dependent differential expression (as for PSBE), and cumulative higher expression through simultaneous expression of multiple gene copies (like LHCA6). Thus, H. incana shows differential regulation of essential photosynthesis genes, with the light-harvesting complex as the first point of deviation. The effect of these expression differences on protein abundance and turnover, and ultimately the high photosynthetic LUE phenotype is relevant for further investigation. Furthermore, this transcriptomic resource of plants fully grown under, rather than briefly exposed to, a very high irradiance, will support the development of highly efficient photosynthesis in crops.
BackgroundBreeding of lettuce (Lactuca sativa L.), the most important leafy vegetable worldwide, for enhanced disease resistance and resilience relies on multiple wild relatives to provide the necessary genetic diversity. In this study, we constructed a super-pangenome based on four Lactuca species (representing the primary, secondary and tertiary gene pools) and comprising 474 accessions. We include 68 newly sequenced accessions to improve cultivar coverage and add important foundational breeding lines.ResultsWith the super-pangenome we find substantial presence/absence variation (PAV) and copy-number variation (CNV). Functional enrichment analyses of core and variable genes show that transcriptional regulators are conserved whereas disease resistance genes are variable. PAV-genome-wide association studies (GWAS) and CNV-GWAS are largely congruent with single-nucleotide polymorphism (SNP)-GWAS. Importantly, they also identify several major novel quantitative trait loci (QTL) for resistance against Bremia lactucae in variable regions not present in the reference lettuce genome. The usability of the super-pangenome is demonstrated by identifying the likely origin of non-reference resistance loci from the wild relatives Lactuca serriola, Lactuca saligna and Lactuca virosa.ConclusionsThe super-pangenome offers a broader view on the gene repertoire of lettuce, revealing relevant loci that are not in the reference genome(s). The provided methodology and data provide a strong basis for research into PAVs, CNVs and other variation underlying important biological traits of lettuce and other crops.
Fusarium head blight (FHB) is one of the most destructive wheat diseases worldwide. To understand the impact of human migration and changes in agricultural practices on crop pathogens, here population genomic analysis with 245 representative strains from a collection of 4,427 field isolates of Fusarium asiaticum, the causal agent of FHB in Southern China is conducted. Three populations with distinct evolution trajectories are identifies over the last 10,000 years that can be correlated with historically documented changes in agricultural practices due to human migration caused by the Southern Expeditions during the Jin Dynasty. The gradual decrease of 3ADON-producing isolates from north to south along with the population structure and spore dispersal patterns shows the long-distance (>250 km) dispersal of F. asiaticum. These insights into population dynamics and evolutionary history of FHB pathogens are corroborated by a genome-wide analysis with strains originating from Japan, South America, and the USA, confirming the adaptation of FHB pathogens to cropping systems and human migration.
Natural populations of Arabidopsis thaliana provide powerful systems to study adaptation of wild plant species. Previous research has predominantly focused on global populations or accessions collected from regions with diverse climates. However, little is known about the genetics underlying adaptation in regions with mild environmental clines. We have examined a diversity panel consisting of 192 A. thaliana accessions collected from the Netherlands, a region with limited climatic variation. Despite the relatively uniform climate, we identified compelling evidence of local adaptation within this population. Notably, semidwarf accessions, due to mutation of the GIBBERELLIC ACID REQUIRING 5 ( GA5 ) gene, occur at a relatively high frequency near the coast and these displayed enhanced tolerance to high wind velocities. Additionally, we evaluated the performance of the population under iron deficiency conditions and found that allelic variation in the FE SUPEROXIDE DISMUTASE 3 ( FSD3 ) gene affects tolerance to low iron levels. Moreover, we explored patterns of local adaptation to environmental clines in temperature and precipitation, observing that allelic variation at LA RELATED PROTEIN 1C ( LARP1c ) likely affects drought tolerance. Not only is the genetic variation observed in a diversity panel of A. thaliana collected in a region with mild environmental clines comparable to that in collections sampled over larger geographic ranges, it is also sufficiently rich to elucidate the genetic and environmental factors underlying natural plant adaptation.
Lettuce (Lactuca sativa L.) is a leafy vegetable crop with ongoing breeding efforts related to quality, resilience, and innovative production systems. To breed resilient and resistant lettuce in the future, valuable genetic variation found in close relatives could be further exploited. Lactuca virosa (2x = 2n = 18), a wild relative assigned to the tertiary lettuce gene pool, has a much larger genome (3.7 Gbp) than Lactuca sativa (2.5 Gbp). It has been used in interspecific crosses and is a donor to modern crisphead lettuce cultivars. Here, we present a de novo reference assembly of L. virosa with high continuity and complete gene space. This assembly facilitated comparisons to the genome of L. sativa and to that of the wild species L. saligna, a representative of the secondary lettuce gene pool. To assess the diversity in gene content, we classified the genes of the 3 Lactuca species as core, accessory, and unique. In addition, we identified 3 interspecific chromosomal inversions compared to L. sativa, which each may cause recombination suppression and thus hamper future introgression breeding. Using 3-way comparisons in both reference-based and reference-free manners, we show that the proliferation of long-terminal repeat elements has driven the genome expansion of L. virosa. Further, we performed a genome-wide comparison of immune genes, nucleotide-binding leucine-rich repeat, and receptor-like kinases among Lactuca spp. and indicated the evolutionary patterns and mechanisms behind their expansions. These genome analyses greatly facilitate the understanding of genetic variation in L. virosa, which is beneficial for the breeding of improved lettuce varieties.
Photosynthesis is the only yield-related trait that has not yet been substantially improved by plant breeding. The limited results of previous attempts to increase yield via improvement of photosynthetic pathways suggest that more knowledge is still needed to achieve this goal. To learn more about the genetic and physiological basis of high photosynthetic light-use efficiency (LUE) at high irradiance, we study Hirschfeldia incana . Here, we compare the transcriptomic response to high light of H. incana with that of three other members of the Brassicaceae, Arabidopsis thaliana, Brassica rapa , and Brassica nigra , which have a lower photosynthetic LUE. First, we built a high-light, high-uniformity growing environment in a climate-controlled room. Plants grown in this system developed normally and showed no signs of stress during the whole growth period. Then we compared gene expression in low and high-light conditions across the four species, utilizing a panproteome to group homologous proteins efficiently. As expected, all species actively regulate genes related to the photosynthetic process. An in-depth analysis on the expression of genes involved in three key photosynthetic pathways revealed a general trend of lower gene expression in high-light conditions. However, H. incana distinguishes itself from the other species through higher expression of certain genes in these pathways, either through constitutive higher expression, as for LHCB8 , ordinary differential expression, as for PSBE , or cumulative higher expression obtained by simultaneous expression of multiple gene copies, as seen for LHCA6 . These differentially expressed genes in photosynthetic path-ways are interesting leads to further investigate the exact relationship between gene expression, protein abundance and turnover, and ultimately the LUE phenotype. In addition, we can also exclude thousands of genes from “explaining” the phenotype, because they do not show differential expression between both light conditions. Finally, we deliver a transcriptomic resource of plant species fully grown under, rather than briefly exposed to, a very high irradiance, supporting efforts to develop highly efficient photosynthesis in crop plants.
P. brasiliense is an important bacterial pathogen causing blackleg (BL) in potatoes. Nevertheless, P. brasiliense is often detected in seed lots that do not develop any of the typical blackleg symptoms in the potato crop when planted. Field bioassays identified that P. brasiliense strains can be categorized into two distinct classes, some able to cause blackleg symptoms and some unable to do it. A comparative pangenomic approach was performed on 116 P. brasiliense strains, of which 15 were characterized as BL-causing strains and 25 as non-causative. In a genetically homogeneous clade comprising all BL-causing P. brasiliense strains, two genes only present in the BL-causing strains were identified, one encoding a predicted lysozyme inhibitor Lprl (LZI) and one encoding a putative Toll/interleukin-1 receptor (TIR) domain-containing protein. TaqMan assays for the specific detection of BL-causing P. brasiliense were developed and integrated with the previously developed generic P. brasiliense assay into a triplex TaqMan assay. This simultaneous detection makes the scoring more efficient as only a single tube is needed, and it is more robust as BL-causing strains of P. brasiliense should be positive for all three assays. Individual P. brasiliense strains were found to be either positive for all three assays or only for the P. brasiliense assay. In potato samples, the mixed presence of BL-causing and not BL-causing P. brasiliense strains was observed as shown by the difference in Ct value of the TaqMan assays. However, upon extension of the number of strains, it became clear that in recent years additional BL-causing lineages of P. brasiliense were detected for which additional assays must be developed.