Land plants underpin civilization and planetary health, yet their genomic diversity remains largely uncharted. Current resources are unstandardized and scarce, lacking reference genomes for 95% of genera, 70% of families, and 51% of orders, impeding evolutionary and functional insight. We thus propose the PLANeT initiative, an international effort to generate high-quality, standardized genomes across the plant tree of life. Integrating artificial intelligence (AI) with genomics, we will decode conserved principles to advance fundamental plant biology, biodiversity conservation, crop improvement, and natural product discovery. Engaging around 100 labs to train 1,000 scientists, we will tackle pivotal questions for a sustainable future.
In October 2024, Parties to the United Nations Convention on Biological Diversity agreed to a new multilateral mechanism to fund biodiversity conservation through the sharing of benefits from open biodiversity data. Biological databases hosting genetic and other biological data, known as digital sequence information (DSI), are central to the implementation of the mechanism. This paper assesses the new international agreement and its implications for DSI databases. We walk through the database provisions in COP16 Decision 16/2, which include notifying users and submitters about the mechanism, improving metadata on geographical location of sample collection, and consistency with open access, as well as consideration of the FAIR, CARE, and TRUST principles. Drawing on surveys, interviews, and a workshop with biological database managers, we identify practical and scalable measures including updating terms of use, revising submission procedures, and strengthening user communication. We also propose approaches to capture and report non-monetary benefits such as capacity building, publications, interoperability, and training. These actions illustrate how DSI databases can remain open, sustainable, and globally connected while supporting benefit-sharing from the use of DSI on genetic resources.
Grass inflorescences are composite structures, featuring complex sets of meristems as stem cell niches that are initiated in a repetitive manner. Meristems differ in identity and longevity, generate branches or split to form flower meristems that finally produce seeds. Within meristems, distinct cell types are determined by positional information and the regional activity of gene regulatory networks. Understanding these local microenvironments requires precise spatio-temporal information on gene expression profiles, which current technology cannot achieve.Here we investigate transcriptional changes during barley development, from the specification of meristem and organ founder cells to the initiation of distinct floral organs, on the basis of an imputation approach integrating deep single-cell RNA sequencing with spatial gene expression data. The expression profiles of more than 40,000 genes can now be analysed at cellular resolution in multiple barley tissues using the new web-based graphical interface BARVISTA, which enables precise virtual microdissection to analyse any sub-ensemble of cells. Our study pinpoints previously inaccessible key transcriptional events in founder cells during primordia initiation and specification, characterizes complex branching mutant phenotypes by barcoding gene expression profiles, and defines spatio-temporal trajectories during flower development. We thus uncover the genetic basis of complex developmental processes, providing novel opportunities for precisely targeted manipulation of barley traits.
MOTIVATION:Trimmomatic is a widely adopted tool for preprocessing high-throughput sequencing data, particularly from Illumina platforms. Since its original publication in 2014, the volume and complexity of sequencing data have increased dramatically, necessitating continuous tool evolution. RESULTS:We present the substantial updates to Trimmomatic over the past decade. Key enhancements include a robust multithreading model for high-performance parallel processing, parallel GZIP/BZIP2 compression, and a suite of new trimming and filtering steps to provide users with more flexible quality control. Usability has been significantly improved through automatic PHRED encoding detection and simplified file handling. The codebase has also been modernized including Maven support, and continuous integration to ensure long-term sustainability and community contributions. These updates solidify Trimmomatic's role as an efficient, flexible, and essential tool in modern bioinformatics pipelines. AVAILABILITY:Trimmomatic remains open-source under the GPL V3 license, with the latest version available at https://github.com/usadellab/Trimmomatic and also on our website https://www.plabipd.de/trimmomatic_main.html (DOI: https://doi.org/10.5281/zenodo.18678155).
Key Performance Indicators (KPIs) are essential for evaluating project success and establishing control mechanisms to monitor development, performance, and user acceptance of services in joint projects. However, the absence of standardized frameworks and effective monitoring tools, combined with service providers' reluctance due to fears of comparability, has limited their adoption in scientific contexts. To address this gap, we developed Scorpion, a flexible tool for KPI monitoring in project management. Scorpion enables service providers to retain control over their metrics while supporting centralized reporting. It offers both web-based and programmatic access, with features for KPI submission, visualization, and user and service management. Initially created for bioinformatics and biodiversity projects, Scorpion is applicable across diverse domains. It is particularly valuable for initiatives like the German National Research Data Infrastructure (NFDI), where funding agencies require KPI reporting for evaluation. We present the Scorpion framework, highlighting its design principles, features, and potential to improve project management practices. Use cases illustrate how Scorpion enhances KPI monitoring efficiency and accuracy, contributing to better impact evaluation, quality assurance, and informed decision-making in project and service management.
Abstract Tomato wild relatives are valuable genetic resources for trait discovery and understanding the genetic basis of fruit metabolism and quality. Yet, only a fraction of naturally occurring variation has been exploited. Here, we performed metabolite profiling of two large Backcross Inbred Line populations derived from crosses between the wild species S. pennellii accession LA5240 (Lost) and cultivated genotypes LEA (determinate) and TOP (indeterminate), including ∼1400 and ∼500 lines, respectively. High-resolution mapping identified enormous metabolic quantitative trait loci (mQTL), including a new locus on chromosome 12 associated with fruit sucrose accumulation that harbours INVERTASE INHIBITOR 3 ( SlINVINH3 ) protein. Comparative analysis indicated that SlINVINH3 is highly expressed in wild S. pennellii 0716 fruit, whereas a six-amino acid deletion is present in its coding sequence compared with S.pennellii LA5240 and S. lycopersicum . We further demonstrated that in SlINVINH3 -overexpressing tomato plants, only the S. pennellii LA5240 allele led to increased sucrose, accompanied by reduced fructose and glucose levels. Furthermore, the large population size enabled us to assess the epistatic interactions, with approximately 40% of interactions being more-than-additive and 60% less-than-additive. Our results demonstrate the power of permanent exotic populations to reveal hidden metabolic diversity and provide an approach for improving fruit quality through targeted breeding and metabolic engineering.
Abstract The plastic pollution crisis urges innovative recycling solutions. Promising approaches especially for polyester-containing wastes include enzymatic hydrolysis and microbial upcycling. For efficient enzymatic hydrolysis of polyesters, elevated temperatures (70–80 °C) are required, necessitating thermophilic microbial chassis for consolidated bioprocessing (CBP). In this study, we engineered Geobacillus thermoleovorans through adaptive laboratory evolution (ALE) for robust growth on adipic acid (AA) and 1,4-butanediol (BDO), two relevant monomers for example derived from poly(butylene adipate-co-terephthalate) (PBAT), enabling growth rates of up to 0.10 h−1 on AA and 0.13 h−1 on BDO. Based on a high-quality annotated genome sequence of the wild type, genomic mutations and gene expression levels were characterized in mutants grown on the respective substrates compared to glucose. For BDO, an alcohol dehydrogenase (Gth_001044) and an aldehyde dehydrogenase (Gth_001082) were identified to be likely responsible for its oxidative degradation. AA uptake appears to be mediated by a dicarboxylate transporter (Gth_003270), followed by CoA activation and β-oxidation involving a CoA transferase (Gth_003192) and several upregulated CoA-family dehydrogenases. To demonstrate applicability of these strains in plastic upcycling, they were co-cultivated with PBAT as the sole carbon source in combination with the cutinase HiC for PBAT hydrolysis. This resulted in growth on the released AA and BDO. Given the potential to purify the remaining terephthalate (TA), this approach highlights the feasibility of selective monomer valorization in bioprocesses. Additional ALE enabled co-utilization of AA and BDO by a single strain and improved AA consumption at lower concentrations, underscoring the strains’ adaptability and high potential for plastic upcycling applications. Key points • G. thermoleovorans evolved for robust growth on adipate and 1,4-butanediol at 60 °C. • Genome and transcriptome analyses revealed underlying pathways and enzymes involved. • Co-cultivation of the evolved strains on PBAT with HiC as the sole carbon source.
Abstract The NFDI-consortium FAIRagro has established a systematic framework for evaluating the FAIRness of Research Data Infrastructures (RDI) within the German agrosystem research landscape. While FAIR principles are widely accepted, their practical implementation by RDIs remains challenging. By operationalizing the FAIR principles into a reproducible multi-dimensional scoring methodology, this initiative addresses the critical need for a transparent and citable benchmark of RDIs that moves beyond simple compliance. This paper details the underlying assessment criteria, comprising 20 aggregated core metrics, the iterative community-driven validation process, and the integration of these metrics into the FAIRagro Search Hub. This framework evaluates RDIs, like repositories or databases, instead of sampling hosted data sets, across the four distinct categories of FAIR independently, yielding granular, pillar-specific ratings. Unlike aggregate scoring models, which can inadvertently mask technical deficiencies by averaging performance across categories, this multi-dimensional approach ensures that a repository’s distinct strengths and bottlenecks remain fully visible. Our findings demonstrate that standardized scoring not only clarifies data accessibility for users but also highlights specific operational gaps, allowing repository providers to identify precisely where the service implementation can be enhanced. By establishing this data-driven service in the agronomy domain, we provide a scalable template for the broader NFDI and EOSC ecosystems to foster a culture of excellence in research data stewardship.
Plant phosphoenolpyruvate carboxylases (PEPCs) are ubiquitously expressed as cytosolic Class-1 PEPC homotetramers composed of 107 kDa plant-type PEPC (PTPC) subunits that are highly sensitive to allosteric inhibition by malate. Class-2 PEPC heterooctameric complexes that are desensitized to malate inhibition also exist in certain sink tissues due to the interaction of a Class-1 PEPC with unrelated 118 kDa bacterial-type PEPC (BTPC) polypeptides. Class-2 PEPCs dynamically associate with the mitochondrial outer envelope and have been hypothesized to support sustained anaplerotic flux and respiratory CO₂ refixation in malate-rich sink tissues, including immature tomato fruit. The current study generated CRISPR-Cas9-edited tomato lines with targeted disruption of the BTPC gene and investigated the impact on fruit development, metabolism, and transcriptional regulation. Immunoblotting and co-immunoprecipitation confirmed the absence of BTPC polypeptides and Class-2 PEPC complexes in the edited lines. Fruits from the edited plants were 25% smaller and 40% lighter and required up to 10 additional days to complete ripening compared to the WT. Metabolomic analysis across ripening stages revealed substantial reductions in malate and citrate, with elevated sugars and amino acids, indicating reprogrammed carbon flux. RNA-seq data showed downregulation of genes for cell wall remodeling, sugar transport, and ethylene-responsive transcription factors. These results provide direct evidence that BTPC is essential for organic acid balance, sugar metabolism, and ripening regulation in tomato. Its absence perturbs metabolic homeostasis and developmental progression, positioning BTPC as a strategic target for enhancing fruit quality traits through genetic engineering.
Tea plants possess a highly heterozygous genome and produce diverse beneficial metabolites, yet the genomic basis of its metabolic diversity remains fully elusive. Here, we show that single-cell sequencing of 107 sperm cells, combined with PacBio HiFi and ONT ultra-long sequencing, enables an accurate haplotype-resolved genome assembly of Fuding Dabaicha (FDDB). Structural variations (SVs) between the two haplotypes comprise 23.8% of the genome and strongly influence crossover patterns. Using these phased genomes, we establish a half-sib-based QTL mapping platform and identify a Gypsy LTR insertion in the promoter of CsDFRb, which associated with the increasing of p-coumaroylquinate levels in young leaves. Moreover, mGWAS reveals 2649 additional loci associated with 2837 metabolites when using the FDDB genome rather than the 'Tieguanyin' reference genome for variant calling. Functional validation of CsC3H and CsST2Ac confirms their role in determining chlorogenic acid and sulfated metabolite levels in tea plant. In addition, we observe that allelic heterogeneity at CsST2Ac affects sulfated metabolite abundance. These phased genomes illuminate how SVs drive metabolic diversity and offer valuable resources for tea breeding.
Connecting the characterization of juvenile (pre-anthesis) plant stress responses in controlled environments to field agronomic performance is a challenge. The oilseed crop Camelina sativa (camelina), with its innate resilience and plasticity, presents an opportunity to understand the underlying mechanisms of juvenile resilience and identify the implications for yield in diverse pedoclimates. A better understanding of camelina's abiotic stress resilience is important in the context of climate change and the development of breeding programs for climate-tolerant crops. In this study, 54 accessions representing the genetic diversity observed in the wider publicly available population were used to investigate the plasticity of camelina's early stage response to drought and heat stress, combined with an evaluation of field performance in multilocation field trials. A combinatorial phenotyping approach of early stage drought and heat stress identified stress-responsive signatures within the diversity panel. The substantial variation in the morphophysiological line-specific responses to stress indicated that juvenile and mature camelina plants have significant plasticity and access different stress response strategies. In response to stress, we observed significant molecular metabolic adjustment alongside significant lipid remodeling and physiological compensation. Camelina was resilient to drought stress, and certain metabolites were identified as indicators of abiotic stress response. Applying an integrated approach, early stage phenotyping and multilocation field trials provided a complete assessment of the camelina stress response and facilitated a connection to crop productivity. This approach facilitates improved breeding programs, addresses the restrictions of limited genetic diversity in camelina, and supports the development of local varieties optimized for climate resilience.
Crop wild relatives are used to improve cultivated plants and precise tracking of genetic introgression requires high-quality genome assemblies. Here we present de novo genome assemblies of two wild tomato species - the broadly stress-resistant Solanum pennellii (LA0716) and the salt-resistant Solanum cheesmaniae (LA1039). The improved S. pennellii genome adds 146 Mbp to the twelve chromosomes compared with the original reference. The alignment of the new assemblies with multiple gold-standard assemblies identified shared and species-specific structural variants. Analysis of repeat content demonstrates independent explosions of Tekay retrotransposons in S. pennellii and S. peruvianum. Genome sequencing of 709 recombinant plants derived from male and female backcrosses of three different hybrids reveals higher crossover rate in female meiosis. Conserved female-enhanced recombination regions were discovered and coldspots were attributed to megabase-scale inversions and insertion-deletion polymorphisms. Our S. pennellii and S. cheesmaniae genome assemblies reveal how repeat content diverged in nature and during breeding, and uncovers how both reproductive gender and structural variants dictate recombination landscapes in tomato hybrids.
Flavour inconsistency in Fragaria x ananassa remains a challenge, largely due to historical breeding focused on yield over sensory traits. Flavour results from organoleptic and bioactive compounds shaped by cultivar (G), environment (E) and their interaction (GxE). Four strawberry cultivars (Clery, Frida, Gariguette, Sonata) were evaluated across five European environments employing an integrative approach combining metabolomics and transcriptomics. This approach revealed elevated temperatures accelerate the start of the harvest season and shorten fruit development duration. GxE interactions critically influenced flavour compounds accumulation, indicating cooler temperatures during fruit development favor the accumulation of sugars and ɣ-decalactone. Integrative analysis identified cultivar-specific and environmentally stable transcriptional patterns, including promising candidate genes involved in the biosynthesis and degradation of flavour relevant compounds, and revealing starting points for strawberry flavour improvement and stabilization. Our findings highlight the central role of GxE interactions in shaping strawberry flavour profile and emphasize the need for multi-environment trials to support resilient flavour enhancement in future breeding programs.
Advances in next-generation sequencing technologies over the last decade have substantially reduced the cost and effort required to sequence plant genomes. Whereas early efforts focused primarily on economically important crops and model species, attention has now turned to a broader range of plants, including those with larger and more complex genomes. In 2024, the genomes of 500 plant species were published, including 370 sequenced for the first time. Tracking and providing access to published plant genomes (now covering more than 1800 species) is an invaluable service for plant researchers. PubPlant is an online resource that serves this purpose by cataloging published plant genome sequences and offering multiple visualizations (https://www.plabipd.de/pubplant_main.html). It includes a chronology of genome publications, and cladograms to display the phylogenetic relationships among the sequenced plants. An overview diagram for seed plants highlights taxonomic orders and families with sequenced species and reveals those that have been overlooked thus far. As a use case for PubPlant, we evaluated the status of sequenced food crops. We found that the five plant families featuring the most food crops were those containing the most sequenced plant species.
Eggplant (Solanum melongena L.) is a major Solanaceous crop of Asian origin, but genomic resources remain limited compared to related species. Here, a core collection of 368 accessions spanning global diversity of S. melongena and wild relatives is phenotyped for agronomic, disease resistance and fruit metabolomic traits and resequenced. Additionally, 40 chromosome-level assemblies of S. melongena, its progenitor S. insanum and the allied species S. incanum enable the construction of two graph-based pangenomes, capturing broad genetic variation. We demonstrate the power of these datasets by identifying major loci controlling prickliness and resistance to Fusarium oxysporum f. sp. melongenae, driven by SVs affecting the LONELY GUY 3 gene and a resistance gene cluster, respectively, as well as a mutation in a GDSL-like esterase/lipase gene altering the levels of dicaffeoyl-quinic acids. These findings provide a cornerstone for pangenome-assisted breeding, enabling detailed analyses of genetic diversity, domestication history, and trait evolution in eggplant.
Cultivar Désirée is an important model for potato functional genomics studies to assist breeding strategies. Here, we present a haplotype-resolved genome assembly of Désirée, achieved by assembling PacBio HiFi reads and Hi-C scaffolding, resulting in a high-contiguity chromosome-level assembly. We implemented a comprehensive annotation pipeline incorporating gene models and functional annotations from the Solanum tuberosum Phureja DM reference genome alongside RNA-seq reads to provide high-quality gene and transcript annotations. Additionally, we provide a genome-wide DNA methylation profile using Oxford Nanopore reads, enabling insights into potato epigenetics. The assembled genome, annotations, methylation and expression data are visualised in a publicly accessible genome browser, providing a valuable resource for the potato research community.
Recent advances in high-quality genome sequencing have revolutionized research in the tomato clade (Solanum section Lycopersicon), enabling the generation of long-read and even chromosome-scale assemblies for cultivated tomato and its wild relatives. These data have shed light on tomato domestication and population genetics and have facilitated breeding using exotic germplasm. This review summarizes progress in tomato genomics, focusing on the diversity of section Lycopersicon and its function as a reservoir of stress-tolerance genes, including drought tolerance from Solanum pennellii and pathogen resistance from Solanum habrochaites and Solanum chilense. We catalog important genetic resources, including introgression lines and multi-parent advanced generation inter-cross (MAGIC) populations, which have allowed the dissection of important traits via the mapping of quantitative trait loci, including those involved in primary and secondary metabolism. We also explore the metabolic diversity of wild and domesticated tomato species and discuss how this has led to gene identification. Finally, we show that tomato genomics will continue to accelerate, given the increasing availability and accessibility of genomics technology, exotic germplasm, and mapping populations, which can be leveraged using advanced genome-editing approaches.
Floral initiation is required for sexual reproduction in angiosperm plants, and has a significant impact on crop yields. In cultivated strawberry, the molecular basis of floral initiation is poorly understood and most studies have focused on a single genotype grown under controlled conditions. To gain more insight into this process, we conducted a study under natural conditions in two countries using two seasonal flowering cultivars. We focused on the early steps of floral initiation by using samples spanning key developmental stages of the shoot apical meristem. The analysis of differential gene expression in leaf and terminal bud tissues revealed enrichment for genes involved in carbohydrate metabolism and phytohormone pathways in leaves. Other protein classes that were enriched during early floral initiation were associated with cytoskeleton organization, cell cycle regulation, and chromatin structure. We also identified genes associated with the photoperiodic pathway, well-characterized floral integrators such as TFL1 and SOC1, and several linked to phytohormone regulation, such as XTH23, PP2 and EIN3. ### Competing Interest Statement The authors have declared no competing interest.
The cultivated strawberry (Fragaria × ananassa) is highly valued for its attractive color and pleasant taste, which is due to the secondary metabolites it accumulates. To clarify which factors influence the content of secondary metabolites in the fruit, F1 hybrids of the commercial strawberry cultivars "Senga Sengana" and "Candonga," which are adapted to different climatic zones, were produced. F1 progeny were grown in five European countries, and the content of secondary metabolites was analyzed. The results show that the genetic factor plays a more important role compared with environmental conditions. Cultivar "Senga Sengana" produced significantly more pelargonidin-3-O-(6'-O-malonyl)glucoside than "Candonga," regardless of the environmental conditions, and the progeny segregated with a ratio of 1:1 in this trait. Twenty-four putative malonyltransferase genes were cloned from F. × ananassa (FaMAT) "Senga Sengana" and "Candonga," and six FaMATs showed enzymatic activity. Three FaMAT enzymes are promiscuous and malonylate flavonoid glucosides and anthocyanins, whereas the others are specific for quercetin-3-O-glucoside and kaempferol-3-O-glucoside. Expression data indicate that the gene products of FaMAT1 and 4 catalyze the malonylation of anthocyanins in "Senga Sengana" and the related hybrids, while the corresponding functional alleles in "Candonga" and the second group of hybrids are only slightly expressed in the mature fruit. Thus, altered FaMAT transcription genetically determines the different levels of malonylated anthocyanins in the progeny.