Modern scientific infrastructure increasingly depends on standardized data governance frameworks, with FAIR (Findable, Accessible, Interoperable, Reusable) principles now widely adopted across research domains. This study investigates how FAIR principles are implemented in plant genetic resources (PGR) information systems by assessing the European Search Catalogue for Plant Genetic Resources (EURISCO) against the Research Data Alliance (RDA) FAIR Data Maturity Model indicators. As a large, federated catalogue containing data on 2.1 million accessions from 43 countries, EURISCO provides a strong empirical basis for evaluating FAIR implementation in distributed biological data systems. The assessment reveals uneven compliance across FAIR dimensions. Human-oriented findability and manual accessibility are comparatively strong, supported by a consolidated catalogue interface, open web access, and standardized passport data using Multi-Crop Passport Descriptors (MCPD v2.1). Persistent identification remains partial because the Digital Object Identifiers (DOI) currently cover only a subset of accessions and primarily identify physical germplasm instead of versioned EURISCO data record. Machine-actionability and semantic interoperability exhibit the most pronounced gaps, reflecting limited use of machine-readable formats, formal vocabularies and explicit links to external knowledge resources. The assessment also highlights domain-specific constraints, including tight coupling of physical collections and digital records, nationally distributed governance, heterogeneous institutional capacities, and long-standing legacy data integration requirements. Evaluation across ex situ passport, in situ crop wild relative, and phenotypic data domains indicates implementation gradients aligned with standardization maturity and documentation complexity. Standardized FAIR metrics only partially capture these constraints inherent to PGR information systems. Effective data governance must therefore balance FAIR ambitions with practical limits, rather than assume that generic frameworks can be applied uniformly across all contexts. Domain-calibrated approaches that prioritize scientific utility and fitness-for-purpose over exhaustive metric-based compliance may ultimately better support complex biological data systems underpinning global agricultural research.
The German Federal Ex Situ Genebank for Agricultural and Horticultural Crops (IPK) harbours over 3000 pea plant genetic resources (PGRs), backed up by corresponding information across 16 key agronomic and economical traits. The unbalanced structure and inconsistent format of this historical data has precluded effective leverage of genebank accessions, despite the opportunities contained in its genetic diversity. Therefore, a three-step statistical approach founded in linear mixed models was implemented to enable a rigorous and targeted data curation. Spring accessions revealed considerable breeding potential, with protein content exceeding market standards by almost one-fifth and with hundred grain weight that could match the upper limits reported for European elite varieties. This variation is embedded within structured populations, comprising five convarieties including sugar snaps and field pea, adding value for breeding across diverse morphotypes and market segments. Winter accessions demonstrated cold resilience, with post-winter survival rate up to 79.27
BACKGROUND:Lupinus albus is a food grain legume recognized for its high levels of seed protein (30-40%) and oil (6-13%), and its adaptability to different climatic and soil conditions. The availability of well-characterized, genetically and phenotypically diverse germplasm will facilitate the development of next-generation L. albus cultivars, encourage biodiversity conservation, and promote the sustainable utilization of this species. RESULTS:We evaluated more than 2000 L. albus accessions based on 35 agro-morphological traits and passport data to establish Intelligent Collections. The Reference-CORE (R-CORE), covering global diversity, exemplified the genotypic variation among accessions differing in biological status (cultivars, breeding/research materials, landraces, and wild relatives). The Training-CORE (T-CORE), a subset of 300 R-CORE accessions, represents the diversity of the entire collection. Principal component analysis showed that the L. albus R-CORE encompasses four phenotypic groups (A1, A2, A3, and B), and that groups A3 and B can be characterized by the main phenotypic traits of pod shattering and seed ornamentation, respectively. The coefficient of total genetic variation differed across morphological traits, phenotypic groups, geographic regions, and according to biological status. CONCLUSIONS:The core collections established in this study will facilitate agricultural research by providing the broad phenotypic data needed for crop improvement programs and by shedding light on the undiscovered biodiversity of L. albus genetic resources. Understanding the variation found in such resources will allow us to develop sustainable tools and technologies that address global challenges such as the provision of healthy and sustainable diets for all and the mitigation of climate change.
Over more than 80 years, the collections of the German Federal Ex Situ Genebank for Agricultural and Horticultural Crops have grown to around 152,000 accessions of 3,000 species preserved at three locations: Gatersleben, Groß Lüsewitz and Malchow/Poel. More than 96% of the material is stored as desiccation-tolerant orthodox seeds according to the active–base–safety (A-B-S) replicate approach at -18°C. Almost 70,000 freshly regenerated safety replicates are stored in the Svalbard Global Seed Vault. However, 4% of the material (2,000 field, 3,000 in vitro and 2,500 cryopreserved accessions) can only be maintained vegetatively, as no or few seeds or no true-breeding seeds are available. Most of the accessions are provided via the standard material transfer agreement (SMTA) and more than 1.2 million samples have been distributed since the genebank was founded. To guarantee the identity of the living plant material, reference samples comprising about 450,000 voucher specimens, 110,000 seed and fruit samples and 57,000 cereal spikes are used for comparisons. Genebank workflows are supported by the Genebank Information System (GBIS), which also manages workflow-independent data to describe the genebank accessions by passport, phenotypic and taxonomic data, thus allowing users to make targeted selections of material. The genebank-related processes, including acquisition, preservation, regeneration, documentation and material distribution, are certified for quality management in accordance with ISO 9001. Nowadays, the genebank is undergoing a transformation process to become a bio-digital resource centre to improve utilization of the genetic resources in research and breeding to address future challenges.
Rapeseed is one of the most important agricultural crops and is used in many ways. Due to the advancing climate crisis, the yield potential of rapeseed is increasingly impaired. In addition to changing environmental conditions, the expansion of cultivated areas also favours the infestation of rapeseed with various pests and pathogens. This results in the need for continuous further development of rapeseed varieties. To this end, the potential of the rapeseed gene pool should be exploited, as the various species included in it contain promising resistance alleles against pests and pathogens. In general, the biodiversity of crops and their wild relatives is increasingly endangered. In order to conserve them and to provide impulses for breeding activities as well, strategies for the conservation of plant genetic resources are necessary. In this study, we investigated to what extent the different species of the rapeseed gene pool are conserved in European genebanks and what gaps exist. In addition, a niche modelling approach was used to investigate how the natural distribution ranges of these species are expected to change by the end of the century, assuming different climate change scenarios. It was found that most species of the rapeseed gene pool are significantly underrepresented in European genebanks, especially regarding representation of the natural distribution areas. The situation is exacerbated by the fact that the natural distributions are expected to change, in some cases significantly, as a result of ongoing climate change. It is therefore necessary to further develop strategies to prevent the loss of wild relatives of rapeseed. Based on the results of the study, as a first step we have proposed a priority list of species that should be targeted for collecting in order to conserve the biodiversity of the rapeseed gene pool in the long term.
Plant genetic resources (PGR) stored at genebanks are humanity's crop diversity savings for the future. Information on PGR contrasted with modern cultivars is key to select PGR parents for pre-breeding. Genotyping-by-sequencing was performed for 7,745 winter wheat PGR samples from the German Federal ex situ genebank at IPK Gatersleben and for 325 modern cultivars. Whole-genome shotgun sequencing was carried out for 446 diverse PGR samples and 322 modern cultivars and lines. In 19 field trials, 7,683 PGR and 232 elite cultivars were characterized for resistance to yellow rust one of the major threats to wheat worldwide. Yield breeding values of 707 PGR were estimated using hybrid crosses with 36 cultivars an approach that reduces the lack of agronomic adaptation of PGR and provides better estimates of their contribution to yield breeding. Cross-validations support the interoperability between genomic and phenotypic data. The here presented data are a stepping stone to unlock the functional variation of PGR for European pre-breeding and are the basis for future breeding and research activities.
Genomic prediction of genebank accessions benefits from the consideration of additive-by-additive epistasis and subpopulation-specific marker effects. Wheat (Triticum aestivum L.) and other species of the Triticum genus are well represented in genebank collections worldwide. The substantial genetic diversity harbored by more than 850,000 accessions can be explored for their potential use in modern plant breeding. Characterization of these large number of accessions is constrained by the required resources, and this fact limits their use so far. This limitation might be overcome by engaging genomic prediction. The present study compared ten different genomic prediction approaches to the prediction of four traits, namely flowering time, plant height, thousand grain weight, and yellow rust resistance, in a diverse set of 7745 accession samples from Germany’s Federal ex situ genebank at the Leibniz Institute of Plant Genetics and Crop Plant Research in Gatersleben. Approaches were evaluated based on prediction ability and robustness to the confounding influence of strong population structure. The authors propose the wide application of extended genomic best linear unbiased prediction due to the observed benefit of incorporating additive-by-additive epistasis. General and subpopulation-specific additive ridge regression best linear unbiased prediction, which accounts for subpopulation-specific marker-effects, was shown to be a good option if contrasting clusters are encountered in the analyzed collection. The presented findings reaffirm that the trait’s genetic architecture as well as the composition and relatedness of the training set and test set are major driving factors for the accuracy of genomic prediction.
The great efforts spent in the maintenance of past diversity in genebanks are rationalized by the potential role of plant genetic resources (PGR) in future crop improvement—a concept whose practical implementation has fallen short of expectations. Here, we implement a genomics-informed prebreeding strategy for wheat improvement that does not discriminate against nonadapted germplasm. We collect and analyze dense genetic profiles for a large winter wheat collection and evaluate grain yield and resistance to yellow rust (YR) in bespoke core sets. Breeders already profit from wild introgressions but PGR still offer useful, yet unused, diversity. Potential donors of resistance sources not yet deployed in breeding were detected, while the prebreeding contribution of PGR to yield was estimated through ‘Elite × PGR’ F 1 crosses. Genomic prediction within and across genebanks identified the best parents to be used in crosses with elite cultivars whose advanced progenies can outyield current wheat varieties in multiple field trials.
Abstract Wheat (Triticum aestivum L.) and other species of the Triticum genus are well represented in genebank collections worldwide. The large genetic diversity harbored by more than 850,000 accessions can be explored for the exploitation in modern breeding programs. Shortcomings in the characterization of accession which limit their use so far might be overcome by engaging genomic prediction. The present report aimed to compare ten different genomic prediction approaches for the prediction of four traits, namely flowering time, plant height, thousand grain weight, and yellow rust resistance, in a diverse set of 7,745 accession samples from Germany’s Federal ex situ genebank at IPK Gatersleben. Approaches were evaluated based on the prediction ability and for the robustness when facing strong population structure. The authors propose the wide application of EG-BLUP due to the observed benefit of incorporating additive-by-additive epistasis. Accounting for subpopulation-specific marker-effects with GSA-RRBLUP was shown to be a good option if contrasting clusters are encountered in the analyzed collection. In general, the trait’s genetic architecture as well as the composition and the relation of training set and test set were revealed as major driving factor in genomic prediction.
Abstract The European Search Catalogue for Plant Genetic Resources (EURISCO) is a central entry point for information on crop plant germplasm accessions from institutions in Europe and beyond. In total, it provides data on more than two million accessions, making an important contribution to unlocking the vast genetic diversity that lies deposited in >400 germplasm collections in 43 countries. EURISCO serves as the reference system for the Plant Genetic Resources Strategy for Europe and represents a significant approach for documenting and making available the world’s agrobiological diversity. EURISCO is well established as a resource in this field and forms the basis for a wide range of research projects. In this paper, we present current developments of EURISCO, which is accessible at http://eurisco.ecpgr.org.
The optimal use of legume genetic resources represents a key prerequisite for coping with current agriculture‐related societal challenges, including conservation of agrobiodiversity, agricultural sustainability, food security, and human health. Among legumes, the common bean (Phaseolus vulgaris) is the most economically important for human consumption, and its evolutionary trajectories as a species have been crucial to determining the structure and level of its present and available genetic diversity. Genomic advances are considerably enhancing the characterization and assessment of important genetic variants. For this purpose, the development and availability of, and access to, well‐described and efficiently managed genetic resource collections that comprise pure lines derived by single‐seed‐descent cycles will be paramount for the use of the reservoir of common bean variability and for the advanced breeding of legume crops. This is one of the main aims of the new and challenging European project INCREASE, which is the implementation of Intelligent Collections with appropriate standardized protocols that must be characterized, maintained, and made available, along with the related data, to users such as breeders and researchers. © 2021 The Authors. Current Protocols published by Wiley Periodicals LLC.
The great efforts spent in the maintenance of past diversity in genebanks are rationalized by the potential role of plant genetic resources in future crop improvement – a concept whose practical implementation has fallen short of expectations. Here, we implement genomics-informed parent selection to expedite pre-breeding without discriminating against non-adapted germplasm. We collect dense genetic profiles for a large winter wheat collection and evaluate grain yield and resistance to yellow rust in representative coresets. Genomic prediction within and across genebanks identified the best parents for PGR x elite derived crosses that outyielded current elite cultivars in multiple field trials.
Well-characterized genetic resources are fundamental to maintain and provide the various genotypes for pre-breeding programs for the production of new cultivars (e.g., wild relatives, unimproved material, landraces). The aim of the current article is to provide protocols for the characterization of the genetic resources of two lupin crop species: the European Lupinus albus and the American Lupinus mutabilis. Intelligent nested collections of lupins derived from homozygous lines (single-seed descent) are being developed, established, and exploited using cutting-edge approaches for genotyping, phenotyping, data management, and data analysis within the INCREASE project (EU Horizon 2020). This will allow us to predict the phenotypic performance of genotyped lines, and will further boost research and development in lupins. Lupins stand out due to their high-quality seed protein (∼40% of seed dry weight) and other primary components in the seeds, which include fatty acids, dietary fiber, and minerals. The potential of lupins as a crop is highlighted by the multiple benefits of plant-based food in terms of food security, nutrition, human health, and sustainable production. The use of lupins in foods, along with other well-studied and widely used food legumes, will also provide a greatly diversified plant-based food palette to meet the Global Goals for Sustainable Development to improve people's lives by 2030. © 2021 The Authors. Current Protocols published by Wiley Periodicals LLC. Basic Protocol 1: Lupin seed phenotypic descriptors Basic Protocol 2: Lupin seed imaging Basic Protocol 3: Standardized phenotypic characterization of lupin genetic resources grown towards primary seed increase (development of single-seed descent genetic resources).
SUMMARY Food legumes are crucial for all agriculture‐related societal challenges, including climate change mitigation, agrobiodiversity conservation, sustainable agriculture, food security and human health. The transition to plant‐based diets, largely based on food legumes, could present major opportunities for adaptation and mitigation, generating significant co‐benefits for human health. The characterization, maintenance and exploitation of food‐legume genetic resources, to date largely unexploited, form the core development of both sustainable agriculture and a healthy food system. INCREASE will implement, on chickpea ( Cicer arietinum ), common bean ( Phaseolus vulgaris ), lentil ( Lens culinaris ) and lupin ( Lupinus albus and L. mutabilis ), a new approach to conserve, manage and characterize genetic resources. Intelligent Collections , consisting of nested core collections composed of single‐seed descent‐purified accessions (i.e., inbred lines), will be developed, exploiting germplasm available both from genebanks and on‐farm and subjected to different levels of genotypic and phenotypic characterization. Phenotyping and gene discovery activities will meet, via a participatory approach, the needs of various actors, including breeders, scientists, farmers and agri‐food and non‐food industries, exploiting also the power of massive metabolomics and transcriptomics and of artificial intelligence and smart tools. Moreover, INCREASE will test, with a citizen science experiment, an innovative system of conservation and use of genetic resources based on a decentralized approach for data management and dynamic conservation. By promoting the use of food legumes, improving their quality, adaptation and yield and boosting the competitiveness of the agriculture and food sector, the INCREASE strategy will have a major impact on economy and society and represents a case study of integrative and participatory approaches towards conservation and exploitation of crop genetic resources.
Summary Enabling data reuse and knowledge discovery is increasingly critical in modern science, and requires an effort towards standardising data publication practices. This is particularly challenging in the plant phenotyping domain, due to its complexity and heterogeneity. We have produced the MIAPPE 1.1 release, which enhances the existing MIAPPE standard in coverage, to support perennial plants, in structure, through an explicit data model, and in clarity, through definitions and examples. We evaluated MIAPPE 1.1 by using it to express several heterogeneous phenotyping experiments in a range of different formats, to demonstrate its applicability and the interoperability between the various implementations. Furthermore, the extended coverage is demonstrated by the fact that one of the datasets could not have been described under MIAPPE 1.0. MIAPPE 1.1 marks a major step towards enabling plant phenotyping data reusability, thanks to its extended coverage, and especially the formalisation of its data model, which facilitates its implementation in different formats. Community feedback has been critical to this development, and will be a key part of ensuring adoption of the standard.
Genebanks play an important role in the long-term conservation of plant genetic resources and are complementary to the conservation of diversity in farmers' fields and in nature. In this context, documentation plays a critical role. Without well-structured documentation, it is not possible to make statements about the value of a resource, especially with regard to its potential for breeding and research. In particular, comprehensive information management is a prerequisite for the further development of genebank collections. This requires detailed information about the composition of a collection, thus allowing statements about which species and/or regions of origin are under-represented. This task is of strategic importance, especially due to the threats to crop plants and their wild relatives caused by advancing climate change. Both the actual conservation management and the fulfilment of legal obligations depend on information. Hence, documentation units have been established in almost all genebanks worldwide. They all face the challenge that knowledge about genebank accessions must be permanently managed and passed on across generations. International standards such as Multi-Crop Passport Descriptors (MCPD) have been established for the exchange of data between genebanks, and allow the operation of international information systems, such as the World Information and Early Warning System on Plant Genetic Resources for Food and Agriculture (WIEWS), the European Search Catalogue for Plant Genetic Resources (EURISCO) or Genesys.
Genebanks are valuable sources of genetic diversity, which can help to cope with future problems of global food security caused by a continuously growing population, stagnating yields and climate change. However, the scarcity of phenotypic and genotypic characterization of genebank accessions severely restricts their use in plant breeding. To warrant the seed integrity of individual accessions during periodical regeneration cycles in the field phenotypic characterizations are performed. This study provides non-orthogonal historical data of 12,754 spring and winter wheat accessions characterized for flowering time, plant height, and thousand grain weight during 70 years of seed regeneration at the German genebank. Supported by historical weather observations outliers were removed following a previously described quality assessment pipeline. In this way, ready-to-use processed phenotypic data across regeneration years were generated and further validated. We encourage international and national genebanks to increase their efforts to transform into bio-digital resource centers. A first important step could consist in unlocking their historical data treasures that allows an educated choice of accessions by scientists and breeders.