Plant genetic resources are considered a treasure trove of valuable, untapped diversity that holds the key to breeding the crops of the future. However, the use of these resources in breeding is often limited due to the lack of comprehensive phenotypic characterization. The present study provides extensive historical phenotypic data from nine genebanks as a MIAPPE compliant data set. We compiled and curated phenotypic data from 43,293 wheat accessions, encompassing 460,399 data points across 52 traits, including the three core traits of plant height, heading time, and thousand kernel weight from seven decades. The exceptional quality of the presented dataset was highlighted by predominantly high heritabilities. Phenotypic data of such quantity and quality is a crucial resource for unlocking the valuable diversity of plant genetic resources for agricultural advancement.
Population growth and the impacts of climate change are placing increasing pressure on global agriculture and breeding programmes. Recent advancements in phenotyping techniques, genotyping technologies, and predictive modelling are accelerating genetic gains in breeding programmes, helping researchers and breeders develop improved crops more efficiently. However, these advancements have also led to an overwhelming torrent of fragmented data, creating significant challenges in data integration and management. To address this issue, the Breeding Application Programming Interface (BrAPI) project was established as a standardized data model for breeding data. BrAPI is an international, community-driven effort that facilitates interoperability among databases and tools, improving the sharing and interpretation of breeding-related data. This open-source standard is software-agnostic and can be used by anyone interested in breeding, phenotyping, germplasm, genotyping, and agronomy data management. This manuscript provides an overview of the BrAPI project, highlighting the significant progress made in the development of the data standard and the expansion of its community. It also presents a showcase of the wide variety of BrAPI-compatible tools that have been built to enhance breeding and research activities, demonstrating how the project is advancing agricultural innovation and data management practices.
Making sense of whole-genome polymorphism data is challenging, but it is essential for overcoming the biases in SNP data. Here we analyze 27 genomes of Arabidopsis thaliana to illustrate these issues. Genome size variation is mostly due to tandem repeat regions that are difficult to assemble. However, while the rest of the genome varies little in length, it is full of structural variants, mostly due to transposon insertions. Because of this, the pangenome coordinate system grows rapidly with sample size and ultimately becomes 70% larger than the size of any single genome, even for n = 27. Finally, we show how short-read data are biased by read mapping. SNP calling is biased by the choice of reference genome, and both transcriptome and methylome profiling results are affected by mapping reads to a reference genome rather than to the genome of the assayed individual.
The generation and analysis of genome-scale data - genomics - is driving a rapid increase in plant biodiversity knowledge. However, the speed and complexity of technological advance in genomics presents challenges for the widescale use of genomics in evolutionary and conservation biology. We introduce and describe a national-scale collaboration conceived to build genomic resources and capability for understanding the Australian flora: the Genomics for Australian Plants (GAP) Framework Initiative. We outline (a) the history of the project including the collaborative framework, partners and funding; (b) GAP principles such as rigour in design, sample verification and documentation, data management and data accessibility; and (c) the structure of the consortium and the four associated activity streams (reference genomes, phylogenomics, conservation genomics and training), with the rationale and aims for each of these. We show, through discussion of successes and challenges, the value of this multi-institutional consortium approach and the enablers, such as well-curated collections and national collaborative research infrastructure, all of which have led to a substantial increase in capacity and delivery of biodiversity knowledge outcomes.
Societal Impact Statement Seedbanks are vital for biodiversity conservation, but their potential remains underutilised due to a limited understanding of the intraspecific genetic diversity they hold. By leveraging digitised data associated with seedbank collections, such as sampling locations, number of maternal plants and seed traits, we can attempt the estimation of genetic variation and identify gaps in collections, enabling better prioritisation of species for conservation efforts. These advancements can inform policy targets like those of the Kunming‐Montreal Global Biodiversity Framework, promoting more effective conservation strategies. Digitisation and emerging machine‐learning technologies offer scalable, cost‐efficient solutions to enhance conservation knowledge, ensuring biodiversity resilience for future generations. Summary Seedbank collections hold significant untapped potential for advancing conservation science and practice, but the intraspecific genetic diversity (i.e. diversity within a species) stored in worldwide seedbank collections remains largely unknown, hindering the effective use of seeds for both informing and implementing in situ interventions. As producing genetic data is time‐consuming and expensive, other data associated with seedbank collections can greatly enhance our understanding of the genetic variation stored in seed collections when genetic data are unavailable. Information such as the location of sampling sites, estimated population size and the number of mother plants from which seeds were collected can facilitate the estimation of the genetic diversity captured in the collections. This information can also be used to estimate the sampling effort required to fill gaps in seedbank collections to better represent genetic diversity, through comparison with existing baselines from species where genetic diversity is characterised, and through simulations. Digitisation of the data associated with seedbank collections makes the approaches above practicable at scale. In addition, digital images of the seeds themselves may identify intraspecific phenotypic variation and can, therefore, be used to prioritise populations for future genetic studies. In this article, we explore the potential of digitised information made available by seedbanks for improving our understanding of the intraspecific genetic diversity preserved in collections. We describe possible improvements that might enhance the predictive power of digital information for genetic studies, and discuss the challenges and opportunities associated with these.
DNA recovered from herbarium specimens represents a vital asset in botanical research, playing a pivotal role in unravelling the evolution, diversity, and ecological dynamics of plants. Despite its importance, challenges such as fragmented DNA and insufficient sequencing yields render molecular data retrieval a high-risk and costly endeavour involving the use of non-replaceable herbarium specimens. Here, we propose a framework based on Artificial Intelligence (AI) to forecast the success of genomic DNA extraction suitable for sequencing from herbarium samples. Our model integrates morphological characteristics and sample colour derived from scanned herbarium images, metadata including sample age and locality, and DNA quantity measurements of samples. We train a deep learning algorithm with ca. 2,000 specimens that have been digitized and sequenced in the framework of the Plant and Fungal Trees of Life (PAFTOL) Project, spanning from year 1832 to the present. As training datasets increase with ongoing digitization and genomic sequencing efforts, our AI predictive model can support researchers in selecting the herbarium samples with the highest likelihood of yielding high-quality genomic DNA from amongst a vast array of globally distributed candidate specimens. Our approach enhances the contribution of herbarium-derived DNA in large-scale studies and facilitates the utilisation of historical collections for a deeper understanding of plant evolution and ecology, with implications for conservation. ### Competing Interest Statement The authors have declared no competing interest.
Biodiversity genomics research requires reliable organismal identification, which can be difficult based on morphology alone. DNA-based identification using DNA barcoding can provide confirmation of species identity and resolve taxonomic issues but is rarely used in studies generating reference genomes. Here, we describe the development and implementation of DNA barcoding for the Darwin Tree of Life Project (DToL), which aims to sequence and assemble high quality reference genomes for all eukaryotic species in Britain and Ireland. We present a standardised framework for DNA barcode sequencing and data interpretation that is then adapted for diverse organismal groups. DNA barcoding data from over 12,000 DToL specimens has identified up to 20% of samples requiring additional verification, with 2% of seed plants and 3.5% of animal specimens subsequently having their names changed. We also make recommendations for future developments using new sequencing approaches and streamlined bioinformatic approaches.
Arabidopsis thaliana was the first plant for which a high-quality genome sequence became available. The publication of the first reference genome sequence almost 25 years ago was already accompanied by genome-wide data on sequence polymorphisms in another accession, or naturally occurring strain. Since then, inventories of genome-wide diversity have been generated at increasingly precise levels. High-density genotype data for A. thaliana , including those from the 1001 Genomes Project, were key to demonstrating the enormous power of GWAS in inbred populations of wild plants, and the comparison of intraspecific polymorphism with interspecific divergence has illuminated many aspects of plant genome evolution. Over the past decade, an increasing number of nearly complete genome sequences have been published for many more accessions. Here, we highlight the diversity of a curated collection of previously published and so far unpublished genome sequences assembled using different types of long reads, including PacBio Continuous Long Reads (CLR), PacBio High Fidelity (HiFi) reads, and Oxford Nanopore Technologies (ONT) reads. This 1001 Genomes Plus (1001G+) resource is being made available at http://1001genomes.org. We invite colleagues with yet unpublished genome assemblies from A. thaliana accessions to contribute to this effort. ### Competing Interest Statement D.W. holds equity in Computomics, which advises plant breeders. D.W. also consults for KWS SE, a globally active plant breeder and seed producer. J.F. is an employee of Tropic TI, Lda. All other authors declare no competing interests.
Our view of genetic polymorphism is shaped by methods that provide a limited and reference-biased picture. Long-read sequencing technologies, which are starting to provide nearly complete genome sequences for population samples, should solve the problem—except that characterizing and making sense of non-SNP variation is difficult even with perfect sequence data. Here, we analyze 27 genomes of Arabidopsis thaliana in an attempt to address these issues, and illustrate what can be learned by analyzing whole-genome polymorphism data in an unbiased manner. Estimated genome sizes range from 135 to 155 Mb, with differences almost entirely due to centromeric and rDNA repeats. The completely assembled chromosome arms comprise roughly 120 Mb in all accessions, but are full of structural variants, many of which are caused by insertions of transposable elements (TEs) and subsequent partial deletions of such insertions. Even with only 27 accessions, a pan-genome coordinate system that includes the resulting variation ends up being 40% larger than the size of any one genome. Our analysis reveals an incompletely annotated mobile-ome: our ability to predict what is actually moving is poor, and we detect several novel TE families. In contrast to this, the genic portion, or “gene-ome”, is highly conserved. By annotating each genome using accession-specific transcriptome data, we find that 13% of all genes are segregating in our 27 accessions, but that most of these are transcriptionally silenced. Finally, we show that with short-read data we previously massively underestimated genetic variation of all kinds, including SNPs—mostly in regions where short reads could not be mapped reliably, but also where reads were mapped incorrectly. We demonstrate that SNP-calling errors can be biased by the choice of reference genome, and that RNA-seq and BS-seq results can be strongly affected by mapping reads to a reference genome rather than to the genome of the assayed individual. In conclusion, while whole-genome polymorphism data pose tremendous analytical challenges, they will ultimately revolutionize our understanding of genome evolution.### Competing Interest StatementD.W.∼holds equity in Computomics, which advises plant breeders. D.W.∼also consults for KWS SE, a plant breeder and seed producer with activities throughout the world. J.F. is an employee of Tropic TI, Lda. All other authors declare no competing interests.
Angiosperms are the cornerstone of most terrestrial ecosystems and human livelihoods(1,2). A robust understanding of angiosperm evolution is required to explain their rise to ecological dominance. So far, the angiosperm tree of life has been determined primarily by means of analyses of the plastid genome(3,4). Many studies have drawn on this foundational work, such as classification and first insights into angiosperm diversification since their Mesozoic origins(5-7). However, the limited and biased sampling of both taxa and genomes undermines confidence in the tree and its implications. Here, we build the tree of life for almost 8,000 (about 60%) angiosperm genera using a standardized set of 353 nuclear genes(8). This 15-fold increase in genus-level sampling relative to comparable nuclear studies(9) provides a critical test of earlier results and brings notable change to key groups, especially in rosids, while substantiating many previously predicted relationships. Scaling this tree to time using 200 fossils, we discovered that early angiosperm evolution was characterized by high gene tree conflict and explosive diversification, giving rise to more than 80% of extant angiosperm orders. Steady diversification ensued through the remaining Mesozoic Era until rates resurged in the Cenozoic Era, concurrent with decreasing global temperatures and tightly linked with gene tree conflict. Taken together, our extensive sampling combined with advanced phylogenomic methods shows the deep history and full complexity in the evolution of a megadiverse clade.
Societal Impact StatementDiverse gene pools are fundamental to crop improvement, biodiversity maintenance and environmental management. The UKCropDiversity‐HPC high‐performance computing resource enables seven UK institutes to perform plant and conservation research with increased efficiency, cost‐effectiveness and environmental sustainability. It supports research across numerous areas, including bioinformatics, genetics, phenomics and conservation ‐ including Artificial Intelligence approaches. Its utilisation supports many United Nations Sustainable Development Goals, including Goals‐2 (Zero Hunger), −13 (Climate Action), −15 (Life on Land), −9 (Industry, Innovation and Infrastructure) and −4 (Quality Education). Accordingly, UKCropDiversity‐HPC helps maximise the societal impact of research undertaken at our seven institutes, driving positive change for future generations.
The group of > 40 cryptic whitefly species called Bemisia tabaci sensu lato are amongst the world's worst agricultural pests and plant-virus vectors. Outbreaks of B. tabaci s.l. and the associated plant-virus diseases continue to contribute to global food insecurity and social instability, particularly in sub-Saharan Africa and Asia. Published B. tabaci s.l. genomes have limited use for studying African cassava B. tabaci SSA1 species, due to the high genetic divergences between them. Genomic annotations presented here were performed using the 'Ensembl gene annotation system', to ensure that comparative analyses and conclusions reflect biological differences, as opposed to arising from different methodologies underpinning transcript model identification.We present here six new B. tabaci s.l. genomes from Africa and Asia, and two re-annotated previously published genomes, to provide evolutionary insights into these globally distributed pests. Genome sizes ranged between 616-658 Mb and exhibited some of the highest coverage of transposable elements reported within Arthropoda. Many fewer total protein coding genes (PCG) were recovered compared to the previously published B. tabaci s.l. genomes and structural annotations generated via the uniform methodology strongly supported a repertoire of between 12.8-13.2 × 103 PCG. An integrative systematics approach incorporating phylogenomic analysis of nuclear and mitochondrial markers supported a monophyletic Aleyrodidae and the basal positioning of B. tabaci Uganda-1 to the sub-Saharan group of species. Reciprocal cross-mating data and the co-cladogenesis pattern of the primary obligate endosymbiont 'Candidatus Portiera aleyrodidarum' from 11 Bemisia genomes further supported the phylogenetic reconstruction to show that African cassava B. tabaci populations consist of just three biological species. We include comparative analyses of gene families related to detoxification, sugar metabolism, vector competency and evaluate the presence and function of horizontally transferred genes, essential for understanding the evolution and unique biology of constituent B. tabaci. s.l species.These genomic resources have provided new and critical insights into the genetics underlying B. tabaci s.l. biology. They also provide a rich foundation for post-genomic research, including the selection of candidate gene-targets for innovative whitefly and virus-control strategies.
UK natural science collections hold over 137 million items, an unrivalled source of data about 4.56 billion years of planetary development and hundreds of years of biological change, including the differences made by humans — but the scientific, commercial, and societal benefits of these collections are constrained by the limits of physical access, and by highly fragmented digitisation efforts with less than 10% digitally available. Following work with Frontier Economics in 2021, which showed potential for £2 billion in benefits to the UK economy from digitising all UK natural science collections, in 2022–23 the Natural History Museum London worked, with analytical support from McKinsey and Company, to understand the impact of what has already been digitised and shared by UK natural science collections — what is the demand for these data, what are they used for, and how does this deliver efficient, effective and impactful research? This study focuses on usage via the Global Biodiversity Information Facility, the largest source of relevant usage data, examining 7.6 million records from twelve UK institutions. While these UK collections data are just 0.3% of total GBIF occurrences, they are cited in 12% of peer reviewed publications citing GBIF data, showing the disproportionate impact of UK collections data and the historical, geographical, and taxonomic richness that they bring. Researchers have already benefited from more than £18 million of efficiency savings from digital UK specimen data. Data from natural science collections held in the UK are uniquely impactful resources, vital to a future in which people and planet thrive, and a step change in the pace of digitisation is needed to unlock their potential for researchers, policymakers, and society.