Soil eukaryotes, including fungi, protists, plants, and animals, are central to biosphere functioning and resilience. The Global Standardised Soil Eukaryome Dataset (GloSED) is the first dataset encompassing the entire spectrum of soil eukaryotes, covering 4,063 sampling sites in 121 countries on all continents, revealing nearly one million operational taxonomic units. All samples were collected and analysed using a standardised protocol minimizing technical biases. Long-read sequencing of full-length ITS and 18S-V9 regions provide broad taxonomic coverage and high-resolution identification supported by specialist curation of "dark taxa". A rigorous bioinformatic processing ensures against homopolymer errors, PCR-mediated chimeras, and index switching providing high data quality. The dataset is supported by raw sequences and an open-source containerised workflow for reproducible analyses. The samples are accompanied by land-cover description and directly measured soil pH, δ13C, δ15N, as well as P, K, Ca, Mg, and total C and N contents. GloSED is the first database that enables ecological and biogeographic studies of entire soil eukaryotic communities from local to global scales.
Estimating fungal geographic ranges and niche potential is limited by the ephemeral nature of fruiting bodies. While environmental DNA offers broader insights, species-level identification remains difficult due to uncertain sequence clustering thresholds, low interspecific variation in barcoding regions, and limited taxonomic resolution. To address this, we analyzed large-scale environmental sequence data to investigate biogeographic patterns and ectomycorrhizal host associations in Russula subsection Xerampelinae, a group widely distributed across boreal and arctic ecosystems. We used maximum likelihood phylogenetic methods to identify internal transcribed spacer sequences from the UNITE database for 13 previously defined Russula species. Additionally, we retrieved sequence-associated metadata on locality and host association from PlutoF. In total, 1363 sequences were resolved in clades with our target species, expanding known distributions and plant host associations. Integrating environmental metadata with phylogenetic analyses revealed widespread distributions and host generalism across the Northern Hemisphere. However, UNITE's species hypotheses clustering thresholds often failed to reflect phylogenetically defined species boundaries, even at conservative levels. Our phylogeny-based assignment approach applied on metabarcoding data improves species resolution and biogeographic inference, providing a step-by-step workflow for sequence-based identification and for addressing ecologically and evolutionarily driven research questions.
Molecular analyses of soil and water commonly reveal large proportions of fungal taxa that cannot be assigned to any taxonomic or functional groups. Some of these so-called dark taxa have been encoded alphanumerically, while others have remained completely overlooked. Using long-read sequencing that covers much of the ribosomal RNA operon, we shed light on the phylogenetic and ecological distribution of fungal dark taxa and formally describe 30 of the most prominent phylum- to order-level lineages based on their characteristic DNA features. This increases the known large-scale fungal phylogenetic diversity by roughly one-third. Formal names will enhance taxonomic reproducibility, facilitate communication among researchers, and enable the estimation of conservation and quarantine needs for uncultivable species and higher-ranking taxa. The new species in the respective highest-level novel taxonomic groups include Pantelleria saittana (Pantelleriomycetes), Paraspizellomyces parrentiae (Paraspizellomycetales), Aquieurochytrium lacustre (Aquieurochytriomycetes), Edaphochytrium valuojaense (Edaphochytriomycetes), Tibetochytrium taylorii (Tibetochytriomycetes), Tropicochytrium toronegroense (Tropicochytriomycetes), Algovorax scenedesmi (comb. nov.) and Solivorax pantropicus (Algovoracomycetes), Aquamastix sanduskyensis (Aquamastigomycetes), Cantoromastix holarctica (Cantoromastigomycetes), Dobrisimastix vlkii (Dobrisimastigomycetes), Palomastix lacustris (Palomastigomycetes), Sedimentomastix tueriensis (Sedimentomastigomycetes), Terrincola waldropii (Terrincolales), Curlevskia holarctica (Curlevskiomycota), Mycosocceria estonica (Mycosocceriales), Maerjamyces jumpponenii (Maerjamycetes), Ruderalia cosmopolita (Ruderaliomycetes), Bryolpidium mundanum (Bryolpidiomycetes), Chthonolpidium enigmatum (Chthonolpidiomycetes), Savannolpidium raadiense (Savannolpidiomycetes), Gelotisporidium boreale (Gelotisporidiomycetes), Sumavosporidium sylvestre (Sumavosporidiomycetes), Parakickxella borikenica (Parakickxellomycetes), Aldinomyces tarquinii (Aldinomycota), Borikenia urbinae (Borikeniomycota), Mirabilomyces abrukanus (Mirabilomycota), Nematovomyces vermicola (comb. nov.) and N. soinasteënsis (Nematovomycota), Viljandia globalis (Viljandiomycota), Waitukubulimyces cliftonii (Waitukubulimycota), and Tartumyces setoi (Tartumyceta).
In navigating the biodiversity crisis, a major uncertainty is the conservation status of inconspicuous, yet megadiverse and functionally crucial, soil organisms. Massive datasets on soil biota are accumulating through molecular sampling approaches, but to date these datasets have provided only limited input into conservation planning and management. We investigated how environmental DNA (eDNA) data of soil macrofungi contribute to regional Red List assessments, which are currently based on fruiting bodies (hereafter, fruit-bodies). In our test region of Estonia (northern Europe), which contained similar to 15,000 fruit-body records for 1583 assessed species, an average soil sample increased the range estimates of Threatened and Near Threatened fungal species by 0.18%. Five hundred soil samples almost doubled their known localities and added 19% previously unrecorded species. However, even after accumulating >1000 soil samples, about half of the Threatened and Near Threatened species known by fruit-bodies remained undetected through eDNA techniques. Effective conservation assessment of soil fungi thus requires both fruit-body and eDNA data; therefore, special efforts are needed to make these data available to conservationists.
Phylogenetic diversity (PD) represents a fundamental measure of biodiversity, encapsulating the extent of evolutionary history within species groups. This measure, pivotal for understanding biodiversity's full dimension, has gained recognition by major environmental and scientific organisations, including the Intergovernmental Science-Policy Platform on Biodiversity and Ecosystem Services. Unlike traditional taxonomic richness, PD offers a comprehensive, evolutionary perspective on biodiversity, essential for conservation planning and biodiversity management. This manuscript describes the development of a BioDT (Biodiversity Digital Twin) prototype, aimed at facilitating the calculation and visualisation of biodiversity metrics from global, dynamic data sources. By utilising the PhyloNext pipeline and integrating with global phylogenetic and species occurrence databases like the Open Tree of Life (OToL) and the Global Biodiversity Information Facility (GBIF), the prototype aims to significantly reduce computation time and enhance user interaction. This enables dynamic visualisation and potentially hypothesis testing, making it a valuable tool for researchers, monitoring initiatives, policy-makers and the public. The prototype's development focuses on improving the PhyloNext pipeline's scalability and creating a more intuitive user interface, expanding its utility for conservation efforts and biodiversity exploration. Our work illustrates the potential impact of the BioDT prototype in supporting diverse user groups in visualising and exploring PD, thus contributing to more informed decision-making in conservation and biodiversity management.
Molecular sequencing data generation is being driven by global and regional efforts to discover, understand and monitor biodiversity. To fully explore this data in biodiversity research we need a network of connected data resources, linking sequence data with natural history collections, taxonomy and literature. The BiCIKL project (Biodiversity Community Integrated Knowledge Library, Penev et al. 2022) has set the groundwork towards creating this network of linked data and fostering FAIR (Findable, Accessible, Interoperable and Reusable) practices in the biodiversity domain. Connecting biodiversity and molecular data along the biodiversity research cycle requires a foundation of well-structured and rich metadata in the molecular sequence databases. Referencing the physical specimens is important as this provides context about the source of the material that was used for generating the molecular sequence data, including information about origin and species identification. To connect biodiversity and molecular data, we developed tools and workflows for improving and standardising metadata, federated searches and validations for specimen reference in sequence data, such as the SpASe tool, which enables the discovery of links between natural history collections and sequences, and the European Nucleotide Archive Source Attribute Helper API, which facilitates the construction of specimen attributes in a structured format. This work was done in close collaboration with DiSSCo (Distributed System of Scientific Collections) and some biodiversity genomics projects (e.g. Biodiversity Genomics Europe, BGE). Furthermore, we enabled community curation of biological source annotations such as specimen references in sequence data through the PlutoF platform and the ELIXIR Contextual Data Clearinghouse (Abarenkov et al. 2021, Balavenkataraman Kadhirvelu et al. 2022) and increased bidirectional linking from sequences in the European Nucleotide Archive (ENA) to collections, taxonomy and literature services (e.g., Plazi TreatmentBank, OpenBioDiv). We also worked closely with the community to enable the structured publication of environmental DNA data, promoting and engaging in the definition of standards and developing tools to facilitate data deposition and retrieval. Overall, the project has contributed significantly to strengthen the connections between the biodiversity and genomics communities towards higher data integration and interoperability. Structured, enriched, accessible and linked sequence data will provide a strong foundation for the application of biodiversity knowledge in the response to global challenges, such as biodiversity loss, ecosystem change and food security. Beyond BiCIKL, we will continue our work as a community to promote a culture of FAIR linked molecular data, towards a fully integrated biodiversity knowledge ecosystem.
Journal impact factors were devised to qualify and compare university library holdings but are frequently repurposed for use in ranking applications, research papers, and even individual applicants in mycology and beyond. The widely held assumption that mycological studies published in journals with high impact factors add more to systematic mycology than studies published in journals without high impact factors nevertheless lacks evidential underpinning. The present study uses the species hypothesis system of the UNITE database for molecular identification of fungi and other eukaryotes to trace the publication history and impact factor of sequences uncovering new fungal species hypotheses. The data show that journal impact factors are poor predictors of discovery potential in systematic mycology. There is no clear relationship between journal impact factor and the discovery of new species hypotheses for the years 2000–2021. On the contrary, we found journals with low, and even no, impact factor to account for substantial parts of the species hypothesis landscape, often discovering new fungal taxa that are only later picked up by journals with high impact factors. Funding agencies and hiring committees that insist on upholding journal impact factors as a central funding and recruitment criterion in systematic mycology should consider using indicators such as research quality, productivity, outreach activities, review services for scientific journals, and teaching ability directly rather than using publication in high impact factor journals as a proxy for these indicators.
BACKGROUND:Understanding biodiversity patterns is a central topic in biogeography and ecology, and it is essential for conservation planning and policy development. Diversity estimates that consider the evolutionary relationships among species, such as phylogenetic diversity and phylogenetic endemicity indices, provide valuable insights into the functional diversity and evolutionary uniqueness of biological communities. These estimates are crucial for informed decision-making and effective global biodiversity management. However, the current methodologies used to generate these metrics encounter challenges in terms of efficiency, accuracy, and data integration.RESULTS:We introduce PhyloNext, a flexible and data-intensive computational pipeline designed for phylogenetic diversity and endemicity analysis. The pipeline integrates GBIF occurrence data and OpenTree phylogenies with the Biodiverse software. PhyloNext is free, open-source, and provided as Docker and Singularity containers for effortless setup. To enhance user accessibility, a user-friendly, web-based graphical user interface has been developed, facilitating easy and efficient navigation for exploring and executing the pipeline. PhyloNext streamlines the process of conducting phylogenetic diversity analyses, improving efficiency, accuracy, and reproducibility. The automated workflow allows for periodic reanalysis using updated input data, ensuring that conservation strategies remain relevant and informed by the latest available data.CONCLUSIONS:PhyloNext provides researchers, conservationists, and policymakers with a powerful tool to facilitate a broader understanding of biodiversity patterns, supporting more effective conservation planning and policy development. This new pipeline simplifies the creation of reproducible and easily updatable phylogenetic diversity analyses. Additionally, it promotes increased interoperability and integration with other biodiversity databases and analytical tools.
Partner specificity is a well-documented phenomenon in biotic interactions, yet the factors that determine specificity in plant-fungal associations remain largely unknown. By utilizing composite soil samples, we identified the predictors that drive partner specificity in both plants and fungi, with a particular focus on ectomycorrhizal associations. Fungal guilds exhibited significant differences in overall partner preference and avoidance, richness, and specificity to specific tree genera. The highest level of specificity was observed in root endophytic and ectomycorrhizal associations, while the lowest was found in arbuscular mycorrhizal associations. The majority of ectomycorrhizal fungal species showed a preference for one of their partner trees, primarily at the plant genus level. Specialist ectomycorrhizal fungi were dominant in belowground communities in terms of species richness and relative abundance. Moreover, all tree genera (and occasionally species) demonstrated a preference for certain fungal groups. Partner specificity was not related to the rarity of fungi or plants or environmental conditions, except for soil pH. Depending on the partner tree genus, specific fungi became more prevalent and relatively more abundant with increasing stand age, tree dominance, and soil pH conditions optimal for the partner tree genus. The richness of partner tree species and increased evenness of ectomycorrhizal fungi in multi-host communities enhanced the species richness of ectomycorrhizal fungi. However, it was primarily the partner-generalist fungi that contributed to the high diversity of ectomycorrhizal fungi in mixed forests.
UNITE (https://unite.ut.ee) is a web-based database and sequence management environment for molecular identification of eukaryotes. It targets the nuclear ribosomal internal transcribed spacer (ITS) region and offers nearly 10 million such sequences for reference. These are clustered into ∼2.4M species hypotheses (SHs), each assigned a unique digital object identifier (DOI) to promote unambiguous referencing across studies. UNITE users have contributed over 600 000 third-party sequence annotations, which are shared with a range of databases and other community resources. Recent improvements facilitate the detection of cross-kingdom biological associations and the integration of undescribed groups of organisms into everyday biological pursuits. Serving as a digital twin for eukaryotic biodiversity and communities worldwide, the latest release of UNITE offers improved avenues for biodiversity discovery, precise taxonomic communication and integration of biological knowledge across platforms.
Fungal metabarcoding of substrates such as soil, wood, and water are uncovering an unprecedented number of fungal species that do not seem to produce tangible morphological structures and that defy our best attempts at cultivation, thus falling outside of the ambit of the International Code of Nomenclature for algae, fungi, and plants. The present study uses the new, ninth release of the species hypotheses of the UNITE database to show that species discovery through environmental sequencing vastly outpaces traditional, Sanger sequencing-based efforts in a strongly increasing trend over the last five years. Our findings challenge the present stance of the mycological community – that “the code” works fine and that these complications will somehow sort themselves out given enough time and a following wind – and suggest that we should be discussing not whether to allow DNA-based descriptions (typifications) of species and by extension higher ranks of fungi, but what the precise requirements for such DNA-based typifications should be. We submit a tentative list of such criteria for further discussion. However, the present authors fear that no waves of change will be lapping the shores of mycology for the foreseeable future, leaving the overwhelming majority of extant fungi without formal names and thus scientific and environmental agency. It is not clear to us who benefits from that, but neither fungi nor mycology are likely to be on the winning side.
How the multiple facets of soil fungal diversity vary worldwide remains virtually unknown, hindering the management of this essential species-rich group. By sequencing high-resolution DNA markers in over 4000 topsoil samples from natural and human-altered ecosystems across all continents, we illustrate the distributions and drivers of different levels of taxonomic and phylogenetic diversity of fungi and their ecological groups. We show the impact of precipitation and temperature interactions on local fungal species richness (alpha diversity) across different climates. Our findings reveal how temperature drives fungal compositional turnover (beta diversity) and phylogenetic diversity, linking them with regional species richness (gamma diversity). We integrate fungi into the principles of global biodiversity distribution and present detailed maps for biodiversity conservation and modeling of global ecological processes.
Fungal metabarcoding of substrates such as soil, wood, and water is uncovering an unprecedented number of fungal species that do not seem to produce tangible morphological structures and that defy our best attempts at cultivation, thus falling outside the scope of the International Code of Nomenclature for algae, fungi, and plants. The present study uses the new, ninth release of the species hypotheses of the UNITE database to show that species discovery through environmental sequencing vastly outpaces traditional, Sanger sequencing-based efforts in a strongly increasing trend over the last five years. Our findings challenge the present stance of some in the mycological community - that the current situation is satisfactory and that no change is needed to "the code" - and suggest that we should be discussing not whether to allow DNA-based descriptions (typifications) of species and by extension higher ranks of fungi, but what the precise requirements for such DNA-based typifications should be. We submit a tentative list of such criteria for further discussion. The present authors hope for a revitalized and deepened discussion on DNA-based typification, because to us it seems harmful and counter-productive to intentionally deny the overwhelming majority of extant fungi a formal standing under the International Code of Nomenclature for algae, fungi, and plants.
Partner specificity is a well-known phenomenon in biotic interactions, but little is known about biotic and abiotic factors that determine specificity in plant-fungal associations. Using PacBio sequencing of soils from monospecific and mixed forest stands, we determined the predictors driving partner specificity in both ectomycorrhizal plants and fungi. Fungal guilds differed strongly in the patterns of partner preference and avoidance, and specificity to particular tree genera. Specialist ectomycorrhizal fungi dominated in belowground communities, and most species preferred one of their partner trees - mostly at the plant genus level. Furthermore, all tree genera (sometimes species) displayed preference towards certain fungal groups. Partner specificity was unrelated to rarity of fungi or plants or environmental conditions except soil pH. Depending on partner taxon, specificity in fungi tended to increase with dominance and optimal pH of the partner tree genus and stand age. Partner tree richness and increased evenness of ectomycorrhizal fungi in multi-host communities promotes species richness. However, mainly partner-generalist fungi contribute to the high diversity in mixed forests. Our results further suggest that reforestation with mixed tree species promotes soil biodiversity, and that besides conserving mixed forests, protection of old pure stands may be particularly important for conserving partner-specific ectomycorrhizal fungi.
The advancements in sequencing technologies have promoted the generation of molecular data for cataloguing and describing biodiversity. The analysis of environmental DNA (eDNA) through the application of metabarcoding techniques enables comprehensive descriptions of communities and their function, being fundamental for understanding and preserving biodiversity. Metabarcoding is becoming widely used and standard methods are being generated for a growing range of applications with high scalability. The generated data can be made available in its unprocessed form, as raw data (the sequenced reads) or as interpreted data, including sets of sequences derived after bioinformatics processing (Amplicon Sequence Variants (ASVs) or Operational Taxonomic Units (OTUs)) and occurrence tables (tables that describe the occurrences and abundances of species or OTUs/ASVs). However, for this data to be Findable, Accessible, Interoperable and Reusable (FAIR), and therefore fully available for meaningful interpretation, it needs to be deposited in public repositories together with enriched sample metadata, protocols and analysis workflows (ten Hoopen et al. 2017). Metabarcoding raw data and associated sample metadata is often stored and made available through the International Nucleotide Sequence Database Collaboration (INSDC) archives (Arita et al. 2020), of which the European Nucleotide Archive (ENA, Burgin et al. 2022) is its European database, but it is often deposited with minimal information, which hinders data reusability. Within the scope of the Horizon 2020 project, Biodiversity Community Integrated Knowledge Library (BiCIKL), which is building a community of interconnected data for biodiversity research (Penev et al. 2022), we are working towards improving the standards for molecular ecology data sharing, developing tools to facilitate data deposition and retrieval, and linking between data types. Here we will present the ENA data model, showcasing how metabarcoding data can be shared, while providing enriched metadata, and how this data is linked with existing data in other research infrastructures in the biodiversity domain, such as the Global Biodiversity Information Facility (GBIF), where data is deposited following the guidelines published in Abarenkov et al. (2023). We will also present the results of our recent discussions on standards for this data type and discuss future plans towards continuing to improve data sharing and interoperability for molecular ecology.
This deliverable report includes description of the work steps towards building a web interface for the reporting of errors and gaps in sequenced material source annotations as part of the Task 8.3 of BiCIKL. Beta version of the web interface has been published and is available for the registered users of PlutoF platform.
SummaryFungi play pivotal roles in ecosystem functioning, but little is known about their global patterns of diversity, endemicity, vulnerability to global change drivers and conservation priority areas. We applied the high-resolution PacBio sequencing technique to identify fungi based on a long DNA marker that revealed a high proportion of hitherto unknown fungal taxa. We used a Global Soil Mycobiome consortium dataset to test relative performance of various sequencing depth standardization methods (calculation of residuals, exclusion of singletons, traditional and SRS rarefaction, use of Shannon index of diversity) to find optimal protocols for statistical analyses. Altogether, we used six global surveys to infer these patterns for soil-inhabiting fungi and their functional groups. We found that residuals of log-transformed richness (including singletons) against log-transformed sequencing depth yields significantly better model estimates compared with most other standardization methods. With respect to global patterns, fungal functional groups differed in the patterns of diversity, endemicity and vulnerability to main global change predictors. Unlike α-diversity, endemicity and global-change vulnerability of fungi and most functional groups were greatest in the tropics. Fungi are vulnerable mostly to drought, heat, and land cover change. Fungal conservation areas of highest priority include wetlands and moist tropical ecosystems.
The international DNA sequence databases abound in fungal sequences not annotated beyond the kingdom level, typically bearing names such as "uncultured fungus". These sequences beget low-resolution mycological results and invite further deposition of similarly poorly annotated entries. What do these sequences represent? This study uses a 767,918-sequence corpus of public full-length fungal ITS sequences to estimate what proportion of the 95,055 "uncultured fungus" sequences that represent truly unidentifiable fungal taxa - and what proportion of them that would have been straightforward to annotate to some more meaningful taxonomic level at the time of sequence deposition. Our results suggest that more than 70% of these sequences would have been trivial to identify to at least the order/family level at the time of sequence deposition, hinting that factors other than poor availability of relevant reference sequences explain the low-resolution names. We speculate that researchers' perceived lack of time and lack of insight into the ramifications of this problem are the main explanations for the low-resolution names. We were surprised to find that more than a fifth of these sequences seem to have been deposited by mycologists rather than researchers unfamiliar with the consequences of poorly annotated fungal sequences in molecular repositories. The proportion of these needlessly poorly annotated sequences does not decline over time, suggesting that this problem must not be left unchecked.
Bistorta vivipara is a widespread herbaceous perennial plant with a discontinuous pattern of distribution in arctic, alpine, subalpine and boreal habitats across the northern Hemisphere. Studies of the fungi associated with the roots of B. vivipara have mainly been conducted in arctic and alpine ecosystems. This study examined the fungal diversity and specificity from root tips of B. vivipara in two local mountain ecosystems as well as on a global scale. Sequences were generated by Sanger sequencing of the internal transcribed spacer (ITS) region followed by an analysis of accurately annotated nuclear segments including ITS1-5.8S-ITS2 sequences available from public databases. In total, 181 different UNITE species hypotheses (SHs) were detected to be fungi associated with B. vivipara, 73 of which occurred in the Bavarian Alps and nine in the Swabian Alps–with one SH shared among both mountains. In both sites as well as in additional public data, individuals of B. vivipara were found to contain phylogenetically diverse fungi, with the Basidiomycota, represented by the Thelephorales and Sebacinales, being the most dominant. A comparative analysis of the diversity of the Sebacinales associated with B. vivipara and other co-occurring plant genera showed that the highest number of sebacinoid SHs were associated with Quercus and Pinus, followed by Bistorta. A comparison of B. vivipara with plant families such as Ericaceae, Fagaceae, Orchidaceae, and Pinaceae showed a clear trend: Only a few species were specific to B. vivipara and a large number of SHs were shared with other co-occurring non-B. vivipara plant species. In Sebacinales, the majority of SHs associated with B. vivipara belonged to the ectomycorrhiza (ECM)-forming Sebacinaceae, with fewer SHs belonging to the Serendipitaceae encompassing diverse ericoid–orchid–ECM–endophytic associations. The large proportion of non-host-specific fungi able to form a symbiosis with other non-B. vivipara plants could suggest that the high fungal diversity in B. vivipara comes from an active recruitment of their associates from the co-occurring vegetation. The non-host-specificity suggests that this strategy may offer ecological advantages; specifically, linkages with generalist rather than specialist fungi. Proximity to co-occurring non-B. vivipara plants can maximise the fitness of B. vivipara, allowing more rapid and easy colonisation of the available habitats.