It is estimated that there are nearly 400 million herbarium specimens held across approximately 3,500 herbaria worldwide (Davis 2023). Over the past decade, many institutions have embarked on large-scale digitisation initiatives to robustly and efficiently digitise herbarium sheet collections. While these processes can be highly productive, reaching daily rates of 600–1,000 herbarium sheets in some programmes, there remains a risk and cost associated with re-digitisation, often caused by complications during the imaging stage. To address these challenges, the Natural History Museum, London (NHM) partnered with the Royal Botanic Garden Edinburgh (RBGE) to develop an Artificial Intelligence (AI)-based workflow aimed at providing enhanced quality control of herbarium sheet digitisation. This initiative was supported by the UK Government’s Department for Culture, Media and Sport (DCMS) through the NHM–DCMS AI Pilot Programme (Poon et al. 2026). The aim of this project was to develop an AI-assisted quality control framework (Zhang et al. 2025). Trained on collections at RBGE, the tool enables automated quality checks of digitised herbarium sheets. This is done by detecting missing objects (such as colour reference cards or barcodes), identifying cropping errors, and highlighting focus issues (Fig. 1). The pipeline integrates several computational techniques, including object detection, decision trees, classical statistics, and deep learning methods. In addition to post-digitisation analysis, a user-friendly interface was also developed to support in situ quality checks during the imaging stages. Model evaluation on an independent test set demonstrated the high accuracy and robustness of the AI-assisted framework (Zhang et al. 2025). The object detection model achieved an overall precision of 0.98, recall of 0.98, and mean Average Precision scores of 0.99 (mAP50) and 0.94 (mAP50–95) across all object categories. Key reference elements, including colour reference cards, barcodes, and rulers, were detected with near-perfect accuracy, ensuring precise object localisation. We note that as model development was done on RBGE collections alone, this meant that reference elements were similar across many herbarium sheets in the sample. Lastly, the issue detection models also performed strongly, achieving a precision of approximately 1.00 for out-of-focus detection, 0.98 for over-cropping, and 0.93 for under-cropping. The proposed workflow underwent further evaluation with herbarium digitisers at RBGE, which will inform refinements to the model, software interface, and the overall pipeline. The next step is to expand on this pilot project by testing the model with collections across different institutions, leading to more refinements for generalisability. This can also help us evaluate and estimate the total proportions of digitised herbarium sheets with issues. Finally, our global aim is to provide an AI-based quality control service that can be integrated into digitisation workflows not only at RBGE but across institutions worldwide.
The digitisation of the world’s natural science collections is expanding massively and providing a unique global resource for answering some of the most fundamental bio- and geodiversity questions. However, digitisation at this scale can only be done in stages, increasing the variation in the level of digitisation both between and within collections. The ability to measure and monitor the level of digitisation of each individual specimen and, by extension, each collection on a national or global scale has never been more important. The Minimum Information about a Digital Specimen (MIDS) standard is being developed to provide an international digitisation standard within the Biodiversity Information Standards (TDWG) organisation. The standard, along with implementations which can calculate the MIDS level of specimens and, by extension, datasets, provide users with tools to help develop a digitisation strategy as well as plan, manage and monitor a digitisation programme, including prioritisation and data enhancement. For researchers, whilst the MIDS level of published specimens does not provide information about the quality of the data present, it does indicate the expected amount of associated data for intended analyses, including geographic coordinates and identifiers. The MIDS website provides access to the current draft of the standard. The four MIDS levels (MIDS0 to MIDS3) are described, and for each level the purpose is included. The purpose has been key to defining the information that is required to be present. The information recorded for each specimen is categorised into information elements. A detailed schema for the information elements includes the label, definition, usage note, purpose and examples as well as the disciplines for which each element is required (Biology, Geology, Palaeontology). The information elements required for each MIDS level are cumulative, with each level adding additional information relevant for the purpose of the level. Specimen data recorded in collection databases and submitted to international aggregators such as the Global Biodiversity Information Facility (GBIF) need to be mapped to the information elements to enable the calculation of the MIDS level. The website explains how the Simple Standard for Sharing Ontology Mappings (SSSOM) (Matentzoglu et al. 2022) is being used to map Darwin Core (DwC) (Wieczorek et al. 2012)and Access to Biological Collection Data (ABCD) (Access to Biological Collection Data task group 2007) terms to MIDS. It provides a tabulated quick reference of MIDS mappings with a filter option to enable users to quickly review the mapping by MIDS level or by information element. Several tools have been developed, implementing MIDS to calculate the MIDS levels of datasets, and the website provides links to these. As the MIDS standard is not yet ratified, there have been updates which are not reflected in all the tools currently available. There is a reference implementation as part of the open Digital Specimen model with the Distributed System of Scientific Collections (DiSSCo) where the MIDS level is calculated for each digital specimen. The MIDS Calculator can be used to calculate the MIDS score for DwC archive and ABCD biological datasets based on the current version of MIDS. Additional functionality for geological and palaeontological datasets is being developed. As the MIDS standard approaches public review we encourage curators and collection managers to test out the functionality on their collection data and provide feedback using the MIDS GitHub repository.
Natural History institutes hold an immense number of specimens and artefacts. For years these collections were not accessible online, remaining inaccessible to researchers from far away and hidden from the general public. Large digitisation projects and cross-institutional agreements aim to bring their collections into the digital era, such as the SYNTHESYS+ project and the Distributed System of Scientific Collections (DiSSCo) Research Infrastructure. As specimens are 3D physical objects with different characteristics many techniques are available to 3D digitise them. For inexperienced users this can be quite overwhelming. Which techniques are already well tested in other institutions and are suitable for a specific specimen or collection? To investigate this, we have set up a dichotomous identification key for digitisation techniques: DIGIT-KEY, (https://digit.naturalheritage.be/digit-key). For each technique, examples used in SYNTHESYS+ Institutions are visualised and training manuals provided. All information can be easily updated and representatives can be contacted if necessary to request more information about a certain technique. This key can be helpful to achieve comparable results across institutions when digitising collections on demand in future DiSSCo research initiatives coordinated through the European Loans and Visits System (ELViS) for Virtual and Transnational Access.
Societal Impact Statement The value of herbarium specimens depends largely on the accuracy and accessibility of the data captured, which is dependent on curation practices. Previous studies have shown high levels of misidentification in collections, which become more problematic with increased access. To evaluate variations in curation practices and assess the impacts of these differences, a survey was sent to herbarium curators worldwide. The results revealed substantial variation in curation, identification and digitisation practices. This study examines how increasing support for herbaria, standardising curation and enhancing information pathways are crucial to improving the accuracy and accessibility of data essential to addressing environmental challenges. Summary This research aimed to assess the current practices and standards in place within institutions of different sizes around the world, in order to develop more effective methods in curation, specimen identification and digitisation. A detailed survey was created and sent to herbaria around the world. The survey comprised five parts: (1) general herbarium framework, (2) curation and collections, (3) specimens without determinations, (4) digitisation and (5) scenario question. The results were analysed based on collection size. The survey results showed that there are many differences among herbaria of varying sizes when comparing curation methods and standards, identification methods and tools, and digitisation efforts. This study found that herbaria require the development of consistent curation practices that adhere to global standards both within and among institutions. To maximise the value of collections, herbaria require both shared resources (tools) and individual resources (funding, time and staff) to provide accurate specimen identification and ongoing curation. Digitisation can support identification and improve access to biodiversity data, but only if these data are shared and regularly synchronised with widely accessible databases.
As the digitisation of natural science collections progresses, the need for tools and skills for the staff who curate those collections becomes increasingly important and urgent. The specimens held in these collections are fundamental to research and are a critical source of data and knowledge urgently needed for the current biodiversity crisis. Without taxonomic curation, the names attached to these specimens may be synonyms, misapplied or incomplete. One study examined specimens from 40 herbaria in 21 countries and found that half of specimens from two tropical plant families had the wrong names applied (Goodwin et al. 2015). These issues may have a significant impact on research undertaken based on these collections. Curation of natural science collections covers a huge range of activities and responsibilities. Taxonomic curation is the management and updating of the taxon names applied to the specimens in collections. Taxonomic curation can be carried out at different levels. A key aspect of taxonomic curation is providing access to accepted classifications across multiple taxonomic levels and enabling comparisons with alternative systems, including those already in use within a collection. At the specimen level, tools to access specimen citations, identifications of duplicates as well as managing online determinations are becoming more important for curation staff and volunteers. These processes require digital skills for curators. For curation staff and volunteers to develop similar skills and be trained in the use of technical tools, adequate and accessible training documentation is required. Before digitisation, collections were frequently curated using regional taxonomic accounts and it was quite common to have specimens from different regions arranged differently, resulting in multiple taxonomies through the collection. With digitisation, there is usually a single taxonomy in the database which pushes us more towards aligning all specimens to a single classification. The digitised specimen record will normally include the name under which the specimen is physically filed, and it is critical that this is always aligned to ensure findability of the specimen. This is a huge task and the combination of the growth of the collections and the lack of growth in staffing levels means that we must look at more efficient options for doing this work. Case studies in the Herbarium of the Royal Botanic Garden Edinburgh were undertaken to review the process of curation of several digitised taxa and one undigitised taxon. This allowed us to explore the impact of digitisation on curation practices. We also investigated the additional training requirements for staff to be able to use the increasing number of online tools and resources available. The case studies included the genera Plectranthus and Coleus in the Lamiaceae (Fig. 2, Streptocarpus in the Gesneriaceae (Fig. 1) and the family Lycopodiaceae. We identified 4 levels of curation: High level to delimit and organise the collection at Kingdom, Phylum and Class level; Family level to delimit and organise families and the genera within them; Genera and species level to delimit and organise genera and species; and Specimen level to identify individual specimens in the collection. The appropriate level of taxonomic curation will depend on the scope of the taxa being curated and the time and resources available. High level to delimit and organise the collection at Kingdom, Phylum and Class level; Family level to delimit and organise families and the genera within them; Genera and species level to delimit and organise genera and species; and Specimen level to identify individual specimens in the collection. The appropriate level of taxonomic curation will depend on the scope of the taxa being curated and the time and resources available. Using the case studies at the Royal Botanic Garden Edinburgh, we are developing more effective and efficient taxonomic curation workflows for digitised collections. For each workflow, we include the tools and resources available as well as consider the skills, knowledge, and training required. The initial workflow for the genera and species level comprises ten steps: Curation requirements and prioritisation; Classification review; Taxonomic scope; Review and comparison of taxonomy in literature; Review and comparison of taxonomy in the collection management system (CMS); Recuration planning and communication across the organisation; Updating of names and taxonomy in the CMS; Reorganisation and updating of specimens and folders (Fig. 2) and specimen records in the CMS; Updating of indexes; and Updating of cabinet labels. Curation requirements and prioritisation; Classification review; Taxonomic scope; Review and comparison of taxonomy in literature; Review and comparison of taxonomy in the collection management system (CMS); Recuration planning and communication across the organisation; Updating of names and taxonomy in the CMS; Reorganisation and updating of specimens and folders (Fig. 2) and specimen records in the CMS; Updating of indexes; and Updating of cabinet labels. The development of workflows has identified requirements for tools and resources, as well as requirements for skills. This approach can also standardise taxonomic curation practices and make it easier to train curation staff, giving them guidance for their work.
The MIDS standard (Minimum Information about a Digital Specimen) aims to indicate the digitisation level of natural history specimens on a simple scale. Much progress has been made over the past year to standardise this process for optimal reusability and scalability. The Simple Standard for Sharing Ontological Mappings (SSSOM) framework (Matentzoglu et al. 2022) has been adopted by the MIDS task group for this purpose and mappings are now available for Darwin Core archives and ABCD. The use of SSSOM mapping sets should facilitate the assessment of digitisation status over time, for different kinds of specimens and for different data standards and their serialisation. In this presentation, we will show the results of calculating MIDS levels for those different use cases when applied to several datasets acquired from the Global Biodiversity Information Facility (GBIF), using the MIDSCalculator and the GBIF MIDS Checker tools. We will determine the reasons why MIDS levels were (not) met, highlighting the impact of MIDS information element definitions and, in particular, their mappings to different standards and the stability of their implementation in different kinds of software. These results should show the feasibility of enabling MIDS across different infrastructures and promote implementation of the MIDS standard. To acquire an overview of the current digitisation status of specimen data, we downloaded an export of specimens from GBIF (GBIF.org 2025f) using its new SQL API service to minimize data storage requirements for unrequired data. To calculate the MIDS scores for this large dataset - 265 million records - we used an adapted script re-using the MIDSCalculator calculation scripts in batches of 10 million (omitting the RShiny interface). We subsequently summarised the results per MIDS information element and level and performed follow-up analyses, again using the GBIF SQL API, to look into the causes of some MIDS information elements often being absent in the GBIF data. Fig. 1 shows the overall results per information element. 4% failed to meet MIDS level 0 and would be considered undigitised. 69% met MIDS level 0, 25% level 1 and 2% level 2. No specimen met level 3, but this was because Darwin Core terms mapped to MediaID are only available in multimedia extensions to Darwin Core and were not available through the SQL API, so achieving this level was impossible through this method. Significant impacts were observed for several information elements: An unmapped Organization caused ca. 10 million specimens to remain at the level of "undigitized" (technically MIDS level -1) (GBIF.org 2025a). More than half of the datasets with specimens seemed to omit a last Modified date, preventing the specimens in them from reaching MIDS level 1 (GBIF.org 2025b). Object Type was even more impactful, missing from 173 million specimens or almost 2 out of every 3 specimens (GBIF.org 2025e). Collecting Number was missing for more than half of the specimens, but this one is difficult to assess properly in the absence of a reliable vocabulary for unknown, blank and non-applicable values (GBIF.org 2025g, GBIF.org 2025d, GBIF.org 2025c). This number is very important for a significant amount of specimens, as it is often used to refer to them in literature, but many specimens do not use these numbers at all. These results show that calculating MIDS even at larger scales is possible and insightful. The technical workflow used could still be optimized in several ways, including parallel computing, more efficient storage methods and more powerful software frameworks than R, and potentially be leveraged to show the evolution of MIDS values, and hence digitization, through time using past GBIF snapshots.
The present corrigendum corrects errors that occurred in Brecko J. et al. (2025). https://doi.org/10.5852/ejt.2025.976.2797
The need to rapidly digitise millions of specimens in natural history collections has seen the development of a staged approach for data capture by institutions that have been carrying out mass digitisation projects. The cost and difficulty of transcribing labels associated with specimens has necessitated an approach that prioritises parts of the data based on research and curatorial requirements as well as ease of transcription and methodology available. This approach means that there is currently a wide variation in the completeness of data capture and imaging of specimens within and between collections. The Biodiversity Information Standards (TDWG) Minimum Information about a Digital Specimen (MIDS) specification has been designed to provide a framework for measuring and monitoring the completeness of the data transcribed from the specimen and the presence of an image or other media file. This framework enables a consistent and standardised approach to the calculation of the level of digitisation of specimens within a collection and, by extension/extrapolation, to the collection itself. Four levels of digitisation have been defined in the MIDS specification including a pre-digitisation level. These have been developed to correspond to practical requirements and existing large-scale digitisation programmes. MIDS level 0 (Bare): A bare or skeletal record making the association between an identifier of a physical specimen and its digital representation, allowing for unambiguous attachment of all other information. MIDS level 1 (Basic): A basic record of specimen information enabling basic discoverability as well as how the user is permitted to use the data. MIDS level 2 (Regular): A regular level of information including data that have been agreed over time as essential for most scientific purposes. MIDS level 3 (Extended): An extended level of information about a specimen including identifiers enabling connections to be made to other data present or known about the specimen. The scope of MIDS does not directly include the quality of data entered. Instead, MIDS can be seen as a tool to support the continuing improvement of data access and quality. It provides guidance for prioritising the data to be captured as well as recommendations for data standards and mapping structures. Here we present the progress in the development of the MIDS specification and its implementation. The MIDS elements for each of the four levels have now been defined and a draft version is available here. A machine-readable mapping to Darwin Core and ABCD has been compiled following the Simple Standard for Sharing Ontological Mappings (SSSOM). The differing data relevances and priorities for different disciplines have been recognised and the inclusion of required elements at each level is discipline specific (biology, palaeontology and geology). In addition, MIDS level scores have been increasingly implemented in existing specimen management systems, allowing an evaluation of the levels achieved in various contexts. There are two implementations of MIDS calculators that have been developed and are available to test. One, developed at the Meise Botanic Garden, uses a zipped Global Biodiversity Information Facility (GBIF) annotated archive (a commonly used GBIF download format based on Darwin Core Archive) and is available here. The other, developed at The Natural History Museum London, interacts directly with the GBIF API and is available here. We encourage users to try these and provide feedback to the TDWG MIDS Task Group.
Caesalpinioideae is the second largest subfamily of legumes (Leguminosae) with ca. 4680 species and 163 genera. It is an ecologically and economically important group formed of mostly woody perennials that range from large canopy emergent trees to functionally herbaceous geoxyles, lianas and shrubs, and which has a global distribution, occurring on every continent except Antarctica. Following the recent re -circumscription of 15 Caesalpinioideae genera as presented in Advances in Legume Systematics 14, Part 1, and using as a basis a phylogenomic analysis of 997 nuclear gene sequences for 420 species and all but five of the genera currently recognised in the subfamily, we present a new higher -level classification for the subfamily. The new classification of Caesalpinioideae comprises eleven tribes, all of which are either new, reinstated or re -circumscribed at this rank: Caesalpinieae Rchb. (27 genera / ca. 223 species), Campsiandreae LPWG (2 / 5-22), Cassieae Bronn (7 / 695), Ceratonieae Rchb. (4 / 6), Dimorphandreae Benth. (4 / 35), Erythrophleeae LPWG (2 /13), Gleditsieae Nakai (3 / 20), Mimoseae Bronn (100 / ca. 3510), Pterogyneae LPWG (1 / 1), Schizolobieae Nakai (8 / 42-43), Sclerolobieae Benth. & Hook. f. (5 / ca. 113). Although many of these lineages have been recognised and named in the past, either as tribes or informal generic groups, their circumscriptions have varied widely and changed over the past decades, such that all the tribes described here differ in generic membership from those previously recognised. Importantly, the approximately 3500 species and 100 genera of the former subfamily Mimosoideae are now placed in the reinstated, but newly circumscribed, tribe Mimoseae. Because of the large size and ecological importance of the tribe, we also provide a clade-based classification system for Mimoseae that includes 17 named lower -level clades. Fourteen of the 100 Mimoseae genera remain unplaced in these lower -level clades: eight are resolved in two grades and six are phylogenetically isolated monogeneric lineages. In addition to the new classification, we provide a key to genera, morphological descriptions and notes for all 163 genera, all tribes, and all named clades. The diversity of growth forms, foliage, flowers and fruits are illustrated for all genera, and for each genus we also provide a distribution map, based on quality-controlled herbarium specimen localities. A glossary for specialised terms used in legume morphology is provided. This new phylogenetically based classification of Caesalpinioideae provides a solid system for communication and a framework for downstream analyses of biogeography, trait evolution and diversification, as well as for taxonomic revision of still understudied genera.
The Royal Botanic Garden Edinburgh Herbarium (RGBE) uses Specify 7 as a Collections Management System (CMS) for managing the collections of herbarium sheets, cryptogam packets, carpological collections, and liquid-preserved and silica-dried material. It is also used to manage the metadata for transactions including accessioning incoming specimens, loans (in and out) and destructive sampling. RBGE migrated to Specify in April 2022. This was undertaken using in-house migration scripts and with the help of the Specify support team. Alongside the migration, the RBGE herbarium team has worked on embedding existing and new workflows, including incorporating permit documentation and restrictions, cultivated material and data exports to aggregators. As part of the migration, we chose to adopt a new model for how our specimens are represented, moving from a CMS where we had one record for each object in the collection, to one where the different preparations (e.g. herbarium sheet, carpological, fluid-preserved) of a collection are all held as part of a single collection object. This has enabled our users to easily see all the preparations of a single collection, including sampling records. This gives the curators a fuller view of our holdings and reduces duplicated effort for data entry. The move to Specify and the new data model for specimens has enabled us to better manage incoming specimens by utilising the Accessions and Permits tables. The restrictions from the permits are recorded as part of the accession record, using two basic categories to allow the curatorial team to quickly determine whether there are any restrictions on the material that has been requested. When the decision was made to migrate to a new CMS, our living collections were also moved, meaning the collections are now managed in two separate CMSs. This has been challenging for the management of our collection of cultivated herbarium specimens, which can be more complicated than a standard herbarium specimen. To help us manage the cultivated specimens, we have the concept of two separate collecting events. Firstly, there can be a specimen collected in the wild. When this collection was made, material could have been collected for cultivation in our living collection. Secondly, we can have a collection made from our living collection. This specimen has its own collecting event, but it is also linked to the specimen collected in the wild. We are creating links between these two separate collection objects to maintain this relationship, along with providing links to our living collection. We have aimed to integrate mass digitisation with curation and have worked with an external contractor to develop tools for large-scale cataloguing, which work directly with Specify 7. Our Minimal Data entry tool allows our digitisation team to rapidly create “stub” records at the MIDS (Minimal Information about a Digital Specimen) Level 1, enabling the cataloguing of the collection ahead of imaging. More recently, we have developed a tool to work alongside Specify, to allow for batch determinations of specimens to support large-scale record updates following research visits, loan returns and re-curation of the collection. Users are able to select the kind of determination, the taxon of the determination, and other optional information. They can then scan the barcodes of the specimens for which this new determination is to be applied, and the relevant records are updated. This has sped up the process of adding new determinations to large numbers of specimens. By moving to Specify 7, we have benefited from a cloud-based solution, which has increased availability of our CMS to both remote and on-site staff. By choosing to have our cloud instance managed by the Specify team, we have lowered maintenance overhead, provided better back-up provision and increased system upgrade frequency. We have also benefitted from an emerging small, but active, community of herbarium-based Specify users.
The Herbarium at the Royal Botanic Garden Edinburgh (RBGE) have run Frankenstein's Plants as part of the Edinburgh Science Festival, which runs during the Easter holidays. Frankenstein's Plants aims to engage both children and the wider public with the work of herbaria, highlighting the Herbarium at RBGE. We first took part in 2019, and continued in 2023 and 2024. The Royal Botanic Garden Edinburgh is a well-known and loved location for visitors and local residents, attracting large numbers of people throughout the year. However, the work of the Herbarium is less well-known and many visitors, including residents, are unaware of the Herbarium collection. The aims of the event include: Raising awareness of the herbarium at RGBE, and how it is used by scientists globally Educating people on how specimens are made Introducing the idea of a scientific name and how it is constructed Raising awareness of the herbarium at RGBE, and how it is used by scientists globally Educating people on how specimens are made Introducing the idea of a scientific name and how it is constructed Frankenstein's Plants takes participants through the 'life' of a specimen, starting with the selection of material for mounting, through to the digitisation of the specimen, contributing to a virtual herbarium. Participants are encouraged to let their imagination go wild and create a monster (or whatever else they feel like) using the material provided. The event is laid out with a series of stations where staff talk through and support each step of the process (Fig. 1). The key steps of specimen creation for this event are: Select plant material: pressed and dried prior to the event, using flowers and foliage bought from an online florist, alongside material gathered from the gardens at RBGE. Mount the specimen: a pre-printed label is attached to a piece of board. The boards we use are approximately A4, allowing enough room for participants to create their creatures. Gummed tape is used to fix the plant material to the sheets, as a relatively low-mess option. The label provides space for recording the species name, a description and ‘collector’ information. Some locality information and a barcode is prefilled. The details on the label aim to give the participant an idea of the types of data that would typically be recorded when collecting. Name the specimen: participants can create their own name for their specimen. A list of options for the genus and species is provided and consists of both real and made-up genera and species epithets. This is an opportunity for the team to talk about how plants are named using the binomial system, and the importance of Latin names for communicating about life on earth. Describe the specimen: basic botanical terms are provided alongside sketches to get the participants thinking about how species can be described. This step can be modified based on the age of the participant, bringing in more information and technical terms where appropriate. Stamp the specimen: whilst not a key part of the processing of a specimen, it is very much enjoyed by the participants! We have a mix of old stamps that were previously used on specimens that they can choose from. Digitise the specimen: the final step is to take a picture of the specimen. We have a copy stand and DSLR (digital single-lens reflex) camera to allow the participant to take a photo of their specimen. The participant scans the barcode, which we use for the file name and later to look up the newly-digitised specimen in the virtual herbarium. They then take the photo, using remote shutter release triggered by a fixed mouse. The laptop is connected to a big screen, so everyone can see the image of their specimen. Select plant material: pressed and dried prior to the event, using flowers and foliage bought from an online florist, alongside material gathered from the gardens at RBGE. Mount the specimen: a pre-printed label is attached to a piece of board. The boards we use are approximately A4, allowing enough room for participants to create their creatures. Gummed tape is used to fix the plant material to the sheets, as a relatively low-mess option. The label provides space for recording the species name, a description and ‘collector’ information. Some locality information and a barcode is prefilled. The details on the label aim to give the participant an idea of the types of data that would typically be recorded when collecting. Name the specimen: participants can create their own name for their specimen. A list of options for the genus and species is provided and consists of both real and made-up genera and species epithets. This is an opportunity for the team to talk about how plants are named using the binomial system, and the importance of Latin names for communicating about life on earth. Describe the specimen: basic botanical terms are provided alongside sketches to get the participants thinking about how species can be described. This step can be modified based on the age of the participant, bringing in more information and technical terms where appropriate. Stamp the specimen: whilst not a key part of the processing of a specimen, it is very much enjoyed by the participants! We have a mix of old stamps that were previously used on specimens that they can choose from. Digitise the specimen: the final step is to take a picture of the specimen. We have a copy stand and DSLR (digital single-lens reflex) camera to allow the participant to take a photo of their specimen. The participant scans the barcode, which we use for the file name and later to look up the newly-digitised specimen in the virtual herbarium. They then take the photo, using remote shutter release triggered by a fixed mouse. The laptop is connected to a big screen, so everyone can see the image of their specimen. Following the event all the images are uploaded to a Flickr gallery, which acts as our virtual herbarium. A link to this gallery is provided on the label, allowing participants to look up their specimen online, as well as providing an approximate count of participants at the event. We have members of staff and volunteers from the PhD and MSc cohorts at the RBGE to talk through each step, as well as providing information about the work of the herbarium and Science Division in general. This is a great opportunity for us to engage with the public about our work, especially as the herbarium team has limited regular engagement with the wider public. Alongside the event we have a display about the Herbarium, including teaching specimens. The display provides information about the collections, without needing a team member to explain it, although we aim to have a member of staff available. Often, we find it is the parents/guardians of the children who are most interested in this, so it allows us to extend the group of people with whom we are engaging. Over the three years we have run the event, we have engaged with over 1,000 'collectors', ranging in age from 3-95 (Fig. 2), comprised of people from Edinburgh and visitors to the garden from farther afield. This figure does not include those additional members of the public who attend the event, as e.g. the parent/guardian of a child, but do not engage in creating a specimen themselves. If we include an estimate of these additional members of the public, the number of people engaged is nearer to 2,000–3,000. Feedback, both during the event and post-event, has been exceptional. We have had comments on how engaging the event is, as well as requests to take pictures so that others can run a similar event themselves.
Expeditions and other collecting events are a major source of objects in natural history museums (e.g., Mesibov 2021). Historically, these trips were often transdisciplinary: biological and Earth science specimens were collected at the same time as ethnological or anthropological objects. As a result, specimens and other material gathered during the same expedition, as well as the related data and metadata, are often distributed across multiple institutions. Many expeditions were driven by colonial agendas, aiming to discover new resources to exploit, and their findings were seldom shared with the source countries and local people. Understanding these expeditions illuminates the colonial origins of museum collections, and contributes to recognizing and addressing their impacts (e.g., Das and Lowe 2018, Ashby and Machin 2021). Research expeditions continue to contribute to natural history collections. There is a need to link historical or contemporary research expeditions to other entities, requiring the unambiguous labelling (and persistent identifiers) of such events. Stable identifiers for expeditions plus the sharing of metadata and descriptions in a wide range of languages will facilitate access to scattered information about the event, the institutions housing specimens and objects, the participants and the locations visited, and assist with the linking of distributed material and related research data. However, structured data for scientific expeditions are currently lacking. While identifier systems have been created for many entities over the last few decades, there is no dedicated identifier for research expeditions and similar events. Several studies have shown the importance of people identifiers for linking collection data (e.g., Groom et al. 2022), and we argue the same is true for expeditions. Wikidata is a multilingual community-curated knowledge base containing data structured in a human- and machine-readable format. It allows easy creation, updating and enriching of items on expeditions, and provides stable identifiers for them that can be used in collection management systems. Expeditions can be linked to participants and other agents, regions, localities, objects, archival material, maps, publications, field notebooks, documentary footage and art works resulting from the expeditions, thus making historical information more easily accessible and assisting with the acknowledgment of any imperial or colonial impact that may have resulted from the expedition. Expeditions in Wikidata can be hierarchical, e.g., linking a series of related events or under an umbrella project together providing a machine-readable way to harvest all project data. Wikidata also can provide links between present day countries and historical names for locations (e.g., former colonial names). Expeditions published as Linked Open Data make datasets more FAIR (Findable, Accessible, Interoperable, Reusable), and are also useful in data transcription and validation processes. Visualisation of itinerary data and travel routes also facilitate data quality checks. An informal working group of people interested in the topic was formed to discuss standards and share best practices and recommendations regarding terminology, data modelling and contextualisation. Building upon previous work (e.g., Bauer et al. 2022, von Mering et al. 2022, Leachman 2023), we aim to work towards the enrichment, linking and standardisation of data about research expeditions. If the Wikidata identifiers of these expeditions and participants are added to the records of the corresponding entities in the collection management system, institutions can link from their own collection metadata to the relations made in Wikidata, including to collections in other institutions. The participants of the expedition can be further linked to specimens gathered during the expedition with the use of tools, such as Bionomia, which can facilitate data round-tripping between these collections and specimen records, the Global Biodiversity Information Facility (GBIF) and Wikidata (Shorthouse 2020). Other initiatives such as the Distributed System of Scientific Collections (DiSSCo) are also interested in incorporating these identifiers as links and annotations.
There are approximately 1.5 billion specimens kept in European Natural History Collections. The mission for the Distributed System of Scientific Collections (DiSSCo) is to unite all these specimens into a one-stop e-science infrastructure of digital specimens. This is a monumental digitisation task and criteria for how to prioritise this effort are, therefore, crucial for the success of the project. In this report, we have reviewed the literature and designed and conducted surveys of the digitisation plans and criteria used by DiSSCo Partners to understand the prioritisation criteria used in the digitisation of natural history collections. As an attempt to provide some guidance for the digitisation of specimens, we suggest that an organisation (e.g. DiSSCo or an individual institution) that is planning to digitise natural history collections considers four categories of prioritisation criteria: Relevance, Data quality, Cost and Feasibility.
The Minimum Information about a Digital Specimen (MIDS) standard is being developed within Biodiversity Information Standards (TDWG) to provide a framework for organisations, communities and infrastructures to define, measure, monitor and prioritise the digitisation of specimen data to achieve increased accessibility and scientific use. MIDS levels indicate different levels of completeness in digitisation and range from Level 0: not yet meeting minimal required information needs for scientific use to Level 3: fulfilling the requirements for Digital Extended Specimens (Hardisty et al. 2022) by inclusion of persistent identifiers (PIDs) that connect the specimen with derived and related data. MIDS Levels 0–2 are generic for all specimens. From MIDS Level 2 onwards we make a distinction between biological, geological and palaeontological specimens. While MIDS represents a minimum specification, defining and publishing more extensive sets of information elements (extensions) is readily feasible and explicitly recommended. The MIDS level of a digital specimen can be calculated based on the availability of certain information elements. The MIDS standard applies to published data. The ability to map from, to and between TDWG standards is key to being able to measure the MIDS level of the digitised specimen(s). Each MIDS term is being mapped across TDWG standards involving Darwin Core (DwC), the Access to Biological Collections Data (ABCD) Schema and Latimer Core (LtC, Woodburn et al. 2022), using mapping properties provided by the Simple Knowledge Organization System (SKOS) ontology. In this presentation, we will show selected case studies that demonstrate the implementation of the MIDS standard supplemented by MIDS mappings to ABCD, to LtC, and to the Distributed System of Scientific Collections' (DISSCo) Open Digital Specimen specification. The studies show the mapping exercise in practice, with the aim of enabling fully automated and accurate calculations. To provide a reliable indicator for the level of digitisation completeness, it is important that calculations are done consistently in all implementations.
The digitisation standard, MIDS (Minimum Information about a Digital Specimen), has been developed to provide a clear definition for the level of digitisation of the specimens held in natural science collections. The standard comprises three levels of digitisation: Level 1 (Basic) is equivalent to a virtual cabinet or drawer in the collection; Level 2 (Regular) provides information needed for most research; Level 3 (Extended) provides all information on the specimen itself as well as links to associated resources. Level 1 (Basic) is equivalent to a virtual cabinet or drawer in the collection; Level 2 (Regular) provides information needed for most research; Level 3 (Extended) provides all information on the specimen itself as well as links to associated resources. A pre-digitisation Level 0 (Bare) allows for the citation of the specimen and the attachment of additional information. The standard has been developed through the collaboration of people representing digitisation, curation, data management, software development and research. The development has been coordinated within the TDWG MIDS Task Group. The aim has been to develop a standard that will fulfill the needs of those managing digitisation programmes (e.g., estimating costs, priorities) as well as the users of the digitised specimens (e.g., informing potential data users of a collection’s data richness).