The digitisation of natural history collections at scale raises a range of logistical and curatorial challenges. One concern is the physical expansion of storage infrastructure. As part of the Distributed System of Scientific Collections UK (Smith et al. 2022), the Natural History Museum London (NHM) will rehouse a significant proportion of its 36 million pinned entomological specimens (~130,000 drawers). Digitisation involves imaging every specimen and attaching a physical, unique, human- and machine-readable identifier (a small card label with a Data Matrix barcode). These barcodes must be readable from above, and their addition during digitisation can increase drawer occupancy if they exceed the existing footprint. Manual estimation of the footprint increase is impractical at this scale, given the range of circumstances associated with different specimens and drawers. To address this, the NHM has developed an AI-driven approach to estimate specimen drawer expansion and support resource planning. The deep learning pipeline can automatically detect drawer objects, referred to here as classes, including specimens, labels, barcodes, notes, unit trays, and drawers, from high-resolution images. Our dataset includes 11,090 digitised pinned Coleoptera drawers from the Index Lot collection (Natural History Museum 2014), representative of typical entomological collections. The AI pipeline supports both "pre-" and "post-" digitisation analysis by calculating the specimen bounding areas, handling overlaps, and estimating the net footprint change. Our object detection models achieved high performance under real-world conditions, with a mean average precision (mAP*1) of 85.45% across all classes. Barcode detection reached 86.63% mAP, while the standard unit tray detection and unit tray type classification model achieved 99.5% accuracy. (An example of the detection outputs is shown in Fig. 1) Three different calculation methods were tested and evaluated to estimate the drawer expansion area, corresponding to three specimen expansion rates for specimens with Data Matrix barcodes attached. Rate 1 (Per-Drawer Area via Polygon Masks*2): calculates expansion rates per drawer by measuring the total area occupied by specimens and their labels, ensuring overlaps are counted only once using polygon masks, as shown in Fig. 2. Rate 2 (Total-Based Expansion Rate): calculates expansion rates by averaging the occupancy areas of barcode-attached specimens across the entire dataset. Rate 3 (Per-Specimen Expansion Rate): averages the expansion rate per specimen using only bounding box*3 areas. Every specimen contributes equally, regardless of its size, to the final value. Rate 1 (Per-Drawer Area via Polygon Masks*2): calculates expansion rates per drawer by measuring the total area occupied by specimens and their labels, ensuring overlaps are counted only once using polygon masks, as shown in Fig. 2. Rate 2 (Total-Based Expansion Rate): calculates expansion rates by averaging the occupancy areas of barcode-attached specimens across the entire dataset. Rate 3 (Per-Specimen Expansion Rate): averages the expansion rate per specimen using only bounding box*3 areas. Every specimen contributes equally, regardless of its size, to the final value. Table 1 compares different expansion rates. Both per-drawer Rate 1 and per-specimen Rate 3 have confidence intervals. Fig. 3 and Fig. 4 show their distributions: Rate 1 is right-skewed, indicating modest area increases from overlap, while Rate 3 is bimodal*4, with a small peak near 0 (large specimens unchanged) and another around 0.3 (barcode additions on smaller specimens). The global total-based Rate 2 is the mean of the ratios without confidence intervals, providing a macro-level view. In addition to estimating drawer expansion and supporting budget planning, this tool has been designed in a modular fashion, allowing individual components to be used independently at different stages of the digitisation and curation workflow. For example, specific models can be used to flag missing labels or barcodes during digitisation, assist in tracking specimen relocation, and support downstream re-curation decisions. Importantly, the pipeline is designed to be reusable across other entomological collections, with all code to be made openly available alongside a forthcoming publication, promoting scalability within the community. Our findings demonstrate that automated spatial analysis not only improves accuracy and speed in collection management but also lays the groundwork for predictive infrastructure modelling across large-scale digitisation efforts.
Sir Joseph Banks is remembered for being a long-standing president of the Royal Society, the unofficial first director of Kew gardens and the pioneering naturalist on Captain James Cook’s great voyage onboard the Endeavour, to observe the transit of Venus and search for an undiscovered southern continent (British Museum (Natural History) 1906). Much of Bank’s life is well documented but his surviving entomology collection has never been accurately catalogued. The Banks Collection at the Natural History Museum (NHM) London is an historic assemblage of insect specimens (Fig. 1, British Museum (Natural History) 1906). It includes specimens collected by Banks and others acquired through a world-wide network of collectors. During his lifetime, Banks shared specimens with his associates and gave many specimens to Dr. William Hunter and Johan Christian Fabricius. After his death, the remaining collection was donated to the Linnean Society and later passed to the British Museum in 1863. The Banks Collection has both historical and cultural value and continues to be a relevant research tool. This is largely because Fabricius, a student of Carl Linnaeus, described many new species from the collection. Consequently, the collection contains many taxonomically important type specimens. The number of specimens in the collection is unknown but estimated to be approximately 4000. The NHM is digitising the collection with the generous support of the Charles Hayward Foundation. The collection is housed in 55 entomological glass-topped drawers, albeit not the original drawers. When conserving historical material, there is an argument to leave everything in its original state, but after much consideration, the curatorial team decided—for the long-term preservation of the collection—the specimens should be rehoused into plastazote®-lined drawers keeping the original layout of the specimens and replacing the current cork-lined drawers as part of the digitisation process. High-resolution images are taken of every specimen with its associated labels. The information on the labels is recorded for each specimen and a barcode added. As of August 2024, 3,300 specimens have been digitised and 30 of the 55 drawers recurated. More than 5000 high-resolution images have been taken (Fig. 2) and 438 labels have been transcribed for almost 1200 specimens. This project is using a new Biodiversity Information Standards (TDWG) data standard, Latimer Core, designed to support the representation and discovery of natural science collections (Woodburn et al. 2022). Latimer Core is intended to be complimentary to specimen-level standards such as Darwin Core (Darwin Core Task Group 2009), providing a way to structure and share higher-level information about groups of collection objects, from whole-museum collections through thematic and historic collections, to the contents of a single drawer. This is useful for collections with lots of associated data, fragile specimens or when displacement and disassociation of information is a concern. Unlike modern specimens, with data labels containing information on collecting event and associated persons, specimens in the Banks collection have few original labels. It was common practice that labels were temporary data storage and disposed of once the data was published. A replacement label was provided after publication (Fig. 2). The publication provides species description, geographic origin and often the name of the collection containing the specimen. Additional labels were subsequently added by curators and researchers. Many specimens have no data labels (Fig. 3). Unravelling important information about individual specimens, species and groups of specimens in this collection, including taxonomic, type status and origin will be supported by Latimer Core, as it allows a more holistic approach compared to only using specimen-level data. We can better incorporate data derived from various sources such as published data, illustrations and personal correspondence to learn about the individual and groups of specimens, the reverse of the usual workflow. Digitising this collection will improve access through digital records (specimen/drawer images, transcribed labels, publication references) reducing the need for physical examination and risk of damage. Whole-drawer digitisation in addition to specimen-level imaging provides information on the organisation and display of a collection. A collection of this nature demands minimal handling and the best storage and collections management procedures to ensure its survival for future generations. However, its significance commands a continued interest by a wide and varied audience. By digitising this collection, we will improve its physical housing, increase its accessibility without compromising the specimens for the future, and support the publication of a comprehensive catalogue of the collection.
IntroductionHistoric museum collections hold a wealth of biodiversity data that are essential to our understanding of the rapidly changing natural world. Novel curatorial practices are needed to extract and digitise these data, especially for the innumerable pinned insects whose collecting information is held on small labels.MethodsWe piloted semi-automated specimen imaging and digitisation of specimen labels for a collection of ~29,000 pinned insects of ground beetles (Carabidae: Lebiinae) held at the Natural History Museum, London. Raw transcription data were curated against literature sources and non-digital collection records. The primary data were subjected to statistical analyses to infer trends in collection activities and descriptive taxonomy over the past two centuries.ResultsThis work produced research-ready digitised records for 2,546 species (40% of known species of Lebiinae). Label information was available on geography in 91% of identified specimens, and the time of collection in 39.8% of specimens and could be approximated for nearly all specimens. Label data revealed the great age of this collection (average age 91.4 years) and the peak period of specimen acquisition between 1880 and 1930, with little differences among continents. Specimen acquisition declined greatly after about 1950. Early detected species generally were present in numerous specimens but were missing records from recent decades, while more recently acquired species (after 1950) were represented mostly by singleton specimens only. The slowing collection growth was mirrored by the decreasing rate of species description, which was affected by huge time lags of several decades to formal description after the initial specimen acquisition.DiscussionHistoric label information provides a unique resource for assessing the state of biodiversity backwards to pre-industrial times. Many species held in historical collections especially from tropical super-diverse areas may not be discovered ever again, and if they do, their recognition requires access to digital resources and more complete levels of species description. A final challenge is to link the historical specimens to contemporary collections that are mostly conducted with mechanical trapping of specimens and DNA-based species recognition.
The primary and secondary types as well as some non-type material donated by Jason Londt (and various collaborators) to the Natural History Museum, London (NHMUK) have been examined and databased which comprises 35 holotypes, 293 paratypes, and 18 non-types (added for completeness), a total of 328 type specimens from 103 species (6% of the total Afro-tropical fauna). All specimen labels were imaged, both frontal and reverse sides, alongside the specimen. Notes were made of any dissections or damage to the specimens. Additional notes were made of any differences between the labels from the species descrip-tions and the actual specimens.
The Natural History Museum, London (NHMUK) has embarked on an ambitious programme to digitise its collections. The first phase of this programme was to undertake a series of pilot projects to develop the workflows and infrastructure needed to support mass digitisation of very large scientific collections. This paper presents the results of one of the pilot projects – iCollections. This project digitised all the lepidopteran specimens usually considered as butterflies, 181,545 specimens representing 89 species from the British Isles and Ireland. The data digitised includes, species name, georeferenced location, collector and collection date - the what, where, who and when of specimen data. In addition, a digital image of each specimen was taken. A previous paper explained the way the data were obtained and the background to the collections that made up the project. The present paper describes the technical, logistical, and economic aspects of managing the project.
BACKGROUND:The Natural History Museum, London (NHMUK) has embarked on an ambitious programme to digitise its collections . The first phase of this programme has been to undertake a series of pilot projects that will develop the necessary workflows and infrastructure development needed to support mass digitisation of very large scientific collections. This paper presents the results of one of the pilot projects - iCollections. This project digitised all the lepidopteran specimens usually considered as butterflies, 181,545 specimens representing 89 species from the British Isles and Ireland. The data digitised includes, species name, georeferenced location, collector and collection date - the what, where, who and when of specimen data. In addition, a digital image of each specimen was taken. This paper explains the way the data were obtained and the background to the collections which made up the project.NEW INFORMATION:Specimen-level data associated with British and Irish butterfly specimens have not been available before and the iCollections project has released this valuable resource through the NHM data portal.