BACKGROUND:Sound collections for singing insects provide important repositories that underpin existing research (e.g. Price et al. 2007 at http://bio.acousti.ca/node/11801; Price et al. 2010) and make bioacoustic collections available for future work, including insect communication (Ordish 1992), systematics (e.g. David et al. 2003), and automated identification (Bennett et al. 2015). The BioAcoustica platform (Baker et al. 2015) is both a repository and analysis platform for bioacoustic collections: allowing collections to be available in perpetuity, and also facilitating complex analyses using the BioVeL cloud infrastructure (Vicario et al. 2011). The Global Cicada Sound Collection is a project to make recordings of the world's cicadas (Hemiptera: Cicadidae) available using open licences to maximise their potential for study and reuse. This first component of the Global Cicada Sound Collection comprises recordings made between 2006 and 2008 of Cicadidae in South Africa and Malawi. NEW INFORMATION:This collection of sounds includes 219 recordings of 133 voucher specimens, comprising 42 taxa (25 identified to species, all identified to genus) from South Africa and Malawi. The recordings have been used to underpin work on the species limits of cicadas in southern Africa, including Price et al. (2007) and Price et al. (2010). The specimens are deposited in the Albany Museum, Grahamstown, South Africa (AMGS). The harvesting of acoustic data as occurrence records by GBIF has been implemented by the Scratchpads Team at the Natural History Museum, London. This link increases the value of individual recordings and the BioAcoustica platform within the global infrastructure of biodiversity informatics by making specimen/occurence records from BioAcoustica available to a wider audience, and allowing their integration with other occurence datasets that also contribute to GBIF.
We describe an online open repository and analysis platform, BioAcoustica (http://bio.acousti.ca), for recordings of wildlife sounds. Recordings can be annotated using a crowdsourced approach, allowing voice introductions and sections with extraneous noise to be removed from analyses. This system is based on the Scratchpads virtual research environment, the BioVeL portal and the Taverna workflow management tool, which allows for analysis of recordings using a grid computing service. At present the analyses include spectrograms, oscillograms and dominant frequency analysis. Further analyses can be integrated to meet the needs of specific researchers or projects. Researchers can upload and annotate their recordings to supplement traditional publication.
We describe an implementation of the Darwin Core Archive (DwC-A) standard that allows for the exchange of biodiversity information contained within the Scratchpads virtual research environment with external collaborators. Using this single archive file Scratchpad users can expose taxonomies, specimen records, species descriptions and a range of other data to a variety of third-party aggregators and tools (currently Encyclopedia of Life, eMonocot Portal, CartoDB, and the Common Data Model) for secondary use. This paper describes our technical approach to dynamically building and validating Darwin Core Archives for the 600+ Scratchpad user communities, which can be used to serve the diverse data needs of all of our content partners.
The Scratchpad Virtual Research Environment (http://scratchpads.eu/) is a flexible system for people to create their own research networks supporting natural history science. Here we describe Version 2 of the system characterised by the move to Drupal 7 as the Scratchpad core development framework and timed to coincide with the fifth year of the project's operation in late January 2012. The development of Scratchpad 2 reflects a combination of technical enhancements that make the project more sustainable, combined with new features intended to make the system more functional and easier to use. A roadmap outlining strategic plans for development of the Scratchpad project over the next two years concludes this article.
Support systems play an important role for the communication between users and developers of software. We studied two support systems, an issues tracker and an email service available for Scratchpads, a Web 2.0 social networking tool that enables communities to build, share, manage and publish biodiversity information on the Web. Our aim was to identify co-learning opportunities between users and developers of the Scratchpad system by asking which support system was used by whom and for what type of questions. Our results show that issues tracker and emails cater to different user mentalities as well as different kind of questions and suggest ways to improve the support system as part of the development under the EU funded ViBRANT programme.
We describe a method to publish nomenclatural acts described in taxonomic websites (Scratchpads) that are formally registered through publication in a printed journal (ZooKeys). This method is fully compliant with the zoological nomenclatural code. Our approach supports manuscript creation (via a Scratchpad), electronic act registration (via ZooBank), online and print publication (in the journal ZooKeys) and simultaneous dissemination (ZooKeys and Scratchpads) for nomenclatorial acts including new species descriptions. The workflow supports the generation of manuscripts directly from a database and is illustrated by two sample papers published in the present issue.
The concept of semantic tagging and its potential for semantic enhancements to taxonomic papers is outlined and illustrated by four exemplar papers published in the present issue of ZooKeys. The four papers were created in different ways: (i) written in Microsoft Word and submitted as non-tagged manuscript (doi: 10.3897/zookeys.50.504); (ii) generated from Scratchpads and submitted as XML-tagged manuscripts (doi: 10.3897/zookeys.50.505 and doi: 10.3897/zookeys.50.506); (iii) generated from an author's database (doi: 10.3897/zookeys.50.485) and submitted as XML-tagged manuscript. XML tagging and semantic enhancements were implemented during the editorial process of ZooKeys using the Pensoft Mark Up Tool (PMT), specially designed for this purpose. The XML schema used was TaxPub, an extension to the Document Type Definitions (DTD) of the US National Library of Medicine Journal Archiving and Interchange Tag Suite (NLM). The following innovative methods of tagging, layout, publishing and disseminating the content were tested and implemented within the ZooKeys editorial workflow: (1) highly automated, fine-grained XML tagging based on TaxPub; (2) final XML output of the paper validated against the NLM DTD for archiving in PubMedCentral; (3) bibliographic metadata embedded in the PDF through XMP (Extensible Metadata Platform); (4) PDF uploaded after publication to the Biodiversity Heritage Library (BHL); (5) taxon treatments supplied through XML to Plazi; (6) semantically enhanced HTML version of the paper encompassing numerous internal and external links and linkouts, such as: (i) vizualisation of main tag elements within the text (e.g., taxon names, taxon treatments, localities, etc.); (ii) internal cross-linking between paper sections, citations, references, tables, and figures; (iii) mapping of localities listed in the whole paper or within separate taxon treatments; (v) taxon names autotagged, dynamically mapped and linked through the Pensoft Taxon Profile (PTP) to large international database services and indexers such as Global Biodiversity Information Facility (GBIF), National Center for Biotechnology Information (NCBI), Barcode of Life (BOLD), Encyclopedia of Life (EOL), ZooBank, Wikipedia, Wikispecies, Wikimedia, and others; (vi) GenBank accession numbers autotagged and linked to NCBI; (vii) external links of taxon names to references in PubMed, Google Scholar, Biodiversity Heritage Library and other sources. With the launching of the working example, ZooKeys becomes the first taxonomic journal to provide a complete XML-based editorial, publication and dissemination workflow implemented as a routine and cost-efficient practice. It is anticipated that XML-based workflow will also soon be implemented in botany through PhytoKeys, a forthcoming partner journal of ZooKeys. The semantic markup and enhancements are expected to greatly extend and accelerate the way taxonomic information is published, disseminated and used.
Metadata are, in essence, information about a resource that allows its retrieval when needed for a particular purpose, functionally equivalent to an index entry. Some resources, such as pictures, video and sound, cannot be searched for particular content so need an associated text element that describes the image and can be used to recover a specific image from amongst many. This is not, of itself, a difficult thing to do, but it does represent a significant task overhead. People building data resources for a particular purpose will include minimal metadata that is sufficient to solve the immediate task in hand and will not, generally, invest the additional time to build more extensive and more broadly useful metadata. Collaboratories Scratchpads are an intuitive web application that enables researchers collaboratively to build, share, manage and publish their biodiversity data online. The key concepts here are that many individuals contribute small items of information that are joined in a flexible architecture. By pooling these resources and by sharing in the development of the architecture, collaborative communities build up and the web site becomes the repository and the resource for further work. Tools for adding metadata Where resources are not immediately machine-readable, such as images, there are no practical alternatives to tagging 'by hand'. Tools have been developed within the Scratchpad environment that allow bulk annotation of images. A group of users define a data structure to contain the metadata, then multiple images can be selected and information common to the group (e.g. locality, expedition, species name) can be added to the relevant fields in one process. The only advantage that this offers is a reduction in the repetitive labour of tagging many pictures, which is not a significant advance in IT strategy, is a major benefit to those actually doing the work. Organisation of Information Individuals will easily create or identify resources that are relevant to a particular study domain using one of a range of tools, including personal bibliographies, Google searches, specialised databases such as EMBL, bibliometric tools, e.g. the Web of Science, and so on. From those resources, a term-list can be built which can be used to create a controlled vocabulary. The controlled vocabulary can be used to create a formal ontology. In the task of organising and recovering information the most immediately useful of these stages is the controlled vocabulary, especially if each term is mapped to a set of synonyms (a thesaurus). Each developmental step, however, requires a significant input of effort, e.g. the extraction of a term list from resources. At each stage the person doing the work has to be confident that the organised product will be of enough utility to repay the labour of it creation. Analysis of Resources Users will identify resources that are in some sense relevant to their domain of interest, as described above. We can use a computer to decipher text resources, such as published papers, and identify key structural elements within the resource. This is most effectively done by using clues not normally incorporated in conventional NLP techniques that generally discard punctuation and typographical cues to leave only the text. Hence, looking for key terms becomes far more difficult than it need be. It is part of the scientific tradition that in formal descriptive writing we use a greater proportion of latinate words that typically occur in general text. Fairly simple rules allow us to identify candidate latin words and we can use the candidates to feed a learning algorithm, especially if we have access to a dictionary describing how those terms are used. Thus we can discriminate between a text that is describing a taxon (high proportion of anatomical terms) from one that describes, for example, ecological impacts (few anatomical terms). Once terms lists start to be accumulated they can be used to lever greater meaning from a text. For example, we already have long lists of latinised species names, so we can look for those names in a text and seek to establish how they are being used: specifically if two names occur in close proximity we can ask the relationship is between them. Briefly, if we could isolate the proper nouns from a text block they would provide the who, where and what metadata, leaving us with the challenge of deducing the why. Building Searchable Resources There are enormous numbers of potentially suitable XML schemas available, but few that encompass taxonomists want recorded in metadata. The Plazi project has developed an extension to the widely used NLM schema called TaxPub. The Plazi project and PenSoft Publishers have developed an assisted workflow, not automatic but using productivity tools to reduce the time needed to process a single page to a few minutes. The use of standard XML schemas is very important because it allows the development of increasingly elegant queries. Whereas the original impetus for the Scratchpads was to mobilise taxonomic information, it quickly became apparent that there are many more uses for, for instance, occurrence data than taxonomy. The focus of our development efforts are, therefore, to extend the application domain into the environmental arena. Building engagement At the end of the day, users will engage with any system that delivers direct and clear benefit to them personally. The underlying Scratchpad database delivers organisational benefits and is vastly easier to maintain than traditional web pages. The authors benefit from increased exposure and international recognition of their expertise. As the consortium behind a particular Scratchpad grows, the underlying database becomes a richer resource that can be used to probe different types of problem. Search structures across many Scratchpads deliver the benefit of access to information originally assembled for a different purpose (taxonomy into ecology and visa versa). The EU project EDIT has demonstrated that existing technology is easily capable of delivering these benefits, but the barriers are sociological. It seems to be easier to build and retain engagement if progress comes as small incremental steps, each delivering a discrete benefit. Release of a complete, polished solution will generally represent a significant learning curve and require the user to change their work-practice in a significant way. The principles described above have been brought together with the recent release of a special issue of the journal ZooKeys. Here data were entered into a Scratchpad, then at the click of a button, rendered into an XML version that was sent to the publisher, who automatically transformed it into a PDF version that was sent to referees. The journal's editorial team only spent time on the paper when referee's comments were received and, when the papers were accepted, further marked up the content to facilitate text-mining. The papers were published only 4 weeks after that first click, in print, PDF, semantically enhanced HTML, and XML versions. The XML version is archived in PubMedCentral and portions of tagged text (e.g., taxon treatments) are automatically harvested and exported to aggregators such as EOL and Plazi.
BACKGROUND:Natural History science is characterised by a single immense goal (to document, describe and synthesise all facets pertaining to the diversity of life) that can only be addressed through a seemingly infinite series of smaller studies. The discipline's failure to meaningfully connect these small studies with natural history's goal has made it hard to demonstrate the value of natural history to a wider scientific community. Digital technologies provide the means to bridge this gap.RESULTS:We describe the system architecture and template design of "Scratchpads", a data-publishing framework for groups of people to create their own social networks supporting natural history science. Scratchpads cater to the particular needs of individual research communities through a common database and system architecture. This is flexible and scalable enough to support multiple networks, each with its own choice of features, visual design, and constituent data. Our data model supports web services on standardised data elements that might be used by related initiatives such as GBIF and the Encyclopedia of Life. A Scratchpad allows users to organise data around user-defined or imported ontologies, including biological classifications. Automated semantic annotation and indexing is applied to all content, allowing users to navigate intuitively and curate diverse biological data, including content drawn from third party resources. A system of archiving citable pages allows stable referencing with unique identifiers and provides credit to contributors through normal citation processes.CONCLUSION:Our framework http://scratchpads.eu/ currently serves more than 1,100 registered users across 100 sites, spanning academic, amateur and citizen-science audiences. These users have generated more than 130,000 nodes of content in the first two years of use. The template of our architecture may serve as a model to other research communities developing data publishing frameworks outside biodiversity research.