The Generalist Repository Ecosystem Initiative (GREI), funded by the NIH, developed an AI taxonomy tailored to data repository roles to guide AI integration across repository management. It categorizes the roles into stages, including acquisition, validation, organization, enhancement, analysis, sharing, and user support, providing a structured framework for implementing AI in repository workflows.
Data preservation has gained momentum and visibility in connection with the growth in digital data and data sharing policies. The Dryad Repository, a curated general–purpose repository for preserving and sharing the data underlying scientific publications, has taken steps to develop a preservation policy to ensure the long–term persistence of this archived data. In 2013, a Preservation Working Group, consisting of Dryad staff and national and international experts in data management and preservation, was convened to guide the development of a preservation policy. This paper describes the policy development process, outcomes, and lessons learned in the process. To meet Dryad’s specific needs, Dryad’s preservation policy negotiates between the ideal and the realistic, including complying with broader governing policies, matching current practices, and working within system constraints.
Data for the 10,108 files in Dryad (datadryad.org) deposited from inception through 20-Sep-2013 for those journals in which the authors had the option of selecting an embargo. See the ReadMe file for column definitions and summary stats.
Editor's Summary HIVE (Helping Interdisciplinary Vocabulary Engineering) is an effort to automatically generate metadata for content, drawing descriptor terms from multiple vocabularies encoded as Simple Knowledge Organization Systems (SKOS). The effort is a response to the challenges of interoperability, cost and usability of multiple terminology sets often needed to adequately describe digital resources. By offering access to more than one vocabulary with useful descriptors for a broad domain, HIVE enables aggregating the best terms to describe resources and automatically apply metadata. HIVE offers knowledge management value for multidisciplinary digital collections while demonstrating the expanded potential use of SKOS. The initiative is headed by the Metadata Research Center at the University of North Carolina's School of Information and Library Science working with several institutional partners. Conferences and workshops are scheduled to inform interested developers and users, who are invited to try out, contribute to and evaluate the system.
Synthetic science promises an unparalleled ability to find new meaning in old data, extant results, or previously unconnected methods and concepts, but pursuing synthesis can be a difficult and risky endeavor. Our experience as biologists, informaticians, and educators at the National Evolutionary Synthesis Center has affirmed that synthesis can yield major insights, but also revealed that technological hurdles, prevailing academic culture, and general confusion about the nature of synthesis can hamper its progress. By presenting our view of what synthesis is, why it will continue to drive progress in evolutionary biology, and how to remove barriers to its progress, we provide a map to a future in which all scientists can engage productively in synthetic research.
Designing and Implementing a Learning Object Repository: Issues of Complexity, Granularity, and User Sense-Making / William E. Moen -- Using DSpace as a Disciplinary Data Repository / Ryan Scherle -- The Sustainability, Preservation and Accessibility of Internal and External Communities by Universities / Jeffrey Trimble -- Introducing Vireo: an ETD Submittal and Management System for DSpace / Adam Mikeal.
Digital data repositories ought to support immediate operational needs and long-term project goals. This paper presents the Dryad repository's metadata best practice balancing of these two needs. The paper reviews background work exploring the meaning of science, characterizing data, and highlighting data curation metadata challenges. The Dryad repository is introduced, and the initiative's metadata best practice and underlying rationales are described. Dryad's metadata approach includes two prongs: one addressing the long-term goal to align with the Semantic Web via a metadata application profile; and another addressing the immediate need to make content available in DSpace via an extensible markup language (XML) schema. The conclusion summarizes limitations and advantages of the two prongs underlying Dryad's metadata effort. KEYWORDS: metadatascientific dataDublin Core Application ProfileSingapore FrameworkSemantic Web ACKNOWLEDGMENT This work is supported by National Science Foundation Grant # EF-0423641. We would like to acknowledge contributions by the Dryad team members Hilmar Lapp and Todd Vision of NESCent; and Michael Whitlock, University of British Columbia. We would also like to thank Stuart Weibel, OCLC, for his thoughtful comments and support of this work. Notes 1. DOE (Department of Energy) Data Explorer (DDE): http://www.osti.gov/dataexplorer/ 2. Knowledge Network for Biocomplexity Data (KNB): http://knb.ecoinformatics.org/ 3. The Dublin Core comprises both the 15 core properties from the DCMES Metadata Element Set (DCMES), Version 1.1. Reference Description: http://dublincore.org/documents/2004/12/20/dces/ and a set of additional properties registered in the DCMI (Dublin Core Metadata Initiative) Metadata Terms namespace: http://dublincore.org/documents/dcmi-terms/ 4. Dublin Core Abstract Model (DCAM): http://dublincore.org/documents/abstract-model/ 5. Dublin Core Application Profile Guidelines: http://dublincore.org/usage/documents/profile-guidelines/. 6. Dryad repository: http://www.datadryad.org/repo/ 7. Dryad repository Partners: http://www.datadryad.org/repo/themes/Dryad/pages/partners.html 8. Joint Data Archiving Policy: http://www.datadryad.org/repo/ 9. Interoperability Levels for Dublin Core Metadata: http://dublincore.org/documents/interoperability-levels/ 10. Dryad Workshop: https://www.datadryad.org/wiki/Dec_5_Workshop_Minutes 11. Collectively the DCMES (http://dublincore.org/documents/2004/12/20/dces/) and DCMI Metadata Terms (http://dublincore.org/documents/dcmi-terms/), as explained in footnote 3. 12. Darwin Core (DwC), Version 1.3: http://digir.sourceforge.net/schema/conceptual/darwin/core/2.0/darwincoreWithDiGIRv1.3.xsd; Version 1.4 being reviewed, see: http://wiki.tdwg.org/twiki/bin/view/DarwinCore/DarwinCoreVersions 13. Publishing Requirements for Industry Standard Metadata (PRISM): http://www.prismstandard.org/specifications/ 14. Journal Publishing Tag Set Tag Library, Version 3.0, November 2008: http://dtd.nlm.nih.gov/publishing/tag-library/ 15. Data Document Initiative (DDI): http://webapp.icpsr.umich.edu/cocoon/DDI-LIBRARY/Version2-1.xsd?section=all 16. Ecological Metadata Language (EML): http://knb.ecoinformatics.org/software/eml/eml-2.0.1/index.html 17. PREMIS Editorial Committee. PREMIS Data Dictionary for Preservation Metadata Version 2.0, 2008: http://www.loc.gov/standards/premis/v2/premis-2-0.pdf 18. Status Element—Dryad: http://www.purl.org/dryad/terms/status 19. Dryad Domain: http://www.purl.org/dryad 20. Text Encoding Initiative (TEI) Header, Chapter 2 (P5: Guidelines for Electronic Text Encoding and Interchange): http://www.tei-c.org/release/doc/tei-p5-doc/en/html/HD.html 21. Tim Berners-Lee on the next Web (TED Conferences, LLC): http://www.ted.com/index.php/talks/tim_berners_lee_on_the_next_web.html 22. GenBank database: http://www.psc.edu/general/software/packages/genbank/genbank.php 23. TreeBASE: http://www.treebase.org 24. Long Term Ecological Research (LTER) Network's Metacat data catalog: http://metacat.lternet.edu/knb 25. Gleaning Resource Descriptions from Dialects of Languages (GRDDL): http://www.w3.org/TR/grddl-primer/
This report presents recent metadata developments for Dryad, a digital repository hosting datasets underlying publications in the field of evolutionary biology. We review our efforts to bring the Dryad application profile into conformance with the Singapore Framework and discuss practical issues underlying the application profile implementation in a DSpace environment. The report concludes by outlining the next steps planned as Dryad moves into the next phase of development.
University music students, teachers, and researchers discover and retrieve musical works and navigate within them, then create annotations and share them with other users.
The Internet contains billions of documents and thousands of systems for searching over these documents. Searching for a useful document can be as difficult as the proverbial search for a needle in a haystack. Each search engine provides access to a different collection of documents. Collections may be large or small, focused or comprehensive. Focused collections may be centered on any possible topic, and comprehensive collections typically have particular topical areas with higher concentrations of documents. Some of these collections overlap, but many documents are available from only a single collection. To find the most needles, one must first select the best haystacks. This dissertation develops a framework for automatic selection of search engines. In this framework, the collection underlying each search engine is examined to determine how properties such as central topic, size, and degree of focus affect retrieval performance. When measured with appropriate techniques, these properties may be used to predict performance. A new distributed retrieval algorithm that takes advantage of this knowledge is presented and compared to existing retrieval algorithms.
Traditional library catalog systems have been extremely effective in providing access to collections of books, films, and other material. However, they have many limitations when it comes to finding musical information, which has significantly different, and in many ways more complex, structure. The Variations2 search system is an alternative system, designed specifically to aid users in searching for music. It leverages a rich set of bibliographic data records, expressing relationships between creators of music and their creations. These records enable musicians to search for music using familiar terms and relationships, rather than trying to decipher the methods libraries typically use to organize musical items. This paper describes the design and implementation of the system that makes these searches possible.
Music information retrieval (MIR) systems tend to fall into two camps: that camp developing cataloging and providing advanced access systems for large collections of music and that camp developing specific query or access mechanisms. We have started to merge these camps by integrating Variations2, which provides access to a digitized portion of Indiana University's vast music library, with Michigan's VocalSearch, which provides a query-by-humming (QBH) search engine. The joint system, V2V, demonstrates how QBH can be used in connection with a large number of holdings in a real-world environment.
A well-known problem for web search is targeting search on information that satisfies users' information needs. User queries tend to be short, and hence often ambiguous, which can lead to inappropriate results from general-purpose search engines. This has led to a number of methods for narrowing queries by adding information. This paper presents an alternative approach that aims to improve query results by using knowledge of a user's current activities to select search engines relevant to their information needs, exploiting the proliferation of high-quality special-purpose search services. The paper introduces the PRISM source selection system and describes its approach. It then describes two initial experiments testing the system's methods.