
Memory institutions and other organizations interested in preserving social media data are using a variety of collection level metadata to represent those materials. The aim of this paper is to start a dialogue within the metadata community about how metadata professionals can describe social media collections in better ways to ensure that the semantic complexity of hashtags remain intact at the collection level. This paper explores how hashtags manifest semantic metadata and how that expression is formally described. A study was conducted using two datasets. The first dataset on hashtags as defined by professional literature was examined and categorized using thematic analysis. The second dataset collected metadata from a selection of Document the Now Twitter datasets and was categorized using Gilliland's (2016) five categories of metadata. Findings delve into the use of collection level metadata to describe social media content.
Wikidata is an open knowledge base that stores structured linked data. It contains over 58 million items ("Wikidata:Statistics," n.d.), but its data reveal a noticeable and prevalent gender disparity. In an effort to contribute to the growth and enhancement of women entries in Wikidata, the Indiana University-Purdue University Indianapolis (IUPUI) University Library and the University of Ottawa Library collaborated to embark on pilot projects that broaden the representation and enhance the visibility of women in STEM (Science, Technology, Engineering, and Mathematics). In this article, we share the methods used at both institutions for collecting faculty data, batch ingesting data using external tools, as well as mapping archival data to existing Wikidata properties. We also discuss the challenges faced during the pilot projects.
The “Japanese Visual Media Graph” project aims to create a research database on Japanese visual media, including, but not limited to anime, manga, computer games and visual novels. It is aimed at researchers in Japan studies who focus on modern media and its expressions, themes, topics, characters and reception. We envision a graph-based, highly interconnected database structure, similar to the Google knowledge graph, that is combined with a flexible search interface and analytic tools. We intend to use the data on Japanese visual media that is being created and curated by the many enthusiast communities on the web. An initial survey of several larger community websites revealed a high level of information granularity, resulting from a deep understanding of the source material and general enthusiasm for recording data by the volunteer contributors. As such, making contact with these communities and learning about their needs and motivations is one of the main project elements. We intend to engage in a meaningful discussion with representatives and administrators of the community sites in order to establish long-term cooperation that benefits both sides.
To advocate open science and knowledge development, the Princess Maha Chakri Sirindhorn Anthropology Centre (SAC) recognizes the significance of collocation of scattered research outputs funded by the SAC for the public use. The SAC Research Database was developed and launched in March 2019 to provide free access to digital full-text research outputs under the creative common license (CC-BY-NC-ND 3.0). The ease of use and interoperability are taken into consideration when selecting the metadata scheme. The Dublin Core Metadata Element Set was chosen with some modified elements for the SAC Research Database. This paper presents the lesson learned from the development of this database.
The University of Calgary library has a unique 40,000 vinyl records collection that is a hidden gem and that is attached to questions on how to make it easily accessible and how to bring it from its physical existence to the digital world. These records include unique classical, folk, jazz, and popular music and are primarily uncatalogued and therefore not accessible to students, faculty and the community. In recent years, vinyl records have regained popularity, especially with millennials and the generation Z. New music has been widely released on vinyl in connection with the opportunity for buyers to download the LP's music digitally. (Harper, 2019) This project's goal is to preserve the library's unique collection but also provide an analog and physical listening experience in a primarily digital music world.
The role of metadata to support research cannot be underestimated; and, yet, it is difficult to develop a systematic understanding of metadata activities throughout the research process. In this paper, we preliminary analyzed how metadata activities were embedded in the research and data lifecycles. Specifically, we identified some key metadata activities associated with the components of the generic research process, from hypothesis formulation to disseminating the results and data management. The exploration raised epistemological questions about the presence of metadata activities in conducting research and managing data. This work conceptualized and grounded the connection between metadata and the lifecycles of research and data processes and presented a high-level mapping identifying the cross section of their activities and established the impression of metadata value in the field of scientific research and data management.
With the shift to a data-driven society, data trading takes on a completely new significance. In the future, data marketplaces will be equivalent to other electronic commerce platforms such as Amazon or eBay. Just like any other online marketplace a data marketplace is a platform that enables convenient buying and selling of products- in this case "data". Metadata is data about data. Metadata plays a significant role in data trading, as it serves as an orientation for all involved parties in the data marketplace. A seller who wants to sell their data on the marketplace needs metadata to describe the selling offer, and the buyer can use it to search and identify relevant data.
Metadata Application Profiles are the elementary blueprints of any Metadata Instance. Efforts like the Singapore Framework for Dublin Core Application Profiles define the framework for designing metadata application profiles to ensure interoperability and reusability. However, the number of publicly accessible, especially machine actionable application profiles are significantly lower. Domain experts find it difficult to create application profiles, considering the technical aspects, costs and disproportionate incentives. Lack of easy-to-use tools for Metadata Application Profile creation is also a reason for lack of larger reach. This paper proposes Yet Another Metadata Application Profile (YAMA) as a user-friendly interoperable preprocessor for creating, maintaining and publishing Metadata Application Profiles. YAMA helps to produce various formats and standards to express the Metadata Application Profiles, changelogs, and different versions, with an expectation of simplifying Metadata Application Profile creation process for domain experts. YAMA includes an integrated syntax for recording application profiles as well as changes between different versions. A proof of concept toolkit, demonstrating the capabilities of YAMA is also being developed. YAMA boasts a human readable yet machine actionable syntax and format, which is seamlessly adaptable to modern version control workflows and expandable for any specific requirements.
Digital collections in cultural heritage institutions are increasingly digitizing physical items, collecting born-digital items, and making these resources available online. Metadata plays a crucial role in the discovery and management of these collections, which makes it important to identify areas of metadata improvement. A number of frameworks and associated metrics support metadata evaluation but this paper focuses on a less-studied aspect of accessibility by using traditional network analysis to understand the connections between metadata records created through shared data values, in elements such as subject or creator. The goal of the research reported in this paper is to investigate potential uses of network analysis and to determine which metrics hold the most promise in effective assessment of metadata at the database or collection level. We introduce the Metadata Record Graph and analyze how it can be used to better understand various-sized collections of metadata.
Japanese Textbook Linked Open Data (LOD) is an LOD dataset of bibliographic and educational information that has been organized over the years by the Library of Education at the National Institute for Educational Policy Research. The dataset consists of bibliographic information for 7,548 volumes of Japanese textbooks authorized from 1992 to 2017, and provides 219,018 Resource Description Framework (RDF) triples as of April 2019. This paper reports a case study of the development and publication of Japanese Textbook LOD.
Though archival resources may be valued for their uniqueness, they do not exist in isolation from each other, and stand to benefit from linked data treatments capable of exposing them to a wider network of resources and potential users. To leverage these benefits, existing, item-level metadata depicting physical materials and their digitized surrogates must be remodeled as linked data. A number of solutions exist, but many current models in this domain are complex and may not capture all relevant aspects of larger, heterogeneous collections of media materials. This paper presents the development of the Linked Archives model, a linked data approach to making item-level metadata available for archival collections of media materials, including photographs, sound recordings, and video recordings. Developed and refined through an examination of existing collection and item metadata alongside comparisons to established domain ontologies and vocabularies, this model takes a modular approach to remodeling archival data as linked data. Current efforts focused on a simplified, user discovery focused module intended to improve access to these materials and the incorporation of their metadata into the wider web of data. This project contributes to work exploring the representation of the range of archival and special collections and how these materials may be addressed via linked data models.
The National Diet Library, Japan (NDL), with support by Xenon Limited Partners, has designed a new metadata schema based on the RDF model while developing a national platform for metadata aggregation and sharing, "Japan Search". Japan Search collects metadata from libraries, museums, archives, and research institutions across the country, and provides an integrated search service as well as APIs (SPARQL Endpoint and REST-API). The aim of this paper is to introduce the new schema, highlighting its dual-layered data model and the normalization of temporal (When), spatial (Where), and agential (Who) information provided in the source data.
The open government movements facilitate the transparency and sharing of government data. Provenance of open government data (OGD) describes source information related to who, how, where, when and other information over the lifecycle of OGD. Provenance of OGD should be tracked for high-quality and trustworthiness of OGD. Currently, OGD portals provide provenance through general metadata elements, such as creator, provider, creation date, publication date, and issued time. In China, local OGD portals define their own metadata profiles. However, these metadata elements in different OGD portals vary and there is no specific and well-defined provenance description scheme for OGD in China. Therefore, this paper is purposed to survey the current provision situation of provenance metadata elements in 42 China OGD portals and conduct the unification of provenance elements based on the survey results. This research is meaningful to facilitate formal description of provenance information in China OGD portals.
As interest in linked data grows throughout the cultural heritage community, it is necessary to critically assess current tools for conversion and creation of linked data "records" and to explore new avenues for creating and encoding data using existing frameworks. This paper discusses the BIBFRAME 2.0 model and Library of Congress conversion specifications from MARC21 through the process of designing and implementing an adapted, minimal-level conversion framework into the cataloging web application, Metadata Maker. In the process of assessment, we identified and addressed local solutions for three key structural issues resulting from the Library of Congress conversion specifications: duplicated data, pervasiveness of blank nodes in RDF/XML, and prevalence of literal data values over URIs stemmed from the current MARC records environment. Additionally, we address concerns with how the BIBFRAME 2.0 model currently conceptualizes Work and linked data as a static "record."
The University of Houston (UH) Libraries, in partnership and consultation with numerous institutions, was awarded an Institute of Museum and Library Services (IMLS) National Leadership/Project Grant to support the creation of the Bridge2Hyku (B2H) Toolkit. Research shows that institutions are inclined to switch from proprietary digital systems to open source digital solutions. However, content migration from proprietary systems to open source repositories remains a barrier for many institutions because of a lack of tools, tutorials, and documentation. The B2H Toolkit includes general migration strategies and use cases as well as tools specifically designed for transitioning from CONTENTdm, a digital collections management software, to the Hyku digital repository. The toolkit acts as a comprehensive resource to guide migration practitioners in migration planning, metadata analysis and harmonization, and to facilitate the repository migration process. This paper focuses on how the toolkit's metadata guidelines and migration tools aid in migration planning, metadata analysis, metadata application profile development, metadata harmonization, and bulk ingest of digital objects into Hyku.
During recent years, cultural heritage institutions have become increasingly interested in participating in open knowledge projects. The most commonly known of these projects is Wikipedia, the online encyclopedia. Libraries and archives in particular, are also showing an interest in contributing their data to Wikidata, the newest project of the Wikimedia Foundation. Wikidata, a sister project to Wikipedia, is a free knowledge base where structured linked data is stored. It aims to be the data hub for all Wikimedia projects. The Wiki community has developed numerous tools and web-based applications to facilitate the contribution of content to Wikidata and to display the data in more meaningful ways. One such web-based application is Scholia which was created to provide users with complete scholarly profiles by making live SPARQL queries to Wikidata and displaying the information in an appealing and effective manner. Scholia provides a comprehensive sketch of the author's scholarship. This presentation will demonstrate our efforts to contribute data related to our faculty members to Wikidata and will provide a demo of Scholia's functionalities. At IUPUI (Indiana University-Purdue University Indianapolis) University Library, we conducted a pilot project where we selected the 19 faculty members identified as core faculty from the IU Lilly Family School of Philanthropy to be included in Wikidata. The School of Philanthropy, located on the IUPUI campus, is the leading school in the subject in the United States. The scholarship produced by its faculty is known to be widely used. The goal of this pilot was not only to provide a presence in Wikidata for our faculty, but also for their publications and co-authors. As a result, we created 110 items to represent some of the works produced by the faculty members and 58 items for all co-authors. Moreover, we selected three publications and worked through their lists of references to contribute 39 cited publication items. Doing the additional work of adding coauthors and cited publications allowed us to start interconnecting works. For the creation of Wikidata items, we used a combination of semi-automated and manual processes. Making use of existing tools such as Source MetaData, QuickStatements, and Resolve Authors alleviated the manual labor, and allowed us to make contributions more efficiently. Once the items were created in Wikidata, we used Scholia to generate the scholarly profiles. By building on existing bibliographic and metadata skills, academic libraries have the capacity to create and curate data about scholars affiliated with their institutions. Our pilot project is just a first step toward more efficient and systematic library-based contributions to Wikidata. We expect that the data sets we build in Wikidata will help our institution better understand and describe the value of its scholarly work in the study of philanthropic giving, nonprofit management, and all other research domains that are a core feature of our campus. In addition to providing value to our local institution, our contributions to Wikidata serve to build an open data platform maintained by the commons. Wikidata provides a welcome, open source alternative to share scholarly and bibliographic data in a marketplace where publishers and other information companies work to capture this data and to profit from selling it back to universities.
The benefits of visualization have been discussed widely and it is already implemented into library services. However, use cases for visualization have been mostly focused on collection analysis to improve collection development policies and budget management, not for discovery services that take full advantage of the rich information contained in library catalog records. One of the challenges of working with library catalog records for visualization is the sheer volume of elements (such as control field, data field, subfield, and indicators) and information included in the MAchine-Readable Cataloging (MARC) format records. As is well-known, there are more than 1,900 fields in the MARC 21, which is just too many to use for effective visualizations (Moen and Benardino, 2003). In addition, some fields are used for recording the same information, for example, the control field 008 positions 7 to 14 and the subfield $c of the data field 264 are used for the production related date information. Instead of showing a clear relationship between resources, the large number of elements and duplicated information included in the catalog record may muddle those relationships in any visualization. The question then is which information added in which fields of the MARC 21 format catalog records should be considered essential information to be included in library catalog data visualizations for discovery. According to Mischo, Schlembach, and Norman's research (2009) on users' search query terms analysis, users tend to use more than three words as search terms (i.e., known item search) rather than simple keyword searches. Many users also use full citations as search terms, thus showing that library users very often already know what they want when they come to the library gateway. Consequently, for the purpose of supporting Functional Requirement for Bibliographic Record (FRBR) User Tasks (IFLA, 2017), such as finding, identifying, selecting, and acquiring (along with browsing), library discovery service systems do not need to index all of the elements included in MARC 21 format catalog records. What is needed instead is only the key information that affects the discovery services, such as access point and authorized access point that connect FRBR Group 1 entities (e.g., work, expression, manifestation, and items) defined by the Resource Description and Access standards (Library of Congress, 2017).
Categorization is a common human behavior and it has many social implications. While categorization helps us make sense of the world around us, it also affects how we perceive the world, what we like and dislike, who we feel comfortable with and who we fear. Categorization is affected by our family, culture and education. But we can take responsibility for our own perceptions, misperceptions can be pointed out and sometimes changed. But what about categorization imposed outside of us that affects us. Should that be allowed? How is that determined? How can it be changed? These are difficult issues. For information aggregators and information analyzers, the guidelines for appropriate behavior are not always clear, nor is the responsibility for outcomes as a result of errors, bias and worse … When errors and bias are commonly held, this can be reflected in the information ecology. The tipping point need not be a majority, truth or based on ethics. It’s easy enough to identify cases of mis-categorization, but when do you do something about it? What can you do about it?
When Tim Berners-Lee published the roadmap for the semantic web in 1998, it was a promising glimpse into what could be accomplished with a standardized metadata system, but nearly 20 years later, adoption of the semantic web has been less than stellar. In those years, web technology has changed drastically, and techniques for implementing semantic web compliant sites have become relatively inaccessible. This poster outlines a JavaScript framework called Beltline.js which seeks to encourage the use of metadata by making it easy to integrate into modern web best-practices.
RDF is short in calculating, especially complex calculating. SPARQL Inferencing Notation (SPIN) has been proposed with a specific capability of returning a value by executing external JavaScript file that in partly performs complex calculating, however it is still far away from accomplishing many practices. This paper investigates SPIN's capability of executing JavaScript, namely SPINx framework, presents a method of equipping RDF data with a new capability of invoking REST API, by which a user who is querying can obtain returned value by invoking the REST API performing complex calculating ,and then the value is semantically annotated for further use .Calculation of lift coefficient of airfoil is taken as a use case ,in which with a given attack angle as input a desired returned value is obtained by invoking a particular REST API while querying the RDF data. Through this use case, it is explicit that RDF data invoking REST API for complex calculating is feasible and profound in both real practice and semantic web.