
The Scholix Framework (SCHOlarly LInk eXchange) is a high level interoperability framework for exchanging information about the links between scholarly literature and data, as well as between datasets. Over the past decade, publishers, data centers, and indexing services have agreed on and implemented numerous bilateral agreements to establish bidirectional links between research data and the scholarly literature. However, because of the considerable differences inherent to these many agreements, there is very limited interoperability between the various solutions. This situation is fueling systemic inefficiencies and limiting the value of these, separated, sets of links. Scholix, a framework proposed by the RDA/WDS Publishing Data Services working group, envisions a universal interlinking service and proposes the technical guidelines of a multi-hub interoperability framework. Hubs are natural collection and aggregation points for data-literature information from their respective communities. Relevant hubs for the communities of data centers, repositories, and journals include DataCite, OpenAIRE, and Crossref, respectively. The framework respects existing community-specific practices while enabling interoperability among the hubs through a common conceptual model, an information model and open exchange protocols. The proposed framework will make research data, and the related literature, easier to find and easier to interpret and reuse, and will provide additional incentives for researchers to share their data.
This article will showcase the aims and research goals of the project entitled Transforming Libraries and Archives through Crowdsourcing, recipient of a 2016 Institute for Museum and Library Services grant. This grant will be used to fund the creation of four bespoke text and audio transcription projects which will be hosted on the Zooniverse, the world-leading research crowdsourcing platform. These transcription projects, while supporting the research of four separate institutions, will also function as a means to expand and enhance the Zooniverse platform to better support galleries, libraries, archives and museums (GLAM institutions) in unlocking their data and engaging the public through crowdsourcing.
Research e-infrastructures are systems of systems, patchworks of tools, services and data sources, evolving over time to address the needs of the scientific process. Accordingly, in such environments, researchers implement their scientific processes by means of workflows made of a variety of actions, including for example usage of web services, download and execution of shared software libraries or tools, or local and manual manipulation of data. Although scientists may benefit from sharing their scientific process, the heterogeneity underpinning e-infrastructures hinders their ability to represent, share and eventually reproduce such workflows. This work presents HyWare, a language for representing scientific process in highly-heterogeneous e-infrastructures in terms of so-called hybrid workflows. HyWare lays in between business process modeling languages, which offer a formal and high-level description of a reasoning, protocol, or procedure, and workflow execution languages, which enable the fully automated execution of a sequence of computational steps via dedicated engines.
This paper provides an overview of the IMLS-funded project Digging Deeper, Reaching Further: Librarians Empowering Users to Mine the HathiTrust Digital Library, and explains how the project team developed a curriculum and workshop series to train librarians on text mining approaches and tools, in order to address the recognized skills gap between the needs of researchers pursuing digital scholarship and the services that librarians are traditionally trained to provide.
Research and practice in digital preservation requires a solid foundation of evidence of what is being protected and what practices are being used. The National Digital Stewardship Alliance (NDSA) storage survey provides a rare opportunity to examine the practices of most major US memory institutions. The repeated, longitudinal design of the NDSA storage surveys offer a rare opportunity to more reliably detect trends within and among preservation institutions rather than the typical surveys of digital preservation, which are based on one-time measures and convenience (Internet-based) samples. The survey was conducted in 2011 and in 2013. The results from these surveys have revealed notable trends, including continuity of practice within organizations over time, growth rates of content exceeding predictions, shifts in content availability requirements, and limited adoption of best practices for interval fixity checking and the Trusted Digital Repositories (TDR) checklist. Responses from new memory organizations increased the variety of preservation practice reflected in the survey responses.
Quality of products is always of concern to users regardless of the type of products. The focus of this paper is on the quality of Earth science data products. There are four different aspects of quality scientific, product, stewardship and service. All these aspects taken together constitute Information Quality. With increasing requirement on ensuring and improving information quality, there has been considerable work related to information quality during the last several years. Given this rich background of prior work, the Information Quality Cluster (IQC), established within the Federation of Earth Science Information Partners (ESIP) has been active with membership from multiple organizations. Its objectives and activities, aimed at ensuring and improving information quality for Earth science data and products, are discussed briefly.
While digital libraries accommodate remote access via the web and mobile devices, their physical presence tends to be minuscule. A locally-developed prototype digital library application called DLib Wall, connected to MultiTaction display hardware from MultiTouch Ltd., is an attempt to create a rich and engaging onsite presence for the digital collections of the University of Nevada, Las Vegas Libraries. DLib Wall is one of the first applications of its kind, and demonstrates several relatively unexplored interaction modalities. In late 2014, it was deployed in the Goldfield Room — a conference room in the Lied Library — on a wall-mounted array of six 42-inch touch displays. DLib Wall represents both application development and the work of a collaborative team charged with creating a large, interactive, and content-rich media wall experience on a limited resource budget. The group's work included: evaluating and recommending technologies, managing the custom development for the DLib Wall application, and creating an initial content plan for the public rollout. This article highlights the technical aspects of the application while sharing key decision points that the team encountered in the process of bringing the project from conception to completion.
Utah Digital Newspapers is a pioneering digital newspapers program at the University of Utah J. Willard Marriott Library. Recently, a small project team completed a successful migration away from CONTENTdm onto a home-grown system called Solphal, built using open-source applications. The migration process is detailed along with examples of scripts used to prepare and enhance metadata. Transitioning away from a limiting vendor-based solution to a home-grown system has enabled the Utah Digital Newspapers program to be more responsive to user requests as well as realizing greater efficiencies in hardware and software. The platform has opened up new possibilities for the future as the collection continues to grow.
Sometimes the best tool for the job isn't the newest. It can be an existing tool that only requires a bit of polishing. This case study describes how a new digital services librarian, using a systematic approach, worked with colleagues to inaugurate a technology program within an academic library. In the process a previously-established open source web application was revived, extending its utility to contemporary development platforms. At the Queens College Libraries, an open source room reservation and scheduling system provided an answer to two important questions, and a means for building a new program.
The Digital Public Library of America brings together the riches of America's libraries, archives, and museums, and makes them freely available to the world. In order to do this, DPLA has had to build elements of the national digital platform to connect to those institutions and to serve their digitized materials to audiences. In this article, we detail the construction of two critical elements of our work: the decentralized national network of "hubs," which operate in states across the country; and a version of the Hydra repository software that is tailored to the needs of our community. This technology and the organizations that make use of it serve as the foundation of the future of DPLA and other projects that seek to take advantage of the national digital platform.
A large portion of scientific results is based on analysing and processing research data. In order for an eScience experiment to be reproducible, we need to able to identify precisely the data set which was used in a study. Considering evolving data sources this can be a challenge, as studies often use subsets which have been extracted from a potentially large parent data set. Exporting and storing subsets in multiple versions does not scale with large amounts of data sets. For tackling this challenge, the RDA Working Group on Data Citation has developed a framework and provides a set of recommendations, which allow identifying precise subsets of evolving data sources based on versioned data and timestamped queries. In this work, we describe how this method can be applied in small scale research data scenarios and how it can be implemented in large scale data facilities having access to sophisticated data infrastructure. We describe how the RDA approach improves the reproducibility of eScience experiments and we provide an overview of existing pilots and use cases in small and large scale settings.
A strong movement towards openness has seized science. Open data and methods, open source software, Open Access, open reviews, and open research platforms provide the legal and technical solutions to new forms of research and publishing. However, publishing reproducible research is still not common practice. Reasons include a lack of incentives and a missing standardized infrastructure for providing research material such as data sets and source code together with a scientific paper. Therefore we first study fundamentals and existing approaches. On that basis, our key contributions are the identification of core requirements of authors, readers, publishers, curators, as well as preservationists and the subsequent description of an executable research compendium (ERC). It is the main component of a publication process providing a new way to publish and access computational research. ERCs provide a new standardisable packaging mechanism which combines data, software, text, and a user interface description. We discuss the potential of ERCs and their challenges in the context of user requirements and the established publication processes. We conclude that ERCs provide a novel potential to find, explore, reuse, and archive computer-based research.
In 2015, in an effort to make the library ebook experience simpler and easier for users, a group of libraries developed SimplyE, a mobile application for finding, borrowing and reading ebooks from the library. The goal of the project is to advance a national digital platform to help library patrons find, borrow, and consume the largest variety and inventory of content possible, demonstrating that improving the user experience, especially in the area of discovery, increases the consumption of library ebooks. This article describes the project to date, and outlines future plans.
The National Digital Stewardship Residency (NDSR) program addresses the need for a dedicated community of professionals with the knowledge and technical skills to ensure the long-term viability of the digital record by matching recent postgraduate degree recipients with cultural heritage institutions to manage digital stewardship projects. Since the initial NDSR DC pilot program, there have been five more iterations of the program — NDSR New York, NDSR Boston, American Archive of Public Broadcasting NDSR, NDSR Art, and Biodiversity Heritage Library NDSR. Although these programs share the same characteristics, each operates independently and no formalized guidelines or standards currently exist to link all the programs together. In the fall of 2015, the Council on Library and Information Resources (CLIR) was awarded an IMLS grant to evaluate the early NDSR programs. By providing a comprehensive picture of the NDSR programs that were completed by 2016, the study was intended to help the NDSR community build connections across initiatives and learn from the experiences of its first participants. IMLS also funded the NDSR Symposium planned for April 2017, which will serve as an opportunity to bring stakeholders together to incorporate the strongest practices of each iteration and develop a standardized model for future programs.
IMLS awarded the Harvard Library Innovation Lab a National Digital Platform grant to further develop the Lab's Perma.cc web archiving service. The funds will be used to provide technical enhancements to support an expanded user base, aid in outreach efforts to implement Perma.cc in the nation's academic libraries, and develop a commercial model for the service that will sustain the free service for the academic community. Perma.cc is a web archiving tool that puts the ability to archive a source in the hands of the author who is citing it. Once saved, Perma.cc assigns the source a new URL, which can be added to the original URL cited in the author's work, so that if the original link rots or is changed the Perma.cc URL will still lead to the original source. Perma.cc is being used widely in the legal community with great success; the IMLS grant will make the tool available to other areas of scholarship where link rot occurs and will provide a solution for those in the commercial arena who do not currently have one.
Access to information plays a critical role in supporting development.Open access to scientific information is one solution.Up to now, the open access movement has been most successful in the Western hemisphere.The demand for open access is great in the developing world as it can contribute to solving problems related to access gaps.Five emerging countries, called BRICS -Brazil, Russia, India, China and South Africaplay a specific and leading role with a significant influence on regional and global affairs because of their large and fast-growing national economies, their demography and geographic situation.What are they doing in the field of open access?The paper presents some elements for a better understanding essentially based on case studies from scientists and professionals from the BRICS.This paper is an updated and enriched synthesis of a recent work on open access in the BRICS countries published by Litwin, Sacramento CA. 1 Open access and developmentAccess to information plays a critical role in supporting development 2 .Open access to scientific information is one solution.The basic idea is simple: "Make research literature available online without price barriers and without most permission barriers" (Suber 2012, p.8). Free availability on the public Internet and in particular on the easily accessible World Wide Web, includes the permission "for any users to read, download, copy, distribute, print, search, or link to the full texts of these articles, crawl them for indexing, pass them as data to software, or use them for any other lawful purpose, without financial, legal, or technical barriers other than those inseparable from gaining access to the Internet itself" (Budapest Declaration 3 ).Up to now, the open access movement has been most successful in the Western hemisphere.The three essential reference papers on open access, i.e. the Budapest, Berlin and Bethesda declarations were mainly prepared and supported by Western institutions, organizations and communities.Two-thirds of 1 Schöpfel, J. (ed.)
Scientific research is published in journals so that the research community is able to share knowledge and results, verify hypotheses, contribute evidence-based opinions and promote discussion. However, it is hard to fully understand, let alone reproduce, the results if the complex data manipulation that was undertaken to obtain the results are not clearly explained and/or the final data used is not available. Furthermore, the scale of research data assets has now exponentially increased to the point that even when available, it can be difficult to store and use these data assets. In this paper, we describe the solution we have implemented at the National Computational Infrastructure (NCI) whereby researchers can capture workflows, using a standards-based provenance representation. This provenance information, combined with access to the original dataset and other related information systems, allow datasets to be regenerated as needed which simultaneously addresses both result reproducibility and storage issues.
An increasing number of qualitative researchers are relying on dedicated software for the analysis of qualitative data, often referred to as CAQDAS — computer-assisted qualitative data analysis — applications. These applications allow users to annotate, analyze, and visualize qualitative data. To understand and work towards solutions to the challenges of sharing CAQDAS data, the Qualitative Data Repository (QDR) convoked a one-day workshop with CAQDAS developers, practitioners, and repository specialists, on October 28, 2016. This report describes the workshop sessions, participants' areas of interest, and future challenges.
Libraries straddle the information needs of the 21st century. The wifi, computers and now mobile hotspots that some libraries provide their patrons are gateways to a broad, important, and sometimes essential information resources. The research summarized here examines how rural libraries negotiate telecommunications environments, and how mobile hotspots might extend libraries' digital significance in marginalized and often resource-poor regions. The Internet has grown tremendously in terms of its centrality to information and entertainment resources of all sorts, but the ability to access the Internet in rural areas typically lags that experienced in urban areas. Not only are networks less available in rural areas, they also often are of lower quality and somewhat more expensive; even mobile phone-based data plans — assuming there are acceptable signals available — may be economically out of reach for people in these areas. With older, lower income and less digitally skilled populations typically living in rural areas, the role of the library and its freely available resources may be especially useful. This research examines libraries' experiences with providing free, mobile hotspot-based access to the Internet in rural areas of Maine and Kansas.