The CARE Principles for Indigenous Data Governance are a seminal advance in the stewardship of Indigenous data. The Data Services for Indigenous Scholarship and Sovereignty (DSISS) project is working to guide how research libraries and data repositories can apply the CARE principles to support scholars of Indigenous culture and language. Building on a set of foundational case studies of Indigenous scholarship, this paper reports on analysis of formal engagement activities with scholars, Indigenous community members, and data professionals. We discuss three prominent themes— ownership, trust, and relational accountability—and their implications for concrete steps toward implementation of the CARE principles in research data services (RDS). The results show that sustaining and furthering Indigenous scholarship and data sovereignty in alignment with CARE requires infrastructure and services that attend to a mix of interrelated, and potentially divergent, interest of scholars, Indigenous communities, and institutions. For data professionals, expertise is needed in Indigenous research methods and the sensitivities and distinctiveness inherent in Indigenous ways of knowing, while stewarding institutions will need to make significant investments in restorative trust- building as genuine extensions of relational accountability.
ABSTRACTIn this interactive event, we will discuss possible futures for LIS education and 21st century librarianship, using librarianship as a case study for the rapidly changing demands for expertise across the information professions. In keeping with the conference theme of an information‐resilient society, we will host and encourage conversation about the education and skills that support resilient information professionals.
Advances in data infrastructure are often led by disciplinary initiatives aimed at innovation in federation and sharing of data and related research materials. In library and information science (LIS), the data services area has focused on data curation and stewardship to support description and deposit of data for access, reuse, and preservation. At the same time, solutions to societal grand challenges are thought to lie in convergence research, characterized by a problem-focused ori-entation and deep cross-disciplinary integration, requiring access to highly varied data sources with differing resolutions or scales. We ar-gue that data curation and stewardship work in LIS should expand to foster convergence research based on a robust understanding of the dynamics of disciplinary and interdisciplinary research methods and practices. Highlighting unique contributions by Dr. Linda C. Smith to the field of LIS, we outline how her work illuminates problems that are core to current directions in convergence research. Draw-ing on advances in data infrastructure in the earth and geosciences and trends in qualitative domains, we emphasize the importance of metastructures and the necessary influence of disciplinary practice on principles, standards, and provisions for ethical use across the evolving data ecosystem.
Data referenced in the Library Journal article "Public Libraries and Open Government Data: Partnerships for Progress".
As a data intensive field that unites researchers from many disciplines, Earth System Science (ESS) is an ideal site for examining evolving cross‐disciplinary data practices. This paper reports on results from a survey examining data sharing, data reuse, and research reproducibility practices of ESS researchers, aimed at informing improvements in data services for interdisciplinary sciences. Data reuse was found to be very high for new and comparative analyses but very limited for reproducing research. Data sharing was also strong, mostly through supplements to published papers, with moderate use of open access repositories. At the same time, there was interesting variability in both data sharing and reuse among ESS disciplines. The most pronounced challenges to reuse and reproducibility stem from limited documentation on how data are collected and managed, practices that are poorly supported by institutions, funders, and publishers. A more refined approach to “reproducibility” is needed that aligns with priorities and practices within the research community. Just as importantly, advances in data service models for ESS and other interdisciplinary fields need to account for the diverse and distributed system of repositories and build a workforce with deeper knowledge of the complex data and methods that drive integrative systems science.
The DCC Curation Lifecycle Model has played a vital role in the field of data curation for over a decade. During that time, the scale and complexity of data have changed dramatically, along with the contexts of data production and use. This paper reports on a study examining factors impacting data curation practices and presents recommendations for updating the DCC Curation Lifecycle Model. The study was grounded in a review of other lifecycle models and informed by a site visit to the Digital Curation Centre and consultation with expert practitioners and researchers. Framed by contemporary conditions impacting the conduct of research and provision of data services, the analysis and proposed recommendations account for the prominence of machine-actionable data, the importance of machine learning for data processing and analytics, growth of integrated research workflows, and escalating concerns with fairness, accountability, and transparency of data and algorithms.
As Earth System Science (ESS) becomes more data-intensive, collaborative, and interdisciplinary, it is important to understand how best to support and advance data reuse. We conducted an online survey of active ESS researchers from 126 U.S. universities and research centers, representing a wide variety of scientific fields. Of the 207 respondents, 51.7% had more than 20 years of research experience. Results indicated that the current primary purposes for reusing data are to conduct new analysis (87%), followed by comparing results (70.4%), with only 18.5% reusing data to reproduce published studies. As expected, data hosted by federally funded data centers were reused most frequently, with open government data and data provided directly from other researchers also widely used. Reuse of data from other types of repositories lags far behind, due in part to a range of service limitations. At the same time, data sharing by respondents is strong—96.6% actively release their data, primarily as supplements to published papers, with moderate use of open access repositories. Of the 45.9% who had attempted to reproduce research, 73.7% failed at least once, often due to the limited detail provided in published papers. Still, 92.3% believe it is the researcher’s responsibility to ensure their work is reproducible. The majority favored traditional modes of documenting research—word processors, text editors, and code commenting over electronic notebooks or workflow systems. Interestingly, 59.9% continue to use hand-written notebooks. Challenges to data reuse and reproducibility specific to ESS included the complex nature of earth systems, increasingly complicated models, lack of data management resources, and limited emphasis on reproducibility in the field. Open-ended responses raised questions about whether “exact replication” is necessary or possible for ESS. Most researchers agreed that data and code should be considered important research products and that outlets are needed for publishing negative results. Taken together, the results suggest a strong data sharing culture in ESS with high levels of reuse and commitment to open science. The research community would benefit greatly from better documentation and sharing of methods and research processes, as well as targeted improvements in data services and tools.
ABSTRACTAs the demand for data science and data‐intensive capabilities grows in all sectors, educators in schools of information and library and information science are working to deepen and expand their programs to meet workforce expectations. This panel will examine current trends and investments in data education and professionalization, with an emphasis on the unique contributions information professionals bring to the data workforce. Educators will present on current initiatives that represent the breadth and complexity of preparing information professionals for work in data intensive environments. Their unique approaches anticipate and respond to workforce demand, address significant educational challenges, and offer models for making progress within the varying contexts of schools in different regions of the U.S. and Asia. Half of the session will be reserved for audience engagement designed to leverage and share the wealth of experience of the educators and students in attendance. The exchange will generate ideas about new directions and successful approaches that educators can apply to set priorities, position their programs, and collaborate to enrich and provide leadership in data education.
A comprehensive record of research data provenance is essential for the successful curation, management, and reuse of data over time. However, creating such detailed metadata can be onerous, and there are few structured methods for doing so. In this case study of data curation in support of geobiology research conducted at Yellowstone National Park, we describe a method of “Research Process Modeling” for documenting noncomputational data provenance in a structured yet flexible way. The method combines systems analysis techniques to model research activities, the World Wide Web Consortium Provenance (PROV) ontology to illustrate relationships between data products, and simple inventory methods to account for research processes and data products. It also supports collaborative data curation between information professionals and researchers, and is therefore a significant step toward producing more useable and interpretable research data. We demonstrate how this method describes data provenance more robustly than “flat” metadata alone and fills a critical gap in the documentation of provenance for field‐based and noncomputational workflows. We discuss potential applications of this approach to other research domains.
Abstract In library and information science (LIS), research on interdisciplinarity is concerned with optimizing information resources, systems, and services for researchers working across disciplinary boundaries. Research libraries are responding to the rapid rise in interdisciplinary scholarship and the advances in digital content, technologies, and infrastructure that accompany an emerging data-intensive research paradigm. This chapter considers two key areas in LIS that inform current practice in research libraries—bibliometrics and information practices research. Bibliometric approaches investigate the patterns and flows of information among disciplines, and information practices research examines the activities and materials involved in the conduct of interdisciplinary work. As technical advances continue to solve problems in navigation and retrieval of information across disciplinary boundaries, the greatest challenge will be to assure the meaning and validity of newly created interdisciplinary knowledge through information systems that can sustain the increasingly long and mutable information paths back to our disciplinary intellectual foundations.
The Open Data for Public Good (ODPG) project began July 1, 2016. The purpose of ODPG’s work is threefold: 1. To prepare future and current public librarians to curate collections of open data of value to local communities, 2. To gain experience in building the necessary infrastructure and preservation environments to sustain open data collections for long-term sustainability of these valued assets, and 3. To make collaborate with open civic data providers on advocacy and outreach activities that increase awareness about, and use of open data by the public. Over the course of ODPG’s first year the goals of the project are being achieved through the development of new LIS and data science curriculum, as well as practical learning experiences for students enrolled at the University of Washington’s iSchool. ODPG’s curriculum development activities have focused specifically on new course modules for curating and responsibly managing open civic data. Practical learning experiences have also been facilitated in collaboration with a network of practicing information professionals from the public sector-- each of whom represent a municipal or state government agency engaged in ongoing open data initiatives that can directly benefit from the data expertise of an LIS workforce. We have completed the first year of our three-year program. During this past year we focused on: 1. Updating graduate-level data curation curriculum 2. Pairing students with external partners for field experiences both as Capstone projects and as summer internships 3. Outreach to external partners, potential partners, and the open data community
The Open Data Literacy project is preparing future and current librarians to advance open data initiatives. This poster will provide an overview of the planned activities and project design, with a focus on strategies that iSchools can implement to collaborate with public sector partners to overcome the current lag in data expertise in the public library workforce. Core activities include new curriculum for master’s students in Library and Information Science, a slate of fieldwork opportunities at institutions managing and publishing open data, and community workshops and open education resources for public librarians and information professionals. The educational framework will improve public accessibility and use of open data while increasing the data capabilities of both new and practicing information professionals in public libraries.
Site-Based Data Curation (SBDC) is an approach to managing research data that prioritizes sharing and reuse of data collected at scientifically significant sites. The SBDC framework is based on geobiology research at natural hot spring sites in Yellowstone National Park as an exemplar case of high value field data in contemporary, cross-disciplinary earth systems science. Through stakeholder analysis and investigation of data artifacts, we determined that meaningful and valid reuse of digital hot spring data requires systematic documentation of sampling processes and particular contextual information about the site of data collection. We propose a Minimum Information Framework for recording the necessary metadata on sampling locations, with anchor measurements and description of the hot spring vent distinct from the outflow system, and multi-scale field photography to capture vital information about hot spring structures. The SBDC framework can serve as a global model for the collection and description of hot spring systems field data that can be readily adapted for application to the curation of data from other kinds scientifically significant sites.
This presentation looks back over ten years of experience advancing data curation education at two Information Schools, highlighting the vital role of earth science case studies, expertise, and collaborations in development of curriculum and internships. We also consider current data curation practices and workforce demand in data centers in the geosciences, drawing on studies conducted in the Data Curation Education in Research Centers (DCERC) initiative and the Site-Based Data Curation project. Outcomes from this decade of data curation research and education has reinforced the importance of key areas of information science in preparing data professionals to respond to the needs of user communities, provide services across disciplines, invest in standards and interoperability, and promote open data practices. However, a serious void remains in principles to guide education and practice that are distinct …
Scientific data centers have provided data services to research communities for decades and are invaluable educational partners for iSchools developing academic programs in data curation. This paper presents analyses from three years of internship placements at the National Center for Atmospheric Research and interviews with managers at prominent data centers across the country, as part of the Data Curation Education in Research Centers project. Key benefits of the internship program are identified, from the perspective of student learning and the contributions made by iSchool students to data center operations. The interviews extend the case results, providing evidence for potential data curation internship programs at data centers. The DCERC education model fosters integration of expertise across the iSchool and data center communities, enriching academic preparation with state-of-the-art practical experience in ways that are vital to the emerging data profession and its ability to meet the future demands of data-intensive research.
In this short paper we summarise our experiences from the EU-funded PATHS project. The aim of this project was to support various types of users with their navigation and exploration of large digital cultural heritage collections. A dataset derived from Europeana was gathered and enriched using techniques from natural language processing and information retrieval. This enabled the design of a system incorporating various navigational and exploratory search aids, such as subject hierarchies, recommendations, links to related Wikipedia articles, a workspace and map-based visualisations. The system also allowed users to create narrative-like structures through the collection through trails/paths that can be used as collection guides or used for educational purposes.