Open Science contributes to the collective building of scientific knowledge and societal progress. However, academic research currently fails to recognise and reward efforts to share research outputs. Yet it is crucial that such activities be valued, as they require considerable time, energy, and expertise to make scientific outputs usable by others, as stated by the FAIR principles. To address this challenge, several bottom-up and top-down initiatives have emerged to explore ways to assess and credit Open Science activities (e.g., Research Data Alliance, RDA) and to promote the assessment of a broad spectrum of research outputs, including datasets and software (e.g., Coalition for Advancing Research Assessment, CoARA). As part of the RDA-SHARC (SHAring Rewards and Credit) interest group, we have developed a set of recommendations to help implement various rewarding schemes at different levels. The recommendations target a broad range of stakeholders. For instance, institutions are encouraged to provide digital services and infrastructure, organise training and cover expenses associated with making data available for the community. Funders should establish policies requiring Open Access to data produced by funded research and provide corresponding support. Publishers should favour open peer-review models and Open Access to articles, data, and software. Government policymakers should set up a comprehensive Open Science strategy, as recommended by UNESCO and followed by a growing number of countries. The present work details different measures that are proposed to the stakeholders. The need to include sharing activities in research evaluation schemes as an overarching mechanism to promote Open Science practices is specifically emphasised.
Computational models are complex scientific constructs that have become essential for us to better understand the world. Many models are valuable for peers within and beyond disciplinary boundaries. However, there are no widely agreed-upon standards for sharing models. This paper suggests 10 simple rules for you to both (i) ensure you share models in a way that is at least “good enough,” and (ii) enable others to lead the change towards better model-sharing practices.
There are global movements aiming to promote reform of the traditional research evaluation and reward systems. However, a comprehensive picture of the existing best practices and efforts across various institutions to integrate Open Science into these frameworks remains underdeveloped and not fully known. The aim of this study was to identify perceptions and expectations of various research communities worldwide regarding how Open Science activities are (or should be) formally recognised and rewarded. To achieve this, a global survey was conducted in the framework of the Research Data Alliance, recruiting 230 participants from five continents and 37 countries. Despite most participants reporting that their organisation had one form or another of formal Open Science policies, the majority indicated that their organisation lacks any initiative or tool that provides specific credits or rewards for Open Science activities. However, researchers from France, the United States, the Netherlands and Finland affirmed having such mechanisms in place. The study found that, among various Open Science activities, Open or FAIR data management and sharing stood out as especially deserving of explicit recognition and credit. Open Science indicators in research evaluation and/or career progression processes emerged as the most preferred type of reward.
Abstract Research increasingly relies on interrogating large-scale data resources. The NIH National Heart, Lung, and Blood Institute developed the NHLBI BioData CatalystⓇ (BDC), a community-driven ecosystem where researchers, including bench and clinical scientists, statisticians, and algorithm developers, find, access, share, store, and compute on large-scale datasets. This ecosystem provides secure, cloud-based workspaces, user authentication and authorization, search, tools and workflows, applications, and new innovative features to address community needs, including exploratory data analysis, genomic and imaging tools, tools for reproducibility, and improved interoperability with other NIH data science platforms. BDC offers straightforward access to large-scale datasets and computational resources that support precision medicine for heart, lung, blood, and sleep conditions, leveraging separately developed and managed platforms to maximize flexibility based on researcher needs, expertise, and backgrounds. Through the NHLBI BioData Catalyst Fellows Program, BDC facilitates scientific discoveries and technological advances. BDC also facilitated accelerated research on the coronavirus disease-2019 (COVID-19) pandemic.
Software and data citation are emerging best practices in scholarly communication. This article provides structured guidance to the academic publishing community on how to implement software and data citation in publishing workflows. These best practices support the verifiability and reproducibility of academic and scientific results, sharing and reuse of valuable data and software tools, and attribution to the creators of the software and data. While data citation is increasingly well-established, software citation is rapidly maturing. Software is now recognized as a key research result and resource, requiring the same level of transparency, accessibility, and disclosure as data. Software and data that support academic or scientific results should be preserved and shared in scientific repositories that support these digital object types for discovery, transparency, and use by other researchers. These goals can be supported by citing these products in the Reference Section of articles and effectively associating them to the software and data preserved in scientific repositories. Publishers need to markup these references in a specific way to enable downstream processes.
Scientific workflows have become integral tools in broad scientific computing use cases. Science discovery is increasingly dependent on workflows to orchestrate large and complex scientific experiments that range from execution of a cloud-based data preprocessing pipeline to multi-facility instrument-to-edge-to-HPC computational workflows. Given the changing landscape of scientific computing and the evolving needs of emerging scientific applications, it is paramount that the development of novel scientific workflows and system functionalities seek to increase the efficiency, resilience, and pervasiveness of existing systems and applications. Specifically, the proliferation of machine learning/artificial intelligence (ML/AI) workflows, need for processing large scale datasets produced by instruments at the edge, intensification of near real-time data processing, support for long-term experiment campaigns, and emergence of quantum computing as an adjunct to HPC, have significantly changed the functional and operational requirements of workflow systems. Workflow systems now need to, for example, support data streams from the edge-to-cloud-to-HPC enable the management of many small-sized files, allow data reduction while ensuring high accuracy, orchestrate distributed services (workflows, instruments, data movement, provenance, publication, etc.) across computing and user facilities, among others. Further, to accelerate science, it is also necessary that these systems implement specifications/standards and APIs for seamless (horizontal and vertical) integration between systems and applications, as well as enabling the publication of workflows and their associated products according to the FAIR principles. This document reports on discussions and findings from the 2022 international edition of the Workflows Community Summit that took place on November 29 and 30, 2022.
The FAIRPoints organization, co-founded by the authors, aims to provide a platform for conversations to take place around realistic and pragmatic implementations of the FAIR (Findable, Accessible, Interoperable, Reusable) principles. The uniqueness of the FAIRPoints effort stems from an additional aim: to capture conversation contributions in the form of “bite-sized” objects – “points” – in a way that facilitates dynamic composition by instructors for the delivery of audience-customized training experiences. Thus, FAIRPoints aims to cultivate pragmatic learning resources to help realize the FAIR principles in practice, both through inviting speakers to prime and lead discussions focused on choices/challenges regarding FAIR, and amplifying downstream value potential by serializing “points” made during such events as FAIR resources. inviting speakers to prime and lead discussions focused on choices/challenges regarding FAIR, and amplifying downstream value potential by serializing “points” made during such events as FAIR resources. Currently, event outcomes are serialized as LearningResource-typed JSON-LD objects in the schema.org sense, i.e. sdo:LearningResource, where @prefix sdo: ., and conform to the bioschemas.org TrainingMaterial profile. However, any differences in participant perspectives must be reconciled, via git revision control, towards a single “view” of a sdo:LearningResource. This situation is at odds with other explicit aims of the FAIRPoints organization such as including diverse voices and collecting heterogeneous input from a global perspective. Using the FAIR Digital Object (FDO) approach, a FAIRPoints sdo:LearningResource instance may be the Object to which an Identifier points, through an FDO Identifier Record, and sdo:LearningResource may be the FDO Type. Crucially, there may be a multiplicity of Metadata records pointed to by an FDO Identifier Record and thus a formal mechanism to cultivate and publish diverse perspectives. This presentation will outline FAIRPoints’ approach to FDO implementation for learning resources and its relation to published practice. Specifically, in relation to the FAIR Digital Twins approach*1, our approach may be seen as the stewardship of a “fluid graph” of learning-resource “knowlets” with support for “qua” projection in service of e.g. an instructor’s dynamic composition of training material for a targeted workshop.
1 Environmental Genomics and Systems Biology Division, E.O. Lawrence Berkeley National Laboratory, Berkeley, California, United States of America, 2 Biosciences Division, Oak Ridge National Laboratory, Oak Ridge, Tennessee, United States of America, 3 American Geophysical Union, Washington, DC, United States of America, 4 U.S. Department of Energy Office of Scientific and Technical Information, Oak Ridge, Tennessee, United States of America
In Earth and Biological sciences, data are often preserved and publicly available in data repositories where the data are citable by DOIs and published under a Creative Commons CC-BY license. Researchers combine many datasets across disciplines, repositories, and regions to better understand processes, patterns, and drivers. Citing these many datasets is difficult as the large number does not fit into the references section of a paper but the licenses of the datasets require that credit is given to their creators. The Data Citation Community of Practice (CoP) was formed to target such challenges in data citation and other scholarly work that will support indexing and measuring the impact. The CoP identified a container as a solution for large numbers of data citations that holds the citations and its internal format, which is referred to as a 'reliquary'. The existing dataset collection methods have been gathered and evaluated using concrete citation use cases. Requirements for the reliquary content have been identified and applied to the use cases. In this presentation, we will report on the current progress on an approach to building a reliquary. Reliquaries are an important part of enabling cross-disciplinary analysis of large amounts of data stored in many repositories. The challenge with a reliquary will be to design a method that works across diverse repositories and domain citation practices and to enhance the indexing system to direct credit to the reliquary content and authors. The CoP is in the process of setting up a Research Data Alliance (RDA) Working Group on Complex Citations in the Earth, Space, and Environmental Sciences to broaden the discussion and to find further use cases for evaluation and interested early adopters.
AGU is launching a community-driven effort, funded by the Alfred P. Sloan Foundation, to support computational notebooks as primary research objects in scholarly publications.
As a researcher, having access to well-documented research datasets and software relevant to your work can vary in difficulty based on your discipline and other factors. When it works well, you benefit from the ability to easily analyze and perhaps use those data and software. When it works poorly, you are sending emails to get access to datasets, asking for more information, hoping you will get responses, and then maybe trusting that you understand the data or software well enough to integrate, rework, or explore further. How do we create more of the “beneficial” experience? How do we create a culture where having better tools, practices, and methods helps us achieve this goal? Well, it takes deliberate intent and patience in taking those initial first steps. Many of the difficult scientific questions still in front of us require access to more usable data and software, in easier ways, enabling us to “see” the complex systems of the universe better. In this talk, we will share the work happening in AGU, their collaborators, and the broader community to take those initial steps, and support the culture of the future.
A gap in community practice on data citation that emerged during the AGU fall meeting 2020 Data FAIR Town Hall, “Why Is Citing Data Still Hard?” with the goal of addressing the use case of citing a large number of datasets such that credit for individual datasets is assigned properly. The discussion included the concept of a “Data Collection” and the infrastructure and guidance still needed to fully implement the capability so it is easier for researchers to use and receive credit when their data are cited in this manner. Such collections of data may contain thousands to millions of elements with a citation needing to include subsets of elements potentially from multiple collections. Such citations will be crucial to enable reproducible research and credit to data and digital object creators. To address this gap, the data citation community of practice formed including members from data centres, research journals, informatics research communities, and data citation infrastructure. The community has the goal of recommending an approach that is realistic for researchers to use and for each stakeholder to implement that leverages existing infrastructure. To achieve data citation of these subsets of large data collections the concept of a “reliquary” is introduced. In this context the reliquary is a container of persistent identifiers (PIDs) or references defining the objects used in a research study. This can include any number of elements. The reliquary can then be cited as a single entity in academic publications. The reliquary concept will enable data citation use cases such as the citation of elements within a data collection that are formed from numerous underlying datasets that have their own PIDs, unambiguous citation of data used in IPCC Assessment Reports, and citing the subsets of collections of research data that contain millions of elements. The discussions over the course of 2021 have developed a theoretical concept, at the time of writing formal use cases and initial applications are being defined. The recommendation developed by this effort will be available for review and comment by communities such as ESIP and RDA. All are welcome.
Data underlying published studies is difficult to find or access, which can hinder new scientific research. Currently, only about 20% of published papers have their supporting data in discoverable and accessible repositories. The AGU, working with our partners (Dryad, CHORUS, ESIP, Wiley), and supported by the National Science Foundation (NSF), will focus on improving guidance and workflows to properly manage, link, and track data and software references throughout the publication pipeline. The resulting best practices will serve as a resource for AGU editors, reviewers, and authors and help advance data and software publication policies. Beyond AGU, this work will serve as a model for linking information across funders, data repositories, and publishers, and improving public access to research outputs. In this talk, current publication practices as they relate to the FAIR principles will be described, together with lessons learned, and how workflows and guidance are being improved.