In the age of big data and open science, what processes are needed to follow open science protocols while upholding Indigenous Peoples’ rights? The Earth Data Relations Working Group (EDRWG), convened to address this question and envision a research landscape that acknowledges the legacy of extractive practices and embraces new norms across Earth science institutions and open science research. Using the National Ecological Observatory Network (NEON) as an example, the EDRWG recommends actions, applicable across all phases of the data lifecycle, that recognize the sovereign rights of Indigenous Peoples and support better research across all Earth Sciences. Using the National Ecological Observatory Network (NEON) as an example, this perspective discusses actions that recognize the sovereign rights of Indigenous Peoples and support better research across all Earth Sciences.
This paper presents the specific process we, as part of the Earth Science Information Partners (ESIP) Semantic Harmonization Cluster, used to harmonize terms from the Global Cryosphere Watch (GCW) compilation of cryospheric terms with the Semantic Web for Earth and Environmental Terminology (SWEET) and the Environment Ontology (ENVO). In addition, we summarize a number of leading practices which may be applied to other projects/domains as well as suggest a generalized process for doing so. This paper describes the history of the effort, the technical and decision-making processes used to resolve differences between semantic resources, and describes a number of the issues encountered, with a focus on those that were addressed during the effort. Lessons learned, examples of the problems encountered and a summary of resulting leading practices growing out of this work is provided as an aid to semantic harmonization efforts in other domains by other groups.
Datasets carry cultural and political context at all parts of the data life cycle. Historically, Earth science data repositories have taken their guidance and policies as a combination of mandates from their funding agencies and the needs of their user communities, typically universities, agencies, and researchers. Consequently, repository practices have rarely taken into consideration the needs of other communities such as the Indigenous Peoples on whose lands data are often acquired. In recent years, a number of global efforts have worked to improve the conduct of research as well as data policy and practices by the repositories that hold and disseminate it. One of these established the CARE Principles for Indigenous Data Governance (Carroll et al. 2020), representing ‘Collective Benefit’, ‘Authority to Control’, ‘Responsibility’, and ‘Ethics”’ hosted by the Global Indigenous Data Alliance (GIDA 2023a). In order to align to the CARE Principles, repositories may need to update their policies, architecture, service offerings, and their collaboration models. The question is how? Operationalizing principles into active repositories is generally a fraught process. This paper captures perspectives and recommendations from many of the repositories that are members of the Earth Science Information Partners (ESIPFed, n.d.) in conjunction with members of the Collaboratory for Indigenous Data Governance (Collaboratory for Indigenous Data Governance n.d.) and GIDA, defines and prioritizes the set of activities Earth and Environmental repositories can take to better adhere to CARE Principles in the hopes that this will help implementation in repositories globally.
This paper reports on the Hackathon Sessions organised at the Polar Data Forum IV (PDF IV) (20–24 September 2021), during which 351 participants from 50 different countries discussed collaboratively about the latest developments in polar data management. The 4th edition of the PDF hosted lively discussions on (i) best practices for polar data management, (ii) data policy, (ii) documenting data flows into aggregators, (iv) data interoperability, (v) polar federated search, (vi) semantics and vocabularies, (vii) Virtual Research Environments (VREs), and (viii) new polar technologies. This paper provides an overview of the organisational aspects of PDF IV and summarises the polar data objectives and outcomes by describing the conclusions drawn from the Hackathon Sessions.
In 1987, NASA sponsored an international workshop that inspired the Directory Interchange Format or DIF – a metadata format to enable "catalog interoperability". The DIF formed the basis of the International Directory Network (IDN) and the Global Change Master Directory (GCMD) and included a set of science keywords. The primary intent was to catalog NASA Earth science and related data, but the keywords have been implemented in many different systems and adopted in varying ways by many different organizations around the world. This review provides an ethnographic examination of how the keywords have evolved and been managed and how they have been adopted over the last 35 years. It illustrates how semantic approaches have evolved over time and provides insights on how standards and associated processes can be sustained and adaptable. Ongoing institutional commitment is essential, but so is transparency and technical flexibility. Understanding and empowering the different roles involved in standards creation, maintenance, and use of standards as well as the services that standards enable is also critical. It is apparent that semantic representations need to be mindful of different contexts and carefully define verbs as well as nouns and categories. Understanding and representing relationships is central to interdisciplinary interoperability.
<p>Community resilience increases a place-based community&#8217;s capacity to respond and adapt to life-changing environmental dynamics like climate change and natural disasters. Timely access to environmental data is an important factor for community resilience. Most Earth science information is created for a particular science community for a specific scientific purpose, without much thought to who else could benefit from it and how they might use it. New approaches are needed to facilitate better data production and integration for community use.</p> <p>In this session, we present the findings of a paper published by ESIP&#8217;s (Earth Science Information Partners) Community Resilience Cluster. As a convening space for over 150 member organizations across different sectors, ESIP&#8217;s biannual meetings, conference calls, and topic-driven clusters provided the infrastructure and expertise to support the Community Resilience cluster&#8217;s examination of the role of Earth science data for community resilience. This presentation highlights the challenges communities face when applying Earth science data to their efforts:</p> <p>&#8226; Inequity in the scientific process,</p> <p>&#8226; Gaps in data ethics and governance,</p> <p>&#8226; A mismatch of scale and focus, and</p> <p>&#8226; Lack of actionable information for communities.</p> <p>Recommendations are made as starting points to address the challenges, along with examples of good practices from across the Earth science community. Given ESIP&#8217;s data stewardship efforts with large organizations and across domains, the recommendations are applicable at scale. We offer actionable steps for the Earth science community to help them produce data to better support community resilience.</p>
Community resilience increases a place-based community's capacity to respond and adapt to life-changing environmental dynamics like climate change and natural disasters. In this paper, we aim to support Earth science's understanding of the challenges communities face when applying Earth science data to their resilience efforts. First, we highlight the relevance of Earth science in community resilience. Then, we summarize these challenges of applying Earth science data to community resilience: Inequity in the scientific process, Gaps in data ethics and governance, A mismatch of scale and focus, and Lack of actionable information for communities. Lastly, we offer the following recommendations to Earth science as starting points to address the challenges presented: Integrate community into the scientific data pathway, Build capacity to bridge science and place-based community needs, Reconcile openness with self-governance, and Improve access to data tools to support community resilience.
Earth and Space Science Open Archive PosterOpen AccessYou are viewing the latest version by default [v1]Developing the Cross-Disciplinary Information Model for NASA’s Science Mission DirectorateAuthorsRuthDuerriDAhmedEleishMarkParsonsiDDanielBerriosiDKaylinBugbeePeterFoxiDSee all authors Ruth DuerriDCorresponding Author• Submitting AuthorRonin Institute for Independent ScholarshipiDhttps://orcid.org/0000-0003-4808-4736view email addressThe email was not providedcopy email addressAhmed EleishRensselaer Polytechnic Instituteview email addressThe email was not providedcopy email addressMark ParsonsiDUniversity of Alabama in HuntsvilleiDhttps://orcid.org/0000-0002-7723-0950view email addressThe email was not providedcopy email addressDaniel BerriosiDNASA Ames Research CenteriDhttps://orcid.org/0000-0003-4312-9552view email addressThe email was not providedcopy email addressKaylin BugbeeUniversity of Alabama in Huntsvilleview email addressThe email was not providedcopy email addressPeter FoxiDRensselaer Polytechnic InstituteiDhttps://orcid.org/0000-0002-1009-7163view email addressThe email was not providedcopy email address
As the problems humanity faces become ever more obvious and dangerous, the need for interdisciplinary, cross-disciplinary and trans-disciplinary research and solutions becomes ever more apparent.The problems themselves are often intertwined in complex ways -for example the impact of climate change on human health, food & water security, disasters, and so on and how all of these are exacerbated by human population growth and a general lack of recognition that humanity is part of the ecosystem upon which Earthly life depends.Underpinning our ability to understand and solve these complex problems are
Earth and Space Science Open Archive This is a preprint and has not been peer reviewed. ESSOAr is a venue for early communication or feedback before peer review. Data may be preliminary.Learn more about preprints preprintOpen AccessYou are viewing the latest version by default [v1]FAIR Research Objects - CREDiT and AttribuionAuthorsRuthDuerriDSee all authors Ruth DuerriDCorresponding Author• Submitting AuthorRonin Institute for Independent ScholarshipiDhttps://orcid.org/0000-0003-4808-4736view email addressThe email was not providedcopy email address
Ongoing stewardship is required to keep data collections and archives in existence. Scientific data collections may face a range of risk factors that could hinder, constrain, or limit current or future data use. Identifying such risk factors to data use is a key step in preventing or minimizing data loss. This paper presents an analysis of data risk factors that scientific data collections may face, and a data risk assessment matrix to support data risk assessments to help ameliorate those risks. The goals of this work are to inform and enable effective data risk assessment by: a) individuals and organizations who manage data collections, and b) individuals and organizations who want to help to reduce the risks associated with data preservation and stewardship. The data risk assessment framework presented in this paper provides a platform from which risk assessments can begin, and a reference point for discussions of data stewardship resource allocations and priorities.
Submitted by inkouper on Thu, 2013-05-09 12:17 Friday, July 12, 2013 08:30 to 10:00 Event: Summer Meeting 2013 [2] Session Type: Breakout [3] Expertise Level: Beginner [4] Collaboration Area: Partnership [5] Abstract/Agenda: The Research Data Alliance (RDA) is an international organization that implements the technology, practice, and connections that make data work across barriers. At its first Plenary in March 2013, the RDA was launched by sponsors from the European Commission, the U. S. and the Australian Government and and leaders in the data community. The Plenary was marked by 3 days of presentations, working sessions and informal discussions aimed at increasing the data sharing and exchange needed to drive global data exchanges.
This is a preprint draft of the paper that was officially published in the Data Science Journal. Please quote from the published version: http://doi.org/10.5334/dsj-2020-010. Abstract: Ongoing stewardship is required to keep data collections and archives in existence. Scientific data collections may face a range of risk factors that could hinder, constrain, or limit current or future data use. Identifying such risk factors to data use is a key step in preventing or minimizing data loss. This paper presents an analysis of data risk factors that scientific data collections may face, and a data risk assessment matrix to support data risk assessments to help ameliorate those risks. The goals of this work are to inform and enable effective data risk assessment by: a) individuals and organizations who manage data collections, and b) individuals and organizations who want to help to reduce the risks associated with data preservation and stewardship. The data risk assessment framework presented in this paper provides a platform from which risk assessments can begin, and a reference point for discussions of data stewardship resource allocations and priorities.
Submitted by superadmin on Wed, 2012-10-31 19:51 Overview: This training module is part of the Federation of Earth Science Information Partners (or ESIP Federation's) Data Management for Scientists Short Course. The subject of this module is “Case Study 1 – National Snow & Ice Data Center (NSIDC) Glacier Photos". The module was authored by Matthew Mayernik from the National Center for Atmospheric Research. Besides the ESIP Federation, sponsors of this Data Management for Scientists Short Course are the Data Conservancy and the United States National Oceanic and Atmospheric Administration (NOAA). We have chosen the NSIDC Glacier Photo Collection as a case study to illustrate the importance of preserving the scientific record because from a big picture perspective, the glacier photos in this collection are very valuable for demonstrating climate change. The two pictures of the Holgate Glacier on this slide have been taken almost 100 years apart. As you see, there is dramatic difference in the extent of ice and distribution of the glacier on this particular mountain location.
In this review, we adopt the definition that ‘Data citation is a reference to data for the purpose of credit attribution and facilitation of access to the data’ (TGDCSP 2013: CIDCR6). Furthermore, access should be enabled for both humans and machines (DCSG 2014). We use this to discuss how data citation has evolved over the last couple of decades and to highlight issues that need more research and attention. Data citation is not a new concept, but it has changed and evolved considerably since the beginning of the digital age. Basic practice is now established and slowly but increasingly being implemented. Nonetheless, critical issues remain. These issues are primarily because we try to address multiple human and computational concerns with a system originally designed in a non-digital world for more limited use cases. The community is beginning to challenge past assumptions, separate the multiple concerns (credit, access, reference, provenance, impact, etc.), and apply different approaches for different use cases.
The Repository Finder tool was developed to help researchers in the domain of Earth, space, and environmental sciences to identify appropriate repositories where they can deposit their research data and to promote practices that implement the FAIR Principles, encouraging progress toward sharing data that are findable, accessible, interoperable, and reusable. Requirements for the design of the tool were gathered through a series of workshops and working groups as a part of the Enabling FAIR Data initiative led by the American Geophysical Union that included the development of a decision tree that researchers may follow in selecting a data repository, interviews with domain repository managers, and usability testing. The tool is hosted on the web by DataCite and enables a researcher to query all data repositories by keyword or to view a list of domain repositories that accept data for deposit, support open access, and provide persistent identifiers. Metadata records from the re3data.org registry of research data repositories and the returned results highlight repositories that have achieved trustworthy digital repository certification through a formal procedure such as the CoreTrust Seal.
Despite growing recognition of the importance of public data to the modern economy and to scientific progress, long-term investment in the repositories that manage and disseminate scientific data in easily accessible-ways remains elusive. Repositories are asked to demonstrate that there is a net value of their data and services to justify continued funding or attract new funding sources. Here, representatives from a number of environmental and Earth science repositories evaluate approaches for assessing the costs and benefits of publishing scientific data in their repositories, identifying various metrics that repositories typically use to report on the impact and value of their data products and services, plus additional metrics that would be useful but are not typically measured. We rated each metric by (a) the difficulty of implementation by our specific repositories and (b) its importance for value determination. As managers of environmental data repositories, we find that some of the most easily obtainable data-use metrics (such as data downloads and page views) may be less indicative of value than metrics that relate to discoverability and broader use. Other intangible but equally important metrics (e.g., laws or regulations impacted, lives saved, new proposals generated), will require considerable additional research to describe and develop, plus resources to implement at scale. As value can only be determined from the point of view of a stakeholder, it is likely that multiple sets of metrics will be needed, tailored to specific stakeholder needs. Moreover, economically based analyses or the use of specialists in the field are expensive and can happen only as resources permit.
The EarthCube Technology & Architecture Committee formed a Resource Registry Working Group (WG) to develop a framework for a registry of EarthCube (EC) resources, enabling users to discover scientific and technical resources (software, tools, vocabularies, etc.) that are relevant to their research. The registry will promote EC investments, reduce time to science, help enable interdisciplinary research, more clearly define what is EC, and provide a vehicle for tool and software producers to notify the community about new products, increase visibility, and gain recognition. A primary requirement is to enable systematic description of EarthCube computational resources in terms of their functionality and interfaces for utilization, to enable users to identify components that can work together in integrated workflows. This requires understanding the specifics of how a software component communicates—both the messaging protocol, and the syntax and semantics of information formats getting data into and out of a component. This registry would work in conjunction with schema.org dataset descriptions being developed by the community to streamline linkage of data and software components for research workflows. The WG created definitions for a set of resources to include in a first iteration of the registry, and a set of properties that should be specified for all resources, as well as properties specific to particular resource types. The suggested resource types are: Software, Interface/API, Interchange format, Dataset, Repository, Service, Platform, Vocabulary/ontology/Information model, Specification, Catalog/registry, and Use Case. Dataset and Use Case resources registration is out of scope for the WG project, to be handled separately. Elaboration of this registry is in the workplan for EarthCube, with the goal maximum reuse of existing vocabularies and technology and compatibility with related registry activities.