The number of data sets, models, capabilities, and tools within Earth and Environmental Systems Sciences Division (EESSD) portfolios is proliferating, with a multitude of data formats, structures, designs, and languages. Given the sheer volume of information and fragmentation of data across multiple repositories, finding relevant data may not be a trivial task; conversely, scientists may also miss important data products or pre-trained models from other repositories that are critical for their research. The current status quo has led to many disparate model repositories and databases that are inaccessible to a wide community of scientific experts. A convergent AI-guided discovery framework can enable search across distributed repositories, provide seamless access through a consistent API, and provide tools for assimilating heterogeneous observations from remote sensing platforms and spatially sparse and distributed networks of sensors. Such a framework holds the key for reliable, accurate, and timely characterization of hydrometeorological conditions such as soil moisture and evapotranspiration fluxes; these are critical for many regional-scale applications, including numerical weather predictions, surface/subsurface hydrology, flood forecasting, drought monitoring, agricultural impacts and planning, and climate change studies.
Data standardization combined with descriptive metadata facilitate data reuse, which is the ultimate goal of the Findable, Accessible, Interoperable, and Reusable (FAIR) principles. Community data or metadata standards are increasingly created through an approach that emphasizes collaboration between various stakeholders. Such an approach requires platforms for collaboration on the development process that centers on sharing information and receiving feedback. Our objective in this study was to conduct a systematic review to identify data standards and reporting formats that use version control for developing data standards and to summarize common practices, particularly in earth and environmental sciences. Out of 108 data standards and reporting formats identified in our review, 32 used GitHub as the version control platform, and no other platforms were used. We found no universally accepted methodology for developing and publishing data standards. Many GitHub repositories did not use key features that could help developers to gather user feedback, or to create and revise standards that build on previous work. We provide guidance for community‐driven standard development and associated documentation on GitHub based on a systematic review of existing practices.
Several factors must be considered in designing a highly accurate, reliable, scalable, and user-friendly geospatial data search interfaces. This paper examines four critical questions that ought to be considered during design phase: (1) Is the search interface or API that provides the search capability useable by both humans and machines? (2) Are the results consistent and reliable? (3) Is the output response format free to use, community-defined, and non-propriety? (4) Does the API clearly state the usage clauses? This paper discusses how certain data repositories at the US Department of Energy's Oak Ridge National Laboratory apply FAIR data principles to enable geospatial searches and address the above-mentioned questions.
The Atmospheric Radiation Measurement (ARM) user facility is a US Department of Energy Office of Science user facility that is managed and operated through a collaborative effort led by nine US Department of Energy national laboratories. The ARM Data Center, located at Oak Ridge National Laboratory, is responsible for the timely collection, processing, and delivery of data products to the scientific community. The ARM Data Center holds more than 11,000 data products, including metadata collected from field campaigns, instruments, value-added products, and principal investigator–contributed data. These data sets are checked for successful transfer (for most data, this transfer is carried out automatically via the network; however, some of the largest data sets and some of the most remote sites require manual shipping of hard disks) and both the data and metadata are processed to a standard format, which is an ARM-standardized structure, via the Network Common Data Form. The Network Common Data Form is a self-describing binary format with many compatible software tools. Once processed, the data are cataloged, stored in the ARM Data Archive, and made discoverable through association with an array of metadata-characterizing information, such as location and measurement classification. These metadata enable powerful search capabilities through the ARM Data Center Data Discovery interface. This paper discusses the workflow of how the new discovery system has been redesigned from user requirements and how the data are distributed to the scientific community.
Focal Areas: (1) Data acquisition and assimilation enabled by machine learning, AI, and advanced methods including experimental/network design/optimization, unsupervised learning (including deep learning), and hardware-related efforts involving AI (e.g., edge computing). Focal areas 2 and 3 have critical dependencies to the modernization described. Key benefits to the focal areas: (1) Modernized observatory framework capable of agile adaptive observation, (2) Advanced instrument and data tagging supporting AI data acquisition for assimilation or validation, and (3) Widespread data interoperability bridging Earth system prediction scales
Prediction and observation of water cycles at various scales involve not only patterns isolated in space and time, but also modeling of complex spatio-temporal relationships across multiple domains. For instance, evapotranspiration (ET) and leaf area indexes (LAI) are two parameters that are needed to accurately model and understand land-atmosphere processes. Accurate assessments of ET and LAI are critical for understanding hydrological processes, deforestation, crop yield, and irrigation impacts. However, ET estimates for global simulations are available at very coarse spatial resolution. They are usually derived from satellite data based on broad plant functional types (PFTs), which fail to capture fine-scale variations because of changes in vegetation type across the globe. Similarly LAI estimates have typically been derived from vegetation indices at global scales or estimated locally using physical models, both of which suffer from a range of uncertainties that impact model sensitivity. The new era of AI model development for Earth systems (ES) calls for data-driven methods that provide domain scientists with uncertainty-aware estimations of biophysical parameters such as ET and LAI in a generalizable, interpretable, and discoverable manner.
Scientific datasets are continuously growing with the amount of raw data being collected worldwide. This amount of data poses the biggest challenge to web search engines on how to retrieve them efficiently. This paper discusses how major scientific data centers are using popular open-source search platforms such as Solr [1] to retrieve structured data stored in data sources such as relational database management systems using its import handler mechanisms [2]. Additionally, we will also focus on how we can configure Solr to serve advanced full-text, faceted search capabilities, along with its key features, which simplify representing and delivering better performance to the scientific search interfaces.
To harness the potential of microbiome science across the broad range of relevant disciplines, new approaches to data infrastructure and transdisciplinary collaboration are necessary. The National Microbiome Data Collaborative (NMDC) is a new initiative to support microbiome data exploration and discovery through a collaborative, integrative data science ecosystem. To harness the potential of microbiome science across the broad range of relevant disciplines, new approaches to data infrastructure and transdisciplinary collaboration are necessary. The National Microbiome Data Collaborative is a new initiative to support microbiome data exploration and discovery through a collaborative, integrative data science ecosystem.
Given the sheer volume of scientific data archived within the data-intensive projects at the US Department of Energy's Oak Ridge National Laboratory, finding precisely what data we are looking for may not be a trivial task; conversely, we may also miss a more prominent data product. To address such issues, we propose improving the data discovery system and using data analytics methods to comprehend what specific users might be interested in based on their physiological state, search patterns, and past data usage history. This work's primary goal is to prune the complexity, increase the visibility of popular data products, and direct users toward the data that best meet their needs. The proposed algorithm constructs a user profile based on the user's explicit or implicit interactions with the system, such as items they are currently looking at on-site and the key metadata mappings related to the data set. The pattern is then used to build a training data set, which will help find relevant data to recommend to the user.
Atmospheric Radiation Measurement (ARM), a U.S. Department of Energy (DOE) scientific user facility, is a key geophysical data source for national and international climate research. Utilizing a standardized schema that has evolved since ARM inception in 1989, the ARM Data Center (ADC) processes over 1.8 petabytes of stored data across over 10,000 data products. Data sources include ARM-owned instruments, as well as field campaign datasets, Value Added Products, evaluation data to test new instrumentation or models, Principal Investigator data products, and external data products (e.g., NASA satellite data). In line with FAIR principles, a team of metadata experts classifies instruments and defines spatial and temporal metadata to ensure accessibility through the ARM Data Discovery. To enhance geophysical metadata collaboration across American and European organizations, this work will summarize processes and tools which enable the management of ARM data and metadata. For example, this presentation will highlight recent enhancements in-field campaign metadata workflows to handle the ongoing Multidisciplinary Drifting Observatory for the Study of Arctic Climate (MOSAiC) data. Other key elements of ARM data center include: the architecture of ARM data transfer and storage processes, evaluation of data quality, ARM consolidated databases. We will also discuss tools developed for identifying and recommending datastreams and enhanced DOI assignments for all data types to assist an interdisciplinary user base in selecting, obtaining, and using data as well as citing the appropriate data source for reproducible atmospheric and climate research.
Data quality assessment, management and improvement is an integral part of any big data intensive scientific research to ensure accurate, reliable, and reproducible scientific discoveries. The task of maintaining the quality of data, however, is non-trivial and poses a challenge for a program like the Department of Energy's Atmospheric Radiation Measurement (ARM) that collects data from hundreds of instruments across the world, and distributes thousands of streaming data products that are continuously produced in near-real-time for an archive 1.7 Petabyte in size and growing. In this paper, we present a computational data processing workflow to address the data quality issues via an easy and intuitive web-based portal that allows reporting of any quality issues for any site, facility or instruments at a granularity down to individual variables in the data files. This portal allows instrument specialists and scientists to provide corrective actions in the form of symbolic equations. A parallel processing framework applies the data improvement to a large volume of data in an efficient, parallel environment, while optimizing data transfer and file I/O operations; corrected files are then systematically versioned and archived. A provenance tracking module tracks and records any change made to the data during its entire life cycle which are communicated transparently to the scientific users. Developed in Python using open source technologies, this software architecture enables fast and efficient management and improvement of data in an operational data center environment.
Atmospheric Radiation Measurement (ARM) is a multi-laboratory/multi-institutional, US Department of Energy Office of Science National User Facility. ARM's data is currently hosted at the ARM Data Center (ADC) in Oak Ridge, Tennessee. The ADC holds more than 12,000 data products, with a total holding of more than 1.8 PB of data that dates back to 1992. This includes data from instruments, value-added products, model outputs, field campaigns, and principle investigator contributed data. In this paper, we discuss how big federal scientific data centers, such as ARM, that use modern and scalable architecture apply findable, accessible, interoperable, and reusable (FAIR) data principles to improve overall efficiency. These principles mainly emphasize machine-to-machine interactions that are directly applicable to ARM because of its data volume.