Large, multidisciplinary projects that collect vast amounts of data are becoming increasingly common in academia. Efficiently managing data across and beyond such projects necessitates a shift from fragmented efforts to coordinated, collaborative approaches. This article presents the data management strategies employed in the Nansen Legacy project (https://doi.org/10.1016/B978-0-323-90427-8.00009-5, Wassmann, 2022), a multidisciplinary Norwegian research initiative involving over 300 researchers and 20 expeditions into and around the northern Barents Sea. To enhance consistency in data collection, sampling protocols were developed and implemented across different teams and expeditions. A searchable metadata catalogue was established, providing an overview of all collected data within weeks of each expedition. The project also implemented a policy that mandates immediate data sharing among members and publishing of data in accordance with the FAIR guiding principles where feasible. We detail how these strategies were implemented and discuss the successes and challenges, offering insights and lessons learned to guide future projects in similar endeavours.
Data products, based on in situ temperature and salinity observations from SeaDataNet infrastructure, have been released within the framework of SeaDataCloud (SDC) project. The data from different data providers are integrated and harmonized thanks to standardized quality assurance and quality control methodologies conducted at various stages of the data value chain. The data ingested within SeaDataNet are earlier validated by data providers who assign corresponding quality flags, but a Quality Assurance Strategy has been implemented and progressively refined to guarantee the consistency of the database content and high quality derived products. Two versions of aggregated datasets for the European marginal seas have been published and used to compute regional high resolution climatologies. External datasets, the World Ocean Database from NOAA and the CORA dataset from the Copernicus Marine Service in situ Thematic Assembly Center, have been integrated with SDC data collections to maximize data coverage and minimize the mapping error. The products are available through the SDC catalogue accompanied by
The Norwegian Scientific Data Network (NorDataNet) is a national e-infrastructure building on the legacy of the International Polar Year. Initially it is focusing on geoscience and establishing interoperability interfaces between existing national data repositories in the areas of discovery metadata and data as well as on harmonised data documentation following the FAIR guiding principles. The technical foundation of NorDataNet is built on data documentation standards, standardised interoperability interfaces and semantic resources. This is now in place and preliminary functionalities are available. These includes the ability to discover and access datasets across the data repositories integrated, as well as visualisation and transformation of datasets served using the requested documentation standards and interfaces. Bottlenecks and achievements while working towards FAIR compliant data and data centres interoperability will be presented.
The Barents Sea, located between the Norwegian Sea and the Arctic Ocean, is one of the main pathways of the Atlantic Meridional Overturning Circulation. Changes in the water mass transformations in the Barents Sea potentially affect the thermohaline circulation through the alteration of the dense water formation process. In order to investigate such changes, we present here a seasonal atlas of the Barents Sea including both temperature and salinity for the period 1965–2016. The atlas is built as a compilation of datasets from the World Ocean Database, the Polar Branch of the Russian Federal Research Institute of Fisheries and Oceanography and the Norwegian Polar Institute using the Data-Interpolating Variational Analysis (DIVA) tool. DIVA allows for a minimization of the expected error with respect to the true field. The atlas is used to provide a volumetric analysis of water mass characteristics and an estimation of the ocean heat and freshwater contents. The results show a recent “Atlantification” of the Barents Sea, that is a general increase in both temperature and salinity, while its density remains stable. The atlas is made freely accessible as user-friendly NetCDF files to encourage further research in the Barents Sea physics (https://doi.org/10.21335/NMDC-2058021735, Watelet et al., 2020).
The data management landscape associated with the Global Ocean Observing System is distributed, complex, and only loosely coordinated. Yet interoperability across this distributed landscape is essential to enable data to be reused, preserved, and integrated and to minimize costs in the process. A building block for a distributed system in which component systems can exchange and understand information is standardization of data formats, distribution protocols, and metadata. By reviewing several data management use cases we attempt to characterize the current state of ocean data interoperability and make suggestions for continued evolution of the interoperability standards underpinning the data system. We reaffirm the technical data standard recommendations from previous OceanObs conferences and suggest incremental improvements to them that can help the GOOS data system address the significant challenges that remain in order to develop a truly multidisciplinary data system.
Temperature and Salinity (TS) historical data collections covering the time period 1900-2013 were created for each European marginal sea (Arctic Sea, Baltic Sea, Black Sea, North Sea, North Atlantic Ocean, Mediterranean Sea) within the framework of SeaDataNet2 (SDN) EU-Project and are available as ODV collections trough the SeaDataNet web catalog at http://sextant.ifremer.fr/en/web/seadatanet/. Two versions have been published and they represent a snapshot of the SDN database content at two different times: V1.1 (January 2014) and V2 (March 2015). A Quality Control Strategy (QCS) has been implemented and continuously refined in order to improve the quality of the SDN database content and to create the best product deriving from SDN data. The QCS was originally implemented in collaboration with MyOcean2 and MyOcean Follow On projects in order to develop a true synergy at regional level to serve operational oceanography and climate change communities. The QCS involved the Regional Coordinators, responsible of the scientific assessment, the National Oceanographic Data Centers (NODC) and the data providers that, on the base of the data quality assessment outcome, checked and eventually corrected anomalies in the original data. The QCS consists of four main phases: 1)data harvesting from the central CDI; 2) file and paramenter aggregation; 3) quality check analysis at regional level; 4) analysis and correction of data anomalies. The approach is iterative to facilitate the upgrade of SDN database content. SDN data collections and the QCS will be presented and the results summarized.
In the frame of the SeaDataNet project, several regional climatologies for the temperature and salinity are being developed by different groups. The data used for these climatologies are distributed by the 40 SeaDataNet data centers. Such climatologies have several uses: 1. the detection of outliers by comparison of the in situ data with the climatological fields, 2. the the optimization of locations of new observations, 3. the initialization of numerical hydrodynamic models. 4. definition of a reference state to identify anomalies and to detect long-term climatic trends Diva (Data Interpolating Variational Analysis) software is adapted to each region by taking into account the geometrical characteristics (coastlines, bathymetry) and the distribution of data (correlation length, signal-to-noise ratio, reference field). The regional climatologies treated in this work are: - JRA5: North Atlantic - JRA6: Mediterranean Sea - JRA7: Baltic Sea - JRA8: North Sea, Arctic Sea Several examples of gridded fields are presented in this work. The validation of the different products is carried out through a comparison with the last release of the widespread World Ocean Atlas 2005.
Temperature and Salinity (TS) historical data collections from 1900 were created for each European marginal sea within the Framework of SeaDataNet2 (SDN) EU-Project. V1 collections of data were created to meet operational oceanography and climate change community requirements that need longer and longer time series of in situ observations to study long term ocean phenomena and their implications in the surrounding environment. This work has been developed in synergy with MyOcean In-Situ Thematic Assemble Centre (INS-TAC) to support and promote monitoring, modeling or downstream service development. The harvesting procedure of temperature and salinity files was performed by a new CDI Robot that used the CDI Data Discovery and Access Service to query, shop and retrieve data sets from the distributed data centers (NODC) in an automatic way. Firstly the shopping mechanism has been tuned through a series of “massive” requests of data. Then the robot retrieved the whole dataset as ODV files including the full CDI metadata, which were then aggregated into a single TS Data Collection using SDN Importer of ODV 4.5.3. It followed the creation of regional and 1900-2012 subsets and the distribution to SDN regional groups, responsible of SDN products, in order to perform quality assessment analysis. Before and during the harvesting and aggregation procedures many efforts have been done to assure the best quality of the V1 collections of data. A screening procedure was applied to identify duplicates and clean SDN infrastructure from redundancies. Numerous files were found to not comply with the ODV/SDN format specification and were rejected. These files have been corrected at the NODCs level. Six TS data collections, one per each European marginal sea (Arctic Sea, Baltic Sea, North Sea, North Atlantic Ocean, Mediterranean Sea, Black Sea), were then analyzed at regional level