This paper describes the maintenance and modernisation of custom data repositories that have supported three long-term research projects since their launch in 2007. It highlights an almost complete rewrite of the code base in 2020 and 2021. To enable research data management (RDM) that adheres to modern and FAIR standards, many features were rebuilt and streamlined for ease of use, with the goal of reducing the friction involved in RDM as much as possible. The update significantly improved the file upload by switching to a fully browser-based solution and completely overhauled the outdated metadata editor with a more interactive version. The ability to search for and find data in the repository has also been enhanced by switching to a flexible, filter-based solution which displays results in real time. It shows that older repositories can be kept in line with the changing landscape of RDM, ensuring that the research data contained therein is not lost. Through these updates and continuing maintenance, these data repositories have stayed available for almost 20 years, even after project funding has ended.
The Helmholtz Metadata Collaboration (HMC) platform was launched in late 2019 to turn FAIR (Findable, Accessible, Interoperable, Reusable) research data into reality within the Helmholtz Association and beyond. The Information Portal was initiated to enable the structured cartography of metadata and FAIR landscape of Helmholtz, providing information for multi-level decision-making and creating a curated knowledge base for research data managers, scientists and other stakeholders. Developed through a top-down approach, 18 categories, and associated metadata schemas were defined and aligned by an HMC taskforce. Data curation followed, with resources collected from different domains based on the aligned metadata schema. The Information Portal is a web application for capturing FAIR data practices across all Helmholtz domains, offering a unified user interface for collecting and exploring results. Built using state-of-the-art technologies, including Python and Docker, the Information Portal leverages GitLab as a database. It offers a public / central read-only version for stakeholders and a personal instance for curation - synchronized to a GitLab repository. Git-based systems offer advantages, such as raw data accessibility, flexible data curation, easy synchronization, and customizable repositories. The single-page web application is user-friendly and developed in multiple iterations for an intuitive and flexible interface. The Information Portal is crucial for creating a sustainable, distributed, semantically enriched Helmholtz data space, promoting seamless data sharing and reuse.
In this document we present our proposal of basic properties that should be part of every PID Kernel Information Profile and PID Record created in the framework of the Helmholtz Metadata Collaboration (HMC). By following these suggestions, we aim to establish a top-level commonality across all research fields in the Helmholtz Association allowing to base cross- community services on top. However, the results presented herein are not limited to the Helmholtz Association, but can also be adopted outside the Helmholtz Association in order to connect contents of data infrastructures. Before reading this document, we recommend to familiarize with basic terms and concepts like Persistent Identifiers, PID Kernel Information Profiles and FAIR Digital Objects as we will touch them only briefly.
Central research data support structures on the university level are usually addressed by the central information-infrastructure providers like university libraries and computing/IT centers. There are several German examples, where also the university research management office and/ or the university's leadership are involved. Besides networking, a critical question that remains is, what model of cooperation and what degree of centralization or decentralization should be chosen. Likely, there will be marked differences between individual universities. The University of Cologne is an old university, which historically developed highly decentralized structures. We will report on first steps to map out the RDM practices at the Cologne campus and to develop a structure of cooperation between loosely coupled information-infrastructure actors.
The Helmholtz Association (Anonymous 2022d), the largest association of large-scale research centres in Germany, covers a wide range of research fields employing more than 43.000 researchers. In 2019, the Helmholtz Metadata Collaboration (HMC) (Anonymous 2022f) Platform as a joint endeavor across all research areas of the Helmholtz Association was started to make the depth and breadth of research data produced by Helmholtz Centres findable, accessible, interoperable, and reusable (FAIR) for the whole science community. To reach this goal, the concept of FAIR Digital Objects (FAIR DOs) has been chosen as top-level commonality for existing and future infrastructures of all research fields.In doing so, HMC follows the original approach of realizing FAIR DOs based on globally unique, Persistent Identifiers (PID), e.g., provided by https://handle.net/, machine actionable PID Records and strong typing using Data Types like https://dtr-test.pidconsortium.eu/#objects/21.T11148/1c699a5d1b4ad3ba4956 registered in a Data Type Registry, e.g., http://dtr-test.pidconsortium.eu/. In all these areas, HMC can build on the great groundwork of the Research Data Alliance and the FAIR DO Forum. However, when it comes to realization, there are still some gaps that will have to be addressed during our work and will be raised in this presentation. For single FAIR DO components like PIDs and Data Types, existing infrastructures are already available. Here, the Gesellschaft für wissenschaftliche Datenverarbeitung mbH Göttingen (GWDG) (Anonymous 2022e) provides strong support with their many years of experience in this field. Within the framework of the ePIC consortium (Anonymous 2022c), the GWDG is offering on the one hand PID prefixes based on a sustainable business model, on the other hand GWDG is very active in terms of providing base services required for realizing FAIR DOs, e.g., different instances of Data Type Registries for accessing, creating, and managing Data Types required by FAIR DOs. Besides that, in the context of HMC we developed a couple of technical components to support the creation and management of FAIR DOs: The Typed PID Maker (Pfeil 2022b) providing machine actionable interfaces for creating, validating, and managing PIDs with machine-actionable metadata stored in their PID record, or the FAIR DO testbed, currently evolving into the FAIR DO Lab (Pfeil 2022a), serving as reference implementation for setting up a FAIR DO ecosystem. However, introducing FAIR DOs is not only about providing technical services, but also requires the definition and agreement on interfaces, policies, and processes.A first step in this direction was made in the context of HMC by agreeing on a Helmholtz Kernel Information Profile (http://dtr-test.pidconsortium.eu/#objects/21.T11148/b9b76f887845e32d29f7). In the concept of FAIR DOs, PID Kernel Information as defined by Weigel et al. (Weigel et al. 2018) is key to machine actionability of digital content. Strongly relying on Data Types and stored in the PID record directly at the PID resolution service, PID Kernel Information can be used by machines for fast decision making. The Helmholtz Kernel Information Profile is an attempt to introduce a top-level commonality across all digital assets produced within the Helmholtz Association and beyond to establish a basis for FAIR research data based on FAIR DOs.Hereby, the Helmholtz Kernel Information Profile integrates the recommendations of the RDA PID Kernel Information Working Group (Anonymous 2022b) as far as possible. By extending the Draft Kernel Information Profile (Weigel et al. 2018) with additional, mostly optional attributes, the Helmholtz Kernel Information Profile allows the adding of contextual information to FAIR DOs, e.g., research topic, or contact information, which is then available for machine decisions. Furthermore, additional properties for representing relationships between FAIR DOs, e.g, hasMetadata and isMetadataFor, were introduced to allow mutual relations between FAIR DOs.Currently, a demonstrator is implemented integrating the above components and services, i.e., PID Service, Data Type Registry, and Typed PID Maker. Fig. 1 outlines the architecture overview of the first version of the demonstrator.In this first version, in a semi-automatic workflow, a user enters a Zenodo (Anonymous 2022a) PID in a graphical Web frontend. A mapping component tries to fill automatically at least the properties required by the Helmholtz Kernel Information Profile using the obtained Zenodo metadata record. In a manual validation loop, the user may add or update certain properties before they are sent to an instance of the Typed PID Maker, validated against the Helmholtz Kernel Information Profile, and stored in the record of a newly registered PID using the services of the ePIC consortium. In addition, registered PID records are made searchable via the graphical frontend on top of a search index, e.g., realized using https://www.elastic.co/.After implementing this generic workflow, additional mappers supporting other repository platforms will be implemented based on the lessons learned, which will lead to a growing number of FAIR DOs and holds potential for providing significant benefits to scientists, e.g., a central point of contact for research data sets stored in different repositories, machine-actionable identification of relevant datasets, and creation of knowledge graphs representing relationships between data sets, repository platforms, researchers and research organizations.Furthermore, the gathered experience and its documentation will help others to apply the FAIR DO concept more easily, which will lead to an ever-growing collection of available FAIR DOs with an increasing quality and level of automation at creation time.
Forschungsdatenmanagement in Institutionen und Forschungsverbünden bringt neue Rollen, Aufgaben- und Berufsprofile hervor, die bisher ganz unterschiedlich realisiert sind. Eins dieser neuen Berufsprofile ist die Position des sogenannten 'Data Stewards'. Mit diesem Arbeitsbereich werden eine große Bandbreite an Tätigkeiten und Verantwortlichkeiten verknüpft. Oftmals werden spezielle Aufgaben übernommen, ohne explizite Arbeitstitel dafür zu vergeben. Im 11. Workshop der DINI/nestor AG Forschungsdaten “Data Stewardship im Forschungsdatenmanagement - Was ist das? Rollen, Aufgabenprofile, Einsatzgebiete” am 16./17.11.2020 wurden die inhaltlichen Aufgabenfelder und Rollen, die sich zurzeit im institutionellen und institutionsübergreifenden Kontext entwickeln, gesammelt, diskutiert und eingeordnet. Ziele des Workshops waren, eine Bestandsaufnahme für den deutschsprachigen Raum anhand zahlreicher Fallbeispiele zu präsentieren und erste Antworten auf Fragen zur praktischen Umsetzung von Data Stewardship-Konzepten aufzuzeigen. Der Workshop wurde von der DINI/nestor-AG Forschungsdaten in Kooperation mit der Universität zu Köln und dem ZB MED Informationszentrum Lebenswissenschaften organisiert und virtuell als Online-Workshop durchgeführt.
Forschungsdatenmanagement in Institutionen und Forschungsverbunden bringt neue Rollen, Aufgaben- und Berufsprofile hervor, die bisher ganz unterschiedlich realisiert sind. Eins dieser neuen Berufsprofile ist die Position des sogenannten 'Data Stewards'. Mit diesem Arbeitsbereich werden eine grose Bandbreite an Tatigkeiten und Verantwortlichkeiten verknupft. Oftmals werden spezielle Aufgaben ubernommen, ohne explizite Arbeitstitel dafur zu vergeben. Im 11. Workshop der DINI/nestor AG Forschungsdaten “Data Stewardship im Forschungsdatenmanagement - Was ist das? Rollen, Aufgabenprofile, Einsatzgebiete” am 16./17.11.2020 wurden die inhaltlichen Aufgabenfelder und Rollen, die sich zurzeit im institutionellen und institutionsubergreifenden Kontext entwickeln, gesammelt, diskutiert und eingeordnet. Ziele des Workshops waren, eine Bestandsaufnahme fur den deutschsprachigen Raum anhand zahlreicher Fallbeispiele zu prasentieren und erste Antworten auf Fragen zur praktischen Umsetzung von Data Stewardship-Konzepten aufzuzeigen. Der Workshop wurde von der DINI/nestor-AG Forschungsdaten in Kooperation mit der Universitat zu Koln und dem ZB MED Informationszentrum Lebenswissenschaften organisiert und virtuell als Online-Workshop durchgefuhrt.
Die Digitalisierung bietet ein neues, effizientes Hilfsmittel, den wissenschaftlichen Fortschritt zu unterstutzen. In nahezu allen Bereichen lassen sich mit Hilfe moderner Informationssysteme Forschungsdaten digital archivieren und bei Bedarf leichter wiederverwenden. In der heutigen Zeit, in der das kollektive Wissen ein enormes Ausmas angenommen hat, ist die Systematisierung von Forschungsdaten zwingender denn je erforderlich. Die E-Science-Tage 2019, aus denen dieser Tagungsband hervorgegangen ist, haben neue Wege der Verarbeitung von Forschungsdaten aufgezeigt und durch den regen Austausch von Erfahrungen und Innovationen die digitale Wissenschaft weiter vorangetrieben.
Diese Sammlung enthalt die PDFs der Vortragsfolien, der Mural-Boards und des Programms des 11. DINI/nestor Workshop Data Stewardship im Forschungsdatenmanagement.
Dislocated boulders are one sign of high-energy wave impacts on coasts. These high-energy impacts, caused by severe storms or tsunamis, can trigger initial cracking and transport of boulders. Monitoring of these boulders, as well as the associated coastal sites is important in distinguishing between gradual coastal processes and high-energy events. Western Greece is a seismically active area, where tsunamis and high-energetic storms might occur and such past events are documented by historic and geoscientific research, making it an ideal location for monitoring dislocated boulders. Since 2008, monitoring of eight different coastal sites in this region was conducted by terrestrial laser scanning and photogrammetric approaches, with low-cost unmanned aerial vehicles. The re-use of similar surveying points in following years, allowed highly accurate monitoring. Point clouds derived from these methods were evaluated for change detection by point cloud comparisons. The data were also used to establish accurate three-dimensional models of dislocated boulders (n = 70). The determined boulder volumes of these accurate three-dimensional models were incorporated in wave transport equations and wave decay curves, and compared with monitoring results. A comprehensive overview of dislocated boulders in western Greece is presented. Three-dimensional boulder reconstruction is compared to an approach which uses a tape-based measuring of boulder axes, with the tape-based measurement showing a mean overestimation of mass by 32%. Accurate monitoring over time by both methods, is achieved by using fixed networks of reference points. Changes for each site over time, detected by direct point cloud comparisons, are fit to the possible inundation calculated by wave decay curves based on computed minimum wave heights for boulder transport. Both storm and tsunami waves may have initiated movement from the cliff edge and further transport is also possible. However, boulders showed no further movement from their current position in the area for the time period of this study.
Forschungsdatenmanagement (FDM), insbesondere an Universitaten, ist gepragt durch eine Vielzahl von Stakeholdern (SH), die auf komplexe Art miteinander wechselwirken. Technologie spielt in der Regel dabei eine untergeordnete Rolle. Es sind soziologische und kulturelle Gegebenheiten, die dazu fuhren, dass institutionelles FDM als eine Art Wicked Problem gesehen werden kann. Hierbei lassen sich klassische Methoden des Projekt- und Kundenbeziehungsmanagements nicht mehr einfach (i. S. v. Ursache und Wirkung) anwenden.Die dezentrale Organisationskultur der Universitat zu Koln (UzK) macht dies deutlich. Hier stehen zentrale Einrichtungen wie Rechenzentrum, Bibliothek, aber auch das Dezernat Forschungsmanagement im Wechselspiel untereinander und mit einer Zahl von weiteren Akteuren auf dem Campus.Mit dem Ziel der Etablierung eines universitatsweiten FDM, hat sich bereits wahrend des Vorprojektes (Dierkes und Curdt 2018) gezeigt, dass den SH auf mehrfache Weise Rechnung getragen und ein Prozess des gemeinsamen Verstehens etabliert werden muss. Erst durch Letzteres scheint eine nachhaltige Weiterentwicklung des institutionellen FDM moglich, indem die vielseitigen Bedarfe abwagend relativiert werden und gleichzeitig eine am Bedarf ausgerichtete Angebotsentwicklung verfolgt werden kann.Im Rahmen des aktuellen Folgeprojektes, welches ein Kompetenzzentrum fur Forschungsdaten an der UzK etablieren will, werden Dialoge auf unterschiedlichen Ebenen initiiert und Arbeitsablaufe hinsichtlich Beratung/Training und Umsetzung von FDM-Masnahmen entwickelt. Diese sollen dazu dienen, Klarheit, Akzeptanz und Effektivitat zwischen den FDM-Akteuren auf dem Campus zu schaffen.Wir berichten in unserem Beitrag uber erste Erfahrungen bei dem Aufbau eines Multi-SH-Managements zur Entwicklung von FDM-Angeboten fur eine der grosten Universitaten Deutschlands. Hinsichtlich der eingangs erwahnten Wickedness wollen wir Methoden wie das Cynefine-Framework einsetzen (Childs und McLeod 2013).
Zitiervorschlag Curdt, Constanze, Dirk Hoffmeister, Tanja Kramm, Ulrich Lang und Georg Bareth. 2019. Etablierung von Forschungsdatenmanagement-Services in geowissenschaftlichen Sonderforschungsbereichen am Beispiel des SFB/Transregio 32, SFB 1211 und SFB/ Transregio 228. Bausteine Forschungsdatenmanagement. Empfehlungen und Erfahrungsberichte für die Praxis von Forschungsdatenmanagerinnen und -managern Nr. 2/2019: S. 61-67. DOI: 10.17192/bfdm.2019.2.8103. Dieser Beitrag steht unter einer Creative Commons Namensnennung 4.0 International Lizenz (CC BY 4.0).
Science conducted in cross-institutional, interdisciplinary, long-term research projects requires active sharing of data, documents and further information. Thus, within the Collaborative Research Centre/Transregio 32 ‘Patterns in Soil-Vegetation-Atmosphere Systems’, funded by the German Research Foundation, research data management (RDM) services have been available since early 2007. These services were established to support all researchers during their entire individual research studies. They cover provision of general guidance, support and training for RDM. To fulfil the scientists’ needs and requests with regard to storage, backup, documentation, search and sharing of data with other project members, a project-specific RDM system was designed and implemented. This system was developed and continuously modified in collaboration with the scientists to facilitate their system acceptance. Besides the mentioned services, the system supports further common services such as controlled access to data, rights management, data publication with DOI and data statistics (on repository and single data level). All RDM services provided for the scientists are thus bundled and available to the users in one system: a ‘one-stop-shop’. After more than ten years of RDM service provision for the CRC/TR32, the repository statistics clearly visualize the use of the diverse RDM system services. Furthermore, it has been shown that an RDM adapted to the needs of interdisciplinary researchers can be fruitful and indispensable when scientists conduct their research study e.g. with a time lag. RDM services established at an early stage can contribute to a successful long-term research project.
Die Universität zu Köln, als eine der größten Hochschulen Deutschlands, nähert sich dem Thema universitätsweites systematisches Forschungsdatenmanagement (FDM) über eine Machbarkeitsstudie an. Im Laufe eines Jahres wurde der Status quo des Umgangs mit Forschungsdaten an der Universität, den Fakultäten, Instituten und Forschungsprojekten ermittelt. Als Grundlage für die weiteren Arbeiten wurde eine Leitlinie zum Umgang mit Forschungsdaten erarbeitet und seitens der Universität verabschiedet. Ausgehend von einem umfänglichen FDM-Service-Portfolio wurden erste Maßnahmenpakete entwickelt, die mit einer realistischen Aufwandsabschätzung eine Grundlage für ein universitätsweites FDM innerhalb der nächsten drei Jahre legen sollen. Die Maßnahmen basieren im Wesentlichen auf dem Aufbau von Informations-, Beratungs- und Schulungsangeboten und sollen die Vernetzung der FDM-Akteure stärken. Ein weiteres Arbeitsgebiet liegt im Aufbau digitaler Services im Bereich Speicherung und Sichtbarmachung von Forschungsergebnissen. The University of Cologne, one of the largest universities in Germany, has approached the topic of university-wide, systematic research data management (RDM) by means of a feasibility study. In the course of a year, the status quo of the handling of research data at the university, the faculties, institutes and research projects was investigated. As a basis for further work, a guideline for the handling of research data was developed and adopted by the university. Based on a comprehensive RDM-service portfolio, first packages of measures were developed in order to provide a basis for a university-wide RDM within the next three years, also giving a realistic estimate of costs. Essentially, the measures are centred on the development of information, consulting and training services and are intended to strengthen the networking of RDM actors. Another field of activity is the development of digital services in the area of storage and visualisation of research results.