The growing demand for open science and data-intensive research highlights the urgent need for robust Research Data Management (RDM) systems in universities. While European institutions such as TU Wien have implemented mature FAIR-compliant infrastructures, Indonesian universities still face challenges of fragmented repositories, limited interoperability, and the absence of institutional policies. This paper presents the development of a FAIR-compliant RDM framework tailored for Institut Teknologi Bandung (ITB), derived from lessons learned at TU Wien. The framework integrates five pillars-Policy & Governance, Infrastructure, Processes & Services, Trust & Quality, and Capacity Building-supported by an implementation roadmap. A prototype system was developed using opensource components (InvenioRDM, DBRepo, JupyterHub) to enable research and business intelligence dashboards. Comparative analysis shows that the proposed framework bridges global best practices with local needs, providing both theoretical contributions to institutional data governance and practical tools for evidence-based decision-making. The outcomes aim to strengthen institutional research transparency, support Indonesia's innovation agenda, and offer a replicable model for other universities.
Trusted Research Environments (TREs) enable the analysis of sensitive data under strict security assertions that protect the data with technical, organizational, and legal measures from (accidentally) being leaked outside the facility. While many TREs exist in Europe, little information is available publicly on the architecture and descriptions of their building blocks and their slight technical variations. To highlight on these problems, an overview of the existing, publicly described TREs and a bibliography linking to the system description are provided. Their technical characteristics, especially in commonalities and variations, are analysed, and insight is provided into their data type characteristics and availability. The literature study shows that 47 TREs worldwide provide access to sensitive data, of which two-thirds provide data predominantly via secure remote access. Statistical offices (SOs) make the majority of sensitive data records included in this study available.
Cyber Situational Awareness (CSA) is an important element in both cyber security and cyber defence to inform processes and activities on strategic, tactical, and operational level. Furthermore, CSA enables informed decision making. The ongoing digitization and interconnection of previously unconnected components and sectors equally affects the civilian and military sector. In defence, this means that the cyber domain is both a separate military domain as well as a cross-domain and connecting element for the other military domains comprising land, air, sea, and space. Therefore, CSA must support perception, comprehension, and projection of events in the cyber space for persons with different roles and expertise. This paper introduces NEWSROOM, a research initiative to improve technologies, methods, and processes specifically related to CSA in cyber defence. For this purpose, NEWSROOM aims to improve methods for attacker behavior classification, cyber threat intelligence (CTI) collection and interaction, secure information access and sharing, as well as human computer interfaces (HCI) and visualizations to provide persons with different roles and expertise with accurate and easy to comprehend mission- and situation-specific CSA. Eventually, NEWSROOM’s core objective is to enable informed and fast decision-making in stressful situations of military operations. The paper outlines the concept of NEWSROOM and explains how its components can be applied in relevant application scenarios.
In the era of big data, research has become increasingly data-driven, with vast amounts of information being generated and analyzed to produce new insights and discoveries. This data deluge requires a combination of methods and technologies to store, process, share and preserve research data. With many of the world’s most valuable data being stored in relational databases where it evolves over time as new knowledge is gained and old knowledge invalidated, current repository systems fail to provide researchers with interfaces to conveniently work with this kind of data within their research environments. For this reason, we have developed DBRepo, an institutional data repository for research data in databases (DBRepo) supporting guidelines of the Working Group on Data Citation of the Research Data Alliance. The system has been in use at TU Wien for almost three years now and provides a variety of data science-related interfaces and can be integrated into many workflows and tools. Further, it assists researchers in depositing their datasets by suggesting the table schema (column names, data types, primary key constraints) and it addresses data interoperability issues by suggesting semantic concepts for dataset columns and units of measurements, where applicable. DBRepo is currently in use by six universities globally who use it as data store for hot and cold research data sets. In the paper, we describe their use-cases and provide lessons learned from the various deployments and workflows. Finally, we show how depositing research data into DBRepo increases the data’s visibility.
The main goals and challenges for the life science communities in the Open Science framework are to increase reuse and sustainability of data resources, software tools, and workflows, especially in large-scale data-driven research and computational analyses. Here, we present key findings, procedures, effective measures and recommendations for generating and establishing sustainable life science resources based on the collaborative, cross-disciplinary work done within the EOSC-Life (European Open Science Cloud for Life Sciences) consortium. Bringing together 13 European life science research infrastructures, it has laid the foundation for an open, digital space to support biological and medical research. Using lessons learned from 27 selected projects, we describe the organisational, technical, financial and legal/ethical challenges that represent the main barriers to sustainability in the life sciences. We show how EOSC-Life provides a model for sustainable data management according to FAIR (findability, accessibility, interoperability, and reusability) principles, including solutions for sensitive- and industry-related resources, by means of cross-disciplinary training and best practices sharing. Finally, we illustrate how data harmonisation and collaborative work facilitate interoperability of tools, data, solutions and lead to a better understanding of concepts, semantics and functionalities in the life sciences.
Meeting the conflicting goals of protecting and maintaining control over sensitive data while also allowing access by third parties constitutes a significant challenge. Secure data infrastructures support data visiting in a highly controlled and monitored environment which, if properly set-up and operated, provide high security guarantees through a combination of technical, legal and procedural mechanisms. To ease the process of deploying such a secure data infrastructure, we present a detailed documentation of the architecture and processes of such an infrastructure and provide a pre-configured reference implementation based entirely on open source software that can be flexibly configured to meet differing security requirements and deployment scenarios. We combine mechanisms for data visiting on secured infrastructure components with optional components of data anonymization and fingerprinting, covered by extensive logging and monitoring functions and embedded in defined processes and contractual frameworks. The set-up is based upon the experience of operating such a secure infrastructure in the medical domain for almost ten years, addressing the emerging need to make such a solution available to a larger set of stakeholders. We show that our system significantly enhances data visiting, offers a higher level of data isolation and present our open source reference implementation thereof.
The growth of the Industry 4.0 initiatives insinuate the needs for the future manufacturing workforce to embrace digital technology skills complimentary to their core skills. One of the main prerequisites to facilitate the learning of such skills is the availability of data processing flow and repository for capturing and managing manufacturing data. To address this challenge and as result of this research, we have identified a set of core requirements, proposed a data repository architecture, and deployed a prototype of the repository based on this architecture in one of the educational factory at UTeM. We conclude the paper with the discussion on the technical and societal impact to stakeholders as well as the outline of our future work.
Data curation is a complex, multi-faceted task. While dedicated data stewards are starting to take care of these activities in close collaboration with researchers for many types of (usually file-based) data in many institutions, this is rarely yet the case for data held in relational databases. Beyond large-scale infrastructures hosting e.g. climate or genome data, researchers usually have to create, build and maintain their database, care about security patches, and feed data into it in order to use it in their research. Data curation, if at all, usually happens after a project is finished, when data may be exported for digital preservation into file repository systems. We present DBRepo, a semantic digital repository for relational databases in a private cloud setting designed to (1) host research data stored in relational databases right from the beginning of a research project, (2) provide separation of concerns, allowing the researchers to focus on the domain aspects of the data and their work while bringing in experts to handle classic data management tasks, (3) improve findability, accessibility and reusability by offering semantic mapping of metadata attributes, and (4) focus on reproducibility in dynamically evolving data by supporting versioning and precise identification/cite-ability for arbitrary subsets of data.
Meeting the conflicting goals of protecting and maintaining control over sensitive data while also allowing access by third parties constitutes a significant challenge. Secure data infrastructures support data visiting in a highly controlled and monitored environment which, if properly set-up and operated, provide high security guarantees through a combination of technical, legal and procedural mechanisms. To ease the process of deploying such a secure data infrastructure, we present a detailed documentation of the architecture and processes of such an infrastructure and provide a pre-configured reference implementation based entirely on open source software that can be flexibly configured to meet differing security requirements and deployment scenarios. We combine mechanisms for data visiting on secured infrastructure components with optional components of data anonymization and fingerprinting, covered by extensive logging and monitoring functions and embedded in defined processes and contractual frameworks based upon the experience of operating such a secure infrastructure in the medical domain for almost ten years, addressing the emerging need to make such a solution available to a larger set of stakeholders. We show that our system significantly enhances data visiting, offers a higher level of isolation and present lessons learned.
Database preservation frequently happens post-factum: databases are transferred and migrated into preservation formats and environments after a project has ended. This increases the risks concerning incompatibility and pushes the preservation burden after the initial lifetime and use of the data. We propose a database repository infrastructure, where databases are created, used and preserved directly in the data curation environment. This increases the FAIRness of the data curated as professional data stewardship activities accompany the databases right from the onset. We present the FAIR Data Austria Database Repository (FDA-DBRepo) infrastructure and provide a first version of an open-source reference implementation.
Databases preserved in archives contain highly valu-able information that frequently cannot be made freely accessible for analyses via standard data portals, be it due to legal, commercial or ethical issues. We present Open Source Secure Data Infrastructure and Processes (OSSDIP), the reference implementation of a high-security data visiting infrastructure initially conceived as a safe-compute environment for medical data. It provides highly controlled and monitored data visiting services while ensuring to the largest degree possible that data cannot be extracted from the infrastructure. This may offer archives a viable alternative for providing restricted access to sensitive data in a more flexible manner.
In the last decade, key-value data storage systems have gained significantly more interest from academia and industry. These systems face numerous challenges concerning storage space- and read optimization. There exists a large potential for improving current solutions by introducing new management techniques and algorithms. In this paper we give an overview of the basic concept of key-value data storage systems and provide an explanation for bottlenecks. Further we introduce two new memory management algorithms and a improved index structure. Finally, these solutions are compared to each other and discussed.