Data-driven innovation has recently changed the mindset in data sharing from centralized architectures and monolithic data exploitation by data providers (data platforms) to decentralized architectures and different data sharing options among all involved participants (data ecosystems). Data sharing is further strengthened through the establishment of several legal frameworks (e.g., European Strategy for Data, Data Act, Data Governance Act) and the emerging initiatives that provide the means to build data ecosystems, which is evident in the formulated communities, established use cases, and the technical solutions. However, the data ecosystems have not been thoroughly studied so far. The differences between the various data ecosystems are not clear, making it hard to choose the most suitable for each use case, negatively impacting their adoption. Since the domain is growing fast, a review of the state-of-the-art data ecosystem initiatives is needed to analyze what each initiative offers, identify collaboration prospects, and highlight features for improvement and open research topics. In this paper, we review the state-of-the-art data ecosystem initiatives, describe their innovative aspects, compare their technical and business features, and identify open research challenges. We aim to assist practitioners in choosing the most suitable data ecosystem for their use cases and scientists to explore emerging research opportunities. Furthermore, we will provide a framework that outlines the key criteria for evaluating these initiatives, ensuring that stakeholders can make informed decisions based on their specific needs and objectives. By synthesizing our findings, we hope to foster a deeper understanding of the evolving landscape of data ecosystems and encourage further advancements in this critical field.
The 2020 European Strategy for Data aims at developing Common European Data Spaces as a means to build a pan-European single market for data, thereby supporting economic growth and maximizing citizens' use of data. It demands data spaces in strategic sectors, with capabilities for effective data management and sharing. Although several initiatives support their adoption, data spaces are still in the early stages of development and face several data management and sharing challenges. To identify the requirements needed to address these challenges, we review the literature in developing a conceptual framework for applying data management and sharing in the context of data spaces. We then evaluate the practical implementation of the proposed solutions by analysing six representative European-funded projects. Focusing on requirements from trust and business models to interoperability, workflow orchestration, energy efficiency, and data quality, the work highlights prominent issues and explains how each project addresses them through technical means. Our evaluation outlines each project's aim and contribution, along with a representative use case from different domains, e.g., water, agriculture, and energy. We recognize widely accepted strategies such as the use of semantic standards, data catalogues, distributed ledger technologies for trust enhancement, and federated identity management. This work highlights recurring patterns, common practices, and key differences in implementation and identifies open research gaps. Thus, it aims to inform future initiatives and provide concrete recommendations to researchers and practitioners on achieving data spaces with best practices. As the field matures, the work hopes to help achieve scalable, stable, and cross-domain data spaces that support sustainable innovation and long-term collaboration throughout Europe.
The 2020 European Strategy for Data aims at developing Common European Data Spaces as a means to build a pan-European single market for data, thereby supporting economic growth and maximizing citizens' use of data. It demands data spaces in strategic sectors, with capabilities for effective data management and sharing. Although several initiatives support their adoption, data spaces are still in the early stages of development and face several data management and sharing challenges. To identify the requirements needed to address these challenges, we review the literature in developing a conceptual framework for applying data management and sharing in the context of data spaces. We then evaluate the practical implementation of the proposed solutions by analysing six representative European-funded projects. Focusing on requirements from trust and business models to interoperability, workflow orchestration, energy efficiency, and data quality, the work highlights prominent issues and explains how each project addresses them through technical means. Our evaluation outlines each project's aim and contribution, along with a representative use case from different domains, e.g., water, agriculture, and energy. We recognize widely accepted strategies such as the use of semantic standards, data catalogues, distributed ledger technologies for trust enhancement, and federated identity management. This work highlights recurring patterns, common practices, and key differences in implementation and identifies open research gaps. Thus, it aims to inform future initiatives and provide concrete recommendations to researchers and practitioners on achieving data spaces with best practices. As the field matures, the work hopes to help achieve scalable, stable, and cross-domain data spaces that support sustainable innovation and long-term collaboration throughout Europe.
Over the past few years, the scale of sensor networks has greatly expanded. This generates extended spatiotemporal datasets, which form a crucial information resource in numerous fields, ranging from sports and healthcare to environmental science and surveillance. Unfortunately, these datasets often contain missing values due to systematic or inadvertent sensor misoperation. This incompleteness hampers the subsequent data analysis, yet addressing these missing observations forms a challenging problem. This is especially the case when both the temporal correlation of timestamps within a single sensor and the spatial correlation between sensors are important. Here, we apply and evaluate 12 imputation methods to complete the missing values in a dataset originating from large-scale environmental monitoring. As part of a large citizen science project, IoT-based microclimate sensors were deployed for six months in 4400 gardens across the region of Flanders, generating 15-min recordings of temperature and soil moisture. Methods based on spatial recovery as well as time-based imputation were evaluated, including Spline Interpolation, MissForest, MICE, MCMC, M-RNN, BRITS, and others. The performance of these imputation methods was evaluated for different proportions of missing data (ranging from 10% to 50%), as well as a realistic missing value scenario. Techniques leveraging the spatial features of the data tend to outperform the time-based methods, with matrix completion techniques providing the best performance. Our results therefore provide a tool to maximize the benefit from costly, large-scale environmental monitoring efforts.
One of the key challenges for (fresh produce) retailers is achieving optimal demand forecasting, as it plays a crucial role in operational decision-making and dampens the Bullwhip Effect. Improved forecasts holds the potential to achieve a balance between minimizing waste and avoiding shortages. Different retailers have partial views on the same products, which—when combined—can improve the forecasting of individual retailers’ inventory demand. However, retailers are hesitant to share all their individual data. Therefore, we propose an end-to-end graph-based time series forecasting pipeline using a federated data ecosystem to predict inventory demand for supply chain retailers. Graph deep learning forecasting has the ability to comprehend intricate relationships, and it seamlessly tunes into the diverse, multi-retailer data present in a federated setup. The system aims to create a unified data view without centralization, addressing technical and operational challenges, which are discussed throughout the text. We test this pipeline using real-world data across large and small retailers, and discuss the performance obtained and how it can be further improved.
Automating network processes without human intervention is crucial for the complex Sixth Generation (6G) environment. Thus, 6G networks must advance beyond basic automation, relying on Artificial Intelligence (AI) and Machine Learning (ML) for self-optimizing and autonomous operation. This requires zero-touch management and orchestration, the integration of Network Intelligence (NI) into the network architecture, and the efficient lifecycle management of intelligent functions. Despite its potential, integrating NI poses challenges in model development and application. To tackle those issues, this paper presents a novel methodology to manage the complete lifecycle of Reinforcement Learning (RL) applications in networking, thereby enhancing existing Machine Learning Operations (MLOps) frameworks to accommodate RL-specific tasks. We focus on scaling computing resources in service-based architectures, modeling the problem as a Markov Decision Process (MDP). Two RL algorithms, guided by distinct Reward Functions (RFns), are proposed to autonomously determine the number of service replicas in dynamic environments. Our proposed methodology is anchored on a dual approach: firstly, it evaluates the training performance of these algorithms under varying RFns, and secondly, it validates their performance after being trained to discern the practical applicability in real-world settings. We show that, despite significant progress, the development stage of RL techniques for networking applications, particularly in scaling scenarios, still leaves room for significant improvements. This study underscores the importance of ongoing research and development to enhance the practicality and resilience of RL techniques in real-world networking environments.
Floods are a recurring natural disaster that pose significant risks to communities and infrastructure. The lack of reliable and accurate data on river systems in developing countries has hindered the development of effective flood early warning systems. This paper presents a data set collected using ultrasonic distance sensors installed at two locations along the Kikuletwa River in the Pangani River Basin, Northern Tanzania. The dataset consists of hourly measurements of river water levels, providing a high-resolution time series that can be used to study trends in water level changes and to develop more accurate flood early warning systems. The Kikuletwa River dataset has significant potential applications for flood management, including the calibration and validation of hydrological models, the identification of critical thresholds for flood warning, and the evaluation of flood forecasting techniques. The dataset can also be used to study the hydrological processes in the basin, such as the relationship between rainfall and river discharge, and to develop more efficient and effective flood management strategies. The ultrasonic distance sensors were configured to record river level data at hourly intervals, providing a continuous time series of river levels. The data was subjected to quality control procedures to ensure accuracy and consistency, and missing or erroneous data was corrected or removed where necessary.
Reliable and accurate flood prediction in poorly gauged basins is challenging due to data scarcity, especially in developing countries where many rivers remain insufficiently monitored. This hinders the design and development of advanced flood prediction models and early warning systems. This paper introduces a multi-modal, sensor-based, near-real-time river monitoring system that produces a multi-feature data set for the Kikuletwa River in Northern Tanzania, an area frequently affected by floods. The system improves upon existing literature by collecting six parameters relevant to weather and river flood detection: current hour rainfall (mm), previous hour rainfall (mm/h), previous day rainfall (mm/day), river level (cm), wind speed (km/h), and wind direction. These data complement the existing local weather station functionalities and can be used for river monitoring and extreme weather prediction. Tanzanian river basins currently lack reliable mechanisms for accurately establishing river thresholds for anomaly detection, which is essential for flood prediction models. The proposed monitoring system addresses this issue by gathering information about river depth levels and weather conditions at multiple locations. This broadens the ground truth of river characteristics, ultimately improving the accuracy of flood predictions. We provide details on the monitoring system used to gather the data, as well as report on the methodology and the nature of the data. The discussion then focuses on the relevance of the data set in the context of flood prediction, the most suitable AI/ML-based forecasting approaches, and highlights potential applications beyond flood warning systems.
In this study, we tackle the challenge of predicting river levels in flood-prone areas, focusing on the Kikuletwa River sub-catchment within the Pangani River Basin, Tanzania, by employing a Long Short-Term Memory (LSTM) neural network model. The model leverages six key environmental parameters as input data: temperature, humidity, precipitation, surface pressure, wind speed, and soil moisture. A TensorFlow-based architecture consisting of a strategic combination of two LSTM layers, dropout layers to mitigate overfitting, and two fully connected layers is presented. A systematic fine-tuning of hyperparameters was conducted to optimize the model's performance. To further enhance generalization, early stopping was integrated into the training process. The model's efficacy was evaluated using the root mean squared error (RMSE) metric. The results showcase the LSTM model's robust capacity to predict river levels, providing valuable insights for flood forecasting and river management in vulnerable regions. This advanced approach holds significant promises towards the improvement of river level forecasting accuracy and minimizing reliance on costly, labour-intensive manual measurements.
In this work, we deal with the Storage Location Assignment Problem, often referred to as the SLAP, in an E-commerce Distribution Center (EDC). With E-commerce steadily increasing in popularity over the past decades, it has become a key part of the logistics industry. Due to the direct link with the customer, EDC's are forced into a significantly more complex and dynamic order picking process compared to conventional Bulk Distribution Centers. As a result of these challenges, many traditional approaches such as genetic algorithms and rule-based methods reach only suboptimal solutions. We propose the use of Reinforcement Learning (RL) to solve the SLAP, leading to a solution that adapts to dynamically changing environment parameters during runtime. For this purpose, we define a model that transforms the SLAP into a sequential decision making problem. We validate this novel approach by training a state-of-the-art RL algorithm within this model and comparing its results with a benchmark genetic algorithm approach. We conclude that the RL algorithm achieves promising results, surpassing benchmark performance and nearing optimal performance in a small-scale warehouse environment.
In the presented research, machine learning methods were applied to the prediction of longitudinal cracks in steel slabs during continuous casting. We employ a deep learning approach to process 68 thermocouple signals as a multivariate time series (MTS) along with 32 static features, which encompass both chemical composition and process information. Our deep learning approach integrates two distinct parallel modules, followed by an aggregation block; a Convolutional Neural Network (CNN) processes the thermocouple MTS, while in parallel, the static data undergo processing via a Fully Connected Network (FCN). To enhance the performance of the CNN, we incorporate two Squeeze and Excitation (SE) blocks, which act as an attention mechanism across different channels. By integrating chemical information with MTS in the detection system, we improve the performance of defect detection by 15% relatively.
In many industries, multiple parties collaborate on a larger project. At the same time, each of those stakeholders participates in multiple independent projects simultaneously. A double patchwork can thus be identified, with a many-to-many relationship between actors and collaborative projects. One key example is the construction industry, where every project is unique, involving specialists for many subdomains, ranging from the architectural design over technical installations to geospatial information, governmental regulation and sometimes even historical research. A digital representation of this process and its outcomes requires semantic interoperability between these subdomains, which however often work with heterogeneous and unstructured data. In this paper we propose to address this double patchwork via a decentralized ecosystem for multi-stakeholder, multi-industry collaborations dealing with heterogeneous information snippets. At its core, this ecosystem, called ConSolid, builds upon the Solid specifications for Web decentralization, but extends these both on a (meta)data pattern level and on microservice level. To increase the robustness of data allocation and filtering, we identify the need to go beyond Solid’s current LDP-inspired interfaces to a Solid Pod and introduce the concept of metadata-generated ‘virtual views’, to be generated using an access-controlled SPARQL interface to a Pod. A recursive, scalable way to discover multi-vault aggregations is proposed, along with data patterns for connecting and aligning heterogeneous (RDF and non-RDF) resources across vaults in a mediatype-agnostic fashion. We demonstrate the use and benefits of the ecosystem using minimal running examples, concluding with the setup of an example use case from the Architecture, Engineering, Construction and Operations (AECO) industry.
Automated Guided Vehicles (AGV) are omnipresent, and are able to carry out various kind of preprogrammed tasks. Unfortunately, a lot of manual configuration is still required in order to make these systems operational, and configuration needs to be re-done when the environment or task is changed. As an alternative to current inflexible methods, we employ a learning based method in order to perform directed exploration of a previously unseen environment. Instead of relying on handcrafted heuristic representations, the agent learns its own environmental representation through its embodiment. Our method offers loose coupling between the Reinforcement Learning (RL) agent, which is trained in simulation, and a separate, on real-world images trained task module. The uncertainty of the task module is used to direct the exploration behavior. As an example, we use a warehouse inventory task, and we show how directed exploration can improve the task performance through active data collection. We also propose a novel environment representation to efficiently tackle the sim2real gap in both sensing and actuation. We empirically evaluate the approach both in simulated environments and a real-world warehouse.
This paper analyses the requirements for managing interoperable building data in a federated Common Data Environment (CDE). We discuss the need for generic (meta)data storage patterns, semantic query interfaces, decentral authentication, data aggregation, and adaptation and prove that their combination is feasible with current-day technologies. We illustrate the mechanisms of such federated CDE by considering the topic of digital Issue Management, one of the primary functions of a CDE. In an exemplary data flow process, we show how generic (federated, Semantic Web-based) data patterns for Issue Management can be aggregated and restructured to match existing industry standards like buildingSMART’s BIM Collaboration Format (BCF) API. Finally, we show the methodology is compatible with current-day practice by implementing this process in a proof of concept. The main contribution of this research is a generic, federated framework for project-related, interdisciplinary collaboration for CDEs.
Advancements in machine learning techniques, availability of more data sets, and increased computing power have enabled a significant growth in a number of research areas. Predicting, detecting, and classifying complex events in earth systems which by nature are difficult to model is one such area. In this work, we investigate the application of different machine learning techniques for detecting and classifying extreme rainfall events in a sub-catchment within the Pangani River Basin, found in Northern Tanzania. Identification and classification of extreme rainfall event is a preliminary crucial task towards success in predicting rainfall-induced river floods. To identify a rain condition in the selected sub-catchment, we use data from five weather stations that have been labeled for the whole sub-catchment. In order to assess which machine learning technique is better suited for rainfall classification, we apply five different algorithms in a historical dataset for the period of 1979 to 2014. We evaluate the performance of the models in terms of precision and recall, reporting random forest and XGBoost as having the best overall performances. However, because the class distribution is imbalanced, a generic multi-layer perceptron performs best when identifying heavy rainfall events, which are eventually the main cause of rainfall-induced river floods in the Pangani River Basin.
In this paper, we describe the methods used to submit our results to the Rest-Mex Sentiment Analysis of the Iberian Languages Evaluation Forum 2022. The addressed challenge proposes a sentiment analysis task of Spanish opinions, categorizing each message into five emotions, and an attraction prediction subtask divided into three categories. Accordingly, our contribution is a hybrid method based on the Estimation of Distribution Algorithms for fine-tuning an mT5-based transformer. For this, we propose the design and development of a deep learning model using the encoder part of the pre-trained mT5-based transformer. The proposed model is trained by dividing the process into two stages using AdamW and the Covariance Matrix Adaptation Evolution Strategy. With this approach, 0.3050 of Mean Absolute Error was obtained for the polarity detection subtask and 0.9781 of Macro F-measure for the attraction prediction subtask, reaching the 10th place out of the 24 teams in competition.
In the last years, several research works have been proposed for the Knowledge Graph Completion task. However, like most Machine Learning models, most Knowledge Graph Completion models are opaque and lack interpretability. In order to achieve transparency, several interpretable and explainable models have been proposed. The Deterministic Local Interpretable Model-Agnostic Explanations (DLIME) was proposed to solve the lack of stability of the Local Interpretable Model-Agnostic Explanations (LIME), one of the most popular surrogate models. However, using DLIME to explain Machine Learning models in graphs becomes an issue due to its experiments being published only with tabular data. Therefore, this work aims to propose an interpretable method for graphs as an extension of DLIME named DLIME-Graphs. As a triple representation, DLIME-Graphs uses triple embeddings computed by SBERT which in turn, are reduced by the UMAP technique. Instead of using Hierarchical Clustering as DLIME, DLIME-Graphs uses HDB-SCAN to get clusters. To explain a test triple, DLIME-Graphs proposes to train two interpretable models: logistic regression and decision tree plus getting the most similar triples by a k-NN algorithm. The demonstration through a study case showed that DLIME-Graphs is able to give explanations for 100% of the triples in the test dataset through the former models offering transparency and interpretability.
An important step in federated query execution frameworks is source selection, determining which endpoints are relevant to evaluate a given query. Source selection process happens as a separate step before the federated query by executing SPARQL ASK queries, updating catalog/index, or collecting heuristic information as a pre-processing stage, however, in domains as the Linked Open University context, these strategies involve some issues. On the other side, the DCAT metadata vocabulary enables a publisher to describe datasets and data services in a catalog using a standard model and vocabulary that facilitates their consumption. In addition, data summarizations are a lightweight form of representing crucial dataset information. Moreover, the Hydra hypermedia vocabulary along with its Hydra API Documentation allow describing RDF Web APIs facilitating the automation of the client-server communication. This work focuses on using the former semantic vocabularies, along with a context-based unified well-accepted vocabulary, in favor of facilitating the source selection process. In order to explain our proposal, a case study in the Linked Open University context was presented. The case study showed that our proposal allows to select the right sources per triple pattern without further processing complexity as the usage of SPARQL ASK queries and, in turn, it is tailored not only to established query interfaces as SPARQL endpoints, but also to new query interfaces.
In this paper, we propose a novel data-driven prediction system for Multivariate Time Series (MTS) in an industrial context, where classic relational data contain keyinformation in order to properly interpret the MTS. Particularly we focus on the accurate endpoint prediction of temperature and chemical composition at the basic oxygen furnace, which is a step in the steel production pipeline where liquid iron is refined to steel. The precise prediction of temperature is important for proper process control while reaching the target chemical composition is essential for quality control. Our deep learning methodology employs two modules followed by an aggregation block; a Convolutional Neural Network (CNN) handles the MTS, while in parallel, the static data is processed by a Fully Connected Network (FCN). We enhance the CNN performance by adding two Squeeze-and-excitation (SE) blocks, which act like an attention module over the different channels. By taking the MTS data into account we improve the prediction by up to 10% relative over the models which only consider the static data. The hybrid FCN-CNN-SE architecture slightly improves the state-of-the-art MTS approaches by 2%, with less outliers on the prediction of final temperature and phosphorus concentration, while being easier to implement and more scalable to larger datasets and input space than current solutions.
The quality of knowledge graphs can be assessed by a validation against specified constraints, typically use-case specific and modeled by human users in a manual fashion. Visualizations can improve the modeling process as they are specifically designed for human information processing, possibly leading to more accurate constraints, and in turn higher quality knowledge graphs. However, it is currently unknown how such visualizations support users when viewing RDF constraints as no scientific evidence for the visualizations’ effectiveness is provided. Furthermore, some of the existing tools are likely suboptimal, as they lack support for edit operations or common constraints types. To establish a baseline, we have defined visual notations to represent RDF constraints and implemented them in UnSHACLed, a tool that is independent of a concrete RDF constraint language. In this paper, we (i) present two visual notations that support all SHACL core constraints, built upon the commonly used visualizations VOWL and UML, (ii) analyze both notations based on cognitive effective design principles, (iii) perform a comparative user study between both visual notations, and (iv) present our open source tool UnSHACLed incorporating our efforts. Users were presented RDF constraints in both visual notations and had to answer questions based on visualization task taxonomies. Although no statistical significant difference in mean error rates was observed, all study participants preferred ShapeVOWL in a self assessment to answer RDF constraint-related questions. Furthermore, ShapeVOWL adheres to more cognitive effective design principles according to our performed comparison. Study participants argued that the increased visual features of ShapeVOWL made it easier to spot constraints, but a list of constraints – as in ShapeUML – is easier to read. However, also that more deviations from the strict UML specification and introduction of more visual features can improve ShapeUML. From these findings we conclude that ShapeVOWL has a higher potential to represent RDF constraints more effective compared to ShapeUML. But also that the clear and efficient text encoding of ShapeUML can be improved with visual features. A one-size-fits-all approach to RDF constraint visualization and editing will be insufficient. Therefore, to support different audiences and use cases, user interfaces of RDF constraint editors need to support different visual notations.
Joaquim Gabarró合作论文数Department of Computer Science, Universitat Politècnica de Catalunya6