AbstractOften underestimated, (semi-)structured textual data sources are an important cornerstone in the manufacturing sector for product and process quality tracking. The ELG pilot project SLAPMAN develops novel methods for industrial text analytics in the form of scalable, reusable, and potentially stateful microservices, which can be easily orchestrated by domain experts in order to define quality anomaly patterns, e. g., by analysing machine states and error logs. The results are fully available as open source and integrated into the IIoT toolbox Apache StreamPipes.
The industrial IoT and its promise to realize data-driven decision-making by analyzing industrial event streams is an important innovation driver in the industrial sector. Due to an enormous increase of generated data and the development of specialized hardware, new decentralized paradigms such as fog computing arised to overcome shortcomings of centralized cloud-only approaches. However, current undertakings are focused on static deployments of standalone services, which is insufficient for geo-distributed applications that are composed of multiple event-driven functions. In this paper, we present StreamPipes Edge Extensions (SEE), a novel contribution to the open source IIoT toolbox Apache StreamPipes. With SEE, domain experts are able to create stream processing pipelines in a graphical editor and to assign individual pipeline elements to available edge nodes, while underlying provisioning and deployment details are abstracted by the framework. The main contributions are (i) a fog cluster management model to represent computing node characteristics, (ii) a node controller for pipeline element life cycle management and (iii) a management framework to deploy event-driven functions to registered nodes. Our approach was validated in a real industrial setup showing low overall overhead of SEE as part of a robot-assisted product quality inspection use case.
Accessing continuous time series data from various machines and sensors is a crucial task to enable data-driven decision making in the Industrial Internet of Things (IIoT). However, connecting data from industrial machines to real-time analytics software is still technically complex and time-consuming due to the heterogeneity of protocols, formats and sensor types. To mitigate these challenges, we present StreamPipes Connect, targeted at domain experts to ingest, harmonize, and share time series data as part of our industry-proven open source IIoT analytics toolbox StreamPipes. Our main contributions are (i) a semantic adapter model including automated transformation rules for pre-processing, and (ii) a distributed architecture design to instantiate adapters at edge nodes where the data originates. The evaluation of a conducted user study shows that domain experts are capable of connecting new sources in less than a minute by using our system. The presented solution is publicly available as part of the open source software Apache StreamPipes.
Distributed publish/subscribe systems are an enabling technology for Industrial Internet of Things applications. While the number of sensors increases, network bandwidth becomes a bottleneck. Existing solutions typically aim to reduce network load either by pre-processing events directly on the edge or by aggregating events into larger batches. However, these approaches are rather static and do not adequately account for the application requirements of subscribers or the actual values of sensor measurements. This paper introduces methods for publish/subscribe systems to dynamically adapt payloads of events at runtime based on i) different data reduction and transformation strategies, ii) a wrapper solution around existing message brokers and iii) a semantics-based event schema registry. Consumers are able to subscribe to various quality levels and receive virtual events, that are reconstructed directly at the subscriber based on knowledge from the semantic model and dynamic decision rules. Our evaluation shows that the concept of virtual events can reduce the network load between publishers, the message broker and subscribers compared to multiple investigated compression techniques.
Newly arising IoT-driven use cases often require low-latency anaiytics to derive time-sensitive actions, where a centralized cloud approach is not applicable. An emerging computing paradigm, referred to as fog computing, shifts the focus away from the central cloud by offloading specific computational parts of analytical stream processing pipelines (SPP) towards the edge of the network, thus leveraging existing resources close to where data is generated. However, in scenarios of mobile edge nodes, the inherent context changes need to be incorporated in the underlying fog cluster management, thus accounting for the dynamics by relocating certain processing elements of these SPP. This paper presents our initial work on a conceptual architecture for context-aware and dynamic management of SPP in the fog. We provide preliminary results, showing the general feasibility of relocating processing elements according to changes in the geolocation.
While todays’ stream processing applications are typically deployed in the cloud, newly arising use cases in the context of Internet of Things (IoT) often require low-latency analytics to derive timesensitive actions. A common approach, referred to as fog computing, shifts the focus away from the cloud by o oading speci c parts of the analytical pipelines in closer proximity of the spatially distributed devices at the edge. However, this requires mechanisms for context-aware deployment, scale, or monitoring both in the cloud as well as the fog landscape. This work explores the challenges with respect to dynamically managing distributed stream processing pipelines in heterogeneous fog computing infrastructures. CCS Concepts • Computer systems organization → Fog computing;
In the era of spatio-temporal big data, geographic information systems have to deal with a myriad of big data induced challenges such as scalability, flexibility or fault-tolerance. Furthermore, the rapid evolution of the underlying, occasionally competing big data ecosystems inevitably needs to be taken into account from the early system design phase. In order to generate valuable knowledge from spatio-temporal big data, a holistic approach manifested in an appropriate architectural design is necessary, which is a non-trivial task with regards to the tremendous design space. Therefore, we present the conceptual architectural framework of BigGIS, a predictive and prescriptive spatio-temporal analytics platform, that integrates big data analytics, semantic web technologies and visual analytics methodologies in our continuous refinement model.
Geographic information systems (GIS) are important for decision support based on spatial data. Due to technical and economical progress an ever increasing number of data sources are available leading to a rapidly growing fast and unreliable amount of data that can be beneficial (1) in the approximation of multivariate and causal predictions of future values as well as (2) in robust and proactive decision-making processes. However, today's GIS are not designed for such big data demands and require new methodologies to effectively model uncertainty and generate meaningful knowledge. As a consequence, we introduce BigGIS, a predictive and prescriptive spatio-temporal analytics platform, that symbiotically combines big data analytics, semantic web technologies and visual analytics methodologies. We present a novel continuous refinement model and show future challenges as an intermediate result of a collaborative research project into big data methodologies for spatio-temporal analysis and design for a big data enabled GIS.
Geographic information systems (GIS) are important for decision support based on spatial data. Due to technical and economical progress an ever increasing number of data sources are available leading to a rapidly growing fast and unreliable amount of data that can be beneficial (1) in the approximation of multivariate and causal predictions of future values as well as (2) in robust and proactive decision-making processes. However, today's GIS are not designed for such big data demands and require new methodologies to effectively model uncertainty and generate meaningful knowledge. As a consequence, we introduce BigGIS, a predictive and prescriptive spatio-temporal analytics platform, that symbiotically combines big data analytics, semantic web technologies and visual analytics methodologies. We present a novel continuous refinement model and show future challenges as an intermediate result of a collaborative research project into big data methodologies for spatio-temporal analysis and design for a big data enabled GIS.