
The dynamic nature of web data brings forward the need for maintaining data versions as well as identifying changes between them. In this paper, we deal with problems regarding understanding evolution, focusing on RDF(S) knowledge bases, as RDF is a de-facto standard for representing data on the web. We argue that revisiting past snapshots or the differences between them is not enough for understanding how and why data evolved. Instead, changes should be treated as first-class citizens. In our view, this involves supporting semantically rich, user-defined changes, called complex changes, as well as identifying the relations between them. In this paper, we present our perspective regarding complex changes, formally define a declarative language for defining complex changes on RDF(S) knowledge bases and present how this language is used to detect complex change instances among dataset versions, which can be queried for analyzing evolution. The approach has been extensively evaluated in terms of language expressivity and detection performance on both artificial and real data.
Geospatial extensions of SPARQL, like GeoSPARQL and stSPARQL, have been defined since 2007, and while several geospatial RDF stores have implemented a substantial part of these extensions, other stores limited their support mostly on point geometry features. A parallel process with the above was that RDF frameworks evolved in an interesting way by presenting a more mature set of geospatial features, such as GeoSPARQL support and including the latest indexing technologies. As a logical consequence, a shift in the use of RDF frameworks is to be expected, from base platforms that users extend to create more complete geospatial RDF stores, to attractive finished RDF solutions for many geospatial applications. Alongside with the ever-increasing size of linked geospatial data that semantic stores need to handle, all the above provided our group the motivation to improve our single-node systems benchmark Geographica, originally defined in 2013. Geographica 2 is more comprehensive, because it now includes new geospatial RDF stores and frameworks, big real-world datasets of many hundred million triples with up to 50 million features of complex geometries, new tests and queries that reveal the scalability of these systems. The augmented and revised real-world workload of Geographica 2 tests the efficiency of primitive spatial functions in RDF stores, their performance in the geocoding scenario against the new Census dataset in addition to many other real use case scenarios and finally includes computation of statistics for geospatial datasets. A more detailed and systematic evaluation is performed using the synthetic workload. The new scalability workload aims at discovering the limits of centralized geospatial RDF stores of various architectures. It employs a set of six well-balanced real-world datasets with highly complex geometries covering many European countries and compares three RDF stores in terms of storage space, bulk loading and query response time. In addition, a special version of the benchmark has been created for systems with limited geospatial functionality and two more systems of this category are introduced along the six systems of the main benchmark, all stressed against point-only subsets of the workloads. Three out of the eight systems use an RDBMS for the persistence layer, while some of them offer a variety of persistence options.
A Cyber-Physical System-of-Systems (CPSoS) can be defined as a System-of-Systems (SoS), where its component systems are Cyber-Physical Systems (CPSs) that have been networked together for achieving a certain higher goal. Therefore, a key viability of any CPSoS is the integration of its CPSs to function as a single integrated system to support a common mission. Although such integration can be achieved relying on the exchange of information among CPSs, only few works have highlighted the importance of considering the quality of such information. Without considering Information Quality (IQ) requirements during the design of CPSoS, CPSs will be vulnerable to faults arising from depending on inaccurate, incomplete, inconsistent, and/or outdated information, which may influence the overall dependability, reliability, and performance of the CPSoS. This paper proposes a model-based approach that offers a novel UML profile, named IQCPSoS (Information Quality for Cyber-Physical System-of-Systems), which contains various stereotypes and tagged values for modeling and analyzing IQ requirements for CPSoS. The profile also proposes a set of constraints expressed in the Object Constraint Language (OCL) to be used for the verification of such models. We evaluate our approach by developing a prototype implementation and test its applicability, usability, and validity for modeling and analyzing IQ requirements for a realistic scenario concerning a Tram system.
Knowledge bases allow data organization and exploration, making easier the data semantic understanding and its use by machines. Traditional strategies for knowledge base construction and augmentation have mostly relied on manual effort or automatic extraction of content from structured and semi-structured sources. In this work, we present DeepEx, a system that autonomously extracts missing attributes of entities in knowledge bases from unstructured text. We use Wikipedia as data source. Given entities on Wikipedia represented by their articles (text and infobox), DeepEx uses a classifier to detect sentences in the articles mentioning the possible missing attributes of the entities and then employs a deep-learning extraction model on those sentences to identify the attributes. The sentence classifier and attribute extractor are built with labels automatically produced by a weak supervision approach using infobox structured information as supervision source. We have compared our strategy with previous approaches to this problem on 29 different attributes from 4 domains. The results showed that our extraction pipeline achieved statistically superior performance in comparison with some baselines and variations of our approach.
We study the combined class of possible keys and functional dependencies over Codd tables under NOT NULL constraints. These constraints can express significant application semantics of data that is compliant with the industry standard SQL. Three major contributions are made. Firstly, the PTIME-complete implication problem is characterized axiomatically by a finite set of Horn rules, algorithmically by a linear time decision procedure, and logically by goal and definite clauses under S-3 logic. Secondly, we establish structural and computational properties of Armstrong tables for the combined class. This includes an algorithm that is guaranteed to compute an Armstrong table of size at most quadratic in that of a minimum-sized table. Thirdly, we conduct an empirical evaluation of a GUI-based implementation of this algorithm. We conclude that Armstrong tables are an effective tool for recognizing domain semantics and should be exploited as early as possible during the acquisition of possible keys and functional dependencies.
This paper describes a program-SPARQL Query Generator (SQG)-which takes as input an OWL ontology, a set of object descriptions in terms of this ontology and an OWL class as the context, and generates relatively large numbers of queries about various types of descriptions of objects expressed in RDF/OWL. The intent is to use SQG in evaluating data representation and retrieval systems from the perspective of OWL semantics coverage. While there are many benchmarks for assessing the efficiency of data retrieval systems, none of the existing solutions for SPARQL query generation focus on the coverage of the OWL semantics. Some are not scalable since manual work is needed for the generation process; some do not consider (or totally ignore) the OWL semantics in the ontology/instance data or rely on large numbers of real queries/datasets that are not readily available in our domain of interest. Our experimental results show that SQG performs reasonably well with generating large numbers of queries and guarantees a good coverage of OWL axioms included in the generated queries.
Ontology alignment plays a key role in the management of heterogeneous data sources and metadata. In this context, various ontology alignment techniques have been proposed to discover correspondences between the entities of different ontologies. This paper proposes a new ontology alignment approach based on a set of rules exploiting the embedding space and measuring clusters of labels to discover the relationship between entities. We tested our system on the OAEI conference complex alignment benchmark track and then applied it to aligning ontologies in a real-world case study. The experimental results show that the combination of word embedding and a measure of dispersion of the clusters of labels, which we call the radius measure, makes it possible to determine, with good accuracy, not only equivalence relations, but also hierarchical relations between entities.
Question Answering based on Knowledge Graphs (KGQA) still faces difficult challenges when transforming natural language (NL) to SPARQL queries. Simple questions only referring to one triple are answerable by most QA systems, but more complex questions requiring complex queries containing subqueries or several functions are still a tough challenge within this field of research. Evaluation results of QA systems therefore also might depend on the benchmark dataset the system has been tested on. For the purpose to give an overview and reveal specific characteristics, we examined currently available KGQA datasets regarding several challenging aspects. This paper presents a detailed look into the datasets and compares them in terms of challenges a KGQA system is facing.
Complaints about finished products are a major challenge for companies in the medical technology industry, where product quality is directly related to public health and therefore strictly regulated. In this paper, we examine how available data can be used to provide automated support to the complaint handling processes in the medical technology companies. We identify the automation potentials in the 8D reference process for complaint management and discuss their organizational and technical challenges. Using data from a large manufacturer of medical products, we show how partial process automation can be achieved in practice by designing, implementing, and evaluating a deep learning-based prototype for automatically suggesting a likely error code for future complaints, given their textual description. Our approach is able to assign the correct error code for more than 75% of all cases and outperforms the conventional classification approaches used as a baseline comparison. Our results show that partial automation of a complaint management process by means of deep learning can be achieved in practice.
The operational backbone of modern organizations is the target of business process management, where business process models are produced to describe how the organization should react to events and coordinate the execution of activities so as to satisfy its business goals. At the same time, operational decisions are made by considering internal and external contextual factors, according to decision models that are typically based on declarative, rule-based specifications that describe how input configurations correspond to output results. The increasing importance and maturity of these two intertwined dimensions, those of processes and decisions, have led to a wide range of data-aware models and associated methodologies, such as BPMN for processes and DMN for operational decisions. While it is important to analyze these two aspects independently, it has been pointed out by several authors that it is also crucial to analyze them in combination. In this paper, we provide a native, formal definition of DBPMN models, namely data-aware and decision-aware processes that build on BPMN and DMN S-FEEL, illustrating their use and giving their formal execution semantics via an encoding into Data Petri nets (DPNs). By exploiting this encoding, we then build on previous work in which we lifted the classical notion of soundness of processes to this richer, data-aware setting, and show how the abstraction and verification techniques that were devised for DPNs can be directly used for DBPMN models. This paves the way towards even richer forms of analysis, beyond that of assessing soundness, that are based on the same technique.
For Read-Write Linked Data, an environment of reasoning and RESTful interaction, we investigate the use of the Guard-Stage-Milestone approach for specifying and executing user agents. We present an ontology to specify user agents. Moreover, we give operational semantics to the ontology in a rule language that allows for executing user agents on Read-Write Linked Data. We evaluate our approach formally and regarding performance. Our work shows that despite different assumptions of this environment in contrast to the traditional environment of workflow management systems, the Guard-Stage-Milestone approach can be transferred and successfully applied on the web of Read-Write Linked Data.
Ontology matching has become one of the main research topics to address problems related to semantic interoperability on the web. The main goal is to find ways to make different ontologies interoperable. Due to the high heterogeneity in the knowledge representation of each ontology, several matchers are proposed in the literature, each seeking to capture a specific aspect of the ontology. Generally, different matchers are complementary and none stand out in all test cases. In this paper, we present a meta-matching approach employing the prey–predator meta-heuristic in order to define a set of weights to find the best possible result from a set of matchers. The approach was evaluated on the Ontology Alignment Evaluation Initiative benchmark and results showed that the prey–predator algorithm is competitive with other popular algorithms as it achieves an average f-measure of 0.91.
The ever-increasing attention of process mining (PM) research to the logs of low structured processes and of non-process-aware systems (e.g., ERP, IoT systems) poses a number of challenges. Indeed, in such cases, the risk of obtaining low-quality results is rather high, and great effort is needed to carry out a PM project, most of which is usually spent in trying different ways to select and prepare the input data for PM tasks. Two general AI-based strategies are discussed in this paper, which can improve and ease the execution of PM tasks in such settings: (a) using explicit domain knowledge and (b) exploiting auxiliary AI tasks. After introducing some specific data quality issues that complicate the application of PM techniques in the above-mentioned settings, the paper illustrates these two strategies and the results of a systematic review of relevant literature on the topic. Finally, the paper presents a taxonomical scheme of the works reviewed and discusses some major trends, open issues and opportunities in this field of research.
A current research problem in the area of business process management deals with the specification and checking of constraints on resources (e.g., users, agents, autonomous systems, etc.) allowed to be committed for the execution of specific tasks. Indeed, in many real-world situations, role assignments are not enough to assign tasks to the suitable resources. It could be the case that further requirements need to be specified and satisfied. As an example, one would like to avoid that employees that are relatives are assigned to a set of critical tasks in the same process in order to prevent fraud. The formal specification of a business process and its related access control constraints is obtained through a decoration of a classic business process with roles, users, and constraints on their commitment. As a result, such a process specifies a set of tasks that need to be executed by authorized users with respect to some partial order in a way that all authorization constraints are satisfied. Controllability refers in this case to the capability of executing the process satisfying all these constraints, even when some process components, e.g., gateway conditions, can only be observed, but not decided, by the process engine responsible of the execution. In this paper, we propose conditional constraint networks with decisions (CCNDs) as a model to encode business processes that involve access control and conditional branches that may be both controllable and uncontrollable. We define weak, strong, and dynamic controllability of CCNDs as two-player games, classify their computational complexity, and discuss strategy synthesis algorithms. We provide an encoding from the business processes we consider here into CCNDs to exploit off-the-shelf their strategy synthesis algorithms. We introduce $$\textsc {Zeta}$$ , a tool for checking controllability of CCNDs, synthesizing execution strategies, and executing controllable CCNDs, by also supporting user interactivity. We use $$\textsc {Zeta}$$ to compare with the previous research, provide a new experimental evaluation for CCNDs, and discuss limitations.
Business processes are often specified in descriptive or normative models. Both types of models should adhere to internal and external regulations, such as company guidelines or laws. Employing compliance checking techniques, it is possible to verify process models against rules. While traditionally compliance checking focuses on well-structured processes, we address case management scenarios. In case management, knowledge workers drive multi-variant and adaptive processes. Our contribution is based on the fragment-based case management approach, which splits a process into a set of fragments. The fragments are synchronized through shared data but can, otherwise, be dynamically instantiated and executed. We formalize case models using Petri nets. We demonstrate the formalization for design-time and run-time compliance checking and present a proof-of-concept implementation. The application of the implemented compliance checking approach to a use case exemplifies its effectiveness while designing a case model. The empirical evaluation on a set of case models for measuring the performance of the approach shows that rules can often be checked in less than a second.
The growth of Web of Data led to the development of dataset recommendation methodologies, which automate the discovery of datasets that may contain same or related instances (i.e., objects), in order to be used as input for several tasks including Link Discovery. The recommendation process takes as input one dataset (or any tripleset) and proposes other datasets which are the most likely to contain related instances. Existing recommenders determine the relevance between datasets by comparing their textual and structural similarity or by examining existing links among them. In this paper, we determine relevancy by comparing the geospatial relatedness of triplesets containing instances belonging to spatial classes (that is, classes containing instances whose locations are georeferenced by point geometries) based on the hypothesis that pairs of classes whose instances present similar spatial distribution are likely to contain semantically related instances. The proposed methodology builds summaries that capture the spatial distribution of classes. It utilizes the summaries, first, to rule out irrelevant (to an input class) classes by applying spatial filters and, then, to rank the remaining classes by applying a geospatial relatedness measure, so as the top ranked classes are more probable to contain related instances. The methodology's evaluation contains an exploration of Web of Data spatial classes characteristics and a discussion of the experiment results that validate our hypothesis. We show that the spatial filtering reduces effectively and efficiently up to 99% the search space for relevant classes in Web of Data and that the proposed geospatial relatedness measures generate ranked lists of recommended classes with 62% mean average precision, approximately 35% higher than simple baselines.
Process event data is usually stored either in a sequential process event log or in a relational database. While the sequential, single-dimensional nature of event logs aids querying for (sub)sequences of events based on temporal relations such as directly/eventually-follows, it does not support querying multi-dimensional event data of multiple related entities. Relational databases allow storing multi-dimensional event data but existing query languages do not support querying for sequences or paths of events in terms of temporal relations. In this paper, we propose a general data model for multi-dimensional event data based on labeled property graphs that allows storing structural and temporal relations in a single, integrated graph-based data structure in a systematic way. We provide semantics for all concepts of our data model, and generic queries for modeling event data over multiple entities that interact synchronously and asynchronously . The queries allow for efficiently converting large real-life event data sets into our data model and we provide 5 converted data sets for further research. We show that typical and advanced queries for retrieving and aggregating such multidimensional event data can be formulated and executed efficiently in the existing query language Cypher, giving rise to several new research questions. Specifically aggregation queries on our data model enable process mining over multiple interrelated entities using off-the-shelf technology.