Insight into differences between different implementations of a process provides valuable information for improvement. Process comparison approaches leverage event data on process executions to provide such insight. However, state-of-the-art procedural methods are often limited to local differences considering activities executed within a limited number of steps (e.g., directly following activities). Thereby, detecting differences which, for instance, relate early steps of a process execution to its outcome remains challenging. In contrast, rule-based declarative approaches can detect global differences with respect to distant activities; yet they are limited by the complexity of the rule templates employed . Moreover, they are prone to yield fragmented diagnostics. If a subprocess occurs more frequently in one process variant, these approaches typically report each activity contained. In this work, we therefore propose a process comparison approach that detects aggregated likelihood differences for global control-flow patterns. To this end, we decompose the difference detection task into subprocesses induced by co-occurring activities. Using Earth Mover’s Distance, we identify differences within individual subprocesses independent of predefined rule templates. We then aggregate and combine subprocesses which distinguish the process variants. By exploiting relations among subprocesses, we retrieve maximal differences affecting many activities. Reducing fragmentation caused by choice-induced frequency differences, we additionally complement these maximal differences. To compare the sensitivity of our difference detection method to existing approaches, we devise a quantitative evaluation framework. Moreover, we demonstrate the effectiveness of our method on a public, real-life event log. Ultimately, the evaluation shows that our method is accurate and capable of providing coherent, global diagnostics.
Business processes drive the value creation at companies requiring them to constantly monitor and improve the former. The field of Process Comparison (PC) offers promising approaches to gain insight into differences between variants of a process that one can leverage to improve the latter. For example, one might consider the same process at different points in time or at different sites. Recent PC methods consider event logs containing data on real-life process executions the single source of truth. However, there often exist additional specifications that can be represented as Petri nets. In this paper, we propose an approach that leverages a given Petri net to compare two event logs in a hierarchical manner. To this end, we decompose the provided net into subprocesses and extract data on their executions from the event logs. Based on these executions, we exemplify how one can flexibly assess different aspects of a process (e.g., control flow, performance, or conformance). Using statistical tests, we eventually detect differences between subprocesses with respect to a selected aspect. Despite the approach is mostly agnostic to the decomposition applied, we present a decomposition strategy that we deem particularly suitable for PC. For this purpose, we consider the ned Process Structure Tree of a Petri net and propose a novel preprocessing approach to improve the final decomposition. We implemented the approach in ProM and evaluate it in a real-life case study.
Traditional process models like Petri nets effectively describe the control flow of processes but fail to capture stochastic information such as choice likelihoods. To address this, Stochastic Labeled Petri Nets (SPNs) have recently gained attention, extending Petri nets with transition weights that allow to associate executions with probabilities. The language of an SPN thereby becomes a probability distribution over traces (i.e., sequences of activities). To assess an SPN's quality, Earth Mover's Stochastic Conformance (EMSC) emerged as a natural metric that measures the similarity of the SPN's trace distribution to the observed real-world distribution. In this paper, we propose a locally optimal approach for fine-tuning (or finding) transitions weights to maximize an SPN's EMSC. Leveraging the relationship between EMSC and the Wasserstein distance, which recently gained attention as a loss function in machine learning, we compute subgradients for EMSC to optimize transition weights via subgradient descent. Besides, we propose a straightforward solution to handle models that allow for infinitely many traces. Our optimization approach is broadly applicable for EMSC that is, for EMSC using arbitrary trace-to-trace distances-unlike existing works that either to not explicitly consider EMSC or only special variants. We demonstrate the applicability of our approach on several real-life event logs and discovery algorithms, comparing it to state-of-the-art stochastic process discovery methods and a recent full automated simulation approach.
AbstractThe Internet of Production (IoP) promises to be the answer to major challenges facing the Industrial Internet of Things (IIoT) and Industry 4.0. The lack of inter-company communication channels and standards, the need for heightened safety in Human Robot Collaboration (HRC) scenarios, and the opacity of data-driven decision support systems are only a few of the challenges we tackle in this chapter. We outline the communication and data exchange within the World Wide Lab (WWL) and autonomous agents that query the WWL which is built on the Digital Shadows (DS). We categorize our approaches into machine level, process level, and overarching principles. This chapter surveys the interdisciplinary work done in each category, presents different applications of the different approaches, and offers actionable items and guidelines for future work.The machine level handles the robots and machines used for production and their interactions with the human workers. It covers low-level robot control and optimization through gray-box models, task-specific motion planning, and optimization through reinforcement learning. In this level, we also examine quality assurance through nonintrusive real-time quality monitoring, defect recognition, and quality prediction. Work on this level also handles confidence, verification, and validation of re-configurable processes and reactive, modular, transparent process models. The process level handles the product life cycle, interoperability, and analysis and optimization of production processes, which is overall attained by analyzing process data and event logs to detect and eliminate bottlenecks and learn new process models. Moreover, this level presents a communication channel between human workers and processes by extracting and formalizing human knowledge into ontology and providing a decision support by reasoning over this information. Overarching principles present a toolbox of omnipresent approaches for data collection, analysis, augmentation, and management, as well as the visualization and explanation of black-box models.
Process discovery algorithms learn process models from executed activity sequences, describing concurrency, causality, and conflict. Concurrent activities require observing multiple permutations, increasing data requirements, especially for processes with concurrent subprocesses such as hierarchical, composite, or distributed processes. While process discovery algorithms traditionally use sequences of activities as input, recently introduced object-centric process discovery algorithms can use graphs of activities as input, encoding partial orders between activities. As such, they contain the concurrency information of many sequences in a single graph. In this paper, we address the research question of reducing process discovery data requirements when using object-centric event logs for process discovery. We classify different real-life processes according to the control-flow complexity within and between subprocesses and introduce an evaluation framework to assess process discovery algorithm quality of traditional and object-centric process discovery based on the sample size. We complement this with a large-scale production process case study. Our results show reduced data requirements, enabling the discovery of large, concurrent processes such as manufacturing with little data, previously infeasible with traditional process discovery. Our findings suggest that object-centric process mining could revolutionize process discovery in various sectors, including manufacturing and supply chains.
Event data, often stored in the form of event logs, serve as the starting point for process mining and other evidence-based process improvements. However, event data in logs are often tainted by noise, errors, and missing data. Recently, a novel body of research has emerged, with the aim to address and analyze a class of anomalies known as uncertainty-imprecisions quantified with meta-information in the event log. This paper illustrates an extension of the XES data standard capable of representing uncertain event data. Such an extension enables input, output, and manipulation of uncertain data, as well as analysis through the process discovery and conformance checking approaches available in literature.
With the advent of Industry 4.0, increasing amounts of data on operational processes (e.g., manufacturing processes) become available. These processes can involve hundreds of different materials for a relatively small number of manufactured special-purpose machines rendering classical process discovery and analysis techniques infeasible. However, in contrast to most standard business processes, additional structural information is often available—for example, Bills of Materials (BOMs), listing the required materials, or Multi-level Manufacturing Bills of Materials (M 2 BOMs), which additionally show the material composition. This work investigates how structural information given by Multi-level Bills of Materials (M 2 BOMs) can be integrated into a top-down operational process analysis framework to improve special-purpose machine manufacturing processes. The approach is evaluated on industrial-scale printer assembly data provided by Heidelberger Druckmaschinen AG .
Among the many sources of event data available today, a prominent one is user interaction data. User activity may be recorded during the use of an application or website, resulting in a type of user interaction data often called click data. An obstacle to the analysis of click data using process mining is the lack of a case identifier in the data. In this paper, we show a case and user study for event-case correlation on click data, in the context of user interaction events from a mobility sharing company. To reconstruct the case notion of the process, we apply a novel method to aggregate user interaction data in separate user sessions-interpreted as cases-based on neural networks. To validate our findings, we qualitatively discuss the impact of process mining analyses on the resulting well-formed event log through interviews with process experts.
Process mining is a scientific discipline that analyzes event data, often collected in databases called event logs. Recently, uncertain event logs have become of interest, which contain non-deterministic and stochastic event attributes that may represent many possible real-life scenarios. In this paper, we present a method to reliably estimate the probability of each of such scenarios, allowing their analysis. Experiments show that the probabilities calculated with our method closely match the true chances of occurrence of specific outcomes, enabling more trustworthy analyses on uncertain data.
Modern software systems are able to record vast amounts of user actions, stored for later analysis. One of the main types of such user interaction data is click data: the digital trace of the actions of a user through the graphical elements of an application, website or software. While readily available, click data is often missing a case notion: an attribute linking events from user interactions to a specific process instance in the software. In this paper, we propose a neural network-based technique to determine a case notion for click data, thus enabling process mining and other process analysis techniques on user interaction data. We describe our method, show its scalability to datasets of large dimensions, and we validate its efficacy through a user study based on the segmented event log resulting from interaction data of a mobility sharing company. Interviews with domain experts in the company demonstrate that the case notion obtained by our method can lead to actionable process insights.
The rapid increase in generation of business process models in the industry has raised the demand on the development of process model matching approaches. In this paper, we introduce a novel optimization-based business process model matching approach which can flexibly incorporate both the behavioral and label information of processes for the identification of correspondences between activities. Given two business process models, we achieve our goal by defining an integer linear program which maximizes the label similarities among process activities and the behavioral similarity between the process models. Our approach enables the user to determine the importance of the local label-based similarities and the global behavioral similarity of the models by offering the utilization of a predefined weighting parameter, allowing for flexibility. Moreover, extensive experimental evaluation performed on three real-world datasets points out the high accuracy of our proposal, outperforming the state of the art.
The discipline of process mining aims to study processes in a data-driven manner by analyzing historical process executions, often employing Petri nets. Event data, extracted from information systems (e.g. SAP), serve as the starting point for process mining. Recently, novel types of event data have gathered interest among the process mining community, including uncertain event data. Uncertain events, process traces and logs contain attributes that are characterized by quantified imprecisions, e.g., a set of possible attribute values. The PROVED tool helps to explore, navigate and analyze such uncertain event data by abstracting the uncertain information using behavior graphs and nets, which have Petri nets semantics. Based on these constructs, the tool enables discovery and conformance checking.
Process discovery from event logs as well as process prediction using process models at runtime are increasingly important aspects to improve the operation of digital twins of complex systems. The integration of process mining functionalities with model-driven digital twin architectures raises the question which models are important for the model-driven engineering of digital twins at designtime and at runtime. Currently, research on process mining and model-driven digital twins is conducted in different research communities. Within this position paper, we motivate the need for the holistic combination of both research directions to facilitate harnessing the data and the models of the systems of the future at runtime. The presented position is based upon continuous discussions, workshops, and joint research between process mining experts and software engineering experts in the Internet of Production excellence cluster. We aim to motivate further joint research into the combination of process mining techniques with model-driven digital twins to efficiently combine data and models at runtime.
The strong impulse to digitize processes and operations in companies and enterprises have resulted in the creation and automatic recording of an increasingly large amount of process data in information systems. These are made available in the form of event logs. Process mining techniques enable the process-centric analysis of data, including automatically discovering process models and checking if event data conform to a given model. In this paper, we analyze the previously unexplored setting of uncertain event logs. In such event logs uncertainty is recorded explicitly, i.e., the time, activity and case of an event may be unclear or imprecise. In this work, we define a taxonomy of uncertain event logs and models, and we examine the challenges that uncertainty poses on process discovery and conformance checking. Finally, we show how upper and lower bounds for conformance can be obtained by aligning an uncertain trace onto a regular process model.
Operational processes in production, logistics, material handling, maintenance, etc., are supported by cyber-physical systems combining hardware and software components. As a result, the digital and the physical world are closely aligned, and it is possible to track operational processes in detail (e.g., using sensors). The abundance of event data generated by today's operational processes provides opportunities and challenges for process mining techniques supporting process discovery, performance analysis, and conformance checking. Using existing process mining tools, it is already possible to automatically discover process models and uncover performance and compliance problems. In the DFG-funded Cluster of Excellence "Internet of Production" (IoP), process mining is used to create "digital shadows" to improve a wide variety of operational processes. However, operational processes are dynamic, distributed, and complex. Driven by the challenges identified in the IoP cluster, we work on novel techniques for comparative process mining (comparing process variants for different products at different locations at different times), object-centric process mining (to handle processes involving different types of objects that interact), and forward-looking process mining (to explore "What if?" questions). By addressing these challenges, we aim to develop valuable "digital shadows" that can be used to remove operational friction.
The real-time prediction of business processes using historical event data is an important capability of modern business process monitoring systems. Existing process prediction methods are able to also exploit the data perspective of recorded events, in addition to the control-flow perspective. However, while well-structured numerical or categorical attributes are considered in many prediction techniques, almost no technique is able to utilize text documents written in natural language, which can hold information critical to the prediction task. In this paper, we illustrate the design, implementation, and evaluation of a novel text-aware process prediction model based on Long Short-Term Memory (LSTM) neural networks and natural language models. The proposed model can take categorical, numerical and textual attributes in event data into account to predict the activity and timestamp of the next event, the outcome, and the cycle time of a running process instance. Experiments show that the text-aware model is able to outperform state-of-the-art process prediction methods on simulated and real-world event logs containing textual data.
Process mining is a discipline which concerns the analysis of execution data of operational processes, the extraction of models from event data, the measurement of the conformance between event data and normative models, and the enhancement of all aspects of processes. Most approaches assume that event data is accurately captured behavior. However, this is not realistic in many applications: data can contain uncertainty, generated from errors in recording, imprecise measurements, and other factors. Recently, new methods have been developed to analyze event data containing uncertainty; these techniques prominently rely on representing uncertain event data by means of graph-based models explicitly capturing uncertainty. In this paper, we introduce a new approach to efficiently calculate a graph representation of the behavior contained in an uncertain process trace. We present our novel algorithm, prove its asymptotic time complexity, and show experimental results that highlight order-of-magnitude performance improvements for the behavior graph construction.
Modern business processes are embedded in a complex environment and, thus, subjected to continuous changes. While current approaches focus on the control flow only, additional perspectives, such as time, are neglected. In this paper, we investigate a more general concept drift detection framework that is based on the Earth Mover's Distance. Our approach is flexible in terms of incorporating additional perspectives thanks to the capability of defining custom feature representations, as well as expressive feature similarity measures. We demonstrate the former by incorporating the time perspective using both a time-binning-based trace descriptor and a suitable similarity measure that considers time and control flow. We evaluate the resulting sliding window detector on different types of control-flow and time drifts, and holistic drifts involving multiple perspectives.
The discipline of process mining deals with analyzing execution data of operational processes, extracting models from event data, checking the conformance between event data and normative models, and enhancing all aspects of processes. Recently, new techniques have been developed to analyze event data containing uncertainty; these techniques strongly rely on representing uncertain event data through graph-based models capturing uncertainty. In this paper we present a novel approach to efficiently compute a graph representation of the behavior contained in an uncertain process trace. We present our new algorithm, analyze its time complexity, and report experimental results showing order-of-magnitude performance improvements for behavior graph construction.
. The increasing digitization of organizations leads to unprecedented amounts of data capturing the behavior of operational processes. On the basis of such data, process mining techniques allow us to obtain a holistic picture of the execution of a company’s processes, and their related events. In particular, production companies aiming at reducing the production cycle time and ensur-ing a high product quality show an increased interest in utilizing process mining in order to identify deviations and bottlenecks in their production processes. In this paper, we present a use case study in which we rigorously investigate how process mining techniques can successfully be applied to real-world data of the car production company e.GO Mobile AG. Furthermore, we present our results facilitating more transparency and valuable insights into the real processes of the company.