ABSTRACTCurrently used software process improvement methods such as the Capability Maturity Model Integration (CMMI) rely in their process assessments on information, which is gathered during interviews, in oral audit sessions, and from quality manuals and process standard reviews. Although valuable information about software processes can be gained in these assessments, the resulting data quality can be improved upon. This paper investigates the potential of process mining to support current software process assessment and improvement approaches. Based on an analysis of CMMI from a process mining perspective, particular CMMI model components are identified for which it is in principle possible to apply process mining techniques. Subsequently, criteria have been defined to select, with respect to these particular CMMI components, software processes for which process mining has an added value. These criteria have been applied in the selection of a particular ‘minable’ software process, that is, a change control process. Subsequently, the results of a case study from industrial practice, on the process mining of a change control board process, are used to illustrate that process mining can provide CMMI assessors with relevant information. This information reflects the actual or ‘real’ software processes in practice, and as such, it offers an excellent basis to support assessors in understanding the ‘actual’ software processes. Copyright © 2014 John Wiley & Sons, Ltd.
Today there are many process mining techniques that, based on an event log, allow for the automatic induction of a process model. The process mining algorithms that are able to deal with incomplete event logs, exceptions, and noise typically have many parameters to tune the algorithm. Therefore, the user needs to select the right parameter setting using a trail-and-error approach. So far, there is no general method available to search for an optimal parameter setting. One of the problems is the lack of negative examples and the omission of a standard measure for the quality of mined process models. Therefore, the so-called k-fold-cv experimental set up as used in the machine learning community cannot be applied directly. This paper describes an adapted version of the k-fold-cv set-up so that it can be used in the context of process mining. Illustrative experimental results of applying this method in combination with the HeuristicsMiner process mining algorithm and three different performance measurements are presented. Using the kfold- cv experimental set-up and an event log with low frequent behavior and noise, it appears possible to find the optimal parameters setting. Another important result is that the simple combination of yes/no parsing of a trace in combination with negative examples based on noise is sufficient for parameter optimization. This makes the framework universal applicable for benchmarking of different process mining algorithms with different process model representation languages.
In this paper the so-called Event Cube is introduced, a multidimensional data structure that can hold information about all business dimensions. Like the data cubes of online analytic processing (OLAP) systems, the Event Cube can be used to improve the business analysis quality by providing immediate results under different levels of abstraction. An exploratory analysis of the application of process mining on multidimensional process data is the focus of this paper. The feasibility and potential of this approach is demonstrated through some practical examples.
One of the aims of process mining is to retrieve a process model from a given event log. However, current techniques have problems when mining processes that contain nontrivial constructs, processes that are low structured and/or dealing with the presence of noise in the event logs. To overcome these problems, a new process representation language is presented in combination with an accompanying process mining algorithm. The most significant property of the new representation language is in the way the semantics of splits and joins are represented; by using so-called split/join frequency tables. This results in easy to understand process models even in the case of non-trivial constructs, low structured domains and the presence of noise. This paper explains the new process representation language and how the mining algorithm works. The algorithm is implemented as a plug-in in the ProM framework. An illustrative example with noise and a real life log of a complex and low structured process are used to explicate the presented approach.
During a software process improvement program, the current state of software development processes is being assessed and improvement actions are being determined. However, these improvement actions are based on process models obtained during interviews and document studies, e.g. quality manuals. Such improvements are scarcely based on the practical way of working in an organization; they do not take into account shortcuts made due to e.g. time pressure. Becoming conscious about the presence of such deviations and understanding their causes and impacts, consequences for particular software process improvement activities in a particular organization could be proposed. This paper reports on the application of process mining techniques to discover shortcomings in the Change Control Board process in an organization during the different lifecycle phases and to determine improvement activities.
A recent trend in technological innovation is towards the development of increasingly multifunctional and complex products to be used within rich socio‐cultural contexts such as the high‐end office, the digital home, and professional or personal healthcare. One important consequence of the development of strongly innovative products is a growing market uncertainty regarding ‘if’, ‘how’, and ‘when’ users can and will adopt such products. Often, it is not even clear to what extent these products are understood and interacted with in the intended manner. The mentioned problems have already become an evident concern in the field, where there is a significant rise in the numbers of seemingly sound products being complained about, signaling a lack of soft reliability. In this paper, we position soft reliability as a growing and critical industrial problem, whose solution requires new academic expertise from various disciplines. We illustrate potential root causes for soft reliability problems, such as discrepancy between the perceptions of users and designers. We discuss the necessary approach to effectively capture subjective feedback data from actual users, e.g. when they contact call centers. Furthermore, we present a novel observation and analysis approach that enables insight into actual product usage, and outline opportunities for combining such objective data with the subjective feedback provided by users. Copyright © 2008 John Wiley & Sons, Ltd.
Process mining techniques attempt to extract non-trivial and useful information from event logs recorded by information systems. For example, there are many process mining techniques to automatically discover a process model based on some event log. Most of these algorithms perform well on structured processes with little disturbances. However, in reality it is difficult to determine the scope of a process and typically there are all kinds of disturbances. As a result, process mining techniques produce spaghetti-like models that are difficult to read and that attempt to merge unrelated cases. To address these problems, we use an approach where the event log is clustered iteratively such that each of the resulting clusters corresponds to a coherent set of cases that can be adequately represented by a process model. The approach allows for different clustering and process discovery algorithms. In this paper, we provide a particular clustering algorithm that avoids over-generalization and a process discovery algorithm that is much more robust than the algorithms described in literature [1]. The whole approach has been implemented in ProM.
A recent trend in technological innovation is towards the developmentof increasingly multifunctional and complex products to be used within rich socio-cultural contexts such asthe high-endoffice,the digitalhome,andprofessionalor personalhealthcare. One important consequence of the development of strongly innovative products is a growing market uncertainty regarding ‘if’, ‘how’, and ‘when’ users can and will adopt such products. Often, it is not even clear to what extent these products are understood and interacted with in the intended manner. The mentioned problems have already become an evident concern in the field, where there is a significant rise in the numbers of seemingly sound products being complained about, signaling a lack of soft reliability. In this paper, we position soft reliability as a growing and critical industrial problem, whose solution requires new academic expertise from various disciplines. We illustrate potential root causes for soft reliability problems, such as discrepancy between the perceptions of users and designers. We discuss the necessary approach to effectively capture subjective feedback data from actual users, e.g. when they contact call centers. Furthermore, we present a novel observation and analysis approach that enables insight into actual product usage, and outline opportunities for combining such objective data with the subjective feedback provided by users. Copyright © 2008 John Wiley & Sons, Ltd.
This demonstration paper describes the ProM process mining tool. Process mining techniques attempt to extract non-trivial and useful process information from so-called event logs. ProM allows for the discovery of different process perspectives (e.g., control-flow, time, resources, and data) and supports related techniques such as control-flow mining, performance analysis, resource analysis, conformance checking, verification, etc. This makes ProM a practical and versatile tool for business process analysis and discovering.
In various application domains there is a desire to compare process models, e.g., to relate an organization-specific process model to a reference model, to find a web service matching some desired service description, or to compare some normative process model with a process model discovered using process mining techniques. Although many researchers have worked on different notions of equivalence (e.g., trace equivalence, bisimulation, branching bisimulation, etc.), most of the existing notions are not very useful in this context. First of all, most equivalence notions result in a binary answer (i.e., two processes are equivalent or not). This is not very helpful because, in real-life applications, one needs to differentiate between slightly different models and completely different models. Second, not all parts of a process model are equally important. There may be parts of the process model that are rarely activated (i.e., “process veins”) while other parts are executed for most process instances (i.e., the “process arteries”). Clearly, differences in some veins of a process are less important than differences in the main arteries of a process. To address the problem, this paper proposes a completely new way of comparing process models. Rather than directly comparing two models, the process models are compared with respect to some typical behavior. This way, we are able to avoid the two problems just mentioned. The approach has been implemented and has been used in the context of genetic process mining. Although the results are presented in the context of Petri nets, the approach can be applied to any process modeling language with executable semantics.
This tool paper describes the functionality of ProM. Version 4.0 of ProM has been released at the end of 2006 and this version reflects recent achievements in process mining. Process mining techniques attempt to extract non-trivial and useful information from so-called "event logs". One element of process mining is control-flow discovery, i.e., automatically constructing a process model (e.g., a Petri net) describing the causal dependencies between activities. Control-flow discovery is an interesting and practically relevant challenge for Petri-net researchers and ProM provides an excellent platform for this. For example, the theory of regions, genetic algorithms, free-choice-net properties, etc. can be exploited to derive Petri nets based on example behavior. However, as we will show in this paper, the functionality of ProM 4.0 is not limited to control-flow discovery. ProM 4.0 also allows for the discovery of other perspectives (e.g., data and resources) and supports related techniques such as conformance checking, model extension, model transformation, verification, etc. This makes ProM a versatile tool for process analysis which is not restricted to model analysis but also includes log-based analysis.
Process mining is to extract business process models from event logs, the mining process is an important learning task. However, the discovery of these processes poses many challenges, including noise, non-local, non-free choice constructs and so on. In the study, we give out the definition of the behavior redundancy degree which is benefit to analyze the behavior conformance. Then, in order to build the optimal the process model, a process mining method based on Discrete Particle Swarm Optimization (DPSO) is presented. The method can take into account the basic Petri net structure and the metrics of behavior conformance and avoid the blindness of building process model. Finally, a DPSO process mining plug-in is developed and a number of event log is tested in the DPSO mining plug-in based on PROM platform. Theoretical analysis and experimental results show that DPSO-based mining method has …
Although there has been a lot of progress in developing process mining algorithms in recent years, no effort has been put in developing a common means of assessing the quality of the models discovered by these algorithms. In this paper, we outline elements of an evaluation framework that is intended to enable (a) process mining researchers to compare the performance of their algorithms, and (b) end users to evaluate the validity of their process mining results. Furthermore, we describe two possible approaches to evaluate a discovered model (i) using existing comparison metrics that have been developed by the process mining research community, and (ii) based on the so-called k-fold-cross validation known from the machine learning community. To illustrate the application of these two approaches, we compared a set of models discovered by different algorithms based on a simple example log.
One of the aims of process mining is to retrieve a process model from an event log. The discovered models can be used as objective starting points during the deployment of process-aware information systems (Dumas et al., eds., Process-Aware Information Systems: Bridging People and Software Through Process Technology. Wiley, New York, 2005) and/or as a feedback mechanism to check prescribed models against enacted ones. However, current techniques have problems when mining processes that contain non-trivial constructs and/or when dealing with the presence of noise in the logs. Most of the problems happen because many current techniques are based on local information in the event log. To overcome these problems, we try to use genetic algorithms to mine process models. The main motivation is to benefit from the global search performed by this kind of algorithms. The non-trivial constructs are tackled by choosing an internal representation that supports them. The problem of noise is naturally tackled by the genetic algorithm because, per definition, these algorithms are robust to noise. The main challenge in a genetic approach is the definition of a good fitness measure because it guides the global search performed by the genetic algorithm. This paper explains how the genetic algorithm works. Experiments with synthetic and real-life logs show that the fitness measure indeed leads to the mining of process models that are complete (can reproduce all the behavior in the log) and precise (do not allow for extra behavior that cannot be derived from the event log). The genetic algorithm is implemented as a plug-in in the ProM framework.
Although there has been a lot of progress in developing pro- cess mining algorithms in recent years, no eort has been put in devel- oping a common means of assessing the quality of the models discovered by these algorithms. In this paper, we outline elements of an evaluation framework that is intended to enable (a) process mining researchers to compare the performance of their algorithms, and (b) end users to evalu- ate the validity of their process mining results. Furthermore, we describe two possible approaches to evaluate a discovered model (i) using existing comparison metrics that have been developed by the process mining re- search community, and (ii) based on the so-called k-fold-cross validation known from the machine learning community. To illustrate the applica- tion of these two approaches, we compared a set of models discovered by dierent algorithms based on a simple example log.