Organizations frequently collaborate on processes to deliver services, such as logistics partners working together to ship goods. Improving such partnerships is challenging in settings where data sharing may compromise a company's competitive advantage or is restricted by regulations. We propose FIDES, an approach that leverages process mining to construct federated system dynamics models, enabling organizations to jointly simulate the impact of strategic decisions. By combining local and multi-party computation for controlled information sharing, participants can tailor their level of involvement and data disclosure to suit their security requirements. Experiments on a logistics event log demonstrate that the proposed federated system dynamics simulation enables cross-organizational what-if analysis, capturing the propagation of a demand-surge.
Once specified and enacted, end-to-end machine learning (ML) pipelines, together with contextual information about how artifacts are consumed and produced during execution, constitute valuable assets that can be shared or published for reuse. Indeed, ML pipelines involve multiple repetitive data transformation steps before, during, and after model training. Evaluating alternative configurations at each stage requires continuous monitoring of the workflow, particularly to support informed model selection for deployment. Although existing monitoring solutions record metrics and configurations, these logs are typically expressed using ad hoc data models, and lack support for tracing complete derivation paths from training data to deployed models. To overcome such limitation, we introduced in a previous work a unified, W3C PROV-compatible meta-model for capturing end- to-end provenance of ML pipelines. In this demonstration, we present an interactive system that operationalizes this meta-model through a lightweight provenance exploration interface. The graphical user interface allows user to load a provenance graph, search for nodes of interest, render provenance subgraphs, inspect node metadata and relationships, and execute predefined end-to-end provenance queries over the full pipeline.
Once specified and enacted, end-to-end machine learning (ML) pipelines, together with contextual information about how artifacts are consumed and produced during execution, constitute valuable assets that can be shared or published for reuse. Indeed, ML pipelines involve multiple repetitive data transformation steps before, during, and after model training, and evaluating alternative configurations at each stage requires continuous monitoring of the workflow, particularly to support informed model selection for deployment. Although existing monitoring and experiment-management solutions record metrics and configurations, these logs are typically expressed using ad hoc data models that capture only limited relationships between artifacts and pipeline steps, and lack support for tracing complete derivation paths from training data to deployed models. To overcome such limitation, we propose, in this paper, a unified, W3C PROV-compatible metamodel for capturing end-to-end provenance of ML pipelines. We demonstrate the applicability of the proposed metamodel through a pipeline for detecting challenging behaviours in children with autism spectrum disorder from wearable physiological signals, and we illustrate its usefulness through decision-support queries that relate model behaviour and evaluation metrics to upstream data transformations, configurations, and execution environments, enabling systematic model selection, debugging, and governance across the ML lifecycle.
Object-centric Predictive Monitoring has recently gained attention due to advances in machine learning and rise of Object-Centric Event Logs (OCELs), which comprehensively capture object interactions. This paper presents a modular framework supporting customizable pipelines for predictive analysis across diverse event logs. The framework comprises three core components: Preprocessing (preserving object relationships via graph structures), Graph Embedding Model, and Prediction Model. We experimentally evaluated various combinations of embeddings and predictors on three public OCELs. Results show that no single configuration consistently dominates. However, GAT and Graph Transformer models perform best for predicting remaining time and the number of events. Performance improves with larger embedding and subgraphs, particularly for neuralbased models. Finally, GAT delivered the most stable and highperforming results across all event logs in generalization tests.
Predictive Process Monitoring (PPM) leverages historical data to forecast information about ongoing business processes. Recent methods have utilized advanced deep learning and classical machine learning models. However, the role of semantic information that can be extracted from event logs has been underexplored, although such information has been demonstrated to have significant advantages for other process mining tasks, such as anomaly detection. Therefore, this paper proposes a novel mechanism that aims to exploit semantic information for PPM, particularly by extracting information regarding the status of business objects associated with process instances from event data. We evaluate this mechanism in outcome-oriented and next activity prediction tasks, using state-of-the-art large language models (LLMs) for semantic extraction. Our results show that integrating semantic information improves prediction performance across these tasks. This work demonstrates that utilizing semantic information in PPM has considerable potential, especially in combination with advanced language models.
Object-centric predictive process monitoring explores and utilizes object-centric event logs to enhance process predictions. The main challenge lies in extracting relevant information and building effective models. In this paper, we propose an end-to-end model that predicts future process behavior, focusing on two tasks: next activity prediction and next event time. The proposed model employs a graph attention network to encode activities and their relationships, combined with an LSTM network to handle temporal dependencies. Evaluated on one reallife and three synthetic event logs, the model demonstrates competitive performance compared to state-of-the-art methods.
Event data from business processes evidence their patterns, behaviors, and dysfunctions. Analytics techniques like clustering and sorting can reveal relevant insights, when data are correlated with a single case identifier. However, when multiple entities are involved, unidimensional models are challenged. We introduce a novel method for analyzing business processes involving multiple interacting entity types. Our approach employs embedding representations to capture pairwise similarities among entity types and their interrelationships. An optimization problem encompasses similarity matrices, cross-entity relationship matrices, and embeddings. An iterative algorithm refines this model, yielding embedding representations and cluster assignments for each entity type. Formulating our method across three diverse business scenarios demonstrates its practicality and potential. Our results, through a proof of concept using real-world data, underscore the value of accounting for the multifaceted nature of business processes, showing substantial improvements and qualitative distinctions compared to unidimensional models.
In data science, a significant portion of time is dedicated to data preparation, a task often challenging for business users with limited technical skills. This research delves into using large language models (LLMs), particularly GPT (Generative Pre-trained Transformer), to streamline these data preparation processes. We study the usage of large language models in data cleaning, standardization, and transformation tasks. Our approach diverges from traditional methods by employing GPT as a direct transformation tool, utilizing its advanced reasoning capabilities through prompt engineering. This research aims to facilitate data transformation tasks, complicated by the diversity of data formats and inputs, for both technical and non-technical users, allowing them to accomplish these tasks through descriptive instructions rather than complex coding. We evaluate GPT’s effectiveness for frequent data transformation tasks and compare it against other data transformation tools on a benchmark. Our findings demonstrate the practical utility of GPT models in data transformation and propose prompt template guidelines for intricate data string transformations.
Process mining provides methods to analyse event logs generated by information systems during the execution of processes. It thereby supports the design, validation, and execution of processes in domains ranging from healthcare, through manufacturing, to e-commerce. To explore the regularities of flexible processes that show a large behavioral variability, it was suggested to mine recurrent behavioral patterns that jointly describe the underlying process. Existing approaches to behavioral pattern mining, however, suffer from two limitations. First, they show limited scalability as incremental computation is incorporated only in the generation of pattern candidates, but not in the evaluation of their quality. Second, process analysis based on mined patterns shows limited effectiveness due to an overwhelmingly large number of patterns obtained in practical application scenarios, many of which are redundant. In this paper, we address these limitations to facilitate the analysis of complex, flexible processes based on behavioral patterns. Specifically, we improve COBPAM, our initial behavioral pattern mining algorithm, by an incremental procedure to evaluate the quality of pattern candidates, optimizing thereby its efficiency. Targeting a more effective use of the resulting patterns, we further propose pruning strategies for redundant patterns and show how relations between the remaining patterns are extracted and visualized to provide process insights. Our experiments with diverse real-world datasets indicate a considerable reduction of the runtime needed for pattern mining, while a qualitative assessment highlights how relations between patterns guide the analysis of the underlying process.
Predictive process monitoring approaches aim to make predictions about the future behavior for running instances of business processes, such as next activity or remaining time. Most of these approaches use single object type event logs as if the business process is operating in isolation. Whereas, in an organization, several instances of different processes related to a set of objects can be executed at the same time and may interact with each other. This paper investigate the use of object-centric event logs as they offer information about events and their related objects, allowing access to a global view about the running processes in an organization. We propose an object-centric predictive approach considering interactions between different object types. The proposed approach is evaluated on a publicly available object-centric log. The analysis of the results shows that using additional features (i.e., several object types’ information) can generally help increase prediction performances.
Emails are more than just a means of communication, as they are a valuable source of information about undocumented business activities and processes. In this paper, we examine a solution that leverages machine learning to i) extract business activities from emails, and ii) construct business process instances, which group together these activities involved in achieving a common goal. In addition, we examine how relational learning can exploit the relationship between sub-problems (i) and (ii) to further improve their results. The research results presented in this paper are reproducible, and the recipe and data sets used are freely available to interested readers.
Event logs that are recorded by information systems provide a valuable starting point for the analysis of processes in various domains, reaching from healthcare, through logistics, to e-commerce. Specifically, behavioral patterns discovered from an event log enable operational insights, even in scenarios where process execution is rather unstructured and shows a large degree of variability. While such behavioral patterns capture frequently recurring episodes of a process’ behavior, they are not limited to sequential behavior but include notions of concurrency and exclusive choices. Existing algorithms to discover behavioral patterns are context-agnostic, though. They neglect the context in which patterns are observed, which severely limits the granularity at which behavioral regularities are identified. In this paper, we therefore present an approach to discover contextual behavioral patterns. Contextual patterns may be frequent solely in a certain partition of the event log, which enables fine-granular insights into the aspects that influence the conduct of a process. Moreover, we show how to analyze the discovered contextual behavioral patterns in terms of causal relations between context information and the patterns, as well as correlations between the patterns themselves. A complete analysis methodology leveraging all the tools presented in the paper and supplemented by interpretations guidelines is also provided. Finally, experiments with real-world event logs demonstrate the effectiveness of our techniques in obtaining fine-granular process insights.
Workflows have gained momentum in the last two decades, in industry as well as in modern sciences, as a means for modeling and enacting processes. The value of workflow specifications does not end once they are enacted. Indeed, such specifications encapsulate knowledge that documents business processes or scientific experiments, and are, therefore, worth storing and sharing to ultimately be reused or repurposed. We focus, in this paper, on service-based workflows, in which the steps are implemented by operations that are provided by services. In doing so, we examine two problems that are inherent to workflow reuse, namely search and preservation. The first problem deals with the issue of finding workflows, or fragments thereof, that are of interest to the user. The second problem, on the other hand, deals with the volatility of services. It examines solutions that can be used to curate a workflow specification that utilizes a service that is no longer available. As well as examining the above two problems, we close the paper by discussing future issues that need to be addressed in the field.
Process innovation is assumed to require a more intrinsic rethinking of business processes, which is typically a creative process. Nevertheless, in this creative, prolific process, there can be artifacts derived from rational practices that are capable to provide insightful recommendations. In this work, the authors claim that an event log, a file that registers the execution of the relevant business processes, can be the source of such an artifact. They describe the fundamental elements of two problem formulations, namely the set of alternatives; the set of potential actions that the decision-maker may undertake; the set of points of view (dimensions) from which the potential actions are observed, analyzed, evaluated, compared, etc.; and the problem statement (what is expected to be done with the alternatives) for two cases.
Techniques for process discovery support the analysis of information systems by constructing process models from event logs that are recorded during system execution. In recent years, various algorithms to discover end-to-end process models have been proposed. Yet, they do not cater for domains in which process execution is highly flexible, as the unstructuredness of the resulting models renders them meaningless. It has therefore been suggested to derive insights about flexible processes by mining behavioral patterns, i.e., models of frequently recurring episodes of a process' behavior. However, existing algorithms to mine such patterns suffer from imprecision and redundancy of the mined patterns and a comparatively high computational effort. In this work, we overcome these limitations with a novel algorithm, coined COBPAM (COmbination based Behavioral Pattern Mining). It exploits a partial order on potential patterns to discover only those that are compact and maximal, i.e. least redundant. Moreover, COBPAM exploits that complex patterns can be characterized as combinations of simpler patterns, which enables pruning of the pattern search space. Efficiency is improved further by evaluating potential patterns solely on parts of an event log. Experiments with real-world data demonstrates how COBPAM improves over the state-of-the-art in behavioral pattern mining.
We describe an approach that is able to discover business process activities from emails. In addition, for each activity type, we extract metadata such as the roles of the people exchanging the email, type of the attached documents, or the domains of the mentioned links.
Emails play, in the personal and particularly in the professional context, a central role in activity management. Emails can be harvested and re-engineered for understanding the undocumented business process activities and their corresponding metadata. Our goal in this paper is to recast emails into business activity centric resources. We describe an approach that is able to discover business process activities from emails. In addition, for each activity type, we extract metadata such as the roles of the people exchanging the email, type of the attached documents, or the domains of the mentioned links. In order to extract activities from emails, we compare several popular non-linear classification techniques. Activities are then clustered according to their types, which allows us to construct the metadata for each activity type. We validate our approach using a public email dataset.
Data as a Service (DaaS) is seen as a promising cloud offering for wrangling the overload of information and making it available across cloud platforms anytime and anywhere. While there exist a large number of DaaS providers in the market, each one has a different way to describe its provided services as well as supplied datasets. The lack of a well-defined machine-readable model strongly hinders the automatic selection and composition of DaaSs. This paper presents MoDaaS, a model-driven framework for the modeling and the description of DaaS services. MoDaaS enables DaaS providers to describe their services capabilities and concerns according to a shared ontology, thereafter it enables them to automatically generate service views in order to assist the integration and data exchange between heterogeneous services.
Farouk Toumani合作论文数Blaise Pascale University;computer science ;LIMOS2