Digital health data quality is a critical concern in the healthcare industry, jeopardizing the secondary use of data for revolutionizing population health, and hindering patient care and organizational outcomes. Limited published evidence exists for explaining why these data quality issues emerge. The Odigos framework is a notable exception asserting that data quality issues emerge from three worlds: material world (e.g., technology artifact), personal world (e.g., technology users/use), and social world (e.g., organizations/ institutions) but has yet to systematically unpack the elements within these worlds. Through deductive and inductive analysis of interview data from a case study of the Emergency Department of Australia's first large digital hospital, we apply and extend the Odigos framework by identifying elements emanating from the three worlds and their interrelationships as root causes of data quality issues. These elements can then be used by hospitals to develop strategies to proactively improve their digital health data quality.
BACKGROUND:The promise of digital health is principally dependent on the ability to electronically capture data that can be analyzed to improve decision-making. However, the ability to effectively harness data has proven elusive, largely because of the quality of the data captured. Despite the importance of data quality (DQ), an agreed-upon DQ taxonomy evades literature. When consolidated frameworks are developed, the dimensions are often fragmented, without consideration of the interrelationships among the dimensions or their resultant impact. OBJECTIVE:The aim of this study was to develop a consolidated digital health DQ dimension and outcome (DQ-DO) framework to provide insights into 3 research questions: What are the dimensions of digital health DQ? How are the dimensions of digital health DQ related? and What are the impacts of digital health DQ? METHODS:Following the PRISMA (Preferred Reporting Items for Systematic Reviews and Meta-Analyses) guidelines, a developmental systematic literature review was conducted of peer-reviewed literature focusing on digital health DQ in predominately hospital settings. A total of 227 relevant articles were retrieved and inductively analyzed to identify digital health DQ dimensions and outcomes. The inductive analysis was performed through open coding, constant comparison, and card sorting with subject matter experts to identify digital health DQ dimensions and digital health DQ outcomes. Subsequently, a computer-assisted analysis was performed and verified by DQ experts to identify the interrelationships among the DQ dimensions and relationships between DQ dimensions and outcomes. The analysis resulted in the development of the DQ-DO framework. RESULTS:The digital health DQ-DO framework consists of 6 dimensions of DQ, namely accessibility, accuracy, completeness, consistency, contextual validity, and currency; interrelationships among the dimensions of digital health DQ, with consistency being the most influential dimension impacting all other digital health DQ dimensions; 5 digital health DQ outcomes, namely clinical, clinician, research-related, business process, and organizational outcomes; and relationships between the digital health DQ dimensions and DQ outcomes, with the consistency and accessibility dimensions impacting all DQ outcomes. CONCLUSIONS:The DQ-DO framework developed in this study demonstrates the complexity of digital health DQ and the necessity for reducing digital health DQ issues. The framework further provides health care executives with holistic insights into DQ issues and resultant outcomes, which can help them prioritize which DQ-related problems to tackle first.
Whilst digital health data provides great benefits for improved and effective patient care and organisational outcomes, the quality of digital health data can sometimes be a significant issue. Healthcare providers are known to spend a significant amount of time on assessing and cleaning data. To address this situation, this paper presents six Digital Health Data Imperfection Patterns that provide insight into data quality issues of digital health data, their root causes, their impact, and how these can be detected. Using the CRISP-DM methodology, we demonstrate the utility and pervasiveness of the patterns at the emergency department of Australia's major tertiary digital hospital. The pattern collection can be used by health providers to identify and prevent key digital health data quality issues contributing to reliable insights for clinical decision making and patient care delivery. The patterns also provide a solid foundation for future research in digital health through its identification of key data quality issues, root causes, detection techniques, and terminology.
Since its emergence over two decades ago, process mining has flourished as a discipline, with numerous contributions to its theory, widespread practical applications, and mature support by commercial tooling environments. However, its potential for significant organisational impact is hampered by poor quality event data. Process mining starts with the acquisition and preparation of event data coming from different data sources. These are then transformed into event logs, consisting of process execution traces including multiple events. In real-life scenarios, event logs suffer from significant data quality problems, which must be recognised and effectively resolved for obtaining meaningful insights from process mining analysis. Despite its importance, the topic of data quality in process mining has received limited attention. In this paper, we discuss the emerging challenges related to process-data quality from both a research and practical point of view. Additionally, we present a corresponding research agenda with key research directions.
Through the application of process mining, organisations can improve their business processes by leveraging data recorded as a result of the performance of these processes. Over the past two decades, the field of process mining evolved considerably, offering a rich collection of analysis techniques with different objectives and characteristics. Despite the advances in this field, a solid statistical foundation is still lacking. Such a foundation would allow analysis outcomes to be found or judged using the notion of statistical significance, thus providing a more objective way to assess these outcomes. This article contributes several statistical tests and association measures that treat process behaviour as a variable. The sensitivity of these tests to their parameters is evaluated and their applicability is illustrated through the use of real-life event logs. The presented tests and measures constitute a key contribution to a statistical foundation for process mining.
Stresses and temptations in the workplace foster employee behaviour that is less than desirable. The term weasel has been used to describe employees that exhibit a variety of undesirable behaviour at work, including taking undeserved credit, performing below expectation, shirking work, and making co-workers look bad. While this behaviour has traditionally been hard to detect, contemporary systems record many of our work-related actions and the resulting event logs can be subjected to process mining analysis. In this paper we focus on detecting weasels through the evidence they leave behind in such event logs. We capture a variety of weasel behaviours in the form of patterns and suggest how process mining can be used to unearth this behaviour. The patterns are validated through a survey with relevant stakeholders.
Process mining provides a range of methods and techniques to analyse business processes through information stored in so-called event logs. The richer these event logs and the higher quality they are, the more insights we can obtain. Till now, information in the form of unstructured text, e.g. notes, comments, reviews, and posts, is not fully and systematically exploited for the purposes of log enrichment. In this paper, we introduce Text2EL, a two-phase event log enrichment approach based on unstructured text. In Phase 1, events, case attributes, and event attributes are extracted from unstructured text associated with organisational processes. In Phase 2, the extracted events and attributes are semantically and contextually validated before enriching the event log. Our approach applies techniques from natural language processing, sentence embeddings, and contextual and expression validation. We evaluated the completeness, concordance, and correctness of an enriched event log through experiments with a real-life healthcare data set. The experiments showed the feasibility and applicability of our approach.
Process mining provides analytical tools and methods which can distil insights about process behaviour from big process-related data. Yet challenges relating to the impact of poor quality data on event logs, the input to process mining analyses, remain. Despite researchers raising concerns about event log data quality, event log preparation is, in practice, generally handled mechanistically, focusing on fixing symptoms rather than on uncovering the root causes of event log data quality issues. To address this, we introduce the Odigos (Greek for "guide") framework. Based on semiotics and Peircean abductive reasoning, the Odigos framework facilitates an informed way of dealing with data quality issues in event logs. Odigos supports both prognostic (foreshadowing potential quality issues) and diagnostic (identifying root causes of discovered quality issues) approaches. We examine in depth how the framework supports a detailed root-cause analysis of a well-known collection of event log imperfection patterns.
By incorporating aspects of coordination and collaboration, workflow implementations of information systems require a sound conceptualisation of business processing semantics. Traditionally, the success of conceptual modelling techniques has depended largely on the adequacy of conceptualisation, expressive power, comprehensibility and formal foundation. An equally important requirement, particularly with the increased conceptualisation of business aspects, is business suitability. In this paper, the focus is on the business suitability of workflow modelling for a commonly encountered class of (operational) business processing, e.g. those of insurance claims, bank loans and land conveyancing. A general assessment is first conducted on some integrated techniques characterising well-known paradigms - structured process modelling, object-oriented modelling, behavioural process modelling and business-oriented modelling. Through this, an insight into business suitability within the broader perspective of technique adequacy, is gained. A specific business suitability diagnosis then follows using a particular characterisation of business processing, i.e. one where the intuitive semantics and inter-relationship of business services and business processes are nuanced. As a result, five business suitability principles are elicited. These are proposed for a more detailed understanding and (synthetic) development of workflow modelling techniques. Accordingly, further insight into workflow specification languages and workflow globalisation in open distributed architectures may also be gained.
Business Process Management Systems ( BPMSs ) provide automated support for the execution of business processes in modern organisations. With the emergence of cloud computing, BPMS deployment considerations are shifting from traditional on-premise models to the Software-as-a-Service ( SaaS ) paradigm, aiming at delivering Business Process Automation as a Service. However, scaling up a traditional BPMS to cope with simultaneous demand from multiple organisations in the cloud is challenging, since its underlying system architecture has been designed to serve a single organisation with a single process engine. Moreover, the complexity in addressing both the dynamic execution environment and the elasticity requirements of users impose further challenges to deploying a traditional BPMS in the cloud. A typical SaaS often deploys multiple instances of its core applications and distributes workload to these application instances via load balancing. But, for stateful and often long-running process instances, standard stateless load balancing strategies are inadequate. In this article, we propose a conceptual design of BPMS capable of addressing dynamically varying demands of end users in the cloud, and present a prototypical implementation using an open source traditional BPMS platform. Both the design and system realisation offer focused strategies on achieving scalability and demonstrates the system capabilities for supporting both upscaling, to address large volumes of user demand or workload, and downscaling, to release underutilised computing resources, in a cloud environment.
, Abstract. The socioeconomic consequences of not successfully completing PhD studies have motivated universities to expend dedicated efforts on improving student journeys. These journeys leave traces in a variety of university IT systems and this trace data can be exploited to derive insights through the application of process mining. Process mining is a form of data-driven process analytics, where process data, collated from different IT systems, is analysed to uncover the real behaviour and performance of processes. Despite its potential application, process mining hitherto has not been applied to visualise, analyse, and improve PhD student journeys, to the best of our knowledge. This paper reports on the findings of a process mining case study conducted at an Australian University that had es-poused a digital transformation initiative to improve PhD student journeys. The case study utilised interactive and comparative process mining techniques and focused on clarifying the way a PhD student journey eventuates, visualising the differences between the real (actual) and prescribed (recommended) processes, comparing the performance of different cohorts, identifying root causes for ad-verse outcomes, and providing evidence-based recommendations for the digital transformation initiative. The findings from this study resulted in restructuring of HDR services and the introduction of a new research management system.
A business process model may be used as both a communication artefact for gathering and sharing knowledge of a business practice among stakeholders and as a specification for the automation of the process by a Business Process Management System (BPMS). For each of these uses, it is desirable to have an ability to visualise the process model from a range of different perspectives and at various levels of granularity. Such views are a common feature of enterprise architecture frameworks, but process modelling and management systems and tools generally have a limited number of available views to offer. This paper presents a taxonomy of process views that are presented in the literature and then proposes the definition and use of a common process model ontology, from which an extensible range of process views may be derived. The approach is illustrated through the realisation of a plug-in component for the YAWL BPMS, although it is by no means limited to that environment. The component illustrates that the process views frequently mentioned in the literature as desirable can be effectively implemented and extended using an ontology-based approach. It is envisaged that the accessibility of a repertoire of views that support business process development will lead to greater efficiencies through more accurate process definitions and improved change management support.
Problem Definition: Queensland’s Compulsory Third-Party (CTP) Insurance Scheme provides a mechanism for persons injured as a result of a motor vehicle accident to receive compensation. Managing CTP claims involves multiple stakeholders with potentially conflicting interests. It is therefore pertinent to investigate whether ‘best practice’ for claims processing can be identified and measured so all claimants receive fair and equitable treatment. The project set out to test the applicability of a mixed-method approach to identify ‘best-practice’ using qualitative, process mining, and data mining techniques in an insurance claims processing domain. Relevance: Existing approaches typically identify ‘best practice’ from literature or surveys of practitioners. The study provides insights into an alternative, mixed-method approach to deriving best practice from historical data and domain knowledge. Methodology: The study is a reflective analysis of insights gained from a practical application of a mixed-method approach to determine ‘best practice’. Results: The mixed-method approach has a number of benefits over traditional approaches in uncovering best practice process behavior from historical data in the real-world context (i.e., can identify process behavior differences between high and low performing cases). The study also highlights a number of challenges with regards to the quality and detail of data that needs to be available to perform the analysis. Managerial Implications: The ‘lessons learned’ from this study will directly benefit others seeking to implement a data-driven approach to understand a ‘best-practice’ process in their own organization.
NoSQL databases disrupted the database market when first introduced. Their contemporary relevance has increased further in the era of big data due to the demands placed on (real-time) analytics. NoSQL databases are well placed to meet these demands due to their performance, availability, scalability, and storage solutions. Unfortunately, to achieve these features, compromises have been made with respect to security and privacy. Growing community awareness and unease combined with increased legislative requirements around data privacy have made such compromises less palatable, risky, or downright unacceptable. And though there is a growing body of knowledge related to data privacy in NoSQL databases, it is diverse and fragmented, and does not adequately address the challenges arising from the current environment. This paper aims to systematically examine various privacy weaknesses of NoSQL databases in the form of patterns. The patterns are shown to manifest themselves in well-known NoSQL databases and this evaluation can be used for benchmarking purposes. Through a survey it is demonstrated that the patterns have been observed in practice and are perceived as relevant. The pattern collection forms a repository of knowledge that can serve as a starting point for future privacy-related research for NoSQL databases through its identification of key problems, trade-offs, existing solution mechanisms, and its provision of terminology.
Workforce analytics brings data-driven methods to organizations for deriving insights from employee-related data and supports decision making. However, it faces an open challenge of lacking the capability to analyze the behavior of employee groups in order to understand organizational performance. This paper proposes a novel notion of work profiles of resource groups, informed by the management literature, for characterizing resource group behavior from multiple aspects relevant to workforce performance. This notion is central to the design of a new, systematic approach that supports resource group analysis by exploiting business process execution data. The approach also provides managers and business analysts with an intuitive means of group-oriented resource analysis by applying visual analytics. We demonstrate the applicability of the approach and usefulness of the proposed notion of resource group work profiles using real datasets from five Dutch municipalities.
The socioeconomic consequences of not successfully completing PhD studies have motivated universities to expend dedicated efforts on improving student journeys. These journeys leave traces in a variety of university IT systems and this trace data can be exploited to derive insights through the application of process mining. Process mining is a form of data-driven process analytics, where process data, collated from different IT systems, is analysed to uncover the real behaviour and performance of processes. Despite its potential application, process mining hitherto has not been applied to visualise, analyse, and improve PhD student journeys, to the best of our knowledge. This paper reports on the findings of a process mining case study conducted at an Australian University that had espoused a digital transformation initiative to improve PhD student journeys. The case study utilised interactive and comparative process mining techniques and focused on clarifying the way a PhD student journey eventuates, visualising the differences between the real (actual) and prescribed (recommended) processes, comparing the performance of different cohorts, identifying root causes for adverse outcomes, and providing evidence-based recommendations for the digital transformation initiative. The findings from this study resulted in restructuring of HDR services and the introduction of a new research management system.
The workflow concept, proliferated through the recently emergent computer supported cooperative work (CSCW) systems and workflow systems, advances information systems (IS) implementation models by incorporating aspects of collaboration and coordination in business processes. Under traditional implementation models, applications are partitioned into discrete units of functionality, with (typically) operational procedures used to describe how human and computerised actions of business processes combine to deliver business services. In this paper, a number of essential modelling concepts and features for business transaction workflows are developed.
Robotic process automation is evolving from robots mimicking human workers in automating information acquisition tasks, to robots performing human decision tasks using machine learning algorithms. In either of these situations, robots or automation agents can have distinct characteristics in their performance, much like human agents. Hence, the execution of an automated task may require adaptations with human participants executing the task when robots fail, to taking a supervisory role or having no involvement. In this paper, we consider different levels of automation, and the corresponding coordination required by resources that include human participants and robots. We capture resource characteristics and define business process constraints that support process adaptations with human-automation coordination. We then use a real-world business process and incorporate automation agents, compute resource characteristics, and use resource-aware constraints to illustrate resource-based process adaptations for its automation.
Robotic Process Automation (RPA) is an emerging technology for automating tasks using bots that can mimic human actions on computer systems. Most existing research focuses on the earlier phases of RPA implementations, e.g. the discovery of tasks that are suitable for automation. To detect exceptions and explore opportunities for bot and process redesign, historical data from RPA-enabled processes in the form of bot logs or process logs can be utilized. However, the isolated use of bot logs or process logs provides only limited insights and not a good understanding of an overall process. Therefore, we develop an approach that merges bot logs with process logs for process mining. A merged log enables an integrated view on the role and effects of bots in an RPA-enabled process. We first develop an integrated data model describing the structure and relation of bots and business processes. We then specify and instantiate a ‘bot log parser’ translating bot logs of three leading RPA vendors into the XES format. Further, we develop the ‘log merger’ functionality that merges bot logs with logs of the underlying business processes. We further introduce process mining measures allowing the analysis of a merged log.