This paper addresses the problem of correcting medical coding errors with respect to some coding recommendations. The problem consists in clustering medical codings and determining for each cluster the set of features to correct in order to maximize the financial benefits subject to coding correction effort constraints. For this purpose, we model the coding recommendation as a disjunction of hypercubes and introduce the concept of correction sets. A mixed integer linear programming model is then proposed to assign medical codes to correction sets in order to maximize the financial benefits. The miscoding is then explained by characterizing optimal clusters with association rules and coding error distribution. A case study on patient stays associated with malnutrition-related ICD codes is presented, and the performance of the proposed methodology is assessed in regard to the current coding staff practice. A significant increase in health services reimbursement is achieved with a limited number of subjects’ features reviewed. Note to Practitioners —Medical miscoding has a significant negative impact on hospitals with a financial loss for under coding and a penalty for over coding. Whether a medical review is necessary for all descriptive features of a miscoded subject? Is it possible to reduce unnecessary medical reviews without compromising the goal of increasing hospital financial benefits? This article attempts to answer these questions with a data-driven optimization approach to determine a limited number of miscoding clusters and the set of features to review for each in order to best balance the financial benefits and the medical review workload. The application to a real-life case study leads to a significant increase in hospital fiscal revenue of nearly 6,992,489.69, while reviewing only a small number of descriptive features (5293 out of 22056 features, or 24% of features). Causes are also provided for each discovered coding error subtype to ameliorate medical coders’ coding practices. Furthermore, the proposed approach allows the decision-maker to balance the cost-benefit and the requirement of public health institutions (i.e., miscoding rate).
BackgroundThe optimization of patient care pathways is crucial for hospital managers in the context of a scarcity of medical resources. Assuming unlimited capacities, the pathway of a patient would only be governed by pure medical logic to meet at best the patient’s needs. However, logistical limitations (eg, resources such as inpatient beds) are often associated with delayed treatments and may ultimately affect patient pathways. This is especially true for unscheduled patients—when a patient in the emergency department needs to be admitted to another medical unit without disturbing the flow of planned hospitalizations.ObjectiveIn this study, we proposed a new framework to automatically detect activities in patient pathways that may be unrelated to patients’ needs but rather induced by logistical limitations.MethodsThe scientific contribution lies in a method that transforms a database of historical pathways with bias into 2 databases: a labeled pathway database where each activity is labeled as relevant (related to a patient’s needs) or irrelevant (induced by logistical limitations) and a corrected pathway database where each activity corresponds to the activity that would occur assuming unlimited resources. The labeling algorithm was assessed through medical expertise. In total, 2 case studies quantified the impact of our method of preprocessing health care data using process mining and discrete event simulation.ResultsFocusing on unscheduled patient pathways, we collected data covering 12 months of activity at the Groupe Hospitalier Bretagne Sud in France. Our algorithm had 87% accuracy and demonstrated its usefulness for preprocessing traces and obtaining a clean database. The 2 case studies showed the importance of our preprocessing step before any analysis. The process graphs of the processed data had, on average, 40% (SD 10%) fewer variants than the raw data. The simulation revealed that 30% of the medical units had >1 bed difference in capacity between the processed and raw data.ConclusionsPatient pathway data reflect the actual activity of hospitals that is governed by medical requirements and logistical limitations. Before using these data, these limitations should be identified and corrected. We anticipate that our approach can be generalized to obtain unbiased analyses of patient pathways for other hospitals.
Process mining techniques can be used to analyse business processes using the data logged during their execution. These techniques are leveraged in a wide range of domains, including healthcare, where it focuses mainly on the analysis of diagnostic, treatment, and organisational processes. Despite the huge amount of data generated in hospitals by staff and machinery involved in healthcare processes, there is no evidence of a systematic uptake of process mining beyond targeted case studies in a research context. When developing and using process mining in healthcare, distinguishing characteristics of healthcare processes such as their variability and patient-centred focus require targeted attention. Against this background, the Process-Oriented Data Science in Healthcare Alliance has been established to propagate the research and application of techniques targeting the data-driven improvement of healthcare processes. This paper, an initiative of the alliance, presents the distinguishing characteristics of the healthcare domain that need to be considered to successfully use process mining, as well as open challenges that need to be addressed by the community in the future.
Quantitative plant biology is a growing field, thanks to the substantial progress of models and artificial intelligence dealing with big data. However, collecting large enough datasets is not always straightforward. The citizen science approach can multiply the workforce, hence helping the researchers with data collection and analysis, while also facilitating the spread of scientific knowledge and methods to volunteers. The reciprocal benefits go far beyond the project community: By empowering volunteers and increasing the robustness of scientific results, the scientific method spreads to the socio-ecological scale. This review aims to demonstrate that citizen science has a huge potential (i) for science with the development of different tools to collect and analyse much larger datasets, (ii) for volunteers by increasing their involvement in the project governance and (iii) for the socio-ecological system by increasing the share of the knowledge, thanks to a cascade effect and the help of ‘facilitators’.
Process mining is increasingly used to discover and analyze health care processes. It is especially powerful in the study and improvement of patient clinical pathways. Combining process mining results and discrete-event simulation is an interesting approach to discover, represent and assess clinical pathways and improve healthcare organizations. The objective of this work is to develop a framework to automate such studies from the data preprocessing stage to use in simulations. We describe the use of Python and the PM4PY package for formatting data and discovering processes. A generic discrete-event simulation model is developed to serve as a base for analyzing and improving the patient flow in a healthcare center. This type of framework enriches the classical simulation model with synthetic pathways based on real patients and should facilitate accessing aggregated patient data and transposing studies on third-party datasets.
Scheduling problems are a subclass of combinatorial problems consisting of a set of tasks/activities/jobs to be processed by a set of resources usually to minimize a time criterion. Some optimization methods used to solve these problems are hybridized with knowledge discovery techniques to extract information during the optimization process and enhance it. However, most of these hybrid techniques are custom-designed and lack generalization. In this paper a module for knowledge extraction in Stochastic Local Searches is designed, aiming to be problem independent and plugged into optimisation methods that relies on multiple Stochastic Local Search replications. The objective is to prune parts of the search space for which the exploration is likely to lead to poor solutions. This is performed through the extraction of high-quality patterns occurring in locally optimal solutions. Benchmarked on two well-known scheduling problems, the Job-shop Problem and the Resource Constrained Project Scheduling Problem, the results show both a speed up in the convergence and the reaching of better local optima solutions.
Predictive process monitoring aims at predicting the evolution of running traces based on models extracted from historical event logs. Standard process prediction techniques are limited to the prediction of the next activity in a running trace. As a consequence, processes with complex topology (i.e. with several events having similar start/end time) are impossible to predict with these classical multinomial classification approaches. In this paper, the goal is to exploit an original features engineering technique which converts the historical event log of a process into different topological and temporal features, capturing the behavior and context of execution of previous events. These features are then used to train an Entity Embeddings Neural Network in order to learn a model able to predict, in a one-shot manner, both the remaining activities until the end in a running trace and the associated timestamp. Experiments show that this approach globally outperforms previous work for both types of predictions.
This paper applies topological data analysis (TDA) techniques to investigate the nature of complex high-dimensional data by extracting global shape information (pat-terns) and gaining novel insights from them. The objective is to characterize miscoding behaviors, identify reasons for miscoding behaviors, and select specific groups of subjects for which the health records are worth giving an additional review. Our method combines a TDA technique and an optimization-based model to provide a geometric representation of inter-related hospital stays while permitting the censoring of miscoded subjects by preferentially selecting subgroups with more coding errors. Through the proposed method, we successfully identified and validated four distinct subtypes of miscoding behaviors that traditional methodologies fail to find. Furthermore, with only 20% of the subjects reviewed, the proposed approach reduces coding errors by 64% of the whole population. Experimental results indicate that the proposed method is promising and can reduce coding errors efficiently, thereby eliminating the negative impacts caused by hospital miscoding.
Medical test selection is a recurring problem in health prevention and consists of proposing a set of tests to each subject for diagnosis and treatment of pathologies. The problem is characterized by the unknown risk probability distribution across the population and two contradictory objectives: minimizing the number of tests and giving the medical test to all at-risk populations. This article sets this problem in a general framework of chance-constrained medical test rationing with unknown subject distribution over an attribute space and unknown risk probability but with a given sample population. A new approach combining decision-tree and Bayesian inference is proposed to allocate relevant medical tests according to the subjects’ profile. Case studies on screening of hypertension and diabetes are conducted, and the performance of the proposed approach is evaluated. Significant savings on unnecessary tests are achieved with limited numbers of subjects needing but not receiving necessary tests. Note to Practitioners—Whether a medical test is needed for all subjects in health prevention? Is it possible to reduce unnecessary tests without jeopardizing the goal of screening at-risk populations? This article attempts to answer these questions by proposing a data-driven approach combining decision trees for subject profiling, Bayesian inference for unknown probability distribution estimation, and combinatorial optimization for test allocation. The application of this approach to a real-case study reduces the number of electrocardiogram (ECG) tests by 90% while keeping the number of hypertensive subjects needing but not receiving ECG tests small (five out of 230). A significant cut of unnecessary tests is also achieved in a second case study of diabetes screening. This approach allows decision-makers to better balance the cost-saving and the level of public health objective. Furthermore, the combination with decision trees makes the practical implementation quite straightforward.
Staying informed is essential for citizens but not everyone is equipped with the right tools to distinguish scientific facts from opinions, well-founded arguments from sensationalist news. Popularization and outreach should be more focused on transmitting methods, demonstrations, and arguments rather than scientific results. It is important to keep concepts' complexity while doing so and to target young audiences. The journal DECODER offers an outreach journal allowing direct exchange between the classroom and researchers to provide critical tools to young minds. It is a duty of researchers to engage into outreach activities, as providers of new knowledge and advocates of the scientific method.
Unplanned readmissions to the emergency department (ED) have been identified as a key factor resulting in a negative effect on subjects’ health and healthcare resource scheduling. Often, the readmission prediction is modeled as a binary classification problem whose objective is to predict if a subject will be readmitted or not. Nevertheless, it ignores the uncertainty nature of readmission and usually results in poor prediction quality. In this paper, the problem is defined as a chance-constrained medical intervention rationing problem: at-risk subjects are targeted and given supplemental medical interventions, while the remaining subjects are treated as outpatients. The objective is to profile subjects, identify at-risk subjects, and select specific groups of subjects to which additional medical interventions are recommended, while addressing the unknown number of at-risk subjects and the unknown subjects’ readmission risks. We propose a white-box approach named Alternating Clustering and Bayesian Inference (ACBI) and investigate its efficiency on a real-life readmission data set. Results are promising and show the method could lead up to a 34.42% reduction in readmission rate.
One mission of a researcher is to share their work and results with the general public but there is a real challenge in accurately and effectively sharing scientific results with a broad audience. Indeed, they are published in scientific journals that are mostly available at high costs; the vocabulary used makes it hard for people outside of the field to understand the concepts; and sometimes there is a language barrier for non-English speakers. However, to make informed decisions on a variety of scientific and societal topics, citizens need to have access to and keep up with these research results. To build critical thinking, this good practise should be developed from an early age. We created the journal DECODER (French for “to decode”, journal-decoder.fr), which enables a researcher and a class to work together on their own simplified research article. The middle and high school students can have the role of active reviewers on the researcher’s shorten article or they can write an outreach article on a given topic in which the researcher is a specialist. Articles are then published under a creative commons license and are freely available on the journal website to benefit a majority. Our partner researchers work in space agencies, in academia, or in industry, in a variety of disciplines from STEM to social sciences. The emphasis is set on multidisciplinarity to raise students’ awareness about research wideness and show them that research is not limited to STEM fields but also exists in economics and humanities. This points out the significance and ubiquity of transdisciplinarity in solving real world’s problems, such as global change issues, biological and physical questions or space exploration from different perspectives. In its first year and a half, the journal has already involved more than ten classes in five different schools and 18 articles have been submitted by ten researchers. The project allows a tight and direct interaction between students and researchers and it makes students responsible for the publication content over a large audience. Thanks to an easy procedure for classes and researchers and small-time requirement, our hope is to mobilize the largest scientific community to help people being more critics and having access to scientific results.
Background: Infections by multidrug-resistant Gram-negative (MDRGN) bacteria are among the greatest contemporary health concerns, especially in intensive care units (ICUs), and may be associated with increased hospitalization time, morbidity, costs, and mortality. Aim: The study aimed to predict carbapenem-resistant MDRGN acquisition in ICUs, to determine its risk factors, and to assess the impact of this acquisition on mortality rate. Methods: A matched case-control study was performed in patients admitted to the ICU at a large Brazilian hospital over a five-year period. Cases were defined as patients who acquired carbapenem-resistant MDRGN bacteria during hospitalization. Controls were defined as patients who had no detection of carbapenem-resistant MDRGN bacteria. Cases were matched to controls according to the admission period. Risk factors were identified by multiple logistic regression using a stepwise selection method. Findings: In total, 343 cases and 1029 controls were analysed. The 30-day mortality rate for subjects with ICU-associated carbapenem-resistant MDRGN was 37.6%. Five variables were identified as statistically significant and more relevant for the acquisition of multidrug-resistant strains: increased Simplified Acute Physiology Score 3, patients with severe chronic obstructive pulmonary disease and exposure to haemodialysis catheter, central venous catheter, or mechanical ventilation. Models developed displayed good results with an accuracy of similar to 90%. Patients who acquired MDRGN were 2.72 times more likely to die than non-MDRGN acquisition patients. Conclusion: Finding risk factors and developing predictive models may benefit patients through early detection and by controlling the spread of MDR. The presence of mechanical ventilation and central venous catheter were the main risk factors demonstrated, and their use requires special attention. (C) 2019 The Healthcare Infection Society. Published by Elsevier Ltd. All rights reserved.
Health preventive medical evaluation programs are strategies commonly implemented as part of national health prevention efforts. Two problems related with the implementation of such strategies are subject profiling and medical test selection. The amount of different types of information that have to be analyzed by physicians in the screening for pathologies and medical test prescription can be overwhelming and significantly increases the complexity of this problem. Two decision-tree-based approaches are proposed in this study for subject profiling and medical test rationing. The proposed models perform well in terms of prediction. Results show that the implementation of these approaches helps to profile consultants into healthy and unhealthy subjects which can be used to ration medical test and design policies for preventive health evaluation programs.
Local Process Models (LPM) describe structured fragments of process behavior occurring in the context of less structured business processes. Traditional LPM discovery aims to generate a collection of process models that describe highly frequent behavior, but these models do not always provide useful answers for questions posed by process analysts aiming at business process improvement. We propose a framework for goal-driven LPM discovery, based on utility functions and constraints. We describe four scopes on which these utility functions and constrains can be defined, and show that utility functions and constraints on different scopes can be combined to form composite utility functions/constraints. Finally, we demonstrate the applicability of our approach by presenting several actionable business insights discovered with LPM discovery on two real life data sets.
Local Process Models (LPMs) describe structured fragments of process behavior occurring in the context of less structured business processes. In contrast to traditional support-based LPM discovery, which aims to generate a collection of process models that describe highly frequent behavior, High-Utility Local Process Model (HU-LPM) discovery aims to generate a collection of process models that provide useful business insights by specifying a utility function. Mining LPMs is a computationally expensive task, because of the large search space of LPMs. In supportbased LPM mining, the search space is constrained by making use of the property that support is anti-monotonic. We show that in general, we cannot assume a provided utility function to be anti-monotonic, therefore, the search space of HU-LPMs cannot be reduced without loss. We propose four heuristic methods to speed up the mining of HU-LPMs while still being able to discover useful HU-LPMs. We demonstrate their applicability on three real-life data sets.
Discovering workflow patterns in event-logs is important for many organizations to understand and optimize organizational processes. Although numerous algorithms have been proposed in the literature to discover patterns in sequences of symbols, most of them are inadequate to discover patterns in rich event-log data. In this paper, motivated by the analysis of patient pathways in the health domain, a rich type of event logs, called activity-cost event logs, is considered where each event is associated with a cost. The paper formalizes the problem of mining interesting low-cost patterns in these logs by combining novel concepts of penalties (activity costs) and consistency of patterns, with traditional measures of confidence, length, and time. Furthermore, to extract these patterns efficiently from event logs, an algorithm named TWINCLE (Time-WINdow, Cost and LEngth constrained sequential rule mining) is proposed. Experiments carried out on benchmark datasets and real-life healthcare event logs show that proposed algorithm is efficient and can discover interesting patterns.
Process Mining aims to extract information from event logs to highlight the underlying business processes. It is useful in situations where there is no detailed and complete knowledge of how an overall system works, such as in a hospital where most processes are complex and ad-hoc. Many Process Mining discovery techniques have been proposed so far, but many challenges are still to be faced. Implicit dependencies are one of them. Choice-related phenomenon, implicit dependencies are not taken into account in most algorithms and graphical representations. In this paper, we propose the Implicit Dependencies Miner, a Process Tree based algorithm able to detect relevant dependencies.