In a large hospital system, a network of hospitals relies on electronic health records (EHRs) to make informed decisions regarding their patients in various clinical domains. Consequently, the dependability of the health information technology (HIT) systems responsible for collecting EHR data is of utmost importance for patient safety. Recently novel methods and tools aimed at identifying anomalies in EHR data to bolster the reliability of HIT systems have been introduced. However, these existing methods and tools primarily concentrate on individual hospitals, which limits our understanding of system-wide anomalous events and their potential impact on patient safety across multiple hospitals. In this article, we introduce a new approach to detecting anomalies in EHR data within a network of hospitals. This is achieved by combining advanced machine learning techniques with graph algorithms to create a tool capable of swiftly identifying and responding to deviations. Our proposed approach employs a combination of five machine learning models, harnessing the unique strengths of each model to provide a more robust detection system. The detected anomalies are then represented as graphs, allowing us to recognize patterns across the hospital network. This aids in identifying anomalies that span multiple medical facilities, potentially indicating broader system-level risks. Extensive real-world testing of our approach demonstrated its ability to offer actionable insights compared to existing methods. Additionally, its scalable design ensures seamless integration into existing HIT infrastructures.
OBJECTIVES:To demonstrate an innovative method combining machine learning with comparative effectiveness research techniques and to investigate a hitherto unstudied question about the effectiveness of common prescribing patterns. DATA SOURCES:United States Veterans Health Administration Corporate Data Warehouse. STUDY DESIGN:For Operation Enduring Freedom/Operation Iraqi Freedom veterans with major depressive disorder, we generate pharmacotherapy pathways (of antidepressants) using process mining and machine learning. We select the medication episodes that were started at subtherapeutic doses by the first assigned primary care physician and observe the paths that those medication episodes follow. Using 2-stage least squares, we test the effectiveness of starting at a low dose and staying low for longer versus ramping up fast while balancing observable and unobservable characteristics of patients and providers through instrumental variables. We leverage predetermined provider practice patterns as instruments. DATA COLLECTION:We collected outpatient pharmacy data for selective serotonin reuptake inhibitors and selective norepinephrine reuptake inhibitors, patient and provider characteristics (as control variables), and the instruments for our cohort. All data were extracted for the period between 2006 and 2020. PRINCIPAL FINDINGS:There is a statistically significant positive effect (0.68, 95% CI 0.11-1.25) of "ramping up fast" on engagement in care. When we examine the effect of "ramping up slow", we see an insignificant negative impact on engagement in care (-0.82, 95% CI -1.89 to 0.25). As expected, the probability of drop-out also seems to have a negative effect on engagement in care (-0.39, 95% CI -0.94 to 0.17). We further validate these results by testing with medication possession ratios calculated periodically as an alternative engagement in care metric. CONCLUSIONS:Our findings contradict the "Start low, go slow" adage, indicating that ramping up the dose of an antidepressant faster has a significantly positive effect on engagement in care for our population.
Objective: Physicians and clinicians rely on data contained in electronic health records (EHRs), as recorded by health information technology (HIT), to make informed decisions about their patients. The reliability of HIT systems in this regard is critical to patient safety. Consequently, better tools are needed to monitor the performance of HIT systems for potential hazards that could compromise the collected EHRs, which in turn could affect patient safety. In this paper, we propose a new framework for detecting anomalies in EHRs using sequence of clinical events. This new framework, EHR-Bidirectional Encoder Representations from Transformers (BERT), is motivated by the gaps in the existing deep -learning related methods, including high false negatives, sub -optimal accuracy, higher computational cost, and the risk of information loss. EHR-BERT is an innovative framework rooted in the BERT architecture, meticulously tailored to navigate the hurdles in the contemporary BERT method; thus, enhancing anomaly detection in EHRs for healthcare applications. Methods: The EHR-BERT framework was designed using the Sequential Masked Token Prediction (SMTP) method. This approach treats EHRs as natural language sentences and iteratively masks input tokens during both training and prediction stages. This method facilitates the learning of EHR sequence patterns in both directions for each event and identifies anomalies based on deviations from the normal execution models trained on EHR sequences. Results: Extensive experiments on large EHR datasets across various medical domains demonstrate that EHR-BERT markedly improves upon existing models. It significantly reduces the number of false positives and enhances the detection rate, thus bolstering the reliability of anomaly detection in electronic health records. This improvement is attributed to the model's ability to minimize information loss and maximize data utilization effectively. Conclusion: EHR-BERT showcases immense potential in decreasing medical errors related to anomalous clinical events, positioning itself as an indispensable asset for enhancing patient safety and the overall standard of healthcare services. The framework effectively overcomes the drawbacks of earlier models, making it a promising solution for healthcare professionals to ensure the reliability and quality of health data.
Background To discover pharmacotherapy prescription patterns and their statistical associations with outcomes through a clinical pathway inference framework applied to real-world data. Methods We apply machine learning steps in our framework using a 2006 to 2020 cohort of veterans with major depressive disorder (MDD). Outpatient antidepressant pharmacy fills, dispensed inpatient antidepressant medications, emergency department visits, self-harm, and all-cause mortality data were extracted from the Department of Veterans Affairs Corporate Data Warehouse. Results Our MDD cohort consisted of 252,179 individuals. During the study period there were 98,417 emergency department visits, 1,016 cases of self-harm, and 1,507 deaths from all causes. The top ten prescription patterns accounted for 69.3% of the data for individuals starting antidepressants at the fluoxetine equivalent of 20-39 mg. Additionally, we found associations between outcomes and dosage change. Conclusions For 252,179 Veterans who served in Iraq and Afghanistan with subsequent MDD noted in their electronic medical records, we documented and described the major pharmacotherapy prescription patterns implemented by Veterans Health Administration providers. Ten patterns accounted for almost 70% of the data. Associations between antidepressant usage and outcomes in observational data may be confounded. The low numbers of adverse events, especially those associated with all-cause mortality, make our calculations imprecise. Furthermore, our outcomes are also indications for both disease and treatment. Despite these limitations, we demonstrate the usefulness of our framework in providing operational insight into clinical practice, and our results underscore the need for increased monitoring during critical points of treatment.
Process mining for conformance analysis focuses on comparing a reference process model against a data-driven process model that is generated via log files from information technology systems. While this approach is helpful when there is an existing process model in an organization, it leaves the question of what to do in the absence of a complete reference process model unanswered. In this paper, we present a comparative assessment approach that combines process mining, process mapping for dimensionality reduction, and statistical analysis. Our goal is to find similarities and dissimilarities in data-driven process models among U.S. Veterans Health Administration (VHA) facilities to assess process conformance among different healthcare facilities, which can help assess the standardization of care. We illustrate our approach by applying it to two clinical radiology order process models generated by two similar facilities. Our results demonstrate statistical similarities in the standardization of care among those two facilities.
When parallel algorithms for simulation were introduced in the 1970s, their development and use interested only experts in parallel computation. This circumstance changed as multi-core processors became commonplace, putting a parallel computer into the hands of every modeler. A natural outcome is growing interest in parallel simulation among persons not intimately familiar with parallel computing. At the same time, parallel simulation tools continue to be developed with the implicit assumption that the modeler is knowledgeable about parallel programming. The unintended consequence is a rapidly growing number of users of parallel simulation tools that are unlikely to recognize when the interaction of race conditions, partitioning strategies, and simultaneous action in their simulation models make results non-reproducible, thereby calling into question the validity of conclusions drawn from the simulation data. We illustrate the potential dangers of exposing parallel algorithms to users who are not experts in parallel computation with example models constructed using existing parallel simulation tools. By doing so, we hope to refocus tool developers on usability, even if this new focus incurs loss of some performance.
To improve clinical care practice, it is important to understand the variability of clinical pathways executed in different contexts (e.g., pathways in different geographical locations, demographics, and phenotypic groups). A common way of representing clinical pathways is through network-based representations that capture trajectories of treatment steps. However, first-order networks, which are based on the Markovian property and the de facto standard model to represent transitions between steps, often fail to capture real trajectories. This paper introduces a visual analytic tool to explore and compare pathways represented in higher-order networks. Because each higher node in the network is a subtrajectory (i.e., partial or full history of treatment steps), the tool can display true sequences of treatment steps and compute the similarity of the two networks in a space of higher-order nodes. The tool also highlights areas in which the two networks are similar and dissimilar and how a certain subtrajectory is realized differently in different pathways. The paper demonstrates the tool's usefulness by applying it to multiple antidepressant pharmacotherapy pathways for veterans diagnosed with major depressive disorder and by illustrating heterogeneity in prescription patterns across pathways.
Detecting anomalous sequences is an integral part of building and protecting modern large-scale health information technology (HIT) systems. These HIT systems generate a large volume of records of patients' state and significant events, which provide a valuable resource to help improve clinical decisions, patient care processes, and other issues. However, detecting anomalous sequences in electronic health records (EHR) remains a challenge in healthcare applications for several reasons, including imbalances in the data, complexity of relationships between events in the sequence, and the curse of dimensionality. Conventional anomaly detection methods use the finite sequence of events to discriminate sequences. They fail to incorporate salient event details under variable higher-order dependencies (e.g., duration between events) that can provide better discrimination of sequences in their models. To address this problem, we propose event sequence and subsequence anomaly detection algorithms that (1) use network-based representations of interactions in the data, (2) account for variable higher-order dependencies in the data, and (3) incorporate events duration for adequate discrimination of the data. The proposed approach identifies anomalies by monitoring the change in the graph after the test sequence is removed from the network. The change is quantified using graph distance metrics so that dramatic changes in the network can be attributed to the removed sequence. Furthermore, the proposed subsequence algorithm recommends plausible paths and salient information for the detected anomalous subsequences. Our results show that the proposed event sequence anomaly detection algorithm outperforms the baseline methods for both synthetic data and real-world EHR data.
Health information technology (HIT) was introduced to streamline administrative medical processes and alleviate errors that may negatively affect patient health. It has helped to some extent in this manner; however, studies have shown that the use of HIT has introduced unexpected new errors in the collection, transmission, use, and processing of electronic health records, as well as potential programming or configuration errors in these HIT systems. This study was initiated to identify HIT anomalies and, from this, potential hazards that may threaten patient safety. The results can be used to identify and prioritize hazard classes, and to assist in the design and implementation of HIT hazard controls and detectors.
Streaming analytics is the process of ingesting and digesting live data from multiple data sources. In the healthcare domain, as the importance of extracting immediate insights while data are streaming into the system grows, the focus is shifting from batch processing to streaming analytics. With data increasing dramatically at high speeds, many informatics designs have been proposed to adapt healthcare domain into this new environment. In our previous work, we introduced a prototype of health informatics technology (HIT) framework that aims to address challenges in adopting state-of-the-art technologies to enable advanced healthcare analytic tasks in new streaming environments. We recently made major updates to the framework so that anomaly from multiple streaming data sources at different granularity levels can be detected in near real-time. In this paper, we detail the implementation and deployment of the framework in Kubernetes clusters and report its performances when tested on electronic health record (EHR) data of Veterans Affairs.
Barrett’s esophagus (BE) is a benign condition of the distal esophagus that initiates a multistage pathway to esophageal adenocarcinoma (EAC). Short of frequent intrusive (and costly) surveillance, effective screening for neoplasia in BE populations is yet to be established since progressors are rare and virtually undetectable without routine biopsies, which often sample only a small portion of the BE tissue. As a result, reliable estimation of the true prevalence of dysplasia in a BE population and evidence-based optimization of screening for at-risk individuals is challenging. Data-driven microsimulations, i.e., model-generated instances of disease history in a predefined virtual population, have found utility in the EAC screening literature as low-overhead alternatives to real-world hypothesis testing of optimal interventions for dysplasia. Despite the successes, computational limitations, paucity of knowledge and data on Barrett’s dysplasia, and the complexities of disease progression as a multiscale multiphysics process have hindered the treatment of disease progression in BE as a spatial process. Agent-based modeling of nucleation and proliferation processes in dysplasia warrants exploration in this context as an approximation that operates at a trade-off between computational tractability and precise representation of the composition and physics of the substrate (tissue). In this study, we describe spatially resolved simulations of premalignant progression toward EAC in a coarse-grained model of Barrett’s tissue that resolves the metaplastic tissue at a length scale of 0.42 mm (~3300 crypts/mm 2 ). The model is calibrated to reproduce historical high-grade dysplasia prevalence when model-generated patients are screened using the Seattle protocol.
In highly configurable health information technology (HIT) systems, such as VistA of the Veterans Health Administration, the variations in how the system is used among different healthcare facilities and how the data are recorded can be significant. Despite the successful standardization of care efforts, some of these variations can be indicative of HIT hazards and demand further investigation. In this work, we implemented a recurrent neural network (RNN) architecture to learn clinical provider order sequences and their temporal dynamics while predicting the orders' terminal state. We demonstrate model performance and provide a use case for the model discerning novel event sequences. This model is proposed to find novel event sequences in an operational environment.
The adoption of health information technology (HIT) has facilitated efforts to increase the quality and efficiency of health care services and decrease health care overhead while simultaneously generating massive amounts of digital information stored in electronic health records (EHRs). However, due to patient safety issues resulting from the use of HIT systems, there is an emerging need to develop and implement hazard detection tools to identify and mitigate risks to patients. This paper presents a new methodological framework to develop hazard detection models and to demonstrate its capability by using the US Department of Veterans Affairs' (VA) Corporate Data Warehouse, the data repository for the VA's EHR. The overall purpose of the framework is to provide structure for research and communication about research results. One objective is to decrease the communication barriers between interdisciplinary research stakeholders and to provide structure for detecting hazards and risks to patient safety introduced by HIT systems through errors in the collection, transmission, use, and processing of data in the EHR, as well as potential programming or configuration errors in these HIT systems. A nine-stage framework was created, which comprises programs about feature extraction, detector development, and detector optimization, as well as a support environment for evaluating detector models. The framework forms the foundation for developing hazard detection tools and the foundation for adapting methods to particular HIT systems.
We have documented the precision and accuracy of the analytical method through an expected working calibration range (3.0 to 60 ppb). The analytical method was statistically evaluated over a range of concentrations to establish a detection limit and quantitation limit for the method. Whenever the true concentration is 8.5 ppb or above, the probability is at least 99.9 percent that the measured concentration will be ppb or above. Thus, 6 ppb could be used as a lower reliability limit for detecting concentrations in excess of 8.5 ppb. In summary, the proposed sample extraction and analyses methods are suitable for quantitative analyses to determine the presence of GD, HD, VX, and GA in wastewater samples. Our findings indicate that we can detect any of these chemical surety materiel (CSM) in water at or below the established U.S. Army Surgeon General's safety levels in drinking water.
Background and Purpose Process mining for conformance analysis consists of comparing a reference process model against a data-driven process model generated via log files from information technology systems. However, in the absence of a complete reference process model, we found no suggested approaches in the literature to address the need for evaluating process conformance among different healthcare facilities to assess standardization of care. Our goal is to find similarities and dissimilarities in data-driven process models among US Veterans Health Administration (VHA) facilities that can be indicative of patient safety issues. Our hypothesis was that the analysis would not produce statistically significant differences in outcome.Methods We present a unique implementation of conformance analysis in process mining that consists of combining process mining, process mapping and statistical metrics. We illustrate our approach by applying it to the analysis of two clinical radiology order process models generated from healthcare data provided by two similar facilities in the VHA.Results The comparative assessment showed that about 70% of the orders completed successfully and 30% were not completed due to policy and duplications. Our analysis found a good statistical correlation between both facilities, as the Spearman’s correlation coefficient between facilities for the frequency of cases per total hours was 0.87879, for the frequency of cases by state transition was 0.79702 and for the throughput time per state transition was 0.63582. Additional statistical analyses using the Mann-Whitney U test and the root mean square error both produced values that were not significant.Conclusions The foregoing approach validated our hypothesis by demonstrating a good statistical correlation of data describing the flow of clinical radiology orders absent a credible reference model. Finding good agreement between both facilities was important in confirming that the clinical orders flow in a similar manner, suggesting standardization of care.
Rule based classifiers that use the presence and absence of key sub-strings to make classification decisions have a natural mechanism for quantifying the uncertainty of their precision. For a binary classifier, the key insight is to treat partitions of the sub-string set induced by the documents as Bernoulli random variables. The mean value of each random variable is an estimate of the classifier's precision when presented with a document inducing that partition. These means can be compared, using standard statistical tests, to a desired or expected classifier precision. A set of binary classifiers can be combined into a single, multi-label classifier by an application of the Dempster-Shafer theory of evidence. The utility of this approach is demonstrated with a benchmark problem.
The process of identifying a cohort of interest is a very challenging task. It requires manually inspecting many patient records of complex structure that might include medical coding errors and missing data. This paper presents a computational pipeline for refining the process of cohort selection based on medical concepts recorded in the electronic health records (EHRs). The pipeline extracts EHR data for a given cohort and normalizes this data using standard vocabularies. Then a stacked denoising autoencoder is used to embed the normalized patient vectors in a low dimensional space, where the patients are subsequently clustered into sub-cohorts. The goal is to represent the cohort in a standard format and abstract variants of sub-populations. As a use-case, we applied the pipeline to 1.8 million Veterans diagnosed with major depressive disorder (MDD), and identified four meaningful sub-cohorts using the features learned by the autoencoder. Then, each sub-cohort was explored using a set of keywords for interpretation.
Over the past decade, health information technology (IT) has enabled the amount of digital information stored in electronic health records (EHRs) to expand greatly. However, according to some studies, hazards in health IT can lead to changes in clinical decisions, care processes, and care outcomes, as well as other issues. Thus, the effects of health IT hazards on patient safety have been at the forefront of recent patient safety research. Nonetheless, hazard detection in health IT remains a challenge. In this paper, the authors assume that safety-related issues in health IT would exhibit anomalous characteristics in EHR data. Although all hazards will exhibit some anomalous characteristics, not all anomalies can be regarded as hazards. The authors hypothesize that errors in health IT could lead to interruptions in the sequence of clinical actions. To this end, the problem of detecting anomalous sequences in big EHR data is considered. This paper focuses on dynamic event sequences, which are a series of clinical actions in motion. The authors propose an adaptive anomaly detection approach that uses higher-order network representation to detect anomalous sequences. Furthermore, the authors propose a contiguous subsequence anomaly detection approach that identifies abnormal subsequences in the detected anomalous sequences. The proposed approaches are tested by using synthetic and real-world EHR data. The proposed methods outperform existing state of the art anomaly detection techniques. To reduce the computational complexity associated with the operational implementation of the proposed approaches, the Apache Spark environment was leveraged, and a much shorter run time together with improved performance were achieved, especially for data with more than 60,000 sequences.
In this work, we aim to enhance the reliability of health information technology (HIT) systems by detection of plausible HIT hazards in clinical order transactions. In the absence of well-defined event logs in corporate data warehouses, our proposed approach identifies relevant timestamped data fields that could indicate transactions in the clinical order life cycle generating raw event sequences. Subsequently, we adopt state transitions of the OASIS Human Task standard to map the raw event sequences and simplify the complex process that clinical radiology orders go through. We describe how the current approach provides the potential to investigate areas of improvement and potential hazards in HIT systems using process mining. The discussion concludes with a use case and opportunities for future applications.
To effectively engage demand-side and distributed energy resources (DERs) for dynamically maintaining the electric power balance, the challenges of controlling and coordinating building equipment and DERs on a large scale must be overcome. Although several control techniques have been proposed in the literature, a significant obstacle to applying these techniques in practice is having access to an effective testing platform. Performing tests at scale using real equipment is impractical, so simulation offers the only viable route to developmental testing at scales of practical interest. Existing power-grid testbeds are unable to model individual residential end-use devices for developing detailed control formulations for responsive loads and DERs. Furthermore, they cannot simulate the control and communications at subminute timescales. To address these issues, this paper presents a novel power-grid simulation testbed for transactive energy management systems. Detailed models of primary home appliances (e.g., heating and cooling systems, water heaters, photovoltaic panels, energy storage systems) are provided to simulate realistic load behaviors in response to environmental parameters and control commands. The proposed testbed incorporates software as it will be deployed, and enables deployable software to interact with various building equipment models for end-to-end performance evaluation at scale.