
We present a novel software tool called CDN (Collaborative Data Network) for large-scale sharing and querying of clinical documents modeled using HL7 v3 standard (e.g., Clinical Document Architecture (CDA), Continuity of Care Document (CCD)). Similar to the caBIG initiative, CDN aims to foster innovations in cancer treatment and diagnosis through large-scale, sharing of clinical data. We focus on cancer because it is the second leading cause of deaths in the US. CDN is based on the synergistic combination of peer-to-peer technology and the extensible markup language XML and XQuery. Using CDN, a user can pose both structured queries and keyword queries on the HL7 v3 documents hosted by data providers. CDN is unique in its design - it supports location oblivious queries in a large-scale, network wherein a user does not explicitly provide the location of the data for a query. A location service in CDN discovers data of interest in the network at query time. CDN uses standard cryptographic techniques to provide security to data providers and protect the privacy of patients. Using CDN, a user can pose clinical queries pertaining to cancer containing aggregations and joins across data hosted by multiple data providers. CDN is implemented with open-source software for web application development and XML query processing. We report the evaluation of CDN in a distributed environment (LAN) using a real dataset of discharge summaries available from the i2b2 project.
The decline of elderly social relationships after retirement has negative impacts on their general health status and longevity. Knowing who elderly people are, what are their behaviors and psychological characteristics, how do they communicate and interact, is the first stage in designing efficient and useful services dedicated to them. This paper reports a methodological plan for conducting user-centered design (UCD) of Interactive TV-based services for elderly within the context of the AAL FoSIBLE project. UCD techniques have been used to identify, classify and model elderly people needs for the purpose of designing innovative social services to increase their well-being and self-esteem.
When automatically mining narrative clinical notes to extract meaningful information, such as medical problems, it is essential to take into account the context of this information. For instance, determining whether medical conditions are negated or not is key information for accurately processing medical reports. This article presents an experiment in adapting the state-of-art NegEx algorithm (Chapman et al.) to the French language and evaluating both algorithms (the original English algorithm and the derived French version) on two clinical corpora (English and French, respectively) annotated for medical problems and their negation status. NegEx is a rule-based algorithm which detects negations of medical problems in English-language medical texts, by looking for specific negation trigger phrases in the context of the medical concepts. Our approach has consisted in designing a new list of trigger phrases in French, by studying examples extracted from French clinical notes and relying on the original English list. We performed an evaluation of the negation detection in both corpora. This study show that the two systems achieve comparable results and good performance (respectively 0.839 and 0.867 F-measure for NegEx and its French adaptation).
The FDA and other national regulatory agencies have expressed their intentions to begin enforcement of medical device regulations on Health Informatics Technology (HIT) vendors. A mechanism which might be employed to achieve this enforcement in the US is the Quality Systems Regulations (QSR), while similar legislation might be employed elsewhere. In order for vendors to achieve conformance with QSR regulations, they must first identify hazards which their products may pose. In order to identify these hazards we have undertaken a literature review from which we have extracted taxonomies of Clinical Informatics Systems (CIS) devices, systems and hazards. We present these three taxonomies, and a discussion of contemporary risk classfication strategies which are being applied to HIT. The taxonomies which have been developed provide actionable hazards for HIT vendors, and provide a potential basis for best practices in the engineering of HIT systems.
Contagions - either pathogens spread through contact networks or societal memes spread through social networks - impact the occurrence and character of both epidemic and endemic diseases. While computational models explore disease parameters in the context of a given contact network, these models are always subject to the caveat that reality may not be consistent with the simplified assumptions regarding contact, contagion or network structure. More - and more accurate - data on the contact dynamics between people and places could alleviate some uncertainties, and make models more robust tools for policy-makers and researchers. Properly applied, consumer electronics can serve as a valuable source of this data. Using smartphones as sensor platforms rather than personal communications devices, it is possible to record high fidelity information on a participant's location, activity level, and contacts between both people and places. This paper describes the design, architecture and a preliminary deployment of a general smartphone-based epidemiological data collection system. The dataset, gathered over one month, contains over 45 million records related to the behavioral patterns of 39 participants. We provide an initial analysis of aggregate level statistics to demonstrate the power and scope of the technique for capturing relevant data. Demonstrating the potential for such data to inform decision-making, we further perform an agent-based simulation of a flu-like illness that uses the dataset to capture aspects of both person-person and environmental (place-person) transmission. We demonstrate that the data collection is possible, valuable, and scalable and that the data can be leveraged to inform detailed models capturing more complex physical interactions than were previously feasible.
In this paper, we present a novel method for studying the deterioration of renal functions after kidney transplant. We track the kidney functions of 111 patients for 24 months after the kidney transplant and use the time series data to group the patients into four clusters. We have developed two graph-based algorithms for analyzing the data as a pre-processing step prior to the formation of the clusters. The resultant clusters thus formed are statistically analyzed to determine the socio-demographic and clinical factors that may provide insights into the renal functions after the transplants. We also compare the cluster formation against other manifold learning techniques. The quality of the clusters was assessed using the silhouette function. We discuss how our findings can be used for effective intervention strategies.
Despite recent progress in high-throughput experimental studies, systems level visualization and analysis of large protein interaction networks (PPI) remains a challenging task, given its scale and high-dimensionality. Specifically, techniques that automatically abstract and summarize PPIs at multiple resolutions to provide high level views of its functional landscape are still lacking. In this demonstration, we present a novel data-driven and generic system called FUSE (Functional Summary Generator) that generates functional maps of a PPI at different levels of organization, from broad process-process level interactions to in-depth complex-complex level interactions. By simultaneously evaluating interaction and annotation data, FUSE abstracts higher-order interaction maps by reducing the details of the underlying PPI to form a functional summary graph of interconnected functional clusters. We demonstrate various innovative features of FUSE which aid users to visualize these summaries in a user-friendly manner and navigate through complex PPIs.
Effective communication between health professionals and patients positively influences chronic health management, as does increased patient awareness of their symptoms and general knowledge of the condition. In this study, we leverage the use of mobile phones by pediatric patients and report results from a four-month randomized controlled trial (RCT). We examined: 1) how a SMS system impacted the health outcomes of asthmatic children; and 2) how physicians used a Web service showing the data gathered from the SMS system. Our results show that 1) the simple act of communicating knowledge and symptom awareness information via SMS leads to improved pulmonary function for pediatric patients; and 2) physicians would use the data sent from the SMS system to monitor their patient's asthma management status.
Social scientists, like those performing research at the Kinsey Institute for Research in Sex, Gender and Reproduction, may use surveys to gather large amounts of sensitive data. Unlike purely medical-related datasets, these social science datasets tend to be sparse and high-dimensional, which presents opportunities to characterize participants in the dataset in unique ways. These unique characterizations may enable individuals to be linked to external data in ways that have not been previously considered. Therefore, traditional approaches to de-identifying data, such as fulfilling HIPAA requirements, may not be sufficient for preventing the re-identification of participants in large social science datasets. In this paper, we evaluate the statistical characteristics of two high-dimensional social science datasets to better understand how unique features impact privacy. We apply a class of statistical de-anonymization attacks in an attempt to achieve theoretical re-identification of participants. We assume that an attacker has exact knowledge of a subset of attribute values for a particular record, and wants to link this subset of data to the actual record to discover the remaining content. We show that although 98% of the records within the dataset are unique given any three attributes, re-identification of the records may not be easily achieved. We attribute limited re-identification to the inherent similarity in the human behavior that the scientists measure. This work is the first to characterize re-identification risks in high-dimensional data that is collected in surveys designed to capture the various behaviors and experiences of groups of individuals.
As identified by recent US Dept. of Health and Human Services rulings [2], cohesive standards are essential to increasing accessibility and consumablity of data from heterogeneous sources. The ability to exchange standardized healthcare information is a critical aspect of this problem and HL7s XML-based Clinical Document Architecture (CDA) is quickly becoming the prevailing format. CDA has clear benefits for healthcare information exchange, but also introduces new technical challenges when one desires to perform in-depth analysis on the information captured by a set of these documents. Healthcare Information Warehouse for Analytics and Sharing (HIWAS), a tooling technology, addresses this challenge in two key ways: one, summarizing data for better understanding by a human user and, two, facilitating conventional analysis by aiding in the creation of and transformation to a simplified model. By making these complex XML documents digestible with conventional business intelligence and analysis technology, HIWAS lowers a key barrier to meaningful use of aggregated clinical data. In this demonstration, we first describe challenges of integrating and analyzing CDA documents, how they are addressed by the HIWAS technology, and finally present an application scenario in which we prepare laboratory data captured in this standard format for analysis by public health officials.
With the growing awareness and enforcement of patient rights, patients are empowered with increasing control on their medical information. In many situations, laws and regulation rules require the acquisition of patients' consent before one can access the patients' health data. However, in practice, patients oftentimes have difficulties determining whether they should permit or deny a certain access request. In this article, we propose an analytical approach to assist patients in the consent management of their medical information. Our consent management system employs a statistical learning method that evaluates the benefits and risks associated with access requests, so as to make personalized recommendation on consent decisions. Multiple factors are considered in the assessment process, including the importance of the request, the sensitivity of the requested information, and correlation information. We have implemented a prototype of our solution and performed evaluation with large-scale medical records.
As asserted by the Institute of Medicine, sound health policy and investment decisions require use of "what if" simulation models to analyze the potential impacts of alternative decisions on health outcomes. The challenge is that high-level health decisions require understanding complex interactions of diverse systems across many disciplines both inside and outside of healthcare, creating a need for experts across widely different domains to combine their data and models. Splash - the Smarter Planet Platform for Analysis and Simulation of Health - is a novel decision support framework that facilitates combining heterogeneous, pre-existing simulation models and data from different domains and disciplines. Splash leverages and extends data integration, search, and scientific-workflow technologies to permit loose coupling of models via data exchange. This approach avoids the need to enforce universal standards for data and models, thereby facilitating both model interoperability and reuse of models and data that were independently created or curated by different individuals or organizations. In this way Splash can help domain experts from different areas collaborate effectively and efficiently to attack complex health problems. We illustrate Splash's architecture and capabilities using a simple, proof-of-concept model of community obesity. We show how models of transportation, eating habits, food-shopping choices, exercise, and human metabolism can be combined with geographic, store location, and population data to play "what if," asking, for instance, how community obesity measures would change if tax incentives are used to encourage grocery chains selling healthy and inexpensive food to open stores near obesity "hot spots."
Bedside clinicians routinely identify temporal patterns in physiologic data in the process of choosing and administering treatments intended to alter the course of critical illness for individual patients. Our primary interest is the study of unsupervised learning techniques for automatically uncovering such patterns from the physiologic time series data contained in electronic health care records. This data is sparse, high-dimensional and often both uncertain and incomplete. In this paper, we develop and study a probabilistic clustering model designed to mitigate the effects of temporal sparsity inherent in electronic health care records data. We evaluate the model qualitatively by visualizing the learned cluster parameters and quantitatively in terms of its ability to predict mortality outcomes associated with patient episodes. Our results indicate that the model can discover distinct, recognizable physiologic patterns with prognostic significance.
Efficient analysis of event sequences and the ability to answer time-related, clinically important questions can accelerate clinical research in several areas such as causality assessments, decision support systems, and retrospective studies. The Clinical Narrative Temporal Reasoning Ontology (CNTRO)-based system is designed for semantically representing, annotating, and inferring temporal relations and constraints for clinical events in Electronic Health Records (EHR) represented in both structured and unstructured ways. The LifeFlow system is designed to support an interactive exploration of event sequences using visualization techniques. The combination of the two systems will provide a comprehensive environment for users to visualize inferred temporal relationships from EHR data. This paper discusses our preliminary efforts on connecting the two systems and the benefits we envision from such an environment.
The Linked Open Data (LOD) community project at the World Wide Web Consortium (W3C) is publishing various open data sets as Resource Description Framework (RDF) on the Web and extending it by setting RDF links between data items from different data sources containing information about genes, proteins, pathways, diseases, and drugs. While this presents a very powerful platform for federated querying and heterogeneous data integration, its true potential can only be realized when combining such information with "real patient" data from electronic health records. In this paper, we report our early experiences in applying Linked Data principles and technologies for representing patient data from electronic health records (EHRs) at Mayo Clinic in RDF. In particular, we demonstrate a proof-of-concept case study leveraging publicly available data from the Linked Open Drug Data cloud to federated querying for type 2 diabetes patients. Our study highlights several challenges and opportunities in using Semantic Web tools and technologies within a healthcare setting for enabling clinical and translational research.
When healthcare organizations plan the introduction of advanced health information systems, they need to envision future use. In this paper we describe four different ways of modeling the flow of information in a healthcare context: normative, indicating how information should flow; descriptive, indicating how information does flow now; formative, indicating how information could flow; and projective, indicating how information will flow with a specific new health information system. All approaches must work together for analysts to envision future use effectively. We illustrate the above distinctions with a case study based in the Department of Diagnostic Radiology (DDR) in a major tertiary hospital. DDR personnel were considering the introduction of software to help them schedule patient porterage (transport) services to, from, and within the department. Our prospective evaluation method let personnel see advantages and disadvantages of different ways of deploying the porterage software and led to the specification and design of the ValuesViewer™ application.
In this paper, we study the use of microblogs as source of information for medical intelligence gathering. The huge amount of irrelevant data available in microblogs requires sophisticated filtering methods in order to identify only relevant postings. Microblogs are characteristically sparse and noisy. This requires additional considerations for selection of features for automatic classification for relevance with respect to medical intelligence gathering. In this paper, we will analyze which features are well suited. The objective of this work is three-fold: 1) Specifying annotation guidelines for creating a dataset for microblog classification, 2) Studying the characteristics of tweets for deciding on a well suited feature set, and 3) making use of that feature set in an automatic classification system for relevance filtering of microblogs. The quality of the classifier is assessed in experiments with various feature sets. The evaluation shows that despite the challenging characteristics of mircoblogs, good accuracy values of up to 89% can be achieved by the classifier. One main outcome of this work is a data set of annotated twitter data which can be used as a "gold standard" benchmark for further research in this domain.
As medical data continues to transition to an electronic format, opportunities arise for researchers to use this microdata to discover patterns and increase knowledge in order to improve patient care. Now more than ever, it is critical to protect the identities of the patients contained in these databases. Even after removing obvious "identifier" attributes, such as social security numbers or first and last names, that clearly identify a specific person, it is possible to join "quasi-identifier" attributes from two or more publicly available databases to identify individuals. K-anonymity is an established approach that has been used to ensure that no one individual can be distinguished within a group of at least k individuals. The majority of the proposed approaches implementing k-anonymity have focused on improving the efficiency of algorithms implementing k-anonymity; less emphasis has been put towards ensuring the "utility" of anonymized data from a researchers' perspective. We propose a data utility measurement, called the research value (RV), which evaluates how well common cutoffs for numerical data or groupings in categorical data are preserved during the anonymization process. The proposed algorithm utilizing the new utility function scales efficiently when the number of attributes is large, while still ensuring that the generalization process is dictated by the data content expert's assessment of the utility of the generalized data.
We study the problem of medical event coreference resolution in clinical text. Clinical text found in clinical narratives and patient case reports usually reflects a sublanguage with medicine specific terminology. It is also frequently characterized by temporal expressions co-occurring with medical events. In this paper, we outline a method for quantifying the similarity between medical events found in the New England Journal of Medicine patient case reports. We believe this method will be valuable in classifying medical events as coreferential. We approach this problem by determining the overlap between pairs of medical events in terms of 1) the relation between medical events in the UMLS graph structure and 2) the temporal relation between the medical events. We demonstrate our ideas on a corpus of New England Journal of Medicine case reports annotated with coreference information. Preliminary results indicate a precision of 78.5% and recall of 95.5% in identifying pairs of coreferential medical events.
In a Mass Casualty Incident (MCI) time and good management are critical. Currently, the first arriving rescue units perform the triage algorithm on paper instead of a mobile device. By using mobile devices, the patients' triage state and position can be instantly shared through a network. We implemented a map application to visualize this data on a rugged tablet PC that is intended to be used by the Ambulant Incident Officer (AIO). Even though, using a mobile device offers more benefits, it also requires some mental efforts from the user. The goal of the SpeedUp project1 is to ensure the speed-up of the rescue process. It is crucial to carefully develop, introduce and evaluate the User Interface (UI) iteratively close to the target group and adapt it to an MCI situation. Thus, multiple UI concepts have been developed to compare, rate and optimize them. This paper represents a follow up study and focuses on approaches to select patients on a digital map displayed on a heavy rugged tablet PC. An evaluation is performed to estimate how intuitive, efficient, and ergonomic the UI is without the need of special training for the target group and to increase the acceptance of new devices.