
This study evaluated an automated approach for appendicitis risk stratification of pediatric Emergency Department patients using Conditional Random Fields, rules and Support Vector Machines. The results show that the approach is very promising for appendicitis risk stratification.
This paper presents an interactive biomedical image retrieval system based on automatic visual region-of-interest (ROI) extraction and classification into visual concepts. In biomedical articles, authors often use annotation markers such as arrows, letters or symbols overlaid on figures and illustrations in the articles to highlight ROIs. These annotations are then referenced and correlated with concepts in the caption text or figure citations in the article text. This association creates a bridge between the visual characteristics of important regions within an image and their semantic interpretation. Our proposed method at first localizes and recognizes the annotations by utilizing a combination of rule-based and statistical image processing techniques. Identifying these assists in extracting ROIs that are likely to be highly relevant to the discussion in the article text. The image regions are then annotated for classification using biomedical concepts obtained from a glossary of imaging terms. Similar automatic ROI extraction can be applied to query images, or user may interactively mark an ROI. As a result of our method, visual characteristics of the ROIs can be mapped to text concepts and then used to search image captions. In addition, the system can toggle the search process from purely visual to a textual one (cross-modal) or integrate both visual and textual search in a single process (multi-modal) based on utilizing user feedback. The hypothesis, that such approaches would improve biomedical image retrieval, is validated through experiments on a biomedical article dataset of thoracic CT scans from the collection of ImageCLEF'2010 medical retrieval track.
The growing availability of electronic clinical data is enabling new opportunities for large-scale distributed data-sharing networks that support comparative effectiveness research (CER). Data stored in electronic health records (EHRs) require substantial processing to be usable in distributed research networks (DRNs). We describe the functional features of ROSITA (Reusable OMOP and SAFTINet Interface Adaptor), a virtual machine package that performs many required functions to transform EHR data for use in distributed CER networks. ROSITA is a """"middleware"""" component of SAFTINet, a multi-institutional DRN focused on CER studies to inform the care of safety net populations.
This pilot study aims to determine how well subjects annotate assertions about problem mentions in clinical text and determine if a statistical difference exists between subjects with and without clinical domain knowledge.
Alan Calvitti Neal Farber Yunan Chen Danielle Zuest Lin Liu Kristin Bell Barbara Gray Zia Agha We develop temporal data mining and visualization methods to quantitatively profile physician Electronic Health Records (EHR) workflow and compare time-at-task versus click count distributions for top-level EHR functionality. The temporal data is based on time-resolved activity during outpatient visits, captured by usability software and audio-video recording and manual coding to physicians' activities.
In clinical notes, medication information follows certain semantic patterns and some medication descriptions contain additional word(s) between medication attributes. Therefore, it is essential to understand the semantic patterns as well as the patterns of the context interspersed among them for natural language processing tools to effectively extract comprehensive medication information. We examined both semantic and context patterns and compared those found in Mayo Clinic and i2b2 challenge data. We found that some variations exist between the institutions but the dominant patterns are common.
Tumours can be considered a set of cells that accumulate genetic and epigenetic alterations. According to the Multi-stage Hit theory, the transformation of a normal into a tumour cell involves a number of limiting events that occur in a number of discrete stages (driver mutations). However, not all mutations that occur in the cell are directly involved in the development of cancer and some probably do not contribute in any way (passenger mutations). Moreover, the process of tumour evolution is punctuated by selection of advantageous mutations and clonal expansions. Actually, it is not known how many limiting-events, i.e., how many driver mutations are necessary or sufficient to promote a carcinogenic process. This conjecture should be explored and tested - mathematically and statistically, with the availability of genomic data on databanks. In this work, we explore the model proposed by Bozic and collaborators (2010) that describes the evolution of the tumour according to a Galton-Watson process. Besides, the model gives the relation between the numbers of passenger mutations giving a specific number of driver mutations. We intend to explore some of the model parameters and test some premises about the number of drive mutations and selective advantage, comparing the simulation results with genomic data from colorectal cancer patients. The genomic data was obtained from the DBMutation (http://www.bioinformatics-brazil.org/dbmutation/), a comprehensive database for genomic mutations in cancer. We expect that correlations between driver mutations and the time evolution of tumour process will facilitate the interpretation of genomic information, to make them useful and applicable to clinical oncology.
Ambulatory care sensitive conditions (ACSCs) are characterized as health conditions for which good outpatient care can potentially prevent the need for hospitalization, or for which early intervention can prevent complications or more severe disease. Currently, there are 16 identified ACSCs within the US health system: diabetes short-term complication, perforated appendix, diabetes long-term complication, pediatric asthma, chronic obstructive pulmonary disease, pediatric gastroenteritis, hypertension, congestive heart failure, low birth weight rate, dehydration, bacterial pneumonia, urinary tract infection, angina admission without procedure, uncontrolled diabetes, adult asthma, and lower-extremity amputation among patients with diabetes. Potentially preventable acute health events (PPEs) for such diagnosis codes represent a straightforward opportunity for reducing medical costs while concomitantly improving quality of care. While claims data have previously been used to predict future health outcomes of patients, we report here a novel approach, using data mining techniques, towards supplementing such data with patients' electronic health records (EHR) to develop a clinical decision support system that satisfactorily predicts the onset of PPEs in a large population of patients.
We sought to examine the frequencies and patterns of nephrotoxicity and neutrophilia due to azathioprine (AZA), and to develop a prototype method for using large de-identified electronic health record (EHR) data to aid in post-market drug surveillance. We leveraged a de-identified database of over 10 million patient EHRs to construct a network of comorbidities induced by administration of AZA, where comorbidities were defined by baseline-controlled laboratory values. To gauge the significance of the identified disease patterns, we calculated the relative risk of developing a comorbidity pair relative to a control cohort of patients taking one of 12 other anti-rheumatic agents. Nephrotoxicity as gauged by elevations in creatinine was present in 11% of patients taking AZA, and this frequency was significantly higher than in patients taking other anti-rheumatic agents (RR: 1.2, 95% CI: 1.04-1.43). Neutrophilia was highly prevalent (45%) in the population and was also unique to AZA (RR: 1.2, 95% CI: 1.17-1.28). Using a comorbidity network analysis, we hypothesized that the joint consideration of anemia (hemoglobin 190 IU/L) may serve as a predictor of impending renal dysfunction. Indeed, these two laboratory values provide approximately 100% sensitivity in predicting subsequent elevations in creatinine. Furthermore, the predictive power is unique to AZA, for jointly considering anemia and an elevated LDH provides only 50% sensitivity in predicting creatinine elevations with other anti-rheumatic agents. Our work demonstrates that the construction of comorbidity networks from de-identified EHR data sets can provide both sufficient insight and statistical power to uncover novel patterns and predictors of disease.
In this study our aim was to present a series of experiments to evaluate the impact of pre-annotation: (1) on the speed of manual annotation of clinical notes and clinical trial announcements; and (2) test for potential bias if pre-annotation is utilized. The gold standard was 900 clinical trial announcements from clinicaltrials.gov website and 1655 clinical notes annotated for diagnoses, signs, symptoms, UMLS CUI and SNOMED CT codes. Two dictionary-based methods were used to pre-annotate the text. Annotation time savings ranged from 2.89% to 29.1% per entity. The pre-annotation did not reduce the IAA or annotator performance but reduced the time to annotate in every experiment. Dictionary-based pre-annotation is a feasible and practical method to reduce cost of annotation without introducing bias in the process.
The deep sequencing of transcriptomes has revolutionized our ability to detect known and novel RNA variants at a never before observed resolution. To capitalize on these ever improving technologies, we require functionally rich methods of annotation to predict and evaluate the consequences of RNA isoform variation at the level of proteins, domains and microRNA binding sites. We introduce a new version of the popular open-source application AltAnalyze, capable of analyzing RNA-Sequencing (RNA-Seq) datasets as well as splicing-sensitive or conventional arrays. This software can be run through an intuitive graphical user interface or command-line. Over 60 species and data from various RNA-Seq alignment workflows are immediately supported without any specialized configuration. AltAnalyze provides multiple options for gene expression quantification, filtering, quality control and biological interpretation. Hierarchical clustering heatmaps, principal component analysis plots, lineage correlation diagrams and visualization of enriched pathways are automatically produced for differentially expressed genes. For detection of alternative splicing, promoter or polyadenylation events, AltAnalyze combines both reciprocal-junction and alternative-exon expression approaches to identify annotated and novel RNA variation. By connecting these regulated splicing-events with optimal inclusion and exclusion isoforms, AltAnalyze is able to evaluate the impact of alternative RNA expression on protein domains, annotated motifs and binding sites for microRNAs. From a broader perspective, AltAnalyze examines the enrichment of effected domains and microRNA binding sites, to highlight the global impact of alternative splicing. Together, AltAnalyze provides an efficient, streamlined and comprehensive set of analysis results, to determine the biological impact of transcriptome regulation.
National guidelines for a number of health conditions recommend that practitioners assess and reinforce patient's adherence to specific diet and lifestyle modifications. Counseling intervention has shown to have a long-term positive effect on patient adherence but the extent to which physicians comply is unknown. Evidence of counseling provided by practitioner is recorded only as free text in electronic medical records. To identify physicians' counseling practices we developed a natural language processing system to detect text documentation of dietary counseling in gout patients.
Diabetes is the seventh leading cause of death in the United States, but careful symptom monitoring can prevent adverse events. A real-time patient monitoring and feedback system is one of the solutions to help patients with diabetes and their healthcare professionals monitor health-related measurements and provide dynamic feedback. However, data-driven methods to dynamically prioritize and generate tasks are not well investigated in the domain of remote health monitoring. This paper presents a wireless health project (WANDA) that leverages sensor technology and wireless communication to monitor the health status of patients with diabetes. The WANDA dynamic task management function applies data analytics in real-time to discretize continuous features, applying data clustering and association rule mining techniques to manage a sliding window size dynamically and to prioritize required user tasks. The developed algorithm minimizes the number of daily action items required by patients with diabetes using association rules that satisfy a minimum support, confidence and conditional probability thresholds. Each of these tasks maximizes information gain, thereby improving the overall level of patient adherence and satisfaction. Experimental results from applying EM-based clustering and Apriori algorithms show that the developed algorithm can predict further events with higher confidence levels and reduce the number of user tasks by up to 76.19 %.
Remote and wearable medical sensing has the potential to create very large and high dimensional datasets. Medical time series databases must be able to efficiently store, index, and mine these datasets to enable medical professionals to effectively analyze data collected from their patients. Conventional high dimensional indexing methods are a two stage process. First, a superset of the true matches is efficiently extracted from the database. Second, supersets are pruned by comparing each of their objects to the query object and rejecting any objects falling outside a predetermined radius. This pruning stage heavily dominates the computational complexity of most conventional search algorithms. Therefore, indexing algorithms can be significantly improved by reducing the amount of pruning. This paper presents an online algorithm to aggregate biomedical times series data to significantly reduce the search space (index size) without compromising the quality of search results. This algorithm is built on the observation that biomedical time series signals are composed of cyclical and often similar patterns. This algorithm takes in a stream of segments and groups them to highly concentrated collections. Locality Sensitive Hashing (LSH) is used to reduce the overall complexity of the algorithm, allowing it to run online. The output of this aggregation is used to populate an index. The proposed algorithm yields logarithmic growth of the index (with respect to the total number of objects) while keeping sensitivity and specificity simultaneously above 98%. Both memory and runtime complexities of time series search are improved when using aggregated indexes. In addition, data mining tasks, such as clustering, exhibit runtimes that are orders of magnitudes faster when run on aggregated indexes.
This paper describes works carried out in the Virtual Imaging Platform (VIP) project to create a comprehensive conceptualization of object models used in medical image simulation and suitable for the major imaging modalities and simulators. The goal is to create an application ontology that can be used to annotate the models in the VIP platform's model repository, to facilitate their sharing and reuse. Such annotations allow making the anatomical, physiological and pathophysiological content of the object models explicit.
We describe a web-based, volumetric image annotation tool that is based entirely on HTML5/CSS3 web presentation technologies. The annotation tool can be used on a wide variety of volumetric medical image formats. The application interfaces with ontology web services so that the annotations are labeled with well-defined terms.
In this paper, an efficient medical image denoising method based on low-rank matrix completion and block matching filtering is proposed. The effectiveness of the algorithm in removing the mixed noise is demonstrated through the results. The results also proved the effectiveness of this algorithm in removing noise from regular structures. This method results in comparable performance with significantly lower computation complexity.
Introduction: Medication safety requires monitoring throughout a drug's market life. Early detection of adverse drug reactions (ADRs) can lead to alerts that prevent patient harm. Recently, electronic medical records (EMRs) have emerged as a valuable resource for pharmacovigilance. This study examines the use of retrospective medication orders and inpatient laboratory results in the EMR to identify ADRs. Methods: Using 12 years of EMR data, we designed a study to correlate abnormal laboratory results with specific drug orders by comparing outcomes of a drug-exposed group and a matched unexposed group. We assessed the relative merits of six pharmacovigilance methods used in spontaneous reporting systems (SRS), including proportional reporting ratio (PRR), reporting odds ratio (ROR), Yule's Q, the Chi-square test, Bayesian confidence propagation neural networks (BCPNN) and a gamma Poisson shrinker (GPS). The time of admission was set as "day zero" and all drug orders and laboratory results timings were represented as days elapsed since that time until discharge. Each patient in the exposed group was randomly matched to four unexposed patients by age group, gender, race, and major diagnoses based on ICD9 codes.
The database of Genotypes and Phenotypes (dbGaP) was developed by the National Heart Lung, and Blood Institute (NHLBI) to archive genome-wide association studies (GWAS) data. As of July 17th 2012, dbGaP contained 305 top-level studies. The metadata for each study (available from the dbGaP website) are organized into distinct sections, including a study description, inclusion/exclusion criteria, policies for authorized access requests, MeSH terms, PubMed identifiers, study histories, and the names of principal and co-investigators. We here tabulate the salient characteristics of dbGaP metadata as part of the Phenotype Discoverer (PhD) project, a research project at the University of California San Diego Division of Biomedical Informatics which aims to enhance the """"searchability"""" of the current dbGaP website through the alignment of phenotypes to a standard information model. In particular, we are interested in using the extracted metadata PubMed identifiers, principal investigator names, associated journal names, etc. as input to a statistical text
The past few years has witnessed rapid development in human genome research, in particular the genome-wide association studies (GWAS) and personalized medicine, which has been made possible by the advance in the Next Generation Sequencing (NGS) technologies that produces a large amount of sequencing data at an exceedingly low cost. New technologies for large-scale meta-analysis on genomic data continue to be developed, enabling the application of human genome research to clinical diagnosis and therapy, a trend dubbed “base pairs to bedside”. However, further progress in this area has been increasingly impeded by the constraints in accessing sequencing data, due in part to privacy concerns involved in data sharing. The current approach to protecting human genomic data is mainly based upon data-use agreements, which involves a time-consuming application/review/agreement process. To enable more convenient data access, this paper proposes a data analysis model that allows biomedical researchers and healthcare practitioners to use the sensitive genomic data that cannot be directly released in an efficient fashion, through the computing service over the data (instead of direct access to the data) provided by a large data center.