OBJECTIVES/GOALS: There is an imperative need to initiate translational genetic studies of hidradenitis suppurativa (HS). Such work requires large cohorts and no HS registries exist. Precision medicine initiatives provide new resources and methods for efficiently constructing cohorts, but empirically informed best practice guidelines are needed. METHODS/STUDY POPULATION: Traditional methods for building cohorts rely on clinical encounters to identify patients and collect phenotype data. Precision medicine initiatives aim to decrease the time and cost of data collection by using alternative sources, including electronic health records (EHR) and remote collection of patient-reported data. The public’s use of the Internet to obtain and exchange health-related information coupled with the success of direct-to-consumer genetic companies suggests that it is feasible to remotely ascertain research participants for genetic studies. Importantly, Internet cohorts provide an opportunity to include research participants who are disconnected from healthcare, and thus remain hidden from research that relies on EHR or clinical services. RESULTS/ANTICIPATED RESULTS: First, to conduct studies in EHR we are developing an analytic pipeline for the automated extraction of an accurate HS diagnosis using natural language processing of clinical notes. In our preliminary work we are also using ICD codes to build cohorts in two EHR systems with and without linked genetic data. Second, we have developed Internet advertising campaigns for symptom-based recruitment. Informed consent and patient-reported data is collected on-line through a series of short surveys. Patients who complete the surveys and express interest in participating in genetic studies are sent saliva collection kits and return mailing material. Finally, we have established an HS biobank that has DNA from 300 participants identified through clinical services. Enrollment is on going. DISCUSSION/SIGNIFICANCE OF IMPACT: Our goal is to assemble an HS cohort that is large enough to power genetic discoveries. Our work is generating empirical evidence for precision medicine guidelines and will improve our knowledge about HS. The methods we are developing can be applied to efficiently create new cohorts for genetic studies of other diseases across different clinical areas.
BACKGROUND:A shareable repository of clinical notes is critical for advancing natural language processing (NLP) research, and therefore a goal of many NLP researchers is to create a shareable repository of clinical notes, that has breadth (from multiple institutions) as well as depth (as much individual data as possible).METHODS:We aimed to assess the degree to which individuals would be willing to contribute their health data to such a repository. A compact e-survey probed willingness to share demographic and clinical data categories. Participants were faculty, staff, and students in two geographically diverse major medical centers (Utah and New York). Such a sample could be expected to respond like a typical potential participant from the general public who is given complete and fully informed consent about the pros and cons of participating in a research study.RESULTS:Two thousand one hundred forty respondents completed the surveys. 56% of respondents were "somewhat/definitely willing" to share clinical data with identifiers, while 89% of respondents were "somewhat (17%)/definitely willing (72%)" to share without identifiers. Results were consistent across gender, age, and education, but there were some differences by geographical region. Individuals were most reluctant (50-74%) sharing mental health, substance abuse, and domestic violence data.CONCLUSIONS:We conclude that a substantial fraction of potential patient participants, once educated about risks and benefits, would be willing to donate de-identified clinical data to a shared research repository. A slight majority even would be willing to share absent de-identification, suggesting that perceptions about data misuse are not a major concern. Such a repository of clinical notes should be invaluable for clinical NLP research and advancement.
This study used Amazon Mechanical Turk to crowdsource public opinions about sharing medical records for clinical research. The 1,508 valid respondents comprised 58.7% males, 54% without college degrees, 41.5% students or unemployed, and 84.3% under 40 years old. More than 74% were somewhat willing to share de-identified records. Education level, employment status, and gender were identified as significant predictors of willingness to share one's own or one's family's medical records (partially identifiable, completely identifiable, or de-identified). Thematic analysis applied to respondent comments uncovered barriers to sharing, including the inability to track uses and users of their information, potential harm (such as identity theft or healthcare denial), lack of trust, and worries about information misuse. Our study suggests that implementing reliable medical record de-identification and emphasizing trust development are essential to addressing such concerns. Amazon Mechanical Turk proved cost-effective for collecting public opinions with short surveys.
BACKGROUND:Manually curating standardized phenotypic concepts such as Human Phenotype Ontology (HPO) terms from narrative text in electronic health records (EHRs) is time consuming and error prone. Natural language processing (NLP) techniques can facilitate automated phenotype extraction and thus improve the efficiency of curating clinical phenotypes from clinical texts. While individual NLP systems can perform well for a single cohort, an ensemble-based method might shed light on increasing the portability of NLP pipelines across different cohorts. METHODS:We compared four NLP systems, MetaMapLite, MedLEE, ClinPhen and cTAKES, and four ensemble techniques, including intersection, union, majority-voting and machine learning, for extracting generic phenotypic concepts. We addressed two important research questions regarding automated phenotype recognition. First, we evaluated the performance of different approaches in identifying generic phenotypic concepts. Second, we compared the performance of different methods to identify patient-specific phenotypic concepts. To better quantify the effects caused by concept granularity differences on performance, we developed a novel evaluation metric that considered concept hierarchies and frequencies. Each of the approaches was evaluated on a gold standard set of clinical documents annotated by clinical experts. One dataset containing 1,609 concepts derived from 50 clinical notes from two different institutions was used in both evaluations, and an additional dataset of 608 concepts derived from 50 case report abstracts obtained from PubMed was used for evaluation of identifying generic phenotypic concepts only. RESULTS:For generic phenotypic concept recognition, the top three performers in the NYP/CUIMC dataset are union ensemble (F1, 0.634), training-based ensemble (F1, 0.632), and majority vote-based ensemble (F1, 0.622). In the Mayo dataset, the top three are majority vote-based ensemble (F1, 0.642), cTAKES (F1, 0.615), and MedLEE (F1, 0.559). In the PubMed dataset, the top three are majority vote-based ensemble (F1, 0.719), training-based (F1, 0.696) and MetaMapLite (F1, 0.694). For identifying patient specific phenotypes, the top three performers in the NYP/CUIMC dataset are majority vote-based ensemble (F1, 0.610), MedLEE (F1, 0.609), and training-based ensemble (F1, 0.585). In the Mayo dataset, the top three are majority vote-based ensemble (F1, 0.604), cTAKES (F1, 0.531) and MedLEE (F1, 0.527). CONCLUSIONS:Our study demonstrates that ensembles of natural language processing can improve both generic phenotypic concept recognition and patient specific phenotypic concept identification over individual systems. Among the individual NLP systems, each individual system performed best when they were applied in the dataset that they were primary designed for. However, combining multiple NLP systems to create an ensemble can generally improve the performance. Specifically, the ensemble can increase the results reproducibility across different cohorts and tasks, and thus provide a more portable phenotyping solution compared to individual NLP systems.
Drug-drug interactions (DDIs) constitute an important concern in drug development and postmarketing pharmacovigilance. They are considered the cause of many adverse drug effects exposing patients to higher risks and increasing public health system costs. Methods to follow-up and discover possible DDIs causing harm to the population are a primary aim of drug safety researchers. Here, we review different methodologies and recent advances using data mining to detect DDIs with impact on patients. We focus on data mining of different pharmacovigilance sources, such as the US Food and Drug Administration Adverse Event Reporting System and electronic health records from medical institutions, as well as on the diverse data mining studies that use narrative text available in the scientific biomedical literature and social media. We pay attention to the strengths but also further explain challenges related to these methods. Data mining has important applications in the analysis of DDIs showing the impact of the interactions as a cause of adverse effects, extracting interactions to create knowledge data sets and gold standards and in the discovery of novel and dangerous DDIs.
Medication regimen may be optimized based on individual drug efficacy identified by pharmacogenomic testing. However, majority of current pharmacogenomic decision support tools provide assessment only of single drug-gene interactions without taking into account complex drug-drug and drug-drug-gene interactions which are prevalent in people with polypharmacy and can result in adverse drug events or insufficient drug efficacy. The main objective of this project was to develop comprehensive pharmacogenomic decision support for medication risk assessment in people with polypharmacy that simultaneously accounts for multiple drug and gene effects. To achieve this goal, the project addressed two aims: (1) development of comprehensive knowledge repository of actionable pharmacogenes; (2) introduction of scoring approaches reflecting potential adverse effect risk levels of complex medication regimens accounting for pharmacogenomic polymorphisms and multiple drug metabolizing pathways. After pharmacogenomic knowledge repository was introduced, a scoring algorithm has been built and pilot-tested using a limited data set. The resulting total risk score for frequently hospitalized older adults with polypharmacy (72.04±17.84) was statistically significantly different (p<0.05) from the total risk score for older adults with polypharmacy with low hospitalization rate (8.98±2.37). An initial prototype assessment demonstrated feasibility of our approach and identified steps for improving risk scoring algorithms.
Integration of detailed phenotype information with genetic data is well established to facilitate accurate diagnosis of hereditary disorders. As a rich source of phenotype information, electronic health records (EHRs) promise to empower diagnostic variant interpretation. However, how to accurately and efficiently extract phenotypes from heterogeneous EHR narratives remains a challenge. Here, we present EHR-Phenolyzer, a high-throughput EHR framework for extracting and analyzing phenotypes. EHR-Phenolyzer extracts and normalizes Human Phenotype Ontology (HPO) concepts from EHR narratives and then prioritizes genes with causal variants on the basis of the HPO-coded phenotype manifestations. We assessed EHR-Phenolyzer on 28 pediatric individuals with confirmed diagnoses of monogenic diseases and found that the genes with causal variants were ranked among the top 100 genes selected by EHR-Phenolyzer for 16/28 individuals (p < 2.2 × 10-16), supporting the value of phenotype-driven gene prioritization in diagnostic sequence interpretation. To assess the generalizability, we replicated this finding on an independent EHR dataset of ten individuals with a positive diagnosis from a different institution. We then assessed the broader utility by examining two additional EHR datasets, including 31 individuals who were suspected of having a Mendelian disease and underwent different types of genetic testing and 20 individuals with positive diagnoses of specific Mendelian etiologies of chronic kidney disease from exome sequencing. Finally, through several retrospective case studies, we demonstrated how combined analyses of genotype data and deep phenotype data from EHRs can expedite genetic diagnoses. In summary, EHR-Phenolyzer leverages EHR narratives to automate phenotype-driven analysis of clinical exomes or genomes, facilitating the broader implementation of genomic medicine.
Pharmacogenetics-related publications, which are increasing rapidly, provide important new pharmacogenetics knowledge. Automated approaches to extract information of new alleles and to identify their impact on metabolic phenotypes from publications are urgently needed to facilitate personalized medicine and improve clinical outcomes. Cytochrome polymorphisms, responsible for a wide variation of drug pharmacodynamics, individual efficacy and adverse effects, have significant potential for optimizing drug therapy. A few studies have addressed specialized efforts to automatically extract cytochrome polymorphisms and their characterizations regarding metabolic phenotypes from the literature. In this paper, we present a novel rule-based text-mining system to extract metabolic phenotypes of polymorphisms from PubMed abstracts with a focus on cytochrome P450. This system is promising as it achieved a precision of 85.71% in a preliminary proof-of-concept evaluation and is expected to automatically provide up-to-date metabolic information for cytochrome polymorphisms, which is critical to advance personalized medicine and improve clinical care.
This study investigated the automated detection of antiretroviral toxicities in structured electronic health records data. The evaluation compared responses generated by 5 clinical pharmacists and 1 prototype knowledge-based application for 15 randomly selected test cases. The main outcomes were inter-subject dissimilarity of responses quantified by the Jaccard distance, and the mean proportion of correct responses by each subject. The statistical differences in inter-subject Jaccard distances suggested that the prototype was inferior to clinical pharmacists in the detection of possible antiretroviral toxicity associations from structured data. The reason for dissimilarities was attributable to inadequate domain coverage by the prototype. The differences in the mean proportion of correct responses between the clinical pharmacists and the prototype were statistically indistinguishable. Overall, this study suggests that knowledge-based applications have the potential to support automated detection of antiretroviral toxicities from structured patient records. Furthermore, the study demonstrates a systematic approach for validating such applications quantitatively.
BACKGROUND:It is beneficial for health care institutions to monitor physician prescribing patterns to ensure that high-quality and cost-effective care is being provided to patients. However, detecting treatment patterns within an institution is challenging, given that medications and conditions are often not explicitly linked in the health record. Here we demonstrate the use of statistical methods together with data from the electronic health care record (EHR) to analyze prescribing patterns at an institution.METHODS:As a demonstration of our method, which is based on regression, we collect EHR data from outpatient notes and use a case/control study design to determine the medications that are associated with hypertension. We also use regression to determine which conditions are associated with a preferential use of one or more classes of hypertension agents. Finally, we compare our method to methods based on tabulation.RESULTS:Our results show that regression methods provide more reasonable and useful results than tabulation, and successfully distinguish between medications that treat hypertension and medications that do not. These methods also provide insight into in which circumstances certain drugs are preferred over others.CONCLUSIONS:Our method can be used by health care institutions to monitor physician prescribing patterns and ensure the appropriateness of treatment.
OBJECTIVE:Improving mechanisms to detect adverse drug reactions (ADRs) is key to strengthening post-marketing drug safety surveillance. Signal detection is presently unimodal, relying on a single information source. Multimodal signal detection is based on jointly analyzing multiple information sources. Building on, and expanding the work done in prior studies, the aim of the article is to further research on multimodal signal detection, explore its potential benefits, and propose methods for its construction and evaluation. MATERIAL AND METHODS:Four data sources are investigated; FDA's adverse event reporting system, insurance claims, the MEDLINE citation database, and the logs of major Web search engines. Published methods are used to generate and combine signals from each data source. Two distinct reference benchmarks corresponding to well-established and recently labeled ADRs respectively are used to evaluate the performance of multimodal signal detection in terms of area under the ROC curve (AUC) and lead-time-to-detection, with the latter relative to labeling revision dates. RESULTS:Limited to our reference benchmarks, multimodal signal detection provides AUC improvements ranging from 0.04 to 0.09 based on a widely used evaluation benchmark, and a comparative added lead-time of 7-22 months relative to labeling revision dates from a time-indexed benchmark. CONCLUSIONS:The results support the notion that utilizing and jointly analyzing multiple data sources may lead to improved signal detection. Given certain data and benchmark limitations, the early stage of development, and the complexity of ADRs, it is currently not possible to make definitive statements about the ultimate utility of the concept. Continued development of multimodal signal detection requires a deeper understanding the data sources used, additional benchmarks, and further research on methods to generate and synthesize signals.
Information Retrieval (IR) and text analytics in the biomedical and healthcare domain has been attracting an abundant amount of research in the past decades. Similar to other domains, the main purpose used to be accurate retrieval of documents. However, in recent years, with the increased access to Electronic Health Records (EHRs), the interests and tasks are expanding (Hersh 2002). Retrieval of biomedical literature always had its unique methods due to the rich knowledge bases, such as the Unified Medical Language System (UMLS), Medical Subject Headings (MeSH), and the Systematized Nomenclature of Medicine (SNOMED) that enable the indexing of documents into concepts, for various purposes, such as retrieval (Hersh & Greens, 1989; Hersh & Hickam, 1992, 1993; Moskovitch et al., 2004; Lin & Demner-Fushman, 2006) and more (Moskovitch et al., 2006; Moskovitch & Shahar 2009; Ruch, 2006). The use of these domain-specific terminologies and vocabularies boosted the development of a large number of domainoriented retrieval methods. An important related thread of research has been the application of Natural Language Processing techniques to extract named entity concepts from data (Ruch, 2006), as well as other important pieces of information. In addition, to evaluate methods in biomedical text analytics and retrieval, several critical test data collections are now available, including TREC (Roberts et al., 2016) and more. In recent years, with the increased access to patients’ data in the form of EHRs, there are new problems and challenges related to the analysis and extraction from clinical notes and discharge summaries, particularly because the notes typically are telegraphic and include abbreviations, meta-data, semistructured text such as tables and billing codes, and use new lines to signal the end of a sentence. Another important retrieval task in the biomedical domain is from images that require image processing and retrieval. In addition to the data accumulated at hospital systems, and to the traditional clinical literature published in scientific journals and conferences and indexed in PubMed, there is an increase in healthrelated discussions in various relevant online forums and social medical sites. These forums span over multiple topics in the medical domain, in which patients share and discuss their experiences and questions. Consequently, the users of medical information retrieval may vary in their clinical knowledge, from physicians, medical students and related experts, to patients, or their relatives. They may also vary in their information needs, from scientific literature for experts to patients’ discussions on the internet. All these characteristics bring many challenges and opportunities to the scientific community. Finally, in the biomedical domain, in addition to improving the access to information through better and more efficient retrieval methodologies, there is the potential to improve the quality of care for patients. Thus, in this special issue, we encouraged participation from researchers in all fields related to medical information research, including mainstream information retrieval, but also natural language processing, multilingual text processing, and medical image analysis and retrieval.
Recent research has suggested that the case-control study design, unlike the self-controlled study design, performs poorly in controlling confounding in the detection of adverse drug reactions (ADRs) from administrative claims and electronic health record (EHR) data, resulting in biased estimates of the causal effects of drugs on health outcomes of interest (HOI) and inaccurate confidence intervals. Here we show that using rich data on comorbidities and automatic variable selection strategies for selecting confounders can better control confounding within a case-control study design and provide a more solid basis for inference regarding the causal effects of drugs on HOIs. Four HOIs are examined: acute kidney injury, acute liver injury, acute myocardial infarction and gastrointestinal ulcer hospitalization. For each of these HOIs we use a previously published reference set of positive and negative control drugs to evaluate the performance of our methods. Our methods have AUCs that are often substantially higher than the AUCs of a baseline method that only uses demographic characteristics for confounding control. Our methods also give confidence intervals for causal effect parameters that cover the expected no effect value substantially more often than this baseline method. The case-control study design, unlike the self-controlled study design, can be used in the fairly typical setting of EHR databases without longitudinal information on patients. With our variable selection method, these databases can be more effectively used for the detection of ADRs.
Academic literature provides rich and up-to-date information concerning adverse drug reactions (ADR), but it is time consuming and labor intensive for physicians to obtain information of ADRs from academic literature because they would have to generate queries, review retrieved articles and summarize the results. In this study, a method is developed to automatically detect and summarize ADRs from journal articles, rank them and present them to physicians in a user-friendly interface. The method studied ADRs for 6 drugs and returned on average 4.8 ADRs that were correct. The results demonstrated this method was feasible and effective. This method can be applied in clinical practice for assisting physicians to efficiently obtain information about ADRs associated with specific drugs. Automated summarization of ADR information from recent publications may facilitate translation of academic research into actionable information at point of care.
An automated, user-friendly and accurate system for retrieving herb-drug interaction (HDIs) related articles in MEDLINE can increase the safety of patients, as well as improve the physicians' article retrieving ability regarding speed and experience. Previous studies show that MeSH based queries associated with negative effects of drugs can be customized, resulting in good performance in retrieving relevant information, but no study has focused on the area of herb-drug interactions (HDI). This paper adapted the characteristics of HDI related papers and created a multilayer HDI article searching system. It achieved a sensitivity of 92% at a precision of 93% in a preliminary evaluation. Instead of requiring physicians to conduct PubMed searches directly, this system applies a more user-friendly approach by employing a customized system that enhances PubMed queries, shielding users from having to write queries, dealing with PubMed, or reading many irrelevant articles. The system provides automated processes and outputs target articles based on the input.