To examine whether complementary and integrative health approaches mitigate opioid prescriptions for pain and whether the relationship differs by post-dramatic stress disorder (PTSD) diagnosis, we followed 1,993,455 Veterans with musculoskeletal disorders during 2005-2017 using Veterans Healthcare Administration electronic health records. Complementary and integrative health (CIH) approaches were defined as >= 1 primary care visits for meditation, Yoga, and acupuncture etc using natural language processing. Opioid prescriptions were ascertained from pharmacy dispensing records. A propensity score was estimated and used to match one control Veteran to each CIH recipient. Over the 2-year follow-up period after the index diagnosis, 140,902 (7.1 %) Veterans received >= 1 modalities. Among the matched analytic sample (272,296 Veterans), the likelihood of dispensing opioid prescriptions was significantly lower for Veterans in the CIH group than their controls [adjusted hazard ratio (aHR), 0.45 (95 % Confidence Intervals (CI): 0.44-0.46)]. The association did not differ between Veterans with [aHR: 0.46 (95 % CI: 0.45-0.47)] and without [aHR: 0.44 (95 % CI: 0.43-0.45)] PTSD. In sensitivity analyses, the exposure group had 3.82 (95 % CI: 3.76-3.87) months longer restricted mean survival time to opioid initiation, 2 % (95 % CI: 4 %-1 %) lower morphine equivalent and 17 % lower total days' supply (95 % CI: 18 %-16 %). The relationship remains significant but was attenuated after eliminating waiting time for the exposure group (aHR, 0.63 (95 % CI: 0.62-0.64)). These observations suggest that CIH approaches may help reduce opioid prescriptions for Veterans with musculoskeletal disorders and related pain. The impact of the timing of receiving such approaches warrants further investigation. Perspective: This article presents a quasi-experimental investigation into potential benefit of complementary and integrative health approaches (CIH) on de-prescribing opioids. The findings may potentially help clinicians who are seeking non-pharmacological alternative options to manage patient pain and opioid dependence".
Purpose We describe pain intensity and opioid prescription jointly over time in Veterans with back pain to better understand their relationship. Methods We performed a retrospective cohort study on electronic health record data from 117 126 Veterans (mean age 49.2 years) diagnosed with back pain in 2015. We used latent class growth analysis to jointly model pain intensity (0-10 scores) and opioid prescriptions over 2 years to identify classes of individuals similar in their trajectory of pain and opioid over time. Multivariable multinomial logit models assessed sociodemographic and clinical predictors of class membership. Results We identified six trajectory classes: a "no pain/no opioid" class (22.2%), a "mild pain/no opioid" class (45.0%), a "moderate pain/no opioid" class (24.6%), a "moderate, decreasing pain/decreasing opioid" class (3.3%), a "moderate pain/high opioid" class (2.6%), and a "moderate, increasing pain/increasing opioid" class (2.3%). Among those in moderate pain classes, being white (vs. non-white) and older were associated with higher odds of being prescribed opioids. Veterans with mental health diagnoses had increased odds of being in the painful classes versus "no pain/no opioid" class. Conclusion We found distinct patterns in the long-term joint course of pain and opioid prescription in Veterans with back pain. Understanding these patterns and associated predictors may help with development of targeted interventions for patients with back pain.
Complementary and integrative health approaches (CIH) are recommended in national guidelines for managing chronic pain and de-prescribing opioids. We followed 1,993,455 opioid-naïve Veterans with musculoskeletal disorders for two years after index diagnosis during 2005-2017. CIH exposure was defined as primary care visits for acupuncture, massage and chiropractic therapy using natural language processing and administratively coded data. Opioid prescriptions dispensed during follow-up period were abstracted from Veterans Health Administration electronic pharmacy records. Propensity score (PS) was used to match one control for each CIH recipient. Overall, 140,902 (7.1%) Veterans received CIH, with those age ≥65y the lowest prevalence (2.7%). Cox proportional hazard model revealed that time to first opioid prescriptions was longer for CIH recipients than PS-matched controls (136,148 matched pairs) and varied across age (p for interaction 0.003). The adjusted Hazard Ratio (HR) was 0.48 (95% Confidence Interval (CI): 0.45-0.51) for Veterans age ≥65y, 0.44 (95% CI: 0.43-0.45) for 50-64y and 0.47 (95% CI: 0.46-0.48) for ≤49y. Restricted mean survival time (RMST) models estimated a smaller CIH benefit for Veterans ≥65y, with an average 3.3 (95% CI: 3.2-3.5) month RMST difference, in contrast to 4.2 (95% CI: 4.1-4.3) and 3.7 (95% CI: 3.6-3.8) months for younger counterparts. Sensitivity analyses in full cohort or modeling total supply and daily dose of opioid prescriptions derived consistent results. These findings suggest potential benefits of CIH in delaying and reducing opioids prescriptions for patients with chronic pain. The observations of lower rate and smaller benefit of CIH use among older Veterans warrants further investigations.
Mental health is an increasing concern in adolescents. Mental health disorders can affect academic performance, affect the cultivation of healthy relationships, and even lead to suicide. Healthy lifestyle can improve mental health, though there are gaps in the research, partly resulted from the lack of detailed longitudinal datasets on lifestyle and mental health. To inform and engage students in the research on adolescent lifestyle and mood, the George Washington University and the T.C. Williams High School in Alexandria, Virginia teamed up in a citizen science project. Students generated questions, collected data on themselves, analyzed the data, and produced research reports relating to their mental health and lifestyle. Student feedbacks suggest that the students find the project to be generally interesting and some students (46%) reported that the participation in the project may influence their college and career plans. The anonymized dataset resulted from the project provides another contribution to science.
BACKGROUND:The data quality of electronic health records (EHR) has been a topic of increasing interest to clinical and health services researchers. One indicator of possible errors in data is a large change in the frequency of observations in chronic illnesses. In this study, we built and demonstrated the utility of a stacked multivariate LSTM model to predict an acceptable range for the frequency of observations. METHODS:We applied the LSTM approach to a large EHR dataset with over 400 million total encounters. We computed sensitivity and specificity for predicting if the frequency of an observation in a given week is an aberrant signal. RESULTS:Compared with the simple frequency monitoring approach, our proposed multivariate LSTM approach increased the sensitivity of finding aberrant signals in 6 randomly selected diagnostic codes from 75 to 88% and the specificity from 68 to 91%. We also experimented with two different LSTM algorithms, namely, direct multi-step and recursive multi-step. Both models were able to detect the aberrant signals while the recursive multi-step algorithm performed better. CONCLUSIONS:Simply monitoring the frequency trend, as is the common practice in systems that do monitor the data quality, would not be able to distinguish between the fluctuations caused by seasonal disease changes, seasonal patient visits, or a change in data sources. Our study demonstrated the ability of stacked multivariate LSTM models to recognize true data quality issues rather than fluctuations that are caused by different reasons, including seasonal changes and outbreaks.
Abstract Family caregivers of community-dwelling older adults have faced unprecedented caregiving challenges during the COVID-19 pandemic. Examining the accumulated impact on family caregivers can help health and aging service providers design resources and supports that are resilient to emergency situations, and reduce negative psychological and physical consequences and risk of abuse within caregiving dyads. Data was collected as part of a pilot intervention in which “Care Coaches” provided telephonic coaching sessions to family caregivers of older adults. We examined Care Coach observations documented after coaching sessions with 24 family caregivers between March 2020 and February 2021. Two coders employed thematic analysis to generate codes and themes. The sample was 70% female, 80% were the spouse or significant other of their care receiver, the mean age was 61, and 53% were Non-Hispanic White. Themes and sub-themes included: (1) increased caregiver burden and diminished care networks due to fear of exposure to or contraction of COVID-19, (2) barriers to accessing in-home personal assistance services and home-delivered meals despite intervention efforts, and (3) the exacerbation of caregiver social isolation due to COVID-19 lockdown policies. Findings highlight the ways in which COVID-19 has amplified caregiver burden through the breakdown of formal and informal support systems. Potential adaptations of community-based services for older adults and their caregivers include remote service liaisons and need assessment of caregiver dyads to assure access to home-based personal assistance services and nutrition support for those at greatest risk of negative consequences during emergency service lapses.
The field of clinical natural language processing (NLP) has been built on the analysis of clinical sublanguage characteristics. It is well recognized that not only does clinical sublanguage differ from general English (or other languages) but also clinical sublanguage differs among clinical subspecialties and among corpora originating from different healthcare systems. A less recognized aspect is that clinical sublanguage, like all languages, evolves over time. This paper analyses the evolution of clinical sublanguage using a large, national clinical text corpus spanning 15 years. Through the analyses of document types, length, ngrams, and concepts, we found strong evidence that clinical sublanguage does evolve and such changes have implications for NLP development and maintenance. Although the analysis is performed on one corpus, our observations of sublanguage changes are generalizable.
Supplementing patient education content with pictographs can improve the comprehension and recall of information, especially patients with low health literacy. Pictograph design and testing, however, are costly and time consuming. We created a Web-based game, Doodle Health, for crowdsourcing the drawing and validation of pictographs. The objective of this pilot study was to test the usability of the game and its appeal to healthcare consumers. The chief purpose of the game is to involve a diverse population in the co-design and evaluation of pictographs. We conducted a community-based focus group to inform the game design. Game designers, health sciences librarians, informatics researchers, clinicians, and community members participated in two Design Box meetings. The results of the meetings were used to create the Doodle Health crowdsourcing game. The game was presented and tested at two public fairs. Initial testing indicates crowdsourcing is a promising approach to pictograph development and testing for relevancy and comprehension. Over 596 drawings were collected and 1,758 guesses were performed to date with 70-90% accuracies, which are satisfactorily high.
OBJECTIVE:Large database research in axial spondyloarthritis (SpA) is limited by a lack of methods for identifying most types of axial SpA. Our objective was to develop methods for identifying axial SpA concepts in the free text of documents from electronic medical records. METHODS:Veterans with documents in the national Veterans Health Administration Corporate Data Warehouse between January 1, 2005 and June 30, 2015 were included. Methods were developed for exploring, selecting, and extracting meaningful terms that were likely to represent axial SpA concepts. With annotation, clinical experts reviewed sections of text containing the meaningful terms (snippets) and classified the snippets according to whether or not they represented the intended axial SpA concept. With natural language processing (NLP) tools, computers were trained to replicate the clinical experts' snippet classifications. RESULTS:Three axial SpA concepts were selected by clinical experts, including sacroiliitis, terms including the prefix spond*, and HLA-B27 positivity (HLA-B27+). With supervised machine learning on annotated snippets, NLP models were developed with accuracies of 91.1% for sacroiliitis, 93.5% for spond*, and 97.2% for HLA-B27+. With independent validation, the accuracies were 92.0% for sacroiliitis, 91.0% for spond*, and 99.0% for HLA-B27+. CONCLUSION:We developed feasible and accurate methods for identifying axial SpA concepts in the free text of clinical notes. Additional research is required to determine combinations of concepts that will accurately identify axial SpA phenotypes. These novel methods will facilitate previously impractical observational research in axial SpA and may be applied to research with other diseases.
In an ideal clinical Natural Language Processing (NLP) ecosystem, researchers and developers would be able to collaborate with others, undertake validation of NLP systems, components, and related resources, and disseminate them. We captured requirements and formative evaluation data from the Veterans Affairs (VA) Clinical NLP Ecosystem stakeholders using semi-structured interviews and meeting discussions. We developed a coding rubric to code interviews. We assessed inter-coder reliability using percent agreement and the kappa statistic. We undertook 15 interviews and held two workshop discussions. The main areas of requirements related to; design and functionality, resources, and information. Stakeholders also confirmed the vision of the second generation of the Ecosystem and recommendations included; adding mechanisms to better understand terms, measuring collaboration to demonstrate value, and datasets/tools to navigate spelling errors with consumer language, among others. Stakeholders also recommended capability to: communicate with developers working on the next version of the VA electronic health record (VistA Evolution), provide a mechanism to automatically monitor download of tools and to automatically provide a summary of the downloads to Ecosystem contributors and funders. After three rounds of coding and discussion, we determined the percent agreement of two coders to be 97.2% and the kappa to be 0.7851. The vision of the VA Clinical NLP Ecosystem met stakeholder needs. Interviews and discussion provided key requirements that inform the design of the VA Clinical NLP Ecosystem.
INTRODUCTION:Substantial amounts of clinically significant information are contained only within the narrative of the clinical notes in electronic medical records. The v3NLP Framework is a set of "best-of-breed" functionalities developed to transform this information into structured data for use in quality improvement, research, population health surveillance, and decision support.BACKGROUND:MetaMap, cTAKES and similar well-known natural language processing (NLP) tools do not have sufficient scalability out of the box. The v3NLP Framework evolved out of the necessity to scale-up these tools up and provide a framework to customize and tune techniques that fit a variety of tasks, including document classification, tuned concept extraction for specific conditions, patient classification, and information retrieval.INNOVATION:Beyond scalability, several v3NLP Framework-developed projects have been efficacy tested and benchmarked. While v3NLP Framework includes annotators, pipelines and applications, its functionalities enable developers to create novel annotators and to place annotators into pipelines and scaled applications.DISCUSSION:The v3NLP Framework has been successfully utilized in many projects including general concept extraction, risk factors for homelessness among veterans, and identification of mentions of the presence of an indwelling urinary catheter. Projects as diverse as predicting colonization with methicillin-resistant Staphylococcus aureus and extracting references to military sexual trauma are being built using v3NLP Framework components.CONCLUSION:The v3NLP Framework is a set of functionalities and components that provide Java developers with the ability to create novel annotators and to place those annotators into pipelines and applications to extract concepts from clinical text. There are scale-up and scale-out functionalities to process large numbers of records.
Background: Cohort identification is important in both population health management and research. In this project we sought to assess the use of text queries for cohort identification. Specifically we sought to determine the incremental value of unstructured data queries when added to structured queries for the purpose of patient cohort identification.Methods: Three cohort identification tasks were evaluated: identification of individuals taking gingko biloba and warfarin simultaneously (Gingko/Warfarin), individuals who were overweight, and individuals with uncontrolled diabetes (UCD). We assessed the increase in cohort size when unstructured data queries were added to structured data queries. The positive predictive value of unstructured data queries was assessed by manual chart review of a random sample of 500 patients.Results: For Gingko/Warfarin, text query increased the cohort size from 9 to 28,924 over the cohort identified by query of pharmacy data only. For the weight-related tasks, text search increased the cohort by 5-29% compared to the cohort identified by query of the vitals table. For the UCD task, text query increased the cohort size by 2-43% compared to the cohort identified by query of laboratory results or ICD codes. The positive predictive values for text searches were 52% for Gingko/Warfarin, 19-94% for the weight cohort and 44% for UCD.Discussion: This project demonstrates the value and limitation of free text queries in patient cohort identification from large data sets. The clinical domain and prevalence of the inclusion and exclusion criteria in the patient population influence the utility and yield of this approach. Published by Elsevier Ltd.
Background: Bodyweight related measures (weight, height, BMI, abdominal circumference) are extremely important for clinical care, research and quality improvement. These and other vitals signs data are frequently missing from structured tables of electronic health records. However they are often recorded as text within clinical notes. In this project we sought to develop and validate a learning algorithm that would extract bodyweight related measures from clinical notes in the Veterans Administration (VA) Electronic Health Record to complement the structured data used in clinical research.Methods: We developed the Regular Expression Discovery Extractor (REDEx), a supervised learning algorithm that generates regular expressions from a training set. The regular expressions generated by REDEx were then used to extract the numerical values of interest.Methods: To train the algorithm we created a corpus of 268 outpatient primary care notes that were annotated by two annotators. This annotation served to develop the annotation process and identify terms associated with bodyweight related measures for training the supervised learning algorithm. Snippets from an additional 300 outpatient primary care notes were subsequently annotated independently by two reviewers to complete the training set. Inter-annotator agreement was calculated.Methods: REDEx was applied to a separate test set of 3561 notes to generate a dataset of weights extracted from text. We estimated the number of unique individuals who would otherwise not have bodyweight related measures recorded in the CDW and the number of additional bodyweight related measures that would be additionally captured.Results: REDEx's performance was: accuracy = 983%, precision = 98.8%, recall = 98.3%, F = 98.5%. In the dataset of weights from 3561 notes, 7.7% of notes contained bodyweight related measures that were not available as structured data. In addition 2 additional bodyweight related measures were identified per individual per year.Conclusion: Bodyweight related measures are frequently stored as text in clinical notes. A supervised learning algorithm can be used to extract this data. Implications for clinical care, epidemiology, and quality improvement efforts are discussed. (C) 2015 Elsevier Inc. All rights reserved.
Ginkgo biloba is a widely used herbal product that could potentially have a severe interaction with warfarin, which is the most frequently prescribed anticoagulant agent in North America. Literature, however, provides conflicting evidence on the presence and severity of the interaction. In this study, we developed text processing methods to extract the ginkgo usage and combined it with prescription data on warfarin from a very large clinical data respository. Our statistical analysis suggests that taking concurrently with warfarin, gingko does significantly increase patients' risk of a bleeding adverse event (hazard ratio = 1.38, 95%CI: 1.20 to 1.58, p<.001). This study also is the first attempt of using a large medical record databaseto confirm a suspected herb-drug interaction.
BackgroundElectronic medical records (EMR) provide an ideal opportunity for the detection, diagnosis, and management of systemic sclerosis (SSc) patients within the Veterans Health Administration (VHA). The objective of this project was to use informatics to identify potential SSc patients in the VHA that were on prednisone, in order to inform an outreach project to prevent scleroderma renal crisis (SRC).MethodsThe electronic medical data for this study came from Veterans Informatics and Computing Infrastructure (VINCI). For natural language processing (NLP) analysis, a set of retrieval criteria was developed for documents expected to have a high correlation to SSc. The two annotators reviewed the ratings to assemble a single adjudicated set of ratings, from which a support vector machine (SVM) based document classifier was trained. Any patient having at least one document positively classified for SSc was considered positive for SSc and the use of prednisone≥10mg in the clinical document was reviewed to determine whether it was an active medication on the prescription list.ResultsIn the VHA, there were 4272 patients that have a diagnosis of SSc determined by the presence of an ICD-9 code. From these patients, 1118 patients (21%) had the use of prednisone≥10mg. Of these patients, 26 had a concurrent diagnosis of hypertension, thus these patients should not be on prednisone. By the use of natural language processing (NLP) an additional 16,522 patients were identified as possible SSc, highlighting that cases of SSc in the VHA may exist that are unidentified by ICD-9. A 10-fold cross validation of the classifier resulted in a precision (positive predictive value) of 0.814, recall (sensitivity) of 0.973, and f-measure of 0.873.ConclusionsOur study demonstrated that current clinical practice in the VHA includes the potentially dangerous use of prednisone for veterans with SSc. This present study also suggests there may be many undetected cases of SSc and NLP can successfully identify these patients.
Integrative medicine including complementary and alternative medicine (CAM) has become more available through mainstream health providers. Acupuncture is one of the most widely used CAM therapies, though its efficacy for treating various conditions requires further investigation. To assist with such investigations, we set out to identify acupuncture patient cohorts using a nationwide clinical data repository. Acupuncture patients were identified using both structured data and unstructured free text notes: 44,960 acupuncture patients were identified using structured data consisting of CPT codes;. Using unstructured free text clinical notes, we trained a support vector classifier with 86% accuracy and was able to identify an additional 101,628 acupuncture patients not identified through structured data (a 226% increase). In addition, characteristics of the patients identified through structured and unstructured data were compared, which show differences in geographic locations and medical service usage patterns. Patients identified with structured data displayed a consistently higher use of the Veterans Health Administration (VHA) medical system.
Query expansion is a commonly used approach to improving search results. Specific expansion methods, however, are expected to have different results. We have developed three different expansion methods using knowledge derived from medical thesaurus, medical literature, and clinical notes. Since the three different sources each have strengths and weaknesses, we hypothesized that combining the three sources will lead to better retrieval performance. Evaluation was performed for the 3 different query expansion techniques and an ensemble method on two sets of clinical notes. 11-point interpolated average precisions, MAP, and P(10) scores were calculated which indicate that topic model based expansion has the best results and the predication method the worst. This finding points to the potential of the topic modeling methods as well as the challenge in integrating different knowledge sources.
: In our TREC participation, we used an ensemble approach in query expansion. Query expansion, such as synonym expansion, had shown promising results in medical literature search. On the other hand, some of the 2011 papers reported worse results from expansion. Since there are multiple knowledge sources available and each resource has clear strengths and weaknesses, we tested the combination of three expansion methods versus each individual method. We found that the ensemble approach performed better (in terms of average infAP, infNDCG, R-prec, and P10) than the individual methods and better than the Lucene baseline. The individual expansion methods, however, did not improve the baseline Lucene performance. We also performed an unofficial run using a concept index to boost the query performance, which led to small improvements in infAP, infNDCG, and R-prec.