Objectives:Medical product safety surveillance efforts, whether using electronic health record (EHR) or claims data, typically rely on structured codes. Utilizing unstructured EHR data, particularly information extracted from clinical text through natural language processing (NLP), enriches information available for data mining, phenotyping, and surveillance. To assess overlapping and distinct information across structured and unstructured EHR data, we mapped both to a common vocabulary (Medical Dictionary for Regulatory Activities, MedDRA). We assess the feasibility of implementing such a mapping and explored similarities and differences at multiple levels of the concept hierarchy. Materials and Methods:We randomly sampled 15,000 encounters (5000 each from ambulatory, emergency, and inpatient encounters). For each encounter, we extracted MedDRA concepts from clinical notes using MetaMap and mapped structured ICD-10-CM diagnoses to MedDRA. We evaluated corroboration between data sources across the MedDRA hierarchy, as well as the unique information contributed by each source. Results:We processed 119,492 clinical notes and mapped 163,254 ICD-10-CM codes to MedDRA. Most encounters (73-98%) had some overlap between MedDRA preferred terms identified from structured and unstructured data. Among MedDRA concepts found in unstructured text, 80-95% were not found in the encounter's associated ICD-10-CM coded data. Discussion and Conclusion:While MedDRA concepts from structured data were mostly corroborated by those extracted from unstructured clinical text, the majority of MedDRA concepts recognized in each encounter were only mentioned in text. Leveraging MedDRA-encoded unstructured text can provide a more comprehensive clinical picture of patients and complement the structured data traditionally used in epidemiological and pharmacovigilance studies.
Objectives:To develop and validate machine learning (ML) models that predict probable cause of death (CoD) using structured electronic health record (EHR) data, unstructured clinical notes, and publicly available sources. Materials and Methods:This multi-institutional retrospective study was conducted across Vanderbilt University Medical Center (VUMC) and Massachusetts General Brigham (MGB), including deceased patients with encounters between October 1, 2015, and January 1, 2021, and confirmed death records. The cohort included 13 708 patients from VUMC and 34 839 from MGB.The primary outcome was underlying CoD categorized into the top 15 National Center for Health Statistics rankable causes, with others grouped as "Other." Performance was assessed using weighted area under the receiver operating characteristic curve (AUC) and F-measure. Results:The XGBoost model using structured EHR data alone achieved weighted AUCs of 0.86 (95% CI, 0.84-0.88) at VUMC and 0.80 (95% CI, 0.79-0.80) at MGB. Adding unstructured notes improved performance, with weighted AUCs of 0.90 (95% CI, 0.88-0.93) at VUMC and 0.92 (95% CI, 0.91-0.92) at MGB. Adding publicly available data did not further improve performance. Cross-institutional validation revealed significant performance degradation. Discussion:Models integrating structured and unstructured EHR data show strong within-institution performance but limited generalizability across healthcare systems, highlighting challenges related to institutional data heterogeneity. Conclusions:Machine learning models combining structured and unstructured EHR data accurately predict CoD within institutions but perform poorly across sites. Health-care institutions may benefit from adopting robust processes for locally tailored models, and future research should focus on enhancing model generalizability while addressing unique institutional data environments.
This study evaluated death ascertainment from publicly available internet sources for patients in two large tertiary care US healthcare systems, Mass General Brigham (MGB) and Vanderbilt University Medical Center (VUMC), benchmarked against state and federal vital statistics data. Names, dates of birth, and dates of death were extracted from 8.1 million internet media records using previously developed natural language processing models. Internet records were matched to 78 848 deceased patients from MGB and VUMC on first name, last name, and date of birth. Dates of death were validated against state vital statistics databases or the National Death Index as reference standards. We calculated sensitivity and positive predicted values (PPV) of internet sources in identifying dates of death within 7 days of the reference standard. Exact matching of records between internet media and reference standards on first name, last name, and date of birth, resulted in 30 067 (38.8%) matches, which showed PPV for death identification (98.2%-MGB; 98.9%-VUMC) in internet media and increased sensitivity of death capture over EHR alone by 24% at MGB and 18% at VUMC. In conclusion, using internet sources to augment mortality data increased capture of death meaningfully over reliance on EHR records alone.
Importance Timely and accurate determination of causes of death (CoD) is essential for public health surveillance, epidemiological research, and healthcare policy development. However, obtaining up-to-date and detailed CoD information is challenging due to delays in official death records and inconsistencies in data reporting across institutions. Objective To develop and validate machine learning (ML) models capable of predicting probable CoD by integrating comprehensive features from structured electronic health record (EHR) data, unstructured clinical notes, and publicly available data. Design, Setting, and Participants This multi-institutional retrospective cohort study was conducted at Vanderbilt University Medical Center (VUMC) and Massachusetts General Brigham (MGB). Deceased patients were included if they had at least one inpatient or outpatient encounter between October 1, 2015, and January 1, 2021, with corresponding death records from state health departments and the National Death Index. The study was comprised of 13,708 deceased patients from VUMC and 34,839 from MGB. Exposures Integration of structured EHR data, unstructured clinical notes processed using advanced language models, and publicly available data into machine learning models to predict CoD. Main Outcomes and Measures The primary outcome was the underlying CoD, classified into one of the top 15 National Center for Health Statistics (NCHS) rankable CoD categories, with all other causes grouped into an “Other” category. Model performance was evaluated using weighted area under the receiver operating characteristic curve (AUC) and weighted F-measure. Results The XGBoost model using structured EHR data alone achieved weighted AUCs of 0.86 (95% CI, 0.84–0.88) at VUMC and 0.80 (95% CI, 0.79-0.80) at MGB. Adding unstructured notes improved performance, with weighted AUCs of 0.90 (95% CI, 0.88–0.93) at VUMC and 0.92(95% CI, 0.91–0.92) at MGB. Adding publicly available data did not further improve performance. Cross-institutional validation revealed significant performance degradation. Conclusions and Relevance ML models integrating EHR structured and unstructured data to predict underlying CoD at the time of the most recent encounter among deceased patients achieved excellent performance within individual institutions. The inclusion of publicly available data did not improve performance, and all versions had poor portability between institutions. Healthcare institutions may benefit from adopting robust processes for locally tailored models, and future research should focus on enhancing model generalizability while addressing unique institutional data environments. ### Competing Interest Statement The authors have declared no competing interest. ### Funding Statement Funding sources: The contents are those of the author(s) and do not necessarily represent the official views of, nor an endorsement, by FDA/HHS, or the U.S. Government. This project was supported by Task Order 75F40119F19002 under Master Agreement 75F40119D10037 from the US Food and Drug Administration (FDA). ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: The Institutional Review Board of Vanderbilt University Medical Center and the Institutional Review Board of Mass General Brigham gave ethical approval for this work. I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes The data used in this study are derived from electronic health records maintained by Vanderbilt University Medical Center (VUMC) and Massachusetts General Brigham (MGB), as well as mortality datasets provided by the National Death Index and the state health departments of Massachusetts, Connecticut, and Vermont. Access to these data is restricted due to institutional data use policies and patient privacy regulations.
Cohort characterization, comparative effectiveness research (CER), and patient-level prediction are limited by patient conditions being incompletely recorded in structured electronic health records (EHRs) and administrative claims data. This is particularly true of mental health (MH) phenotypes. We recently used noisy label learning on administrative claims data in patients with major mental illness to impute uncoded self-harm, 1 and used it as an outcome to enhance statistical power in a comparative effectiveness study. 2 We estimated that only about 1 in 19 self-harm events were coded, but these estimates require evaluation and validation. Noisy label machine learning (ML) can rank-order patients by the probability that they might have an MH condition. 1–3 Classical approaches calibrate probability thresholds for desired positive predictive values and other ML metrics using samples of people who have been clinically assessed as both positive and negative for a condition. We now report evaluation of methods for estimating the true proportion of positives among patients with uncoded or undiagnosed conditions without such “gold standard” assessments. 3–5 We employ positive and unlabeled learning ( PU-learning) , a noisy label ML method that uses: a) a set of known positives [e.g. with diagnoses of post-traumatic stress disorder (PTSD)], b) an unlabeled set of individuals with an unknown proportion of positives and negatives, and c) an ML model to distinguish the two. PU-learning then estimates the class prior (proportion of positives) among the unknowns. We describe our PU-learning method, assess its performance on simulated data to detect rare and common phenotypes, and use it to detect self-harm and PTSD. including were tested on simulated data, then applied to VHA PTSD and self-harm.