Identification of patient cohorts from EHRs is challenging because ICD codes primarily serve billing and may misrepresent disease status, while key information is buried in unstructured notes. Existing computed phenotyping methods also have limitations in maintenance and incomplete modeling. We evaluated GPT-4o’s type II diabetes mellitus (T2DM) phenotyping ability using optimized Retrieval-Augmented Generation (RAG). We built a RAG pipeline and clinical notes were loaded for 275 patients screened by T2DM ICD codes. We optimized chunk size and top-k across seven embedding models, testing 308 RAG configurations using training patients. Prompts (zero-shot and few-shot) were developed via error analysis. GPT-4o’s phenotyping performance was evaluated against ICD codes and PheNorm, within the optimized RAG framework. Token usage and sensitivity to key hyperparameters were also assessed. GPT-4o with optimized RAG significantly outperformed ICD in precision (PPV: 0.940), and PheNorm in sensitivity (0.902), NPV (0.697), and F1 (0.920), while PPV was slightly lower and specificity (0.791) needs improvement compared to PheNorm. General embedding models and zero-shot prompt presented better sensitivity, NPV, and F1-scores, while domain-specific models and a few-shot prompt excelled in specificity and PPV. Optimization enabled lower-ranked embedding models to achieve comparably good performance to the highest ones. Gte-Qwen2-1.5B-instruct and GatorTronS provided the highest token-efficiency in specific metrics. Error analysis revealed contextual misinterpretation and ranking issues. GPT-4 using optimized RAG showed superior in T2DM phenotyping in key metrics. This study provides valuable insights into practical guidance of using RAG, while identifying limitations in errors LLM reasoning and retrieval ranking.
Pragmatic clinical trials (PCTs) evaluate interventions in real-world settings, often using electronic health records (EHRs) for efficient data collection. We report on the challenges in performing EHR analysis of healthcare provider orders in a PCT within the eMERGE consortium, which investigates the impact of reporting genome-informed risk assessments (GIRA) to over 25,000 patients across 10 academic medical centers. Clinical informaticians conducted a landscape analysis to identify approaches for evaluating the outcomes of GIRA reporting through the EHR. Of 98 identified outcomes, 54 (55.1%) were determined to be difficult to extract because they involved provider orders, which are typically documented in free text or proprietary formats within the EHR and only mapped to standardized codes after the service is completed. These findings highlight a critical barrier in using EHRs to support PCTs. The authors recommend closer collaboration between clinicians and informaticians, improved EHR systems that support standardized order entry, and future use of machine learning to automate analysis of provider behavior in clinical trials.
With the burgeoning development of computational phenotypes, it is increasingly difficult to identify the right phenotype for the right tasks. This study uses a mixed-methods approach to develop and evaluate a novel metadata framework for retrieval of and reusing computational phenotypes. Twenty active phenotyping researchers from 2 large research networks, Electronic Medical Records and Genomics and Observational Health Data Sciences and Informatics, were recruited to suggest metadata elements. Once consensus was reached on 39 metadata elements, 47 new researchers were surveyed to evaluate the utility of the metadata framework. The survey consisted of 5-Likert multiple-choice questions and open-ended questions. Two more researchers were asked to use the metadata framework to annotate 8 type-2 diabetes mellitus phenotypes. More than 90% of the survey respondents rated metadata elements regarding phenotype definition and validation methods and metrics positively with a score of 4 or 5. Both researchers completed annotation of each phenotype within 60 min. Our thematic analysis of the narrative feedback indicates that the metadata framework was effective in capturing rich and explicit descriptions and enabling the search for phenotypes, compliance with data standards, and comprehensive validation metrics. Current limitations were its complexity for data collection and the entailed human costs.
BACKGROUND:Type 2 diabetes (T2D) is a worldwide scourge caused by both genetic and environmental risk factors that disproportionately afflicts communities of color. Leveraging existing large-scale genome-wide association studies (GWAS), polygenic risk scores (PRS) have shown promise to complement established clinical risk factors and intervention paradigms, and improve early diagnosis and prevention of T2D. However, to date, T2D PRS have been most widely developed and validated in individuals of European descent. Comprehensive assessment of T2D PRS in non-European populations is critical for equitable deployment of PRS to clinical practice that benefits global populations. METHODS:We integrated T2D GWAS in European, African, and East Asian populations to construct a trans-ancestry T2D PRS using a newly developed Bayesian polygenic modeling method, and assessed the prediction accuracy of the PRS in the multi-ethnic Electronic Medical Records and Genomics (eMERGE) study (11,945 cases; 57,694 controls), four Black cohorts (5137 cases; 9657 controls), and the Taiwan Biobank (4570 cases; 84,996 controls). We additionally evaluated a post hoc ancestry adjustment method that can express the polygenic risk on the same scale across ancestrally diverse individuals and facilitate the clinical implementation of the PRS in prospective cohorts. RESULTS:The trans-ancestry PRS was significantly associated with T2D status across the ancestral groups examined. The top 2% of the PRS distribution can identify individuals with an approximately 2.5-4.5-fold of increase in T2D risk, which corresponds to the increased risk of T2D for first-degree relatives. The post hoc ancestry adjustment method eliminated major distributional differences in the PRS across ancestries without compromising its predictive performance. CONCLUSIONS:By integrating T2D GWAS from multiple populations, we developed and validated a trans-ancestry PRS, and demonstrated its potential as a meaningful index of risk among diverse patients in clinical settings. Our efforts represent the first step towards the implementation of the T2D PRS into routine healthcare.
Introduction Currently, one of the commonly used methods for disseminating electronic health record (EHR)-based phenotype algorithms is providing a narrative description of the algorithm logic, often accompanied by flowcharts. A challenge with this mode of dissemination is the potential for under-specification in the algorithm definition, which leads to ambiguity and vagueness. Methods This study examines incidents of under-specification that occurred during the implementation of 34 narrative phenotyping algorithms in the electronic Medical Record and Genomics (eMERGE) network. We reviewed the online communication history between algorithm developers and implementers within the Phenotype Knowledge Base (PheKB) platform, where questions could be raised and answered regarding the intended implementation of a phenotype algorithm. Results We developed a taxonomy of under-specification categories via an iterative review process between two groups of annotators. Under-specifications that lead to ambiguity and vagueness were consistently found across narrative phenotype algorithms developed by all involved eMERGE sites. Discussion and conclusion Our findings highlight that under-specification is an impediment to the accuracy and efficiency of the implementation of current narrative phenotyping algorithms, and we propose approaches for mitigating these issues and improved methods for disseminating EHR phenotyping algorithms.
AbstractObjectiveIntegrating and harmonizing disparate patient data sources into one consolidated data portal enables researchers to conduct analysis efficiently and effectively.Materials and MethodsWe describe an implementation of Informatics for Integrating Biology and the Bedside (i2b2) to create the Mass General Brigham (MGB) Biobank Portal data repository. The repository integrates data from primary and curated data sources and is updated weekly. The data are made readily available to investigators in a data portal where they can easily construct and export customized datasets for analysis.ResultsAs of July 2021, there are 125 645 consented patients enrolled in the MGB Biobank. 88 527 (70.5%) have a biospecimen, 55 121 (43.9%) have completed the health information survey, 43 552 (34.7%) have genomic data and 124 760 (99.3%) have EHR data. Twenty machine learning computed phenotypes are calculated on a weekly basis. There are currently 1220 active investigators who have run 58 793 patient queries and exported 10 257 analysis files.DiscussionThe Biobank Portal allows noninformatics researchers to conduct study feasibility by querying across many data sources and then extract data that are most useful to them for clinical studies. While institutions require substantial informatics resources to establish and maintain integrated data repositories, they yield significant research value to a wide range of investigators.ConclusionThe Biobank Portal and other patient data portals that integrate complex and simple datasets enable diverse research use cases. i2b2 tools to implement these registries and make the data interoperable are open source and freely available.
Chronic Kidney Disease (CKD) represents a slowly progressive disorder that is typically silent until late stages, but early intervention can significantly delay its progression. We designed a portable and scalable electronic CKD phenotype to facilitate early disease recognition and empower large-scale observational and genetic studies of kidney traits. The algorithm uses a combination of rule-based and machine-learning methods to automatically place patients on the staging grid of albuminuria by glomerular filtration rate (“A-by-G” grid). We manually validated the algorithm by 451 chart reviews across three medical systems, demonstrating overall positive predictive value of 95% for CKD cases and 97% for healthy controls. Independent case-control validation using 2350 patient records demonstrated diagnostic specificity of 97% and sensitivity of 87%. Application of the phenotype to 1.3 million patients demonstrated that over 80% of CKD cases are undetected using ICD codes alone. We also demonstrated several large-scale applications of the phenotype, including identifying stage-specific kidney disease comorbidities, in silico estimation of kidney trait heritability in thousands of pedigrees reconstructed from medical records, and biobank-based multicenter genome-wide and phenome-wide association studies.
Clinical data networks that leverage large volumes of data in electronic health records (EHRs) are significant resources for research on coronavirus disease 2019 (COVID-19). Data harmonization is a key challenge in seamless use of multisite EHRs for COVID-19 research. We developed a COVID-19 application ontology in the national Accrual to Clinical Trials (ACT) network that enables harmonization of data elements that are critical to COVID-19 research. The ontology contains over 50 000 concepts in the domains of diagnosis, procedures, medications, and laboratory tests. In particular, it has computational phenotypes to characterize the course of illness and outcomes, derived terms, and harmonized value sets for severe acute respiratory syndrome coronavirus 2 laboratory tests. The ontology was deployed and validated on the ACT COVID-19 network that consists of 9 academic health centers with data on 14.5M patients. This ontology, which is freely available to the entire research community on GitHub at https://github.com/shyamvis/ACT-COVID-Ontology, will be useful for harmonizing EHRs for COVID-19 research beyond the ACT network.
Clinical data networks that leverage large volumes of data in electronic health records (EHRs) are significant resources for research on coronavirus disease 2019 (COVID-19). Data harmonization is a key challenge in seamless use of multisite EHRs for COVID-19 research. We developed a COVID-19 application ontology in the national Accrual to Clinical Trials (ACT) network that enables harmonization of data elements that that are critical to COVID-19 research. The ontology contains over 50,000 concepts in the domains of diagnosis, procedures, medications, and laboratory tests. In particular, it has computational phenotypes to characterize the course of illness and outcomes, derived terms, and harmonized value sets for SARS-CoV-2 laboratory tests. The ontology was deployed and validated on the ACT COVID-19 network that consists of nine academic health centers with data on 14.5M patients. This ontology, which is freely available to the entire research community on GitHub at https://github.com/shyamvis/ACT-COVID-Ontology, will be useful for harmonizing EHRs for COVID-19 research beyond the ACT network.
The Electronic Medical Records and Genomics (eMERGE) Network, established in 2007, is a consortium of academic and integrated health systems conducting discovery and implementation research in translational genomics. Here, we outline the history of the network, highlight major impacts and lessons learned, and present the tools and resources developed for large-scale genomic analyses and translation into a clinical setting. The network developed methods to extract phenotypes from the electronic medical record to perform genome-wide and phenome-wide association studies. Recruited cohorts were clinically sequenced off a custom panel for targeted sequencing of variants and monogenic disease risks and returned to participants to investigate the impact of return of genomic results. After generating a 105,000 participant-imputed genome-wide association study (GWAS) dataset for discovery, the network enrolled and sequenced 24,998 participants. Integration of these results into the medical record and the effects of results on participants provided key lessons to the field. These learned lessons inform genetic research in diverse populations and provide insights into the clinical impact of return and implementation of genomic medicine using the electronic medical record. The lessons produced by the eMERGE Network can be utilized by other consortia as translational genomic medicine research evolves.
Background/Objectives Melanocortin-4 receptor (MC4R) plays an essential role in food intake and energy homeostasis. More than 170 MC4R variants have been described over the past two decades, with conflicting reports regarding the prevalence and phenotypic effects of these variants in diverse cohorts. To determine the frequency of MC4R variants in large cohort of different ancestries, we evaluated the MC4R coding region for 20,537 eMERGE participants with sequencing data plus additional 77,454 independent individuals with genome-wide genotyping data at this locus. Subjects/Methods The sequencing data were obtained from the eMERGE phase III study, in which multisample variant call format calls have been generated, curated, and annotated. In addition to penetrance estimation using body mass index (BMI) as a binary outcome, GWAS and PheWAS were performed using median BMI in linear regression analyses. All results were adjusted for principal components, age, sex, and sites of genotyping. Results Targeted sequencing data of MC4R revealed 125 coding variants in 1839 eMERGE participants including 30 unreported coding variants that were predicted to be functionally damaging. Highly penetrant unreported variants included (L325I, E308K, D298N, S270F, F261L, T248A, D111V, and Y80F) in which seven participants had obesity class III defined as BMI ≥ 40 kg/m 2 . In GWAS analysis, in addition to known risk haplotype upstream of MC4R (best variant rs6567160 ( P = 5.36 × 10 −25 , Beta = 0.37), a novel rare haplotype was detected which was protective against obesity and encompassed the V103I variant with known gain-of-function properties ( P = 6.23 × 10 −08 , Beta = −0.62). PheWAS analyses extended this protective effect of V103I to type 2 diabetes, diabetic nephropathy, and chronic renal failure independent of BMI. Conclusions MC4R screening in a large eMERGE cohort confirmed many previous findings, extend the MC4R pleotropic effects, and discovered additional MC4R rare alleles that probably contribute to obesity.
BACKGROUND:Implementation of phenotype algorithms requires phenotype engineers to interpret human-readable algorithms and translate the description (text and flowcharts) into computable phenotypes - a process that can be labor intensive and error prone. To address the critical need for reducing the implementation efforts, it is important to develop portable algorithms. METHODS:We conducted a retrospective analysis of phenotype algorithms developed in the Electronic Medical Records and Genomics (eMERGE) network and identified common customization tasks required for implementation. A novel scoring system was developed to quantify portability from three aspects: Knowledge conversion, clause Interpretation, and Programming (KIP). Tasks were grouped into twenty representative categories. Experienced phenotype engineers were asked to estimate the average time spent on each category and evaluate time saving enabled by a common data model (CDM), specifically the Observational Medical Outcomes Partnership (OMOP) model, for each category. RESULTS:A total of 485 distinct clauses (phenotype criteria) were identified from 55 phenotype algorithms, corresponding to 1153 customization tasks. In addition to 25 non-phenotype-specific tasks, 46 tasks are related to interpretation, 613 tasks are related to knowledge conversion, and 469 tasks are related to programming. A score between 0 and 2 (0 for easy, 1 for moderate, and 2 for difficult portability) is assigned for each aspect, yielding a total KIP score range of 0 to 6. The average clause-wise KIP score to reflect portability is 1.37 ± 1.38. Specifically, the average knowledge (K) score is 0.64 ± 0.66, interpretation (I) score is 0.33 ± 0.55, and programming (P) score is 0.40 ± 0.64. 5% of the categories can be completed within one hour (median). 70% of the categories take from days to months to complete. The OMOP model can assist with vocabulary mapping tasks. CONCLUSION:This study presents firsthand knowledge of the substantial implementation efforts in phenotyping and introduces a novel metric (KIP) to measure portability of phenotype algorithms for quantifying such efforts across the eMERGE Network. Phenotype developers are encouraged to analyze and optimize the portability in regards to knowledge, interpretation and programming. CDMs can be used to improve the portability for some 'knowledge-oriented' tasks.
Background: Implementing clinical phenotypes across a network is labor intensive and potentially error prone. Use of a common data model may facilitate the process. Methods: Electronic Medical Records and Genomics (eMERGE) sites implemented the Observational Health Data Sciences and Informatics (OHDSI) Observational Medical Outcomes Partnership (OMOP) Common Data Model across their electronic health record (EHR)-linked DNA biobanks. Two previously implemented eMERGE phenotypes were converted to OMOP and implemented across the network. Results: It was feasible to implement the common data model across sites, with laboratory data producing the greatest challenge due to local encoding. Sites were then able to execute the OMOP phenotype in less than one day, as opposed to weeks of effort to manually implement an eMERGE phenotype in their bespoke research EHR databases. Of the sites that could compare the current OMOP phenotype implementation with the original eMERGE phenotype implementation, specific agreement ranged from 100% to 43%, with disagreements due to the original phenotype, the OMOP phenotype, changes in data, and issues in the databases. Using the OMOP query as a standard comparison revealed differences in the original implementations despite starting from the same definitions, code lists, flowcharts, and pseudocode. Conclusion: Using a common data model can dramatically speed phenotype implementation at the cost of having to populate that data model, though this will produce a net benefit as the number of phenotype implementations increases. Inconsistencies among the implementations of the original queries point to a potential benefit of using a common data model so that actual phenotype code and logic can be shared, mitigating human error in reinterpretation of a narrative phenotype definition.
Background The eMERGE III Network was tasked with harmonizing genetic testing protocols linking multiple sites and investigators.Methods DNA capture panels targeting 109 genes and 1551 variants were constructed by two clinical sequencing centers for analysis of 25,000 participant DNA samples collected at 11 sites where samples were linked to patients with electronic health records. Each step from sample collection, data generation, interpretation, reporting, delivery and storage, were developed and validated in CAP/CLIA settings and harmonized across sequencing centers.Results A compliant and secure network was built and enabled ongoing review and reconciliation of clinical interpretations while maintaining communication and data sharing between investigators. Mechanisms for sustained propagation and growth of the network were established. An interim data freeze representing 15,574 sequenced subjects, informed the assay performance for a range of variant types, the rate of return of results for different phenotypes and the frequency of secondary findings. Practical obstacles for implementation and scaling of clinical and research findings were identified and addressed. The eMERGE protocols and tools established are now available for widespread dissemination.Conclusions This study established processes for different sequencing sites to harmonize the technical and interpretive aspects of sequencing tests, a critical achievement towards global standardization of genomic testing. The network established experience in the return of results and the rate of secondary findings across diverse biobank populations. Furthermore, the eMERGE network has accomplished integration of structured genomic results into multiple electronic health record systems, setting the stage for clinical decision support to enable genomic medicine.
Non-alcoholic fatty liver disease (NAFLD) is a common chronic liver illness with a genetically heterogeneous background that can be accompanied by considerable morbidity and attendant health care costs. The pathogenesis and progression of NAFLD is complex with many unanswered questions. We conducted genome-wide association studies (GWASs) using both adult and pediatric participants from the Electronic Medical Records and Genomics (eMERGE) Network to identify novel genetic contributors to this condition. First, a natural language processing (NLP) algorithm was developed, tested, and deployed at each site to identify 1106 NAFLD cases and 8571 controls and histological data from liver tissue in 235 available participants. These include 1242 pediatric participants (396 cases, 846 controls). The algorithm included billing codes, text queries, laboratory values, and medication records. Next, GWASs were performed on NAFLD cases and controls and case-only analyses using histologic scores and liver function tests adjusting for age, sex, site, ancestry, PC, and body mass index (BMI). Consistent with previous results, a robust association was detected for the PNPLA3 gene cluster in participants with European ancestry. At the PNPLA3-SAMM50 region, three SNPs, rs738409, rs738408, and rs3747207, showed strongest association (best SNP rs738409 p = 1.70 × 10− 20). This effect was consistent in both pediatric (p = 9.92 × 10− 6) and adult (p = 9.73 × 10− 15) cohorts. Additionally, this variant was also associated with disease severity and NAFLD Activity Score (NAS) (p = 3.94 × 10− 8, beta = 0.85). PheWAS analysis link this locus to a spectrum of liver diseases beyond NAFLD with a novel negative correlation with gout (p = 1.09 × 10− 4). We also identified novel loci for NAFLD disease severity, including one novel locus for NAS score near IL17RA (rs5748926, p = 3.80 × 10− 8), and another near ZFP90-CDH1 for fibrosis (rs698718, p = 2.74 × 10− 11). Post-GWAS and gene-based analyses identified more than 300 genes that were used for functional and pathway enrichment analyses. In summary, this study demonstrates clear confirmation of a previously described NAFLD risk locus and several novel associations. Further collaborative studies including an ethnically diverse population with well-characterized liver histologic features of NAFLD are needed to further validate the novel findings.
Electronic health record (EHR) algorithms for defining patient cohorts are commonly shared as free-text descriptions that require human intervention both to interpret and implement. We developed the Phenotype Execution and Modeling Architecture (PhEMA, http://projectphema.org) to author and execute standardized computable phenotype algorithms. With PhEMA, we converted an algorithm for benign prostatic hyperplasia, developed for the electronic Medical Records and Genomics network (eMERGE), into a standards-based computable format. Eight sites (7 within eMERGE) received the computable algorithm, and 6 successfully executed it against local data warehouses and/or i2b2 instances. Blinded random chart review of cases selected by the computable algorithm shows PPV >= 90%, and 3 out of 5 sites had >90% overlap of selected cases when comparing the computable algorithm to their original eMERGE implementation. This case study demonstrates potential use of PhEMA computable representations to automate phenotyping across different EHR systems, but also highlights some ongoing challenges.
Electronic health record (EHR) algorithms for defining patient cohorts are commonly shared as free-text descriptions that require human intervention to both interpret and implement. We developed the Phenotype Execution and Modeling Architecture (PhEMA, http://projectphema.org) to author and execute standardized computable phenotype algorithms. With PhEMA, we converted an algorithm for benign prostatic hyperplasia, developed for the electronic Medical Records and Genomics network (eMERGE), into a standards-based computable format. Eight sites (7 within eMERGE) received the computable algorithm and 6 successfully executed it against local data warehouses and/or i2b2 instances. Blinded random chart review of cases selected by the computable algorithm shows PPV ≥90%, and 3 out of 5 sites had >90% overlap of selected cases when comparing the computable algorithm to their original eMERGE implementation. This case study demonstrates potential use of PhEMA computable representations to automate phenotyping across different EHR systems, but also highlights some ongoing challenges. INTRODUCTION Electronic health records (EHRs) are designed primarily for clinical care and operations, which introduces limitations to the data when used for secondary purposes. These include data availability, accuracy, and information only available in unstructured narrative text.[1-3] To navigate these issues, researchers often develop “phenotype algorithms”– a set of criteria that specifies the analysis of data for identifying and studying patient populations with a given biomedical condition or indication.[4, 5] When replication of results or additional statistical power is needed, a phenotype algorithm may be shared with other institutions for execution on their EHR data, using various methods and tools[6]. A common method used to share algorithms is as free-text descriptions of the algorithmic logic, possibly augmented with flowcharts and lists of codes from medical vocabularies. Implementing this description in a computable form (such as a database query) requires humans to interpret and make decisions for deployment, which can be error-prone and time-consuming.[7] An approach pursued via common data models (CDMs) is to organize EHR data according to a common standard. A CDM, such as those used by Observational Health Data Sciences and Informatics (OHDSI),[8] the National Patient-Centered Clinical Research Network (PCORnet),[9] and Informatics for Integrating Biology & the Bedside’s (i2b2’s)[10] Shared Health Research Informatics NEtwork (SHRINE),[11] can enable cross-site queries since the semantic and syntactic data organization are shared, but require transforming data to the CDM. Within the electronic Medical Records and Genomics (eMERGE) network,[12] the Phenotype KnowledgeBase (PheKB)[13] was created as an online environment to support collaboratively developing, validating, and sharing electronic phenotype algorithms across institutions, with a forum for clarifying implementation details as is often necessary. Truly portable phenotypes requiring little or no human interpretation and data transformation, would facilitate increasing the number of implementing institutions, but are not widely available. To address this gap, we have developed the Phenotype Execution and Modeling Architecture (PhEMA; http://projectphema.org),[14] an open-source infrastructure for standards-based authoring, sharing and execution of phenotyping algorithms. PhEMA uses the National Quality Forum’s (NQF’s) Quality Data Model (QDM) and HL7’s Health Quality Measure Format (HQMF)[15] to unambiguously model phenotype definitions, and enable the execution of these phenotype algorithms against different data representations, including local data warehouses (LDWs), and i2b2 data repositories. We describe a realworld use case for PhEMA: a computable algorithm for identifying patients with benign prostatic hyperplasia (BPH) from EHRs. METHODS BPH Phenotype Algorithm BPH was chosen as a test phenotype for PhEMA in part because the case identification algorithm contains four of the most commonly used EHR data elements (demographics, diagnoses, procedures, and medications), all of the Boolean logical operators (And, Or, Not), a temporal relationship (age), and one aggregation function (count). Since the majority of sites in this study recently executed the algorithm for eMERGE, it also provided a baseline for comparisons against the results using PhEMA. A single institution (Vanderbilt University) defined the original BPH phenotype algorithm for the eMERGE study (referred to as the “original algorithm”), which was modified by the authors for implementation within PhEMA (referred to as the “PhEMA algorithm”). The PhEMA algorithm used a combination of standard medical terminology codes for medications (RxNorm and NDC), diagnostic codes (ICD-9), and procedure codes (CPT). Although the original algorithm used natural language processing (NLP) on patients’ problem lists, NLP was not included in the PhEMA algorithm due to a lack of standard representations of NLP.[16] The study population was defined as males age 40 and older who had no evidence in their EHR of prostate, penile, urethral, or bladder cancer. Within the PhEMA algorithm, the exclusion used just ICD-9 billing codes, although the original algorithm also used ICD-O-3 codes from tumor registries and mentions of these cancers in the problem list. The PhEMA algorithm selected cases only, who were defined as those in the study population with any BPH-related ICD-9 codes on 2 separate days, plus at least one BPH medication or BPH-related surgery code (Figure 1). Translation of Algorithm into a Standardized, Executable Format The modified PhEMA algorithm to detect BPH cases was represented using the PhEMA Authoring Tool (PhAT) (available at: https://github.com/PheMA/bph-use-case), by one of the authors (LVR). Since PhEMA relies on the presence of value sets (groups of medical terminology codes) to represent different medical concepts (e.g., medications, diagnoses), we first looked at the Value Set Authority Center (VSAC)[17] for existing value sets. Outside of the value set of “Male”, no existing value sets met the need of our algorithm; therefore, the authors defined and published the necessary value sets within the VSAC for general use (also available at: https://github.com/PheMA/bph-use-case). Given the use of QDM 4.1 [16, 18] within PhEMA, which accounts for the status of a data element (e.g., an active diagnosis, as opposed to the more abstract “diagnosis”), we included in our definition the relevant statuses for diagnoses, medications and procedures. This approach allows a more precise algorithm definition, as it removes the ambiguity of which statuses are appropriate. The phenotype definition was exported from the PhAT into two executable KNIME workflows (KNIME AG, Zurich Switzerland, Figure 2): one workflow executed query definitions using i2b2 messaging (Figure 2A), and the other executed against an LDW (Figure 2B). KNIME was chosen as the execution engine as it is freely available, runs on multiple operating systems, and, most importantly, can connect directly to external systems that expose a web API. Furthermore, KNIME uses a modular graphical workflow interface that allows encapsulation of algorithm logic (the same across sites) from configuration details (variable between sites), allowing each site to configure the connection to their data without having to edit the algorithm’s logic.[18, 19] Upon review of the exported KNIME workflows, we identified opportunities to optimize their execution before distributing to sites. First, within sites’ i2b2 instances, diagnosis status (active, resolved or inactive) was not explicitly described; thus, the 3 queries (one for each status) were collapsed into a single query. Second, by default the workflow returned multiple types of results within i2b2 for each query – a count, a list of patients, and a list of events. We edited the workflow to only return patient sets, to further reduce execution time. Finally, we removed the temporal relationship between BPH diagnoses and instead required patients have ≥2 BPH diagnoses, as some sites did not have visits associated with diagnoses in their i2b2 data. Once given to the implementation site, each workflow required up to two local customization steps. First, the site must specify the connection details for their repository (i2b2 or LDW). Second, the site performed data customization if needed. For i2b2, this involved updating the ontology mapping so the correct concepts could be found; and for the LDW, editing template data queries to return the required data elements. Two members of the PhEMA team (JAP, LVR) guided this customization via remote
Introduction: In most countries <20% of prevalent cases of familial hypercholesterolemia (FH) are diagnosed. There is an unmet need to develop EHR-based strategies to increase detection and control of FH. Objective: We assessed performance and portability of an electronic phenotyping algorithm for FH in the electronic MEdical Records and GEnomics (eMERGE) Network. Methods: Using the Phenotype KnowledgeBase (PheKB) platform, we deployed an electronic algorithm to ascertain FH at 5 adult sites. Family history of premature ASCVD/FH, findings of tendon xanthomas and corneal arcus were assessed using natural language processing (NLP) systems. LDL cholesterol levels, secondary hypercholesterolemia, and medications were ascertained from structured data sources. Personal history of premature ASCVD was defined using both code-based and NLP systems. Implementation was assessed by a 15-question survey at each site, including a 5-point Likert scale (very difficult/undesirable to excellent/very desirable) for ease of ...