BACKGROUND:Stroke is one of the leading causes of death and disability worldwide. The National Institutes of Health Stroke Scale (NIHSS) scores in electronic health records (EHRs), which quantitatively describe patients' neurological deficits in evidence-based treatment, are crucial in stroke-related clinical investigations. However, the free-text format and lack of standardization inhibit their effective use. Automatically extracting the scale scores from the clinical free text so that its potential value in real-world studies is realized has become an important goal. OBJECTIVE:This study aims to develop an automated method to extract scale scores from the free text of EHRs. METHODS:We propose a two-step pipeline method to identify NIHSS items and numerical scores and validate its feasibility using a freely accessible critical care database: MIMIC-III (Medical Information Mart for Intensive Care III). First, we utilize MIMIC-III to create an annotated corpus. Then, we investigate possible machine learning methods for two subtasks, NIHSS item and score recognition and item-score relation extraction. In the evaluation, we conduct both task-specific and end-to-end evaluations and compare our method with the rule-based method using precision, recall and F1 scores as evaluation metrics. RESULTS:We use all available discharge summaries of stroke cases in MIMIC-III. The annotated NIHSS corpus contains 312 cases, 2929 scale items, 2774 scores and 2733 relations. The results show that the best F1-score of our method was 0.9006, which was attained by combining BERT-BiLSTM-CRF and Random Forest, and it outperformed the rule-based method (F1-score = 0.8098). In the end-to-end task, our method could successfully recognize the item "1b level of consciousness questions", the score "1" and their relation "('1b level of consciousness questions', '1', 'has value')" from the sentence "1b level of consciousness questions: said name = 1", while the rule-based method could not. CONCLUSIONS:The two-step pipeline method we propose is an effective approach to identify NIHSS items, scores and their relations. With its help, clinical investigators can easily retrieve and access structured scale data, thereby supporting stroke-related real-world studies.
Electronic health records provide more insight and possibilities for retrospective studies with the help of data science techniques. However, the free-text in medical records limit the availability and reusability of data. Credentialed data that is contained in free-text also puts more restriction on data sharing and secondary use. We aim to propose a method for guiding EHRs free-text data annotation and sharing from the perspective of data safety and procedure standardization.
[目的]面向复杂疾病临床试验招募的需求,提出一种基于BERT-TextCNN的临床试验疾病亚型识别方法,辅助识别复杂疾病特定亚型的受试人群.[方法]将临床试验疾病亚型识别问题转化为单标签分类问题,应用基于BERT-TextCNN的单标签分类模型进行分类,以卒中为例在临床试验数据集(ClinicalTrials.gov)上开展实验验证.[结果]基于LP法的BERT-TextCNN模型性能最佳,加权宏平均F1值为0.905 3,可以有效判定一项卒中临床试验可纳入卒中亚型受试者情况.[局限]缺乏在其他单病种上的可行性研究,以及在外部数据集上的有效性验证.[结论]本文方法可以有效解决从纳入标准中准确识别复杂疾病亚型的问题.
临床信息模型的可复用性是实现健康医疗信息互联互通和共享的重要基础,检索并识别出临床信息模型中可复用的对象是提高复用性的一种有效途径.以HL7.org发布的HL7 V3 2017标准版本为研究对象,应用扩展的4层贝叶斯网络表征该临床信息模型,在简单贝叶斯网络的基础上增加分层消息描述(HMD)之间语义相似性的扩展层,通过网络的逐层概率推演识别出可复用的临床信息模型.实验设计“就诊预约”、“实验室结果”和“病人实体”等3个检索任务,并应用平均精度均值MAP、平均精度AP和截止点准确率等3个指标以评价检索方法的性能.最终构建含有3 428个节点和22 646条边的4层贝叶斯网络,自上而下依次为数据元素层、HMD层、重复的HMD层和消息类型层,各层节点数量分别为2 177、422、422和407.检索结果显示,MAP值达到了0.382,在第3、第5和第10截止点的平均准确率分别达到77.8%、60.0%和46.7%;该方法不仅可以检索出通用型和领域型两类可复用模型,还能够发现临床语义切实相关的对象(例如检索“就诊预约”时返回的“预约更新通知”对象).所提出的方法将有助于提高HL7 V3信息模型的复用性,并促进临床信息标准的国际化建设,同时对其他临床信息模型检索方法的优化和改进存在一定的借鉴意义.
[目的]面向真实世界数据驱动的临床研究需求,提出一种基于语义对齐的临床量表信息提取方法,辅助识别潜在受试人群.[方法]选取卒中量表NIHSS,分析量表信息在临床试验和真实世界电子病历中的特征,构建基于语义对齐的量表信息提取方法,应用临床试验数据集(ClinicalTrials.gov)和开放电子病历数据集MIMIC-III开展实验验证.[结果]从患者出院小结中抽取NIHSS总评分、检查项评分的F1值分别为0.953 5和0.926 7;围绕两项匹配NIHSS纳排标准的测试任务,可以有效地识别出潜在受试人群.[局限]缺乏在其他量表上的可行性研究,以及在真实临床试验环境中的有效性和可靠性验证.[结论]本方法可以有效地解决临床量表信息在临床研究与电子病历数据的语义一致性问题.
目的:更好地表示临床指南知识,以支持临床决策支持系统的构建.方法:以缺血性卒中为例,在系统梳理临床指南知识表示方法的基础上,提出使用解释节点、数据获取节点、动作节点、复合节点,以及有向边和无向边对指南诊疗流程进行表示,同时考虑实际的电子病历数据格式,设计自然语言处理模块,提高指南与电子病历系统的数据交互能力.结果:在实证研究中,使用8个节点、7条边,以及节点和边中的4个逻辑推理式表示了缺血性卒中患者临床类型判断流程,同时使用9个节点、8条边,以及节点和边中的6个逻辑推理式表示了缺血性卒中急性期的血糖管理,详细说明了临床指南知识表示建模过程.结论:本研究提出的临床指南知识表示方法可以使用节点、边、数据、逻辑推理算法等表示临床指南知识及推理逻辑,以支撑临床决策支持系统的构建.
To facilitate experts to find relevant clinical information models, we took a case study of openEHR. We proposed to use a graphical model to represent EHR archetype sets, aiming to optimize clincal information retrieval performance. In this study, we applied our graphic model to 523 OpenEHR archetypes and represented them as a graph with 5,008 nodes and 6,908 edges, which consists of 3,982 term nodes, 504 concept nodes, and 523 archetype nodes. On basis of the graphical model, it improved the performance for retrieving the clinical queries.
Background Clinical information models (CIMs) enabling semantic interoperability are crucial for electronic health record (EHR) data use and reuse. Dual model methodology, which distinguishes the CIMs from the technical domain, could help enable the interoperability of EHRs at the knowledge level. How to help clinicians and domain experts discover CIMs from an open repository online to represent EHR data in a standard manner becomes important. Objective This study aimed to develop a retrieval method to identify CIMs online to represent EHR data. Methods We proposed a graphical retrieval method and validated its feasibility using an online CIM repository: openEHR Clinical Knowledge Manager (CKM). First, we represented CIMs (archetypes) using an extended Bayesian network. Then, an inference process was run in the network to discover relevant archetypes. In the evaluation, we defined three retrieval tasks (medication, laboratory test, and diagnosis) and compared our method with three typical retrieval methods (BM25F, simple Bayesian network, and CKM), using mean average precision (MAP), average precision (AP), and precision at 10 (P@10) as evaluation metrics. Results We downloaded all available archetypes from the CKM. Then, the graphical model was applied to represent the archetypes as a four-level clinical resources network. The network consisted of 5513 nodes, including 3982 data element nodes, 504 concept nodes, 504 duplicated concept nodes, and 523 archetype nodes, as well as 9867 edges. The results showed that our method achieved the best MAP (MAP=0.32), and the AP was almost equal across different retrieval tasks (AP=0.35, 0.31, and 0.30, respectively). In the diagnosis retrieval task, our method could successfully identify the models covering “diagnostic reports,” “problem list,” “patients background,” “clinical decision,” etc, as well as models that other retrieval methods could not find, such as “problems and diagnoses.” Conclusions The graphical retrieval method we propose is an effective approach to meet the uncertainty of finding CIMs. Our method can help clinicians and domain experts identify CIMs to represent EHR data in a standard manner, enabling EHR data to be exchangeable and interoperable.