Computational linguistics is an interdisciplinary field concerned with the computational modelling of natural language, as well as the study of appropriate computational approaches to linguistic questions. In general, computational linguistics draws upon linguistics, computer science, artificial intelligence, math, logic, philosophy, cognitive science, cognitive psychology, psycholinguistics, anthropology and neuroscience, among others.
BACKGROUND:Data, particularly 'big' data are increasingly being used for research in health. Using data from electronic medical records optimally requires coded data, but not all systems produce coded data.OBJECTIVE:To design a suitable, accurate method for converting large volumes of narrative diagnoses from Australian general practice records to codify them into SNOMED-CT-AU. Such codification will make them clinically useful for aggregation for population health and research purposes.METHOD:The developed method consisted of using natural language processing to automatically code the texts, followed by a manual process to correct codes and subsequent natural language processing re-computation. These steps were repeated for four iterations until 95% of the records were coded. The coded data were then aggregated into classes considered to be useful for population health analytics.RESULTS:Coding the data effectively covered 95% of the corpus. Problems with the use of SNOMED CT-AU were identified and protocols for creating consistent coding were created. These protocols can be used to guide further development of SNOMED CT-AU (SCT). The coded values will be immensely useful for the development of population health analytics for Australia, and the lessons learnt applicable elsewhere.
Artificial intelligence solutions for clinical tasks have been found to be prematurely released to clinical teams and thereby created increased risks and workload for clinicians. This letter discusses the issues that determine good AI practices.
Objective This project examined and produced a general practice (GP) based decision support tool (DST), namely POLAR Diversion, to predict a patient's risk of emergency department (ED) presentation. The tool was built using both GP/family practice and ED data, but is designed to operate on GP data alone. Methods GP data from 50 practices during a defined time frame were linked with three local EDs. Linked data and data mapping were used to develop a machine learning DST to determine a range of variables that, in combination, led to predictive patient ED presentation risk scores. Thirteen percent of the GP data was kept as a control group and used to validate the tool. Results The algorithm performed best in predicting the risk of attending ED within the 30-day time category, and also in the no ED attendance tests, suggesting few false positives. At 0 to 30 days the positive predictive value (PPV) was 74%, with a sensitivity/recall of 68%. Non-ED attendance had a PPV of 82% and sensitivity/recall of 96%. Conclusion Findings indicate that the POLAR Diversion algorithm performed better than previously developed tools, particularly in the 0 to 30 day time category. Its utility increases because of it being based on the data within the GP system alone, with the ability to create real-time "in consultation" warnings. The tool will be deployed across GPs in Australia, allowing us to assess the clinical utility, and data quality needs in further iterations.
The paper applies an artificial intelligence centered method to classify 12 clinical safety incident (CSI) classes. The paper aims to establish a taxonomy that classifies the CSI reports into their correct classes automatically and with high accuracy. The study investigates feasibility of applying the C4.5 decision tree (DT) classifier and the random forest (RF) classifier for this purpose. The classifiers were trained using randomly selected 3600 CSIs from an Incident Information Management System (IIMS) used by seven hospitals. The taxonomies investigated were the Generic Reference Model (GRM) and the World Health Organization (WHO) patient safety classification. The classifiers trained 13 GRM CSI classes and 9 WHO CSI classes using a bag-of-words approach. The overall taxonomies performance on the RF classifier was better than on the DT classifier. The performance achieved by the classifier applying the WHO taxonomy was better than the GRM taxonomy. Four of the five poorly performing classes in the GRM taxonomy significantly improved their performance on changing the taxonomy. To improve the WHO taxonomy performance the improved WHO (WHO-I) taxonomy was built by adding a new class that did not exist in WHO but existed in GRM. The performance of the RF classifier applied to the WHO-I taxonomy further improved.
Much of the important patient information can only be found in patient narratives or in free text fields of structural schema of the Clinical Information System (CIS). So, the integration of free text search facilities will improve question answering on CISs. This paper describes a method for integrating free text search facility to the proposed Data Analytics Language (CliniDAL) to improve its capabilities at answering more common clinical questions. The proposed language constructs in CliniDAL’s grammar enables its parser to recognize the part of the Restricted Natural Language Query (RNLQ) of the CliniDAL interface, which needs a free text resolution mechanism. Then the Natural Language Processing (NLP) approach of the CliniSearch tool finds the correct matches with the query. The search result is integrated into the translated CliniDAL query which can be executed to return a more comprehensive answer to the initial text query. 160 queries are tested in the current work to investigate the improvements on answering more common questions from a CIS, which result in a simple taxonomy of four query categories of: unanswerable queries, queries that require more evidence to be answered, queries requiring user interpretation and queries with suitable answers. Compatibility of query results between the structural schema and patient progress notes is examined which showed the usability of the approach in answering queries, confirming the results from different sources and finding any inconsistency in the stored data in the CIS. The proposed solution provides a simple mechanism for extracting knowledge from CISs.
PURPOSE:This paper reports on a generic framework to provide clinicians with the ability to conduct complex analyses on elaborate research topics using cascaded queries to resolve internal time-event dependencies in the research questions, as an extension to the proposed Clinical Data Analytics Language (CliniDAL).METHODS:A cascaded query model is proposed to resolve internal time-event dependencies in the queries which can have up to five levels of criteria starting with a query to define subjects to be admitted into a study, followed by a query to define the time span of the experiment. Three more cascaded queries can be required to define control groups, control variables and output variables which all together simulate a real scientific experiment. According to the complexity of the research questions, the cascaded query model has the flexibility of merging some lower level queries for simple research questions or adding a nested query to each level to compose more complex queries. Three different scenarios (one of them contains two studies) are described and used for evaluation of the proposed solution.RESULTS:CliniDAL's complex analyses solution enables answering complex queries with time-event dependencies at most in a few hours which manually would take many days.CONCLUSION:An evaluation of results of the research studies based on the comparison between CliniDAL and SQL solutions reveals high usability and efficiency of CliniDAL's solution.
This paper describes a methodology that engages the clinical community into the design process of creating Clinical Information Systems (CISs) under a Clinical Team-Led Design (CTLD) approach in the context of using Immediately Adaptable (IA) system development technology. The methodology is contrasted against the Enterprise Electronic Medical Record (EEMR) model for usability, efficiency, and adaptability. The methodology was tested in a Breast Cancer setting where the CIS went through 4 rapid agile stages. Time and motion statistics, training times, system changes, and user feedback data was collected for assessment. The results showed that the Breast CIS increased time efficiency by 30% in the first 3 months of implementation. Users reported high usability and trainability of the system. Over 95% of system design change requests were satisfied with an average turn-around time of 3 days. The results show that systems designed under a CTLD approach, accompanied by Immediately Adaptable system architecture, provide greater efficiency for staff in clinical settings while enabling the workflow processes to be adapted dynamically as part of continuous process improvement.
BACKGROUND:Every day, patients are admitted to the hospital with conditions that could have been effectively managed in the primary care sector. These admissions are expensive and in many cases are possible to avoid if early intervention occurs. General practitioners are in the best position to identify those at risk of imminent hospital presentation and admission; however, it is not always possible for all the factors to be considered. A lack of shared information contributes significantly to the challenge of understanding a patient's full medical history. Some health care systems around the world use algorithms to analyze patient data in order to predict events such as emergency presentation; however, those responsible for the design and use of such systems readily admit that the algorithms can only be used to assess the populations used to design the algorithm in the first place. The United Kingdom health care system has contributed data toward algorithm development, which is possible through the unified health care system in place there. The lack of unified patient records in Australia has made building an algorithm for local use a significant challenge.OBJECTIVE:Our objective is to use linked patient records to track patient flow through primary and secondary health care in order to develop a tool that can be applied in real time at the general practice level. This algorithm will allow the generation of reports for general practitioners that indicate the relative risk of patients presenting to an emergency department.METHODS:A previously designed tool was used to deidentify the general practice and hospital records of approximately 100,000 patients. Records were pooled for patients who had attended emergency departments within the Eastern Health Network of hospitals and general practices within the Eastern Health Network catchment. The next phase will involve development of a model using a predictive analytic machine learning algorithm. The model will be developed iteratively, testing the combination of variables that will provide the best predictive model.RESULTS:Records of approximately 97,000 patients who have attended both a general practice and an emergency department have been identified within the database. These records are currently being used to develop the predictive model.CONCLUSIONS:Records from general practice and emergency department visits have been identified and pooled for development of the algorithm. The next phase in the project will see validation and live testing of the algorithm in a practice setting. The algorithm will underpin a clinical decision support tool for general practitioners which will be tested for face validity in this initial study into its efficacy.
Text mining in clinical domain is usually more difficult than general domains (e.g. newswire reports and scientific literature) because of the high level of noise in both the corpus and training data for machine learning (ML). A large number of unknown word, non-word and poor grammatical sentences made up the noise in the clinical corpus. Unknown words are usually complex medical vocabularies, misspellings, acronyms and abbreviations where unknown non-words are generally the clinical patterns including scores and measures. This noise produces obstacles in the initial lexical processing step as well as subsequent semantic analysis. Furthermore, the labelled data used to build ML models is very costly to obtain because it requires intensive clinical knowledge from the annotators. And even created by experts, the training examples usually contain errors and inconsistencies due to the variations in human annotators' attentiveness. Clinical domain also suffers from the nature of the imbalanced data distribution problem. These kinds of noise are very popular and potentially affect the overall information extraction performance but they were not carefully investigated in most presented health informatics systems. This paper introduces a general clinical data mining architecture which is potential of addressing all of these challenges using: automatic proof-reading process, trainable finite state pattern recogniser, iterative model development and active learning. The reportability classifier based on this architecture achieved 98.25% sensitivity and 96.14% specificity on an Australian cancer registry's held-out test set and up to 92% of training data provided for supervised ML was saved by active learning.
We consider the task of automatic classification of clinical incident reports using machine learning methods. Our data consists of 5448 clinical incident reports collected from the Incident Information Management System used by 7 hospitals in the state of New South Wales in Australia. We evaluate the performance of four classification algorithms: decision tree, naïve Bayes, multinomial naïve Bayes and support vector machine. We initially consider 13 classes (incident types) that were then reduced to 12, and show that it is possible to build accurate classifiers. The most accurate classifier was the multinomial naïve Bayes achieving accuracy of 80.44% and AUC of 0.91. We also investigate the effect of class labelling by an ordinary clinician and an expert, and show that when the data is labelled by an expert the classification performance of all classifiers improves. We found that again the best classifier was multinomial naïve Bayes achieving accuracy of 81.32% and AUC of 0.97. Our results show that some classes in the Incident Information Management System such as Primary Care are not distinct and their removal can improve performance; some other classes such as Aggression Victim are easier to classify than others such as Behavior and Human Performance. In summary, we show that the classification performance can be improved by expert class labelling of the training data, removing classes that are not well defined and selecting appropriate machine learning classifiers.
STUDY OBJECTIVE:This investigation was initiated after the introduction of a new information system into the Nepean Hospital Emergency Department. A retrospective study determined that the problems introduced by the new system led to reduced efficiency of the clinical staff, demonstrated by deterioration in the emergency department's (ED's) performance. This article is an investigation of methods to improve the design and implementation of clinical information systems for an ED by using a process of clinical team-led design and a technology built on a radically new philosophy denoted as emergent clinical information systems.METHODS:The specific objectives were to construct a system, the Nepean Emergency Department Information Management System (NEDIMS), using a combination of new design methods; determine whether it provided any reduction in time and click burden on the user in comparison to an enterprise proprietary system, Cerner FirstNet; and design and evaluate a model of the effect that any reduction had on patient throughput in the department.RESULTS:The methodology for conducting a direct comparison between the 2 systems used the 6 activity centers in the ED of clerking, triage, nursing assessments, fast track, acute care, and nurse unit manager. A quantitative study involved the 2 systems being measured for their efficiency on 17 tasks taken from the activity centers. A total of 332 task instances were measured for duration and number of mouse clicks in live usage on Cerner FirstNet and in reproduction of the same Cerner FirstNet work on NEDIMS as an off-line system. The results showed that NEDIMS is at least 41% more efficient than Cerner FirstNet (95% confidence interval 21.6% to 59.8%). In some cases, the NEDIMS tasks were remodeled to demonstrate the value of feedback to create improvements and the speed and economy of design revision in the emergent clinical information systems approach. The cost of the effort in remodeling the designs showed that the time spent on remodeling is recovered within a few days in time savings to clinicians. An analysis of the differences between Cerner FirstNet and NEDIMS for sequences of patient journeys showed an average difference of 127 seconds and 15.2 clicks. A simulation model of workflows for typical patient journeys for a normal daily attendance of 165 patients showed that NEDIMS saved 23.9 hours of staff time per day compared with Cerner FirstNet.CONCLUSION:The results of this investigation show that information systems that are designed by a clinical team using a technology that enables real-time adaptation provides much greater efficiency for the ED. Staff consider that a point-and-click user interface constantly interrupts their train of thought in a way that does not happen when writing on paper. This is partially overcome by the reduction of cognitive load that arises from minimizing the number of clicks to complete a task in the context of global versus local workflow optimization.
OBJECTIVE:To detect negations of medical entities in free-text pathology reports with different approaches, and evaluate their performances. METHODS AND MATERIAL:Three different approaches were applied for negation detection: the lexicon-based approach was a rule-based method, relying on trigger terms and termination clues; the syntax-based approach was also a rule-based method, where the rules and negation patterns were designed using the dependency output from the Stanford parser; the machine-learning-based approach used a support vector machine as a classifier to build models with a number of features. A total of 284 English pathology reports of lymphoma were used for the study. RESULTS:The machine-learning-based approach had the best overall performance on the test set with micro-averaged F-score of 82.56%, while the syntax-based approach performed worst with 78.62% F-score. The lexicon-based approach attained an overall average precision of 89.74% and recall of 76.09%, which were significantly better than the results achieved by Negation Tagger with a similar approach. DISCUSSION:The lexicon-based approach benefitted from being customized to the corpus more than the other two methods. The errors in negation detection with the syntax-based approach producing poorest performance were mainly due to the poor parsing results, and the errors with the other methods were probably because of the abnormal grammatical structures. CONCLUSIONS:A machine-learning-based approach has potential advantages for negation detection, and may be preferable for the task. To improve the overall performance, one of the possible solutions is to apply different approaches to each section in the reports.
OBJECTIVEThis paper presents an automated system for classifying the results of imaging examinations (CT, MRI, positron emission tomography) into reportable and non-reportable cancer cases. This system is part of an industrial-strength processing pipeline built to extract content from radiology reports for use in the Victorian Cancer Registry.MATERIALS AND METHODSIn addition to traditional supervised learning methods such as conditional random fields and support vector machines, active learning (AL) approaches were investigated to optimize training production and further improve classification performance. The project involved two pilot sites in Victoria, Australia (Lake Imaging (Ballarat) and Peter MacCallum Cancer Centre (Melbourne)) and, in collaboration with the NSW Central Registry, one pilot site at Westmead Hospital (Sydney).RESULTSThe reportability classifier performance achieved 98.25% sensitivity and 96.14% specificity on the cancer registry's held-out test set. Up to 92% of training data needed for supervised machine learning can be saved by AL.DISCUSSIONAL is a promising method for optimizing the supervised training production used in classification of radiology reports. When an AL strategy is applied during the data selection process, the cost of manual classification can be reduced significantly.CONCLUSIONSThe most important practical application of the reportability classifier is that it can dramatically reduce human effort in identifying relevant reports from the large imaging pool for further investigation of cancer. The classifier is built on a large real-world dataset and can achieve high performance in filtering relevant reports to support cancer registries.
To cite: Nguyen DHM, Patrick JD. J Am Med Inform Assoc 2014;21:893–901. ABSTRACT Objective This paper presents an automated system for classifying the results of imaging examinations (CT, MRI, positron emission tomography) into reportable and non-reportable cancer cases. This system is part of an industrial-strength processing pipeline built to extract content from radiology reports for use in the Victorian Cancer Registry. Materials and methods In addition to traditional supervised learning methods such as conditional random fields and support vector machines, active learning (AL) approaches were investigated to optimize training production and further improve classification performance. The project involved two pilot sites in Victoria, Australia (Lake Imaging (Ballarat) and Peter MacCallum Cancer Centre (Melbourne)) and, in collaboration with the NSW Central Registry, one pilot site at Westmead Hospital (Sydney). Results The reportability classifier performance achieved 98.25% sensitivity and 96.14% specificity on the cancer registry’s held-out test set. Up to 92% of training data needed for supervised machine learning can be saved by AL. Discussion AL is a promising method for optimizing the supervised training production used in classification of radiology reports. When an AL strategy is applied during the data selection process, the cost of manual classification can be reduced significantly. Conclusions The most important practical application of the reportability classifier is that it can dramatically reduce human effort in identifying relevant reports from the large imaging pool for further investigation of cancer. The classifier is built on a large real-world dataset and can achieve high performance in filtering relevant reports to support cancer registries.
Purpose: To elevate the level of care to the community it is essential to provide usable tools for healthcare professionals to extract knowledge from clinical data. In this paper a generic translation algorithm is proposed to translate a restricted natural language query (RNLQ) to a standard query language like SQL (Structured Query Language).Methods: A special purpose clinical data analytics language (CliniDAL) has been introduced which provides scheme of six classes of clinical questioning templates. A translation algorithm is proposed to translate the RNLQ of users to SQL queries based on a similarity-based Top-k algorithm which is used in the mapping process of CliniDAL. Also a two layer rule-based method is used to interpret the temporal expressions of the query, based on the proposed temporal model. The mapping and translation algorithms are generic and thus able to work with clinical databases in three data design models, including Entity-Relationship (ER), Entity-Attribute-Value (EAV) and XML, however it is only implemented for ER and EAV design models in the current work.Results: It is easy to compose a RNLQ via CliniDAL's interface in which query terms are automatically mapped to the underlying data models of a Clinical Information System (CIS) with an accuracy of more than 84% and the temporal expressions of the query comprising absolute times, relative times or relative events can be automatically mapped to time entities of the underlying CIS and to normalized temporal comparative values.Conclusion: The proposed solution of CliniDAL using the generic mapping and translation algorithms which is enhanced by a temporal analyzer component provides a simple mechanism for composing RNLQ for extracting knowledge from CISs with different data design models for analytics purposes. (C) 2014 Elsevier Inc. All rights reserved.
The aim of this project is to use the methods of natural language processing to extract pertinent information from free-text pathology reports to automatically populate structured reports. A processing pipeline has been developed cosseting of a combination of a supervised machine learning based approach using Conditional Random Fields for medical entity recognition and some rule-based methods. In total 477 narrative pathology reports of primary cutaneous melanomas were collected for evaluation. Evaluations on the training set show that system performance can be improved by about 8.7% by refinement of the rules. The overall micro-averaged precision, recall and F-score of end to end evaluation on the test set are 89.44%, 80.60% and 84.79% respectively. Our study indicates the feasibility of this approach to automate the population of structured template from narrative reports with promising results. Error analysis reveals that a single specimen report with standard headings and the presence of simple and concise statements is significantly associated with correct populations. In conclusion, the system can improve pathology reporting, and data mining for cancer registries, clinical audits and epidemiology research.
Objective: To extract pertinent information from narrative pathology reports and automatically populate structured templates. Materials and methods: A processing pipeline system has been developed which consists of: supervised machine learning based approach with conditional random field learner used for medical entity recognition, and rule-based methods for the population of structured templates. In total 612 narrative pathology reports of colorectal cancer were collected for evaluation. Results: The best model of the medical entity recognition experiments with 10- fold cross-validation on the training set achieved the micro-averaged precision with 80.58%, recall with 76.33% and F-score with 78.40%. The overall micro-averaged precision, recall and F-score of end-to-end evaluation on the test set are 85.18%, 78.75% and 81.84% respectively. Discussion: Our study shows that it is feasible to automatically populate structured reports by using a cascaded approach that integrates machine learning and several rule-based methods. It also reveals that the rules designed for structured template population are competent to populate the structured outputs and incorrect results from medical entity recognition such as the low recall on De:Mesorectal Integrity are the major cause of the errors (over 80%). Conclusion: With further improvement (especially for medical entity recognition), the system can contribute to a higher quality of pathology reporting and improve the efficiency for cancer registries, clinical audits and epidemiology research.
This paper reports on the issues in mapping the terms of a query to the field names of the schema of an Entity Relationship (ER) model or to the data part of the Entity Attribute Value (EAV) model using similarity based Top-K algorithm in clinical information system together with an extension of EAV mapping for medication names. In addition, the details of the mapping algorithm and the required pre-processing including NLP (Natural Language Processing) tasks to prepare resources for mapping are explained. The experimental results on an example clinical information system demonstrate more than 84 per cent of accuracy in mapping. The results will be integrated into our proposed Clinical Data Analytics Language (CliniDAL) to automate mapping process in CliniDAL.
Objective: There are abundant mentions of clinical conditions, anatomical sites, medications and procedures in clinical documents. This paper describes use of a cascade of machine learners to automatically extract mentions of named entities about disorders from clinical notes. Tasks: A Conditional Random Field (CRF) machine learner has been used for named entity recognition and to capture more complex (multiple word) named entities we have used Support Vector Machines (SVM). Firstly, the training data was converted to the CRF format. Different feature sets were ap- plied using 10-fold cross validation to find the best feature set for the machine learning model. Secondly, the identified named entities were passed to the SVM to find any relation among the identified disorder mentions to decide whether they are a part of a complex disorder. Approach: Our approach was based on a novel supervised learning model which incorporates two machine learning algorithms (CRF and SVM). Evalua- tion of each step included precision, recall and F-score metrics. Resources: We have used several tools which are created in our lab includ- ing TTSCT (Text to SNOMED CT) service, Lexical Management System (LMS) and Ring-fencing approach. A set of gazetteers was created from the training data and employed in analysis as well. Results: Evaluation results produced a precision of 0.766, recall of 0.726 and F-score of 0.746 for named entity recognition based on 10-fold cross vali- dation; and precision, recall and F-measure of 0.927 for relation extraction based on 5-fold cross validation on the training data. On the official test data on strict mode a precision of 0.686, recall of 0.539 and F-score of 0.604 was achieved. Based on the results our team was the 11 th out of 25 participating teams. In the relaxed mode a precision of 0.912, recall of 0.701 and F-score of 0.793 was recorded and our team was the 12 th . A multi stage supervised ma- chine learning method with mixed computational strategies seems to provide a reasonable strategy for automated extraction of disorders.
Joseph G. Davis合作论文数Director of Information Systems
Associate Head of School
School of Information Technologies J12
University of Sydney NSW 20062
Elisabeth Crawford合作论文数Computer Science Department, Carnegie Mellon University2