Introduction: This study investigates whether it is possible to predict a final diagnosis based on a written nephropathological description—as a surrogate for image analysis—using various NLP methods. Methods: For this work, 1107 unlabelled nephropathological reports were included. (i) First, after separating each report into its microscopic description and diagnosis section, the diagnosis sections were clustered unsupervised to less than 20 diagnostic groups using different clustering techniques. (ii) Second, different text classification methods were used to predict the diagnostic group based on the microscopic description section. Results: The best clustering results (i) could be achieved with HDBSCAN, using BoW-based feature extraction methods. Based on keywords, these clusters can be mapped to certain diagnostic groups. A transformer encoder-based approach as well as an SVM worked best regarding diagnosis prediction based on the histomorphological description (ii). Certain diagnosis groups reached F1-scores of up to 0.892 while others achieved weak classification metrics. Conclusion: While textual morphological description alone enables retrieving the correct diagnosis for some entities, it does not work sufficiently for other entities. This is in accordance with a previous image analysis study on glomerular change patterns, where some diagnoses are associated with one pattern, but for others, there exists a complex pattern combination.
INTRODUCTION:The aim of this study is to evaluate the use of a natural language processing (NLP) software to extract medication statements from unstructured medical discharge letters.METHODS:Ten randomly selected discharge letters were extracted from the data warehouse of the University Hospital Erlangen (UHE) and manually annotated to create a gold standard. The AHD NLP tool, provided by MIRACUM's industry partner was used to annotate these discharge letters. Annotations by the NLP tool where then compared to the gold standard on two levels: phrase precision (whether or not the whole medication statement has been identified correctly) and token precision (whether or not the medication name has been identified correctly within correctly discovered medication phrases).RESULTS:The NLP tool detected medication related phrases with an overall F-measure of 0.852. The medication name has been identified correctly with an overall F-measure of 0.936.DISCUSSION:This proof-of-concept study is a first step towards an automated scalable evaluation system for MIRACUM's industry partner's NLP tool by using a gold standard. Medication phrases and names have been correctly identified in most cases by the NLP system. Future effort needs to be put into extending and validating the gold standard.
Feeding cancer registries with data extracted from textual reports, while maintaining a high level of data quality, has always been a labour-intensive task, due to the heterogeneity of the sources. The support of this task by IT solutions is expected to accelerate and optimise this process. To this end, the commercial text mining system Averbis Health Discovery was tailored to extract information from free text at the cancer registry of the federal state of Baden-Württemberg. The following entity types were extracted from German-language pathology reports: tumour localisation and morphology, pTNM, grading, (sentinel) nodes examined and affected, laterality and R-class. According to the entity type, several machine learning approaches as well as rules were used for the tumour types breast, prostate, colorectal and skin. Whereas for the pilot site, F values ranged between 0.800 and 0.996, values dropped when applying the extraction pipeline to two new sites (cancer registries Rhineland-Palatinate and Lower Saxony), for morphology from 0.950 to 0.657 and 0.933, and for localisation (topography) from 0.902 to 0.675 and 0.768. There was much less difference with R-class and lymph node counts. A thorough error analysis revealed numerous issues that explain these differences, such as different workflows between the sites, disagreements between textual and coded content as well as different handlings of missing values.
Anti-VEGF-Medikamente prägen heute die Therapie von Makulaerkrankungen. In diesem Zusammenhang wird eine Fülle zusätzlicher Daten erhoben. Damit ließen sich Behandlungsverläufe besser verstehen und vorhersagen. Allerdings sind diese Informationen meist nur in freitextlicher Form verfügbar. Wie weit auswertbare Information aus Kliniktexten automatisch gewonnen werden kann, sollte in einer retrospektiven Studie analysiert werden. Ziel war die Einschätzung der Eignung eines zu diesem Zweck parametrierten Text-Mining-Verfahrens. Es standen Daten zu 3683 Patienten zur Verfügung, davon 40.485 Arztbriefe. Für einen Teil waren die interessierenden Daten (Visus, Tensio und Begleitdiagnosen) auch strukturiert erfasst worden und konnten so als Goldstandard für die Textanalyse dienen. Diese wurde mit dem System Averbis Health Discovery durchgeführt. Zur Optimierung auf die Extraktionsaufgabe wurde dieses mit Regelwissen sowie mit einem deutschsprachigen Fachvokabular für die internationale Medizinterminologie SNOMED CT angereichert. Die Übereinstimmung der Datenextrakte mit den strukturierten Datenbankeinträgen wird durch den F1-Wert beschrieben. Hierbei ergab sich eine Übereinstimmung von 94,7 % für den Visus, 98,3 % für die Tensio und 94,7 % für begleitende Diagnosen. Die manuelle Analyse nicht übereinstimmender Fälle zeigte zur Hälfte, dass Textinhalte aus verschiedenen Gründen von Datenbankinhalten abwichen. Nach einer daraus berechneten Adjustierung lagen die F1-Werte noch 1–3 % über den zuvor ermittelten Werten. Für den betrachteten Arztbriefkorpus und die beschriebene Fragestellung sind Text-Mining-Verfahren sehr gut geeignet, um Inhalte zur weiteren Auswertung strukturiert aus Kliniktexten zu extrahieren.
Semantic standards and human language technologies are key enablers for semantic interoperability across heterogeneous document and data collections in clinical information systems. Data provenance is awarded increasing attention, and it is especially critical where clinical data are automatically extracted from original documents, e.g. by text mining. This paper demonstrates how the output of a commercial clinical text-mining tool can be harmonised with FHIR, the leading clinical information model standard. Character ranges that indicate the origin of an annotation and machine generates confidence values were identified as crucial elements of data provenance in order to enrich text-mining results. We have specified and requested necessary extensions to the FHIR standard and demonstrated how, as a result, important metadata describing processes generating FHIR instances from clinical narratives can be embedded.
In this article we investigate whether a custom or an established biomedical terminology is recommended for coding cancer datasets. We first give an overview of biomedical terminology focused on the domain of cancer and introduce three clinical use cases to demonstrate how several cancer aspects can be coded using ICD-10, ICD-O, TNM, MeSH, NCIt, MedDRA, and SNOMED CT. The same collection of terminologies was used in a case study where two dimensions of cancer (anatomy and histology) had already been coded in a dataset using a custom terminology. Although our experiments were limited in terms of coders (2) and coding cases (250 in total, 50 double-coded), they showed that, in most cases, equivalent concepts already existed in standard biomedical terminologies. SNOMED CT and NCIt provided the highest coverage (88% vs. 93%), with NCIt showing a much higher agreement (unweighted Kappa of 71% vs. 49%). As a general conclusion, for annotating cancer datasets we recommend the use of standard terminologies and mappings to local interface terminologies or value sets, instead of building a custom terminology from scratch.
Clinical data have their own peculiarities, as they evolve over time, may be incomplete, and are highly heterogeneous. These characteristics turn a thorough analysis into a challenging task, especially since domain experts are aware of the data flaws, which may impact their trust in the data. As we obtained anonymized clinical data from more than 3,500 patients with retinal diseases, we have to address these challenges. We define a workflow that integrates data cleansing and exploration in an iterative process, so that users are able to easily find anomalies and patterns in the data at any point in their analysis. We implement our workflow in a user-centered visual analytics tool with dedicated visualization and interaction techniques. In collaboration with experts, we apply our tool to examine the interdependency between patients’ visual acuity developments and treatment patterns. We find, that real-life data often have unforeseen incidents which can strongly influence the overall visual acuity development. This differs to study results, which are usually conducted under restrictive conditions and have shown visual acuity improvement with on-schedule treatment.
Background In secondary data there are often unstructured free texts. The aim of this study was to validate a text mining system to extract unstructured medical data for research purposes. Methods From a radiological department, 1,000 out of 7,102 CT findings were randomly selected. These were manually divided into defined groups by 2 physicians. For automated tagging and reporting, the text analysis software Averbis Extraction Platform (AEP) was used. Special features of the system are a morphological analysis for the decomposition of compound words as well as the recognition of noun phrases, abbreviations and negated statements. Based on the extracted standardized keywords, findings reports were assigned to the given findings groups using machine learning methods. To assess the reliability and validity of the automated process, the automated and two independent manual mappings were compared for matches in multiple runs. Results Manual classification was too time-consuming. In the case of automated keywording, the classification according to ICD-10 turned out to be unsuitable for our data. It also showed that the keyword search does not deliver reliable results. Computer-aided text mining and machine learning resulted in reliable results. The inter-rater reliability of the two manual classifications, as well as the machine and manual classification was very high. Both manual classifications were consistent in 93% of all findings. The kappa coefficient is 0.89 [95% confidence interval (CI) 0.87-0.92]. The automatic classification agreed with the independent, second manual classification in 86% of all findings (Kappa coefficient 0.79 [95% CI 0.75-0.81]). Discussion The classification of the software AEP was very good. In our study, however, it followed a systematic pattern. Most misclassifications were found in findings that indicate an increased risk of cancer. The free-text structure of the findings raises concerns about the feasibility of a purely automated analysis. The combination of human intellect and intelligent, adaptive software appears most suitable for mining unstructured but important textual information for research.
Summary Introduction: This article is part of the Focus Theme of Methods of Information in Medicine on the German Medical Informatics Initiative. Similar to other large international data sharing networks (e.g. OHDSI, PCORnet, eMerge, RD-Connect) MIRACUM is a consortium of academic and hospital partners as well as one industrial partner in eight German cities which have joined forces to create interoperable data integration centres (DIC) and make data within those DIC available for innovative new IT solutions in patient care and medical research. Objectives: Sharing data shall be supported by common interoperable tools and services, in order to leverage the power of such data for biomedical discovery and moving towards a learning health system. This paper aims at illustrating the major building blocks and concepts which MIRACUM will apply to achieve this goal. Governance and Policies: Besides establishing an efficient governance structure within the MIRACUM consortium (based on the steering board, a central administrative office, the general MIRACUM assembly, six working groups and the international scientific advisory board), defining DIC governance rules and data sharing policies, as well as establishing (at each MIRACUM DIC site, but also for MIRACUM in total) use and access committees are major building blocks for the success of such an endeavor. Architectural Framework and Methodology: The MIRACUM DIC architecture builds on a comprehensive ecosystem of reusable open source tools (MIRACOLIX), which are linkable and interoperable amongst each other, but also with the existing software environment of the MIRACUM hospitals. Efficient data protection measures, considering patient consent, data harmonization and a MIRACUM metadata repository as well as a common data model are major pillars of this framework. The methodological approach for shared data usage relies on a federated querying and analysis concept. Use Cases: MIRACUM aims at proving the value of their DIC with three use cases: IT support for patient recruitment into clinical trials, the development and routine care implementation of a clinico-molecular predictive knowledge tool, and molecular-guided therapy recommendations in molecular tumor boards. Results: Based on the MIRACUM DIC release in the nine months conceptual phase first large scale analysis for stroke and colorectal cancer cohorts have been pursued. Discussion: Beyond all technological challenges successfully applying the MIRACUM tools for the enrichment of our knowledge about diagnostic and therapeutic concepts, thus supporting the concept of a Learning Health System will be crucial for the acceptance and sustainability in the medical community and the MIRACUM university hospitals.
Hans-Ulrich Prokosch1; Till Acker2; Johannes Bernarding3; Harald Binder4,5; Martin Boeker5; Melanie Boerries6; Philipp Daumke7; Thomas Ganslandt1,8; Jürgen Hesser9; Gunther Höning10; Michael Neumaier11; Kurt Marquardt12; Harald Renz13; Hermann-Josef Rothkötter14; Carmen Schade-Brittinger15; Paul Schmücker16; Jürgen Schüttler17; Martin Sedlmayr1,18; Hubert Serve19; Keywan Sohrabi20; Holger Storf21 1Chair of Medical Informatics, Department of Medical Informatics, Biometrics and Epidemiology, Friedrich-AlexanderUniversity Erlangen-Nürnberg, Erlangen, Germany; 2Institute of Neuropathology, Justus-Liebig-University Giessen, Giessen, Germany; 3Chair of Medical Informatics, Institute for Biometry and Medical Informatics, Otto-von-Guericke-University Magdeburg, Magdeburg, Germany; 4Institute of Medical Biostatistics, Epidemiology and Informatics, University Medical Center of the Johannes Gutenberg University Mainz, Mainz, Germany; 5Institute of Medical Biometry and Statistics, Medical Faculty and Medical Center – University of Freiburg, Freiburg, Germany; 6Institute of Molecular Medicine and Cell Research and Comprehensive Cancer Center Freiburg (CCCF), University Medical Center, Faculty of Medicine, University of Freiburg; German Cancer Research Center (DKFZ), Heidelberg and German Cancer Consortium (DKTK) partner site Freiburg, Freiburg, Germany; 7Averbis GmbH, Freiburg, Germany; 8 Department of Biomedical Informatics, University Medicine Mannheim, Ruprecht-Karls-University Heidelberg, Mannheim, Germany; 9Experimental Radiation Oncology Department, University Medical Center Mannheim, Central Institute for Scientific Computing (IWR), Central Institute for Computer Engineering (ZITI), Heidelberg University, Mannheim, Germany; 10Department of Information Technology, University Medical Center of the Johannes Gutenberg University Mainz, Mainz, Germany; 11Chair for Clinical Chemistry, Medical Faculty Mannheim of Heidelberg University, Mannheim, Germany; 12University Hospital of Giessen and Marburg, Giessen, Germany; 13 Chair for Clinical Chemistry, Philipps University Marburg, Medical Director of the University Clinic Marburg, Marburg, Germany; 14Institute of Anatomy, Otto-von-Guericke-University Magdeburg, Dean of the Medical Faculty, Magdeburg, Germany; 15 Chair of the Coordinating Centre for Clinical Trials, Philipps University Marburg, Marburg, Germany; 16 University of Applied Sciences Mannheim, Institute for Medical Informatics, Mannheim, Germany; 17Department of Anesthesiology, University of Erlangen-Nürnberg, Dean of the Medical Faculty, Erlangen, Germany; 18 Institute of Medical Informatics and Biometrics, Carl Gustav Carus Faculty of Medicine, Technische Universität Dresden, Dresden, Germany; 19 Department of Hematology and Oncology, University Hospital Frankfurt, Goethe University, Frankfurt am Main, Germany; 20 Faculty of Health Sciences, University of Applied Sciences – THM, Giessen, Germany; 21 Medical Informatics Group, University Hospital Frankfurt, Goethe University, Frankfurt am Main, Germany
Alfred Winter1; Sebastian Stäubert1; Danny Ammon2; Stephan Aiche3; Oya Beyan4; Verena Bischoff5; Philipp Daumke6; Stefan Decker4; Gert Funkat7; Jan E. Gewehr8; Armin de Greiff9; Silke Haferkamp10; Udo Hahn11; Andreas Henkel2; Toralf Kirsten12; Thomas Klöss13; Jörg Lippert14; Matthias Löbe1; Volker Lowitsch10; Oliver Maassen15; Jens Maschmann16; Sven Meister17; Rafael Mikolajczyk18; Matthias Nüchter12; Mathias W. Pletz19; Erhard Rahm20; Morris Riedel21; Kutaiba Saleh2; Andreas Schuppert22; Stefan Smers7; André Stollenwerk23; Stefan Uhlig24; Thomas Wendt25; Sven Zenker26; Wolfgang Fleig27,**; Gernot Marx15,**; André Scherag28, 29,**; Markus Löffler1,** 1Leipzig University, Institute of Medical Informatics, Statistics and Epidemiology, Leipzig, Germany; 2University Medical Center Jena, Central Service Provider For Information Technology, Jena, Germany; 3SAP SE, Potsdam, Germany; 4RWTH Aachen University, Chair of Computer Science 5, Aachen, Germany; 5University of Leipzig Medical Center, Division Staff and Justice, Leipzig, Germany; 6Averbis GmbH, Freiburg, Germany; 7University of Leipzig Medical Center, Division Information Management, Leipzig, Germany; 8University Medical Center Hamburg-Eppendorf, Business Division for Information Technology, Hamburg, Germany; 9Essen University Hospital, Central Information Technology, Essen, Germany; 10RWTH Aachen University Hospital, Division Information Technology, Aachen, Germany; 11Friedrich-Schiller-Universität Jena, Language & Information Engineering Lab (JULIE Lab), Jena, Germany; 12Leipzig University, LIFE Research Centre for Civilization Diseases, Leipzig, Germany; 13Martin-Luther-Universität Halle-Wittenberg Medical Center, Medical Director, Halle, Germany; 14Bayer AG, Wuppertal, Germany; 15RWTH Aachen University Hospital, Department of Intensive Care and Intermediate Care, Aachen, Germany; 16University Medical Center Jena, Medical Director, Jena, Germany; 17Fraunhofer Institute for Software and Systems Engineering, Dortmund, Germany; 18Martin-Luther-Universität Halle-Wittenberg, Institute of Medical Epidemiology, Biometry and Informatics, Halle, Germany; 19University Medical Center Jena, Institute of Infectious Diseases and Infection Control, Jena, Germany; 20Leipzig University, Department of Computer Science – Database Group, Leipzig, Germany; 21Forschungszentrum Jülich, Jülich Supercomputing Centre, Jülich, Germany; 22RWTH Aachen University, Institute for Computational Biomedicine II, Aachen, Germany; 23RWTH Aachen University, Informatik 11 – Embedded Software, Aachen, Germany; 24RWTH Aachen University, Medical Faculty, Dean, Aachen, Germany; 25University of Leipzig Medical Center, Data Integration Center, Leipzig, Germany; 26University of Bonn Medical Center, Department of Anesthesiology and Intensive Care Medicine, Bonn, Germany; 27University of Leipzig Medical Center, Medical Director, Leipzig, Germany; 28University Medical Center Jena, Center for Sepsis Control and Care, Jena, Germany; 29University Medical Center Jena, Institute of Medical Statistics, Computer and Data Sciences (IMSID), Jena, Germany
Introduction: This article is part of the Focus Theme of Methods of Information in Medicine on the German Medical Informatics Initiative. Similar to other large international data sharing networks (e.g. OHDSI, PCORnet, eMerge, RD-Connect) MIRACUM is a consortium of academic and hospital partners as well as one industrial partner in eight German cities which have joined forces to create interoperable data integration centres (DIC) and make data within those DIC available for innovative new IT solutions in patient care and medical research. Objectives: Sharing data shall be supported by common interoperable tools and services, in order to leverage the power of such data for biomedical discovery and moving towards a learning health system. This paper aims at illustrating the major building blocks and concepts which MIRACUM will apply to achieve this goal. Governance and Policies: Besides establishing an efficient governance structure within the MIRACUM consortium (based on the steering board, a central administrative office, the general MIRACUM assembly, six working groups and the international scientific advisory board), defining DIC governance rules and data sharing policies, as well as establishing (at each MIRACUM DIC site, but also for MIRACUM in total) use and access committees are major building blocks for the success of such an endeavor. Architectural Framework and Methodology: The MIRACUM DIC architecture builds on a comprehensive ecosystem of reusable open source tools (MIRACOLIX), which are linkable and interoperable amongst each other, but also with the existing software environment of the MIRACUM hospitals. Efficient data protection measures, considering patient consent, data harmonization and a MIRACUM metadata repository as well as a common data model are major pillars of this framework. The methodological approach for shared data usage relies on a federated querying and analysis concept. Use Cases: MIRACUM aims at proving the value of their DIC with three use cases: IT support for patient recruitment into clinical trials, the development and routine care implementation of a clinico-molecular predictive knowledge tool, and molecular-guided therapy recommendations in molecular tumor boards. Results: Based on the MIRACUM DIC release in the nine months conceptual phase first large scale analysis for stroke and colorectal cancer cohorts have been pursued. Discussion: Beyond all technological challenges successfully applying the MIRACUM tools for the enrichment of our knowledge about diagnostic and therapeutic concepts, thus supporting the concept of a Learning Health System will be crucial for the acceptance and sustainability in the medical community and the MIRACUM university hospitals.
Summary Introduction: This article is part of the Focus Theme of Methods of Information in Medicine on the German Medical Informatics Initiative. “Smart Medical Information Technology for Healthcare (SMITH)” is one of four consortia funded by the German Medical Informatics Initiative (MI-I) to create an alliance of universities, university hospitals, research institutions and IT companies. SMITH’s goals are to establish Data Integration Centers (DICs) at each SMITH partner hospital and to implement use cases which demonstrate the usefulness of the approach. Objectives: To give insight into architectural design issues underlying SMITH data integration and to introduce the use cases to be implemented. Governance and Policies: SMITH implements a federated approach as well for its governance structure as for its information system architecture. SMITH has designed a generic concept for its data integration centers. They share identical services and functionalities to take best advantage of the interoperability architectures and of the data use and access process planned. The DICs provide access to the local hospitals’ Electronic Medical Records (EMR). This is based on data trustee and privacy management services. DIC staff will curate and amend EMR data in the Health Data Storage. Methodology and Architectural Framework: To share medical and research data, SMITH’s information system is based on communication and storage standards. We use the Reference Model of the Open Archival Information System and will consistently implement profiles of Integrating the Health Care Enterprise (IHE) and Health Level Seven (HL7) standards. Standard terminologies will be applied. The SMITH Market Place will be used for devising agreements on data access and distribution. 3LGM2 for enterprise architecture modeling supports a consistent development process. The DIC reference architecture determines the services, applications and the standards-based communication links needed for efficiently supporting the ingesting, data nourishing, trustee, privacy management and data transfer tasks of the SMITH DICs. The reference architecture is adopted at the local sites. Data sharing services and the market place enable interoperability. Use Cases: The methodological use case “Phenotype Pipeline” (PheP) constructs algorithms for annotations and analyses of patient-related phenotypes according to classification rules or statistical models based on structured data. Unstructured textual data will be subject to natural language processing to permit integration into the phenotyping algorithms. The clinical use case “Algorithmic Surveillance of ICU Patients” (ASIC) focusses on patients in Intensive Care Units (ICU) with the acute respiratory distress syndrome (ARDS). A model-based decision-support system will give advice for mechanical ventilation. The clinical use case HELP develops a “hospital-wide electronic medical record-based computerized decision support system to improve outcomes of patients with blood-stream infections” (HELP). ASIC and HELP use the PheP. The clinical benefit of the use cases ASIC and HELP will be demonstrated in a change of care clinical trial based on a step wedge design. Discussion: SMITH’s strength is the modular, reusable IT architecture based on interoperability standards, the integration of the hospitals’ information management departments and the public-private partnership. The project aims at sustainability beyond the first 4-year funding period.
The vast amount of clinical data in electronic health records constitutes a great potential for secondary use. However, most of this content consists of unstructured or semi-structured texts, which is difficult to process. Several challenges are still pending: medical language idiosyncrasies in different natural languages, and the large variety of medical terminology systems. In this paper we present SEMCARE, a European initiative designed to minimize these problems by providing a multi-lingual platform (English, German, and Dutch) that allows users to express complex queries and obtain relevant search results from clinical texts. SEMCARE is based on a selection of adapted biomedical terminologies, together with Apache UIMA and Apache Solr as open source state-of-the-art natural language pipeline and indexing technologies. SEMCARE has been deployed and is currently being tested at three medical institutions in the UK, Austria, and the Netherlands, showing promising results in a cardiology use case.
This article is about a new project that combines clinical data intelligence and smart data. It provides an introduction to the “Klinische Datenintelligenz” (KDI) project which is founded by the Federal Ministry for Economic Affairs and Energy (BMWi); we transfer research and development results (R&D) of the analysis of data which are generated in the clinical routine in specific medical domain. We present the project structure and goals, how patient care should be improved, and the joint efforts of data and knowledge engineering, information extraction (from textual and other unstructured data), statistical machine learning, decision support, and their integration into special use cases moving towards individualised medicine. In particular, we describe some details of our medical use cases and cooperation with two major German university hospitals.
Patients with chronic diseases undergo numerous in- and outpatient treatment periods, and therefore many documents accumulate in their electronic records. We report on an on-going project focussing on the semantic enrichment of medical texts, in order to support recall-oriented navigation across a patient's complete documentation. A document pool of 1,696 de-identified discharge summaries was used for prototyping. A natural language processing toolset for document annotation (based on the text-mining framework UIMA) and indexing (Solr) was used to support a browser-based platform for document import, search and navigation. The integrated search engine combines free text and concept-based querying, supported by dynamically generated facets (diagnoses, procedures, medications, lab values, and body parts). The prototype demonstrates the feasibility of semantic document enrichment within document collections of a single patient. Originally conceived as an add-on for the clinical workplace, this technology could also be adapted to support personalised health record platforms, as well as cross-patient search for cohort building and other secondary use scenarios.
Thomas Wittenberg合作论文数Fraunhofer IIS3