Monitoring the occurrence of adverse events in the scientific literature is a mandatory process in drug marketing surveillance. This is a very time-consuming and complex task to fulfill the compliance and, most importantly, to ensure patient safety. Therefore, a machine learning (ML) algorithm has been trained to support this manual intellectual review process, by automatically providing a classification of the literature articles into two types. An algorithm has been designed to automatically classify "relevant articles" which are reporting any kind of drug safety relevant information, and those which are not reporting an adverse drug reaction as "not relevant." The review process is consisted of many rules and aspects which needed to be taken into consideration. Therefore, for the training of the algorithm, thousands of documents from previous screenings have been used. After several iterations of adjustments and fine tuning, the ML approach is definitively a great achievement in pre-sorting the articles into "relevant" and "non-relevant" and supporting the intellectual review process.
The vast amount of clinical data in electronic health records constitutes a great potential for secondary use. However, most of this content consists of unstructured or semi-structured texts, which is difficult to process. Several challenges are still pending: medical language idiosyncrasies in different natural languages, and the large variety of medical terminology systems. In this paper we present SEMCARE, a European initiative designed to minimize these problems by providing a multi-lingual platform (English, German, and Dutch) that allows users to express complex queries and obtain relevant search results from clinical texts. SEMCARE is based on a selection of adapted biomedical terminologies, together with Apache UIMA and Apache Solr as open source state-of-the-art natural language pipeline and indexing technologies. SEMCARE has been deployed and is currently being tested at three medical institutions in the UK, Austria, and the Netherlands, showing promising results in a cardiology use case.
In this paper we describe how the EUCases FP7 project is addressing the problem of lifting Legal Open Data to Linked Open Data to develop new applications for the legal information provision market by enriching structurally the documents (first of all with navigable references among legal texts) and semantically (with concepts from ontologies and classification). First we describe the social and economic need for breaking the accessibility barrier in legal information in the EU, then we describe the technological challenges and finally we explain how the EUCases project is addressing them by a combination of Human Language Technologies.
To ensure the quality of a medical thesaurus is a non-trivial task, due to the inherent complexity of medical terminology. The peculiarities of the medical sublanguage and the subjectivism of lexicographers' choices complicate the thesaurus construction process. Our experience is based on the MorphoSaurus lexicon, the basis of a biomedical cross-language indexing and retrieval system. We describe two complementary maintenance approaches, viz. i) corpus-based error detection, and ii) thesaurus anomaly detection. These techniques were developed to detect so-called dynamic and static errors, which are committed by the lexicographers during the construction and maintenance process. Considering multilingual parallel corpora, the distribution of semantic identifiers should be similar whenever comparing related texts in different languages. In the first approach, those semantic identifiers are identified that exhibit greatest frequency variations when comparing text pairs. A manual review of these search results is supposed to spot content errors, which are subsequently classified and fixed by the lexicographers. The second approach analyses transaction-based anomalies, which are identified by interpreting the log of lexicographers' actions during thesaurus maintenance. This methodology highlights the four most common types of this kind of anomaly and evaluates the effectiveness of the corpus-based detection techniques. The overall quality improvement of the thesaurus was evaluated using the OHSUMED IR benchmark.
To ensure the quality of a medical thesaurus is a non-trivial task, due to the inherent complexity of medical terminology. The peculiarities of the medical sublanguage and the subjectivism of lexicographers' choices complicate the thesaurus construction process. Our experience is based on the MorphoSaurus lexicon, the basis of a biomedical cross-language indexing and retrieval system. We describe two complementary maintenance approaches, viz. i) corpus-based error detection, and ii) thesaurus anomaly detection. These techniques were developed to detect so-called dynamic and static errors, which are committed by the lexicographers during the construction and maintenance process. Considering multilingual parallel corpora, the distribution of semantic identifiers should be similar whenever comparing related texts in different languages. In the first approach, those semantic identifiers are identified that exhibit greatest frequency variations when comparing text pairs. A manual review of these search results is supposed to spot content errors, which are subsequently classified and fixed by the lexicographers. The second approach analyses transaction-based anomalies, which are identified by interpreting the log of lexicographers' actions during thesaurus maintenance. This methodology highlights the four most common types of this kind of anomaly and evaluates the effectiveness of the corpus-based detection techniques. The overall quality improvement of the thesaurus was evaluated using the OHSUMED IR benchmark.
Einleitung/Hintergrund: Leistungserbringer im Gesundheitswesen stehen vor der grosen Herausforderung, durch Innovationen im Bereich Forschung und Entwicklung die Behandlungsqualitat im Gesundheitswesen zu verbessern, die Patientensicherheit zu erhohen und gleichzeitig die Kosten [for full text, please go to the a.m. URL]
Objectives: The increasing amount of electronically available documents in bibliographic databases and the clinical documentation requires user-friendly techniques for content retrieval.Methods: A domain-specific approach on semantic text indexing for document retrieval is presented. It is based on a subword thesaurus and maps the content of texts in different European languages to a common interlingual representation, which supports the search across multilingual document collections.Results: Three use cases are presented where the semantic retrieval method has been implemented: a bibliographic database, a department EHR system, and a consumer-oriented Web portal.Conclusions: It could be shown that a semantic indexing and retrieval approach, the performance of which had already been empirically assessed in prior studies, proved useful in different prototypical and routine scenarios and was well accepted by several user groups.
Semantic interoperability is a major desideratum in health care for computer-based documentation and communication through electronic health records. They require structured data ideally represented via standardized information models, e.g., HL7 RIM or openEHR, connected to standardized terminologies or ontologies, e.g., SNOMED CT or LOINC. But since natural language is seen by health professionals as their most natural and effective form of expression, semantically interoperable architectures must adequately deal with unstructured data and be seamlessly integrated into the workflows of health professionals. Therefore we propose a self-learning natural language processing system, which automatically segments input narratives into sections, detects contexts such as negations, and assigns terminology codes. To be usable in clinical contexts, the system must properly handle the idiosyncratic medical language and grammar and spelling errors of narratives produced in everyday clinical practice. We present user interaction items that are important to make the interaction with the system as easy as possible to reach high acceptance among health professionals.
PURPOSE:Since 2003 the Radiological Society of North America (RSNA) has been developing a lexicon of standardized radiological terms (RadLex) intended to support the structured reporting of imaging observations and the indexing of teaching cases. The aim of this study was to translate the first version of the lexicon (1 - 2007) into German and to implement a language-independent online term browser.MATERIALS AND METHODS:RadLex version 1 - 2007 contains 6303 terms in nine main categories. Two radiologists independently translated the lexicon using medical dictionaries. Terms translated differently were revised and translated by consensus. For the development of an online term browser, a text processing algorithm called morphosemantic indexing was used which splits up words into small semantic units and compares those units to language-specific subword thesauri.RESULTS:In total 6240 of 6303 terms (99 %) were translated. Of those terms 3965 were German, 1893 were Latin, 359 were multilingual, and 23 were English terms that are also used in German and were therefore maintained. The online term browser supports a language-independent term search in RadLex (German/English) and other common medical terminology (e. g., ICD 10). The term browser displays term hierarchies and translations in different frames and the complexity of the result lists can be adapted by the user.CONCLUSION:RadLex version 1 - 2007 developed by the RSNA is now available in German and can be accessed online through a term browser with an efficient search function. This is an important precondition for the future comparison of national and international indexed radiological examination results and the interoperability between digital teaching resources.
Definitory expressions about clinical procedures, findings and diseases constitute a major benefit of a formally founded clinical reference terminology which is ontologically sound and suited for formal reasoning. SNOMED CT claims to support formal reasoning by description-logic based concept definitions.
In the 2006 ImageCLEF Medical Image Retrieval task we evaluate the effects of deep morphological analysis for mono-and cross-lingual document retrieval in the biomedical domain. The morphological analysis is based on the MorphoSaurus system in which subwords are introduced as morphologically meaningful word units. Subwords are organized in language specific lexica that were partly manually and partly automatically generated and currently cover six European languages. They are linked together in a multilingual thesaurus. The use of subwords instead of full words significantly reduces the number of lexical entries that are needed to sufficiently cover a specific language and domain. A further benefit of the approach is its independence from the underlying retrieval system. We combined MorphoSaurus with the open-source search engine Lucene and achieved precision gains of up to 25% over the baseline for a monolingual setting and promising results in a multilingual scenario.
BACKGROUND:An adequate and expressive ontological representation of biological organisms and their parts requires formal reasoning mechanisms for their relations of physical aggregation and containment. RESULTS:We demonstrate that the proposed formalism allows to deal consistently with "role propagation along non-taxonomic hierarchies", a problem which had repeatedly been identified as an intricate reasoning problem in biomedical ontologies. CONCLUSION:The proposed approach seems to be suitable for the redesign of compositional hierarchies in (bio)medical terminology systems which are embedded into the framework of the OBO (Open Biological Ontologies) Relation Ontology and are using knowledge representation languages developed by the Semantic Web community.
We propose an approach to multilingual medical document retrieval in which complex word forms are segmented according to medically relevant morpho-semantic criteria. At its core lies a multilingual dictionary, in which entries are equivalence classes of subwords, i.e. semantically minimal units. Using two different standard test collections for the medical domain, we evaluate our approach for six languages covered by our system.
In this paper we want to describe how the promising technology of biomedical data mining can improve the use of hospital information systems: a large set of unstructured, narrative clinical data from a dermatological university hospital like discharge letters or other dermatological reports were processed through a morpho-semantic text retrieval engine ("MorphoSaurus") and integrated with other clinical data using a web-based interface and brought into daily clinical routine. The user evaluation showed a very high user acceptance - this system seems to meet the clinicians' requirements for a vertical data mining in the electronic patient records. What emerges is the need for integration of biomedical data mining into hospital information systems for clinical, scientific, educational and economic reasons.
We describe an experiment on collecting large language and topic specific corpora automatically by using a focused Web crawler. Our crawler combines efficient crawling techniques with a common text classification tool. Given a sample corpus of medical documents, we automatically extract query phrases and then acquire seed URLs with a standard search engine. Starting from these seed URLs, the crawler builds a new large collection consisting only of documents that satisfy both the language and the topic model. The manual analysis of acquired English and German medicine corpora reveals the high accuracy of the crawler. However, there are significant differences between both languages.
Joachim Wermter合作论文数University Jena
Computational Linguistics Research Group2