OBJECTIVES:To discuss the relationships between ontologies, terminologies and language in the context of Natural Language Processing (NLP) applications in order to show the negative consequences of confusing them.METHODS:The viewpoints of the terminologist and (computational) linguist are developed separately, and then compared, leading to the presentation of reconciliation among these points of view, with consideration of the role of the ontologist.RESULTS:In order to encourage appropriate usage of terminologies, guidelines are presented advocating the simultaneous publication of pragmatic vocabularies supported by terminological material based on adequate ontological analysis.CONCLUSIONS:Ontologies, terminologies and natural languages each have their own purpose. Ontologies support machine understanding, natural languages support human communication, and terminologies should form the bridge between them. Therefore, future terminology standards should be based on sound ontology and do justice to the diversities in natural languages. Moreover, they should support local vocabularies, in order to be easily adaptable to local needs and practices.
OBJECTIVES:The importance of clinical communication between providers, consumers and others, as well as the requisite for computer interoperability, strengthens the need for sharing common accepted terminologies. Under the directives of the World Health Organization (WHO), an approach is currently being conducted in Australia to adopt a standardized terminology for medical procedures that is intended to become an international reference.METHOD:In order to achieve such a standard, a collaborative approach is adopted, in line with the successful experiment conducted for the development of the new French coding system CCAM. Different coding centres are involved in setting up a semantic representation of each term using a formal ontological structure expressed through a logic-based representation language. From this language-independent representation, multilingual natural language generation (NLG) is performed to produce noun phrases in various languages that are further compared for consistency with the original terms.RESULTS:Outcomes are presented for the assessment of the International Classification of Health Interventions (ICHI) and its translation into Portuguese. The initial results clearly emphasize the feasibility and cost-effectiveness of the proposed method for handling both a different classification and an additional language.CONCLUSION:NLG tools, based on ontology driven semantic representation, facilitate the discovery of ambiguous and inconsistent terms, and, as such, should be promoted for establishing coherent international terminologies.
Multilingual natural language processing (NLP), whether it concerns analysis or generation of sentences, requires a sound language-independent representation for grasping the deep meaning of narratives. The formalism of conceptual graphs (CGs), especially designed to cope with natural language semantics, constitutes a good repository for dealing with the compositionality and intricacies of medical language. This paper describes our experiment, as part of the European GALEN project, for exploiting a conceptual graph representation of medical language, upon which multilingual medical language processing is performed.
The CCAM French coding system of clinical procedures was developed between 1994 and 2004 using, in parallel, a traditional domain expert's consensus method on one hand, and advanced methodologies of ontology driven semantic representation and multilingual generation on the other hand. These advanced methodologies were applied under the framework of an European Union collaborative research project named GALEN and produced a new generation of biomedical terminology. Following the interest in several countries and in WHO, the GALEN network has tested the application of the ontology driven tools to the existing reduced Australian ICHI coding system for interventions presently under investigation by WHO to check its ability and appropriateness to become the reference international coding system for procedures. The initial results are presented and discussed in terms of feasibility and quality assurance for sharing and maintaining consistent medical knowledge and allowing diversity in linguistic expressiveness of end users.
In the European Union, the need for systems which are able to accept multiple European languages is of paramount interest, because language barriers can be a strong impediment for large-scale communication in Europe. The use of analysers able to accept different European languages and convert them into a single representation common to all languages would seem to be the ideal solution. The RECIT system presented in this paper, shows an original approach for analysing sentences, understanding their meaning and storing them into a deep representation, available for future querying. The chosen approach, called proximity processing, takes advantage of the typical situation of a closed domain of knowledge (i.e. medicine) and of the structured form of medical reports (discharge summaries), using proximity rules which combine in an integrated way semantic information as well as syntactic information when needed. From the recognition of meaningful components in free text sentences, a knowledge representation is built in the form of conceptual graphs. In this article, we discuss the relevant features of both proximity processing and the subsequent transformation into a language-independent representation. In particular, we highlight the characteristics that enable our system to be easily extended to other European languages, as well as other application domains when pertinent.
This paper presents how acquisition, storage and communication of clinical documents is implemented at the University Hospitals of Geneva. Careful attention has been given to user-interfaces, in order to support complex layouts, spell checking, and templates management with automatic prefilling. A dual architecture has been developed for storage using an entity-attribute-value unified database and a consolidated, patient-centered, layout-respectful file-based storage, providing both representation power and speed of access. This architecture allows a great flexibility for storing a continuum of data types, ranging from simple typed values to complex clinical reports. Finally, communication is entirely based on HTTP-XML internally, and a HL-7 CDA interface V2 is currently studied for external communication. Some of the problems encountered, mostly related to the typology of documents and the ontology of clinical attributes are evoked.
We report on the design of a system for correcting spelling errors resulting in non-existent words. The system aims at improving edition of medical reports. Unlike traditional systems, both semantic and syntactic contexts are considered here. The system is organized along three steps. The first module is based on a context independent string-to-string edit distance calculus. The second module, based on the morpho-syntactic context attempts to rank more relevantly the data set provided by the first module, finally a third contextual module processes words with the same part-of-speech by applying some contextual word-sense disambiguation. Modules 2 and 3 are using both hand written rules and data-driven Markovian matrices. A final evaluation shows a significant improvement compared to context-free spelling correction.
Over the past two decades, two challenging, ongoing domains of medical informatics research have been the construction of models for medical concept representation and the inter-related task of understanding the deep meaning of medical free texts. These domains can take advantage of each other by exploiting the rich semantic content embedded in both concept models and medical texts. This review highlights how these two inter-related domains have evolved by focusing on a number of significant works in this area. The discussion examines one particular aspect: how to employ medical modeling for the purpose of medical language understanding. The understanding process analyzes and extracts the content of medical free texts, and stores the information in a deep semantic representation, useful for future elaborated semantic-driven information retrievals. It is now recognized in the medical informatics community that such understanding processes can be augmented through use of a domain-specific knowledge base that describes what can actually occur in a given domain. For this, a well-balanced representation schema should be developed somewhere between a partial but accurate, versus a complete but complex semantic representation. These observations are illustrated using examples from two major independent efforts undertaken by the authors: the elaboration and the subsequent adjustment of the RECIT multilingual analyzer to a solid model of medical concepts, and the recasting of a frame-based interlingua system, originally developed to map equivalent concepts between controlled clinical vocabularies.
This article considers recent developments in Natural Language Processing of medical texts and is an attempt to figure out the emerging trends in the years to come. New Natural Language Processing (NLP) tools for professionals are soon to be delivered. Once they are on the market, the medical documentation methods may undergo a complete revolution. The question is to know when such a change may occur and what are the expected functionalities.