
SNOMED CT® (SCT) is a large, comprehensive medical terminology with many applications in the health care IT sector. SCT is often mapped to existing billing classifications as well as to proprietary terminologies in order to support access to and from SCT from existing applications. Subsets of SCT are used to reduce complexity and size. These subsets can vary from small sets that will be used to populate drop down lists in electronic medical record applications and larger lists that are used for reference, e.g. the SCT Non-human subset. There are costs and time limitations in maintaining mappings and subsets, particularly after each SCT release when concepts are retired and new concepts are added. There is a need for a careful strategy to identify changes, determine which changes need to be reviewed, and to rank changes so they can be reviewed systematically in order of importance. Here we outline our updating strategies for the Health Language Medical Specialty Subsets, a list of 10,000 SCT terms grouped into 45 subsets. These strategies can be used for any subset of SCT as well as for mappings created to and from SCT. Introduction: Electronic health records (EHR) and other healthcare IT applications rely on controlled medical terminologies to provide well defined concepts for accurate and consistent encoding of records and data mining. SNOMED CT® (SCT) provides broad coverage of all medical domains with approximately 280,000 active concepts. A great deal of attention has been focused on the models and strategies to implement SCT. Most applications will use defined subsets of SCT for specific use cases rather than exposing all of SCT to all users. Subsets are collections, lists, of SCT concepts or terms. Applications using SCT will need to be able to store these lists for purposes of maintenance and delivery to EHR interfaces. Mappings are often created between SNOMED CT and other terminologies such as billing classifications, ICD-9-CM and ICD-10, as well as local and proprietary terminologies. Mappings can be represented as a pair of codes, the source and target of the map or relationship. These mappings can be used for translation of information from one set of codes to another. These mappings and subsets represent work that a local site or user is performing to the standard – in this case SCT. Thus, SCT is distributed from the standards body, and local users such as EHR vendors or hospitals, need to add value to their SCT version with mappings and subsets. It is critical that the local user adopts the next version of SCT in order to prevent semantic drift – multiple versions of a terminology being used that drift in meaning enough to be incompatible. In the paper, Oliver, et. al. discuss some of the principles of localization of terminologies and also explores the impact of migration of localized terminologies to the next version of the standard. Updating subsets and mappings as described in this paper actually only involves a small subset of the many kinds of changes that occur to the terminology. Cimino, et. al. discuss the types of changes that need to be considered such as refinement, name changes, code-reuse and more that impact the updating of terminologies and content based on them. The Semantic Web work is now introducing new issues with regards to ontology versioning. Liang, et. al. discuss the impact of changes in ontologies to existing applications that depend on them. In this semantic web paper, a middle layer to monitor and detect changes is proposed to be used between the underlying ontologies and the dependant applications. The work we present here with subsets uses a terminology service and specialized scripts to serve as this middle layer between the standard ontology, SCT in this case, and the resulting applications – EMRs for example that depend on the subsets. Maintenance of the mappings and subsets are costly and time dependent because the content is often in production. Maintenance can involve changes that are dictated by the applications that use the content but also because the underlying SCT data model has changed. Semiannual updates to SCT can change the SCT model to varying degrees, sometimes substantially. An SCT update involves new concepts and terms, retirement of concepts and terms, and addition and deletion of relationships. Each mapping Representing and sharing knowledge using SNOMED Proceedings of the 3rd international conference on Knowledge Representation in Medicine (KR-MED 2008) R. Cornet, K.A. Spackman (Eds)
Snomed ct is a large-scale medical ontology, which is developed using a variant of the inexpressive Description Logic EL. Description Logic reasoning can not only be used to compute subsumption relationships between Snomed concepts, but also to pinpoint the reason why a certain subsumption relationship holds by computing the axioms responsible for this relationship. This helps developers and users of Snomed ct to understand why a given subsumption relationship follows from the ontology, which can be seen as a first step toward removing unwanted subsumption relationships. In this paper, we describe a new method for axiom pinpointing in the Description Logic EL, which is based on the computation of so-called reachabilitybased modules. Our experiments on Snomed ct show that the sets of axioms explaining subsumption are usually quite small, and that our method is fast enough to compute such sets on demand.
As industrial, governmental, and academic agencies place increasing emphasis on translational research, biomedical researchers are now faced with entirely new challenges in regards to both biomedical data integration and knowledge discovery. There is now both a strong need and a tremendous opportunity to apply translational bioinformatics to address the fundamental challenges in integrating the vast bodies of -omics and clinical data. Here we report on our preliminary work in utilizing SNOMED-CT as both a tool for translational data discovery, and a major component in a framework for the large-scale integration of gene expression microarray data and clinical laboratory data. Annotations from microarray experiments in NCBI GEO were mapped to SNOMED-CT terms using UMLS, and these mappings were joined to clinical laboratory data using ICD9CM to SNOMED-CT mappings within UMLS. We find that microarray experiments characterizing 211 distinct diseases can be mapped to clinical laboratory data measurements for 13,452 distinct patients. We maintain that this work represents critical first steps in providing a foundation for large-scale translational data integration, and underlines the important role that controlled clinical terminologies, such as SNOMEDCT, can play in addressing such problems.
Findings related to developing implementation specifications for the use of SNOMED Clinical Terms (SNOMED CT) in both HL7 and openEHR information models are summarized and compared. Common themes from this work, including overlaps between the expressivity of structure and terminology, are identified and discussed. Distinctions are made between aspects of meaning that are most readily represented by distinct structures, others where terminology offers greater flexibility and a 'gray-area' in which the relative merits are more balanced. Focusing on particular stages in the clinical information life cycle may suggest different points of balance and may lead to different approaches to integration. However, greater consistency is essential if clinical information is to be used effectively in electronic record systems. Consensus guidance documents of the type developed by the work described are only a first step. Mutually aware evolutionary refinement of structural and terminology standards is suggested as an enhancement to independent development.
Work in the field of recording standard, coded data in electronic health records and messages is important to support interoperability of clinical systems. It is also important for reducing medical errors caused by misinterpretation and misrepresentation of data. Standardisation of structured and unstructured data to one or more terminologies such as SNOMED-CT, or ICD requires the help of various integration procedures. We have previously highlighted issues in data models (openEHR Archetypes) when mapping to a terminology model (SNOMED CT) [1]. In this paper, we describe issues with terminology models (SNOMED CT) when aligning the concepts to a data model (openEHR Archetypes). Terminologies and data models play an important role in building structured EHRs and achieving semantic interoperability. Semantic interoperability requires that all recorded data conforms to some reference terminology in order to interpret and reuse it uniformly in all partaking information systems. In the medical domain, standardising data is of great significance, as controlling the vocabulary used to record patient data is critical to making EHRs safe for exchange and reuse. The paper recognises the value of SNOMED CT but demonstrates the difficulties of working with it at an integration level. The difficulties in integration arise primarily due to the semantic gaps in the content of the structured data models and terminology models. The same issues might also arise with data obtained from unstructured sources. Despite the broad coverage that SNOMED offers, there are several concepts that are missing. An efficient process for submission of concepts for inclusion is in need, along with formal rules for post coordination. We believe that in order to achieve the overall objective of semantic interoperability, it is imperative that both data and terminology models are developed with the aim of being able to integrate their clinical content. It is important that both modeling communities are not only cognizant of each others existence but also work closely with each other to ensure that conformance is built into the systems from conception stage. These conformance or compatibility rules should be extended to all other stages of the modeling process i.e. at design time, data integration time, as well as at runtime. It is only then that true interoperability will be achieved, making it possible to build safer health care systems. Reliable and high quality data in these systems will improve the functioning of all health care units heavily dependent on data, reducing medical errors and ultimately providing safer and better patient care. Reference [1] Rahil Qamar, Jay Kola, and Alan Rector. Unambiguous data modeling to ensure higher accuracy term binding to clinical terminologies. AMIA 2007 Annual Symposium, November 2007. Chicago, U.S.A. Representing and sharing knowledge using SNOMED Proceedings of the 3rd international conference on Knowledge Representation in Medicine (KR-MED 2008) R. Cornet, K.A. Spackman (Eds)
The objective of this work is to provide a formalization of the semantics of SNOMED CT’s refinement rules in Description Logics and to exemplify their usage on a real world wound documentation system. The goal of unambiguous documentation and communication of medical information with explicit semantics can be reached by combining standards and terminologies. Information Models (e.g. the HL7 Clinical Document Architecture on Level 3) together with terminology systems (e.g. LOINC and SNOMED CT) are promising candidates for building a semantically interoperable framework for electronic health records. We investigated how LOINC and SNOMED CT concepts can unambiguously and completely cover user interface terms of an existing electronic, formbased documentation system used in clinical dermatology. Especially, the feasibility of postcoordinating complex expressions according to the SNOMED CT terminology model is target of our investigations. Besides analyzing completeness and uniqueness of the mappings and the userfriendliness of the mapping process, we discuss the different ways of post-coordination (refinement types) presented in SNOMED CT’s technical documentation. Where post-coordination was required, we adhered to the SNOMED CT terminology model refinement types and the “SNOMED Compositional Grammar” syntax. The manual mapping process proved to be time consuming and prone to ambiguous solutions where post-coordination of SNOMED CT expressions was necessary. However, for most user interface terms a complete semantic representation could be generated. A coverage of nearly 100% of clinical user interface terms shows the appropriateness of SNOMED CT as a reference terminology for the domain under scrutiny. The natural language descriptions of refinement types in the SNOMED CT documentation were formalized in Description Logics and reduced to four basic patterns. Problems with coding and post-coordination can be explained by weak documentation and poor tool support. The structure of the documentation forces users to collect necessary information from several SNOMED CT reference documents. Although mechanisms for post-coordination allowed to express a substantial amount of terms we suggest that tool support and formalized documentation for post-coordination (refinement) is enhanced. Tool support should reduce browsing complexity, support post-coordination and give clear advice how to use SNOMED CT according to the SNOMED CT compositional grammar and refinement rules. Furthermore, we recommend a thorough redesign of the post-coordination guidelines which entails the clarification of SNOMED CT's logical and ontological foundations. Representing and sharing knowledge using SNOMED Proceedings of the 3rd international conference on Knowledge Representation in Medicine (KR-MED 2008) R. Cornet, K.A. Spackman (Eds)
Objectives: We compared the effects of two semantic terminology models on classification of clinical notes through a study in the domain of heart murmur findings. Methods: One schema was established from the existing SNOMED CT model (S-Model) and the other was from a template model (T-Model) which uses base concepts and non-hierarchical relationships to characterize the murmurs. A corpus of clinical notes (n=309) was collected and annotated using the two schemas. The annotations were coded for a decision tree classifier for text classification task. The standard information retrieval measures of precision, recall, f-score and accuracy and the paired t-test were used for evaluation. Results: The performance of S-Model was better than the original T-Model (p<0.05 for recall and f-score). A revised T-Model by extending its structure and corresponding values performed better than S-Model (p<0.05 for recall and accuracy). Conclusion: We discovered that content coverage is a more important factor than terminology model for classification; however a templatestyle facilitates content gap discovery and completion. Introduction While modern terminologies have advanced well beyond simple one-dimensional subsumption relationships through the introduction of composite expressions, there is an emerging convergence of approaches toward the use of a concept-based clinical terminology with an underlying formal semantic terminology model (STM) [1]. SNOMED CT, the most comprehensive clinically oriented medical terminology system, currently adopts a foundation based on a description logic (DL) model and the underlying DL-based structure to formally represent the meanings of concepts and the interrelationships between concepts [2-3]. The existing SNOMED CT model is mainly pre-coordination oriented, i.e. containing many pre-coordinated terms, and also supports post-coordination. For example, a compositional expression “[ hypophysectomy (52699005) ] + [ transfrontal approach (65519007) ]” could be used to describe a more specific clinical statement than that only using the term “hypophysectomy (52699005)”. For a specific domain, a template model having a semantic structure with a coherent class of terms can be used as a formal representation [4]. This kind of model is mainly post-coordination oriented and a list of atomic terms is organized within a semantic structure. For example, the latest version of the International Classification of Nursing Practice (ICNP) uses a 7-Axis model to support the representation of nursing concepts and integrates the domain concepts of nursing in a manner suitable for computer processing [5]. One of the main goals of the semantic terminology models is to support capturing structured clinical information that is crucial for computer programs such as information retrieval systems and decision support tools [6]. Structured recording has the potential to improve information retrieval from a patient database in response to clinically relevant questions [1]. However, functional difference in retrieval performance has not been clearly demonstrated between these two different semantic terminology models. In this study, we focus upon the specific domain of heart murmur findings. Two schemas were established from two different semantic terminology models for evaluation: one schema is extracted from the existing SNOMED CT model (S-Model) and the other is a template model (T-Model) extracted from a concept-dependent attributes model recently published by Green, et al [7]. The objectives of the study are to annotate the real clinical notes using the two schemas and to compare and evaluate the effects of two models on classification of the clinical notes. Methods and Materials Defining the annotation schemas We defined two schemas for both S-Model and T-Model and represented the two schemas in Protege (version 3.2 beta), which is an ontology editing environment and was developed by Stanford Medical Informatics [8]. For the S-Model, we established a schema by extracting concept trees from the existing sub-hierarchy of heart murmur findings in January 2006 version of SNOMED CT (see Fig. 1). One root concept is “Heart murmur (SCTID_88610006)” which includes 86 sub-concepts of pre-coordinated terms of heart murmur findings. The other root concept is “Anatomical concepts (SCTID_257728006)” which includes two parts relevant to our schema. One part is the concept “Cardiac internal structure (SCTID_277712000)” and its sup-concepts. The other part contains only those anatomical concepts appearing in our clinical notes corpus on the basis of a manual review. For all heart murmur concepts, two semantic attributes derive from SNOMED CT context model for Representing and sharing knowledge using SNOMED Proceedings of the 3rd international conference on Knowledge Representation in Medicine (KR-MED 2008) R. Cornet, K.A. Spackman (Eds)
SNOMED CT is a complex ontology; sophisticated browsers are required to make it understandable and useful. We identified 23 SNOMED CT browsers that have been developed, and inspected 17. We enumerate and provide test criteria for a ‘master list’ of 143 browsing features supported by at least one inspected browser; future work will determine which of these features are implemented by individual browsers. Only 5 features were common to all 17 browsers; 89 were found in less than one third of browsers. We recommend that a core set of browsing features be defined and harmonized across browsers, particularly for text-to-concept search operations.
By constructing local extensions to SNOMED we aim to enrich existing medical and related data stores, simplify the expression of complex queries, and establish a foundation for semantic integration of data from multiple sources. Specifically, a local extension can be constructed from the controlled vocabulary(ies) used in the medical data. In combination with SNOMED, this local extension makes explicit the implicit semantics of the terms in the controlled vocabulary. By using SNOMED as a base ontology we can exploit the existing knowledge encoded in it and simplify the task of reifying the implicit semantics of the controlled vocabulary. Queries can now be formulated using the relationships encoded in the extended SNOMED rather than embedding them ad-hoc into the query itself. Additionally, SNOMED can then act as a common point of integration, providing a shared set of concepts for querying across multiple data sets. Key to practical construction of a local extension to SNOMED is appropriate tool support including the ability to compute subsumption relationships very quickly. Our implementation of the polynomial algorithm for EL+ in Java is able to classify SNOMED in under 1 minute.
General purpose terminology server software facilitates coordinated use of multiple standard medical terminologies for diverse healthcare applications. SNOMED CT is an important clinical reference terminology, whose size and scope make advanced terminology server capabilities particularly useful. Moreover, capabilities tied to SNOMED CT’s special features and requirements can result in substantial further benefits. Enhancements to a general purpose terminology server have been developed to facilitate the tailored creation, validation, organization, deployment, distribution, submission and maintenance of (postcoordinated) extensions to SNOMED CT.
We seek to leverage enhanced expressivity in OWL 1.1 via property chain axioms with right identities in order to organize and constrain anatomic concepts for use in clinical descriptions. Anatomic knowledge represented in SNOMED CT uses SEP triplets; we anticipate that property chains will allow a more parsimonious organization of anatomic concepts. However, these constructs may lead to unanticipated inference, especially when scaling to large numbers of concepts [1]. We used a bottom-up approach based on targeted use case questions to iteratively develop a “micro theory” that both identifies the sensible locations of fractures in long bones and also supports logic-based classification of fractures. Alternative representations of the statement “fractures occur in bone” were explored with the aim of creating rich clinical descriptors that support classification for inference and data mining. The process of creating this micro theory is discussed, where pragmatic decisions were made with an intention of both constraining data entry and enabling inferences within the scope of the use cases.
In this paper a description is presented in which the architectural, lexical and mapping differences are foregrounded between two compositional systems, both operating in the health care domain: LinkBase® and SNOMED. Based on these distinctive features, repercussions on NLP applications are exemplified and briefly discussed.
There has been major progress both in description logics and ontology design since SNOMED was originally developed. The emergence of the standard Web Ontology language in its latest revision, OWL 1.1 is leading to a rapid proliferation of tools. Combined with the increase in computing power in the past two decades, these developments mean that many of the restrictions that limited SNOMED's original formulation no longer need apply. We argue that many of the difficulties identified in SNOMED could be more easily dealt with using a more expressive language than that in which SNOMED was originally, and still is, formulated. The use of a more expressive language would bring major benefits including a uniform structure for context and negation. The result would be easier to use and would simplify developing software and formulating queries.
Auditing biomedical terminologies often results in the identification of inconsistencies and thus helps to improve their quality. In this paper, we present a method based on Semantic Web technologies for auditing biomedical terminologies and apply it to the NCI thesaurus. We stored the NCI thesaurus concepts and their properties in an RDF triple store. By querying this store, we assessed the consistency of both hierarchical and associative relations from the NCI thesaurus among themselves and with corresponding relations in the UMLS Semantic Network. We show that the consistency is better for associative relations than for hierarchical relations. Causes for inconsistency and benefits from using Semantic Web technologies for auditing purposes are discussed.
SNOMED CT (SCT) has been designed and implemented in an era when health computer systems generally required terminology representations in the form of singular precoordinated concepts. Consequently, much of SCT content represents pre-coordinated concepts and their relationships. In this conceptual paper the role of preand post-coordinated terminology expressions are considered in the context of the current development direction of Electronic Health Records and the use of communications and knowledge repositories. The move from current SCT structures to an implementation form of SCT that focuses on “atomic concepts” will support post-coordination and terminology binding to information models. This core or “essential” SNOMED CT called SNOMED Essential Terminology (S-ET) would be smaller in terms of core concept numbers, simpler, easier to maintain and more intuitive for implementers. Our proposed implementation form of SNOMED CT would contain only “atomic concepts” with their attendant hierarchies and relationship data. These would be supported by a strict model for representing current and future pre-coordinated concepts based on the use of an existing specific postcoordination expression, grammar, or representation. The resulting concept expressions would be postcoordinated from a smaller core of atomic components. Using definitional relationships, the proposed implementation form could equate existing pre-coordinated terms with postcoordinated representations, allowing SCT to maintain links with legacy data. A strategy for testing and implementing this approach is discussed and empirical research and feasibility testing is recommended. INTRODUCTION SNOMED CT (SCT) is becoming the international standard clinical terminology with a new international licensing and governance process which makes it widely accessible. The adoption of SCT by multiple countries was influenced by many published studies demonstrating its comprehensive coverage [1-4] and advanced structural features. SNOMED CT has antecedents in the College of American Pathologists family of terminologies, the UK National Health Service Read Codes. As with any living language, it has absorbed content from a number of other terminologies and classifications. SCT contains concepts and terms that describe the “language of use” as well as concepts which define the “language of meaning”[5-7]. Consequently, SCT contains many pre-coordinated concepts that have varying levels of semantic complexity alongside the component or essential concepts which are themselves the building blocks of these complex clinical expressions. While there are sound historical and ongoing pragmatic reasons for this evolutionary development, the resulting mix of concept structures makes implementation within various information models complex and prone to variation. Currently, SCT is “cluttered” with precoordinated terms that are incompletely defined by the internal information model that exists within SCT, making transformations between existing precoordinated terms and postcoordinated representations difficult to achieve. This result limits opportunities for interoperability across systems, [8] which is one of the key objectives of a controlled terminology. This conceptual paper brings to notice issues that Representing and sharing knowledge using SNOMED Proceedings of the 3rd international conference on Knowledge Representation in Medicine (KR-MED 2008) R. Cornet, K.A. Spackman (Eds)
This paper describes the findings of an exploratory study on reverse mapping of ICD-10-CA, the Canadian Adaptation, to SNOMED CT. For this study a set of 5,000 most frequent ICD-10-CA codes from the health ministry of a Canadian province was used. The methods included applying six mapping algorithms to each ICD-10-CA description to find the matching SNOMED CT concepts, and comparing the output against the UK SCT-ICD10 cross map for accuracy. Overall, we found successful SNOMED CT matches for ~63% of the ICD-10-CA codes. Issues requiring further attention include ways to increase successful matches and independent validation of mapping output. This study provides a glimpse of the methods that could lead to a SNOMED CT to ICD10-CA cross map. It should be of interest to those responsible for secondary use of discharge abstracts in epidemiological and statistical reporting.
This paper describes a methodology for encoding problem lists used in general practice with SNOMED CT. Our intent is to help general practitioners to incorporate SNOMED CT into their existing Electronic Medical Record (EMR) systems with minimal disruption as a first step, thus allowing them to assess its impact prior to full-scale conversion. We started with 1,713 original unique terms that made up the problem lists from the general practice EMR used in the study. We ended with 1,468 unique concepts after two cycles of matching and revisions that led to 1,347 or ~92% successful matches. The remaining terms were revised to tease out modifiers or secondary concepts that could be used to provide equivalency through post-coordination. While skeptics of reference terminology systems often balk at their unwieldy size and complexity for local adoption, this study has demonstrated that, using our methodology, it is possible to create a manageable subset of SNOMED concepts for problem lists used in general practice with immediate tangible value.
SNOMED CT medical vocabulary can be used to identify complementary features in a database. This functionality is used to develop a natural language processor (NLP) for PAIRS (Physician Assistant Artificial Intelligence Reference System). Although about 99% of concepts in PAIRS are present in SNOMED CT some features missing in it makes it unacceptable for any diagnostic decision support system (DDSS). Here we show that implementation of another NLP along with SNOMED CT makes it practically useful.
Biomedical domain ontologies could be better put to use for automatic semantic linguistic processing if we could map them to lexical resources that model the linguistic phenomena encountered in this domain, e.g., complex noun phrase structures that reference specific biological entity names and processes. In this paper, we introduce BioFrameNet a domain-specific FrameNet extension. BioFrameNet uses Frame semantics to express the meaning of natural language, is augmented with domain-specific semantic relations, and links to biomedical ontologies like the Gene Ontology all of which are expressed in the Description Logic (DL) variant of OWL. Thus, BioFrameNet annotations of natural-language text precisely map to biomedical ontologies, which in turn facilitates inference using DL reasoners.