The UMLS Metathesaurus is a syntactically uniform, concept-based, semantically enhanced representation of many of the world's authoritative biomedical vocabularies. Released several times a year, the Metathesaurus is becoming a common, longitudinally maintained source of the current versions of these vocabularies. As vocabularies become standards for reimbursement, reporting, interoperation, and use by applications, the vocabulary obtained from the Metathesaurus must be consistent with that obtainable from each vocabulary's authority. Effective with the first 2004 release, the Metathesaurus represents new and updated sources "transparently"--both users and applications are able to "see" each vocabulary in the Metathesaurus without any of the small losses of information introduced by abstractions used in previous versions. Thus, the Metathesaurus can continue to provide its many semantic and lexical value-added features while guaranteeing that original sources will be "visible" in intact form. Vocabulary users and application developers will benefit from the enhancements and economies of scale offered by the Metathesaurus, while preserving distinctions between content provided by external authorities and content added as part of the Metathesaurus development and maintenance process.
Patient descriptors, or "problems," such as "brain metastases of melanoma" are an effective way for caregivers to describe patients. But most problems, e.g., "cubital tunnel syndrome" or "ulnar nerve compression," found in problem lists in an Electronic Medical Record (EMR) are not comparable computationally - in general, a computer cannot determine whether they describe the same or a related problem, or whether the user would have preferred "ulnar nerve compression syndrome." Metaphrase is a scalable, middleware component designed to be accessed from problem-manager applications in EMR systems. In response to caregivers' informal descriptors it suggests potentially equivalent, authoritative, and more formally comparable descriptors. Metaphrase contains a clinical subset of the 1997 UMLS Metathesaurus and some 10,000 "problems" from the Mayo Clinic and Harvard Beth Israel Hospital. Word and term completion, spelling correction, and semantic navigation, all combine to ease the burden of problem conceptualization, entry and formalization.
The nature of healthcare is changing, and with these changes come new information needs. Many Electronic Medical Record (EMR) systems ale appearing which require authoritative names for diseases, therapies, procedures, symptoms and indications. Driving this requirement for authoritative names in the EMR arE a variety of institutional needs and reimbursement imperatives. For example, only if every instance of a disease is correctly coded in the EMR will accurate outcomes analysis be possible.
m | | *l||~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~D5-81840 Hepaticmy IsA D6-94500Amyloidosis, DD-81274 Poisoning bystreptomycin
An architecture built from five software components -a Router, Parser, Matcher, Mapper, and Server -fulfills key requirements common to several point-of-care information and knowledge processing tasks. The requirements include problem-list creation, exploiting the contents of the Electronic Medical Record for the patient at hand, knowledge access, and support for semantic visualization and software agents. The components use the National Library of Medicine Unified Medical Language System to create and exploit lexical closure-a state in which terms, text and reference models are represented explicitly and consistently. Preliminary versions of the components are in use in an oncology knowledge server.
Although considerable effort has been put into creating extensive on-line reference resources for oncology, this compiled knowledge is underutilized in clinical situations. The Mobile Access To Oncology Knowledge (MATOK) project facilitates access to a variety of knowledge sources by providing a system designed to be used at the point-of-care. The system's key characteristics are mobility, homogeneous access, concept-based searching, step-wise refinement, and integration with on-line patient data.
Health care enterprises need enterprise-wide terminologies to compare, reuse and repurpose health care descriptions. But once they are created, these terminologies need to be maintained and enhanced to sustain their utility and that of the descriptions encoded with them. MEME II (Metathesaurus Enhancement and Maintenance Environment, Version II) supports the required activities and enables enterprises to leverage their investment in terminology and descriptions by permitting remote-extra-enterprise-enhancements to terminology to be incorporated locally, and local-intra-enterprise-enhancements to be shared remotely. MEME II represents all changes to terminologies as data, or "actions," that can be interpreted by an "action engine." These actions, or messages, represent semantic "units of work" that can be interpreted by other copies of MEME II. The exchange of update messages increases the likelihood that the comparability of terminology-based health care descriptions can be sustained.
Problem lists assist in organizing patient information in computer based medical records. However, in order to use problem lists for billing, research, decision support and standardization, a categorization of the problems entered is required. We describe the problem list component of our computerized patient record, the On-line Medical Record (OMR), which combines a free-text entry mechanism with a categorization scheme, using a dictionary containing 846 terms. All 118,040 problems entered during the system's six years of use have been analyzed, 477 clinicians have entered a mean +/- S.D. of 238 +/- 604 problems into 22,311 patient records. The average number of problems in each patient's file was 5.1 +/- 3.9. Comments were typed for 80,281 (68%) of the problems, ranging in length from 1 to 2456 characters, with a mean length of 98 +/- 110 characters. Half the problems were entered on the day of the encounter with the patient. Overall, 66% of all problems were categorized in relation to terms from the problem dictionary. Lexical analysis of all problem names showed that 80% could be mapped to Meta 1.4, Snomed 3.0 or a pre-release version of Read 3.0. We conclude that a problem list entry scheme combining free-text entry and optional categorization using a dictionary can result in a high proportion of problems being categorized as desired. Improvement of the system by elimination of unused dictionary terms and addition of 1000 terms identified by the lexical analysis is likely to result in even higher categorization rates.
The barrier word method of identifying nominal phrases in text, using a very long barrier word list, was evaluated in two different sets of text. In a sample of 10 paragraphs from the Medical Knowledge Self-Assessment Program of the American College of Physicians, the yield of nominal phrases as a percent of total chunks isolated was 66%. Some 500,000 chunks were isolated from Principles and Practice of Oncology (PPO). 38% of these chunk-occurrences were of chunks which matched to 10,000 concept names in Meta-1.4, the most recent version of the UMLS Metathesaurus. 50 paragraphs from PPO were chosen at random. Co-occurrences of concepts in those paragraphs were reviewed. 42 of the paragraphs had unique or infrequently occurring co-occurrences which described closely the major thrust of the paragraph.
A terminology is a systematic, authoritative collection of concept names, or terms, in some domain. No single terminology names all the important concepts in biomedicine. One approach to creating a more comprehensive biomedical terminology is to merge existing biomedical terminologies, as the UMLS( Metathesaurus( has done for the last six years. Because existing terminologies may overlap--for example, one terminology may name a concept also named by another terminologyQthe terminologies in the Metathesaurus must be merged. Some terminologies suggest merges through their structure or content e.g., they suggest synonyms or connections to other terminologies; other merges can be suggested by algorithm. Regardless, all merges in the Metathesaurus must be approved by a human editor with appropriate domain knowledge. By the time Meta-U96 is released early in 1996, one prototype and seven released versions of the Metathesaurus will have been produced by a sequence of four qualitatively different methods, named for the way in which they merge terms: #1 "Term Rewrite Rules, #2 "Transitive Closure on Facts," #3 "Fact-at-a-Time Concept Merging," and #4 "Action-at-a-Time Object Processing." The development of each method has been constrained by the annual Metathesaurus release schedule. The first two methods made the best use of limited computational resources, and the last two make better use of human editing resources.
The formality of the Metathesaurus stems from the way that its concepts are related to one another. Thus our first interest is in the relative stability of those concepts. In 1994, we observed that atoms, names in a constituent terminology, are being added to the Metathesaurus faster than concepts, unique, named meanings. Figure 1 shows the number of atoms and concepts in the reviewed portion of the six versions of the Metathesaurus, and the 1995 Metathesaurus continues the trend toward convergence. If this continues for future versions of the Metathesaurus, it may indicate the emergence of an empirical consensus on the identity of biomedical concepts independent of what they are called.
Oncologists' information needs arise at diverse times and settings. For example: "Is superior vena cava syndrome a medical emergency?" Our collaborative group is developing a system that supports an interface with combinations of spoken, gestural, and simulated three-dimensional manipulation to help an oncologist focus on the information need, not the system. The system requires a small amount of input from the oncologist, and then anticipates what information is pertinent to the patient at hand, based on the sources it has available. The system makes use of a "Knowledge Server" to find relevant information. The Knowledge Server uses selected data for the particular patient from a Computer-based Patient Record (CPR) to provide context for the information needs. The Knowledge Server leverages the Unified Medical Language System (UMLS) resources as well as relevant communications standards. A layered, interaction protocol is used to help manage the fulfillment of information needs. Each of the oncology knowledge sources is transformed into a uniform representation that utilizes both its formal schema (e.g., its table of contents) and its concepts and words indexed through the UMLS Metathesaurus. Our focus on the appropriate use of information from a CPR, and on anticipating oncologists' information needs, resulted from our study of several longitudinal patient scenarios. We believe that our use of scenario-based design techniques will help to ensure the system's success.
The Metathesaurus is a machine-created, human edited and enhanced synthesis of authoritative biomedical terminologies. Its formal properties permit it to be a) exploited by computers, and b) modified and enhanced without compromising that usage. If further constraints were imposed on the existence and identity of Metathesaurus relationships, i.e., if every Metathesaurus concept had a "genus" and a "differentia," then the Metathesaurus could be converted into an "Aristotelian Hierarchy." In this sense, a genus is a concept that classifies another concept, and a differentia is a concept that distinguishes the classified concept from all other concepts in the same class. Since, in principle, these constraints would make the Metathesaurus easier to leverage and maintain computationally, it is interesting to ask to what degree the maintenance and enhancement procedures now in place are producing a Metathesaurus that is also an "Aristotelian Hierarchy." Given a liberal interpretation of the current Metathesaurus schema, the proportion of the Metathesaurus that is "Aristotelian" in each annual version is increasing in spite of dramatic concurrent increases in the number of Metathesaurus concepts. Without formality there is no modifiability nor scalability. [1] We need formal methods and computer-based tools that can help us with the task [of controlled medical vocabulary construction]. We need research in which controlled vocabulary development is the focus rather than a stepping stone for work on other theories and applications. [2]
Meta-1.1, the UMLS metathesaurus, represents medical knowledge in the forms of names of concepts and links between those concepts. The representations of the semantic neighborhood of a concept can be thought of as dimensions of the property of semantic locality and include term information (broader, narrower, or otherwise related), the contextual information (parent-child, siblings in a hierarchy), the semantic types, and the co-occurrence data (links discovered empirically from concepts used to index the medical literature.) The degree of redundancy of each of these dimensions was investigated by reviewing the extent of multiple presentations of concepts which appear as related to a given concept. The degree of overlap was surprisingly small. While the co-occurrence data finds some of the links represented by other dimensions, those links are but minute fractions of the vast amount of co-occurrence derived links. Because parent-child relationships are often subsumptive (or categorical) in nature, it might be expected that siblings usually share the same semantic types. While true in the aggregate, the wide variance in percent of types shared may reflect the intended usages of the source vocabularies. Noun phrases were extracted from the definitions of 40 concepts in Meta-1 in order to assess systematically the coverage of important concepts by Meta-1, and to assess whether the links between these definitional concepts, which may have special value, and the concept being defined were indeed present. Out of 161 of these definitional concepts, 29 were not represented in Meta-1, and 37 of those represented in Meta-1 had no direct link to the concept they were defining.(ABSTRACT TRUNCATED AT 250 WORDS)