
Diagnostic and therapeutic decisions in the domain of medicine usually depend on observed or measured data, such as observations made during an examination, or laboratory data. Furthermore, such decisions are often based on abstract domain knowledge. For decision-making, it is necessary to establish relationships between concrete data and abstract knowledge. These relationships can be defined in a formal way by fuzzy set theory. We extended the Arden Syntax for Medical Logic Systems, which can be used for decision support in the medical domain, to include concepts of fuzzy set theory. This paper presents extensions for defining linguistic variables by Arden Syntax, which can formalize different states of an abstract concept, such as different levels of blood glucose. Arden Syntax linguistic variables can be used within conditional expressions in decision rules or within fuzzy control rules for computer-aided diagnosis and therapy.
There is growing recognition of the importance of the Internet and, more generally, information technology to pediatric care. However, acceptance of these technologies has been low. Attitudes of physicians can play a pivotal role in the adoption session. This study tests the extension to a widely used model in the information systems literature: the Technology Acceptance Model (TAM). Data were collected in a survey of pediatricians to see how well the extended model, TAM2, fits in the medical arena. Our results partially confirm the model; significant parts of the model were not confirmed. The primary factors in pediatricians' acceptance of technology applications relate to their usefulness and job relevance. Little weight is given to ease of use and social factors. We discuss possible explanations for the discrepancies and suggest future research.
Flow cytometric systems are being used increasingly in all branches of biological science including medicine. To develop analytic tools for identifying unknown molecules such as the antibodies that recognize different structure in the identical antigens, we explored use of a neural network in flow cytometry data comparison. Peak locations were extracted from flow cytometry histograms and we used the Marquardt backpropagation neural networks to recognize identical or similar binding patterns between antibodies and antigens based on the peak locations. The neural network showed 93.8% to 99.6% correct classification rates for identical or similar molecules. This suggests that the neural network technique can be useful in flow cytometry histogram data analysis.
Advances in optical character recognition (OCR) software and computer hardware have stimulated a reevaluation of the technology and its ability to capture structured clinical data from preexisting paper forms. In our pilot evaluation, we measured the accuracy and feasibility of capturing vitals data from a pediatric encounter form that has been in use for over twenty years. We found that the software had a digit recognition rate of 92.4% (95% confidence interval: 91.6 to 93.2) overall. More importantly, this system was approximately three times as fast as our existing method of data entry. These preliminary results suggest that with further refinements in the approach and additional development, we may be able to incorporate OCR as another method for capturing structured clinical data.
This paper reports the effects of a tailored Web-based delivery system on self-efficacy as it relates to a patients' response to acute myocardial information (AMI) symptoms. The data reported are from MI-HEART, a randomized trial examining ways in which a clinical information system can favorably influence the appropriateness and rapidity of decision-making in patients suffering from symptoms of acute myocardial infarction. Participants were randomized into one of three groups: tailored Web-based, non-tailored Web-based and non-tailored paper based. A theoretically based behavioral-cognitive model was used to identify key variables upon which to tailor education material. A key variable in the model is self-efficacy, operationalized with a three-dimensional scaling. Results show trends in improved self-efficacy scores for all groups at 1-month follow up, with sustained significant increases in baseline to 3-month scores only in the tailored Web-based group. One possible explanation could be related to "hit-count", which was significantly higher in the tailored group. This study is a first step in quantifying the contribution of Web-based tailoring over non-tailoring in changing key determinants of patient delay to AMI symptoms.
Several problems in medicine and biology involve the comparison of two measurements made on the same set of cases. The problem differs from a calibration problem because no gold standard can be identified. Testing the null hypothesis of no relationship using measures of association is not optimal since the measurements are made on the same cases, and therefore correlation coefficients will tend to be significant. The descriptive Bland-Altman method can be used in exploratory analysis of this problem, allowing the visualization of gross systematic differences between the two sets of measurements. We utilize the method on three sets of matched observations and demonstrate its usefulness in detecting systematic variations between two measurement technologies to assess gene expression.
The potential for gene discovery, fueled by DNA microchip technology and the sequencing of hundreds of genomes, is unprecedented. In this context, trying to discover genes that are actually of significance rather than merely appearing so due to noise is of utmost importance. We present a web application, CHIP TUNER, which assists in this gene discovery process. Our system uses evidence-based noise reduction to help delineate candidate target genes of biological importance. Specifically, CHIP TUNER learns from redundant experiments an "identity mask" that defines a region of noise inherent to biological sampling and DNA microarray processing; it then takes this into account during actual sample comparisons. The goal of CHIP TUNER is to improve the chances that newly discovered "important" genes are actually of importance before large amounts of time and resources are invested.
As the amount of data in public genomic databases grows, interoperability among them is becoming an increasingly critical feature. The ability for automated systems to mine and integrate data will be crucial to extracting knowledge from sources of data whose volume far exceeds the capabilities of human researchers. The currently dominant paradigm of presenting information as Web pages and using hyperlinks to describe relationships between pieces of information favors usability, but makes interoperability and automated data exchange more difficult. In this paper we describe how SNPper, a web-based system for the retrieval and analysis of Single Nucleotide Polymorphisms (SNPs), was augmented with a Remote Procedure Call interface, allowing client applications to query our program for SNP data and to receive the response as an XML document. Data represented in this form can be easily parsed by the requesting program, and thus reused for other applications. In this paper we describe the implementation of the interface and we show examples of its usage in a number of existing applications.
In general, it is very straightforward to store concept identifiers in electronic medical records and represent them in messages. Information models typically specify the fields that can contain coded entries. For each of these fields there may be additional constraints governing exactly which concept identifiers are applicable. However, because modern terminologies such as SNOMED CT are compositional, allowing concept expressions to be pre-coordinated within the terminology or post-coordinated within the medical record, there remains the potential to express a concept in more than one way. Often times, the various representations are similar, but not equivalent. This paper describes an approach for retrieving these pre- and post-coordinated concept expressions: (1) Create concept expressions using a logically-well-structured terminology (e.g., SNOMED CT) according to the rules of a well-specified information model (in this paper we use the HL7 RIM); (2) Transform pre- and post-coordinated concept expressions into a normalized form; (3) Transform queries into the same normalized form. The normalized instances can then be directly compared to the query. Several implementation considerations have been identified. Transformations into a normal form and execution of queries that require traversal of hierarchies need to be optimized. A detailed understanding of the information model and the terminology model are prerequisites. Queries based on the semantic properties of concepts are only as complete as the semantic information contained in the terminology model. Despite these considerations, the approach appears powerful and will continue to be refined.
The purpose of this study was to develop and test an interactive computer-mediated smoking cessation program for inner-city women. A non-probability sample of 100 women who receive care at an inner-city community health center in Indianapolis participated in the usability study. Women completed the computer program in the clinic following baseline data collection. Next, participants completed a brief satisfaction instrument. Data on cognitive and behavioral outcomes of the program were obtained by telephone interview one week later. Satisfaction with the program was high (mean satisfaction score was 60.2 with 70 indicating highest possible satisfaction). Average time for completing the computer program was 13.6 minutes. Overall, 79% of the participants reported at least one behavioral change related to smoking. The results indicate that interactive computer technology may be useful for promoting smoking cessation in low-income women.
The ability to access large amounts of de-identified clinical data would facilitate epidemiologic and retrospective research. Previously described de-identification methods require knowledge of natural language processing or have not been made available to the public. We take advantage of the fact that the vast majority of proper names in pathology reports occur in pairs. In rare cases where one proper name is by itself, it is preceded or followed by an affix that identifies it as a proper name (Mrs., Dr., PhD). We created a tool based on this observation using substitution methods that was easy to implement and was largely based on publicly available data sources. We compiled a Clinical and Common Usage Word (CCUW) list as well as a fairly comprehensive proper name list. Despite the large overlap between these two lists, we were able to refine our methods to achieve accuracy similar to previous attempts at de-identification. Our method found 98.7% of 231 proper names in the narrative sections of pathology reports. Three single proper names were missed out of 1001 pathology reports (0.3%, no first name/last name pairs). It is unlikely that identification could be implied from this information. We will continue to refine our methods, specifically working to improve the quality of our CCUW and proper name lists to obtain higher levels of accuracy.
Electronic medical record alerts and reminders are increasingly employed as a means of decreasing medical errors and increasing the quality and cost-effectiveness of care. However, clinicians indicate that alerts and reminders can be either help or hindrance. Discerning the elements that determine which they will be, and the requirements of a helpful alert or reminder, was the focus of this study. We convened three focus groups, comprised of a total of 16 participants. During analysis, five themes emerged: Efficiency, Usefulness, Information Content, User Interface, and Workflow. In addition there were some New Ideas and Surprises. Specific usability and usefulness requirements emerged from within the themes and these are described.
Extracting protein interaction relationships from textual repositories, such as MEDLINE, may prove useful in generating novel biological hypotheses. Using abstracts relevant to two known functionally related proteins, we modified an existing natural language processing tool to extract protein interaction terms. We were able to obtain functional information about two proteins, Amyloid Precursor Protein and Prion Protein, that have been implicated in the etiology of Alzheimer's Disease and Creutzfeldt-Jakob Disease, respectively.
Manually indexed Internet health catalogs such as CliniWeb or CISMeF provide resources for retrieving high-quality health information. Users of these quality-controlled subject gateways are most often referred to them by general search engines such as Google, AltaVista, etc. This raises several questions, among which the following: what is the relative visibility of medical Internet catalogs through search engines? This study addresses this issue by measuring and comparing the visibility of six major, MeSH-indexed health catalogs through four different search engines (AltaVista, Google, Lycos, Northern Light) in two languages (English and French). Over half a million queries were sent to the search engines; for most of these search engines, according to our measures at the time the queries were sent, the most visible catalog for English MeSH terms was CliniWeb and the most visible one for French MeSH terms was CISMeF.
XML has been widely adopted as an important data interchange language. The structure of XML enables sharing of data elements with variable degrees of nesting as long as the elements are grouped in a strict tree-like fashion. This requirement potentially restricts the usefulness of XML for marking up written text, which often includes features that do not properly nest within other features. We encountered this problem while marking up medical text with structured semantic information from a Natural Language Processor. Traditional approaches to this problem separate the structured information from the actual text mark up. This paper introduces an alternative solution, which tightly integrates the semantic structure with the text. The resulting XML markup preserves the linearity of the medical texts and can therefore be easily expanded with additional types of information.
We have analyzed a publicly available dataset consisting of gene-expression measurements from 105 lung carcinomas joined with clinical parameters describing the age, smoking history, and survival statistics for the patients that the tumors originated in. Our aim was to demonstrate how the unsupervised analysis technique embodied in PathlinX allows researchers to quickly gain an intuition for the most significant relationships between heterogeneous data elements. A variety of metrics were evaluated empirically by their ability to distinguish biological signal in the data from random noise; this was accomplished by random permutation of the data rows followed by comprehensive pair-wise comparison of all experimental elements. Thresholds of significance were established based on the metric scores for the permuted data. Sub-threshold associations were then removed. The remaining associations were then grouped by a transitive closure process to generate undirected graphs of associations called PathlinX networks. We discuss the various features of each generated PathlinX network and demonstrate the ability of the technique to highlight biological features in large heterogeneous datasets.
In contrast to existing computerized patient record systems, which merely offer static functionality for storage and presentation, a helpful patient record system is a problem-oriented, knowledge-based system which provides the clinician with situation-specific information from the patient record, relevant to the activity within the patient care process. We suggest extending the data model of current patient record systems with (1) knowledge for recognizing and interpreting care situations, (2) knowledge of how clinicians work and what information they need, and (3) means to rank information according to its relevance in a given situation. We present a framework that enables representation of three prerequisite features for a future helpful patient record system: the primary care workflow process, the problem-oriented information model, and means to identify relevant information to the care process and medical decisions.
Most guidelines developed for cancers screening and for cardiovascular risk management use rules to estimate familial risk. These rules are complex, difficult to memorize, and need to collect a complete pedigree. This paper describes a generic computerized method to estimate familial risks and its implementation in an internet-based application. The program is based on 3 generic models: a model of the family; a model of familial risk; a display model for the pedigree. The model of family allows to represent each member of the family and to construct and display a family tree. The model of familial risk is generic and allows easy update of the program with new diseases or new rules. It was possible to implement guidelines dealing with breast and colorectal cancer and cardiovascular diseases prevention. First evaluation with general practitioners showed that the program was usable. Impact on quality of familial risk estimate should be more documented.
Abstract As the home becomes an increasingly important site for health care, an increasing number of technology applications or devices are being introduced to support health at home. However, introducing new technology into a household raises a number of issues that must be considered prior to, during, and after the technology is implemented. This paper reviews the experiences of the UW-Madison Advanced Technologies for Health@Home Project, summarizing our assessment of household requirements that should be analyzed prior to introducing new technology. The overall goal of the Health@Home project is to improve the functionality and content of information technology innovations for the home. Using Venkatesh and Mazumdar's framework this article will summarize the relevant social, behavioral, technological, and physical dimensions of households that must be carefully assessed and understood to help ensure that the technology fits the needs of home residents.
The combination of a) our changing understanding of genotypic and phenotypic classification of diseases and b) the rapid growth and expansion of the number of entries in two databases targeted toward clinicians resulted in the need to develop a flexible dynamic hierarchical classification system for genetic disorders. The two databases making use of this classification schemas are the GeneClinics (GC) database - www.geneclinics.org and the GeneTests (GT) database - www.genetests.org The GC and GT databases serve respsectively as the users manual and yellow pages of genetic testing. The GeneTests/GeneClinics (GT/GC) classification hierarchy is maintained as a simple set of parent/child relationships in a relational database. The hierarchy is generated in real time in response to a user request. It is not maintained as a set of members with relationships defined by characters that are parsed to determine the structure of the tree. The GT/GC classification hierarchy entries are handled as objects by the data maintenance and search tools and may have a number of attributes and associations that create a rich tool for defining and examining genetic disorders