
Rapid advances in genome sequencing and gene expression microarray technologies are providing unprecedented opportunities to identify specific genes involved in complex biological processes, such as development, signal transduction, and disease. The vast amount of data generated by these technologies has presented new challenges in bioinformatics. To help organize and interpret microarray data, new and efficient computational methods are needed to: (1) distinguish accurately between different biological or clinical categories (e.g., malignant vs. benign), and (2) identify specific genes that play a role in determining those categories. Here we present a novel and simple method that exhaustively scans microarray data for unambiguous gene expression patterns. Such patterns of data can be used as the basis for classification into biological or clinical categories. The method, termed the Characteristic Attribute Organization System (CAOS), is derived from fundamental precepts in systematic biology. In CAOS we define two types of characteristic attributes ('pure' and 'private') that may exist in gene expression microarray data. We also consider additional attributes ('compound') that are composed of expression states of more than one gene that are not characteristic on their own. CAOS was tested on three well-known cancer DNA microarray data sets for its ability to classify new microarray samples. We found CAOS to be a highly accurate and robust class prediction technique. In addition, CAOS identified specific genes, not emphasized in other analyses, that may be crucial to the biology of certain types of cancer. The success of CAOS in this study has significant implications for basic research and the future development of reliable methods for clinical diagnostic tools.
Agreement measures are used frequently in reliability studies that involve categorical data. Simple measures like observed agreement and specific agreement can reveal a good deal about the sample. Chance-corrected agreement in the form of the kappa statistic is used frequently based on its correspondence to an intraclass correlation coefficient and the ease of calculating it, but its magnitude depends on the tasks and categories in the experiment. It is helpful to separate the components of disagreement when the goal is to improve the reliability of an instrument or of the raters. Approaches based on modeling the decision making process can be helpful here, including tetrachoric correlation, polychoric correlation, latent trait models, and latent class models. Decision making models can also be used to better understand the behavior of different agreement metrics. For example, if the observed prevalence of responses in one of two available categories is low, then there is insufficient information in the sample to judge raters' ability to discriminate cases, and kappa may underestimate the true agreement and observed agreement may overestimate it.
Comparative genomics is a large-scale, holistic approach that compares two or more genomes to discover the similarities and differences between the genomes and to study the biology of the individual genomes. Comparative studies can be performed at different levels of the genomes to obtain multiple perspectives about the organisms. We discuss in detail the type of analyses that offer significant biological insights in the comparisons of (1) genome structure including overall genome statistics, repeats, genome rearrangement at both DNA and gene level, synteny, and breakpoints; (2) coding regions including gene content, protein content, orthologs, and paralogs; and (3) noncoding regions including the prediction of regulatory elements. We also briefly review the currently available computational tools in comparative genomics such as algorithms for genome-scale sequence alignment, gene identification, and nonhomology-based function prediction.
The activities of a care providers' team need to be coordinated within a process properly designed on the basis of available best practice medical knowledge. It requires a rethinking of the management of care processes within health care organizations. The current workflow technology seems to offer the most convenient solution to build such cooperative systems. However, some of its present weaknesses still require an intense research effort to find solutions allowing its exploitation in real medical practice. This paper presents an approach to design and build evidence-based careflow management systems, which can be viewed as components of a knowledge management infrastructure each health care organization should be provided with to increase its performance in delivering high quality care by efficiently exploiting the available knowledge resources. The post-stroke rehabilitation process has been taken as a challenging care problem to assess our methodology for designing and developing careflow management systems. Then a system was co-developed with a team of rehabilitation professionals who will be committed to use it in their daily work. The system's main goal is to deliver a full array of rehabilitation services provided by an interdisciplinary team. They are related to identify which patients are most likely to benefit from rehabilitation, manage a rehabilitation treatment plan, and monitor progress both during rehabilitation and after return to a community residence. A model of the rehabilitation process was derived from an international guideline and adapted to the local organization of work. It involves different organizational units, such as wards, rehabilitation units, clinical laboratories, and imaging services. Several organizational agents work within them and play one or more roles. Each role is defined by the goals' set that she/he must fulfill. Special effort has been given to the design and development of a knowledge-based system for managing exceptions, which may occur in daily medical work as any deviation from the normal flow of activities. It allows either avoiding or recovering automatically from expected exceptions. When they are not expected, organizational agents, with enough power to do that, are allowed to modify the scheduled flow of activities for an individual patient under the only constraint of justifying their decision. After an intensive testing in a research laboratory, the system is now in the process of being transferred in a real working setting with the full support of its future users.
We are presenting here a model for processing space-time image sequences and applying them to 3D echo-cardiography. The non-linear evolutionary equations filter the sequence with keeping space-time coherent structures. They have been developed using ideas of regularized Perona-Malik an-isotropic diffusion and geometrical diffusion of mean curvature flow type (Malladi-Sethian), combined with Galilean invariant movie multi-scale analysis of Alvarez et al. A discretization of space-time filtering equations by means of finite volume method is discussed in detail. Computational results in processing of 3D echo-cardiographic sequences obtained by rotational acquisition technique and by real-time 3D echo volumetrics acquisition technique are presented. Quantitative error estimation is also provided.
In this paper, we propose a methodology (in the form of a software package) for automatic extraction of the cancerous nuclei in lung pathological color images. We first segment the images using an unsupervised Hopfield artificial neural network classifier and we label the segmented image based on chromaticity features and histogram analysis of the RGB color space components of the raw image. Then, we fill the holes inside the extracted nuclei regions based on the maximum drawable circle algorithm. All corrected nuclei regions are then classified into normal and cancerous using diagnostic rules formulated with respect to the rules used by experimented pathologist. The proposed method provides quantitative results in diagnosing a lung pathological image set of 16 cases that are comparable to an expert's diagnosis.
Clinical guidelines are intended to improve the quality and cost effectiveness of patient care. Integration of guidelines into electronic medical records and order-entry systems, in a way that enables delivery of patient-specific advice at the point of care, is likely to encourage guideline acceptance and effectiveness. Among the methodologies for modeling guidelines and medical decision rules, the Arden Syntax for Medical Logic Modules and the GuideLine Interchange Format version 3 (GLIF3) emphasize the importance of sharing encoded logic across different medical institutions and implementation platforms. These two methodologies have similarities and differences; in this paper we clarify their roles. Both methods can be used to support sharing of medical knowledge, but they do so in complementary situations. The Arden Syntax is suitable for representing individual decision rules in self-contained units called Medical Logic Modules (MLMs), which are usually implemented as event-driven alerts or reminders. In contrast, GLIF3 is designed for encoding complex multistep guidelines that unfold over time. As a consequence, GLIF3 has several mechanisms for complexity management and additional constructs that may require overhead unnecessary for expressing simple alerts and reminders. Unlike the Arden Syntax, GLIF3 encourages a top-down process of guideline modeling consisting of three levels that are created in order: Level 1 comprises a human-readable flowchart of clinical decisions and actions. Level 2 comprises a computable specification that can be verified for logical consistency and completeness; and Level 3 comprises an implementable specification that includes information required for local adaptation of guideline logic as well as for mapping guideline variables onto institutional medical records. A major emphasis of the current GLIF3 development process has been to create the computable specification that formally represents medical decision and eligibility criteria. We based GLIF3's formal expression language on the Arden Syntax's logic grammar, making the necessary extensions to the Arden Syntax's data structures and operators to support GLIF3's object-oriented data model. We discuss why the process of generating a set of MLMs from a GLIF-encoded guideline cannot be automated, why it can result in information loss, and why simple medical rules are best represented as individual MLMs. We thus show that the Arden Syntax and GLIF3 play complementary roles in representing medical knowledge for clinical decision support.
Information overload is a well-known problem for clinicians who must review large amounts of data in patient records. Concept-oriented views, which organize patient data around clinical concepts such as diagnostic strategies and therapeutic goals, may offer a solution to the problem of information overload. However, although concept-oriented views are desirable, they are difficult to create and maintain. We have developed a general-purpose, knowledge-based approach to the generation of concept-oriented views and have developed a system to test our approach. The system creates concept-oriented views through automated identification of relevant patient data. The knowledge in the system is represented by both a semantic network and rules. The key relevant data identification function is accomplished by a rule-based traversal of the semantic network. This paper focuses on the design and implementation of the system; an evaluation of the system is reported separately.
Linkage of epidemiological registries can provide cost-effective information on the associations between different diseases or exposures in the population under study and on completeness of surveillance system databases. We describe the program SALI (software for automated linkage in Italy) aimed at matching individual records from medium-sized registries (in the order of 100,000 records), where the desired outcome is to miss as few links as possible and, because of low link-likelihood (< 1%), a manual revision of matched pairs is feasible. SALI, developed in CA-Clipper language, uses registry files in dBase format. It requires only name, surname, and date of birth as key fields, and it allows for spelling errors in Italian or other Latin languages through a specific algorithm. Furthermore, a double-blind procedure ensures data confidentiality. The main linkage procedure is based on four stages, two automatic ones, and two where the operator can decide through specific windows whether to accept stage-selected matches. SALI takes into account possible errors in key fields thus reducing false negatives. It was used to solve the problem of linkage between AIDS and cancer registries in Italy. It can be used with every IBM-compatible computer system, assuring uniquely high portability.
Many algorithms have been used to cluster genes measured by microarray across a time series. Instead of clustering, our goal was to compare all pairs of genes to determine whether there was evidence of a phase shift between them. We describe a technique where gene expression is treated as a discrete time-invariant signal, allowing the use of digital signal-processing tools, including power spectral density, coherence, and transfer gain and phase shift. We used these on a public RNA expression set of 2467 genes measured every 7 min for 119 min and found 18 putative associations. Two of these were known in the biomedical literature and may have been missed using correlation coefficients. Digital signal processing tools can be embedded and enhance existing clustering algorithms.
In recent years shared decision making between patients and their health care providers and the inclusion of patient preferences in patient care have been, in theory, embraced as models for good clinical practice. Patients' experiences, values, and preferences are increasingly acknowledged as important pieces of evidence for appropriate health care decision making. To effectively use information about patient preferences in patient care, this information, which is gathered through a process of preference elicitation, needs to be integrated with other types of information, e.g., diagnoses, treatments, and patient status indicators within the context of a longitudinal electronic health record. This integration requires that patient preference-related concepts be represented nonambiguously and in a manner that renders them suitable for computer rather than human processing. In this article, the authors describe important patient preference-related concepts and illustrate the use of the LOINC semantic structure as a terminology model to create fully specified names for a sample of 15 preference elicitations from 8 published research articles.
Over the past decade there have been several attempts to rethink the basic strategies and scope of medical informatics. Meanwhile, bioinformatics has only recently experienced a similar debate about its scientific character. Both disciplines envision the development of novel diagnostic, therapeutic, and management tools, and products for patient care. A combination of the expertise of medical informatics in developing clinical applications and the focused principles that have guided bioinformatics could create a synergy between the two areas of application. Such interaction could have a great influence on future health research and the ultimate goal, namely continuity and individualization of health care. This article summarizes current activities related to facilitating synergy between medical informatics and bioinformatics, emphasizing activities in Europe while relating them to efforts in other parts of the world. The report provides examples of the analysis that European investigators are carrying out, aiming to propose new ideas for collaborations between medical informatics and bioinformatics researchers in a variety of areas.
Medical prognosis has played an increasing role in health care. Reliable prognostic models that are based on survival analysis techniques have been recently applied to a variety of domains, with varying degrees of success. In this article, we review some methods commonly used to model time-oriented data, such as Kaplan-Meier curves, Cox proportional hazards, and logistic regression, and discuss their applications in medical prognosis. Nonlinear, nonparametric models such as neural networks have increasingly been used for building prognostic models. We review their use in several medical domains and discuss different implementation strategies. Advantages and disadvantages of these methods are outlined, as well as pointers to pertinent literature.
Linkage analysis uses information from family pedigrees to map genes and locate disease genes on particular chromosomes. A recombination fraction denoted as θ is estimated as a measure of crossing over between two loci. Genetic linkage calculations are very time-consuming particularly for large family pedigrees, a large number of θ values, and an increased number of markers. This paper reports the implementation of a dynamic master-slave scheme for the parallelization of the Linkmap program on a high-performance cluster such as the Origin 2000 (O2K) consisting of 56 R12000 processors. The Linkmap program is one of four programs in the LINKAGE;shFASTLINK legacy package widely used by the medical research community. Implementations issues are addressed and results are compared with previous results on a cluster of DEC Alphas, and with the sequential execution on the O2K machine.
This paper will present new possibilities for the application of image recognition methods and AI application in biomedical informatics as well as semantically oriented analysis of 2D images of coronary arteries originating from coronography examinations. In particular this paper presents the possibilities for computer analysis and recognition of local stenoses of the lumen of coronary arteries via the application of syntactic methods of pattern recognition. Such stenoses are the result of the appearance of arteriosclerosis plaques, which in consequence lead to different forms of ischemic cardiovascular diseases. Such diseases may be seen in the form of stable or unstable disturbances of heart rhythm or infarction. Analysis of the correct morphology of these artery lumina is made possible with the application of syntactic analysis and pattern recognition methods, in particular with the attribute, context-free grammar of look-ahead LR(1) type.
With the growing use of Natural Language Processing (NLP) techniques for information extraction and concept indexing in the biomedical domain, a method that quickly and efficiently assigns the correct sense of an ambiguous biomedical term in a given context is needed concurrently. The current status of word sense disambiguation (WSD) in the biomedical domain is that handcrafted rules are used based on contextual material. The disadvantages of this approach are (i) generating WSD rules manually is a time-consuming and tedious task, (ii) maintenance of rule sets becomes increasingly difficult over time, and (iii) handcrafted rules are often incomplete and perform poorly in new domains comprised of specialized vocabularies and different genres of text. This paper presents a two-phase unsupervised method to build a WSD classifier for an ambiguous biomedical term W. The first phase automatically creates a sense-tagged corpus for W, and the second phase derives a classifier for W using the derived sense-tagged corpus as a training set. A formative experiment was performed, which demonstrated that classifiers trained on the derived sense-tagged corpora achieved an overall accuracy of about 97%, with greater than 90% accuracy for each individual ambiguous term.
The rapid expansion of biomedical knowledge, reduction in computing costs, and spread of internet access have created an ocean of electronic data. The decentralized nature of our scientific community and healthcare system, however, has resulted in a patchwork of diverse, or heterogeneous, database implementations, making access to and aggregation of data across databases very difficult. The database heterogeneity problem applies equally to clinical data describing individual patients and biological data characterizing our genome. Specifically, databases are highly heterogeneous with respect to the data models they employ, the data schemas they specify, the query languages they support, and the terminologies they recognize. Heterogeneous database systems attempt to unify disparate databases by providing uniform conceptual schemas that resolve representational heterogeneities, and by providing querying capabilities that aggregate and integrate distributed data. Research in this area has applied a variety of database and knowledge-based techniques, including semantic data modeling, ontology definition, query translation, query optimization, and terminology mapping. Existing systems have addressed heterogeneous database integration in the realms of molecular biology, hospital information systems, and application portability.
Clinical guidelines are being developed for the purpose of reducing medical errors and unjustified variations in medical practice, and for basing medical practice on evidence. Encoding guidelines in a computer-interpretable format and integrating them with the electronic medical record can enable delivery of patient-specific recommendations when and where needed. Since great effort must be expended in developing high-quality guidelines, and in making them computer-interpretable, it is highly desirable to be able to share computer-interpretable guidelines (CIGs) among institutions. Adoption of a common format for representing CIGs is one approach to sharing. Factors that need to be considered in creating a format for sharable CIGs include (i) the scope of guidelines and their intended applications, (ii) the method of delivery of the recommendations, and (iii) the environment, consisting of the practice setting and the information system in which the guidelines will be applied. Several investigators have proposed solutions that improve the sharability of CIGs and, more generally, of medical knowledge. These approaches can be useful in the development of a format for sharable CIGs. Challenges in sharing CIGs also include the need to extend the traditional framework for disseminating guidelines to enable them to be integrated into practice. These extensions include processes for (i) local adaptation of recommendations encoded in shared generic guidelines and (ii) integration of guidelines into the institutional information systems.
The paper explores the issues involved in maintaining the logic within a complex computer-based clinical guideline, using as a case study IMM/Serve, an operational guideline whose domain is childhood immunization. For a period of more than a year and a half, we have maintained a log of (1) the national changes to the immunization recommendations, (2) the local customizations of IMM/Serve's logic, and (3) certain logic problems that arose in the process of accommodating these changes and customizations. We describe the nature of these changes, customizations, and problems. We also discuss how different types of domain knowledge might assist in the automated process of validating successive versions of the logic. The paper's goal is to use the immunization domain to provide specific examples of the issues and problems that arise in maintaining a computer-based clinical guideline.