
Background: There are many challenges associated with ontology building, as the process often touches on many different subject areas; it needs knowledge of the problem domain, an understanding of the ontology formalism, software in use and, sometimes, an understanding of the philosophical background. In practice, it is very rare that an ontology can be completed by a single person, as they are unlikely to combine all of these skills. So people with these skills must collaborate. One solution to this is to use face-to-face meetings, but these can be expensive and time-consuming for teams that are not co-located. Remote collaboration is possible, of course, but one difficulty here is that domain specialists use a wide-variety of different "formalisms" to represent and share their data - by the far most common, however, is the "office file" either in the form of a word-processor document or a spreadsheet. Here we describe the development of an ontology of immunological cell types; this was initially developed by domain specialists using an Excel spreadsheet for collaboration. We have transformed this spreadsheet into an ontology using highly-programmatic and pattern-driven ontology development. Critically, the spreadsheet remains part of the source for the ontology; the domain specialists are free to update it, and changes will percolate to the end ontology. Results: We have developed a new ontology describing immunological cell lines built by instantiating ontology design patterns written programmatically, using values from a spreadsheet catalogue. Conclusions: This method employs a spreadsheet that was developed by domain experts. The spreadsheet is unconstrained in its usage and can be freely updated resulting in a new ontology. This provides a general methodology for ontology development using data generated by domain specialists.
Motivation: Data on measured abundances of small molecules from biomaterial is currently accumulating in the literature and in online repositories. Unless formal machine-readable evidence assertions for such metabolite identifications are provided, quality assessment based re-use will be sparse. Existing annotation schemes are not universally adopted, nor granular enough to be of practical use in evidence-based quality assessment. Results: We review existing evidence schemes for metabolite identifications of variant semantic expressivity and derive requirements for a 'compliance-optimized' yet traceable annotation model. We present a pattern-based, yet simple taxonomy of intuitive and self-explaining descriptors that allow to annotate metab-olomics assay results both in literature and data bases with evidence information on small molecule analytics gained via technologies such as mass spectrometry or NMR. We present example annotations for typical mass spectrometry molecule assignments and outline next steps for integration with existing ontologies and metabolomics data exchange formats. Availability: An initial draft and documentation of the metabolite identification evidence code ontology is available at
Kinetic models are increasingly relevant in medical research. In systems biology, more than 10 years of experience with the development of standards and tools to construct and analyse kinetic models exists. This has supported the sharing of kinetic models, increased their reuse, and thereby has helped to reproduce and validate scientific results. Given this expertise, it seems natural to consider the application and development of standards and tools to meet the requirements of medical scientists. In this paper, we discuss challenges and opportunities for standards and tools from systems biology in medical research, and we suggest criteria for the safe use of simulations. We conclude that standards, tools and infrastructure need to be extended to ensure the quality, reliability and safety required when working with medical and patient data. This will foster the adaptation of modelling in the clinic, providing tools for improved diagnosis, prognosis and therapy.
Background: Automatic identification of gene and protein names from biomedical publications can help curators and researchers to keep up with the findings published in the scientific literature. As of today, this is a challenging task related to information retrieval, and in the realm of Big Data Analytics. Objectives: To investigate the feasibility of using word embed-dings (i.e. distributed word representations) from Deep Learning algorithms together with terms from the Cardiovascular Disease Ontology (CVDO) as a step to identifying omics information encoded in the biomedical literature. Methods: Word embeddings were generated using the neural language models CBOW and Skip-gram with an input of more than 14 million PubMed citations (titles and abstracts) corresponding to articles published between 2000 and 2016. Then the abstracts of selected papers from the sysVASC systematic review were manually annotated with gene/protein names. We set up two experiments that used the word embeddings to produce term variants for gene/protein names: the first experiment used the terms manually annotated from the papers; the second experiment enriched/expanded the annotated terms using terms from the human-readable labels of key classes (gene/proteins) from the CVDO ontology. CVDO is formalised in the W3C Web Ontology Language (OWL) and contains 172,121 UniProt Knowledgebase protein classes related to human and 86,792 UniProtKB protein classes related to mouse. The hypothesis is that by enriching the original annotated terms, a better context is provided, and therefore, it is easier to obtain suitable (full and/or partial) term variants for gene/protein names from word embed-dings. Results: From the papers manually annotated, a list of 107 terms (gene/protein names) was acquired. As part of the word embeddings generated from CBOW and Skip-gram, a lexicon with more than 9 million terms was created. Using the cosine similarity metric, a list of the 12 top-ranked terms was generated from word embeddings for query terms present in the generated lexicon. Domain experts evaluated a total of 1968 pairs of terms and classified the retrieved terms as: TV (term variant); PTV (partial term variant); and NTV (non term variant, meaning none of the previous two categories). In experiment I, Skip-gram finds the double amount of (full and/or partial) term variants for gene/protein names as compared with CBOW. Using Skip-gram, the weighted Cohen's Kappa inter-annotator agreement for two domain experts was 0.80 for the first experiment and 0.74 for the second experiment. In the first experiment, suitable (full and/or partial) term variants were found …
Motivation: Medical terminology mapping is a long-standing challenge for projects requiring retrieval, querying and integra-tion of heterogeneous patient data. Current tools fail to fully uti-lise the richness of the underlying coding systems, and can be difficult to install and maintain. For example, National Library of Medicine’s UMLS provides a rich collection of terminology mapping, however, its search results are displayed in a simplistic general purpose interface that cannot easily be navigated and results filtered according to user’s preferences. Specifically, returned results cannot be visualised in a tree to show positions and relationships. BioPortal offers a large number of terminologies and ontologies, each of which can be viewed in a tree struc-ture, however it does not allow for multiple ontologies to be viewed and compared on a single page. Our work aims to ad-dress these issues and provide a simple and easy to use terminology mapping software. Results: MeTMapS was evaluated with academic and clinical research users. The users have tested the mapping between ICD10, Read CTV2, V3 in Hypertension. It was also tested on a list of clinical terms from the inclusion and exclusion criteria of the INFORM clinical trial protocol. Our initial evaluation produced positive results. Availability: We are currently in the process of updating the design based on some improvements suggested by the participants. MeTMapS is developed under Apache V2 license and is currently hosted at KCL for internal use and will shortly be opened to the public once the internal security concerns are resolved. In the meantime, the tool is available from the author upon request.
BACKGROUND:Biological databases store data about laboratory experiments, together with semantic annotations, in order to support data aggregation and retrieval. The exact meaning of such annotations in the context of a database record is often ambiguous. We address this problem by grounding implicit and explicit database content in a formal-ontological framework.METHODS:By using a typical extract from the databases UniProt and Ensembl, annotated with content from GO, PR, ChEBI and NCBI Taxonomy, we created four ontological models (in OWL), which generate explicit, distinct interpretations under the BioTopLite2 (BTL2) upper-level ontology. The first three models interpret database entries as individuals (IND), defined classes (SUBC), and classes with dispositions (DISP), respectively; the fourth model (HYBR) is a combination of SUBC and DISP. For the evaluation of these four models, we consider (i) database content retrieval, using ontologies as query vocabulary; (ii) information completeness; and, (iii) DL complexity and decidability. The models were tested under these criteria against four competency questions (CQs).RESULTS:IND does not raise any ontological claim, besides asserting the existence of sample individuals and relations among them. Modelling patterns have to be created for each type of annotation referent. SUBC is interpreted regarding maximally fine-grained defined subclasses under the classes referred to by the data. DISP attempts to extract truly ontological statements from the database records, claiming the existence of dispositions. HYBR is a hybrid of SUBC and DISP and is more parsimonious regarding expressiveness and query answering complexity. For each of the four models, the four CQs were submitted as DL queries. This shows the ability to retrieve individuals with IND, and classes in SUBC and HYBR. DISP does not retrieve anything because the axioms with disposition are embedded in General Class Inclusion (GCI) statements.CONCLUSION:Ambiguity of biological database content is addressed by a method that identifies implicit knowledge behind semantic annotations in biological databases and grounds it in an expressive upper-level ontology. The result is a seamless representation of database structure, content and annotations as OWL models.
Motivation: We live in the age of Big Data. Data are collected about everything which has a mode of existence; this can be objects , processes, pictures, verbal reports, and many other types of things. The final purpose of data is not to collect more data but to transform data into relevant applications. For this purpose, there is a need to transform data into knowledge which is the basis for a manifold of applications. The current situation of data overload is caused by a lack of methods for abstraction and interpretation of data, but also by an insufficient understanding of the relation between data and knowledge. The overall goal of our work, intended to be realized within a longstanding project, is to establish an ontological framework which may serve as a unifying theory of data and knowledge. We explore various philosophical sources, and ascertain whether they may contribute to the realization of this project. In the present paper we consider White-head's philosophy. Approach: We explore the philosophy of Whitehead, expounded in Process and Reality, with respect to its relation to a recently developed ontology of data called GFO-Data. Whitehead's Process and Reality provides a non-formal approach to the creation of data and knowledge. Results: Basic categories and relations of Whitehead's Process and Reality are analyzed and specified by axioms in FOL. We outline a representation of the informational character of a datum as a prehension. This approach needs to be completed in order to grasp the process of transforming data into knowledge in more detail.
Motivation: SNOMED CT provides about 300,000 codes with fine-grained definitions to support interoperability of health data. However, even experienced human coders tend to disagree about which codes to choose for expressing clinical content. Results: 20 short clinical text fragments were independently annotated with SNOMED CT codes by two terminology experts. We analysed each disagreement instance and classified disagreements into eight categories, for which representative examples are presented. Conclusion: For each disagreement category measures to improve the terminology and to support guidelines for human and machine annotation are proposed and discussed.
Motivation: The previously developed Search Ontology (SO) allows domain experts to formally specify domain concepts, search terms associated to a domain, and rules describing domain concepts. So far, Lucene search queries can be generated from information contained in the SO and can be used for querying literature data bases or PubMed. However, this is still insufficient, since these queries are not well suited for querying XML documents because they are not following their structure. However, in the medical domain, many information items are coded in XML. Thus, querying structured XML documents is crucial for retrieving similar cases or for identifying potential study participants. For example, information items of patients with a similar tumor classification documented in a certain section of the respective pathology report need to be retrieved. This requires a precise definition of queries. In this paper, we introduce a concept for the generation of such queries using a Search Ontology XML extension to enable semantic searches on structured data. Results: For a gain of precision, the paragraph of a document need to be specified, in which a specific information item expressed in a query is expected to appear. The Search Ontology XML Extension (SOX) connects search terms to certain sections in XML documents. The extension consists of a class which represents the XML structure and a relation between search terms and this XML structure. This enables an automatic generation of XPath expressions, which makes an efficient and precise search of structured pathology reports in XML databases possible. The combination of standardized Electronic Health Records with an ontology based query method promises a gain of precision, a high degree of interoperability and long term durability of both, XML documents and queries on XML documents. * Contact: skropf@imise.uni-leipzig.de auciteli@imise.uni-leipzig.de
BACKGROUND:Medical personnel in hospitals often works under great physical and mental strain. In medical decision-making, errors can never be completely ruled out. Several studies have shown that between 50 and 60% of adverse events could have been avoided through better organization, more attention or more effective security procedures. Critical situations especially arise during interdisciplinary collaboration and the use of complex medical technology, for example during surgical interventions and in perioperative settings (the period of time before, during and after surgical intervention).METHODS:In this paper, we present an ontology and an ontology-based software system, which can identify risks across medical processes and supports the avoidance of errors in particular in the perioperative setting. We developed a practicable definition of the risk notion, which is easily understandable by the medical staff and is usable for the software tools. Based on this definition, we developed a Risk Identification Ontology (RIO) and used it for the specification and the identification of perioperative risks.RESULTS:An agent system was developed, which gathers risk-relevant data during the whole perioperative treatment process from various sources and provides it for risk identification and analysis in a centralized fashion. The results of such an analysis are provided to the medical personnel in form of context-sensitive hints and alerts. For the identification of the ontologically specified risks, we developed an ontology-based software module, called Ontology-based Risk Detector (OntoRiDe).CONCLUSIONS:About 20 risks relating to cochlear implantation (CI) have already been implemented. Comprehensive testing has indicated the correctness of the data acquisition, risk identification and analysis components, as well as the web-based visualization of results.
Motivation: This paper extends Röhl & Jansen's (2011) model of dispositions by introducing a relation of parthood between dispositions to formalize multi-track dispositions. Results: We suggest axioms for parthood relations between dispositions and discuss possible applications to life sciences.
In this paper a natural language processing workflow to extract sequential activities from large collections of medical text documents is developed. A graph-based data structure is introduced to merge extracted sequences which contain similar activities in order to build a global graph on procedures which are described in documents on similar topics or tasks. The method describes an information extraction process which will, in the future, enrich or create knowledge bases for process models or activity sequences for the medical domain.
The ability to collect and interlink heterogeneous data and model collections is essential in Systems Biology. Effective data exchange and comparison requires sufficient data annotation. This is particularly apparent in Systems Biology, where data heterogeneity means that multiple community metadata standards are required for the annotation of a whole investigation, including data, models and protocols. Here we describe FAIRDOM (http://fair-dom.org/) strategy in the context of semantic data management in the openSEEK , a webbased resource for sharing and exchanging Systems Biology data and models.