The Matchmaker Exchange application programming interface (API) allows searching a patient's genotypic or phenotypic profiles across clinical sites, for the purposes of cohort discovery and variant disease causal validation. This API can be used not only to search for matching patients, but also to match against public disease and model organism data. This public disease data enable matching known diseases and variant–phenotype associations using phenotype semantic similarity algorithms developed by the Monarch Initiative. The model data can provide additional evidence to aid diagnosis, suggest relevant models for disease mechanism and treatment exploration, and identify collaborators across the translational divide. The Monarch Initiative provides an implementation of this API for searching multiple integrated sources of data that contextualize the knowledge about any given patient or patient family into the greater biomedical knowledge landscape. While this corpus of data can aid diagnosis, it is also the beginning of research to improve understanding of rare human diseases.
Event Abstract Back to Event SciCrunch: A cooperative and collaborative data and resource discovery platform for scientific communities Jeffrey S. Grethe1*, Anita Bandrowski1, Davis E. Banks1, Christopher Condit2, Amarnath Gupta2, Stephen D. Larson1, Yueling Li1, Ibrahim B. Ozyurt1, Andrea M. Stagg1, Patricia L. Whetzel1, Luis Marenco3, Perry Miller3, Rixin Wang3, Gordon M. Shepherd4 and Maryann E. Martone1 1 University of California, San Diego, Center for Research in Biological Systems School of Medicine, United States 2 University of California, San Diego, San Diego Supercomputer Center, United States 3 Yale University, Center for Medical Informatics, United States 4 Yale University, Department of Neurobiology, United States Introduction SciCrunch was designed to help communities of researchers create their own portals to provide access to resources, databases and tools of relevance to their research areas. A data portal that searches across hundreds of databases can be created in minutes. Communities can choose from our existing SciCrunch data sources and also add their own. SciCrunch was designed to break down the traditional types of portal silos created by different communities, so that communities can take advantage of work done by others and share their expertise as well. When a community brings in a data source, it becomes available to other communities, thus ensuring that valuable resources are shared by other communities who might need them. At the same time, individual communities can customize the way that these resources are presented to their constituents, to ensure that their user base is served. To ensure proper credit and to help share expertise, all resources are tagged by the communities that create them and those that access them. Exploring Data SciCrunch is one of the largest aggregations of scientific data and tools available on the Web. One can think of SciCrunch as a “PubMed” for tools and data. Just as you can search across all the biomedical literature through PubMed, regardless of journal, SciCrunch lets you search across hundreds of databases and millions of data records from a single interface. Such databases are considered part of the “hidden web” because their content is not easily accessed by search engines. SciCrunch enhances search with semantic technologies to ensure we bring you all the results. SciCrunch provides three primary searchable collections: • SciCrunch Registry – is a curated catalog of thousands of research resources (data, tools, materials, services, organizations, core facilities), focusing on freely-accessible resources available to the scientific community. Each research resource is categorized by resource type and given a unique identifier. • SciCrunch Data Federation – provides deep query across the contents of databases created and maintained by independent individuals and organizations. Each database is aligned to the SciCrunch semantic framework, to allow users to browse the contents of these databases quickly and efficiently. Users are then taken to the source database for further exploration. SciCrunch deploys a unique data ingestion platform that makes it easy for database providers to make their resources available to SciCrunch. Using this technology, SciCrunch currently makes available over 200 independent databases, comprising ~400 million data records. • SciCrunch Literature – provides a searchable index across literature via PubMed and full text articles from the Open Access literature. SciCrunch Communities SciCrunch currently supports a diverse collection of communities (Figure 1), each with their own data needs: • CINERGI – focuses on constructing a community inventory and knowledge base on geoscience information resources to meet the challenge of finding resources across disciplines, assessing their fitness for use in specific research scenarios, and providing tools for integrating and re-using data from multiple domains. The project team envisions a comprehensive system linking geoscience resources, users, publications, usage information, and cyberinfrastructure components. This system would serve geoscientists across all domains to efficiently use existing and emerging resources for productive and transformative research. • Monarch Initiative (http://monarchinitiative.org; Figure 2) – provides tools that will use semantics and statistical models to support navigation through multi-scale spatial and temporal phenotypes across in vivo and in vitro model systems in the context of genetic and genomic data. These tools will provide basic, clinical, and translational science researchers, informaticists, and medical professionals with an integrated interface and set of discovery tools to reveal the genetic basis of disease, facilitate hypothesis generation, and identify novel candidate drug targets. The goal of the system is to promote true translational research, connecting clinicians with model systems and researchers who might shed light on related phenotypes, assays, or models. • Neuroscience Information Framework (NIF) – is a biological search engine that allows students, educators, and researchers to navigate the Big Data landscape by searching the contents of data resources relevant to neuroscience - providing a platform that can be used to pull together information about the nervous system. Underlying the NIF system is the Neurolex knowledge base. Neurolex seeks to define the major concepts of neuroscience, e.g., brain regions, cell types, in a way that is understandable to a machine. • NIDDK Information Network (dkNET) – serves the needs of basic and clinical investigators by providing seamless access to large pools of data relevant to the mission of The National Institute of Diabetes, Digestive and Kidney Disease (NIDDK). The portal contains information about research resources such as antibodies, vectors and mouse strains, data, protocols, and literature. • Research Identification Initiative (RII) – aims to promote research resource identification, discovery, and reuse. The RII portal offers a central location for obtaining and exploring Research Resource Identifiers (RRIDs) - persistent and unique identifiers for referencing a research resource. A critical goal of the RII is the widespread adoption of RRIDs to cite resources in the biomedical literature. RRIDs use established community identifiers where they exist, and are cross-referenced in our system where more than one identifier exists for a single resource. Figure 1 Figure 2 Acknowledgements This work was partly supported by the NIH Neuroscience Blueprint under contract HHSN27120080035C and the National Institute of Diabetes and Digestive and Kidney Diseases under grant U24DK097771 and the National Institute of Aging under grant 1R03AG043018 Keywords: big data, big data integration, semantic data, Open Data, Data Federation, ontologies, Portal System Conference: Neuroinformatics 2014, Leiden, Netherlands, 25 Aug - 27 Aug, 2014. Presentation Type: Demo, to be considered for oral presentation Topic: Infrastructural and portal services Citation: Grethe JS, Bandrowski A, Banks DE, Condit C, Gupta A, Larson SD, Li Y, Ozyurt IB, Stagg AM, Whetzel PL, Marenco L, Miller P, Wang R, Shepherd GM and Martone ME (2014). SciCrunch: A cooperative and collaborative data and resource discovery platform for scientific communities. Front. Neuroinform. Conference Abstract: Neuroinformatics 2014. doi: 10.3389/conf.fninf.2014.18.00069 Copyright: The abstracts in this collection have not been subject to any Frontiers peer review or checks, and are not endorsed by Frontiers. They are made available through the Frontiers publishing platform as a service to conference organizers and presenters. The copyright in the individual abstracts is owned by the author of each abstract or his/her employer unless otherwise stated. Each abstract, as well as the collection of abstracts, are published under a Creative Commons CC-BY 4.0 (attribution) licence (https://creativecommons.org/licenses/by/4.0/) and may thus be reproduced, translated, adapted and be the subject of derivative works provided the authors and Frontiers are attributed. For Frontiers’ terms and conditions please see https://www.frontiersin.org/legal/terms-and-conditions. Received: 28 Apr 2014; Published Online: 04 Jun 2014. * Correspondence: Dr. Jeffrey S Grethe, University of California, San Diego, Center for Research in Biological Systems School of Medicine, La Jolla, CA, 92093-0446, United States, jgrethe@ucsd.edu Login Required This action requires you to be registered with Frontiers and logged in. To register or login click here. Abstract Info Abstract The Authors in Frontiers Jeffrey S Grethe Anita Bandrowski Davis E Banks Christopher Condit Amarnath Gupta Stephen D Larson Yueling Li Ibrahim B Ozyurt Andrea M Stagg Patricia L Whetzel Luis Marenco Perry Miller Rixin Wang Gordon M Shepherd Maryann E Martone Google Jeffrey S Grethe Anita Bandrowski Davis E Banks Christopher Condit Amarnath Gupta Stephen D Larson Yueling Li Ibrahim B Ozyurt Andrea M Stagg Patricia L Whetzel Luis Marenco Perry Miller Rixin Wang Gordon M Shepherd Maryann E Martone Google Scholar Jeffrey S Grethe Anita Bandrowski Davis E Banks Christopher Condit Amarnath Gupta Stephen D Larson Yueling Li Ibrahim B Ozyurt Andrea M Stagg Patricia L Whetzel Luis Marenco Perry Miller Rixin Wang Gordon M Shepherd Maryann E Martone PubMed Jeffrey S Grethe Anita Bandrowski Davis E Banks Christopher Condit Amarnath Gupta Stephen D Larson Yueling Li Ibrahim B Ozyurt Andrea M Stagg Patricia L Whetzel Luis Marenco Perry Miller Rixin Wang Gordon M Shepherd Maryann E Martone Related Article in Frontiers Google Scholar PubMed Abstract Close Back to top Javascript is disabled. Please enable Javascript in your browser settings in order to see all the content on this page.
The NIF system is a semantic search engine that uses an ontology to improve search quality. In this experience paper we present SKEYQL, our semantic keyword query language and describe a number of ontology-based query reformulation strategies that go beyond standard query expansion techniques. We also present a set of lessons learnt and strategies that did not work. We reaffirm the importance of pre-annotating data to ensure quality query results.
Event Abstract Back to Event The Neuroscience Information Framework (NIF): A Unified Semantic Framework and Associated Tools for Discovery, Integration, and Utilization of Biomedical Data and Resources on the Web Jeffrey S. Grethe1*, Anita Bandrowski1, Davis Banks1, Jonathan Cachat1, Jing Chen1, Christopher Condit1, Amarnath Gupta2, Fahim Imam1, Stephen D. Larson1, Yueling Li1, Ibrahim B. Ozyurt1, Andrea M. Stagg1, Luis Marenco3, Perry Miller3, Rixin Wang3, Gordon M. Shepherd4, Giorgio Ascoli5 and Maryann E. Martone1 1 University of California, San Diego, Center for Research in Biological Systems, School of Medicine, United States 2 San Diego Supercomputer Center, United States 3 Yale University School of Medicine, Center for Medical Informatics, United States 4 Yale University, Department of Neurobiology, United States 5 George Mason University, Krasnow Institute for Advanced Study, United States The Neuroscience Information Framework (NIF; http://neuinfo.org) has recently launched a completely re-designed discovery portal for finding and integrating neuroscience-relevant resources, data, and literature. The new portal provides users with new tools to visualize data content (e.g. through analytics that provide a landscape analysis of where data can be found for topics of interest) to more personalized services via myNIF (e.g. the saving of favorite searches). The portal searches across 3 primary collections: (1) NIF Registry: A human-curated registry of neuroscience-relevant resources annotated with the NIF vocabulary; (2) NIF Literature: A full text indexed corpus derived from the PubMed Open Access subset as well as an entire index of PubMed; (3) NIF Database Federation: A federation of independent databases that enables discovery and access to public research data, contained in databases and structured web resources (e.g. queryable web services) that are sometimes referred to as the deep or hidden web. To further enable the utilization of this vast collection of information, the NIF is applying semantic web technologies to its holdings. By defining a set of standards and best practices for describing and representing data such semantic web and linked data technologies eliminate the barriers between database silos and foster the evolution of the Web into a Web of data. Such technologies are being successfully applied as integration engines for linking biological elements in many domains. Exposing NIF’s content as Linked Open Data will enable further integration with the growing amount of information available from the linked open data cloud – thereby providing much richer resources for the neuroscientist. The publication of this content relies on NIF’s comprehensive ontology (NIFSTD) that covers major domains in neuroscience, including diseases, brain anatomy, cell types, subcellular anatomy, small molecules, techniques and resource descriptors. Over the past year, NIF has continued to grow significantly in content, providing access to over 6,000 resources through the Registry, and more than 200 independent data resources in the data federation, making NIF the largest source of biomedical information on the web. NIF’s tools help people find and utilize neuroscience related resources - provides a consistent and easy to implement framework for those who are providing such resources, e.g., data, and those looking to utilize these data and resources. In this demonstration we will provide a tour of NIF’s suite of services, tools, and data: * Search through NIF’s newly re-designed semantically-enhanced discovery portal * Working with NIF’s linked data (e.g. nervous system connectivity) via SPARQL * Services and tools that provide access to the NIF data federation - the largest collection of Neuroscience relevant information on the web * Contributing to the NeuroLex – a community resource for neuroscience terminology built on a semantic media-wiki platform * Curation and normalization of data utilizing NIF’s Google Refine services * NIF’s semantically enhanced data and tools for its maintenance * myNIF and the NIF Digest – personalized services for researchers Figure 1 Acknowledgements NIH Neuroscience Blueprint HHSN271200800035C via NIDA Keywords: Data Federation, big data, linked data, Search Engine, Registry, data integration, Neuroscience, ontology Conference: Neuroinformatics 2013, Stockholm, Sweden, 27 Aug - 29 Aug, 2013. Presentation Type: Demo Topic: Infrastructural and portal services Citation: Grethe JS, Bandrowski A, Banks D, Cachat J, Chen J, Condit C, Gupta A, Imam F, Larson SD, Li Y, Ozyurt IB, Stagg AM, Marenco L, Miller P, Wang R, Shepherd GM, Ascoli G and Martone ME (2013). The Neuroscience Information Framework (NIF): A Unified Semantic Framework and Associated Tools for Discovery, Integration, and Utilization of Biomedical Data and Resources on the Web. Front. Neuroinform. Conference Abstract: Neuroinformatics 2013. doi: 10.3389/conf.fninf.2013.09.00073 Copyright: The abstracts in this collection have not been subject to any Frontiers peer review or checks, and are not endorsed by Frontiers. They are made available through the Frontiers publishing platform as a service to conference organizers and presenters. The copyright in the individual abstracts is owned by the author of each abstract or his/her employer unless otherwise stated. Each abstract, as well as the collection of abstracts, are published under a Creative Commons CC-BY 4.0 (attribution) licence (https://creativecommons.org/licenses/by/4.0/) and may thus be reproduced, translated, adapted and be the subject of derivative works provided the authors and Frontiers are attributed. For Frontiers’ terms and conditions please see https://www.frontiersin.org/legal/terms-and-conditions. Received: 29 Apr 2013; Published Online: 11 Jul 2013. * Correspondence: Dr. Jeffrey S Grethe, University of California, San Diego, Center for Research in Biological Systems, School of Medicine, La Jolla, CA, 92093-0446, United States, jgrethe@ncmir.ucsd.edu Login Required This action requires you to be registered with Frontiers and logged in. To register or login click here. Abstract Info Abstract The Authors in Frontiers Jeffrey S Grethe Anita Bandrowski Davis Banks Jonathan Cachat Jing Chen Christopher Condit Amarnath Gupta Fahim Imam Stephen D Larson Yueling Li Ibrahim B Ozyurt Andrea M Stagg Luis Marenco Perry Miller Rixin Wang Gordon M Shepherd Giorgio Ascoli Maryann E Martone Google Jeffrey S Grethe Anita Bandrowski Davis Banks Jonathan Cachat Jing Chen Christopher Condit Amarnath Gupta Fahim Imam Stephen D Larson Yueling Li Ibrahim B Ozyurt Andrea M Stagg Luis Marenco Perry Miller Rixin Wang Gordon M Shepherd Giorgio Ascoli Maryann E Martone Google Scholar Jeffrey S Grethe Anita Bandrowski Davis Banks Jonathan Cachat Jing Chen Christopher Condit Amarnath Gupta Fahim Imam Stephen D Larson Yueling Li Ibrahim B Ozyurt Andrea M Stagg Luis Marenco Perry Miller Rixin Wang Gordon M Shepherd Giorgio Ascoli Maryann E Martone PubMed Jeffrey S Grethe Anita Bandrowski Davis Banks Jonathan Cachat Jing Chen Christopher Condit Amarnath Gupta Fahim Imam Stephen D Larson Yueling Li Ibrahim B Ozyurt Andrea M Stagg Luis Marenco Perry Miller Rixin Wang Gordon M Shepherd Giorgio Ascoli Maryann E Martone Related Article in Frontiers Google Scholar PubMed Abstract Close Back to top Javascript is disabled. Please enable Javascript in your browser settings in order to see all the content on this page.
In this short paper, we present early results from an ongoing research on creating a new graph-based representation from NLP analysis of scientific documents so that the graph can be utilized for answering structured queries on NL-processed data. We present a sketch of the data model and the query language to show how scientifically meaningful queries can be posed against this graph structure.
Collisions between ships and whales are an increasing concern for endangered large whale species. After an unusually high number of blue whales (Balaenoptera musculus) were fatally struck in 2007 off the coast of southern California, federal agencies implemented a voluntary conservation program to reduce the likelihood of ship-strikes in the region. This initiative involved seasonal advisory broadcasts requesting vessel operators to voluntarily slow to 10 knots or less when transiting a 75 nm stretch of designated shipping lanes. We monitored ship adherence with those speed advisories using Automatic Identification System data. Daily average speed of cargo and tanker ships and the average speed of individual ship transits before, during, and after the notices were statistically analyzed for changes related to the notices. Whereas a small number of individual ships (1%) traveled significantly slower during the requested periods, speeds were not at or below the recommended 10 knots, nor were daily average speeds reduced during the notices. Voluntary conservation measures are established in a variety of contexts, and may be preferable to regulatory action; in this case, a request to make voluntary changes appeared largely ineffective. Reducing collision risks for whales in this area will require consideration of the various factors that likely explain the lack of adherence when developing an alternative strategy.
This paper presents BIODB, an ontology-enhanced information system to manage heterogeneous data. An ontology-enhanced system is a system where ad hoc data is imported into the system by a user, annotated by the user to connect the data to an ontology or other data sources, and then all data connected through the ontology can be queried in a federated manner. The BIODB system enables multi-model data federation, i.e., it federate data that can be in different data models including, relational, XML and RDF, sequence data and so on. It uses an ontologically enhanced system catalog, an ontological data index, an association index to facilitate cross-model data mapping, and a new algorithm for ontology-assisted keyword queries with ranking. The paper describes these components in detail, and presents an evaluation of the architecture in the context of an actual application.
As increasing volumes and varieties of data are becoming available online, the challenges of accessing and using heterogeneous data resources are growing. We have developed a mediator-based data integration system called Cartel for biological oceanography data. A mediation approach is appropriate in cases where a single central warehouse is not desirable, such as when the needed data sources change frequently through time, or when there are advantages for holding heterogeneous data in their native formats. Through Cartel, data sources of a variety of types can be registered to the system, and users can query against simplified virtual schemas, without needing to know the underlying schema and computational capabilities of each data source. The system can operate on a variety of relational and geospatial data formats, and can perform joins between formats. We tested the performance of the Cartel mediator in two biological oceanography application areas, and found that the system was able to support the variety of data types needed in a typical ecology study, but that the response times were unacceptably slow when very large databases (i.e. Ocean Biogeographic Information System and the World Ocean Atlas) were used. Indexing and caching are currently being added to the system to improve response times. The mediator is an open-source product, and was developed to be a generic, extensible component available to projects developing oceanography data systems.
Annotation is the process of supplementing data with additional information that was not part of the actual observation, but reflects post-facto comments and associations made by a user who analyzes the data. While annotation management systems are emerging in the field of relational data, such systems for scientific applications, where there is a wide heterogeneity in the types of annotable data, are almost nonexistent. In this demonstration paper, we describe Graphitti , a tool that (i) allows a user to annotate a wide variety of scientific data, and (ii) allows a user to query data and their annotations in a seamless manner.
The overarching goal of the NIF (Neuroscience Information Framework) project is to be a one-stop-shop for Neuroscience. This paper provides a technical overview of how the system is designed. The technical goal of the first version of the NIF system was to develop an information system that a neuroscientist can use to locate relevant information from a wide variety of information sources by simple keyword queries. Although the user would provide only keywords to retrieve information, the NIF system is designed to treat them as concepts whose meanings are interpreted by the system. Thus, a search for term should find a record containing synonyms of the term. The system is targeted to find information from web pages, publications, databases, web sites built upon databases, XML documents and any other modality in which such information may be published. We have designed a system to achieve this functionality. A central element in the system is an ontology called NIFSTD (for NIF Standard) constructed by amalgamating a number of known and newly developed ontologies. NIFSTD is used by our ontology management module, called OntoQuest to perform ontology-based search over data sources. The NIF architecture currently provides three different mechanisms for searching heterogeneous data sources including relational databases, web sites, XML documents and full text of publications. Version 1.0 of the NIF system is currently in beta test and may be accessed through http://nif.nih.gov.
The complexity of the nervous system requires high-resolution microscopy to resolve the detailed 3D structure of nerve cells and supracellular domains. The analysis of such imaging data to extract cellular surfaces and cell components often requires the combination of expert human knowledge with carefully engineered software tools. In an effort to make better tools to assist humans in this endeavor, create a more accessible and permanent record of their data, and to aid the process of constructing complex and detailed computational models, we have created a core of formalized knowledge about the structure of the nervous system and have integrated that core into several software applications. In this paper, we describe the structure and content of a formal ontology whose scope is the subcellular anatomy of the nervous system (SAO), covering nerve cells, their parts, and interactions between these parts. Many applications of this ontology to image annotation, content-based retrieval of structural data, and integration of shared data across scales and researchers are also described.
A fishing lure containing one or more light sources includes a guideway along which an electrical contact moves back and forth in response to an oscillatory movement of the lure. A series of spaced-apart stationary electrical contacts are positioned along the guideway to be successively engaged by the movable contact to intermittently complete a circuit and energize the light sources. The light sources are internally mounted for protection by the body of the lure and the light is transmitted to exterior locations by optical conductors.
We present the semantic data model for an ontological database for subcellular anatomy for Neurosciences. The data model builds upon the foundations of OWL and the Basic Formal Ontology, but extends them to include novel constructs that address several unresolved challenges encountered by biologists in using ontological models in their databases. The model addresses the interplay between models of space and objects located in the space, objects that are defined by constrained spatial arrangements of other objects, the interactions among multiple transitive relationships over the same set of concepts and so on. We propose the notion of parametric relationships to denote different multiple ways of parcellating the same space. We also introduce the notion of phantom instances to address the mismatches between the ontological properties of a conceptual object and the actual recorded instance of that object in cases where the observed object is partially visible.