Single-cell RNA-Sequencing (scRNA-Seq) has undergone major technological advances in recent years, enabling the conception of various organism-level cell atlassing projects. With increasing numbers of datasets being deposited in public archives, there is a need to address the challenges of enabling the reproducibility of such data sets. Here, we describe guidelines for a minimum set of metadata to sufficiently describe scRNA-Seq experiments, ensuring reproducibility of data analyses.
Project Website: http://bioschemas.org/ Source Code: https://github.com/BioSchemas/bioschemas License: Creative Commons Attribution-ShareAlike License (version 3.0) Abstract Schema.org provides a way to add semantic markup to web pages. It describes ‘types’ of information, which then have ‘properties’. For example, ‘Event’ is a type that has properties like ‘startDate’, ‘endDate’ and ‘description’. Bioschemas aims to improve data interoperability in life sciences by encouraging people in life science to use schema.org markup. This structured information then makes it easier to discover, collate and analyse distributed data. Bioschemas reuses and extends schema.org in a number of ways: defining a minimum information model for the datatype being described using as few concepts as possible and only where necessary adding new properties, and the introduction of cardinalities and controlled vocabularies. The main outcome of Bioschemas is a collection of specifications that provide guidelines to facilitate a more consistent adoption of schema.org markup within the life sciences for the “Find” part of the FAIR (Findable, Accessible, Interoperable, Reusable) principles. In 2016 Bioschemas successfully piloted with training materials and events to enable the EU ELIXIR Research Infrastructure Training Portal (TeSS) to rapidly and simply harvest metadata from community sites. Encouraged by this in March 2017 we launched a 12 month project to pilot Bioschemas for data repositories and datasets. Specifically we are working on: General descriptions for datasets and data repositories Specific descriptions for prioritised datatypes: Samples, Human Beacons, Plant phenotypes and Protein annotations Facilitating discovery by registries and data aggregators, and by general search engines Facilitate tool development for annotation and validation of compliant resources All work is grounded on describing real data resources for real use cases: to this end large and small dataset are part of the project: Pfam, Interpro, PDBe, UniProt, BRENDA, EGA, COPaKB, and Gene3D. Data aggregators participating include: InterMine, BioSamples and OmicsDI. Registries include Identifiers.org, DataMed, Biosharing and the Beacon Network. Bioschemas operates as an open community initiative, sponsored by the EU ELIXIR Research Infrastructure and is supported by the NIH BD2K programme and Google.
The ability to collect and interlink heterogeneous data and model collections is essential in Systems Biology. Effective data exchange and comparison requires sufficient data annotation. This is particularly apparent in Systems Biology, where data heterogeneity means that multiple community metadata standards are required for the annotation of a whole investigation, including data, models and protocols. Here we describe FAIRDOM (http://fair-dom.org/) strategy in the context of semantic data management in the openSEEK , a webbased resource for sharing and exchanging Systems Biology data and models.
The Ontology for Biomedical Investigations (OBI) is an ontology that provides terms with precisely defined meanings to describe all aspects of how investigations in the biological and medical domains are conducted. OBI re-uses ontologies that provide a representation of biomedical knowledge from the Open Biological and Biomedical Ontologies (OBO) project and adds the ability to describe how this knowledge was derived. We here describe the state of OBI and several applications that are using it, such as adding semantic expressivity to existing databases, building data entry forms, and enabling interoperability between knowledge resources. OBI covers all phases of the investigation process, such as planning, execution and reporting. It represents information and material entities that participate in these processes, as well as roles and functions. Prior to OBI, it was not possible to use a single internally consistent resource that could be applied to multiple types of experiments for these applications. OBI has made this possible by creating terms for entities involved in biological and medical investigations and by importing parts of other biomedical ontologies such as GO, Chemical Entities of Biological Interest (ChEBI) and Phenotype Attribute and Trait Ontology (PATO) without altering their meaning. OBI is being used in a wide range of projects covering genomics, multi-omics, immunology, and catalogs of services. OBI has also spawned other ontologies (Information Artifact Ontology) and methods for importing parts of ontologies (Minimum information to reference an external ontology term (MIREOT)). The OBI project is an open cross-disciplinary collaborative effort, encompassing multiple research communities from around the globe. To date, OBI has created 2366 classes and 40 relations along with textual and formal definitions. The OBI Consortium maintains a web resource (http://obi-ontology.org) providing details on the people, policies, and issues being addressed in association with OBI. The current release of OBI is available at http://purl.obolibrary.org/obo/obi.owl.
ELIXIR explicitly supports the FAIR Principles – Findable, Accessible, Interoperable, Reusable – for its data, software, tools, events and training resources. “Finding” has the significant challenge of effective discovery and indexing of web-based resources across all ELIXIR information providers – this is an issue because there has been no agreement within ELIXIR about how to expose such resources in order to make them discoverable. One solution to this problem is to adopt Schema.org mark-up. Schema.org is a community initiative supported by four major search-engine providers: Google, Bing, Yahoo and Yandex. It provides a simple way to publish data in a standard format. If websites publishing life-science training materials, data, tools, profiles etc. were to use Schema.org mark-up, then their websites could be crawled, and the data could be indexed and exposed in searchable portals. However, this approach has challenges. Bioschemas is a newly formed community group in the life sciences to address these challenges, aiming to make the adoption of Schema.org part of a powerful way to discover and collect life-science information. It produces specifications, one for each information type (‘Training Course’, ‘Event’, 'Sample Data', 'Ontology Term', 'Tool', etc.). Each specification lists the Schema.org minimum properties expected, and the constraints for each property: for example, the ‘Events’ specification that the property ‘topic’ should be an EDAM ontology topic. The specifications are developed openly and are available on GitHub, with the support of existing communities of domain experts. Bioschemas also identifies the types or properties that are needed in the life sciences but not present in Schema.org, and works with the community to encourage the adoption of these types and properties into Schema.org.
BACKGROUND:Making forecasts about biodiversity and giving support to policy relies increasingly on large collections of data held electronically, and on substantial computational capability and capacity to analyse, model, simulate and predict using such data. However, the physically distributed nature of data resources and of expertise in advanced analytical tools creates many challenges for the modern scientist. Across the wider biological sciences, presenting such capabilities on the Internet (as "Web services") and using scientific workflow systems to compose them for particular tasks is a practical way to carry out robust "in silico" science. However, use of this approach in biodiversity science and ecology has thus far been quite limited.RESULTS:BioVeL is a virtual laboratory for data analysis and modelling in biodiversity science and ecology, freely accessible via the Internet. BioVeL includes functions for accessing and analysing data through curated Web services; for performing complex in silico analysis through exposure of R programs, workflows, and batch processing functions; for on-line collaboration through sharing of workflows and workflow runs; for experiment documentation through reproducibility and repeatability; and for computational support via seamless connections to supporting computing infrastructures. We developed and improved more than 60 Web services with significant potential in many different kinds of data analysis and modelling tasks. We composed reusable workflows using these Web services, also incorporating R programs. Deploying these tools into an easy-to-use and accessible 'virtual laboratory', free via the Internet, we applied the workflows in several diverse case studies. We opened the virtual laboratory for public use and through a programme of external engagement we actively encouraged scientists and third party application and tool developers to try out the services and contribute to the activity.CONCLUSIONS:Our work shows we can deliver an operational, scalable and flexible Internet-based virtual laboratory to meet new demands for data processing and analysis in biodiversity science and ecology. In particular, we have successfully integrated existing and popular tools and practices from different scientific disciplines to be used in biodiversity and ecological research.
Bioschemas is an open community initiative that aims to improve data interoperability in the life sciences. It does so by encouraging life scientists to use schema.org mark-up, so that their websites and services contain consistently structured information. This makes it easier to discover, collate and analyse distributed data. The main outcome of Bioschemas is a collection of specifications that provide guidelines to facilitate a more consistent adoption of schema.org mark-up within the life sciences. ELIXIR explicitly supports FAIR (Findable, Accessible, Interoperable, Reusable) principles for its data, software, tools, events and training resources. ‘Finding’ has the challenge of effective discovery and indexing of Web-based resources across ELIXIR information providers – this is hard because there has been no ELIXIR-wide agreement on how to expose such resources to make them more accessible. A consortium of search engines (Google, Bing, Yahoo, Yandex) have developed Schema.org, which provides a set of schemas for describing various Web resources. Schema.org adopters mark-up the contents of their site with hidden fields/attributes (encoded in either RDFa, Microdata or JSON-LD formats), which search engines and other aggregators can then parse and use to index the resources, thereby giving semantically rich search results and facilitating information discovery and access. Bioschemas is a joint initiative of several organisations and life-science communities to extend schema.org to facilitate the description, sharing and promotion of life-science resources and activities. There are currently groups developing several content types including Events, Training Materials, Organizations, Person profiles, Standards, and DATS. More standards are being proposed and developed. Bioschemas proposes amendments to existing schema.org specifications, re-using as many existing types as possible. The proposed specifications contain extra layers of information such as a minimum information model, recommended taxonomic vocabularies, and the cardinality of attributes to improve interoperability. Using Bioschemas to mark-up data improves search-engine optimisation. More than 30% of search results contain schema.org snippets, despite only 0.3% of websites containing this structured data. Marking-up data using Bioschemas schema.org is simple. Dozens of tools for popular Web frameworks and content-management systems are available to expedite the process. The development process is open and backed by a large consortium of search engines, so the longevity of the standard is likely to outlive an equivalent bespoke solution. Systems that aggregate or index metadata need providers to structure their data to make the process simpler. This is easy for large providers who can implement APIs or export their content in bespoke formats, but for the long tail - a large number of small websites containing content that would otherwise be over-looked – schema.org is an accessible and pragmatic way of structuring data for capture by content aggregators.
This report summarizes the proceedings of the 14th workshop of the Genomic Standards Consortium (GSC) held at the University of Oxford in September 2012. The primary goal of the workshop was to work towards the launch of the Genomic Observatories (GOs) Network under the GSC. For the first time, it brought together potential GOs sites, GSC members, and a range of interested partner organizations. It thus represented the first meeting of the GOs Network (GOs1). Key outcomes include the formation of a core group of “champions” ready to take the GOs Network forward, as well as the formation of working groups. The workshop also served as the first meeting of a wide range of participants in the Ocean Sampling Day (OSD) initiative, a first GOs action. Three projects with complementary interests - COST Action ES1103, MG4U and Micro B3 - organized joint sessions at the workshop. A two-day GSC Hackathon followed the main three days of meetings.
The co-authors of this paper hereby state their intention to work together to launch the Genomic Observatories Network (GOs Network) for which this document will serve as its Founding Charter. We define a Genomic Observatory as an ecosystem and/or site subject to long-term scientific research, including (but not limited to) the sustained study of genomic biodiversity from single-celled microbes to multicellular organisms.An international group of 64 scientists first published the call for a global network of Genomic Observatories in January 2012. The vision for such a network was expanded in a subsequent paper and developed over a series of meetings in Bremen (Germany), Shenzhen (China), Moorea (French Polynesia), Oxford (UK), Pacific Grove (California, USA), Washington (DC, USA), and London (UK). While this community-building process continues, here we express our mutual intent to establish the GOs Network formally, and to describe our shared vision for its future. The views expressed here are ours alone as individual scientists, and do not necessarily represent those of the institutions with which we are affiliated.
The study of biodiversity spans many disciplines and includes data pertaining to species distributions and abundances, genetic sequences, trait measurements, and ecological niches, complemented by information on collection and measurement protocols. A review of the current landscape of metadata standards and ontologies in biodiversity science suggests that existing standards such as the Darwin Core terminology are inadequate for describing biodiversity data in a semantically meaningful and computationally useful way. Existing ontologies, such as the Gene Ontology and others in the Open Biological and Biomedical Ontologies (OBO) Foundry library, provide a semantic structure but lack many of the necessary terms to describe biodiversity data in all its dimensions. In this paper, we describe the motivation for and ongoing development of a new Biological Collections Ontology, the Environment Ontology, and the Population and Community Ontology. These ontologies share the aim of improving data aggregation and integration across the biodiversity domain and can be used to describe physical samples and sampling processes (for example, collection, extraction, and preservation techniques), as well as biodiversity observations that involve no physical sampling. Together they encompass studies of: 1) individual organisms, including voucher specimens from ecological studies and museum specimens, 2) bulk or environmental samples (e.g., gut contents, soil, water) that include DNA, other molecules, and potentially many organisms, especially microbes, and 3) survey-based ecological observations. We discuss how these ontologies can be applied to biodiversity use cases that span genetic, organismal, and ecosystem levels of organization. We argue that if adopted as a standard and rigorously applied and enriched by the biodiversity community, these ontologies would significantly reduce barriers to data discovery, integration, and exchange among biodiversity resources and researchers.
The Genomic Standards Consortium (GSC) is an open-membership community that was founded in 2005 to work towards the development, implementation and harmonization of standards in the field of genomics. Starting with the defined task of establishing a minimal set of descriptions the GSC has evolved into an active standards-setting body that currently has 18 ongoing projects, with additional projects regularly proposed from within and outside the GSC. Here we describe our recently enacted policy for proposing new activities that are intended to be taken on by the GSC, along with the template for proposing such new activities.
As biological and biomedical research increasingly reference the environmental context of the biological entities under study, the need for formalisation and standardisation of environment descriptors is growing. The Environment Ontology (ENVO; http://www.environmentontology.org) is a community-led, open project which seeks to provide an ontology for specifying a wide range of environments relevant to multiple life science disciplines and, through an open participation model, to accommodate the terminological requirements of all those needing to annotate data using ontology classes. This paper summarises ENVO's motivation, content, structure, adoption, and governance approach. The ontology is available from http://purl.obolibrary.org/obo/envo.owl - an OBO format version is also available by switching the file suffix to "obo".
“If names be not correct, language is not in accordance with the truth of things. If language be not in accordance with the truth of things, affairs cannot be carried on to success.” - Confucius, Analects, Book XIII, Chapter 3, verses 4-7, translated by James Legge Two workshops (hereafter described as “workshops”) were held in 2012, which brought together domain experts from genomic and biodiversity informatics, information modeling and biology, to clarify concepts and terms at the intersection of these domains. These workshops grew out of efforts sponsored by the NSF funded Resource Coordination Network (RCN) project for GSC [1] (RCN4GSC, hosted at UCSD, with John Wooley as PI) to reconcile terms from the Darwin Core (DwC) [2] vocabulary and with those in the MIxS family of checklists (Minimum Information about Any Type of Sequence) [3]. The original RCN4GSC meetings were able to align many terms between DwC and MIxS, finding both common and complementary terms. However, deciding exactly what constitutes the concept of a sample, a specimen, and an occurrence [4] to satisfy the needs of all use cases proved difficult, especially given the wide variety of sampling strategies employed within and between communities. Further, participants in the initial RCN4GSC workshops needed additional guidance on how to relate these entities to processes that act upon them and the environments in which organisms live. These issues provided the motivation for the workshops described below. The two workshops drew largely from experiences of the Basic Formal Ontology (BFO) [5] and were led by Barry Smith, State University of New York at Buffalo. We chose to interact with Smith based on his successful interactions with the GSC in developing the Environment Ontology (EnvO) [6] and also, on the ability of BFO to unite previously disconnected ontologies in the medical domain [7]. The first workshop addressed term definitions in biodiversity informatics, working within the BFO framework, while the second workshop developed a prototype Bio-Collections Ontology, dealing with samples and processes acting on samples. Concurrent with these workshops were two ongoing efforts involving data acquisition, visualization, and analysis that rely on a solid conceptual understanding of samples, specimens, and occurrences. These implementations are included in this report to show practical applications of term clarification. Finally, this report provides a discussion of some of the next steps discussed during the workshops.