A major goal of the BRAIN Initiative Cell Atlas Network (BICAN) is to create a suite of foundational reference cell atlases and associated standards for human and non-human primate brains. Central to this goal is the creation of cross-species harmonized cellular taxonomies and structural parcellations with formal ontologies that can be mapped into 3D reference frameworks bridging neuroimaging and cellular and histological resolutions. We describe here an iterative approach, focused initially on the basal ganglia, to co-create structural and cellular ontologies in human, macaque and marmoset brains, including a Harmonized Ontology of Mammalian Brain Anatomy (HOMBA), and to map and refine structural parcellations into neuroimaging-based common coordinate frameworks. These references provide the framework for documenting and mapping all experimental sampling in BICAN, allowing analyses of cellular and molecular variation as a function of topographic position, and enabling comparisons of cellular, molecular and neuroimaging-based functional variation within and between primate species.
The rapid growth of scientific publications and evolving experimental paradigms create significant challenges in staying up-to-date with current advances. Assertions are often unstructured and have limited provenance, which hinders reproducibility. Ontologies and knowledge graphs (KGs) offer structured solutions by capturing assertions, evidence, and provenance to support reproducibility. This paper reviews 23 ontologies - 13 focused on assertions and evidence and 10 on provenance - providing an overview of the current landscape while highlighting key challenges and opportunities for improvement.
The ability to extract structured information from unstructured sources-such as free-text documents and scientific literature-is critical for accelerating scientific discovery and knowledge synthesis. Large Language Models (LLMs) have demonstrated remarkable capabilities in various natural language processing tasks, including structured information extraction. However, their effectiveness often diminishes in specialized, domain-specific contexts that require nuanced understanding and expert-level domain knowledge. In addition, existing LLM-based approaches frequently exhibit poor transferability across tasks and domains, limiting their scalability and adaptability. To address these challenges, we introduce StructSense, a modular, task-agnostic, open-source framework for structured information extraction built on LLMs. StructSense is guided by domain-specific symbolic knowledge encoded in ontologies, enabling it to navigate complex domain content more effectively. It further incorporates agentic capabilities through self-evaluative judges that form a feedback loop for iterative refinement, and includes human-in-the-loop mechanisms to ensure quality and validation. We demonstrate that StructSense can overcome both the limitations of domain sensitivity and the lack of cross-task generalizability, as shown through its application to diverse neuroscience information extraction tasks.
Large-scale single-cell ‘omics profiling is being used to define a complete catalogue of brain cell types, something that traditional methods struggle with due to the diversity and complexity of the brain. But this poses a problem: How do we organise such a catalogue - providing a standard way to refer to the cell types discovered, linking their classification and properties to supporting data? Cell ontologies provide a partial solution to these problems, but no existing ontology schemas support the definition of cell types by direct reference to supporting data, classification of cell types using classifications derived directly from data, or links from cell types to marker sets along with confidence scores. Here we describe a generally applicable schema that solves these problems and its application in a semi-automated pipeline to build a data-linked extension to the Cell Ontology representing cell types in the Primary Motor Cortex of humans, mice and marmosets. The methods and resulting ontology are designed to be scalable and applicable to similar whole-brain atlases currently in preparation.
Characterizing cellular diversity at different levels of biological organization and across data modalities is a prerequisite to understanding the function of cell types in the brain. Classification of neurons is also essential to manipulate cell types in controlled ways and to understand their variation and vulnerability in brain disorders. The BRAIN Initiative Cell Census Network (BICCN) is an integrated network of data-generating centers, data archives, and data standards developers, with the goal of systematic multimodal brain cell type profiling and characterization. Emphasis of the BICCN is on the whole mouse brain with demonstration of prototype feasibility for human and nonhuman primate (NHP) brains. Here, we provide a guide to the cellular and spatial approaches employed by the BICCN, and to accessing and using these data and extensive resources, including the BRAIN Cell Data Center (BCDC), which serves to manage and integrate data across the ecosystem. We illustrate the power of the BICCN data ecosystem through vignettes highlighting several BICCN analysis and visualization tools. Finally, we present emerging standards that have been developed or adopted toward Findable, Accessible, Interoperable, and Reusable (FAIR) neuroscience. The combined BICCN ecosystem provides a comprehensive resource for the exploration and analysis of cell types in the brain.
Large scale single cell omics profiling is revolutionising our understanding of cell types, especially in complex organs like the brain. This presents both an opportunity and a challenge for cell ontologies. Annotation of cell types in single cell ‘omics data typically uses unstructured free text, making comparison and mapping of annotation between datasets challenging. Annotation with cell ontologies is key to overcoming this challenge, but this will require meeting the challenge of extending cell ontologies representing classically defined cell types by defining and classifying cell types directly from data. Here we present the Brain Data Standards Ontology (BDSO), a data driven ontology that is built as an extension to the Cell Ontology (CL). It supports two major use cases: cell type annotation, and navigation, search, and organisation of a web application integrating single cell omics datasets for the mammalian primary motor cortex. The ontology is built using a semi-automated pipeline that interlinks cell type taxonomies and necessary and sufficient marker genes, and imports relevant ontology modules derived from external ontologies. Overall, the BDS ontology provides an underlying structure that supports these use cases, while remaining sustainable and extensible through automation as our knowledge of brain cell type expands.
BACKGROUND:There have been relatively few attempts to represent vision or blindness ontologically. This is unsurprising as the related phenomena of sight and blindness are difficult to represent ontologically for a variety of reasons. Blindness has escaped ontological capture at least in part because: blindness or the employment of the term 'blindness' seems to vary from context to context, blindness can present in a myriad of types and degrees, and there is no precedent for representing complex phenomena such as blindness.METHODS:We explore current attempts to represent vision or blindness, and show how these attempts fail at representing subtypes of blindness (viz., color blindness, flash blindness, and inattentional blindness). We examine the results found through a review of current attempts and identify where they have failed.RESULTS:By analyzing our test cases of different types of blindness along with the strengths and weaknesses of previous attempts, we have identified the general features of blindness and vision. We propose an ontological solution to represent vision and blindness, which capitalizes on resources afforded to one who utilizes the Basic Formal Ontology as an upper-level ontology.CONCLUSIONS:The solution we propose here involves specifying the trigger conditions of a disposition as well as the processes that realize that disposition. Once these are specified we can characterize vision as a function that is realized by certain (in this case) biological processes under a range of triggering conditions. When the range of conditions under which the processes can be realized are reduced beyond a certain threshold, we are able to say that blindness is present. We characterize vision as a function that is realized as a seeing process and blindness as a reduction in the conditions under which the sight function is realized. This solution is desirable because it leverages current features of a major upper-level ontology, accurately captures the phenomenon of blindness, and can be implemented in many domain-specific ontologies.
We discuss the applicability of using the OBI assay paradigm for representing patient questionnaires, neuropsychological tests, and neurological exams, as well to annotate data generated from these assessments. We conclude that the specification for OBI 'assay' employs broad enough notion of evaluation to allow for these uses. However, it would be preferable to introduce subclasses of OBI 'planned process' or OBI 'assay' that explicitly addresses these types of use cases and provides clear groupings for general types of assays. Keywords—assay; OBI; questionnaire; neuropsychological test; neurological exam; clinical history I. BACKGROUND The Ontology for Biomedical Investigations (OBI) is an integrated ontology for the description of biological and clinical investigations (1). OBI is domain ontology that provides set of terms and relations to support precise annotation and querying of the data generated in biomedical investigations. It represents the design, types of analyses and assays performed, specifications, and data generated, resulting in classes such as 'assay', 'plan specification', and 'measurement datum'. OBI defines 'assay' as a planned process with the objective to produce information about the material entity that is the evaluant, by physically examining it or its proxies (2). All assays have specified output, an information content entity, which is about the evaluant. Examples of usage are: Assay the wavelength of light emitted by excited Neon atoms. Count of geese flying over house. Subclasses of OBI 'assay' include many laboratory-specific examples, such as 'sequencing assay' and 'metabolite profiling'. However, other types include 'performing clinical assessment', 'age measurement assay', and 'handedness assay'. Several projects are underway which seek to represent and annotate data generated from different types of forms, questionnaires, and tests. Each of these uses-cases broaden the application of OBI 'assay' in one or more ways. Neuropsychological tests are used to assess cognitive domains such as attention, visual-spatial ability, memory, executive function, and language comprehension and expression. In addition to representing the structure of these neuropsychological tests, it is crucial to capture the cognitive processes and functions that they evaluate as well as the data they produce. The neuropsychological Testing Ontology (NPT) utilizes OBI's assay paradigm to represent these tests (3). The handedness assay was used as starting point to model these tests. However, difficulties have been encountered in relating the assay process to the cognitive processes and functions being evaluated. Also, cognitive functions, such as short-term memory, cannot be the bearer of measureable qualities. The solution in NPT is to connect cognitive process to the function it realizes in the assay process using new relationship between data item and function. The Multiple Sclerosis Patient Data Ontology (MSPD) has been developed to represent both clinical measures and patient reported outcomes (PRO) associated with the New York State Multiple Sclerosis Consortium (NYSMSC) patient data registry (4). A PRO is generally considered to be an assessment of any aspect of patient's health status that comes directly from the patient and without any interpretation by clinician (5). The data registry uses standardized forms addressing demographic and clinical information, disease status and progression. It also includes data pertaining to patients' perception of their quality of life and wellbeing, which includes assessment of physical and psychosocial impairment. During the enrollment process patients are asked to rate their perception of their own functional abilities and affective states. A difficulty in using the assay framework has been in reconciling what qualifies as physical examination and subsequent evaluation. An output of survey in which patient is asked to make judgment about his or her perceived limitation in particular limb or visual acuity may indeed qualify in this case as sort of post-hoc physical exam which allows the evaluant to also be the evaluator. The OBI 'self- reported handedness assessment' supports the application of 'assay' to cases where patient self-evaluates outside the context of direct physical exam. However, it is less clear how questionnaires and forms that obtain basic demographic data fit within OBI's account of assays. A patient responding to questions such as date of birth, marital status, insurance provider, etc. pushes one to reconsider what is being evaluated, especially since no physical examination is involved.
We have developed the Multiple Sclerosis Patient Data Ontology (MSPD) to represent data from the patient data registry of the New York State Multiple Sclerosis Consortium (NYSMSC). MSPD is an application ontology that provides a set of classes for the annotation of both clinical measures and patient reported outcome data obtained from the enrollment forms used by the NYSMSC. To do so, we have adopted the paradigm established for representing assays in the Ontology for Biomedical Investigations. Our goal is to compare patient reported outcomes, such as self-reported disability and quality of life perceptions, to objective outcome measures in clinical practice, with reference to diagnoses and treatment modalities. We have begun an ontology-driven retrospective analysis of the patient records in the NYSMSC registry using an ontology term enrichment method in order to spot significant patterns in patient-reported and clinical outcomes in subsets of patients in the NYSMSC patient registry as compared to the NYSMSC patient population as a whole.
There have been relatively few attempts to represent sight or blindness ontologically. This is unsurprising as the related phenomena of sight and blindness are surprisingly difficult to represent ontologically for a variety of reasons. This paper discusses those reasons, explores the current attempts to represent sight or blindness, and how these attempts fail at representing certain types of blindness, viz., color blindness and flash blindness. We then explore a possible solution to representing sight and blindness ontologically. The solution capitalizes on the resources afforded to one who adopts the upper-level Basic Formal Ontology. Roughly, we characterize sight as a function and blindness as a reduction in the conditions under which the sight function is realized. Keywords—ontology; sight; blindness; function; disposition; color blindness; flash blindness; Basic Formal Ontology.
The Ocular Disease Ontology (ODO) is an ontology designed to represent ocular diseases for both clinical and research purposes. This constitutes an effort to build an ontology that represents, to the fullest possible extent, ocular diseases. ODO makes use of the following: Cell Ontology (CL), Ontology for General Medical Science (OGMS), Neuronal CL, Gene Ontology (GO), Uberanatomy Ontology (UBERON), and the Neurological Disease Ontology (ND). Much of the work on ODO focuses efforts on integrating recent work representing retinal cell types from CL and diseases from OGMS. Ocular diseases are generally classified in medical literature depending on the anatomical regions in which they originate (or, more accurately, where the disease has a material basis). The impetus behind the work in ODO rests on the hypothesis that ocular diseases are limited in physical location and physiological effect. Given the somewhat simple anatomical structures that comprise the eye and immediate surrounding area, it is possible to define most ocular diseases using current resources. As there have been no prior attempts to represent ocular disease comprehensively in a formal ontology, we believe that ODO represents a useful and straightforward extension of the OGMS based on these considerations. Long-term goals: Integration with ND, OGMS, and other existing ontologies under the BFO umbrella, complete interoperability, i.e., the ability of terms generated in ODO to function within other BFO ontologies, and coverage of all identified ocular diseases. The work so far: Under the umbrella class ‘ocular disease’ there are eight subclasses that correspond to eight major anatomical regions of the eye (lacrimal apparatus, cornea, lens, pupil, retina, sclera, uvea, vitreous) as well as two subclasses for diseases that escape characterization based on anatomical region (ocular albinism, optic neuropathy). There are currently over 100 terms in ODO (not including synonyms and relations). The motivation behind such methodology is to allow for classification of ocular diseases whose particular physiological mechanisms are as of yet unknown. Employment of current methods allows for classification of ocular diseases where current research has yet to provide a detailed description of disease course but general locational information is available. In this way, ODO terms can be used in the same way that researchers and practitioners currently use them, not merely in the ideal sense where one knows the particular disease course of an ocular disease. In addition, ODO leverages existing work in other ontologies such that ocular diseases are treated as one among many different diseases of the body. This guarantees that ODO will grow and change inasmuch and insofar as the ontologies related to ocular structures and diseases change in general. This interoperability is essential for functioning ontologies. Ideally, one would appeal merely to the anatomical region where the material basis of disease in located as much as possible for ease of representation. However, this approach will not work to differentiate all of the ocular diseases. For this reason, multiple differentia are employed to capture ocular diseases. For example, the two subclasses of ‘ocular albinism’ and ‘optic neuropathy’ are subclasses of ‘ocular disease’ that do not have a specific and predictable associated anatomical region. Future work: Continued development of ODO in conjunction with existing ontologies is slated for the near future. Ultimately, ODO will capture all known existing ocular diseases and remain flexible enough to accurately capture future ocular diseases. One immediate extension is classification of ocular diseases related to neurological diseases. Examination of specific use cases has shown promising results. For example, the class of ocular diseases known as retinal degenerative diseases (retinal degeneration) has provided valuable insight into the structuring of ODO.
Background We are developing the Neurological Disease Ontology (ND) to provide a framework to enable representation of aspects of neurological diseases that are relevant to their treatment and study. ND is a representational tool that addresses the need for unambiguous annotation, storage, and retrieval of data associated with the treatment and study of neurological diseases. ND is being developed in compliance with the Open Biomedical Ontology Foundry principles and builds upon the paradigm established by the Ontology for General Medical Science (OGMS) for the representation of entities in the domain of disease and medical practice. Initial applications of ND will include the annotation and analysis of large data sets and patient records for Alzheimer’s disease, multiple sclerosis, and stroke. Description ND is implemented in OWL 2 and currently has more than 450 terms that refer to and describe various aspects of neurological diseases. ND directly imports the development version of OGMS, which uses BFO 2. Term development in ND has primarily extended the OGMS terms ‘disease’, ‘diagnosis’, ‘disease course’, and ‘disorder’. We have imported and utilize over 700 classes from related ontology efforts including the Foundational Model of Anatomy, Ontology for Biomedical Investigations, and Protein Ontology. ND terms are annotated with ontology metadata such as a label (term name), term editors, textual definition, definition source, curation status, and alternative terms (synonyms). Many terms have logical definitions in addition to these annotations. Current development has focused on the establishment of the upper-level structure of the ND hierarchy, as well as on the representation of Alzheimer’s disease, multiple sclerosis, and stroke. The ontology is available as a version-controlled file at http://code.google.com/p/neurological-disease-ontology along with a discussion list and an issue tracker. Conclusion ND seeks to provide a formal foundation for the representation of clinical and research data pertaining to neurological diseases. ND will enable its users to connect data in a robust way with related data that is annotated using other terminologies and ontologies in the biomedical domain.