The Open Data Commons (ODC) for Traumatic Brain Injury (ODC-TBI) and Spinal Cord Injury (ODC-SCI) are secure online platforms for members of the neurotrauma community to access curated, publicly available domain-specific data, and manage, share, and publish datasets with a citable DOI. Currently, preparing and uploading a dataset and its corresponding data dictionary involve multiple rounds of manual revisions to meet ODC guidelines. Here, we present the ODC Minimum Data Standards (MDS) and the ODC Data Quality App (ODCdqa), an open-source web-based tool that simplifies the revision and curation workflow. The ODCdqa allows users to automate initial data quality checks and provides immediate feedback on whether the dataset meets ODC data specifications or requires edits. All publicly available datasets on the ODC-SCI and ODC-TBI websites were passed into the ODCdqa for analysis. Exclusion criteria were: (1) did not have a data dictionary, and (2) the dataset and/or data dictionary had different headers. The aggregated results were then used to identify common areas of difficulty and inform the future directions of ODC tool development and platform. Out of the 124 publicly available ODC datasets uploaded between 2018 and September 2025, 119 were used (84 from ODC-SCI, 35 from ODC-TBI). Both ODC-SCI and ODC-TBI had a trend of reducing the number of failed checks as the platform matured over the years, and the MDS and ODCdqa were used regularly by curators and users. In summary, the ODCdqa is an automated tool that allows any user to perform initial checks to determine whether a dataset and its corresponding data dictionary are ready for upload to the ODC platform. As the ODC platform evolves, more checks and features will be added to the ODCdqa.
The NIH Common Fund's SPARC (Stimulating Peripheral Activity to Relieve Conditions) program was launched in 2015 to catalyze bioelectronic medicine by advancing foundational knowledge of the autonomic nervous system (ANS) and enabling clinical translation of novel bioelectronic medicines. While the program officially ended in 2025, activities related to the program are still ongoing through no-cost extensions, data sharing, and derivative efforts. We provide here an overview of the major outcomes of SPARC under its four focus areas: Anatomical and Functional Mapping, Technology Development, Translational Research, and Data/Modeling Infrastructure. In the Anatomical and Functional Mapping focus areas, teams generated cross-species, cross-organ ANS connectivity datasets, detailed nerve anatomy, along with detailed maps and connectivity resources to support query, visualization, and hypothesis-generation. Under the Technology Development focus area, teams developed new tools for recording and stimulating autonomic pathways and built computational pipelines for realistic neuromodulation simulations. Within the Translational Research, teams conducted pre-clinical and clinical proof-of-concept studies and used prize-driven efforts to accelerate clinically ready neuromodulation solutions. Data/Modeling Infrastructure teams built a durable open-science ecosystem enabling FAIR sharing, exploration, and reuse of SPARC datasets, models, and workflows by the broader community. Together, SPARC's integrated approach has lowered barriers to precision neuromodulation research and established reusable resources to sustain the field beyond the program.
The BRAIN Initiative Cell Atlas Network (BICAN) is generating large-scale multimodal datasets to profile cell types in the human, non-human primate, and mouse brain. The diversity of single-cell and spatial transcriptomic and epigenomic assays, combined with varied experimental contexts, multiple data-generating laboratories and distributed infrastructure, poses substantial challenges for data integration and reuse in BICAN. To address this, we implemented a standards framework that enables layered integration of these data into knowledge-ready products for interoperable brain cell atlases. This framework organizes data based on three progressively structured layers. First, we introduced an assay-agnostic modeling layer that unifies the representation of single-cell and spatial omics data using a common set of biological entities and processes assessed by diverse experimental techniques. Second, we implemented harmonized metadata standards that capture key experimental features linked to biospecimen provenance across heterogeneous tissue sources, species, and preparations, supporting integration and validation while minimizing burden on data contributors. Third, we present an extensible representation for data-driven cell type taxonomies that integrates molecular data with annotations, ontology mappings, and evidence. Together, these contributions represent an end-to-end framework that transforms heterogeneous datasets into structured, interoperable resources that support broad community reuse via mapping algorithms, annotation systems, and visualization platforms. This approach links biospecimen provenance with cell-level outputs and embeds these in a standardized taxonomy format, enabling downstream applications such as cross-dataset integration, reference mapping, and knowledge-driven analysis. More broadly, our work demonstrates a generalizable strategy for enabling an efficient data-to-knowledge pipeline in a large-scale consortium setting.
Preclinical traumatic brain injury (TBI) research relies on experimental models that vary by mechanism, parameters, surgical procedures, species, strains, and ages, to name a few. While these models are crucial for understanding injury mechanisms and testing therapies, the progress in translating this knowledge to the clinic has been limited. This is in part due to fragmented resources and inconsistent reporting of critical variables. Here, we introduce the PRECISE-TBI model catalog, a centralized, queryable resource that consolidates metadata from published studies. The catalog integrates curated annotations from more than 450 papers, including details such as age, sex, strain, model type, device, and injury parameters. Where available, entries are also linked to protocols and datasets to enhance transparency and reproducibility. The Model Catalog serves as a living resource that enables cross-study comparison, identifies gaps in reporting, and connects the literature to datasets, protocols, device information, and other relevant resources. Analysis of the initial catalog entries revealed gaps in the reporting of device, age, and weight. In contrast, the reporting of sex improved over time, with over 90% of recent studies within the catalog papers reporting sex. Strain was also reported in most studies, with consistent reporting of specificity, especially for the C57 mice substrain. We expect the Model Catalog to serve as a valuable tool to enhance study design and reproducibility in preclinical TBI research while advancing FAIR data principles in the TBI field.
The Open Data Commons for Traumatic Brain Injury (ODC-TBI.org) was launched in 2018 to support data sharing in pre-clinical TBI. As data science and artificial intelligence continue to advance, open sharing of high-quality, FAIR (Findable, Accessible, Interoperable, and Reusable) data has assumed critical importance to propel discovery science and as a countermeasure to some of the rigor and reproducibility problems plaguing translational research across biomedicine. Researcher-led specialist repositories such as ODC-TBI serve as important hubs through which biomedical communities come together to define data sharing requirements for their respective domains in support of new requirements by funders and journals for routine data sharing. ODC-TBI is now the recognized data repository for pre-clinical TBI research, listed on the National Library of Medicine-recommended repository listing, and is supported by the National Institute on Neurological Disorders and Stroke. ODC-TBI forms one of the critical infrastructure components of the PRE C linical I nteragency re S earch resourc E -TBI (PRECISE-TBI) project, an interagency effort to promote and support data sharing, rigor, and reproducibility in pre-clinical TBI research. Through PRECISE-TBI, the ODC-TBI has conducted broad outreach, starting in 2022, resulting in a significant increase in the number of users, datasets uploaded and public data releases. PRECISE-TBI has facilitated the establishment of an Editorial Board providing community oversight of ODC-TBI policies and recommendations, for example, the use of standards such as common data elements (CDEs). Here we describe the current state of the ODC-TBI, including its organization, operation, and governance. We perform a detailed overview of public datasets to provide insight into data sharing practices, including the use of CDEs and ancillary practices such as providing links to publications and citing data. We examine the impact of the ODC-TBI by providing statistics on downloads per datasets and reuse of ODC-TBI data in published studies. The results not only provide insight into the growth trajectory of ODC-TBI and data sharing behaviors in pre-clinical TBI, but also point to areas where increased outreach, communication, and training are needed to firmly establish a culture of data sharing.
Preclinical research in traumatic brain injury (TBI) continues to significantly increase knowledge and yield a large number of peer-reviewed studies, but translation of these results to the clinical setting has been minimal. Rigor and transparency factors such as concealment of group allocation (e.g., "blinding") or ensuring that reagents are identifiable are critical in ensuring that scientific studies are replicable and translatable. Yet, nearly all efforts aimed at measuring these factors have concluded that reporting practices are problematic and incomplete. One way to improve transparency of reporting practices is to require that authors address a set of transparency-related items in some way, such as a checklist or an article section. Recently, Journal of Neurotrauma, a leading publisher of preclinical TBI research, instituted a required rigor-related section, which is explained to authors via a set of transparency, rigor, and reproducibility (TRR) instructions (one example for each article type). These documents include specific transparency sections explaining blinding, power calculations, protocols, code, and data deposition. Experimental Neurology is a journal that is similar in size, impact, and topic, but the journal does not have explicit instructions to authors about transparency items. The purpose of this study was to assess the degree to which transparency reporting items were included in published articles comparing reporting practices in the Journal of Neurotrauma and Experimental Neurology. We used a commercial software, SciScore, which is an AI tool tuned to detect rigor/transparency sentences in published articles and count the number found (roughly dividing by the number expected) to obtain a score. Overall, SciScore found that in six of eight items that were explicitly asked for, such as power calculations, investigator blinding, inclusion criteria, attrition, and data, there were significant differences (more than 10%) compared to Experimental Neurology. However, in Journal of Neurotrauma articles with the extra rigor section, three of four rigor items that were not explicitly asked for in the template rigor documents, such as subject demographics or transparent antibody reporting, were not different from Experimental Neurology. One item, reporting of the sex of subjects, was significantly better in Experimental Neurology. This shows that the Journal of Neurotrauma's required rigor section is effective in improving reporting, but it would be far better if sex as a biological variable and transparent reporting of reagents (items present on major checklists, including NIH rigor criteria) would be included.
For nearly 350 years, the process of disseminating scientific knowledge has remained largely unchanged. Scientists conduct experiments, analyze the data, and publish their findings in the form of scientific articles. Since the turn of the century, this process has been challenged by numerous open science and data sharing efforts to enhance transparency, reproducibility, and replicability of scientific research. Big data approaches, together with machine learning and artificial intelligence, are frequently used to gain insight into the ever-growing complexity of biological systems and biomedical research. To utilize these approaches and harness the continuously increasing computational power requires data to be both machine readable and, ideally, harmonized across studies. Therein lies the challenge: understanding how to organize and describe data is a critical skill for scientists, yet one that is rarely explicitly taught. Common data elements (CDEs), standardized definitions, and reporting structures for data represent a practical solution to this challenge. With the goal of creating a common language to describe and share pre-clinical spinal cord injury (SCI) research data, the open data commons for SCI, in collaboration with the National Institute of Neurological Disorders and Stroke, kicked off this process with the “Preclinical SCI Common Data Elements (CDE) Workshop,” held in conjunction with the National Neurotrauma Symposium in San Francisco, California in June 2024. In this report, we discuss the workshop proceedings, summarize the input provided by the SCI research community, share insights from related CDE efforts, and provide a pragmatic approach to creating CDEs for pre-clinical SCI research.
The Stimulating Peripheral Activity to Relieve Conditions (SPARC) program is a U.S. National Institutes of Health (NIH) funded effort to enhance our understanding of the neural circuitry responsible for visceral control. SPARC's mission is to identify, extract, and compile our overall existing knowledge and understanding of the autonomic nervous system (ANS) connectivity between the central nervous system and end organs. A major goal of SPARC is to use this knowledge to promote the development of the next generation of neuromodulation devices and bioelectronic medicine for nervous system diseases. As part of the SPARC program, we have been developing the SPARC Connectivity Knowledge Base of the Autonomic Nervous System (SCKAN), a dynamic resource containing information about the origins, terminations, and routing of ANS projections. The distillation of SPARC's connectivity knowledge into this knowledge base involves a rigorous curation process to capture connectivity information provided by experts, published literature, textbooks, and SPARC scientific data. SCKAN is used to automatically generate anatomical and functional connectivity maps on the SPARC portal. In this article, we present the design and functionality of SCKAN, including the detailed knowledge engineering process developed to populate the resource with high quality and accurate data. We discuss the process from both the perspective of SCKAN's ontological representation as well as its practical applications in developing information systems. We share our techniques, strategies, tools and insights for developing a practical knowledgebase of ANS connectivity that supports continual enhancement.
Data interoperability is crucial for effectively combining data for scientific inquiry. To facilitate interoperability, data standards such as a common definition of variables are often developed. The Open Data Commons for Spinal Cord Injury (odc-sci.org) has established an initial set of community-based data elements (CoDEs)—a minimal set of variables for sharing—to promote data interoperability in SCI research, aligning with FAIR (Findable, Accessible, Interoperable, and Reusable) data principles. We sought to understand the use of CoDEs by the SCI community to inform current standards adherence and future standards development. We systematically analyzed 39 public datasets in relation to 17 required CoDEs and found variations between reported data and the structure specified by the CoDEs. Overall, we found that the enforcement of data standards improved reporting rates of CoDEs variables. Notably, different variables were found to require different levels of curation to ensure semantic equivalence among datasets. We also uncovered specific reporting habits of researchers such as formatting and naming patterns. A need for different data standards based on the nature of the study (e.g., human study, derivative study) was realized alongside a detailed list of issues that should be addressed when implementing such standards. Among the various approaches to developing data standards, ODC-SCI adopted a semi-formal approach by creating standards that are easy to adopt by the user. Our data-driven evaluation of actual reporting behavior shows that this flexibility can lead to subsequent problems in harmonization. This study serves as a baseline analysis of reporting behaviors for shaping and facilitating data standards.
Exponential scientific data growth presents challenges and opportunities for addressing 56 complex public health issues like the opioid epidemic and chronic pain management.57 Despite the vast amount of research conducted globally, many datasets remain 58 inaccessible or underutilized due to publication access policies and stringent data use 59 agreements.The amount of data generated through research activities is enormous.60 Limited access to scientific datasets stifles discovery and delays the translation of 61 proven scientific advances into real-world applications.(1)To address these challenges, 62 we argue for the critical importance of open science ecosystems, using the National 63 Institutes of Health Helping to End Addiction Long-term ® Initiative (NIH HEAL 64 Initiative ® ) as a case study.We discuss how building community around data can 65 accelerate scientific discovery by enabling dataset integration, increasing statistical 66 power, and fostering interdisciplinary collaboration.67 68 Open science ecosystems represent a fundamental shift in how research is conducted, 69 shared, and utilized.An open science ecosystem is a comprehensive research 70 environment that combines technological infrastructure, standardized protocols, and 71 collaborative networks to enable transparent sharing and integration of scientific data, 72 methods, and findings.These interconnected networks integrate data collection, 73 storage, processing, analysis, and use across organizations while adhering to FAIR 74 (Findability, Accessibility, Interoperability, and Reusability) principles.The ecosystem 75 encompasses not just the technical components for data sharing, but also the human 76 elements: researchers, institutions, funding bodies, and community partners who work 77 together under shared governance frameworks and data standards to accelerate
Hiring, tenure and promotion processes are powerful influences in driving innovation in research practices and in furthering adoption of open data and software. This ‘Ten simple rules’ article provides practical steps to update institutional processes to incorporate data and software, and to recognize those outputs on their own merit.
Neuroscience has made significant strides over the past decade in moving from a largely closed science characterized by anemic data sharing, to a largely open science where the amount of publicly available neuroscience data has increased dramatically. While this increase is driven in significant part by large prospective data sharing studies, we are starting to see increased sharing in the long tail of neuroscience data, driven no doubt by journal requirements and funder mandates. Concomitant with this shift to open is the increasing support of the FAIR data principles by neuroscience practices and infrastructure. FAIR is particularly critical for neuroscience with its multiplicity of data types, scales and model systems and the infrastructure that serves them. As envisioned from the early days of neuroinformatics, neuroscience is currently served by a globally distributed ecosystem of neuroscience-centric data repositories, largely specialized around data types. To make neuroscience data findable, accessible, interoperable, and reusable requires the coordination across different stakeholders, including the researchers who produce the data, data repositories who make it available, the aggregators and indexers who field search engines across the data, and community organizations who help to coordinate efforts and develop the community standards critical to FAIR. The International Neuroinformatics Coordinating Facility has led efforts to move neuroscience toward FAIR, fielding several resources to help researchers and repositories achieve FAIR. In this perspective, I provide an overview of the components and practices required to achieve FAIR in neuroscience and provide thoughts on the past, present and future of FAIR infrastructure for neuroscience, from the laboratory to the search engine.
Effective data management and sharing have become increasingly crucial in biomedical research; however, many laboratory researchers lack the necessary tools and knowledge to address this challenge. This article provides an introductory guide into research data management (RDM), and the importance of FAIR (Findable, Accessible, Interoperable, and Reusable) data-sharing principles for laboratory researchers produced by practicing scientists. We explore the advantages of implementing organized data management strategies and introduce key concepts such as data standards, data documentation, and the distinction between machine and human-readable data formats. Furthermore, we offer practical guidance for creating a data management plan and establishing efficient data workflows within the laboratory setting, suitable for labs of all sizes. This includes an examination of requirements analysis, the development of a data dictionary for routine data elements, the implementation of unique subject identifiers, and the formulation of standard operating procedures (SOPs) for seamless data flow. To aid researchers in implementing these practices, we present a simple organizational system as an illustrative example, which can be tailored to suit individual needs and research requirements.By presenting a user-friendly approach, this guide serves as an introduction to the field of RDM and offers practical tips to help researchers effortlessly meet the common data management and sharing mandates rapidly becoming prevalent in biomedical research.
Abstract Antibodies are ubiquitous key biological research resources yet are tricky to use as they are prone to performance issues and represent a major source of variability across studies. Understanding what antibody was used in a published study is therefore necessary to repeat and/or interpret a given study. However, antibody reagents are still frequently not cited with sufficient detail to determine which antibody was used in experiments. The Antibody Registry is a public, open database that enables citation of antibodies by providing a persistent record for any antibody-based reagent used in a publication. The registry is the authority for antibody Research Resource Identifiers, or RRIDs, which are requested or required by hundreds of journals seeking to improve the citation of these key resources. The registry is the most comprehensive listing of persistently identified antibody reagents used in the scientific literature. Data contributors span individual authors who use antibodies to antibody companies, which provide their entire catalogs including discontinued items. Unlike many commercial antibody listing sites which tend to remove reagents no longer sold, registry records persist, providing an interface between a fast-moving commercial marketplace and the static scientific literature. The Antibody Registry (RRID:SCR_006397) https://antibodyregistry.org.
The December 2022 release of the SPARC Portal ( https://sparc.science ) included the first significant update to the anatomical connectivity flatmaps since the portal first launched. These flatmaps provide an interactive and visual map for the display and exploration of the autonomic nervous system of Human, rat, mouse, pig, and cat ( https://sparc.science/maps ). In addition to improved anatomical ‘cartoons’ for all species and adding a female Human map, these maps, for the first time, automatically render the connectivity knowledge directly retrieved from the SPARC Connectivity Knowledge Base of the Autonomic Nervous System (SCKAN; https://sparc.science/resources/6eg3VpJbwQR4B84CjrvmyD ). SCKAN contains explicit knowledge about CNS-ANS-end organ circuitry derived from SPARC data and scientific literature, in a form that supports computational reasoning. Each flatmap consists of a manually drawn base layer for the species-specific anatomical cartoon and a layer of manually drawn tracts for the large nerves. By annotating these drawings with standard reference SPARC vocabularies consistent with SCKAN usage, software tools are then able to semantically map connectivity circuits retrieved from SCKAN to the visual representation on each species’ flatmap. Beyond the interactive visual rendering of the circuitry on the SPARC Portal, the semantic consistency between flatmaps and SCKAN powers further user interface components on the SPARC Portal to surface additional knowledge for each rendered connection. For example, comprehensive links to the literature and/or data supporting a connectivity statement can be retrieved from SCKAN and provided to Portal users. Now that tools are in place to support this automated workflow to generate the flatmaps from SCKAN knowledge, we are continuing to improve the SPARC Portal to visualise new knowledge as it becomes available. NIH Common Fund, NIH Office of the Director, Awards OT3OD025349 and OT2 OD030541 This is the full abstract presented at the American Physiology Summit 2023 meeting and is only available in HTML format. There are no additional versions or additional content available for this abstract. Physiology was not involved in the peer review process.
Characterizing cellular diversity at different levels of biological organization and across data modalities is a prerequisite to understanding the function of cell types in the brain. Classification of neurons is also essential to manipulate cell types in controlled ways and to understand their variation and vulnerability in brain disorders. The BRAIN Initiative Cell Census Network (BICCN) is an integrated network of data-generating centers, data archives, and data standards developers, with the goal of systematic multimodal brain cell type profiling and characterization. Emphasis of the BICCN is on the whole mouse brain with demonstration of prototype feasibility for human and nonhuman primate (NHP) brains. Here, we provide a guide to the cellular and spatial approaches employed by the BICCN, and to accessing and using these data and extensive resources, including the BRAIN Cell Data Center (BCDC), which serves to manage and integrate data across the ecosystem. We illustrate the power of the BICCN data ecosystem through vignettes highlighting several BICCN analysis and visualization tools. Finally, we present emerging standards that have been developed or adopted toward Findable, Accessible, Interoperable, and Reusable (FAIR) neuroscience. The combined BICCN ecosystem provides a comprehensive resource for the exploration and analysis of cell types in the brain.
The Stimulating Peripheral Activity to Relieve Conditions (SPARC) program is a NIH-funded consortium to improve the understanding of how the autonomic nervous system (ANS) interacts with end organs and the central nervous system. A major goal of SPARC is to use this knowledge to develop the next generation of neuromodulator devices as effective disease therapies. To support neuromodulation planning, the routes of ANS neuron populations (as well as those populations conveying visceral sensing) are to be mapped in terms of the anatomical structures where they pass through or terminate. Such a map would also serve as an organizing scaffold for (i) a growing SPARC repository of electrophysiology recordings and molecular assay data, as well as (ii) the computational simulation of the effect of neural stimulation on any relevant point on the CNS or ANS. The resulting SPARC map represents connectivity as a network of conduits, and is known as the SPARC Connectivity Knowledge Base of the Autonomic Nervous System (SCKAN). SCKAN contains explicit knowledge about CNS-ANS-end organ circuitry derived from SPARC data and scientific literature, in a form that supports computational reasoning. All connections are annotated with standard reference SPARC vocabularies allowing us to integrate datasets that are annotated to the same annotation standard. Circuits on the map represent details of ANS connectivity associated with a particular organ, e.g., bladder control, defensive breathing, modulation of peristalsis or cardiac inotropy/chronotropy. These circuits are created through interviews with SPARC investigators, anatomical experts and the scientific literature. They contain detailed representations of neuron populations giving rise to ANS connections, including mappings of the locations of cell bodies, dendrites, axon segments and synaptic endings. Circuits are modeled using ApiNATOMY, a knowledge model and tool suite specifically created to represent biological connectivity. Furthermore, these detailed circuits are supplemented with general knowledge on connectivity between CNS nuclei, ANS ganglia, nerves, and end organs derived from the scientific literature. To facilitate extracting this knowledge from the literature and maintaining relevance over time, we developed a Natural Language Processing (NLP) pipeline, which extracts connectivity relationships between unique anatomical structures within a sentence from the scientific literature. SCKAN is being used to create a queryable visual atlas of ANS circuitry through the SPARC portal. SPARC Funding: NIH Common Fund, NIH Office of the Director, Award OT2 OD030541 This is the full abstract presented at the American Physiology Summit 2023 meeting and is only available in HTML format. There are no additional versions or additional content available for this abstract. Physiology was not involved in the peer review process.