The field of neuroimaging has witnessed an exponential increase in data production and availability, fueled by technical advances that have enabled large-scale data collection, as well as by changes in policies and funding initiatives that support open data-sharing programs. As the number of openly available datasets grows larger, the need for efficient, standardized, and flexible tools to index, inspect, and fetch neuroimaging data is becoming a pressing need. OpenNeuro is a popular data hosting service in neuroimaging research; it allows to share publicly neuroimaging data stored following the Brain Imaging Data Structure (BIDS) standard. Although tools such as bids2table (https://childmindresearch.github.io/bids2table/bids2table.html) or Neurobagel Query enable interrogating BIDS datasets (including those on OpenNeuro) and the associated metadata, researchers that require using a subset of one or multiple datasets stored in data repositories typically need to fetch the entire dataset, and to use dedicated tools to interrogate and select the relevant data, hindering interoperability and scalability. Furthermore, existing tools to interrogate BIDS datasets only enable interrogating the metadata, and miss other features that can only be obtained by inspecting the actual volumetric data. Conversely, command-line tools allowing to fetch datasets from remote servers (e.g., Cohort Creator) do not provide data inspection or metadata-based filtering capabilities. Thus, a tool to efficiently interrogate and filter neuroimaging data hosted on public repositories is missing. We present NiQuery , a principled, open-source Python library designed to address these challenges by providing a robust framework for querying neuroimaging datasets, inspecting metadata, performing subset selection, and aggregating results. NiQuery seeks to facilitate reproducible research by enabling transparent access to the content of neuroimaging datasets, including volume-specific metadata.
The Brain Imaging Data Structure (BIDS) is a widely adopted, community-driven standard to organize neuroimaging data and metadata. Although numerous extensions have been developed to incrementally extend coverage to new modalities and data types, an unambiguous, granular specification for eye-tracking recordings is lacking. Here, we present how BIDS will structure data and metadata produced by eye-tracking devices, including gaze position and pupil data. In addition to prescribing the organization of the unprocessed (raw) recordings and associated metadata as produced by the device, BEP20 also resolves gaps in current BIDS specifications beyond the scope of eye tracking. In particular, it adds a mechanism for including asynchronous model parameters and messages, such as contextual information, statuses, and events, such as triggers, generated by the device. BEP20 includes examples that illustrate its applicability in various experimental settings. This BIDS extension provides a robust standard that supports the development of self-adaptive, open, and automated eye-tracking data structures, thereby bolstering transparency and reliability of results in this field.
Molecular neuroimaging with positron emission tomography (PET) and single-photon emission computed tomography (SPECT) enables quantification of specific molecular targets in the living brain. Despite its scientific impact, molecular neuroimaging research has historically faced challenges due to high costs, small sample sizes, laboratory-specific analysis pipelines, and limited large-scale data sharing. These factors have hindered reproducibility and the broader reuse of valuable PET datasets. The OpenNeuroPET initiative was established to address these barriers by developing standards, infrastructure, and open-source tools for organizing, sharing, and analyzing molecular neuroimaging data. Through collaborations across Europe and North America, OpenNeuroPET has supported the PET extension of the Brain Imaging Data Structure (PET-BIDS), providing a standardized framework for PET datasets and metadata. Building on PET-BIDS, tools such as PET2BIDS, ezBIDS, and BIDSCoin facilitate data conversion and curation. In parallel, OpenNeuro now hosts PET-BIDS datasets for open sharing, while complementary platforms such as PublicnEUro enable GDPR-compliant controlled access. Emerging open-source workflows and BIDS applications further support automated, reproducible PET preprocessing and quantitative analysis, promoting harmonized processing across centers. Together, these developments mark an important step toward an open molecular neuroimaging ecosystem in which datasets, software, and workflows can be transparently shared, reused, and scaled for collaborative research.
Electromyography (EMG) is fundamental to clinical assessment, rehabilitation, neuromuscular research, and human-machine interfaces. Despite decades of use, no widely adopted standard exists for organizing and sharing EMG data, limiting reusability and large-scale data aggregation. We present EMG-BIDS, an extension to the Brain Imaging Data Structure (BIDS) that standardizes the organization of EMG recordings. EMG-BIDS addresses challenges unique to EMG, including diverse electrode types (surface or intramuscular, single channel to high-density arrays), heterogeneous electrode placements across anatomical locations, montages (e.g., monopolar or bipolar sensor designs), and the critical need for transparent documentation of sensor positioning. The specification introduces hierarchical coordinate systems that link local electrode grids to anatomical landmarks, enabling precise and reproducible placement documentation. EMG-BIDS is now part of BIDS as of version 1.11.0, supported by existing tools, including MNE-BIDS and EEGLAB. We demonstrate the specification through public datasets, including high-density surface EMG recordings. EMG-BIDS provides the foundation for FAIR (Findable, Accessible, Interoperable, Reusable) EMG data sharing, enabling meta-analyses, multi-site studies, and machine learning applications that require standardized, well-documented datasets.
The ethical and legal imperative to share research data without causing harm requires careful attention to privacy risks. While mounting evidence demonstrates that data sharing benefits science, legitimate concerns persist regarding the potential leakage of personal information that could lead to reidentification and subsequent harm. We reviewed metadata accompanying neuroimaging datasets from heterogeneous studies openly available on OpenNeuro, involving participants across the lifespan-from children to older adults-with and without clinical diagnoses, and including associated clinical score data. Using metaprivBIDS (https://github.com/CPernet/metaprivBIDS), a software application for BIDS-compliant tsv/json files that computes and reports different privacy metrics (k-anonymity, k-global, l-diversity, SUDA, PIF), we found that privacy is generally well maintained, with serious vulnerabilities being rare. Nonetheless, issues were identified in nearly all datasets and warrant mitigation. Notably, clinical score data (e.g., neuropsychological results) posed minimal reidentification risk, whereas demographic variables-age, sex assigned at birth, sexual orientations, race, income, and geolocation-represented the principal privacy vulnerabilities. We outline practical measures to address these risks, enabling safer data sharing practices.
Ensuring the long-term reproducibility of data analyses requires results stability tests to verify that analysis results remain within acceptable variation bounds despite inevitable software updates and hardware evolutions. This paper introduces a numerical variability approach for results stability tests, which determines acceptable variation bounds using random rounding of floating-point calculations. By applying the resulting stability test to fMRIPrep , a widely-used neuroimaging tool, we show that the test is sensitive enough to detect subtle updates in image processing methods while remaining specific enough to accept numerical variations within a reference version of the application. This result contributes to enhancing the reliability and reproducibility of data analyses by providing a robust and flexible method for stability testing.
The Brain Imaging Data Structure (BIDS) is an increasingly adopted standard for organizing scientific data and metadata. It facilitates easier and more straightforward data sharing and reuse. BIDS currently encompasses several biomedical imaging and non-imaging techniques, and as more research groups begin to use it, additional experimental techniques are being incorporated into the standard, allowing diverse experimental methods to be stored within the same cohesive structure. Here, we present an extension for magnetic resonance spectroscopy (MRS) data, termed MRS-BIDS.
Functional near-infrared spectroscopy (fNIRS) is an increasingly popular neuroimaging technique that measures cortical hemodynamic activity in a non-invasive and portable fashion. Although the fNIRS community has been successful in disseminating open-source processing tools and a standard file format (SNIRF), reproducible research and sharing of fNIRS data amongst researchers has been hindered by a lack of standards and clarity over how study data should be organized and stored. This problem is not new in neuroimaging, and it became evident years ago with the proliferation of publicly available neuroimaging datasets. To solve this critical issue, the neuroimaging community created the Brain Imaging Data Structure (BIDS) that specifies standards for how datasets should be organized to facilitate sharing and reproducibility of science. Currently, BIDS supports dozens of neuroimaging modalities including MRI, EEG, MEG, PET, and many others. In this paper, we present the extension of BIDS for NIRS data alongside tools that may assist researchers in organizing existing and new data with the goal of promoting public disseminations of fNIRS datasets.
SynopsisMotivation: Create an open-source alternative to FSL’s eddy without commercial limitations thatgeneralizes over problems beyond eddy current distortions of diffusion MRI.Goals: Evaluate the tool’s performance vis-à-vis with corresponding tools.Approach: We implement a Gaussian Process regressor model and a Bayesian optimizationlayer on an existing registration framework, and estimate the alignment of neuroimaging dataenabled by the eddymotion framework.Results: Using publicly available neuroimaging data, we provide evidence about theeffectiveness of our volume-to-volume modeling framework for generalized artifact identificationand correction in neuroimaging.ImpactWe present eddymotion, an open-source framework for volume-to-volume artifact estimationinspired by FSL eddy. Our tool allows for easy alternative model implementation, it can be usedfor non-dMRI imaging modalities and does not have a restrictive license.
The adoption of a standardized preprocessing workflow is vital for fostering community, sharing, and reproducibility. fMRIPrep has been a critical advancement towards this end, however, it is limited in its capacity to be applied to data across the lifespan, starting from infancy. Here, we introduce fMRIPrep Lifespan, an extension of fMRIPrep that extends the standardized processing from childhood to senescence to include neonatal, infant, and toddler structural and functional MRI data preprocessing. This effort involves a NiPreps integration of 1) a workflow akin to fMRIPrep optimized for MRI data in the first years of life (previously NiBabies) and 2) upstream enhancements to the entire NiPreps suite, including multi-echo data processing, modularization of workflow components, and convergence of processing with other popular workflows (ABCD-BIDS, Human Connectome Project Pipelines). Using data from the Baby Connectome Project (participants 1-43 months of age), we demonstrate that fMRIPrep Lifespan produces high-quality outputs across a wide age range. Moving forward, the scalable, modular infrastructure of fMRIPrep Lifespan will ensure adaptability to data from birth to old age while maintaining robust and reproducible frameworks for functional MRI research across the lifespan.
This article describes the rationale, aims, and methodology of the Accelerating Medicines Partnership® Schizophrenia (AMP® SCZ). This is the largest international collaboration to date that will develop algorithms to predict trajectories and outcomes of individuals at clinical high risk (CHR) for psychosis and to advance the development and use of novel pharmacological interventions for CHR individuals. We present a description of the participating research networks and the data processing analysis and coordination center, their processes for data harmonization across 43 sites from 13 participating countries (recruitment across North America, Australia, Europe, Asia, and South America), data flow and quality assessment processes, data analyses, and the transfer of data to the National Institute of Mental Health (NIMH) Data Archive (NDA) for use by the research community. In an expected sample of approximately 2000 CHR individuals and 640 matched healthy controls, AMP SCZ will collect clinical, environmental, and cognitive data along with multimodal biomarkers, including neuroimaging, electrophysiology, fluid biospecimens, speech and facial expression samples, novel measures derived from digital health technologies including smartphone-based daily surveys, and passive sensing as well as actigraphy. The study will investigate a range of clinical outcomes over a 2-year period, including transition to psychosis, remission or persistence of CHR status, attenuated positive symptoms, persistent negative symptoms, mood and anxiety symptoms, and psychosocial functioning. The global reach of AMP SCZ and its harmonized innovative methods promise to catalyze the development of new treatments to address critical unmet clinical and public health needs in CHR individuals.
Reducing contributions from non-neuronal sources is a crucial step in functional magnetic resonance imaging (fMRI) connectivity analyses. Many viable strategies for denoising fMRI are used in the literature, and practitioners rely on denoising benchmarks for guidance in the selection of an appropriate choice for their study. However, fMRI denoising software is an ever-evolving field, and the benchmarks can quickly become obsolete as the techniques or implementations change. In this work, we present a denoising benchmark featuring a range of denoising strategies, datasets and evaluation metrics for connectivity analyses, based on the popular fMRIprep software. The benchmark prototypes an implementation of a reproducible framework, where the provided Jupyter Book enables readers to reproduce or modify the figures on the Neurolibre reproducible preprint server (https://neurolibre.org/). We demonstrate how such a reproducible benchmark can be used for continuous evaluation of research software, by comparing two versions of the fMRIprep. Most of the benchmark results were consistent with prior literature. Scrubbing, a technique which excludes time points with excessive motion, combined with global signal regression, is generally effective at noise removal. Scrubbing was generally effective, but is incompatible with statistical analyses requiring the continuous sampling of brain signal, for which a simpler strategy, using motion parameters, average activity in select brain compartments, and global signal regression, is preferred. Importantly, we found that certain denoising strategies behave inconsistently across datasets and/or versions of fMRIPrep, or had a different behavior than in previously published benchmarks. This work will hopefully provide useful guidelines for the fMRIprep users community, and highlight the importance of continuous evaluation of research methods.
Motivation: Reliable neuroimaging pipelines require the implementation of robust QA/QC protocols. Goal(s): Developing an extension of MRIQC for the QA/QC of diffusion MRI data. Approach: We build on MRIQC's infrastructure to generate individual visual reports of dMRI images and define new image quality metrics (IQMs). Results: We developed a minimal processing pipeline for whole-brain dMRI data of human adults. The processing pipeline generates individual visual reports for the QA of unprocessed inputs. The pipeline also extracts IQMs to train automated decision-making, following MRIQC's established pattern. Impact: MRIQC is a widely-adopted tool for the QA/QC of unprocessed MRI data. However, support for dMRI was previously lacking. This MRIQC extension will improve QA/QC of dMRI by bringing it to the highest standards and will facilitate the implementation of rigorous protocols in multimodal neuroimaging.
In order to support efficient processing, data must be formatted according to standards that are prevalent in the field and widely supported among actively developed analysis tools.The Brain Imaging Data Structure (BIDS) (Gorgolewski et al., 2016) is an open standard designed for computational accessibility, operator legibility, and a wide and easily extendable scope of modalities -and is consequently used by numerous analysis and processing tools as the preferred input format in many fields of neuroscience.HeuDiConv (Heuristic DICOM Converter) enables flexible and efficient conversion of spatially reconstructed neuroimaging data from the DICOM format (quasi-ubiquitous in biomedical image acquisition systems, particularly in clinical settings) to BIDS, as well as other file layouts.HeuDiConv provides a multi-stage operator input workflow (discovery, manual tuning, conversion) where a manual tuning step is optional and the entire conversion can thus be seamlessly integrated into a data processing pipeline.HeuDiConv is written in Python, and supports the DICOM specification for input
The Brain Imaging Data Structure (BIDS) is a community-driven standard for the organization of data and metadata from a growing range of neuroscience modalities. This paper is meant as a history of how the standard has developed and grown over time. We outline the principles behind the project, the mechanisms by which it has been extended, and some of the challenges being addressed as it evolves. We also discuss the lessons learned through the project, with the aim of enabling researchers in other domains to learn from the success of BIDS.
As data sharing has become more prevalent, three pillars - archives, standards, and analysis tools - have emerged as critical components in facilitating effective data sharing and collaboration. This paper compares four freely available intracranial neuroelectrophysiology data repositories: Data Archive for the BRAIN Initiative (DABI), Distributed Archives for Neurophysiology Data Integration (DANDI), OpenNeuro, and Brain-CODE. The aim of this review is to describe archives that provide researchers with tools to store, share, and reanalyze both human and non-human neurophysiology data based on criteria that are of interest to the neuroscientific community. The Brain Imaging Data Structure (BIDS) and Neurodata Without Borders (NWB) are utilized by these archives to make data more accessible to researchers by implementing a common standard. As the necessity for integrating large-scale analysis into data repository platforms continues to grow within the neuroscientific community, this article will highlight the various analytical and customizable tools developed within the chosen archives that may advance the field of neuroinformatics.
Advances in both data acquisition and processing methods have given magnetic resonance imaging researchers (MRI) a plethora of options on how best to clean and standardize data before statistical analysis. Recently, there has been a surge in standardized data processing workflows, but special populations, such infants, require modified techniques not normally found in general pipelines. Here we introduce NiBabies, a robust and open-source structural and functional MRI preprocessing pipeline designed for infant populations.
Reducing contributions from non-neuronal sources is a crucial step in functional magnetic resonance imaging (fMRI) analyses. Many viable strategies for denoising fMRI are used in the literature, and practitioners rely on denoising benchmarks for guidance in the selection of an appropriate choice for their study. However, fMRI denoising software is an ever-evolving field, and the benchmarks can quickly become obsolete as the techniques or implementations change. In this work, we present a fully reproducible denoising benchmark featuring a range of denoising strategies and evaluation metrics, built primarily on the fMRIPrep and Nilearn software packages. We apply this reproducible benchmark to investigate the robustness of the conclusions across two different datasets and two versions of fMRIPrep. The majority of benchmark results were consistent with prior literature. Scrubbing, a technique which excludes time points with excessive motion, combined with global signal regression, is generally effective at noise removal. Scrubbing however disrupts the continuous sampling of brain images and is incompatible with some statistical analyses, e.g. auto-regressive modeling. In this case, a simple strategy using motion parameters, average activity in select brain compartments, and global signal regression should be preferred. Importantly, we found that certain denoising strategies behave inconsistently across datasets and/or versions of fMRIPrep, or had a different behavior than in previously published benchmarks, especially ICA-AROMA. These results demonstrate that a reproducible denoising benchmark can effectively assess the robustness of conclusions across multiple datasets and software versions. Technologies such as BIDS-App, the Jupyter Book and Neurolibre provided the infrastructure to publish the metadata and report figures. Readers can reproduce the report figures beyond the ones reported in the published manuscript. With the denoising benchmark, we hope to provide useful guidelines for the community, and that our software infrastructure will facilitate continued development as the state-of-the-art advances.