Careful evaluation of research methodology is fundamental to scientific progress but represents a significant burden on human experts. The complexity of functional MRI (fMRI) methods makes transparent reporting, as suggested by OHBM COBIDAS guidelines, particularly critical. Large Language Models (LLMs) present a potential solution for rapid, scalable methodological assessment. We evaluated three state-of-the-art LLMs (Gemini 2.5 Pro, Claude 4 Sonnet, ChatGPT-o3-pro) against human expert ratings. Fifty fMRI articles (taken from 2016 to 2025) were independently evaluated by ten human experts and three LLMs using an 82-item COBIDAS based rubric. Human raters demonstrated excellent inter-rater reliability (ICC = 0.801), while LLMs showed poor internal agreement (ICC = 0.254). When comparing total scores across papers, Gemini showed strong positive correlation with human consensus (r = 0.693, p < 0.0001), Claude showed moderate positive correlation (r = 0.394, p = 0.004), while ChatGPT showed negative correlation (r = -0.172, p = 0.233). Gemini maintained high reliability when added to human raters (combined ICC = 0.811), achieving 85.3 % exact agreement and 98.8 % within-1-point agreement. Domain-specific analysis revealed Gemini's consistently high agreement across all six COBIDAS sections (experimental design: 0.915, statistical modeling: 0.880), while ChatGPT and Claude showed weaker, more variable performance. Obvious differences emerged in determining non-applicable items: humans marked 40.5 % as not applicable versus 32.3 % for Gemini, 9.2 % for ChatGPT and 21.1 % for Claude. ChatGPT exhibited extreme score volatility, with papers ranging from 0 to 121 points compared to humans' 44.2-77.7 range. LLM scoring required 1-7 min versus 30-35 min for humans. This proof-of-concept study demonstrates that LLM-assisted methodological evaluation is feasible for complex neuroimaging research and could likely be applied to other research fields.
Large-scale neuroimaging datasets are increasingly used to map relationships between brain structure, function, and behavior across the human lifespan. Routinely, analyses exclude participants who moved too much during imaging. While this decision is framed as quality control, it is increasingly recognized that head motion is not randomly distributed across individuals within a study, and so motion-based exclusion may preferentially remove people with particular characteristics relevant to the scientific goals of the study. Here we survey head motion and how it relates to participant characteristics across six large, publicly available datasets spanning nearly the entire human lifespan, namely the Human Connectome Project (HCP) Young Adult, HCP in Development, HCP in Aging, Adolescent Brain Cognitive DevelopmentSM Study, UK Biobank, and Spatial Topology project. These six datasets comprise more than 50,000 unique participants and 300,000 scans. We further benchmark our findings against motion distributions aggregated by MRIQC across more than 1.5 million scans. Under commonly applied strict exclusion thresholds, large fractions of participants would be removed (exceeding 80% in the UK Biobank task data), and these removals were demographically structured, disproportionately excluding younger and older participants, those with higher BMI, and those with motion-associated clinical conditions. Respiratory pseudo-motion inflated estimates of head motion in adult cohorts, and applying notch filtering to remove respiratory frequencies from these estimates meaningfully reduced exclusion rates. Exclusion also carried downstream consequences. Strict thresholds reduced statistical power, inflated study costs, and altered the apparent predictability of behavioral phenotypes by removing a non-random, behaviorally distinct subgroup. These findings demonstrate that motion exclusion thresholds are not neutral quality-control decisions but structured selection mechanisms that reshape the composition of neuroimaging samples. We recommend that studies report the demographic characteristics of excluded participants, prefer data-driven censoring methods over fixed motion cutoffs, and clarify the target population while considering appropriate weighting techniques.
Abstract Social interaction processing and theory of mind (ToM) frequently co-occur, but their commonalities and distinctions at behavioral and neural levels remain unclear. Participants (N = 231) provided moment-by-moment ratings of four text and four audio narratives on social interactions and ToM engagement, which were reliable (split-half r = 0.98 and 0.92, respectively) but only modestly correlated (r = 0.32). In a second sample (N = 90), we analyzed the co-variation between social interaction and ToM ratings and fMRI activity during text and audio narratives. Activity maps associated with social interaction processing and ToM generalized across text and audio (spatial r = 0.60 and 0.58, respectively) and overlapped in canonical ToM regions (FDR q < 0.01). ToM uniquely engaged the anterior intraparietal sulcus, right lateral occipitotemporal cortex, and right supplementary motor area. These results suggest that observing social interactions automatically engages canonical ToM regions, even without explicit mentalizing, and ToM additionally engages brain regions related to action understanding.
Aggregating neuroimaging data across sites and studies is increasingly common, yet site- and scanner-related batch effects can obscure meaningful biological variation and introduce spurious associations. Although ComBat and its extensions are widely used, they are primarily designed for single-metric (univariate) harmonization. In practice, neuroimaging studies often involve multiple biologically coupled metrics (e.g., cortical thickness, surface area, and gray-matter volume) measured across multiple features (e.g., regional values), with shared covariance structure both within and across metrics. Applying univariate ComBat independently to each metric ignores these dependencies and can leave residual batch effects in cross-metric covariance. Using data from the NIH Acute to Chronic Pain Signatures (A2CPS) program, we show that batch effects occur not only in means and variances but also in covariance across cortical regions and metrics—relationships that univariate ComBat does not fully remove. We propose MV-ComBat, a multivariate extension of ComBat that jointly harmonizes multiple metrics by borrowing strength across them. Both empirical Bayes (EB) and Bayesian Markov Chain Monte Carlo (MCMC) implementations of MV-ComBat effectively reduce batch effects. In our experiments, EB is more robust to measurement error, whereas MCMC more accurately recovers cross-metric correlations when priors are well specified. Recognizing that batch effects can also affect feature-level covariance, CovBat was recently introduced as an extension of ComBat that harmonizes both first- and second-order moments across sites. We extend CovBat to the multivariate framework as MV-CovBat, which performs a second-stage latent-space harmonization to address covariance-related batch effects across features and metrics. Simulations confirm that MV-ComBat improves correlation recovery and biological signal preservation relative to univariate ComBat, particularly for moderate-to-strong effects, and that MV-CovBat further improves separation of true biological variation from batch effects when independence assumptions are violated. Together, these methods provide a flexible and unified framework for harmonizing complex, multi-metric neuroimaging data in large-scale, multi-site studies.
IntroductionCerebellar transcranial direct current stimulation (tDCS) combined with language therapy can aid in chronic aphasia recovery, but the neural mechanisms and biomarkers of treatment efficacy remain uncertain.MethodsIn this secondary analysis of data from a previously conducted clinical trial, we used a randomized, double-blind, sham-controlled, within-subject crossover design with a study sample of 19 participants with post-stroke aphasia. We assessed the degree to which baseline properties of cerebro-cerebellar white matter tracts can predict or moderate longitudinal treatment effects at three time points: post-treatment, 2 weeks post-treatment, and 2 months post-treatment. Tract properties were measured by fractional anisotropy (FA) and mean diffusivity (MD) from diffusion tensor imaging (DTI). We also tested whether there are differential effects between trained and untrained language tasks and between cerebellar tDCS polarity (anodal and cathodal).ResultsBaseline measures of tracts connecting the left lesioned cortex to the right posterolateral cerebellum (stimulation target) influenced treatment gains for untrained tasks, relative to sham control. In contrast, for the trained task, treatment gains were influenced by baseline measures of tracts connecting the non-stimulated left cerebellum with the contralateral right cerebral cortex. Although there were no consistent effects from cerebellar tDCS polarity, a highly consistent pattern emerged across all tasks and tracts. Specifically, language improvements were predicted by a baseline tract profile (i.e., higher FA and lower MD) typically associated with higher white matter integrity, especially within the context of stroke-induced white matter decline.DiscussionThese findings corroborate the potential for baseline tract properties as a biomarker of treatment efficacy and support the notion that adjuvant (cerebellar tDCS + language) therapy preferentially benefits individuals with relatively preserved structural connections within functionally relevant networks.Clinical trial registrationClinicalTrials.gov, identifier (NCT02901574).
Neuroimaging studies typically assume that sensory properties are encoded in response magnitude within fixed neural populations. However, this approach does not capture changes in the spatial extent of activation topography, despite growing evidence for its behavioral relevance. Stimulus intensity provides a powerful test case for the role of activation topography as a coding feature because it is a basic, parametrically varying property shared across sensory modalities. Using a Bayes factor-based approach and four functional magnetic resonance imaging datasets (three large-scale datasets [total N = 609] and one precision dataset [>2300 trials]), we tested whether higher-intensity stimulation is associated with expansion of activation topography. Participants received sensory stimuli of varying intensities in somatosensory (heat, laser, tactile), auditory, and visual modalities. High-versus low-intensity painful stimulation consistently produced topographical expansion in areas including the primary somatosensory, posterior midcingulate, primary visual cortices, and cerebellar lobules V and VI. This result replicated across two independent large-scale datasets and within individual participants in the precision dataset. Expansion was also observed for tactile, auditory, and visual stimulation, and its extent correlated with psychophysical discriminability. Topographical expansion involved both the enlargement of already-activated areas and the recruitment of novel regions. These findings establish topographical expansion as a replicable feature of intensity coding, challenging the prevailing assumption of a fixed neural topography.
Background/Aims:Clinical trials and observational studies support the synthesis and development of clinical guidelines, highlighting the need for strong data quality assurance measures. The Acute to Chronic Pain Signatures (A2CPS) program is a large-scale, multi-site observational study investigating chronic post-surgical pain and opioid dependence. Its primary goal is to identify biomarkers predictive of progression from acute to chronic pain following knee arthroplasty or thoracic surgery. The A2CPS sites collect data across various domains, including brain magnetic resonance imaging, electronic health records, psychosocial measures, multi-omics, Quantitative Sensory testing, and functional testing.While A2CPS is an observational study, its aims, design, and methodology closely align with clinical trial practices. This includes interdisciplinary collaboration, standardized protocols, defined eligibility criteria, and oversight by a Data and Safety Monitoring Committee.In multifaceted studies like A2CPS, high-quality data are paramount to ensure the accuracy of predictive biomarkers. To improve quality assurance, we developed the A2CPS Data Monitoring Web Application (Web App), an interactive R Shiny web app with real-time data monitoring capabilities. Here, we describe the functionality and utility of the A2CPS Data Monitoring Web App in streamlining quality assurance for the A2CPS study. Methods:The Web App is a secure R Shiny web application accessible to authorized A2CPS Data Integration and Resource Center (DIRC) members. It retrieves and preprocesses data from REDCap, which is then fed into the R Shiny framework. The user interface has a navigation bar and six subpanels, providing easy access to the app's modules and enabling users to switch seamlessly among subpanels. Each subpanel addresses a specific use case and has the functionality to generate downloadable error reports for individual sites, making it easy to share quality documents and communicate with data collection sites. The DIRC uses these reports to identify errors, coordinate remediation, and facilitate targeted training for research personnel. Results:Regular use of the Web App, coupled with engagement with the training team, resulted in an overall reduction of 50% in data quality errors over one year in case report form data (i.e., in-person visit data). The decline in errors was consistent across all sites despite steady enrollment rates, indicating that real-time data monitoring enables focused feedback, mitigates recurring errors, and streamlines data quality assurance. Conclusion:The A2CPS Data Monitoring Web App plays a key role in A2CPS data quality assurance. This robust open-source solution reduces data entry errors and provides targeted feedback and training to the data collection sites. Our results demonstrate the potential for using open-source computational frameworks for data monitoring and quality assurance purposes in both clinical trials and observational studies.
Task-based functional magnetic resonance imaging (fMRI) is a powerful tool for studying brain function. However, the reliability and viability of small-sample studies remain a concern. While it is well understood that larger samples are preferable, researchers often need to interpret findings from small studies (e.g., when reviewing the literature, analyzing pilot data, or assessing subsamples). However, quantitative guidance for making these judgments remains scarce. To address this gap, we leverage the UK Biobank and the Human Connectome Project's Young Adult dataset to survey a range of standard task-based fMRI analyses, from obtaining regional activation maps to performing predictive modeling. These analyses are repeated using volumetric and two types of cortical surface data. For classic mass-univariate analyses (e.g., regional activation detection or cluster peak localization), studies with as few as 40 participants can be adequate depending on the effect size. For predictive modeling, similar sample sizes can be used to detect whether a feature is predictable, but developing stable, generalizable models typically requires cohorts at least an order of magnitude larger, and possibly two (hundreds or thousands). Together, these results clarify how reliability depends on the interplay of effect size, sample size, and analysis type, offering practical guidance for designing and interpreting small-scale task-fMRI studies.
Cognitive neuroscience has advanced significantly due to the availability of openly shared datasets. Large sample sizes, large amounts of data per person, and diversity in tasks and data types are all desirable, but are difficult to achieve in a single dataset. Here, we present an open dataset with N = 101 participants and 6 hours of scanning per participant, including 6 multifaceted functional tasks, 2 hours of naturalistic movie viewing, structural T1 images and multi-shell diffusion imaging as well as autonomic physiological data. This dataset’s combination of sample size, extensive data per participant (>600 iso-hours of data), and a wide range of experimental conditions — including cognitive, affective, social, and somatic/interoceptive tasks — positions it uniquely for probing important questions in cognitive neuroscience.
Social interaction perception and theory of mind (ToM) frequently co-occur, but their commonalities and distinctions at behavioral and neural levels remain unclear. Participants (N = 231) provided moment-by-moment ratings of four text and four audio narratives on social interactions and ToM engagement, which were reliable (split-half r = .98 and .92, respectively) but only modestly correlated (r = .32). In a second sample (N = 90), we analyzed the co-variation between social interaction and ToM ratings and fMRI activity during text and audio narratives. Social interaction and ToM activity maps generalized across modalities (spatial r = .83 and .57, respectively), both with significant, overlapping clusters in canonical mentalizing regions (FDR q < .01). ToM uniquely engaged the lateral occipitotemporal cortex, left anterior intraparietal sulcus, and right premotor cortex. These results suggest that perceiving social interactions automatically involves mentalizing, and ToM additionally engages brain regions for action understanding and executive functions.
Subcortical volumes are a promising source of biomarkers and features in biosignatures, and automated methods facilitate extracting them in large, phenotypically rich datasets. However, while extensive research has verified that the automated methods produce volumes that are similar to those generated by expert annotation, the consistency of methods with each other is understudied. Using data from the UK Biobank, we compare the estimates of subcortical volumes produced by two popular software suites: FSL and FreeSurfer. Although most subcortical volumes exhibit good to excellent consistency across the methods, the tools produce diverging estimates of amygdalar volume. Through simulation, we show that this poor consistency can lead to conflicting results, where one but not the other tool suggests statistical significance, or where both tools suggest a significant relationship but in opposite directions. Considering these issues, we discuss several ways in which care should be taken when reporting on relationships involving amygdalar volume.
In order to support efficient processing, data must be formatted according to standards that are prevalent in the field and widely supported among actively developed analysis tools.The Brain Imaging Data Structure (BIDS) (Gorgolewski et al., 2016) is an open standard designed for computational accessibility, operator legibility, and a wide and easily extendable scope of modalities -and is consequently used by numerous analysis and processing tools as the preferred input format in many fields of neuroscience.HeuDiConv (Heuristic DICOM Converter) enables flexible and efficient conversion of spatially reconstructed neuroimaging data from the DICOM format (quasi-ubiquitous in biomedical image acquisition systems, particularly in clinical settings) to BIDS, as well as other file layouts.HeuDiConv provides a multi-stage operator input workflow (discovery, manual tuning, conversion) where a manual tuning step is optional and the entire conversion can thus be seamlessly integrated into a data processing pipeline.HeuDiConv is written in Python, and supports the DICOM specification for input
The Acute to Chronic Pain Signatures (A2CPS) project is a large-scale, multi-site initiative aimed at identifying biomarkers and biosignatures that predict the transition from acute to chronic pain. The project is collecting multimodal, longitudinal data from over 2,500 individuals at risk for developing chronic pain after surgery. Here we describe the neuroimaging component of A2CPS, including the acquisition protocols, processing pipelines, and contents of the initial data release. The imaging protocol includes structural, diffusion, resting-state and task-based functional magnetic resonance imaging (MRI) data. Data are collected across multiple clinical sites using different scanner manufacturers, with attention to protocol harmonization and quality control. The processing pipeline integrates several established neuroimaging tools to extract potential biomarkers, including measures of brain structure, connectivity, and pain-related neural signatures. The first data release includes pre-surgical imaging data for 595 participants, with high quality ratings across modalities (98.7% of sMRI, 99.8% of dMRI, and 94.6% of fMRI images were rated as acceptable or better). Initial analyses demonstrate expected relationships between brain-derived measures and clinical variables, such as associations between brain age and psychological factors. This dataset represents a valuable resource for both pain research and neuroimaging methods development, with future releases planned to include additional participants and expanded analysis pipelines and processed data derivatives. ### Competing Interest Statement The authors have declared no competing interest.
In the "serial dependence" effect, responses to visual stimuli appear biased toward the last trial's stimulus. However, several kinds of serial dependence exist, with some reflecting prior stimuli and others reflecting prior responses. One-factor analyses consider the prior stimulus alone or the prior response alone and can consider both variables only via separate analyses. We demonstrate that one-factor analyses are potentially misleading and can reach conclusions that are opposite from the truth if both dependencies exist. To address this limitation, we developed two-factor analyses (model comparison with hierarchical Bayesian modeling and an empirical "quadrant analysis"), which consider trial-by-trial combinations of prior response and prior stimulus. Two-factor analyses can tease apart the two dependencies if applied to a sufficiently large dataset. We applied these analyses to a new study and to four previously published studies. When applying a model that included the possibility of both dependencies, there was no evidence of attraction to the prior stimulus in any dataset, but there was evidence of attraction to the prior response in all datasets. Two of the datasets contained sufficient constraint to determine that both dependencies were needed to explain the results. For these datasets, the dependency on the prior stimulus was repulsive rather than attractive. Our results are consistent with the claim that both dependencies exist in most serial dependence studies (the two-dependence model was not ruled out for any dataset) and, furthermore, that the two dependencies work against each other.
Many neuroscience theories assume that tuning modulation of individual neurons underlies changes in human cognition. However, non-invasive fMRI lacks sufficient resolution to visualize this modulation. To address this limitation, we developed an analysis framework called Inferring Neural Tuning Modulation (INTM) for "peering inside" voxels. Precise specification of neural tuning from the BOLD signal is not possible. Instead, INTM compares theoretical alternatives for the form of neural tuning modulation that might underlie changes in BOLD across experimental conditions. The most likely form is identified via formal model comparison, with assumed parametric Normal tuning functions, followed by a non-parametric check of conclusions. We validated the framework by successfully identifying a well-established form of modulation: visual contrast-induced multiplicative gain for orientation tuned neurons. INTM can be applied to any experimental paradigm testing several points along a continuous feature dimension (e.g., direction of motion, isoluminant hue) across two conditions (e.g., with/without attention, before/after learning).
In the “serial dependence” effect, responses to visual stimuli appear biased toward the last trial’s stimulus. Fischer and Whitney (2014) proposed that this reflects a “continuity field” that promotes visual stability by biasing perception toward the recent past. However, different kinds of serial dependence exist, with some reflecting prior stimuli and others reflecting prior responses. To untangle the two kinds of dependencies, we used a statistical approach that relies on participants’ naturally occurring, trial-by-trial errors, simultaneously considering the combined effects of the prior response and the prior stimulus. To validate the approach, we collected data in an experiment designed to produce relatively large errors, such that the prior response and prior stimulus were dissociated across trials. We applied the approach to our own data, and to data from previous serial dependence studies, including Fischer and Whitney’s. Whenever these two effects could be disentangled, we found that serial dependencies reflected an attraction to the prior response and repulsion from the prior stimulus. In no case did we find evidence of an attraction to the prior stimulus.
Many cognitive neuroscience theories assume that changes in behavior arise from changes in the tuning properties of neurons (e.g., Dosher & Lu 1998, Ling, Liu, & Carrasco 2009). However, direct tests of these theories with electrophysiology are rarely feasible with humans. Non-invasive functional magnetic resonance imaging (fMRI) produces voxel tuning, but each voxel aggregates hundreds of thousands of neurons, and voxel tuning modulation is a complex mixture of the underlying neural responses. We developed a pair of statistical tools to address this problem, which we refer to as NeuroModulation Modeling (NMM). NMM advances fMRI analysis methods, inferring the response of neural subpopulations by leveraging modulations at the voxel-level to differentiate between different forms of neuromodulation. One tool uses hierarchical Bayesian modeling and model comparison while the other tool uses a non-parametric slope analysis. We tested the validity of NMM by applying it to fMRI data collected from participants viewing orientation stimuli at high- and low-contrast, which is known from electrophysiology to cause multiplicative scaling of neural tuning (e.g., Sclar & Freeman 1982). In seeming contradiction to ground truth, increasing contrast appeared to cause an additive shift in orientation tuning of voxel-level fMRI data. However, NMM indicated multiplicative gain rather than an additive shift, in line with single-cell electrophysiology. Beyond orientation, this approach could be applied to determine the form of neuromodulation in any fMRI experiment, provided that the experiment tests multiple points along a stimulus dimension to which neurons are tuned (e.g., direction of motion, isoluminant hue, pitch, etc.). Significance Statement The spatial resolution afforded by noninvasive neuroimaging in humans continues to improve, but the best available resolution is insufficient for testing theories in cognitive neuroscience; many theories are specified at the level of individual neurons, but magnetic resonance imaging aggregates over hundreds of thousands of neurons. With limited resolution, it is unclear how to test assumptions and predictions of these theories in humans. To bridge this gap, we developed a modeling framework that allows researchers to infer a key property of the neural code -- how stimulus features and cognitive states modulate neural tuning – given only noninvasive neuroimaging data. The framework is broadly applicable to constrain and test theories that link changes in behavior to changes in neural tuning.
Traditional state trace analyses assess the latent dimensionality of a cognitive process by asking whether the means of two dependent variables conform to a monotonic function across a set of conditions. Recently proposed methods test whether a function’s deviation from monotonicity is statistically significant (e.g., Kalish et al. 2016, Davis-Stober et al. 2017). However, these tests assume trial-level independence between the two measures, but violations of this assumption can lead to incorrect conclusions. To address these limitations, we developed a hierarchical Bayesian model that factors out the separate roles of subject dependencies, item dependencies, and trial-level dependence via three separate bivariate normal distributions, capturing each type of dependency between the two measures. This is performed with separate models that do, or do not allow a non-monotonic relation between the condition effects (i.e., same vs. different rank orders). The Widely Applicable Information Criterion (WAIC) – a cross validation measure of model fit – is then used to assess the reliability of the model comparison, providing a statistical conclusion regarding the dimensionality of the latent psychological space. We validated this new state trace analysis technique using model recovery simulation studies, which assumed different ground truths regarding monotonicity and the direction/magnitude of the trial-level dependence. We also provide an application of this new technique to an implicit learning study that compared performance on an implicit retrieval task (forced choice recognition) versus an explicit retrieval task (cued recall).
Knowing the identity of an object can powerfully alter perception. Visual demonstrations of this-such as Gregory's (1970) hidden Dalmatian-affirm the existence of both top-down and bottom-up processing. We consider a third processing pathway: lateral connections between the parts of an object. Lateral associations are assumed by theories of object processing and hierarchical theories of memory, but little evidence attests to them. If they exist, their effects should be observable even in the absence of object identity knowledge. We employed Continuous Flash Suppression (CFS) while participants studied object images, such that visual details were learned without explicit object identification. At test, lateral associations were probed using a part-to-part matching task. We also tested whether part-whole links were facilitated by prior study using a part-naming task, and included another study condition (Word), in which participants saw only an object's written name. The key question was whether CFS study (which provided visual information without identity) would better support part-to-part matching (via lateral associations) whereas Word study (which provided identity without the correct visual form) would better support part-naming (via top-down processing). The predicted dissociation was found and confirmed by state-trace analyses. Thus, lateral part-to-part associations were learned and retrieved independently of object identity representations. This establishes novel links between perception and memory, demonstrating that (a) lateral associations at lower levels of the object identification hierarchy exist and contribute to object processing and (b) these associations are learned via rapid, episodic-like mechanisms previously observed for the high-level, arbitrary relations comprising episodic memories. (PsycINFO Database Record (c) 2019 APA, all rights reserved).