Abstract Neuroscientific studies rely heavily on a-priori hypotheses, which can bias results toward existing theories. Here, we use a hypothesis-neutral approach to study category selectivity in higher visual cortex. Using only stimulus images and their associated fMRI activity, we constrain randomly initialized neural networks to predict voxel activity. Despite no category-level supervision, units in the trained networks act as detectors for semantic concepts like ‘faces’ or ‘words’, providing solid empirical support for categorical selectivity. Importantly, this selectivity is mostly maintained when training the networks without images that contain the preferred category, strongly suggesting that selectivity is not domain-specific machinery, but sensitivity to generic patterns that characterize preferred categories. The ability of the models’ representations to transfer to perceptual tasks further reveals the functional role of their selective responses. Finally, our models show selectivity only for a limited number of categories, all previously identified, suggesting that the essential categories are already known.
Abstract Characterizing the fine-grained functional organization of human higher visual cortex remains a central challenge, as traditional neuroimaging experiments constrain the diversity of stimuli that can be sampled. In prior work we addressed this challenge by developing a novel data-driven tool, termed “BrainDiVE” (Luo et al. 2023), which synthesizes naturalistic images predicted to strongly activate specific brain regions. BrainDiVE leverages pretrained image diffusion models guided by gradients from an image-computable fMRI encoding model. Here, we experimentally validated BrainDiVE by generating images predicted to maximally activate different functional regions of interest and then presenting them to new participants (n=12) in an fMRI study. The model-generated images elicited robust, spatially specific responses in the targeted brain regions, producing significantly greater category selectivity than natural images, validating the method’s ability to capture generalizable neural tuning properties in human ventral visual cortex. We further showed that region-targeted images exaggerate specific sets of low-level and mid-level image statistics, suggesting that category-selective regions are tuned to continuous directions in feature space. Moreover, we demonstrated fine-grained experimental control by differentially activating two face-selective regions, the occipital face area (OFA) and fusiform face area (FFA), providing additional evidence that these regions encode distinct aspects of faces. Finally, we identify a posterior-to-anterior functional gradient within the occipital place area (OPA), suggesting topographic organization based on scene properties such as distance and indoor-outdoor location. These findings enhance our understanding of the representational structure of category-selective regions and introduce a new paradigm for probing neural selectivity in human visual cortex.
Mapping the relationship between neural activity and motor behavior is a central aim of sensorimotor neuroscience and neurotechnology. While most progress to this end has relied on restricting complexity, the advent of foundation models instead proposes integrating a breadth of data as an alternate avenue for broadly advancing downstream modeling. We quantify this premise for motor decoding from intracortical microelectrode data, pretraining an autoregressive Transformer on 2000 hours of neural population spiking activity paired with diverse motor covariates from over 30 monkeys and humans. The resulting model is broadly useful, benefiting decoding on 8 downstream decoding tasks and generalizing to a variety of neural distribution shifts. However, we also highlight that scaling autoregressive Transformers seems unlikely to resolve limitations stemming from sensor variability and output stereotypy in neural datasets. Code: https://github.com/joel99/ndt3.
Current non-invasive neuroimaging techniques trade off between spatial resolution and temporal resolution. While magnetoencephalography (MEG) can capture rapid neural dynamics and functional magnetic resonance imaging (fMRI) can spatially localize brain activity, a unified picture that preserves both high resolutions remains an unsolved challenge with existing source localization or MEG-fMRI fusion methods, especially for single-trial naturalistic data. We collected whole-head MEG when subjects listened passively to more than seven hours of narrative stories, using the same stimuli in an open fMRI dataset (LeBel et al., 2023). We developed a transformer-based encoding model that combines the MEG and fMRI from these two naturalistic speech comprehension experiments to estimate latent cortical source responses with high spatiotemporal resolution. Our model is trained to predict MEG and fMRI from multiple subjects simultaneously, with a latent layer that represents our estimates of reconstructed cortical sources. Our model predicts MEG better than the common standard of single-modality encoding models, and it also yields source estimates with higher spatial and temporal fidelity than classic minimum-norm solutions in simulation experiments. We validated the estimated latent sources by showing its strong generalizability across unseen subjects and modalities. Estimated activity in our source space predict electrocorticography (ECoG) better than an ECoG-trained encoding model in an entirely new dataset. By integrating the power of large naturalistic experiments, MEG, fMRI, and encoding models, we propose a practical route towards millisecond-and-millimeter brain mapping.
Understanding how the human brain processes naturalistic audiovisual information remains a central challenge in cognitive neuroscience. Progress has been limited, however, by the difficulty of modeling complex audiovisual feature spaces - most prior work has therefore relied on short, controlled stimuli, or on stimuli from one modality at a time, leaving cortical mechanisms that support real-world comprehension poorly characterized. Further, while recent advances in artificial intelligence now enable the extraction of high-dimensional, time-resolved features from naturalistic stimuli, how cortical regions dynamically process auditory and visual information as time unfolds remains largely unexplored. Using large-scale fMRI data collected while participants watched movies, we developed two complementary computational approaches relying on prediction performance to map the moment-by-moment dynamics of sensory processing across cortical regions: one tracks when one modality predicts a region's activity substantially better than the other, capturing temporal transitions in modality dominance, while the other identifies periods when both modalities predict the region well, indicating balanced representation of auditory and visual information. Together, these analyses reveal two complementary patterns of audiovisual organization across cortex: a pair of "bows" of modality switching - one posterior bow encircling category-selective visual cortex and another anterior bow spanning dorso-lateral frontal areas, and an arrow-like axis of bimodally predicted regions extending from lateral occipital cortex into the temporal cortex. The coexistence of these systems points to a cortical architecture that flexibly reweights sensory inputs while maintaining balanced multimodal representations, supporting robust comprehension of complex natural events. More broadly, this work illustrates how naturalistic neuroimaging experiments informed by modern machine learning approaches can reveal new principles of dynamic audiovisual processing in the human brain.
Understanding how the human brain processes naturalistic audiovisual information remains a central challenge in cognitive neuroscience. Progress has been limited, however, by the difficulty of modeling complex audiovisual feature spaces - most prior work has therefore relied on short, controlled stimuli, or on stimuli from one modality at a time, leaving cortical mechanisms that support real-world comprehension poorly characterized. Further, while recent advances in artificial intelligence now enable the extraction of high-dimensional, time-resolved features from naturalistic stimuli, how cortical regions dynamically process auditory and visual information as time unfolds remains largely unexplored. Using large-scale fMRI data collected while participants watched movies, we developed two complementary computational approaches relying on prediction performance to map the moment-by-moment dynamics of sensory processing across cortical regions: one tracks when one modality predicts a region’s activity substantially better than the other, capturing temporal transitions in modality dominance, while the other identifies periods when both modalities predict the region well, indicating balanced representation of auditory and visual information. Together, these analyses reveal two complementary patterns of audiovisual organization across cortex: a pair of “bows” of modality switching - one posterior bow encircling category-selective visual cortex and another anterior bow spanning dorso-lateral frontal areas, and an arrow-like axis of bimodally predicted regions extending from lateral occipital cortex into the temporal cortex. The coexistence of these systems points to a cortical architecture that flexibly reweights sensory inputs while maintaining balanced multimodal representations, supporting robust comprehension of complex natural events. More broadly, this work illustrates how naturalistic neuroimaging experiments informed by modern machine learning approaches can reveal new principles of dynamic audiovisual processing in the human brain. ### Competing Interest Statement The authors have declared no competing interest. National Science Foundation Career grant, 237064
We introduce BrainSAIL (Semantic Attribution and Image Localization), a method for linking neural selectivity with spatially distributed semantic visual concepts in natural scenes. BrainSAIL leverages recent advances in large-scale artificial neural networks, using them to provide insights into the functional topology of the brain. To overcome the challenge presented by the co-occurrence of multiple categories in natural images, BrainSAIL exploits semantically consistent, dense spatial features from pre-trained vision models, building upon their demonstrated ability to robustly predict neural activity. This method derives clean, spatially dense embeddings without requiring any additional training, and employs a novel denoising process that leverages the semantic consistency of images under random augmentations. By unifying the space of whole-image embeddings and dense visual features and then applying voxel-wise encoding models to these features, we enable the identification of specific subregions of each image which drive selectivity patterns in different areas of the higher visual cortex. This provides a powerful tool for dissecting the neural mechanisms that underlie semantic visual processing for natural images. We validate BrainSAIL on cortical regions with known category selectivity, demonstrating its ability to accurately localize and disentangle selectivity to diverse visual concepts. Next, we demonstrate BrainSAIL's ability to characterize high-level visual selectivity to scene properties and low-level visual features such as depth, luminance, and saturation, providing insights into the encoding of complex visual information. Finally, we use BrainSAIL to directly compare the feature selectivity of different brain encoding models across different regions of interest in visual cortex. Our innovative method paves the way for significant advances in mapping and decomposing high-level visual representations in the human brain.
Understanding functional representations within higher visual cortex is a fundamental question in computational neuroscience. While artificial neural networks pretrained on large-scale datasets exhibit striking representational alignment with human neural responses, learning image-computable models of visual cortex relies on individual-level, large-scale fMRI datasets. The necessity for expensive, time-intensive, and often impractical data acquisition limits the generalizability of encoders to new subjects and stimuli. BraInCoRL uses in-context learning to predict voxelwise neural responses from few-shot examples without any additional finetuning for novel subjects and stimuli. We leverage a transformer architecture that can flexibly condition on a variable number of in-context image stimuli, learning an inductive bias over multiple subjects. During training, we explicitly optimize the model for in-context learning. By jointly conditioning on image features and voxel activations, our model learns to directly generate better performing voxelwise models of higher visual cortex. We demonstrate that BraInCoRL consistently outperforms existing voxelwise encoder designs in a low-data regime when evaluated on entirely novel images, while also exhibiting strong test-time scaling behavior. The model also generalizes to an entirely new visual fMRI dataset, which uses different subjects and fMRI data acquisition parameters. Further, BraInCoRL facilitates better interpretability of neural signals in higher visual cortex by attending to semantically relevant stimuli. Finally, we show that our framework enables interpretable mappings from natural language queries to voxel selectivity.
How does the brain process language over time? Research suggests that natural human language is processed hierarchically across brain regions over time. However, attempts to characterize this computation have thus far been limited to tightly controlled experimental settings that capture only a coarse picture of the brain dynamics underlying human natural language comprehension. The recent emergence of LLM encoding models promises a new avenue to discover and characterize rich semantic information in the brain, yet interpretable methods for linking information in LLMs to language processing over time are limited. In this work, we develop a low-rank tensor regression method to decompose LLM encoding models into interpretable components of semantics, time, and brain region activation, and apply the method to a Magnetoencephalography (MEG) dataset in which subjects listened to narrative stories. With only a few components, we show improved performance compared to a standard ridge regression encoding model, suggesting the low-rank models provide a good inductive bias for language encoding. In addition, our method discovers a diverse spectrum of interpretable response components that are sensitive to a rich set of low-level and semantic language features, showing that our method is able to separate distinct language processing features in neural signals. After controlling for low-level audio and sentence features, we demonstrate better capture of semantic features. Through use of low-rank tensor encoding models we are able to decompose neural responses to language features, showing improved encoding performance and interpretable processing components, suggesting our method as a useful tool for uncovering language processes in naturalistic settings.
Several recent studies, enabled by advances in neuroimaging methods and large-scale datasets, have identified areas in human ventral visual cortex that respond more strongly to food images than to images of many other categories, adding to our knowledge about the broad network of regions that are responsive to food. This finding raises important questions about the evolutionary and developmental origins of a possible food-selective neural population, as well as larger questions about the origins of category-selective neural populations more generally. Here, we propose a framework for how visual properties of food (particularly color) and nonvisual signals associated with multimodal reward processing, social cognition, and physical interactions with food may, in combination, contribute to the emergence of food selectivity. We discuss recent research that sheds light on each of these factors, alongside a broader account of category selectivity that incorporates both visual feature statistics and behavioral relevance.
The human language system represents both linguistic forms and meanings, but the abstractness of the meaning representations remains debated. Here, we searched for abstract representations of meaning in the language cortex by modeling neural responses to sentences using representations from vision and language models. When we generate images corresponding to sentences and extract vision model embeddings, we find that aggregating across multiple generated images yields increasingly accurate predictions of language cortex responses, sometimes rivaling large language models. Similarly, averaging embeddings across multiple paraphrases of a sentence improves prediction accuracy compared to any single paraphrase. Enriching paraphrases with contextual details that may be implicit (e.g., augmenting "I had a pancake" to include details like "maple syrup") further increases prediction accuracy, even surpassing predictions based on the embedding of the original sentence, suggesting that the language system maintains richer and broader semantic representations than language models. Together, these results demonstrate the existence of highly abstract, form-independent meaning representations within the language cortex.
Do machines and humans process language in similar ways? A recent line of research has hinted in the affirmative, demonstrating that human brain signals can be effectively predicted using the internal representations of language models (LMs). This is thought to reflect shared computational principles between LMs and human language processing. However, there are also clear differences in how LMs and humans acquire and use language, even if the final task they are performing is the same. Despite this, there is little work exploring systematic differences between human and machine language processing using brain data. To address this question, we examine the differences between LM representations and the human brain's responses to language, specifically by examining a dataset of Magnetoencephalography (MEG) responses to a written narrative. In doing so we identify three phenomena that, in prior work, LMs have been found to not capture well: emotional understanding, figurative language processing, and physical commonsense. We further fine-tune models on datasets related to these three phenomena, and find that LMs fine-tuned on tasks related to emotion and figurative language show improved alignment with brain responses. We emphasize the importance of understanding not just similarities between human and machine language processing, but also differences. Our work takes the first steps toward this goal in the context of narrative reading.
Understanding the functional organization of higher visual cortex is a central focus in neuroscience. Past studies have primarily mapped the visual and semantic selectivity of neural populations using hand-selected stimuli, which may potentially bias results towards pre-existing hypotheses of visual cortex functionality. Moving beyond conventional approaches, we introduce a data-driven method that generates natural language descriptions for images predicted to maximally activate individual voxels of interest. Our method -- Semantic Captioning Using Brain Alignments ("BrainSCUBA") -- builds upon the rich embedding space learned by a contrastive vision-language model and utilizes a pre-trained large language model to generate interpretable captions. We validate our method through fine-grained voxel-level captioning across higher-order visual regions. We further perform text-conditioned image synthesis with the captions, and show that our images are semantically coherent and yield high predicted activations. Finally, to demonstrate how our method enables scientific discovery, we perform exploratory investigations on the distribution of "person" representations in the brain, and discover fine-grained semantic selectivity in body-selective areas. Unlike earlier studies that decode text, our method derives *voxel-wise captions of semantic selectivity*. Our results show that BrainSCUBA is a promising means for understanding functional preferences in the brain, and provides motivation for further hypothesis-driven investigation of visual cortex.
Do machines and humans process language in similar ways? Recent research has hinted at the affirmative, showing that human neural activity can be effectively predicted using the internal representations of language models (LMs). Although such results are thought to reflect shared computational principles between LMs and human brains, there are also clear differences in how LMs and humans represent and use language. In this work, we systematically explore the divergences between human and machine language processing by examining the differences between LM representations and human brain responses to language as measured by Magnetoencephalography (MEG) across two datasets in which subjects read and listened to narrative stories. Using an LLM-based data-driven approach, we identify two domains that LMs do not capture well: social/emotional intelligence and physical commonsense. We validate these findings with human behavioral experiments and hypothesize that the gap is due to insufficient representations of social/emotional and physical knowledge in LMs. Our results show that fine-tuning LMs on these domains can improve their alignment with human brain responses.
Language neuroscience currently relies on two major experimental paradigms: controlled experiments using carefully hand-designed stimuli, and natural stimulus experiments. These approaches have complementary advantages which allow them to address distinct aspects of the neurobiology of language, but each approach also comes with drawbacks. Here we discuss a third paradigm—in silico experimentation using deep learning-based encoding models—that has been enabled by recent advances in cognitive computational neuroscience. This paradigm promises to combine the interpretability of controlled experiments with the generalizability and broad scope of natural stimulus experiments. We show four examples of simulating language neuroscience experiments in silico and then discuss both the advantages and caveats of this approach.
Understanding the semantics of visual scenes is a fundamental challenge in Computer Vision. A key aspect of this challenge is that objects sharing similar semantic meanings or functions can exhibit striking visual differences, making accurate identification and categorization difficult. Recent advancements in text-to-image frameworks have led to models that implicitly capture natural scene statistics. These frameworks account for the visual variability of objects, as well as complex object co-occurrences and sources of noise such as diverse lighting conditions. By leveraging large-scale datasets and cross-attention conditioning, these models generate detailed and contextually rich scene representations. This capability opens new avenues for improving object recognition and scene understanding in varied and challenging environments. Our work presents StableSemantics, a dataset comprising 224 thousand human-curated prompts, processed natural language captions, over 2 million synthetic images, and 10 million attention maps corresponding to individual noun chunks. We explicitly leverage human-generated prompts that correspond to visually interesting stable diffusion generations, provide 10 generations per phrase, and extract cross-attention maps for each image. We explore the semantic distribution of generated images, examine the distribution of objects within images, and benchmark captioning and open vocabulary segmentation methods on our data. To the best of our knowledge, we are the first to release a diffusion dataset with semantic attributions. We expect our proposed dataset to catalyze advances in visual semantic understanding and provide a foundation for developing more sophisticated and effective visual models. Website: https://stablesemantics.github.io/StableSemantics
Relating brain activity associated with a complex stimulus to different properties of that stimulus is a powerful approach for constructing functional brain maps. However, when stimuli are naturalistic, their properties are often correlated (e.g., visual and semantic features of natural images, or different layers of a convolutional neural network that are used as features of images). Correlated properties can act as confounders for each other and complicate the interpretability of brain maps, and can impact the robustness of statistical estimators. Here, we present an approach for brain mapping based on two proposed methods: stacking different encoding models and structured variance partitioning. Our stacking algorithm combines encoding models that each uses as input a feature space that describes a different stimulus attribute. The algorithm learns to predict the activity of a voxel as a linear combination of the outputs of different encoding models. We show that the resulting combined model can predict held-out brain activity better or at least as well as the individual encoding models. Further, the weights of the linear combination are readily interpretable; they show the importance of each feature space for predicting a voxel. We then build on our stacking models to introduce structured variance partitioning, a new type of variance partitioning that takes into account the known relationships between features. Our approach constrains the size of the hypothesis space and allows us to ask targeted questions about the similarity between feature spaces and brain regions even in the presence of correlations between the feature spaces. We validate our approach in simulation, showcase its brain mapping potential on fMRI data, and release a Python package. Our methods can be useful for researchers interested in aligning brain activity with different layers of a neural network, or with other types of correlated feature spaces.