Category selectivity is a fundamental principle of organization of perceptual brain regions. Human occipitotemporal cortex is subdivided into areas that respond preferentially to faces, bodies, artifacts, and scenes. However, observers need to combine information about objects from different categories to form a coherent understanding of the world. How is this multicategory information encoded in the brain? Studying the multivariate interactions between brain regions of male and female human subjects with fMRI and artificial neural networks, we found that the angular gyrus shows joint statistical dependence with multiple category-selective regions. Adjacent regions show effects for the combination of scenes and each other category, suggesting that scenes provide a context to combine information about the world. Additional analyses revealed a cortical map of areas that encode information across different subsets of categories, indicating that multicategory information is not encoded in a single centralized location, but in multiple distinct brain regions.SIGNIFICANCE STATEMENT Many cognitive tasks require combining information about entities from different categories. However, visual information about different categorical objects is processed by separate, specialized brain regions. How is the joint representation from multiple category-selective regions implemented in the brain? Using fMRI movie data and state-of-the-art multivariate statistical dependence based on artificial neural networks, we identified the angular gyrus encoding responses across face-, body-, artifact-, and scene-selective regions. Further, we showed a cortical map of areas that encode information across different subsets of categories. These findings suggest that multicategory information is not encoded in a single centralized location, but at multiple cortical sites which might contribute to distinct cognitive functions, offering insights to understand integration in a variety of domains.
People make fast and reasonable predictions about the physical behavior of everyday objects. To do so, people may be using principled approximations, similar to models developed by engineers for the purposes of real-time physical simulations. We hypothesize that people use simplified object approximations for tracking and action (the "body" representation), as opposed to fine-grained forms for recognition (the "shape" representation). We used three classic psychophysical tasks (causality perception, collision detection, and change detection) in novel settings that dissociate body and shape. People's behavior across tasks indicates that they rely on approximate bodies for physical reasoning, and that this approximation lies between convex hulls and fine-grained shapes.
Noise is a major challenge for the analysis of fMRI data in general and for connectivity analyses in particular. As researchers develop increasingly sophisticated tools to model statistical dependence between the fMRI signal in different brain regions, there is a risk that these models may increasingly capture artifactual relationships between regions, that are the result of noise. Thus, choosing optimal denoising methods is a crucial step to maximize the accuracy and reproducibility of connectivity models. Most comparisons between denoising methods require knowledge of the ground truth: of what is the 'real signal'. For this reason, they are usually based on simulated fMRI data. However, simulated data may not match the statistical properties of real data, limiting the generalizability of the conclusions. In this article, we propose an approach to evaluate denoising methods using real (non-simulated) fMRI data. First, we introduce an intersubject version of multivariate pattern dependence (iMVPD) that computes the statistical dependence between a brain region in one participant, and another brain region in a different participant. iMVPD has the following advantages: 1) it is multivariate, 2) it trains and tests models on independent partitions of the real fMRI data, and 3) it generates predictions that are both between subjects and between regions. Since whole-brain sources of noise are more strongly correlated within subject than between subjects, we can use the difference between standard MVPD and iMVPD as a 'discrepancy metric' to evaluate denoising techniques (where more effective techniques should yield smaller differences). As predicted, the difference is the greatest in the absence of denoising methods. Furthermore, a combination of removal of the global signal and CompCorr optimizes denoising (among the set of denoising options tested).
Recognizing face and scene images recruits distinct networks of brain regions. Investigating how information is processed and transformed from region to region within these networks is a critical challenge for visual neuroscience. Recent work has introduced techniques that move towards this direction by studying the multivariate statistical dependence between patterns of response (Coutanche and Thompson-Schill 2013, Anzellotti et al. 2017, Anzellotti and Coutanche 2018). As researchers develop increasingly sophisticated tools to model statistical dependence between the fMRI signal there is a risk that models may increasingly capture artifactual relationships between regions. Choosing optimal denoising methods is a crucial step to maximize the accuracy of connectivity models. A common approach to compare denoising methods uses simulated fMRI data, but it is unknown to what extent conclusions drawn using simulated data generalize to real data. To overcome this limitation, we introduce intersubject multivariate pattern dependence (iMVPD) which computes the statistical dependence between a brain region in one participant, and another brain region in a different participant. IMVPD is multivariate, it trains and tests models on independent partitions of the real fMRI data, and it generates predictions that are both between subjects and between regions. Since whole-brain sources of noise are more strongly correlated within subject than between, we can use the difference between standard MVPD and iMVPD as a ‘discrepancy metric’ to evaluate denoising techniques, where more effective techniques should yield smaller differences. As a sanity check, the ‘discrepancy metric’ is the greatest with no denoising. Furthermore, a combination of CompCorr and removal of the global signal optimizes denoising in face- and scene-selective regions (among all denoising options tested on selected regions). In future work, iMVPD can be used for applications like studying individual differences and analyzing types of data where only one region is measured in each participant (i.e. electrophysiology).
Recently, there have been some attempts to use non-recurrent neural models for language modeling. However, a noticeable performance gap still remains. We propose a non-recurrent neural language model, dubbed graph temporal convolutional network (GTCN), that relies on graph neural network blocks and convolution operations. While the standard recurrent neural network language models encode sentences sequentially without modeling higher-level structural information, our model regards sentences as graphs and processes input words within a message propagation framework, aiming to learn better syntactic information by inferring skip-word connections. Specifically, the graph network blocks operate in parallel and learn the underlying graph structures in sentences without any additional annotation pertaining to structure knowledge. Experiments demonstrate that the model without recurrence can achieve comparable perplexity results in language modeling tasks and successfully learn syntactic information.