The shape and spacing of facial features are essential for the perception of emotional expressions and the identification of individuals. The surface quality of faces, such as skin texture, eye twinkle, and hair gloss, is valuable for estimating health status. A previous fMRI study in monkeys revealed that regions selective for faces overlapped regions selective for object gloss in the central inferior temporal (IT) cortex, suggesting that the surface quality of faces is processed in area TE. However, the representation of facial surface quality has not been directly examined in this area. To understand neural processing of surface quality in face images, neuronal activity was recorded in area TE of three monkeys while face images with different expressions, identities, and surface qualities were presented. The surface qualities included gloss modulations - high-gloss and low-gloss, and “style-transfer” where facial texture was replaced with fabric. These representations were compared to those in several models, including Convolutional Neural Networks and Vision Transformer. Our results revealed that the style-transferred face images strongly modulated activity in TE neurons and that some neurons were tuned to monkey facial expressions, whereas others were more influenced by surface quality. In contrast, the models showed relatively large contributions of surface-quality representation, particularly to the separation between the original and style-transferred faces, compared with those of facial expression and identity. This trend was observed even in models trained on the stylized ImageNet dataset, which was designed to reduce reliance on texture-based classification. These results suggest a strong influence of image style transfer on both TE neurons and models, the diversity of the facial representations by TE neurons, and the tendency of models to emphasize the style-transferred images regardless of training. These findings highlight a potential limitation in current artificial vision models and underscore the value of examining biological representations to inspire more robust and adaptable visual recognition in future model designs.
Object categorization is a fundamental visual function, via which primates group items based on perceptual similarity. Neurons that respond to a class of complex objects, such as faces, can be found in inferior temporal cortex of macaque monkeys, comprising areas TEO and TE. The ability of monkeys to categorize cat/dog images is greatly impaired when both TE and TEO are removed, but is only modestly impaired if either region is left intact. This suggests that both TE and TEO can support object categorization. We investigated what differences exist in category information processing between areas TEO and TE. For cat and dog stimulus images, we found that category decoding performance increased during the initial phase of a stimulus presentation, then remained stable in area TEO for the duration of the presentation in a passive fixation task. In area TE, category decoding performance continued to improve into later in the time window than in TEO. Furthermore, we found that, after cat/dog category training, area TE neuronal populations encode cat and dog category information more strongly than do TEO neurons even in a fixation task (Mann-Whitney U-test, p < 0.05). Together, our results suggest that area TEO processes category information without changing its representation, whereas the category information representation in area TE evolves over time (both within a trial and across category training sessions), indicating that responses in TE may be influenced by top-down feedback.
Decision-making is typically influenced by past choices. Previous studies have shown that the activity of neurons in the orbitofrontal cortex (OFC), which is involved in reward value coding, can be modulated by the reward values of previously chosen options. To further explore how the OFC processes the history of choice-related factors (HCF) from past choices, we analyzed the neuronal activity recorded in the monkey OFC during a decision-making task in which the monkeys could choose between two options based on their reward values. We found that the activity of over 70% of neurons was better explained by models that incorporated HCF than by models without it. The activity of these neurons during current choice was influenced by HCF in the previous trial. Additionally, some types of HCF were represented in these neurons between trials, suggesting that this information was maintained in the OFC and used to guide future decisions. Furthermore, a significant difference was observed in the choice reaction times between trials sorted by HCF. Pharmacological inactivation of the OFC by muscimol infusion eliminated such behavioral differences. These results indicate that HCF-modulated OFC activity contributes to the behavioral bias during current decision-making.
Primates can rapidly categorize images by visual feature similarity. We previously showed that inferior temporal cortex (IT) subregions TEO and TE contribute to visual categorization to differing extents. To investigate the neural plasticity underlying visual categorization, we recorded simultaneously from both areas while two male monkeys learned a visual categorization task. Category specificity and generalization are initially stronger in TEO than in TE but increase with learning only in the TE neuronal population, whose neural representations correlate with behavioral performance. TE and TEO can contribute complementary, partially independent category information. However, TEO does not add learning-relevant variance across days. TE exhibits greater categorical enhancement for correct than error trials compared with TEO. Single-neuron analyses revealed that TE category selectivity strengthens with learning, primarily through enhanced encoding of one category. Combined with lesion evidence, our results suggest that TE plasticity likely plays a more fundamental role in supporting visual category learning.
Area TE is required for normal learning of visual categories based on perceptual similarity. To evaluate whether category learning changes neural activity in area TE, we trained two monkeys (both male) implanted with multielectrode arrays to categorize natural images of cats and dogs. Neural activity during a passive viewing task was compared pre- and post-training. After the category training, the accuracy of abstract category decoding improved. Single units became more category selective, the proportion of single units with category selectivity increased, and units sustained their category-specific responses for longer. Visual category learning thus appears to enhance category separability in area TE by driving changes in the stimulus selectivity of individual neurons and by recruiting more units to the active network.
This paper proposes a time-series data embedding technique that preserves curvature and orientation, with a focus on visualizing temporal manifold-valued data. Manifold-valued data provide pair-wise local distances on which the proposed method is built. First, we introduce a simpler form of our method, conformal folding embedding (CFE), as an interpretable straightforward algorithm to perfectly preserve the angles between adjacent velocity vectors, maintaining local geometric structure while forgoing global structure. Then we introduce the general formulation, dubbed local distance correlation embedding (LDCE), that maximizes distance correlation between input manifoldvalued data and the embedded ones while preserving both global distance structure and local geometric structure. Although the two algorithms are different in formulation, we show their theoretical connection by proving that CFE is a special case of LDCE. We empirically showcase the effectiveness of LDCE in preserving curvature/orientation by visualizing simulated data. The method is also applied to analyze the temporal information encoded at a population level in the inferior temporal cortex of monkeys.
In the canonical view of visual processing the neural representation of complex objects emerges as visual information is integrated through a set of convergent, hierarchically organized processing stages, ending in the primate inferior temporal lobe. It seems reasonable to infer that visual perceptual categorization requires the integrity of anterior inferior temporal cortex (area TE). Many deep neural networks (DNNs) are structured to simulate the canonical view of hierarchical processing within the visual system. However, there are some discrepancies between DNNs and the primate brain. Here we evaluated the performance of a simulated hierarchical model of vision in discriminating the same categorization problems presented to monkeys with TE removals. The model was able to simulate the performance of monkeys with TE removals in the categorization task but performed poorly when challenged with visually degraded stimuli. We conclude that further development of the model is required to match the level of visual flexibility present in the monkey visual system.
Visual short-term memory is an important ability of primates and is thought to be stored in area TE. We previously reported that the initial transient responses of neurons in area TE represented information about a global category of faces, e.g., monkey faces vs. human faces vs. simple shapes, and the latter part of the responses represented information about fine categories, e.g., facial expression. The neuronal mechanisms of hierarchical categorization in area TE remain unknown. For this study, we constructed a combined model that consisted of a deep neural network (DNN) and a recurrent neural network and investigated whether this model can replicate the time course of hierarchical categorization. The visual images were stored in the recurrent connections of the model. When the visual images with noise were input to the model, the model outputted the time course of the hierarchical categorization. This result indicates that recurrent connections in the model are important not only for visual short-term memory but for hierarchical categorization, suggesting that recurrent connections in area TE are important for hierarchical categorization.
Feed-forward deep neural networks have better performance in object categorization tasks than other models of computer vision. To understand the relationship between feed-forward deep networks and the primate brain, we investigated representations of upright and inverted faces in a convolutional deep neural network model and compared them with representations by neurons in the monkey anterior inferior-temporal cortex, area TE. We applied principal component analysis to feature vectors in each model layer to visualize the relationship between the vectors of the upright and inverted faces. The vectors of the upright and inverted monkey faces were more separated through the convolution layers. In the fully-connected layers, the separation among human individuals for upright faces was larger than for inverted faces. The Spearman correlation between each model layer and TE neurons reached a maximum at the fully-connected layers. These results indicate that the processing of faces in the fully-connected layers might resemble the asymmetric representation of upright and inverted faces by the TE neurons. The separation of upright and inverted faces might take place by feed-forward processing in the visual cortex, and separations among human individuals for upright faces, which were larger than those for inverted faces, might occur in area TE.
This study proposes a methodology for constructing a music database for research purposes. We focused on the feasibility of an efficient and reliable technique to collect music stimuli that induce a variety of emotions. We selected iconic phrases from 4 famous classical pieces with 4 controls of scales, respectively. We modified each piece into 4 categories by changing the properties in terms of “mode” (converting a major piece to a minor, or vice versa) and “tempo” (changing the BPM speed faster or slower). We verified our method by a Music Information Retrieval (MIR) system. The MIR analyses showed that most of the pieces were successfully positioned in the intended categories (major/minor mode at a high/low tempo). The result suggests that this may be an efficient method to construct an objective music database that is independent of the psychological evaluations.
Choice reflects the values of available alternatives; more valuable options are chosen more often than less valuable ones. Here we studied whether neuronal responses in orbitofrontal cortex (OFC) reflect the value difference between options, and whether there is a causal link between OFC neuronal activity and choice. Using a decision-making task where two visual stimuli were presented sequentially, each signifying a value, we showed that when the second stimulus appears many neurons encode the value difference between alternatives. Later when the choice occurs, that difference signal disappears and a signal indicating the chosen value emerges. Pharmacological inactivation of OFC neurons coding for choice-related values increases the monkey’s latency to make a choice and the likelihood that it will choose the less valuable alternative, when the value difference is small. Thus, OFC neurons code for value information that could be used to directly influence choice.
Visual object recognition requires both visual sensory information and memory, and its mechanisms are often studied using old-world monkeys. Wittig et al. (2014, 2016) reported that Rhesus monkeys and humans seem to adopt different strategies in a short-term visual memory task. The Rhesus monkeys seemed to rely on recency of stimulus repetition, whereas humans relied on specific memorization. We conducted experiments using a sequential delayed match-to-sample task with random dot visual noise using Rhesus and Japanese monkeys and found that recency effect was observed in both species. There were differences in the noise effect on behavioral performances across species.
There is an on-going debate over whether area TE, or the anatomically adjacent rhinal cortex, is the final stage of visual object processing. Both regions have been implicated in visual perception, but their involvement in non-perceptual functions, such as short-term memory, hinders clear-cut interpretation. Here, using a two-interval forced choice task without a short-term memory demand, we find that after bilateral removal of area TE, monkeys trained to categorize images based on perceptual similarity (morphs between dogs and cats), are, on the initial viewing, badly impaired when given a new set of images. They improve markedly with a small amount of practice but nonetheless remain moderately impaired indefinitely. The monkeys with bilateral removal of rhinal cortex are, under all conditions, indistinguishable from unoperated controls. We conclude that the final stage of the integration of visual perceptual information into object percepts in the ventral visual stream occurs in area TE.
We recognize objects even when they are partially degraded by visual noise. We studied the relation between the amount of visual noise (5, 10, 15, 20, or 25%) degrading 8 black-and-white stimuli and stimulus identification in 2 monkeys performing a sequential delayed match-to-sample task. We measured the accuracy and speed with which matching stimuli were identified. The performance decreased slightly (errors increased) as the amount of visual noise increased for both monkeys. The performance remained above 80% correct, even with 25% noise. However, the reaction times markedly increased as the noise increased, indicating that the monkeys took progressively longer to decide what the correct response would be as the amount of visual noise increased, showing that the monkeys trade time to maintain accuracy. Thus, as time unfolds the monkeys act as if they are accumulating the information and/or testing hypotheses about whether the test stimulus is likely to be a match for the sample being held in short-term memory. We recorded responses from 13 single neurons in area TE of the 2 monkeys. We found that stimulus-selective information in the neuronal responses began accumulating when the match stimulus appeared. We found that the greater the amount of noise obscuring the test stimulus, the more slowly stimulus-related information by the 13 neurons accumulated. The noise induced slowing was about the same for both behavior and information. These data are consistent with the hypothesis that area TE neuron population carries information about stimulus identity that accumulates over time in such a manner that it progressively overcomes the signal degradation imposed by adding visual noise.
In primates, visual recognition of complex objects depends on the inferior temporal lobe. By extension, categorizing visual stimuli based on similarity ought to depend on the integrity of the same area. We tested three monkeys before and after bilateral anterior inferior temporal cortex (area TE) removal. Although mildly impaired after the removals, they retained the ability to assign stimuli to previously learned categories, e.g., cats versus dogs, and human versus monkey faces, even with trial-unique exemplars. After the TE removals, they learned in one session to classify members from a new pair of categories, cars versus trucks, as quickly as they had learned the cats versus dogs before the removals. As with the dogs and cats, they generalized across trial-unique exemplars of cars and trucks. However, as seen in earlier studies, these monkeys with TE removals had difficulty learning to discriminate between two simple black and white stimuli. These results raise the possibility that TE is needed for memory of simple conjunctions of basic features, but that it plays only a small role in generalizing overall configural similarity across a large set of stimuli, such as would be needed for perceptual categorical assignment. SIGNIFICANCE STATEMENT The process of seeing and recognizing objects is attributed to a set of sequentially connected brain regions stretching forward from the primary visual cortex through the temporal lobe to the anterior inferior temporal cortex, a region designated area TE. Area TE is considered the final stage for recognizing complex visual objects, e.g., faces. It has been assumed, but not tested directly, that this area would be critical for visual generalization, i.e., the ability to place objects such as cats and dogs into their correct categories. Here, we demonstrate that monkeys rapidly and seemingly effortlessly categorize large sets of complex images (cats vs dogs, cars vs trucks), surprisingly, even after removal of area TE, leaving a puzzle about how this generalization is done.
The ability to recognize faces is reduced with a picture-plane inversion of the faces, known as the face inversion effect. It has been reported that the configuration of facial features, for example, the distance between the eyes and mouth, becomes less perceptible when the face is inverted. In macaque monkeys, designated cortical areas, i.e., face patches, where face images are processed, have been found in the temporal visual cortex along the ventral visual pathway. Neurons in the anterior face patch (anterior part of the inferior temporal cortex) are known to encode view-invariant identity information. Thus, the anterior face patch is believed to be the final processing stage in the face patch system. A recent study showed that the face-inversion decreases the amount of the information about facial identity and facial expression conveyed by neurons, though it did not affect the information about the global category of the stimulus images (monkey versus human versus shape). The anterior face patch may, therefore, serve as the neural basis underlying the face inversion effect.
When an individual chooses one item from two or more alternatives, they compare the values of the expected outcomes. The outcome value can be determined by the associated reward amount, the probability of reward, and the workload required to earn the reward. Rational choice theory states that choices are made to maximize rewards over time, and that the same outcome values lead to an equal likelihood of choices. However, the theory does not distinguish between conditions with the same reward value, even when acquired under different circumstances, and does not always accurately describe real behavior. We have found that allowing a monkey to choose a reward schedule endows the schedule with extra value when compared to performance in an identical schedule that is chosen by another agent (a computer here). This behavior is not consistent with pure rational choice theory. Theoretical analysis using a modified temporal-difference learning model showed an enhanced schedule state value by self-choice. These results suggest that an increased reward value underlies the improved performances by self-choice during reward-seeking behavior.
To investigate the effect of face inversion and thatcherization (eye inversion) on temporal processing stages of facial information, single neuron activities in the temporal cortex (area TE) of two rhesus monkeys were recorded. Test stimuli were colored pictures of monkey faces (four with four different expressions), human faces (three with four different expressions), and geometric shapes. Modifications were made in each face-picture, and its four variations were used as stimuli: upright original, inverted original, upright thatcherized, and inverted thatcherized faces. A total of 119 neurons responded to at least one of the upright original facial stimuli. A majority of the neurons (71%) showed activity modulations depending on upright and inverted presentations, and a lesser number of neurons (13%) showed activity modulations depending on original and thatcherized face conditions. In the case of face inversion, information about the fine category (facial identity and expression) decreased, whereas information about the global category (monkey vs human vs shape) was retained for both the original and thatcherized faces. Principal component analysis on the neuronal population responses revealed that the global categorization occurred regardless of the face inversion and that the inverted faces were represented near the upright faces in the principal component analysis space. By contrast, the face inversion decreased the ability to represent human facial identity and monkey facial expression. Thus, the neuronal population represented inverted faces as faces but failed to represent the identity and expression of the inverted faces, indicating that the neuronal representation in area TE cause the perceptual effect of face inversion.