We conducted two studies to examine how rapid ensemble perceptions of emotion shape judgments of entitativity. We hypothesize that when a perceiver encounters a crowd, they rapidly discern the degree of emotional variability in that group via ensemble coding, and this input provides a strong basis for both immediate and deliberative entitativity inferences. We test this hypothesis using images of a representative sample of natural groups collected via Instagram (R). Each person in each of similar to 200 group images was independently coded for facial affect, gender, and race. Participants evaluated facial emotion variability in each group. We hypothesized that people perceive group variability in facial emotion after brief exposure (150 ms), and that these rapid perceptions predict entitativity judgments of groups viewed for longer periods (3 s). We further hypothesized that rapid ensemble perceptions of emotion variability support the rapid formation of entitativity judgments: entitativity judgments made after 150 ms will approximate those made after 3 s. In Study 1, participants saw images for 5 s each, evaluating (a) shared emotion among group members or (b) entitativity. Study 2 employed a 3 (exposure: 150 ms, 750 ms, 3 s) x 2 (rating: shared emotion, entitativity) x 2 (faces visible: yes, no) between-subjects design. All hypotheses were supported, and the effects were largely independent of the race and gender composition of the groups. These data suggest that complex conceptual judgments about small groups can be formed in the initial instants of perception. We discuss the implications for ensemble coding theories and methods, as well as for entitativity theories.
Automated Facial Expression Recognition (FER) is challenging due to intra-class variations and inter-class similarities. FER can be especially difficult when facial expressions reflect a mixture of various emotions (aka compound expressions). Existing FER datasets, such as AffectNet, provide discrete emotion labels (hard-labels), where a single category of emotion is assigned to an expression. To alleviate inter- and intra-class challenges, as well as provide a better facial expression descriptor, we propose a new approach to create FER datasets through a labeling method in which an image is labeled with more than one emotion (called soft-labels), each with a different confidence. Specifically, we introduce the notion of soft-labels for facial expression datasets, a new approach to affective computing for more realistic recognition of facial expressions. To achieve this goal, we propose a novel methodology to accurately calculate soft-labels: a vector representing the extent to which multiple categories of emotion are simultaneously present within a single facial expression. Finding smoother decision boundaries, enabling multi-labeling, and mitigating bias and imbalanced data are some of the advantages of our proposed method. Building upon AffectNet, we introduce AffectNet+, the next-generation facial expression dataset. This dataset contains soft-labels, three categories of data complexity subsets, and additional metadata such as age, gender, race, head pose, facial landmarks, valence, and arousal. AffectNet+ will be made publicly accessible to researchers.
It is well known that objects become grouped in perceptual organization when they share some visual feature, like a common direction of motion. Less well known is that grouping can change how people perceive a set of objects. For example, when a pair of shapes consistently share a common region of space, their aspect ratios tend to be perceived as more similar (are attracted toward each other). Conversely, when shapes are assigned to different regions in space their aspect ratios repel from each other. Here we examine whether the visual system produce both attractive and repulsive distortions when the state of grouping between a pair of shapes changes on a moment-to-moment basis. Observers viewed a pair of ellipses that differed in terms of how flat or tall they were and reported the aspect ratio of one ellipse from the pair. Each ellipse was defined by a cloud of coherently-moving dots, and the dots within the two ellipses had either the same or different directions of motion, varying from trial-to-trial. We found that the cued ellipse's aspect ratio was reported to be repelled from the aspect ratio of the uncued ellipse when the shapes had different directions of motion compared to when they had the same direction of motion. These results suggest that the visual system can adaptively alter visual experience based on grouping, in particular, repelling the appearance of objects when they do not appear to go together, and it can do so quickly and flexibly.
Interpretation bias, or the threatening appraisal of ambiguous information, has been linked to anxiety disorder. Interpretation bias has been demonstrated for linguistic (e.g., evaluation of ambiguous sentences) and visual judgments (e.g., categorizing emotionally ambiguous facial expressions). It is unclear how these separate components of bias might be associated. We examined linguistic and visual interpretation biases in youth and emerging adults with (n = 44) and without (n = 40) anxiety disorder, and in youth-parent dyads (n = 40). Linguistic and visual biases were correlated with each other, and with anxiety. Compared to non-anxious participants, those with anxiety demonstrated stronger biases, and linguistic bias was especially predictive of anxiety symptoms and diagnosis. Age did not moderate these relationships. Parent linguistic bias was correlated with youth anxiety but not linguistic bias; parent and youth visual biases were correlated. Linguistic and visual interpretation biases are linked in clinically-anxious youth and emerging adults.
Generative Adversarial Networks (GANs) are capable of synthesizing high-quality facial images. Despite their success, GANs do not provide any information about the relationship between the input vectors and the generated images. Currently, facial GANs are trained on imbalanced datasets, which generate less diverse images. For example, more than 77% of 100K images that we randomly synthesized using the StyleGAN3 are classified as Happy, and only around 3% are Angry. The problem even becomes worse when a mixture of facial attributes is desired: less than 1% of the generated samples are Angry Woman, and only around 2% are Happy Black. To address these problems, this paper proposes a framework, called GANalyzer, for the analysis, and manipulation of the latent space of well-trained GANs. GANalyzer consists of a set of transformation functions designed to manipulate latent vectors for a specific facial attribute such as facial Expression, Age, Gender, and Race. We analyze facial attribute entanglement in the latent space of GANs and apply the proposed transformation for editing the disentangled facial attributes. Our experimental results demonstrate the strength of GANalyzer in editing facial attributes and generating any desired faces. We also create and release a balanced photo-realistic human face dataset. Our code is publicly available on GitHub.
It is well known that objects similar in appearance become bound, or grouped, in perceptual organization. Less well-known is that this relationship works in the opposite direction--grouping (or un-grouping) can make objects look more similar (or dissimilar) to each other. Yet examinations of this latter kind of effect tend to focus on distortions in one direction only (from grouping or un-grouping) that build up over time, across many trials. Here we focus on the flexibility of this process, namely, can the visual system leverage transient cues about objects’ similarity in motion so that they look more like each other in shape, on a trial-by-trial basis? On each trial, observers viewed a pair of briefly presented ellipses that differed in terms of how flat or tall they were, and they reported the aspect ratio of one shape from the pair, cued by an arrow after they disappeared. Crucially, each shape was defined by a cloud of coherently-moving dots on a background of dots with random motion vectors. We manipulated the dots within the two shapes so that they had either the same or different vectors in order to facilitate or disrupt grouping. We found that when the two shapes were grouped by similar motion, the aspect ratio of the cued shape was attracted to the aspect ratio of the uncued shape. We found the opposite pattern when the two shapes had dissimilar motion. In summary, observers reported that the shapes looked more (or less) like each other based on their shared or unshared motion within a brief perceptual moment. We conclude that the visual system can adaptively alter visual experience based on grouping, quickly imposing the appearance of similarity or distinctiveness to fit a perceiver’s belief that objects do or do not belong together.
Social interactions are dynamic and unfold over time. To make sense of social interactions, people must aggregate sequential information into summary, global evaluations. But how do people do this? Here, to address this question, we conducted nine studies (N = 1,583) using a diverse set of stimuli. Our focus was a central aspect of social interaction—namely, the evaluation of others' emotional responses. The results suggest that when aggregating sequences of images and videos expressing varying degrees of emotion, perceivers overestimate the sequence's average emotional intensity. This tendency for overestimation is driven by stronger memory of more emotional expressions. A computational model supports this account and shows that amplification cannot be explained only by nonlinear perception of individual exemplars. Our results demonstrate an amplification effect in the perception of sequential emotional information, which may have implications for the many types of social interactions that involve repeated emotion estimation. Goldenberg et al. show that we tend to overestimate the average intensity of a sequence of emotional expressions and that this is caused by increased memory for stronger expressions.
The study of gaze perception has largely focused on a single cue (the eyes) in two-dimensional settings. While this literature suggests that 2D gaze perception is shaped by atypical development, as in Autism Spectrum Disorder (ASD), gaze perception is in reality contextually-sensitive, perceived as an emergent feature conveyed by the rotation of the pupils and head. We examined gaze perception in this integrative context, across development, among children and adolescents developing typically or with ASD with both 2D and 3D stimuli. We found that both groups utilized head and pupil rotations to judge gaze on a 2D face. But when evaluating the gaze of a physically-present, 3D robot, the same ASD observers used eye cues less than their typically-developing peers. This demonstrates that emergent gaze perception is a slowly developing process that is surprisingly intact, albeit weakened in ASD, and illustrates how new technology can bridge visual and clinical science.
People are good at categorizing the emotions of individuals and crowds of faces. People also make mistakes when classifying emotion. When they do so with judgments of individuals, these errors tend to be negatively biased, potentially serving a protective function. For example, a face with a subtle expression is more likely to be categorized as angry than happy. Yet surprisingly little is known about the errors people make when evaluating multiple faces. We found that perceivers were biased to classify faces as angry, especially when evaluating crowds. This amplified bias depended on uncertainty, occurring when categorization was difficult, and it reached peak intensity for crowds with four members. Drift diffusion modeling revealed the mechanisms behind this bias, including an early response component and more efficient processing of anger from crowds with subtle expressions. Our findings introduce bias as an important new dimension for understanding how perceivers make judgments about crowds. (PsycInfo Database Record (c) 2021 APA, all rights reserved).
How do people go about reading a room or taking the temperature of a crowd? When people catch a brief glimpse of an array of faces, they can focus their attention on only some of the faces. We propose that perceivers preferentially attend to faces exhibiting strong emotions and that this generates a crowd-emotion-amplification effect-estimating a crowd's average emotional response as more extreme than it actually is. Study 1 (N = 50) documented the crowd-emotion-amplification effect. Study 2 (N = 50) replicated the effect even when we increased exposure time. Study 3 (N = 50) used eye tracking to show that attentional bias to emotional faces drives amplification. These findings have important implications for many domains in which individuals must make snap judgments regarding a crowd's emotionality, from public speaking to controlling crowds.
When perceivers view multiple facial expressions shown concurrently, they can quickly and precisely extract the mean emotion from the set. Yet it is not clear how many faces in the set contribute to summary judgments, and how the variance among them influences this process. To address these questions, we used the subset manipulation and varied emotion variance of faces in the sets across three experiments. Sets containing sixteen faces, or a subset of faces randomly selected from the sixteen-face display were presented, and participants judged the average emotion of each face set on a continuous scale. Results showed that when emotion variance was relatively large (Experiments 1 & 2), only two faces in the set contributed to ensemble representations. In Experiment 3 where the emotion variance was smaller, around three to four faces were likely sampled. However, when directly comparing results from Experiments 2 and 3, there was no strong evidence supporting the impact of variance in averaging efficiency. Altogether, these new results suggest that the process of averaging multiple emotional facial expressions can be explained by capacity-limited subsampling. The claim that ensemble representations are capacity unlimited or can overcome the bottlenecks in visual perception might need to be reconsidered.
Most visual scenes contain information at different spatial scales, including the local and global, or the detail and gist. Global processes have become increasingly implicated in research examining summary statistical perception, initially as the output of ensemble coding, and more recently as a gating mechanism for selecting which information is included in the averaging process itself. Yet local and global processing are known to be rapidly integrated by the visual system, and it is plausible that global-level information, like spatial organization, may be included as an input during ensemble coding. We tested this hypothesis using an ensemble shape-perception task in which observers evaluated the mean aspect ratios of sets of ellipses. In addition to varying the aspect ratios of the individual shapes, we independently varied the spatial arrangements of the sets so that they had either flat or tall organizations at the global level. We found that observers made precise summary judgments about the average aspect ratios of the sets by integrating information from multiple shapes. More importantly, global flat and tall organizations were incorporated into ensemble judgments about the sets; summary judgments were biased in the directions of the global spatial arrangements on each trial. This global-to-local integration even occurred when the global organizations were masked. Our results demonstrate that the process of summary representation can include information from both the local and global scales. The gist is not just an output of ensemble representation – it can be included as an input to the mechanism itself.
During ensemble coding, the visual system extracts summary information from input that has been integrated, facilitating gist-level judgments about objects and features that belong together. In contrast, input can be segmented, allowing for quick categorical distinctions between objects. Integration and segmentation usually work in parallel but may sometimes conflict in the context of ensemble coding. To investigate this possibility, we examined summary perception of aspect ratio (i.e., "tallness/flatness"). Aspect ratio has a category boundary (e.g., a circle), and individual aspect ratios may be perceptually exaggerated-segmented-away from this boundary. We predicted that summary perception of multiple aspect ratios would be disrupted when, as a set, they spanned the category boundary, since integration and segmentation would then be at odds. We found that when observers reported the average aspect ratio of a set of ellipses, they were less sensitive to the mean of sets that included both tall and flat ellipses, compared to sets comprised of tall or flat ellipses. Follow-up experiments suggest this occurred because segmentation distorted the appearance of ellipses away from the category boundary, exaggerating set heterogeneity. These experiments advance understanding of how the visual system summarizes information by showing that integration and segmentation can conflict. (PsycInfo Database Record (c) 2020 APA, all rights reserved).
When exposed to others’ emotional responses, people often make rapid decisions as to whether these others are members of their group or not. These group categorization decisions have been shown to be extremely important to understanding group behavior. Yet, despite their prevalence and importance, we know very little about the attributes that shape these categorization decisions. To address this issue, we took inspiration from ensemble coding research and developed a task designed to reveal the influence of the mean and variance of group members’ emotions on participants’ group categorization. In Study 1, we verified that group categorization decreases when the group’s mean emotion is different from the participant’s own emotional response. In Study 2, we established that people identify a group’s mean emotion more accurately when its variance is low rather than high. In Studies 3 and 4, we showed that participants were more likely to self-categorize as members of groups with low emotional variance, even if their own emotions fell outside of the range of group emotions they saw, and that this preference is seen for judgements of both positive and negative group emotions. In Study 5, we showed that this unique preference for low group emotional variance is special to group categorization and does not appear in a more basic face categorization task. Our studies reveal unexplored and important tendencies in group categorization based on group emotions.
While object perception may feel instantaneous, it is an iterative process in which information is accumulated until ambiguity about identity and location is resolved. In theory, awareness of an object should depend on how efficiently this process occurs. Therefore, objects with inherently weak visual representations should be more susceptible to perceptual disruption. We tested this hypothesis by examining the perception of aspect ratio, a 2D feature of shapes with anisotropic representation (circular shapes are less robustly represented than elongated shapes in high-level visual areas). Observers viewed a target shape shown for 20-ms within an array of ellipses. The target, which varied from flat to tall, was either masked or unmasked. Observers indicated the target's aspect ratio and if it was visible. Observers reported seeing elongated shapes far more often than circular shapes, but only on trials with object-substitution masking. This effect replicated across five control experiments, even though the shapes were identical in basic image attributes (e.g., contrast, area). Our findings demonstrate that shapes with extreme aspect ratios are more readily available to awareness than shapes with ambiguous dimensionality. More generally, this work supports theories of object processing which suggest that strength of visual representation gates access to awareness.
Integration and segmentation are fundamental computations in vision, and they serve opposing purposes. Integration occurs, for example, in ensemble coding, whereby perceivers make fast and efficient generalizations about large amounts of information. In contrast, segmentation perceptually exaggerates visual features away from category boundaries, promoting quick-but-crude binary distinctions. Integration and segmentation must work in parallel, yet they are typically examined in isolation. Understanding of how they may conflict is thus surprisingly incomplete. We conducted three experiments examining the ensemble perception of aspect ratio, a visual feature roughly equivalent to “tallness/flatness”, to investigate this potential conflict. In the first two experiments, observers viewed a set of shapes with heterogeneous aspect ratios for 250-ms and used a cursor to adjust a test shape to match the average of the set on each trial. Observers’ distribution of errors across trials served as an index of the precision of ensemble coding. We expected conflict when observers attempted to make a summary judgement about sets of features that spanned a category boundary. Indeed, ensemble coding operated less precisely for sets that included tall and flat shapes, compared to sets that included tall or flat shapes. We suspected that this occurred because sets which spanned the category boundary were perceived as more being heterogeneous than those that did not, even though our sets were carefully matched in terms of physical variability. Replicating previous work (Suzuki & Cavanagh, 1998) we showed in a third experiment that segmentation exaggerated the appearance of individual shapes near the tall/flat category boundary. Segmentation may thus disrupt the integration required for efficient ensemble coding by exaggerating a set’s perceived heterogeneity. This work adds to the understanding of integration by demonstrating aspect ratio integration, and also by showing that integration can be constrained by another fundamental computation in the visual system, segmentation.
Gaze is an emergent visual feature. A person's gaze direction is perceived not just based on the rotation of their eyes, but also their head. At least among adults, this integrative process appears to be flexible such that one feature can be weighted more heavily than the other depending on the circumstances. Yet it is unclear how this weighting might vary across individuals or across development. When children engage emergent gaze, do they prioritize cues from the head and eyes similarly to adults? Is the perception of gaze among individuals with autism spectrum disorder (ASD) emergent, or is it reliant on a single feature? Sixty adults (M = 29.86 years-of-age), thirty-seven typically developing children and adolescents (M = 9.3 years-of-age; range = 7-15), and eighteen children with ASD (M = 9.72 years-of-age; range = 7-15) viewed faces with leftward, rightward, or direct head rotations in conjunction with leftward or rightward pupil rotations, and then indicated whether the face was looking leftward or rightward. All individuals, across development and ASD status, used head rotation to infer gaze direction, albeit with some individual differences. However, the use of pupil rotation was heavily dependent on age. Finally, children with ASD used pupil rotation significantly less than typically developing (TD) children when inferring gaze direction, even after accounting for age. Our approach provides a novel framework for understanding individual and group differences in gaze as it is actually perceived-as an emergent feature. Furthermore, this study begins to address an important gap in ASD literature, taking the first look at emergent gaze perception in this population.
David Wessel合作论文数University of California2