Public datasets have played a significant role in advancing the state-of-the-art in automated facial coding. Many of these datasets contain posed expressions and/or videos recorded in controlled lab conditions with little variation in lighting or head pose. As such, the data do not reflect the conditions observed in many real-world applications. We present AM-FED+ an extended dataset of naturalistic facial response videos collected in everyday settings. The dataset contains 1,044 videos of which 545 videos (263,705 frames or 21,859 seconds) have been comprehensively manually coded for facial action units. These videos act as a challenging benchmark for automated facial coding systems. All the videos contain gender labels and a large subset (77 percent) contain age and country information. Subject self-reported liking and familiarity with the stimuli are also included. We provide automated facial landmark detection locations for the videos. Finally, baseline action unit classification results are presented for the coded videos. The dataset is available to download online: https://www.affectiva.com/facial-expression-dataset/.
Self-report studies have found evidence that cultures differ in the display rules they have for facial expressions (i.e., for what is appropriate for different people at different times). However, observational studies of actual patterns of facial behavior have been rare and typically limited to the analysis of dozens of participants from two or three regions. We present the first large-scale evidence of cultural differences in observed facial behavior, including 740,984 participants from 12 countries around the world. We used an Internet-based framework to collect video data of participants in two different settings: in their homes and in market research facilities. Using computer vision algorithms designed for this dataset, we measured smiling and brow furrowing expressions as participants watched television ads. Our results reveal novel findings and provide empirical evidence to support theories about cultural and gender differences in display rules. Participants from more individualist cultures displayed more brow furrowing overall, whereas smiling depended on both culture and setting. Specifically, participants from more individualist countries were more expressive in the facility setting, while participants from more collectivist countries were more expressive in the home setting. Female participants displayed more smiling and less brow furrowing than male participants overall, with the latter difference being more pronounced in more individualist countries. This is the first study to leverage advances in computer science to enable large-scale observational research that would not have been possible using traditional methods.
Facial coding has become a common tool in media measurement, with large companies (e.g., Unilever) using it to test all of their new video ad content. Facial reactions capture the in-the-moment response of an individual and these data complement self-report measures. Two advancements in affective computing have made measurement possible at scale: 1) computer vision algorithms are used to automatically code sign and message judgments based on facial muscle movements, 2) video data are collected by recording responses in everyday environments via the viewer's own webcam over the Internet. We present results of online facial coding studies of video ads, movie trailers, political content, and long-form TV shows. We explain how these data can be used in market research. Despite the ability to measure facial behavior in a scalable and quantifiable way, the interpretation of these data is still challenging without baselines and comparative measures. Over the past four years we have collected and coded over two million responses to everyday media content. Our huge dataset allows us to calculate reliable normative distributions of responses across different media types. We present these data and argue that this provides a context within which to interpret facial responses more accurately.
There exists a stereotype that women are more expressive than men; however, research has almost exclusively focused on a single facial behavior, smiling. A large-scale study examines whether women are consistently more expressive than men or whether the effects are dependent on the emotion expressed. Studies of gender differences in expressivity have been somewhat restricted to data collected in lab settings or which required labor-intensive manual coding. In the present study, we analyze gender differences in facial behaviors as over 2,000 viewers watch a set of video advertisements in their home environments. The facial responses were recorded using participants’ own webcams. Using a new automated facial coding technology we coded facial activity. We find that women are not universally more expressive across all facial actions. Nor are they more expressive in all positive valence actions and less expressive in all negative valence actions. It appears that generally women express actions more frequently than men, and in particular express more positive valence actions. However, expressiveness is not greater in women for all negative valence actions and is dependent on the discrete emotional state.
The purpose of this panel is to explore issues that will arise in building future personal assistants (PAs), especially for family use. In this regard, we will consider implications of being an "assistant" and those of being "personal." The target timeframe is 3-10 years out, so that very near-term products will not be discussed. We will elaborate briefly on the kinds of communicative and inferential capabilities such PAs will need, and then examine their social and emotional capabilities. We will discuss pros and cons for their evolution and deployment. In this regard, we will discuss the kinds of support that could be provided by the HCI community in building personal assistant systems that are useful, delightful, functional, controllable, educational, ethical, and secure.
We present a real-time facial expression recognition toolkit that can automatically code the expressions of multiple people simultaneously. The toolkit is available across major mobile and desktop platforms (Android, iOS, Windows). The system is trained on the world's largest dataset of facial expressions and has been optimized to operate on mobile devices and with very few false detections. The toolkit offers the potential for the design of novel interfaces that respond to users' emotional states based on their facial expressions. We present a demonstration application that provides real-time visualization of the expressions captured by the camera.
This paper presents large-scale naturalistic and spontaneous facial expression classification on uncontrolled webcam data. We describe an active learning approach that helped us efficiently acquire and hand-label hundreds of thousands of non-neutral spontaneous and natural expressions from thousands of different individuals. With the increased numbers of training samples a classic RBF SVM classifier, widely used in facial expression recognition, starts to become computationally limiting for training and real-time performance. We propose combining two techniques: 1) smart selection of a subset of the training data and 2) the Nystrom kernel approximation method to train a classifier that performs at high-speed (300fps). We compare performance (accuracy and classification time) with respect to the size of the training dataset and the SVM kernel, using either an RBF kernel, a linear kernel or the Nystrom approximation method. We present facial action unit classifiers that perform extremely well on spontaneous and naturalistic webcam videos from around the world recorded over the Internet. When evaluated on a large public dataset (AM-FED) our method performed better than the previously published baseline. Our approach generalizes to many problems that exhibit large individual variability.
Billions of online video ads are viewed every month. We present a large-scale analysis of facial responses to video content measured over the Internet and their relationship to marketing effectiveness. We collected over 12,000 facial responses from 1,223 people to 170 ads from a range of markets and product categories. The facial responses were automatically coded frame-by-frame. Collection and coding of these 3.7 million frames would not have been feasible with traditional research methods. We show that detected expressions are sparse but that aggregate responses reveal rich emotion trajectories. By modeling the relationship between the facial responses and ad effectiveness, we show that ad liking can be predicted accurately (ROC AUC = 0.85) from webcam facial responses. Furthermore, the prediction of a change in purchase intent is possible (ROC AUC = 0.78). Ad liking is shown by eliciting expressions, particularly positive expressions. Driving purchase intent is more complex than just making viewers smile: peak positive responses that are immediately preceded by a brand appearance are more likely to be effective. The results presented here demonstrate a reliable and generalizable system for predicting ad effectiveness automatically from facial responses without a need to elicit self-report responses from the viewers. In addition we can gain insight into the structure of effective ads.
Facial behavior contains rich non-verbal information. However, to date studies have typically been limited to the analysis of a few hundred or thousand video sequences. We present the first-ever ultra large-scale clustering of facial events extracted from over 1.5 million facial videos collected while individuals from over 94 countries respond to one of more that 8000 online videos. We believe this is the first example of what might be described “big data” analysis in facial expression research. Automated facial coding was used to quantify eyebrow raise (AU2), eyebrow lowerer (AU4) and smile behaviors in the 700,000,000+ frames. Facial “events” were extracted and defined by a set of temporal features and then clustered using the k-means clustering algorithm. Verifying the observations in each cluster against human-coded data we were able to identify reliable clusters of facial events with different dynamics (e.g. fleeting vs. sustained and rapid offset vs. slow offset smiles). These events provide a way of summarizing behaviors that occur without prescribing the properties. We examined the how these nuanced facial events were tied to consumer behavior. We found that smile events - particularly those with high peaks - were much more likely to occur during viral ads. This data is cross-cultural, we also examine the prevalence of different events across regions of the globe.
Traditional observational research methods required an experimenter's presence in order to record videos of participants, and limited the scalability of data collection to typically less than a few hundred people in a single location. In order to make a significant leap forward in affective expression data collection and the insights based on it, our work has created and validated a novel framework for collecting and analyzing facial responses over the Internet. The first experiment using this framework enabled 3,268 trackable face videos to be collected and analyzed in under two months. Each participant viewed one or more commercials while their facial response was recorded and analyzed. Our data showed significantly different intensity and dynamics patterns of smile responses between subgroups who reported liking the commercials versus those who did not. Since this framework appeared in 2011, we have collected over three million videos of facial responses in over 75 countries using this same methodology, enabling facial analytics to become significantly more accurate and validated across five continents. Many new insights have been discovered based on crowd-sourced facial data, enabling Internet-based measurement of facial responses to become reliable and proven. We are now able to provide large-scale evidence for gender, cultural and age differences in behaviors. Today such methods are used as part of standard practice in industry for copy-testing advertisements and are increasingly used for online media evaluations, distance learning, and mobile applications.
A supervised machine learning approach to remote video-based heart rate (HR) estimation is proposed. We demonstrate the possibility of training a discriminative statistical model to estimate the Blood Volume Pulse signal (BVP) from the human face using ambient light and any off-the-shelf webcam. The proposed algorithm is 120 times faster than state of the art approach and returns a confidence metric to evaluate the HR estimates plausibility. The algorithm was evaluated against the state-of-the-art on 120 minutes of face videos, the largest video-based heart rate evaluation to date. The evaluation results showed a 53% decrease in the Root Mean Squared Error (RMSE) compared to state-of-the-art.