Which parts or objects are interesting in a content? In this paper we first propose three computational models to automatically predict interestingness rankings of areas/objects inside a 2D picture. We based our modeling on previous experimental findings to ensure reliability of the prediction when compared to the human assessement of interestingness. Our two first models are based on low level features, extracted from image regions, which have been stated as useful in the human interest process. A baseline model is built by estimating a linear regression from a small dataset of 49 images. The second model estimates a rewarding term based on additional experimental observations. By adding image semantics, we then construct a last model, which more generally benefits from a better understanding of the content. It also integrates notions such that unusualness or human beings' presence that have proven to play key roles in the interestingness process. Finally, targeting VR applications, we extend our models to immersive content, both images and videos, and propose an innovative application to guide the viewer in his/her navigation based on intuitive visual or audio cues.
In our present society, the cinema has become one of the major forms of entertainment providing unlimited contexts of emotion elicitation for the emotional needs of human beings. Since emotions are universal and shape all aspects of our interpersonal and intellectual experience, they have proved to be a highly multidisciplinary research field, ranging from psychology, sociology, neuroscience, etc., to computer science. However, affective multimedia content analysis work from the computer science community benefits but little from the progress achieved in other research fields. In this paper, a multidisciplinary state-of-the-art for affective movie content analysis is given, in order to promote and encourage exchanges between researchers from a very wide range of fields. In contrast to other state-of-the-art papers on affective video content analysis, this work confronts the ideas and models of psychology, sociology, neuroscience, and computer science. The concepts of aesthetic emotions and emotion induction, as well as the different representations of emotions are introduced, based on psychological and sociological theories. Previous global and continuous affective video content analysis work, including video emotion recognition and violence detection, are also presented in order to point out the limitations of affective video content analysis work.
The objective of colour mapping or colour transfer methods is to recolour a given image or video by deriving a mapping between that image and another image serving as a reference. These methods have received considerable attention in recent years, both in academic literature and industrial applications. Methods for recolouring images have often appeared under the labels of colour correction, colour transfer or colour balancing, to name a few, but their goal is always the same: mapping the colours of one image to another. In this paper, we present a comprehensive overview of these methods and offer a classification of current solutions depending not only on their algorithmic formulation but also their range of applications. We also provide a new dataset and a novel evaluation technique called ‘evaluation by colour mapping roundtrip’. We discuss the relative merit of each class of techniques through examples and show how colour mapping solutions can have been applied to a diverse range of problems.
Interestingness is the quantification of the ability of an image to induce interest in a user. Because defining and interpreting interestingness remain unclear in the literature, we introduce in this paper two new notions, intra- and inter-interestingness, and investigate a novel set of dedicated experiments. More specifically, we propose four experimental protocols: 1/ object ranking with a pre-defined word list, 2/ pair-wise comparison, 3/ image ranking and 4/ eye-tracking. We take advantage of experimenting on the same dataset to draw potential links between the collected data and to state on the agreement between subjects. While we do not evidence a relationship between the local (intra) and global (inter) notions of interestingness, we do observe correlated outputs throughout the different protocols. Beyond the low or moderate values obtained from inter-rater agreement metrics, we point out the experimental reproducibility to argue about the universal nature of the interestingness notions. In addition, we bring deep insights on the relationships between interestingness and 7 other criteria, some of them already pointed out in the literature as being linked with interestingness. Unusualness and emotion seem to be the strongest enablers for interestingness. These insights are highly relevant for future work on modeling.
This paper provides a description of the MediaEval 2018 “Emotional Impact of Movies task". It continues to build on last year’s edition, integrating the feedback of previous participants. The goal is to create systems that automatically predict the emotional impact that video content will have on viewers, in terms of valence, arousal and fear. Here we provide a description of the use case, task challenges, dataset and ground truth, task run requirements and evaluation metrics.
Since the consumption of digital media exploded in the last decade, making aesthetic pictures quickly - with or without artistic expertise - is more than ever a research topic. Different axis of investigations remain possible: high resolution, high dynamic range or wide color gamut. Additionally to these objective image properties, more perceptual and artistic insights could be of benefit to any user manipulating pictures. In such context, this thesis deals with the topic of Color Harmony. The literature related to this topic is limited, but involves many different scientific areas: color science, image processing and psychology and so on. The validity of collected data is questionable due to their limitation to two- or three-colors patches. The models fitted from these data remain non-exploitable on natural pictures. Other models depicting rules or areas on color wheel lack scientific guidelines for their utilization. Nonetheless, some algorithms employing color harmony theory and models as a core concept showed up in the literature, but suffered from being quantitatively tested and validated. In this thesis, two views are put in perspective in order to respond to the previous statements: an experimental and a computational approaches. The conducted experiment allowed observing some effects with an eye-tracking protocol, never applied before with a task on color harmony assessment. From the collected data of our experimental work, we designed a method to generate a ground truth, which would serve to the validation of the two proposed computational methods. First, we improved an existing architecture for automatic color harmonization and demonstrated exhaustively the benefit of our approach. As a second computational contribution, a novel quality metric is introduced that integrates the concepts of visual masking and color harmony. Thus, we may predict which areas would be perceived harmonious regarding its neighborhood and then the potential masking effects. As a last contribution, two editing tools made accessible the color harmony theory through a hidden formulation of it and a user-friendly and intuitive interface.
We propose a new, fully automatic method for example-based image colorization and a robust color artifact regularization solution. To determine correspondences between the two images, we supplement the PatchMatch algorithm with rich statistical image descriptors. Based on detected matches, our method transfers colors from the reference to the target grayscale image. In addition, we propose a general regularization scheme that can smooth artifacts typical to color manipulation algorithms. Our regularization approach propagates the major colors in image regions, as determined through superpixel-based segmentation of the original image. We evaluate the effectiveness of our colorization for a varied set of images and demonstrate our regularization scheme for both colorization and color transfer applications.
On one hand, the fact that Galvanic Skin Response (GSR) is highly correlated with the user affective arousal provides the possibility to apply GSR in emotion detection. On the other hand, temporal correlation of real-time GSR and self-assessment of arousal has not been well studied. This paper confronts two modalities representing the induced emotion when watching 30 movies extracted from the LIRIS-ACCEDE database. While continuous arousal annotations have been self-assessed by 5 participants using a joystick, real-time GSR signal of 13 other subjects is supposed to catch user emotional response, objectively without user's interpretation. As a main contribution, this paper introduces a method to make possible the temporal comparison of both signals. Thus, temporal correlation between continuous arousal peaks and GSR were calculated for all 30 movies. A global Pearson's correlation of 0.264 and a Spearman's rank correlation coefficient of 0.336 were achieved. This result proves the validity of using both signals to measure arousal and draws a reliable framework for the analysis of such signals.
Research in affective computing requires ground truth data for training and benchmarking computational models for machine-based emotion understanding. In this paper, we propose a large video database, namely LIRIS-ACCEDE, for affective content analysis and related applications, including video indexing, summarization or browsing. In contrast to existing datasets with very few video resources and limited accessibility due to copyright constraints, LIRIS-ACCEDE consists of 9,800 good quality video excerpts with a large content diversity. All excerpts are shared under creative commons licenses and can thus be freely distributed without copyright issues. Affective annotations were achieved using crowdsourcing through a pair-wise video comparison protocol, thereby ensuring that annotations are fully consistent, as testified by a high inter-annotator agreement, despite the large diversity of raters' cultural backgrounds. In addition, to enable fair comparison and landmark progresses of future affective computational models, we further provide four experimental protocols and a baseline for prediction of emotions using a large set of both visual and audio features. The dataset (the video clips, annotations, features and protocols) is publicly available at: http://liris-accede.ec-lyon.fr/.
Recently, mainly due to the advances of deep learning, the performances in scene and object recognition have been progressing intensively. On the other hand, more subjective recognition tasks, such as emotion prediction, stagnate at moderate levels. In such context, is it possible to make affective computational models benefit from the breakthroughs in deep learning? This paper proposes to introduce the strength of deep learning in the context of emotion prediction in videos. The two main contributions are as follow: (i) a new dataset, composed of 30 movies under Creative Commons licenses, continuously annotated along the induced valence and arousal axes (publicly available) is introduced, for which (ii) the performance of the Convolutional Neural Networks (CNN) through supervised finetuning, the Support Vector Machines for Regression (SVR) and the combination of both (Transfer Learning) are computed and discussed. To the best of our knowledge, it is the first approach in the literature using CNNs to predict dimensional affective scores from videos. The experimental results show that the limited size of the dataset prevents the learning or finetuning of CNN based frameworks but that transfer learning is a promising solution to improve the performance of affective movie content analysis frameworks as long as very large datasets annotated along affective dimensions are not available.
Over the past decade, much research experience has been gained in the realm of high dynamic range (HDR) imaging. Because of the significantly improved visual experience that HDR imaging offers, the industry has recently taken a strong interest in this set of technologies. This is evident in the active consideration of HDR by standardization bodies, such as the DVD Forum and European Broadcasting U...
While image editing tends to be popularized, it remains a time-consuming task requiring a minimal expertise. This proposal presents an innovative tool for performing guided image editing. Oriented by a quality metric based on color harmony theory, the user picks up the least harmonious color areas and retouches them based on our recommended color palette. This set of proposed colors ensure the maximization of the global color harmony for the considered picture. In few clicks, anyone is able to retouch incongruous colors while respecting his taste or intent.
Color mapping or color transfer methods aim to recolor a given image or video by deriving a mapping between that image and another image serving as a reference. This class of methods has received considerable attention in recent years, both in academic literature and in industrial applications. Methods for recoloring images have often appeared under the labels of color correction, color transfer or color balancing, to name a few, but their goal is always the same: mapping the colors of one image to another. In this report, we present a comprehensive overview of these methods and offer a classification of current solutions depending not only on their algorithmic formulation but also their range of applications. We discuss the relative merit of each class of techniques through examples and show how color mapping solutions can and have been applied to a diverse range of problems.