
In this paper, we propose an effective superpixel-based saliency model. First, the original image is simplified by performing superpixel segmentation and adaptive color quantization. On the basis of superpixel representation, inter-superpixel similarity measures are then calculated based on difference of histograms and spatial distance between each pair of superpixels. For each superpixel, its global contrast measure and spatial sparsity measure are evaluated, and refined with the integration of inter-superpixel similarity measures to finally generate the superpixel-level saliency map. Experimental results on a dataset containing 1,000 test images with ground truths demonstrate that the proposed saliency model outperforms state-of-the-art saliency models.
This article proposes experiments on the automatic recognition of personality traits and conflict handling style based on nonverbal communication. The tests are performed over the SSPNet-Nokia Corpus, a collection of 60 mobile phone calls (120 subjects in total) based on the Winter Survival Task. Nonverbal behavioral cues are extracted from speech (captured with the phone microphones) and motor activation (captured indirectly via the gyroscopes mounted on the phones). Support Vector Machines are then adopted to map the cues ;into two psychological constructs, namely the "Big Five" personality traits and the dimensions of the "Rahim Organizational Conflict Inventory", the former capturing individual characteristics and the latter accounting for subjects attitude towards conflict and disagreement. The results show that performances higher than chance to a statistically significant extent can be achieved for one personality trait (Neuroticism) and two conflict handling styles (Dominating and Obliging).
The modeling of auditory scenes is a challenging task in Computational Auditory Scene Analysis. A method based on sparse Non-negative Matrix Factorization that can be used with no prior knowledge of the audio content to establish the similarity between scenes is proposed in this work. It is then evaluated on a corpus of soundscapes of train stations from a perceptual study and results are compared with the human perception. The proposed method, by being able to focus on salient events within the scene, achieves better performances than a state-of-the-art Bag-of-Frames approach though not reaching the human performances.
Semantic concept detection in large scale video collections is mostly achieved through a static analysis of selected keyframes. A popular choice for representing the visual content of an image is based on the pooling of local descriptors such as Dense SIFT. However, simple motion features such as optic flow can be extracted relatively easy from such keyframes. In this paper we propose an efficient addition to the DSIFT approach by including information derived from optic flow. Based on optic flow magnitude, we can estimate for each DSIFT patch whether it is static or moving. We modify the bag of words model used traditionally with DSIFT by creating two separate occurrence histograms instead of one: one for static patches and one for dynamic patches. We further refine this method by studying different separation thresholds and soft assign-ment, as well as different normalization techniques. Classifier score fusion is used to maximize the average precision of all these variants. Experimental results on the TRECVID Semantic Indexing collection show that by means of classifier fusion our method increases overall mean average precision of the DSIFT classifier from 0.061 to 0.106.
This paper addresses footstep detection and classification with multiple microphones distributed on the floor. We propose to introduce geometrical features such as position and velocity of a sound source for classification which is estimated by amplitude-based localization. It does not require precise inter-microphone time synchronization unlike a conventional microphone array technique. To classify various types of sound events, we introduce four types of features, i. e., time-domain, spectral and Cepstral features in addition to the geometrical features. We constructed a prototype system for footstep detection and classification based on the proposed ideas with eight microphones aligned in a 2-by-4 grid manner. Preliminary classification experiments showed that classification accuracy for four types of sound sources such as a walking footstep, running footstep, handclap, and utterance maintains over 70% even when the signal-to-noise ratio is low, like 0 dB. We also confirmed two advantages with the proposed footstep detection and classification. One is that the proposed features can be applied to classification of other sound sources besides footsteps. The other is that the use of a multichannel approach further improves noise-robustness by selecting the best microphone among the microphones, and providing geometrical information on a sound source.
Typically, the saliency map of an image is usually inferred by only using the information within this image. While efficient, such single-image-based methods may fail to obtain reliable results, because the information within a single image may be insufficient for defining saliency. In this paper, we propose a novel idea of learning with labeled images and adopt a new paradigm called sample specific late fusion (SSLF). To effectively explore the visual neighborhood information, we propose a semi-supervised learning technique for learning robust sample specific fusion parameters for multiply response maps of generic bottom-up saliency detectors. Different from previous methods, the proposed SSLF method integrates both middle-level image representation and unlabeled data information through an effective graph regularization framework. Extensive experiments have clearly validated its superiority over other state-of-the-art methods.
In the last decade there has been an explosion in the availability of digitized music material, which comprises data of various formats and modalities including textual, symbolic, acoustic and visual rep-resentations. For example, in the case of an opera there typically exist digitized versions of the libretto, different editions of the musical score, as well as a large number of performances given as audio and video recordings. In this paper, we give an overview of various informed approaches to music processing, where the availability of multiple sources of music-related information is used for supporting and improving the analysis of music data. Considering the scenario of the opera “Der Freischüitz” by Carl Maria von Webera work of central musical importance, where one can draw upon a rich body of sources-we highlight how the identification and creation of cross-modal relationships are a key issue in multimedia processing.
Audio source separation consists in recovering different unknown signals called sources by filtering their observed mixtures. In music processing, most mixtures are stereophonic songs and the sources are the individual signals played by the instruments, e.g. bass, vocals, guitar, etc. Source separation is often achieved through a classical generalized Wiener filtering, which is controlled by parameters such as the power spectrograms and the spatial locations of the sources. For an efficient filtering, those parameters need to be available and their estimation is the main challenge faced by separation algorithms. In the blind scenario, only the mixtures are available and performance strongly depends on the mixtures considered. In recent years, much research has focused on informed separation, which consists in using additional available information about the sources to improve the separation quality. In this paper, we review some recent trends in this direction.
In this paper, we present a novel approach to real-time detection of the string number and fretboard position from polyphonic guitar recordings. Our goal is to assess, if a music student is correctly performing guitar exercises presented via music education software or a remote guitar teacher. We combine a state-of-the art approach for multi-pitch detection with a subsequent audio feature extraction and classification stage. Performance of the proposed system is evaluated with manually annotated chords recorded using different guitars.
Different formats and compression algorithms have been proposed for 3D video content, but 3D images are still mostly represented as a stereo pair only. However, for enhanced 3D rendering capabilities, such as depth perception adjustment or display size adaptation, additional depth data is necessary. To facilitate the standardization process of a common 3D format, backward compatibility with legacy technologies is necessary. In this paper, we propose to extend the JPEG file format, as the most popular image format, in a backward compatible manner to represent a stereo pair and additional depth data. We propose an architecture to achieve such backward compatibility with JPEG. The coding efficiency of a simple implementation of the proposed architecture is compared to the state of the art stereoscopic 3D image compression and storage formats.
People with chronic musculoskeletal pain can experience pain-related fear of physical activity and low confidence in their own motor capabilities. These pain-related emotions and thoughts are often communicated through communicative and protective non-verbal behaviours. Studies in clinical psychology have shown that protective behaviours affect well-being not only physically and psychologically, but also socially. These behaviours appear to be used by others to appraise not just a person's physical state but also to make inferences about their personality traits, with protective pain-related behaviour more negatively evaluated than the communicative behaviour. Unfortunately, people with chronic pain may have difficulty in controlling the triggers of protective behaviour and often are not even aware they exhibit such behaviour. New sensing technology capable of detecting such behaviour or its triggers could be used to support rehabilitation in this regard. In this paper we briefly discuss the above issues and present our approach in developing a rehabilitation system.
This article presents a method for automatic tagging of Youtube videos. The proposed method combines an automatic speech recognition (ASR) system, that extracts the spoken contents, and a keyword extraction component that aims at finding a small set of tags representing a video. In order to improve the robustness of the tagging system to the recognition errors, a video transcription is represented in a topic space obtained by a Latent Dirichlet Allocation (LDA), in which each dimension is automatically characterized by a list of weighted terms. Tags are extracted by combining the weighted word list of the best LDA classes. We evaluate this method by employing the user-provided tags of Youtube videos as reference and we investigate the impact of the topic model granularity. The obtained results demonstrate the interest of such model to improve the robustness of the tagging system.
Video content constitutes today a large part of the data traffic on the Internet. This is allowed by the capillary spreading of video codec technologies: nowadays, every computer, tablet and smart phone is equipped with video encoding and decoding technologies. As a matter of fact, the video content often exists in different formats, that, even though they can be incompatible to each other, still have a significant mutual redundancy. The incompatibility prevents an efficient exploitation of the scalability, which on the other hand is a very important characteristic when it comes to efficient network use. An interesting alternative to classical scalable video is to use distributed video coding (DVC) for the enhancement layers. In the envisaged scenario, clients have different decoders for the base layer, adapted to the characteristics of their de-vice. However they can share the same enhancement layer, since DVC allows encoding frames independently from the reference that will be employed at the decoder. This approach has been considered in the past in order to improve temporal and spatial scalability. In this work we review the existing approaches, improve them using more recent DVC techniques and perform a new analysis for the emerging multi-view applications.
Human affect recognition is the field of study associated with using automatic techniques to identify human emotion or human affective state. A person's affective states is often communicated non-verbally through body language. A large part of human body language communication is the use of head gestures. Almost all cultures use subtle head movements to convey meaning. Two of the most common and distinct head gestures are the head nod and the head shake gestures. In this paper we present a robust system to automatically detect head nod and shakes. We employ the Microsoft Kinect and utilise discrete Hidden Markov Models (HMMs) as the backbone to a machine learning based classifier within the system. The system achieves 86% accuracy on test datasets and results are provided.
We present an analysis of existing methods to automatic classification of photos according to aesthetics. We review different components of the classification process: existing evaluation datasets, their properties, most commonly-used image features, qualitative and quantitative, and classification results where comparable. We argue there are methodology gaps in the existing approaches to evaluating the classification results. We introduce the results of our experiments with Random Forest classification applied to image aesthetics classification and compare them to AdaBoost and SVM approaches.
Effective indexing of multimedia documents requires a multi-modal approach in which either the most appropriate modality is selected or different modalities are used in a collaborative fashion. A collaborative pattern is a model of combination between media that defines how and when to combine information coming from different media sources. Fusing information coming from different media seems a natural way to handle multimedia content. We focus on describing fusion strategies where the task is achieved through the use of different modalities. We browse through the literature looking at various state of the art multi-modal fusion techniques varying from naive combination of modalities to more complex methods of machine learning and discuss various issues faced with fusing several modalities having different properties in the context of semantic indexing.
In this paper, we give an overview of recent BBC R&D work on automated affective and semantic annotations of BBC archive content, covering different types of use-cases and target audiences. In particular, after giving a brief overview of manual cataloguing practices at the BBC, we focus on mood classification, sound effect classification and automated semantic tagging. The resulting data is then used to provide new ways of finding or discovering BBC content. We describe two such interfaces, one driven by mood data and one driven by semantic tags.
We describe a semi-automatic video logging system, capable of annotating frames with semantic metadata describing the objects present. The system learns by visual examples provided interactively by the logging operator, which are learned incrementally to provide increased automation over time. Transfer learning is initially used to bootstrap the system using relevant visual examples from ImageNet. We adapt the hard-assignment Bag of Word strategy for object recognition to our interactive use context, showing transfer learning to significantly reduce the degree of interaction required.
Sound field reproduction has been developed for the last twenty years in research laboratories but has only made its way into the professional market and consumer markets a few years ago. Sound field reproduction addresses the major problem of classical channel based reproduction (two channel stereophony and its derivation: 5.1 and 7.1) that can only reproduce accurately the spatial properties of sound scenes within a very restricted listening area. Sound field reproduction enables to expand the size of the listening area up to the entire listening room, thus providing a collective experience. This paper provides a review of existing sound field reproduction technologies and an analysis of their applicability to real life applications. It also outlines the key dependency of these sound field reproduction techniques to the delivery format of the audio content. Various applications are reviewed ranging from sound reinforcement in live scenario to sound reproduction at home where sound field reproduction is already used.
This paper presents a multisource sound localization method based on the generalized cross-correlation (GCC) method weighted by the phase transform (PHAT) and a novel multisource speech tracking method consisting of voice activity detection (VAD) and K-means clustering algorithm for binaural robot audition. The standard K-means clustering algorithm was improved for the purpose of multisource speech tracking by adding two additional steps. Experiments conducted on the SIG-2 humanoid robot in a real environment show that our method can track multiple speakers in real-time with tracking error below 4.35°.