For a socially intelligent robot, different levels of situation as-sessment are required, ranging from basic processing of sensor input to high-level analysis of semantics and intention. However, the attempt to combine them all prompts new research challenges and the need of a co-herent framework and architecture. This paper presents the situation assessment aspect of Romeo2, a unique project aiming to bring multi-modal and multi-layered perception on a single system and targeting for a unified theoretical and functional frame-work for a robot companion for everyday life. It also discusses some of the innovation potentials, which the combination of these various perception abilities adds into the robot's socio-cognitive capabilities.
Multichannel blind source separation performances rapidly degrade when the mixtures are highly reverberated. In fact, blind source separation algorithms usually focus on the separation task without dealing with the dereverberation problem. Some recent studies attempted to reduce the reverberation by introducing a dereverberation module before or after the blind source separation but only limited success was obtained in improving the separation performance in highly reverberant rooms. In this article, we conduct a number of experiments combining state of the art spectral enhancement- based dereverberation and source separation algorithms showing that, in this particular case, speech enhancement does not improve the performance of blind source separation.
We propose an adaptive blind source separation algorithm in the context of robot audition using a microphone array. Our algorithm presents two steps: a fixed beamforming step to reduce the reverberation and the background noise and a source separation step. In the fixed beamforming preprocessing, we build the beamforming filters using the Head Related Transfer Functions (HRTFs) which allows us to take into consideration the effect of the robot's head on the near acoustic field. In the source separation step, we use a separation algorithm based on the l(1) norm minimization. We evaluate the performance of the proposed algorithm in a total adaptive way with real data and varying number of sources and show good separation and source number estimation results.
Cette these propose des algorithmes de separation aveugle de sources audio en utilisant un reseau de capteurs. L'application finale de ces algorithmes est l'audition des robots dans le cadre du projet ROMEO. Dans cette these, nous avons developpe des algorithmes de separation aveugle de sources audio bases sur des criteres de parcimonie. Nous montrons que la minimisation de la norme l1 avec une technique d'optimisation du gradient naturel permet d'elaborer un algorithme se situant au niveau de l'etat de l'art. Nous montrons qu'un critere base sur la parametrisation de la pseudo-norme lp, avec 0
In this article, we present a two-stage blind source separation (BSS) algorithm for robot audition. The first stage consists in a fixed beamforming preprocessing to reduce the reverberation and the environmental noise. Since we are in a robot audition context, the manifold of the sensor array in this case is hard to model due to the presence of the head of the robot, so we use pre-measured head related transfer functions (HRTFs) to estimate the beamforming filters. The use of the HRTF to estimate the beamformers allows to capture the effect of the head on the manifold of the microphone array. The second stage is a BSS algorithm based on a sparsity criterion which is the minimization of the l1 norm of the sources. We present different configuration of our algorithm and we show that it has promising results and that the fixed beamforming preprocessing improves the separation results.
This work addresses the Huawei/3Dlife Grand challenge proposing a set of audio tools for a virtual dance-teaching assistant. These tools are meant to help the dance student develop a sense of rhythm to correctly synchronize his/her movements and steps to the musical timing of the choreographies to be executed. They consist of three main components, namely a music (beat) analysis module, a source separation and remastering module and a dance step segmentation module. These components enable to create augmented tutorial videos highlighting the rhythmic information using, for instance, a synthetic dance teacher voice, but also videos highlighting the steps executed by a student to help in the evaluation of his/her performance.
In this paper, we introduce a modified lp norm blind source separation criterion based on the source sparsity in the time-frequency domain. We study the effect of making the sparsity constraint harder through the optimization process, making the parameter p of the lp norm vary from 1 to nearly 0 according to a sigmoid function. The sigmoid introduces a smooth lp norm variation which avoids the divergence of the algorithm. We compared this algorithm to the regular l1 norm minimization and an ICA based one and we obtained promising results.