This demo will show a prototype of a new software radio enabled broadcast media navigator implemented on an FPGA and quad-core processor, which is able to demodulate simultaneously all channel in the FM band and perform a real time classification of the musical genre. This prototype represents the elementary component of a navigator capable of searching the large quantities of information contained in the radio bands.
Broadcast radio is a rich but underexploited source of multimedia content. To make this available to users, it will be indispensable to develop new types of navigators capable of searching the large quantities of information contained in the radio bands. The article introduces a prototype of a new software radio enabled broadcast media navigator implemented on an FPGA, which is able to demodulate simultaneously all channel in the FM band and perform audio indexing upon them, ultimately using a Graphics Processing Unit.
This paper presents the first public framework for the evaluation of audio fingerprinting techniques. Although the domain of audio identification is very active, both in the industry and the academic world, there is at present no common basis to compare the proposed techniques. This is because corpuses and evaluation protocols differ among the authors. The framework we present here corresponds to a use-case in which audio excerpts have to be detected in a radio broadcast stream. This scenario, indeed, naturally provides a large variety of audio distortions that makes this task a real challenge for fingerprinting systems. Scoring metrics are discussed with regard to this particular scenario. We then describe a whole evaluation framework including an audio corpus, together with the related groundtruth annotation, and a toolkit for the computation of the score metrics. An example of an application of this framework is finally detailed, that took place during the evaluation campaign of the Quaero project. This evaluation framework is publicly available for download and constitutes a simple, yet thorough, platform that can be used by the community in the field of audio identification to encourage reproducible results.
Broadcast radio is a rich but underexploited source of multimedia content. To make this available to users, it will be indispensable to develop new types of navigators capable of searching the large quantities of information contained in the radio bands. The article introduces a prototype of a new software radio enabled broadcast media navigator implemented on an FPGA, which is able to demodulate simultaneously all channel in the FM band and perform audio indexing upon them, ultimately using a Graphics Processing Unit.
Separating multiple tracks from professionally produced music recordings (PPMRs) is still a challenging problem. We address this task with a user-guided approach in which the separation system is provided segmental information indicating the time activations of the particular instruments to separate. This information may typically be retrieved from manual annotation. We use a so-called multichannel nonnegative tensor factorization (NTF) model, in which the original sources are observed through a multichannel convolutive mixture and in which the source power spectrograms are jointly modeled by a 3-valence (time/frequency/source) tensor. Our user-guided separation method produced competitive results at the 2010 Signal Separation Evaluation Campaign, with sufficient quality for real-world music editing applications.
The work presented is this chapter is an introduction to the subject of sin-le sensor source separation dedicated to the case of speech/music audio Mixtures. Approaches related in this study are all based on a full (Bayesian) probabilistic framework for both source modeling, and source estimation. We first present a review of several codebook approaches for single sensor source separation as well as several attempts to enhance the algorithms. All these approaches aim at adaptively estimating the optimal time-frequency masks for each audio component within the mixture. Three strategies for source modeling are presented: Gaussian scaled mixture models, codebooks of autoregressive models, and Bayesian non-negative matrix factorization (BNMF). These models are described in details and two estimators for the time-frequency masks are presented, namely the minimum mean-squared error and the maximum a posteriori. We then propose two extensions and improvements on the BNMF method. The first one suggests to enhance discrimination between speech and music through multi-scale analysis. The second one suggests to constrain the estimation of the expansion coefficients with prior information. We finally, demonstrate the improved performance of the proposed methods on mixtures of voice and music signals before conclusions and perspectives.
In this paper we address the application of single sensor source separation techniques to mixtures of speech and music. Three strategies for source modeling are presented, namely Gaussian scaled mixture models (GSMM), autoregressive (AR) models and amplitude factor (AF). The common ingredient to the methods is the use of a codebook containing elementary spectral shapes to represent non- stationary signals, and to handle separately spectral shape and amplitude information. We propose a new system that employs separate models for the speech and music signals. The speech signal proves to be best modeled with the AR-based codebook, while the music signal is best modeled with the AF-based codebook. Experimental results demonstrate the improved performance of the proposed approach for speech/music separation in some evaluation criteria.
The aim of this paper is to investigate the use of multi- resolution framework for single sensor source separation based on pseudo-Wiener filtering. We propose a scheme in which the signal is iteratively split in target sources and a residual. Each target source is modeled as the sum of elementary components with known Power Spectral Den- sities (PSDs). The approach boils down to perform a non negative decomposition of the spectra of the observed si- gnal in a given frame onto the dictionnary of known PSDs. The resolution of the PSDs (and hence the frame length) is changed at each iteration of the algorithm. The decompo- sition into sources plus residual is done thanks to a confi- dence measure based on the Fisher information matrix of the expansion coefficients. After theoretical developments we compare the mono and multiresolution approaches and a set of audio examples.
The aim of this paper is to investigate the use of multiplewindow Short-Time Fourier Transform (STFT) representation for single sensor source separation. We propose to iteratively split the observed signal into target sources and residuals. Each target source is modeled as the sum of elementary components with known power spectral densities (PSDs). The approach involves a non negative decomposition of the spectra of the observed mixure in a given frame into a dictionary of PSDs. The resolution of the PSDs varies at each iteration of the algorithm. The decomposition into source signals and residual signals employs a confidence measure, which is based on the Fisher information matrix of the expansion coefficients. We demonstrate the improved performance of the proposed method on mixtures of voice and music signals.
The article deals with a technique of voice forgery using the ALISP (automatic language independent speech processing) approach. Such a technique allows the voice of an arbitrary person (the impostor) to be transformed, forging the identity of another person (the client). Our goal is to demonstrate that an automatic speaker recognition system could be seriously threatened by a transformation of this kind. For this purpose, we use a speaker verification system to calculate the likelihood that the forged voice belongs to the genuine client. Experiments on NIST 2004 evaluation data show that the equal error rate for the verification task is significantly increased by our voice transformation.
The aim of Automatic Speaker Verification (ASV) is to detect whether a speech segment has been uttered by the claimed identity or by an impostor. Our contribution includes the distribution of BECARS , a free software based on Gaussian Mixture Models (GMM) for Automatic Speaker Verification (ASV), and the design of a new methodology to estimate the decision score in an ASV system. BECARS in available at http://www.tsi.enst.fr/ blouet/Becars/. The main characteristic of this software is to allow the use of several adaptation techniques including the most common ones such as Maximum A Posteriori (MAP) and Maximum Likelihood Linear Regression (MLLR). The proposed method for score computation is based on the use of a hierarchical Gaussian clusterization method that we describe in details in this paper. We introduce this work with a general summary of Automatic Speaker Verification (ASV), followed by a description of the adaptation technique available in BECARS used in this work. We then present and evaluate our score computation scheme before concluding the paper.
Gerard Chollet合作论文数CNRS (Centre National de la Recherche Scientifique)8
Bernadette Dorizzi合作论文数Institut National des Telecommunications2