This article introduces a class of first-order stationary time-varying Pitman-Yor processes. Subsuming our construction of time-varying Dirichlet processes presented in (Caron et al., 2007), these models can be used for time-dynamic density estimation and clustering. Our intuitive and simple construction relies on a generalized Pólya urn scheme. Significantly, this construction yields marginal distributions at each time point that can be explicitly characterized and easily controlled. Inference is performed using Markov chain Monte Carlo and sequential Monte Carlo methods. We demonstrate our models and algorithms on epidemiological and video tracking data.
The use of Bayesian inference in the inference of time-frequency representations has, thus far, been limited to offline analysis of signals, using a smoothing spline based model of the time-frequency plane. In this paper we introduce a new framework that allows the routine use of Bayesian inference for online estimation of the time-varying spectral density of a locally stationary Gaussian process. The core of our approach is the use of a likelihood inspired by a local Whittle approximation. This choice, along with the use of a recursive algorithm for non-parametric estimation of the local spectral density, permits the use of a particle filter for estimating the time-varying spectral density online. We provide demonstrations of the algorithm through tracking chirps and the analysis of musical data.
A nonparametric approach combining generative models and functional data analysis is presented in this paper for classifying functional data which arise naturally in a wide variety of signal processing applications, such as brain computer interfacing, speech recognition, or image classification. Based on a new and improved family of Bayesian classifiers, we extend hierarchical Bayesian classification methodology from vector to functional settings. We provide theoretical and practical motivations to our approach which relies on Dirichlet process mixtures and Gaussian processes. The performance is evaluated on phoneme recognition task, and compared to that of Functional Support Vector Machines (FSVMs).
This book serves as an ideal starting point for newcomers and an excellent reference source for people already working in the field. Researchers and graduate students in signal processing, computer science, acoustics and music will primarily benefit from this text. It could be used as a textbook for advanced courses in music signal processing. Since it only requires a basic knowledge of signal processing, it is accessible to undergraduate students.
This paper deals with functional regression, in which the input attributes as well as the response are functions. To deal with this problem, we develop a functional reproducing kernel Hilbert space approach; here, a kernel is an operator acting on a function and yielding a function. We demonstrate basic properties of these functional RKHS, as well as a representer theorem for this setting; we investigate the construction of kernels; we provide some experimental insight.
Audio diarization is the process of partitioning an input audio stream into homogeneous regions according to their specific audio sources. These sources can include audio type (speech, music, background noise, ect.), speaker identity and channel characteristics. With the continually increasing number of larges volumes of spoken documents including broadcasts, voice mails, meetings and telephone conversations, diarization has received a great deal of interest in recent years which significantly impacts performances of automatic speech recognition and audio indexing systems. A subtype of audio diarization, where the speech segments of the signal are broken into different speakers, is speaker diarization. It generally answers to the question Who spoke when? and it is divided in two modules: speaker segmentation and speaker clustering. This chapter discusses the problem of automatically detecting speaker change points presented in a given audio stream, without prior acoustic information on the speakers. We introduce a new unsupervised speaker segmentation technique based on One Class Support Vector Machines (1-SVMs) robust to different acoustic conditions. We evaluated the robustness improvements of this method by segmenting different types of audio stream (broadcast news, meetings and telephone conversations) and comparing the results with model selection segmentation techniques based on the Bayesian information criterion (BIC).
In this paper, we discuss concepts and methods of nonlinear regression for functional data. The focus is on the case where covariates and responses are functions. We present a general framework for modelling functional regression problem in the Reproducing Kernel Hilbert Space (RKHS). Basics concepts of kernel regression analysis in the real case are extended to the domain of functional data analysis. Our main results show how using Hilbert spaces theory to estimate a regression function from observed functional data. This procedure can be thought of as a generalization of scalar-valued nonlinear regression estimate.
Chapter 13 Application of Time-Frequency Techniques to Sound Signals: Recognition and Diagnosis Manuel Davy, Manuel DavySearch for more papers by this author Manuel Davy, Manuel DavySearch for more papers by this author Book Editor(s):Franz Hlawatsch, Franz HlawatschSearch for more papers by this authorFrançois Auger, François AugerSearch for more papers by this author First published: 01 January 2008 https://doi.org/10.1002/9780470611203.ch13 AboutPDF ToolsRequest permissionExport citationAdd to favoritesTrack citation ShareShare Give accessShare full text accessShare full-text accessPlease review our Terms and Conditions of Use and check box below to share full-text version of article.I have read and accept the Wiley Online Library Terms and Conditions of UseShareable LinkUse the link below to share a full-text version of this article with your friends and colleagues. Learn more.Copy URL Share a linkShare onFacebookTwitterLinked InRedditWechat Summary Time-frequency representations are powerful analysis tools, especially in a non-stationary context. They can be employed for the classification or detection of signals, in the extremely frequent situation where we have some training signals but no mathematical model of the signal to be processed. We then speak of supervised non-parametric decision. In this chapter, we propose two applications involving sound signals: the verification of speakers and the detection of defects in acoustic loudspeakers. In the approach that we propose, time-frequency signal representations are used as basic information for constructing a decision-making process. Two time-frequency representations are compared using specially designed distance measures. During the training phase, the kernel of the time-frequency representations as well as the distance measure are optimized so as to minimize the probability of error. The latter is estimated for the whole set of training signals. Time‐Frequency Analysis: Concepts and Methods RelatedInformation
This paper presents a method aimed at recognizing environmental sounds for surveillance and security applications. We propose to apply one-class support vector machines (1-SVMs) together with a sophisticated dissimilarity measure in order to address audio classification, and more specifically, sound recognition. We illustrate the performance of this method on an audio database, which consists of 1015 sounds belonging to nine classes. The database used presents high intraclass diversity in temps of signal properties and some kind of interclass similarities. A large discrepancy in the number of items in each class implies nonuniform probability of sound appearances. The method proceeds as follows: first, the use of a set of state-of-the-art audio features is studied. Then, we introduce a set of novel features obtained by combining elementary features. Experiments conducted on a nine-class classification problem show the superiority of this novel sound recognition method. The best recognition accuracy (96.89%) is obtained when combining wavelet-based features, MFCCs, and individual temporal and frequency features. Our 1-SVM-based multiclass classification approach overperforms the conventional hidden Markov model-based system in the experiments conducted, the improvement in the error rate can reach 50%. Besides, we provide empirical results showing that the single-class SVM outperforms a combination of binary SVMs. Additional experiments demonstrate our method is robust to environmental noise.
Standard methods for maximum likelihood parameter estimation in latent variable models rely on the Expectation-Maximization algorithm and its Monte Carlo variants. Our approach is different and motivated by similar considerations to simulated annealing; that is we build a sequence of artificial distributions whose support concentrates itself on the set of maximum likelihood estimates. We sample from these distributions using a sequential Monte Carlo approach. We demonstrate state-of-the-art performance for several applications of the proposed approach.
This paper presents a new technique for segmenting an audio stream into pieces, each one contains speeches of only one speaker. Speaker segmentation has been used extensively in various tasks such as automatic transcription of radio broadcast news and audio indexing. The segmentation method used in this paper is based on a discriminative distance measure between two adjacent sliding windows operating on preprocessed speech. The proposed unsupervised detection method which does not require any pre-trained models is based on the use of the exponential family model and 1-SVMs to approximate the generalized likelihood ratio. Our 1-SVM-based segmentation algorithm provides improvements over baseline approaches which use the Bayesian Information Criterion (BIC). The segmentation results achieved in our experiments illustrate the potential of this method in detecting speaker changes in audio streams containing over-lapped and short speeches.
This paper addresses a speciflc audio classiflcation problem for surveillance and security applications. The developed approach uses original audio features including essentially wavelets-based features, and a procedure of multi-class classiflcation using several One-class SVMs. The originality of our work lies primarily in the proposal of a selection procedure of the optimal audio features combinations, as well as a new method of multi-class classiflcation. These new approaches are evaluated on an original and su-ciently large database.
Generative bayesian supervised classification of signals is made difficult by two main issues. The first is related to the definition of a convenient prior distribution for the signal parameter. The second is in computing the quantities needed to perform the classification. In this paper, we propose a nonparametric prior known as the Dirichlet Process Mixture, as well as a suited Monte Carlo Markov Chain algorithm. These model and algorithm are applied to the classification of radar altimetry data.
This paper presents a Bayesian technique aimed at classifying signals without prior training (clustering). The approach consists of modelling the observed signals, known only through a finite set of samples corrupted by noise, as Gaussian processes. As in many other Bayesian clustering approaches, the clusters are defined thanks to a mixture model. In order to estimate the number of clusters, we assume a priori a countably infinite number of clusters, thanks to a Dirichlet process model over the Gaussian processes parameters. Computations are performed thanks to a dedicated Monte Carlo Markov Chain algorithm, and results involving real signals (mRNA expression profiles) are presented.
With recent and continued increases in the number of available sound archives (radio, TV, Web,...), effective methods must be established to facilitate the process of searching for information within massive databases. Of less complexity than the original sound file but nevertheless containing a summary of important information pertaining to the signal, text files (index files) are linked to the digital sound files. An example of relevant information found in the text file is as follows: 45 minutes of speech, 1 minute of music, 10 speakers (6 men and 4 women). These index files, stored with the original signal, will contribute considerably to the information retrieval process, allowing an immediate and direct access to the information sought. If one would like to know who speaks and when in a sound file, the index key is hence the speaker. A preliminary stage of a speaker indexing system is speaker diarization. State-of-the-art speaker diarization techniques require two main steps: speaker turn detection which consists of detecting speaker turn times, that is boundaries of audio file segments where only one speaker is present, followed by a clustering step which consists of labelling the previous segments in terms of speakers. These two stages require a metric to be defined in order to compare and groups speech segments. This paper presents a novel approach for the speaker diarization of audio recordings. The proposed approach uses a metric based on one-class Support Vector Machines (SVM-I), introduced recently by one of the authors, for the speaker change detection and clustering tasks. Through many experiments using two databases of broadcast recordings, we demonstrate the relevance and superiority of this approach compared to the traditional method based on the generalized likelihood ratio using bayesian information criterion (RVG-BIC).
This paper addresses the joint estimation and detection of time-varying harmonic components in audio signals. We follow a flexible viewpoint, where several frequency/amplitude trajectories are tracked in spectrogram using particle filtering. The core idea is that each harmonic component (composed of a fundamental partial together with several overtone partials) is considered a target. Tracking requires to define a state-space model with state transition and measurement equations. Particle filtering algorithms rely on a so-called sequential importance distribution, and we show that it can be built on previous multipitch estimation algorithms, so as to yield an even more efficient estimation procedure with established convergence properties. Moreover, as our model captures all the harmonic model information, it actually separates the harmonic sources. Simulations on synthetic and real music data show the interest of our approach