This paper addresses the challenge of estimating multiple highly oscillating amplitudes within the nonlinear chirp signal model. The problem is analogous to the mode detection task with fixed instantaneous frequencies, where the oscillating amplitudes signify mechanical vibrations concealing crucial information for predictive maintenance. Existing methods often focus on single-frequency estimation, employ simple amplitude functions, or impose strong noise assumptions. Furthermore, these methods frequently rely on arbitrarily chosen hyperparameters, leading to sub-optimal generalization for a diverse range of amplitudes. To address these limitations, our approach introduces two estimators, based on Capon filters and negative log-likelihood approaches respectively, that leverage locally stationary assumptions and incorporate hyperparameters estimation. The results demonstrate that, even under challenging conditions, these estimators yield competitive outcomes across various noisy scenarios, mitigating the drawbacks associated with existing methods.
In non-linear autoregressive models, the time dependency of coefficients is often driven by a particular time-series which is not given and thus has to be estimated from the data. To allow model evaluation on a validation set, we describe a parametric approach for such driver estimation. After estimating the driver as a weighted sum of potential drivers, we use it in a non-linear autoregressive model with a polynomial parametrization. Using gradient descent, we optimize the linear filter extracting the driver, outperforming a typical grid-search on predefined filters.
In this paper, we introduce new parametric generative driven auto-regressive (DAR) models. DAR models provide a nonlinear and non-stationary spectral estimation of a signal, conditionally to another exogenous signal. We detail how inference can be done efficiently while guaranteeing model stability. We show how model comparison and hyper-parameter selection can be done using likelihood estimates. We also point out the limits of DAR models when the exogenous signal contains too high frequencies. Finally, we illustrate how DAR models can be applied on neuro-physiologic signals to characterize phase-amplitude coupling.
We address the issue of reliably detecting and quantifying cross-frequency coupling (CFC) in neural time series. Based on non-linear auto-regressive models, the proposed method provides a generative and parametric model of the time-varying spectral content of the signals. As this method models the entire spectrum simultaneously, it avoids the pitfalls related to incorrect filtering or the use of the Hilbert transform on wide-band signals. As the model is probabilistic, it also provides a score of the model “goodness of fit” via the likelihood, enabling easy and legitimate model selection and parameter comparison; this data-driven feature is unique to our model-based approach. Using three datasets obtained with invasive neurophysiological recordings in humans and rodents, we demonstrate that these models are able to replicate previous results obtained with other metrics, but also reveal new insights such as the influence of the amplitude of the slow oscillation. Using simulations, we demonstrate that our parametric method can reveal neural couplings with shorter signals than non-parametric methods. We also show how the likelihood can be used to find optimal filtering parameters, suggesting new properties on the spectrum of the driving signal, but also to estimate the optimal delay between the coupled signals, enabling a directionality estimation in the coupling. Author Summary Neural oscillations synchronize information across brain areas at various anatomical and temporal scales. Of particular relevance, slow fluctuations of brain activity have been shown to affect high frequency neural activity, by regulating the excitability level of neural populations. Such cross-frequency-coupling can take several forms. In the most frequently observed type, the power of high frequency activity is time-locked to a specific phase of slow frequency oscillations, yielding phase-amplitude-coupling (PAC). Even when readily observed in neural recordings, such non-linear coupling is particularly challenging to formally characterize. Typically, neuroscientists use band-pass filtering and Hilbert transforms with ad-hoc correlations. Here, we explicitly address current limitations and propose an alternative probabilistic signal modeling approach, for which statistical inference is fast and well-posed. To statistically model PAC, we propose to use non-linear auto-regressive models which estimate the spectral modulation of a signal conditionally to a driving signal. This conditional spectral analysis enables easy model selection and clear hypothesis-testing by using the likelihood of a given model. We demonstrate the advantage of the model-based approach on three datasets acquired in rats and in humans. We further provide novel neuroscientific insights on previously reported PAC phenomena, capturing two mechanisms in PAC: influence of amplitude and directionality estimation.
While most dereverberation methods focus on how to estimate the magnitude of an anechoic signal in the time-frequency domain, we propose a method which also takes the phase into account. By applying a harmonic model to the anechoic signal, we derive a formulation to compute the amplitude and phase of each harmonic. These parameters are then estimated by our method in presence of reverberation. As we jointly estimate the amplitude and phase of the clean signal, we achieve a very strong dereverberation on synthetic harmonic signals, resulting in a significant improvement of standard dereverberation objective measures over the state-of-the-art.
Most dereverberation methods aim to reconstruct the anechoic magnitude spectrogram, given a reverberant signal. Regardless of the method, the dereverberated signal is systematically synthesized with the reverberant phase. This corrupted phase reintroduces reverberation and distortion in the signal. This is why we intend to also reconstruct the anechoic phase, given a reverberant signal. Before processing speech signals, we propose in this paper a method for estimating the anechoic phase of reverberant chirp signals. Our method presents an accurate estimation of the instantaneous phase and improves objective measures of dereverberation.
Room acoustic parameters are key information for dereverberation or speech recognition. Usually, when one needs to assess the level of reverberation, only the reverberation time RT60 or a direct to reverberant sounds index Dτ is estimated. Yet, methods which blindly estimate the reverberation time from reverberant recorded speech do not always differentiate the RT60 from the Dτ to evaluate the level of reverberation. That is why we propose a method to jointly blindly estimate these parameters, from the signal energy decay rate distribution, by means of kernel regression. Evaluation is carried out with real and simulated room impulse responses to generate noise-free reverberant speech signals. The results show this new method outperforms baseline approaches in our evaluation.
Situation assessment is one of the basic abilities for robots to coexist with us in our day-to-day live. For a socially intelligent robot, different levels of situation assessment are required, ranging from basic processing of sensor input to high-level analysis of semantics and intention. The combination of various perception abilities greatly increases the robot’s socio-cognitive capabilities. However, this prompts new research challenges and the need of a coherent framework and architecture. Romeo2 is a unique project, aiming to bring multi-modal and multi-layered perception of situation assessment on a single system and targeting for a unified theoretical and functional framework for robot companion for everyday life. This paper presents different aspects of situation assessment identified and perceived within the Romeo2 project. It aims towards a principled approach to develop different components in a collaborative manner when such basic blocks should be functioning together and discusses about some of the innovation potentials such approach brings for the companion robotics domain.
We present a single channel method for late reverberation suppression. The proposed approach estimates late reverberation as a linear combination of previous time-frequency frames. We impose a sparsity constraint on the predictor in order to select the most relevant signal frames for the estimation. The dataset used for the evaluation is corrupted by background noise, thus we propose to jointly suppress background noise and late reverberation. This leads to an important improvement in the quality of the processed signals as well as an improvement of the automatic speech recognition scores. The method appears to be efficient mainly in far field conditions and in highly reverberant environments. In addition, it is suitable for real time processing.
Reverberation degrades speech intelligibility in telecommunications as well as it increases the word error rate in automatic speech recognition tasks. Several dereverberation methods have been proposed recently in order to counter these effects. In the single microphone case, the dereverberation problem is underdetermined and reverberation suppression approaches are preferred. In this paper we propose a novel method for single channel reverberation suppression. Late reverberation is estimated in the time-frequency domain as a sparse linear combination of previous frames. The predictors associated to the model are determined in a Lasso framework and a spectral subtraction filter is designed to produce the enhanced signal. This model does not require any additional information about the room acoustics and it is well suited for real-time applications. The method has state-of-the-art performance in terms of both reverberation suppression and spectral distortion.
For a socially intelligent robot, different levels of situation as-sessment are required, ranging from basic processing of sensor input to high-level analysis of semantics and intention. However, the attempt to combine them all prompts new research challenges and the need of a co-herent framework and architecture. This paper presents the situation assessment aspect of Romeo2, a unique project aiming to bring multi-modal and multi-layered perception on a single system and targeting for a unified theoretical and functional frame-work for a robot companion for everyday life. It also discusses some of the innovation potentials, which the combination of these various perception abilities adds into the robot's socio-cognitive capabilities.
Multichannel blind source separation performances rapidly degrade when the mixtures are highly reverberated. In fact, blind source separation algorithms usually focus on the separation task without dealing with the dereverberation problem. Some recent studies attempted to reduce the reverberation by introducing a dereverberation module before or after the blind source separation but only limited success was obtained in improving the separation performance in highly reverberant rooms. In this article, we conduct a number of experiments combining state of the art spectral enhancement- based dereverberation and source separation algorithms showing that, in this particular case, speech enhancement does not improve the performance of blind source separation.
We present an original framework for the detection of repeating objects in multimedia streams. This framework is designed so that it can work with any fingerprint model. A fingerprint is extracted for each incoming frame of the multimedia stream. The framework then manages this fingerprint so that if one similar frame comes later in the stream, it will be identified as a repetition. The framework has been tested with two distinct fingerprint models on simulated and `real-world' data. The results show that the framework performs well with both presented models and that it is suitable for industrial use-cases.
We propose an adaptive blind source separation algorithm in the context of robot audition using a microphone array. Our algorithm presents two steps: a fixed beamforming step to reduce the reverberation and the background noise and a source separation step. In the fixed beamforming preprocessing, we build the beamforming filters using the Head Related Transfer Functions (HRTFs) which allows us to take into consideration the effect of the robot's head on the near acoustic field. In the source separation step, we use a separation algorithm based on the l(1) norm minimization. We evaluate the performance of the proposed algorithm in a total adaptive way with real data and varying number of sources and show good separation and source number estimation results.
Gerard Chollet合作论文数CNRS (Centre National de la Recherche Scientifique)2