In studies on artificial bandwidth extension (ABE), there is a lack of international coordination in subjective tests between multiple methods and languages. Here we present the design of absolute category rating listening tests evaluating 12 ABE variants of six approaches in multiple languages, namely in American English, Chinese, German, and Korean. Since the number of ABE variants caused a higher-than-recommended length of the listening test, ABE variants were distributed into two separate listening tests per language. The paper focuses on the listening test design, which aimed at merging the subjective scores of both tests and thus allows for a joint analysis of all ABE variants under test at once. A language-dependent analysis, evaluating ABE variants in the context of the underlying coded narrowband speech condition showed statistical significant improvement in English, German, and Korean for some ABE solutions.
There is a considerable performance gap between the current scalable audio coding schemes and a nonscalable coder operating at the same bitrate. This suboptimality results from the independent coding of the layers in these systems. One of the aspects that plays a role in this suboptimality is the entropy coding. In practical audio coding systems including MPEG advanced audio coding (AAC), the transform domain coefficients are quantized using an entropy-constrained quantizer. In MPEG-4 scalable AAC (S-AAC), the quantization and coding are performed separately at each layer. In case of Huffman coding, the redundancy introduced by the entropy coding at each layer is larger at lower quantization resolutions. Also, the redundancy for the overall coder becomes larger as the number of layers increases. In fact, there is a tradeoff between the overall redundancy and the fine-grain scalability in which the bitrate per layer is smaller and more layers are required. In this paper, a fine-grain scalable coder for audio signals is proposed where the entropy coding of a quantizer is made scalable via joint design of entropy coding and quantization. By constructing a Huffman-like coding tree where the internal nodes can be mapped to the reconstruction points, the tree can be pruned at any internal node to control the rate-distortion (RD) performance of the encoder in a fine-grain manner. A set of metrics and a trellis-based approach is proposed to create a coding tree so that an appropriate path is generated on the RD plane. The results show the proposed method outperforms the scalable audio coding performed based on reconstruction error quantization as used in practical systems, e.g., in S-AAC.
In this paper, we present a new method for noise power spectral density (PSD) matrix estimation based on IMCRA which consists of two parts. For the auto-PSD (diagonal) estimation, we propose a modification to IMCRA where a special level detector is employed to improve the tracking of non-stationary noise backgrounds. For the cross-PSD (offdiagonal) estimation, we propose to calculate a smoothed cross-periodogram by using estimated noise components derived as residuals after the application of a speech enhancement algorithm on the individual microphone signals. Simulation results show the effectiveness of our proposed approach in estimating the noise PSD matrix and its robustness against reverberation when used in combination with an MVDR-based speech enhancement system.
Considering the properties of the residual signal, core-based bit-plane probabilities are provided for MPEG-4 Audio Scalable to Lossless Coding (SLS), which matches the quantization and coding performed in the core layer. Using the same strategy, new probabilities are obtained to consider the clipping effect in bit-plane coding of an unbounded signal, which is useful for non-core mode of SLS coding. Simulations show that considering the core layer parameters and the clipping effect improve the bit-plane probabilities estimation compared to the existing method.
A scalable audio coding method is proposed using a technique, Quantization Index Modulation, borrowed from watermarking. Some of the information of each layer output is embedded (watermarked) in the previous layer. This approach leads to a saving in bitrate while keeping the distortion almost unchanged. This makes the scalable coding system more efficient in terms of Rate-Distortion. The results show that the proposed method outperforms the scalable audio coding based on reconstruction error quantization which is used in practical systems such as MPEG-4 AAC.
A fine grain scalable coding for audio signals is proposed where the entropy coding of the quantizer outputs is made scalable. By constructing a Huffman-like coding tree where internal nodes can be mapped to reconstruction points, we can prune the tree to control the distortion of the quantizer. Our results show the proposed method improves existing similar work and significantly outperforms scalable coding based on reconstruction error quantization as used in practical systems, eg. MPEG-4 audio.
In this paper, we extend our previous work on exploiting speech temporal properties to improve Bandwidth Extension (BWE) of narrowband speech using Gaussian Mixture Models (GMMs). By quantifying temporal properties through information theoretic measures and using delta features, we have shown that narrowband memory significantly increases certainty about highband parameters. However, as delta features are non-invertible, they can not be directly used to reconstruct highband frequency content. In the work presented herein, we embed temporal properties indirectly into the GMM structure through a memorydependent tree-based approach to extend representation of the narrow band. In particular, sequences of past frames are progressively used to grow the GMM in a tree-like fashion. This growth approach results in reliable estimates for the GMM parameters such that Maximum Likelihood estimation is no longer necessary, thus circumventing the complexity accompanying high-dimensionality GMM training.
This paper examines enhancement to ITU-T Recommendation G.711.1 PCM wideband extension speech coder. To further improve the core lower-band coding performance the use of vector quantization and delayed decision coding is studied. A particular case of delayed decision coding, tree encoding, is implemented in the above standard. The bitstream is compatible with both the legacy G.711 and the G.711.1 decoder. PESQ (ITUT P.862, Perceptual Evaluation of Speech Quality) is used to evaluate the performance. Both the vector quantizer and tree encoder have better performance than the original core layer encoder. Index Terms: speech coding, G.711.1, tree encoding
ITU-T G.711.1 is a multirate wideband extension for the well known ITU-T G.711 pulse code modulation of voice frequencies. The extended system is fully interoperable with the legacy narrowband one. In the case where the legacy G.711 is used to code a speech signal and G.711.1 is used to decode it, quantization noise may be audible. For this situation, the standard proposes an optional postfilter. The application of postfiltering requires an estimation of the quatization noise. In this paper we review the process of estimating this coding noise and we propose a better noise estimator.
In Voice-over-IP, the quality of interactive conversation is important to users. Quality-based playout buffering seeks an optimum balance between delay and loss. However, such a scheme still suffers when packet losses are bursty. Path diversity can alleviate the effect of losses and improve perceived quality by providing redundancy. In this paper, a new scheme is proposed which evaluates the performance of both paths. We consider three different path diversity schemes. The playout scheduling algorithms are designed based on conversational quality including both calling quality and interactivity. The simulation results show the efficacy of our algorithms in correcting for losses (isolated and burst) and improving perceived conversational quality.
This paper examines the correlation properties of quantization noise. The quantization noise energy is subtractive if the quantizer output levels are optimized for the probability density of the input signal (pdf optimized). This paper gives a new result that shows that a quantizer (uniform or not) which has quantizer break pointsmidway between output levels (a minimum distance quantizer) and is scaled to minimize the mean-square error, also has this property. Examples are shown that show the correlation properties which determine whether the quantization noise energy is subtractive or additive. This paper also considers a postfilter configuration that compensates for the quantization noise. The postfilter frequency domain gains take the correlation properties of the quantization noise into account. An experiment on reducing the effect of quantization noise in speech gives an indication that taking account of the correlation is useful.
In Voice-over-IP, buffer delay and packet loss are two main factors effecting perceived conversational quality. A quality-based algorithm aims to seek an optimum balancing of delay versus loss. To improve perceived quality further, steps should be taken to mitigate the effect of losses due to network (missing packets) and buffer underflow (late packets) without increasing buffer delays. In this paper, we propose a quality-based playout algorithm with an FEC design based on conversational quality including calling quality and interactivity. The simulation results show our algorithm's efficiency of correcting for losses (isolated and burst) and improving perceived conversational quality.
In this paper, we continue our previous work on improving Bandwidth Extension (BWE) of narrowband speech. We have shown that including memory into the parametrization frontend (through delta features) results in higher highband certainty irrespective of feature type, with MFCCs exhibiting higher correlation, in general, between both bands, reaching twice that using LSFs. By incorporating memory into the frontend of a conventional LP-based BWE system, we were able to translate the higher correlation due to memory into BWE performance improvement. Using high-resolution inverse DCT, we also achieved high quality speech reconstruction from MFCCs, thus enabling MFCC-based BWE with improved performance compared to conventional static LP-based BWE. We continue this work by incorporating the superior correlation properties of frontend memory into our MFCC-based BWE system. Log-Spectral Distortion as well as the more perceptually-correlated Itakura-based measures show that incorporating memory into our MFCC-based BWE system results in BWE performance superior to that of our dynamic LP-based BWE system.
In Voice-over-IP, jitter buffers are introduced at both sides of the sender and the receiver to compensate for delay jitters. A longer buffer reduces the possibility of packet loss and packet disorder at the expense of increasing conversational delays. In this paper, we propose a novel criterion for the calling quality of conversational VoIP, including the effect of delay on interactivity of a conversation. Using this criterion, we propose a quality-based playout scheduling algorithm with improved voice quality and reduced conversational delays. The Simulation results show that the proposed algorithm can achieve the best calling quality compared with other algorithms.
This report examines the time windows used for linear prediction (LP) analysis of speech. The goal of windowing is to create frames of data each of which will be used to calculate an autocorrelation sequence. Several factors enter into the choice of window. The time and spectral properties of Hamming and Hann windows are examined. We also consider windows based on Discrete Prolate Spherical Sequences including multiwindow analysis. Multiwindow analysis biases the estimation of the correlation more than single window analysis. Windows with frequency responses based on the ultraspherical polynomials are discussed. This family of windows includes Dolph-Chebyshev and Saramaki windows. This report also considers asymmetrical windows as used in modern speech coders. The frequency response of these windows is poor relative to conventional windows. Finally, the presence of a “pedestal” in the time window (as in the case of a Hamming window) is shown to be deleterious to the time evolution of the LP parameters. Time Windows for Linear Prediction of Speech 1 TimeWindows for Linear Prediction of Speech
In this paper, we present a perceptual audio coding method that encodes the audio using perceptually salient envelope features. These features are found by passing the audio through a set of gammatone filters, and then computing the Hilbert envelopes of the responses. Relevant points of these envelopes are isolated and transmitted to the decoder. The decoder reconstructs the audio in an iterative manner from these relevant envelope points. Initial experiments suggest that even without sophisticated entropy coding a moderate bitrate reduction is possible while retaining good quality.
We present a novel algorithm for defining the lengths of subcarrier equalizers employed by wireless multicarrier transmission systems operating in frequency-selective fading channels. The equalizer lengths across the subcarriers are incrementally varied in a ldquogreedyrdquo fashion until the global cost function is below some prescribed threshold. By varying the equalizer lengths, the overall complexity of the equalization is constrained while the system meets a minimum error performance. Moreover, we investigate four strategies for terminating the proposed algorithm when an adequate number of equalizer taps have been allocated in this process. The results show that a system that employs variable-length equalizers defined by the proposed algorithm can achieve an improvement in error robustness of as much as an order of magnitude, relative to a system that employs constant length equalizers with the same overall complexity.
In recent years, Automatic Speech Recognition (ASR) systems designed to work in controlled environments using clean speech have reached very high levels of performance. However, the accuracy of speech recognition degrades severely when the systems are operated in noisy environments. In this thesis we address the problem of single-channel speech enhancement. Starting with a study of the state-of-the-art enhancement methods, a comprehensive study of different categories of speech enhancement is presented. As an important class of speech enhancement methods, subspace-based speech enhancement is presented in chapter 2. After a careful study of all forces and drawbacks of this technique, a generalized form of Principal Component Analysis-based (PCA-based) speech enhancement is provided next. As a vital issue in PCA-based enhancement methods, identification of the clean speech signal's model is investigated in chapter 3. Some recent techniques to define the rank of a clean speech signal are presented in this chapter. In the rest of the thesis, a novel technique for rank estimation is developed. We introduce therefore a novel approach for the optimal subspace partitioning using the Variance of the Reconstruction Error (VRE) criterion. This criterion provides consistent parameter estimates and allows us to implement an automatic noise reduction algorithm that can be simply applied to the observed data. This choice also overcomes many limitations encountered with other selection criteria, like overestimation of the signal subspace or the need for empirical parameters. We have also extended our subspace algorithm to take into account the case of colored and babble noise. Informal listening tests and illustrations have confirmed the method to be numerically noise robust regardless of the type of the noise. ii Acknowledgements
We present a novel MFCC-based scheme for the Bandwidth Extension (BWE) of narrowband speech. BWE is based on the assumption that narrowband speech (0.3-3.4 kHz) correlates closely with the highband signal (3.4-7 kHz), enabling estimation of the highband frequency content given the narrow band. While BWE schemes have traditionally used LP-based parametrizations, our recent work has shown that MFCC parametrization results in higher correlation between both bands reaching twice that using LSFs. By employing high-resolution IDCT of highband MFCCs obtained from narrowband MFCCs by statistical estimation, we achieve high-quality highband power spectra from which the time-domain speech signal can be reconstructed. Implementing this scheme for BWE translates the higher correlation advantage of MFCCs into BWE performance superior to that obtained using LSFs, as shown by improvements in log-spectral distortion as well as Itakura-based measures (the latter improving by up to 13%).
Fabrice Labeau合作论文数Electrical and Computer Engineering Department
McGill University11