The mathematical theory of closed form functions for calculating LSFs on the basis of generating functions is presented. Exploiting recurrence relationships in the series expansion of Chebyshev polynomials of the first kind makes it possible to bootstrap iterative LSF-search from a set of characteristic polynomial zeros. The theoretical analysis is based on decomposition of sequences into symmetric and anti-symmetric polynomials defined as a series expansion of reduced Chebyshev polynomials of the first kind. Two variants of closed form functions are presented each characterised by using a recurrence relationship in Chebyshev polynomials. The first exploits the well known three terms recurrence relationships of Chebyshev polynomials. The second hitherto unused recurrence properties of Chebyshev coefficients defining a set of coefficients and zeros used for bootstrapping calculation of LSFs. The theory is tested using bootstrapped calculation of zeros and by evaluating the complexity of the closed form function. The results of the lower complexity calculations show that real axis zeros are within a given iteration tolerance when compared to results of a standard root-finder. Index Term: Recurrence relationships in series expansion of Chebyshev polynomials, generating functions, computational complexity, line spectral frequencies,
In the present paper, a method is proposed for adaptive estimation and tracking of roots of time-varying, complex, and univariate polynomials, e.g. z-transform polynomials that arise from finite signal sequences. The objective with the method is to alleviate the computational burden induced by factorization. The estimation is done by solving a set of linear equations; the number of equations equals the order of the polynomial. To avoid potential drifting of the estimations, it is proposed to verify with Aberth-Ehrlich’s factorization method at given intervals. A numerical experiment supplements theory by estimating roots of time-varying polynomials of different order. As a function of order, the proposed method has a lower run time than Lindsey-Fox and computing eigenvalues of companion matrices. The estimations are quite accurate, but tend to drift slightly in response to increasing coefficient pertubation lengths.
In recent studies, a non-parametric speech waveform representation (rep.) based on zeros of the z-transform (ZZT) has been proposed. The ZZT rep. has successfully been applied in separating mixed phase signals, e.g. pitch-synchronously windowed speech, into min/max phase by using the unit circle as discriminant. As the ZZT rep. is obtained by factorization of the z-transform, relations to the complex cepstrum (CC) exist. The present paper interrelates the ZZT rep. with the CC via factorization of the z-transform, and demonstrates that unit circle discrimination of a ZZT rep. can be formulated as a CC based separation by causality. A numerical experiment supplements theory by separating a range of LF glottal flow waveforms into their opening and closing phase constituents. Further, randomized mixed phase sequences are separated. As the CC based separation also can be obtained via FFT it has a lower time and space complexity than the ZZT based counterpart.
Current research has proposed a non-parametric speech waveform representation (rep) based on zeros of the z-transform (ZZT) [1] [2]. Empirically, the ZZT rep has successfully been applied in discriminating the glottal and vocal tract components in pitch-synchronously windowed speech by using the unit circle (UC) as discriminant [1] [2]. Further, similarity between ZZT reps of windowed speech, glottal flow waveforms, and waveforms of glottal flow opening and closing phases has been demonstrated [1] [3]. Therefore, the underlying cause of the separation on either side of the UC can be analyzed via the individual ZZT reps of the opening and closing phase waveforms; the waveforms are generated by the LF glottal flow model (GFM) [1]. The present paper demonstrates this cause and effect analytically and thereby supplement the previous empirical works. Moreover, this paper demonstrates that immiscibility is variant under changes in frame lengths; lengths that maximize or minimize immiscibility are presented. Index Terms: Zeros of the z-transform, LF glottal flow model, opening/closing phase separation
The objective of this study is to analyze speech signals using the zeros of the z-transform of the signal. Trajectories of the zeros are used to study the characteristics of speech production. The trajectories are obtained by varying the parameters of the window function used on the signal segment. A skew Poisson function is defined with three parameters to control the window function. The proposed method does not assume any model for the analysis. Both synthetic and natural speech signals are analyzed. The goal of this study is to demonstrate, and eventually separate the information of the source part from the vocal tract system part of the speech production process from the signal. This may provide new and additional insights into the speech production process over and above the existing methods such as the spectral and group delay methods. The results from experiments with varying Poisson window parameters and speech signals are presented and discussed.
The nonlocal means (NL-means) algorithm recently proposed for image denoising has proved highly effective for removing additive noise while to a large extent maintaining image details. The algorithm performs denoising by averaging each pixel with other pixels that have similar characteristics in the image. This letter considers the real and imaginary parts of complex speech spectrogram each as a separate image and presents a modified NL-means algorithm to them for denoising to improve the noise robustness of speech recognition. Recognition results on a noisy speech database show that the proposed method is superior to classical methods such as spectral subtraction.
The growth in wireless communication and mobile devices has supported the development of distributed speech recognition (DSR) technology. During the last decade this has led to the establishment of ETSI-DSR standards and an increased interest in research aimed at systems exploiting DSR. So far, however, DSR-based systems executing on mobile devices are only in their infancy. One of the reasons is the lack of easy-to-use software development packages. This chapter presents a prototype version of a configurable DSR system for the development of speech enabled applications on mobile devices.The system is implemented on the basis of the ETSI-DSR advanced front-end and the SPHINX IV recognizer. A dedicated protocol is defined for the communication between the DSR client and the recognition server supporting simultaneous access from a number of clients. This makes it possible for different clients to create and configure recognition tasks on the basis of a set of predefined recognition modes.
In this paper a half frame-rate (HFR) front-end is investigated for distributed speech recognition (DSR). The work is inspired from the need for low bit-rate and is justified by the redundancies known to exist in full frame-rate (FFR) features. At the client-side in the DSR architecture, implementation of the HFR is carried out by using double frame shifting as compared to the FFR resulting in the achievement of half the bit rate. At the server-side, each HFR feature vector is repeated once to construct the FFR features and no changes are therefore required in the recognition back-end. It is experimentally justified that the performance achieved by HFR is comparable to FFR and that repetition of each HFR feature vector is critical for the HFR front-end to maintain the performance. Motivated by the effectiveness of HFR, a number of additional FFR-based DSR schemes are further presented. Finally, this paper introduces an adaptive multi-frame-rate scheme in which the DSR system adapts to the characteristics of the transmission channel by switching between HFR and the FFR-based schemes. This multi-frame-rate scheme is found to be superior to the basic FFR
This paper presents an effective feature processing algorithm for robust speech recognition, based on combined spectral and cepstral processing. The spectral processing consists of FullWave Rectification Spectral Subtraction (FWR-SS) and Likelihood Controlled Instantaneous Noise Estimation (LCINE) while the cepstral processing is based on meanand variance normalisation. The combination is motivated by the fact that the (usually) one frame based spectral subtraction introduces large statistical mismatches between clean and enhanced noisy speech in the cepstral domain, resulting in a degradation of the recognition performance. The introduced cepstral processing is able, to some extent, to mitigate these mismatches and in this sense the two methods are not just combined but shown to be complementary. Statistical analyses as well as recognition experiments are conducted on the Aurora 2 database and a performance comparable to the much more complex ETSI advanced front-end is achieved.
This paper presents research on two aspects of distributed speech recognition (DSR) in the presence of channel transmission errors in wireless network environments. The first is on experiments with a frame-based channel error protection scheme, where in previous research we reported results from experiments using randomly distributed bit-errors. This paper presents results from experiments using three additional, more realistic error distributions: burst-like packet loss, GSM error patterns and UMTS statistics. The second is on exploiting the knowledge about channel transmission errors for the purpose of optimising the Out-of-Vocabulary (OOV) detection. Transmission errors influence the acoustic likelihood, and therefore affect the optimal threshold setting for discrimination between In-Vocabulary (IV) words and OOV words. An OOV-detection method is proposed in which the estimated Frame-Error-Rate (FER) is used to adjust the discrimination threshold. Results from experiments are reported over a range of transmission errors.
This paper describes ongoing research preparing for the widespread deployment of spoken language processing in networks encompassing wired and wireless transmission channels. The paper gives a brief overview of the standardized bit-error protection scheme aimed at minimising channel transmission errors and used within the distributed speech recognition (DSR) paradigm. Within the ETSI-DSR standard, two quantised mel-spectral frames – each of 10 ms duration are grouped together and protected with a 4-bit Cyclic Redundancy Checking (CRC) forming a frame-pair. However, this causes the entire frame-pair erroneous if a one-bit error only occurs in the frame-pair packet. Over an error-prone transmission channel this format will cause severe problems. To overcome this, the paper presents a one-frame architecture in which a 4-bit CRC is calculated to protect each frame independently. This scheme results in that the overall probability of one frame in error is lower, or that an error occurring in one frame does not affect another frame. A number of simple recognition experiments have been conducted to verify the introduction of the one-frame CRC protection scheme for a number of simulated transmission channel biterror rates (BER) ranging from 0 (no transmission channel involved) to 2٠10. Experimental results show that the one-frame protection scheme is more robust to channel errors although a slight increase in the errorprotection overhead is needed.
Kristian G. Olesen合作论文数Department of Computer Science;Aalborg University2