
In glottal source analysis, the phase minimization criterion has already been proposed to detect excitation instants. As shown in this paper, this criterion can also be used to estimate the shape parameter of a glottal model (ex. Liljencrants-Fant model) and not only its time position. Additionally, we show that the shape parameter can be estimated independently of the glottal model position. The reliability of the proposed methods is evaluated with synthetic signals and compared to that of the IAIF and minimum/maximum-phase decomposition methods. The results of the methods are evaluated according to the influence of the fundamental frequency and noise. The estimation of a glottal model is useful for the separation of the glottal source and the vocal-tract filter and therefore can be applied in voice transformation, synthesis, and also in clinical context or for the study of the voice production.
Speaker diarization is the task of determining “who spoke when?” in an audio or video recording that contains an unknown amount of speech and also an unknown number of speakers. Initially, it was proposed as a research topic related to automatic speech recognition, where speaker diarization serves as an upstream processing step. Over recent years, however, speaker diarization has become an important key technology for many tasks, such as navigation, retrieval, or higher level inference on audio data. Accordingly, many important improvements in accuracy and robustness have been reported in journals and conferences in the area. The application domains, from broadcast news, to lectures and meetings, vary greatly and pose different problems, such as having access to multiple microphones and multimodal information or overlapping speech. The most recent review of existing technology dates back to 2006 and focuses on the broadcast news domain. In this paper, we review the current state-of-the-art, focusing on research developed since 2006 that relates predominantly to speaker diarization for conference meetings. Finally, we present an analysis of speaker diarization performance as reported through the NIST Rich Transcription evaluations on meeting data and identify important areas for future research.
This paper presents a theoretical framework to analyze the relative merits of the two most general, dominant approaches to speaker diarization involving bottom-up and top-down hierarchical clustering. We present an original qualitative comparison which argues how the two approaches are likely to exhibit different behavior in speaker inventory optimization and model training: bottom-up approaches will capture comparatively purer models and will thus be more sensitive to nuisance variation such as that related to the speech content; top-down approaches, in contrast, will produce less discriminative speaker models but, importantly, models which are potentially better normalized against nuisance variation. We report experiments conducted on two standard, single-channel NIST RT evaluation datasets which validate our hypotheses. Results show that competitive performance can be achieved with both bottom-up and top-down approaches (average DERs of 21% and 22%), and that neither approach is superior. Speaker purification, which aims to improve speaker discrimination, gives more consistent improvements with the top-down system than with the bottom-up system (average DERs of 19% and 25%), thereby confirming that the top-down system is less discriminative and that the bottom-up system is less stable. Finally, we report a new combination strategy that exploits the merits of the two approaches. Combination delivers an average DER of 17% and confirms the intrinsic complementary of the two approaches.
This paper discusses two behavioural interfaces for reliability analysis: dynamic fault trees, which model the system reliability in terms of the reliability of its components and Arcade, which models the system reliability at an architectural level. For both formalisms, the reliability is analyzed by transforming the DFT or Arcade model to a set of input-output Markov Chains. By using compositional aggregation techniques based on weak bisimilarity, significant reductions in the state space can be obtained.
A novel idea for introducing concurrency in least squares (LS) adaptive algorithms by sacrificing optimality has been proposed. The resultant class of algorithms provides schemes to fill the wide gap in the convergence rates of LS and stochastic gradient (SG) algorithms. It will be particularly useful in the real time implementations of large-order linear and Volterra filters for which both the LS and SG algorithms are unsuited.
This paper is devoted to the maximum likelihood estimation of multiple sources in the presence of unknown noise. With the spatial noise covariance modeled as a function of certain unknown parameters, e.g., an autoregressive (AR) model, a direct and systematic way is developed to find the exact maximum likelihood (ML) estimates of all parameters associated with the direction finding problem, including the direction-of-arrival (DOA) angles /spl Theta/, the noise parameters /spl alpha/, the signal covariance /spl Phi//sub s/, and the noise power /spl sigma//sup 2/. We show that the estimates of the linear part of the parameter set /spl Phi//sub s/ and /spl sigma//sup 2/ can be separated from the nonlinear parts /spl Theta/ and /spl alpha/. Thus, the estimates of /spl Phi//sub s/ and /spl sigma//sup 2/ become explicit functions of /spl Theta/ and /spl alpha/. This results in a significant reduction in the dimensionality of the nonlinear optimization problem. Asymptotic analysis is performed on the estimates of /spl Theta/ and /spl alpha/, and compact formulas are obtained for the Cramer-Rao bounds (CRB's). Finally, a Newton-type algorithm is designed to solve the nonlinear optimization problem, and simulations show that the asymptotic CRB agrees well with the results from Monte Carlo trials, even for small numbers of snapshots. >
The family of lapped orthogonal transforms is extended to include basis functions of arbitrary length. Within this new family, the extended lapped transform (ELT) is introduced, as a generalization of the previously reported modulated lapped transform (MLT). Design techniques and fast algorithms for the ELT are presented, as well as examples that demonstrate the good performance of the ELT in signal coding applications. Therefore, the ELT is a promising substitute for traditional block transforms in transform coding systems, and also a good substitute for less efficient filter banks in subband coding systems.
Eigenstructure methods for estimating angles of arrival of radiation sources generally require complex computations in computing eigencomponents of the covariance matrix and calculating the search function. A unitary transformation method that transforms the complex covariance matrix of an equally spaced linear array, which is Hermitian persymmetric, and the complex search vector into a real symmetric matrix and a real vector, respectively is presented. Both tasks can be accomplished by real computations. The sampled covariance matrix available is not persymmetric. To suit the unitary transformation method, a persymmetrized estimator of the sampled covariance matrix, which is optimal in the sense of Euclidean distance, is proposed.< >
The problem of enhancing speech degraded by stationary and nonstationary additive white noise is addressed. The authors have explored a new idea, enhancing speech based on auditory evidence, for this traditional topic. Distinguishing different objectives for heavy and light noise interference, two related algorithms have been developed. For speech degraded by heavy noise, the improvement in signal-to-noise ratio (SNR) is as high as 12 dB; for lightly noisy speech, the improvement is modest and decreases as the SNR of the noisy speech increases. Quantizing noise is used to assess the capacity of reducing nonstationary noise with the algorithms; a significant reduction of such noise and an improvement in speech quality have been achieved. The advantages of the proposed algorithms for speech enhancement include no need for prior knowledge of the noise and only a modest computational requirement.
The problem of signal parameter estimation of narrowband emitter signals impinging on an array of sensors is addressed. A multidimensional estimation procedure that applies to arbitrary array structures and signal correlation is proposed. The method is based on the recently introduced weighted subspace fitting (WSF) criterion and includes schemes for both detecting the number of sources and estimating the signal parameters. A Gauss-Newton-type method is presented for solving the multidimensional WSF and maximum-likelihood optimization problems. The global and local properties of the search procedure are investigated through computer simulations. Most methods require knowledge of the number of coherent/noncoherent signals present. A scheme for consistently estimating this is proposed based on an asymptotic analysis of the WSF cost function. The performance of the detection scheme is also investigated through simulations.< >
Convergence of frequency-domain adaptive pole-zero IIR (infinite impulse response) filter is studied. The algorithm is shown to converge in probability to an associated ordinary differential equation (ODE) which in turn converges to a local minimum of its performance surface. An analysis of the performance surface shows that the algorithm converges to one of N-factorial members in an equivalence class of global minimum points, where N is the number of adaptive poles. Saddle points exist on manifolds that separate members in the equivalence class. This explains 'shoulders' in the MSE convergence curves and also suggests one way of avoiding these shoulders which cause slow convergence. A second-order simulation example confirms the above results.< >
A fast algorithm for implementation of the QR-factorization-based recursive-least-squares (RLS) adaptive filter is discussed. This fast adaptive rotors (FAR) algorithm can be implemented with a pipelined array of processors called ROTORs and CISORs. The ROTORs compute 2*2 orthogonal (Givens) rotations, and the CISORs compute the cosines and sines of the angles used in the ROTORs. The algorithm requires 4N ROTORs and 2N CISORs at each iteration to compute the solution to the RLS problem. The algorithm is numerically stable. The FAR algorithm is derived using a single generic updating formula for orthogonal matrices, which is introduced and derived. Whereas the generic updating formula is reminiscent of previous fast transversal filters and fast lattice algorithms, the set of internally propogated adaptive filter quantities is entirely different and constitutes yet another complete characterization of the RLS covariance and the forward, backward, and pinning estimation problems.< >
An efficient zooming FFT (fast Fourier transform) algorithm that allows center padding sinc function interpolation of 2-D images is presented. This algorithm avoids the phase shifts that would be introduced if the efficient Skinner interpolation method is used. Output pruning is incorporated to allow efficient determination of a zoomed subimage. Time savings of more than 50% can be achieved. Example images illustrating the use of the algorithm in conjunction with zooming and ARMA (autoregressive moving average) modeling of data are given.< >
The author introduces a scheme for the local processing of visual information, called the Hermite transform. The problem is addressed from the point of view of image coding, and therefore the scheme is presented as an analysis/resynthesis system. The objectives of the present work, however, are not restricted to coding. The analysis part is designed so that it can also serve applications in the area of computer vision. Indeed, derivatives of Gaussians, which have found widespread application in feature detection over the past few years, play a central role in the Hermite analysis. It is also argued that the proposed processing scheme is in close agreement with current insight into the image processing that is carried out by the human visual system. In particular, it is demonstrated that the Hermite transform is in better agreement with human visual modeling than Gabor expansions. >
A class of constrained adaptive filters, called recursive center-frequency adaptive filters, is applied to the tracking of bandpass signals. For this application, these filters form a class of completely digital tracking filters which are shown to have several advantages over existing analog and semidigital tracking filters. A procedure of analyzing the tracking behavior of these filters using the ordinary differential equation approach is presented. The tracking performance of two examples of these filters is analyzed for step, ramp, and sinusoidal variations of the input signal center frequency, and these are shown to track the signals very effectively. The dependency of the performance on the parameters of the algorithm, transfer function of the filter, and the input signal is studied
The solution of l/sub 2/ (minimum variance) and H/sub infinity / estimation problems is considered using a polynomial systems approach. The results for the l/sub 2/ filtering problem, which corresponds with Wiener or Kalman filtering/prediction, are first presented in polynomial matrix form. Attention then turns to the solution of the H/sub infinity / estimation problem for scalar systems. Numerous examples are presented to illustrate the computational procedures. The two types of estimator are appropriate to very different estimation problems and the new H/sub infinity / devices should be valuable in certain application areas. >
An improved subspace approach for high-resolution array processing is presented. The approach is formulated as a constrained least squares minimization problem. The relationship that the proposed approach has to other subspace techniques is explored. Numerical results to illustrate the performance achievable are presented.< >
The authors generalize the weighted redundancy transform (WRT) algorithm for computing the multidimensional discrete Fourier transform (DFT) in the case in which the sample size (blocklength) is not the same on every axis. The proposed algorithm, like the WRT algorithm, is based on the one-dimensional fast Fourier transform (FFT) and, compared to the traditional ways of computing the multidimensional DFT, offers substantial savings in the number of one-dimensional FFT procedure calls. While the algorithm is applicable to transforms of any dimensions, only the two-dimensional case is explored in detail
The conventional Huffman table, which stores a set of code words with different lengths, is difficult to construct. An easily constructed table that includes a code-length array and an encoding mapping array can accompany an encoding procedure to replace the conventional Huffman table. The code-length array records the number of code words of the same length, and the index of this array represents the code length. The encoding mapping array records the indexes of the probability vector for the corresponding symbols, with the elements of the probability vector in decreasing order. The code word generated by the encoding procedure is a new modified Huffman code that can be decoded by a simple procedure. The hardware structure for the encoding and decoding procedures is presented. The required maximum number of clock cycles for encoding and decoding are (N+1) and (N+2), respectively, where N is the number of symbols in the source alphabet. The encoding procedure can accompany the generation of the code-length array, replacing many Huffman tables required in the coding system with a time-varying probability vector. The algorithm has a good performance with respect to speed and memory requirements.<>
An approach to power spectrum estimation that is based on a separable cross-entropy modeling procedure is presented. The authors start with a model of a multichannel, multidimensional, stationary Gaussian random process that is sampled on a nonuniform grid. An approximate separable model in which selected frequency samples of the process are modeled as independent random variables, is then fitted to it. Two cross-entropy-like criteria are used to select optimal separable approximations. One of them yields a spectral estimation algorithm that is a generalized version of Capon's maximum-likelihood method, and the other is similar to classical windowing methods. They discuss different strategies for designing bandpass filters for use with the cross-entropy approach