In shouting, speakers use increased vocal effort to convey spoken messages over distance or above environmental noise. For automatic speaker recognition systems trained using normal speech, shouting causes a severe vocal effort mismatch between the enrollment and test hence reducing the recognition performance. In this study, two compensation methods are proposed to tackle the mismatch in a shouted versus normal speaker recognition task. These techniques are applied in the feature extraction stage of a speaker recognition system to modify the spectral envelopes of shouts to be closer to those in normal speech. The techniques modify the all-pole power spectrum of the MFCC computation chain with shouted-to-normal compensation filtering that is obtained using a GMM-based statistical mapping. In an evaluation using the state-of-the-art i-vector based recognition system, the proposed techniques provided considerable improvements in identification rates compared to the case when shouted speech spectra were not processed.
A novel approach for proximity detection on mobile handsets which does not require any additional transducers is presented. The method is based on transmitting a chirp and processing the received signal by applying Least Mean Square (LMS), where the desired signal is the transmitted chirp. The envelope of three signals (estimated filter taps, estimated output and error signal) are characterized with a set of 12 features which are used to classify a given frame into one of two classes: proximity active or proximity inactive. The classifier employed is based on Support Vector Machine (SVM) with linear kernel. The results show that over 13 minutes of recorded data, the accuracy achieved is 95.28% using 10-fold cross-validation. Furthermore, the feature importance analysis performed on the database indicates that the most relevant feature is based on the estimated filter taps.
Speaker recognition performance degrades substantially in case of vocal effort mismatch (e.g. shouted vs. normal speech) between test and enrollment utterances. Such a mismatch is often encountered, for example, in forensic speaker recognition. This paper introduces a novel spectral mapping method which, when employed jointly with a statistical mapping technique, converts the Mel-frequency band energies of normal speech towards their counterparts in shouted speech. The aim is to obtain more robust performance in speaker recognition by tackling vocal effort mismatch between enrollment and test utterances. The processing is performed on the speech signal before feature extraction. The proposed approach was evaluated by testing the performance of a state-of-the-art i-vector-based speaker recognition system with and without applying the spectral mapping processing to the enrollment data. The results show that pre-processing with the proposed approach results in considerable improvement in correct identification rates.
The 2016 speaker recognition evaluation (SRE'16) is the latest edition in the series of benchmarking events conducted by the National Institute of Standards and Technology (NIST).I4U is a joint entry to SRE'16 as the result from the collaboration and active exchange of information among researchers from sixteen Institutes and Universities across 4 continents.The joint submission and several of its 32 sub-systems were among topperforming systems.A lot of efforts have been devoted to two major challenges, namely, unlabeled training data and dataset shift from Switchboard-Mixer to the new Call My Net dataset.This paper summarizes the lessons learned, presents our shared view from the sixteen research groups on recent advances, major paradigm shift, and common tool chain used in speaker recognition as we have witnessed in SRE'16.More importantly, we look into the intriguing question of fusing a large ensemble of sub-systems and the potential benefit of large-scale collaboration.
A linear predictive spectral estimation method based on higher-lag autocorrelation coefficients is proposed for the noise-robust feature extraction from speech. The method, called higher-lag linear prediction, is derived from a signal prediction model that is optimized in the mean square sense using a cost function that has two prediction error terms, the first of which is similar to that of conventional linear prediction and the second of which is a delayed version introducing an integer delay of M samples. This basic form is developed further into the combined higher-lag linear prediction (CHLLP) model by simultaneously taking advantage of the zero-lag and higher-lag predictions. The CHLLP model was used in the computation of mel-frequency cepstral coefficients and compared with several reference feature extraction methods in speaker recognition. The experiments were conducted by using a modern i-vector-based system. Noise-corruption was done using both additive car, babble, and factory noise in different signal-to-noise ratio conditions as well as speech recordings from real noisy conditions. The results indicate that CHLLP outperformed the reference feature extraction methods in almost all the comparisons in the noise-corrupted conditions and the performance of CHLLP was only slightly inferior to the nonparametric FFT-based spectral modeling in the clean condition.
State-of-the-art language recognition systems involve modeling utterances with the i-vectors. However, the uncertainty of the i-vector extraction process represented by the i-vector posterior covariance is affected by various factors such as channel mismatch, background noise, incomplete transformations and duration variability. In this paper, we propose a new quality measure based on the i-vector posterior covariance and incor-porate it into the recognition process to improve the recognition accuracy. The experimental results with LRE15 database and various duration conditions show a 2 . 9% relative improvement in terms of average performance cost as a result of incorporating the proposed quality measure in language recognition systems.
Wearing a face mask affects the speech production. On top of that, the frequency response and radiation characteristics of the face mask depending on the material and shape of the mask adds to the complexity of analyzing speech under face mask. Our target is to separate the effect of muscle constriction and increased vocal effort in speech produced under face mask from sound transmission and radiation properties of face mask. In this paper, we measure up the far-field effects of wearing four different face masks; motorcycle helmet, rubber mask, surgical mask and scarf inside anechoic chamber. The measurement setup follows the recording configuration of a speech corpus used for speaker recognition experiments. In matching speech under face mask with speech under no mask, the frequency response of the respective face mask is accounted for and compensated for before acoustic feature extraction. The speaker recognition performance is reported using the state-of-the-art i-vector method for mismatched and compensated conditions in order to demonstrate the significance of knowing the type of mask and accounting for its sound transmission properties.
Degraded signal quality and incomplete voice probes have severe effects on the performance of a speaker recognition system.Unified audio characteristics (UACs) have been proposed to quantify multi-condition signal degradation effects into posterior probabilities of quality classes.In previous work, we showed that UAC-based quality vectors (q-vectors) are efficient at the score-normalization stage.Hence, we motivate qvector based calibration by using functions of quality estimates (FQEs).In this work, we examine the robustness of calibration approaches to low-SNR and short-duration conditions utilizing measured and estimated quality indicators.Thereby, comparisons are drawn to quality measure functions (QMFs) employing oracle SNRs and sample duration.In the robustness study, low-SNR and short-duration conditions are excluded from calibration training.The present analysis provides insights on the behavior of calibration schemes in combined conditions of high signal degradation and short segment duration regarding accurate approximation of idealized calibration.We seek calibration methods in order to parsimonious preserve robustness against unseen data.A separate analysis is provided on duration-and noise-only scenarios as well as on combined duration and noise scenarios.QMFs and FQE reduce Cmc costs down to 5 -6% of conventional calibration schemes if all conditions are known, and to 10 -12% in the presence of unseen conditions.
During the past three decades, the issue of processing spectral phase has been largely neglected in speech applications. There is no doubt that the interest of speech processing community towards the use of phase information in a big spectrum of speech technologies, from automatic speech and speaker recognition to speech synthesis, from speech enhancement and source separation to speech coding, is constantly increasing. In this paper, we elaborate on why phase was believed to be unimportant in each application. We provide an overview of advancements in phase-aware signal processing with applications to speech, showing that considering phase-aware speech processing can be beneficial in many cases, while it can complement the possible solutions that magnitude-only methods suggest. Our goal is to show that phase-aware signal processing is an important emerging field with high potential in the current speech communication applications. The paper provides an extended and up-to-date bibliography on the topic of phase aware speech processing aiming at providing the necessary background to the interested readers for following the recent advancements in the area. Our review expands the step initiated by our organized special session and exemplifies the usefulness of spectral phase information in a wide range of speech processing applications. Finally, the overview will provide some future work directions. (C) 2016 Elsevier B.V. All rights reserved.
Linear discriminant analysis (LDA) is a powerful technique in pattern recognition to reduce the dimensionality of data vectors. It maximizes discriminability by retaining only those directions that minimize the ratio of within-class and between-class variance. In this paper, using the same principles as for conventional LDA, we propose to employ uncertainties of the noisy or distorted input data in order to estimate maximally discriminant directions. We demonstrate the efficiency of the proposed uncertain LDA on two applications using state-of-the-art techniques. First, we experiment with an automatic speech recognition task, in which the uncertainty of observations is imposed by real-world additive noise. Next, we examine a full-scale speaker recognition system, considering the utterance duration as the source of uncertainty in authenticating a speaker. The experimental results show that when employing an appropriate uncertainty estimation algorithm, uncertain LDA outperforms its conventional LDA counterpart.
The biometric and forensic performance of automatic speaker recognition systems degrades under noisy and short probe utterance conditions. Score normalization is an effective tool taking into account the mismatch of reference and probe utterances. In an adaptive symmetric score normalization scheme for state-of-the-art i-vector recognition systems, a set of cohort speakers are employed to calculate the mean and variance of impostor scores when compared to reference and probe i-vectors. In dealing with real-life conditions where the quality of audio recordings in test phase does not match enrolment utterance(s) of speakers, we demonstrate the effectiveness of utilizing a condition matched cohort set for score normalization. The cohort set audio material is shortened and degraded by noise in different reasonable and controlled signal-to-noise ratios according to expected test conditions, yielding in multiple set of cohorts. Further, we propose automatic cohort pre-selection based on modeling each degradation category. For each i-vector, a quality vector is assigned as the posterior probability of degradation classes. The cohort set is then formed by i-vectors representing small KL-divergence of respective quality vectors when compared to reference and probe. Further gains are observed by including this quality vector also into the score calibration.
Speech under face cover constitute a case that is increasingly met by forensic speech experts. Wearing face cover mostly happens when an individual strives to conceal his or her identity. Based on the material of face cover and the level of contact with speech production organs, speech production becomes affected by face mask and a part of speech energy gets absorbed in the mask. There has been little research on how speech acoustics is affected by different face masks and how face covers might affect performance of automatic speaker recognition systems. In the present paper, we have collected speech under face mask with the aim of studying the effects of wearing different masks on state-of-the-art text-independent automatic speaker recognition system. The preliminary speaker recognition rates along with mask identification experiments are presented in this paper.
This paper studies the effect of short utterances and noise on the performance of automatic speaker recognition. We focus on calibration aspects, and propose a calibration strategy that uses quality measures to model the calibration parameters. We carry out the proposed calibration by using simple Quality Measure Functions (QMFs) of duration and measured signal-to-noise-ratio from speech segments. We test the effectiveness of the approach using two databases, the development set of the I4U collaboration for the NIST Speaker Recognition Evaluation (SRE) 2012, and the evaluation test material of NIST SRE 2012 itself. In comparison with conventional linear calibration, results show that the proposed QMF approach successfully improves the system performance in terms of both discrimination and calibration. (C) 2015 Elsevier B.V. All rights reserved.
Non-negative matrix factorisations are used in several branches of signal processing and data analysis for separation and classification. Sparsity constraints are commonly set on the model to promote discovery of a small number of dominant patterns. In group sparse models, atoms considered to belong to a consistent group are permitted to activate together, while activations across groups are suppressed, reducing the number of simultaneously active sources or other structures. Whereas most group sparse models require explicit division of atoms into separate groups without addressing their mutual relations, we propose a constraint that permits dynamic relationships between atoms or groups, based on any defined distance measure. The resulting solutions promote approximation with components considered similar to each other. Evaluation results are shown for speech enhancement and noise robust speech and speaker recognition.
Recognition and classification of speech content in everyday environments is challenging due to the large diversity of realworld noise sources, which may also include competing speech. At signal-to-noise ratios below 0 dB, a majority of features may become corrupted, severely degrading the performance of classifiers built upon clean observations of a target class. As the energy and complexity of competing sources increase, their explicit modelling becomes integral for successful detection and classification of target speech. We have previously demonstrated how non-negative compositional modelling in a spectrogram space is suitable for robust recognition of speech and speakers even at low SNRs. In this work, the sparse coding approach is extended to cover the whole separation and classification chain to recognise the speaker of short utterances in difficult noise environments. A convolutive matrix factorisation and coding system is evaluated on 2nd CHiME Track 1 data. Over 98% average speaker recognition accuracy is achieved for shorter than three second utterances at +9 ... -6 dB SNR, illustrating the system’s performance in challenging conditions.
In this paper, a new AM-FM based filter bank analysis for the estimation of spectro-temporal envelope (STE) of speech signals is proposed. The filter bank is simulated by filtering a frequency translated signal using a single resonator centered around the Nyquist frequency. The proposed design of using a single fixed resonator provides distinct advantages over the traditional methods of filter bank design. First, it provides a simple IIR filter with a smooth frequency response with no ripples. Second, the bandwidth of the resonator can be easily controlled by the multiplicity of poles and their proximity to the unit circle on the z-plane. Third, the resonator fixed at the highest possible center frequency provides the best separation between the AM and FM components of the filtered signal. Speaker recognition experiments on noisy and reverberant speech with short test segments show that the proposed AM-FM based filter bank analysis for STE estimation provides consistent improvement over a recently proposed discrete cosine transform based filter bank approach.
Linear prediction is one of the most established techniques in signal estimation, and it is widely utilized in speech signal processing. It has been long understood that the nerve firing rate of human auditory system can be approximated by power law non-linearity, and this has been the motivation behind using perceptual linear prediction in extracting acoustic features in a variety of speech processing applications. In this paper, we revisit the application of power law non-linearity in speech spectrum estimation by compressing/expanding power spectrum in autocorrelation-based linear prediction. The development of so-called LP- α is motivated by a desire to obtain spectral features that present less mismatch than conventionally used spectrum estimation methods when speech of normal loudness is compared to speech under vocal effort. The effectiveness of the proposed approach is demonstrated in a speaker recognition task conducted under severe vocal effort mismatch comparing shouted versus normal speech mode.
One of the biggest challenges in speaker recognition is incomplete observations in test phase caused by availability of only short duration utterances. The problem with short utterances is that speaker recognition needs to be handled by having information from only limited amount of acoustic classes. By considering limited observations from a test speaker, the resulting i-vector as a representative of short utterance will be uncertain; the shorter the duration, the higher the uncertainty. In recent studies, an uncertainty decoding technique has been employed in probabilistic linear discriminant analysis (PLDA) modeling in order to account for uncertain i-vectors. In this paper, we propose to extend uncertainty handling using simplified PLDA scoring and modified imputation. We experiment with a state-of-the-art speaker recognition system focusing on uncertainty caused by controlled utterance duration. The uncertainties after i-vector extraction are being propagated through pre-processing steps and both uncertainty decoding and modified imputation are considered. Our experimental results indicate improved equal error rate and detection cost attained by using uncertainty-of-observation techniques in dealing with short duration utterances.
The vast majority of speaker recognition cross-entropy evaluations are focused on score domain. By examining the generalized relative distance between genuine and impostor sub-spaces, biometric characteristics become comparable to other authentication approaches. In this paper we demonstrate that the i-vector feature space's biometric information measured by relative entropy is comparable to e.g., knowledge-based mechanisms or face recognition. Examining NIST SRE 2004-2010 corpora, short samples of e.g, 5 seconds duration, comprise already 127 bits in a text-independent scenario. Further, the vast majority of short samples does not fall below 50% of the biometric information of samples having a duration of more than 40 seconds. The generalized i-vector feature space entropy of long samples corresponds to 182.1 bits, and the highest lower entropy bound of a subject was observed at 471.6 bits.