This paper summarizes the BUT-AGNITIO system for NIST Language Recognition Evaluation 2009. The post-evaluation analysis aimed mainly at improving the quality of the data (fixing language label problems and detecting overlapping speakers in the training and development sets) and investigation of different compositions of the development set. The paper further investigates into JFA-based acoustic system and reports results for new SVM-PCA systems going beyond BUT-Agnitio original NIST LRE 2009 submission. All results are presented on evaluation data from NIST LRE 2009 task.
We propose a novel design for acoustic feature-based automatic spoken language recognizers. Our design is inspired by recent advances in text-independent speaker recognition, where intraclass variability is modeled by factor analysis in Gaussian mixture model (GMM) space. We use approximations to GMMlikelihoods which allow variable-length data sequences to be represented as statistics of fixed size. Our experiments on NIST LRE’07 show that variability-compensation of these statistics can reduce error-rates by a factor of three. Finally, we show that further improvements are possible with discriminative logistic regression training. Index Terms: acoustic language recognition, intersession variability compensation, discriminative training
In this paper, we have investigated into JFA used for speaker recognition. First, we performed systematic comparison of full JFA with its simplified variants and confirmed superior performance of the full JFA with both eigenchannels and eigenvoices. We investigated into sensitivity of JFA on the number of eigenvoices both for the full one and simplified variants. We studied the importance of normalization and found that genderdependent zt-norm was crucial. The results are reported on NIST 2006 and 2008 SRE evaluation data. Index Terms: speaker recognition, joint factor analysis.
This article presents several techniques to combine between Support vector machines (SVM) and Joint Factor Analysis (JFA) model for speaker verification. In this combination, the SVMs are applied to different sources of information produced by the JFA. These informations are the Gaussian Mixture Model supervectors and speakers and Common factors. We found that using SVM in JFA factors gave the best results especially when within class covariance normalization method is applied in order to compensate for the channel effect. The new combination results are comparable to other classical JFA scoring techniques.
This paper presents BUT system submitted to NIST 2008 SRE. It includes two subsystems based on Joint Factor Analysis (JFA) GMM/UBM and one based on SVM-GMM. The systems were developed on NIST SRE2006 data, and the results arepresented on NIST SRE 2008 evaluation data. We concentrate on the influence of side information in the calibration. Index Terms: speaker recognition, joint factor analysis, NIST SRE 2008.
BUT submitted three systems to NIST SRE 2008 eval- uations, only to the short2-short3 condition. The pri- mary system is a fusion of three sub-systems: 2 based on MFCC and factor analysis and one making use of SVM scoring of CMLLR and MLLR matrices of an ASR sys- tem. The first contrastive systems differs only in calli- bration and the second contrastive system is a simplified version of the primary one (no ASR use).
This paper presents a procedure of acquiring linguistic data from the broadcast media and its use in language recognition. The goal of this work is to answer the question whether the automatically obtained data from broadcasts can replace or augment to the continuous telephone speech. The main challenges are channel compensation issues and great portion of unspontaneous speech in broadcasts. The experimental results are obtained on NIST LRE 2007 evaluation system, using both NIST provided training data and data, obtained from broadcasts.
This paper describes Brno University of Technology (BUT) system for 2007 NIST Language recognition (LRE) evaluation. The system is a fusion of 4 acoustic and 9 phonotactic subsystems. We have investigated several new topics such as discriminatively trained language models in phonotactic systems, and eigen-channel adaptation in model and feature domain in acoustic systems. We also point out the importance of calibration and fusion. All results are presented on NIST 2007 LRE data.
This paper describes the acoustic language recognition subsystems of Brno University of Technology (BUT) which contributed to the BUT main submission to the NIST LRE 2007. Two main techniques are employed in the subsystems discriminative training in terms of Maximum Mutual Information, and channel compensation in terms of eigenchannel adaptation in both, model and feature domain. The complementarity of the approaches is analyzed.
Gender and age estimation based on Gaussian Mixture Models (GMM) is introduced. Telephone recordings from the Czech SpeechDat-East database are used as training and test data set. Mel-Frequency Cepstral Coefficients (MFCC) are extracted from the speech recordings. To estimate the GMMs' parameters Maximum Likelihood (ML) training is applied. Consequently these estimations are used as the baseline for Maximum Mutual Information (MMI) training. Results achieved when employing both ML and MMI training are presented and discussed.
This paper describes "search in speech" techniques developed in the Speech@FIT research group at FIT BUT in the last couple of years. It concentrates on spoken term detection (STD) and presents our system for NIST STD 2006 evaluations in detail. It also briefly mentions our systems for speaker and language recognition.
vutbr .cz ABSTRACT G ender and age estimation based on Gaussian Mixture Models (GMM) is introduced. Records from the Czech SpeechDat(E) database are used as training and test data set. In order to re- duce the data size, Mel-Frequency Cepstral Coefcients (MFCC) are extracted from the speech recordings. Maximum Likelihood (ML) training is applied to estimate the models' parameters and additionaly discriminative training (DT) is applied to the trained models to provide further improvement of the results.
This paper deals with a speaker detection task. In this work, a different interpretation of the problem is introduced than the one used so far. In the standard approach, each speaker is modeled by their own model and the task is to decide whether the test speech segment was generated by the given model or not. In this work, only two models are used: one represents the target trials and the other represents nontarget trials, where the trial is represented by two speech segments, both from the same speaker, and two from different speakers, respectively. As the input features, fixed-length low-dimensional vectors derived from speaker factors generated by Joint Factor Analysis are used. Gaussian Mixture Models framework is used to model the feature distribution. The achieved results are compared to the state of the art systems.
Martin Karafiat合作论文数DCGM4
Lukas Burget合作论文数Department of Computer Graphics and Multimedia (DCGM)2