Dynamic topic modeling facilitates the identification of topical trends over time in temporal collections of unstructured documents. We introduce a novel unsupervised neural dynamic topic model named as Recurrent Neural Network-Replicated Softmax Model (RNNRSM), where the discovered topics at each time influence the topic discovery in the subsequent time steps. We account for the temporal ordering of documents by explicitly modeling a joint distribution of latent topical dependencies over time, using distributional estimators with temporal recurrent connections. Applying RNN-RSM to 19 years of articles on NLP research, we demonstrate that compared to state-of-the art topic models, RNNRSM shows better generalization, topic interpretation, evolution and trends. We also introduce a metric (named as SPAN) to quantify the capability of dynamic topic model to capture word evolution in topics over time.
The goal of our industrial ticketing system is to retrieve a relevant solution for an input query, by matching with historical tickets stored in knowledge base. A query is comprised of subject and description, while a historical ticket consists of subject, description and solution. To retrieve a relevant solution, we use textual similarity paradigm to learn similarity in the query and historical tickets. The task is challenging due to significant term mismatch in the query and ticket pairs of asymmetric lengths, where subject is a short text but description and solution are multi-sentence texts. We present a novel Replicated Siamese LSTM model to learn similarity in asymmetric text pairs, that gives 22% and 7% gain (Accuracy@10) for retrieval task, respectively over unsupervised and supervised baselines. We also show that the topic and distributed semantic features for short and long texts improved both similarity learning and retrieval.
This paper proposes a novel context-aware joint entity and word-level relation extraction approach through semantic composition of words, introducing a Table Filling Multi-Task Recurrent Neural Network (TF-MTRNN) model that reduces the entity recognition and relation classification tasks to a table-filling problem and models their interdependencies. The proposed neural network architecture is capable of modeling multiple relation instances without knowing the corresponding relation arguments in a sentence. The experimental results show that a simple approach of piggybacking candidate entities to model the label dependencies from relations to entities improves performance. We present state-of-the-art results with improvements of 2.0% and 2.7% for entity recognition and relation classification, respectively on CoNLL04 dataset.
This paper presents a comparative study of four different approaches to automatic age and gender classification using seven classes on a telephony speech task and also compares the results with Human performance on the same data. The automatic approaches compared are based on (1) a parallel phone recognizer, derived from an automatic language identification system; (2) a system using dynamic Bayesian networks to combine several prosodic features; (3) a system based solely on linear prediction analysis; and (4) Gaussian mixture models based on MFCCs for separate recognition of age and gender. On average, the parallel phone recognizer performs as well as Human listeners do, while loosing performance on short utterances. The system based on prosodic features however shows very little dependence on the length of the utterance.
Recently an unsupervised learning scheme for Hidden Markov Models (HMMs) used in acoustical Language Identification (LID) based on Parallel Phoneme Recognizers (PPR) was proposed. This avoids the high costs for orthographically transcribed speech data and phonetic lexica but was found to intro-duce a considerable increase of classification errors. Also very recently discriminative Minimum Language Identification Error (MLIDE) optimization of HMMs for PPR based LID was introduced that again only requires language tagged speech data and an initial HMM. The described work shows how to combine both approaches to an unsupervised and discriminative learning scheme. Experimental results on large telephone speech databases show that using MLIDE the relative increase in error rate introduced by unsupervised learning can be reduced from 61% to 26%. The absolute difference in LID error rate due to the supervised learning step is reduced from 4.1% to 0.8%.
In-car automatic speech recognition (ASR) is usually evaluated behaviour for different levels of noise. Yet this is interesting for car manufacturers in order to predict system performances for different speeds and different car models and thus allow to design speech based applications in a better way. It therefore makes sense to split the single WER into SNR dependent WERs, where SNR stands for the signal to noise ratio, which is an appropriate measure for the noise level. In this paper a SNR measure based on the concept of the Articulation Index is developed, which allows the direct comparison with human recognition performance.
This paper describes an extension to an automatic speech recognition system that improves the robustness concerning varying environments. A dedicated control unit tries to derive an optimal set of parameters for a Wiener Filter based noise reduction unit aiming at maximum recognition performances in different environments. The input measure for the control unit is derived from the speech signal. Apart from the SNR level, several other measures are investigated. The controlled parameters are closely related to the strength of the noise reduction. Several non-linear methods such as the Tabulated References and Neural Networks serve as the core of the control unit. Experiments on realistic handsfree as well as non-handsfree speech data show that the word error rate can be reduced by as much as 31% through the proposed methods. An already optimized static configuration of the applied noise reduction hereby serves as the baseline level.
In this paper the implementation of a word-stem based tree search for large vocabulary speaker independent isolated word recognition for embedded systems is presented. Two fast search algorithms combine the effectiveness of the tree structure for large vocabularies and the fast Viterbi search within the regular structures of word-stems. The algorithms are proved to be very effective for workstation and embedded platform realizations. In order to decrease the processing power the word-stem based tree search with frame dropping approach is used. The recognition speed was increased by a factor of 5 without frame dropping and by a factor of 10 with frame dropping in comparison to linear Viterbi search for isolated word recognition task with a vocabulary of 20102 words. Thus, the large vocabulary isolated word recognition becomes possible for embedded systems.
We propose a general methodology to design a robust voice activity detector that suits the needs of the speech enhancement system it is dedicated to. More than imposing rules, we initiate ideas on how to perform the analysis of the requirements for the Voice Activity Detection (VAD) and how to choose a reference, and evaluate the performances of the explored solutions in order to choose the one that best fits. As an example, the methodology is then applied to evaluate five VADs based on features described in the literature in the scope of two typical speech enhancement applications.
In order to make hidden Markov model (HMM) speech recognition suitable for mobile phone applications, Siemens developed a recognizer, Very Smart Recognizer (VSR), for deployment in future mobile phone generations. Typical applications will be name dialling, command and control operations suited for different environments, for example in cars. The paper describes research and development issues of a speech recognizer in mobile devices focusing on noise robustness, memory efficiency and integer implementation. The VSR is shown to reach a word error rate as low as 4.1% on continuous digits recorded in a car environment. Furthermore by means of discriminative training and HMM-parameter coding, the memory requirements of the VSR HMMs are smaller than 64 kBytes.
Abstract Following the objective of the Eurospeech special event,'Noise Robust Recognition', the recognition results of a noiserobust front-end, developed by Siemens, on the Aurora 2 da-tabase [1] are presented in this paper. The front-end wastested with and without a frame dropping algorithm. It isshown that the front-end improves the recognition results inhigh mismatch between training and testing by 43.90% overthe reference front-end and works particularily well in con-ditions with high noise. Furthermore it is shown that theframe dropping mainly increases the performance of the front-end. 1. Introduction The front-end employed is a cepstral analysis scheme withspectral attenuation and spectral subtraction, a channel com-pensation, calculation of the time derivatives, an LDA and aframe dropping algorithm (Figure 1).Details of the front-end can be seen from Table 1. The spe-cific algorithms are shortly described in the following.Parameter settings of the front-endFrame length: 32 msFrame shift: 15 msNumber of samples per frame, NF: 256Preemphasis factor: 0.95Lenght of FFT, FFTL: 256Number of Mel filter banks: 15Number of cepstral coefficients: 12Number of first derivatives: 13Number of second derivatives: 13Size of the feature vector: 24Table 1: Parameter settings of the front-end
This paper shows how the noise robustness of a MFCC feature extraction front-end can be improved by integrating four noise robustness algorithms:a spectral attenuation, a noise level normalisation, a cepstral mean normalization and a frame dropping algorithm. The algorithms were tested separately and in varying combinations on three real world car data sets with different amounts of mismatch between the training and the testing conditions. It was shown that although the algorithms partly have similar effects none of them is completely redundant. Every algorithm can contribute to a further improvement of the recognition results so the best results can be achieved by a combination of all four of them. A relative reduction of the word error rate of up to 57% is achieved.