Condition-dependent training strategy divides a training database into a number of clusters, each corresponding to a noise condition and subsequently trains a hidden Markov model (HMM) set for each cluster. This paper investigates and compares a number of condition-dependent training strategies in order to achieve a better understanding of the effects on automatic speech recogntion (ASR) performance as caused by a splitting of the training databases. Also, the relationship between mismatches in signal-to-noise ratio (SNR) is analyzed. The results show that a splitting of the training material in terms of both noise type and SNR value is advantageous compared to previously used methods, and that training of only a limited number of HMM sets is sufficient for each noise type for robustly handling of SNR mismatches. This leads to the introduction of an SNR and noise classification-based training strategy (SNT-SNC). Better ASR performance is obtained on test material containing data from known noise types as compared to either multicondition training or noise-type dependent training strategies. The computational complexity of the SNT-SNC framework is kept low by choosing only one HMM set for recognition. The HMM set is chosen on the basis of results from noise classification and SNR value estimations. However, compared to other strategies, the SNT-SNC framework shows lower performance for unknown noise types. This problem is partly overcome by introducing a number of model and feature domain techniques. Experiments using both artificially corrupted and real-world noisy speech databases are conducted and demonstrate the effectiveness of these methods.
In this paper, the temporal correlation of speech is exploited in front-end feature extraction, client-based error recovery, and server-based error concealment (EC) for distributed speech recognition. First, the paper investigates a half frame rate (HFR) front-end that uses double frame shifting at the client side. At the server side, each HFR feature vector is duplicated to construct a full frame rate (FFR) feature sequence. This HFR front-end gives comparable performance to the FFR front-end but contains only half the FFR features. Second, different arrangements of the other half of the FFR features creates a set of error recovery techniques encompassing multiple description coding and interleaving schemes where interleaving has the advantage of not introducing a delay when there are no transmission errors. Third, a subvector-based EC technique is presented where error detection and concealment is conducted at the subvector level as opposed to conventional techniques where an entire vector is replaced even though only a single bit error occurs. The subvector EC is further combined with weighted Viterbi decoding. Encouraging recognition results are observed for the proposed techniques. Lastly, to understand the effects of applying various EC techniques, this paper introduces three approaches consisting of speech feature, dynamic programming distance, and hidden Markov model state duration comparison
With the aim of improving noise robustness of speech recognition an approach that exploits the variance information in spectral sub-bands is presented. The variance based features are used in combination with the normally used Mel-frequency cepstral coefficients (MFCC), and experimental results show that the combined features outperform MFCC alone, perceptual linear prediction features and entropy based features.
This paper presents a comparative study of different error concealment (EC) techniques in the context of distributed speech recognition (DSR) that exploits repetition, interpolation or subvector concealment to counteract transmission errors. A number of experiments are conducted and the results demonstrate that repetition is as good as, or even better than, linear interpolation whereas the subvector concealment shows the best performance in terms of recognition accuracy. Further experiments and analyses are conducted with the purpose of uncovering the reasons for the different characteristics of the EC techniques: speech features are inspected, time normalised distances as well as hidden Markov model (HMM) state durations are compared for different EC techniques.
Conventional error concealment (EC) algorithms for distributed speech recognition (DSR) share a common characteristic namely the fact of conducting EC at the vector (or frame) level. This strategy, however, fails to effectively exploit the error-free fraction left within erroneous vectors where a substantial number of subvectors often are error-free. This paper proposes a novel EC approach for DSR encoded by split vector quantization (SVQ) where the detected erroneous vectors are submitted to a further analysis at the subvector level. Specifically, a data consistency test is applied to each erroneous vector to identify inconsistent subvectors. Only inconsistent subvectors are replaced by their nearest neighbouring consistent subvectors whereas consistent subvectors are kept untouched. Experimental results demonstrate that the proposed algorithm in terms of recognition accuracy is superior to conventional EC methods having almost the same complexity and resource requirement.
This paper presents the basic rationale behind the development and testing of a multimodal communication aid especially designed for people suffering from global aphasia, and thus having severe expressive difficulties. The principle of the aid is to trigger patient associations by presenting various multimodal representations of communicative expressions. The aid can in this way be seen as a conceptual continuation of previous research within the field of communication aids based on uni-modal (pictorial) representations of communicative expressions. As patients suffering from global aphasia seldom have identical symptoms, the focus of this paper is placed on the development of a highly dedicated communication aid adaptive to the individual patients’ needs. The paper investigates whether or not such a highly dedicated communication aid based on multimodal representations of communicative expressions can be used to support patients with global aphasia in communicating by means of short sentences with their surroundings. Only a limited evaluation is carried out, and as such no statistically significant results are obtained. The tests however indicate that the aid is capable of supporting a global aphasia patient in participating in conversations based on short sentences which otherwise would be impossible without the use of a communication aid.
A technique for mitigating the effect of packet loss in the context of distributed speech recognition is presented. The proposed packet loss concealment (PLC) technique substitutes packet loss partly by a repetition of neighbouring packets and partly by a splicing in which a number of packets are dropped. Experimental results demonstrate that the proposed PLC technique outperforms existing techniques.
The focus of this paper is to formulate an approach to merging phonemes across languages and to evaluate the resulting cross-language merged speech units on the basis of the traditional acoustic-phonetic descriptions of the phonemes. The methodology is based on the belief that some phonemes across a set of languages may be similar enough to be equated, contrasting traditional phonology which treats phonemes from one language independent from phonemes from another language. The identification of cross-language speech units is performed by an iterative data-driven procedure, which merges acoustically similar phonemes from within one language as well as across languages. The paper interprets a number of merged speech units on the basis of articulatory descriptions.
Two systems (Statistical Trajectory Models (STM) and continuous density HMMs) utilizing three preprocessing methodologies (MFCC, RASTA and FBDYN) were evaluated on two databases, namely CTIMIT and the corresponding down-sampled TIMIT. Within the bounds of the experimental setup the comparative performance analysis showed that the STM significantly outperforms the HMM system on the CTIMIT database. Specifically, the performance of the STM system was found to be at least 10% better as compared to the one obtained by HMM when the RASTA preprocessing was used. The performance of both systems with FBDYN parametrization was found to be inferior to those using MFCC and RASTA. On the other hand, in low-noise conditions on the TIMIT database FBDYN yielded an improved performance for the HMM system, whereas STM achieved the best results with the MFCC parametrization.
decoding, the second transforms the parameters from This work is concerned with the subject of language-the decoding module and classifies the language. identification (LID). Two central issues are addressed. The common acoustic signal preprocessor calculates The first is to analyse the trade-off between detailed 12 RASTA filtered MFCC’s, their first derivatives and acoustic modelling and robust estimation of acoustic the delta-log-energy. The phone and language decoding and language models. The second to find the optimal module consists of three parallel branches. In each of combination of acoustic and language scores for language-these the phone recogniser matches the acoustic identification. parameters to the acoustic models used by that recogniser. Experiments are carried out using the three languages The output from each recogniser is further matched American-English, German and Spanish from the OGI-TS against three language models. database. It is shown that on the average the acoustic The combined output X from all language models modelling is able to recognise 46.3% of the phones correctly and from all recognisers are used as input to the across the three languages. Insertion and deletion rate ‘information combination and the language-classification’ is 35.7% and 6.6%, respectively. Language-identification module (ICLC). This module enforces a transformation performance is 82.6% with the full set of acoustic models. onto the parameters X and estimates the most probable The performance is increased to 83.7% after having language given the acoustic input. conducted 80 iterations of a hierarchical clustering in which phones are merged across the languages.
A database of recordings of D anish E motional S peech, DES, has been recorded and analysed. DES has been collected in order to evaluate how well the emotional state in emotional speech is identified by humans. The results sets a standard for identifying Danish emotional speech. DES contains recordings from four actors, two of each gender. Actors were used for the recordings as they were believed to be able to realistically convey a number of emotions, namely: neutral, surprise, happiness, sadness and anger. The recordings from each actor consist of two isolated words, nine sentences and two passages. The complete database comprises approximately 30 minutes of speech. A listening test with 20 listeners was conducted. The emotions were on the average identified correctly in 67,3% of the cases, with a [66,0 - 68,6] 95% confidence interval. An analysis reveals that most confusion occurred between surprise and happiness and between neutral and sadness.
The paper reports on results from ongoing research on language identification (LID) performed on the three languages: American-English, German and Spanish. The speech material used is from the Oregon Graduate Institute Spontaneous Telephone Speech Corpus, OGI-TS. The baseline LID system consists of three parallel phoneme recognisers, each of which are followed by three language modelling modules each characterising the bigram probabilities. The phoneme models used are derived on the basis of the combined speech corpus comprising the three languages. The phonemes are handled differently in analysis performed in two experiments. In the first experiment they are trained and tested language specifically. In the second, they are separated into a number of groups, one of which contains those language independent speech units which are similar enough to be equated across the training languages, the remaining containing the non combinable language dependent phonemes for each of the languages. A data driven technique has been devised to separate the speech sounds contained within the training corpus into these groups. In order to prepare for an optimal separation between the input classes, a linear discriminant analysis is performed on the training speech material. Results from a number of experiments show that average language identification scores of close to 90% can be retained by the LID system presented here, even for a high number of language independent speech units.
Recently, we described a two step self learning approach for grapheme to phoneme (G2P) conversion (O. Anderson and P. Dalsgaard, 1995). In the first step, grapheme and phoneme strings in the training data are aligned via an iterative Viterbi procedure that may insert graphemic and phonemic nulls where required. In the second step, a Trie structure, encoding pronunciation rules is generated. We describe the alignment module, and give alignment accuracies on the NETtalk database. We also compare transcription accuracies for two approaches to the second step on three databases: the NETtalk database, the CMU dictionary and the French part of the ONOMASTICA lexicon. The two transcription approaches applied in this research are a Trie approach and an approach based on binary decision trees grown by means of the Gelfand-Ravishankar-Delp algorithm (F. Breiman et al., 1984; S. Gelfand et al., 1991; R. Kuhn et al., 1995). We discuss the choice of questions for these decision trees-it may be possible to formulate questions about groups of characters (e.g., “is the next letter a vowel?”) that yield better trees than those that only use questions about individual characters (e.g., “is the next letter an `A' ?”). Finally, we discuss the implications of our work for G2P conversion