We performed automated feature selection for multi-stream (i.e., ensemble) automatic speech recognition, using a hillclimbing (HC) algorithm that changes one feature at a time if the change improves a performance score. For both clean and noisy data sets (using the OGI Numbers corpus), HC usually improved performance on held out data compared to the initial system it started with, even for noise types that were not seen during the HC process. Overall, we found that using Opitz’s scoring formula, which blends single-classifier word recognition accuracy and ensemble diversity, worked better than ensemble accuracy as a performance score for guiding HC in cases of extreme mismatch between the SNR of training and test sets. Our noisy version of the Numbers corpus, our multi-layerperceptron-based Numbers ASR system, and our HC scripts are available online.
We describe the ICSI-SRI-UW team’s entry in the Spring 2004 NIST Meeting Recognition Evaluation. The system was derived from SRI’s 5xRT Conversational Telephone Speech (CTS) recognizer by adapting CTS acoustic and language models to the Meeting domain, adding noise reduction and delay-sum array processing for far-field recognition, and postprocessing for cross-talk suppression. A modified MAP adaptation procedure was developed to make best use of discriminatively trained (MMIE) prior models. These meeting-specific changes yielded an overall 9% and 22% relative improvement as compared to the original CTS system, and 16% and 29% relative improvement as compared to our 2002 Meeting Evaluation system, for the individual-headset and multiple-distant microphones conditions, respectively.
Multi-stream automatic speech recognition (ASR) systems consisting of an ensemble of classifiers working together, each with its own feature vector, are popular in the research literature. Published work on feature selection for such systems has dealt with indivisible blocks of features. I break from this tradition by investigating feature selection at the level of individual features. I use the OGI ISOLET and Numbers speech corpora, including noisy versions I created using a variety of noises and signal-to-noise ratios. I have made these noisy versions available for use by other researchers, along with my ASR and feature selection scripts.I start with the random subspace method of ensemble feature selection, in which each feature vector is simply chosen randomly from the feature pool. Using ISOLET, I obtain performance improvements over baseline in almost every case where there is a statistically significant performance difference, but there are many cases with no such difference. I then try hill-climbing, a wrapper approach that changes a single feature at a time when the change improves a performance score. With ISOLET, hill-climbing gives performance improvements in most cases for noisy data, but no improvement for clean data. I then move to Numbers, for which much more data is available to guide hill-climbing. When using either the clean or noisy Numbers data, hill-climbing gives performance improvements over multi-stream baselines in almost all cases, although it does not improve over the best single-stream baseline. For noisy data, these performance improvements are present even for noise types that were not seen during the hill-climbing process. In mismatched condition tests involving mismatch between clean and noisy data, hill-climbing outperforms all baselines when Opitz's scoring formula is used. I find that this scoring formula, which blends single-classifier accuracy and ensemble diversity, works better for me than ensemble accuracy as a performance score for guiding hill-climbing.
A critical step in encoding sound for neuronal processing occurs when the analog pressure wave is coded into discrete nerve-action potentials. Recent pool models of the inner hair cell synapse do not reproduce the dead time period after an intense stimulus, so we used visual inspection and automatic speech recognition (ASR) to investigate an offset adaptation (OA) model proposed by Zhang et al. [1]. OA improved phase locking in the auditory nerve (AN) and raised ASR accuracy for features derived from AN fibers (ANFs). We also found that OA is crucial for auditory processing by onset neurons (ONs) in the next neuronal stage, the auditory brainstem. Multi-layer perceptrons (MLPs) performed much better than standard Gaussian mixture models (GMMs) for both our ANF-based and ON-based auditory features. Similar results were previously obtained with MSG (Modulation-filtered SpectroGram) auditory features[2]. Thus we believe researchers working with novel features should consider trying MLPs.
Our notion of how speech is processed is still very much dominated by von Helmholtz’s theory of hearing. He deduced that the human inner ear decomposes the spectrum of sound signals. However, physiological recordings of auditory nerve fibers (ANF) showed that the rate-place code, which is thought to transmit spectral information to the brain, is at least complemented by a temporal code. In our paper we challenge the rate-place code using a complex but realistic scenario: speech in noise. We used a detailed model of human auditory processing that closely replicates key aspects of auditory nerve spike trains. We performed quantitative evaluations of coding strategies using standard automatic speech recognition (ASR) tools. Our test data was spoken letters of the whole English alphabet from a variety of speakers, with and without background noise. We evaluated a purely rate-place-based encoding strategy, a temporal strategy based on interspike intervals, and a combination thereof. The results suggest that as few as 4% of the total number of ANFs would be sufficient to code speech information in a rate-place fashion. Rate-place coding performed its best for speech in clean conditions at normal sound level, but broke down at higher-than-normal levels, and failed dramatically in noise at high levels. Low-spontaneous rate fibers improved the rate-place code, mainly for vowels and at higher-than-normal levels. At high speech levels, and in particular in the presence of background noise, combining rate-place coding with the temporal coding strategy greatly improved recognition accuracy. We therefore conclude that the human auditory system does not rely on a rate-place code alone but requires the abundance of fibers for precise temporal coding.
A major difference between the human auditory system and automatic speech recognition (ASR) lies in their representation of sound signals: whereas ASR uses a smoothed low-dimensional temporal and spectral representation of sound signals, our hearing system relies on extremely high-dimensional but temporally sparse spike trains. A strength of the latter representation is in the inherent coding of time, which is exploited by neuronal networks along the auditory pathway. We demonstrate ASR results using features purely derived from simulated spike trains of auditory nerve fibers (ANF) and a layer of octopus neurons. Octopus neurons located in the cochlear nucleus are known for their distinct temporal processing: they not only reject steady-state excitation and fire on signal onsets but also enhance the amplitude modulations of voiced speech. With multi-condition training we do not reach the performance of conventional mel-frequency cepstral coefficients (MFCC) features. With clean training however, our spike-based features performed similarly to MFCCs. Further, recognition scores in noise were improved when features derived from ANFs, which mainly represent spectral characteristics of speech signals, were combined with features derived from spike trains of octopus neurons. This result is promising given the relatively small number of neurons we used and the limitations in how the auditory model was interfaced to the ASR back end.
To reduce inter-speaker variability, vocal tract length normalization (VTLN) is commonly used to transform acoustic features for automatic speech recognition (ASR). The warp factors used in this process are usually derived by maximum likelihood (ML) estimation, involving an exhaustive search over possible values. We describe an alternative approach: exploit the correlation between a speaker’s average pitch and vocal tract length, and model the probability distribution of warp factors conditioned on pitch observations. This can be used directly for warp factor estimation, or as a smoothing prior in combination with ML estimates. Pitch-based warp factor estimation for VTLN is effective and requires relatively little memory and computation. Such an approach is well-suited for environments with constrained resources, or where pitch is already being computed for other purposes.
We describe the ICSI-SRI-UW team’s entry in the Spring 2004 NIST Meeting Recognition Evaluation. The system was derived from SRI’s 5xRT Conversational Telephone Speech (CTS) recognizer by adapting CTS acoustic and language models to the Meeting domain, adding noise reduction and delay-sum array processing for far-field recognition, and postprocessing for cross-talk suppression. A modified MAP adaptation procedure was developed to make best use of discriminatively trained (MMIE) prior models. These meeting-specific changes yielded an overall 9% and 22% relative improvement as compared to the original CTS system, and 16% and 29% relative improvement as compared to our 2002 Meeting Evaluation system, for the individual-headset and multiple-distant microphones conditions, respectively.
In this paper we develop a physiologically motivated model of peripheral auditory processing and evaluate how the different processing steps influence automatic speech recognition in noise. The model features large dynamic compression (>60 dB) and a realistic sensory cell model. The compression range was well matched to the limited dynamic range of the sensory cells and the model yielded surprisingly high recognition scores. We also developed a computationally efficient simplified model of auditory processing and found that a model of adaptation could improve recognition accuracy. Adaptation is a basic principle of neuronal processing, which accentuates signal onsets. Applying this adaptation model to melfrequency cepstral coefficient (MFCC) feature extraction enhanced recognition accuracy in noise (AURORA 2 task, averaged recognition scores) from 56.4% to 75.6% (clean training condition), a relative improvement of 41% in word error rate. Adaptation outperformed RASTA processing by more than 10%, which corresponds to a relative improvement of 31%.
For a connected digits speech recognition task, we have compared the performance of two inexpensive electret microphones with that of a single high quality PZM microphone. Recognition error rates were measured both with and without compensation techniques, where both single-channel and two-channel approaches were used. In all cases the task was recognition at a significant distance (2–6 feet) from the talker’s mouth. The results suggest that the wide variability in characteristics among inexpensive electret microphones can be compensated for without explicit quality control, and that this is particularly effective when both single-channel and two-channel techniques are used. In particular, the resulting performance for the inexpensive microphones used together is essentially equivalent to the expensive microphone, and better than for either inexpensive microphone used alone.
Hands-free use of speech recognition (i.e., not requiring a microphone worn or held by the user) introduces the technical challenges of room acoustics and background noise. In this paper, I will start by describing a possible hands-free application from the SmartKom dialogue system project. I will then describe the acoustical issues involved and the possible technical approaches to dealing with them. I will then discuss one such approach that we have been working with, long-term log spectral subtraction, and give experimental results examining its usefulness for interactive applications. SmartKom The SmartKom project [Smartkom; Wahlster] is intended to create multimodal dialogue systems that combine the use of speech and gesture (both by the system and the user). Figure 1 illustrates the SmartKom Home system concept: a portal to information services such as television for home use. The ‘face’ of the system is the blue character on the left, named Smartakus, who communicates with the user with both synthesized speech and animated gestures. The use of an animated character is intended to make the system easier for novice or naive users. User queries are spoken (for example, “What movies are on TV tonight?”) and the user can use gestures while they are speaking to point to parts of the display. Figure 1: SmartKom Home. Room Acoustics and Hands-Free Speech Recognition The system shown in Figure 1 is envisioned to be usable hands-free without requiring the user to carry or wear a microphone. Making hands-free recognition more reliable is a major research problem. This is because, compared to recordings made near the user’s mouth, recordings made from more distant microphones have higher levels of background noise relative to the speech level and are more affected by reverberation due to echos off reflective surfaces such as walls. Even a modest degree of reverberation, which would present little or no difficulty to a human listener, can substantially decrease the performance of automatic speech recognition systems. The greater distance that sound travels to the microphone also results in higher frequencies in speech being weakened relative to lower frequencies. I will call the combined effect of reverberation and other effects such as this the ‘room response’ from the user to the microphone. It is common to model the room response as a filter applied to the original speech which affects the spectrum of the speech in a consistent way. There are various approaches which could be tried to improve recognition performance in this situation: • single microphone signal processing to reduce the effects of noise, reverberation, and spectral distortion • microphone array signal processing to the same purpose • using noisy or reverberant training data to create the speech models [Stahl] • using adaptation methods to adjust the speech models • using methods of representing the signal that are less sensitive to the distortions caused by reverberation, such as the modulation spectrogram [Kingsbury] Several of these are discussed in [Omologo], which also has a useful discussion of room acoustics. The remainder of this paper will focus on the first approach, which is not to say that I feel it is the single best approach. I think it is likely that the best performance will come from combining approaches. Long-term Log Spectral Subtraction Carlos Avendano [Avendano] developed a method of room response compensation which estimates the room response using the assumption that it does not change and so can be estimated by measuring the unchanging part of the speech spectrum. The estimated room response is then removed from the speech spectrum by subtraction (in the logarithmic magnitude domain). Since speech itself has average characteristics, the room response estimate also contains information about the speech signal. Because of this and other artifacts of this method, we train the recognizer on processed speech. This method is similar to the common technique of cepstral mean subtraction, which is used to compensate for microphone or telephone channel effects. Reverberation effects have a longer extent in time, so to deal with reverberation Avendano’s method calculates the speech spectrum using longer periods of speech than is normally done for cepstral mean subtraction. I will call his method ‘long-term log spectral subtraction’ to emphasize this. The method assumes that the room response is approximately constant. In fact people are rarely perfectly still, and the room response from them to a microphone will change as they move, but we have found this method to be useful in realistic data collected with seated speakers. It is less likely to work if the speaker is walking around a room. Experimental Results In experiments at ICSI last year, we found long-term log spectral subtraction to be useful for offline (non-interactive) recognition [Gelbart01]. In that work we collected several utterances worth of data to estimate the room response before perfoming the subtraction and starting the recognizer. I will now present some of those results and then discuss the use of this method in an interactive system. For our experiments we used the Aurora reference system described in [Hirsch], which is based on the HTK speech recognizer configured to recognize digits. For training data we used four hours of data from the TIDIGITS connected digits corpus, which was collected by close-talking microphone in a quiet environment. To test it, we used connected digits strings (a total of 7704 words) which were read by native English speakers seated around a conference table in a room we are using for recording natural meetings. Simultaneously recordings were made with closetalking microphones and with a table-mounted microphone that was 3-6 feet from each speaker. Table 1 shows the original system’ s performance, measured by word error rate (WER). (Word error rate is the fraction of words omitted, changed, or falsely inserted by the recognizer out of the total number of spoken words.) The first column (“Near”) refers to the close-talking microphones and the second column (“Far”) refers to the table-mounted microphone. Near microphone Far microphone 4.1% 26.3% Table 1: Baseline results. Table 2 on the next page shows the WER results when the long-term log spectral subtraction method was applied to the training and test data. The room response was estimated using 7.168 seconds of data. (This was done separately for each speaker, and if the same speaker re-appeared during different recording sessions we did the estimate separately each time.) Performance improved dramatically for the far test data. It also improved significantly for the near test data. Reverberation may not be an issue with the near microphone, but since the original system did not include any kind of channel compensation (such as cepstral mean subtraction) to compensate for differences in microphone type, etc., between the training and test data, it’ s likely that the method is helping in this regard. Near microphone Far microphone 3.0% 8.2% Table 2: Results with long-term log spectral subtraction. In an interactive application, collecting several utterances worth of data to estimate the room response before beginning processing is not feasible—utterances need to be recognized as soon as the user has finished speaking them. Therefore, to investigate the use of long-term log spectral subtraction in interactive applications I modified the algorithm to process the current utterance using a room response estimate calculated from whatever utterances the user has spoken thus far, so that the current utterance be can passed immediately to the recognizer. The question now is whether the method will still perform well for the first few utterances, where only a few seconds of data are available to estimate the room response. (The average utterance length in the test set was 1.3 seconds.) Table 3 shows the new far microphone WER broken down by whether utterances were the first, second, third, etc. utterance (indicated in the first column) spoken by that speaker in that recording session. The second column gives the total number of words and the third column gives the WER. Utterance number in session Total words Far WER 1-2 679 12.8% 3-4 692 8.4% 5-6 702 8.4% 7-8 670 9.3% 9-10 736 7.5% 11-12 692 7.4% 13-14 661 7.6% 15-16 679 9.3% 17-18 695 8.1% 19-20 697 6.3% 21+ 801 4.7% Table 3: Far microphone results with past-and-present-only long-term log spectral subtraction. The results in Table 3 show that after the first two utterances the mean subtraction is performing well. This is encouraging for the use of it in an interactive system. Incidentally, the especially good result in the last row showed up in the experiment in Table 2 as well (only 6 of the 17 speakers supplied more than 20 utterances in any recording session; perhaps they were easier to recognize than average).
Far-field microphone speech signals cause high error rates for automatic speech recognition systems, due to room reverberation and lower signal-to-noise ratios. We have observed large increases in speech recognition word error rates when using a far-field (3-6 feet) microphone in a conference room, in comparison with recordings from close-talking microphones. In an earlier paper, we showed improvements in far-field speech recognition performance using a longterm log spectral subtraction method to combat reverberation. This method is based on a principle similar to cepstral mean subtraction but uses a much longer analysis window (e.g., 1 s) in order to deal with reverberation. Here we show that a combination of short-term noise filtering and longterm log spectral subtraction can further reduce recognition word error rates.
A novel type of feature extraction for automatic speech recognition is investigated. Two-dimensional Gabor functions, with varying extents and tuned to different rates and directions of spectro-temporal modulation, are applied as filters to a spectro-temporal representation provided by mel spectra. The use of these functions is motivated by findings in neurophysiology and psychoacoustics. Data-driven parameter selection was used to obtain Gabor feature sets, the performance of which is evaluated on the Aurora 2 and 3 datasets both on their own and in combination with the Qualcomm-OGI-ICSI Aurora proposal. The Gabor features consistently provide performance improvements.
In collaboration with colleagues at UW, OGI, IBM, and SRI, we are developing technology to process spoken language from informal meetings. The work includes a substantial data collection and transcription effort, and has required a nontrivial degree of infrastructure development. We are undertaking this because the new task area provides a significant challenge to current HLT capabilities, while offering the promise of a wide range of potential applications. In this paper, we give our vision of the task, the challenges it represents, and the current state of our development, with particular attention to automatic transcription.
Nelson Morgan合作论文数Department of Electrical Engineering and Computer Sciences at UC Berkeley6