“Infinitely clipped” (“IC”) speech is surprisingly intelligible [Licklider and Pollack, J. Acoust. Soc. Am. 20, 42–51 (1948)]. Because infinite clipping so severely distorts the speech spectrum, it is sometimes speculated that a time-domain similarity between natural and IC speech—precise identity of zero-crossing locations—might account for the latter's intelligibility. This hypothesis has practical implications (should speech recognition algorithms look beyond power spectra?) and theoretical ones [cf. Scott, J. Acoust. Soc. Am. 60, 1354–1365 (1976)]. However, the hypothesis would be refuted if it turned out that the small amount of residual similarity between the power spectra of the original and IC signals could alone account for the observed intelligibility. Three types of stimuli—natural vowels, /p,t,k/ in natural-sentence context, and synthetic vowels—were infinitely clipped (“IC stimuli”), then subjected to transformations which preserved the power spectrum but grossly distorted the phase spectrum, thereby destroying the original zero-crossing information. These transformed IC stimuli were no less intelligible than the unaltered IC stimuli, strongly suggesting that power spectrum rather than zero-crossing cues are responsible for IC-speech intelligibility. Further, constant-parameter, synthetic-vowel inteiligibility decreased dramatically upon clipping, natural-vowel intelligibility relatively little, demonstrating the usefulness of spectral change and/or prosodic cues in the absence of accurate formant information.
A speaker-dependent speech recognition system is described for recognizing isolated word utterances using reference templates created by concatenating demisyllable (half-syllable) prototypes. Each word in a vocabulary is specified by one or more entries in a user-supplied lexicon containing a sequence of demisyllables drawn from a corpus of some 1000 units. Experiments were carried out with two talkers using a 1109-word "Basic English" vocabulary to assess the overall effectiveness of demisyllable representations for words. Also, the effects on performance of some simple modifications in demisyllable specifications and adjustments of demisyllable durations were investigated. The recognition error rates obtained for this vocabulary using demisyllable prototypes were 18-33 percent compared with 6-15 percent using whole word prototypes. Although the performance is substantially poorer using demisyllable representations in place of whole words, the approach of using a fixed inventory of smaller-than-word recognition units capable of representing any spoken word in a simple concatenative scheme is clearly an attractive alternative to whole-word prototypes for large-size vocabularies. The approach also has the potential of being effective in representing and recognizing continuous spoken utterances.
It has previously been demonstrated that reliable, speaker-trained, isolated word recognition on a 1109-word Basic English vocabulary can be performed using word templates formed by concatenation of elements from a corpus of demisyllables. Since a dynamic time warping (DTW) algorithm is used to align test and reference patterns, small to moderate differences in duration between test and reference words present no major problem in performing the time alignment. However, improved results (i.e., smaller word distances) are obtained from the DTW algorithm if the syllables of the test and reference words are properly aligned prior to dynamic time warping. In our earlier experiments, each concatenated reference word was linearly prenormalized to the duration of that word in the test set, but this procedure is clearly not applicable for continuous speech recognition. We have now developed a linguistically based set of duration rules which we apply to the demisyllables during the word-creation process (i.e., before DTW), which predict syllable duration as a function of syllable stress level and the position of the syllable within the word. Using the automatic duration rules, we have achieved recognition accuracies comparable to those based on known word durations.
While the two lowest formant frequencies, F1 and F2, are the most important cues in vowel identification, the correspondence between vowels and small regions in F1−F2 space is far from one to one when data from many speakers is considered. This fact suggests that our vowel-identification mechanism, quite reliable in ordinary speech situations, either (a) makes use of additional disambiguating cues, such as the frequencies of higher formants, vowel duration, degree of diphthongization, etc.; or (b) normalizes for individual-speaker differences by deriving information on vocal-tract geometry through exposure to consonantal formant transitions and known vowels; or both. The high vowel-identification error rates observed when isolated vowels are speaker randomized (43% in a recently reported experiment) argue that (a) is relatively unimportant vís-a-vís (b), since in such experiments all the types of information referred to in (a) are retained, yet vowel identification is not reliable. In this paper I argue that there exists a systematic bias favoring hypothesis (b) over hypothesis (a) in past vowel-identification tests and describe an experiment in which near-perfect identification scores were achieved despite speaker randomization and the absence of consonantal transitions.