Two groups of 21 adult subjects with normal hearing viewed the video recordings of the Bamford-Kowal-Bench standard sentence lists issued by the EPI Group in 1986. Each subject viewed all of the 21 lists and attempted to write down the words contained in each sentence. One group lip-read the lists with no sound (the LR:alone condition). The other group also heard a sequence of acoustic pulses which were synchronized to the moments when the talker's vocal folds closed (the LR&Lx condition). Performance was assessed both by loose (KW(L)) and by tight (KW(T)) keyword scoring methods. Both scoring methods produced the same pattern of results: performance was better in the LR&Lx condition; performance in both conditions improved linearly with the logarithm of the list presentation order number; subjects who produced higher overall scores also improved more with experience of the lists. The data were described well by a logistic regression model which provided a formula which can be used to compensate for practice effects and for differences in difficulty between lists. Two simpler, but less accurate, methods for compensating for variation in inter-list difficulty are also described. A figure is provided which can be used to assess the significance of the difference between a pair of scores obtained from a single subject in any pair of presentation conditions.
Two new developments in speech pattern processing hearing aids will be described. The first development is the use of compound speech pattern coding. Speech information which is invisible to the lipreader was encoded in terms of three acoustic speech factors; the voice fundamental frequency pattern, coded as a sinusoid, the presence of aperiodic excitation, coded as a low-frequency noise, and the wide-band amplitude envelope, coded by amplitude modulation of the sinusoid and noise signals. Each element of the compound stimulus was individually matched in frequency and intensity to the listener's receptive range. Audio-visual speech receptive assessments in five profoundly hearing-impaired listeners were performed to examine the contributions of adding voiceless and amplitude information to the voice fundamental frequency pattern, and to compare these codings to amplified speech. In both consonant recognition and connected discourse tracking (CDT), all five subjects showed an advantage from the addition of amplitude information to the fundamental frequency pattern. In consonant identification, all five subjects showed further improvements in performance when voiceless speech excitation was additionally encoded together with amplitude information, but this effect was not found in CDT. The addition of voiceless information to voice fundamental frequency information did not improve performance in the absence of amplitude information. Three of the subjects performed significantly better in at least one of the compound speech pattern conditions than with amplified speech, while the other two performed similarly with amplified speech and the best compound speech pattern condition. The three speech pattern elements encoded here may represent a near-optimal basis for an acoustic aid to lipreading for this group of listeners. The second development is the use of a trained multi-layer-perceptron (MLP) pattern classification algorithm as the basis for a robust real-time voice fundamental frequency extractor. This algorithm runs on a low-power digital signal processor which can be incorporated in a wearable hearing aid. Aided lipreading for speech in noise was assessed in the same five profoundly hearing-impaired listeners to compare the benefits of conventional hearing aids with those of an aid which provided MLP-based fundamental frequency information together with speech+noise amplitude information. The MLP-based pattern element aid gave significantly better performance in the reception of consonantal voicing contrasts from speech in pink noise than that achieved with conventional amplification and consequently, it also gave better overall performance in audio-visual consonant identification.(ABSTRACT TRUNCATED AT 400 WORDS)
A family of prototype speech pattern hearing aids for the profoundly hearing impaired has been compared to amplification. These aids are designed to extract acoustic speech patterns that convey essential phonetic contrasts, and to match this information to residual receptive abilities. In the first study, the presentation of voice fundamental frequency information from a wearable SiVo (sinusoidal voice) aid was compared to amplification in 11 profoundly deafened adults. Intonation reception was often better, and never worse, with fundamental frequency information. Four subjects scored more highly in audio-visual consonant identification with fundamental frequency information, five performed better with amplified speech, and two performed similarly under these two conditions. Five of the 11 subjects continued use of the SiVo aid after the tests were complete. A second study examined a laboratory prototype compound speech pattern aid, which encoded voice fundamental frequency, amplitude envelope, and the presence of voiceless excitation. In five profoundly deafened adults, performance was better in consonant identification when additional speech patterns were present than with fundamental frequency alone; the main advantage was derived from amplitude information. In both consonant identification and connected discourse tracking, performance with appropriately matched compound speech pattern signals was better than with amplified speech in three subjects, and similar to performance with amplified speech in the other two. In nine subjects, frequency discrimination, gap detection, and frequency selectivity were measured, and were compared to speech receptive abilities with both amplification and fundamental frequency presentation. The subjects who showed the greatest advantage from fundamental frequency presentation showed the greatest average hearing losses, and the least degree of frequency selectivity. Compound speech pattern aids appear to be more effective for some profoundly hearing-impaired listeners than conventional amplifying aids, and may be a valuable alternative to cochlear implants.
Control of voice fundamental frequency (Fx) during reading of a standard passage was assessed for four profoundly deafened adults receiving auditory feedback in one of two conditions: (i) with an extended low-frequency response amplifying aid; (ii) with the SiVo aid, which provided only Fx information. A third condition, where the subjects were unaided, was also included. For each feedback condition, quantitative analyses of laryngograph recordings were used to provide measures of Fx mode, 90% Fx range and the regularity of vocal fold vibration. In baseline unaided recordings, three subjects (S1, S2 and S4) showed some aspects of Fx control outside the normal range, while the other (S3) had appropriate Fx control. In the three subjects with impaired control, simplified Fx feedback led to better control than feedback from amplified speech. In two of these subjects, these differences were statistically significant. S3, who showed unimpaired Fx control, did not show any changes in Fx control under the different feedback conditions. Although the patterns of data were different in the individual subjects, simplified Fx feedback led to either improved or unimpaired control of Fx relative to speech feedback or to not feedback. These findings have important implications for speech-processing strategies implemented in hearing aids and cochlear implants, where the effects of different speech-coding strategies on production have been largely ignored.
The effect of three levels of text complexity upon connected discourse tracking rates was investigated in normal listeners who tracked by lipreading alone and by lipreading with auditorily presented voice pitch. Text complexity affected connected discourse tracking under both lipreading conditions, with tracking rates decreasing as the level of text complexity increased. The improvement in tracking rate with the addition of voice pitch information was found to be invariant over changes in text complexity when expressed as a simple difference between the two tracking rates.
Rehabilitation. RehabiUtation consists of training the patient to recognize simple sounds, to discriminate back ground noise and prosodic elements, and to hpread. Finally the patient reaches phonetic recognition without lipread ing (isolated vowel, word closed list, sentence closed list, word semantic list). The rehabilitation program will de pend on whether the patient is prelingually or postlingually deaf.
We have found that the larynx-frequency pattern of speech presented as a sinusoid can be of greater communicative value to profoundly hearing-impaired people than the complete acoustic signal. The presence of higher harmonics can give poorer labelling of isolated intonation contrasts and often minimal gain in segmental spectrally-based distinctions. These observations have led to the development of a practical, body-worn, pattern-processing hearing aid that uses a microprocessor to sense the (analogue-processed) speech fundamental frequency, transform it into an appropriate amplitude and frequency region, and generate digitally the required output sinusoid. Our findings have important implications for the design of other signal-processing hearing aids in demonstrating that a simplification of speech can lead to enhanced speech receptive abilities in persons with impaired hearing.
Although it is generally accepted that single-channel electrical stimulation can significantly improve a deafened patient's speech perceptual ability, there is still much controversy surrounding the choice of speech processing schemes. We have compared, in the same patients, two different approaches: (1) The speech pattern extraction technique of the EPI group, London (Fourcin et al., British Journal of Audiology, 1979,13,85-107) in which voice fundamental frequency is extracted and presented in an appropriate way, and (2) The analogue 'whole speech' approach of Hochmair and Hochmair-Desoyer (Annals of the New York Academy of Sciences, 1983, 405, 268-279) of Vienna, in which the microphone-sensed acoustic signal is frequency-equalized and amplitude-compressed before being presented to the electrode. With the 'whole-speech' coding scheme (which they used daily), all three patients showed an improvement in lipreading when they used the device. No patient was able to understand speech without lipreading. Reasonable ability to distinguish voicing contrasts and voice pitch contours was displayed. One patient was able to detect and make appropriate use of the presence of voiceless frication in certain situations. Little sensitivity to spectral features in natural speech was noted, although two patients could detect changes in the frequency of the first formant of synthesised vowels. Presentation of the fundamental frequency only generally led to improved perception of features associated with it (voicing and intonation). Only one patient consistently showed any advantage (and that not in all tests) of coding more than the fundamental alone.