Although the speech transmission index (STI) is a well-accepted and standardized method for objective prediction of speech intelligibility in a wide range of environments and applications, it is essentially a monaural model. Advantages of binaural hearing in speech intelligibility are disregarded. In specific conditions, this leads to considerable mismatches between subjective intelligibility and the STI. A binaural version of the STI was developed based on interaural cross correlograms, which shows a considerably improved correspondence with subjective intelligibility in dichotic listening conditions. The new binaural STI is designed to be a relatively simple model, which adds only few parameters to the original standardized STI and changes none of the existing model parameters. For monaural conditions, the outcome is identical to the standardized STI. The new model was validated on a set of 39 dichotic listening conditions, featuring anechoic, classroom, listening room, and strongly echoic environments. For these 39 conditions, speech intelligibility [consonant-vowel-consonant (CVC) word score] and binaural STI were measured. On the basis of these conditions, the relation between binaural STI and CVC word scores closely matches the STI reference curve (standardized relation between STI and CVC word score) for monaural listening. A better-ear STI appears to perform quite well in relation to the binaural STI model; the monaural STI performs poorly in these cases.
A procedure was developed for the automated measurement of the speech reception threshold in stationary noise (SRTn), which can be administered by the subjects themselves using a computer. The procedure was based on the SRTn test for Dutch developed by Plomp and Mimpen [(1979). "Improving the reliability of testing the speech reception threshold for sentences," Audiology, 18, 43-52]. Because in the automated procedure the responses were entered on a keyboard, the question of how to deal with typing and spelling errors played a key role. At first the possibility of scoring on keywords only was examined. An experiment was conducted in which the adaptive procedure was varied. Results showed that the combination of scoring each keyword separately and a fixed scheme of the adaptation of the signal-to-noise ratio throughout the procedure yields the highest test-retest reliability. Subsequently, the collection and verification of responses using a keyboard were examined. Two different algorithms were developed and evaluated against the traditional task of verbal repetition and response verification by an experimenter. The results indicated a preference for verification by dynamic alignment over a spelling checker approach. In conclusion, the results show that it is possible to automate the test procedure while maintaining sufficient reliability.
Although the speech transmission index (STI) is a well-accepted and standardized method for objective prediction of speech intelligibility in a wide range of environments and applications, it is essentially a monaural model. Advantages of binaural hearing to the intelligibility of speech are disregarded. In specific conditions this leads to considerable mismatches between subjective intelligibility and the STI. A binaural version of the STI was developed, based on interaural cross correlograms, which shows a considerably improved correspondence with subjective intelligibility in dichotic listening conditions. The new binaural STI is designed to be a relatively simple model which adds only few parameters to the original standardized STI, and changes none of the existing model parameters. For monaural conditions, the outcome is identical to the standardized STI. The new model was validated on set of 39 dichotic listening conditions, featuring anechoic, classroom, listening room, and cathedral environments. For these 39 conditions, subjective intelligibility (CVC-wordscore) was measured, as well as the binaural STI. The relation between binaural STI and CVC-wordscores in dichotic listening conditions closely matches the STI reference curve (standardized relation between STI and CVC-wordscore) for monaural listening. The monaural STI performs poorly in these cases.
The standardized method for determining the speech transmission index (STI) involves the use of a specific intensity-modulated test signal. The STI is obtained from measurements on the transmission channel, usually showing reductions of the modulation depths in the received test signal. Instead of using an artificial signal, various approaches have been suggested in the literature to use speech as a test signal [cf. Payton et al., J. Acoust. Soc. Am. 111, 2431 (2002)]. Such a speech-based STI has several advantages, e.g., predicting intelligibility differences due to speaking style and enabling the evaluation of vocoders. As we encountered shortcomings and inaccuracies in the existing speech-based STI methods, we propose a new procedure for estimating the speech-based modulation transfer function (MTF) which approaches the accuracy of conventional STI implementations. The new procedure uses the cross spectrum between the transmitted and received speech signals, with special phase weighting to address the relative importance of shifted modulations of the temporal envelope. Evaluation of the algorithm on a vocoder database showed promising results, yielding an average correlation coefficient of 0.93 between the subjective CVC scores and the speech-based STI for male speech. Details of this new speech-based STI algorithm will be discussed.
Speech intelligibility was investigated by varying the number of interfering talkers, level, and mean pitch differences between target and interfering speech, and the presence of tactile support. In a first experiment the speech-reception threshold (SRT) for sentences was measured for a male talker against a background of one to eight interfering male talkers or speech noise. Speech was presented diotically and vibro-tactile support was given by presenting the low-pass-filtered signal (0-200 Hz) to the index finger. The benefit in the SRT resulting from tactile support ranged from 0 to 2.4 dB and was largest for one or two interfering talkers. A second experiment focused on masking effects of one interfering talker. The interference was the target talker's own voice with an increased mean pitch by 2, 4, 8, or 12 semitones. Level differences between target and interfering speech ranged from -16 to +4 dB. Results from measurements of correctly perceived words in sentences show an intelligibility increase of up to 27% due to tactile support. Performance gradually improves with increasing pitch difference. Louder target speech generally helps perception, but results for level differences are considerably dependent on pitch differences. Differences in performance between noise and speech maskers and between speech maskers with various mean pitches are explained by the effect of informational masking.
Sentence intelligibility for interfering speech was investigated as a function of level difference, pitch difference, and presence of tactile support. A previous study by the present authors [J. Acoust. Soc. Am. 111, 2432–2433 (2002)] had shown a small benefit of tactile support in the speech-reception threshold measured against a background of one to eight competing talkers. The present experiment focused on the effects of informational and energetic masking for one competing talker. Competing speech was obtained by manipulating the speech of the male target talker (different sentences). The PSOLA technique was used to increase the average pitch of competing speech by 2, 4, 8, or 12 semitones. Level differences between target and competing speech ranged from −16 to +4 dB. Tactile support (B&K 4810 shaker) was given to the index finger by presenting the temporal envelope of the low-pass-filtered speech (0–200 Hz). Sentences were presented diotically and the percentage of correctly perceived words was measured. Results show a significant overall increase in intelligibility score from 71% to 77% due to tactile support. Performance improves monotonically with increasing pitch difference. Louder target speech generally helps perception, but results for level differences are considerably dependent on pitch differences.
Since long, different methods of vibrotactile stimulation have been used as an aid for speech perception by some people with severe hearing impairment. The fact that experiments have shown (limited) benefits proves that tactile information can indeed give some support. In our research program on multimodal interfaces, we wondered if normal hearing listeners could benefit from tactile information when speech was presented in adverse listening conditions. Therefore, we set up a pilot experiment with a male speaker against a background of one, two, four or eight competing male speakers or speech noise. Sound was presented diotically to the subjects and the speech-reception threshold (SRT) for short sentences was measured. The temporal envelope (0–30 Hz) of the speech signal was computed in real time and led to the tactile transducer (MiniVib), which was fixed to the index finger. First results show a significant drop in SRT of about 3 dB when using tactile stimulation in the condition of one competing speaker. In the other conditions no significant effects were found, but there is a trend of a decrease of the SRT when tactile information is given. We will discuss the results of further experiments.
In a 3D auditory display, sounds are presented over headphones in a way that they seem to originate from virtual sources in a space around the listener. This paper describes a study on the possible merits of such a display for bandlimited speech with respect to intelligibility and talker recognition against a background of competing voices. Different conditions were investigated: speech material (words/sentences), presentation mode (monaural/binaural/3D), number of competing talkers (1-4), and virtual position of the talkers (in 45 degrees-steps around the front horizontal plane). Average results for 12 listeners show an increase of speech intelligibility for 3D presentation for two or more competing talkers compared to conventional binaural presentation. The ability to recognize a talker is slightly better and the time required for recognition is significantly shorter for 3D presentation in the presence of two or three competing talkers. Although absolute localization of a talker is rather poor, spatial separation appears to have a significant effect on communication. For either speech intelligibility, talker recognition, or localization, no difference is found between the use of an individualized 3D auditory display and a general display.
In a 3-D auditory display, sounds are presented over headphones in a way that they seem to originate from virtual sources in a space around the listener. The possible merits of such a display were investigated with respect to speech intelligibility and speaker recognition against a background of competing speech. Various conditions were investigated: speech material (words or sentences), presentation mode (monaural, binaural, or 3D), number of competing speakers (1–4), and virtual position of the speakers (in 45° steps around the frontal horizontal plane). Average results for 12 listeners show an increase of speech intelligibility for a 3-D presentation, with more than two competing speakers compared to conventional monaural or binaural presentation. The acuity to recognize a speaker is slightly better and the time required for recognition is significantly shorter for a 3-D presentation in the presence of two or three competing speakers. Although absolute localizability of a speaker is rather poor, spatial separation appears to have a significant effect on communication. For either speech intelligibility, speaker recognition, or localizability, no difference is found between the use of an individualized 3-D auditory display and a general display. [Work supported by the Royal Netherlands Navy.]
In the European-Union-funded project AUDIS (AUditory DISplay), a multipurpose auditory display for 3D hearing applications is being developed. The main goal is to develop a special sound generator in combination with an auditory-symbol database that can, for example, be used in the avionic and automotive industry. As the project relies heavily on binaural technology and, hence, on the availability of reliable human HRTF data, a special program for collecting data has been undertaken. In the course of this program, round-robin tests have been performed in four different laboratories. Differences in the data across the different laboratories have been analyzed. The analysis resulted in recommendations for common measuring procedures to be applied. A catalog of HRTFs has then been collected and perceptual verification tests have been performed. It is planned to make the catalog publicly available. In the presentation, the measuring methods applied and the contents of the catalog will be discussed.
Modulations in the temporal intensity envelope of 24 1/4-octave bands were reduced by proportionally raising the troughs and lowering the peaks relative to the mean intensity in each band. The effect on intelligibility of various degrees of modulation reduction was investigated by measuring the speech-reception threshold (SRT) in noise. For conditions of severe modulation reduction, the number of correctly received sentences in quiet was scored. The effect of this deterministic modulation reduction was compared to the effect of stochastic modulation reduction obtained with addition of noise. Results for 12 normal-hearing subjects show that in the case of deterministic modulation reduction, intelligibility is reduced to 50% when the modulation-transfer factor equals 0.10, whereas in the case of modulation reduction by addition of noise, this intelligibility is reached already at a modulation-transfer factor of 0.27. This confirms that the effect of additive noise on intelligibility cannot be understood completely as a result of only modulation reduction. As suggested by Drullman [J. Acoust. Soc. Am. 97, 585-592 (1995)] two other factors associated with the addition of noise have to be taken into account: (1) the introduction of nonrelevant modulations, and (2) the corruption of the fine structure.
For many people with profound hearing loss conventional hearing aids give only little support in speechreading. This study aims at optimizing the presentation of speech signals in the severely reduced dynamic range of the profoundly hearing impaired by means of multichannel compression and multichannel amplification. The speech signal in each of six 1-octave channels (125-4000 Hz) was compressed instantaneously, using compression ratios of 1, 2, 3, or 5, and a compression threshold of 35 dB below peak level. A total of eight conditions were composed in which the compression ratio varied per channel. Sentences were presented audio-visually to 16 profoundly hearing-impaired subjects and syllable intelligibility was measured. Results show that all auditory signals are valuable supplements to speechreading. No clear overall preference is found for any of the compression conditions, but relatively high compression ratios (> 3-5) have a significantly detrimental effect. Inspection of the individual results reveals that compression may be beneficial for one subject.
In this paper the effect of temporal modulation reduction on spectral contrasts is investigated. First, a spectral modulation transfer function (SMTF) is presented as a method to measure the transfer of spectral ripples (sinusoidal periods/oct) in the short-time spectral envelope by comparing the spectral modulation depth of original and processed speech fragments. Measuring the SMTF for speech subjected to uniform reduction of the temporal modulation depth (i.e., modulation-frequency-independent reduction) in 24 1/4-oct bands showed an almost equal uniform reduction of the spectral modulations. Furthermore, the SMTF was used to measure the reduction of spectral contrasts associated with low-pass and high-pass temporal-envelope filtering [Drullman et al., J. Acoust. Soc. Am.95, 1053-1064 and 2670-2680 (1994a, b)]. For a perceptual evaluation, sentences were processed to reduce spectral contrasts and the speech-reception threshold (SRT) in noise was measured with ten normal-hearing subjects. Comparison of the results with those obtained previously after temporal-envelope filtering revealed that the SRT-effect of temporal high-pass filtering can be completely accounted for by the associated reduction of spectral contrasts. However, this relationship cannot be demonstrated conclusively in the case of temporal low-pass filtering.
This paper describes a number of listening experiments to investigate the relative contribution of temporal envelope modulations and fine structure to speech intelligibility. The amplitude envelopes of 24 1/4-oct bands (covering 100-6400 Hz) were processed in several ways (e.g., fast compression) in order to assess the importance of the modulation peaks and troughs. Results for 60 normal-hearing subjects show that reduction of modulations by the addition of noise is more detrimental to sentence intelligibility than the same degree of reduction achieved by direct manipulation of the envelope; in some cases the benefit in speech-reception threshold (SRT) is almost 7 dB. Two crossover levels can be defined in dividing the temporal envelope into two equally important parts. The first crossover level divides the envelope into two perceptually equal parts: Removing modulations either chi dB below or above that level yields the same intelligibility score. The second crossover level divides the envelope into two acoustically equal peak and trough parts. The perceptual level is 9-12 dB higher than the acoustic level, indicating that envelope peaks are perceptually more important than troughs. Further results showed that 24 intact temporal speech envelopes with noise fine structure retain perfect intelligibility. In general, for the present type of signal manipulations, no one-to-one relation between the modulation-transfer function and the intelligibility scores could be established.
In a previous study [R. Drullman, J. Acoust. Soc. Am. 97, 585–592 (1995)], the relative contribution of temporal modulations and fine structure to sentence intelligibility was investigated. This Letter reports additional listening experiments to assess in more detail the effect of masking noise on the peaks and troughs of the speech signal. For this purpose, the signal structure of each 1/4-oct band in a 24-band filterbank (100–6400 Hz) was altered by manipulating the distribution of speech and noise over the sentences. Results for 12 normal-hearing subjects indicate that removing noise from the peaks has no effect on intelligibility; removing the speech signal from the noisy troughs, however, yields a 2-dB increase of the speech-reception threshold. So, it appears that, even below the noise level, weak speech elements do contribute to intelligibility.
The effect of smearing the temporal envelope on the speech-reception threshold (SRT) for sentences in noise and on phoneme identification was investigated for normal-hearing listeners. For this purpose, the speech signal was split up into a series of frequency bands (width of 1/4, 1/2, or 1 oct) and the amplitude envelope for each band was low-pass filtered at cutoff frequencies of 0, 1/2, 1, 2, 4, 8, 16, 32, or 64 Hz. Results for 36 subjects show (1) a severe reduction in sentence intelligibility for narrow processing bands at low cutoff frequencies (0-2 Hz); and (2) a marginal contribution of modulation frequencies above 16 Hz to the intelligibility of sentences (provided that lower modulation frequencies are completely present). For cutoff frequencies above 4 Hz, the SRT appears to be independent of the frequency bandwidth upon which envelope filtering takes place. Vowel and consonant identification with nonsense syllables were studied for cutoff frequencies of 0, 2, 4, 8, or 16 Hz in 1/4-oct bands. Results for 24 subjects indicate that consonants are more affected than vowels. Errors in vowel identification mainly consist of reduced recognition of diphthongs and of confusions between long and short vowels. In case of consonant recognition, stops appear to suffer most, with confusion patterns depending on the position in the syllable (initial, medial, or final).
The effect of reducing low-frequency modulations in the temporal envelope on the speech-reception threshold (SRT) for sentences in noise and on phoneme identification was investigated. For this purpose, speech was split up into a series of frequency bands (1/4, 1/2, or 1 oct wide) and the amplitude envelope for each band was high-pass filtered at cutoff frequencies of 1, 2, 4, 8, 16, 32, 64, or 128 Hz, or infinity (completely flattened). Results for 42 normal-hearing listeners show: (1) A clear reduction in sentence intelligibility with narrow-band processing for cutoff frequencies above 64 Hz; and (2) no reduction of sentence intelligibility when only amplitude variations below 4 Hz are reduced. Based on the modulation transfer function of some conditions, it is concluded that fast multichannel dynamic compression leads to an insignificant change in masked SRT. Combining these results with previous data on low-pass envelope filtering (temporal smearing) [Drullman et al., J. Acoust. Soc. Am. 95, 1053-1064 (1994)] shows that at 8-10 Hz the temporal modulation spectrum is divided into two equally important parts. Vowel and consonant identification with nonsense syllables were studied for cutoff frequencies of 2, 8, 32, 128 Hz, and infinity, processed in 1/4-oct bands. Results for 12 subjects indicate that, just as for low-pass envelope filtering, consonants are more affected than vowels. Errors in vowel identification mainly consist of reduced recognition of diphthongs and of durational confusions. For the consonants there are no clear confusion patterns, but stops appear to suffer least. In most cases, the responses tend to fall into the correct category (stop, fricative, or vowel-like).
In this paper, three experiments are reported that were run in order to assess the quality of Dutch synthetic speech using accented and unaccented diphones, i.e., diphones extracted from accented and unaccented syllables, respectively. In a paired-comparison design, subjects were asked to evaluate the naturalness and fluency of different versions of an utterance. The results of the first two experiments, in which isolated polysyllabic words were used, indicate that the use of accented or unaccented diphones has a perceptual effect on phonologically long vowels only. In a third experiment, the use of the different diphone types in short sentences with a fixed temporal structure was evaluated. Results suggest that using unaccented diphones in unaccented or secondarily accented syllables does not result systematically in more natural-sounding speech.