Automatic speech recognition (ASR) for wideband (WB) telephone speech services must cope with a lack of matching speech databases for acoustic model training. This paper investigates the impact of mixing insufficient WB and additional narrowband (NB) speech training data. It turns out that decimation and interpolation techniques, reducing the bandwidth mismatch between the NB speech material in training and the WB speech data to be recognized, do not succeed in outperforming the pure NB ASR baseline. However, true WB ASR training supported by artificial bandwidth extension (ABE) reveals a performance gain. A new ABE approach that makes use of robust dynamic features and a Viterbi path decoder exploiting phonetic a priori knowledge proves to be superior. It yields a reduction of 1.9 % word error rate relative to the NB ASR baseline and 9.3 % relative to a WB ASR experiment trained on only a limited amount of WB speech data.
In telephony applications, artificial bandwidth extension (ABE) can be applied to narrowband (NB) calls for speech quality and intelligibility enhancement. However, high-band extension is challenging due to insufficient mutual information between the lower and upper frequency band in speech. Estimation errors particularly of fricatives /s, z/ are the consequence leading to annoying artifacts, such as lisping. In this paper, two neural networks are employed to support an HMM-based ABE: The first one detects /s, z/ phonemes to assist the estimation process, while the second one corrects the estimated high-band energy. In an absolute category rating test the proposed ABE attains a significantly improved speech quality vs. NB speech. This is confirmed by a comparison category rating test pointing out a speech quality gain of 1.0 CMOS points over NB speech.
During the transition to wideband speech telephony, artificial bandwidth extension (ABE) could help to preserve customer satisfaction by enhancing speech quality in case of narrowband (NB) calls. However, the assessment of speech quality for ABE systems is still an open question. In the literature, instrumental measures are often used to judge the quality of ABE solutions. When subjective listening tests are considered, they most often use a comparison category rating (CCR) scale and, more rarely, an absolute category rating (ACR) scale. This paper investigates the relevance of instrumental and subjective assessment methods for ABE systems. An ACR and a CCR test are organized. Their results are compared and discussed. Discrepancies between these two tests open the discussion for the design of a proper subjective listening test for ABE systems. Some instrumental measures are also evaluated. A poor correlation between these measures and the subjective results is observed.
During the transition period from narrowband to wideband speech transmission services, Artificial Bandwidth Extension (ABE) algorithms are able to reduce the perceptual degradation of narrowband-transmitted speech signals by extending the audio bandwidth. In this paper, we analyze whether the resulting speech quality can be predicted reliably with instrumental models. Estimations from the new ITU standard POLQA, its predecessor WB-PESQ and the diagnostic DIAL model are compared to subjective listener judgments. This comparison reveals that the instrumental measures are not fully able to cope with ABE-processed speech, particularly in predicting ABE rank orders reliably. Reasons for this finding and corresponding diagnoses are discussed.
Because of its limited bandwidth, telephone speech is poorly intelligible. Artificial bandwidth extension (ABWE) reconstructs themissing frequencies aiming at, e.g., higher intelligibility. It was recently demonstrated that hearing-impaired persons wearing a hearing aid benefit from ABWE-enhanced telephone speech. However, it is unclear, whether persons without hearing impairment also take profit from ABWE in the same test conditions and if so, to what extent. This paper presents a subjective listening test with normal-hearing subjects based on meaningless German syllables simulating narrowband (NB), ABWE-enhanced and wideband (WB) telephone speech in two noisy listening conditions. The test results reveal a clear impact of hearing impairment on the ABWE capability to improve telephone intelligibility. For a signal-to-noise ratio (SNR) of 0 dB, subjects with and without hearing impairment similarly benefit from ABWE. At 20 dB SNR, hearing-impaired subjects take even more profit in contrast to normal-hearing subjects.
Today’s instrumental speech quality measures are limited in their use as they ”do not yet sufficiently include processing steps beyond the periphery of the auditory sys-tem” [1]. This becomes particularly obvious when using reference-based instrumental methods to assess the quality of artificial speech bandwidth extension (ABWE) approaches. While Blauert and Jekosch [1] have not proposed particular schemes, they advocate a model of sound quality representing layers of abstraction. In fact, once subjects are asked for opinion scores following any of ITU-T’s definitions, they have already understood (or not) what was spoken. It is our firm conviction that in not-too-bad testing conditions this knowledge serves as internal reference for judging speech quality – which in consequence asks for a paradigm shift of reference-based instrumental speech quality measures. In consequence, not only some (direct wideband) reference speech data is useful, but also a phonetic transcription of the speech, serving as human-internal representation of what was spoken. The paper will give thoughts to support this thesis, along with a proof that not all sounds are equal, asking for a phoneme-specific processing of future reference-based instrumental speech quality assessment methods.
Due to its limited acoustic bandwidth, conventional telephone speech is poorly intelligible, particularly for older and/or hearing impaired persons. Artificial bandwidth extension (ABWE) aims at improving the intelligibility of narrowband (NB) speech by estimating and reconstructing the missing frequency components. Being employed at the receiver side, it does not require changes in the speech transmission system. Hence, it could be integrated into a telephone or directly into the hearing aid. This paper investigates the potential of ABWE to improve the intelligibility of NB telephone speech for hearing impaired persons wearing a hearing aid. Subjective listening tests on meaningless German logatomes demonstrate a significant improvement for critical fricatives, particularly for /s/.
In anticipation of upcoming mobile telephony services with higher speech quality, a wideband (50 Hz to 7 kHz) mobile telephony derivative of TIMIT has been recorded called WTIMIT. It opens up various scientific investigations; e.g., on speech quality and intelligibility, as well as on wideband upgrades of network-side interactive voice response (IVR) systems with retrained or bandwidth-extended acoustic models for automatic speech recognition (ASR). Wideband telephony could enable network-side speech recognition applications such as remote dictation or spelling without the need of distributed speech recognition techniques. The WTIMIT corpus was transmitted via two prepared Nokia 6220 mobile phones over T-Mobile's 3G wideband mobile network in The Hague, The Netherlands, employing the Adaptive Multirate Wideband (AMR-WB) speech codec. The paper presents observations of transmission effects and phoneme recognition experiments. It turns out that in the case of wideband telephony, server-side ASR should not be carried out by simply decimating received signals to 8 kHz and applying existent narrowband acoustic models. Nor do we recommend just simulating the AMR-WB codec for training of wideband acoustic models. Instead, real-world wideband telephony channel data (such as WTIMIT) provides the best training material for wideband IVR systems.
In the past, artificial bandwidth extension (ABWE) has primarily been investigated to enhance transmitted narrowband speech signals at the receiving side. State-of-the-art schemes show improved quality versus narrowband speech; however, a clear gap to wideband speech is still reported. This is largely due to the insufficient ABWE performance on fricatives, particularly /s/. We asked ourselves to what extent the speech quality could be improved, if we knew the currently spoken phoneme. In this paper we present a framework using phonetic transcriptions as a-priori knowledge besides the speech waveform. Possible applications are high-quality offline ABWE of telephone, pilot, or historic speech recordings, memory efficient narrowband speech synthesis followed by ABWE, and extension of narrowband telephone databases to train wideband acoustic models for automatic speech recognition. For the classical conversational telephony application, an improved ABWE scheme is also proposed making use of transcription information only during training.
Artificial bandwidth extension techniques can be employed in mobile terminals to improve the intelligibility and quality of the far-end speaker’s speech signal at the receiver. To accomplish this, usually statistical models are trained requiring wideband speech material from the conversational partner, or at least from the language that is expected to be used in the conversation. In practice however, both, the speaker and language of a certain phone conversation are not known to the user equipment. Therefore we investigated the performance of an HMMbased multilingually trained artificial bandwidth extension on speech signals of which the speaker and language were unseen in training. The cross-language training and test turned out to cause only minor degradations compared to the use of monolingually trained acoustic models of the language used in test. The experimental results further showed that both of these speaker-independent methods could even keep up with the speaker-dependent technique to a large extent. Our findings indicate that artificial bandwidth extension can be effciently trained with speakerand language-independent speech data without significant losses in speech intelligibility and quality.
Artificial bandwidth extension techniques can be employed in mobile terminals to improve the quality of the far-end speaker's signal at the receiver. To accomplish this, usually statistical models are trained requiring wideband speech material from a language that is expected to be used in the conversation. In practice however, the language of a certain phone conversation is not known to the user equipment. Therefore we investigated the performance of an HMM- based multilingually trained artificial bandwidth extension on speech signals of which the language was unseen in training. The cross-language training and test turned out to cause only minor degradations compared to the use of monolingually trained acoustic models of the language used in test. Our findings indicate that artificial bandwidth extension can be efficiently trained with multilingual speech data without significant losses in speech quality.
In modern and innovative videoconference systems and human machine interfaces, localisation techniques play an important role for automatic camera and beamformer steering. Conventional acoustical and visual localisation techniques can be combined to form an audiovisual joint location estimate providing a more robust localisation of the active person. Tracking algorithms such as the well-known Kalman or extended Kalman filter and also particle filters can serve to further improve the location estimates. This paper is about a problem-specific SIR particle filtering algorithm applied to an existing audiovisual speaker localisation. The performance of the suggested algorithm will be evaluated using real audiovisual data based on experiments. It turns out that the proposed algorithm is able to improve the audiovisual location estimates.