In this paper methods for determination of perceptual quality of speech transmission and speech enhancement systems are discussed. The rating categories introduced in the listening test (7.) are not meant to replace existing "objective" or "technical" measures. They are meant to complement the description of a device under test and place more emphasize on the end-users impression of the system. The rating categories introduced in this paper are complete and intuitive, even non-experts can understand their importance for successful and agreeable speech communication. Benchmark tests using these categories are much more informative for the non-expert and can still be presented "at a glance". Furthermore, the comparison ratings serve to clarify the responsibility for quality losses, which the device under test cannot avoid, since their origin lies in the mobile network.
Voice and speech parameters for a single speaker vary widely over different contexts, in particular in situations in which speakers are affected by stress or emotion or in which speech styles are used strategically. This high degree of intra-speaker variability presents a major challenge for speaker verification systems. Based on a large-scale study in which different kinds of affective states were induced in over 100 speakers from three language groups, we use a statistical approach to identify speech and voice parameters that are likely to strongly vary as a function of the respective situation and affective state as well as those that tend to remain relatively stable. In addition, we evaluate the latter with respect to their potential to differentiate individual speakers.
It is argued that reliable acoustic profiles of speech under stress can only be found if different types of stress are clearly distinguished and experimentally induced. We report first results of a study with 100 speakers from three language groups, using a computer-based induction procedure that allows distinguishing cognitive load due to task engagement from psychological stress. Findings show significant effects of load, and partly of stress, for speech rate, energy contour, F0, and spectral parameters. It is further suggested that the mean results for the complete sample of speakers do not reflect the amplitude of stress effects on the voice. Future research should isolate and focus on speakers for whom the psychological stress induction has been successful.
Speech databases used in studies on emotional expression are often too small to represent a realistic survey on how humans express their emotions in spoken language. Large databases give a better survey, but with them acoustic analyses can hardly be performed manually. In consequence all those effects a trained phonetician might discover using acoustic or visual representations of the signals have to be defined in programs before any “automatic” measurement of these features can be carried out.
The ongoing work described in this contribution attempts to demonstrate the need to train ASV algorithms on emotional speech, in addition to neutral speech, in order to achieve more robust results in real life verification situations. A computerized induction program with 6 different tasks, producing different types of stressful or emotional speaker states, was developed, pretested, and used to record French, German, and English speaking participants. For a subset of these speakers, physiological data were obtained to determine the degree of physiological arousal produced by the emotion inductions and to determine the correlation between physiological responses and voice production as revealed in acoustic parameters. In collaboration with a commercial ASV provider (Ensigma Ltd.), a standard verification procedure was applied to this speech material. This paper reports the first set of preliminary analyses for the subset of 30 German speakers. It is concluded that an evaluation of the promise of training ASV material on emotional speech requires in-depth analyses of the individual differences in vocal reactivity and further exploration of the link between acoustic changes under stress or emotion and verification results.
During the last years a growing interest in automatic speaker verification (ASV) systems developed. Recent ASV systems still produce two kinds of errors to a considerable extent: false acceptances and false rejections of speakers. The systems generally react extremely sensitive to intra-speaker variability of voice and speaking style, caused for example by different psychological states of the speaker. To improve the performance of ASV systems, different attempts can be made to modify the underlying mathematical or statistical models. Furthermore, most of the systems still use a set of spectral parameters as originally developed for speech recognition systems. The performance of ASV systems could be improved by selecting a set of acoustic parameters which show both, minimal intraspeaker variation and maximal inter-speaker variation.