We present a meta-analysis of results from experimental studies on attitude reception in seven languages (Brazilian Portuguese, Japanese, French, German, Cantonese, American English, Hindi). The studies involved free-labeling of perceived attitudes in audio-visual stimuli. The productions of 88 speakers from the seven languages were obtained using the same elicitation methodology, allowing to record sixteen audiovisual attitudes. These performances, rated in preceding works using a free-labeling paradigm, were grouped and analyzed to compare how the attitudinal performances spread along the main dimensions of the shared cognitive representation. A hierarchical clustering then regrouped attitudinal expressions as a function of their cognitive proximity. A large cluster solution showed the main dimension that organizes these expressions may be interpreted as the “Unpredictability” dimension proposed by Fontaine et al. (2007) for emotions, followed by the Evaluation-Pleasantness one; Activation-Arousal and Potency-Control arrived later but played a determinant role in the organization of attitudes. A fine-grained 13-cluster solution showed most attitudes were singled out by the listeners despite the variations in speakers and elicitation contexts. The analysis of each of these clusters brings insight into the cultural similarities and differences in the reception of these different attitudinal expressions. A notable result is the variation in valence attributed to the expression of Irony that underlines the potential communication problems that may be linked to interaction routines. On the other hand, Surprise was clearly identified. The existence and importance of the Unpredictability dimension and its relation to the illocutionary opposition between assertive and interrogative acts underlines the pertinence of Mello & Raso’s (2011) analysis.
This paper presents a corpus of attitudinal expressions in Hindi and their evaluation by native Hindi speaking raters. The paradigm is adapted from Rilliard et al. (2013), and forms part of an intercultural endeavour aimed at studying prosodic and facial expressions of social affect across languages. Our corpus includes a total of 16 attitudes, such as arrogance, surprise, politeness etc. portrayed by 19 speakers (10f, 9m) of which three experts selected the best four males and females and the better of two turns for subsequent analysis. A follow-up experiment with more participants employed a total of 512 stimuli, 256 hereof full audio-visual stimuli, 128 audio-only and 128 silent video stimuli. Results indicate higher ratings for emotionally loaded attitudes such as irritation and doubt as compared to, for instance, irony or seductiveness. Reduced modality stimuli were rated more poorly.
A low-resource emotional speech synthesis system for empathetic speech synthesis based on modelling prosody features is presented here. Secondary emotions, identified to be needed for empathetic speech, are modelled and synthesised in this investigation. As secondary emotions are subtle in nature, they are difficult to model compared to primary emotions. This study is one of the few to model secondary emotions in speech as they have not been extensively studied so far. Current speech synthesis research uses large databases and deep learning techniques to develop emotion models. There are many secondary emotions, and hence, developing large databases for each of the secondary emotions is expensive. Hence, this research presents a proof of concept using handcrafted feature extraction and modelling of these features using a low-resource-intensive machine learning approach, thus creating synthetic speech with secondary emotions. Here, a quantitative-model-based transformation is used to shape the emotional speech’s fundamental frequency contour. Speech rate and mean intensity are modelled via rule-based approaches. Using these models, an emotional text-to-speech synthesis system to synthesise five secondary emotions-anxious, apologetic, confident, enthusiastic and worried-is developed. A perception test to evaluate the synthesised emotional speech is also conducted. The participants could identify the correct emotion in a forced response test with a hit rate greater than 65%.
This paper reports on the role of technology in state-of-the-art pronunciation research and instruction, and makes concrete suggestions for future developments. The point of departure for this contribution is that the goal of second language (L2) pronunciation research and teaching should be enhanced comprehensibility and intelligibility as opposed to native-likeness. Three main areas are covered here. We begin with a presentation of advanced uses of pronunciation technology in research with a special focus on the expertise required to carry out even small-scale investigations. Next, we discuss the nature of data in pronunciation research, pointing to ways in which future work can build on advances in corpus research and crowdsourcing. Finally, we consider how these insights pave the way for researchers and developers working to create researchinformed, computer-assisted pronunciation teaching resources. We conclude with predictions for future developments.
In earlier works we examined four types of speech acts in several Gallo-Romance dialects: statements, polar questions, incredulous questions (expecting a negative response), and queries for confirmation (expecting a positive response). In a perception experiment we found that the first two types are generally reliably identified whereas for the latter two, only a limited number of utterances yielded recognition rates far above chance. We identified confusions mostly between polar and incredulous questions, as well as between the other two. We also established that differences that facilitated discrimination mostly concerned the fundamental frequency ( F0 ) contours. In the current work we conduct a perceptual experiment employing synthetic stimuli, involving pairs of reliably recognized, but by type confusable utterances and attempt to morph one type of speech act into the confusable other. Our results indeed show, that manipulating the utterance-final F0 contour appropriately actually increases the probability of one type of speech act being confused with the other, and what the categorical perception thresholds are. as low vs. high onsets. In conclusion our study shows that the resynthesis paradigm is very well suited to investigate and confirm our earlier categorical observations on natural speech data and a useful tool in subsequent works.
In an earlier exploratory study we examined the prosody of speeches by two IT industry leaders, Steve Jobs and Marc Zuckerberg, whose perceived charisma differs greatly, with Jobs usually regarded as the much more captivating speaker.This previous study focused mainly on fundamental frequency contours as well as on the perceived local speech rate.Instead of analyzing the raw F0 data directly, we modeled the F0 contours using the Fujisaki model and examined the differences in the respective model components.Whereas in our comparison between Jobs and Zuckerberg we were only able to examine distributions of Fujisaki model parameters, in the current study we decided to systematically vary some of the Fujisaki model parameters on a fixed set of utterances and investigate their effects on perceived charisma.We found that in general pitch range extensions are beneficial, especially when connected to accented syllables, but also that effects differ considerably between male and female speakers.
In the current study four types of speech acts in several Gallo-Romance dialects are examined: statements, polar questions, incredulous questions (expecting a negative response), and queries for confirmation (expecting a positive response). In a perception experiment we found that the first two types are generally reliably identified whereas for the latter two, only a limited number of utterances yielded recognition rates far above chance. We identified confusions mostly between polar and incredulous questions, as well as between the other two. When examining prosodic differences between the four classes that facilitated discrimination, we found that they mostly concerned the fundamental frequency ( F 0 ) contours, especially in the latter part of the utterances.
Following up on earlier experiments on the cross-cultural and cross-language perception of short audio-visual utterances produced with varying attitudinal expressions, we compare the verbal responses of native speakers of Hindi with those of German and Cantonese-speaking evaluators to stimuli in the latter two languages. Contrary to our expectations, however, most Indian participants felt most confident rating the stimuli in English and not Hindi. As we had already previously translated all reply terms by Germans and Hong Kong raters to English, we decided to stay in the same language for the cross-language evaluation and draw on ratings of valence, arousal and dominance from a study of almost 14,000 lemmas. We converted our pre-existing labels to this reference system and compared them to the responses of the Hindi speakers. We found that the type of attitude, the rater language but also the stimulus language had significant influence on the raters’ responses that differed in at least two of the dimensions. When we calculated correlations within and between rater groups, we found that the speakers of Hindi were better able to replicate the judgments of the other two groups on stimuli in their own languages than the group ignorant of that language. Semantic analysis of responses revealed that attitudes associated with strong negative emotions such as doubt and anger are picked up well by the non-speakers, whereas more complex attitudes, viz. seductiveness and irony are not.
This chapter reviews commonly recurring tendencies in the phonetic realization of tones, both in intonation and in lexical tone systems. It discusses local interactions between tonal targets, such as tonal coarticulation, dissimilatory H-raising, and rightward target displacement. Non-coarticulatory patterns include globally oriented patterns such as declination, look-ahead upstep, and final lowering as well as interactions between tone and the segmental skeleton, such as segmental anchoring, timing adjustments based on syllable structure or segmental features, and patterns of duration-driven truncation and compression of tone melodies. The chapter also considers morphosyntactically, pragmatically, and metalinguistically conditioned hyperarticulation effects arising from prominence or the Lombard effect. Lastly, it discusses issues relating to contour shape, such as the convexity or concavity of f0 movements, plateau versus sharp peak shapes for f0 maxima, and the propensity for L tones to be accompanied by a falling or dipping f0.
Bernd J. Kroger, Jim Kannampuzha, Dominik Bauer, Peter Birkholz, Philippe Dreuw, Hermann Ney An Action-Based Concept for the Phonetic Annotation of Sign Language Gestures 33 Sascha Fagel, Gérard Bailly Speech, Gaze and Head Motion in a Face-to-Face Collaborative Task 40 Ralf Winkler, Gunter Uhlmann, Gerd Schneider Maschinelle Klassifikation von Artikulationsbewegungen im Rahmen einer visuellen Artikulationsschulung für gehörlose und schwerhöriger Kinder 48 Prosody and Affect Benjamin Weiss, Sebastian Möller, Tim Polzehl Wirkung menschlicher Stimme auf die wahrgenommene SympathieEinfluss der Stimmanregung anhand von Laryngogrammen 56 Jürgen Trouvain Affektäußerungen in Sprachkorpora 64 Sören Wittenberg, Oliver Jokisch Das Prosodisch-Phonetische Annotationssystem PROPHANO 71
This paper evaluates results from a cross-cultural and crosslanguage experiment series employing short audio-visual utterances produced with varying attitudinal expressions. German and Cantonese-speaking participants freely labeled such utterances in the two languages and assigned to each stimulus a verbal label. Based on the results of the four experiments we were able to establish to what degree the attitudinal frames of reference of the two groups overlap and how they differ. Verbal labels were assessed regarding their emotional content in terms of valence, activation and dominance, and for the linguistic opposition between assertive and interrogative speech act, and hence permit to abstract from the language of the rater and ultimately even abstract from the attitudinal categories used when eliciting the stimuli. Instead we regard each utterance as a data-point in the emotional space. We found that the judgments of the two rater groups agree well with respect to the valence of attitudinal expressions and diverge most as to the perceived activation of the stimulus presenter. Cantonese speaking participants seem to mirror Germans’ ratings of German stimuli better than vice versa, which suggests an interesting asymmetry of attitudinal perception. As for the modality of presentation, the audio channel primarily transmits linguistically relevant information regarding the opposition of assertion and interrogation while the visual information signals the emotional content.
Prominence is a perceptual attribute employed to communicate focus, contrasts and expressive nuances. This article ex-plores the automatic detection of segments considered prominent by native listeners, using a corpus of Argentinean Spanish. The prominence detection is modeled as a binary classification problem over syllabic units. From perceptual assessments by a group of native listeners, we obtained a set of prominent syllable annotations, which are used as the gold standard to train and evaluate automatic classifiers. We study the performance of the classifiers under different sets of acoustic features, under various combinations of syllabic contexts, and using different classification algorithms. The best overall performance using leave-one speaker out cross validation had a mean precision rate of 94.75%, and was obtained using an SVM classifier, with two context syllables around each side of the central syllable, and applying the complete set of acoustic features considered.
This paper reports results from a free labeling experiment employing short audio-visual utterances of Cantonese produced with varying attitudinal expressions. It is part of a series of such experiments with a cross-language setting between German and Cantonese. Cantonese-speaking perceivers were asked to specify a single word that best described these stimuli, which were presented in audio-visual, audio-only, and video-only modalities. The resulting terms were analyzed with respect to the emotional dimensions of valence, activation and dominance, as well as the linguistic dimension of assertion/interrogation. The analysis results are compared with the outcomes from similar experiments employing German stimuli with Cantonese perceivers, as well as German perceivers assessing both German and Cantonese stimuli. It is found that Cantonese perceivers judge the Cantonese stimuli as more activated than German perceives do. The valence judgments agree relatively well, however, “polite” stimuli were judged less positively by Cantonese perceivers. Generally speaking, valence judgments are mostly influenced by the stimuli whereas activation and dominance judgments depend more on the perceiver group.
This study examines at a new level of quantitative detail the intonation and timing properties of charismatic speech by comparing two popular CEOs, Steve Jobs and Mark Zuckerberg, who are known from informal observations and formal perception experiments alike to be more or less charismatic speakers, respectively. By applying the Fujisaki model we decomposed F0 contours into baseline frequency, phrasal F0 excursions and pitch accent-associated F0 excursions. Timing details are examined by applying Pfitzinger’s model of perceived local speech rate to phone and syllable segmentations. Results suggest that high pitch not only involves generally higher F0 levels, but that these increases in F0 are not the same for every prosodic domain or level of the Fujisaki model. In addition we found significant differences depending on whether customers or investors are addressed.
Focus or prominence is an important linguistic function of prosody. The acoustic realisation of prominence in an utterance, in most languages, involves one or more acoustic dimensions while affecting one or more words in the utterance. It is of interest to identify the acoustic correlates as well as their possible interaction in the production and perception of focus. In this article, we consider the acoustics of focus in Marathi. Previous studies on Hindi, the more researched member of the Indo-Aryan family, have reported that the well-known rising F0 pattern on non-final content words in an utterance becomes hyper-articulated when the word is in focus. The associated F0 excursion, duration and intensity increase and are accompanied by post-focal compression of pitch range. A preliminary goal of the present study was to verify whether Marathi exhibits similar behaviour. We used Subject-Object-Verb (SOV) structured utterances with elicited focus on each word by 12 native Marathi speakers. We observed that each narrow focus location is accompanied by a distinct set of local and global acoustic correlates in F0, duration and intensity which closely parallel previous observations on Hindi. F0 cues were also examined via the accent command amplitudes of the Fujisaki model. F0 range, duration and intensity were found to vary significantly with focus condition prompting a study to examine their relative importance in the perceptual judgement of focus. Perception testing with synthetically manipulated utterances revealed that duration cues are interpreted in a categorical manner, relatively uninfluenced by the pitch cues. Only when duration is ambiguous, does the on-focus F0 cue appear to play a role. An explanation for this may lie in the normal F0-rise characteristic of the content words in Marathi, making pitch a less dependable functional cue for focus. (C) 2017 Elsevier Ltd. All rights reserved.
Based on the paradigm by Rilliard et al. we collected audio-visual expressions of attitudes such as arrogance, irony, sincerity and politeness in German. In the experimental design subjects are immersed in sixteen different communicative situations in which they are supposed to portray a certain attitude in a short dialog. Attitudes can be propositional, that is, reactions to a factual situation and/or social, that is, with respect to the relationship with the collocutor. Furthermore, attitudes can be of positive or negative valence or neutral. Undeniably there is a large repertory of subtle differences in the way certain talkers express certain attitudes. The important question is, however, whether collocutors either from the same language or a different one can actually decode these attitudes reliably. On that account we carried out three perceptual experiments in which we presented our recordings of the portrayed attitudes audio-visually, audio-only and video-only. In the first study, German perceivers rated the expressions given the intended attitude, in the second study, they had to choose the most suitable in a choice of five attitudes, and in the third study raters were able to assign freely the term best matching each attitudinal expression. This last experiment was recently replicated by native speakers of Cantonese in Hong Kong. The current article reviews and reevaluates the results from the first three experiments with the German subjects under the premise that perceivers actually have a more limited set of attitudinal registers which they can reliably draw on. This means that expressions can be sorted into a much smaller number of categories than the projected sixteen. In addition we compare and contrast these resulting clusters with the new data from the Cantonese speaking group. Our results indicate indeed a small number of readily decoded attitudes forming four clusters depending on the experiment design - which are also distinct acoustically. Clusters from the statistical analysis are very similar for the German and the Cantonese perceivers and overlap with basic emotions. This result suggests that expressions of attitudes with low identification rates are more complex to decode and require more pragmatic information, that is, more contextual and possibly idiosyncratic information to be interpreted correctly.