We formalize and extend the contribution model of Clark and Schaefer (1987, 1989) so that it can be represented computationally; we then present a method for combining the turns of two individual agents into one incrementally determined, coherent representation of the processes of dialog. This representation is intended to approximate what a participant might represent about the dialog so far, for the immediate purpose of referring, making contextual inferences, and repairing problems of understanding, as well as for the longer term purpose of storing the products of dialog in memory. Such an approach, we argue, is necessary for enabling a computer-based partner to converse in a way that seems natural to a human partner.
This paper links prosody to the information in a text and how it is processed by the speaker. It describes the operation and output of LOQ, a text-to-speech implementation that includes a model of limited attention and working memory. Attentional limitations are key. Varying the attentional parameter in the simulations varies in turn what counts as given and new in a text, and therefore, the intonational contours with which it is uttered. Currently, the system produces prosody in three different styles: child-like, adult expressive, and knowledgeable. This prosody also exhibits differences within each style -- no two simulations are alike. The limited resource approach captures some of the stylistic and individual variety found in natural prosody.
I describe a limited-resource approach to generating prosody that mediates text-based information through a model of attention and working memory, whose simulation parameters are quantitative. The main parameter quanties recall. Varying it varies what counts as given and new in a text, and therefore, the pitch accents with which the text is uttered. Currently, the system produces prosody in three di erent styles of read speech { child-like, adult expressive, and knowledgeable { and individual variation within each. A comparison with natural data shows clear and predictable stylistic similarities, although not at signi cance. However, informal feedback is more forgiving, indicating that the prosody is both natural and expressive for consecutive phrases, but that work is still needed to make this e ect consistent throughout the text.
Statement of the Panel The purpose of this panel is to provide members of the IUI community with a look at where speech is heading in the near and not so near term. At present speech research has made great strides in speech recognition (to the point that large vocabulary, continuous dictation products are commercially available), some strides in speech understanding for limited tasks, and progress on synthesis (where products have long been available and continue to improve). Because of these
This paper introduces Linguistic Style Improvisation, a theory and set of algorithms for improvisation of spoken utterances by artificial agents, with applications to interactive story and dialogue systems. We argue that linguistic style is a key aspect of character, and show how speech act representations common in AI can provide abstract representations from which computer characters can improvise. We show that the mechanisms proposed introduce the possibility of socially oriented agents, meet the requirements that lifelike characters be believable, and satisfy particular criteria for improvisation proposed by Hayes-Roth.
This paper introduces Linguistic Style Improvisation, a theory and algorithms for improvisation of spoken utterances by artificial agents, with applications to interactive story and dialogue systems. We argue that linguistic style is a key aspect of character, and show how speech act representations common in AI can provide abstract representations from which computer characters can improvise. We show that the mechanisms proposed introduce the possibility of socially oriented agents, meet the requirements that lifelike characters be believable, and satisfy particular criteria for improvisation proposed by Hayes-Roth.
Expectations about the correlation of cue phrases, the duration of unfilled pauses and the structuring of spoken discourse are framed in light of Grosz and Sidner's theory of discourse and are tested for a directions-giving dialogue. The results suggest that cue phrase and discourse structuring tasks may align, and show a correlation for pause length and some of the modifications that speakers can make to discourse structure.
By strictest interpretation, theories of both centering and intonational meaning fail to predict the existence of pitch accented pronominals. Yet they occur felicitously in spoken discourse. To explain this, I emphasize the dual functions served by pitch accents, as markers of both propositional (semantic/pragmatic) and attentional salience. This distinction underlies my proposals about the attentional consequences of pitch accents when applied to pronominals, in particular, that while most pitch accents may weaken or reinforce a cospecifier's status as the center of attention, a contrastively stressed pronominal may force a shift, even when contraindicated by textual features.
This document is a revised version of my master's thesis, submitted in May, 1989 to the Media Arts and Sciences Section of theDepartment of Architecture, at the Massachusetts Institute of Technology. The revisions are as follows: grammatical and factualcorrections, particularly in Chapter 2; revised table formats to better conform to IPA standards; the addition of a table in Appendix B;and the addition of Appendix C 1 containing pitch tracks of, energy tracks and spectrograms of synthesized...
IntroductionWhen compared to human speech, synthesized speech is distinguished by insufficient intelligibility,inappropriate prosody and inadequate expressiveness. These are serious drawbacks for conversationalcomputer systems. Intelligibility is basic --- intelligible phonemes are necessary for word recognition.Prosody --- intonation (melody) and rhythm --- clarifies syntax and semantics and aids in discourseflow control. Expressiveness, or affect, provides information about the...
Synthesized English speech is readily distinguished from human speech on the basis of inappropriate intonation and insu cient expressiveness. This is a drawback for conversational computer systems. Intonation is the carrier of emphasis or de-emphasis, serving to clarify meaning for the spoken word much as variations in typeface and punctuation do for the written word. Expressiveness is not tied to word or phrase meaning but is global in scope. It provides the context in which the intonation occurs, and reveals the speaker's intentions and general mental state. In synthesized speech, intonation makes the message easier to understand; enhanced expressiveness contributes to dramatic e ect, making the message easier to listen to.
Synthesized speech need not be expressionless. By identifying the e ects of emotion on speech and choosing an appropriate representation, the generation of a ect is possible and can become computational. I describe a program | the A ect Editor | which implements an acoustical model of speech and generates synthesizer instructions to produce the desired a ect. The authenticity of the a ect is limited by synthesizer capabilities and by incomplete descriptions of the acoustical and perceptual phenomena. However, the results of an experiment show that this approach produces synthesized speech with recognizable, and, at times, natural, a ect.
Julia Hirschberg合作论文数Department of Computer Science, Columbia University2