An accessible explanation of the technologies that enable such popular voice-interactive applications as Alexa, Siri, and Google Assistant. Have you talked to a machine lately? Asked Alexa to play a song, asked Siri to call a friend, asked Google Assistant to make a shopping list? This volume in the MIT Press Essential Knowledge series offers a nontechnical and accessible explanation of the technologies that enable these popular devices. Roberto Pieraccini, drawing on more than thirty years of experience at companies including Bell Labs, IBM, and Google, describes the developments in such fields as artificial intelligence, machine learning, speech recognition, and natural language understanding that allow us to outsource tasks to our ubiquitous virtual assistants. Pieraccini describes the software components that enable spoken communication between humans and computers, and explains why it's so difficult to build machines that understand humans. He explains speech recognition technology; problems in extracting meaning from utterances in order to execute a request; language and speech generation; the dialog manager module; and interactions with social assistants and robots. Finally, he considers the next big challenge in the development of virtual assistants: building in more intelligence—enabling them to do more than communicate in natural language and endowing them with the capacity to know us better, predict our needs more accurately, and perform complex tasks with ease.
chapter Natural Language Understanding in Socially Interactive Agents Share on Author: Roberto Pieraccini View Profile Authors Info & Claims The Handbook on Socially Interactive Agents: 20 years of Research on Embodied Conversational Agents, Intelligent Virtual Agents, and Social Robotics Volume 1: Methods, Behavior, CognitionSeptember 2021 Pages 147–172https://doi.org/10.1145/3477322.3477328Online:02 October 2021Publication History 0citation2DownloadsMetricsTotal Citations0Total Downloads2Last 12 Months2Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access
These tutorials/keynote speeches: Z-numbers, a new direction in the analysis of uncertain and complex systems; human space computing and cyber-physical systems; digital ecosystems, the resources for future humanity and society; massive data analytics for smart planet; social media for sustained digital ecosystems; new era of civilization, technology understand human and human; SAP co-innovation, envision the future, crossroots innovation; prediction markets, virtual currencies and social scores applied and technology innovation for networked life.
Start of the above-titled section of the conference proceedings record.
The eight papers in this special issue cover advances in spoken dialogue systems and mobile interface.
Stanley Kubrick's 1968 film 2001: A Space Odyssey famously featured HAL, a computer with the ability to hold lengthy conversations with his fellow space travelers. More than forty years later, we have advanced computer technology that Kubrick never imagined, but we do not have computers that talk and understand speech as HAL did. Is it a failure of our technology that we have not gotten much further than an automated voice that tells us to "say or press 1"? Or is there something fundamental in human language and speech that we do not yet understand deeply enough to be able to replicate in a computer? In The Voice in the Machine, Roberto Pieraccini examines six decades of work in science and technology to develop computers that can interact with humans using speech and the industry that has arisen around the quest for these technologies. He shows that although the computers today that understand speech may not have HAL's capacity for conversation, they have capabilities that make them usable in many applications today and are on a fast track of improvement and innovation. Pieraccini describes the evolution of speech recognition and speech understanding processes from waveform methods to artificial intelligence approaches to statistical learning and modeling of human speech based on a rigorous mathematical model--specifically, Hidden Markov Models (HMM). He details the development of dialog systems, the ability to produce speech, and the process of bringing talking machines to the market. Finally, he asks a question that only the future can answer: will we end up with HAL-like computers or something completely unexpected?
A lot. Since inception of Contender, a machine learning method tailored for computer-assisted decision making in industrial spoken dialog systems, it was rolled out in over 200 instances throughout our applications processing nearly 40 million calls. The net effect of this data-driven method is a significantly increased system performance gaining about 100,000 additional automated calls every month.
An examination of more than sixty years of successes and failures in developing technologies that allow computers to understand human spoken language. Stanley Kubrick's 1968 film 2001: A Space Odyssey famously featured HAL, a computer with the ability to hold lengthy conversations with his fellow space travelers. More than forty years later, we have advanced computer technology that Kubrick never imagined, but we do not have computers that talk and understand speech as HAL did. Is it a failure of our technology that we have not gotten much further than an automated voice that tells us to “say or press 1”? Or is there something fundamental in human language and speech that we do not yet understand deeply enough to be able to replicate in a computer? In The Voice in the Machine, Roberto Pieraccini examines six decades of work in science and technology to develop computers that can interact with humans using speech and the industry that has arisen around the quest for these technologies. He shows that although the computers today that understand speech may not have HAL's capacity for conversation, they have capabilities that make them usable in many applications today and are on a fast track of improvement and innovation. Pieraccini describes the evolution of speech recognition and speech understanding processes from waveform methods to artificial intelligence approaches to statistical learning and modeling of human speech based on a rigorous mathematical model—specifically, Hidden Markov Models (HMM). He details the development of dialog systems, the ability to produce speech, and the process of bringing talking machines to the market. Finally, he asks a question that only the future can answer: will we end up with HAL-like computers or something completely unexpected?
In a recent publication [1], we laid out the mathematical foundations for an optimization technique—Contender—applicable to commercially deployed spoken dialog systemssimilar to what the research community would refer to as a light version of reinforcement learning.
Signal Processing Society (SPS) and to the SP community. When working as an editorial board member and area editor under the leadership of my predecessor, Prof. Shih-Fu Chang, I witnessed the immeasurable vibrancy, invigorating energy, and unbounded intellectual landscape of our SP community. During 2007–2008, with Prof. Chang’s guidance, I initiated the effort in expanding the scope and technical fields of SP [2]. This led to substantial broadening of the article coverage in SPM along the two axes of “signal” and “processing” [3]. In the meantime, while helping Prof. Chang to solicit potential articles for SPM, I interacted with several pioneers in various technical areas pertinent to SP. These interactions provided me with the opportunity to learn, analyze, and appreciate a wide range of SP-enabled future wants and needs (e.g., [4]). In my own work environment within a major computer software company, SP methods and applications as defined in the expanded scope had also permeated every corner. Our community had clearly come to realize that while SP played an integral part in the technological development of television, telephone, communication, multimedia, space travel, and computers, more exciting challenges and opportunities would lie ahead for SP in broad areas such as intelligent communication; natural human-machine interface; universal language translation; biomolecular information processing; automated navigation; efficient generation/ distribution/consumption of “green” energy; intelligent sensor and human
Spoken language understanding (SLU) plays a key role in every spoken dialogue system. This chapter shows how different SLU techniques are integrated into dialogue systems, focusing on the significant differences between research and commercial implementations. After discussing about how to make this integration robust against speech recognition glitches, the chapter reviews example projects, architectures, and corpora associated with the application of SLU to spoken dialogue systems. The chapter provides an overview about both research and commercial aspects of SLU in dialogue systems. It reviews approaches, deals with deficiencies regarding speech input, discusses about example datasets and applications, shows how to measure performance, and finally draws the conclusion that research and commercial worlds, despite of major differences in their approaches to SLU and dialogue management, are slowly growing together. Controlled Vocabulary Terms interactive systems; learning (artificial intelligence); speech processing