The Natural Conversation Benchmark (NC-Bench) introduce a new approach to evaluating the general conversational competence of large language models (LLMs). Unlike prior benchmarks that focus on the content of model behavior, NC-Bench focuses on the form and structure of natural conversation. Grounded in the IBM Natural Conversation Framework (NCF), NC-Bench comprises three distinct sets. The Basic Conversation Competence set evaluates fundamental sequence management practices, such as answering inquiries, repairing responses, and closing conversational pairs. The RAG set applies the same sequence management patterns as the first set but incorporates retrieval-augmented generation (RAG). The Complex Request set extends the evaluation to complex requests involving more intricate sequence management patterns. Each benchmark tests a model's ability to produce contextually appropriate conversational actions in response to characteristic interaction patterns. Initial evaluations across 6 open-source models and 14 interaction patterns show that models perform well on basic answering tasks, struggle more with repair tasks (especially repeat), have mixed performance on closing sequences, and find complex multi-turn requests most challenging, with Qwen models excelling on the Basic set and Granite models on the RAG set and the Complex Request set. By operationalizing fundamental principles of human conversation, NC-Bench provides a lightweight, extensible, and theory-grounded framework for assessing and improving the conversational abilities of LLMs beyond topical or task-specific benchmarks.
With generative AI acquiring the right training data is a critical part of designing the user experience. Training large language models to talk like humans requires exposing them to the interaction patterns distinctive of natural conversation. Although models are typically fine-tuned on question-answer or instruction pairs, they are less often trained on real-time human conversations. Natural conversation data are hard to find and "conversation" is used to mean very different kinds of interaction or content. We demonstrate a method for scoring language content using generic conversational phrase detection. We generate three scores: 1) range of unique features, 2) density of features within sections of the content, and 3) overall score combining these. Using our method, we score over 27,000 documents from 6 datasets, which vary widely in terms of whether or not they contain conversation content. Our results show this approach is effective in distinguishing conversation content from non-conversation and from conversation-like content.
Generative AI has caused a paradigm shift in the area of Artificial Intelligence (AI) and as such has inspired much new research, especially on Large Language Models (LLMs). LLMs are transforming how people interact with computers in service-oriented fields in both the consumer (for example: retail, travel, education, healthcare) and enterprise (customer care, field service, sales, marketing, etc.) spaces. One barrier to widespread adoption is the current unpredictability of LLM behavior: users must trust that LLM-based services and systems are accurate, fair, and unbiased. Model responses that exhibit biases related to race, social status, and other sensitive topics can have serious consequences, ranging from lack of trust in the model to adverse social implications for consumers, all the way to damage to the reputations of the corporations that provide them. This study explores how to uncover biases related to social stigmas in LLM output, by using an adversarial prompt-based approach. Discovering model vulnerabilities of this type is a nontrivial task due to the large search space, making it resource-intensive. We present an evaluation framework for probing and analyzing the behaviors of multiple LLMs systematically. We use a curated set of adversarial prompts with a focus on uncovering biased responses to prompts associated with social attributes.
Troubles in speaking, hearing, and understanding occur routinely in any kind of conversational setting. The natural flow of conversation includes methods for “repairing” such troubles by repeating or paraphrasing all or parts of prior turns. In the case of conversational AI systems, these troubles occur due to failure of different components of the system such as the speech recognition, natural language understanding, and natural language generation. Such errors may occur infrequently, but still often enough to have a significant impact on key performance indicators (KPIs). Identifying the root cause of these errors is a complex task that requires a team to meticulously examine and interpret the interaction between the voice agent and customers. In this work, we present an interactive system, DTTool, that surfaces system-generated annotations that hint at anomalous events that lead to candidate errors that impact KPIs and demonstrate how the team could discover unknown errors using DTTool.
Although methods for repairing prior turns in natural conversation are critical for enabling mutual understanding, or successful communication, these methods are seldom built into conversational user interfaces systematically. Chatbots and voice assistants tend to ask users to paraphrase what they said if it was not understood, but users cannot do the same if they encounter trouble in understanding what the agent said. Understanding is a one-way street in most (intent-based) conversation-like interfaces. An exception to this is Moore and Arar (2019), who demonstrate nine types of user-initiated repair on agent responses that are common in natural conversation and who have shown that users will employ these repair features correctly in text-based interfaces if taught. In this small-scale study, we test these user-initiated repairs (in second position) in a voice-based interface. With understanding-oriented repairs, we found that participants employed them much the same way in text and voice. In addition, we examine some hearing- and speaking-oriented repairs that emerged from the use of our novel multi-modal interface. We found that participants used them to manage troubles specific to the voice modality. Analysis of user logs and transcripts suggests that user-initiated repair features are valuable components of conversational interfaces.
In business-to-business (B2B) commerce, the provider organization’s collective understanding of the client organization is distributed across the communications among a multitude of individuals. However, the substance of many of these communications typically goes unrecorded even when using standard Customer Relationship Management (CRM) tools and therefore their value remains largely untapped. In an attempt to capture these valuable client interaction data, we built an intuitive data capture tool that fits in with sellers’ client-facing workflows. But despite giving positive feedback on our prototype, our study participants informed us that the company was transitioning to a new CRM system, an industry standard. Therefore, we pivoted to creating visualizations for client insights. Based on extensive feedback from our participants, we designed four novel B2B interaction visualizations: 1) Interactions, 2) Timeline, 3) Topics and 4) Network. In this paper, we report on our design exploration for using CRM data to visualize client relationships in B2B context and discuss the challenges we encountered.
extended-abstract Share on A Special Interest Group on Developing Theories of Language Use in Interaction with Conversational User Interfaces Authors: Paola Raquel Peña School of Information and Communication Studies, University College Dublin, Ireland School of Information and Communication Studies, University College Dublin, Ireland 0000-0002-6371-8074View Profile , Philip R Doyle School of Information and Communication Studies, University College Dublin, Ireland School of Information and Communication Studies, University College Dublin, Ireland 0000-0002-2686-8962View Profile , Emily Yj Ip School of Computer Science and Statistics, Trinity College Dublin, Ireland School of Computer Science and Statistics, Trinity College Dublin, Ireland 0000-0001-6307-9666View Profile , Giovanni Di Liberto School of Computer Science and Statistics, Trinity College Dublin, Ireland School of Computer Science and Statistics, Trinity College Dublin, Ireland 0000-0002-7361-0980View Profile , Darragh Higgins School of Computer Science and Statistics, Trinity College, Dublin, Ireland School of Computer Science and Statistics, Trinity College, Dublin, Ireland 0000-0003-1008-0051View Profile , Rachel Mcdonnell School of Computer Science and Statistics, Trinity College Dublin, Ireland School of Computer Science and Statistics, Trinity College Dublin, Ireland 0000-0002-1957-2506View Profile , Holly Branigan Department of Psychology, University of Edinburgh, United Kingdom Department of Psychology, University of Edinburgh, United Kingdom 0000-0002-7845-8850View Profile , Joakim Gustafson Division of Speech, Music and Hearing, KTH Royal Institute of Technology, Sweden Division of Speech, Music and Hearing, KTH Royal Institute of Technology, Sweden 0000-0002-0397-6442View Profile , Donald Mcmillan Department of Computer and Systems Sciences, Stockholm University, Sweden Department of Computer and Systems Sciences, Stockholm University, Sweden 0000-0002-1392-5737View Profile , Robert J Moore IBM Research, United States IBM Research, United States 0000-0002-5636-9822View Profile , Benjamin R. Cowan School of Information and Communication Studies, University College Dublin, Ireland School of Information and Communication Studies, University College Dublin, Ireland 0000-0002-8595-8132View Profile Authors Info & Claims CHI EA '23: Extended Abstracts of the 2023 CHI Conference on Human Factors in Computing SystemsApril 2023Article No.: 509Pages 1–4https://doi.org/10.1145/3544549.3583179Published:19 April 2023Publication History 1citation101DownloadsMetricsTotal Citations1Total Downloads101Last 12 Months101Last 6 weeks101 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access
KEYWORDS: Conversational user interfacesconversational UX designConversation Analysischatbotsvoice assistants
Healthcare and wellbeing are two main interconnected application areas of conversational agents (CAs). There is a significant increase in research, development, and commercial implementations in this area. In parallel to the increasing interest, new challenges in designing and evaluating CAs have emerged. This study aims to identify key design, development, and evaluation challenges of CAs in healthcare and wellbeing research. The focus is on the very recent projects with their emerging challenges. A review study was conducted with 17 invited studies, most of which were presented at the ACM CHI2020 conference workshop on CAs for health and wellbeing. Eligibility criteria required the studies to involve a CA applied to a health or wellbeing project in an ongoing or recently finished project. The participating studies were asked to report on their projects' design and evaluation challenges. We used thematic analysis to review the studies. The findings include a range of topics from primary care to caring for older adults to health coaching. We identified four major themes: i) domain information and integration, ii) user-system interaction and partnership, iii) evaluation, and iv) conversational competence. While some challenges are shared with other CA application areas, safety and privacy remain the major challenges in the healthcare and wellbeing domains. An increased level of collaboration across different institutions and entities may be a promising direction to address some of the major challenges which otherwise would be too complex to be addressed by the projects with their limited scope and budget.
Background Health care and well-being are 2 main interconnected application areas of conversational agents (CAs). There is a significant increase in research, development, and commercial implementations in this area. In parallel to the increasing interest, new challenges in designing and evaluating CAs have emerged. Objective This study aims to identify key design, development, and evaluation challenges of CAs in health care and well-being research. The focus is on the very recent projects with their emerging challenges. Methods A review study was conducted with 17 invited studies, most of which were presented at the ACM (Association for Computing Machinery) CHI 2020 conference workshop on CAs for health and well-being. Eligibility criteria required the studies to involve a CA applied to a health or well-being project (ongoing or recently finished). The participating studies were asked to report on their projects’ design and evaluation challenges. We used thematic analysis to review the studies. Results The findings include a range of topics from primary care to caring for older adults to health coaching. We identified 4 major themes: (1) Domain Information and Integration, (2) User-System Interaction and Partnership, (3) Evaluation, and (4) Conversational Competence. Conclusions CAs proved their worth during the pandemic as health screening tools, and are expected to stay to further support various health care domains, especially personal health care. Growth in investment in CAs also shows the value as a personal assistant. Our study shows that while some challenges are shared with other CA application areas, safety and privacy remain the major challenges in the health care and well-being domains. An increased level of collaboration across different institutions and entities may be a promising direction to address some of the major challenges that otherwise would be too complex to be addressed by the projects with their limited scope and budget.
research-article Share on Special Issue on Conversational Agents for Healthcare and Wellbeing Authors: A. Baki Kocaballi The School of Computer Science, The University of Technology Sydney The School of Computer Science, The University of Technology SydneySearch about this author , Liliana Laranjo The Westmead Applied Research Centre (WARC), The University of Sydney The Westmead Applied Research Centre (WARC), The University of SydneySearch about this author , Leigh Clark The Department of Computer Science, Swansea University The Department of Computer Science, Swansea UniversitySearch about this author , Rafał Kocielnik The California Institute of Technology The California Institute of TechnologySearch about this author , Robert J. Moore IBM Research - Almaden IBM Research - AlmadenSearch about this author , Q. Vera Liao Microsoft Research Montréal Microsoft Research MontréalSearch about this author , Timothy Bickmore The Khoury College of Computer Sciences, Northeastern University The Khoury College of Computer Sciences, Northeastern UniversitySearch about this author Authors Info & Claims ACM Transactions on Interactive Intelligent SystemsVolume 12Issue 2June 2022 Article No.: 9pp 1–3https://doi.org/10.1145/3532860Published:12 July 2022Publication History 0citation255DownloadsMetricsTotal Citations0Total Downloads255Last 12 Months255Last 6 weeks54 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my Alerts New Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access
While interest in AI conversational agents has grown rapidly in recent years, creating agents capable of holding recipient-specific conversations remains a challenge. In other words, an automated agent should be able to engage in recipient design or tailoring the form of its utterances and selection of utterances based on the user's knowledge, or assumptions about what the user knows. Without recipient design, the agent formulates its utterances the same way for every user. While an agent might still fall back on conversational repair practices, such as paraphrase or definition requests, recipient design can minimize the necessity for repair in the first place. In this work, we leverage recipient design methods from natural conversation, as identified in the field of Conversation Analysis (CA) and implement three of them (self-reports of knowledge, prior difficulty in understanding, and prior exposure to a reference) as part of our conversational agent, the Alma Assistant.
In one general aspect, a computer-implemented method includes identifying current choices with different verbosity levels for a current turn in a conversation; normalizing multi-dimensional verbosity vectors for each of the current choices to obtain a normalized value for each of the current choices; determining a state definition for the current turn in the conversation, utilizing the normalized values for each of the current choices; providing the state definition for the current turn in the conversation and the normalized values for each of the current choices to a trained reinforcement learning module; receiving, from the trained reinforcement learning module, a score associated with each of the current choices for the current turn in the conversation; and selecting one of the current choices to be entered for the current turn in the conversation, based on the score associated with each of the current choices for the current turn in the conversation.
As CUIs become more prevalent in both academic research and the commercial market, it becomes more essential to design usable and adoptable CUIs. Though research on the usability and design of CUIs has been growing greatly over the past decade, we see that many usability issues are still prevalent in current conversational voice interfaces, from issues in feedback and visibility, to learnability, to error correction, and more. These issues still exist in the most current conversational interfaces in the commercial market, like the Google Assistant, Amazon Alexa, and Siri. The aim of this workshop therefore is to bring both academics and industry practitioners together to bridge the gaps of knowledge in regards to the tools, practices, and methods used in the design of CUIs. This workshop will bring together both the research performed by academics in the field, and the practical experience and needs from industry practitioners, in order to have deeper discussions about the resources that require more research and development, in order to build better and more usable CUIs.
Design systems for creating user experience reduce effort, scaffold learning and increase collaboration [2], yet they are currently available primarily for applications with visual interfaces. Design systems for conversational UX are still few and far between. Any design system should provide a distinctive philosophy, interaction patterns and content format. The Alma Design System, from IBM Research, provides these. Its design philosophy takes human conversation as its metaphor and adapts principles from conversation science, such as mutual understanding, recipient design, minimization and repair. Its pattern language provides 100 interaction patterns and components adapted from the literature of Conversation Analysis. It includes formal patterns for conversational activities and their management. And its content format consists of six slots and can be extended to any number of parts.