
Modern TTS systems are capable of creating highly realistic and natural-sounding speech, making them an important tool when designing virtual agents. While sounding highly realistic, the process of customizing such TTS voices remains a complex task, mostly requiring the expertise of specialists within the field. One reason for this is the utilization of deep learning models, which are characterized by their expansive, non-interpretable parameter spaces, restricting the feasibility of manual voice customization. In this paper, we present a novel human-in-the-loop paradigm based on an evolutionary algorithm for directly interacting with the parameter space of a neural TTS model. We integrated our approach into a user-friendly graphical user interface that allows users to efficiently create original voices. Those voices can then be used to equip virtual agents with highly customized TTS capabilities by using an open-source programming interface provided by us. Further, in a first pilot study, we show that VoiceX is an appropriate tool for creating individual, custom voices.
Evaluating the quality of generated nonverbal behavior remains a major challenge in the development of generative models. While human evaluations are reliable, they are costly and impractical for large-scale or iterative optimization. In this work, we propose an objective evaluation framework based on aggregated ranks across multiple fidelity and diversity metrics, computed from both raw features and learned latent representations. Compared to existing works, our framework emphasizes consistency across multiple metrics, aiming to provide a more holistic assessment.
AI companions (AICs; persistent digital personas driven by large language models) are increasingly popular, yet there is limited understanding of how non-experts think about the notion of companionship with machine agents. This descriptive work builds on recent scholarship by exploring topics salient to users as they discuss AI companionship in online forums. From a corpus of curated entries (N = 3,621) across 14 subreddits, LLM-assisted analysis induced four high-level themes associated with socioemotionality, control, limitations, and imaginaries. The topical hierarchy suggests this form of machine companionship both converges with and diverges from human relations.
This paper presents the creation and perceptual validation of the EVE corpus, a resource, in English and French, for speech emotion recognition of audio and audiovisual content and for the generation of verbal and non-verbal behaviour. For each language, ten actors performed 10 linguistically and semantically neutral sentences with different emotions (fear, anger, happiness, sadness, disgust, surprise, confidence, confusion, contempt, empathy) and a neutral condition. Each was expressed at two arousal levels, with two trials per level. The emotional content of the corpus was perceptually validated by 600 participants per language. The corpus is accessible through this link: https://doi.org/10.58119/ULG/VREIOB.
This paper presents an exploratory study of users’ first impressions of personality traits and preference for virtual receptionist agents with different body types and genders. The results show that gender and body type interact, since a female muscular agent was rated high on many traits while a male muscular agent was rated low. Another finding is that even though agents with large bodies were rated higher or as high on important traits, e.g. helpful and knowledgeable, the muscular female agent was still preferred over them. This implies that further research on body types in agent design is needed and may be important to support representation and norm critical design.
Mutual trust between humans and interactive artificial agents is crucial for effective human-agent teamwork. This involves not only the human appropriately trusting the artificial teammate, but also the artificial teammate assessing the human’s trustworthiness for different tasks (i.e., artificial trust in human partners). Literature indicated that transparency and explainability is generally beneficial for human-agent collaboration. However, communicating artificial trust potentially affects human trust and satisfaction, which impact team dynamics. Towards studying these effects, we developed an artificial trust model and implemented five distinct communication approaches which varied in modality (visual/graphical and/or text), level (communication and/or explanation), and timing (real-time or occasional). We evaluated the effects of the different communication styles through a user study (N=120) in a 2D grid-world Search and Rescue scenario. Our results show that all our artificial trust explanations improved human trust and satisfaction, but the mere graphical communication of it did not. These results are bound to the specific scenario and context in which this study was run and require further exploration. As such, this work presents a first step towards understanding the consequences of communicating and explaining to a human teammate their assessed trustworthiness.
Creating high-quality virtual agents has traditionally required a time-intensive authoring process. This demo presents a rapid development pipeline for building embodied virtual career coaches using LLMs and D-ID avatars. The system enables educators to deploy personalized agents without scripting dialogue, leveraging a retrieval-augmented knowledge base. We demonstrate how the system answers student questions with synchronized speech and facial animation, highlighting its implications for scalable, accessible coaching in education.
Negotiation is pivotal for conflict resolution in human-agent interactions, where emotional and behavioral dynamics can significantly shape the outcomes. However, many existing strategies prioritize time- or behavior-based tactics and overlook the dynamic role of emotional awareness. This paper presents the Solver Agent, which integrates real-time facial expression recognition into a hybrid strategy incorporating time- and behavior-based approaches. It is deployed on a humanoid robot with multimodal interaction capabilities (speech, gestures, facial expression analysis) to dynamically refine its bidding and concession strategies based on an opponent’s emotional cues and negotiation patterns. In user studies with 28 participants, the Solver Agent achieved higher agent scores, improved social welfare, and faster agreements than a baseline hybrid strategy without compromising participant satisfaction. Participants also viewed the Solver Agent as more attuned to their preferences and goals. These findings highlight that embodied emotion-aware negotiation can foster equitable and efficient collaboration, pointing to new opportunities in human-agent interaction research.
We present an interactive virtual agent that employs politeness strategies to suggest compromises during a joint decision-making task with a human partner. In our approach, politeness is expressed not only through language, but also through the agent’s willingness to compromise by proposing options that better align with the human partner’s preferences. We conducted a within-participant experiment (n = 97) in which users interacted with two types of agents: a performance-optimizing agent and a socially considerate agent. The former prioritized decision quality, while the latter occasionally made socially motivated quality concessions to the user as a form of face-work to defer to the human partner. Our results show that participants reported higher satisfaction with the final decision and consistent intention to collaborate with the socially considerate agent on both tasks. However, the social agent was less effective than the performance-optimizing agent at improving the users’ initial decisions. These findings highlight trade-offs in human-agent interaction between optimizing task performance and managing the social dynamics of collaboration. We conclude by discussing the ethical and cultural considerations that may be involved in navigating this trade-off.
This research explores a new opportunity for Video-mediated Relay Interpreting (VRI) for Deaf/Hard of Hearing (DHH) users and sign language interpreters, namely the use of Augmented Reality (AR) glasses in VRI. The perspectives of DHH users and interpreters are investigated using the User Acceptance Model (TAM) and the extended, more holistic version, the Unified Theory of Acceptance and Use of Technology (UTAUT) in combination with on-site user testing and evaluation. Results show that interpreters preferred AR VRI over VRI. While DHH users recognized the potential of future advancements, they considered current technology inadequate. The results of interactive unscripted testing in this study contrasted outcomes from previously scripted controlled experiments. Suggestions for future research include examination of tailored hardware for DHH users, studying specific use cases, and investigating the ethical concerns of this technology.
Healing Agent (HA) is a speculative Artificial Intelligence (AI)-powered interactive Virtual Companion (VC) designed to support individuals coping with the loss of a loved one. The system enables users to recreate a photorealistic avatar of the deceased by integrating photographs, voice recordings, and video clips to replicate appearance, speech patterns, personality traits, and communication style. Users can input shared memories via voice or text. Grounded in established understandings of bereavement, this intervention aims to offer emotional comfort without disrupting the grieving process. The design prioritizes privacy, ensuring all personal data remains locally stored with full user control over data retention and access. While potentially beneficial for emotional regulation and meaning-making, concerns about emotional dependency, avoidance, and social withdrawal are discussed through psychological and ethical considerations. This proposal explores whether such AI-mediated remembrance can aid or complicate human grief trajectories in the short and long term.
The exchange of affective information lies at the core of social interactions and allows inference of the intentions and states of an interactive partner. In experimental settings, this exchange is often limited to nonverbal expressions, resulting in a reduced form of complex social behavior. The implementation of embodied conversational agents in Virtual Reality promises highly naturalistic and interactive verbal communication with AI-controlled agents. However, the degree to which affective information can be generated in such paradigms remains unclear. The goal of the present study is to evaluate a behavioral paradigm in which an embodied conversational agent provides affective information in a conversation and to test whether target emotions can be induced and maintained in real-time interactions. Data from 48 human-AI interactions demonstrate that AI-controlled agents were able to generate specific affective speech content for different conversational contexts (happy, angry, and sad events). Analysis within conversations shows that context-specific emotions were most frequent at the beginning of a conversation and then decreased with turns. Overall, the results suggest that AI-controlled embodied conversational agents are a promising way to include affective information in the simulation of naturalistic conversations.
This paper investigates the requirements for intelligible high-quality motion capture of sign language, with a particular focus on face animation. There are several methods for capturing facial animation, yet there is no established standard for which type of facial animation is the most optimal for sign languages. We compare two facial animation approaches — ARKit and Metahuman Animator (MHA) in terms of intelligibility. As expected, a human evaluation study with deaf Swedish Sign Language signers showed that MHA outperforms ARKit because MHA has a higher number of facial controls and an advanced depth data solver. In addition, we introduce a biLSTM-based occlusion infilling technique for MHA data as opposed to linear interpolation. Although biLSTM infilled MHA animations showed no improvement in human evaluation compared to original MHA animations, the model produces linguistically reasonable infillings upon qualitative exploration of the renders.
Intelligent virtual agents (IVAs) are increasingly designed with proactive behavior such as initiating context-aware interactions and anticipating user needs. While previous studies have examined agents across many contexts, there remains a lack of ecologically valid paradigms that capture the effects of IVAs’ proactive behaviors; this gap is particularly evident for behaviors that build a relational bond between the user and the agent, commonly known as rapport. This position paper proposes a VR-based methodological framework that leverages physiological and behavioural measurements to study user interactions with LLM-driven IVAs over time. To illustrate how this framework supports both implementation and evaluation of proactive behaviors, we present a case study focused on IVAs’ rapport-building behaviors in the context of health counseling. This VR-based approach offers researchers a flexible and multimodal experimental platform that combines high realism and exquisite measurement opportunities to examine social and cognitive processes underlying human-IVA interactions.
This paper introduces a preprocessing pipeline for keypoints extracted using MediaPipe, aiming to improve pose annotation consistency in sign language datasets. We evaluate its effectiveness using a sign similarity task based on phonological features, without relying on gloss annotations. Similarity is measured using Dynamic Time Warping (DTW) across videos from sign language dictionaries. Although such similarity analyses can support various sign language processing applications - such as lexical search, clustering, and data enrichment - the main contribution of this work is to standardise pose features across heterogeneous sources, including different signers and backgrounds. Experiments on two dictionary datasets show that our pipeline significantly improves similarity measurements, with promising benefits for other sign language processing tasks.
Computer-assisted Sign Language (SL) systems offer a promising solution for real-time communication and education within Deaf and Hard-of-Hearing (DHH) communities. In this paper we tackle the production of Sign Language videos from text sentences by proposing a novel two-way transformer-based translation system. We use extended 2D skeletal poses from MediaPipe as an intermediate representation step. Following our Sign Language Production (SLP) pipeline, we generate a synthetic, highly-detailed SL output, that mimics signers from the original dataset, by performing neural rendering on the transformer-generated skeletal poses. This enables the creation of signer-specific SLP systems from a relatively limited training set. We evaluate on a large-scale Greek Sign Language Dataset and conduct an extended user evaluation study, proving our pipeline’s effectiveness across different signers and a broad Greek vocabulary.
Sign languages are the primary means of communication for deaf communities, but the development of effective automatic recognition systems remains a significant challenge. In this work, we focus on the task of Isolated Sign Language Recognition (ISLR) using a multimodal approach grounded in a Large Language Model (LLM) architecture. We merge modalities, including visual characteristics into the linguistic representation space of LLMs, and perform ablation studies to evaluate the individual contributions of each visual modality to the recognition performance. Experiments are conducted on the AVASAG100 dataset, where our method achieves a weighted F1-score (W-F1) of 70.36 +/- 3.00 and a macro F1-score (M-F1) of 62.34 +/- 3.18 projecting landmarks extracted from the pose into the LLM's emebdding-space. These results underscore the value of multimodal integration in ISLR and provide guidelines for future research directions.
Augmented Reality (AR) technologies are increasingly integrated into exposure therapy and related interventions for anxiety disorders. As the field rapidly expands, a diverse body of research has emerged, spanning clinical trials, technology development, and application in real-world and therapeutic contexts. However, a comprehensive synthesis of research patterns and thematic directions within this literature is lacking. This systematic review aims to map and characterize the general research trends, thematic clusters, and methodological patterns of AR-based interventions for anxiety and related disorders. Rather than focusing on isolated mechanisms or specific clinical factors, the review provides an overarching view of how the field has evolved, the principal areas of focus, and the gaps that remain. The review identified studies investigating AR exposure therapy and related AR interventions for anxiety. Included articles were analyzed for publication trends, targeted disorders, clinical and technical approaches, and thematic content. Unsupervised clustering and qualitative synthesis were used to group studies by major research themes. The review reveals a substantial increase in AR exposure therapy research over the past decade. Four primary thematic clusters were identified: (1) AR for stress management and real-world applications; (2) integration of AR within cognitive behavioral therapy and broader anxiety treatment; (3) AR exposure therapy for specific phobias, often supported by clinical validation; and (4) the use of virtual avatars and social interaction within AR environments, reflecting technical innovation and new therapeutic paradigms. The literature is characterized by rapid growth, increasing interdisciplinary collaboration, and a shift toward more rigorous empirical designs, though significant gaps remain in sample diversity, standardized reporting, and long-term evaluation.
The integration of emotional intelligence into virtual health assistance systems has the potential to revolutionize the way users interact with healthcare applications. By adapting to users’ emotional states, virtual characters or avatars can provide more personalized and engaging interactions, leading to improved user satisfaction and overall well-being. This study presents the development of photorealistic virtual assistants using Unreal Engine’s Metahuman Creator, which can adapt their speech and facial expressions to match the user’s emotional state. The system, which synchronizes Metahuman avatars with an AI-powered chatbot, enables the display of emotions such as happiness, excitement, boredom, and sadness, creating a more natural and engaging conversation. A user study with 21 participants evaluated the system’s effectiveness, revealing that emotion-adaptive avatars notably enhance experience quality and emotional engagement compared to neutral virtual assistants. The findings of this study have significant implications for the development of more effective and user-friendly healthcare applications and highlight the need for further research into the impact of emotion-adaptive virtual assistants on user engagement, well-being, and adherence to health recommendations.
Virtual agents have been studied for application in various fields, especially in healthcare. Since the emergence of large language models, virtual agents are equipped with these models to enable more natural interaction. However, these models often show overconfidence in decision-making, which can lead to either over- or under-trust and subsequently affect user acceptance. This proposal aims to shape appropriate trust through self-calibration. Self-calibration involves the process of estimating model uncertainty and communicating uncertainty. A role-play method was used to gain insights from human-to-human interaction in uncertain situations. Potential factors (e.g., task type, user preferences, etc.) were investigated to identify correlations within the decision-making process and to enable appropriate interaction, whether reactive or proactive. By developing a well-calibrated system, our goal is to build a more transparent and reliable system.