Social anxiety disorder has a considerable negative impact on educational attainment, employment status, and relationship development. Exposure therapy is an established method used to treat social anxiety by getting patients used to talking to other people. However, this procedure requires patients to have conversations with strangers, which may impose a psychological burden. This study aimed to examine whether the burden could be reduced during a first-time face-to-face interaction by using a human digital twin avatar technology that we call "Another Me." We conducted an exploratory, randomized, open-label, controlled study with 30 young participants who self-reported nervousness during conversations with strangers. Participants were divided into an intervention group and a control group. The intervention group (n = 15) watched a generated video of a conversation between their Another Me and the Another Me of an assigned interlocutor. After viewing the video, they engaged in an online conversation with the assigned interlocutor. The control group (n = 15) watched a video of a conversation between the avatars of two strangers and then engaged in a conversation with an assigned interlocutor. Results showed that the primary endpoint-the between-group difference in heart rate change from before to after the online conversation-was not statistically significant. However, exploratory secondary analyses suggested that heart rate decreased in the intervention group after viewing the pre-dialogue simulation video and remained relatively stable until the completion of the online conversation. In contrast, the control group exhibited an increase in heart rate before the online conversation. Additionally, the intervention group reported lower anxiety levels and a greater willingness to converse until the end of the face-to-face conversation. These findings should be interpreted as preliminary and hypothesis-generating. Further confirmatory trials with rigorous physiological monitoring are required.
This paper introduces the Multi-Modal Speed Dating (MMSD) dataset, a large-scale corpus designed to support research on romantic impression formation and interpersonal interaction. MMSD contains 1,250 dyadic interactions between 147 Japanese-speaking men and women, along with synchronized multimodal recordings (audio, video, transcripts), detailed profile data, and responses to 33 psychometric scales. Participants engaged in a structured series of speed dates mimicking real-world dating scenarios and rated their impressions—including love and like—at four time points. Final preferences and mutual contact decisions were also recorded. The dataset is suitable for a wide range of research applications, from predictive modeling of romantic impressions based on pre-date attributes to exploratory analysis of verbal and nonverbal behaviors that contribute to romantic outcomes. Prior studies using MMSD have shown that features such as psychological profiles, facial traits, and interaction behaviors are effective for predicting romantic attraction, and that linguistic behavior is particularly predictive of final outcomes. We detail the construction of the dataset, summarize findings from previous analyses, and outline future research opportunities enabled by its richness—including time-series modeling, personalized prediction, and cross-cultural comparisons. MMSD represents a novel benchmark for interdisciplinary research across psychology, affective computing, and human-centered AI, offering deeper insights into the dynamics of romantic connection.
Graphic recording (GR) visually externalizes dialogue content to support understanding and sharing in meetings and workshops. While recent approaches use large language models (LLMs) to automatically generate GRs, many rely on end-to-end generation and fail to explicitly address semantic alignment between dialogue content and visual representations, often leading to breakdowns in topic structure and role expression. This study argues that the core challenge of GR generation lies not in model expressiveness, but in the lack of explicit treatment of the judgments required to understand dialogue and map it onto appropriate visual structures. Based on an analysis of expert graphic recorders’ practices, we formalize these judgments as an intent structure and propose a stepwise GR generation process using it as an intermediate representation. Qualitative analysis suggests that this design helps maintain more stable semantic alignment between dialogue content and visual representations, offering broader implications for dialogue visualization and meaning-bearing generation tasks.
Large language models (LLMs) can predict interpersonal attraction from conversation transcripts, but it remains unclear what a speech predictor can add beyond transcript-only LLM prediction. Using Japanese speed-dating conversations, we combine predictions from a transcript-only LLM and a supervised speech predictor to estimate participants' reported liking of their partners. We show that speech can complement transcript-only LLM prediction, but that this complementarity is conditional rather than universal. Combining the two predictions significantly improves pairwise ranking accuracy over the transcript-only LLM alone in all evaluated conditions. By contrast, gains in per-participant Pearson $r$ vary across conversation rounds and rating directions, with none significant after correction. Retrospectively, these $r$ gains are concentrated among participants for whom the speech predictor is more accurate. Speech can therefore retain predictive value even when an LLM predicts attraction from transcripts. The relevant question is not simply whether speech helps, but where its complementarity emerges.
This paper focuses on simulating text dialogues in which impressions between speakers improve during speed dating. This simulation involves selecting an utterance from multiple candidates generated by a text generation model that replicates a specific speaker's utterances, aiming to improve the impression of the speaker. Accurately selecting an utterance that improves the impression is crucial for the simulation. We believe that whether an utterance improves a dialogue partner's impression of the speaker may depend on the personalities of both parties. However, recent methods for utterance selection do not consider the impression per utterance or the personalities. To address this, we propose a method that predicts whether an utterance improves a partner's impression of the speaker, considering the personalities. The evaluation results showed that personalities are useful in predicting impression changes per utterance. Furthermore, we conducted a human evaluation of simulated dialogues using our method. The results showed that it could simulate dialogues more favorably received than those selected without considering personalities.
We present a novel system that automatically generates and visualizes 3DCG dance animations based on the user’s preferred music and dance composition. The key technology of the system is a transformer-based diffusion model that produces dance choreographies conditioned on arbitrary inputs of music audio and dance composition. Integrated into a user-friendly GUI, the system allows users to instantly generate and preview multiple dance sequences simply by selecting their desired music and dance composition. This capability supports both creative choreography ideation and effective dance practice.
Advancements in science and technology have enabled communication beyond the constraints of time, distance, and physical limitations. In particular, the metaverse has facilitated interactions through avatars, and increasing attention is being given to "digital twins," which are nearly identical digital replicas of real-world entities. Digital twins allow for simulations, future predictions, and iterative analyses that would otherwise be costly in the real world. Among these, "human digital twins" are particularly promising, with potential applications in personalized treatments based on individual characteristics, medical history, and real-time physiological data, as well as behavioral simulations. This study examines the potential application of human digital twin technology in interpersonal relationship-building. Establishing positive human relationships is crucial for well-being, contributing to reduced feelings of isolation and an improved quality of life. In first-time interactions, mutual self-disclosure is typically a key process in deepening relationships. Human digital twins, which can communicate through both verbal and non-verbal cues, may help reduce psychological barriers in initial conversations. In this research, we investigated whether watching a conversation between digital twins before an initial meeting influences relationship formation, using a Speed Dating scenario. The experimental results suggest that viewing digital twin interactions promotes psychological readiness before face-to-face meetings and helps establish a foundation for familiarity with the other person. This study demonstrates the potential of digital twin technology in facilitating interpersonal relationship-building and opens new possibilities for digital communication.
In brief interactions like speed dating, a listener’s active listening behaviors are crucial for successful dialogue and positive impression formation. Quantitatively analyzing active listening behaviors in short first-time interactions, with their complex multimodal nature, has remained challenging. To address this, this study divided the listener’s active listening behaviors in Speed dating into three dimensions as defined by existing research, and analyzed each dimension based on verbal and nonverbal behaviors. We annotated their manifestation intensity in the Multi-Modal Speed Dating corpus. Subsequently, we developed models to predict each dimension’s intensity using visual, acoustic, and linguistic features. Experimental results showed that models incorporating acoustic features yielded the highest prediction performance. This suggests acoustic features are particularly vital in perceiving active listening behaviors in short first-time interactions. Our findings provide an analysis of the multimodal features of human-perceived active listening behaviors in short first-time interactions based on empirical data, contributing to the automated evaluation of such behaviors.
In this study, we examine how incorporating personality traits into a nonverbal behavior generation model for upper-body motion (head, arms, and posture) affects the quality and characteristics of the generated behaviors. We first constructed a multimodal dialogue corpus containing speech audio, transcripts, 3D upper-body skeleton data, and participants' Big Five personality scores, and then used the corpus to develop a model that predicts 3D skeleton coordinates from speech, text, and personality traits. Objective evaluation showed that the model with personality input more accurately reproduced individualized behaviors aligned with personality traits. The generated gestures also reflected the relationship between gesture expressivity and personality. Subjective evaluation further showed that observers could reliably perceive intended differences in personality levels-specifically, high vs. low Big Five scores-based only on the generated movements. These findings demonstrate that modeling personality traits enables the generation of agent behaviors that are both personality-consistent and perceptible to users.
Choosing a life partner is a significant decision requiring understanding preferences and compatibility. Speed dating provides a structured setting for exploring these aspects through direct interactions within limited time. Understanding specific behaviors during speed dating that influence outcomes could benefit participants and service providers significantly. However, the precise timing and behavior patterns affecting speed dating outcomes remain underexplored. This paper reports exploratory experiments analyzing multimodal features over time using machine learning to uncover when and which behaviors impact speed dating success. The results suggest multimodal information from 2-4 and 9-10 minute segments predicts outcomes as accurately as data from the entire session. Additionally, linguistic features play a particularly significant role among multimodal features.
This study proposes methods for identifying participant roles for the Task area and Socio-Emotional area in four-person online group discussions using language, audio, and facial features extracted through deep learning-based pretrained models. Two modeling approaches were examined: a BiLSTM-based model and a Transformer-based model incorporating a crossperson attention mechanism. Both models leverage behavioral information from all participants to capture interpersonal dynamics. While both models demonstrated the ability to estimate the two most frequent role labels in both the Task and Socio-Emotional areas with high F1 scores, the BiLSTM model consistently outperformed the Transformer-based model in overall performance. An ablation study revealed a performance degradation of the cross-person attention model in multimodal settings. This finding indicates that a further investigation is warranted to determine how crossperson attention mechanisms can be more effectively applied. Future work will therefore focus on enhancing model performance through these improvements.
Automatic rapport estimation in social interactions is a central component of affective computing. Recent reports have shown that the estimation performance of rapport in initial interactions can be improved by using the participant's personality traits as the model's input. In this study, we investigate whether this findings applies to interactions between friends by developing rapport estimation models that utilize nonverbal cues (audio and facial expressions) as inputs. Our experimental results show that adding Big Five features (BFFs) to nonverbal features can improve the estimation performance of self-reported rapport in dyadic interactions between friends. Next, we demystify how BFFs improve the estimation performance of rapport through a comparative analysis between models with and without BFFs. We decompose rapport ratings into perceiver effects (people's tendency to rate other people), target effects (people's tendency to be rated by other people), and relationship effects (people's unique ratings for a specific person) using the social relations model. We then analyze the extent to which BFFs contribute to capturing each effect. Our analysis demonstrates that the perceiver's and the target's BFFs lead estimation models to capture the perceiver and the target effects, respectively. Furthermore, our experimental results indicate that the combinations of facial expression features and BFFs achieve best estimation performances not only in estimating rapport ratings, but also in estimating three effects. Our study is the first step toward understanding why personality-aware estimation models of interpersonal perception accomplish high estimation performance.
Significant research attention has recently been focused on the automatic generation of human dance choreography from music. While several generation models have been proposed, they cannot specify what kind of movements to generate, and as a result, random movements are generated. We therefore propose a generation model called Composition-specified Dance Choreography Generation from Music (CDCGM) that enables creators to specify which dance composition (i.e., type of movement) to take when generating a dance at each time step. We implemented CDCGM by first constructing a new dataset that includes motion captures of breakdancing and time-series annotation data of representative movement types. Evaluation experiments using our corpus showed that CDCGM can generate dances that faithfully reflect the specified dance composition with high quality. Compared to conventional state-of-the-art models, CDCGM is capable of generating quality dances that improve the expressiveness and the degree to which the dance matches the content and timing of the music. We also propose a new application for CDCGM in which users watch newly generated dance choreography simply by entering music and dance composition. The results of a user study evaluation of the application demonstrated that users found the experience of generating dance by specifying any dance composition for any music extremely fun, that it has the potential to greatly contribute to dance choreography and learning, and that there is a strong desire to use this application on a daily basis.
Praising behavior is an important part of human communication. However, people who are unfamiliar with often praising have difficulty improving their praising skills. To solve this problem, we aim to construct a system for evaluating praising skills. So far, we have attempted to predict the degree of praising skills from verbal and nonverbal behaviors. However, our previous studies were focused on scenes in which the praiser was actually praising, and we have not dealt with scenes in which the praiser was not praising. In this paper, we attempt to detect whether the praiser is actually praising the receiver by including scenes in the study in which the praiser is not praising the receiver. First, we extract features related to the verbal and nonverbal behaviors of the praiser and receiver. Second, we construct machine learning models that utilize these features to detect whether or not the praiser is actually praising the receiver. Our results show that the machine learning model utilizing the acoustic and embedding-based linguistic behaviors of the praiser and the visual and acoustic behaviors of the receiver has the best detection performance.
We propose an advanced new system that automatically generates dances according to the user’s favorite music and dance style and then presents 3DCG animations of the generated dances on a naked-eye 3D display so that the user can have the experience of actually dancing with the displayed dancer. For automatic dance generation, we developed a new generation model based on the transformer-based diffusion model that generates a dance conditioned on any music audio and dance style of the user’s choice. We implemented this generation model in a GUI system that automatically generates dance animations according to the users’ specifications. The CG animations are then presented on a naked-eye 3D display system we recently proposed.
This paper addresses user-specific dialogs. In contrast to previous research on personalized dialogue focused on achieving virtual user dialogue as defined by persona descriptions, user-specific dialogue aims to reproduce real-user dialogue beyond persona-based dialogue. Fine-tuning using the target user's dialogue history is an efficient learning method for a user-specific model. However, it is prone to overfitting and model destruction due to the small amount of data. Therefore, we propose a learning method for user-specific models by combining parameter-efficient fine-tuning with a pre-trained dialogue model that includes user profiles. Parameter-efficient fine-tuning adds a small number of parameters to the entire model, so even small amounts of training data can be trained efficiently and are robust to model destruction. In addition, the pre-trained model, which is learned by adding simple prompts for automatically inferred user profiles, can generate speech with enhanced knowledge of the user's profile, even when there is little training data during fine-tuning. In experiments, we compared the proposed model with large-language-model utterance generation using prompts containing users' personal information. Experiments reproducing real users' utterances revealed that the proposed model can generate utterances with higher reproducibility than the compared methods, even with a small model.
Emotion recognition plays a crucial role in computer science, particularly in enhancing human-computer interactions. The process of emotion labeling remains time-consuming and costly, thereby impeding efficient dataset creation. Recently, large language models (LLMs) have demonstrated adaptability across a variety of tasks without requiring task-specific training. This indicates the potential of LLMs to recognize emotions even with fewer emotion labels. Therefore, we assessed the performance of an LLM in emotion recognition using two established datasets: MELD and IEMOCAP. Our findings reveal that for emotion labels with few training samples, the performance of the LLM approaches or even exceeds that of SPCL, a leading model specializing in text-based emotion recognition. In addition, inspired by the Chain of Thought, we incorporated a multi-step prompting technique into the LLM to further enhance its discriminative capacity between emotion labels. The results underscore the potential of LLMs to reduce the time and costs of emotion data labeling.