Emotional head motion plays an important role in human-computer interaction (HCI), which is one of the important factors to improve users' experience in HCI. However, it is still not clear how head motions are influenced by speech features in different emotion states. In this study, we aim to construct a bimodal mapping model from speech to head motions, and try to discover what kinds of prosodic and linguistic features have the most significant influence on emotional head motions. A two-layer clustering schema is introduced to obtain reliable clusters from head motion parameters. With these clusters, an emotion related speech to head gesture mapping model is constructed by a Classification and Regression Tree (CART). Based on the statistic results of CART, a systematical statistic map of the relationship between speech features (including prosodic and linguistic features) and head gestures is presented. The map reveals the features which have the most significant influence on head motions in long or short utterances. We also make an analysis on how linguistic features contribute to different emotional expressions. The discussions in this work provide important references for realistic animation of speech driven talking-head or avatar.
In this study, we create a 3D interactive virtual character based on multi-modal emotional recognition and rule based emotional synthesize techniques. This agent estimates users' emotional state by combining the information from the audio and facial expression with CART and boosting. For the output module of the agent, the voice is generated by TTS (Text-to-Speech)system by freely given text. The synchronous visual information of agent, including facial expression, head motion, gesture and body animation, are generated by multi-modal mapping from motion capture database. A kind of high level behavior markerup language(hBML) which contains five keywords is used to drive the animation of virtual agent for emotional expression. Experiments show that this virtual character is considered natural and realistic in multimodal interaction environments.
This paper introduces a multimodal framework of generating a 3D human-like talking agent which can communicate with user through speech, lip movement, head motion, facial expression and body animation. In this framework, lip movements are obtained by searching and matching acoustic features which are represented by Mel-frequency cepstral coefficients (MFCC) in audio-visual bimodal database. Head motion is synthesized by visual prosody which maps textual prosodic features into rotational and translational parameters. Facial expression and body animation are generated by transferring motion data to skeleton. A simplified high level Multimodal Marker Language (MML), in which only a few fields are used to coordinate the agent channels, is introduced to drive the agent. The experiments validate the effectiveness of the proposed multimodal framework.
引言 自计算机问世以来,人们就梦想着有朝一日能与计算机进行自然的对话,便捷地获取计算机提供的各种信息和服务.为此,国内外学术界对自然人机对话的研究给予了很大关注和投入,如美国国防高级研究计划局计划、欧盟框架计划、日本学术振兴会计划以及我国的863计划.863计划和国家自然科学基金项目长期以来都十分关注人机对话研究,均设立了相关的研究项目.
During tracking process in optical flow, some points are normally easy to be lost or become outliers, due to the fact that the object undergoes changes of illumination or becomes partially occluded. The paper presents a novel scheme, Heterogeneity Elimination Individually (HEI), for outlier rejections from the tracking results of optical flow. HEI determines the most poisonous element by the distance between the tracked point and the projected point which is determined by its complements set from the source to the target one. Then HEI eliminates the farthest point every time from the element heterogeneous sequence until all remaining points are within the error tolerance and ensures the accuracy of the optical flow tracking results. Finally, we make an experiment by comparing our method with a popularly used method, RANSAC. The comparison results show the validity and effectiveness of our proposed method for outlier rejection.
This paper creates a Chinese interactive virtual character based on multi-modal mapping and rules, which receives information from the input modules and generates audio and visual speech, face expressions and body animations. The audio and visual speech are synthesized from the input text by multi-modal mapping, while face expressions and body movements are rule-based driven by emotion states. All of the original animations are captured by a motion capture system and plotted into a character model, which is created by the 3D creation software. We use a skeletal open source animation engine to create the scene in which the virtual character can talk like human communicating with users. The whole expression of this virtual character is considered very natural and realistic.
Emotion expressions sometimes are mixed with the utterance expression in spontaneous face-to-face communication, which makes difficulties for emotion recognition. This article introduces the methods of reducing the utterance influences in visual parameters for the audio-visual-based emotion recognition. The audio and visual channels are first combined under a Multistream Hidden Markov Model (MHMM). Then, the utterance reduction is finished by finding the residual between the real visual parameters and the outputs of the utterance related visual parameters. This article introduces the Fused Hidden Markov Model Inversion method which is trained in the neutral expressed audio-visual corpus to solve the problem. To reduce the computing complexity the inversion model is further simplified to a Gaussian Mixture Model (GMM) mapping. Compared with traditional bimodal emotion recognition methods (e.g., SVM, CART, Boosting), the utterance reduction method can give better results of emotion recognition. The experiments also show the effectiveness of our emotion recognition system when it was used in a live environment.
In this paper, we present a feature-based cartoon system, which can generate a realistic cartoon face form an input picture. The cartoon face can be multi-style by using different cartoon template. The whole process requires little user interaction. The system consists of four components, ASM-based facial feature locating, cartoon texture mapping, ears adding and hair adding. The advantages of method we propose are easy to accomplish, low time complexity and the good ability of retaining the real facial feature.
Speech-driven lip synchronization, an important part of facial animation, is to animate a face model to render lip movements that are synchronized with the acoustic speech signal. It has many applications in human-computer interaction. In this paper, we present a framework that systematically addresses multimodal database collection and processing and real-time speech-driven lip synchronization using collaborative filtering which is a data-driven approach used by many online retailers to recommend products. Mel-frequency cepstral coefficients (MFCCs) with their delta and acceleration coefficients and Facial Animation Parameters (FAPs) supported by MPEG-4 for the visual representation of speech are utilized as acoustic features and animation parameters respectively. The proposed system is speaker independent and real-time capable. The subjective experiments show that the proposed approach generates a natural facial animation.
Emotion recognition has been one of the most important issues in human computer interaction (HCI). In this paper, we propose a novel bimodal emotion recognition approach by using the boosting-based framework, in which we can automatically determine the adaptive weights for audio and visual features. In this way, we balance the dominances of audio and visual features dynamically in feature-level to obtain better performance. To ensure the tracking accuracy of facial feature points, the traditional KLT algorithm is integrated with Point Distribution Model (PDM) to guide the deformation of facial features. Experiments show the validity and effectiveness of our method.
Natural head motion is an indispensable part of realistic facial animation. This paper presents a novel approach to synthesize natural head motion automatically based on grammatical and prosodic features, which are extracted by the text analysis part of a Chinese Text-to-Speech (TTS) system. A two-layer clustering method is proposed to determine elementary head motion patterns from a multimodal database which covers six emotional states. The mapping problem between textual information and elementary head motion patterns is modeled by Classification and Regression Trees (CART). With the emotional state specified by users, results from text analysis are utilized to drive corresponding CART model to create emotional head motion sequence. Then, the generated sequence is interpolated by spineand us ed to drive a Chinese text-driven avatar. The comparison experiment indicates that this approach provides a better head motion and an engaging human-computer comparing to random or none head motion.
Local field potentials (LFPs) arise from dendritic currents that are summed by the brain tissue's impedance. By assuming that the rhythms existing in the LFPs result from the coordinated neural activity of sparse and transient neural assemblies transformed by the neural tissue, we propose to recover these neural assemblies sources using an independent component analysis on segments of a single LFP channel. The corresponding source signals and the set of temporal filters that operate on them constitute an efficient time-frequency decomposition of the LFP. This decomposition has the potential to identify sources that are more statistically dependent with stimuli or single-cell activity than the raw signal. In this work we show preliminary results on a synthetic dataset and a real dataset recorded from a rats nucleus accumbens during a reward administering experiment. When compared with the standard time-frequency analysis, this computational model for LFP analysis is totally data-driven because the filters , which form the basis for the decomposition, are estimated directly from the data.
Nikolay Chumerin合作论文数K.U.Leuven1