The goal of this work is to develop a statistical model of the lip movements for a specific speaker. The system is based on a Markov model trained on the visual part of the speech production. Such a model allows for an automatic generation of mouth shapes having fairly good natural looking dynamics. Applications might be both in image analysis and in image synthesis
Speech-driven facial animation combines techniques from different disciplines such as image analysis, computer graphics, and speech analysis. Active shape models (ASM) used in image analysis are excellent tools for characterizing lip contour shapes and approximating their motion in image sequences. By controlling the coefficients for an ASM, such a model can also be used for animation. We design a mapping of the articulatory parameters used in phonetics into ASM coefficients that control nonrigid lip motion. The mapping is designed to minimize the approximation error when articulatory parameters measured on training lip contours are taken as input to synthesize the training lip movements. Since articulatory parameters can also be estimated from speech, the proposed technique can form an important component of a speech-driven facial animation system.
Advances in joint acoustical/visual analysis for model-based lip motion synthesis is presented. The 2D lip motion field is modeled as a linear combination of a low dimensional motion basis computed through principal component analysis (PCA). The vector of PCA coefficients is expressed as a function of a limited set of articulatory parameters which describe the external appearance of the mouth. The acoustical processing estimates these articulatory parameters from the direct analysis of the speech waveform based on a neural processing stage, i.e., through a bank of time delay neural networks. The achieved results have been subjectively evaluated by visualizing the estimated motion on a wire-frame mouth template presented in synchronization with speech. The experiments carried out so far deal with single-speaker trained TDNNs and with single-speaker PCA, but suitable algorithms for generalizing the techniques are currently under investigation.
Very high compression in videophone coding can be reached successfully only if model-based segmentation is performed to allow suitable bit allocation. The exploitation of a priori knowledge suggests the application of fast and simple segmentation algorithms oriented at partitioning the image into variable resolution domains for subsequent texture encoding. In this paper we describe, together with the achieved preliminary results, a model-based approach to image segmentation relying on the estimation of the face symmetry axis and of the primary facial features. Through these parameters a flexible lattice is adapted frame by frame on the image, identifying a time-varying net of triangular patches whose texture is eventually encoded via Legendre basis functions. Target applications for videophone coding of QCIF color sequences at very low bitrate, less than 16 Kbit/sec., are foreseen.< >
This paper describes a new approach to very low bit-rate interpersonal visual communication based on a suitable scene model, i.e. a flexible structure adapted to the specific characteristics of the speaker's face. The face model is dynamically adapted to time-varying facial expressions by means of few parameters, estimated from the analysis of the real image sequence, which are used to apply knowledge-based deformation rules on a simplified muscle structure. Facial muscles are distributed in correspondence to the primary facial features and can be activated through the direct stimulation of each individual fiber or, indirectly, by interaction with adjacent stimulated fibers. The analysis algorithms performed at the transmitter to estimate the model parameters are based on feature-oriented operators aimed at segmenting the real incoming frames and at the extraction of the primary facial descriptors. The analysis/synthesis algorithms have been developed on a Silicon Graphics workstation and have been tested on various ‘head-and-shoulder’ sequences: the obtained results are very promising for applications both in videophone coding and in picture animation, where the facial expressions of a synthetic actor is reproduced according to the parameters extracted from a real speaking face.
This paper describes an approach to use artificial reality techniques for real-time interpersonal visual communication at very low bitrate. A flexible structure is suitably adapted to the specific characteristics of the speaker’s head by means of few parameters estimated from the analysis of the real image sequence, while head motion and facial mimics are synthesized on the model by means of knowledge-based deformation rules acting on a simplified muscle structure. The analysis algorithms performed at the transmitter to estimate the model parameters are based on feature-oriented operators aimed at segmenting the real incoming frames and at the extraction of the primary facial descriptors. The system performances have been evaluated on different “head-and-shoulder” sequences and the precision, robustness and complexity of the employed analysis/synthesis algorithms have been tested. Promising results have been achieved for applications both in videophone coding and in picture animation where the facial mimics of a synthetic actor is reproduced according to the parameters extracted from a real speaking face.
An innovative knowledge-based scheme is presented for videophone sequence coding, with a reasonable subjective reconstruction quality. The videophone images are segmented and approximated by means of a suitably adapted 3D wire-frame model capable of faithfully reproducing the signal temporal evolution through a limited number of updating parameters extracted from the real sequence. The muscle modeling represents the basic issue for the analysis of facial expressions at the encoder and for their synthesis at the decoder. Facial muscles are modeled either as linear fibers or as circular fibers and are characterized by descriptive parameters such as thickness, length, position and direction. The animation is performed by simulating the muscle contraction and relaxation over the domain of influence defined by the muscle parameters. The interactive procedure for the construction and adaptation of the 3D model and the analysis/synthesis algorithms, implemented on a multiprocessor AT&T Pixel Machine, are described, together with some examples of the achieved experimental results
In this paper an innovative approach to videophone coding is described, based on a 3D model of a human face with a complex muscle structure capable to reproduce faithfully the basic facial expressions and mimics. The muscle structure is organized in a set of interconnected fibers covering the whole surface of the face and characterized by predefined mechanical properties. Muscles can be activated through the direct stimulation of each individual fiber or, indirectly, by interaction with adjacent stimulated fibers. Through the analysis algorithms performed at the transmitter, the input video sequence is processed to extract a set of suitable facial control features and to estimate the muscle parameters. These parameters, namely the value of the estimated muscle stimula, are then quantized, coded and transmitted to the receiver where they are applied to the model to synthesize the corresponding facial expression. In the paper the muscle structure is described together with the coding/decoding algorithms which have been implemented on a multiprocessor AT&T Pixel Machine. Some samples of the achieved experimental results are also presented to show the significant subjective quality of the reconstructed sequence.