This paper describes an approach to use artificial reality techniques for real-time interpersonal visual communication at very low bitrate. A flexible structure is suitably adapted to the specific characteristics of the speaker’s head by means of few parameters estimated from the analysis of the real image sequence, while head motion and facial mimics are synthesized on the model by means of knowledge-based deformation rules acting on a simplified muscle structure. The analysis algorithms performed at the transmitter to estimate the model parameters are based on feature-oriented operators aimed at segmenting the real incoming frames and at the extraction of the primary facial descriptors. The system performances have been evaluated on different “head-and-shoulder” sequences and the precision, robustness and complexity of the employed analysis/synthesis algorithms have been tested. Promising results have been achieved for applications both in videophone coding and in picture animation where the facial mimics of a synthetic actor is reproduced according to the parameters extracted from a real speaking face.
An innovative algorithm is presented for the design and implementation of a self-adaptive vector quantizer for videophone coding. The codebook is organized in a binary tree structure capable to modify its topology and content in order to adaptively track the input varying statistics. The resulting tree structure is unbalanced, in the sense that some branches are extended more than others, depending on the input vector distribution, and therefore the suboptimality of tree structured VQ with respect to the full-search approach is greatly reduced. As it is known a priori that the class of input video data consists of a human head slowly moving against a static background, knowledge based segmentation of incoming frames can be efficiently performed. Exhaustive simulations have shown the effectiveness of the algorithm in standard CIF videophone sequences. Some procedures have been parallelized and implemented on a AT&T Pixel Machine for frame-rate coding. Results and comments on both the algorithm and the parallel implementation are presented.
The reconstruction of 3D information in a scene in relative motion with respect to a visual sensor is a basic issue in many applications, like image processing and computer graphics, object recognition, computer vision, robotics. The technique presented in this paper exploits the information of 2D multiple perspective views of an object to first compute the spatial position and orientation of the camera, for each view, on the base of a motion recovery from corresponding points approach, and then to build a volumetric model of the object from occluding contours. Finally, the model is integrated with the pictorial information extracted from the views, at a resolution independent from the volumetric one. Some experimental results are also presented.
Estimating the 3-D motion from a sequence of images is one of the means to efficiently code time-varying scenes and to infer 3-D shape information for scene analysis. Linear and nonlinear approaches have been proposed in the literature to the problem of estimating the motion parameters of a rigid body from a set of corresponding points. In both approaches, errors on the input data reflect into approximations on the estimated parameters. in particular, in the case of the linear approach considered in this paper, the input errors affect the so-called matrix of the essential motion parameters, in such a way that the estimated motion does no longer correspond to the motion of a rigid body. We show that the estimate of the motion can be improved by constraining the above matrix to satisfy the properties corresponding to the motion of a rigid body. This implies adding nonlinear constraints to the original set of linear equations. On the other hand, the unconstrained solution of these latter turns out to be a convenient initial guess for the iterative procedure used to solve the resulting least squares problem. Some simulation ana experimental results are presented, showing how the improvements of the estimated motion depend on the image resolution, the number of corresponding points, their spatial distribution, the amount of motion and the sensor calibration.
In this paper an approach to the analysis of boundary details in planar shapes is presented. It is possible to distinguish, in the structure of an object, between a "main structure" and the overlying "details" or "textures"; at this purpose, a complementarity principle for the description of shape is proposed. It is then suggested that the two entities can be separated in terms of frequency analysis, the low frequency component being associated with the main structure, and high frequency components with details/textures. A representation for high-frequency components is proposed, based on the "zero-crossing of details/textures" which extends to the contours of shapes the theorem by Logan.
A 2-D linear system is said to be a generalized scale-invariant filter if its weighting function is such that the effect of a scale change of the input image is merely a (generally different) scale change of the output. This property cannot be satisfied by shift-invariant systems. After a discussion on the use of such filters in the framework of scale-independent pattern recognition, the most general class of generalized scale-invariant shift-variant 2-D systems is presented and two different implementation techniques are described, one based on the use of the Mellin transform and the other one exploiting suitable coordinate transformations in conjunction with linear shift-invariant systems.
Pietro G. Morasso合作论文数Biomedical Engineering
University of Genova1
Salvatore Gaglio合作论文数Artificial Intelligence;Italian National Research Council ( C.N.R.)1