
We study the effect of development set size on system performance, as measured by verification error. The study was performed using the FERET and FRGC2 databases to construct development training sets of varying size, while XM2VTS was used to test the system. Surprisingly, the achievable performance levels off relatively quickly. Increasing the size of the development set does not bring any benefit. On the contrary it may result in performance degradation. This finding appears to be development set independent. However, the choice of the development set size is protocol dependent.
We propose a new method for synthesizing an illumination normalized image from a face image including diffuse reflection, specular reflection, attached shadow and cast shadow. The method is derived from the self-quotient image (SQI) which is defined by the ratio of albedo at the pixel value to a locally smoothed pixel value. However, the SQI is not synthesized from an image containing shadows or specular reflections. Since these regions correspond to areas of high or low albedo, they cannot be discriminated from diffuse reflection by using only a single image. To classify the appearances, we utilize a simple model defined by a number of basis images which represent diffuse reflection on a generic face. Through experimental results we show the effectiveness of this method for face identification on the Yale Face Database B and on a real-world database, using only a single image for each individual in training
A new method for face recognition, Landmark Model Matching, is proposed in this paper. It is based on the concepts of Elastic Bunch Graph Matching and Active Shape Model, and optimised with Particle Swarm Optimisation. It is a fully automatic algorithm and can be used for face databases where only one image per person is available. A face is represented by a Landmark Model consisting of nodes labelled with jets and gray-level profiles. A Landmark Distribution Model is created from a few training images. The model similarity between the Landmark Distribution Model and the deformable Landmark Model that has to be fitted to the face in the image is maximised by Particle Swarm Optimisation, to find the optimal model to represent the face. Improved results were obtained by this method compared with Elastic Bunch Graph Matching without Optimisation.
Gabor feature has been widely recognized as one of the best representations for face recognition. However, traditionally, it has to be reduced in dimension due to curse of dimensionality. In this paper, an ensemble based Gabor Fisher classifier (EGFC) method is proposed, which is an ensemble classifier combining multiple Fisher discriminant analysis (FDA)-based component classifiers learnt using different segments of the entire Gabor feature. Since every dimension of the entire Gabor feature is exploited by one component FDA classifier, we argue that EGFC makes better use of the discriminability implied in all the Gabor features by avoiding the dimension reduction procedure. In addition, by carefully controlling the dimension of each feature segment, small sample size (3S) problem commonly confronting FDA is artfully avoided. Experimental results on FERET show that the proposed EGFC significantly outperforms the known best results so far. Furthermore, to speed up, hierarchical EGFC (HEGFC) is proposed based on pyramid-based Gabor representation. Our experiments show that, by using the hierarchical method, the time cost of the HEGFC can be dramatically reduced without much accuracy lost
Gabor filters based features, with their good properties of space-frequency localization and orientation selectivity, seem to be the most effective features for face recognition currently. In this paper, we propose a kind of weighted Gabor complex features which combining Gabor magnitude and phase features in unitary space. Its weights are determined according to recognition rates of magnitude and phase features. Meanwhile, subspace based algorithms, PCA and LDA, are generalized into unitary space, and a rarely used distance measure, unitary space cosine distance, is adopted for unitary subspace based recognition algorithms. Using the generalized subspace algorithms our proposed weighted Gabor complex features (WGCF) produce better recognition result than either Gabor magnitude or Gabor phase features. Experiments on FERET database show good results comparable to the best one reported in literature
We present a novel method to relight video sequences given known surface shape and illumination. The method preserves fine visual details. It requires single view video frames, approximate 3D shape and standard studio illumination only, making it applicable in studio production. The technique is demonstrated for relighting video sequences of faces
As an extension to prior work by the authors in the area of photometric normalisation for face verification, we apply these algorithms in a component-based framework. In particular, we investigate how the requirement for complexity of the normalisation changes when smaller image patches are used. We show that for smaller image patches, a simpler normalisation can out-perform a more complicated method. In addition, we show that a method that applies a simpler normalisation to a number of smaller face image components that are then fused, out-performs a more complicated method applied to the full face image
An intelligent robot requires natural interaction with humans. Visual interpretation of gestures can be useful in accomplishing natural human-robot interaction (HRl). Previous HRI researches were focused on issues such as hand gesture, sign language, and command gesture recognition. However, automatic recognition of whole body gestures is required in order to operate HRI naturally. This can be a challenging problem because describing and modeling meaningful gesture patterns from whole body gestures are complex tasks. This paper presents a new method for spotting and recognizing whole body key gestures at the same time on a mobile robot. Our method is simultaneously used with other HRI approaches such as speech recognition, face recognition, and so forth. In this regard, both of execution speed and recognition performance should be considered. For efficient and natural operation, we used several approaches at each step of gesture recognition; learning and extraction of articulated joint information, representing gesture as a sequence of clusters, spotting and recognizing a gesture with HMM. In addition, we constructed a large gesture database, with which we verified our method. As a result, our method is successfully included and operated in a mobile robot
A set of techniques is presented for extracting essential shape information from image sequences. Presented methods are (i) human detection, (ii) human body parts detection, and (iii) hand shape analysis, all based on depth image streams. In particular, representative types of hand shapes used in Japanese sign language (JSL) are recognized in a non-intrusive manner with a high recognition rate. An experimental JSL recognition system is built that can recognize over 100 words by using an active sensing hardware to capture a stream of depth images at a video rate. Experimental results are shown to validate our approach and characteristics of our approach are discussed
This paper presents the application of a kernel particle filter for 3D body tracking in a video stream acquired from a single uncalibrated camera. Using intensity-based and color-based cues as well as an articulated 3D body model with shape represented by cylinders, a real-time body tracking in monocular cluttered image sequences has been realized. The algorithm runs at 7.5 Hz on a laptop computer and tracks the upper body of a human with two arms. First, experimental results show that the proposed approach has good tracking as well as recovering capabilities despite using a small number of particles. The approach is intended for use on a mobile robot to improve human robot interaction
Active statistical models including active shape models and active appearance models are very powerful for face alignment. They are composed of two parts: the subspace model(s) and the search process. While these two parts are closely correlated, existing efforts treated them separately and had not considered how to optimize them overall. Another problem with the subspace model(s) is that the two kinds of parameters of subspaces (the number of components and the constraints on the components) are also treated separately. So they are not jointly optimized. To tackle these two problems, an unified subspace optimization method is proposed. This method is composed of two unification aspects: (1) unification of the statistical model and the search process: the subspace models are optimized according to the search procedure; (2) unification of the number of components and the constraints: the two kinds of parameters are modelled in an unified way, such that they can be optimized jointly. Experimental results demonstrate that our method can effectively find the optimal subspace model and significantly improve the performance.
The registration of 3D scans of faces is a key step for many applications, in particular for building 3D morphable models. Although a number of algorithms are already available for registering data with neutral expression, the registration of scans with arbitrary expressions is typically performed under the assumption of a known, fixed identity. We present a novel algorithm which breaks this restriction, allowing to register 3D scans of faces with arbitrary identity and expression. Furthermore, our algorithm can process incomplete data, yielding results which are both continuous and with low reconstruction error. Even in the case of complete, expression-less data, our method can yield better results than previous algorithms, due to an adaptive smoothing, which regularizes the results surface only where the estimated correspondence is unreliable
We present a novel tracking algorithm that uses dynamic programming to determine the path of target objects and that is able to track an arbitrary number of different objects. The traceback method used to track the targets avoids taking possibly wrong local decisions and thus reconstructs the best tracking paths using the whole observation sequence. The tracking method can be compared to the nonlinear time alignment in automatic speech recognition (ASR) and it can analogously be integrated into a hidden Markov model based recognition process. We show how the method can be applied to the tracking of hands and the face for automatic sign language recognition
In this work, we address a method that is able to track simultaneously 3D head movements and facial actions like lip and eyebrow movements in a video sequence. In a baseline framework, an adaptive appearance model is estimated online by the knowledge of a monocular video sequence. This method uses a 3D model of the face and a facial adaptive texture model. Then, we consider and compare two improved models in order to increase robustness to occlusions. First, we use robust statistics in order to downweight the hidden regions or outlier pixels. In a second approach, mixture models provides better integration of occlusions. Experiments demonstrate the benefit of the two robust models. The latter are compared under various occlusions
In this paper, a layered deformable model (LDM) is proposed for human body pose recovery in gait analysis. This model is inspired by the manually labeled silhouettes in (Z. Liu, et al., July 2004) and it is designed to closely match them. For fronto-parallel gait, the introduced LDM model defines the body part widths and lengths, the position and the joint angles of human body using 22 parameters. The model consists of four layers and allows for limb deformation. With this model, our objective is to recover its parameters (and thus the human body pose) from automatically extracted silhouettes. LDM recovery algorithm is first developed for manual silhouettes, in order to generate ground truth sequences for comparison and useful statistics regarding the LDM parameters. It is then extended for automatically extracted silhouettes. The proposed methodologies have been tested on 10005 frames from 285 gait sequences captured under various conditions and an average error rate of 7% is achieved for the lower limb joint angles of all the frames, showing great potential for model-based gait recognition
Good registration (alignment to a reference) is essential for accurate face recognition. The effects of the number of landmarks on the mean localization error and the recognition performance are studied. Two landmarking methods are explored and compared for that purpose: (1) the most likely-landmark locator (MLLL), based on maximizing the likelihood ratio, and (2) Viola-Jones detection. Both use the locations of facial features (eyes, nose, mouth, etc) as landmarks. Further, a landmark-correction method (BILBO) based on projection into a subspace is introduced. The MLLL has been trained for locating 17 landmarks and the Viola-Jones method for 5. The mean localization errors and effects on the verification performance have been measured. It was found that on the eyes, the Viola-Jones detector is about 1% of the interocular distance more accurate than the MLLL-BILBO combination. On the nose and mouth, the MLLL-BILBO combination is about 0.5% of the inter-ocular distance more accurate than the Viola-Jones detector. Using more landmarks will result in lower equal-error rates, even when the landmarking is not so accurate. If the same landmarks are used, the most accurate landmarking method gives the best verification performance
This paper introduces a new representation of hand motions for tracking and recognizing hand-finger gestures in an image sequence. A human hand has 15 joints and its high dimensionality makes it difficult to model hand motions. To make things easier, it is important to represent a hand motion in a low dimensional space. Principle component analysis (PCA) has been proposed to reduce the dimensionality. However, the PCA basis vectors only represent global features, which are not optimal to represent intrinsic features. This paper proposes an efficient representation of hand motions by independent component analysis (ICA). The ICA basis vectors represent local features, each of which corresponds to the motion of a particular finger. This representation is more efficient in modeling hand motions for tracking and recognizing hand-finger gestures in an image sequence. This paper demonstrates the effectiveness of our method by tracking hands in real image sequences
Skin segmentation is the cornerstone of many applications such as gesture recognition, face detection, and objectionable image filtering. In this paper, we attempt to address the skin segmentation problem for gesture recognition. Initially, given a gesture video sequence, a generic skin model is applied to the first couple of frames to automatically collect the training data. Then, an SVM classifier based on active learning is used to identify the skin pixels. Finally, the results are improved by incorporating region segmentation. The proposed algorithm is fully automatic and adaptive to different signers. We have tested our approach on the ECHO database. Comparing with other existing algorithms, our method could achieve better performance.
This paper proposes a new approach for recognition of task-oriented actions based on stochastic context free grammar (SCFG). Our attention puts on actions in the Japanese tea ceremony, where the action can be described by context free grammar. Our aim is to recognize the action in the tea services. Existing SCFG approach consists of generating symbolic string, parsing it and recognition. The symbolic string often includes uncertainty. Therefore, the parsing process needs to recover the errors at the entry process. This paper proposes a segmentation method error-less as much as possible to segment an action into a string of finer actions. This method, based on an acceleration of the body motion, can produce the fine action corresponding to a terminal symbol with little error. After translating the sequence of fine actions into a set of symbolic strings, SCFG-based parsing of this set leaves small number of ones to be derived. Among the remaining strings, Bayesian classifier answers the action name with a maximum posterior probability. Giving one SCFG rule the multiple probabilities, one SCFG can recognize multiple actions.
We describe an accurate and robust method of locating facial features. The method utilises a set of feature templates in conjunction with a shape constrained search technique. The current feature templates are correlated with the target image to generate a set of response surfaces. The parameters of a statistical shape model are optimised to maximise the sum of responses. Given the new feature locations the feature templates are updated using a nearest neighbour approach to select likely feature templates from the training set. We find that this template selection tracker (TST) method outperforms previous approaches using fixed template feature detectors. It gives results similar to the more complex active appearance model (AAM) algorithm on two publicly available static image sets and outperforms the AAM on a more challenging set of in-car face sequences