We introduce a new framework, namely tensor canonical correlation analysis (TCCA) which is an extension of classical canonical correlation analysis (CCA) to multidimensional data arrays (or tensors) and apply this for action/gesture classification in videos. By tensor CCA, joint space-time linear relationships of two video volumes are inspected to yield flexible and descriptive similarity features of the two videos. The TCCA features are combined with a discriminative feature selection scheme and a nearest neighbor classifier for action classification. In addition, we propose a time-efficient action detection method based on dynamic learning of subspaces for tensor CCA for the case that actions are not aligned in the space-time domain. The proposed method delivered significantly better accuracy and comparable detection speed over state-of-the-art methods on the KTH action data set as well as self-recorded hand gesture data sets.
Local spatiotemporal features or interest points provide compact but descriptive representations for efficient video analysis and motion recognition. Current local feature extraction approaches involve either local filtering or entropy computation which ignore global information (e.g. large blobs of moving pixels) in video inputs. This paper presents a novel extraction method which utilises global information from each video input so that moving parts such as a moving hand can be identified and are used to select relevant interest points for a condensed representation. The proposed method involves obtaining a small set of subspace images, which can synthesise frames in the video input from their corresponding coefficient vectors, and then detecting interest points from the subspaces and the coefficient vectors. Experimental results indicate that the proposed method can yield a sparser set of interest points for motion recognition than existing methods.
Current approaches to motion category recognition typically focus on either full spatiotemporal volume analysis (holistic approach) or analysis of the content of spatiotemporal interest points (part-based approach). Holistic approaches tend to be more sensitive to noise e.g. geometric variations, while part-based approaches usually ignore structural dependencies between parts. This paper presents a novel generative model, which extends probabilistic latent semantic analysis (pLSA), to capture both semantic (content of parts) and structural (connection between parts) information for motion category recognition. The structural information learnt can also be used to infer the location of motion for the purpose of motion detection. We test our algorithm on challenging datasets involving human actions, facial expressions and hand gestures and show its performance is better than existing unsupervised methods in both tasks of motion localisation and recognition.
An appearance-based approach to track an object that may undergo appearance change is proposed. Unlike recent methods that store a detailed representation of object's appearance, this method allows an appearance feature with a reduced dimension to be used. Through the use of a sparse Bayesian classifier, high classification and detection accuracy can be maintained even if a reduced feature vector is used. In addition, the classifier allows online-training which enables online-updating of the original classification model and provides better adaptability. Experiments show that the method can be used to track targets undergo appearance change due to the change in view-point, facial expression and lighting direction.
An approach to recognise and segment 9 elementary gestures from a video input is proposed and it can be applied to continuous sign recognition. An isolated gesture is recognised by first converting a portion of video into a motion gradient orientation image and then classifying it into one of the 9 gestures by a sparse Bayesian classifier. The portion of video used is decided by using a sampling technique based on condensation framework. By doing so, gestures can be segmented from the video in a probabilistic manner. Experiments show that the proposed method can achieve accuracy around 90% in both isolated and continuous gesture recognition without using special equipment such as glove devices and the system can run in real-time
An approach to recognise 10 elementary gestures is proposed and it can be applied to sign language recognition. In this work, a motion gradient orientation image is extracted directly from a raw video input and transformed to a motion feature vector. This feature vector is then classified into one of the 10 elementary gestures by a sparse Bayesian classifier. A training set of 628 samples and a testing set of over 1000 samples have been obtained to evaluate the proposed method. A real-time system was built and trained with the training set. From the experiment, the reported classification accuracy is 90% and the system can run in around 25 frames per second. Compared with other recently proposed methods that involve the use of hand tracking, the system can work reliably in real-time without relying on accurate tracking, and give a probabilistic output that is useful in complex motion analysis.
An approach to increase adaptability of a recognition system, which can recognise 10 elementary gestures and be extended to sign language recognition, is proposed. In this work, recognition is done by firstly extracting a motion gradient orientation image from a raw video input and then classifying a feature vector generated from this image to one of the 10 gestures by a sparse Bayesian classifier. The classifier is designed in a way that it supports online incremental learning and it can be thus re-trained to increase its adaptability to an input captured under a new condition. Experiments show that the accuracy of the classifier can be boosted from less than 40% to over 80% by re-training it using 5 newly captured samples from each gesture class. Apart from having a better adaptability, the system can work reliably in real-time and give a probabilistic output that is useful in complex motion analysis.
Low back pain becomes one of the significant problem in the industrialized world. Efficient and effective spinal motion analysis is required to understand low back pain and to aid the diagnosis. Videofluoroscopy provides a cost effective way for such analysis. However, common approaches are tedious and time consuming due to the low quality of the images. Physicians have to extract the vertebrae manually in most cases and thus continuous motion analysis is hardly achieved. In this paper, we propose a system which can perform automatic vertebrae segmentation and tracking. Operators need to define exact location of landmarks in the first frame only. The proposed system will continuously learn the texture pattern along the edge and the dynamics of the vertebrae in the remaining frames. The system can estimate the location of the vertebrae based on the learnt texture and dynamics throughout the sequence. Experimental results show that the proposed system can segment vertebrae from videofluoroscopic images automatically and accurately.
Face detection has potential applications in a wide range of commercial products such as automatic face recognition system. Commonly used face detection algorithms can extract faces from images accurately and reliably, but they often take a long time to finish the detection process. Recently, there is an increasing demand of real time face detection algorithm in applications like video surveillance system. This paper aims at proposing a multi-scale face detection scheme using Quadtree so that the time complexity of the face detection process can be reduced. By performing analysis from coarse to fine scales, the proposed scheme uses skin color as a heuristic feature, and support vector machine as a verification tool to detect face. Experimental results show that the proposed scheme can detect faces from images reliably and quickly.
Object tracking is useful in applications like computer-aided medical diagnosis, video editing, visual surveil- lance etc. Commonly used approaches usually involve the use of filter (e.g. Kalman filter) to predict the location of the ob ject in next image frame. Such approaches actually borrow ideas from signal theory and are limited to applications where dynamic model is known. In this paper, a flexible and reliable estimat ion algorithm using wavelet network (or wavenet) is proposed to build an object tracking system. This system simulates the perception of motion that occurs in primates. Neural-based filters will be used for color, shape and motion analysis. Experimental results show that object can be tracked accurately without fi xing any dynamic model compare with commonly used Kalman filter.
Recognition of human motion provides hints to understand hu man activities and gives opportunities to the development of new human-com puter interface. Recent studies, however, are limited to extracting motion h istory image and recognizing gesture or locomotion of human body parts. Alth ough the approach employed, i.e. the transformation of the 3D space-ti m (x-y-t) analysis to the 2D image analysis, is faster than analyzing 3D motion f eature, it is less accurate and less robust in nature. In this paper, a fast traj ectory-classification algorithm for interpreting movement of human body parts usi ng wavelet analysis is proposed to increase the accuracy and robustness of h uman motion recognition. By tracking human body in real time, the motion rajectory (x-y-t) can be extracted. The motion trajectory is then broken down i nto wavelets that form a set of wavelet features. Classification based on the wa velet features can then be done to interpret the human motion. An online hand drawing digit recognition system was built using the proposed algorithm. Experiments show that the proposed algorithm is able to recognize digits from human movement accurately in real time.
Robust image segmentation plays an important role in a wide range of daily applications, like visual surveillance system, computer-aided medical diagnosis, etc. Although commonly used image segmentation methods based on pixel intensity and texture can help finding the boundary of targets with sharp edges or distinguished textures, they may not be applied to images with poor quality and low contrast. Medical images, images captured from web cam and images taken under dim light are examples of images with low contrast and with heavy noise. To handle these types of images, we proposed a new segmentation method based on texture clustering and snake fitting. Experimental results show that targets in both artificial images and medical images, which are of low contrast and heavy noise, can be segmented from the background accurately. This segmentation method provides alternatives to the users so that they can keep using imaging device with low quality outputs while having good quality of image analysis result.
Regression analysis is an essential tools in most research fields such as signal processing, economic forecasting etc. In this paper, an regression algorithm using probabilistic wavelet network is proposed. As in most neural network (NN) regression methods, the proposed method can model nonlinear functions. Unlike other NN approaches, the proposed method is much robust to noisy data and thus over-fitting may not occur easily. This is because the use of wavelet representation in the hidden nodes and the probabilistic inference on the value of weights such that the assumption of smooth curve can be encoded implicitly. Experimental results show that the proposed network have higher modeling and prediction power than other common NN regression methods.
Robust image segmentation plays an important role in a wide range of daily applications, like visual surveillance system, computer-aided medical diagnosis, etc. Although commonly used image segmentation methods based on pixel intensity and texture can help finding the boundary of targets with sharp edges or distinguished textures, they may not be applied to images with poor quality and low contrast. Medical images, images captured from web cam and images taken under dim light are examples of images with low contrast and with heavy noise. To handle these types of images, we proposed a new segmentation method based on texture clustering and snake fitting. Experimental results show that targets in both artificial images and medical images, which are of low contrast and heavy noise, can be segmented from the background accurately. This segmentation method provides alternatives to the users so that they can keep using imaging device with low quality outputs while having good quality of image analysis result.
Video fluoroscopy provides a cost effective way for the diagnosis of low back pain. Backbones or vertebrae are usually segmented manually from fluoroscopic images of low quality during such a diagnosis. In this paper, we try to reduce human workload by performing automatic vertebrae detection and segmentation. Operators need to provide the rough location of landmarks only. The proposed algorithm would perform edge detection, which is based on pattern recognition of texture, along the snake formed from the landmarks. The snake would then attach to the edge detected. Experimental results show that the proposed system can segment vertebrae from video fluoroscopic image automatically and accurately.
Human body tracking is useful in applications like medical diagnostic, human computer interface, visual surveillance etc. In most cases, only rough position of the target is needed, and blob tracking can be used. The blob region is located within a searching window, which is shifted and resized in each frame based on previous observations. The observations are the locations of the blob in the frames, and are fed into an estimator for predicting the position and the size of the searching window. However, a blob region is regarded as a noisy observation, and the information provided by the blob observation is deficient for most estimators to work well. In this paper, a reliable and efficient estimation algorithm using wavelet is proposed to track human body under information deficiency. The human body is located roughly within a small searching window using color and motion as heuristics. The location and the size of the searching window are estimated using the proposed wavelet estimation scheme. Experimental results show that human body can be tracked accurately and efficiently using the proposed method. The tracker works well in various conditions like clutter background, and background with distractors.