•A latent graphical model integrating multi-target tracking, group discovery, and activity recognition is proposed.•Performance of activity recognition improves when multi-target tracking and group clustering are incorporated.•Group activities are better recognized based on the structured relations within the group and group–group compatibilities.•Increasing the connectivity of different groups improves the overall performance.•Incorporating activity information leads to robust group localization in the video.
•To construct one unified benchmark dataset with various scenes and challenges.•To standardize uniform evaluation criteria for comparing different kinds of algorithms.•To carefully design a baseline method for human running detection.•To open up a new research direction for the specific human motion.
The recently proposed l(2)-norm based collaborative representation for classification (CRC) model has shown inspiring performance on face recognition after the success of its predecessor - the l(1) - norm based sparse representation for classification (SRC) model. Though CRC is much faster than SRC as it has a closed-form solution, it may have the same weakness as SRC, i. e., relying on a "good" (properly controlled) training dataset for serving as its dictionary. Such a weakness limits the usage of CRC in real applications because the quality requirement is not easy to verify in practice. Inspired by the encouraging progress on dictionary learning for sparse representation, which can much alleviate this problem, we propose the discriminative collaborative representation (DCR) model. It has a novel classification model well fitting its discriminative learning model. As a result, DCR has the same advantage of being efficient as CRC, while at the same time showing even stronger discriminative power than existing dictionary learning methods. Extensive experiments on nine widely used benchmark datasets for both controlled and uncontrolled classification tasks demonstrate its consistent effectiveness and efficiency.
Bag detection in pedestrian images is a very practical visual surveillance problem. It is challenging because bag appearance may vary greatly. In this paper, we propose a novel two-stage approach for bag detection in pedestrian images. Firstly, we utilize two stripe vocabulary forests to check whether a pedestrian is with a bag. Secondly, we locate the bag location by ranking the generated bottom-up region proposals. The ranker is learned with a convolutional neural network (CNN). Experiments are performed on a subset of CUHK person re-identification dataset that show the effectiveness of our approach for bag detection in pedestrian images. Although developed for a specific problem, our approach could be applied to detect other carrying objects in pedestrian images.
Face recognition with still face images has been widely studied, while the research on video-based face recognition is inadequate relatively, especially in terms of benchmark datasets and comparisons. Real-world video-based face recognition applications require techniques for three distinct scenarios: 1) Videoto-Still (V2S); 2) Still-to-Video (S2V); and 3) Video-to-Video (V2V), respectively, taking video or still image as query or target. To the best of our knowledge, few datasets and evaluation protocols have benchmarked for all the three scenarios. In order to facilitate the study of this specific topic, this paper contributes a benchmarking and comparative study based on a newly collected still/video face database, named COX 1 Face DB. Specifically, we make three contributions. First, we collect and release a largescale still/video face database to simulate video surveillance with three different video-based face recognition scenarios (i.e., V2S, S2V, and V2V). Second, for benchmarking the three scenarios designed on our database, we review and experimentally compare a number of existing set-based methods. Third, we further propose a novel Point-to-Set Correlation Learning (PSCL) method, and experimentally show that it can be used as a promising baseline method for V2S/S2V face recognition on COX Face DB. Extensive experimental results clearly demonstrate that video-based face recognition needs more efforts, and our COX Face DB is a good benchmark database for evaluation.
Clothing attributes, of which color plays an important role, are receiving more and more interests in machine vision researches and applications because of their uses and effectiveness in tasks like pedestrian analysis. However, color description is a challenging problem due to complex environments such as illumination variations. Most prior works describe color attributes using only low-level features or mid-level descriptors, which results in a marked drop of the discriminative power or photometric invariance. In this paper we introduce a new efficient joint representation that aims to overcome the shortcomings of using low-level features or mid-level descriptors alone and present a novel hybrid approach to pedestrian clothing color attribute extraction. As a necessary preprocessing step, a novel processing pipeline is also proposed. We evaluate our approach on the task of color classification on both the public dataset VIPeR and our own newly-built pedestrian dataset. Experimental results have demonstrated the effectiveness of our approach and have shown its great potential for further researches and applications.
Beyond recognizing the actions of individuals, activity group localization aims to determine ‘‘who participates in each group’’ and ‘‘what activity the group performs’’. In this paper, we propose a latent graphical model to group participants while inferring each group’s activity by exploring the relations among them, thus simultaneously addressing the problems of group localization and activity recognition. Our key insight is to exploit the relational graph among the participants. Specifically, each group is represented as a tree with an activity label while relations among groups are modeled as a fully connected graph. Inference of such a graph is reduced into an extended minimum spanning forest problem, which is casted into a max-margin framework. It therefore avoids the limitation of high-ordered hierarchical model and can be solved efficiently. Our model is able to provide strong and discriminative contextual cues for activity recognition and to better interpret scene information for localization. Experiments on three datasets demonstrate that our model achieves significant improvements in activity group. localization and state-of-the-arts performance on activity recognition.
Due to the misalignment of image features, the performance of many conventional face recognition methods degrades considerably in across pose scenario. To address this problem, many image matching-based methods are proposed to estimate semantic correspondence between faces in different poses. In this paper, we aim to solve two critical problems in previous image matching-based correspondence learning methods: 1) fail to fully exploit face specific structure information in correspondence estimation and 2) fail to learn personalized correspondence for each probe image. To this end, we first build a model, termed as morphable displacement field (MDF), to encode face specific structure information of semantic correspondence from a set of real samples of correspondences calculated from 3D face models. Then, we propose a maximal likelihood correspondence estimation (MLCE) method to learn personalized correspondence based on maximal likelihood frontal face assumption. After obtaining the semantic correspondence encoded in the learned displacement, we can synthesize virtual frontal images of the profile faces for subsequent recognition. Using linear discriminant analysis method with pixel-intensity features, state-of-the-art performance is achieved on three multipose benchmarks, i.e., CMU-PIE, FERET, and MultiPIE databases. Owe to the rational MDF regularization and the usage of novel maximal likelihood objective, the proposed MLCE method can reliably learn correspondence between faces in different poses even in complex wild environment, i.e., labeled face in the wild database.
High accurate face recognition is of great importance for real-world applications such as identity authentication, watch list screening, and human-computer interaction. Despite tremendous progress made in the last decades, fully automatic face recognition systems are still far from the goal of surpassing the human vision system, especially in uncontrolled conditions. In this paper, we propose an approach for robust face recognition by fusing two complementary features: one is the Gabor magnitude of multiple scales and orientations and the other is Fourier phase encoded by Local Phase Quantization (LPQ). To further reduce the high dimensionality of both features, patch-wise Fisher Linear Discriminant Analysis is applied respectively and further combined by score-level fusion. In addition, multi-scale face models are exploited to make use of more information and improve the robustness of the proposed approach. Experimental results show that the proposed approach achieves 96.09%, 95.64% and 95.15% verification rates (when FAR=0.1%) on ROC1/2/3 of Face Recognition Grand Challenge (FRGC) version 2 Experiment 4, impressively surpassing the best known results, i.e. 93.91%, 93.55%, and 93.12%.
Recently, more and more approaches are emerging to solve the cross-view matching problem where reference samples and query samples are from different views. In this paper, inspired by Graph Embedding, we propose a unified framework for these cross-view methods called Cross-view Graph Embedding. The proposed framework can not only reformulate most traditional cross-view methods (e.g., CCA, PLS and CDFE), but also extend the typical single-view algorithms (e.g., PCA, LDA and LPP) to cross-view editions. Furthermore, our general framework also facilitates the development of new cross-view methods. In this paper, we present a new algorithm named Cross-view Local Discriminant Analysis (CLODA) under the proposed framework. Different from previous cross-view methods only preserving inter-view discriminant information or the intra-view local structure, CLODA preserves the local structure and the discriminant information of both intra-view and inter-view. Extensive experiments are conducted to evaluate our algorithms on two cross-view face recognition problems: face recognition across poses and face recognition across resolutions. These real-world face recognition experiments demonstrate that our framework achieves impressive performance in the cross-view problems.
This paper proposes a novel hierarchical summarization system for surveillance video in which synopsis video and original video are at the lower layer and dynamic collage at the higher layer. A user can efficiently browse the video content through dynamic zooming into collage images and further skimming synopsis video or original video for details, enabling an interactive interface for video navigation. We call this system Dynamic VideoBook because its structure is akin to a book, that is, dynamic collage images resemble the cover and the contents while synopsis video gives an abridged edition of the video. Unlike previous collage techniques for films or entertainment videos, collage for surveillance video is challenging since such video imposes restrictions on spatiotemporal relationship and focus on multiple foreground objects. So we put forward a new collage method which better conveys the spatiotemporal motion and structurally extracts the visual contents from the video simultaneously. And we also enable dynamic zooming in to reveal more information in real-time. In addition, by mapping the upper layer of sketchy visualization to the lower layer of visual data, this system can efficiently explore and browse the video content, highlight one motion or interactively search one specific fragment.
We propose a feature, the Histogram of Oriented Normal Vectors (HONV), designed specifically to capture local geometric characteristics for object recognition with a depth sensor. Through our derivation, the normal vector orientation represented as an ordered pair of azimuthal angle and zenith angle can be easily computed from the gradients of the depth image. We form the HONV as a concatenation of local histograms of azimuthal angle and zenith angle. Since the HONV is inherently the local distribution of the tangent plane orientation of an object surface, we use it as a feature for object detection/classification tasks. The object detection experiments on the standard RGB-D dataset [1] and a self-collected Chair-D dataset show that the HONV significantly outperforms traditional features such as HOG on the depth image and HOG on the intensity image, with an improvement of 11.6% in average precision. For object classification, the HONV achieved 5.0% improvement over state-of-the-art approaches.
In this paper, we explore the real-world Still-to-Video (S2V) face recognition scenario, where only very few (single, in many cases) still images per person are enrolled into the gallery while it is usually possible to capture one or multiple video clips as probe. Typical application of S2V is mug-shot based watch list screening. Generally, in this scenario, the still image(s) were collected under controlled environment, thus of high quality and resolution, in frontal view, with normal lighting and neutral expression. On the contrary, the testing video frames are of low resolution and low quality, possibly with blur, and captured under poor lighting, in non-frontal view. We reveal that the S2V face recognition has been heavily overlooked in the past. Therefore, we provide a benchmarking in terms of both a large scale dataset and a new solution to the problem. Specifically, we collect (and release) a new dataset named COX-S2V, which contains 1,000 subjects, with each subject a high quality photo and four video clips captured simulating video surveillance scenario. Together with the database, a clear evaluation protocol is designed for benchmarking. In addition, in addressing this problem, we further propose a novel method named Partial and Local Linear Discriminant Analysis (PaLo-LDA). We then evaluated the method on COX-S2V and compared with several classic methods including LDA, LPP, ScSR. Evaluation results not only show the grand challenges of the COX-S2V, but also validate the effectiveness of the proposed PaLo-LDA method over the competitive methods.