With the rapid development of video technology more and more video applications gradually enter people's lives Therefore conducting research on video quality is very meaningful Herein a no-reference video quality assessment algorithm based on the powerful feature-extraction capabilities of convolutional neural networks and recurrent neural networks combined with the attention mechanism is proposed This algorithm first extracts the spatial features of the distorted videos by using the Visual Geometry Group VGG network the distortion of video airspace feature extraction Further we use cycle time-domain features of neural networks to extract the video distortion Then the introduced attention mechanism important degree for the space-time characteristics of the video is calculated according to the important degree of the overall characteristics of the video Finally regression of the entire connection layer is performed to obtain the evaluation score of the video quality Experiment results on three public video databases show that the predicted results are in good agreement with human subjective quality scores and have better performance than the latest video quality evaluation algorithms
With the rapid development of video technology, more and more video applications gradually enter people's lives, Therefore, conducting research on video quality is very meaningful. Herein, a no-reference video quality assessment algorithm based on the powerful feature-extraction capabilities of convolutional neural networks and recurrent neural networks combined with the attention mechanism is proposed. This algorithm first extracts the spatial features of the distorted videos by using the Visual Geometry Group (VGG) network, the distortion of video airspace feature extraction. Further, we use cycle time-domain features of neural networks to extract the video distortion. Then the introduced attention mechanism important degree for the space-time characteristics of the video is calculated according to the important degree of the overall characteristics of the video. Finally, regression of the entire connection layer is performed to obtain the evaluation score of the video quality. Experiment results on three public video databases show that the predicted results arc in good agreement with human subjective quality scores and have better performance than the latest video quality evaluation algorithms.
In video quality assessment, most researchers manually extract the features first, and then use machine learning to predict video quality score, which leads to unideal result. Since the VGG-16 net has excellent robustness in feature extraction, we use the network model and migrate parameters to construct the end-to-end video quality assessment network. The experimental results on LIVE video database show that the assessment score of this method is consistent with the subjective assessment score, and its assessment indexes of Spearman rank correlation coefficient and Pearson correlation coefficient reached 0.867 and 0.843, respectively, which indicated that the performance of the proposed method is better than most of the current video quality assessment algorithms based on manual feature extraction.
As face recognition progresses from constrained scenarios to unconstrained scenarios, new challenges such as large pose, bad illumination, and partial occlusion, are encountered. While 3D or multi-modality RGB-D sensors are helpful for face recognition systems to achieve robustness against these challenges, the requirement of new sensors limits their application scenarios. In our paper, we propose a discriminative face depth estimation approach to improve 2D face recognition accuracies under unconstrained scenarios. Our discriminative depth estimation method uses a cascaded FCN and CNN architecture, in which FCN aims at recovering the depth from an RGB image, and CNN retains the separability of individual subjects. The estimated depth information is then used as a complementary modality to RGB for face recognition tasks. Experiments on two public datasets and a dataset we collect show that the proposed face recognition method using RGB and estimated depth information can achieve better accuracy than using RGB modality alone.
RGB-D face recognition has attracted increasing attentions in recent years because of its robustness in unconstrained environment. However, existing approaches either handle individual modalities using completely separate pipelines or treat all the modalities equally using the same pipeline. Such approaches did not adequately consider the modality differences and exploit the modality correlations. We propose a novel approach for RGB-D face recognition that is able to learn complementary features from multiple modalities and common features between different modalities. Specifically, we introduce a joint loss taking activation from both modality-specific feature learning networks, and enforcing the features to be learned in a complementary way. We further extend the capability of this multi-modality (e.g., RGB-D vs. RGB-D) matcher into cross-modality (e.g., RGB vs. RGB-D) scenarios by learning a common feature transformation mapping different modalities into the same feature space. Experimental results on a number of public RGB-D face databases (e.g., EURECOM, VAP, IIIT-D, and BUAA), and a large RGB-D database we collected, show the impressive performance of the proposed approach.
Traditional sparse image models treat color image pixel as a scalar, which represents color channels separately or concatenate color channels as a monochrome image. In this paper, we propose a vector sparse representation model for color images using quaternion matrix analysis. As a new tool for color image representation, its potential applications in several image-processing tasks are presented, including color image reconstruction, denoising, inpainting, and super-resolution. The proposed model represents the color image as a quaternion matrix, where a quaternion-based dictionary learning algorithm is presented using the K-quaternion singular value decomposition (QSVD) (generalized K-means clustering for QSVD) method. It conducts the sparse basis selection in quaternion space, which uniformly transforms the channel images to an orthogonal color space. In this new color space, it is significant that the inherent color structures can be completely preserved during vector reconstruction. Moreover, the proposed sparse model is more efficient comparing with the current sparse models for image restoration tasks due to lower redundancy between the atoms of different color channels. The experimental results demonstrate that the proposed sparse image model avoids the hue bias issue successfully and shows its potential as a general and powerful tool in color image analysis and processing domain.
The term soft biometrics typically refers to attributes of people such as their gender, the shape of their head or the color of their hair. There is growing interest in soft biometrics as a means of improving automated face recognition since they hold the promise of significantly reducing recognition errors, in part by ruling out illogical choices. This paper concentrates specifically on soft biometrics as opposed to extended attributes, and presents the results from three experiments quantifying performance gains on a difficult face recognition task when standard face recognition algorithms are augmented using soft biometrics. These experiments include (1) a best-case analysis using perfect knowledge of gender and race, (2) support vector machine-based soft biometric classifiers and (3) face shape expressed through an active shape model. All three experiments indicate small improvements may be made when soft biometrics augment an existing algorithm. However, in all cases, the gains were modest. One reason is that false matches are more likely between faces of people sharing the same soft biometric traits. This is to be expected, since face recognition algorithms utilize appearance information, which is the same information used by algorithms designed to assign soft biometric labels to face images.
This report presents results from the Video Person Recognition Evaluation held in conjunction with the 11th IEEE International Conference on Automatic Face and Gesture Recognition. Two experiments required algorithms to recognize people in videos from the Point-and-Shoot Face Recognition Challenge Problem (PaSC). The first consisted of videos from a tripod mounted high quality video camera. The second contained videos acquired from 5 different handheld video cameras. There were 1401 videos in each experiment of 265 subjects. The subjects, the scenes, and the actions carried out by the people are the same in both experiments. Five groups from around the world participated in the evaluation. The video handheld experiment was included in the International Joint Conference on Biometrics (IJCB) 2014 Handheld Video Face and Person Recognition Competition. The top verification rate from this evaluation is double that of the top performer in the IJCB competition. Analysis shows that the factor most effecting algorithm performance is the combination of location and action: where the video was acquired and what the person was doing.
The Point-and-Shoot Face Recognition Challenge (PaSC) is a performance evaluation challenge including 1401 videos of 265 people acquired with handheld cameras and depicting people engaged in activities with non-frontal head pose. This report summarizes the results from a competition using this challenge problem. In the Video-to-video Experiment a person in a query video is recognized by comparing the query video to a set of target videos. Both target and query videos are drawn from the same pool of 1401 videos. In the Still-to-video Experiment the person in a query video is to be recognized by comparing the query video to a larger target set consisting of still images. Algorithm performance is characterized by verification rate at a false accept rate of 0.01 and associated receiver operating characteristic (ROC) curves. Participants were provided eye coordinates for video frames. Results were submitted by 4 institutions: (i) Advanced Digital Science Center, Singapore; (ii) CPqD, Brasil; (iii) Stevens Institute of Technology, USA; and (iv) University of Ljubljana, Slovenia. Most competitors demonstrated video face recognition performance superior to the baseline provided with PaSC. The results represent the best performance to date on the handheld video portion of the PaSC.
A new algorithm for learning binary codes is presented using randomized initial assignments of bit labels to classes followed by iterative refinement to minimize intraclass Hamming distance. This Randomized Intraclass-Distance Minimizing Binary Codes (RIDMBC) algorithm is introduced in the context of face recognition, an area of biometrics where binary codes have rarely been used (unlike iris recognition). A cross-database experiment is presented training RIDMBC on the Labeled Faces in the Wild (LFW) and testing it on the Point-and-Shoot Challenge (PaSC). The RIDMBC algorithm performs better than both PaSC baselines. RIDMBC is compared with the Predictable Discriminative Binary Codes (DBC) algorithm developed by Rastegari et al. The DBC algorithm has an upper bound on the number of bits in a binary code; RIDMBC does not. RIDMBC outperforms DBC when using the same bit code length as DBC's upper bound and RIDMBC further improves when more bits/features are added.
This paper identifies important factors for face recognition algorithm performance in video.The goal of this study is to understand key factors that affect algorithm performance and to characterize the algorithm performance.We evaluate four factor metrics for a single video as well as two comparative metrics for pairs of videos.This study carried out an investigation of the effect of nine factors on three algorithms using the Point-and-Shoot Challenge (PaSC) video dataset.These factors can be categorized into three groups: 1) image/video (pose yaw, pose roll, face size, and face detection confidence); 2) environment (environmental condition with person's activity and sensor model); and 3) subject (subject ID, gender, and race).For videobased face recognition, the analysis shows that the distribution-based methods were generally more effective in quantifying factor values.For predicting face recognition performance in a video, we observed that face detection confidence and face size serve as potentially useful quality measure metrics.We also find that male faces are easier to identify than female faces, and Asians are easier than Caucasians.Further, on the PaSC video dataset, the performance of face recognition algorithms are primarily driven by environment and sensor factors.
EVALUATING SOFT BIOMETRICS IN THE CONTEXT OF FACE RECOGNITION Soft biometrics typically refer to attributes of people such as their gender, the shape of their head, the color of their hair, etc. There is growing interest in soft biometrics as a means of improving automated face recognition since they hold the promise of significantly reducing recognition errors, in part by ruling out illogical choices. Here four experiments quantify performance gains on a difficult face recognition task when standard face recognition algorithms are augmented using information associated with soft biometrics. These experiments include a best-case analysis using perfect knowledge of gender and race, support vector machine-based soft biometric classifiers, face shape expressed through an active shape model, and finally appearance information from the image region directly surrounding the face. All four experiments indicate small improvements may be made when soft biometrics augment an existing algorithm. However, in all cases, the gains were modest. In the context of face recognition, empirical evidence suggests that significant gains using soft biometrics are hard to come by.
Inexpensive "point-and-shoot" camera technology has combined with social network technology to give the general population a motivation to use face recognition technology. Users expect a lot; they want to snap pictures, shoot videos, upload, and have their friends, family and acquaintances more-or-less automatically recognized. Despite the apparent simplicity of the problem, face recognition in this context is hard. Roughly speaking, failure rates in the 4 to 8 out of 10 range are common. In contrast, error rates drop to roughly 1 in 1,000 for well controlled imagery. To spur advancement in face and person recognition this paper introduces the Point-and-Shoot Face Recognition Challenge (PaSC). The challenge includes 9,376 still images of 293 people balanced with respect to distance to the camera, alternative sensors, frontal versus not-frontal views, and varying location. There are also 2,802 videos for 265 people: a subset of the 293. Verification results are presented for public baseline algorithms and a commercial algorithm for three cases: comparing still images to still images, videos to videos, and still images to videos.
In this paper, we propose a quaternion-based sparse representation model for color images and its corresponding dictionary learning algorithm. Differing from traditional sparse image models, which represent RGB channels separately or process RGB channels as a concatenated real vector, the proposed model describes the color image as a quaternion vector matrix, where each color pixel is encoded as a quaternion unit and thus the inter-relationship among RGB channels is well preserved. Correspondingly, we propose a quaternion-based dictionary learning algorithm using a socalled K-QSVD method. It conducts the sparse basis selection in quaternion vector space, providing a kind of vectorial representation for the inherent color structures rather than a scalar representation via current sparse image models. The proposed sparse model is validated in the applications of color image denoising and inpainting. The experimental results demonstrate that our sparse image model avoids the hue bias phenomenon successfully and shows its potential as a powerful tool in color image analysis and processing domain.
We investigate the existence of quality measures for face recognition. First, we introduce the concept of an oracle for image quality in the context of face recognition. Next we introduce greedy pruned ordering (GPO) as an approximation to an image quality oracle. GPO analysis provides an estimated upper bound for quality measures, given a face recognition algorithm and data set. We then assess the performance of 12 commonly proposed face image quality measures against this standard. In addition, we investigate the potential for learning new quality measures via supervised learning. Finally, we show that GPO analysis is applicable to other biometrics.
Proposed a method for object feature points matching used in active machine vision technique. By projecting pseudo-random coded structured light onto the object to add code information, its feature points can be easily identified exclusively taking advantage of the window unique property of pseudo-random array. Then, the 3D coordinates of object both on camera plane and code plane can be obtained by decoding process, which will provide the foundation for further 3D reconstruction. Result of simulation shows that this method is easy to operate, calculation simple and of high matching precision for object feature points matching.