Video data are of two different intrinsic modes, in‐frame and temporal. It is beneficial to incorporate static in‐frame features to acquire dynamic features for video applications. However, some existing methods such as recurrent neural networks do not have a good performance, and some other such as 3D convolutional neural networks (CNNs) are both memory consuming and time consuming. This study proposes an effective framework that takes the advantage of deep learning on the static image feature extraction to tackle the video data. After extracting in‐frame feature vectors using a pretrained deep network, the authors integrate them and form a multi‐mode feature matrix, which preserves the multi‐mode structure and high‐level representation. They propose two models for follow‐up classification. The authors first introduce a temporal CNN, which directly feeds the multi‐mode feature matrix into a CNN. However, they show that characteristics of the multi‐mode features differ significantly in distinct modes. The authors therefore further propose the multi‐mode neural network (MMNN), in which different modes deploy different types of layers. They evaluate their algorithm with the task of human action recognition. The experimental results show that the MMNN achieves a much better performance than the existing long short‐term memory‐based methods and consumes far fewer resources than the existing 3D end‐to‐end models.
Recently deep neural networks have been successfully used for natural image deconvolution. Whereas the existing methods usually involve an inversion of the blur followed by a denoising step. In this paper we propose a pure learning approach to learn a mapping from a blurred patch to a clean patch directly with a deep dual-pathway rectifier neural network. The experimental results show that our approach outperform the state-of-the-art methods on non-blind image deconvolution within reasonable training time. By analyzing the learned representations, we empirically show that our model works by efficiently detecting the blurry input patterns and then reconstructing the clean patch with the corresponding dictionary atoms.
This paper deals with automatic human action recognition in videos. Rather than considering traditional hand-craft features such as HOG, HOF and MBH, we explore how to learn both static and motion features from CNNs trained on large-scale datasets such as ImagNet and UCF101. We propose a novel method named multi-resolution latent concept descriptor (mLCD) to encode two-stream CNNs. Entensive experiments are conducted to demonstrate the performance of the proposed model. By combining our mLCD features with the improved dense trajectory features, we can achieve comparable performance with state-of-the-art algorithms on both Hollywood2 and Olympic Sports datasets.
Depth of Field is an important factor to synthesize realistic photography effects. In our paper, we discuss a new image-based approach that can synthesize DoF effects without user iterations. Our approach can produce the based saliency maps with realtime performance. In particular, there is not depth information needed to capture from cameras. The depth information is approximated by using saliency maps. In particular, we take advantages of the flash non-flash image pairs to refine the DoF synthesis quality. The focused regions can be segmented using GrabCut. The experimental results show that the our image based DoF synthesis can simulate high-quality DoF effects efficiently.
Depth of Field (DoF) is an indispensable feature of photo realistic rendering and photography retouching. In this paper, we propose an image-based rendering technique which can simulate the depth-of-field effect. The proposed technique can render the depth-of-field effect automatically without any interactions. Compared to the ordinary depth-of-field rendering technique, our algorithm is less time-consuming and needs no additional depth maps to assist the depth-of-field rendering. In our proposed algorithm, the saliency detection technique is employed to simulate the depth information. The flash-based technique is also introduced to promote the final depth-of-field rendering visual effect.
We present a new approach to simulate depth-of-field effect with non-local means filtering in an unsupervised fashion which requires no user input at all. Unlike traditional rendering methods which handle the depth-of-field effect in 3D scenes, our novel approach handles the depth-of-field effect in 2D scenes. Our proposed approach mainly consists of two steps: first we extract the depth of field (DoF) region from the background in the form of alpha matte, which based on saliency detection technique and spectral matting. Then we blur the background with non-local means depth filtering, which is heavily used in image de-noising. We demonstrate our approach can precisely and efficiently simulate depth-of-field effect of images without prior knowledge of their content.
Liqing Zhang (张丽清)合作论文数Department of Computer Science and Engineering, Shanghai Jiao Tong University;Center for Brain-like Computing and Machine Intelligence, Shanghai Jiao Tong University3