We present an algorithm to get high-speed video using camera array with good perceptual quality in realistic scenes that may have clutter and complex background. We synchronize the cameras such that each captures an image at a different time offset. The algorithm processes the jittery interleaved frames and produces a stabilized video. Our method consists of: synthesis of views from a virtual camera to correct for differences in cameras perspectives, and video compositing to remove remaining artifacts especially around disocclusions. More explicitly, we process the optical flow of the raw video to estimate, for each raw frame, the disparity to the target virtual frame. We input these disparities to content-aware warping to synthesize the virtual views, significantly alleviating the jitter. Yet, while the warping fills the disocclusion holes, the filling may not be coherent temporally, leading to small jitter still visible in static/slow regions around large disocclusions. However, these regions don't benefit from high rate in high-speed video. Therefore, we extract low frame rate regions from only one camera and video composite them with the remaining highly moving regions taken by all cameras. The final video is smooth and efficiently has high frame rate in high motion regions.
We propose a novel highly efficient method for filling disparity holes, regions where disparity estimation fails to produce correct result, with the most plausible values. While the filling values may not exactly match the missing disparities, the filled disparity map has cohesive and smooth areas. Such disparity map enables many applications, such as refocusing and layer effects and other new 3D photography apps, to overcome artifacts due to holes and be visually pleasant for the user. To solve for the filling disparities, we incorporate the visual saliency in our model and decouple the solution complexity from the resolution of the original disparity map. Hence, our technique strikes a good balance between perceptual quality and computational efficiency. Overall, our method produces high quality results fulfilling or exceeding the requirements of practical applications that use depth. Moreover, it is fast and hence adequate to run on the ubiquitous mobile platforms.
The two volume set LNCS 8887 and 8888 constitutes the refereed proceedings of the 10th International Symposium on Visual Computing, ISVC 2014, held in Las Vegas, NV, USA. The 74 revised full papers an
In this paper a fast triangular mesh based registration method is proposed. Having Template and Reference images as inputs, the template image is triangulated using a content adaptive mesh generation algorithm. Considering the pixel values at mesh nodes, interpolated using spline interpolation method for both of the images, the energy functional needed for image registration is minimized. The minimization process was achieved using a mesh based discretization of the distance measure and regularization term which resulted in a sparse system of linear equations, which due to the smaller size in comparison to the pixel-wise registration method, can be solved directly. Mean Squared Difference (MSD) is used as a metric for evaluating the results. Using the mesh based technique, higher speed was achieved compared to pixelbased curvature registration technique with fast DCT solver. The implementation was done in MATLAB without any specific optimization. Higher speeds can be achieved using C/C++ implementations.
The widespread success of Kinect enables users to acquire both image and depth information with satisfying accuracy at relatively low cost. We leverage the Kinect output to efficiently and accurately estimate the camera pose in presence of rotation, translation, or both. The applications of our algorithm are vast ranging from camera tracking, to 3D points clouds registration, and video stabilization. The state-of-the-art approach uses point correspondences for estimating the pose. More explicitly, it extracts point features from images, e. g., SURF or SIFT, and builds their descriptors, and matches features from different images to obtain point correspondences. However, while features-based approaches are widely used, they perform poorly in scenes lacking texture due to scarcity of features or in scenes with repetitive structure due to false correspondences. Our algorithm is intensity-based and requires neither point features' extraction, nor descriptors' generation/matching. Due to absence of depth, the intensity-based approach alone cannot handle camera translation. With Kinect capturing both image and depth frames, we extend the intensity-based algorithm to estimate the camera pose in case of both 3D rotation and translation. The results are quite promising.
The popularity of mobile photography paves the way to create new ways of viewing, interacting and enabling a user's creative expression with personal media. In this paper, we describe an instantaneous and automatic method to localize the camera and enable segmentation of foreground objects such as people from an input image, assuming knowledge of the environment in which the image was taken. Camera localization is performed by comparing multiple views of the 3D environment against the uncalibrated input image. Following localization, selected views of the 3D environment are aligned, color-mapped and compared against the input image to segment the foreground content. We demonstrate results using our proposed system in two illustrative applications: a virtual game played between multiple users involving virtual projectiles and a group shot of multiple people who may not be available simultaneously at the same time or place created against a background of their choice.
We describe an augmented reality prototype for exploring a 3D urban environment on mobile devices. Our system utilizes the location and orientation sensors on the mobile platform as well as computer vision techniques to register the live view of the device with the 3D urban data. In particular, the system recognizes the buildings in the live video, tracks the camera pose, and augments the video with relevant information about the buildings in the correct perspective. The 3D urban data consist of 3D point clouds and corresponding geo-tagged RGB images of the urban environment. We also discuss the processing steps to make such 3D data scalable and usable by our system.
La presente invention se rapporte a des systemes, a des dispositifs et a des procedes adaptes pour recevoir une image source comprenant une partie d'avant-plan et une partie d'arriere-plan, la partie d'arriere-plan comprenant un contenu d'image d'un environnement en trois dimensions (3D). Selon la presente invention, une pose de camera de l'image source peut etre determinee en comparant des caracteristiques de l'image source a des caracteristiques d'image d'une image cible de l'environnement en 3D. D'autre part, l'utilisation de la pose de la camera dans le but de segmenter la partie d'avant-plan de la partie d'arriere-plan peut generer une image source segmentee. Ensuite, l'image source segmentee ainsi obtenue et la pose associee de la camera peuvent etre enregistrees dans une base de donnees en reseau. Enfin, la pose de la camera et l'image source segmentee peuvent etre utilisees afin de realiser une simulation de la partie d'avant-plan dans un environnement virtuel en 3D.
In this paper, we present a large-scale mobile augmented reality system that recognizes the buildings in the mobile device's live video and registers this live view with the 3-dimensional models of the buildings. Having the camera pose estimated and tracked, the system adds relevant information about the buildings to the video in the correct perspective. We demonstrate the system on a large database of geo-tagged panoramic images of an urban environment with associated 3-dimensional planar models. The system uses the capabilities of emerging mobile platforms such as location and orientation sensors, and computational power to detect, track, and augment buildings in urban scenes.
We present an augmented reality tourist guide on mobile devices. Many of latest mobile devices contain cameras, location, orientation and motion sensors. We demonstrate how these devices can be used to bring tourism information to users in a much more immersive manner than traditional text or maps. Our system uses a combination of camera, location and orientation sensors to augment live camera view on a device with the available information about the objects in the view. The augmenting information is obtained by matching a camera image to images in a database on a server that have geotags in the vicinity of the user location. We use a subset of geotagged English Wikipedia pages as the main source of images and augmenting text information. At the time of publication our database contained 50 K pages with more than 150 K images linked to them. A combination of motion estimation algorithms and orientation sensors is used to track objects of interest in the live camera view and place augmented information on top of them.
Sensitivity analysis attacks aim at estimating a watermark from multiple observations of the detector's output. Subsequently, the attacker removes the estimated watermark from the watermarked signal. In order to measure the vulnerability of a detector against such attacks, we evaluate the fundamental performance limits for the attacker's estimation problem. The inverse of the Fisher information matrix provides a bound on the covariance matrix of the estimation error. A general strategy for the attacker is to select the distribution of auxiliary test signals that minimizes the trace of the inverse Fisher information matrix. The watermark detector must trade off two conflicting requirements: (1) reliability, and (2) security against sensitivity attacks. We explore this tradeoff and design the detection function that maximizes the trace of the attacker's inverse Fisher information matrix while simultaneously guaranteeing a bound on the error probability. Game theory is the natural framework to study this problem, and considerable insights emerge from this analysis.
Despite their popularity, spread spectrum schemes are vulnerable against sensitivity analysis attacks on standard deterministic watermark detectors. A possible defense is to use a randomized watermark detector. While randomization sacrifices some detection performance, it might be expected to improve detector security to some extent. This paper presents a framework to design randomized detectors with exponentially large randomization space and controllable loss in detection reliability. We also devise a general procedure to attack such detectors by reducing them into equivalent deterministic detectors. We conclude that, contrary to prior belief, randomization of the detector is not the ultimate answer for providing security against sensitivity analysis attacks in spread spectrum systems. Instead, the randomized detector inherits the weaknesses of the equivalent deterministic detector.
Igor Kozintsev合作论文数Intel Microprocessor Research Lab12
Chandra Kambhamettu合作论文数University of Delaware;Department of Computer and Information Sciences1