Foreground segmentation technologies play an important role in applications such as free-viewpoint video (FVV) and sports video analysis. In this situation, we propose a new method that achieves accurate foreground silhouette extraction method using dynamic object presence probability (DOPP). Our main contributions are as follows. 1) Object presence probability for each pixel is calculated from the object recognition results based on deep learning. After that, background subtraction is implemented by changing the threshold and the update rate of the background model in response to the object presence probability. Parameter tuning of background subtraction is executed by using the object recognition results to improve the silhouette extraction quality. 2) To calculate more accurate silhouette i mages, parameters of background subtraction are adjusted by monitoring optical flows between consecutive frames. The object presence probability of the current frame is dynamically updated by using the object presence probability of the previous frame with optical flows. In the experiments, we confirmed that the proposed method achieved more accurate silhouette extraction than conventional methods in three sports sequences.
针对视觉式头部姿态测量系统存在编码特征点安装空间需求大的问题,设计一种结构简单、需求空间小的瞄准线测量方法.使用定位摄像机拍摄头盔上的编码特征点图像,计算机对图像进行解析和计算,结合正交迭代(OI)算法解算头部姿态.根据编码特征点随意布局特点进行了多组多次的仿真分析,仿真数据表明此方法是合理、有效的,利用小空间布局的编码特征点可以解算出高精度瞄准线.
In this paper, we report on a parallel free-viewpoint video synthesis algorithm that can efficiently reconstruct a high-quality 3D scene representation of sports scenes. The proposed method focuses on a scene that is captured by multiple synchronized cameras featuring wide-baselines. The following strategies are introduced to accelerate the production of a free-viewpoint video taking the improvement of visual quality into account: (1) a sparse point cloud is reconstructed using a volumetric visual hull approach, and an exact 3D ROI is found for each object using an efficient connected components labeling algorithm. Next, the reconstruction of a dense point cloud is accelerated by implementing visual hull only in the ROIs; (2) an accurate polyhedral surface mesh is built by estimating the exact intersections between grid cells and the visual hull; (3) the appearance of the reconstructed presentation is reproduced in a view-dependent manner that respectively renders the non-occluded and occluded region with the nearest camera and its neighboring cameras. The production for volleyball and judo sequences demonstrates the effectiveness of our method in terms of both execution time and visual quality.
Free viewpoint technologies that synthesize virtual viewpoint by using multiple actual videos are one of the hottest topics in the video processing field, and would provide immersive experiences for users. To realize this concept, many conventional methods have been proposed. However, these methods require high computational cost to synthesize a virtual viewpoint because they have to calculate huge data to express three-dimensional information, e.g. the shapes of objects, only by using two-dimensional video. This makes them inadequate for end-to-end (from video capture to virtual view rendering) live streaming. To overcome this problem, we propose a simple and fast free-viewpoint synthesis method based on a visual hull, which is a general concept in this field. We calculate the silhouette of an object along planes in virtual space by simple projection from video images to the 3D space, and express the whole shape of the object by integrating the planes. This scheme works very quickly while providing fine quality because it consists of standard functions in general GPU architecture. The experimental results show our method can generate a fine virtual view of an object by using multiple videos in real time.
A free viewpoint application has been developed that yields an immersive user experience. The free viewpoint approach called the "billboard methodis" suitable for displaying a synthesized 3D view in a mobile device, but it suffers from the limitation that a billboard cannot present an accurate impression of depth for a foreground object, and it gives users an unacceptable impression from certain virtual viewpoints. To solve this problem, we propose the optimal deformation of the billboard. The deformation is designed as a mapping of grid points in the input billboard silhouette to produce an optimal silhouette from an accurate voxel model of the object. We formulate and solve this procedure as a nonlinear optimization problem based on a grid-point constraint and some a priori information. Our results show that the optimal deformation is produced by the proposed method, which generates a synthesized virtual image having a natural appearance.
This paper proposes a new parallel approach to solve connected components on a 2-D binary image. The following strategies are employed to accelerate neighborhood exploration after dividing an input image into independent blocks: 1) in the local labeling stage, a coarse-labeling algorithm, including row-column connection and unification, is applied first to reduce the complexity of an initialized local label map; a refinement algorithm is then introduced to merge separated sub-regions from a single component; and 2) in the block merge stage, we scan the pixels on the block boundary instead of solving the connectivity of all the pixels. With the proposed method, the length of label-equivalence lists in both the local labeling stage and global labeling stage are compressed and the number of memory accesses is reduced. Thus, the efficiency of connected component labeling is improved. The proposed strategies are illustrated using 4-neighbor connectivity, and the case of 8-neighbor connectivity is also discussed. The YACCLAB data sets, including both synthetic and real images, are used to evaluate the new algorithm and compare it to existing algorithms. The comparative results show that the proposed new algorithm outperforms the other approaches in both the 4-neighbor connectivity and 8-neighbor connectivity cases.
A free-viewpoint application has been developed that yields an immersive user experience. One of the simple free-viewpoint approaches called “billboard methods” is suitable for displaying a synthesized 3D view in a mobile device, but it suffers from the limitation that a billboard should be positioned in only one position in the world. This fact gives users an unacceptable impression in the case where an object being shot is situated at multiple points. To solve this problem, we propose optimal deformation of the billboard. The deformation is designed as a mapping of grid points in the input billboard silhouette to produce an optimal silhouette from an accurate voxel model of the object. We formulate and solve this procedure as a nonlinear optimization problem based on a grid-point constraint and some a priori information. Our results show that the proposed method generates a synthesized virtual image having a natural appearance and better objective score in terms of the silhouette and structural similarity.
The paper proposes an algorithm to robustly reconstruct an accurate billboard model of an individual object including an occluded one in each camera. Each billboard model is utilized to synthesize high-quality, free-viewpoint video especially for outdoor sport scenes in which roughly calibrated cameras are sparsely placed. The two main contributions of the proposed algorithm are (1) robustness to occlusions caused by overlaps of multiple objects in every camera, that is one of the biggest issues for billboard-based method, and (2) applicability to challenging shooting conditions in which accurate 3D model cannot be reconstructed because of calibration errors, small number of cameras and so on. In order to achieve the contributions above, the algorithm does not try to reproduce an accurate 3D model of each object but utilize a "rough 3D model". The algorithm precisely extracts an individual object region in every camera by reconstructing a "rough 3D model" of each object and back-projecting it to every camera. The 3D coordinate for each billboard to be located is calculated based on the position of a rough 3D model. Experimental results compare the visual quality of free-viewpoint videos synthesized with our proposed method and conventional methods and show the effectiveness of our proposed method in terms of the naturalness of positional relationships and the fineness of the surface textures of all the objects.
In this paper, we report an optimized union-find (UF) algorithm that can label the connected components on a 2D image efficiently by employing the GPU architecture. The proposed method contains three phases: UF-based local merge, boundary analysis, and link. The coarse labeling in local merge reduces the number atomic operations, while the boundary analysis only manages the pixels on the boundary of each block. Evaluation results showed that the proposed algorithm speed up the average running time by more than 1.3X.
This paper describes a super high-speed vision platform (HSVP) that can process a 12-bit, gray level, 1024 ×1024 image at 12500 frames per second (fps) simultaneously by implementing parallel hardware logic on a large size field programmable gate array (FPGA) platform. Multiple experimental results demonstrate that our platform can calculate a 4096-level brightness-histogram and 4096-level edge-intensity-histogram for 1024×1024 images at 12500 fps, including the zeroth and first moment features for centroid calculation.
Over-segmentation of a grayscale image is a typical problem in existing watershed algorithms. To overcome this problem, preprocessing is mainly applied to the grayscale image before performing the watershed transformation to generate a gradient or binary image. In this paper, a novel watershed algorithm based on the concept of connected-component labeling and chain code is proposed, which generates a final label map in just four scans of a preprocessed binary image. The low memory consumption, low complexity, and simple data structure of the algorithm make it highly suitable for hardware implementation. Evaluation results showed that the proposed algorithm decreases the average running time by more than 39% without loss of accuracy.
This paper reports on the development of a fast 3D shape scanner that can output 3D video at 250 fps using two high-frame-rate camera-projector systems with an implementation of 10-bit Gray code light pattern encoded in both horizontal and vertical. The 3D data, which is extracted by the two camera-projector systems and accelerated by installing a GPU board for parallel processing of structured light illumination, are registered together to obtain an entire shape. To avoid the interference of projection patterns, the high-speed vision platform used in this paper is dedicated and improved for dual-camera frame-straddling by developing a hardware logics for the time delay control between the two cameras, while the two projectors are synchronized with the two cameras using impulse signal. The effectiveness is demonstrated through several 3-D shape measurements when several moving objects and static objects are observed by the proposed scanner.
Blink-spot projection method We present a blink-spot projection method for observing moving three-dimensional (3D) scenes. The proposed method can reduce the synchronization errors of the sequential structured light illumination, which are caused by multiple light patterns projected with different timings when fast-moving objects are observed. In our method, a series of spot array patterns, whose spot sizes change at different timings corresponding to their identification (ID) number, is projected onto scenes to be measured by a high-speed projector. Based on simultaneous and robust frame-to-frame tracking of the projected spots using their ID numbers, the 3D shape of the measuring scene can be obtained without misalignments, even when there are fast movements in the camera view. We implemented our method with a high-frame-rate projector-camera system that can process 512 × 512 pixel images in real-time at 500 fps to track and recognize 16 × 16 spots in the images. Its effectiveness was demonstrated through several 3D shape measurements when the 3D module was mounted on a fast-moving six-degrees-of-freedom manipulator.
In this paper, the authors report on the development of a projection-mapping system that can project RGB light patterns that are enhanced for three dimensional (3D) scenes using a graphics processing unit (GPU) based high-frame-rate (HFR) vision system synchronized with HFR projectors. The proposed system can acquire 512 × 512 depth-images in real time at 500 fps. The depth-images processing is accelerated by installing a GPU board for parallel processing of Gray-code structured light illumination using infrared (IR) light patterns projected from an IR projector. Using the computed depth-image, suitable RGB light patterns to be projected are generated in real time for enhanced application tasks. They are projected from an RGB projector as augmented information onto a 3D scene with pixel-wise correspondence even when the 3D scene is time-varied. Experimental results obtained from enhanced application tasks for time-varying 3D scenes such as (1) depth-based color mapping, (2) augmented reality (AR) spirit level and (3) AR wristwatch confirm the efficacy of our system.
A high-frame-rate (HFR) structured light vision is developed for observing moving three-dimensional (3-D) scenes; it is mountable on the end of a robot manipulator for 3-D shape inspection. Our system can simultaneously obtain depth images of 512×512 pixels at 500 fps by implementing a motion-compensated coded structured light method on an HFR camera-projector platform; the 3-D computation is accelerated using the parallel processing on a GPU board. This method can remarkably reduce the synchronization errors in the structured-light-based measurement, which are encountered in the projection of multiple light patterns with different timings; such synchronization errors become larger as the ego-motion of a manipulator becomes larger. We demonstrate the performance of our system by showing several 3-D shape measurement results when the 3-D module is mounted on a fast-moving 6-DOF manipulator as a sensing head.
In this paper, we report on the development of a projection mapping system that can project RGB light patterns that are enhanced for three-dimensional (3-D) scenes using a GPU-based high-frame-rate (HFR) vision system synchronized with HFR projectors. Our system can acquire 512×512 depth images in real time at 500 fps. The depth image processing is accelerated by installing a GPU board for parallel processing of a gray-code structured light method using infrared (IR) light patterns projected from an IR projector. Using the computed depth images, suitable RGB light patterns to be projected are generated in real time for enhanced application tasks. They are projected from an RGB projector as augmented information onto a 3-D scene with pixel-wise correspondence even when the 3-D scene is time-varied. Experimental results obtained from enhanced application tasks for time-varying 3-D scenes such as (1) depth-based color mapping and (2) augmented reality (AR) spirit level, confirm the efficacy of our system.
We propose a novel dot-pattern-projection three-dimensional (3-D) shape measurement method that can measure 3-D displacements of blink dots projected onto a measured object accurately even when it moves rapidly or is observed from a camera as moving rapidly. In our method, blinking dot patterns, in which each dot changes its size at different timings corresponding to its identification (ID) number, are projected from a projector at a high frame rate. 3-D shapes can be obtained without any miscorrespondence of the projected dots between frames by simultaneous tracking and identification of multiple dots projected onto a measured 3-D object in a camera view. Our method is implemented on a field-programmable gate array (FPGA)-based high-frame-rate (HFR) vision platform that can track and recognize as much as 15×15 blink-dot pattern in a 512×512 image in real time at 1000 fps, synchronized with an HFR projector. We demonstrate the performance of our system by showing real-time 3-D measurement results when our system is mounted on a parallel link manipulator as a sensing head.