
Plenoptic cameras enable the capturing of spatial as well as angular color information which can be used for various applications among which are image refocusing and depth calculations. However, these cameras are expensive and research in this area currently lacks data for ground truth comparisons. In this work we describe a flexible, easy-to-use Blender model for the different plenoptic camera types which is on the one hand able to provide the ground truth data for research and on the other hand allows an inexpensive assessment of the cameras usefulness for the desired applications. Furthermore we show that the rendering results exhibit the same image degradation effects as real cameras and make our simulation publicly available.
Video synchronization is a fundamental computer-vision task that is necessary for a wide range of applications. A 3D video involves two streams, which show the scene from different angles concurrently, but many cases exhibit desynchronization between them. This paper investigates the problem of synchronizing the left and right stereoscopic views. We assume the temporal shift (time difference) and geometric distortion between the two streams are constant throughout each scene. We propose a temporal-shift estimation method with subframe accuracy based on a block-matching algorithm.
Occlusion filling is a basic problem for multiview video generation from existing monocular video. The essential goal of this problem is to recover missing information about a scenes 3D structure and corresponding texture. We propose a method for content-aware deformation of the source view that ensures no disoccluded regions are visible in the synthesized views while also keeping visible distortions to a minimum. We formulate this problem in terms of global energy minimization. Furthermore, we introduce a similar variable-rejection algorithm that, along with other known optimization techniques, allows us to accelerate the energy function minimization by nearly 30 times and still maintain the visual quality of the synthesized views.
There are many basic ways of providing a glasses-free 3D display and the three methods considered most likely to succeed commercially were chosen for our current research, these are; multi-layer light field, head tracked and super multiview displays. Our multi-layer light field display enables a far smaller form factor than other types, and faster algorithms along with horizontal parallax-only will considerably speed-up computation time. A spin-off of this technology is a near-eye display that provides focus cues for maximizing user comfort. Head tracked displays use liquid crystal display panels illuminated with a directional backlight to produce multiple sets of exit pupil pairs that follow the user's eyes under the control of a head position tracker. Our super multiview display (SMV) system uses high frame-rate projectors for spatio-temporal multiplexing that give dense viewing zones with no accommodation/convergence (A/C) conflict. Bandwidth reduction is achieved by discarding redundant information at capture. The status of the latest prototypes and their performance is described; and we conclude by indicating the future directions of our research.
We propose an iterative closest point (ICP) based calibration for time of flight (ToF) multiple depth sensors. For the multiple sensor calibrations, we usually use 2D patterns calibration with IR images. The depth sensor output depends on calibration parameters at a factory; thus, the re-calibration must include gaps from the calibration in the factory. Therefore, we use direct correspondences among depth values, and the calibrating extrinsic parameters by using ICP. Usually, simultaneous localization and mapping (SLAM) uses ICP, such as KinectFusion. The case of multiple sensor calibrations, however, is harder than the SLAM case. In this case, the distance between cameras is too far to apply ICP. Therefore, we modify the ICP based calibration for multiple sensors. The proposed method uses specific calibration objects to enforce the matching ability among sensors. Also, we proposed a compensation method for ToF depth map distortions.
Light-field visualization is continuously emerging in industrial sectors, and the appearance on the consumer market is approaching. Yet this process is halted, or at least slowed down, by the lack of proper display-independent light-field formats. Such formats are necessary to enable the efficient interchange between light-field content creation and visualization, and thus support potential future use case scenarios of this technology. In this paper, we introduce the results of a perceived quality assessment research, performed on our own novel light-field visualization format. The subjective tests, which compared conventional linear camera array visualization to our format, were completed by experts only, thus quality assessment was an expert evaluation. We aim to use the findings gathered in this research to carry out a large-scale subjective test series in the future, with non-expert observers.
With the rapid advances in light field displays and cameras, research in light field content creation, visualization, coding and quality assessment is now beyond a state of emergence; it has already emerged and started attracting a significant part of the scientific community. The capability of light field displays to offer glasses-free 3D experience simultaneously for multiple users has opened new avenues in subjective and objective quality assessment of light field image content, and video is also becoming research target of such quality evaluation methods. Yet it needs to be stated that while static light field contents have evidently received relatively more attention, the research on light field video content still remains largely unexplored. In this paper, we present results of the objective quality assessment of key frames extracted from light field video content. To this end, we use our own full-reference 3D objective quality metric.
In this paper we present a novel method for generating dense reconstructions by applying only structure-from-motion(SfM) on large-scale datasets without the need for multi-view stereo as a post-processing step. A state-of-the-art optical flow technique is used to generate dense matches. The matches are encoded such that verification for correctness becomes possible, and are stored in a database on-disk. The use of this out-of-core approach transfers the requirement for large memory space to disk, therefore allowing for the processing of even larger-scale datasets than before. We compare our approach with the state-of-the-art and present the results which verify our claims.
Channel mismatch (the result of swapping left and right views) is a 3D-video artifact that can cause major viewer discomfort. This work presents a novel high-accuracy method of channel-mismatch detection. In addition to the features described in our previous work, we introduce a new feature based on a convolutional neural network; it predicts channel-mismatch probability on the basis of the stereoscopic views and corresponding disparity maps. A logistic-regression model trained on the described features makes the final prediction. We tested this model on a set of 900 stereoscopic-video scenes, and it outperformed existing channel-mismatch detection methods that previously served in analyses of full-length stereoscopic movies.
The design of micro photon sieve arrays (PSAs) is investigated for light-field capture with high spatial resolution in plenoptic cameras. A commercial very high-resolution full-frame camera with a manual lens is converted into a plenoptic camera for high-resolution depth image acquisition by using the designed PSA as an add-on diffractive optical element in place of an ordinary refractive microlens array or a diffractive micro Fresnel Zone Plate (FZP) array, which is used in integral imaging applications. The noise introduced by the diffractive nature of the optical element is reduced by standard image processing tools. The light-field data is also used for computational refocusing of the 3D scene with wave propagation tools.
Recording and imaging the 3D world has led to the use of light fields. Capturing, distributing and presenting light field data is challenging, and requires an evaluation platform. We define a framework for real-time processing, and present the design and implementation of a light field evaluation system. In order to serve as a testbed, the system is designed to be flexible, scalable, and able to model various end-to-end light field systems. This flexibility is achieved by encapsulating processes and devices in discrete framework systems. The modular capture system supports multiple camera types, general-purpose data processing, and streaming to network interfaces. The cloud system allows for parallel transcoding and distribution of streams. The presentation system encapsulates rendering and display specifics. The real-time ability was tested in a latency measurement; the capture and presentation systems process and stream frames within a 40 ms limit.
Sign language recognition (SLR) is a challenging, but highly important research field for several computer vision systems that attempt to facilitate the communication among the deaf and hearing impaired people. In this work, we propose an accurate and robust deep learning-based methodology for sign language recognition from video sequences. Our novel method relies on hand and body skeletal features extracted from RGB videos and, therefore, it acquires highly discriminative for gesture recognition skeletal data without the need for any additional equipment, such as data gloves, that may restrict signer's movements. Experimentation on a large publicly available sign language dataset reveals the superiority of our methodology with respect to other state of the art approaches relying solely on RGB features.
A key challenge when displaying and processing sensed real-time 3D data is efficiency of generating and post-processing algorithms in order to acquire high quality 3D content. In contrast, our approach focuses on volumetric generation and processing volumetric data using an efficient low-cost hardware setting. Acquisition of volumetric data is performed by connecting several Kinect v2 scanners to a single PC that are subsequently calibrated using planar pattern. This process is by no means trivial and requires well designed algorithms for fast processing and quick rendering of volumetric data. This can be achieved by fusing efficient filtering methods such as Weighted median filter (WM), Radius outlier removal (ROR) and Laplace-based smoothing algorithm. In this context, we demonstrate the robustness and efficiency of our technique by sensing several scenes.
The plenoptic camera is gaining more and more attention as it captures the 4D light field of a scene with a single shot and enables a wide range of post-processing applications. However, the pre-processing steps for captured raw data, such as demosaicing, have been overlooked. Most existing decoding pipelines for plenoptic cameras still apply demosaicing schemes which are developed for conventional cameras. In this paper, we analyze the sampling pattern of microlens-based plenoptic cameras by ray-tracing techniques and ray phase space analysis. The goal of this work is to demonstrate guidelines and principles for demosaicing the plenoptic captures by taking the unique microlens array design into account. We show that the sampling of the plenoptic camera behaves differently from that of a conventional camera and the desired demosaicing scheme is depth-dependent.
A key role in the advancement of 3 Dimensional TV services is played by the development of 3D video quality metrics used for the assessment of the perceived quality. Moreover, this key role can only be supported when the features associated with the 3D video nature is reliably and efficiently characterized in these metrics. In this study, z-direction motion incorporated with significant depth levels in depth map sequences are considered as the main characterizations of the 3D nature. The 3D video quality metrics can be classified into three categories based on the need for the reference video during the assessment process at the user end: Full Reference (FR), Reduced Reference (RR) and No Reference (NR). In this study we propose a NR quality metric, PNRM, suitable for on-the-fly 3D video services. In order to evaluate the reliability and effectiveness of the proposed metric, subjective experiments are conducted in this paper. Observing the high correlation with the subjective experimental results, it can be clearly stated that the proposed metric is able to mimic the Human Visual System (HVS).
Many factors can cause color distortions between stereoscopic views during 3D-video shooting. Numerous viewers experience discomfort and headaches when watching stereoscopic videos that contain such distortions. In addition, 3D videos with color differences are hard to process because many algorithms assume brightness constancy. We propose an automatic method for correcting color distortions between stereoscopic views and compare it with analogs. The comparison shows that our proposed method combines high color-correction accuracy with relatively low computational complexity.
In the paper an adaptive color correction method for virtual view synthesis is presented. It deals with the typical problem in free navigation systems - different illumination in views captured by different cameras acquiring the scene. The proposed technique adjusts the local color characteristics of objects visible in two real views. That approach allows to significantly reduce number and visibility of color artifacts in the virtual view. Proposed method was tested on 12 multiview test sequences. Obtained and presented in the paper results show, that proposed color correction provides increase of the virtual view quality measured by PSNR, SSIM and subjective evaluation.
We present an accurate model of integral imaging display based on wave optics. The model enables accurate characterization of the display through simulated perceived images by the human visual system. Thus, it is useful to investigate the capabilities of the display in terms of various quality factors such as depth of field and resolution, as well as delivering visual cues such as focus. Furthermore, due to the adopted wave optics formalism, simulation and analysis of more advanced techniques such as wavefront coding for increased depth of field are also possible.
The capturing of angular and spatial information of the scene using single camera is made possible by new emerging technology referred to as plenoptic camera. Both angular and spatial information, enable various post-processing applications, e.g. refocusing, synthetic aperture, super-resolution, and 3D scene reconstruction. In the past, multiple traditional cameras were used to capture the angular and spatial information of the scene. However, recently with the advancement in optical technology, plenoptic cameras have been introduced to capture the scene information. In a plenoptic camera, a lenslet array is placed between the main lens and the image sensor that allows multiplexing of the spatial and angular information onto a single image, also referred to as plenoptic image. The placement of the lenslet array relative to the main lens and the image sensor, results in two different optical designs of a plenoptic camera, also referred to as plenoptic 1.0 and plenoptic 2.0. In this work, we present a novel dataset captured with plenoptic 1.0 (Lytro Illum) and plenoptic 2.0 (Raytrix R29) cameras for the same scenes under the same conditions. The dataset provides the benchmark contents for various research and development activities for plenoptic images.
The simplicity of the holographic stereogram (HS) makes it an attractive option in comparison to the more complex coherent computer generated hologram (CGH) methods. The cost of its simplicity is that the HS cannot accurately reconstruct deep scenes due to the lack of correct accommodation cues. The exact nature of the accommodation cues present in HSs, however, has not been investigated. In this paper, we provide analysis of the relation between the hologram sampling properties and the perceived accommodation response. The HS can be considered as a generator of a discrete light field (LF) and can thus be examined by considering the light ray oriented nature of the hologram diffracted light. We further support the analysis by employing a numerical reconstruction tool simulating the viewing process of the human eye. The simulation results demonstrate that HSs can provide accommodation cues depending on the choice of hologram segmentation size. It is further demonstrated that the accommodation response can be enhanced at the expense of loss in perceived spatial resolution.