We present a fast and efficient volumetric capture and reconstruction system that processes either RGB-D or RGB-only input to generate 3D representations in the form of point clouds and Gaussian splats. For Gaussian splat reconstructions, we took the GPS-Gaussian regressor and improved it, enabling high-quality reconstructions with minimal overhead. The system is designed for easy setup and deployment, supporting in-the-wild operation under uncontrolled illumination and arbitrary backgrounds, as well as flexible camera configurations, including sparse setups, arbitrary camera numbers and baselines. Captured data can be exported in standard formats such as PLY, MPEG V-PCC, and SPLAT, and visualized through a web-based viewer or Unity/Unreal plugins. A live on-location preview of both input and reconstruction is available at 5-10 FPS. We present qualitative findings focused on deployability and targeted ablations. The complete framework is open-source, facilitating reproducibility and further research.
Volumetric Video (VV) allows audiences to experience real-world scenery in full 6° of freedom, i.e., selecting individual viewpoints and directions. It is often produced as point clouds or meshes with shape and texture data. This leads to a substantial volume of data that requires efficient compression. However, there is a lack of universally accepted standards for VV, resulting in many organizations implementing their own approaches. Hence, we suggest Volumetric Video Graphics Library Transmittal Format (VVglTF), an extension of the open-source 3D content format glTF, which is optimized for the Internet. We utilize glTF's rendering workflow to play VV of any duration or file size efficiently. Custom extensions in glTF expand the functionality of the basic model and offer customization for VV playback. We also developed a simple and efficient VVglTF streaming system based on HTTP. It is designed to play VV content across different network conditions efficiently. In our experiments, we validated the efficiency of our approach.
Video quality assessment is generally important to assess any kind of immersive media technology, for example, to evaluate a captured content, algorithms for encoding and projection, and systems, as well as for technology optimization. This chapter provides an overview of the two types of video quality assessment, subjective testing with human viewers, and quality prediction or estimation using video quality metrics or models. First, viewing tests with humans as the gold standard for video quality are reviewed in light of their instantiation for omnidirectional video (ODV). In the second part of the chapter, the less time-consuming, better scalable second type of assessment with objective video quality metrics and models is discussed, considering the specific requirements of ODV. Often they incorporate computational models of human perception and content properties. ODV introduces the challenges of interactivity compared to standard 2D video and typically spherical projection distortions due to its omnidirectional, “point-of-view” (in terms of camera-shot type) nature. Accordingly, subjective tests for ODV include specific considerations of the omnidirectional nature of the presented content and dedicated head-rotation or even additional eyetracking data capture. In the last part of the chapter, it is shown how to improve objective video quality prediction by taking into account user behavior and projection distortions.
Color mismatch in stereoscopic 3D (S3D) images can create visual discomfort and affect the performance of S3D image processing algorithms, e.g., for depth estimation. In this paper, we propose a new deep learning-based solution for the problem of color mismatch correction. The proposed solution consists of a multi-task convolutional neural network, where color correction is the primary task and correspondence estimation is the secondary task. For the training and evaluation of the proposed network, a new S3D image dataset with color mismatch was created. Based on this dataset, experiments were conducted showing the effectiveness of our solution.
Omnidirectional video (ODV) enables viewers to look at every direction from a fixed point and provides a much more immersive experience than traditional 2D video. Assessing the video quality is important for delivering ODV to the end-user with the best possible quality. For this goal, two aspects of ODV should be considered. The first is the spherical nature of ODV and the related projection distortions when the ODV is stored in a planar format.The second is the interactive look-around consumption nature of ODV.Related to this aspect, visual attention, that identifies the regions that attract the viewer’s attention, is important for ODV quality assessment. Considering these aspects, in this paper, we study in particular objective full-reference quality assessment for ODV. To this end, we propose a quality assessment framework based on the spherical Voronoi diagram and visual attention. In this framework, a given ODV is subdivided into multiple planar patches with low projection distortions using the spherical Voronoi diagram. Afterwards, each planar patch is analyzed separately by a quality metric for traditional 2D video, obtaining a quality score for each patch. Then, the patch scores are combined based on visual attention into a final quality score. To validate the proposed framework, we create a dataset of ODVs with scaling and compression distortions, and conduct subjective experiments in order to gather the subjective quality scores and the visual attention data for our ODV dataset. The evaluation of the proposed framework based on our dataset shows that both the use of the spherical Voronoi diagram and visual attention are crucial for achieving state-of-the-art performance.
Omnidirectional video (ODV) represents one of the latest and most promising trends in immersive media. The success of ODV depends on the ability to deliver high-quality ODV to the viewers. For this reason, new methods are needed to measure ODV quality that takes into account the interactive look around nature and the spherical representation of ODV. In this paper, we study full-reference objective quality metrics for ODV based on typical encoding distortions in adaptive streaming systems, namely, scaling and compression. The contribution of this paper is three-fold. First, we propose new objective metrics that take into account the unique aspects of ODV. The proposed metrics are based on the subdivision of a given ODV into multiple patches using the spherical Voronoi diagram. Second, we introduce a new dataset of 75 impaired ODVs with different resolutions and compression levels, together with the subjective quality scores gathered during an experiment with 21 participants. Third, we evaluate the proposed Voronoi-based objective metrics using our dataset. The evaluation of the proposed objective metrics and the comparison with existing metrics show that the proposed metrics achieve a better correlation with the subjective scores. The ODV dataset together with the subjective quality scores and the code of the proposed quality metrics are available with this paper.
Shooting a live-action immersive 360-degree experience, i.e. omnidirectional content (ODC) is a technological challenge as there are many technical limitations which need to be overcome, especially for capturing and post-processing in stereoscopic 3D (S3D). In this paper, we introduce a novel approach and entire system for stitching and color mismatch correction and detection in S3D omnidirectional content, which consists of three main modules: pre-processing, spherical color correction and color mismatch evaluation. The system and its individual modules are evaluated on two datasets, including a new dataset which will be publicly available with this paper. We show that our system outperforms the state of the art in color correction of S3D ODC and demonstrate that our spherical color correction module even further improves the results of the state of the art approaches.
In this paper, we study an artifact of stereoscopic 3D (S3D) video called sharpness mismatch (SM), that occurs when one view is more blurred than the other. SM beyond a certain level can create visual discomfort, and consequently degrade the quality of experience. Therefore, it is important to measure the just noticeable sharpness mismatch (JNSM), i.e., the minimal level of SM that is perceived by the human visual system and creates discomfort. The knowledge of the JNSM can be used in the evaluation of the quality of S3D video, and more in general when processing S3D video, like in asymmetric compression. In this paper, we focus in particular on the detection of SM. For this goal, we organized a psychophysical experiment with 23 subjects and a crosstalk-free stereoscopic display in order to gather psychophysical data necessary for the development of a SM detection method. Based on the gathered experiment data, we propose a new SM detection method. The evaluation of this method shows that its performance is close but not better than that of the state-of-the-art methods. Therefore, our goal in the near future is to improve the proposed method.
This paper presents a novel sharpness mismatch detection method for stereoscopic images based on the comparison of edge width histograms of the left and right view. The new method is evaluated on the LIVE 3D Phase II and Ningbo 3D Phase I datasets and compared with two state-of-the-art methods. Experimental results show that the new method highly correlates with user scores of subjective tests and that it outperforms the current state-of-the-art. We then extend the method to stereoscopic omnidirectional images by partitioning the images into patches using a spherical Voronoi diagram. Furthermore, we integrate visual attention data into the detection process in order to weight sharpness mismatch according to the likelihood of its appearance in the viewport of the end-user's virtual reality device. For obtaining visual attention data, we performed a subjective experiment with 17 test subjects and 96 stereoscopic omnidirectional images. The entire dataset including the viewport trajectory data and resulting visual attention maps are publicly available with this paper.
In this paper, we present a novel framework for quality control in cinematic VR (360-video) based on Voronoi patches and saliency which can be used in post-production workflows. Our approach first extracts patches in stereoscopic omnidirectional images (ODI) using the spherical Voronoi diagram. The subdivision of the ODI into patches allows an accurate detection and localization of regions with artifacts. Further, we introduce saliency in order to weight detected artifacts according to the visual attention of end-users. Then, we propose different artifact detection and analysis methods for sharpness mismatch detection (SMD), color mismatch detection (CMD) and disparity distribution analysis. In particular, we took two state of the art approaches for SMD and CMD, which were originally developed for conventional planar images, and extended them to stereoscopic ODIs. Finally, we evaluated the performance of our framework with a dataset of 18 ODIs for which saliency maps were obtained from a subjective test with 17 participants.
Digital restoration of film content that has historical value is crucial for the preservation of cultural heritage. Also, digital restoration is not only a relevant application area of various video processing technologies that have been developed in computer graphics literature but also involves a multitude of unresolved research challenges. Currently, the digital restoration workflow is highly labor intensive and often heavily relies on expert knowledge. We revisit some key steps of this workflow and propose semiautomatic methods for performing them. To do that we build upon state-of-the-art video processing techniques by adding the components necessary for enabling (i) restoration of chemically degraded colors of the film stock, (ii) removal of excessive film grain through spatiotemporal filtering, and (iii) contrast recovery by transferring contrast from the negative film stock to the positive. We show that when applied individually our tools produce compelling results and when applied in concert significantly improve the degraded input content. Building on a conceptual framework of film restoration ensures the best possible combination of tools and use of available materials. (C) 2017 SPIE and IS&T
With the release of new head-mounted displays (HMDs) and new omni-directional capture systems, 360-degree video is one of the latest and most powerful trends in immersive media, with an increasing potential for the next decades. However, especially creating 360-degree content in 3D is still an error-prone task with many limitations to overcome. This paper describes the critical aspects of 3D content creation for 360-degree video. In particular, conflicts of depth cues and binocular rivalry are reviewed in detail, as these cause eye fatigue, headache, and even nausea. Both the reasons for the appearance of the conflicts and how to detect some of these conflicts by objective image analysis methods are detailed in this paper. The latter is the main contribution of this paper and part of long-term research roadmap of the authors in order to provide a comprehensive framework for artifact detection and correction in 360-degree videos. Then, experimental results are demonstrating the performance of the proposed approaches in terms of objective measures and visual feedback. Finally, the paper concludes with a discussion and future work.
Evaluation methodologies provide a better understanding of the relationship between a technique and the image attributes. Metrics are used to evaluate the similarities between images. They may use different approaches depending on what needs to be achieved. If only objective values need to be compared statistics-based metrics are suitable. A number of image comparison metrics have been proposed in the literature that is based mostly on images statistics. This chapter shows the MATLAB code for the root mean square error (RMSE) calculation. It presents the MATLAB code for the computation of mean square error (MSE) between two images. Peak signal to noise ratio (PSNR) is another widely used metric, which takes into account the maximum value of the signal, and can be defined based on MSE. Metrics enable automation and the use of metrics can also be applied directly to methods, for example, to …
In this paper, we present a novel sharpness mismatch detection (SMD) approach for stereoscopic omnidirectional images (ODI) for quality control within the post-production workflow, which is the main contribution. In particular, we applied a state of the art SMD approach, which was originally developed for traditional HD images, and extended it to stereoscopic ODIs. A new efficient method for patch extraction from ODIs was developed based on the spherical Voronoi diagram of evenly distributed points on the sphere. The subdivision of the ODI into patches allows an accurate detection and localization of regions with sharpness mismatch. A second contribution of the paper is the integration of saliency into our SMD approach. In this context, we introduce a novel method for the estimation of saliency maps from viewport data of head-mounted displays (HMD). Finally, we demonstrate the performance of our SMD approach with data collected from a subjective test with 17 participants.
Professional TV studio footage often poses specific challenges to camera calibration due to lack of features and complex camera operation. As available algorithms often fail, we propose a novel approach based on robust tracking of ellipse and line features of a predefined logo. We further devise a predictive and iterative estimation algorithm, which incorporates confidence measures and filtering. Our results validate accuracy and reliability of our approach, demonstrated with challenging professional footage.
Subjective studies showed that most HDR video tone mapping operators either produce disturbing temporal artifacts, or are limited in their local contrast reproduction capability. Recently, both these issues have been addressed by a novel temporally coherent local HDR tone mapping method, which has been shown, both qualitatively and through a subjective study, to be advantageous compared to previous methods. However, this method's high-quality results came at the cost of a computationally expensive workflow that could only be executed offline. In this paper, we present a modified algorithm which builds upon the previous work by redesigning key components to achieve real-time performance. We accomplish this by replacing the optical flow based per-pixel temporal coherency with a tone-curve-space alternative. This way we eliminate the main computational burden of the original method with little sacrifice in visual quality.
We present a novel, end-to-end workflow for content creation and distribution to a multitude of displays that have different dynamic ranges. The emergence of new, consumer level HDR displays with various peak luminances expected in 2015 gives rise to two new research questions: (i) how can the raw source content be graded for a diverse set of displays both efficiently and without restricting artistic freedom, and (ii) how can an arbitrary number of graded video streams be represented and encoded in an efficient way. In this work we propose a new editing paradigm which we call dynamic range mapping to obtain a novel Continuous Dynamic Range (CDR) video representation, where the luminance of the video content, instead of being a scalar value, is defined as a continuous function of the display dynamic range. We present an interactive interface where CDR videos can be efficiently created while providing full artistic control. In addition, we discuss the efficient approximation of CDR video using a polynomial series approximation, and its encoding and distribution to an arbitrary set of target displays. We validate our workflow in a subjective study, which suggests that a visually lossless CDR video representation can be achieved with little bandwidth overhead. Our solution can be implemented easily in the current distribution infrastructure and consists of transmitting two gradings and an additional meta-data stream, which occupies less than 13% current standard video distribution bandwidth.
We present a novel, end-to-end workflow for content creation and distribution to a multitude of displays that have different dynamic ranges. The emergence of new, consumer level HDR displays with various peak luminances expected in 2015 gives rise to two new research questions: (i) how can the raw source content be graded for a diverse set of displays both efficiently and without restricting artistic freedom, and (ii) how can an arbitrary number of graded video streams be represented and encoded in an efficient way. In this work we propose a new editing paradigm which we call dynamic range mapping to obtain a novel Continuous Dynamic Range (CDR) video representation, where the luminance of the video content, instead of being a scalar value, is defined as a continuous function of the display dynamic range. We present an interactive interface where CDR videos can be efficiently created while providing full artistic control. In addition, we discuss the efficient approximation of CDR video using a polynomial series approximation, and its encoding and distribution to an arbitrary set of target displays. We validate our workflow in a subjective study, which suggests that a visually lossless CDR video representation can be achieved with little bandwidth overhead. Our solution can be implemented easily in the current distribution infrastructure and consists of transmitting two gradings and an additional meta-data stream, which occupies less than 13% current standard video distribution bandwidth. & 2015 Elsevier Ltd. All rights reserved.