Research in light field (LF) processing has heavily increased over the last decade. This is largely driven by the desire to achieve the same level of immersion and navigational freedom for camera-captured scenes as it is currently available for CGI content. Standardization organizations such as MPEG and JPEG continue to follow conventional coding paradigms in which viewpoints are discretely represented on 2-D regular grids. These grids are then further decorrelated through hybrid DPCM/transform techniques. However, these 2-D regular grids are less suited for high-dimensional data, such as LFs. We propose a novel coding framework for higher-dimensional image modalities, called Steered Mixture-of-Experts (SMoE). Coherent areas in the higher-dimensional space are represented by single higher-dimensional entities, called kernels. These kernels hold spatially localized information about light rays at any angle arriving at a certain region. The global model consists thus of a set of kernels which define a continuous approximation of the underlying plenoptic function. We introduce the theory of SMoE and illustrate its application for 2-D images, 4-D LF images, and 5-D LF video. We also propose an efficient coding strategy to convert the model parameters into a bitstream. Even without provisions for high-frequency information, the proposed method performs comparable to the state of the art for low-to-mid range bitrates with respect to subjective visual quality of 4-D LF images. In case of 5-D LF video, we observe superior decorrelation and coding performance with coding gains of a factor of 4x in bitrate for the same quality. At least equally important is the fact that our method inherently has desired functionality for LF rendering which is lacking in other state-of-the-art techniques: (1) full zero-delay random access, (2) light-weight pixel-parallel view reconstruction, and (3) intrinsic view interpolation and super-resolution.
A novel image approximation framework called steered mixture-of-experts (SMoE) was recently presented. SMoE has multiple applications in coding, scale-conversion, and general processing of image modalities. In particular, it has strong potential for coding and streaming higher dimensional image modalities that are necessary to leverage full translational and rotational freedom (6 degrees-of-freedom) in virtual reality for camera captured images. In this paper, we analyze the rendering performance of SMoE for 2D images and 4D light fields. Two different GPU implementations that parallelize the SMoE regression step at pixel-level are presented, including experimental evaluations based on rendering performance and quality. In this paper it is shown that on appropriate hardware, an OpenCL implementation can achieve 85 fps and 22 fps for, respectively, 1080p and 4K renderings of large models with more than 100,000 of Gaussian kernels.
Steered Mixture-of-Experts (SMoE) is a novel framework for approximating multidimensional image modalities. Our goal is to provide full Six Degrees-of-Freedom capabilities for camera captured content. Previous research concerned only limited translational movement for which the 4D light field representation is sufficient. However, our goal is to arrive at a representation that allows for unlimited translational-rotational freedom, i.e. our goal is to approximate the full 5D plenoptic function. Until now, SMoE was only applied on Euclidean spaces. However, the plenoptic function contains two spherical coordinate dimensions. In this paper, we propose a methodology to extend the SMoE framework to spherical dimensions. Furthermore, we propose a method to reduce the parameter space to the same two dimensional Euclidean space as for planar 2D images by using a projection of the covariance matrices onto tangent spaces perpendicular to the unit sphere. Finally, we propose a novel training technique for spherical dimensions based on these observations. Experiments performed on omnidirectional 360° images show that the introduction of the dimensionality-reduction projection step results in very low quality loss.
Steered Mixture-of-Experts (SMoE) is a novel framework for the approximation, coding, and description of image modalities. The future goal is to arrive at a representation for Six Degrees-of-Freedom (6DoF) image data. The goal of this paper is to introduce SMoE for 4D light field videos by including the temporal dimension. However, these videos contain vast amounts of samples due to the large number of views per frame. Previous work on static light field images mitigated the problem by hard subdividing the modeling problem. However, such a hard subdivision introduces visually disturbing block artifacts on moving objects in dynamic image data. We propose a novel modeling method that does not result in block artifacts while minimizing the computational complexity and which allows for a varying spread of kernels in the spatio-temporal domain. Experiments validate that we can progressively model light field videos with increasing objective quality up to 0.97 SSIM.
Steered Mixture-of-Experts (SMoE) is a novel framework for representing multidimensional image modalities. In this paper, we propose a coding methodology for SMoE models that is readily extendable to any dimensional SMoE model, thus representing any image modality of any dimension. We evaluate the coding performance of SMoE models of light field video, a 5D image modality, i.e. time, two angular, and two spatial dimensions. The coding consists of the exploiting the redundancy between the parameters of SMoE models, i.e. a set of multivariate Gaussian distributions. We compare the performance of three multi-view HEVC (MV-HEVC) con figurations that differ in terms of random access. Each subaperture view from the light field video is interpreted as a single view in MV-HEVC. Experiments validate that excellent coding performance compared to MV-HEVC for low-to midrange bitrates in terms of PSNR and SSIM with bitrate savings up to 75%.
Previous research showed highly efficient compression results for low bit-rates using Steered Mixture-of-Experts (SMoE), higher rates still pose a challenge due to the non-convex optimization problem that becomes more difficult when increasing the number of components. Therefore, a novel estimation method based on Hidden Markov Random Fields is introduced taking spatial dependencies of neighboring pixels into account combined with a tree-structured splitting strategy. Experimental evaluations for images show that our approach outperforms state-of-the-art techniques using only one robust parameter set. For video and light field modeling even more gain can be expected.
Steered Mixture-of-Experts (SMoE) is a novel framework for the approximation, coding, and description of image modalities such as light field images and video. The future goal is to arrive at a representation for Six Degrees-of-Freedom (6DoF) image data. Previous research has shown the feasibility of real-time pixel-parallel rendering of static light field images. Each pixel is independently reconstructed by kernels that lay in its vicinity. The number of kernels involved forms the bottleneck on the achievable framerate. The goal of this paper is twofold. Firstly, we introduce pixel-level rendering of light field video, as previous work only rendered static content. Secondly, we investigate rendering using a predefined number of most significant kernels. As such, we can deliver hard real-time constraints by trading off the reconstruction quality.
We propose a novel approach for modeling and coding color in images and video. Luminance is linearly correlated with chrominance locally, as such we can predict color given the luma value. Using the Steered Mixture-of-Experts (SMoE) approach, the image is viewed as a stochastic process over 5 random variables including the 2-D pixel locations, 1 luminance and 2 chrominance values. We model this process as a continuous joint density function by fitting a K-modal 5-D Gaussian Mixture Model (GMM). As such, the chroma values are predicted as the expectation of the conditional density. To validate, the technique was integrated within JPEG showing PSNR gains in the lower bitrate regions. A deeper analysis of the tolerance of the activation function is given through recycling color models in video sequences, yielding a high quality reconstruction over a considerable range of frames.
The proposed framework, called Steered Mixture-ofExperts (SMoE), enables a multitude of processing tasks on light fields using a single unified Bayesian model. The underlying assumption is that light field rays are instantiations of a non-linear or non-stationary random process that can be modeled by piecewise stationary processes in the spatial domain. As such, it is modeled as a space-continuous Gaussian Mixture Model. Consequently, the model takes into account different regions of the scene, their edges, and their development along the spatial and disparity dimensions. Applications presented include light field coding, depth estimation, edge detection, segmentation, and view interpolation. The representation is compact, which allows for very efficient compression yielding state-of-the-art coding results for low bit-rates. Furthermore, due to the statistical representation, a vast amount of information can be queried from the model even without having to analyze the pixel values. This allows for "blind" light field processing and classification
Our challenge is the design of a “universal” bit-efficient image compression approach. The prime goal is to allow reconstruction of images with high quality. In addition, we attempt to design the coder and decoder “universal”, such that MPEG-7-like low-and mid-level descriptors are an integral part of the coded representation. To this end, we introduce a sparse Mixture-of-Experts regression approach for coding images in the pixel domain. The underlying stochastic process of the pixel amplitudes are modelled as a 3-dimensional and multi-modal Mixture-of-Gaussians with K modes. This closed form continuous analytical model is estimated using the Expectation-Maximization algorithm and describes segments of pixels by local 3-D Gaussian steering kernels with global support. As such, each component in the mixture of experts steers along the direction of highest correlation. The conditional density then serves as the regression function. Experiments show that a considerable compression gain is achievable compared to JPEG for low bitrates for a large class of images, while forming attractive low-level descriptors for the image, such as the local segmentation boundaries, direction of intensity flow and the distribution of these parameters over the image.
In this paper, we introduce a novel approach for video compression that explores spatial as well as temporal redundancies over sequences of many frames in a unified framework. Our approach supports "compressed domain vision" capabilities. To this end, we developed a sparse Steered Mixture-of- Experts (SMoE) regression network for coding video in the pixel domain. This approach drastically departs from the established DPCM/Transform coding philosophy. Each kernel in the Mixture-of-Experts network steers along the direction of highest correlation, both in spatial and temporal domain, with local and global support. Our coding and modeling philosophy is embedded in a Bayesian framework and shows strong resemblance to Mixture-of-Experts neural networks. Initial experiments show that at very low bit rates the SMoE approach can provide competitive performance to H.264.
This paper introduces a novel approach for coding luminance images using kernel-based adaptive filtering and context-adaptive arithmetic coding. This approach tackles the problem that is present in current image and video coders; these coders depend on assumptions of the image and are constrained by the linearity of their predictors. The efficacy of the predictors determines the compression gain. The goal is to create a generic image coder that learns and adapts to the characteristics of the signals and handles nonlinearity in the prediction. Results show that pixel luminance prediction using the Kernel Least Mean Squares (KLMS) yields a significant gain compared to the standard Least Mean Squares algorithm. By coding the residual using a Context-Adaptive Arithmetic Coder (CAAC), the codec is able to outperform the current industry standards of lossless image coding. An average bitrate reduction of more than 2.5% is found for the used test set.
Kernel regression has been proven successful for image de-noising, deblocking and reconstruction. These techniques lay the foundation for new image coding opportunities. In this paper, we introduce a novel compression scheme: Sparse Steering Kernel Synthesis Coding (SSKSC). This pre- and postprocessor for JPEG performs non-uniform sampling based on the smoothness of an image, and reconstructs the missing pixels using adaptive kernel regression. At the same time, the kernel regression reduces the blocking artifacts from the JPEG coding. Crucial to this technique is that non-uniform sampling is performed while maintaining only a small overhead for signalization. Compared to JPEG, SSKSC achieves a compression gain for low bits-per-pixel regions of 50% or more for PSNR and SSIM. A PSNR gain is typically in the 0.0-0.5 bpp range, and an SSIM gain can mostly be achieved in the 0.0-1.0 bpp range.
Bart Goossens合作论文数Ghent University
TELIN-IPI-IBBT1