Many challenges that deal with processing of HDR material remain very much open for the film industry, whose extremely demanding quality standards are not met by existing automatic methods. Therefore, when dealing with HDR content, substantial work by very skilled technicians has to be carried out at every step of the movie production chain. Based on recent findings and models from vision science, we propose in this work effective tone mapping and inverse tone mapping algorithms for production, post-production and exhibition. These methods are automatic and real-time, and they have been both fine-tuned and validated by cinema professionals, with psychophysical tests demonstrating that the proposed algorithms outperform both the academic and industrial state-of-the-art. We believe these methods bring the field closer to having fully automated solutions for important challenges for the cinema industry that are currently solved manually or sub-optimally. Another contribution of our research is to highlight the limitations of existing image quality metrics when applied to the tone mapping problem, as none of them, including two state-of-the-art deep learning metrics for image perception, are able to predict the preferences of the observers.
The responses of visual neurons, as well as visual perception phenomena in general, are highly nonlinear functions of the visual input, while most vision models are grounded on the notion of a linear receptive field (RF). The linear RF has a number of inherent problems: it changes with the input, it presupposes a set of basis functions for the visual system, and it conflicts with recent studies on dendritic computations. Here we propose to model the RF in a nonlinear manner, introducing the intrinsically nonlinear receptive field (INRF). Apart from being more physiologically plausible and embodying the efficient representation principle, the INRF has a key property of wide-ranging implications: for several vision science phenomena where a linear RF must vary with the input in order to predict responses, the INRF can remain constant under different stimuli. We also prove that Artificial Neural Networks with INRF modules instead of linear filters have a remarkably improved performance and better emulate basic human perception. Our results suggest a change of paradigm for vision science as well as for artificial intelligence.
The statistics of real world images have been extensively investigated, but in virtually all cases using only low dynamic range image databases. The few studies that have considered high dynamic range (HDR) images have performed statistical analyses categorizing images as HDR according to their creation technique, and not to the actual dynamic range of the underlying scene. In this study we demonstrate, using a recent HDR dataset of natural images, that the statistics of the image as received at the camera sensor change dramatically with dynamic range, with particularly strong correlations with dynamic range being observed for the median, standard deviation, skewness, and kurtosis, while the one over frequency relationship for the power spectrum breaks down for images with a very high dynamic range, in practice making HDR images not scale invariant. Effects are also noted in the derivative statistics, the single pixel histograms, and the Haar wavelet analysis. However, we also show that after some basic early transforms occurring within the eye (light scatter, nonlinear photoreceptor response, center-surround modulation) the statistics of the resulting images become virtually independent from the dynamic range, which would allow them to be processed more efficiently by the human visual system.
In 1986, Paul Whittle investigated the ability to discriminate between the luminance of two small patches viewed upon a uniform background. In 1992, Paul Whittle asked subjects to manipulate the luminance of a number of patches on a uniform background until their brightness appeared to vary from black to white with even steps. The data from the discrimination experiment almost perfectly predicted the gradient of the function obtained in the brightness experiment, indicating that the two experimental methodologies were probing the same underlying mechanism. Whittle introduced a model that was able to capture the pattern of discrimination thresholds and, in turn, the brightness data; however, there were a number of features in the data set that the model couldn't capture. In this paper, we demonstrate that the models of Kane and Bertalmío (2017) and Kingdom and Moulden (1991) may be adapted to predict all the data but only by incorporating an accurate model of detection thresholds. Additionally, we show that a divisive gain model may also capture the data but only by considering polarity-dependent, nonlinear inputs following the underlying pattern of detection thresholds. In summary, we conclude that these models provide a simple link between detection thresholds, discrimination thresholds, and brightness perception.
The dynamic range of real world scenes may vary from around 102 to greater than 107, whilst the dynamic range of monitors may vary from 102 to 105. In this paper, we investigate the impact of the dynamic range ratio (DRratio) between the captured scene and the displayed image, upon the value of system gamma preferred by subjects (a simple global power law transformation applied to the image). To do so, we present an image dataset with a broad distribution of dynamic ranges upon various sub-ranges of a SIM2 monitor. The full dynamic range of the monitor is 105 and we present images using either the full range, 75% or 50% of this, while maintaining a fixed mid-luminance level. We find that the preferred system gamma is inversely correlated with the DRratio and importantly, is one (linear) when the DRratio is one. This strongly suggests that the visual system is optimized for processing images only when the dynamic range is presented correctly. The DRratio is not the only factor. By using 50% of the monitor dynamic range and using either the lower, middle or upper portion of the monitor, we show that increasing the overall luminance level also increases the preferred system gamma, although to a lesser extent than the DRratio.
High Dynamic Range (HDR) technologies support the capture and presentation of a wider range of luminance values than conventional systems.An important element of video processing is the transfer function which should emulate human perception and this needs to be revisited for HDR content and displays.In the paper, we adapt a nonlinearity designed for the tone-mapping problem to the problem of video coding.We test the nonlinearity using the Motion Picture Experts Group methodology and find it can outperform existing methods in terms of HDR video quality measure.
Sensitivity to variations in luminance has been extensively studied via detection thresholds to give the well-known threshold versus intensity (TvI) function. The function is expansive – the lower the luminance level, the lower the threshold. However, when a pedestal is introduced such that the task is to discriminate between the luminance of two patches superimposed upon a uniform background, the results are substantially more complex. Thresholds are both low at the lowest luminance levels tested and additionally around the background luminance level, an effect termed 'crispening'. This has lead authors to propose two separate mechanisms, one more sensitive to low luminance levels and another to contrast around the background luminance level. In this paper, we model discrimination thresholds via a single mechanism. We assume that the maximal sensitivity of the HVS is well modeled by the shape of the TvI function, but in the case of non-uniform backgrounds this function is modulated by a gain control mechanism that increases thresholds away from the background luminance level. We evaluate our model upon the data of Paul Whittle (1986) who examined discrimination threshold over a broad luminance range and also upon new and old data for functions exhibiting various levels of 'crispening' (Nagy and Kamholz, 1995). Second, in keeping with the work of Paul Whittle (1992) we investigate whether this model can predict supra-threshold brightness functions. We find that as long as a realistic (a non-Weber) TvI function is used, that the brightness functions can be accurately estimated. In the case of non-uniform backgrounds (salt and pepper noise or the inclusion of an annulus), the model requires an additional gain term for each background luminance level. Although this adds to the complexity of the model, it offers the possibility of extending the model to arbitrarily complex stimuli. Meeting abstract presented at VSS 2017
AVA Christmas Meeting, Queen Mary University of London, December 19, 2016 Keynote Lectures 1. The Lawful Relation Between Discrimination Threshold and Perceptual Bias
'Crispening' is an effect whereby subjects perception of luminance is biased away from the background luminance level. The effect is strong, but may be reduced or abolished by the addition of a hue shift or an annulus that separates the tests stimuli from the background [16, 18]. In this paper we investigate whether the 'crispening' effect may arise from a simple gain mechanism that decreases sensitivity at luminance levels away from the background luminance level. The model takes as input the threshold versus intensity function, then decreases sensitivity via a gain mechanism. The supra-threshold percept is then estimated via Fechnerian integration of the resulting thresholds. We find that the model can predict subjects' luminance nonlinearities in all conditions as long as a parameter that controls the degree of gain is allowed to vary. Perhaps more interestingly, we find that the model can explain the luminance nonlinearity in the case where an annulus is present by treating the annulus as an additional background luminance level that also mediates gain. When multiple background luminance levels are included, the gain no longer produces the distinctive 'crispening' effect, although the gain still substantially affects the shape of the luminance nonlinearity. This may account for why 'crispening' is not observed when complex, real world scenes are investigated [2].
System gamma is the end-to-end exponent that describes the relationship between the relative luminance values at capture and the reproduced image. The system gamma preferred by subjects is known to vary with the background luminance condition and the image in question. We confirm the previous two findings using an image database with both high and low dynamic range images (from 10(2) to 10(7)), but also find that the preferred system gamma varies with the dynamic range of the monitor (CRT, LCD, or OLED). We find that the preferred system gamma can be predicted in all conditions and for all images by a simple model that searches for the value that best flattens the lightness distribution, where lightness is modeled as a power law of onscreen luminance. To account for the data, the exponent must vary with the viewing conditions. The method presented allows the inference of lightness perception in natural scenes without direct measurement and makes testable predictions for how lightness perception varies with the viewing condition and the distribution of luminance values in a scene. The data from this paper has been made available online.
We investigate the role of lightness perception in determining the perceived contrast of simple textures. It is known that the background luminance of a display affects the relationship between onscreen-luminance and perceived lightness. This effect can be approximated by a power-law with an exponent that increases with the background luminance level (Bartleson & Breneman 1967; Bartleson, 1975; Stevens & Stevens, 1963). Moreover, recent work has demonstrated that an adaptive lightness model is critical to understanding the perceived contrast of natural textures (Kane & Bertalmío, Submitted). However, a simple power-law cannot account for the effect of ‘crispening’, whereby subjects are more sensitive to luminance variations around the background luminance (Whittle, 1992). In this study we empirically estimated subject's luminance-to-lightness functions via a bisection paradigm (Munsell, Sloan, & Godlove, 1933) for five background luminance levels, from 0 to 100%. The results reveal complex functions with clear evidence of ‘crispening’. We then computed the point-of-subjective-equality (PSE) for the contrast of a reference and test patches with a mean luminance of 25% and 75% respectively, using the background luminance conditions from experiment one. The PSE's as a function of background luminance exhibit a peak and a trough around the mean luminance of the test and reference, respectively. We find that subjects' PSE can only be modeled by first passing the stimulus through the empirically estimated lightness functions before estimating contrast. Future work will investigate whether the demonstrated impact of ‘crispening’ generalizes to more complex stimuli, and stimuli that subtend a greater angle.
We investigate the role of lightness perception in determining the perceived contrast of images. In particular, it is known that the background luminance of a display affects the relationship between onscreen luminance and perceived lightness. Stevens & Stevens (1963) modeled this effect using a simple power law. However, Whittle (1992) observed a more local effect, whereby subjects are more sensitive to lightness variations around the background luminance (the crispening effect). We probe lightness perception by asking subjects to manipulate the contrast of small patches on a uniform background until they appear to vary, from black to white, in a perceptually linear manner. In a second experiment, we estimate the contrast required to match the contrast of a light patch to that of a dark patch. Both experiments are conducted using five background luminance conditions, from 0 to 100% luminance. We find that subjects contrast judgments can only be modeled by first estimating the perceived lightness in a scene, using the empirically estimated lightness functions, before computing contrast. We conclude that models of contrast perception must include sophisticated models of lightness perception.
Comunicacio presentada al IS&T International Symposium on Electronic Imaging, celebrat del 14 al 18 de febrer de 2016 a San Francisco (CA, USA) i organitzat per la Society for Imaging Science and Technology.
Cameras automatically apply non-linear transformations to the sensor data, allowing for perceptually-uniform quantization suited to standard dynamic range displays in dim conditions. In the cinema industry, data is recorded in raw (linear) format and non-linearly corrected in post-production by a skilled technician who optimizes image appearance for cinema (dark) conditions. We propose a method that automatically performs this non-linear transformation taking into account the intended viewing conditions. It is based on visual perception models and produces results that look natural, without any spatio-temporal artifacts. User preference tests show that our method outperforms state of the art approaches. The technique is fast and could be implemented on camera hardware. It can be used for on-set monitoring on regular displays, as a substitute for gamma-correction, and as a way of providing the colorist with content that is both natural looking and has a crisp, clear image.
The optimization or falsification of vision science models can require time-consuming experimentation. This is especially true for models of artifact detection that require large databases of thresholds judgements or subjective image quality scores. Wang&Simoncelli (JoV, 2005, 2008) proposed a novel psychophysical method to avoid such experimental burden: the MAximum Differentiation (MAD) competition. This technique computes a pair of maximally different images according to each vision model under investigation, and the subject then selects the pair of images that they perceive to have greater difference. This paradigm is able to reduce the falsification of competing models to one experiment. As a result, MAD has been used to simplify the optimization of divisive-normalization contrast perception models (Malo&Simoncelli, SPIE 2015). The MAD paradigm is proposed in a context-independent manner and used on complex, unconstrained datasets. However, as a proof-of-concept, we demonstrate that the MAD paradigm can produce contradictory results in different surround conditions: these computational examples (based on luminance adaptation and the associated crispening effect, see supplementary material) show that the decision between models cannot be reduced to a single image comparison. On the contrary, it is mandatory to extend MAD, either by (1) doing a number of surround-dependent comparisons with the same images, which would reduce the conceptual advantage of MAD, or by (2) including the effects of the surround in the models considered in the MAD competition, which would give surround-dependent image pairs. REFERENCES Wang Z. & Simoncelli E.P. (2005) MAD competition: comparing quantitative models of perceptual discriminability. VSS Abstract. J. Vision, 5(8): 230-230. Wang Z. & Simoncelli E. P. (2008). MAximum Differentiation competition: A methodology for comparing computational models of perceptual quantities. J. Vision, 8(12): 8, 1–13 Malo J. & Simoncelli E. P. (2015). Geometrical and statistical properties of vision models obtained via MAximum Differentiation. Proc. SPIE, Human Vision Electr. Imag. Vol. 9394 Meeting abstract presented at VSS 2016
We investigate the impact of the background luminance upon the perceived image quality of real world scenes. To do so, we generate a set of small image patches that span the full range of mean luminance values and contrasts that may be displayed upon a monitor with a finite luminance range. Subjects viewed the images on a uniform black, grey or white surround and were asked to rate the perceived quality on a scale from 0 to 9. We find that that the maximum image quality scores occur for images with a mean luminance of less than half, consistent with the image being passed through a compressive non-linearity before contrast is computed. Moreover, the maximum image quality scores occur at lower mean luminance levels when the background luminance is darker, a pattern consistent with investigations into lightness perception. We conclude that models of contrast perception require an adaptive model of lightness perception. However, we also note the considerable challenge of developing a model of lightness perception that can generalize to any given display configuration.
We propose a fast, local denoising method where the Euclidean curvature of the noisy image is approximated in a regularizing manner and a clean image is reconstructed from this smoothed curvature. User preference tests show that when denoising real photographs with actual noise our method produces results with the same visual quality as the more sophisticated, non local algorithms Non-local Means and BM3D, but at a fraction of their computational cost. These tests also highlight the limitations of objective image quality metrics like PSNR and SSIM, which correlate poorly with user preference.
Jesus Malo合作论文数Dpt. of Optics,
School of Physics.
Universitat de Valencia2