We describe a system to remove real-world reflections from images for consumer photography. Our system operates on linear (RAW) photos, and accepts an optional contextual photo looking in the opposite direction (e.g., the "selfie" camera on a mobile device). This optional photo disambiguates what should be considered the reflection. The system is trained solely on synthetic mixtures of real RAW photos, which we combine using a reflection simulation that is photometrically and geometrically accurate. Our system comprises a base model that accepts the captured photo and optional context photo as input, and runs at 256p, followed by an up-sampling model that transforms 256p images to full resolution. The system produces preview images at 1K in 4.5-6.5s on a MacBook or iPhone 14 Pro. We show SOTA results on RAW photos that were captured in the field to embody typical consumer photos, and show that training on RAW simulation data improves performance more than the architectural variations among prior works.
Common editing operations performed by professional photographers include the cleanup operations: deemphasizing distracting elements and enhancing subjects. These edits are challenging, requiring a delicate balance between manipulating the viewer's attention while maintaining photo realism. While recent approaches can boast successful examples of attention attenuation or amplification, most of them also suffer from frequent unrealistic edits. We propose a realism loss for saliency-guided image enhancement to maintain high realism across varying image types, while attenuating distractors and amplifying objects of interest. Evaluations with professional photographers confirm that we achieve the dual objective of realism and effectiveness, and outperform the recent approaches on their own datasets, while requiring a smaller memory footprint and runtime. We thus offer a viable solution for automating image enhancement and photo cleanup operations.
According to the efficient coding hypothesis, sensory systems are adapted to maximize their ability to encode information about the environment. Sensory neurons play a key role in encoding by selectively modulating their firing rate for a subset of all possible stimuli. This pattern of modulation is often summarized via a tuning curve. The optimally efficient distribution of tuning curves has been calculated in variety of ways for one-dimensional (1-D) stimuli. However, many sensory neurons encode multiple stimulus dimensions simultaneously. It remains unclear how applicable existing models of 1-D tuning curves are for neurons tuned across multiple dimensions. We describe a mathematical generalization that builds on prior work in 1-D to predict optimally efficient multidimensional tuning curves. Our results have implications for interpreting observed properties of neuronal populations. For example, our results suggest that not all tuning curve attributes (such as gain and bandwidth) are equally useful for evaluating the encoding efficiency of a population.
In this paper, we tackle the problem of spatio-temporal tagging of self-driving scenes from raw sensor data. Our approach learns a universal embedding for all tags, enabling efficient tagging of many attributes and faster learning of new attributes with limited data. Importantly, the embedding is spatio-temporally aware, allowing the model to naturally output spatio-temporal tag values. Values can then be pooled over arbitrary regions, in order to, for example, compute the pedestrian density in front of the SDV, or determine if a car is blocking another car at a 4-way intersection. We demonstrate the effectiveness of our approach on a new large scale self-driving dataset, SDVScenes, containing 15 attributes relating to vehicle and pedestrian density, the actions of each actor, the speed of each actor, interactions between actors, and the topology of the road map.
Detecting the intention of drivers is an essential task in self-driving, necessary to anticipate sudden events like lane changes and stops. Turn signals and emergency flashers communicate such intentions, providing seconds of potentially critical reaction time. In this paper, we propose to detect these signals in video sequences by using a deep neural network that reasons about both spatial and temporal information. Our experiments on more than a million frames show high per-frame accuracy in very challenging scenarios.
In this paper, we present LaserNet, a computationally efficient method for 3D object detection from LiDAR data for autonomous driving. The efficiency results from processing LiDAR data in the native range view of the sensor, where the input data is naturally compact. Operating in the range view involves well known challenges for learning, including occlusion and scale variation, but it also provides contextual information based on how the sensor data was captured. Our approach uses a fully convolutional network to predict a multimodal distribution over 3D boxes for each point and then it efficiently fuses these distributions to generate a prediction for each object. Experiments show that modeling each detection as a distribution rather than a single deterministic box leads to better overall detection performance. Benchmark results show that this approach has significantly lower runtime than other recent detectors and that it achieves state-of-the-art performance when compared on a large dataset that has enough data to overcome the challenges of training on the range view.
A theoretical framework is developed to describe the transformation that distributes probability density functions uniformly over space. In one dimension, the cumulative distribution can be used, but does not generalize to higher dimensions, or non-separable distributions. A potential function is shown to link probability density functions to their transformation, and to generalize the cumulative. A numerical method is developed to compute the potential, and examples are shown in two dimensions.
We describe a photo forensic technique based on detecting inconsistencies in lighting. This technique explicitly measures the 3-D lighting properties for individual people, objects, or surfaces in a single image. We show that with minimal training, an analyst can accurately specify 3-D shape in a single image from which 3-D lighting can be automatically estimated. A perturbation analysis on the estimated lighting is performed to yield a probabilistic measure of the location of the illuminating light. Inconsistencies in lighting within an image evidence photo tampering.
We describe a method for detecting physical inconsistencies in lighting from the shading and shadows in an image. This method imposes a multitude of shading- and shadow-based constraints on the projected location of a distant point light source. The consistency of a collection of such constraints is posed as a linear programming problem. A feasible solution indicates that the combination of shading and shadows is physically consistent, while a failure to find a solution provides evidence of photo tampering.
A variety of forensic methods have been developed to identify falsified photos, each unified by the ability to estimate and detect properties of a photo that are perturbed by forgery. There exist, however, many photos in which the required properties cannot be estimated. We present an approach to detect forgery in these photos. We use this approach to detect physically inconsistent shadows and shading in photos for which it is not possible to estimate the associated lighting properties. Specifically, we develop a method to detect inconsistent shadows cast by point and area light sources when a strict shadow-to-object correspondence cannot be estimated. We further develop a method to detect inconsistencies between shadows and the shading on objects when object geometry is only partially known, and when objects are photographed under unknown perspective. We conclude by describing prior methods that can be generalized to analyze photos in which estimation is not possible.
It is often desirable to determine if an image has been modified in any way from its original recording. The JPEG format affords engineers many implementation trade-offs which give rise to widely varying JPEG headers. We exploit these variations for image authentication. A camera signature is extracted from a JPEG image consisting of information about quantization tables, Huffman codes, thumbnails, and exchangeable image file format (EXIF). We show that this signature is highly distinct across 1.3 million images spanning 773 different cameras and cell phones. Specifically, 62% of images have a signature that is unique to a single camera, 80% of images have a signature that is shared by three or fewer cameras, and 99% of images have a signature that is unique to a single manufacturer. The signature of Adobe Photoshop is also shown to be unique relative to all 773 cameras. These signatures are simple to extract and offer an efficient method to establish the authenticity of a digital image.
In recent years, advertisers and magazine editors have been widely criticized for taking digital photo retouching to an extreme. Impossibly thin, tall, and wrinkle- and blemish-free models are routinely splashed onto billboards, advertisements, and magazine covers. The ubiquity of these unrealistic and highly idealized images has been linked to eating disorders and body image dissatisfaction in men, women, and children. In response, several countries have considered legislating the labeling of retouched photos. We describe a quantitative and perceptually meaningful metric of photo retouching. Photographs are rated on the degree to which they have been digitally altered by explicitly modeling and estimating geometric and photometric changes. This metric correlates well with perceptual judgments of photo retouching and can be used to objectively judge by how much a retouched photo has strayed from reality.
Photo deblurring has been a major research topic in the past few years. So far, existing methods have focused on removing the blur due to camera shake and object motion. In this paper, we show that the optical system of the camera also generates significant blur, even with professional lenses. We introduce a method to estimate the blur kernel densely over the image and across multiple aperture and zoom settings. Our measures show that the blur kernel can have a non-negligible spread, even with top-of-the-line equipment, and that it varies nontrivially over this domain. In particular, the spatial variations are not radially symmetric and not even left-right symmetric. We develop and compare two models of the optical blur, each of them having its own advantages. We show that our models predict accurate blur kernels that can be used to restore photos. We demonstrate that we can produce images that are more uniformly sharp unlike those produced with spatially-invariant deblurring techniques.
We describe how to exploit the formation and storage of an embedded image thumbnail for image authentication. The creation of a thumbnail is modeled with a series of filtering operations, contrast adjustment, and compression. We automatically estimate these model parameters and show that these parameters differ significantly between camera manufacturers and photo-editing software. We also describe how this signature can be combined with encoding information from the underlying full resolution image to further refine the signature's distinctiveness.
When creating a photographic composite, it can be difficult to match lighting conditions. We describe a technique for measuring lighting conditions in an image, and describe its use in detecting photographic composites. Specifically, we describe how to approximate a 3-D lighting environment with a low-dimensional model and how to estimate the model's parameters from a single image. Inconsistencies in the lighting model are then used as evidence of tampering.
Photos are commonly falsified by compositing two or more people into a single image. We describe how such composites can be detected by estimating a camera’s intrinsic parameters. Differences in these parameters across the image are then used as evidence of tampering. Expanding on earlier work, this approach is more applicable to low-resolution images, but requires a reference image of each person in the photo as they are directly facing the camera. When considering composites of famous people, such a reference photo is easily obtained from an on-line image search.
We describe a technique for authenticating printed and scanned text documents. This technique works by modeling the degradation in a document caused by printing. The resulting printer profile is then used to detect inconsistencies across a document, and for ballistic purposes - that of linking a document to a printer.