We propose Coded-E2LF (coded event to light field), a computational imaging method for acquiring a 4-D light field using a coded aperture and a stationary event-only camera. In a previous work, an imaging system similar to ours was adopted, but both events and intensity images were captured and used for light field reconstruction. In contrast, our method is purely event-based, which relaxes restrictions for hardware implementation. We also introduce several advancements from the previous work that enable us to theoretically support and practically improve light field reconstruction from events alone. In particular, we clarify the key role of a black pattern in aperture coding patterns. We finally implemented our method on real imaging hardware to demonstrate its effectiveness in capturing real 3-D scenes. To the best of our knowledge, we are the first to demonstrate that a 4-D light field with pixel-level accuracy can be reconstructed from events alone. Our software is included in the supplementary material.
To time-efficiently and stably acquire the intensity information for phase retrieval under a coherent illumination, we leverage an event-based vision sensor (EVS) that can detect changes in logarithmic intensity at the pixel level with a wide dynamic range. In our optical system, we translate the EVS along the optical axis, where the EVS records the intensity changes induced by defocus as events. To recover phase distributions, we formulate a partial differential equation, referred to as the transport of event equation, which presents a linear relationship between the defocus events and the phase distribution. We demonstrate through experiments that the EVS is more advantageous than the conventional image sensor for rapidly and stably detecting the intensity information, defocus events, which enables accurate phase retrieval, particularly under low-lighting conditions.
We propose a hybrid reconstruction method for coded light-field imaging. Most previous methods utilized pretrained reconstruction, in which the reconstruction process was first pre-trained on a light-field dataset taken from various 3-D scenes and then applied to new target 3-D scenes. However, pre-trained reconstruction is not necessarily optimal for a specific 3-D scene and sometimes results in insufficient reconstruction quality for the fine details. To address this issue, we first introduce a method of self-supervised reconstruction that focuses on the data observed from a specific 3-D scene. To this end, we incorporate a learning-based 3-D representation technique called neural radiance fields (NeRFs) into the framework of coded light-field imaging. Moreover, we combine pre-trained and self-supervised approaches seamlessly to synergize the strengths of both. Experimental results demonstrate that our method can achieve better reconstruction quality consistently over various 3-D scenes than the previous pre-trained methods.
We introduce a novel phase-shifting digital holography (PSDH) method leveraging a hybrid event-based vision sensor (EVS). The key idea of our method is the phase shift during a single exposure. The hybrid EVS records a hologram blurred by the phase shift, together with the events corresponding to blur variations. We present analytical and optimization-based methods that theoretically support the reconstruction of full-complex wavefronts from the blurred hologram and events. The experimental results demonstrate that our method achieves a reconstruction quality comparable to that of a conventional PSDH method while enhancing the acquisition efficiency.
The vast volume of medical image data necessitates efficient compression techniques to support remote healthcare services. This paper explores Region of Interest (ROI) coding to address the balance between compression rate and image quality. By leveraging UNET segmentation on the Brats 2020 dataset, we accurately identify tumor regions, which are critical for diagnosis. These regions are then subjected to High Efficiency Video Coding (HEVC) for compression, enhancing compression rates while preserving essential diagnostic information. This approach ensures that critical image regions maintain their quality, while non-essential areas are compressed more. Our method optimizes storage space and transmission bandwidth, meeting the demands of telemedicine and large-scale medical imaging. Through this technique, we provide a robust solution that maintains the integrity of vital data and improves the efficiency of medical image handling.
To efficiently compress the sign information of images, we address a sign retrieval problem for the block-wise discrete cosine transformation (DCT): reconstruction of the signs of DCT coefficients from their amplitudes. To this end, we propose a fast sign retrieval method on the basis of binary classification machine learning. We first introduce 3D representations of the amplitudes and signs, where we pack amplitudes/signs belonging to the same frequency band into a 2D slice, referred to as the sub-band block. We then retrieve the signs from the 3D amplitudes via binary classification, where each sign is regarded as a binary label. We implement a binary classification algorithm using convolutional neural networks, which are advantageous for efficiently extracting features in the 3D amplitudes. Experimental results demonstrate that our method achieves accurate sign retrieval with an overwhelmingly low computation cost.
We propose a compression method for phase-only holograms. We first apply a deep neural network-based phase unwrapping algorithm to a target phase map. The unwrapped phase map becomes spatially smooth like natural images. On the basis of this characteristic, we then compress the unwrapped phase map using a deep image compression method. We demonstrate that phase unwrapping achieves high rate-distortion performance.
This research addresses challenges in visible light communication (VLC) systems that use two orthogonally aligned rolling shutters (RS) image sensors as receivers and an LED array as a transmitter. The study focuses on overcoming burst signal loss caused by unsensed periods between frames in RS image sensors. To improve VLC performance, we propose and compare three different schemes. The first is a conventional approach that superimposes a Barker code synchronization signal on the transmission signal using pulse width modulation (PWM). The second introduces a new method of spatial synchronization by dedicating specific LEDs in an array for synchronization signals. The third scheme extends this spatial approach by incorporating a 4 -level PWM for information transmission to increase data rates. The study aims to evaluate and compare the error rate characteristics of these three schemes, assessing their effectiveness in mitigating burst errors and enhancing overall VLC system performance. This research has potential applications, including vehicle communication, Internet of Things devices, and smartphones.
We propose a data compression method for a light field using a compact and computationally efficient neural representation. We first train a neural network with learnable parameters to reproduce the target light field. We then compress the set of learned parameters as an alternative representation of the light field. Our method is significantly different in concept from the traditional approaches where a light field is encoded as a set of images or a video (as a pseudo-temporal sequence) using off-the-shelf image/video codecs. We experimentally show that our method achieves a promising rate-distortion performance.
A light field is represented as a set of multi-view images captured from a dense 2-D array of viewpoints. To treat a light field as being continuous, we represent it as a neural radiance field (NeRF), which is a learned representation of a 3-D scene. NeRFs are renowned for their ability to reconstruct a target 3-D scene with compelling visual quality, but they are slow to train. A solution for this problem is to use a tiny neural network and trainable volumetric features as the scene representation, which is considered the baseline of our research. For further acceleration, we propose a method for warm-starting the per-scene training by setting good initial values for the trainable parameters. To this end, we introduce another encoder network to obtain the initial volumetric features from the target light field. Starting with the appropriate initial values, our method can achieve better rendering quality with fewer training iterations than the baseline.
We propose a computational imaging method for time-efficient light-field acquisition that combines a coded aperture with an event-based camera. Different from the conventional coded-aperture imaging method, our method applies a sequence of coding patterns during a single exposure for an image frame. The parallax information, which is related to the differences in coding patterns, is recorded as events. The image frame and events, all of which are measured in a single exposure, are jointly used to computationally reconstruct a light field. We also designed an algorithm pipeline for our method that is end-to-end trainable on the basis of deep optics and compatible with real camera hardware. We experimentally showed that our method can achieve more accurate reconstruction than several other imaging methods with a single exposure. We also developed a hardware prototype with the potential to complete the measurement on the camera within 22 msec and demonstrated that light fields from real 3-D scenes can be obtained with convincing visual quality. Our software and supplementary video are available from our project website.
We propose a new gradient method for holography, where a phase-only hologram is parameterized by not only the phase but also amplitude. The key idea of our approach is the formulation of a phase-only hologram using an auxiliary amplitude. We optimize the parameters using the so-called Wirtinger flow algorithm in the Cartesian domain, which is a gradient method defined on the basis of the Wirtinger calculus. At the early stage of optimization, each element of the hologram exists inside a complex circle, and it can take a large gradient while diverging from the origin. This characteristic contributes to accelerating the gradient descent. Meanwhile, at the final stage of optimization, each element evolves along a complex circle, similar to previous state-of-the-art gradient methods. The experimental results demonstrate that our method outperforms previous methods, primarily due to the optimization of the amplitude.
A light field is usually represented as a set of multi-view images captured from a two-dimensional (2-D) array of viewpoints and requires a large amount of data compared with a standard 2-D image. We propose a 2-D compatible light-field compression method for encoding a light field as a 2-D monocular image and subsidiary data. In terms of the image quality, we prioritize the central image (regarded as the 2-D monocular image) over the other images in the light field, because the light field is considered an extension of the 2-D monocular image. To this end, we encode and decode the monocular image using a standard image codec and introduce a learned encoder and decoder pair for the subsidiary data. Experimental results indicate that our method achieved promising rate-distortion performance, especially for extremely low bit-rate ranges. Even though our method requires only a small amount of subsidiary data compared with those for the monocular image, the entire light field can be reconstructed with reasonable visual quality.
We propose a method for compressively acquiring a light field video using a single camera equipped with an optical aperture-exposure coding mechanism. The aperture-exposure coding is applied to each exposure time, enabling the embedding of the information of a light field video (a 5-D volume) into a single observed image (a 2-D measurement). Temporally-successive images obtained from the camera are used to computationally reconstruct the light field video at a faster frame rate than that of the camera. We also developed a hardware prototype to validate our method on real 3-D time-varying scenes. Using our method, we can obtain a light field video with 5 x 5 viewpoints over 4 temporal sub-frames (100 views in total) per each observed image. By repeating the capture and reconstruction processes over time, we can acquire a light field video of arbitrary length at 4 x the frame rate of the camera. To the best of our knowledge, we are the first to propose a method of joint angular-temporal compression for light-field acquisition, achieving a finer temporal resolution than that of the camera. A supplementary video is available from https://youtu.be/FAujrak8Dok.
An event camera adopts a bio-inspired sensing mechanism that can record the luminance changes over time. The recorded information, called events, are detected asynchronously at each pixel in the order of microseconds. Events are quite useful for framerate upsampling of a video, because the information between the low-framerate video frames (key-frames) can be supplemented from the events. We propose a method for framerate upsampling from events on the basis of an unsupervised approach; our method does not require ground-truth high-framerate videos for pre-training but can be trained solely on the key-frames and events taken from the target scene. We also report some promising experimental results with a fast moving scene captured by a DAVIS346 event camera.
Being a general representation format of the dense light field, lenslet video, where each frame consists of a 2D grid of micro-images, shows high potential for applications in immersive media such as glasses-free 3D displays and virtual reality. However, its distinct spatial-temporal-angular distribution places a significant challenge on conventional video coding. In July 2021, the Moving Picture Experts Group (MPEG) established an Ad-Hoc group, Lenslet Video Coding (LVC), to explore use cases, efficient compression methods, testing sequences, conversion tools, and coding architectures towards a new compression standard. This paper provides an overview of recent progress in the LVC Ad Hoc group and presents the compression efficiency of state-of-the-art codec agnostic coding tools to encourage contributions in the future.
Visible light communication (VLC) and visible light ranging are applicable techniques for intelligent transportation systems (ITS). They use every unique light-emitting diode (LED) on roads for data transmission and range estimation. The simultaneous VLC and ranging can be applied to improve the performance of both. It is necessary to achieve rapid data rate and high-accuracy ranging when transmitting VLC data and estimating the range simultaneously. We use the signal modulation method of pulse-width modulation (PWM) to increase the data rate. However, when using PWM for VLC data transmission, images of the LED transmitters are captured at different luminance levels and are easily saturated, and LED saturation leads to inaccurate range estimation. In this paper, we establish a novel simultaneous visible light communication and ranging system for ITS using PWM. Here, we analyze the LED saturation problems and apply bicubic interpolation to solve the LED saturation problem and thus, improve the communication and ranging performance. Simultaneous communication and ranging are enabled using a stereo camera. Communication is realized using maximal-ratio combining (MRC) while ranging is achieved using phase-only correlation (POC) and sinc function approximation. Furthermore, we measured the performance of our proposed system using a field trial experiment. The results show that error-free performance can be achieved up to a communication distance of 55 m and the range estimation errors are below 0.5m within 60m.