
Coherent anti-Stokes Raman scattering (CARS) microspectroscopy is a powerful tool for label-free cell imaging thanks to its ability to acquire a rich amount of information. An important family of operations applied to such data is multivariate curve resolution (MCR). It aims to find main components of a dataset and compute their spectra and concentrations in each pixel. Recently, autoencoders began to be studied to accomplish MCR with dense and convolutional models. However, many questions, like the results variability or the reconstruction metric, remain open and applications are limited to hyperspectral imaging. In this article, we present a nonlinear convolutional encoder combined with a linear decoder to apply MCR to CARS microspectroscopy. We conclude with a study of the result variability induced by the encoder initialization.
We propose two automatic parameter tuning methods for Plug-and-Play (PnP) algorithms that use CNN denoisers. We focus on linear inverse problems and propose an iterative algorithm to calculate generalized cross-validation (GCV) and Stein’s unbiased risk estimator (SURE) functions for a half-quadratic splitting-based PnP (PnP-HQS) algorithm that uses a state-of- the-art CNN denoiser. The proposed methods leverage forward mode automatic differentiation to calculate the GCV and SURE functions and tune the parameters of a PnP-HQS algorithm automatically by minimizing the GCV and SURE functions using grid search. Because linear inverse problems appear frequently in computational imaging, the proposed methods can be applied in various domains. Furthermore, because the proposed methods rely on GCV and SURE functions, they do not require access to the ground truth image and do not require collecting an additional training dataset, which is highly desirable for imaging applications for which acquiring data is costly and time-consuming. We evaluate the performance of the proposed methods on deblurring and MRI experiments and show that the GCV-based proposed method achieves comparable performance to that of the oracle tuning method that adjusts the parameters by maximizing the structural similarity index between the ground truth image and the output of the PnP algorithm. We also show that the SURE-based proposed method often leads to worse performance compared to the GCV-based proposed method.
Ultrasound elasticity images, which enable the visualization of quantitative maps of tissue stiffness, can be reconstructed by solving an inverse problem. Classical model-based approaches for ultrasound elastography use deterministic finite element methods (FEMs) to incorporate the governing physical laws leading to poor performance in low SNR conditions. Moreover, these approaches utilize approximate linear forward models discretized by FEMs to describe the underlying physics governed by partial differential equations (PDEs). To achieve highly accurate stiffness images, it is essential to compensate the error induced by noisy measurements and inaccurate forward models. In this regard, we propose a joint model-based and learning-based framework for estimating the elasticity distribution by solving a regularized optimization problem. To address noise, we introduce a statistical representation of the imaging system, which incorporates the noise statistics as a signal-dependent correlated noise model. Moreover, in order to compensate for the model errors, we introduce an explicit data-driven correction model, which can be integrated with any regularization term. This constrained optimization problem is solved using fixed-point gradient descent where the analytical gradient of the inaccurate data-fidelity term is corrected using a neural network, while regularization is achieved by data-driven unrolled regularization by denoising (RED). Both networks are jointly trained in an end-to-end manner.
A lightweight learning-based exposure bracketing strategy is proposed in this paper for high dynamic range (HDR) imaging without access to camera RAW. Some low-cost, power-efficient cameras, such as webcams, video surveillance cameras, sport cameras, mid-tier cellphone cameras, and navigation cameras on robots, can only provide access to 8-bit low dynamic range (LDR) images. Exposure fusion is a classical approach to capture HDR scenes by fusing images taken with different exposures into a 8-bit tone-mapped HDR image. A key question is what the optimal set of exposure settings are to cover the scene dynamic range and achieve a desirable tone. The proposed lightweight neural network predicts these exposure settings for a 3-shot exposure bracketing, given the input irradiance information from 1) the histograms of an auto-exposure LDR preview image, and 2) the maximum and minimum levels of the scene irradiance. Without the processing of the preview image streams, and the circuitous route of first estimating the scene HDR irradiance and then tone-mapping to 8-bit images, the proposed method gives a more practical HDR enhancement for real-time and on-device applications. Experiments on a number of challenging images reveal the advantages of our method in comparison with other state-of-the-art methods qualitatively and quantitatively.
A novel iterative linear classification algorithm is developed from a maximum likelihood (ML) linear classifier. The main contribution of this paper is the discovery that a well-known maximum likelihood linear classifier with regularization is the solution to a contraction mapping for an acceptable range of values of the regularization parameter. Hence, a novel iterative scheme is proposed that converges to a fixed point, the globally optimum solution. To the best of our knowledge, this formulation has not been discovered before. Furthermore, the proposed iterative solution converges to a fixed point at a rate faster than the traditional gradient descent technique. The performance of the proposed iterative solution is compared to conventional gradient descent methods on linear and non-linearly separable data in terms of both convergence speed and overall classification performance.
In this paper, we propose a multimodal unsupervised video learning algorithm designed to incorporate information from any number of modalities present in the data. We cooperatively train a network corresponding to each modality: at each stage of training, one of these networks is selected to be trained using the output of the other networks. To verify our algorithm, we train a model using RGB, optical flow, and audio. We then evaluate the effectiveness of our unsupervised learning model by performing action classification and nearest neighbor retrieval on a supervised dataset. We compare this triple modality model to contrastive learning models using one or two modalities, and find us-ing all three modalities in tandem provides a 1.5% improvement in UCF101 classification accuracy, a 1.4% improvement in R@1 retrieval recall, a 3.5% improvement in R@5 retrieval recall, and a 2.4% improvement in R@10 retrieval recall as compared to using only RGB and optical flow, demonstrating the merit of utilizing as many modalities as possible in a cooperative learning model.
The light-field display (LfD) radiance image is a raster description of a light-field where every pixel in the image represents a unique ray within a 3D volume. The LfD radiance image can be projected through an array of micro-lenses to project a perspective-correct 3D aerial image visible for all viewers within the LfDs projection frustum. The synthetic LfD radiance image is comparable to the radiance image as captured by a plenoptic/light-field camera but is rendered from a 3D model or scene. Synthetic radiance image rasterization is an example of extreme multi-view rendering as the 3D scene must be rendered from many (1,000s to millions) viewpoints into small viewports per update of the light-field display. However, GPUs and their accompanying APIs (OpenGL, DirectX, Vulkan) generally expect to render a 3D scene from one viewpoint to a single large viewport/framebuffer. Therefore, LfD radiance image rendering is extremely time consuming and compute intensive. This paper reviews the novel, full-parallax, BowTie Radiance Image Rasterization algorithm which can be embedded within an LfD to accelerate light-field radiance image rendering for real-time update.
Phase retrieval (PR) concerns the recovery of complex phases from complex magnitudes. We identify the connection between the difficulty level and the number and variety of symmetries in PR problems. We focus on the most difficult far-field PR (FFPR), and propose a novel method using double deep image priors. In realistic evaluation, our method outperforms all competing methods by large margins. As a single-instance method, our method requires no training data and minimal hyperparameter tuning, and hence enjoys good practicality.
One of the main problems of neural network-based no-reference metrics design for image visual quality assessment is small size of image databases with mean opinion scores (MOS). For large networks which can memorize key features of several thousands of images, usage of the databases for metrics training may lead to overlearning. Since data augmentation for image quality assessment is limited by a horizontal image flipping only, the main way to decrease overlearning is to use transfer learning which can significantly speed up training process. In theis paper, we propose a new technique of transfer learning between networks of different architectures using a large set of images without MOS. We implemented the technique for transfer learning between pre-trained KonCept512 metric and a IMQNet metric proposed in this paper. An effectiveness of the transfer learning is estimated in a numerical analysis. It is shown that the trained IMQNet metric provides significantly better correlation with KonCept512 metric (0.89) than other modern metrics. It is also shown that IMQNet pre-trained by the proposed transfer learning shows better correlation with MOS of KonIQ-10k database (0.86) than IMQNet pre-trained using directly the MOS of KonIQ10k (0.73).
Diagnosing ligament injuries using MRI scans is a labor-intensive task that requires an expert. In this paper, we propose a fully recurrent neural network (RNN) for detecting Anterior Cruciate Ligament (ACL) tears using MRI scans. The proposed network localizes the ACL and classifies it into several categories: ACL tear, normal tear, and healthy. Existing detection methods use deep learning networks based on single MRI sections, and in this way lose 3D spatial context. To address this, we propose a fully recurrent neural network that processes a sequence of 3D sections and so captures 3D spatial context. The proposed network is based on a YOLOv3 backbone and can produce a sequence of decisions which are then combined by majority voting. Experimental results show improvement over state-of-the-art methods.
In this paper, a convolutional neural network for joint image demosaicing, denoising, deblurring, super-resolution and clarity enhancement is proposed. The network inputs are four-channel Bayer CFA image (R, G, G, B) and three channels of the same size containing distortions maps, namely, noise level map, blur level map, and clarity degradation map. It is shown that the designed network FiveNet can effectively process images with the mix of five different distortions. It is also demonstrated that adding clarity enhancement into the processing chain can additionally increase image quality (by up to 3-4 dB in PSNR). A small dataset ClarityDegr120 of color images with different clarity degradations and enhancements is designed using images processed by FiveNet. Mean opinion scores (MOS) for the test set are collected. The MOS prove that clarity enhancement can significantly increase image visual quality. A comparative analysis using the MOS demonstrates a low correspondence between image quality metrics and human perception for the clarity enhancement task.
Learned image compression methods generally optimize a rate-distortion loss, trading off improvements in visual distortion for added bitrate. Increasingly, however, compressed imagery is used as an input to deep learning networks for various tasks such as classification, object detection, and super-resolution. We propose a recognition-aware learned compression method, which optimizes a rate-distortion loss alongside a task-specific loss, jointly learning compression and recognition networks. We augment a hierarchical autoencoder-based compression network with an EfficientNet recognition model and use two hyperparameters to trade off between distortion, bitrate, and recognition performance. We characterize the classification accuracy of our proposed method as a function of bitrate and find that for low bitrates our method achieves as much as 26% higher recognition accuracy at equivalent bitrates compared to traditional methods such as Better Portable Graphics (BPG).
When an image is captured using an electronic sensor, statistical variations introduced by photon shot and other noise introduce errors in the raw value reported for each pixel sample. Earlier work found that modest improvements in raw image data quality reliably could be obtained by using empirically-determined pixel value error bounds to constrain texture synthesis. However, the prototype software implementation, KREMY (KentuckY Raw Error Modeler, pronounced “creamyâ€), was not effective in processing very noisy images. In comparison, the current work has reimplemented KREMY to make it capable of credibly improving far noisier raw DNG images. The key is a new approach that uses a simpler, but statistical, model for pixel value errors rather than simple bounds constraints.
Noise parameters estimation is needed for many tasks of digital image processing. Many efficient algorithms of noise variance estimation were proposed during last two decades. However, most of those estimators are efficient only for a specific kind of noise for which they were designed. For example, methods of estimation of variance of white additive Gaussian noise (AWGN) fail in the case of additive colored Gaussian noise (ACGN) or for noises with other distributions. In this paper a new fully blind method of noise level estimation is proposed. For a given image, a distorted image with a removed part of pixels (around 10%) is generated. Then an inpainting (or impulse noise removal) method is used to recover missed pixels values. The difference between true and recovered values is used for a robust estimation of noise level. The algorithm is applied for different image scales to estimate noise spectrum. In the paper we propose a convolutional neural network PIXPNet for effective prediction of values of missing pixels. A comparative analysis shows that the proposed PIXPNet provides smallest error of recovered pixels values among all existing methods. A good efficiency of usage of the proposed approach in both AWGN and spatially correlated noise suppression is demonstrated.
Measuring the shape, motion, and physical properties of oscillating fluids is critical for understanding the physics of fluid systems and optimizing and controlling them in real-time. Conventional surface measurement techniques such as profile analysis or stereo reconstruction are not effective for monitoring fluids in industrial processes due to occluding structures, extreme heat, and complex light interactions at the fluid surface. We pro-pose a video-based method comprising forward and inverse transforms. The forward transform employs a physics-based fluid surface model combined with a ray-traced renderer to map shape and motion parameters to synthetic video frames. The inverse transform uses machine learning models to recover surface parameters from video. The inverse models are trained on synthetic data generated by the forward transform. We illustrate the method on an industrial 3D printer for which we recover the motion and surface of a molten aluminum alloy oscillating inside a microscopic nozzle. The inverse transform is ill-posed but can be regularized. We show that surface properties can be reliably inferred with ei-ther a suitably regularized nearest neighbor regressor or a deep convolutional network whose results are less stable but faster to compute.
Detection of moving foreground objects is essential to many image-sequence-analysis applications. However, preexisting methods tend to work best when the foreground is visually distinct from the background, suffering when objects are camouflaged. To address this shortcoming, a foreground-extraction algorithm resilient to camouflage is proposed by incorporating a redundant discrete wavelet transform into the well-known DECOLOR technique based on a sparse and low-rank model of the foreground-extraction problem. Detection of camouflaged moving objects is enhanced as a result of the combination of multiple background estimates in independent wavelet subbands into an overall estimate of the background, leveraging the known robustness of redundant wavelet transforms to additive noise. Experimental results demonstrate that the proposed method offers robustness to camouflage superior to that of other competing methods for image sequences containing snow leopards in the wild.
tromagnetic field distributions of electronic and magnetic materials. The signal-to-noise ratio of electron hologram decreases when the electron beam irradiation is reduced to avoid unnecessary charging and damage. Noise in the hologram causes phase errors. To obtain accurate phase information, we propose an aperture optimization of Fourier based phase reconstruction. Our method effectively separates the signal from the noise using an extended Fourier ring correlation. From the experimental results using a simulated electron hologram with low signal-to-noise, it was found that the proposed method reduces the phase error to 41% of the conventional method. When applied to a real hologram, the proposed method achieved smoother phase reproduction.
With the recent advance in video super-resolution (VSR) techniques, there have been many requests for super-resolve real-world old analog TV series into high-definition digital content. As excellent classical TV series may receive little to no attention due to their poor video quality, restoring them would open new business opportunities for reusing old TV contents. A problem with restoring real-world old TV series is in the complex artifacts introduced by the old interlaced scanning and compression artifacts during the digitization of old analog videos. Though recent DNN-based VSR models perform nicely on clean videos, due to the artificial nature of interlacing and compression artifacts, they fail to restore old videos into a high-definition counterpart free from noticeable artifacts. In this work, we propose OldVSR for restoring old real-world TV series with artifacts of artificial nature. The proposed model implements a bidirectional recurrent structure with first and second-order propagation where each recurrent layer implements two main functions, i.e., Feature alignment (FA) and Pyramid feature aggregation (PFA). The outputs of the forward and backward layers are merged and upsampled to produce a High-Definition (HD) frame of the input standard-definition (SD) frame. We demonstrate through experiments that our proposed OldVSR can effectively remove artifacts of artificial nature from old videos and successfully restores old TV series.
Phase unwrapping is an integral part of multiple imaging techniques, and as a result, a wide range of algorithms have been created to unwrap phases. One such algorithm is the minimum Lp-norm phase unwrapping algorithm. This algorithm transforms the phase unwrapping problem into a minimization problem of a certain functional, which it solves with an iterative method. However, the problem is usually not convex, and when there are many sharp edges in the data to be unwrapped, the algorithm often produces a local minimum with new discontinuities in originally smooth areas. To prioritize solutions which minimize the functional better in smooth areas, we use weights to deprioritize data lying along edges in the ground-truth image. This requires a method to find ground-truth edges using the wrapped image, which we describe. When using the modified algorithm, we generally obtain improved results on images with multiple edges (both lower errors and more correct edge placement).
We measured the contrast of standard charts using two different types of retro-reflectors in an AIRR (Aerial imaging by retro-reflection) system, and examined the results to be reproduced by optical simulation. As a result, it became possible to reproduce the effect of retro-reflector diffraction on the Aerial image quality of the AIRR system by optical simulation.