We present a novel neural algorithm for performing high-quality, highresolution, real-time novel view synthesis. From a sparse set of input RGB images or videos streams, our network both reconstructs the 3D scene and renders novel views at 1080p resolution at 30fps on an NVIDIA A100. Our feed-forward network generalizes across a wide variety of datasets and scenes and produces state-of-the-art quality for a real-time method. Our quality approaches, and in some cases surpasses, the quality of some of the top offline methods. In order to achieve these results we use a novel combination of several key concepts, and tie them together into a cohesive and effective algorithm. We build on previous works that represent the scene using semi-transparent layers and use an iterative learned render-and-refine approach to improve those layers. Instead of flat layers, our method reconstructs layered depth maps (LDMs) that efficiently represent scenes with complex depth and occlusions. The iterative update steps are embedded in a multi-scale, UNet-style architecture to perform as much compute as possible at reduced resolution. Within each update step, to better aggregate the information from multiple input views, we use a specialized Transformer-based network component. This allows the majority of the per-input image processing to be performed in the input image space, as opposed to layer space, further increasing efficiency. Finally, due to the real-time nature of our reconstruction and rendering, we dynamically create and discard the internal 3D geometry for each frame, generating the LDM for each view. Taken together, this produces a novel and effective algorithm for view synthesis. Through extensive evaluation, we demonstrate that we achieve state-of-the-art quality at real-time rates.
This paper presents a novel approach to inpainting 3D regions of a scene, given masked multi-view images, by distilling a 2D diffusion model into a learned 3D scene representation (e.g. a NeRF). Unlike 3D generative methods that explicitly condition the diffusion model on camera pose or multi-view information, our diffusion model is conditioned only on a single masked 2D image. Nevertheless, we show that this 2D diffusion model can still serve as a generative prior in a 3D multi-view reconstruction problem where we optimize a NeRF using a combination of score distillation sampling and NeRF reconstruction losses. Predicted depth is used as additional supervision to encourage accurate geometry. We compare our approach to 3D inpainting methods that focus on object removal. Because our method can generate content to fill any 3D masked region, we additionally demonstrate 3D object completion, 3D object replacement, and 3D scene completion.
Differentiable simulations of optical systems can be combined with deep learning-based reconstruction networks to enable high performance computational imaging via end-to-end (E2E) optimization of both the optical encoder and the deep decoder. This has enabled imaging applications such as 3D localization microscopy, depth estimation, and lensless photography via the optimization of local optical encoders. More challenging computational imaging applications, such as 3D snapshot microscopy which compresses 3D volumes into single 2D images, require a highly non-local optical encoder. We show that existing deep network decoders have a locality bias which prevents the optimization of such highly non-local optical encoders. We address this with a decoder based on a shallow neural network architecture using global kernel Fourier convolutional neural networks (FourierNets). We show that FourierNets surpass existing deep network based decoders at reconstructing photographs captured by the highly non-local DiffuserCam optical encoder. Further, we show that FourierNets enable E2E optimization of highly non-local optical encoders for 3D snapshot microscopy. By combining FourierNets with a large-scale multi-GPU differentiable optical simulation, we are able to optimize non-local optical encoders 170$\times$ to 7372$\times$ larger than prior state of the art, and demonstrate the potential for ROI-type specific optical encoding with a programmable microscope.
Current 3D localization microscopy approaches are fundamentally limited in their ability to image thick, densely labeled specimens. Here, we introduce a hybrid optical-electronic computing approach that jointly optimizes an optical encoder (a set of multiple, simultaneously imaged 3D point spread functions) and an electronic decoder (a neural-network-based localization algorithm) to optimize 3D localization performance under these conditions. With extensive simulations and biological experiments, we demonstrate that our deep-learning-based microscope achieves significantly higher 3D localization accuracy than existing approaches, especially in challenging scenarios with high molecular density over large depth ranges.
We propose a neural point spread function (PSF) engineering approach for 3D single-molecule localization microscopy. Our approach jointly optimizes the PSFs of multiple optical paths in a microscope with a differentiable localization algorithm.
This Immersive Pavilion installation introduces our new system for capturing, reconstructing, compressing, and rendering light field video content. By leveraging DeepView, a recently introduced view synthesis algorithm, our system can reconstruct challenging scenes with view-dependent reflections, semi-transparent surfaces, and near-field objects as close as 34 cm to the surface of our 46 camera capture rig. Improving upon past light field video systems that required specialized storage and graphics hardware for playback, our compressed videos can be rendered in a web browser or on mobile VR headsets while being streamed over a gigabit network connection. This makes ours the first system to encode high quality light field video at sufficiently low bandwidth for internet streaming.
We present a system for capturing, reconstructing, compressing, and rendering high quality immersive light field video. We accomplish this by leveraging the recently introduced DeepView view interpolation algorithm, replacing its underlying multi-plane image (MPI) scene representation with a collection of spherical shells that are better suited for representing panoramic light field content. We further process this data to reduce the large number of shell layers to a small, fixed number of RGBA+depth layers without significant loss in visual quality. The resulting RGB, alpha, and depth channels in these layers are then compressed using conventional texture atlasing and video compression techniques. The final compressed representation is lightweight and can be rendered on mobile VR/AR platforms or in a web browser. We demonstrate light field video results using data from the 16-camera rig of [Pozo et al. 2019] as well as a new low-cost hemispherical array made from 46 synchronized action sports cameras. From this data we produce 6 degree of freedom volumetric videos with a wide 70 cm viewing baseline, 10 pixels per degree angular resolution, and a wide field of view, at 30 frames per second video frame rates. Advancing over previous work, we show that our system is able to reproduce challenging content such as view-dependent reflections, semi-transparent surfaces, and near-field objects as close as 34 cm to the surface of the camera rig.
We present a portable multi-camera system for recording panoramic light field video content. The proposed system captures wide baseline (0.8 meters), high resolution (>15 pixels per degree), large field of view (>220°) light fields at 30 frames per second. The array contains 47 time-synchronized cameras distributed on the surface of a hemispherical, 0.92 meter diameter plastic dome. We use commercially available action sports cameras (Yi 4k) mounted inside the dome using 3D printed brackets. The dome, mounts, triggering hardware and cameras are inexpensive and the array itself is easy to fabricate. Using modern view interpolation algorithms we can render objects as close as 33-cm to the surface of the array.
Whole-brain recordings give us a global perspective of the brain in action. In this study, we describe a method using light field microscopy to record near-whole brain calcium and voltage activity at high speed in behaving adult flies. We first obtained global activity maps for various stimuli and behaviors. Notably, we found that brain activity increased on a global scale when the fly walked but not when it groomed. This global increase with walking was particularly strong in dopamine neurons. Second, we extracted maps of spatially distinct sources of activity as well as their time series using principal component analysis and independent component analysis. The characteristic shapes in the maps matched the anatomy of subneuropil regions and, in some cases, a specific neuron type. Brain structures that responded to light and odor were consistent with previous reports, confirming the new technique’s validity. We also observed previously uncharacterized behavior-related activity as well as patterns of spontaneous voltage activity.
We present a novel approach to view synthesis using multiplane images (MPIs). Building on recent advances in learned gradient descent, our algorithm generates an MPI from a set of sparse camera viewpoints. The resulting method incorporates occlusion reasoning, improving performance on challenging scene features such as object boundaries, lighting reflections, thin structures, and scenes with high depth complexity. We show that our method achieves high-quality, state-of-the-art results on two datasets: the Kalantari light field dataset, and a new camera array dataset, Spaces, which we make publicly available.
Various aspects of the present disclosure are directed toward optics and imaging. As may be implemented with one or more embodiments, an apparatus includes one or more phase masks that operate with an objective lens and a microlens array to alter a phase characteristic of light travelling in a path from a specimen, through the objective lens and microlens array and to a photosensor array. Using this approach, the specimen can be imaged with spatial resolution characteristics provided via the altered phase characteristic, which can facilitate construction of an image with enhanced resolution.
We present a variety of new compositing techniques using Multi-plane Images (MPI's) [Zhou et al. 2018] derived from footage shot with an inexpensive and portable light field video camera array. The effects include camera stabilization, foreground object removal, synthetic depth of field, and deep compositing. Traditional compositing is based around layering RGBA images to visually integrate elements into the same scene, and often requires manual 2D and/or 3D artist intervention to achieve realism in the presence of volumetric effects such as smoke or splashing water. We leverage the newly introduced DeepView solver [Flynn et al. 2019] and a light field camera array to generate MPIs stored in the DeepEXR format for compositing with realistic spatial integration and a simple workflow which offers new creative capabilities. We demonstrate using this technique by combining footage that would otherwise be very challenging and time intensive to achieve when using traditional techniques, with minimal artist intervention.
Prolonged behavioral challenges can cause animals to switch from active to passive coping strategies to manage effort-expenditure during stress; such normally adaptive behavioral state transitions can become maladaptive in psychiatric disorders such as depression. The underlying neuronal dynamics and brainwide interactions important for passive coping have remained unclear. Here, we develop a paradigm to study these behavioral state transitions at cellular-resolution across the entire vertebrate brain. Using brainwide imaging in zebrafish, we observed that the transition to passive coping is manifested by progressive activation of neurons in the ventral (lateral) habenula. Activation of these ventral-habenula neurons suppressed downstream neurons in the serotonergic raphe nucleus and caused behavioral passivity, whereas inhibition of these neurons prevented passivity. Data-driven recurrent neural network modeling pointed to altered intra-habenula interactions as a contributory mechanism. These results demonstrate ongoing encoding of experience features in the habenula, which guides recruitment of downstream networks and imposes a passive coping behavioral strategy.
Deconvolution is widely used to improve the contrast and clarity of a 3D focal stack collected using a fluorescence microscope. But despite being extensively studied, deconvolution algorithms can introduce reconstruction artifacts when their underlying noise models or priors are violated, such as when imaging biological specimens at extremely low light levels. In this paper we propose a deconvolution method specifically designed for 3D fluorescence imaging of biological samples in the low-light regime. Our method utilizes a mixed Poisson-Gaussian model of photon shot noise and camera read noise, which are both present in low light imaging. We formulate a convex loss function and solve the resulting optimization problem using the alternating direction method of multipliers algorithm. Among several possible regularization strategies, we show that a Hessian-based regularizer is most effective for describing locally smooth features present in biological specimens. Our algorithm also estimates noise parameters on-the-fly, thereby eliminating a manual calibration step required by most deconvolution software. We demonstrate our algorithm on simulated images and experimentally-captured images with peak intensities of tens of photoelectrons per voxel. We also demonstrate its performance for live cell imaging, showing its applicability as a tool for biological research.
Tracking the coordinated activity of cellular events through volumes of intact tissue is a major challenge in biology that has inspired significant technological innovation. Yet scanless measurement of the high-speed activity of individual neurons across three dimensions in scattering mammalian tissue remains an open problem. Here we develop and validate a computational imaging approach (SWIFT) that integrates high-dimensional, structured statistics with light field microscopy to allow the synchronous acquisition of single-neuron resolution activity throughout intact tissue volumes as fast as a camera can capture images (currently up to 100 Hz at full camera resolution), attaining rates needed to keep pace with emerging fast calcium and voltage sensors. We demonstrate that this large field-of-view, single-snapshot volume acquisition method—which requires only a simple and inexpensive modification to a standard fluorescence microscope—enables scanless capture of coordinated activity patterns throughout mammalian neural volumes. Further, the volumetric nature of SWIFT also allows fast in vivo imaging, motion correction, and cell identification throughout curved subcortical structures like the dorsal hippocampus, where cellular-resolution dynamics spanning hippocampal subfields can be simultaneously observed during a virtual context learning task in a behaving animal. SWIFT’s ability to rapidly and easily record from volumes of many cells across layers opens the door to widespread identification of dynamical motifs and timing dependencies among coordinated cell assemblies during adaptive, modulated, or maladaptive physiological processes in neural systems.
BACKGROUND:The determination and regulation of cell morphology are critical components of cell-cycle control, fitness, and development in both single-cell and multicellular organisms. Understanding how environmental factors, chemical perturbations, and genetic differences affect cell morphology requires precise, unbiased, and validated measurements of cell-shape features.RESULTS:Here we introduce two software packages, Morphometrics and BlurLab, that together enable automated, computationally efficient, unbiased identification of cells and morphological features. We applied these tools to bacterial cells because the small size of these cells and the subtlety of certain morphological changes have thus far obscured correlations between bacterial morphology and genotype. We used an online resource of images of the Keio knockout library of nonessential genes in the Gram-negative bacterium Escherichia coli to demonstrate that cell width, width variability, and length significantly correlate with each other and with drug treatments, nutrient changes, and environmental conditions. Further, we combined morphological classification of genetic variants with genetic meta-analysis to reveal novel connections among gene function, fitness, and cell morphology, thus suggesting potential functions for unknown genes and differences in modes of action of antibiotics.CONCLUSIONS:Morphometrics and BlurLab set the stage for future quantitative studies of bacterial cell shape and intracellular localization. The previously unappreciated connections between morphological parameters measured with these software packages and the cellular environment point toward novel mechanistic connections among physiological perturbations, cell fitness, and growth.
We describe a method to record near-whole brain activity in behaving adult flies. Pan-neuronal calcium and voltage sensors were imaged at high speed with light field microscopy. Functional maps were then extracted by principal component analysis and independent component analysis. Their characteristic shapes match the anatomy of sub-neuropil regions and in some cases a specific neuron type. Responses to light and odor produced activity patterns consistent with previous techniques. Furthermore, the method detected activity linked to behavior as well as a previously uncharacterized pattern of spontaneous activity in the central complex.
We provide a method to record near-whole brain activity in behaving adult flies and extract signals from specific anatomical structures. We image pan-neuronal calcium or voltage sensors fluorescence at high speed with light field microscopy. We then apply computational methods to extract functional maps, and find that their characteristic shape match the anatomy of sub-neuropile regions and sometimes small population of neurons. Associated time series are also consistent with the literature.