Four-dimensional (4D; 3D + time) microscopic imaging has emerged as a powerful technique for investigating dynamic phenomena in complex systems, enabling direct visualization of structural evolution in space and time. However, when pushing the limits of spatiotemporal resolution, most time-resolved imaging techniques yield inherently sparse 4D datasets. While deep learning-based reconstruction methods have shown promise in reconstructing 4D from sparse spatiotemporal measurements, a practical approach for evaluating their performance in the absence of a 4D reference has, to the best of our knowledge, been lacking. Here, we present a bootstrapped cross-validation framework that estimates reconstruction performance by quantifying correlations between reconstructions generated from independently sampled subsets of the acquired data, as inspired by the 3D validation strategy in cryo-electron microscopy, where reconstructions from split datasets are compared to assess resolutions. This enables both qualitative and quantitative assessment in the absence of ground truth. We investigate two representative scenarios with sparse and ultra-sparse X-ray datasets and validate this approach using 4D-ONIX, a 4D deep-learning reconstruction method, on simulated water droplet collision experiments. The proposed approach provides a reference-free framework for performance estimation and support for better-informed experimental strategies across a wide range of ultrafast imaging applications.
Decomposing 3D assets into material parts is a common task for artists, yet remains a highly manual process. In this work, we introduce Select Any Material (SAMa), a material selection approach for in-the-wild objects in arbitrary 3D representations. Building on SAM2's video prior, we construct a material-centric video dataset that extends it to the material domain. We propose an efficient way to lift the model's 2D predictions to 3D by projecting each view into an intermediary 3D point cloud using depth. Nearest-neighbor lookups between any 3D representation and this similarity point cloud allow us to efficiently reconstruct accurate selection masks over objects' surfaces that can be inspected from any view. Our method is multiview-consistent by design, alleviating the need for costly per-asset optimization, and performs optimization-free selection in seconds. SAMa outperforms several strong baselines in selection accuracy and multiview consistency and enables various compelling applications, such as replacing the diffuse-textured materials on a text-to-3D output with PBR materials or selecting and editing materials on NeRFs and 3DGS captures.
Volumetric inverse rendering seeks to recover the optical properties of participating media from images. Existing approaches either rely on differentiable stochastic light transport simulation, which require substantial algorithmic effort, or use simplified models that fail to capture global illumination. We propose a formulation that reconciles physically complete light transport with general-purpose neural optimization. The optical properties of the medium and the full light field are represented as neural fields and estimated through a joint optimization process. Global illumination is enforced via a residual objective derived from the Radiative Transfer Equation in local differential form, complemented by a volume rendering term along primary viewing rays to mitigate low-frequency bias. We demonstrate reconstruction of spatially varying, color-resolved scattering, absorption, and phase function parameters from multi-view images. Beyond reconstruction, the same framework supports learning generative models of participating media with physical optical properties under global illumination.
We propose a system to optimize parametric designs subject to radiation pressure, the effect of light on the motion of objects. This is most relevant in the design of spacecraft, where radiation pressure presents the dominant non-conservative forcing mechanism, which is the case beyond approximately 800 km altitude. Despite its importance, the high computational cost of high-fidelity radiation pressure modeling has limited its use in large-scale spacecraft design, optimization, and space situational awareness applications. We enable this by offering three innovations in the simulation, in representation and in optimization: First, a practical computer graphics-inspired Monte-Carlo (MC) simulation of radiation pressure. The simulation is highly parallel, uses importance sampling and next-event estimation to reduce variance and allows simulating an entire family of designs instead of a single spacecraft as in previous work. Second, we introduce neural networks as a representation of forces from design parameters. This neural proxy model, learned from simulations, is inherently differentiable and can query forces orders of magnitude faster than a full MC simulation. Third, and finally, we demonstrate optimizing inverse radiation pressure designs, such as finding geometry, material or operation parameters that minimizes travel time, maximizes proximity given a desired end-point, minimize thruster fuel, trains mission control policies or allocated compute budget in extraterrestrial compute.
We derive methods to compute higher order differentials (Hessians and Hessian-vector products) of the rendering operator. Our approach is based on importance sampling of a convolution that represents the differentials of rendering parameters and shows to be applicable to both rasterization and path tracing. We further suggest an aggregate sampling strategy to importance-sample multiple dimensions of one convolution kernel simultaneously. We demonstrate that this information improves convergence when used in higher-order optimizers such as Newton or Conjugate Gradient relative to a gradient descent baseline in several inverse rendering tasks.
Neural fields offer continuous, learnable representations that extend beyond traditional discrete formats in visual computing. We study the problem of learning neural representations of repeated antiderivatives directly from a function, a continuous analogue of summed-area tables. Although widely used in discrete domains, such cumulative schemes rely on grids, which prevents their applicability in continuous neural contexts. We introduce and analyze a range of neural methods for repeated integration, including both adaptations of prior work and novel designs. Our evaluation spans multiple input dimensionalities and integration orders, assessing both reconstruction quality and performance in downstream tasks such as filtering and rendering. These results enable integrating classical cumulative operators into modern neural systems and offer insights into learning tasks involving differential and integral operators.
The X-ray flux from X-ray free-electron lasers and storage rings enables new spatiotemporal opportunities for studying in-situ and operando dynamics, even with single pulses. X-ray multi-projection imaging is a technique that provides volumetric information using single pulses while avoiding the centrifugal forces induced by conventional time-resolved 3D methods like time-resolved tomography, and can acquire 3D movies (4D) at least three orders of magnitude faster than existing techniques. However, reconstructing 4D information from highly sparse projections remains a challenge for current algorithms. Here we present 4D-ONIX, a deep-learning-based approach that reconstructs 3D movies from an extremely limited number of projections. It combines the computational physical model of X-ray interaction with matter and state-of-the-art deep learning methods. We demonstrate its ability to reconstruct high-quality 4D by generalizing over multiple experiments with only two to three projections per timestamp on simulations of water droplet collisions and experimental data of additive manufacturing. Our results demonstrate 4D-ONIX as an enabling tool for 4D analysis, offering high-quality image reconstruction for fast dynamics three orders of magnitude faster than tomography.
Real camera footage is subject to noise, motion blur (MB) and depth of field (DoF). In some applications these might be considered distortions to be removed, but in others it is important to model them because it would be ineffective, or interfere with an aesthetic choice, to simply remove them. In augmented reality applications where virtual content is composed into a live video feed, we can model noise, MB and DoF to make the virtual content visually consistent with the video. Existing methods for this typically suffer two main limitations. First, they require a camera calibration step to relate a known calibration target to the specific cameras response. Second, existing work require methods that can be (differentiably) tuned to the calibration, such as slow and specialized neural networks. We propose a method which estimates parameters for noise, MB and DoF instantly, which allows using off-the-shelf real-time simulation methods from e.g., a game engine in compositing augmented content. Our main idea is to unlock both features by showing how to use modern computer vision methods that can remove noise, MB and DoF from the video stream, essentially providing self-calibration. This allows to auto-tune any black-box real-time noise+MB+DoF method to deliver fast and high-fidelity augmentation consistency.
Creation of chirality through screw dislocation-driven growth for highly crystalline nano-WSe 2 by chemical vapor transport based on thermodynamic simulations.
We demonstrate generating HDR images using the concerted action of multiple black-box, pre-trained LDR image diffusion models. Common diffusion models are not HDR as, first, there is no sufficiently large HDR image dataset available to re-train them, and, second, even if it was, re-training such models is impossible for most compute budgets. Instead, we seek inspiration from the HDR image capture literature that traditionally fuses sets of LDR images, called "exposure brackets", to produce a single HDR image. We operate multiple denoising processes to generate multiple LDR brackets that together form a valid HDR result. To this end, we introduce a brackets consistency term into the diffusion process to couple the brackets such that they agree across the exposure range they share. We demonstrate HDR versions of state-of-the-art unconditional and conditional as well as restoration-type (LDR2HDR) generative modeling.
The unprecedented x-ray flux density provided by modern x-ray sources offers new spatiotemporal possibilities for x-ray imaging of fast dynamic processes. Approaches to exploit such possibilities often result in either (i) a limited number of projections or spatial information due to limited scanning speed, as in time-resolved tomography, or (ii) a limited number of time points, as in stroboscopic imaging, making the reconstruction problem ill-posed and unlikely to be solved by classical reconstruction approaches. Four-dimensional (4D) reconstruction from such data requires sample priors, which can be included via deep learning (DL). State-of-the-art 4D reconstruction methods for x-ray imaging combine the power of artificial intelligence and the physics of x-ray propagation to tackle the challenge of sparse views. However, most approaches do not constrain the physics of the studied process, i.e. a full physical model. Here we present 4D physics-informed optimized neural implicit x-ray imaging, a novel physics-informed 4D x-ray image reconstruction method combining the full physical model and a state-of-the-art DL-based reconstruction method for 4D x-ray imaging from sparse views. We demonstrate and evaluate the potential of our approach by retrieving 4D information from ultra-sparse spatiotemporal acquisitions of simulated binary droplet collisions, a relevant fluid dynamic process. We envision that this work will open new spatiotemporal possibilities for various 4D x-ray imaging modalities, such as time-resolved x-ray tomography and more novel sparse acquisition approaches like x-ray multi-projection imaging, which will pave the way for investigations of various rapid 4D dynamics, such as fluid dynamics and composite testing.
We propose a novel generative video model to robustly learn temporal change as a neural Ordinary Differential Equation (ODE) flow with a bilinear objective which combines two aspects: The first is to map from the past into future video frames directly. Previous work has mapped the noise to new frames, a more computationally expensive process. Unfortunately, starting from the previous frame, instead of noise, is more prone to drifting errors. Hence, second, we additionally learn how to remove the accumulated errors as the joint objective by adding noise during training. We demonstrate unconditional video generation in a streaming manner for various video datasets, all at competitive quality compared to a conditional diffusion baseline but with higher speed, i.e., fewer ODE solver steps.
The X-ray flux provided by X-ray free-electron lasers and storage rings offers new spatiotemporal possibilities to study in-situ and operando dynamics, even using single pulses of such facilities. X-ray Multi-Projection Imaging (XMPI) is a novel technique that enables volumetric information using single pulses of such facilities and avoids centrifugal forces induced by state-of-the-art time-resolved 3D methods such as time-resolved tomography. As a result, XMPI offers the potential to acquire 3D movies (4D) at least three orders of magnitude faster than current methods. However, it is exceptionally challenging to reconstruct 4D from highly sparse projections as acquired by XMPI with current algorithms. Here, we present 4D-ONIX, a Deep Learning (DL)-based approach that learns to reconstruct 3D movies (4D) from an extremely limited number of projections. It combines the computational physical model of X-ray interaction with matter and state-of-the-art DL methods. We demonstrate the potential of 4D-ONIX to generate high-quality 4D by generalizing over multiple experiments with only two to three projections per timestamp for binary droplet collisions and additive manufacturing. We envision that 4D-ONIX will become an enabling tool for 4D analysis, offering new spatiotemporal resolutions to study processes not possible before.
A Neural Radiance Field (NeRF) encodes the specific relation of 3D geometry and appearance of a scene. We here ask the question whether we can transfer the appearance from a source NeRF onto a target 3D geometry in a semantically meaningful way, such that the resulting new NeRF retains the target geometry but has an appearance that is an analogy to the source NeRF. To this end, we generalize classic image analogies from 2D images to NeRFs. We leverage correspondence transfer along semantic affinity that is driven by semantic features from large, pre-trained 2D image models to achieve multi-view consistent appearance transfer. Our method allows exploring the mix-and-match product space of 3D geometry and appearance. We show that our method outperforms traditional stylization-based methods and that a large majority of users prefer our method over several typical baselines.
Bounding volumes are an established concept in computer graphics and vision tasks but have seen little change since their early inception. In this work, we study the use of neural networks as bounding volumes. Our key observation is that bounding, which so far has primarily been considered a problem of computational geometry, can be redefined as a problem of learning to classify space into free or occupied. This learning-based approach is particularly advantageous in high-dimensional spaces, such as animated scenes with complex queries, where neural networks are known to excel. However, unlocking neural bounding requires a twist: allowing -- but also limiting -- false positives, while ensuring that the number of false negatives is strictly zero. We enable such tight and conservative results using a dynamically-weighted asymmetric loss function. Our results show that our neural bounding produces up to an order of magnitude fewer false positives than traditional methods.
Differentiable rasterization changes the standard formulation of primitive rasterization -- by enabling gradient flow from a pixel to its underlying triangles -- using distribution functions in different stages of rendering, creating a "soft" version of the original rasterizer. However, choosing the optimal softening function that ensures the best performance and convergence to a desired goal requires trial and error. Previous work has analyzed and compared several combinations of softening. In this work, we take it a step further and, instead of making a combinatorial choice of softening operations, parameterize the continuous space of common softening operations. We study meta-learning tunable softness functions over a set of inverse rendering tasks (2D and 3D shape, pose and occlusion) so it generalizes to new and unseen differentiable rendering tasks with optimal softness.
Gradient-based optimization is now ubiquitous across graphics, but unfortunately can not be applied to problems with undefined or zero gradients. To circumvent this issue, the loss function can be manually replaced by a “surrogate” that has similar minima but is differentiable. Our proposed framework, ZeroGrads, automates this process by learning a neural approximation of the objective function, which in turn can be used to differentiate through arbitrary black-box graphics pipelines. We train the surrogate on an actively smoothed version of the objective and encourage locality, focusing the surrogate's capacity on what matters at the current training episode. The fitting is performed online, alongside the parameter optimization, and self-supervised, without pre-computed data or pre-trained models. As sampling the objective is expensive (it requires a full rendering or simulator run), we devise an efficient sampling scheme that allows for tractable run-times and competitive performance at little overhead. We demonstrate optimizing diverse non-convex, non-differentiable black-box problems in graphics, such as visibility in rendering, discrete parameter spaces in procedural modelling or optimal control in physics-driven animation. In contrast to other derivative-free algorithms, our approach scales well to higher dimensions, which we demonstrate on problems with up to 35k interlinked variables.
Deep-learning methods using unpaired datasets hold great potential for image reconstruction, especially in biomedical imaging where obtaining paired datasets is often difficult due to practical concerns. A recent study by Lee et al. ( Nature Machine Intelligence 2023) has introduced a parameterized physical model (referred to as FMGAN) using the unpaired approach for adaptive holographic imaging, which replaces the forward generator network with a physical model parameterized on the propagation distance of the probing light. FMGAN has demonstrated its capability to reconstruct the complex phase and amplitude of objects, as well as the propagation distance, even in scenarios where the object-to-sensor distance exceeds the range of the training data. We performed additional experiments to comprehensively assess FMGAN’s capabilities and limitations. As in the original paper, we compared FMGAN to two state-of-the-art unpaired methods, CycleGAN and PhaseGAN, and evaluated their robustness and adaptability under diverse conditions. Our findings highlight FMGAN’s reproducibility and generalizability when dealing with both in-distribution and out-of-distribution data, corroborating the results reported by the original authors. We also extended FMGAN with explicit forward models describing the response of specific optical systems, which improved performance when dealing with non-perfect systems. However, we observed that FMGAN encounters difficulties when explicit forward models are unavailable. In such scenarios, PhaseGAN outperformed FMGAN.
We propose a method to reproduce dynamic appearance textures with space-stationary but time-varying visual statistics. While most previous work decomposes dynamic textures into static appearance and motion, we focus on dynamic appearance that results not from motion but variations of fundamental properties, such as rusting, decaying, melting, and weathering. To this end, we adopt the neural ordinary differential equation (ODE) to learn the underlying dynamics of appearance from a target exemplar. We simulate the ODE in two phases. At the "warm-up" phase, the ODE diffuses a random noise to an initial state. We then constrain the further evolution of this ODE to replicate the evolution of visual feature statistics in the exemplar during the generation phase. The particular innovation of this work is the neural ODE achieving both denoising and evolution for dynamics synthesis, with a proposed temporal training scheme. We study both relightable (BRDF) and non-relightable (RGB) appearance models. For both we introduce new pilot datasets, allowing, for the first time, to study such phenomena: For RGB we provide 22 dynamic textures acquired from free online sources; For BRDFs, we further acquire a dataset of 21 flash-lit videos of time-varying materials, enabled by a simple-to-construct setup. Our experiments show that our method consistently yields realistic and coherent results, whereas prior works falter under pronounced temporal appearance variations. A user study confirms our approach is preferred to previous work for such exemplars.
Atomic layer deposition (ALD) is an effective technique for depositing thin films with precise control of layer thickness and functional properties. In this work, Sb2Te3-Sb2Se3 nanostructures were synthesized using thermal ALD. A decrease in the Sb2Te3 layer thickness led to the emergence of distinct peaks from the Laue rings, indicative of a highly textured film structure with optimized crystallinity. Density functional theory simulations revealed that carrier redistribution occurs at the interface to establish charge equilibrium. By carefully optimizing the layer thicknesses, we achieved an obvious enhancement in the Seebeck coefficient, reaching a peak figure of merit (zT) value of 0.38 at room temperature. These investigations not only provide strong evidence for the potential of ALD manipulation to improve the electrical performance of metal chalcogenides but also offer valuable insights into achieving high performance in two-dimensional materials.