Precise measurements of sickness symptoms induced during a virtual reality (VR) experience are essential for evaluating VR systems and developing designs oriented toward usability, safety and user acceptance. However, VR sickness assessment typically relies either on discrete self-report questionnaires (which lack temporal resolution, interrupt the experience, thus reducing immersion, and provide coarse snapshots of symptom evolution) or on objective signals obtained with biosensors, which typically require extensive post-processing and interpretation. To address these shortcomings, we propose a continuous interface for real-time self-reporting of VR sickness, designed following a human-centered methodology. We design and evaluate three interface prototypes that allow users to report symptom intensity while remaining fully immersed in the virtual scene. Our findings demonstrate that users significantly prefer the continuous nature of our interfaces over the discrete Likert Scales of traditional questionnaires, identifying them as a more intuitive and less cognitively demanding alternative. In addition, the study allows us to identify the most suitable design according to user-centered criteria. Our contribution is an empirically evaluated continuous interface for real-time VR sickness assessment.
Time-of-Flight non-line-of-sight (ToF NLOS) imaging techniques provide state-of-the-art reconstructions of scenes hidden around corners by inverting the optical path of indirect photons scattered by visible surfaces and measured by picosecond resolution sensors. The emergence of a wide range of ToF NLOS imaging methods with heterogeneous formulae and hardware implementations obscures the assessment of both their theoretical and experimental aspects. We present a comprehensive study of a representative set of ToF NLOS imaging methods by discussing their similarities and differences under common formulation and hardware. We first outline the problem statement under a common general forward model for ToF NLOS measurements, and the typical assumptions that yield tractable inverse models. We discuss the relationship of the resulting simplified forward and inverse models to a family of Radon transforms, and how migrating these to the frequency domain relates to recent phasor-based virtual line-of-sight imaging models for NLOS imaging that obey the constraints of conventional lens-based imaging systems. We then evaluate performance of the selected methods on hidden scenes captured under the same hardware setup and similar photon counts. Our experiments show that existing methods share similar limitations on spatial resolution, visibility, and sensitivity to noise when operating under equal hardware constraints, with particular differences that stem from method-specific parameters. We expect our methodology to become a reference in future research on ToF NLOS imaging to obtain objective comparisons of existing and new methods.
Transient rendering simulates light in motion, measuring the time of flight from the light source to the camera. However, the stochastic nature of Monte Carlo is aggravated in transient rendering, since samples are now spread along the temporal domain. In our work, we propose to denoise transient Monte Carlo renders by exploiting the spatio-temporal correlation of transient light transport, extending a recent statistical denoising formulation. By relying on statistics, we achieve a near-optimal tradeoff between reduced variance and introduced bias. We efficiently collect per-time-bin statistics in the temporal domain while avoiding impractical memory requirements, and use these collected statistics to analyze the spatio-temporal correlation and discriminate which time bins should be combined. Our statistics-based transient denoiser does not hallucinate, guarantees convergence of the result, is efficient, does not require any training and naturally handles participating media. We believe that the generality of our method might pave the way for denoising time-resolved Monte Carlo simulations in other domains, such as non-line-of-sight imaging, acoustic rendering, or absorption microscopy.
Galaxy clusters are powerful probes of astrophysics and cosmology through gravitational lensing: the clusters' mass, dominated by 85
Non-line-of-sight imaging employs ultra-fast illumination and sensing devices to reconstruct scenes outside their line of sight by analyzing the temporal profile of indirect scattered illumination on a secondary relay surface. Commonly, the NLOS methods transform the temporal domain into the frequency domain and operate on it, and then identify surface locations by locating the maxima in amplitude along the reconstruction volume. Phase information, which is virtual as it results from a Fourier transform, is very often discarded or ignored. We incorporate phase information into our novel Zero-Phase Phasor Fields imaging technique, which we derive for a confocal capture configuration. We show how, at positions that belong to the hidden geometry, we can ensure the phase is zero, so we can locate the hidden geometry with great precision by locating the zero crossings in the phase. This allows us to reconstruct at widely spaced locations and still achieve up to 125 micrometer depth precision, as our experimental validation shows with both synthetic and captured data, the latter publicly available. Moreover, the phase is robust to noise, as we demonstrate with decreasing signal-to-noise ratio using publicly available dataset captures of the same scene.
Images as an artistic medium often rely on specific camera angles and lens distortions to convey ideas or emotions; however, such precise control is missing in current text-to-image models. We propose an efficient and general solution that allows precise control over the camera when generating both photographic and artistic images. Unlike prior methods that rely on predefined shots, we rely solely on four simple extrinsic and intrinsic camera parameters, removing the need for pre-existing geometry, reference 3D objects, and multi-view data. We also present a novel dataset with more than 57,000 images, along with their text prompts and ground-truth camera parameters. Our evaluation shows precise camera control in text-to-image generation, surpassing traditional prompt engineering approaches.
Virtual reality (VR) experiences often leverage rich and spatialized multimodal environments to increase immersion and engagement. This demands a consistent spatial perception of audiovisual stimuli, since perceived discrepancies can disrupt the sense of presence. In this work, we investigate the consequences of two types of spatial audiovisual disparities: true disparity, where there is a measurable spatial offset between auditory and visual cues, and perceptual disparity, where users report misalignment despite cues being colocated. Unlike most previous studies that employed controlled but simplified experimental setups, our research focuses on complex, realistic VR environments, allowing us to assess the actual implications for VR content design. Our experiments indicate that users are highly sensitive to true audiovisual disparities in controlled environments, detecting even minor misalignments. However, when engaged in additional tasks within realistic settings, their ability to notice such discrepancies diminishes significantly. We also observed that previously found perceptual disparities persist in complex audiovisual environments. However, we identify self-initiated head rotations as a key factor; its absence prevents the effect entirely. We hope our findings offer practical insights for designing more immersive and perceptually coherent VR experiences.
Time-gated non-line-of-sight (NLOS) imaging methods reconstruct scenes hidden around a corner by inverting the optical path of indirect photons measured at visible surfaces. These methods are, however, hindered by intricate, timeconsuming calibration processes involving expensive capture hardware. Simulation of transient light transport in synthetic 3D scenes has become a powerful but computationally-intensive alternative for analysis and benchmarking of NLOS imaging methods. NLOS imaging methods also suffer from high computational complexity. In our work, we rely on dimensionality reduction to provide a real-time simulation framework for NLOS imaging performance analysis. We extend steady-state light transport in self-contained 2D worlds to take into account the propagation of time-resolved illumination by reformulating the transient path integral in 2D. We couple it with the recent phasorfield formulation of NLOS imaging to provide an end-to-end simulation and imaging pipeline that incorporates different NLOS imaging camera models. Our pipeline yields real-time NLOS images and progressive refinement of light transport simulations. We allow comprehensive control on a wide set of scene, rendering, and NLOS imaging parameters, providing effective real-time analysis of their impact on reconstruction quality. We illustrate the effectiveness of our pipeline by validating 2D counterparts of existing 3D NLOS imaging experiments, and provide an extensive analysis of imaging performance including a wider set of NLOS imaging conditions, such as filtering, reflectance, and geometric features in NLOS imaging setups.
How viewers interpret different pictorial projections has been a longstanding question affecting many disciplines, including psychology, art, computer science, and vision science. The most-prominent theories assume that viewers interpret pictures according to a single linear perspective projection. Yet, no existing theory accurately describes viewers’ perceptions across the wide variety of projections used throughout art history. Recently, Hertzmann hypothesized that pictorial 3D shape perception is interpreted according to a separate linear perspective for each eye fixation in a picture. We performed four experiments based on this hypothesis. The first two experiments found that viewers consider object depictions as more accurate when an object is projected according to its own local linear projection, rather than one consistent with the rest of the picture. In the third experiment, viewers exhibited change blindness to projections in peripheral vision, suggesting that perception of shape primarily occurs around fixations. The fourth experiment found surface slant compensation to be dependent on fixation. We conclude that pictorial shape perception operates according to per-fixation perspective.
Selection is the first step in many image editing processes, enabling faster and simpler modifications of all pixels sharing a common modality. In this work, we present a method for material selection in images, robust to lighting and reflectance variations, which can be used for downstream editing tasks. We rely on vision transformer (ViT) models and leverage their features for selection, proposing a multi-resolution processing strategy that yields finer and more stable selection results than prior methods. Furthermore, we enable selection at two levels: texture and subtexture, leveraging a new two-level material selection (DuMaS) dataset which includes dense annotations for over 800,000 synthetic images, both on the texture and subtexture levels.
Large diffusion models have made a remarkable leap synthesizing high-quality artistic images from text descriptions. However, these powerful pre-trained models still lack control to guide key material appearance properties, such as gloss. In this work, we present a threefold contribution: (1) we analyze how gloss is perceived across different artistic styles (i.e., oil painting, watercolor, ink pen, charcoal, and soft crayon); (2) we leverage our findings to create a dataset with 1,336,272 stylized images of many different geometries in all five styles, including automatically-computed text descriptions of their appearance (e.g., "A glossy bunny hand painted with an orange soft crayon"); and (3) we train ControlNet to condition Stable Diffusion XL synthesizing novel painterly depictions of new objects, using simple inputs such as edge maps, hand-drawn sketches, or clip arts. Compared to previous approaches, our framework yields more accurate results despite the simplified input, as we show both quantitative and qualitatively.
In their seminal experiment in 1944, Heider and Simmel revealed that humans have a pronounced tendency to impose narrative meaning even in the presence of simple animations of geometric shapes. Despite the shapes having no discernible features or emotions, participants attributed strong social context, meaningful interactions, and even emotions to them. This experiment, run on traditional 2D displays has since had a significant impact on fields ranging from psychology to narrative storytelling. Virtual Reality (VR), on the other hand, offers a significantly new viewing paradigm, a fundamentally different type of experience with the potential to enhance presence, engagement and immersion. In this work, we explore and analyze to what extent the findings of the original experiment by Heider and Simmel carry over into a VR setting. We replicate such experiment in both traditional 2D displays and with a head mounted display (HMD) in VR, and use both subjective (questionnaire-based) and objective (eye-tracking) metrics to record the observers' visual behavior. We perform a thorough analysis of this data, and propose novel metrics for assessing the observers' visual behavior. Our questionnaire-based results suggest that participants who viewed the animation through a VR headset developed stronger emotional connections with the geometric shapes than those who viewed it on a traditional 2D screen. Additionally, the analysis of our eye-tracking data indicates that participants who watched the animation in VR exhibited fewer shifts in gaze, suggesting greater engagement with the action. However, we did not find evidence of differences in how subjects perceived the roles of the shapes, with both groups interpreting the animation's plot at the same level of accuracy. Our findings may have important implications for future psychological research using VR, especially regarding our understanding of social cognition and emotions.
The non-line-of-sight (NLOS) imaging field encompasses both experimental and computational frameworks that focus on imaging elements that are out of the direct line-of-sight, for example, imaging elements that are around a corner.Current NLOS imaging methods offer a compromise between accuracy and reconstruction time as experimental setups have become more reliable, faster, and more accurate.However, all these imaging methods implement different assumptions and light transport models that are only valid under particular circumstances.This paper lays down the foundation for a cohesive theoretical framework which provides insights about the limitations and virtues of existing approaches in a rigorous mathematical manner.In particular, we adopt Dirac notation and concepts borrowed from quantum mechanics to define a set of simple equations that enable: i) the derivation of other NLOS imaging methods from such single equation (we provide examples of the three most used frameworks in NLOS imaging: back-propagation, phasor fields, and f-k migration); ii) the demonstration that the Rayleigh-Sommerfeld diffraction operator is the propagation operator for wave-based imaging methods; and iii) the demonstration that back-propagation and wave-based imaging formulations are equivalent since, as we show, propagation operators are unitary.We expect that our proposed framework will deepen our understanding of the NLOS field and expand its utility in practical cases by providing a cohesive intuition on how to image complex NLOS scenes independently of the underlying reconstruction method.
The atmospheric dynamics of the Amazon, critical for global environmental stability, have faced increasing influence from numerous El Niño and La Niña events in recent decades. While reanalysis data has incorporated these events through models and measurements, the intricate mechanics of spatial and temporal water vapor transport remain unclear. In this study, we present a preliminary analysis of these dynamics, utilizing over twenty years of ERA5 monthly data over atmospheric layer. Our investigation was constructed on two primary scales, each offering unique insights. The first scale aims to replicate and validate the system's seasonality concerning the Intertropical Convergence Zone (ITCZ), those patterns allow evaluating some patterns and its effects in land hydrological process observed along the basin integrating specific methods and models to clarify how the seasonality was replicated and validated. On the second scale, we delve into smaller hydrological sub-units of the Amazon, identifying their contribution to water recycling, and net fluxes across the basin boundaries. We provide an innovative estimation of transport paths using a 4 cardinal directional approach (brubaker box scheme modified) that makes possible identificate, analyse and understanding, water sources and sinks and its relevance in the normal hydrological production and synergically systems The findings indicate the system's relatively stable dynamics in terms of water vapor sources and altitudinal variation across atmospheric layers. Our methodology introduces a novel framework for calculating comprehensive trajectories of water vapor transport from a hybrid lagrangian-eulerian approach, significantly enhancing our understanding of the Amazon's hydrological cycle from an atmospheric perspective. To provide more precision, we specify that the stability observed in the system pertains to water vapor sources and altitudinal variation. These stable dynamics contribute valuable insights into the intricate water vapor transport mechanisms in the Amazon. Additionally, we highlight the implications of our findings for future research in understanding how to create alternatives to mitigate some impacts of El Niño and La Niña events on the Amazon's atmospheric dynamics.
The light field in an underwater environment is characterized by complex multiple scattering interactions and wavelength-dependent attenuation, requiring significant computational resources for the simulation of underwater scenes. We present a novel approach that makes it possible to simulate multi-spectral underwater scenes, in a physically-based manner, in real time. Our key observation is the following: In the vertical direction, the steady decay in irradiance as a function of depth is characterized by the diffuse downwelling attenuation coefficient, which oceanographers routinely measure for different types of waters. We rely on a database of such real-world measurements to obtain an analytical approximation to the Radiative Transfer Equation, allowing for real-time spectral rendering with results comparable to Monte Carlo ground-truth references, in a fraction of the time. We show results simulating underwater appearance for the different optical water types, including volumetric shadows and dynamic, spatially varying lighting near the water surface.
Humans perceive the world by integrating multimodal sensory feedback, including visual and auditory stimuli, which holds true in virtual reality (VR) environments. Proper synchronization of these stimuli is crucial for perceiving a coherent and immersive VR experience. In this work, we focus on the interplay between audio and vision during localization tasks involving natural head-body rotations. We explore the impact of audio-visual offsets and rotation velocities on users' directional localization acuity for various viewing modes. Using psychometric functions, we model perceptual disparities between visual and auditory cues and determine offset detection thresholds. Our findings reveal that target localization accuracy is affected by perceptual audio-visual disparities during head-body rotations, but remains consistent in the absence of stimuli-head relative motion. We then showcase the effectiveness of our approach in predicting and enhancing users' localization accuracy within realistic VR gaming applications. To provide additional support for our findings, we implement a natural VR game wherein we apply a compensatory audio-visual offset derived from our measured psychometric functions. As a result, we demonstrate a substantial improvement of up to 40% in participants' target localization accuracy. We additionally provide guidelines for content creation to ensure coherent and seamless VR experiences
We propose a novel method to reconstruct non-line-of-sight (NLOS) scenes that combines polarization and time-of-flight light transport measurements. Unpolarized NLOS imaging methods reconstruct objects hidden around corners by inverting time-gated indirect light paths measured at a visible relay surface, but fail to reconstruct scene features depending on their position and orientation with respect to such surface. We address this limitation (known as the missing cone problem) by capturing the polarization state of light in time-gated imaging systems at picosecond time resolution, and introducing a novel inversion method that leverages directionality information of polarized measurements to reduce directional ambiguities in the reconstruction. Our method is capable of imaging features of hidden surfaces inside the missing cone space of state-of-the-art NLOS methods, yielding fine reconstruction details even when using a fraction of measured points on the relay surface. We demonstrate the benefits of our method in both simulated and experimental scenarios.
Generative models have enabled intuitive image creation and manipulation using natural language. In particular, diffusion models have recently shown remarkable results for natural image editing. In this work, we propose to apply diffusion techniques to edit textures, a specific class of images that are an essential part of 3D content creation pipelines. We analyze existing editing methods and show that they are not directly applicable to textures, since their common underlying approach, manipulating attention maps, is unsuitable for the texture domain. To address this, we propose a novel approach that instead manipulates CLIP image embeddings to condition the diffusion generation. We define editing directions using simple text prompts (e.g., "aged wood" to "new wood") and map these to CLIP image embedding space using a texture prior, with a sampling-based approach that gives us identity-preserving directions in CLIP space. To further improve identity preservation, we project these directions to a CLIP subspace that minimizes identity variations resulting from entangled texture attributes. Our editing pipeline facilitates the creation of arbitrary sliders using natural language prompts only, with no ground-truth annotated data necessary.