
Several uncertainties emerge in each component of the visual analytics (VA) cycle that hinder the user’s ability to make efficient and effective decisions, e.g., through missing data, model approximations, or visual mappings. Well-known VA strategies aim first at making users aware of these uncertainties, often through visual means. When visuals alone are insufficient to accurately quantify or communicate uncertainties, VA designers may rely on guidance to support users’ understanding of these uncertainties throughout the VA cycle. While prior VA research has attempted to conceptualize guidance, the ability and mechanisms for guidance to comprehensively address different sources of uncertainties remain an open question. In this survey, we characterize the relationships between uncertainties and guidance in VA literature. Our key contribution is a taxonomic framework that relates uncertainty sources to relevant guidance strategies and their respective profiles, i.e., roles, scopes, and features. Through this taxonomy, we discuss how guidance addresses uncertainties, identify research gaps, and promote a more comprehensive understanding of guidance strategies to support effective design for uncertainty in VA. Our survey underscores the effectiveness of guidance in navigating uncertainties with context-aware and multi-scope strategies. We highlight challenging opportunities for further research in this space, especially in the adjacent areas of accessibility and onboarding, and suggest new research areas, such as narrative and persuasion guidance to support uncertainty awareness in VA.
Power diagrams are fundamental weighted partitioning structures in computational geometry and computer graphics. Existing studies have extended Voronoi diagrams from Euclidean domains to networks by replacing straight-line distances with shortest-path metrics. However, network counterparts of power diagrams that preserve connected power cells and satisfy maximum-capacity constraints remain insufficiently explored. This paper presents maximum-capacity constrained network power diagrams (MCCNPD) for embedded network partitioning. We first define a network power diagram using a squared shortest-path power distance on an embedded network graph. To generate connected network power cells, we construct them through a rooted connectivity-preserving growth process. Based on this, maximum-capacity constraints are imposed on network power cells, and a two-stage optimization strategy is developed to compute capacity-feasible partitions. The first stage performs weight-driven global deformation of network power cells. The second stage applies frozen-weight local refinement to resolve residual capacity violations while preserving cell connectivity. The formulation is further extended from discrete demands to continuous embedded network demand densities via line-integral capacity accounting. Experiments on representative synthetic networks and a real urban road network show that the proposed method produces connected network power partitions and effectively reduces capacity violations under both discrete and continuous demand models.
The flexibility of label placement directly affects the number and quality of labels placed around the graphic. Many domains benefit from label placement on the tight outline around the labeled graphic. However, such label placement suffers from “label crowding” in certain regions, preventing the placement of the required number of labels. We present a method capable of assessing conflicts among label candidates and decreasing the “label crowding” by placing certain labels at less desirable positions, but outside of the crowded area. Consequently, this results in a larger number of placed labels while also aiming to improve label spacing in less dense regions.The proposed method operates in real time and produces results comparable to those of longer-running optimization methods.
Progressive compression of triangle mesh geometry typically exploits spatial coherence to reduce data size while preserving surface detail. In applications where lossy compression is permissible, an effective strategy is to align distortion with the limitations of human visual perception. Existing methods rely on perceptual metrics to steer refinement, but doing so typically incurs additional data overhead to specify where each refinement occurs. We propose a progressive geometry compression algorithm that leverages a perceptually informed model of normal uncertainty to predict where distortion is most likely to be noticeable. This enables the encoder to focus refinements in those regions without explicitly transmitting their locations at each step, thereby reducing overhead.While our initial approach produced reconstructions ranked higher by established perceptual metrics compared to a baseline of Edgebreaker with weighted parallelogram prediction, its high computational cost limited practical deployment. In this extended work, we address this bottleneck by introducing an Adaptive Quasi-Monte Carlo sampling method utilizing a Halton sequence and an early-stopping heuristic. This strategy substantially accelerates the compression pipeline and improves the overall perceptual quality. Furthermore, we explore a neural estimator with the potential to further speed up the normal uncertainty evaluation process. However, while this approach shows potential, further research is required to achieve sufficient accuracy and reliability for practical integration.
Road-surface-centred novel-view synthesis remains challenging under vehicle-mounted acquisition. Near-linear forward motion and grazing-angle observations leave road geometry weakly constrained. Sparse or repetitive texture, reflections, and far-range perspective compression further degrade image correspondence, stereo depth, and photometric optimisation. We formulate this setting as a metric-consistency problem and introduce a three-phase stereo–LiDAR neural rendering framework that keeps metric information explicit from pose estimation to ray-level supervision. The first stage builds a SuperPoint–SuperGlue stereo-temporal graph that combines temporal tracking edges with fixed-baseline stereo edges to improve SfM coverage under forward road motion. The second stage converts RAFT-Stereo disparity into a sensor-supported dense Z-depth prior through online LiDAR-based scale estimation and log-domain residual propagation. The final stage trains Instant-NGP with anchor-distance confidence-weighted structural losses, so completed-depth pixels close to direct LiDAR support impose stronger ray-level constraints than weakly supported regions. On RSRD (Zhao et al., 2023), the pose front end produces full-sequence reconstructions for all 15 evaluated sequences. Under a near-view held-out-frame protocol over six scenes, the confidence-aware renderer achieves 30.25 dB mean PSNR and 0.7883 mean SSIM, and reduces rendered-depth MAE at held-out LiDAR pixels from 1.81 m with RGB-only training to 0.48 m. The depth prior reduces held-out LiDAR sample AbsRel from 0.0141 for RAFT-Stereo with only per-frame metric conversion to 0.0088 after log-domain calibration under a 10% anchor/90% held-out protocol. This held-out LiDAR evaluation measures metric consistency within the same stereo–LiDAR system rather than independent dense-depth accuracy. Together, the results show improved near-view image quality and rendered-depth consistency under road-centric stereo–LiDAR acquisition.
Neural signed distance functions for sparse point cloud reconstruction exhibit strong directional asymmetry, with far greater sensitivity to normal perturbations than tangential ones. Existing distributionally robust methods (e.g., SDRO) ignore this asymmetry, using isotropic neighborhoods that assign the same perturbation scale to the first-order sensitive normal direction and the tangent plane, which can overemphasize normal excursions in the robust aggregation. We propose Geometry-Consistent Robust Learning (GCRL). Geometry-Induced Anisotropic SDRO (GIA-SDRO) replaces isotropic neighborhoods with curvature-adaptive anisotropic ones, compressing normal perturbations while preserving tangential exploration. Structure-Aware Hard Query Mining (SHQM) concentrates optimization on thin structures and high-curvature regions via a saliency field combining curvature and field inconsistency. Together, these components reduce surface drift and topological artifacts in failure-prone regions. On ShapeNet, FAUST, and SRB, GCRL outperforms recent baselines under sparse and noisy inputs.
Advancements in computational power have increasingly enabled the application of deep learning to enhance fluid simulation pipelines. In this work, we introduce a multi-task deep learning framework designed to simultaneously address boundary particle detection and normal vector estimation in SPH fluids. Our primary contribution is the Sparse Voxel Fluid Convolutional Neural Network (Sparse VFCNN), an architecture that leverages spatially sparse 3D convolutions to overcome the computational bottlenecks associated with traditional volumetric networks. Developed as an optimized successor to an initial dense architecture (Dense VFCNN), the sparse variant achieves competitive performance with a significant reduction in processing time. Our results demonstrate that the Sparse VFCNN approximates the precision of state-of-the-art methods while providing unified outputs — a feature essential for consistent geometric representation in fluid simulations. Furthermore, the framework is agnostic to the ground-truth data source, allowing for seamless integration with various detection and estimation methods. By combining high efficiency with multi-task learning, the Sparse VFCNN offers a robust and versatile alternative for enhancing geometric fidelity in large-scale fluid simulation frameworks.
The reconstruction of High Dynamic Range (HDR) images from a single Low Dynamic Range (LDR) image is a challenging problem in computer vision. State-of-the-art methods still have limitations in recovering regions degraded by severe clipping and noise in overexposed and underexposed areas. To overcome these limitations, this paper proposes SAFHDR, a U-Net-based architecture composed of DAMU-Net, which incorporates deformable convolutions guided by the Optimized Attention Block and efficient multiscale processing via the Lightweight Residual Features Block, and the Information Aggregation Module, dedicated to preserving the global structure of the scene. Additionally, a two-stage training strategy is adopted to improve perceptual fidelity and suppress chromatic and structural artifacts. Experiments conducted on the NTIRE 2021 dataset demonstrate that the proposed method outperforms state-of-the-art methods on HDR-oriented perceptual metrics, with high detail recovery capability in critical regions.
With the increasing demand for high-fidelity interactions in virtual reality and digital human systems, the development of accurate and generalizable grasp synthesis has emerged as a critical challenge. Existing methods typically exhibit limited generalization across different mesh hand models and suffer from latency. We propose FitGrab, a novel grasp synthesis framework designed to support diverse skinned mesh hands through a unified representation based on biological-angle encoding. To ensure precise hand–object interaction, FitGrab employs a sensor-guided geometric adaptation method that captures spatial context through three complementary sensor modalities: joint sensor, inter-joint gap sensor, and ambient sensor, each implemented using Basis Point Set (BPS) distance encoding. FitGrab leverages conditional autoencoders to infer biological angles from the current sensor states, enabling real-time synthesis without long temporal contexts. The proposed framework is versatile, accommodating both single-hand and dual-hand grasp scenarios, and is compatible with diverse skinned mesh hand models. Experiments across multiple evaluation settings show that FitGrab produces precise and adaptable grasp poses for unseen objects while maintaining competitive frame-wise quality, robust cross-hand-model transfer, and interactive runtime. Code and resources will be released to support future research.
Skinning decomposition-based facial auto rigging aims to construct an actor-specific Linear Blend Skinning (LBS) facial rig from an actor’s example expressions and a predefined rigging template. However, facial rigs recovered by existing skinning decomposition methods may accurately reproduce the target expressions yet yield visible local distortions when the solved joint transformations are interpolated from the neutral configuration to each target expression. We attribute this issue to two factors. First, target-state fitting formulations often admit multiple near-equivalent joint-rotation solutions that reproduce the target expressions similarly. Second, single-step closed-form rotation updates are path-independent and therefore provide little stable preference among these candidates. To address this issue, we propose a rotation-stabilized skinning decomposition method for generating actor-specific LBS facial rigs that produce more stable intermediate deformations under interpolation. Specifically, we improve rotation solving through explicit rotation-space regularization and the implicit bias induced by path-dependent continuous optimization with identity initialization. Furthermore, without changing the alternating optimization framework between joint transformations and skinning weights, we reformulate the optimization under a unified differentiable objective. This objective incorporates the surface normal consistency constraint and allows it to directly supervise joint transformation updates through a shared differentiable optimization process. Experiments show that the proposed method produces a substantially more contracted rotation-angle distribution and more stable intermediate deformations, with a trade-off in target-expression vertex reconstruction error. In addition, the unified differentiable formulation reduces surface normal error.
We present ConGAT, CONtext-aware Graph ATtention Network, a visual analysis system for identifying and analyzing cell–cell interactions in large tissue volumes, which leverages graph attention networks and self-supervised learning. Advanced multichannel microscopy techniques, such as Cyclic Immunofluorescence (CycIF), enable detailed, biomarker-based single-cell mapping, but yield large, densely packed volumes that are difficult to explore. ConGAT implements a machine learning-assisted visual analysis approach for these data. First, ConGAT constructs 3D spatial graphs from volumetric image data and applies a 3D graph attention network to model cell–cell relationships directly in 3D space. Next, we introduce a self-supervised feature refinement strategy that learns discriminative region-of-interest representations without labeled data. Through an interactive visual interface, ConGAT allows experts to define arbitrary combinations of biomarkers, enabling flexible ROIs identification and iterative exploration of spatial patterns. We assess ConGAT on synthetic ground truth and on real-world data. An evaluation with domain expert validation demonstrates that ConGAT identifies spatial patterns and candidate ROIs that show substantial agreement with expert assessments of biological relevance.
3D Gaussian Splatting (3DGS) enables high-quality novel view synthesis with real-time rendering, but it remains fragile when only a few input views are available. Under sparse-view supervision, standard Gaussian primitives often form unstable combinations for coarse geometry and overfit high-frequency details, producing view-dependent artifacts in novel views. We revisit this problem through a wavelet-domain co-adaptation analysis and observe that low-frequency components in test views are especially sensitive to perturbations among overlapping Gaussian primitives. This suggests that sparse-view 3DGS should treat coarse geometry and high-frequency details differently rather than applying priors or regularization uniformly. We propose Frequency-Aware Student Splatting (FASS), a sparse view synthesis method that combines a more adaptive primitive representation with frequency-aware regularization strategies. Specifically, our FASS replaces Gaussians with Student’s t primitives to reduce dependence on fragile primitive combinations, while discarding the scooping operation to avoid sparse-view holes and flickering artifacts. It further introduces the Low-Frequency Depth Regularization (LFDR) strategy, which applies monocular depth supervision only to reliable low-frequency depth components, and the High-Frequency Weighted Dropout (HFWD) strategy, which assigns larger dropout probabilities to primitive clusters with stronger high-frequency responses. Experiments on LLFF and DTU under 2-view, 3-view, and 4-view settings show that our FASS improves synthesis quality over existing 3DGS-based methods while preserving real-time rendering.
We present a reliable, fully combinatorial method for UV mapping that leverages a Voronoi-based decomposition of a triangulated surface mesh. Given a sparse set of sample points on the input shape, we construct the corresponding Voronoi partition and iteratively refine it to ensure that all regions are topologically equivalent to disks. From the disk decomposition, we propose two directions. On one side, adjacent regions are merged following a greedy approach that minimizes some objective function, all while preserving disk equivalence. Alternatively, the regions can be used to guide a cut that makes the whole surface topologically equivalent to a disk. In both cases, this topological guarantee enables straightforward and reliable UV parametrization. Our method exhibits an extremely low failure rate, making it suitable for practical use. In quantitative experiments on standard UV mapping benchmarks, we achieve performance comparable to state-of-the-art techniques. Furthermore, we analyze robustness and efficiency across different sampling densities, providing insights into the computational cost of each step of the pipeline.
The geometric reconstruction of biological membranes from electron microscopy data, such as three-dimensional images acquired by cryo-electron tomography, is of great importance for the analysis of cellular environments. Recent advances in deep learning-based approaches enable the segmentation of biological membranes in those data with unprecedented quality and completeness, resulting in highly complex morphological structures. However, in order to fully understand the geometric properties of membranes and how they relate to other cellular structures, including membrane proteins, an explicit geometric representation in the form of triangle meshes is necessary. Here, we present a ridge surface-based approach that is able to transform a wide variety of membrane morphologies, given as segmented voxel representations of membranes, into high-quality triangle meshes. We compare our approach with MidSurfer, a recently published method, which not only fails for those complex morphologies but is also at least one magnitude slower than the presented approach. In addition, we test a crease surface extraction algorithm and a medial surface skeleton extraction method on the data presented, and discuss the pros and cons of each approach.
As virtual reality (VR) systems gain acceptance into critical domains, such as healthcare, education, banking, and military, sensitive user data has to be protected from malicious users. Existing security measures such as usernames/passwords, PINs, or multi-factor approaches do not provide any protection when the attacker gains access to the credentials or when the user intentionally hands over their credentials to an ally, for instance to take an exam on their behalf. Recognizing these challenges, a large body of work has emerged over the past decade on behavioral biometrics for single VR systems. However, cross-system behavioral biometrics for VR remains at a nascent stage. Cross-system behavioral biometrics is necessary to enable users to seamlessly transition between their office, home, job site, or clinic issued system. Early work in cross-system behavioral biometrics for VR showed that while deep learning techniques such as Siamese neural networks were effective, they required near complete trajectories for high assurance user authentication or identification. The emergence of motion forecasting approaches for VR biometrics enabled the use of limited data from the start of the user’s action, thereby preventing the attacker from gaining access to the full user trajectory. However, such forecasting models were designed for a single VR system that did not enable cross-system authentication. In recent work, we showed that cross-system forecasting-based authentication for VR biometrics can be performed using an Informer-based model to train the forecasting component and a fully convolutional network to train the authenticator. Using a publicly available dataset of 41 users performing a ball throwing task using the Meta Quest, HTC Vive, and HTC Vive Cosmos, we showed that in comparison to non-forecasted Siamese networks, our approach reduces the equal error rate (EER) by an average of 53.16% across all VR system combinations over prior cross-system authentication work. In this paper, we compare the performance of the Informer-based cross-system forecasting model, which operates on a point-wise input and treats each timestamp as a separate token, against a patching architecture, namely PatchTST, that that provides lower runtime by splitting the input time series into individual channels, operates on each channel independently, and encodes channel features into patches. Using the same ball throwing dataset, we show that PatchTST shows an average EER reduction of 30.30% over prior cross-system authentication. Using PatchTST, we obtain a speedup of more than 3 times compared to Informer when using a GPU with half the core count. Link to code: https://github.com/Terascale-All-sensing-Research-Studio/Cross_System_Authentication_using_Forecasting/ .