Generating dynamic 4D objects from sparse inputs is difficult because it demands joint preservation of appearance and motion coherence across views and time while suppressing artifacts and temporal drift. We hypothesize that the view discrepancy arises from supervision limited to pixel- or latent-space video-diffusion losses, which lack explicitly temporally aware, feature-level tracking guidance.We present \emph{Track4DGen}, a two-stage framework that couples a multi-view video diffusion model with a foundation point tracker and a hybrid 4D Gaussian Splatting (4D-GS) reconstructor. The central idea is to explicitly inject tracker-derived motion priors into intermediate feature representations for both multi-view video generation and 4D-GS. In Stage One, we enforce dense, feature-level point correspondences inside the diffusion generator, producing temporally consistent features that curb appearance drift and enhance cross-view coherence. In Stage Two, we reconstruct a dynamic 4D-GS using a hybrid motion encoding that concatenates co-located diffusion features (carrying Stage-One tracking priors) with Hex-plane features, and augment them with 4D Spherical Harmonics for higher-fidelity dynamics modeling.\emph{Track4DGen} surpasses baselines on both multi-view video generation and 4D generation benchmarks, yielding temporally stable, text-editable 4D assets. Lastly, we curate \emph{Sketchfab28}, a high-quality dataset for benchmarking object-centric 4D generation and fostering future research.
Multimodal in-context learning (ICL) is becoming a key capability that allows large vision-language models (LVLMs) to adapt to novel tasks without parameter updates, which expands their usefulness in many real-world applications. However, ICL performance remains unstable even when the in-context demonstrations (ICDs) are well matched, showing that LVLMs still struggle to make full use of the provided context. While existing work mainly focuses on prompt engineering or post-hoc logit calibration, we study the attention mechanisms inside LVLMs to address their inherent limitations. We identify two important weaknesses in their self-attention that hinder effective ICL. To address these weaknesses, we propose Context-Aware Modulated Attention (CAMA), a training-free and plug-and-play method that dynamically adjusts attention logits based on the input in-context sequence. CAMA uses a two-stage modulation process that strengthens attention to semantically important tokens, especially visual ones. Across four LVLMs and seven benchmarks, CAMA consistently outperforms vanilla models and baselines, showing clear effectiveness and generalization. It can also activate the intended benefits of prompt engineering methods and remains robust across different sequence configurations. Therefore, CAMA opens up new directions for improving multimodal reasoning through a deeper understanding of attention dynamics.
Forestry inventory, the task of systematically measuring each tree within a given area, constitutes the basis for sustainable forest management and addresses a significant need in modern digital forestry. To automate the inventory process, existing methods mainly rely on LiDAR systems to scan forest plots. However, these LiDAR scanning sensors are often expensive and not readily accessible. To address these limitations, we propose VidFin, a novel video-based forest inventory system architecture that simultaneously generates DBH (Diameter at Breast Height) estimation, tree tracking, and tree positions from a consumer-level, monocular drone video. Within the system, each component is tightly interconnected; for example, DBH measurements depend on accurate distance estimation, which is jointly trained with UAV trajectory calculation, while tree positions are optimized using both UAV trajectory data and distance measurements. This temporal association in the video data and the interconnection between components results in improved inventory performance across diverse forestry scenes. To validate the system's performance, we developed a benchmark encompassing various forestry scenes on 13 flight paths. The experimental results demonstrate the superiority of our approach, achieving an average measurement error of 2.26 cm in DBH and a recall of 98.3%.
Interacting with small objects in virtual reality (VR) can be challenging due to the physical limitations of controllers and headsets, which often lead to unintended collisions and tracking loss when devices come too close, thereby disrupting the user's immersive experience. While researchers have developed techniques like translational gain, hand remapping, and specialized interaction to address these challenges, these approaches are often task-specific or insufficient for precise, detailed interactions or observations. To address these challenges, in this paper, we introduce a novel interaction technique called dynamic translational gains manipulation (DTGM), which adjusts scaling in real-time based on the user's proximity to objects. We conducted a user study to evaluate the effectiveness of improving precision during object manipulation and understanding the subjective mental workload of the proposed DTGM technique. Our results revealed that the DTGM technique improved interaction efficiency, making it suitable for various VR applications where precision and space optimization are crucial.
Knowledge graphs encode semantic relationships among entities and are an emerging technique in text analysis. Yet their complexity hampers the analysis of relationships in graphs with thousands of nodes. We design BlossomNet, a conceptual visual analytics system for community, temporal, and influence analysis of large knowledge graphs. BlossomNet integrates three coordinated views: (1) a Cluster-based Overview that reveals community structure, (2) a Theme River View that traces thematic change and migration, and (3) an Influence View that uses floral glyphs and arcs to depict influence and growth. BlossomNet contributes a design framework for visual exploration of complex, person-centric social and semantic networks, with applications in domains where lineage, trend formation, and influence pathways are central.
In recent years, two-dimensional (2D) superconducting materials have garnered significant interest due to their unique properties and potential applications. Here, we conducted thermodynamic and dynamic stability studies on 51 metal-intercalated hexagonal boron carbon (h-BC) compounds, and ultimately identified 22 stable compounds. Among these 22 compounds, 18 materials are metals, while the remaining 4 materials include 1 semiconductor ( MgB_2C_2 ) and 3 semimetals ( TiB_2C_2 , ZrB_2C_2 , and HfB_2C_2 ). The possible superconductivity of eighteen metals is studied by solving the Allen–Dynes modified McMillan equation to estimate their superconducting transition temperature ( T_c ). The highest T_c is observed in KB_2C_2 ( T_c = 53.47 K), followed by NaB_2C_2 ( T_c = 48.30 K), while the lowest T_c is in AlB_2C_2 ( T_c = 0.04 K). Due to the high T_c of alkali metal intercalation compounds, this work mainly focuses on them. For alkali metal intercalation compounds, we found that the T_c rises with the increase of the main group atomic number, mainly due to the degree of metalization of the σ -bonding band at the Fermi level. Another important reason is the softening of the phonon spectrum. These findings enrich the family of 2D superconductors, providing new theoretical insights for experimental synthesis and opening research ideas for 2D superconducting electronic devices.
Potential risk stock assessments of heavy metals are essential for informing management strategies and mitigating environmental and human health risks. This process enables cost-effective decision-making, supports regulatory compliance, and promotes sustainable land and ecosystem management; however, such studies remain limited. To investigate chromium (Cr) distribution and estimate its potential risk stock under varying soil layers, soil types, parent materials, and land-use conditions, a large dataset of surface soils, deep soils, and soil columns from a karst region in southwestern China was analyzed. The concentration of Cr in surface soils, deep soils, and soil columns was 23.3 - 1126.9, not detected - 315.6, and 36.9 - 684.0 mg/kg, respectively. The consistency in the Cr distribution across the surface and deep layers, geological structures, and mine locations indicated a main geological origin for Cr rather than anthropogenic pollution. This was further supported by principal component analysis, which identified ore stockpiling and parent-rock weathering as the primary sources. Variance analysis showed significant influences of soil type, land-use type, parent rock, landform, soil pH, and SOC on spatial distributions. Cr exhibited exponential variations with depth. Using an exponential fitting model with multiple integrations, the Cr potential risk stock was estimated at 1.78 × 105 tons with middle and deep layers accounting for the highest, underscoring the complexity and challenges of managing Cr risks. Given the limited soil resources in karst regions, it is crucial for local authorities to prioritize Cr risk management and implement targeted land-use strategies.
Recently, the concept of higher-order topological insulators has aroused widespread attention and research interest. However, current studies have predominantly focused on the domain of acoustic waves. Compared to acoustic waves, elastic waves are vector waves, making their study more complex and challenging. Therefore, achieving higher-order topological states in elastic waves holds significant research value. In this paper, we proposed the design of an intelligent topological metamaterial, which is composed of magneto-rheological thin layers and an elastic substrate. First, by adjusting the topological structure, we successfully excited first-order topological states of Lamb waves in numerical simulations. Subsequently, we constructed a two-dimensional topological structure to excite zero-order topological corner states. Given the unique advantages of magnetic fields in regulating material properties and behaviors, we investigated the effects of magnetic fields as an external control mechanism on Lamb waves in magneto-rheological materials. Our analysis focused on the regulation of Lamb wave topological edge states and corner states via magnetic fields. The results demonstrate that by varying the magnetic field strength, we can precisely control the characteristics of the topological states. Magnetic field modulation of the topological states in Lamb waves enables the realization of non-contact, controllable phononic devices, which is of great significance for the development of topological acoustics.
Most existing Dynamic Gaussian Splatting methods for complex dynamic urban scenarios rely on accurate object-level supervision from expensive manual labeling, limiting their scalability in real-world applications. In this paper, we introduce SplatFlow, a Self-Supervised Dynamic Gaussian Splatting within Neural Motion Flow Fields (NMFF) to learn 4D space-time representations without requiring tracked 3D bounding boxes, enabling accurate dynamic scene reconstruction and novel view RGB/depth/flow synthesis. SplatFlow designs a unified framework to seamlessly integrate time-dependent 4D Gaussian representation within NMFF, where NMFF is a set of implicit functions to model temporal motions of both LiDAR points and Gaussians as continuous motion flow fields. Leveraging NMFF, SplatFlow effectively decomposes static background and dynamic objects, representing them with 3D and 4D Gaussian primitives, respectively. NMFF also models the correspondences of each 4D Gaussian across time, which aggregates temporal features to enhance cross-view consistency of dynamic components. SplatFlow further improves dynamic object identification by distilling features from 2D foundation models into 4D space-time representation. Comprehensive evaluations conducted on the Waymo and KITTI Datasets validate SplatFlow's state-of-the-art (SOTA) performance for both image reconstruction and novel view synthesis in dynamic urban scenarios.
We present Melody Way, a visual analytics system for exploring influence, collaboration, and genre evolution in the music industry. The system integrates temporal and network perspectives in a coordinated multi-view environment, enabling analysts to move fluidly between high-level trend analysis and detailed relationship exploration. Interactive filters, search, and an AI chatbot support targeted investigation and predictive queries, with results visually verified in context. Applied to the 2025 VAST Challenge MC-1 dataset, Melody Way revealed evolving genre dynamics, distinctive artist trajectories, and potential future leaders in Oceanus Folk, demonstrating how integrated visual and AI-assisted analysis can illuminate complex patterns in large cultural datasets.
With the development of high-performance computers, cloud storage, and advanced sensors, people’s ability to gather complex learning data has greatly improved. However, analyzing these data remains a significant challenge. Especially for spatiotemporal learning data such as eye-tracking and mouse movement, understanding and analyzing these data to identify the learning insights behind them is a difficult task. We propose a visualization platform called “MultiScaleAnalyzer”, which employs hierarchical structure to illustrate spatiotemporal learning data in multiple views. From high-level overviews to detailed analyses, “MultiScaleAnalyzer” provides varying resolutions of data tailored to educators’ need. To demonstrate the platform’s effectiveness, we applied “MultiScaleAnalyzer” to a mathematical word problem-solving dataset, showcasing how the visualization platform facilitates the exploration of student problem-solving patterns and strategies.
The usability of virtual reality (VR) training applications is crucial for their success, but examining the usability in the early development stages remains challenging. A realistic and plausible solution would be revisiting and reconciling Heuristics Evaluation (HE) methods among the most widely used usability inspection methods in the human-computer interaction (HCI) domain. While research on studying and using HE methods is growing within the VR domain, few studies have considered the novel VR environment challenges new requirements for fitting HE methods to the context and applying them effectively. To this end, we conducted a user study with 14 evaluators using the standard HE methods to complete two HE sessions for a VR training application. We identified five critical challenges that evaluators encountered in the HE process by observing and interviewing them. Based on our findings, we discuss the importance of considering an easy-to-use heuristic set, how we can facilitate the HE procedures in the VR context, and the opportunities for developing HE-supporting tools.
OBJECTIVE:Our objective is to develop and validate TrajVis, an interactive tool that assists clinicians in using artificial intelligence (AI) models to leverage patients' longitudinal electronic medical records (EMRs) for personalized precision management of chronic disease progression. MATERIALS AND METHODS:We first perform requirement analysis with clinicians and data scientists to determine the visual analytics tasks of the TrajVis system as well as its design and functionalities. A graph AI model for chronic kidney disease (CKD) trajectory inference named DisEase PrOgression Trajectory (DEPOT) is used for system development and demonstration. TrajVis is implemented as a full-stack web application with synthetic EMR data derived from the Atrium Health Wake Forest Baptist Translational Data Warehouse and the Indiana Network for Patient Care research database. A case study with a nephrologist and a user experience survey of clinicians and data scientists are conducted to evaluate the TrajVis system. RESULTS:The TrajVis clinical information system is composed of 4 panels: the Patient View for demographic and clinical information, the Trajectory View to visualize the DEPOT-derived CKD trajectories in latent space, the Clinical Indicator View to elucidate longitudinal patterns of clinical features and interpret DEPOT predictions, and the Analysis View to demonstrate personal CKD progression trajectories. System evaluations suggest that TrajVis supports clinicians in summarizing clinical data, identifying individualized risk predictors, and visualizing patient disease progression trajectories, overcoming the barriers of AI implementation in healthcare. DISCUSSION:The TrajVis system provides a novel visualization solution which is complimentary to other risk estimators such as the Kidney Failure Risk Equations. CONCLUSION:TrajVis bridges the gap between the fast-growing AI/ML modeling and the clinical use of such models for personalized and precision management of chronic diseases.
In this paper, we present a novel indoor 3D reconstruction method with occluded surface completion, given a sequence of depth readings. Prior state-of-the-art (SOTA) methods only focus on the reconstruction of the visible areas in a scene, neglecting the invisible areas due to the occlusions, e.g., the contact surface between furniture, occluded wall and floor. Our method tackles the task of completing the occluded scene surfaces, resulting in a complete 3D scene mesh. The core idea of our method is learning 3D geometry prior from various complete scenes to infer the occluded geometry of an unseen scene from solely depth measurements. We design a coarse-fine hierarchical octree representation coupled with a dual-decoder architecture, i.e., Geo-decoder and 3D Inpainter, which jointly reconstructs the complete 3D scene geometry. The Geo-decoder with detailed representation at fine levels is optimized online for each scene to reconstruct visible surfaces. The 3D Inpainter with abstract representation at coarse levels is trained offline using various scenes to complete occluded surfaces. As a result, while the Geo-decoder is specialized for an individual scene, the 3D Inpainter can be generally applied across different scenes. We evaluate the proposed method on the 3D Completed Room Scene (3D-CRS) and iTHOR datasets, significantly outperforming the SOTA methods by a gain of 16.8% and 24.2% in terms of the completeness of 3D reconstruction. 3D-CRS dataset including a complete 3D mesh of each scene is provided on project webpage(1).
Forest monitoring and education are key to forest protection, education and management, which is an effective way to measure the progress of a country's forest and climate commitments. Due to the lack of a large-scale wild forest monitoring benchmark, the common practice is to train the model on a common outdoor benchmark (e.g., KITTI) and evaluate it on real forest datasets (e.g., CanaTree100). However, there is a large domain gap in this setting, which makes the evaluation and deployment difficult. In this paper, we propose a new photorealistic virtual forest dataset and a multimodal transformer-based algorithm for tree detection and instance segmentation. To the best of our knowledge, it is the first time that a multimodal detection and segmentation algorithm is applied to a large-scale forest scenes. We believe that the proposed dataset and method will inspire the simulation, computer vision, education and forestry communities towards a more comprehensive multi-modal understanding.
The popularity of LiDAR devices and sensor technology has gradually empowered users from autonomous driving to forest monitoring, and research on 3D LiDAR has made remarkable progress over the years. Unlike 2D images, whose focused area is visible and rich in texture information, understanding the point distribution can help companies and researchers find better ways to develop point-based 3D applications. In this work, we contribute an unreal-based LiDAR simulation tool and a 3D simulation dataset named LiDAR-Forest, which can be used by various studies to evaluate forest reconstruction, tree DBH estimation, and point cloud compression for easy visualization. The simulation is customizable in tree species, LiDAR types and scene generation, with low cost and high efficiency.
Probabilistic-based non-linear dimensionality reduction (PB-NL-DR) methods, such as t-SNE and UMAP, are effective in unfolding complex high-dimensional manifolds, allowing users to explore and understand the structural patterns of data. However, due to the trade-off between global and local structure preservation and the randomness during computation, these methods may introduce false neighborhood relationships, known as distortion errors and misleading visualizations. To address this issue, we first conduct a detailed survey to illustrate the design space of prior layout enrichment visualizations for interpreting DR results, and then propose a node-link visualization technique, ManiGraph. This technique rethinks the neighborhood fidelity between the high- and low-dimensional spaces by constructing dynamic mesoscopic structure graphs and measuring region-adapted trustworthiness. ManiGraph also addresses the overplotting issue in scatterplot visualization for large-scale datasets and supports examining in unsupervised scenarios. We demonstrate the effectiveness of ManiGraph in different analytical cases, including generic machine learning using 3D toy data illustrations and fashion-MNIST, a computational biology study using a single-cell RNA sequencing dataset, and a deep learning-enabled colorectal cancer study with histopathology-MNIST.
The increasing advancement of human-machine interaction (HMI) technology has brought the modes of vehicle HMI into focus, as they are closely related to driver and passenger safety and directly affect the travel experiences. This study compared the usability and safety of three vehicle HMI modes: hardware interaction (HI), hardware and software interaction (HSI), and software interaction (SI). The evaluation comprised two dimensions: usability and safety. Sixty participants' performance on these tasks was evaluated at two driving speeds (30 km/h and 60 km/h). The results of the nonparametric tests indicated significant differences between the three interaction modes: (1) HI was the highest safety-oriented interaction mode with participants had the highest average vehicle speed and maximum acceleration measured at 60 km/h and the lowest glance frequency at both speeds; (2) HSI was the most usable interaction mode. Participants had the shortest task-completion time measured at 60 km/h and the highest score on the NASA-TLX and SUS scales taken for both speeds; (3) SI was the lowest secure and usable in-vehicle interaction mode. Participants had the longest task-completion time at 60 km/h, the highest error frequency under 30 and 60 km/h and the highest glance frequency, the longest total glance duration and the longest average glance time. In conclusion, HI and HSI were more secure and usable invehicle interaction modes than SI. From a theoretical exploration perspective, this paper elaborates on some exploratory thoughts and innovative ideas for practical application to the screen HMI mode selection and design in intelligent vehicle cabins.