Generating realistic human-object interactions (HOI) remains challenging due to the complex temporal coordination required between human actions and object responses. Current diffusion-based methods excel at spatial modeling but fail to capture the critical temporal dynamics that govern physically plausible interactions, leading to artifacts such as premature object movement and poor action-response synchronization. To tackle this issue, we introduce V-HOI, a novel velocity-aware framework that explicitly models temporal dynamics through enhanced motion representations and adaptive constraints. A fundamental principle underlying realistic interactions is the coordinated motion dynamics where contact regions and objects exhibit coupled velocity behaviors that vary systematically across interaction phases. Our framework employs a ControlNet-inspired three-branch architecture featuring a dedicated motion branch and dual interaction branches, trained via a two-stage strategy that preserves motion generation capabilities while learning interaction dynamics. Extensive experiments on the FullBodyManipulation dataset demonstrate substantial improvements over existing methods across both motion quality and interaction metrics, with particularly notable advances in contact modeling accuracy. Ablation studies validate the necessity of each component and confirm that our velocity-aware approach successfully generates temporally coordinated and physically plausible interaction sequences.
Spatiotemporal attention learning has always been a challenging research task in video question answering (VideoQA). It needs to consider not only the modelling of local neighbourhood dependencies between the adjacent frames in a video but also the modelling of long-term dependencies between nonadjacent frames. Although the existing methods are usually good at modelling temporal dependencies in one aspect, they cannot simultaneously and effectively model the temporal dependencies between adjacent and nonadjacent frames. To address this issue, we first derive a novel statistic-driven difference-aware generation function, which can efficiently calculate the difference between a sequence feature value and the whole mean value to identify the significance of the feature. Subsequently, we design a novel parameter-free spatiotemporal attention mechanism (PSAM), which captures the most relevant cues scattered in the context of a spatiotemporal video by generating functions and utilizes a gating mechanism to adaptively integrate and filter relevant and irrelevant information. Finally, we use the PSAM and hierarchical modelling to construct a lightweight multiscale context fusion- and reasoning-based VideoQA model. Extensive experimental research results obtained on five benchmark datasets for the VideoQA task show that our VideoQA model has high Q&A performance and lightweight characteristics. Simultaneously, comprehensive ablation experimental results show that the PSAM can not only improve the performance of the model but also significantly reduce the number of model parameters. In addition, extensive experimental findings obtained on the benchmark dataset of joint tasks (video moment retrieval and video highlight detection) further demonstrate that the PSAM is a general and effective spatiotemporal attention mechanism.
Monocular video human mesh recovery faces fundamental challenges in maintaining metric consistency and temporal stability due to inherent depth ambiguities and scale uncertainties. While existing methods rely primarily on RGB features and temporal smoothing, they struggle with depth ordering, scale drift, and occlusion-induced instabilities. We propose a comprehensive depth-guided framework that achieves metric-aware temporal consistency through three synergistic components: A Depth-Guided Multi-Scale Fusion module that adaptively integrates geometric priors with RGB features via confidence-aware gating; A Depth-guided Metric-Aware Pose and Shape (D-MAPS) estimator that leverages depth-calibrated bone statistics for scale-consistent initialization; A Motion-Depth Aligned Refinement (MoDAR) module that enforces temporal coherence through cross-modal attention between motion dynamics and geometric cues. Our method achieves superior results on three challenging benchmarks, demonstrating significant improvements in robustness against heavy occlusion and spatial accuracy while maintaining computational efficiency.
Precise measurement and quantification of semantic relationships between visual and textual data streams is essential for video-based measurement systems, particularly in automated video inspection and analysis applications. Current measurement methodologies face significant challenges in calibrating and aligning semantic distributions across different measurement modalities, resulting in reduced measurement accuracy and system reliability. This article presents a novel measurement approach called the semantic distance-aware cross-modal attention mechanism (SDCAM), which introduces a cross-modal calibration function to quantify semantic disparities between visual and textual features. The proposed system employs the query text as a reference standard and utilizes semantic relative distance as a measurement metric, enabling precise cross-modal calibration and measurement. Building on SDCAM, we develop an advanced video measurement system that achieves multiscale semantic alignment through a three-stage measurement process: collaborative representation measurement (CRM), hierarchical feature aggregation (HFA), and contextual measurement synthesis (CMS). This approach generates comprehensive contextual measurements tailored to specific query parameters, significantly improving measurement accuracy across diverse video analysis tasks. Experiments conducted across four widely used video question answering (VideoQA) datasets demonstrate the superiority of the proposed model over existing state-of-the-art methods, while reducing parameters by 62%. The system's effectiveness is further validated through integration testing with video retrieval and highlight detection tasks, where it improves measurement precision by 2% compared to baseline systems, while maintaining this performance advantage despite a 42% reduction in model parameters.
The widespread adoption of digital technology has ushered in a new era of digital transformation across all aspects of our lives. Online learning, social, and work activities, such as distance education, videoconferencing, interviews, and talks, have led to a dramatic increase in speech-rich video content. In contrast to other video types, such as surveillance footage, which typically contain abundant visual cues, speech-rich videos convey most of their meaningful information through the audio channel. This poses challenges for improving content consumption using existing visual-based video summarization, navigation, and exploration systems. In this paper, we present VisAug, a novel interactive system designed to enhance speech-rich video navigation and engagement by automatically generating informative and expressive visual augmentations based on the speech content of videos. Our findings suggest that this system has the potential to significantly enhance the consumption and engagement of information in an increasingly video-driven digital landscape.
Stylized human motion synthesis remains a fundamental challenge in computer animation and graphics, with a wide spectrum of applications spanning gaming, film production, virtual reality, and beyond. While recent advances in text-driven motion generation have shown promise, existing approaches face critical limitations including the inability to maintain consistent trajectory control, the lack of fine-grained stylization intensity adjustment, and inadequate generalization across diverse motion styles. To address these challenges, We introduce MCSMoG, a novel framework for controllable stylized motion synthesis through multi-conditional guidance. First, a new Multi-Conditional Motion Latent Diffusion (MC-MLD) model is proposed to introduce additional trajectory guidance and achieve trajectory decoupling. Second, we develop a Style and Non-Style Feature Fusion Module that dynamically blends motion features through an adjustable parameter, providing control over stylization intensity. Third, we integrate MotionCLIP as our style encoder, enhancing the model’s generalization capability across diverse and unseen motion styles. Extensive experiments conducted on the combined HumanML3D and 100STYLE datasets demonstrate that our approach outperforms state-of-the-art methods, achieving a 4.6% reduction in FID scores and a 4.1% increase in motion diversity. User studies further confirm the superiority of our method in style fidelity, semantic consistency, and motion naturalness.
Recent advances in LiDAR technology have enabled privacy-preserving gait recognition using 3D point cloud data. However, existing approaches struggle with the inherent challenges of point cloud processing and understanding such as spatial sparsity, irregular sampling, and complex temporal dynamics. In this paper, we present MSPoint-Gait, a novel framework that addresses these challenges through multi-scale analysis and cross-modal learning. At the core of our framework lies a Depth-Aware Attention Module (DAAM) that leverages rich 3D geometric information to generate attention-weighted depth representations, enabling fine-grained feature extraction from point cloud sequences. We further introduce a Multi-Scale Spatio-Temporal (MSST) network that hierarchically captures both local and global gait patterns through adaptive convolution kernels across multiple spatial and temporal scales. These components are unified through a novel cross-modal learning strategy that effectively bridges the semantic gap between raw point clouds and structured depth representations. The proposed frame-work achieves state-of-the-art performance on the challenging SUSTech1K dataset, with 91.9% Rank-1 and 98.0% Rank-5 accuracy, demonstrating significant improvements over existing methods across various walking conditions and viewpoints.
Robust gait recognition requires highly discriminative representations, which are closely tied to input modalities. While binary silhouettes and skeletons have dominated recent literature, these 2D representations fall short of capturing sufficient cues that can be exploited to handle viewpoint variations, and capture finer and meaningful details of gait. In this paper, we introduce a novel framework, termed DepthGait, that incorporates RGB-derived depth maps and silhouettes for enhanced gait recognition. Specifically, apart from the 2D silhouette representation of the human body, the proposed pipeline explicitly estimates depth maps from a given RGB image sequence and uses them as a new modality to capture discriminative features inherent in human locomotion. In addition, a novel multi-scale and cross-level fusion scheme has also been developed to bridge the modality gap between depth maps and silhouettes. Extensive experiments on standard benchmarks demonstrate that the proposed DepthGait achieves state-of-the-art performance compared to peer methods and attains an impressive mean rank-1 accuracy on the challenging datasets.
Cross-video question answering presents significant challenges beyond traditional single-video understanding, particularly in establishing meaningful connections across video streams and managing the complexity of multi-source information retrieval. We introduce VideoForest, a novel framework that addresses these challenges through person-anchored hierarchical reasoning, enabling effective cross-video understanding without requiring end-to-end training. VideoForest integrates three key innovations: 1) a human-anchored feature extraction mechanism that employs ReID and tracking algorithms to establish robust spatiotemporal relationships across multiple video sources; 2) a multi-granularity spanning tree structure that hierarchically organizes visual content around person-level trajectories; and 3) a multi-agent reasoning framework that efficiently traverses this hierarchical structure to answer complex queries. To evaluate our method, we develop CrossVideoQA, a comprehensive benchmark specifically designed for person-centric cross-video analysis. Experimental results demonstrate VideoForest's superior performance in cross-video reasoning tasks, achieving 71.93% accuracy in person recognition, 83.75% in behavior analysis, and 51.67% in summarization and reasoning.
Text-driven human motion generation is gaining momentum lately thanks to its great potential in shaping the new pathway of interactive computer graphics in the era of AI. Despite the enormous efforts made so far, existing methods still struggle to ensure fluidity and body coordination when generating motions, which seriously hinders its application in a wide spectrum of areas such as gaming, animation, and the emerging metaverse. One of the many causes is, that learning directly from motion data is prone to interference from noise within the data, resulting in reduced quality of the generated motions. In this study, we for the first time propose to promote text-to-motion generation via out-of-distribution detection in the embedding space. Leveraging the Z-score-based outlier detection algorithm, we apply masking to motion data within the motion encoder and replace target data with means, ensuring the consistency of data distribution. To verify the effectiveness of the proposed method, we have conducted extensive experiments on the widely used KIT-ML dataset. Experimental results indicate that compared to previous frameworks, our solution significantly improves the quality of text-driven human motion generation.
The heat and moisture transfer simulation about human body and clothing is a method to simulate the regulation mechanism of human body, and heat and moisture performance of clothing in thermal environment by using computer technology. According to the simulation results, it can be used to the monitor and predicate the key physiological data of human body, and the functional clothing design. Some mathematical models about this area have been built. However, a large number of differential equations are involved to be used to the numerical computation in the simulation process, which leads to high computational complexity, large memo-ry consumption and long computing time. With the expansion of computing scale, these problems will become more obvious and directly affect the practical application of simulation models. The main work of this paper is to study the efficient solution method based on neural network, and proposed a deep learning-based differential equation computing network. The network uses differential equation as the constraint of the loss function of neural network, and integrates the differential equation into the training and prediction process of the network. The experimental results show that the proposed network can reduce the computational complexity of the simulation model and improve the computational efficiency.
Text imageability is often used to quantize the ease with which a natural language description can invoke a mental image in a reader. With the proliferation of artificial intelligence powered text-toimage generation models, it will likely play an even more significant role in bridging the gap between language and visual representation. Unfortunately, automatically suggesting proper imageable textural prompts from a piece of plain text has scarcely been systematically investigated. In this paper, we narrow the gap by introducing a novel framework for text imageability assessment to automatically predict whether a piece of plain text and a prompt is highly imageable to be fed into a textto-image model for faithful image generation. We have also developed a new visual-text dataset, named Ted1.6k, to facilitate model training and validation. Experiment results demonstrate the effectiveness of the proposed method in promoting prompt-guided image generation.
Video question answering is a challenging task that requires models to recognize visual information in videos and perform spatio-temporal reasoning. Current models increasingly focus on enabling objects spatio-temporal reasoning via graph neural networks. However, the existing graph network-based models still have deficiencies when constructing the spatio-temporal relationship between objects: (1) The lack of consideration of the spatio-temporal constraints between objects when defining the adjacency relationship; (2) The semantic correlation between objects is not fully considered when generating edge weights. These make the model lack representation of spatio-temporal interaction between objects, which directly affects the ability of object relation reasoning. To solve the above problems, this paper designs a heuristic semantics-constrained spatio-temporal heterogeneous graph, employing a semantic consistency-aware strategy to construct the spatio-temporal interaction between objects. The spatio-temporal relationship between objects is constrained by the object co-occurrence relationship and the object consistency. The plot summaries and object locations are used as heuristic semantic priors to constrain the weights of spatial and temporal edges. The spatio-temporal heterogeneity graph more accurately restores the spatio-temporal relationship between objects and strengthens the model's object spatio-temporal reasoning ability. Based on the spatio-temporal heterogeneous graph, this paper proposes Heuristic Semantics-constrained Spatio-temporal Heterogeneous Graph for VideoQA (HSSHG), which achieves stateof-the-art performance on benchmark MSVD-QA and FrameQA datasets, and demonstrates competitive results on benchmark MSRVTT-QA and ActivityNet-QA dataset. Extensive ablation experiments verify the effectiveness of each component in the network and the rationality of hyperparameter settings, and qualitative analysis verifies the object-level spatio-temporal reasoning ability of HSSHG
Groundwater is a vital resource in the Chuoshui River alluvial plain (CSAP), a key agricultural area in Taiwan. Understanding groundwater recharge is crucial for sustainable water management amidst changing climatic conditions and increasing water demand. This study investigates the major ion composition, solute Sr concentrations, and 87Sr/86Sr ratios in groundwater and stream water from the Choushui River (CSR) to trace groundwater recharge sources. The Piper diagram reveals that most groundwater samples are of the freshwater Ca–HCO3 type, aligning with the total dissolved solids (TDS) classification. TDS and major ion compositions indicate that groundwater near Baguashan Terrace (BGT) and Douliu Hill (DLH) primarily derives from stream water and rainwater. Na+ and Cl− enrichment in some aquifers of BGT and DLH is attributed to the dissolution of paleo-sea salt and mixing with paleo-seawater from sedimentary porewater. Elevated dissolved Sr concentrations and lower 87Sr/86Sr ratios in these aquifers further support the intrusion of paleo-seawater. Groundwater in the proximal fan shows high TDS due to intensive weathering, complicating the use of TDS as a tracer. Sr isotopic compositions and solute Sr2+ concentrations effectively distinguish recharge sources, revealing that the CSR mainstream primarily recharges the proximal fan and BGT region, while CSR tributaries and rainwater mainly recharge the DLH region. This study concludes that Sr isotopic compositions and solute Sr2+ concentrations are more reliable than TDS and major ion compositions in identifying groundwater recharge sources, enhancing our understanding of groundwater origins and the processes affecting water quality.
Joint video moment retrieval and highlight detection is a video understanding task that requires the model to construct multimodal interaction between heterogeneous features. Recent Transformer-based models mainly focus on promoting global interaction between features. However, local interaction and temporal asynchronism modeling are not deeply considered. To solve this problem, this paper proposes a dual-branch complementary multimodal interaction mechanism (DCMI), which consists of a global difference feature activation module (GDFA) and a local information dynamic aggregation module (LIDA). GDFA measures the difference between the target element and the global features, thus activating important information. LIDA designs a multimodal heterogeneous graph and constructs asynchronous interaction between heterogeneous features to dynamically aggregate local information. DCMI adaptively fuses the complementary dual branches to improve the model's cognitive and decision-making abilities of global and local information. Comprehensive comparisons with existing methods on public datasets verify the superiority of the proposed model. Extensive ablation experiments and qualitative analysis show the effectiveness and rationality of DCMI, which can promote the interaction between multimodal features.
Data augmentation is serving as a critical and fundamental technology to improve model generalization and performance in a wide spectrum of machine learning tasks. Despite the increasing interest in developing various pathways to artificially generate new data to reduce the overfitting issue during model training, enriching the diversity of training data in the field of medicine remains facing enormous challenges. By virtue of recent advancements in generative artificial intelligence, we present a novel data augmentation framework, CLIP-MedFake, to address the shortage of training data used in medical image classification. The proposed method first employs the Stable Diffusion model to generate new fake data based on a small amount of training data, and then adopts the paradigm of few-shot learning and uses the CLIP architecture as the backbone to pre-train the model with synthetic data and then fine-tune it with real medical images. Extensive experiment results on two publicly available datasets demonstrate the effectiveness of the proposed method in promoting medical image classification.
Multi-modal attention learning in video question answering (VideoQA) is a challenging task, as it requires consideration of information recognition within modalities and information interaction and fusion between modalities. Existing methods employs the cross-attention mechanism to compute feature similarity between modalities, thereby aggregating relevant information in a shared space. However, heterogeneous features have different distributions in the shared space, making it difficult to directly match semantics, which may affect similarity calculation. To address this issue, a novel enhanced cross-modal attention mechanism (ECAM) is proposed in this paper that pre-fuses two modalities to generate an enhanced key with feature importance distributions to effectively solve the semantic mismatch. Compared with the existing cross-attention mechanism, ECAM can realize the semantic matching between multiple modalities more accurately and pay more attention to the relevant feature regions. In the multi-modal fusion phase, a two-stage fusion strategy is proposed to exploit the advantages of the two fusion methods to deeply explore the complex and diverse dependency relationships between the multi-modal features. Collectively supported by these two newly designed modules, we proposed the VideoQA solution based on two-stage deep exploration of temporally-evolving features with enhanced cross-modal attention mechanism which is able to conquer challenging semantic understanding and question answering tasks. Extensive experiments on four VideoQA datasets show that the new approach attains superior results in comparison with state-of-the-art peer methods. Moreover, experiments on the latest joint task datasets prove that ECAM is a general mechanism that can be easily adapted to solve other visual-linguistic tasks.
Outfit collocation requires considering the interrelationship and adaptability among the attributes of component items. However, with the numerous and diverse attributes of fashion items, accurately capturing attribute features and modeling the complex relationships between attributes become the key challenges. To address these challenges, we propose a novel scheme Decoupling-driven Multi-level Attribute Parsing for interpretable outfit collocation. First, we decouple a series of attribute features from the item's visual feature by fully supervised, which can improve the robustness of the model in processing both relevant and irrelevant attributes of items. Furthermore, employing a deep deconvolution neural network with attention mechanisms to reconstruct the decoupled attribute features into a visual image that is close to the original item image. It ensures all attribute features can be combined to contain complete item information. Next, graph attention networks are constructed to parse multi-level attribute compatibility relationships from three perspectives: intra-attribute, inter-attribute, and item integration relationships. Finally, we use multi-layer perceptrons to fuse the score distributions of the three and output the outfit compatibility score. Experiments conducted on the IQON3000 dataset demonstrate that our model outperforms existing state-of-the-art methods and exhibits good interpretability.
Recent advancement in learning and teaching methodology experimented with virtual reality (VR)-based presentation form to create immersive learning and training environment. The quality of such educational VR applications not only relies on the virtual model, but the 2D presentation materials such as text, diagrams and figures. However, manual designing or seeking these educational resources is both labor intensive and time-consuming. In this paper, we introduce a new automatic algorithm to detect and extract presentation slides in educational videos, which will provide abundant resources for creating slide-based immersive presentation environment. The proposed approach mainly involves five core components: shot boundary detection, training instances collection, shot classification, slide region detection and slide transition detection. We conducted comparison experiment to evaluate the performance of the proposed method. The results indicate that, in comparison with peer method, the proposed method improves the precision of slide detection from 81.6 to 92.6% and recall from 74.7 to 86.3% on average. With the detected slides, content analyzer can be employed to further extract reusable elements, which can be used for developing VR-based educational applications.
The joint task of video moment retrieval and video highlight detection is a challenging study, which requires building a model that not only captures contextual information between sequences in time but also has the ability to understand and judge significance. This paper solves these problems from three aspects. Firstly, we design a parameter-free cross-modal statistical correlation interaction method. A novel saliency enhancement function is defined to quantify the saliency differences between the important features associated with the query and other features to achieve parameter-free cross-modal fusion. Secondly, we propose a novel modality-aware heterogeneous graph reasoning mechanism (MHGR). MHGR can effectively capture the global context information between sequences, enhance the local association relationship between sequences, and deal with the complexity of multi-modal data better through the organic combination of two key modules: parameter-free cross-modal statistical correlation interaction, and heterogeneous graph reasoning mechanism. Thirdly, a lightweight solution for the joint task of video moment retrieval and highlight detection is designed based on the above two novel algorithm modules. Comprehensive experiments are conducted on publicly available benchmark data to validate the advantages of the new solution in comparison with a series of state-of-the-art peer methods. Quantitative results consistently demonstrate that the new solution is lightweight and has high inference performance so the remarkable improvement in accuracy achieved by the new solution with respect to peer methods. An extended ablation study is further conducted to show the usefulness of each module of the solution in acquiring its computational capabilities.