Visual intelligence seeks to perceive, interpret, and synthesize the visual world and is central to modern computer vision. Human-centered visual intelligence is especially demanding because it studies people as expressive, socially situated subjects whose meaning is rarely conveyed by appearance alone. It couples vision with audio and language across four representative tasks: human emotion recognition, human video generation, human voice cloning, and human video matting. Yet existing resources remain task-specific, providing modalities and annotations for individual problems rather than a shared foundation coordinating understanding and generation. This limits multimodal signal use and broader research. We address this gap with HUG-VIS, a unified benchmark for Human-centered Understanding and Generation in Visual Intelligence. It contains 8,400 seated half-body videos of 30 professional actors, each performing the same 280 emotion-action-prompt assignments under a controlled Mandarin studio protocol, with synchronized video, audio, text, and alpha mattes. We evaluate diverse open- and closed-source models across the four tasks under a unified zero-shot protocol using automatic metrics, criterion-specific mean opinion scores, and multiple cross-task analyses. Results show that (i) linguistic content dominates current emotion recognition, while purely visual affect recognition is weakest; (ii) in video generation and voice cloning, automatic metrics and human judgment agree overall but differ in their top rankings, requiring joint reporting; (iii) boundary fidelity under motion is the main remaining obstacle for human matting; and (iv) task difficulty varies across emotions, models, and metrics, with notable cross-task correlations. The dataset and results are available at https://github.com/GML-MMGroup/HUG-VIS.
Scale-free feature is a popularly observed characteristic in various kinds of complex systems. This feature is often determined by verifying that degree distribution follows power-law distribution when using complex network to depict the underlying structure of complex systems under consideration. Previous works have proven that power-law exponent y plays a key role in the analysis of topological structure of networks of this type. In this work, we propose a principled framework by cycle-based renewing operation to construct scale-free networks that allow for an arbitrary positive integer to be power-law exponent candidate. That is to say, exponent y now belongs to N* (positive integers set). Next, we study the relationship between exponent y and other fundamental structural parameters of the proposed scale-free networks, and verify that some pre-existing statements or declarations associated with topological structure of scale-free network can be further improved. For instance, dense scale-free networks can have a larger diameter, and thus have no small-world property. In addition, some new findings are also obtained. For example, given an arbitrary exponent y is an element of N*, there exists scale-free network that turns out to have a Pearson correlation coefficient nearly close to the theoretical upper bound in the limit of large graph size. The results obtained may shed lights on more comprehensive understanding of scale-free networks.
This article addresses high-precision trajectory tracking for fully actuated hexarotor unmanned aerial vehicles (UAVs) subject to model uncertainty and external disturbances. To exploit the repetitive nature of UAV missions, an adaptive iterative learning control (AILC) scheme is developed, which incrementally enhance tracking accuracy using historical data form previous executions. The controller is designed based on the decoupled translational dynamics of a six-rotor tilted-propeller platform and incorporates adaptive laws that jointly estimate and compensate for reference input errors, aerodynamic drag, residual attitude deviations, and lumped disturbances. Rigorous convergence and boundedness are established via a composite energy function (CEF) analysis in the iteration domain. The approach is validated via MATLAB simulations, PX4-based software-in-the-loop (SITL) in Gazebo and real-flight experiments on a laboratory hexarotor platform. To the best of current knowledge, this work: 1) represents the first application of AILC to a fully actuated hexarotor UAV, showing significant improvements in trajectory-tracking accuracy over benchmark methods; and 2) demonstrates the platform’s potential for attitude-stable logistics tasks such as liquid transport through real-flight tests.
In recent years, land-air bimodal robots, which leverage the extended endurance of wheeled robots and the maneuverability of quadrotor robots, have garnered significant attention. This paper addresses the challenges of motion control and autonomous mode-switching in land-air bimodal robots navigating sloping terrains with unknown slope angles and varying surface materials. Firstly, we provide a comprehensive analysis of the robot’s dynamics models in both modes, representing them as a uniform second-order nonlinear state-space formulation incorporating unknown uncertainties. This uniformly represented formulation is beneficial to controller design for two modes and facilitates the deployment of the bimodal controller in a low-cost microcontroller. Then, a terrain-aware controller is proposed for land mode without the accurate dynamics models, and it can overcome the influence of unknown uncertainties brought by unstructured sloping terrains. Subsequently, in order to improve the robot’s endurance, an autonomous mode-switching control strategy is explored for maximizing the use of wheeled land mode to cross sloping terrains. Finally, we conduct the first real-world experiment with a self-developed land-air robot in some sloping terrains with different slope angles and surface materials to validate the accessibility in three-dimensional space of the land-air bimodal robot with the proposed switching control strategy.
Emotion recognition from electroencephalography (EEG) signals remains challenging due to high inter-subject variability, limited labeled data, and the lack of interpretable reasoning in existing approaches. While recent multimodal large language models (MLLMs) have advanced emotion analysis, they have not been adapted to handle the unique spatiotemporal characteristics of neural signals. We present E^2-LLM (EEG-to-Emotion Large Language Model), the first MLLM framework for interpretable emotion analysis from EEG. E^2-LLM integrates a pretrained EEG encoder with Qwen-based LLMs through learnable projection layers, employing a multi-stage training pipeline that encompasses emotion-discriminative pretraining, cross-modal alignment, and instruction tuning with chain-of-thought reasoning. We design a comprehensive evaluation protocol covering basic emotion prediction, multi-task reasoning, and zero-shot scenario understanding. Experiments on the dataset across seven emotion categories demonstrate that E^2-LLM achieves excellent performance on emotion classification, with larger variants showing enhanced reliability and superior zero-shot generalization to complex reasoning scenarios. Our work establishes a new paradigm combining physiological signals with LLM reasoning capabilities, showing that model scaling improves both recognition accuracy and interpretable emotional understanding in affective computing.
Time series forecasting remains a critical challenge across various domains, often complicated by high-dimensional data and long-term dependencies. This paper presents a novel transformer architecture for time series forecasting, incorporating two key innovations: parameter sharing (PS) and Spatial-Temporal Segment Attention (SegAtt). We also define the time series segment as the concatenation of sequence patches from the same positions across different variables. The proposed model, PSformer, reduces the number of training parameters through the parameter sharing mechanism, thereby improving model efficiency and scalability. The introduction of SegAtt could enhance the capability of capturing local spatio-temporal dependencies by computing attention over the segments, and improve global representation by integrating information across segments. The combination of parameter sharing and SegAtt significantly improves the forecasting performance. Extensive experiments on benchmark datasets demonstrate that PSformer outperforms popular baselines and other transformer-based approaches in terms of accuracy and scalability, establishing itself as an accurate and scalable tool for time series forecasting.
This paper presents a multilayer network game model to explore the dynamics of cooperation in structured populations. The model integrates a static lower-layer relationship network, where nodes represent agents and edges reflect the strength of their social ties, with a dynamic upper-layer game network, where interactions are determined probabilistically based on relationship strengths. The model employs a Prisoner’s Dilemma framework to capture strategic decision-making and also introduces two mechanisms — recommendation and vigilance — to simulate trust propagation and reputation effects. Through extensive simulations, we demonstrate that cooperation levels are significantly influenced by network structure, with scale-free networks exhibiting greater resilience to defection compared to regular or small-world networks. Additionally, we analyze the impact of some key parameters, including the temptation to defect, interaction thresholds, and fitness weights, on cooperation outcomes. The results underscore the critical role of network topology and social mechanisms in shaping cooperative behavior, providing valuable insights into the interplay between network structure and strategic interactions.
Gait recognition has achieved remarkable success in constrained environments, yet its performance often degrades significantly in cross-domain and cross-vertical-view scenarios. This is primarily due to the fact that domain-specific silhouette geometry causes models to overfit to extrinsic geometric characteristics rather than learning generalizable motion patterns. To address this issue, we present ViSA-Gait, which utilizes Vision Foundation Models (VFMs) as a “Semantic Compass” for stable universal human body knowledge guidance, enabling a lightweight backbone to robustly capture fine-grained gait dynamics. Specifically, we introduce a token-based distillation mechanism where a Spatial Token Learner (STL) and a Temporal Token Learner (TTL) filter dense VFM features into motion-consistent descriptors. These descriptors are then adaptively injected into a lightweight 3D-CNN backbone via a Gated Cross-Attention (GCA) mechanism, functioning as “Semantic Anchors” that regularize the feature space and guide the model to focus on intrinsic body motion. Extensive experiments demonstrate that ViSA-Gait achieves SOTA performance on cross-domain and cross-vertical-view benchmarks while remaining competitive within-domain, offering a new perspective on bridging the gap between geometry-based analysis and universal semantic understanding.
It is of increasing interest to study various dynamics occurring on network models that reliably display some properties observed in real-world networks. In this work, we first introduce two graph operations and then propose a generative framework for creating an ensemble of new stochastic tree networks T-m,T- p(t) where p is probability parameter and m is a tunable parameter belonging to positive integer set. For given m, all the realizations in stochastic networks T-m,T- p(t) turn out to follow the same power-law degree distribution with exponent larger than 3, and thus have scale-free feature. In addition, fractal feature is found on our networks. Also, we study degree-degree correlation of stochastic networks T-m,T- p(t) by determining assortativity, and show that networks T-m,T- p(t) can have assortative structure by tuning parameters p and m. More importantly, networks 7(m,1)(t) are proved to be always disassortative independent of both the choice of seminal model and value for parameter m. Besides that, we obtain that there exists a structural transition in degree-degree correlation of tree network 7(m),0(t) by properly selecting original model and parameter m. These degree-degree correlation phenomena have not been reported in previous scale-free tree networks with fractal feature. Then, we study random walks on stochastic networks T-m,T- p(t) in detail and obtain analytical solution to mean hitting time in a mathematically rigorous manner. The results mean that two fundamental structural parameters, namely, diameter and fractal dimension, of networks T-m,T- p(t) have a remarkable influence on the quantity. This further provides a guideline for optimizing the topological structure of tree networks. Next, we consider consensus problem on networks T-m,T- p(t) and analytically determine the solution of first-order noise coherence. Similarly, diameter and fractal dimension are proved to play a key role in determination of this parameter. Finally, we conduct extensive experiments, which demonstrates that computer simulations are in perfect agreement with the theoretical analysis.
Graph Neural Networks (GNNs) have shown efficacy in graph node classification, but face computational challenges on large-scale graphs. Although existing graph reduction methods address these issues, they still require high computational resources and fail to prioritize robust performance on out-of-distribution data. To tackle these challenges, we introduce the subgraph invariant learning paradigm, inspired by the small-world phenomenon. This approach enables models trained on specific subgraphs to generalize across diverse subgraphs, reducing computational demands, and enhancing scalability. To promote generalization, we maximize the invariance log-likelihood by deriving a theoretical lower bound of it and formulating the InVar loss. This loss minimizes the discrepancy between node representations and their corresponding invariance representations while maximizing the entropy of the node representation. In response to InVar loss, we propose the Invariance Facilitation Model (IFM), comprising the Invariance Representation Encoder (IRE) and Node Representation Encoder (NRE). IRE, capturing the invariance representations, utilizes Invariance ATTention (InvarATT) to compress long-range dependencies, while NRE learns the node representation, by integrating invariance representations via Telematic ATTention (TeleATT) and exchanging local information within each subgraph through GNNs. Evaluations on four large-scale graph datasets demonstrate the effectiveness, computational efficiency, and interpretability of IFM for large-scale graph node classification.
Talking face video generation with arbitrary speech audio is a significant challenge within the realm of digital human technology. The previous studies have emphasized the significance of audio-lip synchronization and visual quality. Currently, limited attention has been given to the learning of visual uncertainty, which creates several issues in existing systems, including inconsistent visual quality and unreliable performance across different input conditions. To address the problem, we propose a Joint Uncertainty Learning Network (JULNet) for high-quality talking face video generation, which incorporates a representation of uncertainty that is directly related to visual error. Specifically, we first design an uncertainty module to individually predict the error map and uncertainty map after obtaining the generated image. The error map represents the difference between the generated image and the ground truth image, while the uncertainty map is used to predict the probability of incorrect estimates. Furthermore, to match the uncertainty distribution with the error distribution through a KL divergence term, we introduce a histogram technique to approximate the distributions. By jointly optimizing error and uncertainty, the performance and robustness of our model can be enhanced. Extensive experiments demonstrate that our method achieves superior high-fidelity and audio-lip synchronization in talking face video generation compared to previous methods.
Tactile perception is essential for embodied agents to understand physical attributes of objects that cannot be determined through visual inspection alone. While existing approaches have made progress in visual and language modalities for physical understanding, they fail to effectively incorporate tactile information that provides crucial haptic feedback for real-world interaction. In this paper, we present VTV-LLM, the first multi-modal large language model for universal Visuo-Tactile Video (VTV) understanding that bridges the gap between tactile perception and natural language. To address the challenges of cross-sensor and cross-modal integration, we contribute VTV150K, a comprehensive dataset comprising 150,000 video frames from 100 diverse objects captured across three different tactile sensors (GelSight Mini, DIGIT, and Tac3D), annotated with four fundamental tactile attributes (hardness, protrusion, elasticity, and friction). We develop a novel three-stage training paradigm that includes VTV enhancement for robust visuo-tactile representation, VTV-text alignment for cross-modal correspondence, and text prompt finetuning for natural language generation. Our framework enables sophisticated tactile reasoning capabilities including feature assessment, comparative analysis, scenario-based decision making and so on. Experimental evaluations demonstrate that VTV-LLM achieves superior performance in tactile video understanding tasks, establishing a foundation for more intuitive human-machine interaction in tactile domains.
Human emotion synthesis is a crucial aspect of affective computing. It involves using computational methods to mimic and convey human emotions through various modalities, with the goal of enabling more natural and effective human-computer interactions. Recent advancements in generative models, such as Autoencoders, Generative Adversarial Networks, Diffusion Models, Large Language Models, and Sequence-to-Sequence Models, have significantly contributed to the development of this field. However, there is a notable lack of comprehensive reviews in this field. To address this problem, this paper aims to address this gap by providing a thorough and systematic overview of recent advancements in human emotion synthesis based on generative models. Specifically, this review will first present the review methodology, the emotion models involved, the mathematical principles of generative models, and the datasets used. Then, the review covers the application of different generative models to emotion synthesis based on a variety of modalities, including facial images, speech, and text. It also examines mainstream evaluation metrics. Additionally, the review presents some major findings and suggests future research directions, providing a comprehensive understanding of the role of generative technology in the nuanced domain of emotion synthesis.
Embodied Multimodal Large Models (EMLMs) have gained significant attention in recent years due to their potential to bridge the gap between perception, cognition, and action in complex, real-world environments. This comprehensive review explores the development of such models, including Large Language Models (LLMs), Large Vision Models (LVMs), and other models, while also examining other emerging architectures. We discuss the evolution of EMLMs, with a focus on embodied perception, navigation, interaction, and simulation. Furthermore, the review provides a detailed analysis of the datasets used for training and evaluating these models, highlighting the importance of diverse, high-quality data for effective learning. The paper also identifies key challenges faced by EMLMs, including issues of scalability, generalization, and real-time decision-making. Finally, we outline future directions, emphasizing the integration of multimodal sensing, reasoning, and action to advance the development of increasingly autonomous systems. By providing an in-depth analysis of state-of-the-art methods and identifying critical gaps, this paper aims to inspire future advancements in EMLMs and their applications across diverse domains. Project resources are accessible via https://github.com/BurryChen/Embodied-Multimodal-Large-Models.
With the integration of multimodal large language models (MLLMs) into robotic systems and AI applications, embedding emotional intelligence (EI) capabilities is essential for enabling these models to perceive, interpret, and respond to human emotions effectively in real-world scenarios. Existing static, text-based, or text-image benchmarks overlook the multimodal complexities of real interactions and fail to capture the dynamic, context-dependent nature of emotional expressions, rendering them inadequate for evaluating MLLMs' EI capabilities. To address these limitations, we introduce EmoBench-M, a systematic benchmark grounded in established psychological theories, designed to evaluate MLLMs across 13 evaluation scenarios spanning three hierarchical dimensions: foundational emotion recognition (FER), conversational emotion understanding (CEU), and socially complex emotion analysis (SCEA). Evaluation was conducted on 27 state-of-the-art MLLMs, using both objective task-specific metrics and LLM-based evaluation, revealing a substantial performance gap relative to human-level competence. Even the best performing models, Gemini-3.0-Pro and GPT-5.2, achieve the highest scores on EmoBench-M, 70.5 and 66.5 points respectively. Specialized models such as AffectGPT exhibit uneven performance across EmoBench-M, demonstrating strengths in certain scenarios but generally lacking comprehensive emotional intelligence. By providing a comprehensive, multimodal evaluation framework, EmoBench-M captures both the strengths and weaknesses of current MLLMs across diverse emotional contexts. All benchmark resources, including datasets and code, are publicly available at https://emo-gml.github.io/, facilitating further research and advancement in MLLM emotional intelligence.
Precise audio-visual synchronization in speech videos is crucial for content quality and viewer comprehension. Existing methods have made significant strides in addressing this challenge through rule-based approaches and end-to-end learning techniques. However, these methods often rely on limited audio-visual representations and suboptimal learning strategies, potentially constraining their effectiveness in more complex scenarios. To address these limitations, we present UniSync, a novel approach for evaluating audio-visual synchronization using embedding similarities. UniSync offers broad compatibility with various audio representations (e.g., Mel spectrograms, HuBERT) and visual representations (e.g., RGB images, face parsing maps, facial landmarks, 3DMM), effectively handling their significant dimensional differences. We enhance the contrastive learning framework with a margin-based loss component and cross-speaker unsynchronized pairs, improving discriminative capabilities. UniSync outperforms existing methods on standard datasets and demonstrates versatility across diverse audio-visual representations. Its integration into talking face generation frameworks enhances synchronization quality in both natural and AI-generated content.
In this paper, the popularly discussed topic, i.e., how to construct available theoretical networked models that certainly capture some structural features popularly observed on realistic networks, is still our focus. Specifically, we first propose an evolving deterministic network \(N(t)\) using three types of growth ways. Then, we study some topological structural parameters including degree distribution, diameter and clustering coefficient on network \(N(t)\). The results demonstrate that the proposed network has scale-free feature and small-world property. In the meantime, we obtain an interesting finding, i.e., the first handshake between Fibonacci series and the “pure” preferential attachment mechanism. Next, we enumerate spanning trees on network \(N(t)\), and derive the closed-form solution of spanning trees number. Secondly, we introduce randomness into the growth process of network \(N(t)\) to further establish evolving stochastic networks \(\mathfrak{N}(t)\) that follow the same degree distribution as network \(N(t)\), and also determine some topological structural parameters so as to investigate effect of randomness on structural properties. We show analytically that such a randomization approach makes the resulting stochastic networks not only to greatly inherit some fundamental structural properties from deterministic network \(N(t)\), but also to considerably improve the robustness of network when encountering deliberate removal of edge. Lastly, we list out some open problems.
Existing NeRF-based head avatar reconstruction methods utilize expression coefficients as driving signals. Despite significant advancements, they fail to accurately capture facial feature deformations under complex expression changes. To address this issue, we integrate prior information from public facial feature dictionaries with expression coefficients as driving signals, and employ a region attention mechanism to more accurately capture facial deformations. First, when reconstructing a single face, we extract the facial features from a single image and obtain prior information on these features from a public facial feature dictionary. This prior information is integrated with expression coefficients as driving signals to more accurately drive the deformation of facial features. Second, we introduce a region attention mechanism that learns the explicit relationship between local spatial regions and driving signals during training. This allows for differential driving effects on various facial regions with the same driving signal, achieving more precise local motion modeling. Experimental results show that our method can precisely capture subtle deformations across all facial regions and outperforms state-of-the-art methods in both qualitative and quantitative aspects.
In untrimmed video tasks, identifying temporal boundaries in videos is crucial for temporal video grounding. With the emergence of multimodal large language models (MLLMs), recent studies have focused on endowing these models with the capability of temporal perception in untrimmed videos. To address the challenge, in this paper, we introduce a multimodal large language model named MLLM-TA with precise temporal perception to obtain temporal attention. Unlike the traditional MLLMs, answering temporal questions through one or two words related to temporal information, we leverage the text description proficiency of MLLMs to acquire video temporal attention with description. Specifically, we design a dual temporal-aware generative branches aimed at the visual space of the entire video and the textual space of global descriptions, simultaneously generating mutually supervised consistent temporal attention, thereby enhancing the video temporal perception capabilities of MLLMs. Finally, we evaluate our approach on both video grounding task and highlight detection task on three popular benchmarks, including Charades-STA, ActivityNet Captions and QVHighlights. The extensive results show that our MLLM-TA significantly outperforms previous approaches both on zero-shot and supervised setting, achieving state-of-the-art performance.