Referring Audio-Visual Segmentation (Ref-AVS) aims to segment specific objects in videos based on natural language expressions involving audio, vision, and text information. This task poses significant challenges in cross-modal reasoning and fine-grained object localization. In this paper, we propose a simple framework, SimToken, that integrates a multimodal large language model (MLLM) with the Segment Anything Model (SAM). The MLLM is guided to generate a special semantic token representing the referred object. This compact token, enriched with contextual information from all modalities, acts as a prompt to guide SAM to segment objectsacross video frames. To further improve semantic learning, we introduce a novel target-consistent semantic alignment loss that aligns token embeddings from different expressions but referring to the same object. Experiments on the Ref-AVS benchmark demonstrate that our approach achieves superior performance compared to existing methods.Code will be available at https://github.com/DianJin-HFUT/SimToken
Action recognition aims to identify an action from video frames. The action/background information is diverse in different frames, which hinders learning the implicit action patterns. In this work, we propose a Mask-aware Kernel Model (MKM), which ensures implicit action pattern learning by integrating kernel learning with proper cluster relations. The MKM provides novel cluster-aware kernels to enhance the action representation for frame patches. The MKM introduces a kernel clustering learner, kernel masking filter, and a kernel attention selector. First, to learn temporal features, the temporal Vision Transformer uses temporal correlation to ensure the action features for kernel learning. Second, to analyze the action kernels for frame patches, we design a kernel clustering learner module. This module learns cluster relations with patch-wise convolutions to describe the common action among patches. The cluster relations are learned in each frame, which ensures cluster-aware kernel learning with input frame adaptivity. Third, to analyze the action kernels with spatial adaptivity, we design a kernel masking filter module. This module introduces a location mask by analyzing the region patterns with spatial convolution. The patch-level mask ensures the kernel learning with region-aware selection. Fourth, after learning multiple channel features by convolution with multiple kernels, we design a kernel attention selector module. This module excites kernel-aware features by learning channel-wise attention with channel-wise convolutions, which ensures the kernel learning with channel-wise selection for effective action representation. Extensive experiments demonstrate that our method achieves state-of-the-art performance on Something-Something V1 & V2, Kinetics-400, UAV-human, and Diving 48 datasets.
Micro-Actions (MAs) are subtle and spontaneous human behaviors that provide important non-verbal cues in social interaction and affective communication. However, their short duration, weak motion patterns, and fine-grained semantic differences make them difficult to annotate, model, and evaluate in a standardized manner. To promote academic research on micro-action analysis, we proposed and have annually organized the Micro-Action Analysis Grand Challenge (MAC) as a public benchmark platform for this emerging field. The first two editions of MAC established standardized evaluation settings for micro-action recognition and detection, providing publicly accessible datasets and protocols. Building upon these editions, this paper presents the 3rd MAC, held in conjunction with ACM Multimedia 2026. Under the theme of moving from recognition to fine-grained micro-action understanding, this edition further expands the scope of the challenge beyond conventional recognition and detection. In particular, we introduce a new task named fine-grained micro-action understanding, evaluated with the assistance of multimodal large language models, aiming to assess models' ability to capture fine-grained semantic cues and interpret subtle human micro-actions at a deeper level. We summarize the datasets, task settings, evaluation protocols, competition results, and representative solutions from top-performing teams. Finally, we discuss future directions for micro-action analysis and its broader role in human-centric video understanding.
Mental health assessment is crucial for early intervention and effective treatment, yet traditional clinician-based approaches are limited by the shortage of qualified professionals. Recent advances in artificial intelligence have sparked growing interest in automated psychological assessment, yet most existing approaches are constrained by their reliance on static text analysis, limiting their ability to capture deeper and more informative insights that emerge through dynamic interaction and iterative questioning. Therefore, in this paper, we propose a multi-agent framework for mental health evaluation that simulates clinical doctor-patient dialogues, with specialized agents assigned to questioning, adequacy evaluation, scoring, and updating. In detail, we introduce an adaptive questioning mechanism in which an evaluation agent assesses the adequacy of user responses to determine the necessity of generating targeted follow-up queries to address ambiguity and missing information. Additionally, we employ a tree-structured memory in which the root node encodes the user's basic information, while child nodes (e.g., topic and statement) organize key information according to distinct symptom categories and interaction turns. This memory is dynamically updated throughout the interaction to reduce redundant questioning and enhance the information extraction and contextual tracking capabilities. Experimental results on the DAIC-WOZ dataset illustrate the effectiveness of our proposed method, which achieves better performance than existing approaches. Our code is released at \url{https://github.com/MindIntLab-HFUT/AgentMental}.
Multimodal recommendation (MMRec) typically integrates the multimodal content of items with user-item interactions to capture user preferences. Existing models predominantly encapsulate intricate multimodal content within an integrated, albeit implicit, embedding space, relying on a single set of latent representations to assimilate and process information across heterogeneous modalities. However, this representation learning approach oversimplifies the distinct and shared semantics of different modalities, leading to the entanglement and confounding of diverse preference cues. Indeed, different modalities characterize distinct and complementary aspects of items (e.g., images reveal the overall silhouette; text conveys usage), while also capturing shared and overlapping semantics (e.g., both modalities reflect style and color). Disentangling these multi-aspect features (i.e., distinct and overlapping ones) and learning fine-grained representations is crucial for complete capture and coherent expression of user preferences. To address this, we develop a novel model, Cross-modal Feature DisenTangling via Bidirectional Distillation(CFDTBD), which explicitly disentangles multimodal preference representations into differential feature learning and common feature learning across modalities. The former captures modality-specific, complementary signals, while the latter extracts shared, consistent semantics. A bidirectional distillation strategy is formulated to ensure multi-aspect feature learning by pushing the embedding spaces of differential features further apart while bringing common features closer together. This disentangled framework allows CFDTBD to extract fine-grained cross-modal preference cues from otherwise entangled and coarse item embeddings, based on the relative contributions of distinct and overlapping aspects in multimodal content. The effectiveness of CFDTBD is validated through extensive experiments on three public e-commerce datasets, including Beauty and Arts_crafts_and_Sewing from Amazon (May 1996 - Oct 2018), and Taobao from Alibaba (released in 2018 for the Tianchi competition). The source code is available at https://github.com/hfutmars/CFDTBD.
Micro-gesture recognition (MGR) has recently emerged as an important research direction in affective computing and human-computer interaction, aiming to decode subtle and unconscious bodily movements that reflect hidden emotions. Unlike illustrative gestures, which are intentional, expressive, and long in duration, micro-gestures are subtle, spontaneous, and short-lived, making their recognition far more challenging. MGR has made remarkable progress with the emergence of several public datasets. However, existing reviews mostly focus on conventional gesture or facial micro-expression analysis, leaving MGR as a distinct field that is insufficiently summarized. In this paper, we present the first comprehensive survey of the MGR method. It covers several key aspects: 1) datasets of two diverse modalities and their collection protocols; 2) recognition methods across supervised, unsupervised, contrastive, multimodal fusion, and multimodal large language model (MLLM) paradigms; and 3) challenges such as long-tail distribution, cross-dataset generalization, and bridging recognition with emotion understanding. This survey aims to provide both an overview and future perspectives to advance the development of micro-gesture recognition. Our project is available at Github: https://github.com/timwang2001/Awesome_Micro_Gesture .
In this paper, we present XInsight Lab's solution to the micro-gesture classification track of the 4th MiGA Challenge at IJCAI 2026, in which our solution ranked first and achieved a new state-of-the-art result. We propose a multimodal ensemble framework that integrates a self-supervised RGB-based model with supervised multi-stream models from previous solutions. The self-supervised RGB model is pretrained on 120K unlabeled clips via masked video modeling and then fine-tuned on iMiGUE. This simple yet effective RGB baseline achieves 69.224
Remote photoplethysmography (rPPG) enables contactless physiological monitoring by capturing subtle skin-color variations from facial videos. However, most existing methods predominantly rely on time-domain modeling, making them vulnerable to motion artifacts and illumination fluctuations, where weak physiological clues are easily overwhelmed by noise. To address these challenges, we propose FreqPhys, a frequency-guided rPPG framework that explicitly leverages physiological frequency priors for robust signal recovery. Specifically, FreqPhys first applies a Physiological Bandpass Filtering module to suppress out-of-band interference, and then performs Physiological Spectrum Modulation together with adaptive spectral selection to emphasize pulse-related frequency components while suppress residual in-band noise. A Cross-domain Representation Learning module further fuses these spectral priors with deep time-domain features to capture informative spatial–temporal dependencies. Finally, a frequency-aware conditional diffusion process progressively reconstructs high-fidelity rPPG signals. Extensive experiments on six benchmarks demonstrate that FreqPhys yields significant improvements over state-of-the-art approaches, particularly under challenging motion conditions. It highlights the importance of explicitly modeling physiological frequency priors. The source code will be released.
Molecular dynamics (MD) simulation is a key method for studying complex molecular systems, and deep learning has been widely used to accelerate it. Recently, diffusion models have shown excellent performance in generating high-quality molecular conformations and modeling complex distributions. However, their discrete noise-adding process limits the ability to capture smooth oscillatory behavior between consecutive frames and also poses challenges in maintaining spatial structural consistency and effectively processing molecular graph features. To address this, we propose a two-module approach: a molecular graph interaction module, enhanced with classical potential functions, and a diffusion module that uses the Discrete Cosine Transform (DCT) to better capture smooth molecular motions. These improvements enable our model to achieve strong performance across experiments.
Electroencephalography (EEG) offers a noninvasive approach for examining neurophysiological correlates of dimensional psychopathology, yet systematic evidence across EEG paradigms and feature granularities remains limited. Here, we develop a granularity-aware EEG feature pipeline that organizes multi-scale descriptors into global, regional, and channel levels. Using the Healthy Brain Network (HBN) cohort, we evaluate the prediction of four psychopathology dimensions: p-factor, internalizing, externalizing, and attention problems, across four EEG paradigms. Given the heterogeneity of pediatric psychopathology and the moderate reliability of questionnaire-derived scores, this setting represents a challenging feasibility test rather than a clinical screening scenario. Tree-based models and granularity-balanced feature selection showed promising improvements over conventional approaches in selected conditions, although effect sizes remained modest. Visualization of selected markers revealed dimension-specific spatial and spectral patterns that were broadly aligned with existing neurophysiological knowledge. An exploratory cross-dataset sanity check on the independent PEARL cohort suggested that the proposed selection principle remains technically feasible under protocol shifts, without claiming cross-dataset generalizability. Overall, multi-scale EEG features contain weak but detectable signals related to dimensional psychopathology, and granularity-aware selection may serve as a useful feature-reduction strategy for future EEG-based phenotyping studies.
Deception detection is a critical yet challenging task in forensic analysis, security, and social interaction. The complexity of deceptive behaviors has motivated growing interest in multimodal deception detection (MMDD), which integrates diverse signals to improve reliability. This survey provides a comprehensive overview of recent advances in MMDD, covering research background, benchmark datasets, evaluation metrics, feature fusion methods, and deception detection architectures that have evolved from traditional machine learning to deep learning. We conclude with a discussion of existing challenges, future research directions, and the ethical issues involved. An open-source Github repository ( https://github.com/open-code-and-source/awesome-MMDD ) is maintained alongside this survey, offering curated datasets and an awesome list of related works for MMDD.
The growing adoption of robotics and augmented reality in real-world applications has driven considerable research interest in 3D object detection based on point clouds. While previous methods address unified training across multiple datasets, they fail to model geometric relationships in sparse point cloud scenes and ignore the feature distribution in significant areas, which ultimately restricts their performance. To deal with this issue, a unified 3D indoor detection framework, called UniGeo, is proposed. To model geometric relations in scenes, we first propose a geometry-aware learning module that establishes a learnable mapping from spatial relationships to feature weights, which enabes explicit geometric feature enhancement. Then, to further enhance point cloud feature representation, we propose a dynamic channel gating mechanism that leverages learnable channel-wise weighting. This mechanism adaptively optimizes features generated by the sparse 3D U-Net network, significantly enhancing key geometric information. Extensive experiments on six different indoor scene datasets clearly validate the superior performance of our method.
Web-based platforms are becoming a primary channel for psychological support, yet most LLM-driven chatbots remain opaque, single-stage, and weakly grounded in established therapeutic practice, limiting their usefulness for web applications that promote digital well-being. To address this gap, we present XInsight, a counseling-inspired multi-agent framework that models psychological support as a stage-consistent workflow aligned with the classical Exploration-Insight-Action paradigm. Building on structured client representations, XInsight orchestrates specialized agents under a unified Reason-Intervene-Reflect cycle: an Exploration agent organizes background and concerns into a structured Case Conceptualization Form, a Routing agent performs Adaptive Therapeutic Routing (ATR) across SFBT, CBT, and MBCT, a unified Therapeutic agent executes school-consistent submodules, and a Consolidation agent guides review, skill integration, and relapse-prevention planning. A Recording agent continuously transforms open-ended web dialogues into standardized psychological artifacts, including case formulations, therapeutic records, and relapse-prevention plans, enhancing interpretability, continuity, and accountability. To support rigorous and transparent assessment, we introduce XInsight-Bench with a Scale-Guided LLM Evaluation (SGLE) protocol that combines therapy-specific clinical scales with general counseling criteria. Experiments show improved paradigm alignment, multi-therapy integration, interaction depth, and interpretability over existing multi-agent counseling systems, indicating that XInsight provides a practical blueprint for integrating counseling-inspired support agents into web applications for digital well-being.
Despite the high accuracy of EEG-based emotion recognition, existing models remain opaque "black boxes", lacking semantic grounding between abstract neural features and human-interpretable states. In this paper, we reframe EEG explainability as a cross-modal generation task, shifting the paradigm from feature attribution to behavioral visualization. We introduce Facial Emoji Proxy Modeling, a novel framework that translates high-dimensional EEG signals into identity-agnostic facial emojis. Guided by the neuroscientific prior of neural-facial consistency, this approach grounds neural representations in the manifold of observable facial dynamics. Technically, our framework integrates FMENet, a specialized backbone modeling expression-relevant spatial synergies, and the Facial Emoji Learning Branch (FELB), which treats emoji reconstruction as a structured semantic regularizer. Extensive experiments on EAV and MMER benchmarks demonstrate that our method achieves state-of-the-art accuracy among EEG-only models. Crucially, it generates semantically faithful facial animations that provide a transparent, privacy-preserving window into the brain's emotional evolution, effectively allowing users to "see the emotion" directly from neural signals.
Remote photoplethysmography (rPPG) has emerged as a promising contactless technology for physiological monitoring, yet its widespread adoption is hindered by the prohibitive cost of acquiring large-scale labeled data. To alleviate this limitation, we propose a semi-supervised remote physiological measurement framework that leverages both labeled and unlabeled facial videos via frequency-guided pseudo-labeling, named FrePL. Specifically, we introduce a pseudo-labeling mechanism to select high-confidence rPPG pseudo-labels via spectral energy concentration. To further mitigate confirmation bias and reduce label noise in the early stages, we adopt a dynamic thresholding strategy with a progressively decaying confidence threshold, which retains more pseudo-labels as the model becomes more robust. A subsequent peak frequency filtering pseudo-label refinement enhances pseudo-label quality by preserving physiological frequency components while suppressing noise. Additionally, we adopt a physiological consistency contrastive learning objective that exploits the intrinsic physiological coherence across augmented views of the same sample, while enforcing discriminative representations between different individuals. Extensive experiments on four benchmark datasets demonstrate that our FrePL not only achieves performance comparable to fully-supervised methods (4.57 mean absolute error (MAE) on VIPL-HR with 20% labeled data), but also significantly outperforms existing semi-supervised approaches.
Audio-Visual Question Answering (AVQA) requires reasoning over temporally evolving audio and visual signals to answer natural-language questions about dynamic scenes. Most existing methods assume that both modalities are available during training and testing. In practice, however, an audio or visual stream may be unavailable because of passive signal loss, such as hardware or transmission failures, or intentional removal, such as withholding visual information for privacy. We formulate Training-time Modality-Missing AVQA (TM-AVQA), a setting in which modality-complete, audio-missing, and visual-missing samples may occur during both training and testing. To address this new setting, we develop an AVQA-specific two-stage framework that adapts established cross-modal reconstruction and dependency-modeling principles to the supervision constraints of TM-AVQA. In Stage-I, a reconstruction network infers task-oriented feature-level representations of the missing modality from the available modality and question text. Its Dense Temporal-scale Reconstruction (DTR) module aggregates complementary information across multiple temporal granularities, while its Multimodal Dependency Modeling (MDM) module captures dependencies among the audio, visual, and question modalities. To supervise reconstruction when the target-modality features are unavailable, we adapt two complementary learning objectives to TM-AVQA: Cross-Modal Relation-based Contrastive Learning (CMR-CL), which exploits within-video and cross-video multimodal relations, and Cross-Sample Relation-based Pseudo-label Learning (CSR-PL), which constructs feature-level pseudo targets from relevant modality-complete samples. In Stage-II, the frozen reconstruction network is combined with existing AVQA backbones for answer prediction. We construct controlled TM-AVQA variants of MUSIC-AVQA, MUSIC-AVQA-R, and AVQA datasets by deleting one modality from selected samples. Experiments across multiple AVQA backbones, missing-modality conditions, and missing rates show consistent and robust improvements over the evaluated AVQA and missing-modality baselines. Additional ablations and analyses examine the contribution of the proposed task-specific designs, reconstruction quality, efficiency, and transfer to audio-visual recognition task.
Gaze target detection aims to localize a person’s gaze target. During gaze transition in video, the absence of accurate temporal variation modeling (TVM) may lead to errors in gaze target localization. In this work, we propose a Transition-aware Gaze Model (TGM), which focuses on analyzing temporal differences to achieve accurate location variation modeling. The TGM contains four key components: a frame gaze model, and three transition-aware modules (path variation, direction variation, and fusion). First , the frame Transformer extracts gaze location and direction features. Second , to analyze the feature difference among transition frames, we introduce TVM guided by transition-aware loss. TVM analyzes the location features to capture the moving trajectory of targets (defined as path variation ), which facilitates the search for target locations near the path. Third , TVM also analyzes the direction features to capture the transition-aware direction area (defined as direction variation ), which facilitates the search for target locations within this area. Fourth , since gaze directions dynamically adjust to track gaze targets, path variation, and direction variation are inherently aligned with the natural movement of a person’s gaze. Thus, these two variations are fused into a unified transition-aware feature, which helps cover all potential target locations. To search for accurate target locations, we embed this transition-aware feature into frame features with cross-attention, which can enhance gaze target detection in transition frames. Extensive experiments demonstrate that our method achieves state-of-the-art performance on two datasets, namely VideoAttentionTarget and VideoCoAtt.
The Dense Audio-Visual Event Localization (DAVEL) task aims to temporally localize events in untrimmed videos that occur simultaneously in both the audio and visual modalities. This paper explores DAVEL under a new and more challenging weakly-supervised setting (W-DAVEL task), where only video-level event labels are provided and the temporal boundaries of each event are unknown. We address W-DAVEL by exploiting cross-modal salient anchors, which are defined as reliable timestamps that are well predicted under weak supervision and exhibit highly consistent event semantics across audio and visual modalities. Specifically, we propose a Mutual Event Agreement Evaluation module, which generates an agreement score by measuring the discrepancy between the predicted audio and visual event classes. Then, the agreement score is utilized in a Cross-modal Salient Anchor Identification module, which identifies the audio and visual anchor features through global-video and local temporal window identification mechanisms. The anchor features after multimodal integration are fed into an Anchor-based Temporal Propagation module to enhance event semantic encoding in the original temporal audio and visual features, facilitating better temporal localization under weak supervision. We establish benchmarks for W-DAVEL on both the UnAV-100 and ActivityNet1.3 datasets. Extensive experiments demonstrate that our method achieves state-of-the-art performance.
Remote photoplethysmography (rPPG) measurement enables non-contact physiological monitoring but suffers from accuracy degradation under head motion and illumination changes. Existing deep learning methods are mostly heuristic and lack theoretical grounding, limiting robustness and interpretability. In this work, we propose a physics-informed rPPG paradigm derived from the Navier–Stokes equations of hemodynamics, showing that the pulse signal follows a second-order dynamical system whose discrete solution naturally leads to a causal convolution, justifying the use of a Temporal Convolutional Network (TCN). Based on this principle, we design the PHASE-Net, a lightweight model with three key components: 1) Zero-FLOPs Axial Swapper module to swap or transpose a few spatial channels to mix distant facial regions, boosting cross-region feature interaction without changing temporal order; 2) Adaptive Spatial Filter to learn a soft spatial mask per frame to highlight signal-rich areas and suppress noise for cleaner feature maps; and 3) Gated TCN, a causal dilated TCN with gating that models long-range temporal dynamics for accurate pulse recovery. Extensive experiments demonstrate that PHASE-Net achieves state-of-the-art performance and strong efficiency, offering a theoretically grounded and deployment-ready rPPG solution.
Non-transferable learning (NTL) aims to restrict a model's generalization ability to a target domain. Existing methods primarily focus on coarse-grained tasks. However, the problem of NTL at the fine-grained level remains largely unexplored, mainly due to the difficulty in distinguishing subtle inter-class variations within the source domain while simultaneously suppressing transferable patterns that may generalize to the target domain. To this end, this paper proposes CNIE for this challenge, and the framework consists of three key components. The Fine-grained Prompt Representation (FPR) module effectively activates the rich semantic knowledge embedded in the pre-trained model by introducing mixed-granularity prompts and transfers it to downstream tasks. The Content-Aware Decoupling Model (CDM) disentangles visual features into two orthogonal components: content and context. The context captures latent domain-specific information to guide domain-level suppression and enhance non-transferability. The Non-Transferable Information Extraction (NIE) module encodes subclass-specific content, which is constrained by an information bottleneck to learn a minimal sufficient representation that preserves the intrinsic causal semantics of fine-grained categories. Both qualitative results and quantitative visualizations on three FGVC-related NTL datasets consistently demonstrate that the learned fine-grained feature representations achieve optimal non-transferability to the target domain with minimal degradation in the source domain, while also establishing new state-of-the-art results in non-transferable fine-grained representation learning.