Recent advances in video multimodal models have significantly improved VideoQA performance. However, these systems often rely on spurious statistical correlations rather than answer-relevant causal evidence, resulting in unfaithful and brittle reasoning, especially in complex real-world scenarios. Existing methods either rely on cross-modality correlations, costly curated training resources, or insufficient causal assumptions and constraints, and typically operate at the time-interval level. As a result, they fail to explicitly disentangle causal visual cues from confounders and provide limited fine-grained evidence localization. To address this issue, we propose a Counterfactual Reasoning framework for fine-grained Evidence Disentanglement (CREDiT). CREDiT formulates the VideoQA process using a structural causal model and learns cross-modality representations that are explicitly decomposed into causal and non-causal components under independence and minimality constraints. To facilitate faithful disentanglement, we introduce feature-level causal interventions and construct counterfactual inputs that approximate causal effects while suppressing non-causal correlations. Extensive experiments on NExT-GQA, SportsQA, and SPORTU-video demonstrate that CREDiT consistently improves answer accuracy and reasoning reliability across both generic and complex sports scenarios, leading to more trustworthy VideoQA systems.
Compositional Zero-Shot Learning (CZSL) addresses the challenge of recognizing unseen attribute-object compositions in images, representing a fundamental challenge in artificial intelligence. Current approaches, which primarily focus on semantic alignment or distribution independence of primitives, have not achieved effective state-object decoupling and causal interventional invariance, limiting their performance on unseen compositions. To tackle this challenge, this study introduces I2CD (Invertible Causal framework via Disentangle-Compose-Disentangle), a novel framework that integrates invertible neural networks with causal intervention techniques to achieve state-object disentanglement. The framework employs a disentangle-compose-disentangle mechanism for counterfactual generation within the disentangled representation space, ensuring that modifications to one primitive (attribute or object) maintain independence from the other, thus enabling robust causal disentanglement. Representational consistency is maintained through semantic alignment between initial disentangled representations and their recomposed-then-disentangled counterparts with corresponding textual concepts. Comprehensive evaluations on three benchmark datasets—MIT-States, UT-Zappos, and C-GQA—demonstrate the framework's effectiveness in achieving both disentanglement and compositional generalization in CZSL tasks.
Precise environmental perception is critical for the reliability of autonomous driving systems. While collaborative perception mitigates the limitations of single-agent perception through information sharing, it encounters a fundamental communication-performance trade-off. Existing communication-efficient approaches typically assume MB-level data transmission per collaboration, which may fail due to practical network constraints. To address these issues, we propose InfoCom, an information-aware framework establishing the pioneering theoretical foundation for communication-efficient collaborative perception via extended Information Bottleneck principles. Departing from mainstream feature manipulation, InfoCom introduces a novel information purification paradigm that theoretically optimizes the extraction of minimal sufficient task-critical information under Information Bottleneck constraints. Its core innovations include: i) An Information-Aware Encoding condensing features into minimal messages while preserving perception-relevant information; ii) A Sparse Mask Generation identifying spatial cues with negligible communication cost; and iii) A Multi-Scale Decoding that progressively recovers perceptual information through mask-guided mechanisms rather than simple feature reconstruction. Comprehensive experiments across multiple datasets demonstrate that InfoCom achieves near-lossless perception while reducing communication overhead from megabyte to kilobyte-scale, representing 440-fold and 90-fold reductions per agent compared to Where2comm and ERMVP, respectively.
Class Incremental Learning (CIL) requires models to acquire knowledge from sequential tasks containing non-overlapping classes while avoiding catastrophic forgetting. While vision-language foundation models like CLIP demonstrate remarkable potential for CIL through their pre-trained cross-modal alignment capabilities, existing CLIP-based approaches critically overlook the progressive degradation of visual representations in incremental scenarios . Through feature space analysis, we identify a crucial dichotomy : textual embeddings maintain stable discriminative power across sequential tasks, whereas visual features exhibit progressive deterioration manifested by intra-task confusion (ambiguous decision boundaries between co-occurring classes) and inter-task interference (semantic collision between historical and novel categories). To address these dual challenges, we propose T ask-guided H ierarchical M ulti-modal align M ent ( THMM-CLIP ), a framework that establishes persistent visual-textual coherence through Hierarchical Multi-modal Alignment (HMA) and Robust Prompt Selection (RPS) . HMA adapts lightweight task-specific prompt vectors to dynamically recalibrate the CLIP image encoder, thereby achieving: (i) intra-task alignment; (ii) inter-task discriminability alignment; and (iii) global structural alignment with textual features. RPS incorporates a dual-level task identifier that integrates class-level and task-level representative features to ensure precise prompt retrieval during inference. Ablation studies validate all components’ contributions, while t-SNE visualizations, confusion matrices, and Grad-CAM analyses confirm strengthened cross-modal alignment.
Collaborative perception is a powerful paradigm that addresses the limitations of single-agent perception caused by occlusion and distance. However, in practice, sparse-agent collaborative perception (SCP) faces similar vulnerabilities to single-agent perception, leading to significant performance degradation compared to dense-agent collaborative perception (DCP). In this paper, we investigate the negative impact of SCP and propose a unified framework, CoTrans, to bridge the SCP-DCP gap for existing collaborative perception methods. This framework relies on a conceptually simple yet efficient collaboration translation, which simulates the gap in a learnable manner. Furthermore, in CoTrans, a drop-agent strategy is introduced to support the learning process without any communication burden, and a lightweight injection network equipped with localization and injection capabilities is proposed to fuse the collaboration translation with SCP features. To the best of our knowledge, this is the first unified framework designed to enhance SCP. Extensive experiments across five large-scale collaborative perception datasets demonstrate that CoTrans consistently enhances existing collaborative systems in SCP scenarios, while incurring only 1 ms of inference latency. In particular, it achieves a 6.6% AP gain for AttFuse on the real-world DAIR-V2X dataset under sparse-agent conditions. The project page is at https://weiquanmin.github.io/cotrans.
Few-shot class-incremental learning (FSCIL) requires a model to learn the knowledge of new categories incrementally, using only a few samples, after being trained on a base session with ample categories and sample sizes. This task presents two major challenges: catastrophic forgetting and overfitting. Current approaches primarily enhance the model’s ability to extract knowledge during the base stage to improve adaptability to new tasks. Large-scale pre-trained models, known for their high robustness and zero-shot transfer capabilities, have demonstrated promising performance in FSCIL. The key to solving FSCIL lies in effectively fine-tuning such large models to balance the learning of new knowledge and the retention of old knowledge. Inspired by human-like knowledge retrieval mechanisms, we propose Class-specific Knowledge-Guided Prompt Tuning (CKGPT), which leverages class-specific prompts to guide the model in learning targeted knowledge reuse and integration effectively. When faced with novel tasks, the model selectively activates previously learned knowledge that is the most relevant, improving performance on new tasks while minimizing updates to irrelevant knowledge to reduce forgetting. By incorporating mechanisms that balance knowledge retention and transfer, CKGPT ensures a more robust adaptation to sequential tasks. Extensive experiments on multiple benchmarks validate the effectiveness of our method in achieving superior performance.
Accurate reconstruction of 3D vehicle pose and shape from monocular images is challenging, particularly for distant objects in autonomous driving. Existing methods often suffer from geometric ambiguity in depth estimation and structural hollowness in shape recovery, primarily due to inadequate multi-scale feature aggregation and unflexible prior modeling. To overcome these limitations, MonoVPR is proposed, a novel framework integrating dynamic context adaptation and progressive geometry refinement. Specifically, a Hierarchical Dual-Context Attention (HDCA) module is introduced to resolve scale-dependent degradation through gated cross-attention across multi-resolution feature maps, dynamically fusing object-centric geometric cues with scene-centric semantics. For shape refinement, the Bounded Iterative Mesh Refiner (BIMR) progressively optimizes template-guided deformations via multi-head attention and a tanh-bounded correction loop, ensuring physically plausible reconstructions.Extensive experiments on the ApolloCar3D benchmark demonstrate MonoVPR achieves state-of-the-art performance, showing exceptional capability in reconstructing geometrically consistent shapes and precise poses for challenging long-range scenarios.
Efficient convolutional neural network (CNN) architecture design has attracted growing research interests. However, they typically apply single receptive field (RF), small asymmetric RFs, or pyramid RFs to learn different feature representations, still encountering two significant challenges in medical image classification tasks: i) They have limitations in capturing diverse lesion characteristics efficiently, e.g., tiny, coordination, small and salient, which have unique roles on the classification results, especially imbalanced medical image classification. ii) The predictions generated by those CNNs are often unfair/biased, bringing a high risk when employing them to real-world medical diagnosis conditions. To tackle these issues, we develop a new concept, Expert-Like Reparameterization of Heterogeneous Pyramid Receptive Fields (ERoHPRF), to simultaneously boost medical image classification performance and fairness. This concept aims to mimic the multi-expert consultation mode by applying the well-designed heterogeneous pyramid RF bag to capture lesion characteristics with varying significances effectively via convolution operations with multiple heterogeneous kernel sizes. Additionally, ERoHPRF introduces an expertlike structural reparameterization technique to merge its parameters with the two-stage strategy, ensuring competitive computation cost and inference speed through comparisons to a single RF. To manifest the effectiveness and generalization ability of ERoHPRF, we incorporate it into mainstream efficient CNN architectures. The extensive experiments show that our proposed ERoHPRF maintains a better trade-off than state-of-the-art methods in terms of medical image classification, fairness, and computation overhead. The code of this paper is available at https://github.com/XiaoLing12138/Expert-Like-Reparameterization-of-Heterogeneous-Pyramid-Receptive-Fields.
Deploying neural networks on microcontroller units (MCUs) is critical for edge intelligence but remains challenging due to tight memory, storage, and computation constraints. Existing approaches, such as model compression and hardware-aware neural architecture search (HW-NAS), often depend on proxy metrics, incur high search cost, and do not fully bridge the gap between architecture design and verified deployment. This paper presents AutoMCU, a feasibility-first large language model (LLM)-based multi-agent system for automated neural network customization under MCU constraints. Given natural-language task requirements and hardware specifications, AutoMCU iteratively generates structured architecture candidates, filters infeasible designs through vendor toolchain feedback before training, evaluates feasible models under a controlled protocol, and verifies deployability through backend-grounded deployment analysis. AutoMCU includes two key mechanisms: 1) hardware-in-the-loop architecture generation for early elimination of undeployable candidates under RAM and Flash constraints, and 2) state-isolated multi-agent scheduling for stable coordination of proposal, training, evaluation, and deployment stages. Experiments on CIFAR-10 and CIFAR-100 under strict MCU constraints show that AutoMCU achieves competitive accuracy while reducing customization time to about 1–2 hours, compared with hundreds of GPU hours for representative MCU-oriented HW-NAS baselines. Comparisons with ColabNAS and the LLM-based NAS method GENIUS on NAS-Bench-201 further demonstrate the effectiveness and stability of AutoMCU. Real-device deployments on multiple STM32 microcontrollers validate its practical applicability to MCU-scale edge intelligence.
Evaluation metrics are essential tools for quantifying the performance of crowd localization models. GAME ignores substantial localization information and is rarely used for crowd localization evaluation. The commonly used localization metrics Precision, Recall, and F-score relies on a greedy algorithm, which can lead to ambiguity matching. Moreover, the distance threshold for the boolean matching matrix lacks both scale sensitivity and universality. To overcome the limitations of matching accuracy, scale invariance, and universality in crowd localization evaluation, a novel metric termed Scale-aware Optimal Transportation Cost (S-OTC) is proposed. To ensure globally optimal matching, it leverages optimal transport theory to compute the cost of transporting weights from ground truth points to predicted points, which serves as a measure of accuracy. First, S-OTC models predicted and ground truth points as two weighted discrete measures, thereby establishing a stable and consistent optimal transportation plan. Then, an adaptive scale-sensitive method is proposed, which produces a scale prior with only point annotations and refines the measure of ground truth to adapt to varying head region scales. This not only ensures insensitivity to head size but also guarantees the universality of both point-annotated and box-annotated datasets. Third, a straightforward but effective normalization technique is presented to refine the cost matrix in transportation plan, ensuring invariance to changes in image resolution. In addition, a new dataset for validating evaluation metrics is constructed, containing 1,800 pairs of perturbed images annotated with human preference choices. This dataset can be used to evaluate the sensitivity of evaluation metrics for count errors and spatial deviations. Extensive experiments demonstrate that S-OTC outperforms existing metrics in terms of stability, sensitivity, and universality. This indicates that S-OTC can serve as the new standard for crowd localization evaluation.
As the final link in the fuel supply chain, gas stations integrate transportation, storage, and retail operations. From a consumer technology perspective, safety, efficiency, and service are key determinants of gas station usability and ensure smooth fueling processes. To boost the intelligent management and control capabilities of gas stations, this article presents AiGuard, an artificial intelligence-driven supervision platform designed to enhance safety management, improve service quality, and optimize operational efficiency at gas stations. It integrates existing algorithms for object detection, real-time tracking, and pose estimation with a purpose-designed model for interactive behavior recognition, enabling comprehensive analysis of key gas station zones, like the fuel unloading area, forecourt, convenience store, and financial office. The platform enhances safety through real-time monitoring and early hazard detection, improves service quality by reducing waiting time and transactional errors, and increases operational efficiency through intelligent scheduling and compliance tracking. To the best of our knowledge, AiGuard is the first comprehensive platform to provide full-site monitoring and intelligent management for gas station operations, offering a practical and scalable solution for future consumer-oriented smart stations.
Modern optical flow estimation, though empowered by deep neural architectures, remains rooted in the discrete correspondence paradigm inherited from classical vision. Most networks infer frame-to-frame displacements or correlation volumes, capturing where pixels move but not how motion evolves continuously through time. Yet physical motion in the real world follows smooth dynamics governed by underlying velocity fields, as long established in fluid mechanics and transport theory. To bridge this gap, we introduce Optical Flow Matching (OFM), a continuous formulation that learns a time-dependent velocity field to transport pixel coordinates along motion distribution coherent trajectories. A key component of our OFM is Triangle Velocities Synergy (TVS), a lightweight geometric mechanism that provides a stable and physically meaningful velocity construction, ensuring that continuous transport remains well-defined. Combined with an Euler-based ODE solver, OFM yields flow fields that are temporally smooth, geometrically consistent, and process-interpretable. Experiments on Sintel, KITTI, and Spring demonstrate that OFM achieves state-of-the-art accuracy, enhanced temporal stability, and notably stronger cross-dataset generalization, advancing optical flow estimation from correspondence inference to continuous dynamical reasoning. All code and trained models will be released upon acceptance to facilitate further research.
The realistic and controllable generation of pure smoke is critical for smoke image editing, smoke visual special effects generation, and smoke data synthesizing within security scenarios. It is a relatively underexplored topic and continues to present significant challenges. Existing methods face challenges in the generation of smoke with intricate details and the regulation of various smoke styles. In this paper, a Pure Smoke image Generation Network (PSGNet) is proposed with a gradient and style learning approach to generate realistic and controllable smoke images. To achieve flexibility in control across the spatial dimension, the smoke shape mask is used to encode spatial details, such as the location and contour of the smoke, along with other related properties. To enhance the physical realism of synthesized smoke, a novel gradient-based learning framework is proposed to generate smoke gradient features, highlighting a special focus on explicitly encoding and exploiting gradient information. This framework uses a smoke gradient learning architecture that captures the subtle structures and patterns characteristic of real smoke, enabling the generation of highly realistic smoke with rich, fine-scale detail. In addition, a spatially aware style learning strategy is proposed to provide fine-grained control over smoke attributes such as density, color, and overall look. It is able to effectively model style features across both channel and spatial dimensions, thereby enabling spatially aware style manipulation. By combining the gradient module with this style learning framework, the method produces smoke that exhibits rich visual details and customizable image styles. Experiments conducted on six benchmark datasets demonstrate that the proposed PSGNet significantly outperforms the state-of-the-art approaches.
Open-World Object Detection (OWOD) aims to detect unseen objects as “unknown” while incrementally learning them without catastrophic forgetting. This problem presents two major challenges: (1) the lack of annotations for unknown objects during training, and (2) the risk of catastrophic forgetting during model updates. To address these issues, we propose the COmpositional and Bidirectional low-Rank Adaptive open-world detection transformer (COBRA)-a novel framework built upon a pre-trained Deformable DETR model. Specifically, COBRA first employs an attentional filtering mechanism that prunes previously known (P-Known) and currently known (C-Known) objects, yielding a purified set of candidate unknowns. To system-atically pseudo-label these unknowns, we introduce a Primitive Composition Recognition (PCR) module, which evaluates set-level similarity between candidate objects and learned primitives, enabling accurate labeling of pseudo-unknowns. To mitigate catastrophic forgetting during incremental updates, COBRA leverages Bidirectional Low-Rank Adaptation (Bi-LoRA)-a parameter-efficient mechanism that supports forward knowledge transfer and stable backward integration. Together, these components form a synergistic pipeline for continual object discovery and knowledge consolidation. Extensive experiments on MS COCO and PASCAL VOC demonstrate that our rehearsal-free COBRA framework outperforms SAM-powered methods in unknown recall while achieving lower forgetting compared to rehearsal-based competitors.
Edge computing holds great potential to enhance the real-time capabilities of intelligent vehicles by deploying resources at the network edge. However, due to the dynamic and heterogeneous nature of vehicular networks, several challenges remain. These include workload imbalances caused by the interplay between mobility patterns and varying resource availability, the fluctuating decision spaces resulting from the continuous entry and exit of vehicles, and the absence of effective incentive mechanisms for resource preservation. To address these challenges, this paper formulates the task offloading and resource allocation (TSRM) problem within the vehicle-edge cloud cooperation framework by characterizing transmission and computation times accounting for the characteristics of heterogeneous resources and vehicle mobility. We prove that this problem is NP-hard, with the objective of optimizing the cumulative task service ratio to enhance overall system efficiency. We propose a Pre-Matching enhanced Hierarchical Double Auction (PMHDA) algorithm based on hierarchical auction and matching theory, which consists of two main auctions: the vehicle-edge auction (VEA) and the edge-cloud auction (ECA). In the VEA, client vehicles must decide whether to process tasks locally or delegate them to server vehicles or edge nodes. A pre-matching algorithm is designed to identify the optimal offloading candidates, taking into account vehicle mobility and dynamic network conditions. In the ECA, edge nodes determine whether to process tasks locally or forward them to other edge nodes or the cloud. Extensive simulations on real-world datasets demonstrate that PMHDA provides substantial service-rate gains over baseline methods, achieving an average improvement of 75.47% under diverse scenarios.
Source-Free Domain Adaptive Object Detection transfers knowledge from a labeled source domain to an unlabeled target domain while preserving data privacy by restricting access to source data during adaptation. Existing approaches predominantly leverage the Mean Teacher framework for self-training in the target domain. The exponential moving average (EMA) mechanism in the Mean Teacher stabilizes the training by averaging the student weights over training steps. However, in domain adaptation, its inherent lag in responding to emerging knowledge can hinder the rapid adaptation of the student to target-domain shifts. To address this challenge, Dual-rate Dynamic Teacher (DDT) with Asynchronous EMA (AEMA) is proposed, which implements group-wise parameter updates. In contrast to traditional EMA, which simultaneously updates all parameters, AEMA dynamically decomposes teacher parameters into two functional groups based on their contributions to capture the domain shift. By applying a distinct smoothing coefficient to two groups, AEMA simultaneously enables fast adaptation and historical knowledge retention. Comprehensive experiments carried out on three widely used traffic benchmarks have demonstrated that the proposed DDT achieves superior performance, outperforming SOTA methods by a clear margin. The codes are available at https://github.com/qih96/DDT.
Gesture recognition plays a crucial role in a wide range of consumer electronics applications, including human-computer interaction and virtual reality, by enabling the identification and interpretation of human gestures. In recent times, WiFi-based gesture recognition has garnered significant attention due to its privacy protection and unobtrusive nature. However, this approach heavily depends on neural network-based models and is notably influenced by environmental conditions, such as specific location and orientation. To address these environmental impacts, we propose a network framework that leverages transfer learning and meta-learning. The primary focus is on utilizing transfer learning to train a convolutional neural network (ResNet) for feature extraction, enabling the extraction of domain-independent features specifically related to gestures, rather than being influenced by environmental factors. Additionally, we employ the meta-learning algorithm MAML to train the fully connected network for gesture classification. Following training, a small set of samples can be utilized for rapid adaptation to diverse domains, maintaining high accuracy across different domains to achieve cross-domain gesture recognition. Our performance evaluation is based on the Widar 3.0 dataset, encompassing both in-domain and cross-domain scenarios. The simulation results demonstrate that our proposed algorithm surpasses Widar3.0 and WiGRUNT by approximately 7.9% and 1.12% in gesture recognition accuracy across all scenarios.
Grounded situation recognition (GSR) is a comprehensive structured scene understanding task that predicts the salient activity (verb), entities (nouns) involved in the activity with their roles, as well as the corresponding bounding-box groundings of the entities from the given image. Existing I.I.D.-based methods for GSR are limited in their ability to recognize novel verb-noun combinations. To address this problem, in this paper, we novelly consider GSR as a Non-I.I.D. task and focus on learning verb-invariant and role-specific representations for verb and noun predictions. Based on the causality, a novel Contrastive Invariant Risk Minimization (CIRM) model for GSR is proposed. In the proposed CIRM, invariant risk minimization is integrated into a transformer architecture to learn invariant representations for verb prediction. To enhance the intra-verb compactness and the inter-verb separability, contrastive learning is utilized to learn discriminative features. As far as we know, this is the first work that regards the task of GSR as a problem of out-of-distribution generalization. Extensive experiments on the benchmark SWiG dataset demonstrate the effectiveness of our proposed CIRM over other state-of-the-art methods in all evaluation metrics.
Achieving a balance between accuracy and efficiency is a critical challenge in facial landmark detection (FLD). This paper introduces Parallel Optimal Position Search (POPoS), a high-precision encoding-decoding framework designed to address the limitations of traditional FLD methods. POPoS employs three key contributions: (1) Pseudo-range multilateration is utilized to correct heatmap errors, improving landmark localization accuracy. By integrating multiple anchor points, it reduces the impact of individual heatmap inaccuracies, leading to robust overall positioning. (2) To enhance the pseudo-range accuracy of selected anchor points, a new loss function, named multilateration anchor loss, is proposed. This loss function enhances the accuracy of the distance map, mitigates the risk of local optima, and ensures optimal solutions. (3) A single-step parallel computation algorithm is introduced, boosting computational efficiency and reducing processing time. Extensive evaluations across five benchmark datasets demonstrate that POPoS consistently outperforms existing methods, particularly excelling in low-resolution heatmaps scenarios with minimal computational overhead. These advantages make POPoS as a highly efficient and accurate tool for FLD, with broad applicability in real-world scenarios.