
Modern face recognition paradigms predominantly rely on angular margin penalties to enforce strict geometric boundaries between classes. However, these methods suffer from Margin Saturation, where increasing margin penalties yield negligible gains due to an intrinsic performance ceiling. In this work, we identify a statistical contributing mechanism of this bottleneck: the uncontrolled variance growth of negative proxies. Through theoretical analysis, we show that the stochastic noise inherent in SGD training drives negative proxy similarities to converge to Normal distributions with accumulating variance, a critical issue we term the Variance Amplification Defect. This defect increases inter-class confusion, which can be particularly detrimental in high-security scenarios sensitive to tail distributions. To address this, we propose NormalFace, a corrective framework centered on variance-aware loss modulation that requires no additional manually tuned hyperparameters. Unlike heuristic approaches requiring complex hyperparameter searches, NormalFace derives a closed-form weight function that explicitly counteracts the SGD-induced variance drift. Experimental results across multiple benchmarks demonstrate that NormalFace effectively rectifies distribution distortion. While achieving performance on par with strong baselines on standard metrics, it shows particular advantages in stringent high-security regimes (e.g., FAR=1e-6), where controlling false-accept risk is especially important, supporting its role as a useful correction to the softmax-based paradigm.
Multi-object tracking is a crucial task in intelligent surveillance, yet it remains challenging in crowded environments due to severe occlusion, perspective distortion, and appearance similarity. To address these challenges, this work proposes Spatial Object Permanence Track (SOPERM-Track), which aims to endow a motion-based tracker with human-like spatial object permanence by reconstructing spatiotemporal consistency through physical geometric constraints. Specifically, short-term height stability maintains scale continuity when object observations become unreliable due to occlusion, ground contact distance models ground-plane spatial relationships from the bottom edges of detection boxes, and a global perspective map learns scene-level perspective patterns online for global scale calibration. These complementary spatial cues form an appearance-free association mechanism to improve identity consistency under occlusion, interaction, and non-linear motion. On the MOT17, MOT20, and DanceTrack benchmarks and on the shuttlecock trajectory set, SOPERM-Track demonstrates competitive tracking performance and runtime efficiency in scenarios involving long-term occlusion, similar appearances, complex motion patterns, and fast tiny-object trajectory tracking.
In multimedia understanding, transferring knowledge across domains with different data distributions remains a fundamental challenge. Universal Domain Adaptation (UniDA) addresses this problem by transferring knowledge from a labeled source domain to an unlabeled target domain without prior knowledge of the target label set. UniDA involves two key tasks: cross-domain alignment of shared-class samples and accurate identification of unknown-class samples. For shared-class alignment, conventional explicit alignment methods often fail to preserve intra-class compactness under large domain shifts in visual appearance, leading to negative transfer. For unknown-class identification, existing approaches mainly rely on inter-class feature discrepancies while neglecting explicit modeling of the semantic boundary between known and unknown categories, resulting in limited discriminative capability for novel classes with uncertain quantity and distribution. To address these issues, we propose CASE (Cross-modal semantic Anchoring alignment and Structure Enhancement), a novel UniDA framework that incorporates cross-modal semantic information. CASE includes two core modules: (1) a Cross-modal Semantic Anchoring Alignment (CSA) module that constructs a shared semantic anchor space between source and target domains to enable semantic-guided cross-domain alignment; and (2) a Semantic Structure Enhancement (SSE) module that integrates class-level semantic information to enhance image representations and improve the separability between known and unknown classes in semantic space. Experiments on several standard UniDA benchmarks demonstrate the effectiveness of CASE. The demo code is publicly available at https://github.com/yuefengw/CASE-UniDA.
The Contrastive Vision-Language Pre-training (CLIP) methodology represents a groundbreaking advancement in visual representation learning, employing large-scale image-text contrastive learning to align image and text representations directly. Few-shot adaptation techniques based on CLIP focus primarily on two key areas: firstly, fine-tuning the CLIP model with limited samples or incorporating additional adaptable modules to optimize compatibility with specific datasets; secondly, enhancing model predictions through trainable text prompts to boost few-shot performance. Despite these advances, a significant challenge remains: the model often overfits during fine-tuning with limited samples, leading to considerably decayed generalization across unseen categories. To address these issues, we introduce CMA-CLIP, a pair of lightweight cross-modal adapters that bolster CLIP capabilities online using few-shot samples. We extract both global and local features from input image-text pairs, refining text features with guidance from the image's global feature. Concurrently, the global feature and text feature collaboratively identify the semantically relevant local features, which are fused with the global feature to create an enriched representation. This cross-modal attention module fosters mutual awareness between images and texts, effectively bridging the feature gap between two modalities. Our approach demonstrates robust few-shot adaptation capabilities, generalizing to novel classes without further fine-tuning, thus alleviating the over-fitting issue and achieving superior performance. Across eight classification datasets, our method outperforms the state-of-the-art approach by an average margin of approximately 8.33%. Extensive experiments across various visual classification tasks and thorough ablation studies validate the efficacy of our approach.
Recent sign language translation (SLT) models have achieved significant improvements under the assumption that source data and target data follow the same distributions. However, this assumption rarely holds in practical scenarios, as models often encounter distribution shifts induced by common corruptions in target data during inference, which disrupt temporal dynamics and semantic information essential for understanding sign language videos, leading to performance degradation. Test-time adaptation (TTA) emerges as a research paradigm to mitigate distribution shifts by adapting models on unlabeled target data at test time, and has shown promise in other multimedia tasks (e.g., action recognition, video classification). However, TTA remains unexplored in SLT. In this paper, we present the first study on the problem of corruption-induced distribution shifts in SLT, and propose a novel TTA method: temporal disentanglement and regularization (entitled TRANSLATE), which leverages intrinsic temporal dynamics and semantic cues in sign language videos as self-supervised signals. TRANSLATE includes three key components: (i) A parameter-free feature disentanglement method, which separates dynamic sign features (e.g., hand movements) and temporally stable non-sign features (e.g., other body parts) based on temporal variance; (ii) A temporal dual regularization strategy on disentangled features to enhance temporal awareness, which promotes motion distinctiveness and ensures temporal stability; (iii) A textual consistency regularization based on motion-aware augmentation, which preserves motion information and enforces semantic consistency of predictions across original and temporally augmented views. Extensive experiments demonstrate that TRANSLATE consistently improves SLT performance under diverse corruptions, validating its effectiveness and robustness.
Large Vision-Language Models (VLMs) have demonstrated exceptional performance in knowledge-based visual question answering, while small-scale VLMs, owing to their distinct training paradigms, retain complementary strengths that large VLMs lack. Consequently, collaboration between large and small VLMs has emerged as a promising research direction. However, existing collaborative frameworks predominantly rely on unidirectional prompting, which limits the depth of synergistic reasoning across model scales. To address this limitation, we propose SyMRR, a large-small model Synergy framework based on Multimodal Reflective Reasoning, which enables deep, reflective, and visually grounded collaboration between compact and large VLMs. Specifically, we introduce a multimodal reflective chain-of-thought mechanism that guides the large VLM to critically assess the small model's predictions, generating both textual rationales and visually grounded evidence. These reflective signals are then integrated back into the small model's features via a dual reflective interaction strategy, utilizing modality-specific cross-attention to refine the initial multimodal cues. The resulting interaction-enhanced representations guide the large VLM toward more reliable and interpretable answer generation. Notably, SyMRR adopts a parameter-efficient design by freezing both backbone models and optimizing only lightweight interaction modules. Extensive experiments on the OK-VQA and A-OKVQA benchmarks demonstrate that SyMRR achieves state-of-the-art performance among collaborative frameworks, highlighting the effectiveness of deep reflective interaction for complex, knowledge-intensive multimodal reasoning.
In multi-view clustering, there has been a significant increase in recent interest in learning with high-order relations. Tensors have been suggested as the inherent means of describing these relationships. Regrettably, arbitrary signal corruptions, such as noise and missing instances in particular views, frequently accompany multi-view data, which might potentially result in an unreliable or inexact tensor. How to recover the characteristics of the corrupted tensor to make it fully explore the high-order relationships across different views for incomplete multi-view clustering (IMVC) remains a challenging task. To overcome this challenge, this paper proposes a unified tensor learning framework for IMVC, which jointly considers high-order learning with low-rank tensor, sparse noise removal, and missing entries completion by seamlessly integrating the tensor robust principal component analysis (TRPCA) and low-rank tensor completion (LRTC). Specifically, completing the corrupted tensor and removing sparse noise of the completed tensor are simultaneously conducted in the unified framework, where the obtained low-rank tensor can more effectively capture the intrinsic high-order structure information of multi-view data due to enhanced low-rank recovery with the interaction of TRPCA and LTRC. Extensive studies conducted on various benchmark datasets reveal that the proposed framework performs better than the state-of-the-art approaches.
Instruction-driven image editing can compromise the integrity of embedded watermarks, raising concerns regarding copyright protection. Existing image watermarking methods typically leverage diffusion models to simulate editing distortions, yet remain largely limited to single-turn editing scenarios. To address this issue, we propose Reinforcement-Optimized Diversified Watermarking (RODW), a novel framework to tackle the unexplored challenge of maintaining watermark robustness under multi-turn image editing. RODW guides the watermarking model to better learn the complex characteristics of multi-turn image editing by simulating real-world instruction-driven image editing scenarios. Specifically, RODW introduces the diversified sampling generation strategy that constructs editing instructions in various styles. These instructions provide richer editing features for watermarking model training, enabling the watermarking model to learn from a more comprehensive set of simulated editing operations. Furthermore, we design a Reinforcement Learning Optimization (RLO) strategy to guide the watermark embedding process, which allows the watermarking model to learn fine-grained editing features, ultimately enhancing watermark robustness under multi-turn image editing. Extensive experiments on benchmark datasets demonstrate that RODW exhibits significant advantages over existing state-of-the-art methods.
Existing 3D defect detection methods suffer from poor transferability to unknown defects due to the uncertain shapes and small sizes. Furthermore, such high-precision 3D annotations are often expensive and time-consuming. Therefore, we propose a domain adaptation method based on coupling consistency for semi-supervised 3D defect detection (DACC). First, more potential soft pseudo-labels are generated through a dual-threshold strategy in the teacher network, blending them with a small number of ground truth labels to enhance the effective transfer of supervision information. Then, the domain adaptation student network focuses on aligning the feature distributions between the regular view domain and the normal view domain, where the feature representations of the aligned domains are comparatively learned by introducing coupling consistency. In addition, we focus on small defect instances, and the weights of these instances are incrementally updated through exponential moving average. Experiments on the CPS3D-Det dataset demonstrate that our DACC achieves state-of-the-art performance compared with existing fully supervised methods, and it also has a competitive advantage on the public dataset, the KITTI 3D detection dataset.
Deep learning methods for CNV segmentation primarily rely on supervised learning, which requires significant effort to produce pixel-level annotations. Semi-supervised learning (SSL) is a promising approach in cases of limited annotation. However, in clinical practice, imaging factors such as different vendor devices or imaging protocols often degrade OCT image quality, impacting the performance of SSL methods: 1) In supervised learning, existing SSL methods often fail to properly handle low-quality OCT images, as the training data are dominated by regular OCT images. 2) In pseudo-label learning, all unlabeled images are utilized simultaneously, leading to incorrect labels and reduced performance. To address these challenges, we propose a Difficulty and Stability-Aware Semi-supervised Network (DSASS-Net) for CNV segmentation. To better segment low-quality OCT images, we develop a Two-Phase Adaptive Difficulty-aware Knowledge Distillation method: first, a Difficult Prior-aware Module is designed to capture prior knowledge from low-quality OCT images; second, an Adaptive Consistency Regularization is introduced to distill difficult prior knowledge adaptively. To reduce the impact of incorrect pseudo labels caused by low-quality OCT images, we further design a Holistic Stability Prior-aware Module to effectively select reliable unlabeled images. Intensive experimental results on the collected dataset demonstrate that our DSASS-Net outperforms state-of-the-art semi-supervised segmentation methods by 1.14%, 1.3%, and 3.35% in terms of TPR, Dice, and IoU, respectively.
Parameter-efficient fine-tuning (PEFT) has emerged as a popular solution for adapting pre-trained models to Referring Image Segmentation (RIS) by updating only a small set of parameters. Although these approaches improve training efficiency, they still incur significant inference costs due to redundant vision tokens. While current token compression (TC) techniques offer a way to alleviate computational burdens, they inevitably lead to a dramatic decline in performance for RIS, a dense prediction task related to text descriptions. In this work, we explore token compression specifically tailored for RIS. Unlike merging-based or pruning-based TC methods, which disrupt the complete spatial structure of images, we propose SkipRIS, a novel approach that leverages text semantics to accurately identify and skip redundant vision tokens during complex computations. Our experiments reveal that SkipRIS significantly outperforms existing token compression techniques. Most impressively, SkipRIS achieves performance comparable to full token training while skipping 50% of vision tokens.
Task planning is a fundamental capability for household robots to accomplish complex tasks. However, it remains highly challenging due to environmental variability, object uncertainty, and execution anomalies. Existing approaches either rely on rule-based static representations with limited adaptability or employ large language models (LLMs) for reasoning, which often fail to guarantee physical executability and robust recovery under changing conditions. To address these limitations, we propose a dynamic task-planning framework for household robots based on a hierarchical object–scene association mechanism. Specifically, Hierarchical Object–Scene Association Modeling (HOSAM) captures multi-level associations between objects and scenes to enable effective home environment representation with both global scene awareness and fine-grained semantic grounding. On this basis, Association-Guided Intelligent Substitution (AGIS) infers suitable alternative objects according to functional and attribute similarity when the target object is missing, unavailable, or abnormal, thereby preserving task continuity. Furthermore, Association-Aware Task Planning Generation (ATPG) integrates association structures and substitution results into the planning process to generate fine grained and executable task sequences under scene constraints, improving both physical feasibility and adaptability. Extensive experiments in virtual simulations and real-world scenarios demonstrate that the proposed framework outperforms baseline approaches in task completion efficiency, adaptability to environmental changes, and error recovery capability, highlighting its potential for practical deployment in dynamic home environments.
Due to the lack of large-scale labeled Thermal InfraRed (TIR) datasets, most TIR trackers usually use the finetuning paradigm. However, existing methods based on Full Fine- Tuning (FFT) and Parameter-Efficient Fine-Tuning (PEFT) often do not explicitly model TIR-specific prior knowledge, thereby limiting their performance improvement. To this end, we propose a new FFT-based TIR tracking framework based on visual and semantic priors, termed VSPT. Specifically, this framework consists of three main components: a pre-trained ViT-based tracker, a visual prior branch, and a semantic prior branch. Given that visual features such as multi-scale and shape are crucial for TIR tracking tasks, we first propose a visual prior module employing a dynamic multi-scale CNN architecture to model this prior information. Second, considering the challenges faced by TIR tracking, like target deformation, occlusion, and background distractors, we propose a semantic prior module based on a pre-trained MAE-based infrared large model. This branch captures semantic information in TIR images through a dynamic masking mechanism. Third, we design a feature injector and a dynamic feature extractor based on the cross-attention mechanism. The feature injector integrates these priors into each Transformer block of ViT via dynamic residual connections, while the feature extractor simulates multiple levels of prior features injected into the prior branches. Since our method is a parallel injection structure, it does not destroy the original ViT architecture, therefore, it can effectively utilize the pretrained ViT features and fine-tune on a small-scale TIR dataset to learn TIR-specific features simultaneously. We conduct extensive experiments on four TIR tracking benchmarks, and the results show that our method sets a new state-of-the-art. The code is available at https://github.com/honggg-source/VSPT.
Image Change Captioning (ICC) aims to generate natural language descriptions that accurately reflect semantic differences between two visually similar images. Most existing methods adopt Transformer-based models, or reframe ICC as a prompt-driven generation task using multimodal large language models (MLLMs). However, these approaches often struggle with limited generalization and low sensitivity to subtle changes. To overcome these challenges, we propose a unified and efficient data-driven framework that combines synthetic supervision with gradient-guided perceptual enhancement. Technically, we first introduce a scalable and diverse synthetic data synthesis method that leverages MLLMs and controllable image editing models to automatically generate diverse, high-quality, change-oriented training samples across various domains, covering both semantic change and no-change cases while reducing the reliance on manual annotation. To further enhance fine-grained change perception, we introduce a Difference-Aware Visual Cropping mechanism during inference. This mechanism reweights attention maps using gradients from the language prediction loss, enabling the model to dynamically crop and focus on the most critical change regions before caption generation. Extensive experiments demonstrate that our method achieves competitive performance with State-of-the-Art (SOTA) methods, highlighting its broader generalization and broad applicability. Code is available at https://github.com/user-jinhong/UEM-MLLM.
Diffusion models (DMs) have recently shown remarkable results in image deblurring, a vital task in computer vision. Despite their progress, the existing methods suffer from local inconsistency, leading to unrealistic reconstructed image details. To address this challenge, we observe that high-frequency information in the existing DMs can exaggerate and generate unrealistic image details, leading to local inconsistency. On the other hand, low-frequency information usually preserves the global structure of the image, which is more conducive in diffusion to generate realistically deblurred images. Motivated by this observation, we propose a novel framework Diffusion in Low- Frequency latent space (DiffLF) for image deblurring to enhance DM's generative capability while reducing unnatural details. Specifically, we propose to apply DM in the low-frequency latent space, which is naturally formed by progressively downsampling in a U-Net architecture. Further, we realize the U-Net architecture by proposing low-frequency latent space diffusion (LFD) blocks and effective state space (ESS) blocks. Our method integrates the advantages of DMs and Mamba to preserve long-sequence characteristics while maintaining the deblurring ability. To evaluate our proposed method, we conduct extensive experiments on benchmark image deblurring datasets. Experimental results show that our proposed DiffLF outperforms the state-of-the-art approaches.
Recent category-level pose estimation methods leverage text prompts describing object appearance or category attributes to enhance performance. However, appearance attributes are task-irrelevant and susceptible to instance or environmental variations, which may introduce spurious statistical correlations between learned features and pose labels, impairing cross-scenario adaptability. Moreover, the category attribute only provides coarse-grained and pose-insensitive semantics, limiting its effectiveness in guiding precise pose estimation. To mitigate these issues, we propose OAG-Pose, a framework that integrates orientation description into text prompts. The orientation description is a task-relevant attribute that encapsulates the rotational relationship between the object and the camera coordinate system, providing fine-grained and robust semantics against instance and environmental variations, thereby enhancing the guidance of text prompts for pose estimation. Additionally, we propose a parameter-efficient Normalized Object Coordinate Space (NOCS) reconstruction method to alleviate the limitations of orientation guidance in geometric supervision. Unlike previous methods that rely on shape priors or direct regression, we introduce coordinate consistency constraints through cross-projection interactions, enabling the model to learn geometric features from multiple perspectives and reduce optimization complexity. OAG-Pose achieves competitive performance on the REAL275 and CAMERA25 datasets while exhibiting superior generalization on the Wild6D dataset and various real-world scenarios.
Open-vocabulary Video Question Answering (OVQA) is a challenging task as it requires models to demonstrate generalizability when encountering novel, undefined categories in videos, which is highly relevant to the openness and diversity of video content in real-world scenarios. Existing OVQA approaches primarily enhance model understanding by incorporating external knowledge, such as integrating pretrained GloVe embeddings into Vision-Language Models, aiming to bridge the semantic gap between video content and questions. However, these text-based methods suffer from two fundamental limitations: (1) inherent inability to capture fine-grained visual details, and (2) ambiguous textual descriptions due to natural language polysemy, collectively constraining performance on rare/unseen answer categories. To address the above issue, this paper tackles OVQA from a different perspective by employing a knowledge-based multi-modal prompts learning framework. The proposed method, named MMQA, adopts two innovative components to pursue a more robust embedding of category cues. A novel heterogeneous graph is first introduced to integrate video visual features, question semantic representations, and external knowledge, utilizing nodes and edges to facilitate the exploration of potential connections between rare/unseen answers and external knowledge. To address textual ambiguity, we propose a “text+image exemplar” complementary prompting strategy to generate multi-modal prompts by retrieving image exemplars from knowledge graphs, thereby mitigating text coarseness. Extensive experiments on the MSVD-QA, ActivityNet-QA, TGIF-QA, and MSRVTT-QA datasets have demonstrated the promising performance of MMQA, e.g., it outperforms existing methods across rare and unseen answers.
Recent single-image super-resolution methods leverage the generative priors of diffusion models to achieve superior perceptual quality in the reconstructed images. However, these perceptual-driven methods suffer from severe fidelity distortion in stereo super-resolution tasks due to the stereo inconsistency introduced by the stochastic nature of the diffusion generation process. We address this challenge by leveraging stereo constraints to guide diffusion at both noise sampling and diffusion feature fusion, allowing the model to handle complex low-resolution degradation while maintaining cross-view consistency, achieving a favorable perceptual-fidelity balance. To be specific, we address the fundamental problem that the inherent stochasticity of diffusion disrupts stereo consistency by proposing a stereo-consistent noise sampling optimization strategy, where we obtain the leftview noise by sampling an approximate optimal initial noise and optimizing the variance noise via a perceptual-consistency loss, while the right-view noise is obtained by disparity-aware warping. This approach drives the diffusion process toward outputs that satisfy stereo high-frequency detail consistency while delivering improved perceptual quality. To enhance stereo consistency at the structural and semantic levels, we design a Disparity-Aware State-Space Module, which performs stereo feature fusion via a sequentially bi-directional cross-scanning approach that is more suitable for the diffusion feature space than traditional attention operation. Together, we build a diffusion-based stereo superresolution framework, DiffSSR+, which progressively integrates stereo guidance into the diffusion process to achieve a better fidelity-perception trade-off, and is plug-and-play to a broad range of diffusion SR architectures.
This paper presents a reconstruction-based framework for autonomous driving testing in urban dynamic scenes. To the best of our knowledge, it is the first to transform reconstructed urban dynamic scenes into a vectorized Gym-style observation environment for training reinforcement learning driven added vehicles (RLPlanner). For urban dynamic scene reconstruction, we propose a Learnable Time Window (LTW) mechanism for 4D GS, where a lightweight neural module adaptively estimates temporal windows from local spatiotemporal context. Shorter windows are assigned to dynamic regions to suppress motion blur, while longer windows are used for static regions to improve temporal consistency. For added vehicle testing, the framework bridges the gap between visual scene representation and downstream control tasks by transforming the reconstructed scenes into a vectorized RL training environment. This enables closed-loop control of added vehicles via reinforcement learning, achieving autonomous obstacle avoidance and trajectory planning. Extensive experiments on challenging urban dynamic datasets demonstrate that our reconstruction method achieves competitive visual quality compared to state-of the-art frameworks. Moreover, the RLPlaner successfully performs obstacle avoidance and trajectory planning across multiple reconstructed scenes, and multiple reconstructed scenes with reinforcement learning-driven added vehicles are obtained. Project page: https://github.com/fangthawzw/VecSIM.
The performance of cross-modal retrieval models is heavily dependent on the quality of the training dataset. However, the collected data inevitably suffers from the noisy correspondence problem, which prevents the models from learning an accurate common representation space. Most existing methods handle noisy data by first dividing the dataset to compute soft correspondence labels, which then adjust sample weights during training, enhancing the model's robustness and discriminative alignment. However, these existing methods divide the dataset within a single representation space, failing to accurately distinguish between matched pairs and weakly matched pairs with subtle differences. Furthermore, they overlook the importance of identifying images and texts with similar semantic content for constructing an accurate common representation space. To address these limitations, we propose a novel Dual Division Network (DDN). Our method divides the dataset into a non-noisy subset, where image-text pairs possess semantic relevance, and a noisy subset, where such relevance is absent. Then, the image-text pairs from the non-noisy subset are mapped into a multiple representation space constructed based on Graph Convolutional Networks. A hierarchical difference division strategy is designed to differentiate between matched pairs and weakly matched pairs within the non-noisy subset. Additionally, by leveraging the non-noisy probability and the clean probability, we identify images and texts with similar semantic content and propose a homogeneity enhancement loss to minimize the distance between these cross-modal samples in the common representation space. Experimental results validate the effectiveness of our method on widely-used benchmark datasets, including Flickr30K, MS-COCO, and Conceptual Captions.