Hyperspectral images (HSIs) and multispectral images (MSIs) fusion is a hot topic in the remote sensing society. A high-resolution HSI (HR-HSI) can be obtained by fusing a low-resolution HSI (LR-HSI) and a high-resolution MSI (HR-MSI) or RGB image. However, most deep learning-based methods require a large amount of HR-HSIs for supervised training, which is very rare in practice. In this paper, we propose a coupled diffusion posterior sampling (CDPS) method for HSI and MSI fusion in which the HR-HSIs are no longer required in the training process. Because the LR-HSI contains the spectral information and HR-MSI contains the spatial information of the captured scene, we design an unsupervised strategy that learns the required diffusion priors directly and solely from the input test image pair (the LR-HSI and HR-MSI themselves). Then, a coupled diffusion posterior sampling method is proposed to introduce the two priors in the diffusion posterior sampling which leverages the observed LR-HSI and HR-MSI as fidelity terms. Experimental results demonstrate that the proposed method outperforms other state-of-the-art unsupervised HSI and MSI fusion methods. Additionally, this method utilizes smaller networks that are simpler and easier to train without other data.
Hyperspectral and multispectral image fusion (HMF) enhances spatial-spectral quality by fusing low-resolution hyperspectral images (LR-HSI) with high-resolution multispectral images (HR-MSI). Although recent fusion methods have shown promise in preserving the multi-mode structure of high-dimensional data, existing fusion methods still face some challenges. For tensor-based approaches, conventional mode-wise decomposition, such as order-3 CP or Tucker decomposition, may disrupt intrinsic spatial consistency. Furthermore, although deep learning exhibits powerful feature representation ability, existing deep fusion methods either rely on 'data-driven' deep fusion networks remain insufficiently interpretability with large training data. To address these issues, a novel Self-Expressive High-Order Tensor Unrolling Network (SHOTUN) is proposed for unsupervised HSI-MSI fusion. Within the sparse core tensor decomposition framework, we introduce the intrinsic self-expressive relationships among overlapping image patches as a form of high-order mode representation to preserve spatial structure of the fusion model. During optimization, we adopt an alternative optimizing strategy and design dedicated modules for each sub-problem, yielding an interpretable end-to-end training pipeline. Furthermore, to improve generalization across different sensors, we introduce a pre-training strategy into the unsupervised training for the more accurate estimation of unknown degraded parameters. Extensive experimental results on simulated and real datasets demonstrate the effectiveness of our proposed method. The source code is publicly available at https://github.com/Shawn-H-Wang/SHOTUN.
Small object detection is critical to remote sensing image interpretation, yet extremely small size and weak texture impair spatial representation, while complex backgrounds aggravate cross-scale ambiguity and redundant responses during multi-branch fusion. To address these issues, we propose MRC-YOLO, a multi-granularity representation and context fusion network. The multi-granularity spatial enhancement (MGSE) module jointly models local texture, regional patch-level structure, and long-range strip-shaped dependencies. The cross-scale context collaborative attention (CCCA) module coordinates global, deformable-neighborhood, and patch-wise attention to improve contextual discrimination. The redundancy-aware channel calibration fusion (RCCF) module learns channel-wise coefficients to suppress redundant responses before aggregation. Experiments on AI-TODv2, LEVIR-Ship, and RSOD demonstrate the effectiveness of MRC-YOLO.The source code is available at https://github.com/Nuist-HSI-Group/MRC-YOLO.
Change detection in remote sensing imagery is crucial for monitoring temporal variations in surface characteristics; nevertheless, it presents significant challenges owing to indistinct boundaries, limited semantic differentiation, and inadequate incorporation of multi-scale contextual information. To solve these problems, we propose EIMDGNet (Edge-Induced and Multi-Dimensional Grouped Difference Network), a novel architecture that enhances boundary representation and cross-scale feature interaction for accurate and robust change detection. EIMDGNet adopts a dual-branch ResNet18 backbone to extract multi-scale features from bi-temporal images, capturing both fine spatial detail and high-level semantic context. To improve boundary awareness and reduce pseudo-change interference, we introduce the Edge-Induced Differential Multi-Dimensional Group Enhancement Module (EID-MDGEM). This module enriches fine-grained spatial features through grouped pooling across spatial and channel dimensions, enabling precise localization of change contours. Within EID-MDGEM, the Edge Feature Enhancement Module (EFEM) integrates a parameter-free attention mechanism to generate edge-saliency maps, highlighting true change regions while suppressing background noise and irrelevant variations. To further enhance semantic consistency across feature scales, we design the Multi-Scale Hierarchical Progressive Fusion Module (MSHPM). This component employs a bottom-up progressive strategy to hierarchically integrate low-level spatial details with high-level semantic abstractions, thus increasing the continuity and completeness of detected change regions. By tightly coupling edge-aware enhancement with multi-scale hierarchical fusion, EIMDGNet effectively addresses major obstacles in change detection, including boundary ambiguity, inconsistent scale information, and feature misalignment. We evaluated EIMDGNet on five remote sensing change detection datasets: LEVIR-CD, DSIFN-CD, S2Looking, CLCD-CD and GVLM-CD. Our method consistently outperformed state-of-the-art approaches, achieving 91.49% F1 and 82.93% IoU on LEVIR-CD, 77.32% F1 and 69.39% IoU on DSIFN-CD, the highest 49.19% IoU and 99.20% OA on S2Looking, 81.65% F1 and 72.91% IoU on CLCD-CD, and 85.49% F1 and 76.08% IoU on GVLM-CD. These results demonstrate the superior accuracy and robustness of EIMDGNet across diverse change detection scenarios.
Hyperspectral image (HSI) restoration tasks including super-resolution, denoising, and inpainting, present significant challenges due to intrinsic spectral-spatial coupling and limited training data availability. Recent advances in RGB image restoration demonstrate that models pretrained on large-scale datasets acquire exceptional generalization capabilities, suggesting potential cross-modal knowledge transfer solutions for HSI recovery. However, existing approaches exhibit two critical limitations: (i) prohibitive computational costs from mandatory fine-tuning procedures, and (ii) inadequate cross-modal adaptation causing spectral distortions. To address these challenges, we propose a Two-Stage Cross-Modal Decoupling Network (CMDN) achieves spectral-faithful HSI restoration without fine-tuning the pretrained RGB prior; instead, we perform unsupervised test-time learning only on a lightweight spectral rectifier for sample-specific spectral calibration. Our methodology introduces two fundamental innovations: First, we develop a theoretically grounded framework using Singular Value Decomposition (SVD) to decouple HSIs into orthogonal spatial coefficients and spectral bases. This decomposition enables strategic reconfiguration of spatial coefficients into pseudo-RGB formats through band reorganization, facilitating direct deployment of frozen RGB-pretrained models for spatial textures recovery while preserving spectral integrity. Second, we propose a Physics-Motivated Spectral Rectifier (PMSR) that dynamically adjusts spectral reconstruction weights using spatial gradient priors, correcting spectral deviations through physics-consistent optimization rather than explicit error modeling, thereby achieving superior spectral fidelity. Comprehensive experiments confirm our method’s superiority in both spatial reconstruction accuracy and spectral consistency over state-of-the-art techniques. Code is available at: https://github.com/QYo-Liu/CMDN.
Deep learning (DL) has been extensively applied to hyperspectral image target detection (HTD) with notable success. However, many existing DL-based methods focus on expanding the training samples to capture richer information, resulting in high computational costs and overfitting risks. Additionally, challenges such as complex data distributions and limited model transferability remain significant obstacles. To address these issues, we propose an unbalanced episode meta-learning with Bi-sparse contrastive network (UEML) for HTD. In contrast to directly modeling the target dataset, our approach leverages meta-learning to pre-train the model on a categorical dataset rich in label information, resulting in a universal detection model. Specifically, an unbalanced episode training paradigm is proposed for meta-task construction, which simulates the category-imbalance scenarios inherent to HTD by adaptively adjusting the support set, enabling the acquisition of content-agnostic yet task-relevant transferable meta-knowledge. Additionally, elastic sparsity constraints are imposed on the feature extraction process across both spatial and spectral dimensions, enhancing the model's generalization and discriminative capabilities. During the fine-tuning phase, we employ a pseudo-sample generation strategy based on segmented sampling and spatial-spectral hybrid augmentation to construct the training set, allowing for more accurate and comprehensive sample extraction from complex background regions. This strategy effectively mitigates underfitting caused by insufficient information. Furthermore, contrastive learning is incorporated to address complexities arising by multi-class background characteristics in the pseudo-binary classification task, improving the stability of the detection model. Our proposed algorithm demonstrates rapid target detection capabilities, and experiments on six public datasets indicate that it performs significantly better than existing state-of-the-art methods. Code is available at: https://github.com/QYo-Liu/ UEML.
Anomaly object detection is a critical task for ensuring the safety of high-speed railway system (HSRS). With the advancement of imaging technologies and hardware such as visual sensors, large quantities of high-resolution inspection images are available in HSRS. Against this backdrop, vision-based measurement (VBM) has emerged as the mainstream technology for detecting anomalous objects. In HSRS, VBM is divided into four stages-1) data preprocessing; 2) image analysis; 3) measurand identification; and 4) manual maintenance. Mainstream VBM methods center on the second stage, treating image analysis as the most crucial element. However, recent approaches have been dedicated to locating anomaly objects in images, addressing only "where anomaly objects lie in images." This gap greatly exacerbates the inherent lags between anomaly detection and manual preventive maintenance in HSRS. In this article, we propose an image-based detection and real-world positioning (ID-RP) method, which integrates anomaly localization in both images and the real world into a unified system. As the unique identifier of positional information in HSRS, pillar number plate (PNP) allows to uniquely locate defects in terms of specific section, line orientation, and precise positions within high-speed rail operational routes. Therefore, our ID-RP framework unifies PNP detection and anomaly detection into a complete system. Moreover, we develop a multigeometric feature extraction (MGFE) network, which bridges the feature extraction discrepancy caused by the divergent appearance and morphology of PNP and anomaly objects, enabling joint learning of the two tasks. Our method further constructs a multitask and multimodal learning (M2L) pipeline to simultaneously improve localization precision and classification accuracy of PNP and anomaly objects. We compare our proposed ID-RP with many state-of-the-art common and anomaly object detectors. Extensive experimental comparisons demonstrate that our complete system not only achieves more accurate performance on anomaly detection but also precisely calibrates the real-world position of anomaly objects, providing a more practical and reliable solution for anomaly detection and preventive maintenance in HSRS.
Hyperspectral unmixing (HU) is a fundamental task in the analysis and interpretation of hyperspectral images. A large number of HU methods have been developed rapidly, and several review papers have summarized these developments. More recently, spectral variability (SV) has attracted increasing attention in HU, leading to the emergence of various SV-aware hyperspectral unmixing (SVHU) methods. While existing reviews cover general HU, dedicated surveys on SVHU remain scarce. Moreover, most relevant reviews focus on a specific methodological branch and fail to address the latest advances. To address this gap, we present a comprehensive review of SVHU methods. To the best of our knowledge, this is the first work to propose a unified SVHU framework that covers both single-temporal and multi-temporal settings. Existing SVHU methods are systematically categorized into three groups: model-driven, data-driven, and model–data-driven approaches, each of which is discussed in detail. Subsequently, we describe the datasets and evaluation metrics, and compare several representative state-of-the-art methods on multiple commonly used public datasets. In addition, we summarize related downstream applications, including hyperspectral super-resolution, classification, and change detection. Finally, we discuss the main challenges and outline several promising directions for future research.
In intelligent traffic monitoring, traditional object tracking relies on frame-by-frame detection and tracking. However, the limited computing power of edge devices makes running intensive detection on every frame challenging, hindering real-time analysis. To reduce overhead, sparse detection on key frames with lightweight tracking on intermediate frames is commonly employed. Yet rapid changes in traffic density, vehicle speed, and network bandwidth introduce uncertainty in detection demand and computational load. As a result, end devices with fixed scheduling are often insufficient to maintain tracking accuracy and real-time performance. To tackle these challenges, we propose an end–end–edge collaborative, content-aware video frame scheduling system. It constructs a multi-layer collaborative architecture for efficient resource allocation and supports computation offloading among devices and to the edge server. The scheduling problem is formulated as a nonlinear integer program with delay constraints to maximize average tracking accuracy across cameras. A meta-reinforcement learning approach further enhances adaptability, enabling rapid policy updates with minimal interaction for online optimization in complex traffic scenarios, thereby improving both generalization and stability.
Fusing a low-resolution hyperspectral image (LR-HSI) with a high-resolution multispectral image (HR-MSI) is a widely adopted strategy for hyperspectral image super-resolution (HISR), for which diffusion models have recently shown strong potential. However, existing methods still suffer from two critical limitations: 1) spectral models based on multilayer perceptrons (MLPs) lack the necessary inductive bias for sequential data, structurally disregarding the local correlations between adjacent spectral bands and 2) spatial models struggle to simultaneously leverage the specific structures from the observed image with the generic priors learned from large-scale datasets. To overcome these challenges, we propose a hybrid-prior guided coupled diffusion (HPGC-Diff) for the unsupervised HISR task, which effectively integrates a low-rank prior, a spectral local correlation prior, and a generic prior. Leveraging the low-rank representation, we decompose the target HR-HSI into two low-dimensional components and establish separate processes for their joint reconstruction. Specifically, we design a 1-D U-Net spectral diffusion model that effectively learns the structured spectral distribution from the LR-HSI. For spatial modeling, we introduce a dual-source spatial model that integrates a generic prior from multiple pretrained diffusion models to provide a powerful generative capability, with conditional features extracted from the HR-MSI by a dedicated network to inject fine-grained structural details. Finally, spatial and spectral diffusion sampling is jointly guided and alternated with conditional feature optimization to ensure stable convergence under observational constraints. Extensive experiments on simulated and real-world datasets demonstrate that HPGC-Diff achieves superior performance compared to state-of-the-art methods.
Weakly-supervised object detection (WSOD) learns detectors with only image-level classification annotations. Without precise instance-level labels, most previous WSOD methods in remote sensing images (RSIs) select the highest-scoring proposals as the final detection results, which are confronted by two major challenges: (1) instances with small scale or rare poses are easily neglected; (2) optimizing network by the top-scoring region inevitably overlooks many valuable candidate proposals. To mitigate the above-mentioned challenges, we propose a data-driven bidirectional spatial-adaptive network (BSANet). It contains a forward-reverse spatial dropout (FRSD) module to reduce instance ambiguity induced from extreme scales and poses, as well as crowded scene, and to better excavate the entire instances. From attention learning perspective, the proposed FRSD is conceptually similar to a data-driven hard attention mechanism, which adaptively samples and reconstructs the spatially related regions for mining more latent feature responses. Meanwhile, our FRSD effectively alleviates the inherent problem that non-parametric hard attention learning fashion cannot adapt to different datasets. In addition, we build a soft attention branch to simultaneously model soft pixel-level and hard region-level attention information for exploring the complementary benefit between soft and hard attention learning. We evaluate our BSANet on the challenging NWPU VHR-10.v2 and DIOR datasets. Experimental results demonstrate that our method sets a new state-of-the-art.
Recent researches perform aerial object detection from particular perspectives, e.g., small resolution and arbitrary orientations of instances. However, objects with large aspect ratios are actually very common in aerial scenarios, and recent methods seldom carry out their work on the aforementioned topic. In this article, we propose a large aspect ratio object-related method, named see hidden insight from transposition (SHIFT), to elevate the performance of aerial object detection. First, we regard the common convolutional operation as the forward-view feature modeling process, which generates the initially comprehensive descriptors for most aerial objects. This step serves as the foundation for capturing the basic features of different objects. By transposing image features into top- and left-view, our SHIFT can better capture the overall shape and spatial structure of large aspect ratio objects, which may be partially missed or inaccurately represented in the forward-view. Obviously, the top- and left-view features are all strip-like maps, which implicitly express the height and width of objects in the image, rather than the texture and appearance. In order to bridge the representational gaps among the aforementioned multiview features, we further construct an adaptive feature fusion (AFF) module to fuse them in an intelligent way, which models the implicit relationships among these descriptors through the soft attention weights. Our AFF module can be treated as a plug-and-play component, which is end-to-end trainable in the whole detection network. Experimental results on DOTA-v1.0, DOTA-v1.5, and DIOR-R datasets demonstrate that our method achieves superior performance than many recent object detection approaches by a significant margin, including various large aspect ratio related methods.
Weakly misalignment severely degrades multimodal remote sensing object detection, particularly when the displacement is spatially non-uniform and small objects are involved. To address these issues, we propose a Region-based Progressive Alignment Network (RPANet) through regional explicitimplicit alignment, realizing reliable local offset estimation and offset-guided implicit alignment. Specifically, we propose a Redundancy-aware Regional Offset Modeling Module (R2OMM) to estimate reliable regional offsets, where the Regional Spatial Modulation (RSM) enhances informative spatial structures within local windows, and the Channel-recalibrated Offset Modeling (COM) suppresses redundant responses while predicting explicit regional offsets for coarse geometric correction. Additionally, we design a Dual-Path Synergistic Implicit Alignment Module (DPSIAM) to perform offset-guided implicit alignment by jointly modeling intra-modal refinement and inter-modal interaction. Experiments on the DroneVehicle, VEDAI, DVTOD and LLVIP datasets verify the effectiveness of RPANet and show that it achieves state-of-the-art detection performance, especially under weakly misalignment scenarios.
The perpetual challenge in hyperspectral image (HSI) classification lies in the scarcity of labeled samples. Through training the model with source scene data and corresponding labels, and subsequently applying this learning to generalize across target domain data, the cross-scene HSI classification yields favorable outcomes. A novel approach to tackle this issue involves leveraging text information regarding prior knowledge of land cover classes to guide the learning of image features, thereby enhancing the generalization capability of the model. However, the finiteness of both image and text information poses difficulties in improving the training accuracy of the model. To address these limitations, this paper proposes a Cross-Modal Generative Network (CMGnet). This network aims to expand available samples by generating diverse samples of images and text, thus mitigating the challenge of inadequate training samples in cross-scene scenarios with limited data. The method employs two branches, one for images and the other for text, to conduct sample diversity generation. To ensure the accuracy of the generated samples, the label consistency constraints are incorporated. The image encoder is trained by the generated images alongside the source domain images, and the generated text features together with the original text features form a cross-domain shared space. In this space, text-image alignment is achieved through supervised contrast learning, thereby facilitating the learning of domain-invariant features guided by linguistic modal. Experimental results on three datasets demonstrate the superiority of CMGnet in domain generalization (DG). The codes will be available from the website: https://github.com/wbq330/CMGnet.
Cross-domain few-shot learning aims to recognize target-domain samples with only a few labeled examples, despite distribution discrepancies from the source domain. While recent methods achieve promising results on natural-image benchmarks, their performance often degrades in remote-sensing scenarios, where domain discrepancies are more pronounced. We identify two key challenges underlying this issue: (1) existing methods often lack sufficiently strong domain-invariant semantic priors under severe remote-sensing domain shifts, and (2) prototype estimation and decision consistency are easily degraded by limited or noisy support samples. To address these challenges, we propose Outer–Inner dual-path Co-Distillation (OICD) for cross-domain few-shot remote sensing classification. The outer path introduces DINO-Guided Patch Alignment, which leverages a frozen DINOv2 teacher to inject domain-invariant semantic structures through patch-level alignment. The inner path performs task-specific refinement by combining Adaptive Instance Reweighting to stabilize prototypes with Cross Multi-View Consistency to impose robust local–global relational constraints. These two paths jointly enhance both the stability and discriminability of the learned representations under a unified optimization objective. Experiments on five cross-domain remote-sensing tasks show that OICD achieves strong overall performance compared with existing methods.
Knowledge distillation provides an effective paradigm for developing lightweight remote sensing object detectors. However, existing methods predominantly concentrate on the design of foreground feature masks while neglecting the suppression of small-scale targets by dominant large-scale counterparts in cluttered backgrounds during the distillation, which severely degrades detection robustness. To address this, we propose scale-conscious knowledge distillation. Multiscale feature distillation (MSFD) in this framework decouples the coupled features in conventional methods into hierarchical multiscale representation spaces, enabling the student model to capture cross-granularity features from micro-details to macro-structures through differentiated receptive fields. Scale-adaptive output distillation (SAOD) innovatively introduces a dynamic weighting mechanism based on the normalized object area, effectively alleviating the gradient vanishing issue for small objects. Experimental results show that our method consistently improves the performance of student models with different architectures on both the DOTA and DIOR datasets. Source codes are available at https://github.com/RQ-W/SCD.git
Fusing the spectral information of low-resolution hyperspectral image (LR-HSI) with the spatial details of high-resolution multispectral image (HR-MSI) is an effective strategy for reconstructing high-resolution hyperspectral image (HR-HSI). Existing methods perform well when the degradation characteristics of training and test data are consistent. However, their generalization ability often deteriorates when faced with unknown degradations. In this article, we treat the degradation conditions as confounders in the causal relationship, and introduce diverse degradation schemes to mitigate the spurious correlations of the fusion process, compelling the model to learn latent causal patterns independent of specific degradation mechanisms. For this, we propose a degradation-guided blind hyperspectral image fusion network with spatial-frequency attention to better reveal these underlying causal characteristics. Specifically, two U-shaped subnetworks adaptively extract multiscale spatial and spectral degradation features from LR-HSI and HR-MSI, which are then embedded as guidance features into the reconstruction network for more accurate HR-HSI reconstruction. In addition, we design a spatial-frequency attention module, which integrates spatial window self-attention and frequency-domain attention in a parallel and interactive manner to fully exploit the complementary characteristics of the spatial and frequency domains. Extensive experiments on multiple public datasets demonstrate the effectiveness and superior generalization of the proposed method. The code of the proposed approach is available on https://github.com/liuofficial/CDGN
Hyperspectral image plays an indispensable role in the field of change detection, yet its application still faces numerous challenges. On one hand, traditional attention mechanisms are often constructed based on local information, making them prone to overlooking long-range contextual relationships hidden within global information. On the other hand, fine-grained features are susceptible to interference from irrelevant changes. To address these issues, this paper proposes a network framework named Global Spatial-spectral and Frequency-domain Mamba (GSSFDM) to enhance the accuracy and robustness of change detection in hyperspectral images. The proposed framework comprises three core modules. Global spectral-spatial attention module constructs a global spectral-spatial self-attention mechanism based on global information, effectively capturing long-range correlations between the spatial and spectral dimensions, thereby enhancing the model's ability to understand global information. Fine-grained change feature extraction module utilizes flow field estimation techniques to filter out irrelevant change features, building prior knowledge to enhance the model's ability to distinguish changes. Concurrently, it combines Fourier transform to separate phase detail changes, further enhancing the representation of fine-grained features, which aids in more precise localization of change boundaries. Local enhancement Mamba module combines the Mamba architecture with an effective channel attention mechanism, specifically designed to strengthen the capture and representation of fine-grained change features, significantly improving the model's performance in perceiving and recognizing subtle changes. Extensive experiments on three benchmark datasets demonstrate that the proposed GSSFDM consistently outperforms existing state-of-the-art methods.
Hyperspectral image (HSI) super-resolution plays a crucial role in enhancing the spatial resolution of HSIs while preserving spectral integrity. A common approach for HSI super-resolution involves fusing low-resolution HSIs (LR-HSIs) with high-resolution multispectral images (HR-MSIs). In this article, we propose a novel spectral degradation guided Diffusion Schrodinger Bridge (SDG-DSB) model specifically designed for HSI super-resolution. Recognizing that LR-HSIs contain complete spectral information, we leverage a subspace model to decompose LR-HSIs into endmembers and LR abundance maps. To reconstruct the HR-HSI, we propose a DSB, transforming LR abundance maps into HR counterparts. This process is supported by multiple pretrained DSB generators trained on natural images, integrating spatial priors for improved accuracy. Additionally, the HR-MSI guides the diffusion process, maintaining consistency with the spectral degradation path from HR-HSI to HR-MSI. This spectral degradation captures the inherent correlations between spectral bands, allowing for a more accurate restoration of high-frequency spatial details while maintaining the integrity of spectral information. Finally, the HR abundance maps are combined with endmembers extracted from LR-HSI to generate the HR-HSI. Extensive experiments on benchmark HSI datasets show that the proposed SDG-DSB outperforms state-of-the-art methods, demonstrating its effectiveness in both spatial and spectral reconstruction.
Compositional Zero-Shot Learning (CZSL) aims to recognize novel compositions of objects and states by transferring knowledge from seen compositions. A critical omission in prior studies is the uniform penalty imposed on all incorrect compositions, ignoring their inherent affinities with the ground-truth labels. This oversight leads to severe overfitting on seen classes and impedes the discovery of genuine visual-semantic relationships. To address this, we propose Clique-based Interclass Affinity (CIA), a framework that introduces hierarchical semantic supervision by grouping compositions into affinity cliques. CIA encodes both semantic affinity and visual affinity to construct multi-level cliques. These cliques guide a one-to-many alignment between visual and semantic features, enabling the model to learn generalizable class prototypes through structured constraints, rather than treating all incorrect classes equally. Unlike prior works focusing on direct classification, CIA emphasizes unveiling intrinsic compositional structures by analyzing inter-semantic and visual relationships. Extensive experiments on MIT-States, UT-Zappos, and C-GQA demonstrate CIA’s superiority, showcasing its robustness in both closed-world and open-world settings. Our code is available at https://github.com/LanchJL/CIA-CZSL.