Polarimetric synthetic aperture radar (PolSAR), as an advanced remote sensing technology, can provide rich polarization information and is of great significance for target recognition in complex environments. However, deep learning methods based on neural networks often fall into the dilemma of convergence or over-fitting due to insufficient samples and weak physical perception. To address this limitation, we propose a PolSAR vehicle target recognition method via scattering mechanism-driven hybrid attention. Firstly, we introduce a physics-inspired neural module, termed the Polarization Knowledge Extraction Module (PKEM), which retrieves physical scattering information from the Discrete Null-Pol Synthesis Pattern (DNSP). We then propose a Polarization Guided Feature Enhancement (PGFE) module, which dynamically re-weights deep features according to the physical prior information, thereby enhancing the model’s physical perception capabilities. Finally, we constructed a PolSAR vehicle target recognition dataset that can significantly facilitate the advancement of PolSAR ATR. Extensive experimental results on our self-collected measured dataset and Gotcha dataset confirm the superior effectiveness of the proposed physics-inspired target recognition framework. Our data will be available at https://github.com/DJJJJJJ/SMG_PolSAR.
Restricted by view sensitivity and limited data volume, current polarimetric SAR vehicle target recognition still faces bottlenecks including inadequate characterization of scattering properties and poor generalization performance. To address these limitations, this paper proposes a rotation-domain similarity feature (RSF) to precisely characterize the angular dependence of target scattering characteristics, and introduces an RSF-Span Overview-Focusing Network (RSOFNet). The RSF matches the target coherency matrix with canonical scattering templates across the continuous rotation domain, yielding angle-dependent responses with explicit physical meaning. A Span branch extracts sample-level scene context from the total-power image to guide RSF feature learning in a sample-adaptive, top-down manner. A physics-constrained fusion layer integrates RSF-derived physical statistical priors, Span-modulated mechanism-discriminative weights, and context-conditioned template reliability weights, ensuring physically interpretable feature embeddings. Experiments on the ATR-HAPV and GOTCHA datasets indicate that the proposed method achieves performance superior to data-driven, single-feature, and multi-feature fusion baselines under standard, cross-depression-angle, and reduced-training-data conditions, with the advantage becoming more pronounced as sample size decreases.
Deep learning-based automatic target recognition is a research hotspot in the field of synthetic aperture radar image interpretation, and convolutional neural networks (CNNs) have made significant advancements in this area. However, CNNs suffer from feature confusion during feature extraction, which neglects the affiliation relationships between features. In addition, these methods hardly apply the attribute features of targets, leading to a lack of physical interpretation. To incorporate interpretable features containing attribute features into the recognition framework, we propose a multilevel capsule fusion network that integrates global and local attribute features. The network is comprised of a multilevel feature extraction module and an attribute attention reconstruction module (AARM). The MLFM hierarchical fusion capsule is utilized to extract multidimensional attribute features of targets, enhancing the physical interpretability of the. The AARM module employs a two-dimensional attribute attention mechanism to emphasize the distinctions between target attribute parameters, improving the separability of similar targets. Furthermore, we demonstrate the rationality of capsule form attribute features by applying perturbations to the reconstructed images which are obtained after the AARM module, encompassing attributes such as amplitude, azimuth, depression angle, and shape. Finally, we design a corresponding attribute loss function by measuring the L2 distance between various attribute parameters, achieving more precise target recognition. Experimental results indicate that the proposed method is superior to the state-of-the-art methods in recognition performance on both self-constructed measured datasets and the MSTAR dataset, demonstrating excellent generalization capabilities.
Synthetic aperture radar (SAR) is an effective imaging and observation sensor that has been widely applied in both military and civilian fields. Deep learning approaches have gained prominence in SAR target recognition and received extensive attention. However, these methods often struggle when data is scarce, leading to insufficient training and challenges in effective feature extraction. To address this limitation, we propose a less data-dependent feature extraction framework. Specifically, we introduce the dual-tree complex wavelet transform (DTCWT) to capture multi-frequency feature details of SAR images, integrated with convolutional neural network. This approach enables effective extraction of high- and low-frequency information. By leveraging the characteristics of these frequency features, low-frequency subbands are used to emphasize the global structural features in the images, and high-frequency subbands are employed to identify the significance of different regions in the images. In response to the aforementioned characteristics, we introduced an attention mechanism to effectively incorporate high-frequency local information into low-frequency global information, thereby enhancing feature representation and recognition efficiency. Moreover, we propose an adaptive rotational convolution, and apply it to the high-frequency feature extraction. The adaptive rotational convolution can adapt to the directionally selective subbands with a single convolution kernel. Experiments conducted on the MSTAR and SAR car datasets demonstrate that the proposed method can achieve better recognition performance with fewer parameters, especially on small-scale datasets. The ablation study also confirms the effectiveness of the introduced DTCWT and rotational convolution.
The accurate detection of unmanned aerial vehicles (UAVs) in various sizes played an important role in the practical applications. Yet the preceding works suffered from the missing inference, the false alarms, and the poor accuracy due to the the adverse scene conditions, as well as the mutable scales. To solve the problems, a hierarchical attention promoted cross-scale learning framework was proposed in this paper. First, the hierarchical attention mechanism was introduced in the backbone to generate the multi-scale features of targets, so they can be discerned and located at different scales. The resulting features were further delivered to the neck, in which two branches of features were built, respectively. The former was obtained by the target-specific feature operator, while the latter was generated by the upsampling operation. The dual branches were further connected in the quasi-residual structure. So the content of targets can be protected well, and the detail information can be reconstructed. Finally, the dynamic focusing loss measurement was presented to regress the bounding box of the target, so the learning effectiveness of presented the architecture can be promoted. To verify the proposed method, multiple rounds of experiments were performed. The results demonstrated that small and weak drones can be detected accurately, especially in adverse lighting and weather conditions. The evaluation metric of mean average precision rate (mAP) can be improved by 18.5% (YOLO6) on the collected dataset.
The distinctive electromagnetic scattering characteristics of SAR target images are crucial for SAR target recognition. However, existing purely data-driven paradigms often overlook inherent physical scattering laws, which limits model robustness and interpretability. To address this issue, we propose a physics-aware interpretable framework, termed the attribute scattering center (ASC)-guided deformable capsule network. The proposed network adopts a two-stage learning framework. In the first stage, an ASC-guided capsule network is constructed to extract ASC capsules with explicit electromagnetic scattering semantics. Within the capsule architecture, we design a multiscale densely connected deformable module to adaptively capture target features at different scales. In addition, we design a polarity attribute feature perception module to enhance the exploitation of target attribute information. In the second stage, ASC capsules are embedded into deep attribute capsules to obtain a fused representation of physical scattering features and deep semantic features, thereby achieving reliable SAR target recognition. Furthermore, to further advance research on SAR target recognition, we construct and publicly release a new high-resolution benchmark dataset, termed ATR-SARVehicle-1.0 (https://doi.org/10.57760/sciencedb.j00001.01475), which covers multiple frequency bands and terrain backgrounds. Extensive experiments conducted on both the proposed benchmark and the MSTAR dataset demonstrate that the proposed method outperforms existing approaches in terms of recognition accuracy and interpretability.
In the past few years, polarimetric synthetic aperture radar (PolSAR) as an advanced technology has been widely exploited to Earth observation, among which ship detection is an active research topic. Taking the sub-look decomposition technology as the basis, this article proposes a new ship detection method, abbreviated to amplitude-based ship detection metric (ASM). In brief, two single-look complex (SLC) images I(1)and I(2 )are first obtained from the original PolSAR image O for forming the data group { I-1 , O , and I-2 }. Then, the H/A/alpha decomposition is performed on { I-1, O , and I-2} so as to yield the H/alpha plane group { P-1 , P-0 , and P-2 }, which is subsequently used to suppress sea clutter and generate another filtered data group { F-1, F-0, and F-2} that, respectively, corresponds to I-1, O , and I-2 . Thereafter, a new 3x3 spatial-spectral coherence difference matrix [ ST ] is further constructed by { F-1 , F-0 , and F-2 }, wherein the spatial information and spectral information are simultaneously used. Therefore, [ ST ] can effectively highlight ships from sea clutter. To verify this point, an ASM is finally built by multiplying the amplitude values of the terms ST13 , ST23 , and ST33 together. Extensive experiments demonstrate that: 1) [ ST ] is more appropriate for ship detection than [ T ] and 2) ASM has a better ability to improve ships' target-to-clutter ratio (TCR) in comparison with other state-of-the-art (SOTA) methods. Last but not least, experimental results also show that the detection performance of ASMR (i.e., calculating ASM along the range direction) is similar to that of ASMA (i.e., calculating ASM along the azimuth direction). For example, the average TCR values of ASMA and ASMR are, respectively, 7.5 and 7.06 dB higher than the second place. Therefore, this means that the frequency characteristics of ships should be considered in the practical detection process as well.
Training of synthetic aperture radar (SAR) target detection and recognition methods based on deep learning heavily relies on a large amount of data. As one of the significant approaches to address the scarcity of SAR data, SAR image intelligent generation methods have witnessed rapid development. However, these methods often require many data samples for learning and are prone to deviating from the physical scattering characteristics. To address these issues, this letter proposes a two-stage SAR image generation method based on attribute feature decoupling within a generative adversarial network (GAN) architecture. In the first stage, the original SAR target image undergoes feature extraction and reconstruction, yielding generated images highly similar to real images. The attribute features decoupled during this process correlate with the scattering characteristics of SAR target, providing guiding information for generating target images in the second stage. In the second stage, by applying perturbations to specific dimensions of the decoupled features, we can reconstruct target images with altered attributes, achieving diverse data augmentation. Multitask discrimination based on pixel intensity, authenticity, and feature distance differences enhances the quality of generated images across multiple levels. The decoupled representation-driven generation paradigm simplifies the network's mapping learning task through task decomposition, diminishing the dependency on the volume of data. The experimental results demonstrate that the generated images possess higher quality and superior application performance, with an improvement of 5.23% in recognition accuracy.
Optimal transport (OT) has recently been shown as a promising criterion for unsupervised restoration when no explicit prior model is available. Despite its theoretical appeal, OT still significantly falls short of supervised methods on challenging tasks such as super-resolution, deraining, and dehazing. In this paper, we propose a sparsity-aware optimal transport (SOT) framework to bridge this gap by leveraging a key observation: the degradations in these tasks exhibit distinct sparsity in the frequency domain. Incorporating this sparsity prior into OT can significantly reduce the ambiguity of the inverse mapping for restoration and substantially boost performance. We provide analysis to show exploiting degradation sparsity benefits unsupervised restoration learning. Extensive experiments on real-world super-resolution, deraining, and dehazing demonstrate that SOT offers notable performance gains over standard OT, while achieving superior perceptual quality compared to existing supervised and unsupervised methods. In particular, SOT consistently outperforms existing unsupervised methods across all three tasks and narrows the performance gap to supervised counterparts.
Object 6D pose estimation, a crucial task for robotics and augmented reality applications, becomes particularly challenging when dealing with novel objects whose 3D models are not readily available. To reduce dependency on 3D models, recent studies have explored one-reference-based pose estimation, which requires only a single reference view instead of a complete 3D model. However, existing methods that rely on real-valued coordinate regression suffer from limited global consistency due to the local nature of convolutional architectures and face challenges in symmetric or occluded scenarios owing to a lack of uncertainty modeling. We present CoordAR, a novel autoregressive framework for one-reference 6D pose estimation of unseen objects. CoordAR formulates 3D-3D correspondences between the reference and query views as a map of discrete tokens, which is obtained in an autoregressive and probabilistic manner. To enable accurate correspondence regression, CoordAR introduces 1) a novel coordinate map tokenization that enables probabilistic prediction over discretized 3D space; 2) a modality-decoupled encoding strategy that separately encodes RGB appearance and coordinate cues; and 3) an autoregressive transformer decoder conditioned on both position-aligned query features and the partially generated token sequence. With these novel mechanisms, CoordAR significantly outperforms existing methods on multiple benchmarks and demonstrates strong robustness to symmetry, occlusion, and other challenges in real-world tests.
The divertor target plate inside the Experimental Advanced Superconducting Tokamak (EAST) device is a core component for device operation, and its surface condition directly affects the operational stability and safety of the device. Accurate online detection of surface defects on the target plate is of great significance for long-term reliable operation of the device. To this end, this paper constructs the first EAST divertor target plate surface defect dataset annotated with Oriented Bounding Boxes (OBB), and proposes an improved YOLOv11-OBB algorithm to achieve efficient and real-time defect detection, supporting the autonomous maintenance of the target plate. Aiming at the specific detection scenario of divertor target plate defects, a Multi-scale local feature enhancement and contrast module (MSC) is introduced into the backbone network to improve the network's ability to focus on defect features; the deformable convolution is used to replace the standard convolution module in the original C3K2 module, which enhances the performance of network learning and adapts to irregular defect feature boundaries; an Adaptive weighted feature fusion module (ACF) with weight learning is adopted in the neck network to optimize the fusion efficiency of multi-scale features. The results of ablation experiments show that the performance of the improved YOLOv11-OBB model is significantly better than that of the baseline model; the comparative experiments with other mainstream models further verify the effectiveness and real-time performance of the improved model.
Ship structure recognition using polarimetric information is of great significance for maritime remote sensing. However, accurate and reasonable ship structure recognition remains challenging, primarily due to the insufficiency of information exploration in target characterization space, and the lack of intelligent recognizer based on polarimetric-physical coupling. To address these limitations, we propose a novel Polarimetric Second-Moment Response (PSMR) and a polarimetric prior-based structure recognition framework. First, a three-dimensional polarization state vector (Q Vector) is derived, establishing a fundamental base for electromagnetic wave polarization state in complex space that simultaneously describes polarization orientation and ellipse angles and phase information. Second, for the first time, the PSMR concept is proposed, which provides a benchmark for target characterization. The PSMR spectrums are proposed to represents the polarimetric response, which reflects the polarimetric characteristics of each structure. And a designed average polarization similarity measure (APSM) quantifies the difference between the target and canonical structures by comparing their spectrums. Third, a polarimetric prior-based structure recognition framework PSAResNet is constructed, in which a polarimetric similarity attention (PSA) branch module and a similarity approximation loss function (SAL) are incorporated based on the APSM. The improved framework enhances the fusion of the neural network with polarimetric priors, enabling superior and more generalized performance of the structure recognition under measured environment. Experimental results on three measured PolSAR data of partial-coherent and distributed targets demonstrate that the proposed method achieves precise and physically consistent ship structure recognition. Furthermore, this study discusses the correlation between structure recognition and characteristics represented by model-based polarimetric decomposition.
In recent years, the use of UAVs has become increasingly widespread, and the public safety risks posed by unauthorized UAV flights have become increasingly prominent, creating an urgent need for effective detection and identification of UAV targets. However, such targets are small in size, have low contrast, and exhibit an extremely low signal-to-noise ratio; conventional detection methods generally suffer from insufficient feature discrimination, missed detections, and false alarms in complex backgrounds. To address these challenges, this paper proposes a Frequency-Guided Cross-scale Refinement Network (FGCR-Net). Based on an encoder-decoder architecture, this network achieves end-to-end collaborative optimization through cross-layer feature fusion, side-channel prediction refinement, and frequency-domain background suppression. First, a multi-path selective cross-layer fusion module (SCFM) is designed. This module employs coordinated modeling via both channel and spatial paths, supplemented by adaptive weighting with learnable coefficients, to perform differentiated selective fusion of the encoder’s fine-grained features and the decoder’s semantic features, thereby bridging the semantic gap at jump connections; Second, we designed a Cross-Scale Adaptive Fusion Enhancement Attention Module (CAFEM), which cascades multi-receptive-field hollow convolutions, strip pooling, and a bidirectional semantic guidance mechanism to perform cross-scale refinement on the side outputs of each decoder layer, thereby alleviating the issues of blurred boundaries and false alarms caused by inconsistent quality of multi-scale prediction maps and insufficient cross-layer consistency; finally, we design a Frequency-Guided Semantic Enhancement Module (FGSEM), which uses the Fast Fourier Transform (FFT) to decouple encoder features into the frequency domain. By leveraging low-frequency energy to predict the background confidence map and applying spatially selective suppression to high-frequency components, this module distinguishes, from a frequency-domain perspective, the high-frequency responses of complex backgrounds and targets that are highly similar in the spatial domain. Experiments on MSDS-UAV, a self-built multi-scenario UAV dataset for small targets, demonstrate that our method consistently outperforms existing state-of-the-art methods across multiple performance metrics, with Pixel Accuracy, Mean Intersection over Union, and Probability of Detection reaching 92.76%, 70.91%, and 92.69%, respectively; Compared to the baseline model, these three metrics improved by 1.90, 3.20, and 3.76 percentage points, respectively, fully validating the effectiveness and superiority of the proposed method.
Polarimetric synthetic aperture radar (PolSAR) provides physically interpretable observations through polarization-dependent scattering responses. Developing reliable PolSAR interpretation methods is often constrained by data scarcity due to limited coverage of acquisition conditions and scene states. Generative models can mitigate this scarcity, but most treat PolSAR as generic multichannel imagery and neglect interchannel polarimetric correlations that encode scattering characteristics. Consequently, generated samples may be visually plausible yet exhibit limited polarimetric fidelity, reducing their utility for PolSAR interpretation. To address this issue, we propose a scattering-aware diffusion model (SADM) that enforces scattering consistency to improve polarimetric fidelity. Specifically, we devise a scattering-decoupling strategy within SADM that learns representations for each polarimetric channel independently to preserve physical meaning and captures scattering mechanisms via interchannel correlations among these representations. This strategy then employs the resulting scattering descriptors to modulate the spatial features in the main generation pathway, endowing them with explicit scattering semantics. Furthermore, we incorporate a scattering mechanism consistency loss to facilitate learning the joint distribution of PolSAR data and polarimetric features, enhancing polarimetric fidelity beyond pixel-level appearance. Experiments on a publicly available Ku-band full-polarimetric 10-class civilian vehicle dataset (KuPolVehicle-10) show that SADM achieves a Yamaguchi-based polarimetric Fr & eacute;chet inception distance (Pol-FID) of 3.506 versus 5.245 for a standard denoising diffusion probabilistic model (DDPM) baseline. Moreover, SADM-based augmentation improves classification accuracy by 4.94%, indicating that enforcing polarimetric fidelity translates to improved utility in PolSAR interpretation. Additional evaluations on OpenSARShip 2.0 confirm that SADM generalizes well across polarization modes, frequency bands, and target classes.
In synthetic aperture radar (SAR) ship detection, distribution shifts between train and test data arise from sensor bands, resolutions, and imaging principles. These shifts distort domain-invariant features and slow down the convergence of the detection network, ultimately leading to severe performance degradation in cross-sensor scenarios. To address this, we propose a physics-driven domain-adaptive SAR ship detection method termed DDA-SSD. Firstly, Physical Scatter Adaptive Gradient Reversal Layer (PSA-GRL) is proposed to focus domain alignment on target scattering centers while suppressing background interference. Then, we propose the Domain-Aware Detail Enhance Module (DADEM) to improve domain classification accuracy through adaptive channel-group gating. Finally, we construct the Scattering-aware Dynamic Fusion Feature Pyramid Network (SDF-FPN) to reconstruct robust, domain-invariant target signatures via scattering-guided upsampling. Evaluated experimental results on six cross-domain tasks (S2H, S2A, H2S, H2A, A2S, A2H) from three datasets (SSDD, HRSID, AIRSAR-SHIP), DDA-SSD achieves AP of 70.4%, 53.6%, 72.1%, 50.7%, 67.3%, and 53.6%, respectively, demonstrating significant superiority over existing domain-adaptive detection methods. The source code is publicly available at https://github.com/GY-monkey/DDA-SSD.
This paper presents an experimental framework for the recognition and localization of vehicle and personnel targets, involving the production of datasets and the experimental validation of the framework. The approach utilizes the short-time Fourier transform (STFT) to extract feature information from acoustic vibration signals. A residual neural network is then applied to recognize these features, addressing gradient vanishing and explosion problems. For localization, a six-element acoustic array signal is processed using the Multiple Signal Classification Algorithm (MUSIC). The framework commences with the extraction of feature information from sound and vibration signals collected by acoustic shock sensors via STFT. Subsequently, a target recognition algorithm employing residual neural networks is proposed. For localization, the method employs MUSIC to divide the signal and noise spaces by constructing a six-element array signal receiving matrix and performing eigenvalue computation, followed by spectral peak search to obtain target localization. Experimental trials validate the recognition and localization of moving targets. The overall recognition rate, taking environmental factors into account, is 96.80
Deep learning methods have been widely applied in the field of synthetic aperture radar (SAR) target recognition. However, these methods typically rely only on the amplitude information of SAR images, without exploiting the electromagnetic scattering (ES) characteristics of targets. This limitation reduces the reliability and robustness of target recognition. To address this issue, we propose a component scattering feature fusion network (CSFF-Net) that integrates component-level scattering features with image features. Initially, the attribute scattering center (ASC) model is employed to extract scattering center parameters of the target. Based on the partition and reconstruction of attribute parameters, more interpretable global and component scattering features are obtained. Subsequently, a Gabor hierarchical dense connection (GHDC) module is proposed to capture multi-scale hierarchical features from SAR images. Finally, a scattering feature fusion (SFF) module is designed to achieve effective fusion of scattering features and image features, enhancing the discriminability and robustness of the learned features. Extensive experiments are conducted on a measured dataset containing 10 types of vehicles. The results demonstrate the superior performance of the proposed method.
The existing synthetic aperture radar (SAR) ship detectors perform well on data with consistent distributions but degrade significantly when faced with domain shifts and the absence of labeled data. Moreover, the traditional convolutional neural networks (CNNs) struggle with global feature extraction due to the local receptive fields, while transformer approaches struggle with computational efficiency when extracting global features from complex SAR images. Designing an effective cross-domain SAR ship detector that can handle unlabeled data with domain shifts remains a challenge. In this letter, we propose a novel Mamba-based unsupervised domain adaptation (UDA) SAR ship detection model integrated with pseudolabels optimization strategy. First, we propose the domain adaptive state-space module (DASSM) to construct the Mamba mean teacher (MMT) framework for the first time, enhancing the capture of both global and local SAR image features at a linear time complexity and facilitating domain-invariant feature learning. To enhance the quality of pseudolabels, we design the adaptive pseudolabel optimizer (APLO) module with wise-IoU (WIoU) and dynamic dual-threshold pseudolabel selector (DDPLS). The WIoU is utilized to improve the generation of pseudolabels, while DDPLS is further employed to categorize and optimize pseudolabels. Extensive experiments on public datasets illustrate the effectiveness and superiority of the proposed method for cross-domain detection of unlabeled SAR data.
In recent years, convolutional neural networks (CNNs) have been extensively utilized for synthetic aperture radar (SAR) ship detection tasks. The fixed square shape of convolutional kernels in traditional convolutional limits the ability to extract features. Moreover, the large number of parameters in CNNs restricts their deployment on platforms with limited resources. To address these challenges, this article proposes a novel SAR ship detection network called DFES-Net, which incorporates deformable feature fusion and accurate anchor prediction to enhance detection performance. DFES-Net employs depthwise separable deformable convolution (DWDCN) and a lightweight feature enhancement module to improve detection accuracy in complex scenes. Specifically, DWDCN is integrated into the backbone network to adapt convolutional positions to the shape of ship targets, thereby enhancing feature extraction and improving detection accuracy in complex scenarios. In addition, a receptive field enhancement module based on dilated convolutions is introduced to increase the effective receptive field and improve detection performance for nearshore and densely packed small targets. An effective regression loss function, CADIoU, is proposed to generate accurate bounding boxes for SAR ship targets. Finally, a magnitude-based dynamic hierarchical pruning algorithm is introduced to dynamically prune parameters across various network structures, thereby reducing the model's parameters. Extensive experiments conducted on the SSDD and HRSID datasets demonstrate that our method achieves a 2.0x speed up with a model memory of only 4.7 MB (73.4% smaller than YOLOv7-tiny), while attaining mAP scores of 97.9% and 92.2% on the SSDD and HRSID datasets, respectively.
Traditional time-frequency analysis methods often fall short in analyzing non-stationary signals characterized by wide frequency spans and complex variations. To slove this problem, this paper advances a modified S-transform algorithm based on an optimized gaussian window function. We achieve enhanced time-frequency resolution by controlling the shape of the window function through reconfiguring the scaling factor of the gaussian window. This adaptation allows the algorithm to tailor its time-frequency analysis to the characteristics of different signals, thereby preventing the window from becoming excessively narrow at low frequencies or overly broad at high frequencies. As a result, a balanced resolution is maintained across the entire time-frequency plane, significantly improving the overall energy concentration. Finally, simulation experiments verify the effectiveness of the proposed algorithm and its applicability to various types of signals.
Zhaowen Zhuang (庄钊文)合作论文数National University of Defense Technology3