In GNSS-denied underwater environments, individual unmanned underwater vehicles (UUVs) suffer from unbounded dead-reckoning drift, making collaborative navigation crucial for accurate state estimation. However, the severe communication delay inherent in underwater acoustic channels poses serious challenges to real-time state estimation. Traditional filters, such as Extended Kalman Filters (EKF) or Unscented Kalman Filters (UKF), usually block the main control loop while waiting for delayed data, or completely discard Out-of-Sequence Measurements (OOSM), resulting in serious drift. To address this, we propose an Asynchronous Two-Speed Kalman Filter (TSKF) enhanced by a novel projection mechanism, which we term Variational History Distillation (VHD). The proposed architecture decouples the estimation process into two parallel threads: a fast-rate thread that utilizes Gaussian Process (GP) compensated dead reckoning to guarantee high-frequency real-time control, and a slow-rate thread dedicated to processing asynchronously delayed collaborative information. By introducing a finite-length State Buffer, the algorithm applies delayed measurements (t-T) to their corresponding historical states, and utilizes a VHD-based projection to fast-forward the correction to the current time without computationally heavy recalculations. Simulation results demonstrate that the proposed TSKF maintains trajectory Root Mean Square Error (RMSE) comparable to computationally intensive batch-optimization methods under severe delays (up to 30 s). Executing in sub-millisecond time, it significantly outperforms standard EKF/UKF. The results demonstrate an effective control, communication, and computing (3C) co-design that significantly enhances the resilience of autonomous marine automation systems.
Super-resolution (SR) and despeckling for Synthetic Aperture Radar (SAR) images are critical tasks. However, these tasks are challenging due to speckle noise and low resolution. Speckle noise, caused by the coherent imaging mechanism, severely degrades image quality, while low resolution limits the preservation of structural details and textures. Existing methods often fail to effectively balance noise suppression and texture reconstruction, especially when addressing both tasks simultaneously. To tackle these challenges, this article proposes an unsupervised framework that combines training for region-specific diffusion model (RSDM) and latent space integration for reconstruction (LSIR). RSDM uses low-rank adaptation to train diffusion models tailored for homogeneous and inhomogeneous regions, enabling it to capture statistical uniformity and low-frequency features in homogeneous regions, while focusing on structural complexity and high-frequency details in inhomogeneous regions. LSIR integrates these region-specific models through B & eacute;zier interpolation for latent noise and linear interpolation for functional integration, allowing simultaneous high-quality despeckling and SR. Experiments conducted on multiple SAR datasets confirm the effectiveness of the proposed framework. The results demonstrate significant improvements over state-of-the-art methods in structural detail preservation, noise reduction, and overall visual quality.
Current multimodal image fusion methods typically cannot precisely locate and fuse features of key regions, leading to artifacts in fusion results and relatively weak local feature expression, failing to ensure the integrity of image structures or target contours. To address this issue, we propose a universal fusion framework called AdverFuse. Based on the Mamba module, this framework introduces an adaptive weight mixed attention mechanism module specifically targeting fusion artifact problems. The module enhances cross-modal feature complementarity through channel branches, precisely locates target regions using spatial attention, and adjusts modal contributions based on the confidence scores output by spatial and channel attention mechanisms, making feature enhancement more aligned with the characteristics of different modal data. Additionally, we design an adversarial network feature enhancement registration module to improve the local feature expression capability of fused images. Combined with the adversarial training mechanism of the discriminator, this module adaptively balances the contribution weights of modalities such as infrared and visible light, maps features of different modalities to a unified semantic space, and extracts richer semantic features while preserving local detail information. We design a convolutional attention module to achieve more comprehensive feature interaction and reduce computational complexity. Experimental results across multiple datasets demonstrate that this method has significant advantages and outperforms SOTA methods.
Solving PDEs on changing geometries and with varying parameters poses a significant challenge in fields such as materials science, engineering, manufacturing, and design. The primary difficulty lies in the high computational cost required by the conventional solvers in order to recompute the solution whenever the geometry or parameters change. Recent neural operators have shown promising results in learning PDE operators and rapidly predicting PDE solutions. However, they still face limitations in handling either varying domain geometries or varying PDE parameters. To address these limitations, we introduce the graph-based geometry-aware neural operator (GGNO), a neural operator learning framework designed to generalize simultaneously across varying domain geometries and PDE parameters. GGNO uses a dual-graph architecture, with separate geometry and parameter graphs, to encode domain shape and parameter information and couple them through message passing and interpolation. Across Darcy flow, a 2D plate problem, and a mechanics benchmark, GGNO achieves higher accuracy than neural operator baselines and generalizes robustly to unseen domains. These results suggest that GGNO is a promising surrogate solver for PDEs defined on diverse geometries and parameter settings.
Despite its all-weather superiority over optical sensors, the dense deployment of millimeter-wave radar in automotive applications causes severe vehicle-to-vehicle interference, degrading detection integrity and introducing safety risks. However, existing interference suppression methods are often limited by over-smoothing, we propose a Diffusion-based framework for real-time Radar Interference Mitigation (DiffRIM). Its stochastic forward process models the incremental addition of interference, and the learned reverse process conducts iterative removement, thereby significantly enhancing interpretability. We further introduce a Lightweight Autoencoder (LWAE) with Mobile Encoder and Decoder (ME/MD) modules, which extracts multi-scale spatial features through point-wise and hierarchical processing. A dual-residual (DualRes) connection mitigates gradient issues, while depthwise separable convolutions (DSC) and channel attention form an efficient spatio-temporal attention (STA) mechanism to extract sparse patterns. Extensive experimental evaluations demonstrate the superior mitigation and generalization performance of our approach in both synthetic and real-world datasets. To support the community, our implementation will be made publicly accessible on https://github.com/luluisthebest/DiffRim.
Existing Synthetic Aperture Radar (SAR) image generation methods still lack reliable controllability over key imaging parameters, particularly azimuth angle, depression angle, and polarization mode. Our preliminary GeoDiff-SAR supported limited azimuth completion, but remained ineffective for large missing azimuth sectors and did not provide unified control over multiple imaging conditions. To address this problem, we propose GeoDiff-SAR II, a 3D model-guided decoupled framework for controllable SAR image generation. The proposed framework imposes controllability through physically grounded geometric-electromagnetic cues rather than image intensity alone. We introduce a Geometric-Electromagnetic Conditioning Map (GECM), a structured intermediate representation that encodes the target pose map and dominant scattering centers, thereby decoupling macroscopic geometry from microscopic scattering responses. During training, GECMs are derived from real sparse-azimuth SAR images. During inference, the same representation is rendered directly from a 3D CAD model under specified azimuth, depression angle, and polarization conditions, enabling physically consistent control across large viewpoint gaps. The imaging parameters are further converted into text conditions, while the GECM is injected through ControlNet to provide explicit spatial guidance. Combined with Low-Rank Adaptation (LoRA) on a FLUX backbone, the proposed framework unifies geometric-electromagnetic conditioning and parameter-aware generation within a single process. Experiments on simulated and real datasets demonstrate controllable generation over key SAR imaging parameters, stable generalization across large azimuth gaps, and consistent improvements in image fidelity, physical consistency, and downstream Automatic Target Recognition (ATR) performance.
The two-channel cost of Compact Polarimetric (CP) Synthetic Aperture Radar (SAR) with close performance compared to four-channel Full Polarimetric (FP) SAR, has great potential for observing and interpreting ship targets. In the existing literature of ship targets using CP SAR data, the ships are often studied as a whole target for detection purpose, and the refinement of the scattering mechanism of ships remains a challenge. CP SAR measurements depend on the transmitted wave, and there are different interpretations of the two-channel polarization ratio for the same scatterer in different CP modes. The diversity of interpretations inconveniences the use of CP data and leads to non-uniformity in target decomposition algorithms. Moreover, the existing CP target decomposition methods cannot provide a refined interpretation for the inner structure of ships. In response to the above issues, this study formalizes CP data, constructs a joint feature space, and applies iterative clustering to achieve refined ship interpretation in this paper. First, we propose a novel polarimetric formalism method that establishes a unified CP covariance matrix, ensuring a consistent interpretation of scatterer physics across different CP modes. Second, we construct a discriminative joint feature space, i.e. the R-CP/alpha(BCP) plane, and further introduce the surface-to-volume ratio for dynamic adaptive threshold adjustment. Finally, we use the constructed classification space for initialization and then use the Generalized Hybrid Polarimetric Scattering Similarity (GHPSS) based Wishart classifier to iteratively classify the CP ship target. Experiments were conducted using CP ship data simulated from GF-3 FP data, and the results show that the method proposed can extract the ship's scattering mechanisms in a refined manner and preliminarily determine the ship type based on the scattering mechanism. The results support the potential of using CP data for refined ship interpretation.
Synthetic aperture radar (SAR) image generation can mitigate data scarcity, but controllablegeneration under sparse observation angles remains difficult. Recent SAR generative studies im-prove texture realism, yet explicit geometry-aware control is still limited. This paper studiesthe focused and verifiable setting of intermediate-azimuth completion: 3D-model-derived geo-metric priors guide a diffusion model to synthesize the views missing from sparse-angle trainingdata. GeoDiff-SAR constructs a lightweight multi-bounce ray-tracing prior, encodes the result-ing point cloud, and fuses it with text conditioning while adapting Stable Diffusion 3.5 Mediumthrough low-rank adaptation. On a real four-category aircraft dataset, GeoDiff-SAR reaches anSSIM of 0.812 and azimuth consistency of 0.940, compared with 0.738 and 0.782 for the text-conditioned SD3.5 Medium baseline. The same sparse-angle protocol on five MSTAR vehicleclasses yields an SSIM of 0.878 and azimuth consistency of 0.917. These results support theconclusion that a lightweight 3D geometric prior improves viewpoint adherence for controllableSAR generation; it is intended as generation guidance rather than high-fidelity electromagneticreconstruction.
The reliable operation of Unmanned Underwater Vehicle (UUV) clusters is highly dependent on continuous acoustic communication. However, this communication method is highly susceptible to intermittent interruptions. When communication outages occur, standard state estimators such as the Unscented Kalman Filter (UKF) will be forced to make open-loop predictions. If the environment contains unmodeled dynamic factors, such as unknown ocean currents, this estimation error will grow rapidly, which may eventually lead to mission failure. To address this critical issue, this paper proposes a Variational History Distillation (VHD) approach. VHD regards trajectory prediction as an approximate Bayesian reasoning process, which links a standard motion model based on physics with a pattern extracted directly from the past trajectory of the UUV. This is achieved by synthesizing “virtual measurements” distilled from historical trajectories. Recognizing that the reliability of extrapolated historical trends degrades over extended prediction horizons, an adaptive confidence mechanism is introduced. This mechanism allows the filter to gradually reduce the trust of virtual measurements as the communication outage time is extended. Extensive Monte Carlo simulations in a high-fidelity environment demonstrate that the proposed method achieves a 91% reduction in prediction Root Mean Square Error (RMSE), reducing the error from approximately 170 m to 15 m during a 40-second communication outage. These results demonstrate that VHD can maintain robust state estimation performance even under complete communication loss.
Weakly Supervised Semantic Segmentation (WSSS) with image-level labels has gained attention for its cost-effectiveness. Most existing methods emphasize inter-class separation, often neglecting the shared semantics among related categories and lacking fine-grained discrimination. To address this, we propose Contrastive Prompt Clustering (CPC), a novel WSSS framework. CPC exploits Large Language Models (LLMs) to derive category clusters that encode intrinsic inter-class relationships, and further introduces a class-aware patch-level contrastive loss to enforce intra-class consistency and inter-class separation. This hierarchical design leverages clusters as coarse-grained semantic priors while preserving fine-grained boundaries, thereby reducing confusion among visually similar categories. Experiments on PASCAL VOC 2012 and MS COCO 2014 demonstrate that CPC surpasses existing state-of-the-art methods in WSSS.
Traditional spaceborne synthetic aperture radar (SAR) radiometric calibration predominantly depends on artificial calibrators with precisely known radar cross-sections (RCS). Nevertheless, deploying and maintaining these artificial calibrators is costly and subject to limitations. Man-made point targets (MMPTs) present a practical substitute owing to their stability and widespread presence. However, their RCS values typically stem from the historical data of a single SAR satellite, restricting their application in radiometric stability analysis. Another calibrated, system-similar SAR can supply RCS references for the uncalibrated SAR, thereby ensuring calibration reliability. However, the anisotropic RCS causes substantial calibration errors when the incidence angles vary. To tackle this problem, this paper proposes an electromagnetic (EM) simulation approach. In this approach, multi-angle RCS responses are modeled and utilized to rectify the RCS affected by incidence angle differences. Experiments, which employ offshore windmills and data from the Gaofen-3 series, indicate that the proposed method decreases the average calibration error from 2.57 dB to 1.18 dB. This reduction showcases an enhancement in cross-calibration accuracy.
Interferometric Synthetic Aperture Radar (InSAR) is a promising payload for Unmanned Aerial Vehicle (UAV) scene matching navigation due to the rich textures in interferogram images compared to SAR intensity images. However, geometric parameter estimation errors during reference interferogram image generation cause significant textural discrepancies with real-time data. Compounded by inherent non-local similarity of InSAR images, these issues render conventional matching algorithms ineffective, degrading navigation accuracy. To address these challenges, this paper proposes a Morpho-Phase feature-based InSAR image matching method to mitigate the impact of parameter errors. Firstly, a Phase-Robust Keypoint (PRK) detection method is proposed, which overcomes the impact of parameter errors on keypoint detection by introducing a compensated phase and extracting phase extrema. Secondly, a Hierarchical Morphological-Phase Descriptor (HMPD) is constructed to resolve the feature ambiguity caused by the non-local similarity of interferograms by combining morphological features with phase statistics. Experimental results based on real-world InSAR data demonstrate that the proposed matching method effectively mitigates the impact of parameter errors on InSAR image matching, enhances navigation positioning accuracy, and provides stable, high-precision positioning capabilities in practical scene matching navigation tasks.
Cross-Domain Sequential Recommendation (CDSR) plays a crucial role in modern consumer electronics and e-commerce platforms, where users interact with diverse services such as books, movies, and online retail products. These systems must accurately capture both domain-specific and cross-domain behavioral patterns to provide personalized and seamless consumer experiences. To address this challenge, we propose TEMA-LLM (Tag-Enriched Multi-Attention with Large Language Models), a practical and effective framework that integrates Large Language Models (LLMs) for semantic tag generation and enrichment. Specifically, TEMA-LLM employs LLMs to assign domain-aware prompts and generate descriptive tags from item titles and descriptions. The resulting tag embeddings are fused with item identifiers as well as textual and visual features to construct enhanced item representations. A Tag-Enriched Multi-Attention mechanism is then introduced to jointly model user preferences within and across domains, enabling the system to capture complex and evolving consumer interests. Extensive experiments on four large-scale e-commerce datasets demonstrate that TEMA-LLM consistently outperforms state-of-the-art baselines, underscoring the benefits of LLM-based semantic tagging and multi-attention integration for consumer-facing recommendation systems. The proposed approach highlights the potential of LLMs to advance intelligent, user-centric services in the field of consumer electronics.
Aerial object detection is vital for remote sensing applications, yet remains challenged by the prohibitive costs of large-scale data annotation. While current zero-shot detection (ZSD) and open vocabulary detection (OVD) reduce annotation dependency, their reliance on base-class training limits adaptability in complex scenarios. We propose KGCS, a novel zero-annotation framework that achieves accurate detection without any training data through synergistic integration of domain knowledge with foundation models. Our approach introduces two key innovations: a structure-aware dual-path proposal strategy that maintains semantic coherence for composite structures and a tripartite feature dictionary enabling semantic-level alignment beyond conventional category-based matching. Extensive evaluations on DIOR and DOTA datasets demonstrate KGCS's superior performance in both precision and recall, highlighting its potential as an efficient and scalable solution for aerial imagery analysis. The code will be available at https://github.com/HSH55/KGCS
Few-shot class-incremental learning (FSCIL) in synthetic aperture radar imagery presents unique challenges due to severe data scarcity and SAR-specific variability. In particular, strong azimuth sensitivity in SAR induces large intra-class variation and inter-class confusion, and FSCIL sequential updates further lead to catastrophic forgetting of previously learned classes. Inspired by neural collapse, we propose an optical-guided SAR FSCIL framework, which derives orthogonal feature subspaces from a data-rich optical ATR dataset and uses them as geometric priors to guide SAR feature learning. SAR features are projected onto these orthogonal subspaces via principal angle constraints, effectively transferring discriminative structure from the optical to the SAR domain. Specifically, our projection loss and the classifier loss optimized with a frozen simplex-ETF geometry jointly induce neural collapse by concentrating features around class means while maintaining large inter-class angles. We evaluate the approach on a benchmark comprising an optical ATR dataset and a SAR ATR dataset with 24 target classes, organized into a base training session and seven incremental sessions. Compared with recent FSCIL methods including NCFSCIL and so on, our method achieves the highest final accuracy and a favorable trade-off between final performance and performance degradation. Moreover, neural collapse metrics show improved intra-class compactness and inter-class separability, indicating that the learned features more closely approximate the ideal simplex-ETF geometry.
Existing deep learning methods for optical and synthetic aperture radar (SAR) image registration have attempted to incorporate structural or gradient information. Although capable of simultaneously extracting low-level structural features and high-level semantic features, most still primarily rely on grayscale or intensity information. As a result, explicit guidance for accurately modeling structural features within the network is often limited. Under cross-modal conditions characterized by strong nonlinear radiometric variations and noise interference, these limitations in feature description partially constrain registration performance. To address this issue, this article proposes a registration framework guided by hierarchical structural information. We introduce a structure-aware weight extraction module to explicitly capture multiscale edge and curvature features. By integrating channel and spatial attention mechanisms, we translate structural priors into guidance for the feature learning process. Furthermore, within the joint detection-description coupled constraint framework, we introduce a differentiated edge prior information computation strategy and embed it into the constraints of keypoint detection to improve structural consistency constraints under cross-modal conditions. Concurrently, we design an enhanced description alignment loss that leverages multilevel constraints to enhance the learning capability of cross-modal similarity features. Finally, matching and registration performance comparisons with other algorithms were conducted on two datasets. The experimental results demonstrate that our method achieves effective improvements in enhancing the robustness and accuracy of optical-SAR image registration.
Deep learning-based object detection has become a crucial component in the interpretation of Synthetic Aperture Radar (SAR) imagery, widely applied in monitoring and surveillance tasks. However, the robustness of these data-driven models against adversarial perturbations in complex electromagnetic environments remains an open question. This paper investigates the vulnerability of SAR object detectors to Universal Adversarial Patch (UAP) attacks. We design a compact, location-agnostic patch and optimize it using a detection-aware objective specifically formulated to suppress objectness confidence and bounding box regression scores. This allows the patch to degrade performance effectively without requiring precise target alignment or scene-specific adaptation. Experiments on the SAR-AIRcraft-1.0 dataset indicate a noticeable decline in detection metrics, particularly in Recall rates and mAP, with observed black-box transferability across different architectures. Supported by ablation studies on patch scale and position, these findings reveal the potential susceptibility of SAR detectors to localized interference, serving as a reference for future security assessments.
InSAR-based scene matching navigation is a cutting-edge autonomous navigation scheme for GNSS-denied environments. However, due to the inherent speckle noise of SAR systems and texture differences between real-time and reference interferograms, the extracted matching feature points inevitably contain observation noise and outliers. Meanwhile, current mainstream platform localization inversion primarily relies on simplified airborne SAR imaging geometric models. Lacking error suppression mechanisms, these models cause the matching observation errors to directly propagate and amplify into platform positioning errors. To address these issues, this paper proposes a robust localization estimation model named Inv-RD. By incorporating time-dependent payload orbital equations, the inherently underdetermined Range-Doppler (RD) inversion problem is reframed into an overdetermined nonlinear optimization problem based on redundant observations. Subsequently, the Levenberg-Marquardt (L-M) algorithm is employed for global optimization, which significantly mitigates the impact of matching noise on localization accuracy. Experiments using measured flight data demonstrate that, even in the presence of matching errors, the proposed Inv-RD model significantly outperforms traditional geometric models in terms of positioning precision.