The accurate vehicle navigation plays a critical role in vehicle-to-everything (V2X) applications, including connected transportation systems, intelligent traffic management, and autonomous driving. To address the stringent demands of these scenarios, the integration of the Global Navigation Satellite Systems (GNSS) with the visual-inertial navigation system (VINS) has emerged as a pivotal advancement. Despite these strides, navigation systems remain susceptible to abnormal data. This data, originating from unpredictable external environments and internal device fallibility, poses a threat of substantial errors and system drift. In this article, we present robust fusion-based navigation (RF-Nav), a robust fusion-based GNSS-VINS navigation system with enhanced data processing and dynamic factor correction. The framework innovates with a dual-pronged approach: it first applies adaptive gamma correction with bilateral filtering and contrast-limited adaptive histogram equalization (AGCBF-CLAHE) to refine raw images; then, it deploys a long short-term memory (LSTM) denoising network enhanced with an advanced wavelet threshold for inertial measurement unit (IMU) data refinement. This dual enhancement of visual and IMU data integrity is further bolstered by a dynamic factor confidence correction mechanism, rooted in factor graph optimization (FGO), designed to counteract the adverse effects of abnormal data. Extensive experiments on large-scale public and real-field datasets demonstrate that RF-Nav exhibits superior robustness and accuracy in various environments.
The multimodal image enhancement technology for autonomous aerial vehicles (AAVs) significantly improves full-time perception and recognition capabilities in complex environments. However, existing multimodal image enhancement models typically utilize features from high-quality modalities to cross-modal guide the restoration of low-quality images. When multimodal images simultaneously suffer from degradation, these models often have obvious performance loss. To address these issues, we propose a visible-infrared bidirectional collaborative enhancement network (BCENet). The network introduces a bidirectional degradation-aware integration mechanism to dynamically compensate and enhance cross-modal features under varying degradation levels. Specifically, we propose an adaptive degradation perception method that identifies degradation types and decouples features through frequency spectral analysis of degraded visible images, thereby enhancing the model's ability to represent different degradation features. We also propose a multimodal wavelet interaction enhancement method, which leverages the multiscale decomposition capability of wavelet transform to decouple multifrequency subband signals, effectively suppressing blur and noise. Subsequently, the multimodal features are dynamically and interactively fused in both spatial and channel dimensions, enabling deep complementary integration of cross-modal information. Furthermore, to advance image enhancement research in complex AAV scenarios, we construct a new multimodal image enhancement dataset based on VGTSR1.0 and VGTSR2.0 datasets. It contains visible images processed with haze, motion blur, and low-light conditions, along with their corresponding low-resolution (LR) infrared images. Extensive experiments demonstrate that our method outperforms current state-of-the-art methods in multiple tasks. It also shows significant potential for practical applications in night monitoring and AAV remote sensing.
WiFi sensing offers passive and privacy-preserving perception that complements vision-based sensing, but its performance degrades sharply under domain shifts caused by changes in environment, users, or hardware. This challenge is exacerbated in real-world deployments where source data are unavailable, motivating test-time adaptation (TTA) as a practical solution for self-calibration using only unlabeled target samples. We introduce WiTTA-Bench, the first comprehensive benchmark for WiFi TTA, covering 20 representative methods, two adaptation protocols (OTTA and TTDA), and three major physics-induced shifts in WiFi: cross-environment, cross-subject, and cross-device. Furthermore, we contribute a new dataset featuring paired recordings from heterogeneous devices to bridge the cross-device gap. Extensive experiments reveal three key insights unique to WiFi sensing: (i) WiFi domain shifts exhibit a physics-induced hierarchy: environmental changes alter multipath statistics, subject variation perturbs temporal–spectral geometry, and hardware differences reshape the entire feature manifold; (ii) OTTA and TTDA are complementary: lightweight OTTA handles mild statistical drift, while TTDA is necessary to correct deep, hardware-induced structural distortions; and (iii) OTTA is hyperparameter-robust and scales linearly with source quality, whereas TTDA is more sensitive due to recursive self-training. WiTTA-Bench establishes the first systematic foundation for adaptive, robust, and deployable WiFi sensing under realistic wireless conditions.
Accurate remaining useful life (RUL) prediction for aircraft engines is essential for reducing maintenance costs and preventing failures. However, existing methods struggle to capture both spatial relationships and complex multifactor interactions, such as those between temperature and pressure. To address these challenges, we propose the knowledge-guided dual-path multifeature fusion (KDMFusion) framework that combines spatial-temporal feature extraction and multifeature fusion. A convolutional neural network (CNN) enhanced with domain knowledge captures spatial dependencies, while a gated recurrent unit (GRU) with a self-attention mechanism models long-term temporal relationships. By integrating spatial, temporal, and engineered features, our model offers a comprehensive representation of engine states, improving performance under varying operational conditions and noisy data. Extensive tests on the NASA C-MAPSS and N-CMAPSS datasets demonstrate the effectiveness of our method. On the C-MAPSS dataset, our approach reduces the root mean square error (RMSE) by at least 3.0% and the Score by at least 12.4%. Similarly, on the N-CMAPSS dataset, it decreases the RMSE by at least 4.8% and the Score by at least 5.0%. These results highlight the robustness and reliability of our method.
Regular monitoring of marine life is essential for preserving the stability of marine ecosystems. However, underwater target detection presents several challenges, particularly in balancing accuracy with model efficiency and real-time performance. To address these issues, we propose an innovative approach that combines the Structured Space Model (SSM) with feature enhancement, specifically designed for small target detection in underwater environments. We developed a high-accuracy, lightweight detection model-UWNet. The results demonstrate that UWNet excels in detection accuracy, particularly in identifying difficult-to-detect organisms like starfish and scallops. Compared to other models, UWNet reduces the number of model parameters by 5% to 390%, substantially improving computational efficiency while maintaining top detection accuracy. Its lightweight design enhances the model's applicability for deployment on underwater robots.
Domain shift, arising from varying imaging devices and data sources, poses a significant challenge segmentation models, particularly in clinical applications where accurate segmentation is essential for diagnosis and treatment. While Source-Free Domain Adaptation (SFDA) is advantageous in clinical settings adapts pre-trained models without requiring source data, it can generate inaccurate pseudo-labels that performance. To address this challenge, we propose a two-stage Progressive Pseudo-Labels Enhancement approach for SFDA in medical image segmentation. In Stage I, we integrate clustering into the SFDA framework and devise an Adaptive Feature Enhancement Clustering (AFEC) module to generate robust pseudo-labels. enhancing features, we address the limitations of adapting pre-trained models, improving feature learning the target domain. Furthermore, we analyze the differences between source and enhanced target domains select a suitable strategy adaptively for clustering neighboring pixels in the target domain. This results significant pseudo-labels quality improvement in the average Dice score for challenging optic cup segmentation, from 74.11% to 83.04%, ensuring more reliable segmentation for subsequent clinical analysis. In Stage refine pseudo-labels using a Clinical-Prior Contrastive Loss (CPCL), which incorporates clinical knowledge to enhance boundary delineation and improve segmentation accuracy. Additionally, we introduce Feature Awareness Contrastive Loss (FACL) to address domain shift by improving intra-class consistency and reducing inter-class discrepancies, resulting in a performance boost from 83.04% to 85.00%. Our framework outperforms leading SFDA methods on four benchmark fundus image datasets across seven domain shifts for optic disc cup segmentation. Code is available at https://github.com/ggllllll/ESFDA.git.
Change detection techniques, which extract different regions of interest from bi-temporal remote sensing images, play a crucial role in various fields such as environmental protection, damage assessment, and urban planning. However, visual style interferences stemming from varying acquisition times, such as radiation, weather, and phenology changes, often lead to false detections. Existing methods struggle to robustly measure background similarity in the presence of such discrepancies and lack quantitative validation for assessing their effectiveness. To address these limitations, we propose Representation Consistency Change Detection (RCCD), a novel deep learning framework that enforces global style and local spatial consistency of features across encoding and decoding stages for robust cross-visual style change detection. RCCD leverages large-kernel convolutional supervision for local spatial context awareness and global content-aware style transfer for feature harmonization, effectively suppressing interference from background variations. Extensive evaluations on S2Looking and LEVIR-CD+ datasets demonstrate RCCD’s superior performance, achieving state-of-the-art F1-scores. Furthermore, on dedicated subsets with large visual style differences, RCCD exhibits more substantial improvements, highlighting its effectiveness in mitigating interference caused by visual style errors. The code has been open-sourced on GitHub.
Deep learning methods, renowned for their ability to discern physical features from images, are frequently used in the semantic segmentation of remote sensing images. However, objects with different functional attributes may exhibit similar physical characteristics, resulting in comparable spectral reflectance and visual features. This issue, known as the "different categories with the same spectra" problem, limits the ability to discriminate between objects, thereby increasing the difficulty of differentiation. Studies based on Euclidean space often struggle to distinguish between objects due to the limited information available. To improve differentiation, additional information-such as neighborhood relationships-needs to be incorporated. The geographical scenario, i.e., the environmental context of the object-which includes crucial neighborhood categories and their spatial relationships, provides spatial relationship information for objects. This information provides an important context for object differentiation and becomes the key to distinguishing these similar objects. Following this idea, geographical scenarios are represented as graphs in graph space, with categories as nodes and adjacency relationships as edges, and a geographical knowledge graph is created based on all the scenarios. We convert the remote sensing images into graphs to match with the geographical scenarios and propose a knowledge-based semantic segmentation network for remote sensing, graph structure attention network (GSAN). In GSAN, a graph structure attention (GSAT) is designed based on the graph kernel. This allows it to discern graph structures corresponding to different geographical scenarios. GSAT serves as a link between the fine-grained visual objects and the coarse-grained semantic knowledge. Experiment results indicate that GSAN outperforms other attention networks in semantic segmentation on our sea and land remote sensing (SLRS) dataset. This demonstrates its advantages in geographical scenario recognition and remote sensing semantic segmentation.
Underwater object detection plays a crucial role in applications such as marine ecological monitoring and underwater rescue operations. However, challenges such as limited underwater data availability and low scene diversity hinder detection accuracy. In this paper, we propose the Underwater Layout-Guided Diffusion Framework (ULGF), a diffusion model-based framework designed to augment underwater detection datasets. Unlike conventional methods that generate underwater images by integrating in-air information, ULGF operates exclusively on a small set of underwater images and their corresponding labels, requiring no external data. We have publicly released the ULGF source code and the generated dataset for further research. Our approach enables the generation of high-fidelity, diverse, and theoretically infinite underwater images, substantially enhancing object detection performance in real-world underwater scenarios. Furthermore, we evaluate the quality of the generated underwater images, demonstrating that ULGF produces images with a smaller domain gap. Zhuang Yaoming and colleagues propose a diffusion model-based framework for generating underwater detection datasets. This approach requires only a small set of underwater images with corresponding annotations to produce high-quality, diverse underwater images, thereby enhancing object detection performance in real-world underwater scenarios.
The Segment Anything Model (SAM) excels in general segmentation but encounters difficulties in medical imaging due to few-shot learning challenges, particularly with extremely limited annotated data. Existing approaches often suffer from insufficient feature extraction and inadequate loss function balancing, resulting in decreased accuracy and poor generalization. To address these issues, we propose BiASAM, which uniquely incorporates two bidirectional attention mechanisms into SAM for medical image segmentation. Firstly, BiASAM integrates a spatial-frequency attention module to improve feature extraction, enhancing the model's ability to capture both fine and coarse details. Secondly, we employ an attention-based gradient update mechanism that dynamically adjusts loss weights, boosting the model's learning efficiency and adaptability in data-scarce scenarios. Additionally, BiASAM utilizes the point and box fusion prompt to enhance segmentation precision at both global and local levels. Experiments across various medical datasets show BiASAM achieves performance comparable to fully supervised methods with just two labeled samples.
Precipitation nowcasting predicts future radar sequences based on current observations, which is a highly challenging task driven by the inherent complexity of the Earth system. Accurate nowcasting is of utmost importance for addressing various societal needs, including disaster management, agriculture, transportation, and energy optimization. As a complementary to existing non-autoregressive nowcasting approaches, we investigate the impact of prediction horizons on nowcasting models and propose SimCast, a novel training pipeline featuring a short-to-long term knowledge distillation technique coupled with a weighted MSE loss to prioritize heavy rainfall regions. Improved nowcasting predictions can be obtained without introducing additional overhead during inference. As SimCast generates deterministic predictions, we further integrate it into a diffusion-based framework named CasCast, leveraging the strengths from probabilistic models to overcome limitations such as blurriness and distribution shift in deterministic outputs. Extensive experimental results on three benchmark datasets validate the effectiveness of the proposed framework, achieving mean CSI scores of 0.452 on SEVIR, 0.474 on HKO-7, and 0.361 on MeteoNet, which outperforms existing approaches by a significant margin.
Temporal action detection (TAD) is a vital challenge in computer vision and the Internet of Things, aiming to detect and identify actions within temporal sequences. While TAD has primarily been associated with video data, its applications can also be extended to sensor data, opening up opportunities for various real-world applications. However, applying existing TAD models to sensory signals presents distinct challenges such as varying sampling rates, intricate pattern structures, and subtle, noise-prone patterns. In response to these challenges, we propose a Sensory Temporal Action Detection (STADe) model. STADe leverages Fourier kernels and adaptive frequency filtering to adaptively capture the nuanced interplay of temporal and frequency features underlying complex patterns. Moreover, STADe embraces adaptability by employing deep fusion at varying resolutions and scales, making it versatile enough to accommodate diverse data characteristics, such as the wide spectrum of sampling rates and action durations encountered in sensory signals. Unlike conventional models with unidirectional category-to-proposal dependencies, STADe adopts a cross-cascade predictor to introduce bidirectional and temporal dependencies within categories. To extensively evaluate STADe and promote future research in sensory TAD, we establish three diverse datasets using various sensors, featuring diverse sensor types, action categories, and sampling rates. Experiments across one public and our three new datasets demonstrate STADe's superior performance over state-of-the-art TAD models in sensory TAD tasks.
Human activity recognition (HAR) is a crucial task in IoT systems with applications ranging from surveillance and intruder detection to home automation and more. Recently, non-invasive HAR utilizing WiFi signals has gained considerable attention due to advancements in ubiquitous WiFi technologies. However, recent studies have revealed significant privacy risks associated with WiFi signals, raising concerns about bio-information leakage. To address these concerns, the decentralized paradigm, particularly federated learning (FL), has emerged as a promising approach for training HAR models while preserving data privacy. Nevertheless, FL models may struggle in end-user environments due to substantial domain discrepancies between the source training data and the target end-user environment. This discrepancy arises from the sensitivity of WiFi signals to environmental changes, resulting in notable domain shifts. As a consequence, FL-based HAR approaches often face challenges when deployed in real-world WiFi environments. Albeit there are pioneer attempts on federated domain adaptation, they typically require non-trivial communication and computation cost, which is prohibitively expensive especially considering edge-based hardware equipment of end-user environment. In this paper, we propose a model to democratize the WiFi-based HAR system by enhancing recognition accuracy in unannotated end-user environments while prioritizing data privacy. Our model leverages the hypothesis transfer and a lightweight hypothesis ensemble to mitigate negative transfer. We prove a tighter theoretical upper bound compared to existing multi-source federated domain adaptation models. Extensive experiments shows our model improves the average accuracy by approximately 10 absolute percentage points in both cross-person and cross-environment settings comparing several state-of-the-art baselines.
WiFi-based smart human sensing technology enabled by Channel State Information (CSI) has received great attention in recent years. However, CSI-based sensing systems suffer from performance degradation when deployed in different environments. Existing works solve this problem by domain adaptation using massive unlabeled high-quality data from the new environment, which is usually unavailable in practice. In this paper, we propose a novel augmented environment-invariant robust WiFi gesture recognition system named AirFi that deals with the issue of environment dependency from a new perspective. The AirFi is a novel domain generalization framework that learns the critical part of CSI regardless of different environments and generalizes the model to unseen scenarios, which does not require collecting any data for adaptation to the new environment. AirFi extracts the common features from several training environment settings and minimizes the distribution differences among them. The feature is further augmented to be more robust to environments. Moreover, the system can be further improved by few-shot learning techniques. Compared to state-of-the-art methods, AirFi is able to work in different environment settings without acquiring any CSI data from the new environment. The experimental results demonstrate that our system remains robust in the new environment and outperforms the compared systems.
In recent times, notable advancements have been achieved in amalgamating heterogeneous remote sensing imagery to facilitate Earth observation through the adoption of convolutional neural networks. Nonetheless, due to the variety in imaging mechanisms and imbalanced information prevalent among heterogeneous data, the efficacious exploitation of semantic correlation across different modalities for generating discriminative features continues to pose a formidable challenge. Moreover, not all modalities included in the training dataset can be obtained in real-world test scenarios. Hence, the following inquiry arises: How to explicitly leverage semantic correlation between heterogeneous images to construct discriminative features and maintain performance in test scenarios with missing modalities? To address this pressing concern, we propose an innovative assisted learning framework that employs a “teacher-student” architecture equipped with local and global distillation schemes. We partition the framework into two distinct segments, where each segment acquires specific knowledge independently. In terms of local distillation, the teacher network fosters discriminative feature extraction in the student network using a pixel-wise approach, augmented by the inclusion of a regularization factor to ensure the accuracy of knowledge transfer. For global distillation, the student network is motivated to assimilate category-related information derived from the teacher network, thereby further enriching the knowledge encoding. Extensive evaluations of the datasets, utilizing either optical image/Digital Surface Model (DSM) or optical image/Synthetic Aperture Radar (SAR) for land use classification, provide evidence favoring the effectiveness of the proposed method. Code is available at https://github.com/WHUlwb/Assisted_learning.
In the above article [1] , the legends of Figs. 7 – 9 and 11 – 20 , and Tables I , II , and IV should have used the word “carrier” instead of “battleship.” Each figure is provided here with the corrected terminology.
Forest resources have important ecological and environmental values, and monitoring forest changes using remote sensing images is essential for resource management and ecological protection. However, current forest change detection methods fail to simultaneously integrate fine spatial information with temporal dynamic data, making them susceptible to pseudo changes induced by seasonal factors. In this paper, we propose a forest change detection method called STFNet that integrates multi-source spatiotemporal information. By combining fine spatial details of high-resolution images with dynamic information from time-series images, STFNet enhances the accuracy of forest change detection, alleviating the problem of information fusion difficulties caused by inconsistent granularity in spatiotemporal spectral features from different sources. In STFNet, we propose a cross-attention-based temporal differential feature fusion module (CATFF) to capture spatiotemporal dependencies within time-series images and a multiresolution contextual differential feature fusion module (MCDF) to achieve efficient spatial contexture fusion across multiresolution images. To validate our method, we conduct experiments using Gaofen and Sentinel-2 satellite images. Experimental results demonstrate that STFNet achieves excellent performance with an F1-score of 87.65%, outperforming state-of-the-art methods by at least 2.02%. Our ablation study further confirms the effectiveness of our method in leveraging time-series information to detect forest changes and suppress seasonal interference.
In the storage of hazardous chemicals, due to space limitations, various hazardous chemicals are usually mixed stored when their chemical properties do not conflict. In a fire or other accidents during storage, the emergency response includes two key steps: first, using fire extinguishers like dry powder and carbon dioxide to extinguish the burning hazardous chemicals. In addition, hazardous chemicals around the accident site are often watered to cool down to prevent the spread of the fire. But both the water and extinguishers may react chemically with hazardous chemicals at the accident site, potentially triggering secondary accidents. However, the existing research about hazardous chemical domino accidents only focuses on the pre-rescue stage and ignores the simulation of rescue-induced accidents that occur after rescue. Aiming at the problem, a quantitative representation algorithm for the spatial correlation of hazardous chemicals is first proposed to enhance the understanding of their spatial relationships. Subsequently, a graph neural network is introduced to simulate the evolution process of hazardous chemical cascade accidents. By aggregating the physical and chemical characteristics, the initial accident information of nodes, and bi-temporal node status information, deep learning models have gained the ability to accurately predict node states, thereby improving the intelligent simulation of hazardous chemical accidents. The experimental results validated the effectiveness of the method.
Recently, WiFi-based human activity recognition (HAR) has emerged as a promising technique for human-computer interactions, owing to its widespread availability and non-invasiveness. However, deploying WiFi-based HAR systems in new environments often results in performance degradation. Existing WiFi-based HAR systems across different environments typically assume identical category spaces between the source and target, an assumption challenged by practical scenarios. In this paper, we present WiCAU, a comprehensive adaptation with uncertainty-awareness for WiFi-based HAR across environments, designed to tackle the challenges of environments with unequal category spaces—a scenario known as Partial Domain Adaptation (PDA). Different from conventional PDA methods that usually focus on training the feature extractor to align feature distributions or implement separate reweighting models to adjust source domain feature weights,WiCAU integrates feature alignment and source data reweighting to mitigate the risk of negative transfer. It also introduces an uncertain complement entropy to effectively handle uncertainty within the source environment. Moreover, WiCAU employs a hybrid network that combines wavelet analysis with deep neural networks to capture both the temporal and spatial dynamics present in WiFi Channel State Information (CSI) data. WiCAU’s superior performance in PDA scenarios for HAR is demonstrated through comprehensive experiments with both self-built and publicly available datasets.