In recent decades, medium-resolution satellites have played a pivotal role in ecosystem research, providing global daily coverage. The Moderate-Resolution Spectral Imager (MERSI), a key component of Fengyun-3 (FY-3) series, exhibits excellent data continuity over three generations, delivering large-scale spatial-continuous documents. However, images from sensors with wide view angles often exhibit considerable artifacts due to restriction of multidetector parallel scanning system and multi-interferences from space. Furthermore, MERSI introduces known observation redundancy across pixels and tracks, presenting significant yet untapped potential for image quality optimization and restoration from Level-1 (L1) to downstream products. We proposed a spatial quality enhancement (SPQE) algorithm for MERSI-like sensors. It has demonstrated effectiveness in generating detailed spatial information images without artifacts while retaining spectral information with four main procedures: 1) atmospheric correction inspired by Moderate-Resolution Imaging Spectroradiometer (MODIS); 2) optimization of geolocation sample methods through linear extrapolation, elimination of bow-tie effect through empirical screening; 3) utilization of observation redundancy from adjacent pixels through a restoration algorithm based on Thomas method; and 4) exploitation of observation redundancy from adjacent tracks via relative registration and information stack sampling. To illustrate the effectiveness and robustness of SPQE, the widely used MERSI-II images are tested. Examples of SPQE-derived surface reflectance (SR) imagery across global scenarios are provided. We evaluate the potential effect of SPQE on reflectance and ecological applications. Extensive validation demonstrates that downstream products generated by SPQE-derived images exhibit a stronger correlation with classic products and public datasets compared to nonenhanced images. Given its performance and versatility, we anticipate the widespread use of SPQE for similar sensors.
In recent times, notable advancements have been achieved in amalgamating heterogeneous remote sensing imagery to facilitate Earth observation through the adoption of convolutional neural networks. Nonetheless, due to the variety in imaging mechanisms and imbalanced information prevalent among heterogeneous data, the efficacious exploitation of semantic correlation across different modalities for generating discriminative features continues to pose a formidable challenge. Moreover, not all modalities included in the training dataset can be obtained in real-world test scenarios. Hence, the following inquiry arises: How to explicitly leverage semantic correlation between heterogeneous images to construct discriminative features and maintain performance in test scenarios with missing modalities? To address this pressing concern, we propose an innovative assisted learning framework that employs a “teacher-student” architecture equipped with local and global distillation schemes. We partition the framework into two distinct segments, where each segment acquires specific knowledge independently. In terms of local distillation, the teacher network fosters discriminative feature extraction in the student network using a pixel-wise approach, augmented by the inclusion of a regularization factor to ensure the accuracy of knowledge transfer. For global distillation, the student network is motivated to assimilate category-related information derived from the teacher network, thereby further enriching the knowledge encoding. Extensive evaluations of the datasets, utilizing either optical image/Digital Surface Model (DSM) or optical image/Synthetic Aperture Radar (SAR) for land use classification, provide evidence favoring the effectiveness of the proposed method. Code is available at https://github.com/WHUlwb/Assisted_learning.
Building change detection (CD) aims to detect changes in buildings from bi-temporal pairwise images obtained at different times. Typically, a deep learning-based building CD algorithm requires bi-temporal samples with significant building changes for training. However, obtaining such bi-temporal samples is challenging because building changes have a low probability of occurrence. Fortunately, it is relatively simple to obtain single-temporal samples that include a substantial number of buildings. By using these single-temporal building samples, pseudo bi-temporal building change samples can be generated, which can effectively address the problem of limited available bi-temporal building change samples. In view of that, this study proposes a metric guided single-temporal supervised learning framework that uses single-temporal building samples for building CD. In the proposed framework, patch-pairing single-temporal supervised learning (PPSL) adopts a patch-pairing method to construct pseudo bi-temporal building change samples, while equipping the network to effectively suppresses the negative impact of geometric offset and radiation difference in real samples. To further suppress the impact of radiation difference and enhance the effectiveness of our framework, a metric-guided spatial attention module (MGSAM) is designed to minimize the intra-class feature differences between temporal samples and augment the spatial context modeling ability. The proposed method is verified by experiments on different datasets, and the results demonstrate that the proposed method can outperform the existing methods and achieve superior performance.
Change Detection (CD) is a vital monitoring method in Earth observation, especially pertinent for land-use analysis, city management, and disaster damage assessment. However, in the era of constellation interconnection and air-sky collaboration, the changes in the Regions Of Interest (ROI) cause many false detections due to geometric perspective rotation and temporal style difference. In response to these challenges, we introduce CDNeXt, this framework elucidates a robust and efficient method for combining Siamese networks based on the pre-trained backbone with the innovative Temporospatial Interactive Attention Module (TIAM) for remote sensing imagery. The CDNeXt can be categorized into four primary components: Encoder, Interactor, Decoder, and Detector. Notably, the Interactor, powered by TIAM, queries and rebuilds spatial perspective dependencies and temporal style correlations from binary temporal features extracted by the Encoder to enlarge the difference of ROI change. Culminating the process, the Detector integrates the hierarchical features generated by the Decoder, subsequently producing a binary change mask. We have achieved new State-Of-The-Art (SOTA) performance in change detection, with our method surpassing existing techniques on four benchmark datasets: an F1 score of 82.63% on SYSU-CD, 87.14% on LEVIR-CD+, 66.71% on S2Looking, and 71.11% BANDON. To further validate the effectiveness of the TIAM, we compared it to other attention modules in both interactive and non-interactive modes. Our code is available on GitHub: https://github.com/wjj282439449/CDNeXt.
Change Detection (CD) in remote sensing imagery is a valuable technique to identify urban expansion and anomalies on the Earth’s surface. However, the viewpoint difference between bi-temporal images often leads to significant view-point errors, resulting in false detection. Off-nadir datasets and global attention mechanisms have been commonly used to address this issue, but their effectiveness is challenging to validate. To tackle these challenges, we propose a novel Dual View Interactive Network (DVINet) specifically designed for CD in off-nadir images. Our approach incorporates the Viewpoint Interactive Attention Mechanism (VIM), which computes semantic correlations among dual view features extracted by the encoder to establish feature correspondences. In addition, we pioneer a viewpoint error validation method by evaluating the performance of CD in building facade regions. Experimental results on two off-nadir datasets, BANDON and S2Looking, show that DVINet achieves 70.65% and 66.26% F1, respectively, outperforming six mainstream methods in terms of CD performance.
当前武汉市城市更新行动,从大拆大建,进入"留改拆"并举的2.0时代,改造方式也从局部改造向成片连片更新转变.在当前2.0时代中,如何利用人工智能(Artificial Intelligence,AI)技术智能识别出城市"留改拆"单元显得尤为重要.当前AI和遥感技术已在自然资源典型地物类型识别、耕地保护和执法监察中得到广泛应用,本文第一次将AI和遥感技术用于"留改拆"单元的智能识别中,以辅助智能化城市更新行动.建立"留改拆"单元的样本,利用深度学习网络建立AI+遥感技术的智能化识别模型,选择遥感数据,进行武汉市更新片区"留改拆"单元智能化识别.通过遥感技术与深度学习算法的融合,提升了城市更新行动中"留改拆"单元识别的工作效率,为城市更新行动中的难点问题提供了科学依据.
Image segmentation aims to partition an image according to the objects in the scene and is a fundamental step in analysing very high spatial-resolution (VHR) remote sensing imagery. Current methods struggle to effectively consider land objects with diverse shapes and sizes. Additionally, the determination of segmentation scale parameters frequently adheres to a static and empirical doctrine, posing limitations on the segmentation of large-scale remote sensing images and yielding algorithms with limited interpretability. To address the above challenges, we propose a deep-learning-based region merging method dubbed DeepMerge to handle the segmentation of complete objects in large VHR images by integrating deep learning and region adjacency graph (RAG). This is the first method to use deep learning to learn the similarity and merge similar adjacent super-pixels in RAG. We propose a modified binary tree sampling method to generate shift-scale data, serving as inputs for transformer-based deep learning networks, a shift-scale attention with 3-Dimension relative position embedding to learn features across scales, and an embedding to fuse learned features with hand-crafted features. DeepMerge can achieve high segmentation accuracy in a supervised manner from large-scale remotely sensed images and provides an interpretable optimal scale parameter, which is validated using a remote sensing image of 0.55 m resolution covering an area of 5,660 km^2. The experimental results show that DeepMerge achieves the highest F value (0.9550) and the lowest total error TE (0.0895), correctly segmenting objects of different sizes and outperforming all competing segmentation methods.
Optical and synthetic aperture radar (SAR) images, two standard Earth observation tools, can reflect the characteristics of the surface from different perspectives and provide complementary information for land use classification. However, because they belong to different modes and express land objects differently, it is challenging to effectively fuse and use them to perform a pixel-wise classification. Current methods only focus on the local receptive field to fuse deep features in a single dimension, which is too simple to fully exploit the correlation between different modes. Moreover, the appearance disparities between each modality may induce semantic misalignment and disrupt the conditions of features’ fusion. To overcome the above problems, we introduce a spatial-aware circular module to generate a cross-modality receptive field and globally enhance the interaction between each pixel in the spatial dimension. Additionally, we recalibrate the features in the channel dimension to selectively refine and retain the essential things during the process, which can further achieve feature refinement. To reduce the impact of modal appearance disparities, we transform their high-level features into a common latent space and align their distributions to correlate the complementary cues hidden in each modality. The experimental results for the WHU-OPT-SAR dataset show that our method performed better than other state-of-the-art methods, with a mean intersection over union (mIoU) of 58.5% and an overall accuracy (OA) of 84.2%. Furthermore, the method obtained competitive results in Ezhou and Panjin, China. The results demonstrate our method’s applicability.
In surveying, mapping and geographic information systems, building extraction from remote sensing imagery is a common task. However, there are still some challenges in automatic building extraction. First, using only single-scale depth features cannot take into account the uncertainty of features such as the hue and texture of buildings in images, and the results are prone to missed detection. Moreover, extracted high-level features often lose structural information and have scale differences with low-level features, which results in less accurate extraction of boundaries. To simultaneously address these problems, we propose pyramid feature extraction (PFE) to construct multi-scale representations of buildings, which is inspired by the feature extraction of scale-invariant feature transform. We also apply attention modules in channel dimension and spatial dimension to PFE and low-level feature maps. Furthermore, we use the structural-cue-guided feature alignment module to learn the correlation between feature maps at different levels, obtaining high-resolution features with strong semantic representation and ensuring the integrity of high-level features in both structural and semantic dimensions. An edge loss is applied to get a highly accurate building boundary. For the WHU Building Dataset, our method achieves an F1 score of 95.3% and an Intersection over Union (IoU) score of 90.9%; for the Massachusetts Buildings Dataset, our method achieves an F1 score of 85.0% and an IoU score of 74.1%.
The detection and monitoring of changes in urban buildings, as a major place for human activities, have been considered profoundly in the field of remote sensing. In recent years, comparing with other traditional methods, the deep learning-based methods have become the mainstream methods for urban building change detection due to their strong learning ability and robustness. Unfortunately, often, it is difficult and costly to obtain sufficient samples for the change detection method development. As a result, the application of the deep learning-based building change detection methods is limited in practice. In our work, we proposed a novel multi-task network based on the idea of transfer learning, which is less dependent on change detection samples by appropriately selecting high-dimensional features for sharing and a unique decoding module. Different from other multi-task change detection networks, with the help of a high-accuracy building mask, our network can fully utilize the prior information from building detection branches and further improve the change detection result through the proposed object-level refinement algorithm. To evaluate the proposed method, experiments on the publicly available WHU Building Change Dataset were conducted. The experimental results show that the proposed method achieves F1 values of 0.8939, 0.9037, and 0.9212, respectively, when 10%, 25%, and 50% of change detection training samples are used for network training under the same conditions, thus, outperforming other methods.
Cloud, one of the poor atmospheric conditions, significantly reduces the usability of optical remote-sensing data and hampers follow-up applications. Thus, the identification of cloud remains a priority for various remote-sensing activities, such as product retrieval, land-use/cover classification, object detection, and especially for change detection. However, the complexity of clouds themselves make it difficult to detect thin clouds and small isolated clouds. To accurately detect clouds in satellite imagery, we propose a novel neural network named the Pyramid Contextual Network (PCNet). Considering the limited applicability of a regular convolution kernel, we employed a Dilated Residual Block (DRB) to extend the receptive field of the network, which contains a dilated convolution and residual connection. To improve the detection ability for thin clouds, the proposed new model, pyramid contextual block (PCB), was used to generate global information at different scales. FengYun-3D MERSI-II remote-sensing images covering China with 14,165 × 24,659 pixels, acquired on 17 July 2019, are processed to conduct cloud-detection experiments. Experimental results show that the overall precision rates of the trained network reach 97.1% and the overall recall rates reach 93.2%, which performs better both in quantity and quality than U-Net, UNet++, UNet3+, PSPNet and DeepLabV3+.