Cloud contamination is a common problem in Earth observation that hinders various remote sensing applications. To address this problem, recent studies have employed deep neural networks and multi-modal data fusion to reconstruct cloud-free optical imagery. However, this task faces many challenges, such as: (1) the scarcity of suitable multi-modal datasets; (2) the ineffective use of feature correlations; and (3) the limited applicability of existing models. To overcome these challenges, this study proposes a novel solution that fuses high-spatial SAR and low-spatial optical data to reconstruct high-quality cloud-free multi-spectral optical products. First, a curated benchmark dataset, named SMILE-CR, is created with a realistic cloud simulation strategy. The SMILE-CR serves as a global and multi-modal cloud removal dataset for the Landsat-8 sensor, with Sentinel-1 and MODIS data as additional supplementary data. Second, a Transformer-based cloud removal network, abbreviated as CRformer, is developed with two novel modules: multi-head dense and sparse attention and multi-scale gated-dconv feed-forward network. The CRformer achieves global attention while suppressing the weak correlations and enhancing the multi-scale cloud features by filtering out invalid features. The performance of the proposed method is evaluated through extensive experiments. The results show that the CRformer surpasses the state-of-the-art cloud removal methods with significant improvements in both quantitative and qualitative metrics. The fusion of MODIS and Sentinel-1 data is shown to be effective and necessary in reconstructing Landsat-8 observations. Moreover, the CRformer model can be readily applied to reconstruct time-series cloud-free Landsat-8 products in Wuhan city, which can improve the average accuracy of land cover mapping by over 3%.
Landsat optical sensor is crucial for the long-term observations of the Earth’s surface with a 30 m spatial resolution. However, the 16-day revisit cycle and severe atmospheric interference have impeded the monitoring of rapid surface changes. Spatiotemporal fusion (STF) is a classic method of predicting Landsat surface reflectance with multi-temporal and multi-source data, but it is limited by unpredictable temporal changes and cloudy Landsat-MODIS image pairs. Another emerging solution is synthetic aperture radar (SAR)-to-optical image translation (S2OIT), which always produces spectral distortions. To tackle these defects, we propose a new data-driven solution, SAR-optical data-based spatial–spectral fusion (SOSSF), which combines the high-spatial and cloud-free advantages of Sentinel-1 data and the high-spectral and high-temporal advantages of MODIS images to synthesize high-spatial and high-temporal Landsat-8 images. To achieve this solution, we first establish a worldwide benchmark dataset, namely SMILE, with various land cover types and all meteorological seasons, satisfying the big data requirements of deep learning. Second, we design an attention-based dual-path fusion network (ADFNet) to respectively extract and fully fuse spatial and spectral information from SAR-optical data. Extensive experiments suggest that the proposed SOSSF solution outperforms the state-of-the-art STF and S2OIT solutions, robustly performing in the continuously changing and frequently cloudy regions. The proposed ADFNet model achieves the best visual effect and the highest accuracy in different scenes, seasons, and bands. Furthermore, the proposed SOSSF solution is proven to be a practical way to simulate time-series and large-scale Landsat-8 surface reflectance, considerably enriching raw Landsat-8 products.
As a challenge in the field of smart medicine, medical picture segmentation gives important decisions and is the basis for future diagnosis by doctors. In the past decade, FCN-based network topologies have made amazing progress in the field. However, the limited perceptual capacity of convolutional kernels in FCN network topologies limits the network's ability to acquire a global field of view. We propose BSANet, a 3D medical image segmentation network based on self-focus and multi-scale information fusion with a high-performance feature extraction module. BSANet can help the network to extract deeper features by obtaining a larger range of perceptual capabilities by using its self-focus and multi-scale information aggregation pooling modules. Brain tumor segmentation dataset and multi-organ segmentation dataset are used to train and evaluate our model. BSANet produces excellent results with its high-performance feature extraction network with an attention module and multi-scale information fusion module.
It is of great practical significance to quickly, accurately, and effectively identify the effects of rice diseases on rice yield. This paper proposes a rice disease identification method based on an improved DenseNet network (DenseNet). This method uses DenseNet as the benchmark model and uses the channel attention mechanism squeeze-and-excitation to strengthen the favorable features, while suppressing the unfavorable features. Then, depth wise separable convolutions are introduced to replace some standard convolutions in the dense network to improve the parameter utilization and training speed. Using the AdaBound algorithm, combined with the adaptive optimization method, the parameter adjustment time reduces. In the experiments on five kinds of rice disease datasets, the average classification accuracy of the method in this paper is 99.4%, which is 13.8 percentage points higher than the original model. At the same time, it is compared with other existing recognition methods, such as ResNet, VGG, and Vision Transformer. The recognition accuracy of this method is higher, realizes the effective classification of rice disease images, and provides a new method for the development of crop disease identification technology and smart agriculture.
The evaluation of rice disease severity is a quantitative indicator for precise disease control, which is of great significance for ensuring rice yield. In the past, it was usually done manually, and the judgment of rice blast severity can be subjective and time-consuming. To address the above problems, this paper proposes a real-time rice blast disease segmentation method based on a feature fusion and attention mechanism: Deep Feature Fusion and Attention Network (abbreviated to DFFANet). To realize the extraction of the shallow and deep features of rice blast disease as complete as possible, a feature extraction (DCABlock) module and a feature fusion (FFM) module are designed; then, a lightweight attention module is further designed to guide the features learning, effectively fusing the extracted features at different scales, and use the above modules to build a DFFANet lightweight network model. This model is applied to rice blast spot segmentation and compared with other existing methods in this field. The experimental results show that the method proposed in this study has better anti-interference ability, achieving 96.15% MioU, a speed of 188 FPS, and the number of parameters is only 1.4 M, which can achieve a high detection speed with a small number of model parameters, and achieves an effective balance between segmentation accuracy and speed, thereby reducing the requirements for hardware equipment and realizing low-cost embedded development. It provides technical support for real-time rapid detection of rice diseases.
In this article, we elaborate on the scientific outcomes of the 2021 Data Fusion Contest (DFC2021), which was organized by the Image Analysis and Data Fusion Technical Committee of the IEEE Geoscience and Remote Sensing Society, on the subject of geospatial artificial intelligence for social good. The ultimate objective of the contest was to model the state and changes of artificial and natural environments from multimodal and multitemporal remotely sensed data towards sustainable developments. DFC2021 consisted of two challenge tracks: Detection of settlements without electricity (DSE) and multitemporal semantic change detection. We focus here on the outcome of the DSE track. This article presents the corresponding approaches and reports the results of the best-performing methods during the contest.
In this paper, a multi-model fusion framework is proposed for automatic detection of settlements without electricity (DSE) based on the multimodal and multitemporal remote sensing data. To settle the problems of data noise and data redundancy, the data preprocessing step, which consists of band selection, cloud removal, grayscale stretch and data augmentation, is firstly applied. Two models of single-task and dual-task are further constructed for DSE. The single-task model builds a global context convolutional neural network (GC-CNN) for the detection of settlements without electricity and the dual-task model employs the GC-CNN for settlement detection and the random forest classifier for electricity detection. Moreover, a model fusion principle and a post-processing method is designed to integrate and improve the results above, thus producing the final segmentation result. Verified through the competition website, the proposed method achieved a F1-score of 0.8806, ranking second in the first track of 2021 IEEE GRSS Data Fusion Contest.
The evaluation of geometric accuracy of high-resolution satellite images (HRSIs) has been increasingly recognized in recent years. The traditional approach is to verify each satellite individually. It is difficult to directly compare the difference in their accuracy. In order to evaluate geometric accuracy for multiple satellite images based on the same ground control benchmark, a reliable test field in Xianning (China) was utilized for geometric accuracy validation of HRSIs. Our research team has obtained multiple HRSIs in the Xianning test field, such as SPOT-6, Pleaides, ALOS, ZY-3 and TH-1. In addition, ground control points (GCPs) were acquired with GPS by field surveying, which were used to select the significant feature area on the images. We assess the orientation accuracy of the HRSIs with the single image and stereo models. Within this study, the geometrical performance of multiple HRSIs was analyzed in detail, and the results of orientation are shown and discussed. As a result, it is feasible and necessary to establish such a geometric verification field to evaluate the geometric quality of multiple HRSIs.