As a fundamental indicator of surface thermal conditions, land surface temperature (LST) provides essential information for characterizing land-atmosphere exchanges and surface energy dynamics. Owing to sensor design constraints, satellite observations with high-temporal resolution are typically limited to coarse spatial resolutions, making spatial downscaling of LST products essential for capturing fine-scale thermal heterogeneity and supporting detailed thermal environment analyses. In this study, a Geographically Weighted Gradient Boosting Regression (GWRBoost) model is developed by integrating geographically weighted regression (GWR) with gradient boosting to explicitly characterize spatial nonstationarity and complex nonlinear interactions between LST and environmental factors. The normalized difference vegetation index (NDVI), normalized difference water index (NDWI), and digital elevation model (DEM) were incorporated to generate used as predictors to downscale high-resolution LST with the Moderate Resolution Imaging Spectroradiometer (MODIS) LST from a spatial resolution of 990-90 m, with Landsat 8 LST serving as reference observations for validation. The proposed GWRBoost model was evaluated against the thermal data sharpening (TsHARP) method and the conventional GWR model. The results indicate that the GWRBoost model outperforms the comparative approaches, achieving a root-mean-squared error (RMSE) of 2.07 K and a mean absolute error (MAE) of 1.75 K. Additional site-based validation further demonstrates the stability and consistency of the LST downscaling outputs produced by the proposed model.
In response to the challenges faced by traditional deep learning (DL) models in adaptively extracting multiscale features from remote sensing images and establishing long-range dependencies across both channels and spatial dimensions-issues that contribute to a decline in classification performance-this study introduces a novel network architecture that integrates dynamic multiscale attention with spatiotemporal fusion, designated as multiscale dynamic convolution and attention fusion network (MDCA-Net). The research presents a dynamic multiscale large kernel attention, termed D-MLKA, which is engineered to dynamically weight and amalgamate multiscale convolutional features derived from 7 & times; 7 and 5 & times; 5 kernels, thereby facilitating a spectral-spatial adaptive perception mechanism. Subsequently, a spectral-spatial convolution and attention fusion module (SS-CAFM) is developed to leverage depthwise separable convolutions for the extraction of cross-channel spectral correlations, in conjunction with multihead attention to effectively capture long-range spatial dependencies. Moreover, a pseudo-temporal enhancement unit is introduced to bolster the model's robustness against interannual variations by generating virtual temporal features. Comparative experiments conducted on the Indian Pines and Houston University public datasets yielded overall accuracies (OAs) of 77.38% and 83.34%, respectively. Furthermore, when validated against the 2024 multispectral dataset from Rong'an County, Guangxi Autonomous Region, the model achieved a classification accuracy of 98.66%, representing a 2.15% improvement over existing methodologies such as Morphformer. Ultimately, the model demonstrated average accuracy (AA) of 98.12% and 97.79% on datasets from 2020 and 2022, respectively, thereby sustaining high performance levels while mitigating the risk of performance degradation.
Jointclassification of hyperspectral imagery with LiDAR and synthetic aperture radar (SAR) multimodal remote sensing data holds significant value for enhancing feature recognition accuracy. However, it still faces challenges, such as severe spectral interference, modal imbalance, and insufficient feature fusion. To address these challenges, this article proposes a multisource fusion network-DuFANet. This approach introduces a frequency-domain attention module in the hyperspectral branch to explicitly decompose and enhance spectral-spatial features. In the LiDAR/SAR branch, a depthwise convolutional self-attention module is designed to efficiently model modal local geometric features and global contextual information. A parallel transformer encoder is employed after both hyperspectral and LiDAR/SAR branches to capture long-range dependencies and global semantics. The fusion stage introduces a dual-branch cross-modal fusion attention module, achieving stable complementarity through dual-layer mechanisms at both the channel and spatial dimensions. Experiments on five datasets-Berlin, Muufl, Houston2013, Augsburg, and Trento, DuFANet achieves overall classification accuracies of 79.29%, 92.31%, 99.44%, 97.59%, and 99.75%, respectively, significantly outperforming existing methods. Furthermore, the model demonstrates exceptional performance in complex urban scenes and under few-shot conditions, achieving a favorable balance between accuracy and computational cost.
Crop identification with remote sensing images has currently emerged as a prominent research topic with a focus on flat regions dominated by crop cultivation. However, farmland is also distributed in vast complex terrain areas where the study of crop identification is necessary and of importance. This study aims to propose an innovative method to address the challenges by investigating what complex terrain poses to the feature extraction and classification algorithms. Shouyang Loess Hill and Jinzhong Basin were chosen as the study area. Three indices, topographic indices, vegetation indices, and texture indices were extracted from ALOS DEM and time-series Sentinel-2 spectral bands, respectively. Various feature combinations of spectral bands and three indices were ingested into the algorithm of random forest (RF), and then five models were built to investigate the impact of various feature combinations on the accuracy of crop identification in both hilly and flat areas. The results indicate that the combination of multiple features can significantly improve the accuracy of crop identification in complex terrain areas, while the differences in flat regions were less pronounced. This study revealed that the combination of diverse features with topographic features is an effective method to improve the accuracy of crop identification in complex terrain areas. The high-accuracy crop maps can provide valuable insights for optimizing agricultural practices and resource management in complex terrain areas. The future study should extend this method to other, more complex terrain areas with efficient optimization of topographic features.
Purposes In this study, the possibility of improving classification accuracy for crop type identification is examined with data fusion technology. Methods The GF-1 and Landsat 9 images were used to perform data fusion for crop type classification in Jinzhong region of Shanxi Province, China. The combination of PC Spectral Sharpening (PC), Gram-Schmidt Pan Sharpening (GS), and NNDiffuse Pan Sharpening (NN) fusion models facilitated the integration of the red, green, blue, and near-infrared bands from GF-1 WFV and Landsat 9 satellite images. Assessment of fusion outcomes according to mean, standard deviation, and information entropy identified the optimal fusion bands. By employing the Random Forest classification algorithm, crop classification was conducted on GF-1 WFV images, Landsat 9 images, and the best-fused images. Results Results demonstrate a significant enhancement in crop classification accuracy and stability for the fused GF-1 and Landsat 9 images, achieving an overall classification accuracy of 92.9%, a Kappa coefficient of 0.92, and an F1 Score of 87.4%. Furthermore, the overall accuracy, Kappa coefficient, and F1 Score of crop classification in the fused image are increased by 1.7%, 0.2, and 0.6%, respectively, compared with classification solely based on the GF-1 WFV image. Similarly, compared with Landsat 9 image classification, improvements are 3.2%, 0.4, and 4.4%, respectively. The utilization of GF-1 WFV near-infrared band and application of the NN algorithm to fuse Landsat 9 data demonstrate promising results in crop classification, highlighting its potential for widespread utilization in accurately extracting agricultural information across extensive geographical areas.
Highlights What are the main findings? A novel dual-branch collaborative Transformer network (PST-Net) is proposed, which effectively addresses the challenge of difficult collaborative modeling between global dependencies and local details in hyperspectral image classification through the collaboration between an adaptive spectral-spatial Token module and a parallel attentionenhanced CNN branch. PST Net achieved excellent classification performance on four challenging datasets (Salinas, Whuhh, Qingyun, Houston) with only 2% labeled samples, outperforming multiple comparison methods. What is the implication of the main finding? This method demonstrates strong robustness and generalization ability even with minimal training samples, providing an effective solution for hyperspectral image classification tasks where labeled data is scarce in practical applications. The proposed adaptive Token construction and cross-level interaction fusion mechanism provides a general and efficient framework for designing hybrid models that synergistically leverage the advantages of Transformer and CNN.Highlights What are the main findings? A novel dual-branch collaborative Transformer network (PST-Net) is proposed, which effectively addresses the challenge of difficult collaborative modeling between global dependencies and local details in hyperspectral image classification through the collaboration between an adaptive spectral-spatial Token module and a parallel attentionenhanced CNN branch. PST Net achieved excellent classification performance on four challenging datasets (Salinas, Whuhh, Qingyun, Houston) with only 2% labeled samples, outperforming multiple comparison methods. What is the implication of the main finding? This method demonstrates strong robustness and generalization ability even with minimal training samples, providing an effective solution for hyperspectral image classification tasks where labeled data is scarce in practical applications. The proposed adaptive Token construction and cross-level interaction fusion mechanism provides a general and efficient framework for designing hybrid models that synergistically leverage the advantages of Transformer and CNN.Abstract Hyperspectral image classification holds significant applications across multiple domains due to its rich spectral and spatial information. However, it faces challenges such as spectral variation within the same object, spectral variation across different objects, and noise interference. Existing methods like convolutional neural networks perform well in local feature extraction but inadequately model long-range dependencies. While Transformers can capture global relationships, they struggle to effectively coordinate spectral and spatial information modeling. To address these limitations, this paper proposes a dual-branch collaborative Transformer network (PST-Net). This architecture integrates an adaptive spectral-spatial token (ASST) module, a Parallel Attention-Augmented lightweight CNN branch (PA-SSCNN), and a collaborative fusion layer. The ASST constructs joint representation tokens through local spectral smoothing and learnable spatial embedding. PA-SSCNN employs 3D-2D cascaded convolutions and channel-spatial attention mechanisms to enhance local texture and spatial feature extraction; CHIB enables deep interaction and synergistic fusion of dual-branch features across different levels and scales. Experimental results demonstrate that with only 2% labeled samples, PST-Net achieves overall classification accuracies of 96.31%, 96.59%, 95.27%, and 89. 06% on the Salinas and Whuhh, and the two complex urban scene datasets Qingyun and Houston. Especially in fine-grained categories and complex scenes, it exhibits strong robustness. The ablation experiment further validated the effectiveness and complementarity of each module. This study provides an efficient collaborative modeling framework for hyperspectral image classification that balances global dependencies and local details.
With increasing demand for natural gas, the construction of natural gas extraction-related facilities has increased significantly. Accurate identification of these facilities is crucial for guiding spatial planning and evaluating environmental impacts. Existing research has primarily concentrated on offshore facilities, with limited attention to onshore facilities. This scarcity stems from identification challenges due to their dispersed distribution and complex environments. To address this gap, this study proposes a method combining a multimodal convolutional neural network (CNN) with object-based segmentation for onshore facility extraction. Experiments were conducted in northern Sichuan, China, with high-resolution Chinese satellite images, GF-2. Performance was compared between machine learning and CNN using sequentially cropped imageries. The proposed method achieved a precision of 59.97%, a recall of 94.87%, and an F1-score of 73.49%. The high recall indicates that most facilities were successfully detected, and the F1-score reflects the overall performance. These results suggest that the proposed method can effectively extract onshore facilities. Compared with machine learning and CNN using sequentially cropped imageries, the F1-score of the proposed method increased by 20.16% and 51.49%, respectively. The experimental results reveal that the proposed method can accurately identify onshore facilities, offering a scientific basis for assessing the environmental impact of greenhouse gases.
In recent years, convolutional neural networks (CNNs) and graph neural networks (GNNs) have been widely applied to hyperspectral remote sensing image classification. However, CNNs rely on fixed-size convolutional kernels for feature extraction, resulting in limited receptive fields that struggle to capture relationships between distant pixels in hyperspectral images. While GNN-based hyperspectral classification methods effectively mine features in non-Euclidean hyperspectral data, they face significant challenges in constructing graph structures. For instance, traditional graph convolutional networks typically rely on local neighborhood connections, making it difficult to directly model spatial relationships between pixels at extreme distances. This limits their ability to effectively extract long-range interaction information within images. To address these challenges, this article proposes GASE-Net, a network for hyperspectral image classification. Adopting a parallel architecture, it integrates a CNN with channel-spatial attention and a gated graph attention network, while incorporating noise reduction and feature enhancement modules at the front end to boost feature representation capabilities. The proposed network fuses features from two branches through an adaptive weighting mechanism. On the one hand, it models and captures spectral-spatial features from non-Euclidean hyperspectral images based on graph structures, effectively characterizing relationships between pixels at ultra-long distances. On the other hand, it employs a pixel-level CNN with a fused channel-spatial attention mechanism to extract fine-grained spatial features, thereby overcoming the inherent limitations of CNNs in modeling long-range relationships. Experiments on the Salina, WHU-Hi-HanChuan, WHU-Hi-HongHu, Pingan, Qingyun, and Indian Pines dataset, the proposed method achieves overall classification accuracies of 99.33%, 98.24%, 98.57%, 98.41%, 97.83%, and 94.02%, respectively. Compared with other models, GASE-Net demonstrates strong generalization capabilities and excellent classification performance in hyperspectral remote sensing image classification.
Based on studies using high-medium resolution images, convolutional neural networks (CNNs) and semantic segmentation have shown superiority over classical machine learning (ML), particularly in small-scale mapping. However, few/no studies have assessed the techniques on coarse resolution image classification for extensive area land cover mapping. In this study, we evaluated the performance and feasibility of three CNN models (1-D CNN, 2-D CNN, and 3-D CNN), and U-net for coarse-resolution satellite image classification and compared them to a random forest (RF) classifier. We utilized time-series, coarse resolution (1 km) composite imageries acquired by FengYun-3C visible and infrared radiometer. Labeled datasets were collected as shapefiles and split into three independent datasets: training, validation, and test datasets, and preprocessed to meet each model's input format requirements. We conducted several experiments to optimize models and select the best models. Then, the best models were evaluated on an unseen dataset. Among the DL models, one-dimensional (1-D) CNN achieved the highest overall accuracy (OA) 0. 87 and kappa (k) 0.84, 2% higher than the best results attained by 2-D CNN, 3-D CNN, and U-net models. However, 1-D CNN is outperformed by RF which achieved 0.89 (OA) and 0.87 (k). Achieving the best and the second-best results using RF and 1-D CNN models, respectively, indicates the superiority of the pixel-based method and the insignificance of spatial information in coarse-resolution image classification. Furthermore, although the DL models can yield high accuracy, especially 1-D CNN, they are less feasible than RF classifiers for coarse-resolution satellite image classification in extensive area land cover mapping.
Reconstruction of land surface temperature (LST) under clouds has been an area of significant research interest in recent years. Solar-cloud-satellite geometry has significant impacts on satellite-derived land surface biophysical parameters, such as radiation flux and LST; however, current studies often neglect these influences on reconstruction of cloudy LST. To address this challenge, we developed an integrated methodology for generating seamless all-weather LST based on surface energy balance (SEB) theory with consideration of the solar-cloud-satellite geometry effects both on LST and radiation. Cloudy pixels were categorized (radiation-unobstructed and radiation-obstructed clouds) and reconstructed separately to account for geometry effects. Moreover, corrections were incorporated to mitigate geometry effects on net surface shortwave radiation (NSSR), the crucial intermediate input data for estimating cloudy LST. Compared to the existing method, validation results using ground measurements from the Surface Radiation Budget (SURFRAD) network demonstrate significant improvements, with average errors decreasing from 5.62 to 1.86 K under radiation-unobstructed conditions and from 3.26 to 1.33 K under radiation-obstructed conditions, respectively. This study contributes valuable insights to reconstructing LST under varying cloudy conditions, indicating the importance of considering geometry effects for robust and reliable cloudy LST assessments.
Drought has been putting enormous pressure on agriculture and food security, especially for African countries with inadequate monitoring. However, there is still a lack of standardization in the synthesized drought index, which could better monitor drought than normalized indices, as well as insufficient attention to the geographical backgrounds of weights. We aim to develop a standardized optimal synthesized drought index (SOSDI) that will improve drought monitoring and reflect dominant factors. The standardization transforms the input variable distribution functions, with the constraint optimization to compute the weight. Weight matrices are suggested to examine the geospatial heterogeneity of the input variable. The assessment of SOSDI from several dimensions was implemented. SOSDI was applied to monitor the periodic drought in East Africa, where rain-fed agriculture was susceptible to seasonal alternation. The results confirmed that standardization was more effective than normalization in the synthesized drought index, and the average correlation between SOSDI with the existing indices, soil moisture, and crop yield reached about 0.6, 0.7, and 0.7, respectively. The cyclical drought is well captured on the seasonal and annual scales, mostly during the long-dry season. SOSDI could capture drought evolution and had good resistance to disturbances. Overall, SOSDI demonstrated a strong capability in drought monitoring and enhanced the drought expression derived from vegetation and temperature more effectively, rather than just precipitation.
Convolutional neural networks have made significant progress in multimodal remote sensing image classification, but traditional convolutional neural networks are limited by fixed-size convolutional kernels, which are unable to effectively model and adequately extract contextual information; hyperspectral imagery and LiDAR data have comparatively large information differences, which do not allow for effective information interaction and fusion. Based on this, this paper proposes a multimodal dual fusion network (CSTC) based on the Vision Transformer for the collaborative classification of HSI and LiDAR data. The model is designed through a two-branch architecture: the HSI branch extracts spectral–spatial features by dimensionality reduction using principal component analysis and inputs them into the cross-connectivity feature fusion module; the LiDAR branch mines spatial elevation features through the stacked MobileNetV2 module. The features of the two branches are encoded by a Transformer, and the modal interaction fusion is realized by the cross-attention module for the first time. Then, the features are spliced and input into the secondary Transformer for deep cross-modal fusion, and finally, the classification is completed by the multilayer perceptron. Experiments show that the CSTC model achieves overall classification accuracies of 92.32%, 99.81%, 97.90%, and 99.37% on the publicly available MUUFL dataset, Trento dataset, Augsburg dataset, and Houston2013 dataset, respectively, which is superior to the latest HSI–LiDAR separate classification algorithms. The ablation experiments and model performance evaluation experiments further show that the proposed CSTC model achieves excellent results in terms of robustness, adaptability, and parameter scale.
Multimodal remote sensing image classification is widely used in land refinement classification, but cross modal feature fusion faces challenges. To address this issue, this paper proposes a three branch attention fusion network (TMFN) for HSI and LiDAR/SAR images. The network first achieves early information exchange through feature extraction branches and lightweight fusion modules, while maintaining model lightweighting while ensuring complementary spectral spatial features; Obtain local and global spatial dependencies through the backbone network of parallel Transformer and multi-scale dual attention convolution module; Introducing an adaptive SE attention fusion module when integrating multimodal features to achieve cross modal semantic deep level feature fusion. Our model achieved overall classification accuracy of $88.99 \%, 99.63 \%$, and 74.89 % on three typical open-source datasets, Muufl, Trento, and Berlin, respectively, outperforming representative algorithms in recent years. The ablation experiment further validated the effectiveness of each module. This method not only improves classification accuracy but also considers model running efficiency and scalability, providing a new solution for multi-source remote sensing image fusion.
The precise estimation of sugarcane yield at the field scale is urgently required for harvest planning and policy-oriented management. Sugarcane yield estimation from satellite remote sensing is available, but satellite image acquisition is affected by adverse weather conditions, which limits the applicability at the field scale. Secondly, existing approaches from remote sensing data using vegetation parameters such as NDVI (Normalized Difference Vegetation Index) and LAI (Leaf Area Index) have several limitations. In the case of sugarcane, crop yield is actually the weight of crop stalks in a unit of acreage. However, NDVI’s over-saturation during the vigorous growth period of crops results in significant limitations for sugarcane yield estimation using NDVI. A new sugarcane yield estimation is explored in this paper, which employs allometric variables indicating stalk magnitude (especially stalk height and density) rather than vegetation parameters indicating the leaf quantity of the crop. In this paper, UAV images with RGB bands were processed to create mosaic images of sugarcane fields and estimate allometric variables. Allometric equations were established using field sampling data to estimate sugarcane stalk height, diameter, and weight. Additionally, a stalk density estimation model at the pixel scale of the plot was created using visible light vegetation indices from the UAV images and ground survey data. The optimal stalk density estimation model was applied to estimate the number of plants at the pixel scale of the plot in this study. Then, the retrieved height, diameter, and density of sugarcane in the fields were combined with stalk weight data to create a model for estimating the sugarcane yield per plot. A separate dataset was used to validate the accuracy of the yield estimation. It was found that the approach presented in this study provided very accurate estimates of sugarcane yield. The average yield in the field was 93.83 Mg ha−1, slightly higher than the sampling yield. The root mean square error of the estimation was 6.63 Mg ha−1, which was 5.18% higher than the actual sampling yield. This study offers an alternative approach for precise sugarcane yield estimation at the field scale.
To investigate effective techniques for estimating rice production in hilly and mountainous areas, in this study, we collected yield data at the field level, agro-meteorological data, and Sentinel-2/MSI remote sensing data in Chongqing, China, between 2020 and 2023. The integral values of vegetation indicators from the rice greening up to heading–filling stages were determined using the Newton–trapezoidal integration method. Using correlation analysis and importance analysis of permutation features, the effects of agro-meteorological variables and vegetation index integrals on rice yield were assessed. The chosen characteristics were then combined with three machine learning techniques—random forest (RF), support vector machine (SVM), and partial least squares regression (PLSR)—to create six rice yield estimate models. The results showed that combined vegetation indices were more effective than indices used in separate development phases. Specifically, the correlation coefficients between the integral values of eight vegetation indices from rice greening up to heading–filling stages and rice yield were all above 0.65. By introducing agro-meteorological factors as new independent variables and combining them with vegetation indices as input parameters, the predictive capability of the model was evaluated. The results showed that the performance of PLSR remained stable, while the prediction accuracies of SVM and RF improved by 13% to 21.5%. After feature selection, the inversion performance of all three machine learning models improved, with the RF model coupled with variables selected during permutation feature importance analysis achieving the optimal inversion effect, which was characterized by a coefficient of determination of 0.85, a root mean square error of 529.1 kg/hm2, and a mean relative error of 5.63%. This study provides technical support for improving the accuracy of remote sensing-based crop yield estimation in hilly and mountainous regions, facilitating precise agricultural management and informing agrarian decision making.
Land surface temperature (LST) is a critical parameter in global long-term meteorological and climatological studies. The Visible and Infrared Radiometer (VIRR) sensor aboard the Chinese Fengyun-3 (FY-3) series satellites provides a continuous collection of thermal infrared (TIR) data, facilitating the generation of global long-term LST products. Notably, the FY-3C VIRR has served as a key instrument in collecting global TIR data since 2013. In this study, we proposed an operational split-window (SW) algorithm for retrieving LST from FY-3C VIRR TIR data. Initially, the Thermodynamic Initial Guess Retrieval 2000 atmospheric profile library, the atmospheric transfer model MODTRAN, and the Advanced Spaceborne Thermal Emission and Reflection Radiometer (ASTER) spectral library were employed to construct a simulation database for fitting the algorithm coefficients. To enhance the accuracy of the SW algorithm, LST, atmospheric water vapor content (WVC), and average emissivity were segmented into various subranges. Subsequently, land surface emissivity (LSE) was dynamically estimated by combining the ASTER Global Emissivity Database (GED) with the normalized difference vegetation index threshold approach. Finally, in situ measurements from the Surface Radiation Budget (SURFRAD) network, along with the Moderate Resolution Imaging Spectroradiometer (MODIS) LST products (MOD11A1 and MOD21A1), were utilized to evaluate the accuracy of the retrieved LSTs. The results indicate: 1) the retrieved LSTs showed a high correlation with the in situ LSTs, with a coefficient of determination of 0.94, a root-mean-square error (RMSE) of 2.6 K, and a bias of 0.3 K; 2) the retrieved LSTs were consistent with MODIS LST products, showing a root-mean-square difference (RMSD) of approximately 2.4 K; and 3) compared to the result of MOD11A1, MOD21A1 exhibited a significantly smaller bias. These results indicated that the proposed algorithm is effectively capable of estimating global LST from FY-3C VIRR TIR data with reasonable accuracy.
Temporally incomparability across the scan lines in polar-orbiting satellite-derived land surface temperature (LST) affects their widespread application. Some challenges persist in the existing research on this issue, such as the absence of a universal algorithm applicable for the partly clear-sky condition in the daytime and scale inconsistency of the used datasets, when LST varies nonlinearly over time. Given this situation, we proposed an improved approach for temporal normalization, integrating ensemble regression models and a new LST variation rate model (RM), which captures typical LST variation characteristics over time during polar-orbiting satellite overpass periods. The Aqua Moderate Resolution Imaging Spectroradiometer (MODIS) LST data across the contiguous United States (CONUS) were collected to investigate its effectiveness. Moreover, cross-validation was conducted using the time-interpolated Geostationary Operational Environmental Satellite 16 (GOES-16) advanced baseline imager (ABI) LST. The normalized LST had remarkable consistency with the GOES-16 LST, with superior accuracy in contrast with the original LST. The root-mean-squared error (RMSE) was improved by approximately 1.56 K, and bias was enhanced up to 1.80 K. This study exhibited relatively superior performance in terms of quantitative outcomes and spatial distribution of LST compared with the previous studies. These evaluations indicate that the proposed method could be a dependable and general solution for addressing temporal inconsistencies in clear-sky LST during polar-orbiting satellite overpass periods.
The prompt acquisition of precise land cover categorization data is indispensable for the strategic development of contemporary farming practices, especially within the realm of forestry oversight and preservation. Forests are complex ecosystems that require precise monitoring to assess their health, biodiversity, and response to environmental changes. The existing methods for classifying remotely sensed imagery often encounter challenges due to the intricate spacing of feature classes, intraclass diversity, and interclass similarity, which can lead to weak perceptual ability, insufficient feature expression, and a lack of distinction when classifying forested areas at various scales. In this study, we introduce the DASR-Net algorithm, which integrates a dual attention network (DAN) in parallel with the Residual Network (ResNet) to enhance land cover classification, specifically focusing on improving the classification of forested regions. The dual attention mechanism within DASR-Net is designed to address the complexities inherent in forested landscapes by effectively capturing multiscale semantic information. This is achieved through multiscale null attention, which allows for the detailed examination of forest structures across different scales, and channel attention, which assigns weights to each channel to enhance feature expression using an improved BSE-ResNet bilinear approach. The two-channel parallel architecture of DASR-Net is particularly adept at resolving structural differences within forested areas, thereby avoiding information loss and the excessive fusion of features that can occur with traditional methods. This results in a more discriminative classification of remote sensing imagery, which is essential for accurate forest monitoring and management. To assess the efficacy of DASR-Net, we carried out tests with 10m Sentinel-2 multispectral remote sensing images over the Heshan District, which is renowned for its varied forestry. The findings reveal that the DASR-Net algorithm attains an accuracy rate of 96.36%, outperforming classical neural network models and the transformer (ViT) model. This demonstrates the scientific robustness and promise of the DASR-Net model in assisting with automatic object recognition for precise forest classification. Furthermore, we emphasize the relevance of our proposed model to hyperspectral datasets, which are frequently utilized in agricultural and forest classification tasks. DASR-Net’s enhanced feature extraction and classification capabilities are particularly advantageous for hyperspectral data, where the rich spectral information can be effectively harnessed to differentiate between various forest types and conditions. By doing so, DASR-Net contributes to advancing remote sensing applications in forest monitoring, supporting sustainable forestry practices and environmental conservation efforts. The findings of this study have significant practical implications for urban forestry management. The DASR-Net algorithm can enhance the accuracy of forest cover classification, aiding urban planners in better understanding and monitoring the status of urban forests. This, in turn, facilitates the development of effective forest conservation and restoration strategies, promoting the sustainable development of the urban ecological environment.
Convolutional Neural Networks (CNNs) have been applied effectively to classify high to medium-resolution imageries, especially for local scale land cover mapping, and attained high accuracy by outperforming the conventional machine learning techniques as the models incorporate spatial information as well as spectral reflectance values to differentiate objects.However, few studies have explored the effectiveness of CNN models for coarse-resolution satellite image classification. In this work, therefore, we applied 1D CNN and 3D CNN for time-series, coarse resolution (1km) FY-3C image classification in extensive area land cover mapping of a part of Eastern and North-East Africa.The result indicates both 1D CNN and 3D CNN models achieved high overall accuracy (OA >=85), although the former outperformed the latter significantly by 2-4%, indicating the superiority of pixel-based classification in time-series coarse resolution image classification.
In remote sensing image processing, when categorizing images from multiple remote sensing data sources, the deepening of the network hierarchy is prone to the problems of feature dispersion, as well as the loss of semantic information. In order to solve this problem, this paper proposes to integrate a parallel network architecture HDAM-Net algorithm with a hybrid dual attention mechanism Hybrid dual attention mechanism for forest land cover change. Firstly, we propose a fusion MCA + SAM (MS) attention mechanism to improve VIT network, which can capture the correlation information between features; secondly, we propose a multilayer residual cascade convolution (MSCRC) network model using Double Cross-Attention Module (DCAM) attention mechanism, which is able to efficiently utilize the spatial dependency between multiscale encoder features: the spatial dependency between multiscale encoder features. Finally, the dual-channel parallel architecture is utilized to solve the structural differences and realize the enhancement of forestry image classification differentiation and effective monitoring of forest cover changes. In order to compare the performance of HDAM-Net, mountain urban forest types are classified based on multiple remote sensing data sources, and the performance of the model is evaluated. The experimental results show that the overall accuracy of the algorithm proposed in this paper is 99.42%, while the Transformer (ViT) is 96.92%, which indicates that the proposed classifier is able to accurately determine the cover type.The HDAM-Net model emphasizes the effectiveness in terms of accurately classifying the land, as well as the forest types by using multiple remote sensing data sources for predicting the future trend of the forest ecosystem. In addition, the land utilization rate and land cover change can clearly show the forest cover change and support the data to predict the future trend of the forest ecosystem so that the forest resource survey can effectively monitor deforestation and evaluate forest restoration projects.