Large-scale algal blooms attack Taihu Lake frequently in recent years. In this study, the hourly coverage area of algal blooms in Taihu Lake is detected using Geostationary Ocean Color Imager (GOCI) images and its relationship with environmental factors is analyzed based on the measurements including fluorescent dissolved organic matter(fDOM), water temperature (Tw), hourly air temperature (Ta), hourly maximum air temperature (Tm) and wind speed (V). Other remote sensing products are also employed, including photosynthetically active radiation (PAR). The data in this study covers a period from March 2020 to March 2021 with a temporal resolution of one hour. Tw, Ta, and Tm exhibit a positive relationship with the coverage area of algal blooms in spring, autumn, and winter, and a negative relationship in summer. In winter, the coverage area of algal blooms is more sensitive to water temperature (r = 0.85). A method was proposed to investigate the coverage area of algal blooms in Taihu Lake based on environmental factors. The maximum likelihood estimation method was adopted in this proposed method, and the results were determined using BIC methods. The coefficients of determination (R2) for the regression model are 0.75, 0.74, 0.57, 0.83 in spring, summer, autumn and winter, respectively. The algal blooms happen in a fDOM ranging from 12 to 20 QSU, with highest Tw of 33 °C (37 °C for Ta and Tm) and the highest wind speed of 9 m/s. The wind speed is negatively correlated with the coverage area of algal blooms, while the PAR is positively correlated. When water temperature is below 26.86 °C, an increase in temperature promotes the growth of algal blooms(r = 0.71). However, the growth of the algal bloom becomes inhibited when water temperature exceeds 26.86 °C (r = -0.33) or the fDOM concentration exceeds 15.97 QSU.
This study proposes an integrated pixel-level reflectance adjustment (IPRA) method using Sentinel-2 MSI as the reference to address radiometric discrepancies in GF-1/6 WFV imagery, particularly caused by sensor decay and geometric distortions. The proposed IPRA method leverages time-series data and a spatial heterogeneity detection mechanism to effectively mitigate geometric distortions. Furthermore, it incorporates a weighted linear regression (WLR) model to weight pixels based on their temporal decay characteristics. The results demonstrate that IPRA outperforms existing methods (i.e., IRMAD, HM, and TRA) in radiometric consistency, yielding smaller radiometric discrepancies relative to Sentinel-2 MSI. Specifically, NAE decreased by 42.9% (from 0.319 to 0.182), RMSE decreased by 37.3% (from 0.051 to 0.032), PSNR improved from 25.906 dB to 30.195 dB, and the SC value approached the ideal value of 1 (from 1.540 to 1.001). In conclusion, the IPRA method provides a robust solution for normalizing GF-1/6 WFV imagery and thus facilitates its cross-sensor applications.
Satellite observations of tropospheric formaldehyde (HCHO) have been widely used to support diagnosing atmospheric environmental quality. As one of the most classic trace gas payloads, the Ozone Monitoring Instrument (OMI) releases HCHO Level-2 data, while coarse resolution and relatively high uncertainty reduce the potential value of the data. We report a global multi-scale HCHO oversampling dataset version 1.0 (referred to as OMHCHOS V1.0) produced by NASA Level-2 OMI HCHO product using self-developed oversampling algorithm, with data from 2005 to 2023 as of the date of submission. This comprehensive dataset encompasses seven distinct spatial resolutions (up to 0.05°) and twelve temporal resolutions (monthly and months), enabling precise quantification of uncertainty propagation and relative uncertainties. To facilitate on-demand retrieval by users, we have developed a matching spatio-temporal scale optimisation model that integrates three critical parameters of HCHO column: temporal resolution (TR), spatial resolution (SR), and relative uncertainty (UR). This dataset will provide researchers with more reliable sources for conducting high-resolution, high-accuracy studies on HCHO-related atmospheric environmental implications.
China faces significant carbon emission challenges. Given the unbalanced development among cities and the significant differences in emission characteristics across sectors, timely and high-resolution monitoring of sectoral carbon emissions is necessary for achieving China's carbon peaking and carbon neutrality goals. However, most existing carbon emission products rely heavily on statistical inventories, leading to time lags, low spatial resolution, and limited capability to monitor carbon emissions from specific sectors. This study proposes a carbon emission estimation model that enables inventory-independent sectoral emission estimation by leveraging readily accessible, stable, and continuous remote sensing data. First, provincial carbon emissions from BeijingTianjinHebei (BTH) in 2018 were spatialized to a 500-m monthly resolution using nighttime light (NTL), electricity consumption (EC), and land-use data. Then, a carbon emission estimation model, a remote sensing and machine-learning-based model for timely updating sectoral carbon emission (RSML-TUSCE), was developed based on extreme gradient boosting (XGBoost), which extracts emissions from 12 indicators encompassing spatial location, human activity metrics, and environmental and climatic variables. The model achieves an R-2 of 0.98 and demonstrates strong temporal robustness. Finally, without relying on inventory data, RSML-TUSCE was used to estimate the sectoral carbon emissions of the BTH region in 2019 with a 500-m resolution. The results show that the model can effectively identify sector-wise carbon emission hotspots, capture the spatiotemporal variation of emissions, and enable timely updates of high-resolution data. The model application enhances our understanding of the spatiotemporal variations in sector-specific carbon emissions, providing support for formulating emission reduction policies across cities and sectors.
Highlights What are the main findings? In-depth characterization of the spatial scale effects inherent in high-resolution surface BRDF. Explicit analysis of the linkage between BRDF scale dependence and surface spatial heterogeneity. Quantitative determination of optimal observation scales for distinct terrestrial targets. What are the implications of the main findings? Construction of prior BRDF knowledge for typical land features at their optimal scales to fulfill diverse remote sensing applications. These findings advance the practical application of remote sensing technology, offering significant implications for enhancing target detection accuracy and the authenticity of remote sensing products.Highlights What are the main findings? In-depth characterization of the spatial scale effects inherent in high-resolution surface BRDF. Explicit analysis of the linkage between BRDF scale dependence and surface spatial heterogeneity. Quantitative determination of optimal observation scales for distinct terrestrial targets. What are the implications of the main findings? Construction of prior BRDF knowledge for typical land features at their optimal scales to fulfill diverse remote sensing applications. These findings advance the practical application of remote sensing technology, offering significant implications for enhancing target detection accuracy and the authenticity of remote sensing products.Abstract Land surface bidirectional reflectance distribution functions (BRDF) are critical for quantitative remote sensing but are significantly constrained by scale effects, limiting the interoperability of multi-resolution data and the accuracy of quantitative inversion, thereby rendering the investigation of BRDF multi-scale effects increasingly urgent. This study utilized UAV (Unmanned Aerial Vehicle)-based multi-angular observations and the RPV model to retrieve the BRDF of typical land covers, employing the Window Averaging Method to simulate multi-scale responses and systematically investigate the relationship between BRDF characteristics and spatial scale. The results indicate the following key findings: (1) The RPV (Rahman-Pinty-Verstraete) model demonstrated high robustness and inversion accuracy, yielding RMSE (Root Mean Square Error) below 0.06 and RRMSE (Relative RMSE) below 25% across all land covers, with the 840 nm band exhibiting superior performance. (2) Significant spatial scale effects were observed, where BRDF characteristics varied distinctively with scale but eventually stabilized at specific thresholds; specifically, the stabilization scales were identified as 1.3 m for bare soil, 1.5 m for tea plantations, 1 m for rice, and 2 m for forests. (3) The scale evolution of BRDF features exhibited a parallel trend with spatial heterogeneity, a correlation that enables the quantitative identification of optimal observation scales for different land cover types.
This paper addresses the critical challenge of semantic segmentation for remote sensing images (RSIs) under extremely limited labeled data. A high-quality initial model is paramount for downstream semi-supervised or weakly supervised learning paradigms, as it mitigates error propagation from the outset. We conducted a systematic investigation into self-supervised pretraining to serve this precise need. Within the low-label regime, we identify and tackle two pivotal factors limiting performance: (1) the domain shift between large-scale pretraining data and specific target tasks, and (2) the deficiency in local feature learning caused by large-window masking in visual foundation model (VFM) pretraining. To resolve these issues, we first benchmark various pretraining strategies, demonstrating that a two-phase General-Purpose Pretraining (GPPT) followed by Domain-Adaptive Pretraining (DAPT) framework is optimal, significantly outperforming both single-phase methods and the existing two-phase paradigm initialized from ImageNet. Subsequently, we propose an Edge-Guided Masked Image Modeling (EGMIM) method for the DAPT phase, which explicitly integrates edge priors to guide the masking and reconstruction process, thereby enhancing the model’s capability to capture fine-grained local structures. Extensive experiments on four RSI benchmarks validate the effectiveness of our approach, showing consistent and substantial gains, particularly in extreme low-label scenarios. Beyond empirical results, we provide in-depth mechanistic analyses to explain the synergistic roles of GPPT and DAPT.
Fusing optical imagery with complementary modalities (X-modality), such as light detection and ranging (LiDAR) and synthetic aperture radar (SAR), is essential for robust semantic segmentation in complex environments. Although recent modality-agnostic models improve generalizability beyond fixed-pair methods, they still lack a unified and efficient architecture to process diverse combinations, including RGB-X, multispectral (MSI-X), and hyperspectral (HSI-X). To address this, we propose UniMamba, a unified cross-modal fusion Mamba framework for general multimodal semantic segmentation. Specifically, Uni- Mamba employs a unified encoder that processes diverse 2D–3D inputs without modality-specific modifications. The encoder integrates HyperMamba blocks to jointly model spatial, spectral, and frequency features while capturing global context with linear complexity. Each encoding stage further incorporates a Cross HyperMamba (CroHMa) module for explicit cross-modal interaction, seamlessly followed by a Fusion Mamba (FuseMa) module for complementary feature fusion. Extensive experiments on three benchmarks demonstrate that UniMamba outperforms state-of-the-art methods and generalizes robustly across diverse multimodal settings. Code is available at .
The rapid advancement of remote sensing (RS) technology has posed increasing demands for secure and efficient processing of multispectral data. However, conventional joint image encryption and compression schemes, originally developed for natural images, are not well suited to the specific requirements of multispectral RS scenarios, such as managing interband redundancy, capturing spatial texture variation, and preserving spectral consistency for downstream applications. To address these challenges, we propose a joint encryption and compression algorithm for multispectral images (JECA-MS), the first joint encryption and compression framework specifically designed for multispectral RS images with support for ciphertext domain recompression. The JECA-MS incorporates four key innovations: 1) an adaptive two-size texture block decision (TBD) strategy that classifies image regions into strong and weak texture blocks (WTBs), reducing data volume in weak-texture areas by up to fourfold; 2) a modified 3-D discrete cosine transform (3D-MDCT) that enhances spatial-spectral decorrelation, particularly in homogeneous regions such as clouds and water; 3) a ciphertext domain recompression mechanism that enables flexible adjustment of compression ratios (CRs) without decryption; and 4) a dedicated JECA-MS coding format (JECA-MS-CF) for efficient data encapsulation and compatibility with RS data structures. Extensive experiments show that the JECA-MS achieves 55% and 36% improvements in CRs for water and cloud images, while reducing encoding and decoding time by 39% and 68%, compared to state-of-the-art methods. Security evaluation shows that the JECA-MS can resist statistical attacks, achieve a tradeoff between lightweight encryption and compression performance. This work offers a flexible solution for secure and efficient RS data management.
High-resolution remote sensing images often suffer from inadequate fusion between global and local features, leading to the loss of long-range dependencies and blurred spatial details, while also exhibiting limited adaptability to multi-scale object segmentation. To overcome these limitations, this study proposes RST-Net, a semantic segmentation network featuring a dual-branch encoder structure. The encoder integrates a ResNeXt-50-based CNN branch for extracting local spatial features and a Shunted Transformer (ST) branch for capturing global contextual information. To further enhance multi-scale representation, the multi-scale feature enhancement module (MSFEM) is embedded in the CNN branch, leveraging atrous and depthwise separable convolutions to dynamically aggregate features. Additionally, the residual dynamic feature fusion (RDFF) module is incorporated into skip connections to improve interactions between encoder and decoder features. Experiments on the Vaihingen and Potsdam datasets show that RST-Net achieves promising performance, with MIoU scores of 77.04% and 79.56%, respectively, validating its effectiveness in semantic segmentation tasks.
Download This Paper Open PDF in Browser Add Paper to My Library Share: Permalink Using these links will ensure access to this page indefinitely Copy URL Copy DOI
Large-area medium-resolution land cover classification is a key information for monitoring land use, human activity influence on ecological environments, etc. However, the existing land cover classification products lack insights of the utilization of multi-source data and regional characteristics, thus facing accuracy shackles. In this work, we supplemented SAR and DEM data into classification procedure, and gathered a comprehensive set of features, including global ecological zones (GEZs) announced by FAO as an indicator. With a multi-source land cover point label dataset for Mekong basin (LanCoMe) with a revised classification system, a random forest model was trained to produce Mekong land cover classification mappings. The model achieves validation accuracy of 91.3%, and the product is well correlated with other published global land cover products. The result also indicates that GEZs can be a prior feature for large-area land cover tasks.
Download This Paper Open PDF in Browser Add Paper to My Library Share: Permalink Using these links will ensure access to this page indefinitely Copy URL Copy DOI
Land surface temperature (LST) plays a crucial role in characterizing land surface processes and the energy balance. Accurate monitoring of spatial-temporal variations in LST holds great significance for global and regional climate change research. However, with the increasing of multi-source remote sensing data, it has become a challenging task to construct an LST model that reduces dependence on auxiliary data (such as land surface emissivity (LSE) and atmospheric water vapor content (WVC) synchronized with satellite observations. This study proposed an LST retrieval model based on random forest model (RFLSTR). The results of 10-fold cross-validation showed that the RFLSTR model had high modelling accuracy, the coefficient of determination (R2) was 0.99, and the root mean square error (RMSE) and mean absolute error (MAE) were lower than 2 K. Landsat 8 images from the Beijing area and South Korea, were utilized to assess the spatial-temporal migration performance of the RFLSTR model. Additionally, Landsat 9 images from South Korea were employed to analyse the sensor migration performance of the RFLSTR model. In conclusion, the findings highlight that the proposed RFLSTR model achieve high retrieval accuracy without relying on auxiliary data (such as LSE and WVC) synchronized with satellite observations. Moreover, the RFLSTR model exhibits robust spatial-temporal capabilities, enhancing the efficiency of LST retrieval and facilitating the utilization of satellite data.
Hyperspectral image (HSI) classification has recently reached its performance bottleneck. Multimodal data fusion is emerging as a promising approach to overcome this bottleneck by providing rich complementary information from the supplementary modality (X-modality). However, achieving comprehensive cross-modal interaction and fusion that can be generalized across different sensing modalities is challenging due to the disparity in imaging sensors, resolution, and content of different modalities. In this study, we propose a Local-to-Global Cross-modal Attention-aware Fusion (LoGoCAF) framework for HSI-X classification that jointly considers efficiency, accuracy, and generalizability. LoGoCAF adopts a pixel-to-pixel two-branch semantic segmentation architecture to learn information from HSI and X modalities. The pipeline of LoGoCAF consists of a local-to-global encoder and a lightweight multilayer perceptron (MLP) decoder. In the encoder, convolutions are used to encode local and high-resolution fine details in shallow layers, while transformers are used to integrate global and low-resolution coarse features in deeper layers. The MLP decoder aggregates information from the encoder for feature fusion and prediction. In particular, two cross-modality modules, the feature enhancement module (FEM) and the feature interaction and fusion module (FIFM), are introduced in each encoder stage. The FEM is used to enhance complementary information by combining the feature from the other modality across direction-aware, position-sensitive, and channel-wise dimensions. With the enhanced features, the FIFM is designed to promote cross-modality information interaction and fusion for the final semantic prediction. Extensive experiments demonstrate that our LoGoCAF achieves superior performance and generalizes well. The code will be made publicly available.
Aerosol is an important atmospheric component that severely influences the global climate and air quality of our planet [...]
The utilization of remote sensing soil moisture products in agricultural and hydrological studies is on the rise. Conducting a regional applicability analysis of these soil moisture products is essential as a preliminary step for their effective utilization. The triple collocation (TC) method enables the estimation of the standard deviation of errors in products when true soil moisture values are unavailable. It assesses data uncertainty and mitigates the influence of product errors on fusion, thereby enhancing product accuracy significantly. In this study, the TC uncertainty error analysis was employed to integrate Soil Moisture Active Passive (SMAP), the Advanced Microwave Scanning Radiometer 2 (AMSR-2), and the European Space Agency Climate Change Initiative (ESA CCI) active (ESA CCI A) and passive (ESA CCI P) products, with ground-based measurements serving as a reference. Traditional evaluation metrics, such as the correlation coefficient (R), bias, root mean square error (RMSE), and unin situed root mean square error (ubRMSE), were employed to evaluate the accuracy of the product. The findings indicate that SMAP and ESA CCI P products demonstrate strong spatiotemporal continuity within the research area and exhibit low uncertainty across various land types. The products derived from the Advanced Microwave Scanning Radiometer 2 (AMSR-2) exhibit a high level of temporal and spatial continuity; however, there is a requirement for enhancing their accuracy. The products of ESA CCI A exhibit notable spatiotemporal disjunction, contributing significantly to their elevated level of uncertainty. After fusion with TC analysis, the correlation coefficient (R = 0.7) of the TC-2 product derived from the fusion of SMAP, AMSR-2, and ESA CCI P products is significantly higher than the correlation coefficient of the TC-1 product (R = 0.65) obtained from the fusion of SMAP, AMSR-2, and ESA CCI A products at a 95% confidence level. The integration of data can efficiently mitigate the challenges associated with spatiotemporal gaps and inaccuracies in products, offering a dependable foundation for the subsequent utilization of remote sensing products.
Direct estimation of PM2.5 concentration using the satellite top-of-atmosphere (TOA) reflectance has become a research hotspot for satellite sensing monitoring of PM2.5. However, the optimal feature parameter selection in current studies on PM2.5 estimation based on TOA reflectance is still unclear, and the model construction is mostly based on a single model, and the accuracy is limited. Therefore, this study used various combinations of full-spectrum data, angle information and meteorological data from the Himawari-8 satellite containing TOA reflectance to explore the optimal feature parameter selection for estimating PM2.5, and removed the band data with multicollinearity in the full-spectrum data by calculating the variance inflation factor (VIF) of the full-spectrum satellite data. Concurrently, based on the ensemble stacking algorithm, a PM2.5 ensemble stacking estimatiion model (PMISES) with a double-layer structure and multiple machine learning models was constructed. The results indicated that the PM2.5 concentration estimation by satellite full-spectrun data after the multicollinearity test combined with angle information and meteorological data could improve the model accuracy. Compared with using satellite full-spectral data alone, the model R2 increased by 0.1 and RMSE decreased by 7 μg/m3. The combination was determined to be the optimal parameter combination for estimating PM2.5. The PMISES model developed in this study demonstrates superior accuracy when compared to other standalone models (RF, LightGBM, XGBoost). The R2 reaches 0.87 and the RMSE is 17.09 μg/m3. The model was utilized for the estimation and analysis of PM2.5 concentration in the Beijing-Tianjin-Hebei (BTH) region, and the outcomes exhibited consistent concordance with ground-based monitoring data. Obviously, the optimal PM2.5 concentration estimation method proposed in this study is reliable and can provide a valuable reference for PM2.5 concentration monitoring in BTH region.
Spatiotemporal fusion is a method of fusing high spatial resolution low temporal resolution remote sensing images and low spatial resolution high temporal resolution in order to obtain high spatiotemporal resolution remote sensing images, which can provide data support for temporal observation of fine objects, and plays an important role in the fields of Earth sciences, environmental monitoring, and so on. This article reveals an issue that is often overlooked in the field of deep learning-based spatiotemporal fusion: the discontinuity between image blocks and image blocks. This discontinuity may have an impact on the visualization of remote sensing images and subsequent applications. In this regard, this article proposes a spatially seamless stitching approach to optimize the spatiotemporal fusion model based on deep learning. By using this method, we successfully obtain high-quality fused remote sensing images with smoother transitions. The spatiotemporal fusion model used in the experiment is a generative adversarial network-based spatiotemporal fusion model (GAN-STFM) and the data are from the Beijing Gaofen-6 dataset (BJGF6). After our splicing method, the ratio of root-mean-square error (RMSE) at the splicing seam to the overall RMSE is reduced from 1.28 to 0.99, which effectively improves the continuity of the image. This new image splicing method has the potential to improve the utility of deep learning-based spatiotemporal fusion algorithms, which has application value for generating large-scale long time series remote sensing datasets with high temporal and high spatial resolution.
Intelligently sharing and reusing the knowledge developed by application practices are the keys to break through technical barriers and fully activate and release the effectiveness of Earth Observations(EO).The Earth Observation Knowledge Hub(EOKH),which is used to couple decentralized knowledge bases organically,is a research frontier for the governance and intelligent service of global EO applications.In the design of Group on Earth Observations(GEO),the GEO Knowledge Hub(GKH)is intended to provide authoritative,validated,and reproducible content for evidence-based reporting on policy commitments and decision-making.Thus,the GKH offers a platform for users to discover,learn about,and employ methods,analytical tools,and applications;it also provides opportunities for the GEO community to collaborate and provide mutual assistance related to GKH contents.However,important lessons,such as the sensitive issue of intellectual ethics and how to profit GKH from the recent technological advances in information technologies,have been learned during the implementations.In response to the problems and insights encountered by GEO in developing GKH,we systematically analyzed the connotation and characteristics of EOKH and sorted out the fundamental needs and challenges for the development of EOKH in China. First,we systematically analyzed the connotation and characteristics of EOKH.The study argues that EOKH is the intersection node of high-throughput trusted EO knowledge in knowledge-sharing networks.It has three typical features,i.e.,connectivity,high throughput,and lightweight computing.Its core mission is to identify and transfer valuable research in a timely manner and to promote high throughput of application packages.Second,we sorted out the fundamental needs and challenges for the development of EOKH in China.Considering the latest progress in the study of EO ontology,we also analyzed the possible key technical problems and gave strategies to cope with them.On this basis,the system architecture prototype of EOKH,which is drawn on the system design concept of representational state transfer,is proposed,and an ontology model of conceptual EO knowledge and a formalization model of process-oriented EO knowledge are established. The study argues that EOKH should be in an open collaborative environment where humans are in the loop.The key technologies are system metrics,knowledge transfer,knowledge reuse,and knowledge exploration and visualization.The transfer and reuse of knowledge packages can greatly enhance the ease of development and reuse of EO application practices.Ontology modeling helps formalize the intrinsic connection between the human-cyber-physical systems of EO application and enhances the interpretability of higher-order complex problems.EOKH transforms knowledge sharing from point to point at the element level to group collaborative at the system level,which not only reduces the cost of cross-industry and socially integrated EO applications but also avoids repetitive research inputs,helps break through technical barriers and cognitive obstacles,and releases the effectiveness of satellite digital economy services comprehensively.The study argues that connecting EO knowledge in the human-cyber-physical systems and exploring high-throughput knowledge coproduction and transfer technology in a"humans-in-the-loop"environment are necessary to enhance the interpretability of tacit EO knowledge and promote the competitiveness and activeness of EOKH.