The super-resolution (SR) of Sentinel-2 (S2) imagery has been constrained by the absence of reliable high-resolution reference data for its 20- and 60-meter spectral bands. In this letter, we introduce SEN2NEON, a cross-sensor benchmark dataset that rigorously harmonizes high-resolution AVIRIS-NG hyperspectral imagery with real Sentinel-2 Level-2A reflectances through spectral response convolution and radiometric correction. SEN2NEON provides physically aligned low- and high-resolution (LR–HR) pairs for all Sentinel-2 bands (except the cirrus band) at the original S2 and 2.5m reference resolution, respectively. Compared to existing cross-sensor datasets, it offers a larger number of spectrally aligned bands, a finer effective reference resolution, and stronger LR–HR consistency. We quantify harmonization quality via a cross-dataset alignment analysis based on LR–HR resampling consistency metrics. By preserving the original Sentinel-2 observations as the LR input and deriving the HR reference via sensor-aware harmonization rather than synthetic degradation, SEN2NEON establishes a reproducible and physically grounded benchmark for multispectral Sentinel-2 super-resolution research. The sensor-aware harmonization strategy underlying SEN2NEON further provides a transferable and physically grounded framework for super-resolution benchmarking that can be extended to a broader range of multispectral satellite sensors.
Deep learning-based super-resolution (SR) models offer a promising approach to enhancing the effective spatial resolution of optical satellite images. However, existing SR implementations have shown that, while these models can reconstruct fine-scale details, they often introduce undesirable artifacts, such as nonexistent local structures, reflectance distortions, and geometric misalignment. To mitigate these issues, fully synthetic data approaches have been explored for training, as they provide complete control over the degradation process and allow precise supervision and ground-truth availability. However, challenges in domain transfer have limited their effectiveness when applied to real satellite images. In this work, we propose SEN2SR, a new deep learning framework trained to super-resolve Sentinel-2 images while preserving spectral and spatial alignment consistency. Our approach harmonizes synthetic training data to match the spectral and spatial characteristics of Sentinel-2, ensuring realistic and artifact-free enhancements. SEN2SR generates 2.5-meter resolution images for Sentinel-2, upsampling the 10-meter RGB and NIR bands and the 20-meter Red Edge and SWIR bands. To ensure that SR models focus exclusively on enhancing spatial resolution, we introduce a low-frequency hard constraint layer at the final stage of SR networks that always enforces spectral consistency by preserving the original low-frequency content. We evaluate a range of deep learning architectures, including Convolutional Neural Networks, Mamba, and Swin Transformers, within a comprehensive assessment framework that integrates Explainable AI (xAI) techniques. Quantitatively, our framework achieves superior PSNR while maintaining near-zero reflectance deviation and spatial misalignment, outperforming state-of-the-art SR frameworks. Moreover, we demonstrate maintained radiometric fidelity in downstream tasks that demand high-fidelity spectral information and reveal a significant correlation between model performance and pixel-level model activation. Qualitative results show that SR networks effectively handle diverse land cover scenarios without introducing spurious high-frequency details in out-of-distribution cases. Overall, this research underscores the potential of SR techniques in Earth observation, paving the way for more precise monitoring of the Earth’s surface. Models, code, and examples are publicly available at https://github.com/ESAOpenSR/SEN2SR.
Earth observation (EO) is essential to understanding the Earth system, enabling the transformation of planetary properties into measurable variables that can be analysed, compared, and modelled. In recent decades, EO capabilities have grown rapidly, accompanied by an even faster expansion in the number and variety of available EO instruments. Today, EO includes instruments deployed on satellites, airborne platforms, and terrestrial or in-situ systems. However, despite this proliferation of instruments, users often lack a single, reliable source describing their existence and key characteristics. Although existing data catalogues have substantially improved dataset discovery, they primarily describe data products rather than providing persistent, curated metadata about the instruments that produced them. Here we present Awesome Earth Observation Instruments, an open, standardized, and community-oriented registry providing machine-readable metadata for EO instruments. The catalogue is hosted on GitHub and allows contributors to submit instrument metadata following a common schema. The schema combines a lightweight core with modular extensions covering spectral, geometric, and data access-related metadata, enabling both standardization and flexibility across diverse EO systems. All submissions undergo automated schema validation and human review. Because the schema is open, versioned, and extensible, the catalogue can continuously evolve as new instruments and metadata requirements emerge. This facilitates the discovery, interpretation, and analysis of EO data in light of instrument characteristics. To support programmatic access and interoperability, we further envisage an API for integration within common EO analysis environments. The catalogue is openly available at https://github.com/awesome-spectral-indices/awesome-earth-observation-instruments.
Super-resolution (SR) of remote-sensing imagery is commonly assessed through spatial fidelity or visual sharpness, although index-based geomatics requires preservation of cross-band spectral relationships. We compare five Sentinel-2 full-band SR configurations (LDSR-S2, SRGAN, SPAN, Mamba, and SWIN) in two hazard-mapping cases: flood-water detection during the 2024 Valencia flood and burn-scar mapping after the 2025 Palisades wildfire. Each model refines four native RGB–NIR bands, while SEN2SR reconstructs the remaining bands to produce a ten-band product at 2.5 m. We evaluate native-grid reconstruction, the introduction of high-frequency details, and downstream thematic, boundary, and edge-region metrics against a bilinear-interpolation baseline. Dynamic-threshold MNDWI and dNBR detectors are applied independently to each output. Among the learned configurations, SWIN achieves the strongest native-grid reconstruction and task-specific spectral consistency and the strongest fire agreement, but adds the least high-frequency content. Flood full-ROI gains are modest, with LDSR-S2 increasing the F1-score from 0.085 for bilinear interpolation to 0.091. All learned configurations increase flood-edge recall, F1-score, and IoU while reducing edge precision and balanced accuracy. LDSR-S2 gives the strongest final edge F1-score and IoU in both tasks, Mamba gives the lowest learned-model symmetric flood-boundary distance. SPAN yields the best spatial consistency and lowest learned flood-edge spectral error and SRGAN adds the most high-frequency content and the largest combined edge-region gain, alongside the greatest task-specific spectral deviation and the largest symmetric boundary-distance increases. Thus, increased edge activation does not establish uniformly improved delineation or recovered sub-pixel detail. These cases demonstrate feasibility rather than generalization. Operational validation requires more diverse, time-synchronous, high-resolution, spectrally compatible references.
Remote sensing super-resolution aims to enhance the spatial details of satellite images by introducing meaningful high-frequency features while avoiding hallucinations and spectral distortions. High-resolution imagery is usually not publicly available, whereas low-resolution imagery is freely available with a much higher revisit rate, such as the Sentinel-2 multispectral imaging mission. Cross-sensor super-resolution has the potential to bridge this gap, providing high spatial and temporal resolution imagery which are otherwise unavailable for many remote sensing users and applications. With the recent advancements in diffusion models, many methodologies have emerged which take advantage of their generative power to perform super-resolution. We propose an adapted latent diffusion approach, since image diffusion is computationally prohibitive to be applied to large Earth observation datasets. Contrary to standard latent diffusion, we encode the low-resolution image to condition the diffusion process, forcing better spectral consistency with the input imagery. The model includes visible and near-infrared bands. To ensure trustworthy results, we utilize the probabilistic nature of diffusion models to generate pixel-level uncertainty maps. This confidence metric is crucial for real-world applications, such as environmental monitoring, land cover classification, and change detection, where accurate surface feature reconstruction and spectral consistency are essential. The uncertainty map allows users to evaluate the reliability of the product for these tasks. The proposed model super-resolves Sentinel-2 imagery at 10 to 2.5 m and is the first multispectral remote sensing (RS) super-resolution diffusion model efficient enough to process large-scale RS datasets, as well as the only model providing a pixelwise uncertainty metric.
The acquisition of near-infrared (NIR) imagery is crucial for various remote sensing (RS) applications but is often unavailable in RS datasets, providing only the red-green-blue (RGB) bands. Most classical RS workflows rely on NIR information, such as the normalized difference vegetation index and normalized difference water index, for vegetation and soil monitoring or land cover classification workflows. To overcome this limitation, we introduce NIR-generative adversarial network (GAN), a conditional GAN designed to synthesize the NIR band directly from RGB imagery through an image-to-image translation framework. Crucially, the proposed model integrates geographical and climatic context using location embeddings, and incorporates application-specific index-derived loss functions to optimize synthesized outputs for downstream tasks. NIR-GAN enables the generation of partly synthetic datasets, facilitating the training and evaluation of multispectral machine learning models. The developed model demonstrates the capability to generate realistic NIR images from various sources of RS data, regardless of the spatial resolution and spectral characteristics of the sensor, reaching up to 26.6 dB PSNR on Sentinel-2 images. Although synthetic data will never be able to replace the acquisition of real NIR information, this approach can fill the gap in situations where this information would otherwise be absent, including but not limited to the assembly of large-scale training datasets.
In recent years, research in flood mapping from remote sensing satellite imagery has predominantly focused on deep learning methods. While new flood segmentation models are increasingly being proposed, most of these works focus on advancing architectures trained on single datasets. Therefore, these studies overlook the intrinsic limitations and biases of the available training and evaluation data. This often leads to poor generalization and overconfident predictions when these models are used in real-world scenarios. To address this gap, the objective of this work is twofold. First, we train and evaluate flood segmentation models on five publicly available datasets including data from Sentinel-1, Sentinel-2, and both SAR and multispectral modalities. Our findings reveal that models achieving high detection accuracy on a single dataset (intra-dataset validation) do not necessarily generalize well to unseen datasets. In contrast, models trained on more diverse samples from multiple datasets demonstrate greater robustness and generalization ability. Furthermore, we present a dual-stream multimodal architecture that can be independently trained and tested on both single-modality and dual-modality datasets. This enables the integration of all the diversity and richness of the available data into a single unified framework. The results emphasize the need for a more comprehensive validation using diverse and well-designed datasets, particularly for multimodal approaches. If not adequately addressed, the shortcomings of current datasets can significantly limit the potential of deep learning-based operational flood mapping approaches.
We present OpenSR-SRGAN, an open and modular framework for single-image super-resolution in Earth Observation. The software provides a unified implementation of SRGAN-style models that is easy to configure, extend, and apply to multispectral satellite data such as Sentinel-2. Instead of requiring users to modify model code, OpenSR-SRGAN exposes generators, discriminators, loss functions, and training schedules through concise configuration files, making it straightforward to switch between architectures, scale factors, and band setups. The framework is designed as a practical tool and benchmark implementation rather than a state-of-the-art model. It ships with ready-to-use configurations for common remote sensing scenarios, sensible default settings for adversarial training, and built-in hooks for logging, validation, and large-scene inference. By turning GAN-based super-resolution into a configuration-driven workflow, OpenSR-SRGAN lowers the entry barrier for researchers and practitioners who wish to experiment with SRGANs, compare models in a reproducible way, and deploy super-resolution pipelines across diverse Earth-observation datasets.
In this work, we introduce a synergistic framework to perform multitemporal and multimodal flood extent mapping from multispectral and SAR satellite image time series. Two state-of-the-art flood datasets, Kuro Siwo and WorldFloods, have been used to create a dataset with pre-flood and postflood images from Sentinel-1 and Sentinel-2, along with hand-labeled reference flood masks for each data source. We develop deep learning based fusion models achieving an Intersection over Union (IoU) similar to 70% in our best fusion model, which outperforms single modality models by similar to 10%. With this work, we pave the road to develop improved flood detection services by exploiting all available flood segmentation datasets and available satellite images.
In recent years, there has been a growing interest in using image super-resolution (SR) techniques in remote sensing. These techniques aim to reconstruct high-resolution (HR) imagery from low-resolution (LR) sources. Despite the development of sophisticated SR methodologies, determining what constitutes ‘good’ SR is still a matter of debate. Present-day literature often presents SR models through a strong computer vision perspective, heavily relying on synthetic datasets. Moreover, commonly used metrics often prioritize attributes that do not necessarily correspond to improvements in spatial resolution. To address this challenge, we present OpenSR-test , a comprehensive benchmark designed exclusively for evaluating SR of remote sensing images. Our framework incorporates specific quality metrics and curated cross-sensor datasets, each spanning various scale factors with consistent metadata. Utilizing OpenSR-test , we evaluate state-of-the-art SR algorithms from a remote sensing perspective. The OpenSR-test framework and datasets are publicly available at https://esaopensr.github.io/opensr-test/.
Detecting and screening clouds is the first step in most optical remote sensing analyses. Cloud formation is diverse, presenting many shapes, thicknesses, and altitudes. This variety poses a significant challenge to the development of effective cloud detection algorithms, as most datasets lack an unbiased representation. To address this issue, we have built CloudSEN12 +, a significant expansion of the Cloud- SEN12 dataset. This new dataset doubles the expert-labeled annotations, making it the largest cloud and cloud shadow detection dataset for Sentinel-2 imagery up to date. We have carefully reviewed and refined our previous annotations to ensure maximum trustworthiness. We expect CloudSEN12+ will be a valuable resource for the cloud detection research community. (c) 2024 The Author(s). Published by Elsevier Inc. This is an open access article under the CC BY license (http://creativecommons.org/licenses/by/4.0/)
The increasing demand for high spatial resolution in remote sensing has underscored the need for super-resolution (SR) algorithms that can upscale low-resolution (LR) images to high-resolution (HR) ones. To address this, we present SEN2NAIP, a novel and extensive dataset explicitly developed to support SR model training. SEN2NAIP comprises two main components. The first is a set of 2,851 LR-HR image pairs, each covering 1.46 square kilometers. These pairs are produced using LR images from Sentinel-2 (S2) and corresponding HR images from the National Agriculture Imagery Program (NAIP). Using this cross-sensor dataset, we developed a degradation model capable of converting NAIP images to match the characteristics of S2 imagery ( $$S{2}_{like}$$ ). This led to the creation of a second subset, consisting of 35,314 NAIP images and their corresponding $$S{2}_{like}$$ counterparts, generated using the degradation model. With the SEN2NAIP dataset, we aim to provide a valuable resource that facilitates the exploration of new techniques for enhancing the spatial resolution of Sentinel-2 imagery.
We present a deep learning model, CH4Net, for automated monitoring of methane super-emitters from Sentinel-2 data. When trained on images of 23 methane super-emitter locations from 2017–2020 and evaluated on images from 2021, this model detects 84 % of methane plumes compared with 24 % of plumes for a state-of-the-art baseline while maintaining a similar false positive rate. We present an in-depth analysis of CH4Net over the complete dataset and at each individual super-emitter site. In addition to the CH4Net model, we compile and make open source a hand-annotated training dataset consisting of 925 methane plume masks as a machine learning baseline to drive further research in this field.
This paper introduces DTACSNet, a Convolutional Neural Network (CNN) model specifically developed for efficient onboard atmospheric correction and cloud detection in optical Earth observation satellites. The model is developed with Sentinel-2 data. Through a comparative analysis with the operational Sen2Cor processor, DTACSNet demonstrates a significantly better performance in cloud scene classification (F2 score of 0.89 for DTACSNet compared to 0.51 for Sen2Cor v2.8) and a surface reflectance estimation with average absolute error below 2% in reflectance units. Moreover, we tested DTACSNet on hardware-constrained systems similar to recent deployed missions and show that DTACSNet is 11 times faster than Sen2Cor with a significantly lower memory consumption footprint. These preliminary results highlight the potential of DTACSNet to provide enhanced efficiency, autonomy, and responsiveness in onboard data processing for Earth observation satellite missions.
Methane is the second most important greenhouse gas contributor to climate change; at the same time its reduction has been denoted as one of the fastest pathways to preventing temperature growth due to its short atmospheric lifetime. In particular, the mitigation of active point-sources associated with the fossil fuel industry has a strong and cost-effective mitigation potential. Detection of methane plumes in remote sensing data is possible, but the existing approaches exhibit high false positive rates and need manual intervention. Machine learning research in this area is limited due to the lack of large real-world annotated datasets. In this work, we are publicly releasing a machine learning ready dataset with manually refined annotation of methane plumes. We present labelled hyperspectral data from the AVIRIS-NG sensor and provide simulated multispectral WorldView-3 views of the same data to allow for model benchmarking across hyperspectral and multispectral sensors. We propose sensor agnostic machine learning architectures, using classical methane enhancement products as input features. Our HyperSTARCOP model outperforms strong matched filter baseline by over 25% in F1 score, while reducing its false positive rate per classified tile by over 41.83%. Additionally, we demonstrate zero-shot generalisation of our trained model on data from the EMIT hyperspectral instrument, despite the differences in the spectral and spatial resolution between the two sensors: in an annotated subset of EMIT images HyperSTARCOP achieves a 40% gain in F1 score over the baseline.
Clouds play a key role in regulating climate change but are difficult to simulate within Earth system models (ESMs).Improving the representation of clouds is one of the key tasks toward more robust climate change projections.This study introduces a new machine-learning-based framework relying on satellite observations to improve understanding of the representation of clouds and their relevant processes in climate models.The proposed method is capable of assigning distributions of established cloud types to coarse data.It facilitates a more objective evaluation of clouds in ESMs and improves the consistency of cloud process analysis.The method is built on satellite data from the Moderate Resolution Imaging Spectroradiometer (MODIS) instrument labeled by deep neural networks with cloud types defined by the World Meteorological Organization (WMO), using cloud-type labels from CloudSat as ground truth.The method is applicable to datasets with information about physical cloud variables comparable to MODIS satellite data and at sufficiently high temporal resolution.We apply the method to alternative satellite data from the Cloud_cci project (ESA Climate Change Initiative), coarse-grained to typical resolutions of climate models.The resulting cloud-type distributions are physically consistent and the horizontal resolutions typical of ESMs are sufficient to apply our method.We recommend outputting crucial variables required by our method for future ESM data evaluation.This will enable the use of labeled satellite data for a more systematic evaluation of clouds in climate models.
The future Copernicus Imaging Microwave Radiometer (CIMR) mission is planned to be launched in the 2027+ time frame. At its present phase, the first version of each Algorithm Theoretical Basis Document (ATBD) must be defined. CIMR will provide observations at L (1.4 GHz), C (6.9 GHz), X (10.65 GHz), Ku (18.7 GHz) and Ka (36.5 GHz) microwave frequencies. These observations will be relevant to develop high resolution land surface products. Here we present a preliminary study with the aim of exploring the future capabilities that the synergy of CIMR frequencies can provide. Focused on the 0 th -order Tau-Omega (τ-ω) model, we analysed the influence of soil roughness (H) and scattering albedo (ω) to retrieve soil moisture (SM) and vegetation optical depth (VOD) at L-band and how these parameters can be potentially estimated from higher frequency bands. We evaluated our results over CONUS, concluding that the soil roughness (H) parameter is affecting VOD and ω mainly in non-forested areas: in those areas, the increase of H produces a decrease in VOD. Our maps of ω revealed dependence with land cover type: generally, the lowest ω values were found in forested areas. Instead, our H map yielded patterns that could be mostly associated with topographic effects. Furthermore, by utilizing a depolarization index, TB dep , we discovered that its values were constrained to nearly zero (indicating minimal soil impact) in areas with vegetation, whereas in bare soils, topography had a significant influence on TB dep . We hypothesize that the use of this index could help in finding relationships among the multi-frequency information from CIMR, allowing us to understand the degree of sensitivity of each band to vegetation and topography.
Floods are among the most destructive extreme events that exist, being the main cause of people affected by natural disasters. In the near future, estimated flood intensity and frequency are projected to increase. In this context, automatic and accurate satellite-derived flood maps are key for fast emergency response and damage assessment. However, current approaches for operational flood mapping present limitations due to cloud coverage on acquired satellite images, the accuracy of flood detection, and the generalization of methods across different geographies. In this work, a machine learning framework for operational flood mapping from optical satellite images addressing these problems is presented. It is based on a clouds-aware segmentation model trained in an extended version of the WorldFloods dataset. The model produces accurate and fast water segmentation masks even in areas covered by semitransparent clouds, increasing the coverage for emergency response scenarios. The proposed approach can be applied to both Sentinel-2 and Landsat 8/9 data, which enables a much higher revisit of the damaged region, also key for operational purposes. Detection accuracy and generalization of proposed model is carefully evaluated in a novel global dataset composed of manually labeled flood maps. We provide evidence of better performance than current operational methods based on thresholding spectral indices. Moreover, we demonstrate the applicability of our pipeline to map recent large flood events that occurred in Pakistan, between June and September 2022, and in Australia, between February and April 2022. Finally, the high-resolution (10-30m) flood extent maps are intersected with other high-resolution layers of cropland, building delineations, and population density. Using this workflow, we estimated that approximately 10 million people were affected and 700k buildings and 25,000 km $$^2$$ of cropland were flooded in 2022 Pakistan floods.
In Earth observation, deep learning models rely heavily on comprehensive datasets for training and evaluation. However, the relevance of data quality is often underestimated, leading to subpar generalization in real-world remote sensing scenarios. This study aims to bridge this gap by proposing a straightforward method to identify critical human annotation errors in semantic segmentation datasets. The approach is based on two indices: trustworthiness and hardness. By implementing these indices, we estimate the extent of human annotation errors in CloudSEN12, a global dataset specifically designed for cloud detection in Sentinel-2 imagery. Considering only the trustworthiness index, our approach identified 1794 potential labelling errors among 10,000 image patches. Out of these, 106 were confirmed as human errors, resulting in a true positive rate of 9.86%. When this method was applied to other extensive cloud masking datasets, such as KappaSet and Sentinel-2 Cloud Mask Catalogue, it was found that over 44% of the human labels were inaccurate. These results do not imply the inferior quality of these datasets, instead, they highlight the considerable shift between the annotation protocols, making inter-dataset benchmarking exercises inequitable.
Jesus Malo合作论文数Dpt. of Optics,
School of Physics.
Universitat de Valencia9