Peatlands cover just 3% of Earth's land surface, yet store an estimated 600-700 Pg carbon (PgC), approximately one-third of Earth's soil carbon, making them critical regulators of the global carbon cycle. However, peatland spatial extent remains highly uncertain, particularly at fine spatial scales and in data-sparse regions. Existing global peatland datasets rely on heterogeneous inventories and regional products, leading to large inconsistencies in both total peat area and spatial distribution. These limitations hinder accurate assessments of peatland-climate feedbacks, carbon budgets, national policy development, and restoration efforts. We propose a machine learning framework that combines a priori information from existing peat databases (PEATMAP, Global Peatland Database, and CORINE Land Cover) with satellite observations in the visible, together with topographic and hydrological information. Our methodology employs a neural network trained with 17 input variables including Landsat-8 surface reflectance, topographic attributes from the MERIT database (elevation, slope, distance to drainage, height above drainage), and water table depth data. The model first generates a continuous Peatland Index (PI) at 3 arc-second (~90m) resolution, that can be thresholded to obtain a binary peat classification. In regions with reliable coarse resolution peat information, the PI can be used to downscale it and obtain a coherent high resolution peat classification. The obtained pan-boreal/Northern Hemisphere peatland map at 90m was evaluated through both quantitative and qualitative approaches. Fully independent validation using the Peat-DBase field dataset (over 180,000 peat and non-peat observations) demonstrates an overall accuracy of 68.4% and an F1-score of 0.80. Regional assessments show 69.2% overall accuracy (F1=0.81) in Eurasia and 63.8% (F1=0.74) in North America. Qualitative spatial evaluation across multiple case-study regions reveals that the proposed map successfully captures fine-scale spatial details absent in existing inventories, including explicit delineation of open water bodies, river networks, and topographic constraints on peatland distribution. The product exhibits improved spatial coherency with high-resolution imagery while remaining consistent with large-scale patterns from current peat databases. This work provides a spatially coherent, high-resolution peatland dataset spanning the Northern Hemisphere, offering improved capabilities for carbon stock estimation, hydrological modeling, and monitoring peatland degradation. Future improvements will incorporate SAR data, additional environmental drivers, and deep learning-based feature extraction to further enhance classification accuracy, spatial details, time-evolution, and peat information.
Clouds cover approximately 60% of the globe and are therefore an obstacle to observing the atmosphere and surface of the Earth from space. To limit their impact on Infrared Atmospheric Sounding Interferometer (IASI)-based atmospheric and surface property retrievals, it is important to obtain an IASI-coherent cloud detection/classification. Many cloud retrievals, whether physical or statistical, are performed at the pixel-level. However, since clouds are spatially structured, using the spatial coherency across the IASI footprints should improve cloud detection. Infrared Atmospheric Sounding Interferometer orbits (restructured as rectangular images) are collocated with the cloud classification (clear, water, ice, and two-level ice) extracted from SEVIRI-based Optimal Cloud Analysis to train a machine learning model. The training is performed over the SEVIRI disk, but the resulting model can be applied at the global scale (i.e., transfer learning). We use a partial-Convolutional Neural Networks (p-CNN), a new image-scale model able to deal with a large amount of spatially missing pixels. This p-CNN model correctly classifies the four cloud types with accuracy 77%; and this number increases to 88% when considering only spatially homogeneous IASI pixels, which shows the importance of subpixel heterogeneity. The other main source of differences is the IASI/SEVIRI resolution discrepancy. Thanks to our image-processing approach and the cloud spatial coherency, two-layer clouds are better retrieved than with pixel-wise processing. Our new IASI cloud product not only classifies the cloud phase at a global scale, but also estimates the cloud-type fractions in each IASI pixel. It therefore has a potential for subpixel downscaling.
The Central Highlands of Vietnam is the largest Robusta coffee ( Coffea canephora Pierre ex A.Froehner) growing region in the world. This study identifies the most important climatic variables that determine the current distribution of coffee in the Central Highlands and builds a ‘coffee suitability’ model to assess changes in this distribution due to climate change scenarios. A neural network-based suitability model was trained on coffee occurrence data derived from national statistics on coffee-growing areas. Bias-corrected regional climate models, adjusted to reduce systematic deviations from observed patterns, were used for two climate change scenarios (RCP8.5 and RCP2.6) to assess changes in suitability for three future time periods (2038–2048, 2059–2069, 2060–2070) relative to the 2009–2019 baseline. Average expected losses in suitable areas were 62% and 27% for RCP8.5 and RCP2.6, respectively. The loss in suitability due to RCP8.5 is particularly pronounced after 2060. Increasing mean minimum temperature during the harvest (October–November) and growing season (March–September), and decreasing precipitation during the late growing season (July–September) mainly determined the loss in suitable areas. Given these risks, adaptation strategies such as shade management, soil conservation, and the development of climate-resilient varieties are essential to sustain coffee production in the region.
Neural networks have been used for the retrieval of soil moisture (SM) from microwave observations over the last 20 years. The exploitation of the SM and Ocean Salinity (SMOS) observations has largely benefited from such statistical models. However, these retrievals are currently done at the pixel level, ignoring spatial context and using a constant incidence angle configuration approach. While pixels are seen with a varying number of angles from the SMOS instrument, only pixels monitored by a fixed preselected angle configuration are considered. These two limitations can have a negative impact on the quality of the retrievals. This paper introduces a new neural network (NN) architecture that combines two powerful innovations. First, the new model is image based: It ingests the entire SMOS orbit swath and thus leverages the strong spatial pattern present in the satellite observations. Second, a "partial convolutional layer" is tested. It allows being flexible, in the retrieval, on the angle configuration: More incidence angles can be exploited when they are available. Finally, a concept called "localization" is also exploited, helping the NN retrieval to adapt its behavior to specific local conditions. Experiments are conducted at the SMOS orbit scale over the contiguous United States (CONUS) region using 5 years of SMOS data (2016-19). A temporal correlation of 0.74 (unitless) with respect to ERA5 reanalysis and 0.62 with respect to in situ SM measurement network is obtained (to be compared to, respectively, 0.63 and 0.60 with the legacy pixel and fixed angle based approach). Furthermore, the use of partial convolutions results in enlarging the retrieval domain by +240% versus legacy retrieval and by +140% versus the operational SMOS level-2 product.
Land surfaces are characterised by strong heterogeneities of, among other variables, soil texture, orography, land cover, snow, or Soil Moisture (SM). SM is of broad scientific interest due to its role in the Earth system and its capital practical value for a wide range of applications from flood forecasting to agriculture. The scientific community has made significant progress in estimating SM from satellite-based passive MicroWave (MW). Most of SM estimates relie on a physical-based inversion to retrieve SM from passive MW. As an alternative to physical-based inversions, Neural Network (NN) retrieval algorithms have been successfully implemented for several sensors in recent years (Aires et al., 2005; Kolassa et al., 2016). The Soil Moisture and Ocean Salinity (SMOS) L3BT product (Al Bitar et al. 2017) uses an angle-binning scheme to organize the measured Brightness Temperature (BT) data. Three points that could improve SM retrieval will be considered in this presentation. (1) For coarse resolution MW instruments such as SMOS, NN algorithms are currently defined at the pixel level. Using the strong spatial patterns at the surface should help the SM retrieval, and we intend here to use an image-processing-based retrieval to investigate its potential. (2) Despite the important scanning angle information available on SMOS, not all angles are available for every pixel: The need to specify a limited angle configuration can drastically reduce the number of retrieved pixels, and the potential use of some large angle information is lost (Rodriguez-Fernandez et al. 2015). These missing data (both pixels and some angle configurations) could impede the use of image-based retrieval approaches. To tackle this issue, innovative machine learning techniques, such as “partial convolutional layers”, have been suggested very recently (Boucher et al. 2023), where missing data can be managed for both the spatial and the angle dimensions. This expands significantly the spatial coverage of the SMOS retrieval, especially for pixels with incomplete angle information. (3) A concept called “Localization” is also exploited, helping the ML retrieval to adapt its behaviour to specific local conditions to reduce local retrieval biases. By specializing its behaviour to local conditions, the relation between passive MW and SM is “simplified” over a particular pixel, this allows to reduce the impact of missing local information needed for a truly global model. We propose several NN and ML architectures to incorporate localization information into the networks, reducing significantly local biases. Experiments are conducted over the CONUS using several years of SMOS data. Impacts of the image- versus the pixel-scale processing is measured, as well the spatial extension of the SM retrieval due to better missing-data handling, and the effect of the localization is analysed too. The best configuration for a global-scale retrieval is yet to be found because the spatial domain to consider is strategic for image-processing schemes, but original and important technical solutions are proposed here that could pave the way for the next generation of SM retrievals.
In recognition of the importance of inland waters, numerous datasets mapping their extents, types, or changes have been created using sources ranging from historical wetland maps to real-time satellite remote sensing. However, differences in definitions and methods have led to spatial and typological inconsistencies among individual data sources, confounding their complementary use and integration. The Global Lakes and Wetlands Database (GLWD), published in 2004, with its globally seamless depiction of 12 major vegetated and non-vegetated wetland classes at 1 km grid cell resolution, has emerged over the last few decades as a foundational reference map that has advanced research and conservation planning addressing freshwater biodiversity, ecosystem services, greenhouse gas emissions, land surface processes, hydrology, and human health. Here, we present a new iteration of this map, termed GLWD version 2, generated by harmonizing the latest ground- and satellite-based data products into one single database. Following the same design principle as its predecessor, GLWD v2 aims to avoid double counting of overlapping surface water features while differentiating between natural and non-natural lakes, rivers of multiple sizes, and several other wetland types. The classification of GLWD v2 incorporates information on seasonality (i.e., permanent vs. intermittent vs. ephemeral); inundation vs. saturation (i.e., flooding vs. waterlogged soils), vegetation cover (e.g., forested swamps vs. non-forested marshes), salinity (e.g., salt pans), natural vs. non-natural origins (e.g., rice paddies), and stratification of landscape position and water source (e.g., riverine, lacustrine, palustrine, coastal/marine). GLWD v2 represents 33 wetland classes and – including all intermittent classes – depicts a maximum of 18.2 ×106 km2 of wetlands (13.4 % of the global land area excluding Antarctica). The spatial extent of each class is provided as the fractional coverage within each grid cell at a resolution of 15 arcsec (approximately 500 m at the Equator), with cell fractions derived from input data at resolutions as small as 10 m. The upgraded GLWD v2 offers an improved representation of inland surface water extents and their classification for contemporary conditions (∼ 1984–2020). Despite being a static map, it includes classes that denote intrinsic temporal dynamics. GLWD v2 is designed to facilitate large-scale hydrological, ecological, biogeochemical, and conservation applications, aiming to support the study and protection of wetland ecosystems around the world. The GLWD v2 database is available at https://doi.org/10.6084/m9.figshare.28519994 (Lehner et al., 2025).
Satellite remote sensing is commonly used to observe the hydrologic cycle at spatial scales ranging from river basins to the globe. Yet it remains difficult to obtain a balanced water budget using remote sensing data, which highlights the errors and uncertainties in earth observation (EO) data. Various methods have been proposed to correct EO datasets to make them more coherent, so that they result in a more balanced water budget. This study aimed to improve estimates of water budget components (precipitation, evapotranspiration, runoff, and total water storage change) at the global scale using the methods of optimal interpolation (OI) and neural network (NN) modeling. We trained a set of NNs on a set of 1,358 river basins and validated them on an independent set of 340 basins and in-situ observations of evapotranspiration and river discharge. We extended the models to make pixel-scale predictions in 0.5° grid cells for near-global coverage. Calibrated datasets result in lower water budget residuals in validation basins: the mean and standard deviation of the imbalance is 11 ± 44 mm/mo when calculated with uncorrected EO data and 0.03 ± 24 mm/mo after calibration by the NN models. This study suggests to data producers where corrections should be made to the EO datasets, and demonstrates the benefits of physically-driven NN models for studying the hydrologic cycle at the global scale.
Water resources play a crucial role in the global water cycle and are affected by human activities and climate change. However, the impacts of hydropower infrastructures on the surface water extent and volume cycle are not well known. We used a multi-satellite approach to quantify the surface water storage variations over the 2000-2020 period and relate these variations to climate-induced and anthropogenic factors over the whole basin. Our results highlight that dam operations have strongly modified the water regime of the Mekong River, exhibiting a 55 % decrease in the seasonal cycle amplitude of inundation extent (from 3178 km2 to 1414 km2) and a 70 % decrease in surface water volume (from 1109 km3 to 327 km3) over 2000-2020. In the floodplains of the Lower Mekong Basin, where rice is cultivated, there has been a decline in water residence time by 30 to 50 days. The recent commissioning of big dams (2010 and 2014) has allowed us to choose 2015 as a turning point year. Results show a trend inversion in rice production, from a rise of 40 % between 2000 and 2014 to a decline of 10 % between 2015 and 2020, and a strong reduction in aquaculture growth, from +730 % between 2000 and 2014, to +53 % between 2015 and 2020. All these results show the negative impact of dams on the Mekong basin, causing a 70 % decline in surface water volumes, with major repercussions for agriculture and fisheries over the period 2000-2020. Therefore, new future projects such as the Funan Techo canal in Cambodia, scheduled to start construction at the end of 2024, will particularly affect 1300 km2 of floodplains in the lower Mekong basin, with a reduction in the amount of water received, and other areas will be subjected to flooding. The human, material and economic damage could be catastrophic.
Estimating river discharge Q at global scale from satellite observations is not yet fully satisfactory in part because of limited space/time resolution. Furthermore, on highly anthropized basins, it is essential to anchor the analysis to reliable Q measurements. Gauge networks are however very sparse and limited in time, and SWOT (Surface Water Ocean Topography) river discharge estimates at global scale are not yet available. The method proposed here is able to obtain continuous daily Q estimates at 1 km/daily resolution, using indirect satellite data and ground-based estimates. We focus here on the Ebro. Over such an anthropized basin (e.g. change of land use, irrigation), the exploitation of 205 available gauges at their nominal resolution (i.e., daily point measurements) is a necessity. The hydrological Continuum model is used to help interpolate spatially and temporally the observations into our optimal interpolation scheme. The proposed Q -mapping is similar to an assimilation scheme were Earth observations (precipitation, evapotranspiration and total water storage change) and model simulations are constrained by in situ gauge measurements. The Q estimates are evaluated using a rigorous leave-one-out experiment, showing a good agreement with the in situ data: a correlation of 0.72 (median), and a 75th percentile of Nash-Sutcliffe Efficiency up to 0.62. Our spatio-temporal continuous Q estimates at high spatial/temporal resolution can describe complex continental water dynamics, including extreme events. SWOT estimates will soon be available, at the global scale but with irregular space/time sampling: our method should help exploit them to obtain a regular space-temporal description of the water cycle at high resolution.
GRACE (Gravity Recovery And Climate Experiment) is one of the most important satellite instrument for terrestrial hydrology and water cycle analysis. It is the unique source of information for the Total Water Storage Change (TWSC) that includes the mostly unknown groundwater storage change. Its use for many socio-economic applications is however drastically reduced by its spatial (≃200 km) and temporal (≃ monthly) resolutions. A statistical/physical dynamical downscaling is proposed here to obtain High Resolution (HR) 1 km and daily TWSC estimates, using auxiliary information from other satellite and modeling data (i.e. precipitation, evaporation and river network direction from topography). This scheme utilizes the Water Budget (WB) closure, an hydrological model and a set of statistical tools. The quality of the results is demonstrated over the Po river basin, a very challenging basin for GRACE observations due to its relatively small size and complex contrasting surface types. It is shown that terrestrial runoff and river discharge are well disentangled in the obtained HR TWSC fields. The new TWSC data can also show the impact of weather events such as strong precipitations. This study demonstrates how the optimal combination of multiple Earth data can help us better describe the terrestrial water cycle at higher space/time resolutions.
The aim of CERISE is to develop new and innovative coupled land-atmosphere data assimilation approaches and land initialisation techniques to pave the way for the next generations of the Copernicus Climate Change Service (C3S) reanalysis and seasonal prediction systems. These developments are combined with innovative work on machine-learned observation operator to ensure optimal data fusion fully integrated in coupled assimilation systems. The project aims at improving the quality and consistency of the C3S reanalysis systems and of the components of the seasonal prediction multi-system, directly addressing the evolving user needs for improved and more consistent C3S Earth system products.This presentation gives an overview of the objectives of the CERISE project with a focus on developments of coupled land-atmosphere data assimilation to improve the climate consistency of the next generation of C3S Earth system global and regional reanalyses. It describes ensemble-based unified land data assimilation and coupling infrastructure and methodology developments conducted in the first 18 months of the project. Work on machine-learning based observation operator to enhance the exploitation of passive microwave data is introduced, presenting the training databases, the machine learning approaches developed, and results comparing simulated and observed brightness temperature from AMSR2. Results from numerical experiments show the benefits of using ensemble-based land data assimilation for surface and near-surface weather representation both at regional and global scales. The first CERISE global land reanalysis prototype is presented. Its results are compared to state-of-the-art operational reanalysis using a set of newly developed diagnostic tools. Infrastructure, methodology and scientific results presented highlight the feasibility and the added value of the integration of the CERISE developments in the existing C3S core service.
Evapotranspiration (E) is one of the most uncertain components of the global water cycle (WC). Improving global E estimates is necessary to improve our understanding of climate and its impact on available surface water resources. This work presents a methodology for deriving monthly corrections to global E datasets at 0.25∘ resolution. A principled approach is proposed to firstly use indirect information from the other water components to correct E estimates at the catchment level, and secondly to extend this sparse catchment-level information to global pixel-level corrections using machine learning (ML). Several E satellite products are available, each with its own errors (both random and systematic). Four such global E datasets are used to validate the proposed approach and highlight its ability to extract seasonal and regional systematic biases. The resulting E corrections are shown to accurately generalize WC closure constraints to unseen catchments. With an average deviation of 14% from the original E datasets, the proposed method achieves up to 20% WC residual reduction on the most favorable dataset.
A surface water extent downscaling framework was developed in the past using a floodability index based on topography. We presented here a new downscaling approach including several improvements. (1) The use of a new Floodability Index (FI), including better integration of auxiliary permanent waters (i.e., presence of water during the whole time record). By using this updated FI, the new downscaling became a true data-fusion with permanent water databases originating mainly from visible observations. (2) Some discontinuities between low resolution cells have been reduced thanks to a new smoothing algorithm. (3) Finally, a coastal extrapolation scheme has been presented to deal with coarse resolution data contaminated by the ocean. This new and complex downscaling framework was tested here on the GIEMS (Global Inundation Extent from Multi-Satellite) database but the approach is generalizable and any surface water database could be used instead. It was shown that this new downscaling procedure (including several processing steps, algorithms and data sources) is a significant improvement compared to the previous version thanks to the new floodability index and the downscaling processing chain. Compared to the previous version, the downscaling results (GIEMS-D) were more coherent with the permanent water database and preserved better the original low-resolution information (e.g., mean scare error water fraction (0–1) of 0.0041 for the old version, and 0.0018 for the new version, over flooded areas in the Amazon). GIEMS-D has also been evaluated at the global scale and over the Amazon basin using independent datasets, showing an overall good performance of the downscaling.
Firstly, a new global floodability index with a resolution of 3 arc-second is built from topography-based information provided by the MERIT database, using a neural network approach. The topography and permanent water were defined in a coherent way, ensuring the coherency between the resulting floodability index and permanent water, which is unprecedented in previous versions. The evaluation of the floodability index is done with independent observation datasets on surface water and land cover, showing good performances in areas where surface water is naturally driven by topography conditions and limitation in human-affected areas and some specific environments like peatland. Secondly, some of the applications that the floodability index can serve are introduced, including downscaling low-resolution data, analyzing and comparing datasets at different resolution, and data fusion.
The Qinghai–Tibet Plateau is rich in water resources with numerous lakes, rivers, and glaciers, and, as a source of many rivers in Central Asia, it is known as the Asian Water Tower. Under global climate change, it is critical to understand the current influencing factors on surface water area in this region. Although there are numerous studies on surface water mapping, they are still limited by temporal/spatial resolution and record length. Moreover, the complicated topographic condition makes it challenging to map the surface water accurately. Here, we proposed an automatic two-step annual surface water classification framework using long time-series Landsat images and topographic information based on the Google Earth Engine (GEE) platform. The results showed that the producer accuracy (PA) and user accuracy (UA) of the surface water map in the Qinghai–Tibet Plateau in 2020 were 99% and 90%, respectively, and the Kappa coefficient reached 0.87. Our dataset showed high consistency with high-resolution images, indicating that the proposed large-scale water mapping method has great application potential. Furthermore, a new annual surface water area dataset on the Qinghai–Tibet Plateau from 2000 to 2020 was generated, and its relationship with climate, vegetation, permafrost, and glacier factors was explored. We found that the mean surface water area was about 59 481 km2, and there was a significant increasing trend (=322 km2/year, $p < 0.01$ ) during 2000–2020 in the plateau. Greening, warming, and wetting climate conditions contributed to the increase of surface water area. Active layer thickness and permafrost types may be the most related to the decrease of surface water area. This study provides important information for ecological assessment and protection of the plateau and promotes the implementation of sustainable development goals related to surface water resources.
Climate models are widely used in climate change impact studies. However, these simulations often cannot be used directly due to inherent limitations, such as structural biases or parametric uncertainties. Nevertheless, several so-called “bias correction” (B-C) or “bias adjustment” methods have been proposed to get these simulations closer to real observations. Various studies have reviewed available methods; however, numerous innovative methods have been developed in recent years. An up-to-date review of the B-C methods is presented here. To compare these complex methods, a focus is placed on the pedagogy of the presentation. The main lines of thought are presented based on the method assumptions, mathematical form, properties, and applicative purposes. Six representative quantile-based methods are compared for temperature and precipitation monthly time series over the European area, for a climate change scenario with a strong CO2 forcing which is chosen here to facilitate the analysis of the differences among the methods. New, simple, and easy-to-understand diagnostic tools are recommended to measure the impact of the adjustment on the ability of B-C methods to: (1) bring the model outputs closer to observations over the historical record, (2) exploit as much as possible the climate change signal provided by the model. Each B-C method is intended to find the best compromise between these two objectives. A discussion on potential pathways for future developments is finally proposed.
Multilayer perceptrons have been popular in the remote-sensing community for the last 30 years, in particular for the infrared atmospheric sounding interferometer instrument. However, for coarse-resolution infrared instruments such as the infrared atmospheric sounding interferometer, these algorithms are currently used at the pixel level and are trained at a global scale. This can result in regional biases that not only affect the quality of the retrieval but can potentially also propagate into model forecasts if these results were to be assimilated. To help reduce these biases, we try to help the neural network (NN) adjusting its behaviour to local conditions; we call this "localization". We investigate how to localize traditional multilayer perceptrons applied at the pixel scale and compare this with novel artificial intelligence image-processing techniques, more particularly convolutional NNs and so-called "localized convolutional NNs". These approaches are tested for the retrieval of surface temperature over a fixed domain. Different techniques are proposed to localize both approaches. Further evaluation of the retrieval methods will be covered in a part II companion paper.
When using neural networks (NNs), the lack of input information characterizing the radiative transfer can result in regional biases, especially when retrieving surface properties. In the Part I companion article we explored localization techniques in an attempt to help the NN adjust its behaviour to local conditions. In this article we analyze results from an image‐processing approach, the novel localized convolutional NN (CNN) model for the retrieval of surface temperature (TS) over a fixed domain using infrared atmospheric sounding interferometer (IASI) observations. An in‐depth evaluation is performed. The localized‐CNN architecture is a promising artificial intelligence solution that provides retrievals similar to, if not better than, those of the European Organisation for the Exploitation of Meteorological Satellites' PWLR3 retrieval algorithm that also uses IASI observations, collocated with microwave data too. This shows the benefits of localizing the CNN retrieval. This image‐processing retrieval scheme allows interpolation of the TS below the clouds, and thus a preliminary analysis of the cloud impact on the TS is performed. The possibility to estimate retrieval uncertainties is investigated, and a practical solution, based on the binning of the input space, is proposed for CNN architectures. The best strategy for a global‐scale retrieval is yet to be found for such an image‐processing scheme, but potential solutions and their respective advantages and disadvantages are discussed.