Remote sensing in combination with AI models can provide near real time observations of land use changes like deforestation. We propose a data processing and assessment platform that can detect deforestation and land cover from the latest satellite observations. The platform contains multiple AI models pre-trained on satellite observation for Tree Cover, Canopy Height, Above Ground Biomass and Land Use mapping at scale in order to assess provenance of a commodity and compliance with the European Union Deforestation Regulation (EUDR). We demonstrate an integrated verification and monitoring tool, that can receive as input a polygon of interest, a commodity specification, and triggers an automated pipeline to download and to process satellite data and run any of the AI models in near real time to assess if a region is associated with deforestation or degradation of forests. We demonstrate the performance of such a monitoring tool in Bolivia, Brazil, and Mynamar.
Automated analysis of vast Earth observation data via interactive Vision-Language Models (VLMs) can unlock new opportunities for environmental monitoring, disaster response, and resource management. Existing generic VLMs do not perform well on Remote Sensing data, while the recent Geo-spatial VLMs remain restricted to a fixed resolution and few sensor modalities. In this paper, we introduce EarthDial, a conversational assistant specifically designed for Earth Observation (EO) data, transforming complex, multi-sensory Earth observations into interactive, natural language dialogues. EarthDial supports multi- spectral, multi-temporal, and multi-resolution imagery, enabling a wide range of remote sensing tasks, including classification, detection, captioning, question answering, visual reasoning, and visual grounding. To achieve this, we introduce an extensive instruction tuning dataset comprising over 11.11M instruction pairs covering RGB, Synthetic Aperture Radar (SAR), and multispectral modalities such as Near-Infrared (NIR) and infrared. Furthermore, EarthDial handles bi-temporal and multi-temporal sequence analysis for applications like change detection. Our extensive experimental results on 44 downstream datasets demonstrate that EarthDial outperforms existing generic and domain-specific models, achieving better generalization across various EO tasks. Our source codes and pre-trained models are at https://github.com/hiyamdebary/EarthDial.
Maximizing carbon sequestration in trees across different ecoregions has the potential to support carbon markets while improving forest restoration and preserving biodiversity. Tree height is a predictor of carbon stored in its biomass, but estimating tree height using publicly available remote sensing datasets remains challenging. Artificial intelligence can play an important role in improving estimates of vegetation canopy height, especially in regions with limited local measurements. This study compares a transformer-based Geospatial Foundation Model (GFM) and a baseline deep learning model (U-Net) for predicting tree canopy height across Kenya’s diverse ecoregions. The models use cloudfree mosaics from the Harmonized Landsat and Sentinel-2 (HLS) product as predictors and space-borne GEDI laser data for canopy height reference. Both models had similar root mean square error (RMSE) scores: GFM at 6.05 m and UNet at 5.80 m for the most prevalent small to medium-sized trees. In a second experiment, the models trained in Kenya were applied to a Mozambique study area. In this challenging set-up, GFM generalized better to the different ecoregions.
Assessing and understanding the urban scale impacts of extreme climate events is a global necessity. Risks associated with heat, where intra-urban dynamics and rural/urban boundary conditions greatly impact its distribution, are of particular interest as the evolution of climate change and urbanization persists. Characterizing Urban Heat Island (UHI) effects is dependent on the availability of high-resolution near-surface air temperature maps and a description of the Local Climate Zones (LCZs). This study assesses the applicability of state-of-the-art (SOTA) Artificial Intelligence (AI) techniques for UHI detection and characterization. A Geospatial Foundation Model (GFM) is fine-tuned to predict 2m air temperature at a 1 km resolution for the urban areas of Johannesburg, South Africa, with mean absolute error measures less than 1.5 degrees C. UHI characterization is further enabled through a Fully Connected Network (FCN) model for LCZs classification for the same region of interest.
This paper considers the optimal sensor allocation for estimating the emission rates of multiple sources in a two-dimensional spatial domain. Locations of potential emission sources are known (e.g., factory stacks), and the number of sources is much greater than the number of sensors that can be deployed, giving rise to the optimal sensor allocation problem. In particular, we consider linear dispersion forward models, and the optimal sensor allocation is formulated as a bilevel optimization problem. The outer problem determines the optimal sensor locations by minimizing the overall Mean Squared Error of the estimated emission rates over various wind conditions, while the inner problem solves an inverse problem that estimates the emission rates. Two algorithms, including the repeated Sample Average Approximation and the Stochastic Gradient Descent based bilevel approximation, are investigated in solving the sensor allocation problem. Convergence analysis is performed to obtain the performance guarantee, and numerical examples are presented to illustrate the proposed approach.
Global vegetation structure mapping is critical for understanding the global carbon cycle and maximizing the efficacy of nature-based carbon sequestration initiatives. Moreover, vegetation structure mapping can help reduce the impacts of climate change by, for example, guiding actions to improve water security, increase biodiversity and reduce flood risk. Global satellite measurements provide an important set of observations for monitoring and managing deforestation and degradation of existing forests, natural forest regeneration, reforestation, biodiversity restoration, and the implementation of sustainable agricultural practices. In this paper, we explore the effectiveness of fine-tuning of a geospatial foundation model to estimate above-ground biomass (AGB) using space-borne data collected across different eco-regions in Brazil. The fine-tuned model architecture consisted of a Swin-B transformer as the encoder (i.e., backbone) and a single convolutional layer for the decoder head. All results were compared to a U-Net which was trained as the baseline model Experimental results of this sparse-label prediction task demonstrate that the fine-tuned geospatial foundation model with a frozen encoder has comparable performance to a U-Net trained from scratch. This is despite the fine-tuned model having 13 times less parameters requiring optimization, which saves both time and compute resources. Further, we explore the transfer-learning capabilities of the geospatial foundation models by fine-tuning on satellite imagery with sparse labels from different eco-regions in Brazil.
Nature-based carbon sequestration solution have the potential to capture carbon dioxide from the atmosphere and store it in vegetation biomass or soil. Forests are covering around 30% of Earth’s land surface and combined with forest longevity, trees/soil have the potential to store carbon from decades to centuries. One key challenge is to develop methodologies for high-resolution measurements of carbon sequestered and assess year to year change. Here, we use deep neural network to generate a wall-to-wall map of AGB within the Continental USA (CONUS) with 30-meter spatial resolution for the year 2021. We combine radar and optical multispectral imagery, with a physical climate parameter of Solar Induced Fluorescence (SIF)-based Growth Primary Production (GPP). Validation results show that a masked variation of UNet has the lowest validation RMSE of 37.93 ± 1.36 Mg C/ha, as compared to 81.95 ± 0.01 Mg C/ha (linear regressor), 53.37 ± 0.05 Mg C/ha (gradient boosting), and 52.30 ± 0.03 Mg C/ha for random forest algorithm. Furthermore, models that learn from SIF-based GPP in addition to radar and optical imagery reduce validation RMSE by almost 10% and the standard deviation by 40%.
To mitigate global warming, greenhouse gas sources need to be resolved at a high spatial resolution and monitored in time to ensure the reduction and ultimately elimination of the pollution source. However, the complexity of computation in resolving high-resolution wind fields left the simulations impractical to test different time lengths and model configurations. This study presents a preliminary development of a physics-informed super-resolution (SR) generative adversarial network (GAN) that super-resolves the three-dimensional (3D) low-resolution wind fields by upscaling x9 times. We develop a pixel-wise self-attention (PWA) module that learns 3D weather dynamics via a self-attention computation followed by a 2D convolution. We also employ a loss term that regularizes the self-attention map during pretraining, capturing the vertical convection process from input wind data. The new PWA SR-GAN shows the high-fidelity super-resolved 3D wind data, learns a wind structure at the high-frequency domain, and reduces the computational cost of a high-resolution wind simulation by x89.7 times.
Methane emissions from oil and gas infrastructure, wetlands, and livestock contribute to the greenhouse gas inventory. The analysis of satellite short-wave infrared imagery offers opportunities for screening large areas to detect methane leaks. Deep learning algorithms excel at analyzing these data, however, they require large annotated datasets for model calibration that are difficult to get. To overcome this limitation, we explore a methodology to spot methane plumes using deep binary classifiers trained on a large dataset of synthetically created methane plumes, customized for this specific task, using publicly available images of the Sentine1-2 satellites. To build the database, we simulate plume patterns using the Hybrid Single-Particle Lagrangian Integrated Trajectory model (HYSPLIT) and use a simple stochastic model to account for reflectance attenuation due to methane in band 12 centered at 2190 nm. To help distinguish methane plumes from the image background, we compute a methane signature image based on a background subtraction technique. Once calibrated, the classification model is applied to image patches centered in the local minima of the methane signature within the satellite image, scoring a value ranging from 0 to 1 associated with the presence of a methane plume. We compare experimentally the general-purpose ResNet architecture and MethaNet, a domain-specific convolutional neural network, using simulated data. Then, we evaluate the feasibility of our approach in detecting large methane leaks at two study sites located in the Hassi Messaoud oil field in Algeria and the Permian Basin in the US, each covering an area of 0.25$\times$ 0.25 degrees. We found that ResNet is effective in identifying large, known methane plumes that were set aside for testing purposes. This method could be considered as a component of a solution for planning mitigation activities.
Maintaining and, ultimately, increasing vegetation coverage is likely the most impactful approach to globally capture carbon. Biomass is a crucial parameter for quantifying carbon stored in vegetation, and estimating it poses challenges as statistical models need to be customized to specific biomes. This study compares the prediction of aboveground biomass using various regression methods that were locally fitted in three distinct study sites located in Texas and Louisiana, USA. These sites (biomes) had average aboveground biomass densities of 4.1, 17.3, and 94.6 Mg/ha. The predictions obtained from these localized models were then compared to those derived from a general model that pooled data from all three sites together. Optical and radar imagery acquired from Sentinel satellites were used as predictors, while biomass density from GEDI served as the reference. In most experiments, Random Forest scored best, and the results indicate that the biome-specific models exhibited slightly higher accuracy. Specifically, the root mean square error (RMSE) values for the biome-specific models were 8.8, 16.8, and 54.8 Mg/ha, respectively. In comparison, the general model exhibited approximately 1 Mg/ha higher RMSE. The results indicate that the locally fitted models tailored to specific biomes generally outperformed the general model tested.
Soil organic carbon (SOC) sequestration is the transfer and storage of atmospheric carbon dioxide in soils, which plays an important role in climate change mitigation. SOC concentration can be improved by proper land use, thus it is beneficial if SOC can be estimated at a regional or global scale. As multispectral satellite data can provide SOC-related information such as vegetation and soil properties at a global scale, estimation of SOC through satellite data has been explored as an alternative to manual soil sampling. Although existing studies show promising results, they are mainly based on pixel-based approaches with traditional machine learning methods, and convolutional neural networks (CNNs) are uncommon. To study the use of CNNs on SOC remote sensing, here we propose the FNO-DenseNet based on the Fourier neural operator (FNO). By combining the advantages of the FNO and DenseNet, the FNO-DenseNet outperformed the FNO in our experiments with hundreds of times fewer parameters. The FNO-DenseNet also outperformed a pixel-based random forest by 18% in the mean absolute percentage error.
Source localization and emission strength quantification is an ongoing challenge for distributed pollution sources. Here we outline a wireless sensor approach to localize all potential emission sources on an oil and gas well pad under well controlled experimental conditions. Using backtracking algorithms and time synchronized methane and wind measurements, sources are attributed to equipment on the well pad. After localizing the sources, we estimate source magnitude and uncertainty using a Bayesian inference method. The approach outlined in this work can identify and quantify leaks in the close proximity of the sources under dynamic plume dispersion taking into account the site layout, potential source locations and the characteristics of the sensor network. Localization of the system is within a meter from the emission location and the Bayesian approach yields rates that are within a factor 3 of the actual rate. Further, the actual rates are generally within the 95% confidence intervals for the prediction.
Inferring the source information of greenhouse gases, such as methane, from spatially sparse sensor observations is an essential element in mitigating climate change. While it is well understood that the complex behavior of the atmospheric dispersion of such pollutants is governed by the Advection-Diffusion equation, it is difficult to directly apply the governing equations to identify the source location and magnitude (inverse problem) because of the spatially sparse and noisy observations, i.e., the pollution concentration is known only at the sensor locations and sensors sensitivity is limited. Here, we develop a multi-task learning framework that can provide high-fidelity reconstruction of the concentration field and identify emission characteristics of the pollution sources such as their location, emission strength, etc. from sparse sensor observations. We demonstrate that our proposed framework is able to achieve accurate reconstruction of the methane concentrations from sparse sensor measurements as well as precisely pin-point the location and emission strength of these pollution sources.
We present and evaluate a weakly-supervised methodology to quantify the spatio-temporal distribution of urban forests based on remotely sensed data with close-to-zero human interaction.Successfully training machine learning models for semantic segmentation typically depends on the availability of highquality labels.We evaluate the benefit of high-resolution, three-dimensional point cloud data (LiDAR) as source of noisy labels in order to train models for the localization of trees in orthophotos.As proof of concept we sense Hurricane Sandy's impact on urban forests in Coney Island, New York City (NYC) and reference it to less impacted urban space in Brooklyn, NYC.
For large scale monitoring of the environment, the number of possible pollution sources can be larger than the number of sensors. For optimal sensor placement under various wind fields in source inversion problems, this paper proposes a framework under non-Gaussian priors for the detection and inversion estimate of emission rates. The optimization framework with non-Gaussian prior utilizes a bi-level optimization expression with inner quadratic programming. The proposed truncated Gaussian prior is to incorporate non-negativity of emission rates, but it poses a challenge in optimization. We preliminarily investigate the bi-level optimization with a Gaussian plume model example. The Karush–Kuhn–Tucker conditions of the inner quadratic programming are considered for solving the bi-level optimization. The efficiency of the proposed optimization framework is demonstrated by numerical results to optimally place sensors and quantify emission rates.
Nature-based carbon sequestration is currently the most viable solutions to extract CO 2 from the atmosphere and convert it into carbon. Oceans, soils and forests have the potential to capture and store large amount of carbon for decades. There is an ongoing debate about the permanence of the carbon sequestered by nature-based processes and the precise techniques required to monitor these carbon pools. Remote sensing plays a crucial role in the large scale observations of the Earth surface and provides a scalable method to monitor land use that can affect carbon sequestration. Optical spectral information and radar signals are the best candidates as proxy data to quantify and monitor the change in carbon sequestered. Here we outline the design of an AI enabled framework to monitor, verify, and quantify carbon sequestration in nature-based carbon sequestration processes.
Vegetation, trees in particular, sequester carbon by absorbing carbon dioxide from the atmosphere. However, the lack of efficient quantification methods of carbon stored in trees renders it difficult to track the process. We present an approach to estimate the carbon storage in trees based on fusing multi-spectral aerial imagery and LiDAR data to identify tree coverage, geometric shape, and tree species—key attributes to carbon storage quantification. We demonstrate that tree species information and their three-dimensional geometric shapes can be estimated from aerial imagery in order to determine the tree’s biomass. Specifically, we estimate a total of 52, 000 tons of carbon sequestered in trees for New York City’s borough Manhattan.
A key challenge of supervised learning is the availability of human-labeled data. We evaluate a big data processing pipeline to auto-generate labels for remote sensing data. It is based on rasterized statistical features extracted from surveys such as e.g. LiDAR measurements. Using simple combinations of the rasterized statistical layers, it is demonstrated that multiple classes can be generated at accuracies of ~ 0.9.As proof of concept, we utilize the big geo-data platform IBM PAIRS to dynamically generate such labels in dense urban areas with multiple land cover classes. The general method proposed here is platform independent, and it can be adapted to generate labels for other satellite modalities in order to enable machine learning on overhead imagery for land use classification and object detection.
We present a super-resolution model for an advection-diffusion process with limited information. While most of the super-resolution models assume high-resolution (HR) ground-truth data in the training, in many cases such HR dataset is not readily accessible. Here, we show that a Recurrent Convolutional Network trained with physics-based regularizations is able to reconstruct the HR information without having the HR ground-truth data. Moreover, considering the ill-posed nature of a super-resolution problem, we employ the Recurrent Wasserstein Autoencoder to model the uncertainty.