Nitrogen dioxide (NO2) is one among several constituents of air pollution. To restrict its surface-level concentration to within the limits prescribed by regulatory authorities, dedicated monitoring of its spatiotemporal spread is needed. Satellite-based remote sensing of the tropospheric composition can be used to estimate NO2 concentration. However, this is not a direct measurement of the surface-level NO2 concentration, though several studies have shown that the tropospheric vertical column density (VCD) estimated by the satellite sensor is correlated with surface-level concentration. This review article covers various aspects related to the estimation of surface-level NO2 using remotely sensed data. It provides detailed literature, tracing the evolution of the various methods developed for the estimation, from scaling methods to the initial linear regression (LR) models onward to the more recent deep learning (DL) architectures. The performance of these models is critically reviewed.
The High Mountain Asia (HMA) continues to witness an increased frequency of glacial lake outburst floods (GLOFs), which is likely in response to continued global warming. In situ measurements to understand the triggers all across the region will remain inadequate given the vastness and lack of accessibility of the region. This work explores a data-driven logistic regression-based framework to evaluate potential GLOF triggers, such as the lake dam type, its surface area, aspect, distance, freeboard, slope, precipitation, and temperature. A comprehensive inventory of past GLOF events in the region has been compiled, with 25 events verified using pre-& post-event multispectral images acquired between 2016 and 2022. The logistic regression model is developed using samples of the positive class (lakes with confirmed GLOFs) and of the negative class (potentially dangerous lakes that have not experienced a GLOF event). The samples of the negative lake class were collected with resembling characteristics from the nearby areas of the positive class. We randomly keep 80% of the samples for training. The models performance is assessed using adjusted R 2 and the Akaike Information Criterion (AIC) on the test samples, which are 79% and 21.5, respectively. The classification accuracy is 90%, which is promising. In short, the proposed method is a useful tool to investigate risk of outburst flooding of glacial lakes.
Detecting firearms and accurately localizing individuals carrying them in images or videos is of paramount importance in security, surveillance, and content customization. However, this task presents significant challenges in complex environments due to clutter and the diverse shapes of firearms. To address this problem, we propose a novel approach that leverages human–firearm interaction information, which provides valuable clues for localizing firearm carriers. Our approach incorporates an attention mechanism that effectively distinguishes humans and firearms from the background by focusing on relevant areas. Additionally, we introduce a saliency-driven locality-preserving constraint to learn essential features while preserving foreground information in the input image. By combining these components, our approach achieves exceptional results on a newly proposed dataset. To handle inputs of varying sizes, we pass paired human–firearm instances with attention masks as channels through a deep network for feature computation, utilizing an adaptive average pooling (AAP) layer. We extensively evaluate our approach against existing methods in human–object interaction (HOI) detection and achieve significant results (AP $=$ 77.8%) compared to the baseline approach (AP $=$ 63.1%). This demonstrates the effectiveness of leveraging attention mechanisms and saliency-driven locality preservation for accurate human–firearm interaction detection. Our findings contribute to advancing the fields of security and surveillance, enabling more efficient firearm localization and identification in diverse scenarios.
Oil spillages on a sea’s or an ocean’s surface are a threat to marine and coastal ecosystems. They are mainly caused by ship accidents, illegal discharge of oil from ships during cleaning and oil seepage from natural reservoirs. Synthetic-Aperture Radar (SAR) has proved to be a useful tool for analyzing oil spills, because it operates in all-day, all-weather conditions. An oil spill can typically be seen as a dark stretch in SAR images and can often be detected through visual inspection. The major challenge is to differentiate oil spills from look-alikes, i.e., low-wind areas, algae blooms and grease ice, etc., that have a dark signature similar to that of an oil spill. It has been noted over time that oil spill events in Pakistan’s territorial waters often remain undetected until the oil reaches the coastal regions or it is located by concerned authorities during patrolling. A formal remote sensing-based operational framework for oil spills detection in Pakistan’s Exclusive Economic Zone (EEZ) in the Arabian Sea is urgently needed. In this paper, we report the use of an encoder–decoder-based convolutional neural network trained on an annotated dataset comprising selected oil spill events verified by the European Maritime Safety Agency (EMSA). The dataset encompasses multiple classes, viz., sea surface, oil spill, look-alikes, ships and land. We processed Sentinel-1 acquisitions over the EEZ from January 2017 to December 2023, and we thereby prepared a repository of SAR images for the aforementioned duration. This repository contained images that had been vetted by SAR experts, to trace and confirm oil spills. We tested the repository using the trained model, and, to our surprise, we detected 92 previously unreported oil spill events within those seven years. In 2020, our model detected 26 oil spills in the EEZ, which corresponds to the highest number of spills detected in a single year; whereas in 2023, our model detected 10 oil spill events. In terms of the total surface area covered by the spills, the worst year was 2021, with a cumulative 395 sq. km covered in oil or an oil-like substance. On the whole, these are alarming figures.
In recent years, there has been a noticeable increase in the inclination towards digitizing our surroundings, encompassing various domains such as virtual reality, cultural heritage conservation, and architectural representation. The computation of high-resolution three-dimensional (3D) colored point clouds and meshes holds significant importance for such applications. However, traditional structure-from-motion (SfM) techniques may produce sparse 3D point clouds when low-resolution input images are used, resulting in a low-quality mesh generation. Traditional point cloud upsampling techniques that improve the 3D point cloud resolution typically work on LiDAR-generated point clouds devoid of color information. Furthermore, most learned point cloud upsampling techniques compute graph features that capture local information by identifying a local neighborhood in a limited region around a point and hence may result in sub-optimal representation. To address these limitations, we propose CloudUP, a colored 3D point cloud upsampling approach that utilizes multi-scale spatial attention. Specifically, we design a novel Multi-Scale Point-Cloud Feature Extractor (MPFE) by employing attention across the scales to extract point cloud features and effectively capture 3D shape information of the points relative to its neighborhood. We further extract spatial neighborhood-guided color features used to predict the color for the upsampled points. The color prediction is trained with a content-preserving loss function that aims to maintain intricate details and vivid colors. Our color refinement pipeline is guided by a vibrant colored dataset (collected by us) to assist in preserving the 3D contents.
A glacial lake outburst flood (GLOF) is typically a natural phenomenon caused by rapid discharge of water from a glacier, leading to a flood. The frequency of GLOFs has increased significantly in the northern areas of Pakistan, which demands identification and continuous monitoring of potentially dangerous glacial lakes. In this paper, an up-to-date inventory of glacial lakes in this region is presented. This inventory (HKH-PK-2020) has been prepared using high resolution PlanetScope imagery acquired in 2020 over northern Pakistan. It contains a total of 8808 lakes. We compare our database with the High Mountain Asia (HMA) glacial lakes inventory over northern Pakistan, prepared in 2018 using Landsat imagery. The new inventory contains 6537 more glacial lakes than the HMA inventory. Furthermore, we have prepared an annotated dataset containing 3525 images (of high resolution PlanetScope imagery over a selected number of lakes from the inventory). Each image comprises 4 bands, namely red, green, blue, and near infrared. The annotations are binary: lake or background. Finally, we have performed an ablation study with two encoder-decoder based convolutional neural networks (CNNs) trained on this dataset for pixel-based classification. Our results show an intersection over union (IoU) score of 72.81% for the lake class, which is a promising first result indicating a use of deep learning for automated inventory updates in future.
Vegetation cover classification using mixed or low-resolution scalar images is challenging. Fortunately, recently deep learning object detection methods have emerged as a replacement to the conventional machine learning methods for the detection and classification of land use and land cover. This paper presents a deep learning object detection approach for land use and land cover detection using low/mixed resolution satellite images acquired from Google Earth satellite images. Google Earth images are accessible freely using the Google Earth Pro desktop application. Our dataset consists of two (02) classes (vegetation and non-vegetation) with a total of 450 labeled images captured from different parts of Pakistan. We present a comparison of the recent anchor-free object detection model YOLOX with the anchor-based object detection model YOLOR for solving real-time problems. The end-to-end differentiability, efficient GPU utilization, and absence of hand-crafted parameters make anchor-free models a compelling choice in object detection, and yet not been explored on Land cover classification using satellite images. Our experimental study shows that YOLOX delivers an overall accuracy of 83.50% on Vegetation and 86% on Non-Vegetation classes, which outperformed YOLOR by 30% on Vegetation classes and 34% on non-Vegetation classes for our dataset. We also show how an object detection system can be used for Vegetation and Non-Vegetation classification tasks, which can then be used for change monitoring and assisting in developing geographical maps using low/mixed resolution freely available satellite images.
Nitrogen dioxide (NO 2 ) is one of the six gaseous air pollutants that need regular monitoring in big cities around the world. It contributes to particle pollution and can trigger chemical reactions that lead to increased concentration of ozone in the troposphere. Lahore, a metropolitan city of Pakistan is among the most polluted cities in the world. Area-wide monitoring of NO 2 is necessary in this region to devise a long-term emission control policy. However, it lacks a dense network of ground-based air quality monitoring stations (AQMS), which is need of the hour. The installation of AQMS requires huge financial resources. In this paper, we investigate a machine learning-based approach to estimate surface level concentration of NO 2 using remote sensing and modeled meteorological data. We use multiple linear regression (M1) and a polynomial fitted regression (M2) techniques to model ambient NO 2 , using remotely sensed vertical column density (VCD) of NO 2 , acquired by tropospheric monitoring instrument (TROPOMI), onboard Sentinel 5P satellite, and modeled meteorological parameters such as surface pressure, dew point temperature, and wind speed. Results show that M2 outperformed M1 with an $\mathbf{R}^{2}$ value of 0.49 and root mean square error (RMSE) value of $\mathbf{19}.\mathbf{27}\ \mu \mathbf{g}/\mathbf{m}^{3}$ . There is a moderate positive correlation between in-situ measurements and remotely sensed VCD of NO 2 , which makes it an interesting problem that needs to be explored further to achieve desirable results.
Glacial lake outburst floods (GLOFs) are a major threat to the local communities and important infrastructures in the high mountain regions. This paper focuses on the development of a benchmark dataset for glacial lakes classification in Sentinel 2 multi-spectral data and subsequent detection of glacial lakes prior to a glacial lake outburst flood (GLOF). Towards this end, we collected Sentinel 2 true color scenes of High-Mountain Asia (HMA) region using glacial lakes inventory of this region. It covers an area of 2080.12 km 2 with nearly 30,121 glacial lakes. After data collection, we retained 1200 cloud free true color images and manually generated their ground truth masks. The dataset covers lakes with different shapes, sizes and radiometric signatures. For detection of glacial lakes, we used an encoder-decoder based convolutional neural network (CNN). The model is trained on the labelled dataset of glacial lakes for semantic segmentation of true color images into two relevant classes: lake and no lake. The performance of the proposed model is evaluated using intersection over union (IoU) score. It classifies glacial lakes correctly with an IoU score of 79.90%, which is quite good as far as complexity of the problem is concerned.
Oil spillage over a sea or ocean surface is a threat to marine and coastal ecosystems. Spaceborne synthetic aperture radar (SAR) data have been used efficiently for the detection of oil spills due to their operational capability in all-day all-weather conditions. The problem is often modeled as a semantic segmentation task. The images need to be segmented into multiple regions of interest such as sea surface, oil spill, lookalikes, ships, and land. Training of a classifier for this task is particularly challenging since there is an inherent class imbalance. In this work, we train a convolutional neural network (CNN) with multiple feature extractors for pixel-wise classification and introduce a new loss function, namely, “gradient profile” (GP) loss, which is in fact the constituent of the more generic spatial profile loss proposed for image translation problems. For the purpose of training, testing, and performance evaluation, we use a publicly available dataset with selected oil spill events verified by the European Maritime Safety Agency (EMSA). The results obtained show that the proposed CNN trained with a combination of GP, Jaccard, and focal loss functions can detect oil spills with an intersection over union (IoU) value of 63.95%. The IoU value for sea surface, lookalikes, ships, and land class is 96.00%, 60.87%, 74.61%, and 96.80%, respectively. The mean intersection over union (mIoU) value for all the classes is 78.45%, which accounts for a 13% improvement over the state of the art for this dataset. Moreover, we provide extensive ablation on different convolutional neural networks (CNNs) and vision transformers (ViTs)-based hybrid models to demonstrate the effectiveness of adding GP loss as an additional loss function for training. Results show that GP loss significantly improves the mIoU and F1 scores for CNNs as well as ViTs-based hybrid models. GP loss turns out to be a promising loss function in the context of deep learning with SAR images.
Oil spills cause a significant threat to marine and coastal ecosystems. It is one of the major causes of water pollution. This research focuses on the use of deep learning for oil spills detection and classification. UNet is a convolutional neural network, originally proposed for biomedical image segmentation and modified for the discrimination of oil spills and look-alikes. The model is trained on a publicly available benchmark oil spill detection dataset of Sentinel-1 synthetic aperture radar (SAR) images. The images have been semantically segmented into multiple regions of interest such as sea surface, oil spills, look-alikes, ships and land. The proposed UNet-based model achieves intersection over union (IoU) value of 95.69% for sea surface, 60.85% for oil spills, 54.90% for look-alikes, 70.27% for ships and 96.79% for land class. The mean intersection over union (mIoU) value for all the classes is 75.70% which consitutes a nearly 10% increase compared to state of the art for this dataset.
Visual identification of gunmen in a crowd is a challenging problem, that requires resolving the association of a person with an object (firearm). We present a novel approach to address this problem, by defining human-object interaction (and non-interaction) bounding boxes. In a given image, human and firearms are separately detected. Each detected human is paired with each detected firearm, allowing us to create a paired bounding box that contains both object and the human. A network is trained to classify these paired-bounding-boxes into human carrying the identified firearm or not. Extensive experiments were performed to evaluate the effectiveness of the algorithm, including exploiting full pose of the human, hand-keypoints, and their association with the firearm. The knowledge of spatially localized features is key to the success of our method by using multi-size proposals with adaptive average pooling. We have also extended a previously existing firearm detection dataset, by adding more images and tagging in the extended dataset the human-firearm pairs (including bounding boxes for firearms and gunmen). The experimental results (78.5 AP(hold)) demonstrate effectiveness of the proposed method.