Continuous wave indirect time-of-flight cameras obtain depth images by emitting a modulated continuous light wave and measuring the delay of the received signal. The wave is modulated sequentially using multiple frequencies. In this article, a new algorithm for the identification of the optimum set of frequencies is presented. The optimization problem is modeled with multiple objective criteria: minimizing the number of unwrapping errors and maximizing the distance estimation precision under given constraints such as the maximum modulated frequency or the minimum unambiguous range of working distances. The dependence of the proposed optimization on each constraint and the performance of the method are analyzed using Monte Carlo simulations.
In the context of Earth observation, the trade-off between spatial, spectral, and temporal resolution often limits the versatility of remote sensing images in many important applications. In response, this paper introduces a novel deep learning diffusion model, specifically tailored to improve the spatial resolution of the optical products acquired by the Sentinel-3 (S3) satellite. Our framework employs a diffusion probabilistic model, benefiting from the higher spatial resolution of the Sentinel-2 satellite during training via a new multi-modal loss formulation. This ensures consistency with the original S3 images while enhancing the spatial details. Two distinct conditional low-resolution encoders were experimented with, providing insights into their respective contributions to the diffusion process. The efficacy of the proposed model is demonstrated through extensive ablation studies and comparisons with state-of-the-art methods, using both synthetic and real S3 products. The findings indicate that our model successfully improves spatial resolution while maintaining the integrity of the spectral information, contributing to the field of remote sensing single-image super-resolution.
Accurate depth estimation is crucial in various computer vision applications, such as robotics, augmented reality, or autonomous driving. Despite the common use of Time-ofFlight (ToF) sensing systems, they still face challenges such as invalid pixels and missing depth values, particularly with low light reflectance, distant objects, or light-saturated conditions. Cameras using indirect ToF technology provide depth maps along with active infrared brightness images, which can offer a potential guide for depth restoration in fusion approaches. This study proposes a method for depth completion by combining depth and active infrared images in ToF systems. The approach is based on a belief propagation strategy to extend valid nearby information in missing depth regions, using the infrared gradient for depth consistency. Emphasis is placed on considering object edges, especially those coinciding with depth discontinuities, to approximate missing values. Empirical results demonstrate the efficiency and simplicity of the proposed algorithm, showcasing superior outcomes compared to other reference guided depth inpainting methods.
Motivated by the increasing demand for robust segmentation in unlabeled remote sensing data, we propose domain adaptation multimodal and multi-temporal transformer (DAM-Former), a novel unsupervised domain adaptation (UDA) model that fuses high-resolution (HR) multimodal imagery with multi-temporal multispectral data. Current UDA approaches in remote sensing rarely exploit the complementary strengths of spatial and temporal features. To address this gap, our framework integrates two interconnected branches: a transformer-based network for HR multimodal data and a lightweight convolutional network with temporal attention for multi-temporal imagery. To improve segmentation accuracy and lower noise, the extracted features are robustly combined through a deep temporal fusion (DTF) module and a new mixed loss (ML) with an ensemble pseudo-label (EP) strategy. Extensive experiments and an ablation study on the FLAIR-2 dataset demonstrate that DAM-Former outperforms state-of-the-art methods, marking the first in-depth study of temporal information fusion in UDA segmentation for remote sensing data. Code available at https://github.com/ibanezfd/DAM_Former.
Detecting rainfall-induced shallow landslides in data-sparse regions has become increasingly important for effective landslides disaster management. Previous studies have predominantly focused on automated methods for deep-seated, earthquake-triggered landslides. This study addresses this gap by employing a U-net Convolutional Neural Network (CNN) model to detect rainfall-induced shallow landslides using multi-temporal, high-resolution PlanetScope (3m spatial resolution), medium-resolution Sentinel-2 (10m spatial resolution) imagery, and ALOS-PALSAR-provided digital elevation model (DEM). Four datasets were created: Datasets A and B using PlanetScope, and Datasets C and D using Sentinel-2, with Datasets B and D also including DEM data. A total of 181 manually delineated landslide polygons were used as ground truth masks. Each dataset was tested using repeated stratified hold-out validation. Performance metrics included precision, recall, F1 score, loss, and accuracy. Results indicated that Datasets A and B outperformed the others; however, integrating DEM with Dataset B did not enhance model accuracy. The best mean precision, recall, F1 score, loss, and accuracy were 1, 0.625, 0.625, 0.380, and 0.999, respectively, for both Datasets A and B. This study demonstrates the U-net model's potential for detecting rainfall-induced shallow landslides in various geographic and temporal contexts globally.
In this work, we present an overview of human gesture recognition in degraded environments with multi-dimensional integral imaging. It is shown that for human gesture recognition in degraded environments such as low light, and occlusion, we can gain substantial improvements in performance over conventional imaging.
Earth data collection from satellites and aircraft has exponentially grown, but a substantial portion of it remains unlabeled. This has prompted the remote sensing community to explore effective methods for leveraging unlabeled data. In our prior investigation, we evaluated various deep semi-supervised learning algorithms on two very high-resolution (VHR) optical datasets (UCM and AID). Notably, the CoMatch algorithm demonstrated the highest accuracy, motivating further exploration. This letter extends our earlier work by integrating the established class-aware contrastive semi-supervised learning framework (CoMatch + CCSSL) into CoMatch and introducing a new triplet metric learning loss (CoMatch + Triplet). CoMatch + Triplet excelled with 93.2% accuracy on UCM, while CoMatch led with 92.19% on AID. The addition of the triplet loss can produce a clearer separation of the samples from different classes in the embedding space at very early learning stages, being able to learn faster and getting maximum performance with few iterations. The exploration of diverse semi- and self-supervised training methodologies presented in this work sheds light on the strengths and limitations of these approaches, enhancing our understanding of their applicability in remote sensing applications.
Current and upcoming Sun-Induced chlorophyll Fluorescence (SIF) satellite products (e.g., GOME, TROPOMI, OCO, FLEX) have medium-to-coarse spatial resolutions (i.e., 0.3–80 km) and integrate radiances from different sources into a single ground surface unit (i.e., pixel). However, intrapixel heterogeneity, i.e., different soil and vegetation fractional cover and/or different chlorophyll content or vegetation structure in a fluorescence pixel, increases the challenge in retrieving and quantifying SIF. High spatial resolution Sentinel-2 (S2) data (20 m) can be used to better characterize the intrapixel heterogeneity of SIF and potentially extend the application of satellite-derived SIF to heterogeneous areas. In the context of the COST Action Optical synergies for spatiotemporal SENsing of Scalable ECOphysiological traits (SENSECO), in which this study was conducted, we proposed direct (i.e., spatial heterogeneity coefficient, standard deviation, normalized entropy, ensemble decision trees) and patch mosaic (i.e., local Moran’s I) approaches to characterize the spatial heterogeneity of SIF collected at 760 and 687 nm (SIF760 and SIF687, respectively) and to correlate it with the spatial heterogeneity of selected S2 derivatives. We used HyPlant airborne imagery acquired over an agricultural area in Braccagni (Italy) to emulate S2-like top-of-the-canopy reflectance and SIF imagery at different spatial resolutions (i.e., 300, 20, and 5 m). The ensemble decision trees method characterized FLEX intrapixel heterogeneity best (R2 > 0.9 for all predictors with respect to SIF760 and SIF687). Nevertheless, the standard deviation and spatial heterogeneity coefficient using k-means clustering scene classification also provided acceptable results. In particular, the near-infrared reflectance of terrestrial vegetation (NIRv) index accounted for most of the spatial heterogeneity of SIF760 in all applied methods (R2 = 0.76 with the standard deviation method; R2 = 0.63 with the spatial heterogeneity coefficient method using a scene classification map with 15 classes). The models developed for SIF687 did not perform as well as those for SIF760, possibly due to the uncertainties in fluorescence retrieval at 687 nm and the low signal-to-noise ratio in the red spectral region. Our study shows the potential of the proposed methods to be implemented as part of the FLEX ground segment processing chain to quantify the intrapixel heterogeneity of a FLEX pixel and/or as a quality flag to determine the reliability of the retrieved fluorescence.
The increasing availability of remote sensing data allows dealing with spatial-spectral limitations by means of pan-sharpening methods. However, fusing inter-sensor data poses important challenges, in terms of resolution differences, sensor-dependent deformations and ground-truth data availability, that demand more accurate pan-sharpening solutions. In response, this paper proposes a novel deep learning-based pan-sharpening model which is termed as the double-U network for self-supervised pan-sharpening (W-NetPan). In more details, the proposed architecture adopts an innovative W-shape that integrates two U-Net segments which sequentially work for spatially matching and fusing inter-sensor multi-modal data. In this way, a synergic effect is produced where the first segment resolves inter-sensor deviations while stimulating the second one to achieve a more accurate data fusion. Additionally, a joint loss formulation is proposed for effectively training the proposed model without external data supervision. The experimental comparison, conducted over four coupled Sentinel-2 and Sentinel-3 datasets, reveals the advantages of W-NetPan with respect to several of the most important state-of-the-art pan-sharpening methods available in the literature. The codes related to this paper will be available at https://github.com/rufernan/WNetPan.
This paper introduces an optimization energy framework based on infrared guidance to improve depth consistency in Time of Flight image systems. The primary objective is to formulate the problem as an image energy optimization task, aimed at maximizing the coherence between the depth map and the corresponding infrared image, both captured simultaneously from the same Time of Flight sensor. The concept of depth consistency relies on the underlying hypothesis concerning the correlation between depth maps and their corresponding infrared images. The proposed optimization framework adopts a weighted approach, leveraging an iterative estimator. The image energy is characterized by introducing spatial conditional entropy as a correlation measure and spatial error as image regularization. To address the issue of missing depth values, a preprocessing step is initially applied, by using a depth completion method based on infrared guided belief propagation, which was proposed in a previous work. Subsequently, the proposed framework is employed to regularize and enhance the inpainted depth. The experimental results demonstrate a range of qualitative improvements in depth map reconstruction, with a particular emphasis on the sharpness and continuity of edges.
Continuous wave Time of Flight cameras obtain depth images by emitting a modulated continuous light wave and measuring the delay of the received signal. One of the main sources of error is shot noise. In this letter we analyze and provide insights about the effect of shot noise when obtaining the phase delay with a new architecture, based on differentiating the phase integrations in the analog domain, and using any number of points in the calculation of the Discrete Fourier Transform (DFT). One of the main findings is that the error depends on the integration time regardless the number of points in the DFT used to calculate the phase offset. Simulated experiments are provided which support the proposed theoretical analysis.
Image registration is an essential task in image processing, where the final objective is to geometrically align two or more images. In remote sensing, this process allows comparing, fusing, or analyzing data, especially when multimodal images are used. In addition, multimodal image registration becomes fairly challenging when the images have a significant difference in scale and resolution, together with local small image deformations. For this purpose, this letter presents a novel optical flow (OF)-based image registration network, named the FloU-Net, which tries to further exploit intersensor synergies by means of deep learning. The proposed method is able to extract spatial information from resolution differences and through a U-Net backbone generate an OF field estimation to accurately register small local deformations of multimodal images in a self-supervised fashion. For instance, the registration between Sentinel-2 (S2) and Sentinel-3 (S3) optical data is not trivial, as there are considerable spectral–spatial differences among their sensors. In this case, the higher spatial resolution of S2 results in S2 data being a convenient reference to spatially improve S3 products, as well as those of the forthcoming Fluorescence Explorer (FLEX) mission, since image registration is the initial requirement to obtain higher data processing level products. To validate our method, we compare the proposed FloU-Net with other state-of-the-art techniques using 21 coupled S2/S3 optical images from different locations of interest across Europe. The comparison is performed through different performance measures. Results show that the proposed FloU-Net can outperform the compared methods. The code and dataset are available in https://github.com/ibanezfd/FloU-Net .
Depth information has been successfully used in many computer vision applications, but depth imaging sensors frequently provide missing values, mainly around objects boundaries. These invalid values and image gaps cause serious problems in some applications. In order to estimate missing depth values and fill gaps in depth images (D), we propose a new algorithm for depth completion based on belief propagation. The rationale of the proposed technique is based on the idea that missing values must be estimated by taking into account object boundaries, mainly those related with depth discontinuities. Time of Flight (ToF) cameras provide depth information and some additional data, such as active infrared (IR) brightness images. Therefore, object boundaries information for depth missing areas can be reconstructed by using auxiliary IR information or by RGB images in RGB-D systems. These auxiliary images are used as a guidance for the depth completion, also known as depth inpainting. Experimental results show that our algorithm is very simple to implement, fast and produces better results than other more complex, and usually slower, existing methods.
Continuous wave indirect Time-of-Flight cameras obtain depth images by emitting a modulated continuous light wave and measuring the delay of the received signal. In this paper we generalize the estimation of the effect of the shot noise when obtaining the phase delay with an arbitrary number of points in the Discrete Fourier Transform, extending and generalizing the analysis done in previous works for the case of four points. For that particular case, we compare our analysis with the state of art. Moreover, we extend the error model using a second order approximation in the error propagation analysis, which provides more accurate estimations according to the Montecarlo simulation experiments. The analysis, based on both analytical and numerical methods, shows that the phase error is, in general, related to the exposure time and weakly to the number of points in the Discrete Fourier Transform. It also depends on the background illumination level, on the amplitude of the received signal, and, when using a three point DFT, on the distance to the objects.
Continuous wave indirect Time-of-Flight cameras obtain depth images by emitting a modulated continuous light wave and measuring the delay of the received signal. In this paper we generalize the estimation of the effect of the shot noise when obtaining the phase delay with an arbitrary number of points in the Discrete Fourier Transform, extending and generalizing the analysis done in previous works for the case of four points. For that particular case, we compare our analysis with the previous analysis. Moreover, we extend the error model using a second order approximation in the error propagation analysis, which provides more accurate estimations according to the Montecarlo simulation experiments. The analysis, based on both analytical and numerical methods, shows that the phase error is, in general, related to the exposure time and weakly to the number of points in the Discrete Fourier Transform. It also depends on the background illumination level, the square of the amplitude of the received signal, and on the distance to the objects when using a three point DFT.
Deep semi-supervised learning (DSSL) is a rapidly-growing field that takes advantage of a limited number of labeled examples to leverage massive amounts of unlabeled data. The underlying idea is that training on small yet well-selected examples can perform as effectively as a predictor trained on a larger number chosen at random [14]. In this study, we explore the most relevant approaches in DSSL literature like FixMatch [19], CoMatch [13], and, the class aware contrastive SSL (CCSSL) [25]. Our objective is to perform an initial comparative study of these methods and assess them on two remote sensing (RS) datasets: UCM [27] and AID [22]. The performance of these methods was determined based on their accuracy in comparison to a supervised benchmark. The results highlight that the CoMatch framework achieves the highest accuracy for both the UCM and AID datasets, with accuracies of 95.52% and 93.88% respectively. Importantly, all DSSL algorithms outperform the supervised benchmark, emphasizing their effectiveness in leveraging a limited number of labeled examples to enhance classification accuracy for remote sensing scene classification tasks. The code used in this study was adapted from CCSSL [25] and the detailed implementation will be accessible at https://github.com/itzahs/SSL-for-RS .
Deep learning has certainly become the dominant trend in hyperspectral (HS) remote sensing (RS) image classification owing to its excellent capabilities to extract highly discriminating spectral–spatial features. In this context, transformer networks have recently shown prominent results in distinguishing even the most subtle spectral differences because of their potential to characterize sequential spectral data. Nonetheless, many complexities affecting HS remote sensing data (e.g., atmospheric effects, thermal noise, quantization noise) may severely undermine such potential since no mode of relieving noisy feature patterns has still been developed within transformer networks. To address the problem, this article presents a novel masked auto-encoding spectral–spatial transformer (MAEST), which gathers two different collaborative branches: 1) a reconstruction path, which dynamically uncovers the most robust encoding features based on a masking auto-encoding strategy, and 2) a classification path, which embeds these features onto a transformer network to classify the data focusing on the features that better reconstruct the input. Unlike other existing models, this novel design pursues to learn refined transformer features considering the aforementioned complexities of the HS remote sensing image domain. The experimental comparison, including several state-of-the-art methods and benchmark datasets, shows the superior results obtained by MAEST. The codes of this article will be available at https://github.com/ibanezfd/MAEST.
Monitoring water bodies from remote sensing data is certainly an essential task to supervise the actual conditions of the available water resources for environment conservation, sustainable development, and many other applications. Being Sentinel-2 images some of the most attractive data, existing traditional index-based and deep learning-based water extraction methods still have important limitations in effectively dealing with large heterogeneous areas since many types of water bodies with different spatial-spectral complexities are logically expected. Note that, in this scenario, optimal feature abstraction and neighborhood information may certainly vary from water to water pixel, however existing methods are generally constrained by a fix abstraction level and amount of land cover context. To address these issues, this article presents a new attentional dense convolutional neural network (AD-CNN) especially designed for water body extraction from Sentinel-2 imagery. On the one hand, the AD-CNN exploits dense connections to allow uncovering deeper features while simultaneously characterizing multiple data complexities. On the other hand, the proposed model also implements a new residual attention module to dynamically put the focus on the most relevant spatial-spectral features for classifying water pixels. To test the performance of the AD-CNN, a new water database of Nepal (WaterPAL) is also built. The conducted experiments reveal the competitive performance of the proposed architecture with respect to several traditional index-based and state-of-the-art deep learning-based water extraction models.
This paper lies at the intersection of three research areas: human action recognition, egocentric vision, and visual event-based sensors. The main goal is the comparison of egocentric action recognition performance under either of two visual sources: conventional images, or event-based visual data. In this work, the events, as triggered by asynchronous event sensors or their simulation, are spatio-temporally aggregated into event frames (a grid-like representation). This allows to use exactly the same neural model for both visual sources, thus easing a fair comparison. Specifically, a hybrid neural architecture combining a convolutional neural network and a recurrent network is used. It is empirically found that this general architecture works for both, conventional gray-level frames, and event frames. This finding is relevant because it reveals that no modification or adaptation is strictly required to deal with event data for egocentric action classification. Interestingly, action recognition is found to perform better with event frames, suggesting that these data provide discriminative information that aids the neural model to learn good features.
Depth estimation is one of the fundamentals tasks in 3D imaging. In this work we present a new method based on a weighted energy optimization to enhance depth images captured from Time of Flight (ToF) image systems. We propose a fusion approach for ToF depth image enhancement, with the purpose of removing noise, solve depth discontinuities and filling missing values. The proposed method is based on a combination of the raw depth image and its corresponding near-infrared image provided by the same ToF image sensor. Experiments on the guided depth image enhancement and noisy depth restoration provide satisfactory results compared to conventional guide image filters in a fusion context.