Spaceborne LiDAR systems are crucial for Earth observation but face hardware constraints, thus limiting resolution and data processing. We propose integrating compressed sensing and diffusion generative models to reconstruct high-resolution satellite LiDAR data within the Hyperheight Data Cube (HHDC) framework. Using a randomized illumination pattern in the imaging model, we achieve efficient sampling and compression, reducing the onboard computational load and optimizing data transmission. Diffusion models then reconstruct detailed HHDCs from sparse samples on Earth. To ensure reliability despite lossy compression, we analyze distortion metrics for derived products like Digital Terrain and Canopy Height Models and evaluate the 3D reconstruction accuracy in waveform space. We identify image quality assessment metrics—ADD_GSIM, DSS, HaarPSI, PSIM, SSIM4, CVSSI, MCSD, and MDSI—that strongly correlate with subjective quality in reconstructed forest landscapes. This work advances high-resolution Earth observation by combining efficient data handling with insights into LiDAR imaging fidelity.
Spaceborne LiDAR reconstructions from compressed HyperHeight Data Cubes (HHDCs) are increasingly obtained with diffusion models, but reliable quality assessment (QA) tools tailored to LiDAR-derived rasters are lacking. This paper closes that gap by 1) building a subjective database for HHDC percentile products-Canopy Height Models (CHMs), median height maps, and Digital Terrain Models (DTMs)-with 342 reconstruction/reference pairs and Mean Opinion Scores (MOS) collected from multiple observers in both pseudo-colored and grayscale renderings; 2) benchmarking 70+ full-reference image quality assessment (FR-IQA) metrics against MOS; and 3) proposing a compact nonlinear fusion of top performers. Data-driven FR IQA methods (DISTS, TOPIQ trained on PIPAL, LPIPS/LPIPS+/ST-LPIPS) exhibit the best correlation with human opinion on LiDAR-derived rasters, achieving absolute PLCC/SROCC values of around 0.85-0.86 on aggregate. A simple weighted sum-product fusion of five complementary metrics further increases alignment to PLCC/SROCC of 0.9221/0.9204 across all percentiles and modalities. Cross-modal consistency of the subjective scores is high (e.g., grayscale vs. pseudo-color for the same percentile), while pseudo-colored DTMs exhibit the most outliers, indicating colormap- and homogeneity-induced instability in some metrics. These findings enable practical QA for generative LiDAR reconstruction: the fused metric serves as a reliable MOS surrogate for evaluation and as a perceptual loss to regularize diffusion-based reconstructions. The released dataset and benchmarks provide a reference for future LiDAR-specific IQA and extensions to tensor-level HHDC assessment.
This paper presents the interesting results of applying compression ratio (CR) in the prediction of the boundary between visually lossless and visually lossy compression, which is of particular importance in perceptual image compression. The prediction is carried out through the objective quality (peak signal-to-noise ratio, PSNR) and image representation in bits per pixel (bpp). In this analysis, the results of subjective tests from four publicly available databases are used as ground truth for comparison with the results obtained using the compression ratio as a predictor. Through a wide analysis of color and grayscale infrared JPEG and Better Portable Graphics (BPG) compressed images, the values of parameters that control these two types of compression and for which CR is calculated are proposed. It is shown that PSNR and bpp predictions can be significantly improved by using CR calculated using these proposed values, regardless of the type of compression and whether color or infrared images are used. In this paper, CR is used for the first time in predicting the boundary between visually lossless and visually lossy compression for images from the infrared part of the electromagnetic spectrum, as well as in the prediction of BPG compressed content. This paper indicates the great potential of CR so that in future research, it can be used in joint prediction based on several features or through the CR curve obtained for different values of the parameters controlling the compression.
A tendency to increase the number of acquired remote sensing images and to make their average size larger has been observed. To manage such data, compression is needed, and lossy compression is often preferable. Since lossy compression introduces distortions, this results in worse classification and object detection. Therefore, lossy compression must be controlled, i.e., the introduced distortions must be under a certain limit. The distortions and the limit can be characterized by different metrics (quantitative criteria). Here, we consider the case of using the HaarPSI metric, which has a very high correlation with visual quality and human attention (saliency map), for three-channel optical band images compressed by the better portable graphics (BPG) encoder, one of the best modern compression techniques. We analyze a two-step procedure of providing a desired visual quality and show its peculiarities for the modes 4:4:4, 4:2:2, and 4:2:0 of image compression. We show how the HaarPSI metric relates to other known metrics of image visual quality and thresholds of distortion visibility. It is demonstrated that the two-step procedure provides about three times better accuracy in providing the desired visual quality compared to the fixed setting of parameter Q that controls compression for the BPG encoder. The provided accuracy is close to the reachable limit determined by the integer value setting of the Q parameter. We also briefly analyze the influence of compression on the classification accuracy of real-life remote sensing data.
This chapter contains the results obtained during execution of Ukrainian-Polish Project in 2020 intended on design of methods and means for processing grayscale and multichannel images and video using visual quality metrics. The combined metrics have been
This chapter contains the results obtained during execution of Ukrainian-French Project within “Dnipro” framework in 2022 intended on design of methods and means for processing multichannel remote sensing data using the trained neural networks. Neural net
Unmanned aerial vehicle (UAV) imaging is a dynamically developing field, where the effectiveness of imaging applications highly depends on quality of the acquired images.Noreference image quality assessment is widely used for quality control and image processing management.However, there is a lack of accuracy and adequacy of existing quality metrics for human visual perception.In this paper, we demonstrate that this problem persists for typical applications of UAV images.We present a methodology to improve the efficiency of visual quality assessment by existing metrics for images obtained from UAVs, and introduce a method of combining quality metrics with the optimal selection of the elementary metrics used in this combination.A combined metric is designed based on a neural network trained to utilize subjective assessments of visual quality.The metric was tested using the TID2013 image database and a set of real UAV images with embedded distortions.Verification results have demonstrated the robustness and accuracy of the proposed metric.
This paper deals with providing the desired quality in the Better Portable Graphics (BPG)-based lossy compression of color and three-channel remote sensing (RS) images. Quality is described by the Mean Deviation Similarity Index (MDSI), which is proven to be one of the best metrics for characterizing compressed image quality due to its high conventional and rank-order correlation with the Mean Opinion Score (MOS) values. The MDSI properties are studied and three main areas of interest are determined. It is shown that quite different quality and compression ratios (CR) can be observed for the same values of the quality parameter Q that controls compression, depending on the compressed image complexity. To provide the desired quality, a modified two-step procedure is proposed and tested. It has a preliminary stage carried out offline (in advance). At this stage, an average rate-distortion curve (MDSI on Q) is obtained and it is available until the moment when a given image has to be compressed. Then, in the first step, an image is compressed using the starting Q determined from the average rate-distortion curve for the desired MDSI. After this, the image is decompressed and the produced MDSI is calculated. In the second step, if necessary, the parameter Q is corrected using the average rate-distortion curve, and the image is compressed with the corrected Q. Such a procedure allows a decrease in the MDSI variance by around one order after two steps compared to variance after the first step. This is important for the MDSI of approximately 0.2–0.25 corresponding to the distortion invisibility threshold. The BPG performance comparison to some other coders is performed and examples of its application to real-life RS images are presented.
An expansion of the use of UAV images requires the improvement of methods and means for image analysis and processing in order to effectively solve various problems. Visual quality metrics play a key role in this sense, since their use allows determining the need in different image processing operations, their type and parameters, automating the entire process. The paper considers the problem of using no-reference visual quality metrics and test image databases to solve such problems. The effectiveness of more than 40 existing visual quality metrics for images with typical distortions for UAVs has been evaluated. The paper proposes a combined metric based on an artificial neural network to improve the accuracy of visual quality assessment and the possibility of its application in practice with sufficient efficiency.
Visual quality is important for remote sensing data presented as grayscale, color or pseudo-color images. Although several visual quality metrics (VQMs) have been used to characterize such data, only a limited analysis of their applicability in remote sensing applications has been done so far. In this paper, we study correlation factors for a wide set of VQMs for color images with distortion types typical for remote sensing. It is demonstrated that there are many metrics that have very high Spearman rank order correlation, e.g. PSNR-based and SSIM-based metrics. Meanwhile, there are also metrics that are practically uncorrelated with others. A detailed analysis of VQMs that have the largest SROCC values and belong to different groups is presented in this paper.
No-reference image quality assessment is one of the most demanding areas of image analysis for many applications where the results of the analysis should be strongly correlated with the quality of an input image and the corresponding reference image is unavailable. One of the examples might be remote sensing since the transmission of such obtained images often requires the use of lossy compression and they are often distorted, e.g., by the presence of noise and blur. Since the practical usefulness of acquired and/or preprocessed images is directly related to their quality, there is a need for the development of reliable and adequate no-reference metrics that do not need any reference images. As the performance and universality of many existing metrics are quite limited, one of the possible solutions is the design and application of combined metrics. Several possible approaches to their composition have been previously proposed and successfully used for full-reference metrics. In the paper, three possible approaches to the development and optimization of no-reference combined metrics are investigated and verified for the dataset of images containing distortions typical for remote sensing. The proposed approach leads to good results, significantly improving the correlation of the obtained results with subjective quality scores.
The paper describes a new image database HTID for verification and training of no-reference image visual quality metrics. The database contains 2880 color images of size 1536×1024 pixels cropped from the real-life photos produced by the mobile phone cameras with various shooting and post-processing settings. Mean opinion scores for images of the database are obtained. Peculiarities of the database are considered. A comparative analysis of the state-of-the-art no-reference image visual quality metrics is carried out. It is shown that the proposed database takes its own unique place in the existing image databases and can be effectively used for metrics' verification.
Traditional approach to collect mean opinion score (MOS) values for evaluation of full-reference image quality metrics has two serious drawbacks. The first drawback is a nonlinearity of MOS, only partially compensated by the use of rank order correlation coefficients in a further analysis. The second drawback are limitations on number of distortion types and distortion levels in image database imposed by a maximum allowed time to carry out an experiment. One of the largest of databases used for this purpose, TID2013, has almost reached these limitations, which makes an extension of TID2013 within the boundaries of this approach to be practically unfeasible. In this paper, a novel methodology to collect MOS values, with a possibility to infinitely increase a size of a database by adding new types of distortions, is proposed. For the proposed methodology, MOS values are collected for pairs of distortions, one of them being a signal dependent Gaussian noise. A technique of effective linearization and normalization of MOS is described. Extensive experiments for linearization of MOS values to extend TID2013 database are carried out.
Remote sensing images are subject to complex noise and distortions, which reduce their quality and can lead to loss of information. Visual quality metrics can be applied to automate and optimize remote sensing image processing. The problem of taking into account distortions typical for remote sensing data in existing test image databases is considered in this paper. The efficiency of 50 full-reference quality metrics for typical remote sensing distortions is estimated. The paper proposes a robust combined metric based on alpha-trimmed mean. The influence of linearization and the use of various correlation coefficients on the formation of a robust metric are studied. The effectiveness of the solution is tested using the cross-database approach based on the test image databases including TID2013, KADID10k, MDID and LIVE Multiply Distorted Image Quality Database.
Remote sensing images are subject to different types of degradations. The visual quality of such images is important because their visual inspection and analysis are still widely used in practice. To characterize the visual quality of remote sensing images, the use of specialized visual quality metrics is desired. Although the attempts to create such metrics are limited, there is a great number of visual quality metrics designed for other applications. Our idea is that some of these metrics can be employed in remote sensing under the condition that those metrics have been designed for the same distortion types. Thus, image databases that contain images with types of distortions that are of interest should be looked for. It has been checked what known visual quality metrics perform well for images with such degradations and an opportunity to design neural network-based combined metrics with improved performance has been studied. It is shown that for such combined metrics, their Spearman correlation coefficient with mean opinion score exceeds 0.97 for subsets of images in the Tampere Image Database (TID2013). Since different types of elementary metric pre-processing and neural network design have been considered, it has been demonstrated that it is enough to have two hidden layers and about twenty inputs. Examples of using known and designed visual quality metrics in remote sensing are presented.
The problem of increasing efficiency of blind image quality assessment is considered. No-reference image quality metrics both independently and as components of complex image processing systems are employed in various application areas where images are the main carriers of information. Meanwhile, existing noreference metrics have a significant drawback characterized by a low adequacy to image perception by human visual system (HVS). Many well-known no-reference metrics are analyzed in our paper for several image databases. A method of combining several noreference metrics based on artificial neural networks is proposed based on multi-database verification approach. The effectiveness of the proposed approach is confirmed by extensive experiments.
This paper is devoted to assessing effectiveness of verification of no-reference quality metrics. The analysis of the impact of image visual quality on accuracy of verification of quality metrics is carried out. The problem of estimating metrics adequacy for images with almost unnoticeable distortions is shown. Modifications of rank correlation coefficients are proposed to address this problem. The results are presented for several good particular quality metrics as well as the designed neural network based combined visual quality metric.
This paper deals with analysis of component-wise and three-dimensional (3D) lossy compression of hyperspectral data. The latter approach is known to be more efficient due to possibility to employ inter-band correlation of images in neighbor sub-bands. However, there are certain peculiarities for this approach which are subject of our study. Methods of lossy compression that employ discrete cosine transform (DCT) in blocks are studied. It is shown that 3D compression is able to provide considerable benefits in compression ratio and/or mean square error (MSE) of introduced distortions but only under certain conditions. Experiments have been done for real-life hyperspectral data with high input signal-to-noise ratio. Their main results are presented.
A task of time delay estimation using two or more sensors that receive and process noise-like wideband signals embedded in spatially uncorrelated noise can be treated as classical. A problem arises if environment noise is non-Gaussian and has heavy tails. Then, due to several factors, probability of appearing abnormal elementary time delay estimates can be large, especially if conventional method of cross-correlation signal processing is applied. Different methods to cope with the problem have been proposed including the use of robust DFT at initial stages. However, this reduces computational efficiency of processing radically. Here, we propose a specific modification of classical processing that presumes pre-processing of noisy received signals using center weighted median filter. We show that this easy and fast step allows providing better performance than procedures based on robust DFT. Considerable reduction of abnormal error probability is observed and the method can operate in conditions of limited a priori information about characteristics of additive non-Gaussian noise.
This paper studies the problem of full reference visual quality assessment of denoised images with a special emphasis on images with low contrast and noise-like texture. Denoising of such images together with noise removal often results in image details loss or smoothing. A new test image database, FLT, containing 75 noise-free 'reference' images and 300 filtered ('distorted') images is developed. Each reference image, corrupted by an additive white Gaussian noise, is denoised by the BM3D filter with four different values of threshold parameter (four levels of noise suppression). After carrying out a perceptual quality assessment of distorted images, the mean opinion scores (MOS) are obtained and compared with the values of known full reference quality metrics. As a result, the Spearman Rank Order Correlation Coefficient (SROCC) between PSNR values and MOS has a value close to zero, and SROCC between values of known full-reference image visual quality metrics and MOS does not exceed 0.82 (which is reached by a new visual quality metric proposed in this paper). The FLT dataset is more complex than earlier datasets used for assessment of visual quality for image denoising. Thus, it can be effectively used to design new image visual quality metrics for image denoising.