Traditional subjective quality assessment protocols are designed to produce global quality scores. However, in some use cases such as watermarking, there is a need for local subjective assessment since the traditional methods lack the granularity required to generate precise local quality maps. In this paper, we propose LISA (Local Impairment Scale Annotator), a new subjective protocol and supporting annotation tool that enable pixel-wise local quality assessment. We present the complete LISA protocol specifications, from the temporal organization of the test to the user interface design. Finally, we illustrate LISA's deployment in a use case scenario evaluating highfidelity digital watermarking. The annotation tool is available at https://github.com/edemezet-nagra/LISA-subjective-protocol
The application of methods based on Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3D GS) have steadily gained popularity in the field of 3D object segmentation in static scenes. These approaches demonstrate efficacy in a range of 3D scene understanding and editing tasks. Nevertheless, the 4D object segmentation of dynamic scenes remains an underexplored field due to the absence of a sufficiently extensive and accurately labelled multi-view video dataset. In this paper, we present MUVOD, a new multi-view video dataset for training and evaluating object segmentation in reconstructed real-world scenarios. The 17 selected scenes, describing various indoor or outdoor activities, are collected from different sources of datasets originating from various types of camera rigs. Each scene contains a minimum of 9 views and a maximum of 46 views. We provide 7830 RGB images (30 frames per video) with their corresponding segmentation mask in 4D motion, meaning that any object of interest in the scene could be tracked across temporal frames of a given view or across different views belonging to the same camera rig. This dataset, which contains 459 instances of 73 categories, is intended as a basic benchmark for the evaluation of multi-view video segmentation methods. We also present an evaluation metric and a baseline segmentation approach to encourage and evaluate progress in this evolving field. Additionally, we propose a new benchmark for 3D object segmentation task with a subset of annotated multi-view images selected from our MUVOD dataset. This subset contains 50 objects of different conditions in different scenarios, providing a more comprehensive analysis of state-of-the-art 3D object segmentation methods. Our proposed MUVOD dataset is available at https://volumetric-repository.labs.b-com.com/#/muvod.
Autism spectrum is very wide, but autistic people share some traits such as social or language impairments. Another common characteristic is a strong passion, called an affinity, which can be anything from a movie character to a topic like history, or even a specific object. Numerous testimonies attest to the support provided by affinities for autistic individuals, offering a reassuring space in a sometimes frightening world. Clinical psychologists consider affinities as keys that can unlock language and learning for people with ASD. However, objective evidence of this role is still lacking. In this study, we used eye-tracking technology to explore the visual attention patterns of autistic individuals when presented with their affinity compared to neutral stimuli. We recruited 52 autistic participants and showed them 38 images: 10 featuring their affinity and 28 neutral ones, while recording their eye movements. Eye-tracking data provided crucial insights into how visual attention is modulated by affinities. Our results reveal significant variability in visual engagement depending on the specific autistic traits and affinities of the participants. Some showed heightened visual engagement with affinity images, while others withdrew their gaze. Some exhibited a mixed response, with both increased engagement and gaze withdrawal, and a few showed no difference between the two sets of images. These findings highlight the complex relationship between visual attention and affinities in autistic individuals, highlighting the potential of eye-tracking as a tool for understanding and leveraging these affinities in therapeutic and educational settings.
The continuous breakthroughs in 3D reconstruction and rendering technologies have enabled synthesized 3D models to achieve exceptional realism. In particular, Neural Radiance Fields (NeRF) and 3D Gaussian Splattings (3DGS) have gained significant attention due to their impressive ability to deliver high-quality 3D models. However, research on the quality assessment of NeRF/3DGS models remains an underexplored area, hindering the development of relevant generation, compression, and transmission algorithms. To fill this gap, in this paper, we propose the first no-reference quality assessment metric for NeRF/3DGS models rendered in Processed Video Sequences (PVS). Considering the uniqueness and diversity of distortions introduced in NeRF/3DGS models, the core idea behind the proposed metric is to extract universal quality-aware features that are generalizable across various distortions, rather than targeting a specific one. Specifically, inspired by the fact that spatial distortions typically alter the statistical distributions, we first measure the spatial fidelity of the rendered PVS by analyzing the spatial explicit statistics of textural variation and naturalness. Then, motivated by the ability of implicit energy composition changes to reflect temporal distortions, we propose to evaluate the temporal consistency of the rendered PVS through inter-frame discrepancy energy and multi-frame motion energy in the Singular Value Decomposition (SVD) domain. Finally, the spatial explicit statistics and temporal implicit energy are combined as perceptual features to evaluate the quality of NeRF/3DGS models via Support Vector Regression (SVR). Extensive experimental results on three representative databases demonstrate the superiority of the proposed metric in various aspects, such as predictive accuracy, performance stability, and cross-database generalizability. The source code will be publicly available at https://github.com/ZhengyuZhang96/PVS-3DMQA.
Gaussian Splatting (GS) and Neural Radiance Fields (NeRF) are two groundbreaking technologies that have revolutionized the field of Novel View Synthesis (NVS), enabling immersive photorealistic rendering and user experiences by synthesizing multiple viewpoints from a set of images of sparse views. The potential applications of NVS, such as high-quality virtual and augmented reality, detailed 3D modeling, and realistic medical organ imaging, underscore the importance of quality assessment of NVS methods from the perspective of human perception. Although some previous studies have explored subjective quality assessments for NVS technology, they still face several challenges, especially in NVS methods selection, scenario coverage, and evaluation methodology. To address these challenges, we conducted two subjective experiments for the quality assessment of NVS technologies containing both GS-based and NeRF-based methods, focusing on dynamic and real-world scenes. This study covers concentric, front-facing, and single-view videos while providing a richer and greater number of real scenes. Meanwhile, it is the first time to explore the quality of content produced by NVS methods in dynamic scenes with moving objects. The two types of subjective experiments help to fully comprehend the influences of different visual paths from a human perception perspective and pave the way for future development of full-reference and no-reference quality metrics. In addition, we established a comprehensive benchmark of various state-of-the-art objective metrics on the proposed database, highlighting that existing methods still struggle to accurately capture subjective quality. The results give us some insights into the limitations of existing NVS methods and may promote the development of new NVS methods.
Low-dose computed tomography(CT)imaging reduces radiation dose,but the associated artifacts and noise can compromise diagnostic accuracy.This study proposes a method for denoising low-dose CT images,based on a convolutional neural network and the dual-tree complex wavelet transform(LDTNet).The method exploits the robust information extraction capabilities of convolutional neural networks in the spatial domain and integrates the multi-scale decomposition characteristics of the dual-tree complex wavelet transform in the frequency domain to mitigate information loss typically encountered in conventional single-domain denoising approaches.In addition,the method features a large receptive field,thereby facilitating the capture of subtle structures and edge information.Consequently,the denoising process is further refined.Validation on the AAPM dataset demonstrates that LDTNet achieves significant improvements in both peak signal-to-noise ratio and structural similarity metrics,while visual assessments further confirm its superior performance in noise suppression and image detail restoration.
The growing usage of video conferencing and online streaming applications is raising research interest in real-time compression solutions for visual face content. Meanwhile, the performance brought by deep generative network-based compression methods is beginning to get attention from standardization organizations such as Moving Picture Expert Group (MPEG). This paper focuses on generative methods in use for efficient compression of face videos. We analyze the bitrate contribution corresponding to different coded entities in the purpose of optimizing coding efficiency for a recently proposed Generative Face Video Coding (GFVC) scheme. Based on this study, we propose a simple but efficient enhancement of the existing hybrid face video compression, leading to an average of $10-20{\%}$ Bjøntegaard Delta-rate (BD-Rate) gain over the reference Hybrid Deep Animation Codec (HDAC) method, without further training of the initial model.
Light Field Image (LFI) has garnered remarkable interest and fascination due to its burgeoning significance in immersive applications. Although the abundant information in LFIs enables a more immersive experience, it also poses a greater challenge for Light Field Image Quality Assessment (LFIQA), especially when reference information is inaccessible. In this paper, inspired by the holistic visual perception of high-dimensional LFIs and neuroscience studies on the Human Visual System (HVS), we propose a novel Blind Light Field image quality assessment metric by exploring MultiPlane Texture and Multilevel Wavelet Information, abbreviated as MPT-MWI-BLiF. Specifically, considering the texture sensitivity of the secondary visual cortex (V2), we first convert LFIs into multiple individual planes and capture textural variations from these planes. Then, the statistical histogram of textural variations for all planes is calculated as holistic textural variation features. In addition, motivated by the fact that neuronal responses in the visual cortex are frequency-dependent, we simulate this visual perception process by decomposing LFIs into multilevel wavelet subbands with Four-Dimensional Discrete Haar Wavelet Transform (4D-DHWT). After that, the subband geometric features of first-level 4D-DHWT subbands and the coefficient intensity features of second-level 4D-DHWT subbands are computed respectively. Finally, we combine all the extracted quality-aware features and employ the widely-used Support Vector Regression (SVR) to predict the perceptual quality of LFIs. To fully validate the effectiveness of the proposed metric, we perform extensive experiments on five representative LFIQA databases with two cross-validation methods. Experimental results demonstrate the superiority of the proposed metric in quality evaluation, as well as its low time complexity compared to other state-of-the-art metrics. The full code will be publicly available at https://github.com/ZhengyuZhang96/MPT-MWI-BLiF
Image enhancement algorithms are essential for improving visual quality but often introduce new distortions, highlighting the need for reliable image quality assessment (IQA). However, existing IQA methods typically focus on semantic information or distortion-prone regions while ignoring their interactions, resulting in unsatisfactory performance. To address this issue, we propose to integrate semantic information with edge residual learning and design a semantic-guided residual learning IQA framework tailored for enhanced images across diverse scenarios. Specifically, the proposed framework utilizes a covariance-guided encoder to extract semantic information, which is then enhanced using a semantic refinement module. The refined semantic information is subsequently utilized to guide edge residual feature learning in the decoder. Extensive experiments on multiple tasks such as deraining, dehazing, and low-light enhancement demonstrate that our method outperforms state-of-the-art approaches.
Adrenal lesions are common incidental findings in clinical practice, which are mostly benign and harmless; adenomas are the most common benign adrenal tumors, representing more than 75% adrenal lesions. Medical information provided by CT scans such as lesions' dimension, attenuation values, etc., are crucial for the diagnosis of adenoma. Measurements of percentage washout of injected contrast material from contrast-enhanced CT provide reproducible means to distinguish adenomas from malignant masses. Despite of the 3D volume CT data, only selected 2D slices are used in this diagnosis process, which introduces uncertainty of clinical decision and requires high expertise of medical professionals. To alleviate this problem and to facilitate the diagnosis, we proposed an region-growing based 3-Slices washout calculation method, as a preliminary study of our further work of automatic 3D adrenal lesion characterisation. Comparing with the expert's diagnosis, our method showed a significant (more than 10%) improvement on the accuracy, revealing that computer-based 3D lesion characterisation could become a promising and reliable tool for the diagnosis of adrenal lesions.
We present a multi-view camera and spatialized audio microphone capture system designed for computer vision applications in free navigation immersive experiences. We propose a dataset of two long and complex in-situ training situations in the medical field. The scenarios in the dataset feature precise gestures for the learner to reproduce during complex situations with multiple simultaneous visual and auditory cues important for training. 3D computer vision techniques are used to reconstruct a 4D scene model from a set of videos to render novel views from unseen viewpoints. However, the quality of the rendered objects is directly dependent on the density of coverage by reference views. To ensure maximum Quality of Experience, we propose a dual rig of cameras, a central rig that captures the details of the gesture zone of the training scenarios and a peripheral rig that captures the environment of the room and the interactions occurring around the gesture zone. The central rig provides dense coverage of the central content, facilitating high-quality reconstruction on novel views of the captured gestures. Recordings include audio interactions of multiple actors, captured by Ambisonic microphones spatially distributed around the scene. The captured scenes are real-world educational content for medical courses, so this dataset provides a rare opportunity to assess the Quality of Experience of volumetric video techniques on realistic content, and to compare their pedagogical capabilities with standard multi-view video content.
Evaluating the quality of medical images is a critical step in developing precise, trustworthy, and clinically applicable image processing algorithms and machine learning models for medical imaging applications. This evaluation serves to both benchmark and optimize algorithms. However, there is a lack of standard or guidelines for conducting subjective Medical Image and Video Quality Assessment (MIVQA) tests. Although there are existing standards for natural image and video quality assessment that provide information on selecting images and video sequences, assessor types and numbers, subjective testing procedures (test environment, participant selection, methodology, etc.) and model performance evaluation, they are not tailored to medical images. Although this study does not aim to propose a comprehensive method for MIVQA subjective tests, it addresses several aspects that may be worth considering when a MIVQA subjective test is conducted.
By recording scenes from multiple viewpoints, Light Field Image (LFI) encompasses both angular and spatial information, thereby offering users a more immersive experience. Since LFIs may be distorted at various stages from acquisition to visualization, Light Field Image Quality Assessment (LFIQA) is of vitally important to monitor the potential impairments of LFI quality. However, existing objective LFIQA metrics fail to establish a reasonable correlation between spatial and angular information in LFIs, especially ignoring the imbalance problem of large spatial variations and subtle angular variations, which results in unsatisfactory quality evaluation performance. To alleviate this imbalance, in this paper, we propose a novel Blind LFIQA metric based on Angular-Spatial Effect Modeling, abbreviated as ASEM-BLiF. Specifically, the proposed metric consists of two branches. In the principal branch, we first present an Angular Effect Modeling (AEM) module to capture the angular information independently of spatial information. Based on AEM, we further design an Angular-Spatial Quality Learning (ASQL) module to model the local angular-spatial effect and establish the global relationship between different local regions for quality assessment via Transformer. In the auxiliary branch, a Discriminative Region Selection (DRS) module is proposed for auxiliary learning to improve the learning efficiency and prediction accuracy from a local perspective. Moreover, we present a Dynamic Weighting Loss (DWLoss) to achieve an optimal balance between principal and auxiliary learning throughout training. To demonstrate the effectiveness of the proposed metric, extensive experiments are conducted on five publicly available LFIQA databases with a variety of metrics. The experimental results show that compared to our previous work DeeBLiF, the current state-of-the-art LFIQA metric, our proposed ASEM-BLiF metric achieves 5.67%, 7.75%, 5.96%, 4.44%, and 0.33% SROCC performance improvements in quality assessment on the Win5-LID, NBU-LF1.0, LFDD, VALID10bit, and SHU databases, respectively. The code will be publicly available.
Quality assessment is a key element for the evaluation of hardware and software involved in image and video acquisition, processing, and visualization. In the medical field, user-based quality assessment is still considered more reliable than objective methods, which allow the implementation of automated and more efficient solutions. Regardless of increasing research in this topic in the last decade, defining quality standards for medical content remains a non-trivial task, as the focus should be on the diagnostic value assessed from expert viewers rather than the perceived quality from na\"{i}ve viewers, and objective quality metrics should aim at estimating the first rather than the latter. In this paper, we present a survey of methodologies used for the objective quality assessment of medical images and videos, dividing them into visual quality-based and task-based approaches. Visual quality based methods compute a quality index directly from visual attributes, while task-based methods, being increasingly explored, measure the impact of quality impairments on the performance of a specific task. A discussion on the limitations of state-of-the-art research on this topic is also provided, along with future challenges to be addressed.
Prior point cloud provides 3D environmental context, which enhances the capabilities of monocular camera in downstream vision tasks, such as 3D object detection, via data fusion. However, the absence of accurate and automated registration methods for estimating camera extrinsic parameters in roadside scene point clouds notably constrains the potential applications of roadside cameras. This paper proposes a novel approach for the automatic registration between prior point clouds and images from roadside scenes. The main idea involves rendering photorealistic grayscale views taken at specific perspectives from the prior point cloud with the help of their features like RGB or intensity values. These generated views can reduce the modality differences between images and prior point clouds, thereby improve the robustness and accuracy of the registration results. Particularly, we specify an efficient algorithm, named neighbor rendering, for the rendering process. Then we introduce a method for automatically estimating the initial guess using only rough guesses of camera's position. At last, we propose a procedure for iteratively refining the extrinsic parameters by minimizing the reprojection error for line features extracted from both generated and camera images using Segment Anything Model (SAM). We assess our method using a self-collected dataset, comprising eight cameras strategically positioned throughout the university campus. Experiments demonstrate our method's capability to automatically align prior point cloud with roadside camera image, achieving a rotation accuracy of 0.202 degrees and a translation precision of 0.079m. Furthermore, we validate our approach's effectiveness in visual applications by substantially improving monocular 3D object detection performance.
Numerous roadside perception datasets have been introduced to propel advancements in autonomous driving and intelligent transportation systems research and development. However, it has been observed that the majority of their concentrates is on urban arterial roads, inadvertently overlooking residential areas such as parks and campuses that exhibit entirely distinct characteristics. In light of this gap, we propose CORP, which stands as the first public benchmark dataset tailored for multi-modal roadside perception tasks under campus scenarios. Collected in a university campus, CORP consists of over 205k images plus 102k point clouds captured from 18 cameras and 9 LiDAR sensors. These sensors with different configurations are mounted on roadside utility poles to provide diverse viewpoints within the campus region. The annotations of CORP encompass multi-dimensional information beyond 2D and 3D bounding boxes, providing extra support for 3D seamless tracking and instance segmentation with unique IDs and pixel masks for identifying targets, to enhance the understanding of objects and their behaviors distributed across the campus premises. Unlike other roadside datasets about urban traffic, CORP extends the spectrum to highlight the challenges for multi-modal perception in campuses and other residential areas.
Abstract Video Coding for Machines (VCM) is gaining momentum in applications like autonomous driving, industry manufacturing, and surveillance, where the robustness of machine learning algorithms against coding artifacts is one of the key success factors. This work complements the MPEG/JVET standardization efforts in improving the resilience of deep neural network (DNN)-based machine models against such coding artifacts by proposing the following three advanced fine-tuning procedures for their training: (1) the progressive increase of the distortion strength as the training proceeds; (2) the incorporation of a regularization term in the original loss function to minimize the distance between predictions on compressed and original content; and (3) a joint training procedure that combines the proposed two approaches. These proposals were evaluated against a conventional fine-tuning anchor on two different machine tasks and datasets: image classification on ImageNet and semantic segmentation on Cityscapes. Our joint training procedure is shown to reduce the training time in both cases and still obtain a 2.4% coding gain in image classification and 7.4% in semantic segmentation, whereas a slight increase in training time can bring up to 9.4% better coding efficiency for the segmentation. All these coding gains are obtained without any additional inference or encoding time. As these advanced fine-tuning procedures are standard-compliant, they offer the potential to have a significant impact on visual coding for machine applications.
People with autism spectrum disorder usually exhibit heterogeneous gaze patterns. Universal saliency prediction, which generates salient regions based on high fixations across all observers, is limited to analyzing the visual attention of autism spectrum disorders. To solve the problem, we propose a learning-based method named PSMANet to predict the personalized saliency map based on personal information. Collecting personal information and collecting large-scale datasets are challenging tasks for people with autism spectrum disorders, since they often suffer from deficits in social communication and interaction. The proposed approach introduces the image-similarity-measure based embedding to extract personal information and transfers the saliency distribution knowledge from universal saliency prediction to personalized saliency prediction. For evaluating our network, two popular metrics, Normalized Scanpath Salience (NSS) and Area Under Curve (AUC), are used. The experimental results show that it achieves good performance on the databases of people with autism spectrum disorder.
Novel view synthesis has recently been approached with Neural radiance fields, for high quality rendering. While those methods initially addressed the whole scene with a single global function, many state-of-the-art methods decompose the scene by spatially encoding feature vectors. High speed and quality are achieved, but reconstruction stability is still fragile. This usually requires human supervision through data pre-processing, preventing a robust, fully automatic chain for volumetric reconstruction of the visual scene. We observe that stability and quality can be improved by interpreting the scene beforehand, in order to properly set its bounding box. We propose a simple and robust approach to fit the scene volume bounds based on sparse point clouds processed by a Structure From Motion (SfM) pipeline. The benefit of this method is shown over multiple scenes and two state-of-the-art radiance field reconstruction methods.
Autism spectrum disorders affect the way people perceive their environment and interact with it. Many autistic people have a passion for an object or a topic, such as a film, planes, or geography maps to name very few of them, which is called an affinity. This affinity is sometimes described as an obsession that prevents the ASD subjects to connect with the surrounding world, but it is also considered as a key to the autistic world and a way to make a connection. In this paper we investigate the specific role of affinity in the autistic person’s attention. We have conducted eye tracking experiments over 44 autistic subjects from 3 different institutions. We have shown them neutral images and images with their own affinity and recorded their gaze position. Results are not conclusive in all the 3 institutions, but in the 2 first ones we got significant differences between the 2 sets of images indicating a higher visual attention for the affinity.