In recent years, live video streaming has gained widespread popularity across various social media platforms. Quality of experience (QoE), which reflects end-users' satisfaction and overall experience, plays a critical role for media service providers to optimize large-scale live compression and transmission strategies to achieve perceptually optimal rate-distortion trade-off. Although many QoE metrics for video-on-demand (VoD) have been proposed, there remain significant challenges in developing QoE metrics for live video streaming. To bridge this gap, we conduct a comprehensive study of subjective and objective QoE evaluations for live video streaming. For the subjective QoE study, we introduce the first live video streaming QoE dataset, TaoLive QoE, which consists of 42 source videos collected from real live broadcasts and 1,155 corresponding distorted ones degraded due to a variety of streaming distortions, including conventional streaming distortions such as compression, stalling, as well as live streaming-specific distortions like frame skipping, variable frame rate, etc. Subsequently, a human study was conducted to derive subjective QoE scores of videos in the TaoLive QoE dataset. For the objective QoE study, we benchmark existing QoE models on the TaoLive QoE dataset as well as publicly available QoE datasets for VoD scenarios, highlighting that current models struggle to accurately assess video QoE, particularly for live content. Hence, we propose an end-to-end QoE evaluation model, Tao-QoE, which integrates multi-scale semantic features and optical flow-based motion features to predicting a retrospective QoE score, eliminating reliance on statistical quality of service (QoS) features.
Point cloud is one of the most widely used digital representation formats for three-dimensional (3D) contents, the visual quality of which may suffer from noise and geometric shift distortions during the production procedure as well as compression and downsampling distortions during the transmission process. To tackle the challenge of point cloud quality assessment (PCQA), many PCQA methods have been proposed to evaluate the visual quality levels of point clouds by assessing the rendered static 2D projections. Although such projection-based PCQA methods achieve competitive performance with the assistance of mature image quality assessment (IQA) methods, they neglect that the 3D model is also perceived in a dynamic viewing manner, where the viewpoint is continually changed according to the feedback of the rendering device. Therefore, in this paper, we evaluate the point clouds from moving camera videos and explore the way of dealing with PCQA tasks via using video quality assessment (VQA) methods. First, we generate the captured videos by rotating the camera around the point clouds through several circular pathways. Then we extract both spatial and temporal quality-aware features from the selected key frames and the video clips through using trainable 2D-CNN and pre-trained 3D-CNN models respectively. Finally, the visual quality of point clouds is represented by the video quality values. The experimental results reveal that the proposed method is effective for predicting the visual quality levels of the point clouds and even competitive with full-reference (FR) PCQA methods. The ablation studies further verify the rationality of the proposed framework and confirm the contributions made by the quality-aware features extracted via the dynamic viewing manner. The code is available at https://github.com/zzc-1998/VQA_PC.
Face images play a crucial role in numerous applications; however, real-world conditions frequently introduce degradations such as noise, blur, and compression artifacts, affecting overall image quality and hindering subsequent tasks. To address this challenge, we organized the VQualA 2025 Challenge on Face Image Quality Assessment (FIQA) as part of the ICCV 2025 Workshops. Participants created lightweight and efficient models (limited to 0.5 GFLOPs and 5 million parameters) for the prediction of Mean Opinion Scores (MOS) on face images with arbitrary resolutions and realistic degradations. Submissions underwent comprehensive evaluations through correlation metrics on a dataset of in-the-wild face images. This challenge attracted 127 participants, with 1519 final submissions. This report summarizes the methodologies and findings for advancing the development of practical FIQA approaches.
Accurate preoperative prediction of pathological complete response (pCR) following neoadjuvant immunochemotherapy is crucial for refining and customising perioperative treatment decisions. However, the challenge persists in developing reliable, interpretable, and intelligent imaging markers using early-stage computed tomography (CT). Considering the dynamic evolution of tumours during treatment can offer additional discriminative insights, modelling the spatial and temporal relationships in pre- and post-neoadjuvant treatment CT images may help address these challenges. This study proposes an adaptive spatio-temporal aware graph inference framework (A2ST-GCM) to learn potential correlations between tumour regions in the same period and across periods, generating an adaptive spatio-temporal topology, and effectively aggregating features. Specifically, the model first utilises the adaptive spatial graph convolution module (AS-GConv) to capture irregular spatial dependencies within tumour regions before and after treatment. Subsequently, the adaptive spatio-temporal graph convolution module (AST-GConv) is employed to model dynamic dependencies over time, effectively exploring the complementarity and correlation mechanisms among multiple spatio-temporal features, ultimately obtaining a spatio-temporal graph representation containing tumour heterogeneity information. Furthermore, this paper introduces a novel weighted loss function, effectively alleviating the class imbalance issue in predicting pCR. Quantitative experimental results demonstrate the outstanding performance of our model in pCR prediction. To the best of our knowledge, this paper represents the first attempt to apply graph networks to predicting pathological response in neoadjuvant immunochemotherapy. It provides a potential auxiliary tool to evaluate pathological response, aiming to identify individuals who could benefit from surgery avoidance, thereby offering significant clinical benefits for patients undergoing personalised organ-preserving treatment.
Epidermal growth factor receptor (EGFR) is the key to targeted therapy with tyrosine kinase inhibitors in lung cancer. Traditional identification of EGFR mutation status requires biopsy and sequence testing, which may not be suitable for certain groups who cannot perform biopsy. In this paper, using easily accessible and non-invasive CT images, the residual neural network (ResNet) with mixed loss based on batch training technique is proposed for identification of EGFR mutation status in lung cancer. In this model, the ResNet is regarded as the baseline for feature extraction to avoid the gradient disappearance. Besides, a new mixed loss based on the batch similarity and the cross entropy is proposed to guide the network to better learn the model parameters. The proposed mixed loss utilizes the similarity among batch samples to evaluate the distribution of training data, which can reduce the similarity of different classes and the difference of the same classes. In the experiments, VGG16Net, DenseNet, ResNet18, ResNet34 and ResNet50 models with the mixed loss are trained on the public CT dataset with 155 patients including EGFR mutation status from TCIA. The trained networks are employed to the collected preoperative CT dataset with 56 patients from the cooperative hospital for validating the efficiency of the proposed models. Experimental results show that the proposed models are more appropriate and effective on the lung cancer dataset for identifying the EGFR mutation status. In these models, the ResNet34 with mixed loss is optimal (accuracy = 81.58%, AUC = 0.8861, sensitivity = 80.02%, specificity = 82.90%).
As a critical indicator of how easily the human immune system recognizes tumour cells, tumour mutational burden (TMB) is widely used to identify the potential effectiveness of immune checkpoint inhibitor therapy. However, the difficulties associated with the whole exome sequencing (WES) process, such as high tissue sampling requirements, high costs, and long turnaround times, have hindered the widespread clinical use of WES. Furthermore, the mutation landscape varies across cancer types, and the distribution of TMBs varies across cancer subtypes. Therefore, there is an urgent clinical need to develop a small cancer-specific panel to estimate TMB accurately, predict immunotherapy response cost-effectively and assist physicians in precise decisionmaking. This paper uses a graph neural network framework (Graph-ETMB) to address the cancer specificity problem in TMB. The correlation and tractability between mutated genes are described through message-passing and aggregation algorithms between graph networks. Then the graph neural network is trained in the lung adenocarcinoma data through a semi-supervised approach, resulting in a mutation panel containing 20 genes with a length of only 0.16 Mb. The number of genes to be detected is smaller than most commercial panels currently in clinical use. In addition, the efficacy of the designed panel in predicting immunotherapy response was further determined in an independent validation dataset, exploring the association between TMB and immunotherapy efficacy.
While video enhancement has drawn significant interest and has been extensively studied by academia and industry, the corresponding research on video quality assessment (VQA) for enhanced video has not been widely addressed. Video enhancement methods normally change the relevant metrics like brightness, contrast, color, etc., leading to the fluctuation of perceptual quality and challenging the related VQA task. In this paper, we propose a novel approach for VQA task based on Swin Transformer with improved spatio-temporal feature fusion, which precisely mines the stage-wise feature concatenation and provides competitive assessment performance. In addition, we propose an efficient data augmentation strategy to improve data diversity and further enhance assessment accuracy. Experimental results demonstrate that the proposed approach achieves state-of-the-art performance on two benchmark VQA datasets, and ranks first in CVPR NTIRE 2023 Quality Assessment for Video Enhancement Challenge, which proves that the proposed approach is not only promising in VQA for enhanced video but also ubiquitous in general VQA tasks.
This paper reports on the NTIRE 2023 Quality Assessment of Video Enhancement Challenge, which will be held in conjunction with the New Trends in Image Restoration and Enhancement Workshop (NTIRE) at CVPR 2023. This challenge is to address a major challenge in the field of video processing, namely, video quality assessment (VQA) for enhanced videos. The challenge uses the VQA Dataset for Perceptual Video Enhancement (VDPVE), which has a total of 1211 enhanced videos, including 600 videos with color, brightness, and contrast enhancements, 310 videos with deblurring, and 301 deshaked videos. The challenge has a total of 167 registered participants. 61 participating teams submitted their prediction results during the development phase, with a total of 3168 submissions. A total of 176 submissions were submitted by 37 participating teams during the final testing phase. Finally, 19 participating teams submitted their models and fact sheets, and detailed the methods they used. Some methods have achieved better results than baseline methods, and the winning methods have demonstrated superior prediction performance.
Stereo matching has attracted substantial research attention, as it plays an important role in many visual tasks such as autonomous driving. Despite the remarkable progress has been made by all kinds of stereo matching methods, they still suffer in some challenging scenes like disparity discontinuities. The root cause of blurred depth boundaries lies in information diffusion across edge boundaries. Aiming at this problem, this paper presents a novel edge-preserving stereo matching method by enforcing anisotropy, where information is allowed to propagate throughout the image except through large discontinuities. Specifically, we propose an intra-scale cost aggregation algorithm to smooth out noise in homogeneous regions while well preserving strong edges in the guided image. And weighted averaging scheme, where weights are calculated according to the variances of pixels in different overlapping support windows, is utilized for enhancing anisotropy to improve accuracy at disparity discontinuities. We carry out comprehensive experiments on Middlebury public datasets to demonstrate the accuracy and edge-preserving property of our method. Qualitative and quantitative performance evaluation on Middlebury data sets demonstrate the superior performance of our method for strong edge preservation with 20.77% decreases in wrong matching ratios in discontinuous regions.
Recent years have witnessed the rapid development of image storage and transmission systems, in which image compression plays an important role. Generally speaking, image compression algorithms are developed to ensure good visual quality at limited bit rates. However, due to the different compression optimization methods, the compressed images may have different levels of quality, which needs to be evaluated quantificationally. Nowadays, the mainstream full-reference (FR) metrics are effective to predict the quality of compressed images at coarse-grained levels (the bit rates differences of compressed images are obvious), however, they may perform poorly for fine-grained compressed images whose bit rates differences are quite subtle. Therefore, to better improve the Quality of Experience (QoE) and provide useful guidance for compression algorithms, we propose a full-reference image quality assessment (FR-IQA) method for compressed images of fine-grained levels. Specifically, the reference images and compressed images are first converted to $YCbCr$ color space. The gradient features are extracted from regions that are sensitive to compression artifacts. Then we employ the Log-Gabor transformation to further analyze the texture difference. Finally, the obtained features are fused into a quality score. The proposed method is validated on the fine-grained compression image quality assessment (FGIQA) database, which is especially constructed for assessing the quality of compressed images with close bit rates. The experimental results show that our metric outperforms mainstream FR-IQA metrics on the FGIQA database. We also test our method on other commonly used compression IQA databases and the results show that our method obtains competitive performance on the coarse-grained compression IQA databases as well.
User-generated content (UGC) live videos are often bothered by various distortions during capture procedures and thus exhibit diverse visual qualities. Such source videos are further compressed and transcoded by media server providers before being distributed to end-users. Because of the flourishing of UGC live videos, effective video quality assessment (VQA) tools are needed to monitor and perceptually optimize live streaming videos in the distributing process. In this paper, we address UGC Live VQA problems by constructing a first-of-a-kind subjective UGC Live VQA database and developing an effective evaluation tool. Concretely, 418 source UGC videos are collected in real live streaming scenarios and 3,762 compressed ones at different bit rates are generated for the subsequent subjective VQA experiments. Based on the built database, we develop a Multi-12imensional VQA (MD-VQA) evaluator to measure the visual quality of UGC live videos from semantic, distortion, and motion aspects respectively. Extensive experimental results show that MD-VQA achieves state-of-the-art performance on both our UGC Live VQA database and existing compressed UGC VQA databases.
PURPOSE:Identifying the stage of lung cancer accurately from histopathology images and gene is very important for the diagnosis and treatment of lung cancer. Despite the substantial progress achieved by existing methods, it remains challenging due to large intra-class variances, and a high degree of inter-class similarities.METHODS:In this paper, we propose a phased Multimodal Multi-scale Attention Model (MMAM) that predicts lung cancer stages using histopathology image data and gene data. The model consists of two phases. In Phase1, we propose a Staining Difference Elimination Network (SDEN) to eliminate staining differences between different histopathology images, In Phase2, it utilizes the image feature extractor provided by Phase1 to extract image features, and sends the multi-scale image features together with gene features into our Adaptive Enhanced Attention Fusion (AEAF) module for multimodal multi-scale features fusion to enable prediction of lung cancer staging.RESULTS:We evaluated the proposed MMAM on the TCGA lung cancer dataset, and achieved 88.51% AUC and 88.17% accuracy on classification prediction of lung cancer stages I, II, III, and IV.CONCLUSION:The method can help doctors diagnose the stage of lung cancer patients and can benefit from multimodal data.
Background and purpose: Accurate identification of lung cancer subtypes in medical images is of great significance for the diagnosis and treatment of lung cancer. Despite substantial progress in existing methods, they remain challenging due to limited annotated datasets, large intra-class differences, and high inter-class similarities. Methods: To address these challenges, we propose a Frequency Domain Transformer Model (FDTrans) to identify patients' lung cancer subtypes using the TCGA lung cancer dataset. We add a pre-processing process to transfer histopathological images to the frequency domain using a block-based discrete cosine transform and design a coordinate Coordinate-Spatial Attention Module (CSAM) to obtain critical detail information by reassigning weights to the location information and channel information of different frequency vectors. Then, a Cross-Domain Transformer Block (CDTB) is designed for Y, Cb, and Cr channel features, capturing the long-term dependencies and global contextual connections between different component features. At the same time, feature extraction is performed on the genomic data to obtain specific features. Finally, the image branch and the gene branch are fused, and the classification result is output through the fully connected layer. Results: In 10-fold cross-validation, the method achieves an AUC of 93.16% and overall accuracy of 92.33%, which is better than similar current lung cancer subtypes classification detection methods. Conclusion: This method can help physicians diagnose the subtypes classification of lung cancer in patients and can benefit from both spatial and frequency domain information.
Mobile games have played an increasingly significant role in people's leisure lives in recent years, thanks to the fast expansion of the gaming industry and the widespread use of mobile devices. The aesthetic quality of game pictures is a very important factor that attracts users' interest. However, evaluating the aesthetic quality of mobile game pictures is difficult since the painting styles of games vary greatly and the evaluation criteria are also diversified. In this article, we propose a multitask deep learning-based method, which is able to predict the aesthetic quality of mobile game images in multiple dimensions. The proposed model consists of two modules, a feature extraction module and a quality regression module. We extract quality-aware features from intermediate layers of the deep convolution neural network and then incorporate them into the final feature representation in the feature extraction module, allowing the model to fully use visual information from low to high levels. The quality regression module uses fully connected layers to map quality-aware features into quality scores across multiple dimensions. The multidimensional aesthetic quality scores are trained using a multitask learning approach, in which quality-aware features are shared across multiple dimensional quality prediction tasks. Finally, several key factors which help the proposed model perform better are analyzed. The experimental results indicate that our proposed method not only achieves the greatest performance on mobile game images, but also is applicable to natural scene images.
Recently, increasing interest has been drawn in Transformer-based models for No-reference Image Quality Assessment (NR-IQA), especially for the hybrid approach. The hybrid approach tend to apply Transformer to aggregate quality information from feature maps extracted by Convolutional Neural Networks (CNN). However, existing methods cannot fully utilize the information of hierarchical features extracted by the deep neural network, resulting in the limited performance of image quality evaluation. In this work, we propose a novel Hierarchical Feature Fusion Transformer for NR-IQA (HiFFTiq), which is able to effectively exploit complementary strengths of features extracted by different layers. Further, we propose a new Uniform Partition Pooling (UPP) which can reduce the resolution of input features via uniform partitions and can well retain the quality-related information compared to the traditional pooling method Sliding Window Pooling (SWP). The results of experiment demonstrate that HiFFTiq leads to improvements of performance over the state-of-the-art methods on three large scale NR-IQA datasets.
Wireless fidelity (WIFI) can transmit data efficiently, but it has the advantages of low computing cost and simple deployment. However, the WIFI signal is susceptible to interference from human movement, multipath propagation, and indoor temperature changes, resulting in problems such as low positioning accuracy and poor stability. To improve positioning accuracy and stability, we propose a novel fingerprint positioning technology applying vision-guided definition for WIFI-based localization, defined as KV. There are two stages in localization: WIFI-based coarse localization and fusion localization. In the WIFI-based coarse positioning stage, we propose a two-metric adaptive k-nearest neighbor (KNN) localization method to improve the accuracy and robustness of WIFI-based localization. Unlike the traditional KNN method, a difference-based search approach is introduced to obtain the K value intelligently and reasonably instead of manual determination. In the fusion localization stage, we employ an effective multiangle unsupervised fusion positioning method applying the fusion of WIFI and vision to enhance positioning accuracy and positioning stability further. Experimental results show that our proposed fusion localization method achieves 1.24 m in localization accuracy and the positioning error is 60% within 0.8 m. The performance of the proposed method surmounts other state-of-the-art methods, which proves its effectiveness. This method has important practical significance in multisource indoor positioning technology.
This paper presents a novel Adaptive Mutation Quantum-inspired Squirrel Search Algorithm (AM-QSSA). Firstly, based on the population mutation, a location-update of quantum state correlative and attractor method is proposed. By introducing a random process to modify the sliding factor of the local attractor, the inevitable lack of diversity in the population renewal method is solved. The premature convergence problem of adaptive mutation rate improvement algorithm based on squirrel position update mode is introduced. Meanwhile, the paper decomposes the location update process of the SSA, and improves it with quantum-behavior. Furthermore, it proposes a novel quantum-inspired squirrels search algorithm. This method finds the complementary effect between quantum behavior and squirrel search algorithm, and solves the problem of premature convergence probability of SSA. In addition, it improves population diversity, and achieves the balance between global and local search. The efficiency of the proposed AM-QSSA is evaluated using exploitation analysis, exploration analysis, success rate analysis, convergence rate analysis on classical benchmark functions as well as Congress on Evolutionary Computation (CEC) 2017 test functions. For further study, AM-QSSA optimizes an image registration problem for an extensive study to check its applicability. The results reveal that AM-QSSA is more efficient and stable than SSA. And it is comparable to the most advanced optimization algorithms.
Tracking the 6-degree-of-freedom (6D) object pose in video sequences is gaining attention because it has a wide application in multimedia and robotic manipulation. However, current methods often perform poorly in challenging scenes, such as incorrect initial pose, sudden re-orientation, and severe occlusion. In contrast, we present a robust 6D object pose tracking method with a novel hierarchical feature fusion network, refer it as HFF6D, which aims to predict the object’s relative pose between adjacent frames. Instead of extracting features from adjacent frames separately, HFF6D establishes sufficient spatial-temporal information interaction between adjacent frames. In addition, we propose a novel subtraction feature fusion (SFF) module with attention mechanism to leverage feature subtraction during feature fusion. It explicitly highlights the feature differences between adjacent frames, thus improving the robustness of relative pose estimation in challenging scenes. Besides, we leverage data augmentation technology to make HFF6D be used more effectively in the real world by training only with synthetic data, thereby reducing manual effort in data annotation. We evaluate HFF6D on the well-known YCB-Video and YCBInEOAT datasets. Quantitative and qualitative results demonstrate that HFF6D outperforms state-of-the-art (SOTA) methods in both accuracy and efficiency. Moreover, it is also proved to achieve high-robustness tracking under the above-mentioned challenging scenes.
A reliable video quality assessment (VQA) algorithm is essential for evaluating and optimizing video processing pipelines. In this paper, we propose a quality aggregation network (QAN) for full-reference VQA, which models the characteristics of human visual perception of video quality in both spatial and temporal domain. The proposed QAN is composed of two mod-ules, the spatial quality aggregation (SQA) network and the tem-poral quality aggregation (TQA) network. Specifically, the SQA network models the quality of video frames using 3D CNN, taking both spatial and temporal masking effects into consideration for the modeling of the perception of human visual system (HVS). In the TQA network, considering the memory effect of HVS facing the temporal variation of frame-level quality, an LSTM-based temporal quality pooling network is proposed to capture the nonlinearities and temporal dependencies involved in the process of quality evaluation. According to the experimental results on two well-established VQA databases, the proposed model could outperform the state-of-the-art metrics. The code of the proposed method is available at: https://github.com/lorenzowu/QAN.
In this paper, the defects and deficiencies of the recently proposed whale optimization algorithm (WOA) are improved. A whale optimization algorithm mixed with an artificial bee colony (ACWOA) is proposed to solve the WOA problems of slow convergence, low precision, and easy to fall into local optimum. The ACWOA algorithm integrates the artificial bee colony algorithm and chaotic mapping, effectively avoiding the local optimal situation and improving the quality of the initial solution. Also, nonlinear convergence factors and adaptive inertia weight coefficients are added to accelerate the convergence rate. To verify the performance of the improved algorithm, 20 benchmark functions and CEC2019 multimodal multi-objective benchmark functions have been used to compare ACWOA with the classical intelligent population algorithms (PSO, MVO, and GWO) and the recent state-of-the-art algorithms (CWOA, HWPSO, and HIWOA) in recent years. The proposed algorithm is applied to two well-known engineering mathematical models and a real application (the quality process control). The experiments show that the ACWOA algorithm has strong competitiveness in convergence speed and solution accuracy and has certain practical value in complex mathematical model scenarios.