
We explore 3D mesh saliency detection as a nonlinear regression problem, leveraging geometric features derived from spiral patches. The saliency prediction involves a multi-scale analysis, utilizing roughness, geometric, and spectral saliencies, which are calculated through spiral patches at three levels of mesh decimation. These features are processed using a multi-layer perceptron to predict vertex-level saliency. Evaluation on the Schelling dataset demonstrates the approach’s efficacy, achieving competitive results compared to state-of-the-art methods, particularly in saliency prediction and extraction of keypoints. The method emphasizes the integration of advanced geometric descriptors with machine learning for enhanced 3D content analysis.
The widespread use of 3D objects in many fields poses new challenges for their secure transmission, storage and visualization. While selective encryption methods allow format compliance and enable adjustable security by selectively encrypting 3D geometry, they produce visible noise and disrupt geometric coherence which reduces compression efficiency. In this paper, we propose a reversible deformation method for 3D object protection that breaks with state-of-the-art 3D security approaches by offering more natural and visually coherent protection. The deformation is applied directly to the geometry, controlled by a secret key, and can be adjusted to achieve the desired level of visual security. Experimental results presented on a large database demonstrate that our method offers protection comparable to selective encryption, while providing significantly greater resilience to reconstruction attacks such as Laplacian smoothing, particularly at low security levels where selective encryption is more vulnerable. This work introduces a new direction for format-compliant 3D object protection, designed for secure visualization in untrusted environments.
The total variation models are popular in various image processing tasks such as smoothing, decomposition and depth estimation. However, solving these models are challenging. They can be solved by traditional iterative algorithms that require a large number of iterations to converge or deep neural networks that have a large number of trainable parameters. In this paper, we propose a novel Anderson accelerated residual solver (AARS) for these models. Although our method is iterative, it requires much less iteration numbers than the traditional iterative methods, thanks to the Anderson acceleration. Mean-while, it can theoretically guarantee to converge to the global optimal solution. This is theoretically proved and numerically confirmed. Several numerical experiments are conducted to show the effectiveness and efficiency of the proposed solver. It can be applied in various image processing tasks where solving the total variation models is necessary.
Egocentric continual action recognition faces severe challenges such as sudden viewpoint changes, occlusions, and complex backgrounds. In such scenarios, relying solely on visual modalities is susceptible to interference and lacks sufficient recognition robustness. To overcome the limitations of unimodal approaches, multimodal fusion methods are widely adopted, significantly enhancing recognition performance. However, existing multimodal schemes generally suffer from insufficient exploration of cross-modal complementarity and the vulnerability of modal independence. To address this, this paper proposes a Dual-path Decoupling-Distillation NetWork (D3Net), aiming to achieve more effective dynamic fusion of modal information and knowledge transfer.D3Net first explicitly separates the shared and private features of modalities through a dual-path decoupling module, combined with a dynamic gating mechanism to adaptively adjust the modal fusion weights. Secondly, it designs a complementary distillation module, leveraging cross-modal contrastive learning to effectively mitigate the issues of poor unimodal robustness and vulnerability to interference. Finally, through a cross-task distillation mechanism, it efficiently extracts knowledge from old tasks, alleviating the catastrophic forgetting problem during learning. Experimental results demonstrate that D3Net achieves an average accuracy of 83.97% under the 8×4 task configuration on the UESTC MMEA CL dataset, surpassing baseline method by 5.17%.
Anomaly detection in industrial manufacturing is vital for ensuring product quality and operational efficiency. However, supervised deep learning approaches often struggle due to the scarcity of defective samples and class imbalance in real-world datasets. In this work, we propose a generative defect synthesis framework that aims to enhance industrial anomaly detection by producing realistic and diverse defective samples. Our approach leverages generative models to synthesize high-fidelity anomalies while preserving the underlying texture and structural patterns of normal samples. We evaluate the proposed method on the MVTec AD dataset, a benchmark for unsupervised industrial anomaly detection, and investigate how the realism of synthetic data affects detection performance. Experimental results demonstrate that augmenting training with generated defects significantly improves model robustness, particularly in low-data regimes. Furthermore, we explore model compression techniques, including quantization and pruning, showing that 8-bit quantization and moderate pruning yield a 3-4x reduction in model size with minimal performance degradation. A detailed case study across various object and texture categories highlights substantial gains in both image-level and pixel-level AUC scores, and explores the optimal trade-off between synthetic and real training data for efficient deployment.
Few-Shot Class-Incremental Learning (FSCIL) requires models to progressively learn novel classes with limited samples while mitigating catastrophic forgetting of base classes. Existing methods face dual challenges: novel classes are prone to misclassification into base classes because the strong discriminability of base classes distracts the classification of novel classes, and the feature space lacks sufficient generalization capacity. This paper proposes the OrthCal framework built on the deep integration of orthogonal contrastive learning and a prototype calibration strategy to improve the performance during incremental sessions. Our three-stage optimization includes: 1) pretraining with hybrid supervised and self-supervised contrastive learning to construct geometrically constrained orthogonal pseudo-targets. 2) Dynamic prototype calibration, using semantic similarity among base classes to adjust novel class prototypes without additional training. 3) Hybrid loss design optimizing orthogonality constraints, perturbation-sensitive contrastive loss, and calibrated prototypes jointly to address challenges arising from data limitations during incremental sessions. Experiments on miniImageNet and CIFAR100 demonstrate that OrthCal achieves state-of-the-art performance. The framework provides a unified solution for feature space optimization and prototype calibration in FSCIL.
3D scene modeling is essential for immersive experiences in Virtual, Augmented, and Mixed Reality (VR/AR/MR) applications. Neural Radiance Fields (NeRF) have emerged as a strong alternative to traditional representations such as meshes and point clouds for 6-DoF rendering. However, maintaining high visual quality while enabling efficient transmission in dynamic environments remains a significant challenge. In this paper, we propose NeRFCompressor, a novel compression framework for dynamic scene representation using NeRF-like models. Building on tensor decomposition-based 3D reconstruction, NeRFCompressor improves transmission efficiency by leveraging existing video codecs to exploit both intra-scene and inter-scene redundancies. It maintains high QoE with minimal degradation in reconstruction quality. Experiments show that NeRFCompressor outperforms state-of-the-art methods in compressing both static and dynamic scene representations.
In large-scale camera systems, the automatic operations of Pan-Tilt-Zoom (PTZ) cameras introduce two new video quality problems that are not present in fixed cameras. First, during zoom operations, the improper adjustment of the focal length often causes sustained blurred frames. Second, command losses in the system can lead to uncontrolled movement of the cameras. This paper proposes PTZ-SRDD (PTZ-targeted Static Rule-based Distortions Detector), a real-time framework for detecting these distortion events for PTZ cameras. A dedicated video dataset is collected to ensure comprehensive coverage of relevant scenarios. PTZ-SRDD integrates frame-type-based downsampling to reduce computational overhead, frame-level blur detection based on gradient features, and sliding window analysis to identify long-span events. Experimental results show that PTZ-SRDD achieves real-time detection of both problems with high accuracy, making it a practical solution for large-scale PTZ camera systems.
3D Gaussian splatting representations have demonstrated outstanding potential for dense scene reconstruction and simultaneous localization and mapping (SLAM). However, existing multi-view filtering techniques frequently overlook critical edge information, resulting in blurred reconstruction boundaries. To address these challenges, we propose High-fidelity Gaussian SLAM based on Optical Flow Assisted Tracking (HGS_OFAT). Our approach analyzes the relationships between edge cues and pixels selected by multi-view pixel selection strategy, and employs a prudent initialization strategy to optimize Gaussian parameters—thereby reducing redundancy and improving reconstruction fidelity. Furthermore, HGS_OFAT integrates an opticalflow tracker before Gaussian-based reconstruction to obtain accurate camera poses, significantly enhancing both localization and reconstruction accuracy. Experiments on synthetic and realworld datasets demonstrate that HGS_OFAT outperforms existing methods in tracking robustness and reconstruction quality.
Weakly supervised temporal action localization (WTAL) targets the joint classification of action categories and precise delineation of their temporal boundaries in untrimmed videos while relying only on video-level labels. The absence of frame-level supervision inevitably causes two key difficulties: (i) incomplete localization of action segments and (ii) confusion between foreground and background frames. To overcome these challenges, we propose the Consensus-Guided Selective Multimodal Fusion Network (CG-SMFNet). First, a Selective Fusion Module (SFM) exploits the complementarity of multimodal cues to distill rich semantic representations. Second, a Consensus Attention Mechanism (CAM) dynamically assigns fusion weights to the three modality branches and enables bidirectional information exchange, ensuring a more holistic capture of action content. Finally, a Discrepant Expansion Mechanism (DEM) introduces a semantic contrast loss that enlarges the distance between foreground segments and semantically similar background regions, further sharpening localization accuracy. Extensive experiments on public benchmarks verify that CG-SMFNet achieves state-of-the-art performance under weak supervision.
Immersive visual communication has many important applications and Gaussian Splatting (GS) is a recent breakthrough that uses learnable geometry and color representation to capture 3D world with a very efficient parallelizable rendering pipeline. However, the compression of GS data still lacks efficiency and is quite complex in computation which prevents its adoption and deployment as a streaming solution in the real world. In this work, we develop a lightweight scalable GS coding scheme that exploits the correlation between adjacent quality layers and come up with a lightweight novel prediction and residual coding scheme that creates layered representation and is friendly to the MPEG DASH-like receiver-driven scalable streaming solutions. Simulation results demonstrate the efficiency of the proposed compression solution, as well as low latency/complexity in the decoding and rendering process. To the best of our knowledge, this is the first high-efficiency and low-complexity scalable GS coding solution that can be deployed with the existing MPEG DASH framework.
User identification is essential for securing electronic devices, especially in immersive environments such as virtual and augmented reality (VR/AR). Traditional identification methods typically rely on proactive, one-time authentication, rendering them susceptible to common security threats and inadequate for providing continuous protection after login. Furthermore, existing user identification solutions for VR/AR head-mounted displays (HMDs) often compromise user convenience or require additional hardware, increasing cost and complexity. Although significant research has explored biometric-based identification in VR/AR, many proposed methods face practical challenges: they often depend on extensive sensor data or demand intrusive user interactions, limiting their suitability for diverse real-world scenarios. In this work, we present a lightweight and practical biometric user identification algorithm that utilizes only head movement patterns—eliminating the need for auxiliary sensors such as external cameras or controller-based inputs, and imposing no extra tasks on the user. Our approach is platform-independent and can generalize to new tasks without retraining. We validate its effectiveness through comprehensive experiments on multiple public datasets. Results demonstrate that our method achieves high identification accuracy with minimal data and computational overhead, making it a scalable and viable solution for continuous user identification in next-generation AR/VR applications.
Carbon-efficient computing has drawn significant attention recently aiming to achieve environmental sustainability. Several computation intensive applications, such as deep learning, have been studied in the community for carbon optimizations. However, video streaming, one of the most popular and resource-consuming Internet applications, is under-explored for carbon efficiency. In this paper, we for the first time investigate the carbon efficiency of video streaming systems with the goal of reducing carbon emissions caused by hosting the video streaming service. We develop a dynamic workload migration mechanism utilizing real-time carbon intensity data to select the hosting data center. Furthermore, to minimize the impact on the end-user experience, we consider the migration frequency to maintain service stability when making the migration decisions. The evaluation results indicate significant carbon reductions and acceptable stream switching overhead.
Deploying steel plate surface defect detection models to edge devices necessitates lightweight and real-time capabilities. To address this, we propose an Ultra Lite Residual Module with a smaller parameter count, which we then use to construct an Ultra Lite ResNet (UL-ResNet). Compared to ResNet-50, our UL-ResNet achieves a 7.3× inference speedup and a 142.6× reduction in model parameters. To ensure that this highly simplified model can still effectively learn feature extraction and construction methods, we employ model distillation. This technique allows the feature extraction and classification capabilities of existing steel plate surface defect detection models to be transferred to our smaller model. Recognizing the disparities in feature maps and hierarchical feature levels between the teacher and student models, we also utilize a trainable attention module to facilitate knowledge transfer from the teacher model. Experimental results demonstrate that our student model effectively learns from the teacher model, achieving a classification accuracy of 99.44%.
High-resolution (HR) thermal imaging is essential in domains such as autonomous navigation and remote sensing, but its deployment is often limited by the cost of thermal sensors and the bandwidth required to transmit HR thermal data. This paper presents a cross-modal thermal image compression framework that leverages high-resolution RGB images as side information to enhance and efficiently encode low-resolution (LR) thermal inputs. Our method first applies a Guided Thermal Image Super-Resolution (GTISR) network to generate an initial HR thermal estimate from the LR thermal image and its aligned RGB counterpart. Rather than compressing the super-resolved output directly, we compute and encode only the residual using a learned image compression (LIC) model, while the LR thermal image is encoded separately using a standard codec such as Versatile Video Coding (VVC). Because the residual has substantially lower entropy, this two-stage approach significantly improves rate–distortion performance. Evaluated on the Perception Beyond the Visible Spectrum (PBVS) 2024 GTISR Challenge dataset, our framework achieves a 17.09% BD-rate reduction over a thermal-only baseline. These results demonstrate the benefit of integrating RGB side information into the compression pipeline and highlight the potential of cross-modal strategies for efficient thermal imaging under bandwidth constraints.
With the continuous growth in display technologies, computer vision and graphics, Volumetric Video is emerging as a promising frontier in research. However, the increasing interest in this field has led to growing confusion regarding video representations and terminologies. To address this, we present a taxonomy of videos based on interaction degrees of freedom and stereoscopic properties. Our work provides detailed insights into the technological foundations of each immersive media format, identifies their respective strengths and weaknesses, and offers a foundational understanding of volumetric video capture, reconstruction, and display. We hope this survey will stimulate future research and applications in immersive media.
Auditory models are useful tools for estimating perceptual attributes of a sound field. Integrating such auditory models in the optimisation of immersive sound systems is a promising strategy when listeners’ perception is central to the application. To that end, differentiability is key to allowing the perceptual model to be included in gradient-based optimisation loops. Existing differentiable models, however, are black-box deep-learning based, which limits their interpretability. In this paper, we propose an analytical white-box differentiable model of auditory localisation based on an existing non-differential model. Our evaluations show that the model produces outputs that are highly correlated with the outputs of the non-differential model and data collected in subjective listening tests. The proposed model also enables optimisation of amplitude panning laws in a stereophonic spatial sound field rendering through gradient descent. This study therefore demonstrates, more generally, the feasibility of designing and optimising immersive sound systems using white-box differentiable models of auditory perception.
Document layout analysis, a critical process in automated document processing, traditionally relies on object detection techniques, primarily focusing on the structural segmentation of documents. However, these approaches often fall short in comprehensively understanding the semantic content within the text, leading to a disjointed analysis of document structure and content. To address this, we propose a novel methodology that combines text clustering with multi-modal graph convolution networks, aiming to integrate structural detection with semantic understanding. Our approach starts with text detection, followed by encoding using a large language model. Subsequently, we integrate visual and positional data using Graph Neural Networks to perform clustering, creating a synergy between the textual and structural aspects of documents. Extensive experiments on mainstream datasets demonstrate that our method significantly outperforms existing approaches, especially in understanding text-centric document layouts. This paper contributes to the field by offering a novel, semantically-enriched approach to document layout analysis, enhancing the capabilities of automated document processing systems in handling diverse and complex document formats.
Recent advancements in large multimodal models (LMMs) have significantly enhanced both text-to-image (T2I) generation and image-to-text (I2T) interpretation. However, critical challenges in perceptual quality and text-image correspondence remain hindering the practicality of AI-generated images (AIGIs). Therefore, a reliable benchmark and automatic model for AIGI evaluation is desirable, which heavily relies on the scale and quality of human annotations. To this end, we present CompBench, the largest dataset for benchmarking and comparing image generation, which features: (i) the largest AIGI pair comparison dataset, comprising 616,346 carefully curated image pairs generated by 24 state-of-the-art AIGI models annotated with 1.6M+ human annotations, enabling robust relative quality assessment through pairwise comparison, (ii) multi-dimensional pairwise comparison from perceptual and text-image correspondence perspectives across three difficulty levels, and (iii) bidirectional benchmarking and evaluating for both T2I generation models and AIGI comparison models. Based on CompBench, we propose LMM4Comp, a LMM-based evaluation metric that learns nuanced quality distinctions from multiple dimensions for pairwise comparison at both instance level and model level. Experiments demonstrate that LMM4Comp achieves state-of-the-art performance, highly aligning to human preference. Both of the CompBench dataset and LMM4Comp metric will be released at https://github.com/IntMeGroup/CompBench.
Re-buffering is one of the most common factors that reduce the user’s Quality of Experience in online streaming video. In this paper, we introduce a novel solution to tackle the rebuffering problem in video streaming using generative Artificial Intelligence (genAI). In our proposed solution, a generative AI model is trained and executed to generate video frames in case of sudden network bandwidth drops. For that purpose, we first formulate the genAI-assisted adaptation problem in adaptive video streaming as an optimization problem. We then present a lightweight adaptation algorithm featuring a network reduction detection module and an AI-based video frame generation module. Experiment results show that the proposed method can effectively reduce the number of re-buffering events under challenging network conditions.