Dynamic Computed Tomography (4DCT) is widely used to visualize internal dynamic processes in medical and industrial contexts. However, 4DCT usually necessitates continuous multi-view scanning, which leads to high radiation exposure and low scanning speed. Sparse-view CT reconstruction has emerged as a promising solution by reducing scanning views. However, traditional methods often suffer from artifacts and noise under sparse-view setting. To overcome these limitations, we introduce a novel dynamic CT reconstruction framework that combines Gaussian Splatting with prior transfer. The proposed method utilizes a fast differentiable renderer and incorporates prior scene information to produce high-quality, spatio-temporally continuous reconstructions from sparse-view data. Furthermore, with self-supervised strategy of Gaussian Splatting model, our method avoids the need for paired CT-projection data, which are often difficult to acquire. This paper represents the first application of Gaussian Splatting to dynamic CT reconstruction, demonstrating its effectiveness under sparse-view conditions and achieving a 6.75 dB improvement compared to state-of-the-art methods.
Nowadays, high-quality images are pursued by both humans for better viewing experience and by machines for more accurate visual analysis. However, images are usually compressed before being consumed, decreasing their quality. It is meaningful to predict the perceptual quality of compressed images for both humans and machines, which guides the optimization for compression. This issue remains underexplored. In this paper, we propose a unified approach to fill this gap. Specifically, we create a deep learning-based model to predict Satisfied User Ratio (SUR) and Satisfied Machine Ratio (SMR) of compressed images simultaneously. We first pre-train a feature extractor network on a large-scale SMR-annotated dataset with human perception-related quality labels generated by diverse image quality models, which simulates the acquisition of SUR labels. Then, we leverage and fuse the extracted multi-layer features to predict SUR and SMR with a more capable network architecture. We propose a Difference Feature Residual Learning (DFRL) module to learn more discriminative difference features. We also design a Multi-Head Attention Aggregation and Pooling (MHAAP) layer to aggregate difference features and reduce their redundancy. We further introduce an MLP-Mixer module to integrate global spatial and channel information for subsequent SUR and SMR regression. Experimental results indicate that the proposed model significantly outperforms state-of-the-art SUR and SMR prediction methods by up to 48% on several challenging datasets. Moreover, our joint learning scheme of human and machine perceptual quality prediction tasks is effective at improving the performance of both.
Omnidirectional videos (ODVs) provide an immersive visual experience by capturing the 360° scene. With the rapid advancements in virtual/augmented reality, metaverse, and generative artificial intelligence, the demand for high-quality ODVs is surging. However, ODVs often suffer from low resolution due to their wide field of view and limitations in capturing devices and transmission bandwidth. Although video super-resolution (SR) is a capable video quality enhancement technique, the performance ceiling and practical generalization of existing methods are limited when applied to ODVs due to their unique attributes. To alleviate spatial projection distortions and temporal flickering of ODVs, we propose a Spatio-Temporal Distortion Aware Network (STDAN) with joint spatio-temporal alignment and reconstruction. Specifically, we incorporate a spatio-temporal continuous alignment (STCA) to mitigate discrete geometric artifacts in parallel with temporal alignment. Subsequently, we introduce an interlaced multi-frame reconstruction (IMFR) to enhance temporal consistency. Furthermore, we employ latitude-saliency adaptive (LSA) weights to focus on regions with higher texture complexity and human-watching interest. By exploring a spatio-temporal jointly framework and real-world viewing strategies, STDAN effectively reinforces spatio-temporal coherence on a novel ODV-SR dataset and ensures affordable computational costs. Extensive experimental results demonstrate that STDAN outperforms state-of-the-art methods in improving visual fidelity and dynamic smoothness of ODVs.
The growing popularity of virtual reality (VR) provides new opportunities for cultural heritage protection and online education. Many studies have investigated how to guide users to move to a specific view/position. However, the coguided interaction of position and gaze is still under researched. In this article, a novel interaction model named GazeTance Guidance (GTG) is proposed. This model utilizes head rotation and viewing distance to improve the effect of room-scale user guidance. To explore the efficiency of GTG, we conducted two within-subject studies with 25 and 31 participants navigating through virtual museums: first, we evaluated user acceptance of audio commentary and virtual guidance markers in terms of presence and user experience, and assessed the feasibility of memory tasks. second, we separately investigated the effects of gaze and distance guidance on motion sickness, cognitive load, memory efficiency, and user behavior. The experimental results prove that the guidance markers do not affect user experience and sense of presence. Compared with independent effects, the combination of gaze and distance guidance can reduce the user's motion sickness and significantly improve users' memory effect. These findings provide new inspiration for the interaction design of complex VR tours.
Region of Interest (ROI)-based image compression optimizes bit allocation by prioritizing ROI for higher-quality reconstruction. However, as the users (including human clients and downstream machine tasks) become more diverse, ROI-based image compression needs to be customizable to support various preferences. For example, different users may define distinct ROI or require different quality trade-offs between ROI and non-ROI. Existing ROI-based image compression schemes predefine the ROI, making it unchangeable, and lack effective mechanisms to balance reconstruction quality between ROI and non-ROI. This work proposes a paradigm for customizable ROI-based deep image compression. First, we develop a Text-controlled Mask Acquisition (TMA) module, which allows users to easily customize their ROI for compression by just inputting the corresponding semantic text. It makes the encoder controlled by text. Second, we design a Customizable Value Assign (CVA) mechanism, which masks the non-ROI with a changeable extent decided by users instead of a constant one to manage the reconstruction quality trade-off between ROI and non-ROI. Finally, we present a Latent Mask Attention (LMA) module, where the latent spatial prior of the mask and the latent Rate-Distortion Optimization (RDO) prior of the image are extracted and fused in the latent space, and further used to optimize the latent representation of the source image. Experimental results demonstrate that our proposed customizable ROI-based deep image compression paradigm effectively addresses the needs of customization for ROI definition and mask acquisition as well as the reconstruction quality trade-off management between the ROI and non-ROI. Additionally, even by using the uniform mask as input, our method still outperforms the anchor methods in image reconstruction and machine vision tasks (such as object detection and instance segmentation). Our source code will be available at: https://github.com/hccavgcyv/Customizable-ROI-Based-Deep-Image-Compression.
Video compression has recently benefited from implicit neural representations (INRs), which model videos as continuous functions. INRs offer compact storage and flexible reconstruction, providing a promising alternative to traditional codecs. However, most existing INR-based methods treat the temporal dimension as an independent input, limiting their ability to capture complex temporal dependencies. To address this, we propose a Hierarchical Temporal Neural Representation for Videos, TeNeRV. TeNeRV integrates short- and long-term dependencies through two key components. First, an Inter-Frame Feature Fusion (IFF) module aggregates features from adjacent frames, enforcing local temporal coherence and capturing fine-grained motion. Second, a GoP-Adaptive Modulation (GAM) mechanism partitions videos into Groups-of-Pictures and learns group-specific priors. The mechanism modulates network parameters, enabling adaptive representations across different GoPs. Extensive experiments demonstrate that TeNeRV consistently outperforms existing INR-based methods in rate-distortion performance, validating the effectiveness of our proposed approach.
With the growing demand for immersive visual experiences, high-quality omnidirectional images (ODIs) have become increasingly important. However, limitations in imaging devices and transmission bandwidth often lead to low-resolution ODIs, hindering the rendering of fine-grained 360° details, especially in the presence of real-world degradations and geometric distortions. Existing real-world super-resolution (Real-SR) methods are inadequate for ODIs, as their degradation models fail to account for the complex imaging pipeline involving fisheye capture and Equirectangular Projection (ERP), introducing severe aliasing and projection-specific distortions. To address these challenges, we propose D^2R^2OSR, a Degradation-Disentangled Representation framework for Real-world Omnidirectional image Super-Resolution. D^2R^2OSR explicitly models degradations arising from both fisheye imaging and ERP projection, guided by two key insights: (1) projection priors play a critical role in shaping real-world degradations, and (2) human perception in immersive environments is inherently viewpoint-centric. Accordingly, we introduce a Perspective Projection Representation (PPR) operating alongside the ERP branch to capture viewpoint-aware features, together with a Degradation-Specific Module (DSM) that jointly models ERP-induced geometric distortions and PPR-specific real-world degradations. Extensive experiments demonstrate that D^2R^2OSR achieves state-of-the-art performance and produces visually compelling, high-fidelity omnidirectional Real-SR results while maintaining favorable computational efficiency for low-resource deployment.
Efficient point cloud compression (PCC) is essential for immersive 3D applications, such as 3D telepresence and remote collaboration, requiring high visual quality despite limited network bandwidth. However, existing PCC methods often prioritize data fidelity while sacrificing perceptual quality, particularly in low-bitrate scenarios. To address these issues, we innovatively propose GS-PCC, the first coding framework that incorporates 3D Gaussian Splatting (3DGS) for PCC. This approach enables high-quality reconstruction with superior perceptual fidelity and optimized bitrate efficiency. Furthermore, GS-PCC provides a unified compression framework that jointly encodes both geometry and attributes. Experimental results demonstrate that GS-PCC significantly outperforms MPEG G-PCC and Google Draco, offering an integrated and efficient solution for human-centric, bandwidth-constrained 3D applications.
As a retina-inspired sensor with ultra-high temporal resolution, spike camera can continuously capture dynamic scenes with high-speed motion. It is a key task to restore clear images from spike streams. The quantization effects in spike readout bring degradation to the visual quality of restored images. To tackle the degradation without introducing motion blur, existing methods often employ a short-term temporal window to infer the light intensity at a certain time point. However, these methods only focus on the spike signals within the current window, which limits their performance. Motivated by the human-like memory mechanism for visual signals from the retina, we explore Spike Stream Memory Transfer (SSMT) to restore the dynamic scenes, considering spike signals beyond the window. Specifically, we design a framework that leverages temporal memory by transferring previously inferred light intensity and motion to enhance current reconstruction. The framework enables a long-term temporal perception of spike streams to handle the spike quantization effects. Besides, we utilize the estimated motion to suppress the potential blur from inter-stream clips, considering the underlying motion of spike streams. We also develop a spike interval-guided alignment module to tackle the blur from intra-stream clips. Experimental results on both synthetic and real-captured data demonstrate that our method can restore high-quality images from spike streams.
Implicit Neural Representations (INRs) have demonstrated significant potential in video compression by representing videos as neural networks. However, as the number of frames increases, the memory consumption for training and inference increases substantially, posing challenges in resource-constrained scenarios. Inspired by the success of traditional video compression frameworks, which process video frame by frame and can efficiently compress long videos, we adopt this modeling strategy for INRs to decrease memory consumption, while aiming to unify the frameworks from the perspective of timeline-based autoregressive modeling. In this work, we present a novel understanding of INR models from an autoregressive (AR) perspective and introduce a Unified AutoRegressive Framework for memory-efficient Neural Video Compression (UAR-NVC). UAR-NVC integrates timeline-based and INR-based neural video compression under a unified autoregressive paradigm. It partitions videos into several clips and processes each clip using a different INR model instance, leveraging the advantages of both compression frameworks while allowing seamless adaptation to either in form. To further reduce temporal redundancy between clips, we treat the corresponding model parameters as proxies for these clips, and design two modules to optimize the initialization, training, and compression of these model parameters. In special, the Residual Quantization and Entropy Constraint (RQEC) module dynamically balances the reconstruction quality of the current clip and the newly introduced bitrate cost using the previously optimized parameters as conditioning. In addition, the Interpolation-based Initialization (II) module flexibly adjusts the degree of reference used during the initialization of neighboring video clips, based on their correlation. UAR-NVC supports adjustable latencies by varying the clip length. Extensive experimental results demonstrate that UAR-NVC, with its flexible video clip setting, can adapt to resource-constrained environments and significantly improve performance compared to different baseline models. The project page: https://wj-inf.github.io/UAR-NVC-page/.
Implicit Neural Representation (INR) has emerged as a promising paradigm for video compression, offering a continuous and compact representation of video. Existing INR-based video compression methods do not fully exploit the representation capability of INR within limited parameter constraints, leaving the potential for optimal rate-distortion performance underexplored. This work introduces RNeRV, a recurrent neural representation for video compression. RNeRV enhances the expressive power of INR by reusing parameters and improves the extraction of spatiotemporal features for accurate video reconstruction.We further provide theoretical analysis to support the effectiveness of parameter reuse. Moreover, we introduce a mixed spatiotemporal grid, which fully integrates spatial and temporal features in one grid to jointly compress spatiotemporal information. Experimental results demonstrate that RNeRV achieves superior performance compared to advanced INR-based video compression methods, providing a step forward in INR-based video compression.
Recent advances in video compression introduce implicit neural representation (INR) based methods, which effectively capture global dependencies and characteristics of entire video sequences. Unlike traditional and deep learning based approaches, INR-based methods optimize network parameters from a global perspective, resulting in superior compression potential. However, most current INR methods utilize a fixed and uniform network architecture across all frames, limiting their adaptability to dynamic variations within and between video sequences. This often leads to suboptimal compression outcomes as these methods struggle to capture the distinct nuances and transitions in video content. To overcome these challenges, we propose Hierarchically Adaptive Neural Representation for Video Compression (HANeRV), an innovative INR-based video compression network that adaptively conducts structure optimisation based on the specific content of each video sequence. To better capture dynamic information across video sequences, we propose a dynamic architecture-level adjustment (DAA). Furthermore, to enhance the capture of dynamics between frames within a sequence, we implement a dynamic frame-level adjustment (DFA). Finally, to effectively capture spatial structural information within video frames, thereby enhancing the detail restoration capabilities of HANeRV, we devise a structure level hierarchical structural adaptation (HSA). Experimental results show that HANeRV achieves state-of-the-art performance among INR-based video compression methods and surpasses the H.266/VVC (x266, medium preset) anchor on diverse datasets.
Implicit Neural Representations (INRs) have emerged as a promising paradigm for video compression. However, existing INR-based frameworks typically suffer from inherent spectral bias, which favors low-frequency components and leads to over-smoothed reconstructions and suboptimal rate-distortion performance. In this paper, we propose FaNeRV, a Frequency-aware Neural Representation for videos, which explicitly decouples low- and high-frequency components to enable efficient and faithful video reconstruction. FaNeRV introduces a multi-resolution supervision strategy that guides the network to progressively capture global structures and fine-grained textures through staged supervision . To further enhance high-frequency reconstruction, we propose a dynamic high-frequency injection mechanism that adaptively emphasizes challenging regions. In addition, we design a frequency-decomposed network module to improve feature modeling across different spectral bands. Extensive experiments on standard benchmarks demonstrate that FaNeRV significantly outperforms state-of-the-art INR methods and achieves competitive rate-distortion performance against traditional codecs.
Learning video information only by their category names limited the development of the generalized zero-shot video classification (GZSVC) task. By analyzing the way that humans learn new things, we found that people can utilize knowledge such as textual concepts and visual fundamentals to construct new video cognition. Taking this as inspiration, we propose a multi-modal knowledge-driven approach to solve the GZSVC task by searching and learning various knowledge. In the real world, it is hard to guarantee that important components of new videos can be covered by existing knowledge. To bridge this knowledge gap, our method constructs a reliable knowledge supplement from multi-modal information for categories, which can also establish connections between classes. In order to fuse the information from different modalities, we propose a multi-modal generative model to synthesize visual features that are rich in content and closer to the true distribution of videos. Since training process lacks real unseen visual information, we propose that the model should pay more attention to semantic information in this task, and we strengthen the constraint and utilization of semantic information in the proposed framework. Extensive experimental results on various databases show that our proposed method outperforms the state-of-the-art GZSVC methods.
Point cloud sequences play a pivotal role in immersive applications as they provide 6 degree of freedom experience into dynamic 3D environments. However, the dynamic point cloud sequences have an overwhelming amount of data, which poses challenges for storage and transmission. Therefore, effective point cloud sequence compression is crucial for immersive applications. In this paper, we propose a dynamic point cloud compression framework, which enhances the compression efficiency of point cloud sequences from the perspectives of representation and prediction. To improve the representation and reconstruction ability of a single frame, we propose a specialized pre-training mechanism for compression, which enables a well-designed autoencoder to acquire substantial prior knowledge. The pre-training mechanism is based on the proposed mask operation and is trained on a large amount of data of multiple categories. This enables the pre-trained encoder to identify and extract features with strong representational capabilities, and allows the pre-trained decoder to accurately reconstruct the complete point cloud. For precise prediction, we develop tailored methods for different frame types. For intra-frames (I frames), we introduce a dual-branch spatial intra-prediction network by leveraging coordinate and occupancy information from downsampled sparse point clouds to predict representative features. For inter-frames (P frames), we propose a multi-scale context-based spatio-temporal inter-prediction network. Specifically, for each scale, we utilize the proposed Transformer-style spatio-temporal modeling module to analyze and model the temporal characteristics among multiple reference frames, providing rich context information. Integrating three-scale contexts, we achieve comprehensive feature prediction, significantly improving accuracy and eliminating spatio-temporal redundancy. Experimental results show that our method outperforms existing compression methods in both I-frame compression and P-frame compression.
Spike camera is a kind of bio-inspired neuromorphic camera which is designed for capturing dynamic scenes of ultra-high speed motion at extremely high temporal resolution. It adopts an “integrate-and-fire” mechanism to convert the dynamic arrivals of photons at each sensor pixel to a stream of asynchronously fired spikes. The occurrence of spike firing may be disturbed by multiple factors, including the Poisson nature of photon arrivals and the quantization effect in spike readout. Therefore, recovering high-quality visual images from the recorded spike stream is an important yet challenging problem. This paper presents a reconstruction scheme for spike camera based on Bayesian imaging framework. We discuss the probability model of photon arrival in the imaging process and derive the likelihood model of the Bayesian framework. To fully exploit the dynamic information of spike streams, temporal correlation is used to model the image prior of reconstructed scenes. To be specific, we propose an auto-regressive model based on a non-stationary Laplacian distribution to model the temporal correlation. An efficient way to solve the optimization problem of the proposed Bayesian framework is further given. Experimental results show that the proposed method improves the reconstruction quality on both real-captured and synthesized spike datasets.
Pulmonary embolism (PE) is a life-threatening condition for which computed tomography pulmonary angiography (CTPA) is the standard diagnostic modality. However, conventional CTPA protocols require relatively high iodine contrast and radiation doses, raising concerns about renal injury and radiation exposure. In this study, we propose a deep learning-based framework for PE diagnosis under low-iodine and low-radiation CTPA conditions. The proposed two-stage framework integrates image enhancement and classification by jointly leveraging original low-exposure images and their super-resolved counterparts. We further construct and publicly release a low-iodine, low-radiation CTPA dataset developed in collaboration with a clinical institution to support reproducible research in safe imaging. Experimental results demonstrate that the proposed method substantially improves diagnostic performance compared with single-branch baselines, achieving an area under the ROC curve (AUC) of 0.928 while maintaining balanced sensitivity and specificity. These findings suggest that the proposed framework enables accurate and safer PE diagnosis under reduced contrast and radiation exposure, offering a practical solution for improving diagnostic safety in clinical CTPA imaging.
Human-centric multi-view video has a clear semantic structure: a static background and dynamic human motion. We propose a generative compression framework that explicitly decouples these components. The background is modeled once with 3D Gaussian Splatting, while the human is represented by a personalized Gaussian avatar reconstructed from a sparse set of key views that are transmitted only once and driven by compact per-frame pose parameters from the Skinned Multi-Person Linear (SMPL) model. The encoder sends only three elements: the background, the key views, and the SMPL parameters, enabling high-fidelity multi-viewpoint synthesis at dramatically reduced bitrates. This shifts compression from low-level redundancy removal to semantics-aware generative modeling. Experiments across multiple human-centric datasets demonstrate superior rate–distortion performance, particularly for long and densely captured sequences, and naturally enable semantic editing.
Light field images capture multi-view scene information and play a crucial role in 3D scene reconstruction. However, their high-dimensional nature results in enormous data volumes, posing a significant challenge for efficient compression in practical storage and transmission scenarios. Although neural representation-based methods have shown promise in light field image compression, most approaches rely on direct coordinate-to-pixel mapping through implicit neural representation (INR), often neglecting the explicit modeling of scene structure. Moreover, they typically lack end-to-end rate-distortion optimization, limiting their compression efficiency. To address these limitations, we propose SANR, a Scene-Aware Neural Representation framework for light field image compression with end-to-end rate-distortion optimization. For scene awareness, SANR introduces a hierarchical scene modeling block that leverages multi-scale latent codes to capture intrinsic scene structures, thereby reducing the information gap between INR input coordinates and the target light field image. From a compression perspective, SANR is the first to incorporate entropy-constrained quantization-aware training (QAT) into neural representation-based light field image compression, enabling end-to-end rate-distortion optimization. Extensive experiment results demonstrate that SANR significantly outperforms state-of-the-art techniques regarding rate-distortion performance with a 65.62% BD-rate saving against HEVC.
Computed tomography (CT) is a significant clinical detection method without invasive procedures. However, high-dose radiation is associated with elevated risk of cancer. Low-dose CT reconstruction techniques have the potential to significantly reduce this risk. These techniques reduce radiation dose by two means: low radiation intensity and incomplete scans. The focus of this review is incomplete-scan CT reconstruction techniques, including sparse-view and limited-angle CT reconstruction. In the context of incomplete scans, CT images obtained by conventional CT reconstruction techniques exhibit noises and streak artifacts. Deep learning techniques are changing this situation. However, there are few reviews that recently focus on the fast developing area. This paper reviews recent research on CT reconstruction of incomplete scans via deep learning and classifies them into two categories including image domain and reconstruction domain processing methods. This study also evaluates the strengths and limitations of these approaches and discusses future prospects, highlighting promising technologies and directions.