Existing image compression methods tend to suffer from texture distortion, structural blurring, and loss of semantic details at ultra-low bitrates. To address these issues, this paper proposes a Semantic Token-guided Generative Latent Coding (STGLC) framework for ultra-low bitrate image compression, which leverages high-level semantic information to guide latent representations toward preserving semantic structures. Specifically, we introduce visual semantic tokens as priors and construct a cross-level feature interaction module to guide latent features in prioritizing the encoding of semantic information. Additionally, to maximize the retention of semantic details at ultra-low bitrates, we design a context-semantic-aware patch-wise entropy model that dynamically allocates bitrates during quantization and encoding based on the contribution of latent variables to semantic integrity. Extensive experiments on the DIV2K and CLIC2020 datasets demonstrate that the proposed framework consistently achieves state-of-the-art visual reconstruction performance, maintaining high quality even under extreme compression conditions below 0.04 bpp.
The Internet of Bodies (IoB) requires efficient transmission of correlated physiological signals under strict bandwidth constraints. We propose DiSC2, a bandwidth-efficient distributed semantic communication framework for cross-modal redundancy suppression, which combines distributed semantic encoders, cloud-side cross-modal attention, a hybrid task-anddistortion loss, and a lightweight quantization-aware edge encoder. Unlike independent waveform transmission, DiSC2 prioritizes complementary task-relevant information across modalities and shifts redundancy exploitation to the cloud. Experiments on PttPPG and ScientISST MOVE show that DiSC2 preserves downstream classification performance more effectively than SSC and DWT+LDPC baselines, especially under low-SNR and low-CBR conditions, while remaining suitable for resource-constrained wearable sensing.
Hyperspectral images (HSIs), with rich spatial–spectral information, are widely used in remote sensing, precision agriculture, and environmental monitoring. In satellite and aerial platforms, HSIs enable large-scale surface observation and land-cover classification, but their high dimensionality creates severe challenges for bandwidth-limited and time-varying communication channels. Most existing reconstruction methods assume ideal links, while joint source–channel coding (JSCC) has shown robustness for natural images but does not account for hyperspectral dependencies. This paper presents HSI-JSCC, the first end-to-end framework that integrates hyperspectral reconstruction with JSCC to jointly optimize spatial–spectral modeling and transmission robustness. The framework introduces two key components: (i) a Cumulative Hybrid Attention Block (CHAB) that combines Mamba and Transformer for multi-scale feature extraction, and (ii) a spectral-aware modulation strategy that guides bandwidth allocation based on spectral importance and channel states. Experiments on the CAVE and KAIST datasets show that HSI-JSCC achieves higher fidelity than existing methods, particularly under low-rate and low-SNR conditions, demonstrating its potential for reliable hyperspectral communication in practical remote sensing systems.
Learning visual concept tokens from multi-object scenes faces a fundamental challenge of feature entanglement, where tokens inadvertently capture information from each other, leading to degraded object fidelity and unreliable compositional generation. We identify the root cause as the lack of hierarchical learning strategies and explicit decoupling constraints during training. To address these limitations, we propose a decoupled visual concept token learning framework based on two-stage hierarchical optimization. In Stage I, we perform embedding-level token learning by integrating new tokens into the pretrained text-embedding distribution while enforcing inter-token separation via Token-space Concept Separation and Alignment Loss (TCSAL). In Stage II, we impose attention-level decoupling during image generation using Visual-Concept Semantic Alignment Loss (VCSAL) to suppress cross-token interference. To support stable and prior-preserving training, we further introduce a VLM-guided Adaptive Data Augmentation (ADA) module to increase training diversity while mitigating language drift and preserving pretrained knowledge. This hierarchical design enables visual concept tokens to function as independent semantic units, facilitating fine-grained attribute control and reliable compositional generation. We also introduce an End-of-Text (EOT) embedding-based token-level evaluation metric to directly assess the semantic quality of learned visual concept tokens. Extensive experiments demonstrate the effectiveness of the proposed framework and show that it achieves strong and competitive performance on visual concept token learning tasks.
The growing prevalence of stereo cameras necessitates efficient Stereo Image Compression (SIC). However, existing SIC methods mainly model redundancy between the two observed 2D views and rarely exploit the continuous cross-view structure underlying a stereo pair. To address this limitation, we introduce a novel geometric-aware stereo image compression framework, GeoSIC. We construct a continuous implicit 3D feature field in a feature-domain planar-viewpoint coordinate space, which models the continuous transition across views and provides a geometry-compatible prior for SIC. We further design a geometric-aware fusion module to perform feature refinement and quantization-loss compensation. In addition, to obtain informative novel-reference features that facilitate conditional coding, we introduce an unobserved reference-viewpoint prediction to query the implicit field. Extensive experiments show that GeoSIC consistently outperforms state-of-the-art methods. Compared with the BiSIC baseline, GeoSIC achieves an average BD-rate reduction of 13.7% and an average BD-PSNR/BD-MSSSIM gain of 0.454 dB.
Integrating generative diffusion models with contrastive learning effectively mitigates data sparsity in recommender systems. However, existing strategies predominantly rely on random Gaussian noise, which overlooks inherent topological structures and induces distribution shifts. Furthermore, simplistic linear fusion leads to feature entanglement between user intents and noise. To address these challenges, we propose AdaDCL, a structure-aware framework based on structured priors and disentangled modulation. First, a Variational Autoencoder extracts latent parameters to construct a generative prior enriched with structural information. Subsequently, a Feature-wise Linear Modulation module transforms auxiliary semantics into dynamic parameters, enabling disentangled control over latent features. This ensures that augmented signals integrate semantic information while preserving crucial collaborative relations. Finally, the fused signals are injected into a diffusion model for reconstruction to generate authentic user preferences as contrastive views, employing a multi-task collaborative optimization strategy to align generative and discriminative gradients.Extensive experiments on three datasets demonstrate that AdaDCL achieves the best overall mean performance against state-of-the-art baselines, with statistical significance on the majority of reported metrics and consistent numerical gains across all metrics.
Underwater image transmission demands efficient compression due to severe visual degradation and the stringent bitrate limitations of acoustic communication. Existing learning-based approaches often rely on auxiliary networks to extract underwater physical priors and manually designed dictionaries, leading to a less integrated, not strictly end-to-end design. We propose a novel underwater image compression framework that combines an Implicit Neural Representation (INR)-assisted decoder with an adaptive Underwater Dictionary-based Entropy Model (UDEM). The INR branch captures low-frequency degradations, leveraging its inductive bias to align with underwater distortions. Meanwhile, UDEM integrates frequency-domain enhancement with a learnable underwater feature dictionary to improve probability modeling for entropy coding. Experiments across five benchmark datasets verify that the proposed framework achieves superior rate-distortion performance in the low-bitrate regime.
Generative Semantic Communication (GSC) is a promising solution for image transmission over narrow-band and high-noise channels. However, existing GSC methods rely on long, indirect transport trajectories from a Gaussian to an image distribution guided by semantics, causing severe hallucination and high computational cost. To address this, we propose a general framework named Schrödinger Bridge-based GSC (SBGSC). By leveraging the Schrödinger Bridge (SB) to construct optimal transport trajectories between arbitrary distributions, SBGSC breaks Gaussian limitations and enables direct generative decoding from semantics to images. Within this framework, we design Diffusion SB-based GSC (DSBGSC). DSBGSC reconstructs the nonlinear drift term of diffusion models using Schrödinger potentials, achieving direct optimal distribution transport to reduce hallucinations and computational overhead. To further accelerate generation, we propose a self-consistency-based objective guiding the model to learn a nonlinear velocity field pointing directly toward the image, bypassing Markovian noise prediction to significantly reduce sampling steps. Simulation results demonstrate that DSBGSC outperforms state-of-the-art GSC methods, improving FID by at least 38
Existing image compression models often lack personalization capabilities, treating all image regions equally and failing to meet the compression needs of different users for specific Regions of Interest (ROI). To address this challenge, we propose an innovative variable rate image compression framework that achieves user-centric dynamic compression by introducing visual in-context learning. Our method extracts cross-image semantics from user-provided visual examples to understand their intent. This semantic information is then converted into visual semantic query tokens and spatial masks to effectively guide the bit allocation of the compression model. Furthermore, we design a novel Semantic Spatial Control Block (SSCB) to fully leverage these semantic and spatial cues, thereby achieving a balance between preserving user-specified details and overall image quality. Experimental results demonstrate that our method significantly improves performance on the ROI, achieving a 31.54 % BD-Rate reduction and a 2.7479 dB BD-PSNR gain over the baseline model.
In recent years, Text-to-Image (T2I) models have made remarkable advancements, yet accurate accurate association of attributes remains a key challenge. This paper presents FreeAlign, a novel training-free framework designed to enhance attribute alignment in T2I generation. By modulating attention and adapting U-Net components, FreeAlign achieves precise alignment between image attributes and textual descriptions. It strengthens attribute-target associations through refining attention maps, adjusts U-Net’s backbone and skip connections based on energy ratios, and reorders prompts to balance attribute focus. Large Language Models (LLMs) enrich prompts with diverse, contextually relevant text, enhancing diffusion models’ generative power and quality. Extensive experiments show that FreeAlign delivers superior alignment for diverse prompts while preserving intricate details and ensuring structural integrity, establishing a new benchmark for attribute precision in T2I generation.
Although diffusion models advance condition-based visual generation, they suffer from speed and cost issues, unlike faster AutoRegressive methods that are limited in performance. To address these, we introduce the Stable Control Visual AutoRegressive Model (SCVAR). SCVAR ensures stable control by aligning visual conditions on multiple scales. Rather than unfolding the 2D image into a 1D raster, SCVAR decouples it into multiple scales. This shifts the sequential representation in SCVAR from tokens to scales, satisfying the unidirectional dependency of the AR model while preserving the 2D structure of the image. Compared to indiscriminate conditional guidance, cross-scale alignment provides more precise constraints, enabling SCVAR to achieve state-of-the-art performance in experiments against diffusion models, with 10x faster generation speed. The decoupled condition also reduces training costs. Compared to end-to-end conditional computation, experiments demonstrate that SCVAR matches performance with only 40% additional parameters.
In image semantic communication, the granularity of semantic descriptions required for different objects within an image varies based on the communication intent and the importance of the objects. However, current semantic codecs optimized for global assessment metrics fail to adapt to user intent and cannot provide differentiated semantic granularity for objects of different importance. Generative semantic codecs using representations such as edge maps or semantic segmentation maps are insufficient for capturing fine-grained semantic information. This paper proposes dividing the transmitted image semantics into global coarse-grained and key object fine-grained semantics to better align with sender intent and optimize bandwidth usage. We introduce a novel semantic codec scheme based on a pre-trained text-to-image diffusion model. Global coarse-grained semantics are represented using short textual descriptions. Fine-grained semantic information of key objects is extracted using the Denoising Diffusion Implicit Model (DDIM) inversion and compressed in the frequency domain. Experimental results demonstrate that the proposed semantic codec enables high-quality recovery of coarse- and fine-grained semantics in image transmission while significantly reducing data transmission requirements.
随着移动通信与人工智能技术的飞速发展,当今社会对于信息智能化处理的需求迅速攀升,信息处理技术正在经历着由信号处理向内容处理的深刻变革。以傅里叶分析为代表的传统信号分析方法在分析和刻画信号时-频特性方面取得了巨大的成功,而在内容层面的信号语义特性分析方法却存在着一定的研究空白。因此,本文致力于为语义信息分析提供新的视角,提出了一种基于语义分解的语义信息分析方法。该方法认为,一个大颗粒度的语义信息可以分解为多个小颗粒度的语义信息,正如傅里叶变换可以将时域信号分解成多个不同频率分量的信号。本文所提出的语义分解方法关注到现实世界中语义信息天然存在的层级结构,其核心思路在于将复杂的语义信息分解为基本的单元组合结构,从而进行表征与理解。为此,本文提出了一系列核心概念,包括语义基元、语义分解树和语义颗粒度等,并定义了语义基元提取的相关运算和方法。通过实验验证,本文提出的语义信息分解方法能够有效地分解和表征语义信号,为信息处理技术的进一步发展提供了理论基础和实践指导。
Semantic communication, leveraging deep network-based Joint Source-Channel Coding (JSCC), has garnered increasing attention in recent years. However, existing methods are primarily suitable for transmitting single-scale semantics such as pixels, but not for adaptively fusing multi-scale semantics such as objects and scenes. Owing to substantial variations in data volume across different semantic scales, selecting the appropriate semantic scale for transmission based on varying Channel State Information (CSI) can significantly enhance the efficiency of conveying semantic information. This letter introduces a cross-scale Generative Semantic Communication (GSC) method for image transmission, named BriGSC. Under the constraints of CSI, our method can jointly perceives textual and visual features to represent semantics at different scales, achieves rate-adaptive encoding, transmission, decoding and image generation. The experiment results show that compared with semantic communication methods based on deep learning (SwinJSCC) and generative models (SGD-JSCC), our method has better competitive noise resistance and coding efficiency through jointly encoding multi-scale semantic features. Under various channel conditions, the average values of FID and LPIPS were 35% and 25% lower than SwinJSCC and SGD-JSCC respectively. The code is available at https://github.com/AsanoSaki/BriGSC.
Recent advancements in deep learning for semantic communication have been significant, yet fixed-length encoding techniques struggle to capture the complex and variable nature of semantic content, potentially compromising detail and transmission efficiency. This letter presents a novel semantic encoding framework featuring scalable coding to enhance the processing of diverse semantic targets within images. Our approach decomposes raw, unstructured images into hierarchically structured semantic features, thereby enriching semantic representation accuracy. The proposed scalable coding method employs a dynamic resource allocation strategy, guided by a semantic knowledge base and user intent, to selectively encode each semantic target. This approach yields substantial improvements in communication efficiency. Additionally, we introduce a semantic alignment strategy that optimizes reconstructed image edges and quality through specialized loss functions. Experimental results demonstrate that, compared with the latest methods, our approach achieves user-intent-aware flexible encoding and decoding without sacrificing performance. This highlights the effectiveness of our framework in maintaining high coding efficiency and image reconstruction quality, while enabling adaptive semantic communication.
Unlike traditional bit-level data transmission methods, semantic communication focuses on conveying the meaning behind the data. Though promising results have been achieved, existing end-to-end learning-based semantic communication frameworks often require a synchronization of deep models between the transmitter and the receiver. Such design leads to tens of thousands models to be stored at receiver since different manufactures may optimize their own models. To address this problem, we propose a novel model-unaware generative image compression framework for semantic communication. It features at employing human-understandable multi-modality representations as an intermediate layer to enhance information transmission efficiency and semantic consistency. Our framework introduces a mask-based rate-distortion optimization module, which effectively removes low-relevance information for image generation and reduces the bit rate while maintaining semantic consistency. Experimental results demonstrate that the framework can still reconstruct high-quality images at very low bit rates, showcasing its potential for applications in modern communication systems.
In recent years, semantic communication based on deep learning for source-channel joint encoding has garnered significant attention. It utilizes network models trained end-to-end to represent signals as embedding vectors and has demonstrated superior performance compared to traditional methods. However, due to the significant disparity between embedding vectors and human language, it can be challenging to succinctly capture abstract semantics such as scenes. In this paper, we introduce the Scene Graph-based Generative Semantic Communication (SG2SC) framework, built upon structured semantics and conditional generative models for image transmission. SG2SC aims to faithfully convey abstract semantics like scenes. It begins by detecting object categories, spatial attributes, and inter-category relationships in the image, representing scene semantics in a graph structure. Subsequently, it employs graph neural networks for scene graph encoding, decoding, and transmission, and finally utilizes a conditional diffusion model for semantic decoding. Benefiting from its concise graph structure semantics, SG2SC outperforms traditional method, semantic communication based on deep joint source-channel coding, and segmentation-based generative semantic communication in terms of noise resistance and encoding efficiency.
Despite deep learning's progress in semantic communication, traditional fixed-length encoding does not adequately address the variable complexity of semantic content, often leading to loss of critical nuances and reduced communication accuracy. Current methods also introduce unnecessary redundancy, compromising transmission efficiency. Addressing these challenges, our work introduces an innovative adaptive rate encoding mechanism that captures the intrinsic semantics of images and fine-tunes the coding rate based on semantic interconnection probability. We employ a cross-attention model to construct a layered semantic probability graph parsed into a hierarchical semantic tree, which represents the probabilistic relationships of image semantics and unravels the latent semantic structure. This not only delineates the image's semantic architecture but also enables our adaptive encoding to dynamically allocate resources, minimizing redundancy and enhancing efficiency. Our experiments confirm that our approach provides a more judicious bit allocation to complex image features and allocates more bits to semantically rich features while achieving superior compression of simpler content. The proposed method not only improves upon existing semantic fidelity metrics but also reduces the bit demand for transmitting complex images. Our adaptive encoding strategy represents a significant stride in leveraging the endogenous semantic information of images for more accurate and efficient communication.
Real-time electrocardiogram (ECG) monitoring and diagnosis through Internet of Things (IoT) are crucial for addressing the severity and timely treatment of cardiovascular diseases, enabling timely intervention and preventing life-threatening complications. However, current ECG monitoring research predominantly focuses on individual aspects such as signal compression, diagnostic analysis, or secure transmission, lacking joint optimization of various modules in IoT scenarios. To address this gap, this work proposes a novel framework based on superimposed semantic communication for real-time ECG monitoring in IoT. The framework comprises three hierarchical levels: the edge level for data collection and processing, the relay level for signal compression and coding, and the cloud level for data analysis and reconstruction. The proposed framework offers several unique advantages. By employing semantic encoding guided by ECG classification tasks, it selectively extracts crucial features within and between signals, improving compression ratio and adaptability to channel noise. The superimposed semantic encoding achieves content encryption without requiring any additional operations. Moreover, the framework utilizes lightweight anomaly detection neural networks, reducing edge device power consumption and conserving communication resources. Simulation and real experimental results demonstrate that the proposed method achieves real-time encoding and transmission of ECG signals with a compression ratio of 0.019 on the MIT-BIH dataset. Furthermore, it attains a heartbeat classification accuracy of 0.988 and a reconstruction error of 0.061.
Monitoring multimodal signals provides a more comprehensive understanding of health conditions compared to singlemode monitoring. In the face of the significant volumes of multimodal signals, existing IoT health monitoring systems primarily focus on high-fidelity signal transmission by encoding multimodal signals separately. However, due to the lack of consideration for the downstream applications and correlation between multimodal signals, a portion of bandwidth resources is wasted on task-irrelevant information and intermodal redundancy. To address this issue, we propose the Multimodal Semantic Integration Communication (MoSIC) framework composed of three levels: At the sensor level, multiple wearable sensors collect and send different modal signals to a mobile terminal; at the mobile terminal level, the terminal employs deep source-channel joint encoding for the received multimodal signals, extracting single-modal embedded features using a backbone network, and obtaining cross-modal features through a feature fusion network with contrastive constrain; at the cloud level, a decoding network symmetric to the encoding network reconstructs the multimodal signals, which are then used for downstream applications such as human activity recognition. MoSIC focuses on semantically integrating multimodal signals for downstream applications, resulting in improved encoding and transmission efficiency. It also reduces the radio-frequency power consumption and bandwidth requirements.