The advancement of unmanned aerial vehicle (UAV) technology has driven their widespread deployment as aerial relays for ground nodes (GNs) in jamming environments. However, dynamic jamming patterns and the inherent mobility of UAVs result in differentiated quality-of-service (QoS) for GNs in UAV-assisted data collection systems. We formulate a joint optimization problem that aims to maximize the fairness throughput of the system, subject to constraints on GNs association, UAV trajectory planning, and channel allocation. To address this problem, we develop a heuristic attention-enhanced QMIX (HAEQMIX) algorithm, which achieves GNs association through a heuristic fairness access method, and the attention-enhanced QMIX (AEQMIX) scheme adaptively aggregates heterogeneous node features for trajectory design and channel selection. Simulation results demonstrate that our algorithm exhibits superior anti-jamming capabilities and achieves the highest fairness throughput across various node densities, outperforming state-of-the-art works.
To achieve optimal performance in large-scale Internet of Things (IoT) deployment, urban networks necessitate advanced optimization strategies, which, in turn, place increasing demands on the accuracy and reliability of path loss (PL) prediction models. However, the inherent complexity of urban environments introduces substantial challenges for effective modeling. In this work, a propagation-aware PL prediction framework is proposed. Fresnel zone theory is used to scientifically define the propagation region, which better captures key scattering structures. The proposed model employs a dual-path learning strategy, integrating expert-engineered statistical features with latent spatial representations extracted via a variational autoencoder (VAE). Furthermore, a cross-attention fusion mechanism is introduced to enable adaptive interaction and effective integration of these heterogeneous features. Validated on a large-scale measurement campaign in dense urban environment, the proposed model demonstrates robust predictive performance across diverse urban blocks and outperforming established baseline models. This research provides a powerful new framework for PL prediction and offers valuable insights for intelligent network planning and optimization in urban IoT systems.
UAV(unmanned aerial vehicle)swarms hold broad application prospects in both military and civilian domains,and are developing rapidly.Communication and networking technologies tailored for swarms are crucial for enabling their formation systems and synergistic effectiveness.To this end,key technologies,research status,and development trends in UAV swarm communication and networking were systematically reviewed.Specifically,typical application scenarios and networking requirements of UAV swarms were analyzed.Key challenges and corresponding solutions were thoroughly discussed at the physical layer,data link layer,and network layer,respectively.Furthermore,leveraging the cyber-physical fusion characteristics of UAV,the deep coupling effects among communication,computation,and control,along with their joint optimization methods,were analyzed and summarized.Regarding intelligent empowerment,the intelligent architecture integrating"intent understanding,environment adaptation,and resource scheduling"as a trinity,alongside the latest research progress,were examined.For engineering applications,critical issues such as high-frequency band communication hardware optimization and antenna lightweighting were explored.The prospects for integrating UAV swarms into future integrated space-air-ground networks and the directions for next-step development were outlined.This review aims to provide references for top-level planning and design,as well as for future research directions regarding key technological breakthroughs in the field of UAV swarm communication and networking.
This paper focuses on the challenge of achieving anti-jamming communication in multi-UAV networks without a common channel. We formulate an optimization problem that maximizes communication utility, which is defined by normalized throughput and synchronization overhead, through the joint optimization of the available channel set (ACS), channel and power. Based on the idea of multi-scale decision, we propose an adaptive ACS decision algorithm and a branching QMIX algorithm (ABQMIX), where cluster head (CH) and cluster members (CMs) are treated as heterogeneous agents. At the time frame scale, CH adaptively selects a high-quality and appropriately sized ACS based on observations of channel status and the evaluation of the communication utility. At the time slot scale, CMs make distributed decisions on channel and power allocation via an agent network enhanced by a branching architecture. Simulation results demonstrate the effectiveness of the adaptive ACS scheme in improving network performance. Moreover, the branching QMIX reduces the action space dimensionality and enhances communication utility.
In 6 G directional communications, high-frequency millimeter-wave (mmWave) and terahertz (THz) signals suffer from severe propagation loss and are highly susceptible to blockage, while the high mobility of the unmanned aerial vehicle (UAV) further complicates link establishment. To address these challenges, this paper proposes a novel UAV beam alignment scheme that leverages multi-modal environment semantics information to improve beam alignment accuracy by integrating multi-modal semantics into highly dynamic UAV environments. Specifically, a dual-path feature extraction module that acquires visual semantics via target detection and positional semantics through GPS/INS fusion is presented. A neural network based on bidirectional cross-attention mechanism is designed to effectively fuse multi-modal environment semantics information to output the optimal beam. Simulation results demonstrate significant improvements in beam alignment accuracy through multi-modal fusion. Compared to existing vision-only and position-only schemes, the proposed scheme achieves increases in Top-1 accuracy of 16.25% and 74.66%, respectively.
Semantic communications, powered by the ability to efficiently extract and transmit crucial semantic features, have shown attractive potential in supporting intelligent tasks and personalized communication. Although existing researches have achieved remarkable performance by leveraging local semantic relationships to reconstruct data, the important role of global semantic information in semantic recovery remains unexplored. In this paper, we introduce a vital global semantic information, topic, into semantic communications to enhance semantic recovery, which plays a crucial role in ensuring contextual consistency and semantic understanding. Specifically, we propose a topic enhanced semantic communication system (TESC) that integrates topic and content representations to ensure reliable communication. First, to realize effective information fusion, we develop two topic-semantic fusion mechanisms that generate a more comprehensive semantic representation. Second, to obtain topic information at the receiver without increasing transmission overhead, we introduce a pre-decoder based topic extractor that pre-decodes the sentences and extracts the corresponding topic embeddings. Third, a two-phase training algorithm has been devised to guarantee the convergence of each network module. Simulation results demonstrate the effectiveness of incorporating topic information into semantic communications. Compared to other benchmarks, the proposed TESC significantly improves communication reliability without increasing the amount of transmitted data.
Training generative models for semantic communication requires vast, distributed datasets, but this process faces critical challenges in distributed settings, including strict user privacy, limited client resources, and complex channel heterogeneity. Most existing works typically deploy the whole models on clients, leading to substantial overhead and limited adaptability to channel variations. To this end, we propose a training framework for the federated generative adversarial latent diffusion models (FGAL-DM). This framework employs an asymmetric architecture with a lightweight encoder on each client and a powerful signal-to-noise ratio (SNR)-aware generative model on the server. To overcome inherent privacy problem, we introduce a two-stage training strategy: supervised pre-training on public data for robust initialization, followed by adversarial fine-tuning in the latent space using a public data reference to avoid accessing private data. Additionally, a layer-wise contribution-aware aggregation mechanism is employed to further enhance training efficiency. Experimental results show that our framework achieves up to a 0.8 dB PSNR gain and a 1.4% LPIPS reduction, which facilitates the construction of a high-fidelity federated semantic communication model in practical scenarios.
In scenarios with extremely harsh channel conditions and severely limited communication resources, the reliability and effectiveness of semantic communication require urgent enhancement to satisfy the increasing demands of 6G technology. To address this issue, we propose an importance prioritized framework that integrates both message importance and feature importance to identify critical semantics for reliable and efficient semantic transmission. Considering service personalization and task intelligence, we analyze the message importance by factoring the receiver’s preferences and the communication tasks requirements. Specifically, a large AI model is introduced to quantify message importance, while an importance-based metric for semantic accuracy is established to evaluate the overall reliability of semantic communication. To safeguard significant messages in harsh channel conditions, an unequal error protection strategy based on message importance is employed. Furthermore, we propose a novel approach for feature importance analysis based on loss variation to accurately identify critical features. A feature importance prediction network is designed for algorithm deployment. Additionally, a semantic compression strategy based on feature importance is utilized to prioritize the transmission of essential features in limited communication resources scenarios. Extensive experimental results demonstrate substantial performance advantages of our framework and methods, especially in low signal-to-noise ratio and communication resource shortages, providing a reliable and efficient solution for semantic communication in adverse communication environments.
Semantic communication (SC) enables bandwidth-efficient wireless image transmission, but most existing SC schemes are user-agnostic and ignore receiver-dependent semantics. To address this issue, we propose a personalized digital semantic communication (PDSC) framework that integrates a vision-language model (VLM)-based semantic encoder with a latent diffusion model (LDM)-based semantic decoder. Specifically, the semantic encoder extracts source-aware personalized semantic tokens from both the source image and the receiver's historical interactions. These tokens are vector-quantized into discrete semantic indices and further encoded into a compact fixed-length bitstream, enabling compatibility with digital transmission. At the receiver, the semantic decoder reconstructs a personalized image conditioned on the recovered semantic tokens. Furthermore, we formulate a capacity-constrained personalized semantic rate-distortion problem and introduce a semantic distortion metric that jointly characterizes source-semantic fidelity and user-preference alignment. Experiments show that PDSC achieves superior source-semantic consistency and personalization over state-of-the-art SC baselines, including CDDM and MoS, under bandwidth-limited wireless transmission.
High-precision channel simulation is critical for modeling and evaluating communication systems. Fading sequences in complex time-varying channels must match both the target amplitude distribution and the Doppler Power Spectral Density (DPSD). However, traditional Sum-of-Sinusoids (SOS) and filtering methods often fail to maintain consistent first- and second-order statistics for non-Gaussian channels. To address this, we propose the Copula-Based Joint Statistical Consistency Channel Simulation (CB-SCS) method. Our innovation utilizes the copula invariance principle to build the simulation framework. We introduce a strictly monotonic homologous mapping that transforms the reference Gaussian process amplitude to an arbitrary target distribution. Crucially, this mapping maximally preserves the original temporal correlation structure. This approach resolves the consistency conflict between amplitude distribution and DPSD without requiring complex iterations. Experimental results demonstrate superior joint statistical consistency. The method achieves an average Kolmogorov-Smirnov (K–S) distance of only 0.012 and a Probability Density Function (PDF) Normalized Mean Square Error (NMSE) of −25.10 dB. Compared to typical baselines, CB-SCS offers higher robustness and provides an efficient solution for simulating complex time-varying channels.
The channel nonlinearity and high-speed mobility in low-earth orbit (LEO) constellation communication are key bottlenecks that constrain the performance of waveform transmission. To address these challenges, we propose a universal continuous phase modulation (CPM) spread-spectrum communication method with high spectral efficiency, and a Doppler-insensitive signal detection mechanism. First, an enhanced dual-phase CPM spread-spectrum (DP-CPM-SS) waveform is investigated. By analyzing the principles and power spectrum characteristics of DP-CPM-SS, we correct the modulation index and design the optimalWiener filtering reception for CPM, improving the spectral efficiency and noise resilience. Subsequently, a CPM noncoherent detection integrating frequency estimation and phase pre-compensation is developed. The performance lower bound of the partial matched filter-fast fourier transform (PMF-FFT) frequency estimation algorithm is derived, revealing the relationship among the length and number of partial matched filters, the normalized frequency offset and the mean square error (MSE) of frequency estimation. Additionally, error probability performance and frequency offset adaptation range are analyzed. Numerical results show that the proposed method exhibits superior and stable performance under severe Doppler effects, which is a promising scheme for LEO constellation communication.
Multi-cluster Unmanned Aerial Vehicle (UAV) networks operating in hostile environments face the intertwined challenges of external malicious jamming and severe inter-cluster co-channel interference. Existing anti-jamming research primarily addresses single-link or isolated-swarm scenarios, overlooking the spatial coupling and resource contention inherent in large-scale multi-cluster deployments. To bridge this gap, we propose a fully distributed spectrum coordination and anti-jamming framework. The core idea is to formulate spatial resource allocation as a local altruistic potential game, enhanced by a dynamic priority metric that jointly captures asymmetric anti-jamming demand and historical delivery satisfaction. We prove that this cooperative mechanism constitutes a generalized exact potential game, guaranteeing finite-step convergence under asynchronous updates. Building on this foundation, we design the Spatial Adaptive Best Response (SABR) algorithm, which achieves O(1) per-node complexity. Extensive simulations demonstrate that SABR attains near-optimal utility, a robust packet delivery ratio, and near-perfect fairness in ultra-dense networks.
Channel prediction is vital to mitigate demodulation degradation due to outdated channel state information (CSI) in fast time-varying channels. However, existing predictors suffer from high computational complexity and degraded accuracy when applied to high-dimensional multiple-input multiple-output (MIMO) channels. This paper proposes ChannelMamba, which leverages the selective state space model (SSM) of Mamba to capture spatial-frequency correlations of channels. The selective SSM adeptly distills hidden patterns within historical time series CSI to forecast future CSI while maintaining near-linear complexity. Most conventional predictors forecast future CSI in a sequential manner, which often leads to error accumulation. To overcome this, we tokenize the time stamp of each CSI through a linear layer as parallel input and employ a lightweight temporal encoding layer to forecast multi-step future CSI in parallel. Meanwhile, the bidirectional Mamba block enables ChannelMamba to capture global correlations across high-dimensional CSI, enhancing accuracy under high-mobility scenarios. Simulation results demonstrate that ChannelMamba outperforms transformer-based counterparts by more than 1.2 dB in normalized mean squared error (NMSE) under 28 GHz urban macrocell scenarios ranging from 30 to 70 km/h.
Adapting deep-learning receivers after deployment is challenging because ground-truth channel state and transmitted payload symbols are unavailable, while sparse pilots alone provide insufficient supervision. The orthogonal frequency-division multiplexing waveform, however, embeds physical structure that can provide label-free supervision through physics-informed constraints. Here, we advance Physics-Embedded Inverse Learning (PEIL) toward deployment in two phases. Phase I develops Improved PEIL, a compact receiver with 0.093M trainable parameters that coordinates carrier frequency offset estimation and channel state information refinement through symbol-recovery supervision. In Phase II, Physics-Embedded Label-free Adaptation (PELA) adapts the same receiver using pilot reconstruction together with cyclic-prefix phase consistency and delay–symbol smoothness of the channel impulse response tail. Trained through 16-QAM symbol recovery, Improved PEIL generalizes without retraining to unseen QPSK and 64-QAM payloads. Under the EVA-to-ETU transition at 35 dB, PELA reduces symbol error rate from 0.1501 to 0.0258 and block error rate from 0.7106 to 0.0591 relative to static Improved PEIL. Across the evaluated mobility shifts, PELA becomes beneficial when deployment conditions differ substantially from the training environment. These results show that protocol-native pilots and waveform physics can supply the missing supervision for selective post-deployment adaptation.
Semantic communication (SemCom) is expected to play a key role in 6G and digital twin network (DTN) environments, where reliability depends on preserving meaning and user intent rather than reconstructing bit sequences. This paper presents a GenAI-driven, contextual-knowledge-enhanced SemCom framework that strengthens semantic recovery by integrating relational, intra-sentential, inter-sentential, and historical knowledge. At the transmitter, a context-attentive (CA) mechanism combines sentence-level semantic features with relational knowledge generated by a relation-guided knowledge base. At the receiver, a knowledge enhancement (KE) mechanism employs Transformer-XL to incorporate historical semantic context and utilizes a relational knowledge modeling module to refine semantic inference. Experimental results demonstrate that the proposed framework achieves robust semantic recovery across a wide range of signal-to-noise ratios (SNRs), with significant improvements in low-SNR conditions. Ablation studies further highlight the complementary roles of contextual knowledge at the transmitter and receiver. These results confirm the effectiveness of contextual knowledge in enhancing deep learning-based SemCom and underscore its potential for deployment within future DTN-enabled 6G architectures.
By focusing on the intrinsic meaning of information, semantic communication (SemCom) marks a fundamental paradigm shift from physical bit transmission to personalized semantic service. Considering the importance of personalized features related to the speaker in speech for source recovery and understanding, we propose a semantic-driven framework for personalized speech transmission, named PerSemCom, which combines speaker acoustic features with semantic information. Specifically, we first introduce an efficient semantic extraction mechanism to achieve the conversion from speech to text transcriptions, and design a semantic corrector coupled with multi-domain knowledge to mitigate the effects of wireless channel distortion. Building upon the reliable transcriptions at receiver, we further establish a speaker embedding vector knowledge base and achieve high-fidelity speech reconstruction through quantitative modeling of speaker-specific acoustic features. Extensive experimental results demonstrate that our proposed framework outperforms existing schemes in terms of subjective perception at harsh channel conditions. Complexity analysis and latency measurements also show competitive advantages in computational efficiency and real-time capabilities. Reconstructed personalized speech samples have been publicly available at https://kwtankw.github.io/PerSemCom/.
The existing backoff mechanism of the SPMA(statistic priority-based multiple access)protocol relies on static function models and has single dimension of optimization parameters,making it unable to adapt to dynamic transmission and multi-priority requirements in UAV ad hoc networks.To address this issue,the dynamic decision-making process of node selection for backoff time in the SPMA protocol was modeled as a Markov decision process,and an intelligent backoff strategy based on the DDQN(double deep Q-network)was innovatively proposed.This strategy comprehensively consideres factors such as service priority,thresholds,and channel load,and adopts the DDQN algorithm to select backoff time within a finite and discrete action space.Simulation results show that,compared to traditional binary exponential backoff strategies and logarithmic function-based backoff strategies,the proposed strategy can reduce the transmission delay for low-priority services by up to 33.3%,increase the initial backoff success rate by 18%,and effectively improve the transmission success rate and adapt to the variation of network scale well.
Learning robust speaker representations under noisy conditions presents significant challenges, which requires careful handling of both discriminative and noise-invariant properties. In this work, we proposed an anchor-based stage-wise learning strategy for robust speaker representation learning. Specifically, our approach begins by training a base model to establish discriminative speaker boundaries, and then extract anchor embeddings from this model as stable references. Finally, a copy of the base model is fine-tuned on noisy inputs, regularized by enforcing proximity to their corresponding fixed anchor embeddings to preserve speaker identity under distortion. Experimental results suggest that this strategy offers advantages over conventional joint optimization, particularly in maintaining discrimination while improving noise robustness. The proposed method demonstrates consistent improvements across various noise conditions, potentially due to its ability to handle boundary stabilization and variation suppression separately.
The advancement of communication technologies has imposed increasingly stringent requirements on the precision and generalization capabilities of channel models. Traditional statistical and deterministic models frequently encounter challenges in achieving an optimal balance between these two aspects. To address this challenge, this letter introduces a channel prediction model based on deep learning and Fresnel zone propagation theory to achieve accurate prediction of path loss (PL) and delay spread (DS). To improve the extraction of environmental information from environmental images for channel prediction, an enhanced image environmental feature representation method is proposed. This method adaptively adjusts the dimensions of the input image based on the distance between the transmitter and receiver. Furthermore, an image importance distribution map is defined to delineate the variations in significance across the different regions of the environmental image. Subsequently, a model architecture capable of processing images of variable sizes is developed utilizing spatial pyramid pooling techniques. Finally, using channel measurement data obtained at 5.9 GHz in Changsha, the performance of the proposed model is validated employing a leave-one-out cross-validation method. Compared to existing models, the proposed prediction model achieves a reduction in the average root-mean-square error for PL and DS prediction by approximately 12.78% and 8.2%, respectively.
Multi-numerology transmissions give rise to challenges in terms of high peak-to-average power ratio (PAPR) and inter-numerology interference (INI) arising from the coexistence of non-orthogonal numerologies. Because PAPR reduction and INI mitigation interact in signal processing, we jointly optimize them to improve performance over separate processing. The optimization problem is formulated based on the injection of cancellation signals within the guard band, where these signals are optimized for each numerology to mitigate INI and then superposed to reduce signal peaks. Since different numerologies are modulated over separate subbands, the joint optimization problem is formulated as a separable problem that can be decomposed into several subproblems. Building on this, we develop an iterative alternating-optimization procedure and then unroll it into a layer-wise network with trainable parameters. Within the network, both the penalty factors and PAPR threshold can be trained via data-driven. This design enables faster convergence with reduced unfolding depth, and the soft-threshold method can adaptively achieve suboptimal solutions when strict PAPR thresholds preclude feasible sets. Theoretical analysis shows that the proposed alternating-optimization procedure converges to a Karush–Kuhn–Tucker (KKT) point of the joint optimization problem. Simulation results demonstrate that the proposed method effectively improves both PAPR and INI performance, while maintaining scalability for multi-numerology cases with two or more numerologies.