Affine frequency division multiplexing (AFDM) is a recently proposed multicarrier waveform whose bit error rate (BER) performance in doubly selective channels is comparable to that of orthogonal time-frequency space (OTFS) and superior to that of orthogonal frequency division multiplexing (OFDM). In this paper, the impacts of joint transmitter (Tx) and receiver (Rx) in-phase and quadrature imbalance (IQI) on AFDM signals are investigated, where we show that AFDM suffers more severe IQI than OFDM and OTFS due to the inherent feature of complicated chirp-assisted modulation. We further derive analytical expressions for the pairwise and average bit error probability as a function of the IQI parameters. These indicate that such distortions significantly limit the achievable operating signal-to-noise ratio at the receiver side and data rates. To this end, we propose a cascade compensation scheme to mitigate these effects. Specifically, we first compensate for Rx IQI to convert the improper Gaussian noise into additive white Gaussian noise, and then apply a judicious design to eliminate the Tx IQI. Both analytical and simulation results reveal that joint Tx and Rx IQI introduce an error floor in the BER performance of AFDM systems, whereas the proposed approach effectively compensates such impairments.
Index modulation (IM) techniques have been widely studied over the past decade for their ability to enhance spectral and energy efficiency by exploiting additional degrees of freedom (DoF) in waveforms. In this paper, we propose a novel subcarrier filtering orthogonal frequency-division multiplexing (OFDM) scheme, named waveform index modulation (WIM), to boost the spectral efficiency (SE) of OFDM systems without compromising other performance metrics. In WIM-OFDM, information is conveyed not only by the modulated constellation symbols but also by altering the subcarrier filter shapes, thereby utilizing an additional DoF in the OFDM signaling process. An SE-enhanced version, referred to as generalized WIM-OFDM (GWIM-OFDM), is also designed to further boost the index transmission rate by maximizing the DoF for filter selection on each subcarrier. Additionally, the optimization of subcarrier filter shapes is formulated, and a special case of subcarrier filter pair can be optimized by utilizing the proposed non-convex to convex scaling method. At the receiver, a low-complexity interference cancellation algorithm is proposed to eliminate the introduced inter-carrier interference caused by the non-orthogonal subcarrier shapes. Finally, to validate the proposed scheme, closed-form expressions for the achievable rates and the upper bound on the average bit error rate are derived to prove the superiority of our WIM-OFDM and GWIM-OFDM schemes theoretically. Monte Carlo simulation results corroborate the benefits of the proposed scheme, that is, the GWIM-OFDM scheme exhibits 4.7 to 6.1 dB performance gain, considering both bit error rate and peak-to-average power ratio, compared with the traditional OFDM and other IM benchmarking schemes under the same spectrum mask at different transmission rate scenarios.
In affine frequency division multiplexing (AFDM) systems, the severe peak-to-average power ratio (PAPR) signals exist in the time domain due to the coherent superposition of numerous modulated symbols. Eventually, high PAPR signals require sophisticated and expensive power amplifiers with a very large linear range. To this end, a neural network (NN) aided intelligent transceiver optimization framework is proposed for suppressing PAPR based on the spreading AFDM structure. Specifically, the transceiver jointly optimizes the constellation geometry and associated bit labeling, the precoding NN, as well as the NN based detector. Moreover, the precoding NN is learned from a precoding approach which minimizes the variance of the instantaneous power of output signals at the transmitter. The joint optimization framework aims to achieve maximum PAPR reduction under the constraints of unit energy and spectral emission mask. Besides, to mitigate the potential inter-carrier interference during the offline training, a long short term memory based detector is designed within the optimization framework. Simulation results demonstrate that the conceived NN based optimization method achieves a significant enhancement on PAPR reduction compared with conventional approaches, while slightly improving the bit error ratio performance.
Affine frequency division multiplexing (AFDM) achieves full diversity but faces multiple-access challenges due to signal dispersion. To address this issue, we propose a power domain non-orthogonal multiple access AFDM (PD-NOMA-AFDM) system, which enables parallel transmission of multi-user signals on the same resource block through power-domain multiplexing. Furthermore, we design a binary-loop maximal ratio combining-message passing (BLMM)-based successive interference cancellation (SIC) scheme. Specifically, the inner loop fully leverages the sparsity of the AFDM equivalent channel to effectively eliminate inter-symbol interference and achieve reliable initial symbol estimation; the outer loop iteratively updates extrinsic information to compensate for performance degradation caused by banded-matrix approximation. We prove the convergence of the inner loop to the MMSE fixed point and the local convergence of the outer loop. Subsequently, by combining Lipschitz continuity and perturbation theory, we demonstrate the convergence of the overall BLMM detector to a neighborhood of the exact fixed point. The pairwise error probability analysis is then used to characterize its diversity gain and performance gap to maximum likelihood (ML) detection. To further narrow this gap, a graph neural network (GNN) is incorporated into the BLMM multi-user detection framework. This approach dynamically captures the multi-user interference (MUI) characteristics through node message interactions, thereby improving the accuracy of the approximate a posteriori probability distribution. Simulation results show that the proposed BLMM-GNN achieves near-ML performance with strong robustness.
Affine frequency division multiplexing (AFDM) exhibits excellent Doppler robustness and the ability to characterize doubly selective channels. However, its signal dispersion characteristics make it challenging to directly adopt traditional time-frequency multiple access schemes. To address this issue, we introduce cooperative rate splitting multiple access (RSMA) for AFDM systems. The flexible configuration of AFDM chirp parameters can reduce the correlation between users' equivalent channels, which decreases the interference from RSMA private streams. We conduct a theoretical analysis of the cooperative RSMA-AFDM system and demonstrate that minimizing the overlap in the channel column spaces among users can effectively enhance the system performance. Guided by this analysis, we design a chirp parameter optimization scheme that reduces multi-user interference and maximizes diversity gain. To fully exploit the diversity gain brought by the proposed chirp parameter optimization, two expectation propagation (EP)-based distributed cooperative detection schemes are proposed. First, a decision-fusion-based method is developed, where local information and cooperative information are fused by maximum ratio combining, achieving a globally consistent estimate of the common stream. Second, we develop a belief-consensus EP-based detection scheme. In each iteration, user nodes exchange and fuse the first- and second-order statistics of the common stream, and the resulting beliefs gradually converge to a consistent global decision, which significantly improves the overall reliability.
Affine frequency division multiplexing (AFDM) has emerged as a promising technology for high-mobility scenarios, offering reliable performance in doubly selective channels. However, the computational complexity of the maximum likelihood (ML) detection scheme renders it impractical for real-time AFDM applications. To address this, we propose a low-complexity AFDM symbol detection algorithm based on expectation propagation (EP) in this paper. The proposed EP-based detection scheme iteratively updates messages to approximate the ML result, reducing computational complexity from exponential to cubic order. By exploiting the sparse and quasi-banded structure of the channel in the discrete affine Fourier transform (DAFT) domain and employing matrix block decomposition, lower-upper factorization, and upper triangular matrix forward substitution, we further reduce the complexity of the EP algorithm to linear order. Additionally, we optimize the EP algorithm’s performance by incorporating deep learning-based moment matching, making the algorithm more adaptive with trainable parameters for both positive and negative components. Moreover, we propose a DAFT-domain iterative detection and decoding scheme, where external information from the decoder is fed back to the detector, resulting in improved system reliability. Simulation results show that the proposed scheme achieves near-ML performance while reducing complexity by dozens of orders of magnitude compared to the ML detector, striking a balance between performance enhancement and computational complexity.
Orthogonal time frequency space (OTFS) modulation has emerged as a promising candidate to overcome the performance degradation of orthogonal frequency division multiplexing (OFDM), which are commonly encountered in high-mobility wireless communication scenarios. However, conventional OTFS transceivers rely on multiple separately designed signal-processing modules, whose isolated optimization often limits global optimal performance. To overcome limitations, this paper proposes a modular deep learning (DL) based end-to-end OTFS transceiver framework that consists of trainable and interchangeable neural network (NN) modules, including constellation mapping/demapping, superimposed pilot placement, inverse Zak (IZak)/Zak transforms, and a U-Net-enhanced NN tailored for joint channel estimation and detection (JCED), while explicitly accounting for the impact of the cyclic prefix. This physics-informed modular architecture provides flexibility for integration with conventional OTFS systems and adaptability to different communication configurations. Simulations demonstrate that the proposed design significantly outperforms baseline methods in terms of both normalized mean squared error (NMSE) and detection reliability, maintaining robustness under integer and fractional Doppler conditions. The results highlight the potential of DL-based end-to-end optimization to enable practical and high-performance OTFS transceivers for next-generation high-mobility networks.
In this paper, we propose SpectralNeRF, an end-to-end Neural Radiance Field (NeRF)-based architecture for high-quality physically based rendering from a novel spectral perspective. We modify the classical spectral rendering into two main steps, 1) the generation of a series of spectrum maps spanning different wavelengths, 2) the combination of these spectrum maps for the RGB output. Our SpectralNeRF follows these two steps through the proposed multi-layer perceptron (MLP)-based architecture (SpectralMLP) and Spectrum Attention UNet (SAUNet). Given the ray origin and the ray direction, the SpectralMLP constructs the spectral radiance field to obtain spectrum maps of novel views, which are then sent to the SAUNet to produce RGB images of white-light illumination. Applying NeRF to build up the spectral rendering is a more physically-based way from the perspective of ray-tracing. Further, the spectral radiance fields decompose difficult scenes and improve the performance of NeRF-based methods. Comprehensive experimental results demonstrate the proposed SpectralNeRF is superior to recent NeRF-based methods when synthesizing new views on synthetic and real datasets. The codes and datasets are available at https://github.com/liru0126/SpectralNeRF.
RAW-to-sRGB mapping, or the simulation of the traditional camera image signal processor (ISP), aims to generate DSLR-quality sRGB images from RAW data captured by smartphone sensors. Despite achieving comparable results to sophisticated handcrafted camera ISP solutions, existing CNN-based methods still suffer from detail disparity and color distortion due to their inherent locality restrictions. In this paper, we present ISPFormer, a novel Transformer-based framework utilizing self-attention to tackle the learnable ISP problem. Specifically, we propose a Wavelet-based Transformer Block (WTB) with two loss functions to correct color and enhance high-frequency details, where WTB integrates wavelet transformation into window-based self-attention to perform self-attention in sub-bands across frequency domains, enabling larger sliding window modeling without additional computational overhead. Based on the proposed WTB, we construct the ISPFormer as a multi-stage network, where a multi-scale window adjustment strategy is further proposed to flexibly assign varying window sizes for each stage, reconstructing visually satisfactory results in a coarse-to-fine manner. Extensive experiments demonstrate that our ISPFormer achieves competitive quantitative and qualitative results. Code is available at https://github.com/RenYangSCU/ISPFormer.
Adaptive bit-loading algorithms typically rely on greedy iterative searches and are effective primarily in orthogonal frequency-division multiplexing (OFDM) systems. Nonorthogonal multicarrier waveforms offer higher spectral efficiency but suffer from inter-carrier interference, making traditional iterative methods impractical. In this study, we propose integrating bit-loading with a neural network (NN)-based non-orthogonal multicarrier system. Initially, an NN-based transceiver structure is proposed, which is trained end-to-end to match the channel characteristics and increase the transmission rate while maintaining a low bit-error-rate (BER). To mitigate the high latency associated with iterative searches, we further introduce a bitloading policy based on deep reinforcement learning (DRL). The DRL agent is trained offline to maximize throughput subject to a predefined BER constraint, allowing it to generate the optimal bit allocation map with just a single forward inference during deployment. Simulation results demonstrate that the proposed approach achieves up to a 66% throughput improvement compared to a baseline greedy bit-loading scheme employing OFDM.
Robots can acquire complex manipulation skills by learning policies from expert demonstrations, which is often known as vision-based imitation learning. Generating policies based on diffusion and flow matching models has been shown to be effective, particularly in robotic manipulation tasks. However, recursion-based approaches are inference inefficient in working from noise distributions to policy distributions, posing a challenging trade-off between efficiency and quality. This motivates us to propose FlowPolicy, a novel framework for fast policy generation based on consistency flow matching and 3D vision. Our approach refines the flow dynamics by normalizing the self-consistency of the velocity field, enabling the model to derive task execution policies in a single inference step. Specifically, FlowPolicy conditions on the observed 3D point cloud, where consistency flow matching directly defines straight-line flows from different time states to the same action space, while simultaneously constraining their velocity values, that is, we approximate the trajectories from noise to robot actions by normalizing the self-consistency of the velocity field within the action space, thus improving the inference efficiency. We validate the effectiveness of FlowPolicy in Adroit and Metaworld, demonstrating a 7× increase in inference speed while maintaining competitive average success rates compared to state-of-the-art methods.
This paper proposes a novel joint transceiver optimization framework for multi-carrier (MC) waveform design. Unlike conventional orthogonal frequency division multiplexing, which employs memoryless modulation and fixed inverse discrete Fourier transform-based waveform generation, our approach utilizes neural network (NN)-based modulation with memory and NN-driven waveform generation at the transmitter. On the receiver side, a large-kernel attention-based NN replaces the traditional demodulation process, effectively mitigating large-span inter-carrier interference. This architecture provides enhanced flexibility for MC waveform optimization, allowing better adaptation to spectral emission mask constraints and maximizing the utilization of allocated spectrum resources. Additionally, it achieves significant spectral efficiency gains across diverse channel conditions, including additive white Gaussian noise (AWGN) and linear time-varying (LTV) channels with delay and Doppler spread. Numerical evaluations demonstrate significant bit error rate performance improvements, with up to 10 dB signal-to-noise ratio gain in LTV channels and approximately 6 dB gain in AWGN channels, underscoring the superiority of the proposed framework over state-of-the-art schemes.
In this paper, we address the key challenges of faster-than-Nyquist signaling (FTNS) in doubly selective fading channels (DSCs) by proposing an end-to-end architecture that jointly optimizes FTN waveform, data detector, and channel estimator. Unlike traditional model-driven approaches, our solution introduces a data-driven optimization strategy, enhancing its ability to optimize parameters on a larger scale. In the proposed architecture, constellation geometry, precoder, and pilot sequences are considered as optimizable variables. Specifically, with the goal of maximizing spectral efficiency (SE), the constellation geometry and precoder are co-trained with a neural network (NN)-based data detector, showing superior performance in mitigating inter-symbol interference (ISI) caused by FTNS and channel spreading. Concurrently, the pilot sequences are optimized alongside the channel estimator to improve the estimation accuracy of CIR at pilot blocks. To track rapidly time-varying CIRs, we apply an attention mechanism to the channel estimator that efficiently leverages the channel's temporal correlation through dynamically assigning weights to the pilot blocks to accurately predict the CIRs for each data block. Simulation results show that our scheme achieves a signal to noise ratio gain of more than 6 dB under a maximum Doppler shift of 10 kHz compared with the state-of-the-art linear pre-equalized FTN system with 64-ary quadrature amplitude modulation.
Filter Shapes Index Modulation (FSIM) is a novel single-carrier waveform scheme with increased spectral efficiency by utilizing different pulse-shaping filters to implicitly convey extra index bits. However, except for introducing inter-symbol interference (ISI), the high peak-to-average power ratio (PAPR) of the FSIM waveform, which can significantly impact the power efficiency of the FSIM system and counteract the achieved signal-to-noise ratio (SNR) gains, has not been investigated thus far. In this paper, we analyze the PAPR performance of the FSIM signaling and propose a PAPR reduction method based on filter bank optimization. And since the ISI constraint is relaxed in the process of optimization, an enhanced ISI reconstruction and cancellation scheme is also proposed to address the extra introduced ISI. Simulation results exhibit that, compared to the traditional single-carrier system, the proposed schemes reduce the PAPR of the FSIM signal and achieve 0.45 to 2 dB net SNR gains at different transmission rates under the multipath fading channel.
To fully obtain the time-frequency diversity gain of the affine frequency division multiplexing (AFDM) system, detection algorithms that offer high performance and low complexity are essential. The maximum likelihood (ML) algorithm can achieve theoretically optimal performance. However, the exponential complexity limits its practical application. This paper designs an AFDM signal detection algorithm based on expectation propagation (EP). The proposed EP-based scheme achieves effective AFDM signal detection by iteratively updating messages to approximate the true a posterior distribution. In addition, this paper further reduces the complexity of the proposed EP-based algorithm by utilizing the characteristics of the discrete affine Fourier transform (DAFT) domain equivalent channel. Specifically, the sparsity and quasi-banded structure of the DAFT domain channel are first utilized for block processing. Subsequently, a low complexity matrix inversion operation is realized by combining the lower-upper (LU) factorization and the upper triangular matrix forward substitution algorithm. With typical AFDM system parameters, the proposed scheme reduces the complexity by 35.6 times compared to the traditional EP algorithm, while the performance is virtually unaffected. Simulation results show that the proposed scheme has a performance gain of up to 5 dB over the conventional algorithm.
In this paper, we propose a novel waveform index modulation orthogonal frequency-division multiplexing (WIM-OFDM) scheme to increase spectral efficiency for multicarrier systems. More specifically, the proposed WIM-OFDM scheme conveys not only the classic constellation symbols but also extra index bits by changing the subcarrier filtering shape for each symbol. As a key point, the optimization of the subcarrier filter shapes is formulated and a preliminary subcarrier filter pair is given to verify the performance of the proposed scheme. Our simulation results demonstrate that the proposed WIM-OFDM scheme exhibits superior performance in both peak-to-average power ratio and bit error ratio compared to conventional OFDM-IM and its dual-mode counterparts without increasing the out-of-band emission.
For fine-grained recognition, capturing distinguishable features and effectively utilizing local information play a key role, since the objects of recognition exhibit subtle differences in different subcategories. Finding subtle differences between subclasses is not straightforward. To address this problem, we propose a weakly supervised fine-grained classification network model with Local Diversity Guidance (LDGNet). We designed a Multi-Attention Semantic Fusion Module (MASF) to build multi-layer attention maps and channel–spatial interaction, which can effectively enhance the semantic representation of the attention maps. We also introduce a random selection strategy (RSS) that forces the network to learn more comprehensive and detailed information and more local features from the attention map by designing three feature extraction operations. Finally, both the attention map obtained by RSS and the feature map are employed for prediction through a fully connected layer. At the same time, a dataset of ancient towers is established, and our method is applied to ancient building recognition for practical applications of fine-grained image classification tasks in natural scenes. Extensive experiments conducted on four fine-grained datasets and explainable visualization demonstrate that the LDGNet can effectively enhance discriminative region localization and detailed feature acquisition for fine-grained objects, achieving competitive performance over other state-of-the-art algorithms.
Affine frequency division multiplexing (AFDM) is a promising waveform for future high mobility wireless communication, modulating as chirps-based multicarrier to resist time-varying channels. This paper firstly analyzes the peak-to-average power ratio (PAPR) performance of AFDM signal and the nonlinear effect caused by power amplifier (PA). The reason for nonlinearity occurred is that AFDM signal generates a high PAPR level resulting in PA working at saturation point. The nonlinear distortion deteriorates target performance in a large degree. To this end, we introduce an equalization approach consisting of maximal-ratio combining (MRC) and super-resolution network to tackle the distortion challenges. Specifically, the MRC algorithm is used to exploit the sparse representation of communication channel and the proposed retrieval network (RE-Net) based super-resolution model alleviates the nonlinearity induced by PA. Simulation results reveal the designed equalization scheme executes a pronounced performance promotion compared with conventional solutions.
Most current point cloud super-resolution reconstruction requires huge calculations and has low accuracy when facing large outdoor scenes; a Dense Feature Pyramid Network (DenseFPNet) is proposed for the feature-level fusion of images with low-resolution point clouds to generate higher-resolution point clouds, which can be utilized to solve the problem of the super-resolution reconstruction of 3D point clouds by turning it into a 2D depth map complementation problem, which can reduce the time and complexity of obtaining high-resolution point clouds only by LiDAR. The network first utilizes an image-guided feature extraction network based on RGBD-DenseNet as an encoder to extract multi-scale features, followed by an upsampling block as a decoder to gradually recover the size and details of the feature map. Additionally, the network connects the corresponding layers of the encoder and decoder through pyramid connections. Finally, experiments are conducted on the KITTI deep complementation dataset, and the network performs well in various metrics compared to other networks. It improves the RMSE by 17.71%, 16.60%, 7.11%, and 4.68% compared to the CSPD, Spade-RGBsD, Sparse-to-Dense, and GAENET.
This paper proposes a point-by-point weighted fusion algorithm based on an improved random sample consensus (RANSAC) and inverse distance weighting to address the issue of low-resolution point cloud data obtained from light detection and ranging (LiDAR) sensors and single technologies. By fusing low-resolution point clouds with higher-resolution point clouds at the data level, the algorithm generates high-resolution point clouds, achieving the super-resolution reconstruction of lidar point clouds. This method effectively reduces noise in the higher-resolution point clouds while preserving the structure of the low-resolution point clouds, ensuring that the semantic information of the generated high-resolution point clouds remains consistent with that of the low-resolution point clouds. Specifically, the algorithm constructs a K-d tree using the low-resolution point cloud to perform a nearest neighbor search, establishing the correspondence between the low-resolution and higher-resolution point clouds. Next, the improved RANSAC algorithm is employed for point cloud alignment, and inverse distance weighting is used for point-by-point weighted fusion, ultimately yielding the high-resolution point cloud. The experimental results demonstrate that the proposed point cloud super-resolution reconstruction method outperforms other methods across various metrics. Notably, it reduces the Chamfer Distance (CD) metric by 0.49 and 0.29 and improves the Precision metric by 7.75% and 4.47%, respectively, compared to two other methods.