
Successive cancellation (SC) decoding of polar codes suffers from high decoding latency due to its sequential nature. Fast-SC decoding alleviates this by identifying special nodes that enable concurrent multi-bit decoding. Recently, a more general special node for high-rate polar codes has been introduced, offering reduced decoding latency but at the cost of substantial memory overhead for storing precomputed flipping sets. Alternative flipping set generation methods eliminate memory overhead but rely on iterative procedures, undermining decoding latency benefits. This paper proposes a reduced-memory fast-SC (RMFSC) decoding method based on the inherent parity constraints of sequence single-parity-check (SSPC) nodes, without the need for flipping set storage. Simulation results and complexity analysis show that the proposed RM-FSC reduces memory requirements by up to $93.9 \%$ compared to the prior arts, while preserving low latency and incurring only negligible performance degradation.
This paper introduces a quantization-aware optimization framework for MIMO detection that jointly optimizes algorithmic parameters and fixed-point quantization across layers in Breadth-First Search Detection (BFSD). Unlike conventional approaches that separate algorithm design from hardware implementation, our framework integrates these perspectives to achieve superior performance-complexity trade-offs. Experimental results on an $8 \times 8$ MIMO system with 16-QAM modulation demonstrate up to $37.7 \%$ reduction in computational complexity while maintaining BER performance comparable to floating-point implementations, with notable advantages at high SNR regions. The proposed framework is algorithm-agnostic, supporting various optimization algorithms, and can be readily extended to other detection problems requiring hardware-efficient implementations.
Automorphism ensemble decoding with the successive cancellation constituent decoder (AED-SC) can improve the decoding performance of the successive cancellation decoder when decoding polar codes with short-to-medium code lengths. To reduce the number of automorphisms used by the AEDSC, simplified repetition handling early-stopping is proposed in this work. Compared to the original repetition handling earlystopping, our simplified version compares the path metric instead of the codeword. Up to a $2 \times$ reduction in the number of automorphisms used is observed when decoding polar codes with different code lengths and dimensions while having little degraded decoding performance. Also, the proposed simplification reduces the number of bit comparisons from a number equal to the code length to the number of bits required to represent the path metric.
This paper proposes an end-to-end radio simultaneous localization and mapping (SLAM) algorithm that directly leverages channel impulse response (CIR) to overcome fundamental limitations in existing approaches. Traditional radio SLAM algorithms assume pre-estimated channel parameters, making performance highly sensitive to estimation accuracy, while recent end-to-end methods jointly perform parameter estimation and SLAM but suffer from high computational complexity and model mismatch vulnerability. The proposed algorithm minimizes information loss by operating directly on raw CIR measurements and utilizes end-to-end learning for enhanced robustness. Simulation results in the 3GPP TR 38.857 indoor factory scenario demonstrate that the proposed algorithm achieves comparable performance to conventional radio SLAM while reducing computational time by less than $2 \%$, confirming its strong potential for practical deployment.
With 5G/6G development, sophisticated scenarios and diverse requirements challenge decoding choices in MultipleInput Multiple-Output (MIMO) system. Traditional methods struggle to rapidly identify optimal algorithms tailored to unique requirements, as parameters vary drastically across scenarios. To tackle those challenges, we propose a Transformer-based intelligent algorithm prediction method, introducing TokenMIMO-a MIMO algorithm classifier for complex scenarios to efficiently predict optimal decoding algorithms across algorithm prediction spaces. It achieves an accuracy of over 96%, outperforms state-of-the-art modules. Its exceptional generalization capability enables it to recommend optimal decoding algorithms to users with a policy distribution exceeding $94 \%$ accuracy within an extremely large algorithm selection space.
A tensorial Hankel reconstruction method for underdetermined direction of arrival (DOA) estimation with multifrequency sparse arrays is proposed in this paper. Firstly, virtual arrays are extended into a uniform linear array (ULA) across all frequencies and structured as a 3-D tensor. Then, it is transformed into a 4-D Hankel representation via spatial dimension augmentation, where missing elements are dispersed while dimensional information is enriched. After low-rank tensor completion, the restored 4-D Hankel tensor is inversely mapped to the 3-D space for DOA estimation. Simulation results are provided to verify the effectiveness of the proposed method in handling severe data loss from sensor failures.
Human-motion energy harvesting is a promising solution for wearable electronics and devices, offering a sustainable power source that extends operational longevity and enhances durability. However, current techniques and prototypes have yet to achieve fully interactive, battery-free functionality. This paper presents a battery-free interactive gaming system tailored for geriatric rehabilitation, powered solely by energy harvested from transient fingertip motion. To ensure reactivity and stability under intermittent power, we employ a multistable fingertip motion harvester (FMH) and a non-volatile interaction design. The FMH leverages dynamically varying potential wells to provide reliable energy, while non-volatile memory and checkpointing enable continuous computation and seamless gaming, even with frequent power interruptions. Beyond advancing fundamental research, this work demonstrates a practical paradigm for battery-free, motion-powered interactive gaming, highlighting its potential in engaging and sustainable rehabilitation for elderly users.
Guessing Random Additive Noise Decoding (GRAND) algorithms can decode any moderate redundancy code of any structure, with integrated circuits demonstrating that it is possible to execute GRAND efficiently in hardware. By using GRAND as a component decoder, its remit has been extended to efficiently decode long, high redundancy product codes via the Iterative GRAND (IGRAND) algorithm in the hard detection setting. Here, by leveraging the property of even codebooks to split the noise effect search space into even and odd Hamming weights for the rows and columns in product code decoding, we demonstrate significant reductions in the number of queries IGRAND makes without any degradation in decoding performance. For IGRAND, we establish that the reduction in query numbers can be significantly higher than for the component code alone, resulting in significant efficiency gains. In particular, simulation results show that the average number of guesses reduces by a factor of 4 at a bit flip probability of $10^{-1.5}$ when the CRC $(64,53)^{2}$ codebook is used, without impacting on its decoding accuracy.
This paper considers the design of phase-only beampatterns with minimum peak sidelobe level (PSL). Unlike existing approaches that assume fixed array element positions, we attempt to achieve a lower PSL by jointly optimizing the element positions and weights. However, the resulting problem is nonconvex and difficult to handle. In order to deal with the problem, a computationally efficient algorithm is developed by leveraging the alternate direction method of multipliers (ADMM) framework, which decomposes the problem into several tractable subproblems. Numerical results verify that the proposed method achieves improved performance compared to other methods.
The modulo analog-to-digital converter (ADC) offers a promising solution to address the dynamic range (DR) limitation of conventional ADCs and provide resolution enhancement given a fixed quantization bit budget. However, the distortion due to the modulo operation necessitates an unfolding mechanism for accurate signal reconstruction. This paper presents a fast method for unfolding the modulo ADC output using compressed sensing and sliding discrete Fourier Transform (DFT). More precisely, we show that the first-order difference of the modulo residue samples is sparse under a specific choice of the folding threshold. Using this sparsity result, the modulo residue signal can be recovered from its out-of-band DFT measurements by formulating a sparse recovery problem. Unlike existing DFT-based reconstruction methods for modulo ADCs, the proposed sliding DFT-based approach uses shorter observation windows for faster unfolding. In addition, the proposed algorithm works with modulo ADCs without additional folding signal. Our numerical results demonstrate that low-resolution modulo ADCs equipped with our proposed recovery method can achieve lower mean squared error (MSE) than conventional ADCs without modulo.
Large-scale multiple-input multiple-output (LSMIMO) is crucial for enhancing capacity and performance in fifth generation (5 G) and beyond. Achieving high-performance data detection with high-order modulation in LS-MIMO systems poses a significant challenge due to complex posterior distributions. While recent diffusion-based detection approaches like the annealed Langevin dynamics (ALD) show promise, their high computational cost resulting from extensive sampling iterations limits their real-time applicability. To address this limitation and further improve detection accuracy, we propose an expectation propagation-guided diffusion (EPGD) detector for LS-MIMO. EPGD leverages a novel conditional diffusion framework, integrating our proposed expectation propagation-guided data prediction model and initialization method. Experimental results demonstrate that EPGD significantly reduces the symbol error rate (SER) compared with ALD, concurrently achieving a substantial computational complexity reduction of up to $84 \%$ in $32 \times 32$ MIMO systems. Furthermore, the SER performance of EPGD also outperforms that of other baseline detectors.
In multiple-input multiple-output (MIMO) systems, the hybrid analog-digital structure (HADS) offers an effective solution to reduce transmission loss and power consumption. However, practical scenarios involving HADS in MIMO systems face challenges such as array mutual coupling effects and limited available snapshots (comparable to the number of sensors). To address these issues, this paper proposes a direction-of-arrival (DOA) estimation method based on Toeplitz rectification (TR). First, the spatial covariance matrix (SCM) is reconstructed by adjusting the switching states and phase shift vectors of the reconfigurable phase shifters. Leveraging the structural characteristics of mutual coupling, the mutual coupling effects are mitigated using a middle subarray strategy. For low signal-tonoise ratio (SNR) scenarios, DOA estimation is performed by integrating TR to enhance robustness. In high-SNR regimes where TR alone underperforms, an empirical threshold is determined through extensive experiments to adaptively activate the MUSIC algorithm for refined estimation. Numerical simulations validate the effectiveness of the proposed method, demonstrating significant improvements in estimation accuracy and robustness under mutual coupling and snapshot-limited conditions.
Denoising diffusion error correction code (DDECC) is a recent state-of-the-art neural decoding framework by casting decoding as a deterministic reverse diffusion process conditioned on parity-check errors. However, its heuristic step parameterization and lack of sample diversity limit decoding performance. We propose a time-conditioned diffusion-based decoder that reformulates the decoding task as a discrete-time reverse diffusion process driven by an advanced diffusion sampler. Our method introduces two key components: (i) a Transformerbased denoiser conditioned on diffusion time steps, and (ii) an initialization that maps the received signal to a valid diffusion state. To further enhance error correction, we propose a multiple diffusion sampling (MDS) algorithm that leverages stochasticity to generate L parallel decoding trajectories for enhanced solution exploration. Experiments on $(63,45) \mathbf{B C H}$ and $(64,42)$ polar codes show that our decoder can achieve superior decoding accuracy over DDECC with reduced complexity. As L increases, MDS approaches maximum-likelihood performance.
This paper presents a quantization-aware implementation of a hybrid beamforming (HBF) receiver on a field-programmable gate array (FPGA) for millimeter-wave (mmWave) communications, addressing practical hardware constraints such as fixed-point arithmetic and the limited resolution of analog phase shifters. The proposed design incorporates input normalization and a baseband refinement strategy to enhance computational robustness and index selection accuracy under low bit-width settings, thereby preserving spectral efficiency. The system adopts a co-optimized architecture that partitions computation between the FPGA and the ARM processor, leveraging the structure of the HBF algorithm to enable efficient hardware acceleration without compromising accuracy. We design and evaluate multiple FPGA implementations with varying degrees of parallelism, highlighting clear trade-offs between latency and resource usage. Even in the slowest configuration and including data transfer overhead, the proposed system significantly outperforms ARM-based software processing in execution time.
This paper addresses the joint design of transmit beamforming and Active Intelligent Reflecting Surface (AIRS) reflection in a Directional Modulation (DM) assisted Simultaneous Wireless Information and Power Transfer (SWIPT) system. The resulting optimization problem is non-convex and involves coupled variables with discrete constraints, making it challenging to solve. To address this challenge, we propose a novel algorithm named the Refined Element-Wise Greedy Search (RE-GS) algorithm to enhance physical layer security. The core of RE-GS is a nested optimization framework: for each candidate discrete state of a single AIRS element, the algorithm re-optimizes the transmit beamforming vector and the power splitting ratio. This meticulous search process ensures a robust solution. Simulation results show that the proposed scheme can effectively suppress signal leakage to eavesdropping users, enhance communication security and energy transfer efficiency. The RE-GS algorithm is shown to achieve its security objectives, validating its effectiveness for such performance-critical applications.
The Ziv-Zakai bound (ZZB) for two-dimensional (2D) wideband direction-of-arrival (DOA) estimation of a single source under subband model is investigated. The regularity conditions and the minimum error probability for 2-D wideband ZZB are derived first, along with its closed-form expression. In the asymptotic region with high signal-to-noise ratio (SNR), the derived ZZB approaches the Cramér-Rao bound (CRB). In contrast, in the prior performance region with a low SNR, the estimation error is distributed throughout the entire prior parameter space, and thus ZZB depends on the prior covariance matrix. Simulation results indicate that the derived ZZB provides a better performance benchmark than the widely accepted CRB in terms of tightness and effectiveness.
This paper investigates an unmanned aerial vehicle (UAV)-enabled dual-function radar and communication (DFRC) system, equipped with a movable antenna (MA) array which allows dynamic adjustment of antenna positions to enhance spatial beamforming flexibility. A joint optimization framework, which simultaneously considers the UAV trajectory, antenna position, transmit radar and communication beamforming, is developed. The objective is to maximize the achievable communication data rate, subject to the constraints on the UAV mobility, radar beampattern gain and power budget. To efficiently solve the resulting non-convex problem, an alternating optimization (AO) algorithm is proposed. Specifically, the particle swarm optimization (PSO) scheme is used to optimize the antenna positions, while beamforming and trajectory variables are updated via a successive convex approximation (SCA) method. Simulation results confirm that the proposed design achieves significant performance improvements over conventional UAVDFRC systems with fixed antenna arrays (FPA).
Systolic arrays (SAs) for matrix multiplication are commonly used in machine learning (ML), wireless communication, and signal processing. Inherently offering high throughput with good data reuse, they are well-positioned for both low-power edge devices and accelerator applications in high-performance computing. Current realizations suffer from startup latency, defined as the time required to fully utilize all processing elements (PEs). In this work, this issue is addressed by introducing bidirectional systolic arrays with connected edges that form toroidal dataflows. The proposed systolic arrays significantly reduce computational and readout latency from $4 n-2$ to $2.5 n-1$ clock cycles for an $n \times n$ matrix multiplication, while simultaneously reducing energy per operation by up to $43 \%$ compared to conventional SAs. Moreover, a variety of differently shaped SAs are synthesized in a 22 nm CMOS technology, and it is shown that the toroidal designs offer a $5 \%-12 \%$ lower silicon area cost.
Direction-of-arrival (DOA) estimation is a fundamental problem in radar, sonar, and other related fields. In this paper, the key difference between wideband and narrowband signal models is analyzed, and a wideband DOA estimation method based on the theory of matched filtering is proposed. After matched filtering together with a power-peak detection process, the narrowband DOA estimation method can be applied to the formed time-domain model, indicating that the matched filtering serves as a bridge, transforming the wideband timedomain model into a narrowband-like counterpart under certain conditions. The numerical simulations show that the proposed method approaches the optimal performance of the frequencydomain focusing without prior DOA information and with lower computational complexity.
Guessing Random Additive Noise Decoding (GRAND) is a recently proposed universal decoding technique for short linear block codes. GRAND-based hardware implementations range from those aimed at low resource utilization at the expense of higher decoding latency to those aimed at achieving low latency at the cost of higher resource utilization. To the best of our knowledge, this work offers the first High-Level Synthesis (HLS)-based open-source hardware implementation framework for GRAND. The proposed framework simplifies Design Space Exploration (DSE) for GRAND hardware implementation and facilitates selecting the implementation parameters to achieve the optimal tradeoff between the decoding latency and hardware resources. The HLS-based framework is employed for developing the baseline hard-input GRAND hardware, and the VLSI architecture is modified to increase the parallelization factor. Hardware implementation results demonstrate that improved GRAND with a parallelization factor of 4 requires $34 \%$ more hardware resources; however, the worst-case decoding latency is $\frac{1}{4} \times$ the latency of the baseline GRAND.