It is important to estimate the 3-D human shape from a single image, as it is utilized in virtual reality, tactile games, and human-computer interaction. Most previous studies attempted to establish a mapping from the input image to 3-D shape coordinate spaces using neural network optimization based on Euclidean distance. However, such a strategy fails to maintain the aligned 2-D-3-D surface of the human body because structural topological associations have not been effectively considered. We propose a surface-aligned 3-D human shape estimation via the Gaussian curvature structure learning approach. Specifically, we use positive or negative Gaussian curvature to model the convex and concave appearance of the human body. Then, we use them to guide our 3-D human shape estimation network, capturing the structural relationships between the vertices and the local surfaces of the human body. The optimization of our network is simultaneously processed via coordinate and curvature features in both Euclidean and Riemann spaces to maintain the 2-D-3-D consistency of the surface and vertex on the human body. Furthermore, the designed curvature completion module can obtain a more 3-D surface prior, which mitigates the performance degradation problem with more complex deformation of the human body in an outdoor scene. Extensive testing on some selected representative benchmarks demonstrates the efficacy of our approach. Especially, it shows superiority in the detailed texture reconstruction.
Recently, a marginal area that compensates for inaccurate viewport prediction is encouraged to compete for network resources, resulting in improved quality of experience (QoE) of 360-degree video. However, 360-degree video may suffer from a degradation in quality if the marginal area abuse network resources. To address this issue, we propose a multi-priority multi-path transmission framework (M2PT) for delay-constrained 360-degree video streaming in heterogeneous time-varying wireless networks. The proposed framework analyzes the importance of multi-priorities under dynamic network conditions to solve the trade-off between load balancing of different links and quality allocation of different viewports. Besides, a semi-physical simulation platform is established to verify the performance of M2PT. The experiment results show that the proposed M2PT provides better performances: a lower ratio of overdue frames, a reduction of the freeze ratio, and a higher average QoE quality, up to 3–6 dB for different video sequences compared to several reference schemes.
LoRa is a widely recognized modulation technology in the field of low power wide area networks (LPWANs). However, the data rate of LoRa is too low to satisfy the requirements of Internet of Things applications. To address this issue, we propose a novel high-data-rate LoRa scheme based on the spreading factor index (SFI). In the proposed SFI-LoRa scheme, the starting frequency bin of a chirp signal is used to transmit information bits, while the combinations of spreading factors are exploited as a set of indices to convey additional information bits. Moreover, the theoretical symbol error rate, data rate, transmission throughput, complexity and energy efficiency of the proposed SFI-LoRa scheme are carefully analyzed. Simulation results not only verify the accuracy of our theoretical analysis, but also demonstrate that the proposed SFI-LoRa scheme can improve the transmission throughput of existing LoRa schemes without sacrificing the BER performance over additive white Gaussian noise, Rayleigh fading, and multipath flat-fading channels. Therefore, the proposed SFI-LoRa scheme is a potential solution for applications requiring a high data rate in the LPWAN domain.
In this paper, we investigate the performance of protograph-based low-density parity-check (LDPC) codes for rate-diverse two-user Gaussian multiple access channel (GMAC). Considering the characteristic of rate-diverse GMAC, we first analyze the message passing model of the coding system. Then, we derive a rate-diverse joint protograph-based extrinsic information transfer (RDJP-EXIT) chart to analyze the rate-diverse joint user messages decoding (RDJD) behavior. Guided by the RDJP-EXIT, we design the protograph base matrices of different channel conditions for two users, respectively. Simulation results show that the proposed protograph-based LDPC codes of low decoding threshold significantly outperform the conventional AR4JA codes and irregular LDPC codes in GMAC channels, which are also close to optimal in terms of sum rate.
In this paper, we propose an Adaptive Gabor filter-guided Feature Enhancement Network (AGNet) for fast scene text detection. Specifically, we introduce Gabor Orientation Filters (GoFs) into the shallow layers of the model to extract spatial frequency and orientation features from arbitrarily shaped text. The Gabor-guided model enhances sensitivity to texture features, improving text feature representation in lightweight networks. Additionally, considering the limited receptive field in lightweight feature extraction, we propose the Weighted Feature Enhancement Network (WFEN) to integrate multi-level features. This network combines bidirectional feature maps and spatial weighting to effectively filter the Gabor-guided backbone features. The dual-stage fusion ensures efficient text feature representation with low parameters. Extensive experiments on the CTW1500, Total-Text, and MSRA-TD500 datasets demonstrate that AGNet achieves state-of-the-art performance in balancing accuracy and inference speed. Notably, it achieves a competitive F-measure of 86.8% at a speed of 40.3 FPS on the CTW1500 dataset. Code is available at:https://github.com/Wabbb1/AGNET.
With the rapid development of intelligent sensing technology, unmanned aerial vehicles (UAVs) have shown their potential in mobile non-cooperative target (MNCT) tracking tasks. Collaborative learning of target motion characteristics is an effective means of improving tracking stability for UAV swarms. However, UAV swarm-based MNCT tracking faces critical challenges in real-world environments, including sensing interference and data heterogeneity. To address this and realize per-sensing-round updates of the models deployed on the UAV swarm, we propose FedMES, a federated learning framework that jointly optimizes data reliability and model performance. First, we design a data filtering mechanism using polynomial fitting and moving average correction to suppress destructive noise. Second, dynamic aggregation weights are assigned based on historical data confidence to prioritize high-reliability clients. Third, clients selectively adopt the global or local model via a dual-metric evaluation post-aggregation. Evaluated on real-world MNCT trajectories, FedMES reduces mean estimation errors by 10%-22% compared to FedAvg/FedProx, demonstrating superior robustness under high noise and data loss.
This letter proposes a novel multilevel polar-coded differential spatial modulation (MLP-DSM) scheme. Specifically, the proposed structure intelligently integrates the polar coding and DSM by cascading signal polarization and channel polarization to establish the MLP-DSM system with a more significant overall polarization effect. Furthermore, we put forward the concept of matrix Hamming distance (MHD) to quantify the intrinsic difference between antenna activation order matrices (AAOMs). Based on the MHD metric, we design a novel set partitioning order (SPO) constellation tailored for MLP-DSM systems. The designed SPO constellation exploits the reduced Latin rectangles (RLRs) to increase the structural difference among AAOMs, thereby improving the constellation-constrained capacity. Moreover, the set partitioning (SP) mapping criterion is adopted to further strengthen the system polarization effect. Simulation results demonstrate that the proposed MLP-DSM schemes obtain excellent error performance gains over the conventional DSM system and other counterparts.
Split learning (SL) transfers most of the training workload to the server, which alleviates computational burden on client devices. However, the transmission of intermediate feature representations, referred to as smashed data, incurs significant communication overhead, particularly when a large number of client devices are involved. Existing SL methods apply uniform compression across all channels, over-compressing smashed data associated with important channels while under-compressing smashed data corresponding to less important channels. To address this challenge, we propose a communication-efficient adaptive channel pruning-aided SL (ACP-SL) scheme. In ACP-SL, a label-aware channel importance scoring (LCIS) module is designed to generate channel importance scores, distinguishing important channels from less important ones. Based on these scores, an adaptive channel pruning (ACP) module is developed to prune less important channels, thereby compressing the corresponding smashed data and reducing the communication overhead. Experimental results show that ACP-SL outperforms most benchmark schemes in test accuracy. Furthermore, it reaches a target test accuracy in fewer training rounds, reducing communication overhead.
Scene Text Recognition (STR) is challenging in extracting effective character representations from visual data when text is unreadable. Permutation language modeling (PLM) is introduced to refine character predictions by jointly capturing contextual and visual information. However, in PLM, the use of random permutations causes training fit oscillation, and the iterative refinement (IR) operation also introduces additional overhead. To address these issues, this paper proposes the Hierarchical Attention autoregressive Model with Adaptive Permutation (HAAP) to enhance position-context-image interaction capability, improving autoregressive LM generalization. First, we propose Implicit Permutation Neurons (IPN) to generate adaptive attention masks that dynamically exploit token dependencies, enhancing the correlation between visual information and context. Adaptive correlation representation helps the model avoid training fit oscillation. Second, the Cross-modal Hierarchical Attention mechanism (CHA) is introduced to capture the dependencies among position queries, contextual semantics and visual information. CHA enables position tokens to aggregate global semantic information, avoiding the need for IR. Extensive experimental results show that the proposed HAAP achieves state-of-the-art (SOTA) performance in terms of accuracy, complexity, and latency on several datasets.
As the process node is scaled down, the spin-transfer-torque magnetic random-access memory (STT-MRAM) exhibits higher memory density than the static random-access memory (SRAM), making it one of the more promising successors of the low-level on-chip cache memory. However, the low read margin (RM) of the magnetic tunnel junction (MTJ) in STT-MRAM can limit the achievable read accuracy. We implemented 2-bit channel quantization for error-correcting code (ECC) schemes and explored the trade-offs between improved read accuracy and factors such as circuit area, power consumption, and latency. The proposed quantization scheme consists of a sensing amplifier-based 2-bit quantizer and MTJ resistor-based soft-decision thresholds. Compared to 1-bit channel quantization using the Bose-Chaudhuri-Hocquenghem (BCH) code, the proposed 2-bit quantization architecture achieves a fourfold reduction in frame error rate (FER) from 8.0 & times;10-4 to 2.0 & times;10-4 when paired with polar codes and successive cancellation (SC) decoding. Additionally, this approach results in decoding complexity that is only 1/13th of that required for BCH at a 0.7 code rate.
In this paper, we propose a physical-layer shaped network coding (PLSNC) scheme that integrates physical-layer network coding (PLNC) with constellation shaping for amplitude phase-shift keying (APSK) systems. Two source nodes employ low-density parity-check (LDPC) and linear shaping codes to prevent ambiguous detection. By biasing a subset of the LDPC output toward zeros, lower-energy symbols are transmitted more frequently. This characteristic enables the relay to directly recover the shaped network codeword without decoding individual nodes' shaping bits. Furthermore, we investigate the design criteria for linear shaping codes and introduce an iterative decoding architecture comprising the APSK demodulator, shaping decoder, and LDPC decoder. Simulation results demonstrate that the proposed scheme significantly improves both error performance and shaping gain compared to conventional unshaped systems.
Low-Power Wide Area Networks (LPWANs) support large-scale IoT connectivity but are capacity-limited, motivating concurrent transmissions to improve spectral efficiency. However, existing concurrent LPWAN solutions typically rely on a single base waveform or quasi-orthogonal waveforms, resulting in limited concurrency. In this paper, we propose Zadoff-Chu Random Access (ZCRA), a low-overhead framework for massive random access. ZCRA introduces a Zadoff-Chu-sequence-based waveform, which modulates multiple bits via cyclic shifts of Zadoff-Chu root sequences, exploiting their zero cyclic autocorrelation property. By allowing users to randomly select root sequences, ZCRA effectively mitigates inter-user interference without waveform pre-assignment based on low cross-correlation property. To resolve same-root collisions under low SINR, we further develop a constellation-set-match-based decoding scheme (CSMatch) that leverages spectral correlation structure to separate superimposed transmissions and remains robust to carrier frequency offsets and multipath channels. Extensive evaluation demonstrates that ZCRA-CSMatch significantly improves concurrency, achieving a lower symbol error rate compared with state-of-the-art collision resolution approaches, along with a $\mathbf{2. 1 3} \boldsymbol{\times}$ network throughput improvement over current concurrent systems and supporting 200 nodes (40 concurrency).
In this paper, we propose a polarization-aided belief propagation (PA-BP) decoding for M-user uplink non-orthogonal multiple access (NOMA). In this scheme, polar code of each user is regarded as a nested code within a longer polar code with BP decoding, which has the length M times than that of the original user codes. It enables the PA-BP decoder to obtain all user messages with the enhanced channel polarization effect, as compared to the conventional successive interference cancellation BP (SIC-BP) decoder at finite block length. To improve the performance of PA-BP decoder, we further propose interference cancellation (IC) aided-BP (ICA-BP) and ICA-BPL list decoders that incorporates IC during PA-BP iterations. We then analyze the channel capacity of the proposed scheme and use extrinsic information transfer (EXIT) chart to show the convergence behavior of these iterative decoders. By utilizing the Monte Carlo method to optimize this long polar code, simulation results show that the proposed PA-BP outperforms the conventional SIC-BP by 0.8 dB over Gaussian channels, and by 0.7 dB over Rayleigh fading channels. In particular, ICA-BP can achieve more performance gain of 0.4 dB over PA-BP over both channels.
Non-binary successive cancellation list (NB-SCL) decoding expands each surviving path into q candidate branches at every information symbol, which causes high path expansion, sorting, and pruning complexity. To address this issue, this paper proposes low-complexity list decoding algorithms for 2×2 kernel non-binary polar codes (NBPCs). First, we design a split-reduced non-binary successive cancellation list (SR-NBSCL) decoder that skips path splitting when the current symbol is sufficiently reliable. We then exploit the final Rate-1 node structure and switch the last group of information symbols to simplified non-binary successive cancellation (NB-SC) decoding, resulting in the enhanced split-reduced non-binary successive cancellation list (ESR-NBSCL) decoder. To further reduce branch expansion at unreliable symbols, we introduce an accumulated reliability-deviation (ARD) metric and propose an adaptive branch-pruning non-binary successive cancellation list (ABP-NBSCL) decoder, which prunes unreliable candidate branches before sorting and then reduces the dominant sorting complexity. Simulation results show that the proposed decoders achieve frame-error-rate (FER) performance close to that of conventional NB-SCL decoding with much lower complexity. In particular, the ABP-NBSCL decoder reduces the path splitting number (PSN) by more than 80% at several tested signal-to-noise ratios (SNRs), with a negligible performance loss.
Haze, aerosols, and thin cloud-like atmospheric interference reduce contrast and obscure structural cues in optical remote sensing images, degrading both visual quality and downstream interpretation. Direct-regression networks may not sufficiently exploit physically interpretable haze and frequency cues, whereas full-image diffusion requires iterative reconstruction of the complete clean image. This paper presents PFDiff, a physics- and frequency-guided residual diffusion framework for remote sensing image dehazing. A haze-aware restoration U-Net first predicts a coarse restoration using a dark-channel map, a transmission-like proxy, and a Laplacian high-frequency cue. Conditional diffusion then models the lower-energy residual between the clean target and the coarse output, conditioned on an 11-channel tensor comprising the hazy input, coarse restoration, and the three priors. A learned gate applies spatially varying residual fusion. We further construct HazeRS45 using real-haze-derived masks, multiple transmission patterns, and atmospheric-light variations. Experiments on paired remote sensing dehazing datasets, a cross-degradation benchmark, unpaired real-world images, and downstream tasks demonstrate strong fidelity and structural preservation. Across RSID and HazeRS45, PFDiff achieves the highest average PSNR and SSIM of 26.17 dB and 0.94, respectively, together with the best average DISTS among the compared methods. It also shows favorable cross-degradation and no-reference real-world performance. Controlled ablations support the residual target and identify 10-step sampling as a balanced quality–efficiency setting.
Deoxyribonucleic acid (DNA) data storage is recognized as a transformative archival technology due to its ultrahigh density and exceptional longevity. However, its practical capacity is severely limited by the conventional separated architecture, where lossy source coding and DNA constraint satisfaction are implemented as independent, sequential stages. This work proposes a unified framework that integrates lossy compression with DNA base composition constraints. A protograph-based low-density parity-check code structure is employed for efficient quantization. To enable joint optimization, a novel multi-point central difference method is devised to approximate gradients through non-differentiable operations. Experimental results demonstrate that the proposed system directly generates biochemically feasible DNA sequences while approaching the rate-distortion bound. This will advance the integration of information theory with physics-dependent system design.
Real-time shipping container code spotting (CCS) plays a vital role in enabling intelligent inspection to optimize smart port logistics. Recent methods suffer from error accumulation and have complex pipelines, resulting in high computational complexity that limits the generalization ability for multidirectional, including horizontal, vertical, and multiline (HVM) codes. Moreover, fixing the parameters of visual sensors (e.g., angle, position, and focal length) to avoid noise interference limits deployment on mobile devices. To address them, we propose an end-toend framework for HVM code spotting. First, an inter-domain feature fusion network that includes the self-spatial enhancement module (SSEM) and self-channel enhancement module (SCEM) is proposed to reduce noise interference and improve the efficiency of the lightweight model. Next, a transformerbased branch with character contrastive earning (CCL) loss is designed to enhance the representation of character features. Experimental results indicate that the proposed method achieves state-of-the-art performance, in which the F-measure and recognition accuracy reach 91.2% and 93.2%, respectively, in real-time.
In unsourced random access (URA) system, user pilot collisions are a key challenge when multiple users select identical pilot sequences, leading to poor activity detection and decoding performance. To address this issue, this paper develops a two-stage segmented pilot (TSSP) design for collision-resilient URA in massive multiple-input multiple-output (MIMO) system. The proposed scheme enables a substantial enlargement of the effective pilot space without increasing the physical codebook size and detection complexity. First, we construct a probabilistic collision model by leveraging the birthday-problem analogy, which provides theoretical characterization of collision events and closely matches the simulation results. Second, we propose a segmented pilot mapping and decoding framework in which the pilot bits are divided into two interconnected parts linked by common pilot bits. This structured segmentation allows the system to resolve partial collisions and reduce missed detection in high-density access scenarios. Third, we develop a joint-covariance linear minimum mean square error (JC-LMMSE) scheme that exploits the segmented pilot structure to suppress pilot contamination and improve estimation accuracy. Monte Carlo simulations show that the proposed TSSP framework achieves up to 3.5 dB gain in collision scenarios and approximately 1 dB gain in collision-free cases over existing methods, while maintaining comparable computational complexity.
The high spectral efficiency of spatial modulation-based schemes is generally achieved at the expense of requiring a larger number of antennas. Additionally, existing permutation matrix modulation-based schemes transmit $M$-ary phase shift keying ($M$-PSK) symbols using the same modulation order, which limits their flexibility. To address these issues, a novel multiple-mode permutation matrix modulation (MM-PMM) scheme is proposed in this paper. In MM-PMM, the permutation index, along with multiple-mode constellation and its index, are utilized to transmit information bits, thereby enhancing spectral efficiency. Furthermore, joint maximum likelihood detection (J-MLD) and low-complexity detection (LCD) methods are proposed to recover the transmitted bits. The bit error probability of the proposed MM-PMM scheme is also derived. Finally, the bit error rate (BER) performance of the proposed MM-PMM scheme is compared to that of benchmark schemes, and comparison results demonstrate that the proposed MM-PMM scheme achieves at least a 2 dB gain in BER performance over benchmark schemes.
In this paper, a quadruple code index modulation (QCIM) technique is proposed for low-complexity and spectral-efficient wireless transmission. The proposed QCIM scheme utilizes quadruple code indices in addition to M-ary phase shift keying symbols to transmit more information bits compared to conventional code index modulation (CIM) schemes. To recover the transmitted information bits, we propose maximum likelihood detection (MLD) and greedy detection (GD) algorithms for the proposed QCIM, with the GD algorithm offering almost the same bit error rate (BER) performance to the MLD algorithm but with significantly lower system complexity. Additionally, we extend QCIM into a reconfigurable intelligent surface (RIS)-assisted QCIM (RIS-QCIM) system, which significantly enhances BER performance and eliminates the need for channel state information at the receiver. Numerical results demonstrate that the proposed QCIM and RIS-QCIM achieve higher spectral efficiency and superior BER performance in contrast to benchmark schemes.