
Massive MIMO (Multiple-Input Multiple-Output) is an advanced wireless communication technology, using a large number of antennas to improve the overall performance of wireless communication systems in terms of capacity, spectral efficiency, and energy efficiency. The performance of MIMO systems is highly dependent on the quality of channel state information (CSI). Predicting CSI is, therefore, essential for improving communication system performance, particularly in MIMO systems, since it represents key characteristics of a wireless channel, including propagation, fading, scattering, and path loss. This study proposes a BERT-inspired transformer architecture, called BERT4MIMO, which is specifically designed to process high-dimensional CSI data from massive MIMO systems. BERT4MIMO offers superior performance in reconstructing CSI under varying mobility scenarios and channel conditions through deep learning and attention mechanisms. The experimental results demonstrate the effectiveness of BERT4MIMO in a variety of wireless environments.
Integrated sensing and communication (ISAC) has emerged as a promising technology for improving spectrum utilization by enabling radar sensing and wireless communication within a unified platform. In this paper, we consider an FMCW-based Radar and Communication (FRaC) system employing index modulation (IM) for autonomous vehicle applications and address two important challenges: the grid mismatch in radar parameter estimation and the high computational complexity of communication decoding. For radar sensing, we propose an off-grid estimation framework that successively refines the range, velocity, and angle estimates using local first-order approximations and block-sparse recovery. For the communication receiver, we propose a low-complexity iterative decoder for the IM symbols that estimates the transmitted symbol vector in the continuous domain through gradient-descent iterations and subsequently maps the continuous-domain estimates to the corresponding symbols in the discrete codebook via a quantization step, thereby avoiding the exhaustive search required by maximum-likelihood (ML) detection. In addition, essential theoretical properties of the proposed methods are investigated, including the derivation of the Cramér-Rao lower bound (CRLB) together with the analyses of local identifiability and algorithm behavior. Simulation results demonstrate that the proposed off-grid estimation method substantially improves parameter estimation accuracy, while the proposed iterative decoder achieves a significant reduction in computational complexity with only a modest degradation in bit-error-rate performance compared with ML detection.
Light Detection and Ranging (LiDAR)-based three-dimensional (3D) object detection plays a critical role in autonomous driving because it enables accurate perception of surrounding objects from sparse point-cloud measurements. However, point clouds are inherently irregular and sparsely distributed, especially for distant or occluded objects, making it difficult to extract discriminative and robust geometric representations. Existing convolution-based detectors mainly rely on local receptive fields and therefore struggle to capture long-range dependencies, whereas Transformer-based methods improve global interaction but often overlook structural sparsity and noisy attention responses in voxel features. To address this gap, this paper aims to develop an efficient and robust Transformer-based detector that can adaptively perceive informative sparse structures while suppressing unreliable feature interactions. We propose DiffFormer, a Differential Transformer framework, for LiDAR-based 3D object detection. DiffFormer consists of an Adaptive Structural Perception (ASP) module and a Hybrid Differential Attention (HDA) mechanism. The ASP module enhances structural awareness through foreground voxel selection and hierarchical region expansion, while the HDA mechanism models positive-negative attention differences to suppress noisy responses and strengthen reliable long-range dependency modeling. Experiments on the KITTI validation set demonstrate that DiffFormer achieves 84.77% and 75.25% 3D AP at 40 recall positions for the Car and Cyclist categories, respectively, under the Moderate difficulty level.
Multi-focus image fusion (MFIF) aims to generate an all-in-focus image from multiple partially focused source images, preserving the optimal information from each input. However, when processing scenes with complex textures and noise interference, existing methods often produce artifacts at the focus boundaries, making it difficult to balance brightness and contrast consistency. To address the aforementioned issues, this paper proposes a multi-focus image fusion method based on uncertainty-aware dual-channel coupled neural P systems (UA-DCNPS).After the source images are decomposed by the stationary wavelet transform (SWT), high-frequency subbands and low-frequency subbands are obtained. For the high-frequency sub-bands, the UA-DCNPS models the local clarity and its uncertainty, generating a weight map by converting hard decisions into continuous mappings to attenuate boundary artifacts. For the low-frequency subbands, an adaptive weighting strategy based on regional energy and structure preservation constraints is designed to maintain the overall brightness and contrast consistency of the fused image. Experimental results demonstrate that the proposed method outperforms nine representative fusion methods in terms of FMI, SSIM, and PSNR on the Lytro, MFFW, and MFI-WHU datasets. Compared with the corresponding second-best results, the maximum improvements in these three metrics reach 1.18%, 3.44%, and 1.26%, respectively. Ablation experiments and complementary graphical analyses further validate the effectiveness of the uncertainty-aware mechanism and the proposed fusion strategy. In addition, visible–infrared image fusion experiments on the TNO dataset indicate that the proposed method has potential for cross-modal applications.
Active target velocity estimation in the deep-sea bottom-bouncing area is a critical challenge, which requires overcoming the performance degradation of estimation methods caused by low received signal-to-noise ratio and reverberation interference. The velocity estimation method based on the matched filter (MF) is a conventional approach in active sonar system. However, it is vulnerable to the aforementioned environmental interference and fails to exploit the characteristics of the deep-sea bottom-bouncing area. To solve this problem, a target velocity estimation method based on the improved parameterized centroid frequency-chirp rate distribution (PCFCRD) is proposed in this paper. It takes advantage of the multipath characteristics of the deep-sea bottom-bouncing area and the features of PCFCRD in multipath condition. Using the approximate delay differences obtained through a constrained search, the improved PCFCRD achieves multipath energy aggregation by translating and superimposing the PCFCRD results of different paths, thereby enhancing the ability to resist environmental interference. During superposition, point-wise processing is implemented to further enhance noise and reverberation resistance and improve the ability to suppress sidelobes. Simulations and validation based on sea trial data indicate that the improved PCFCRD has higher velocity estimation accuracy and better robustness than the MF-based velocity estimation method and the conventional PCFCRD under environmental interference.
Analyzing non-stationary respiratory sounds requires capturing brief, high-frequency acoustic transients. However, standard high-resolution time-frequency representations exhibit significant temporal redundancy and impose memory constraints for edge-based deployment, while traditional downsampling methods often omit or alias transient biomarkers. To address this redundancy-fidelity problem, we propose MelHilbertLDS, a lightweight acoustic feature representation framework. The method integrates quasi-random temporal downsampling via Low-Discrepancy Sequences (LDS) with spatial structuring using generalized Hilbert space-filling curves. LDS minimizes discrepancy bounds to provide principled coverage of transient events, acting as a temporal regularizer. Hilbert folding translates 1D temporal proximity into 2D spatial adjacency, which expands the effective receptive field of convolutional neural networks (CNNs) and compresses spatial dimensions. We evaluated the framework across three acoustic datasets, covering controlled event detection, data-scarce environments, and noisy clinical recordings. Results show that MelHilbertLDS reduces performance variance in low-resource scenarios and improves training stability. It achieves up to a 50-fold reduction in activation memory, a 44% reduction in CPU inference latency on MobileNetV3, and a 96% reduction in FLOPs on high-capacity CNNs, while maintaining competitive diagnostic performance. MelHilbertLDS offers a promising approach for automated analysis of transient acoustic events under memory constraints.
Medical image classification plays a vital role in clinical diagnosis. Although existing hybrid architectures have mitigated the challenges caused by scale variation and background interference, they remain inadequate in effectively fusing local details with global semantics. Moreover, deep feature fusion may lead to an imbalance between low-frequency structural information and high-frequency detail cues. To address these limitations, we propose DHSC-Net, a Dynamic Hybrid Network with Spectral Consistency Regularization. Specifically, we introduce a Dynamic Scale-Aware Hybrid Fusion (DSHF) module at each fusion stage to enable sample-adaptive local-global integration. Within this module, Dynamic Scale-Aware Gating (DSG) performs cross-branch dynamic weighting, while Dual-Dimension Feature Recalibration (DDFR) further refines the fused representations through sequential recalibration. Meanwhile, a spectral consistency loss is imposed during training. By explicitly decomposing fused features into low- and high-frequency components, the proposed constraint guides low-frequency structures toward the global branch and high-frequency details toward the local branch, thereby encouraging branch-wise spectral alignment in the fused features. Extensive experiments were conducted on four medical image classification datasets. The results show that DHSC-Net achieved accuracies of 91.44%, 85.78%, 94.17%, and 89.20% on Kvasir v2, ISIC 2018, COVID-19, and private UC datasets, respectively, achieving the best overall performance among the compared advanced models. Additional ablation studies and mechanism analyses support the contributions of spectral regularization and dynamic fusion design, further demonstrating the robustness and practical relevance of DHSC-Net.
This paper focuses on the time-of-arrival measurement model, and investigates the problem of robust indoor positioning in impulsive noise environments. To enhance the resistance of the conventional positioning methods to the impulsive noise, the optimum external weighting matrix is derived, and a new robust approach, which is based on the weighted ℓp-norm (WLP) minimization, is devised. To overcome the difficulty in solving the resultant nonconvex optimization problem, the majorization-minimization technique is utilized to relax it as several successive convex counterparts, which are tackled in an iterative way. Furthermore, to improve the positioning accuracy without increasing the number of base stations, a multi-snapshot framework of indoor positioning is proposed, which fuses the multiple time-of-arrival measurements to suppress impulsive noise in a statistically efficient way. The convergence, asymptotic mean square error, and computational complexity of the WLP-based approach are analyzed theoretically. Simulation results demonstrate the robustness of the proposed indoor positioning approach under the Gaussian mixture model noise by comparing with other methods.
In this paper, we present a new method for parameter estimation of multi-component linear frequency modulated (LFM) signals, using fractional Fourier transform (FrFT) and an improved super-resolution generative adversarial network (SRGAN). Different from existing estimation methods, the proposed method innovatively transforms parameter estimation into image processing, and effectively overcomes the impact of spectral overlap in FrFT domain by performing high-resolution reconstruction of signal peaks. This method initially maps multi-component LFM signals into a two-dimensional FrFT domain with coarse-step FrFT, in which each component exhibits a distinct energy peak. SRGAN-based image super-resolution (SR) is then introduced to perform denoising and high-resolution reconstruction on the low-resolution FrFT spectrum. Finally, the signal parameters are estimated using the peak locations in the high-resolution FrFT spectrum. Considering that conventional SRGAN mainly focuses on improving image perceptual quality, directly applying SR to the FrFT spectrum makes it difficult to recover peak features. To address this issue, we incorporate attention mechanisms into the network architecture to enhance signal energy concentration at peak locations. Meanwhile, an estimation loss is defined to avoid peak shifts and spurious peaks in the SR results. Extensive simulations demonstrate that the proposed method provides superior estimation accuracy and exhibits strong noise robustness.