This letter presents a novel generative semantic communications framework for wireless image transmission, designed to excel in dynamic channel conditions and at low signal-to-noise ratios (SNRs). The proposed framework, dubbed DiffMoECom (Diffusion Mixture-of-Experts Communications), leverages a Mixture-of-Experts (MoE) architecture integrated with prompt learning to dynamically select optimal channel adaptation strategies by activating the most relevant prompt experts. A Joint Source and Channel Coding (JSCC) backbone is employed to transmit the low-frequency components of the source image. Furthermore, we introduce a range-space inversion mechanism, which projects the coarse JSCC output into a noisy latent space. This facilitates faithful semantic reconstruction of high-frequency details through conditional diffusion-based null space compensation. Extensive experiments across various fading channels and SNR conditions demonstrate that our method consistently surpasses state-of-the-art DeepJSCC and generative approaches, achieving superior performance in distortion metrics with competitive inference efficiency.
Accurate maritime radar target detection plays a critical role in ensuring maritime surveillance and navigation safety. However, reliable detection in sea clutter environments remains challenging due to the strong non-stationarity of clutter and the low signal-to-clutter ratio (SCR) of small targets. Although convolutional neural networks (CNNs) capture local spatial features from time–frequency (TF) representations and Transformers model long-range temporal dependencies, their combined use for target detection often leads to excessive computational cost. To address this issue, we propose an Efficient CNN–Mamba Network (ECMNet), a hybrid architecture that integrates CNN-based local feature extraction with Mamba-based global sequence modeling. Specifically, a 3-layer lightweight CNN module captures discriminative local spatial structures, while a Mamba-based module, composed of four Mamba blocks, models long-range temporal dependencies via selective state-space modeling with linear time complexity. By jointly leveraging both components, ECMNet achieves effective local–global feature modeling while maintaining lower computational efficiency, improving detection performance in complex maritime scenarios. Extensive experiments on the IPIX maritime radar datasets demonstrate that ECMNet consistently outperforms state-of-the-art (SOTA) methods, achieving detection probabilities exceeding 90 10^-3 . Furthermore, compared with SOTA approaches, ECMNet not only achieves superior detection performance but also demonstrates enhanced computational efficiency in terms of parameter scale, computational complexity, and inference latency.
Accurate and robust foliage penetration (FOPEN) target recognition is crucial for practical outdoor sensing applications. Compared to traditional sensing technologies, Device-Free Sensing (DFS) has emerged as a promising solution due to its low deployment cost. However, existing DFS-based approaches suffer from significant performance degradation under cross-weather domain shifts and typically rely on extensive relabeling or online fine-tuning when climate conditions change. Such dependence on parameter updating severely limits their applicability in dynamic foliage environments. To overcome this limitation, we propose PNFOPENet, a meta-learning framework that employs Prototypical Networks (PN) to enable FOPEN recognition across different weather conditions using minimal labeled samples, without requiring any model parameter updates. Specifically, the model is trained through episodic meta-learning tasks under given weather conditions to learn class prototypes using a small number of samples. This facilitates the model’s adaptation to unseen weather conditions without updating, by using few samples to recalibrate class prototypes, with samples classified based on a metric learning approach. Extensive experiments conducted on a real-world FOPEN dataset collected under multiple weather conditions demonstrate that the proposed method consistently outperforms two other selected metric-based meta-learning methods. More importantly, PNFOPENet maintains high recognition accuracy under cross-weather settings without fine-tuning, validating its effectiveness for few-shot FOPEN target recognition.
The growing demands of artificial intelligence and immersive media require communication beyond bit-level accuracy to meaning awareness. Conventional optical systems that focus on syntactic precision suffer significant inefficiencies. We introduce a multidimensional semantic communication framework that bridges this gap by directly mapping high-level semantic features onto the orthogonal physical dimensions of light, frequency, polarization, and intensity, within a multimode fiber. By densely multiplexing information across these physical degrees of freedom, the system achieves an unprecedented semantic equivalent spectral efficiency approaching 1000 bit s(-1) Hz(-1). Moreover, it demonstrates profound resilience, maintaining high-fidelity reconstruction even when the physical-layer symbol error rate exceeds 36%, a condition under which conventional communication systems fail completely. Crucially, this deeply integrated codesign of semantic encoding and physical-layer modulation enables full semantic demodulation with only single-ended intensity detection, therefore significantly reducing system complexity and cost. This work establishes a validated pathway toward hyper-efficient, error-resilient optical networks for the next generation of data-intensive computing.
The development of satellite communication is driving the demand for device identification that can respond to security threats from illegal uplink terminals. Specific emitter identification (SEI) is a promising technology of physical layer device identification. Deep learning (DL)-based SEI can effectively learn features for distinguishing emitters through a large number of samples. In the satellite communication scenario, due to the high cost of labeling, wireless signals often face the dilemma of few labeled samples, leading to a decrease in recognition accuracy of existing methods. Therefore, we introduce an innovative FS-SEI method leveraging multi-domain fusion and metric learning (MDF-ML), eliminating the reliance on an auxiliary dataset. Specifically, MDF-ML is proposed to explore additional implicit samples within the sample space, enhancing generalization performance through multi-domain representation and phase rotation. It then uses metric learning to constrain feature distances in the feature space, thereby boosting discriminability. Our simulation results demonstrate that the accuracy of our proposed SEI method exceeds 90%, which outperforms the existing method by 24.21% under 5 sample shots.
Human action recognition (HAR) based on Wi-Fi signals has become a research hotspot due to its advantages of privacy protection, a comfortable experience, and a reliable recognition effect. However, the performance of existing Wi-Fi-based HAR systems is vulnerable to changes in environments and shows poor system generalization capabilities. In this paper, we propose a cross-environment HAR system (CHARS) based on the channel state information (CSI) of Wi-Fi signals for the recognition of human activities in different indoor environments. To achieve good performance for cross-environment HAR, a two-stage action recognition method is proposed. In the first stage, an HAR adversarial network is designed to extract robust action features independent of environments. Through the maximum–minimum learning scheme, the aim is to narrow the distribution gap between action features extracted from the source and the target (i.e., new) environments without using any label information from the target environment, which is beneficial for the generalization of the cross-environment HAR system. In the second stage, a self-training strategy is introduced to further extract action recognition information from the target environment and perform secondary optimization, enhancing the overall performance of the cross-environment HAR system. The results of experiments show that the proposed system achieves more reliable performance in target environments, demonstrating the generalization ability of the proposed CHARS to environmental changes.
The frequency-dependent beamforming technology has demonstrated outstanding performance in two-dimensional beam training for wideband THz systems, owing to its synchronized and high-resolution scanning capabilities. However, extending this technique to three-dimensional beam training poses challenges, as linearly split beams cannot provide comprehensive angular coverage in the candidate area. To this end, we propose a two-stage 3D beam training scheme, consisting of an initial stage and a refinement stage. In the initial stage, the angular coverage of frequency-dependent beamforming is modeled as a controllable rectangular region, and the codebook is designed by arranging this rectangular region within the candidate area. In the refinement stage, frequency-dependent beamforming is densely arranged around the candidate positions estimated in the initial stage to achieve precise localization. Additionally, we design a true-time delay network to facilitate frequencydependent beamforming for uniform planar arrays. Simulation results show that the proposed method approaches near-optimal performance in terms of achievable sum-rate while maintaining acceptable pilot overhead.
Human action recognition (HAR) based on Wi-Fi plays a critical support in the Internet of Things (IoT). Recently, Wi-Fi-based HAR using deep learning models achieves remarkable performance. However, existing HAR models have poor generalization capacity, where the multipath effects and the recognition tasks diversity in different environments would affect the model performance at a great level. This article proposes a cross-environment HAR system based on the federated learning named WiFed-CHAR. This system collaboratively learns the action feature from source environments and generate a feature extraction knowledge base on the cloud. In addition, a HAR module assignment and optimization strategy is proposed to guide the new environment to inherit the most suitable feature extraction knowledge from knowledge base and achieve high performance even with limited data. Extensive experiments are conducted to validate the effectiveness of WiFed-CHAR. When given one sample/action, the HAR of new environments reaches 80.14%, surpassing other competitive baselines.
The rapid proliferation of emerging video services has substantially increased the demand for real-time transmission of high-definition video content. However, optimizing the trade-off between video transmission quality and communication bandwidth remains a significant challenge. To address this issue, we propose RAJSCC, a rate-adaptive video deep joint source-channel coding (DeepJSCC) framework. RAJSCC incorporates a spatio-temporal token merging mechanism to aggregate semantically similar tokens, effectively reducing redundancy both within and across frames. Additionally, an adaptive token merging predictor, designed based on simple statistical features of the input videos, enables dynamic rate control at the group of pictures (GoP) level, ensuring smooth and continuous variation in the overall video coding rate. Extensive experiments demonstrate that RAJSCC significantly outperforms traditional video transmission schemes, such as H.264 and H.265 with low-density parity-check (LDPC), as well as existing DeepJSCC methods, in terms of reconstruction quality. More importantly, the proposed adaptive spatio-temporal token merging mechanism reduces bandwidth consumption by 63.5% and computational cost by 13.0%, while incurring only a marginal 1–2 dB degradation in reconstruction quality. These findings highlight the effectiveness of RAJSCC in achieving a superior balance between transmission efficiency and video quality, making it a promising solution for real-time high-definition video communication in bandwidth-constrained environments.
In space-air-ground integrated systems, radar signal analysis is crucial for effective spectrum management. In recent years, time-frequency transforms (TFT) have gained significant attention for radar signal detection and identification. However, challenges such as cross-term interference and low signal-to-noise ratio (SNR) limit the effectiveness in multicomponent signal analysis. Therefore, this paper proposes a novel multitask learning-based TFT framework, named One-Stage TFT (OSTFT), which directly generates high-quality time-frequency representations (TFR) from raw in-phase and quadrature signals. OSTFT incorporates a generative network combined with classification and localization tasks to enhance feature extraction and image clarity. Experimental results demonstrate that OSTFT achieves superior performance in TFR quality and radar signal recognition, with an 81.4% detection rate using the You Only Look Once (YOLO) frame-work, outperforming existing TFT methods under various noise conditions. Compared to the best-performing TFR-denoising method, OSTFT improves the detection rate by 2.7%.
We propose an ultrafast C-band spectrometer based on MMF and MCF with 5 MHz detection speed. We further enhance precision to 0.3 pm within a 10 pm range, demonstrating broad bandwidth and high precision.
In recent years, the emerging technique of device-free sensing (DFS) has gained popularity for foliage penetration (FOPEN) target recognition. This popularity is primarily attributed to its inherent advantage of not requiring specialized sensing equipment beyond wireless transceivers. Concerning weather variations, DFS heavily relies on labeled data for model training, which necessitates the annotation of samples for each weather environment. However, this annotation process proves impractical for real-world applications, especially under adverse weather conditions. To address this issue, this paper presents an unsupervised domain adaptation (UDA)-based cross-weather FOPEN target recognition system (CW-FTRS). Experimental results validate that the proposed method achieves an average accuracy of over 72
By applying fractional Fourier transformation for time-frequency representation followed by Vision Transformer for impairment analysis, we achieve precise estimation of all-order polarization mode dispersion in high baud rate long-haul optical communication systems.
The research on proximity sensing electronic skin has garnered significant attention. This electronic skin technology enables detection without physical contact and holds vast application prospects in areas such as human-robot collaboration, human-machine interfaces, and remote monitoring. Especially in the context of the spread of infectious diseases like COVID-19, there is a pressing need for non-contact detection to ensure safe and hygienic operations. This article comprehensively reviews the significant advancements in the field of proximity sensing electronic skin technology in recent years. It covers the principles, as well as single-type proximity sensors with characteristics such as a large area, multifunctionality, strain, and self-healing capabilities. Additionally, it delves into the research progress of dual-type proximity sensors. Furthermore, the article places a special emphasis on the widespread applications of flexible proximity sensors in human-robot collaboration, human-machine interfaces, and remote monitoring, highlighting their importance and potential value across various domains. Finally, the paper provides insights into future advancements in flexible proximity sensor technology.
To comprehensively assess optical fiber communication system conditions, it is essential to implement joint estimation of the following four critical impairments: nonlinear signal-to-noise ratio (SNRNL), optical signal-to-noise ratio (OSNR), chromatic dispersion (CD) and differential group delay (DGD). However, current studies only achieve identifying a limited number of impairments within a narrow range, due to limitations in network capabilities and lack of unified representation of impairments. To address these challenges, we adopt time-frequency signal processing based on fractional Fourier transform (FrFT) to achieve the unified representation of impairments, while employing a Transformer based neural networks (NN) to break through network performance limitations. To verify the effectiveness of the proposed estimation method, the numerical simulation is carried on a 5-channel polarization-division-multiplexed quadrature phase shift keying (PDM-QPSK) long haul optical transmission system with the symbol rate of 50 GBaud per channel, the mean absolute error (MAE) for SNRNL, OSNR, CD, and DGD estimation is 0.091 dB, 0.058 dB, 117 ps/nm, and 0.38 ps, and the monitoring window ranges from 0~20 dB, 10~30 dB, 0~51000 ps/nm, and 0~100 ps, respectively. Our proposed method achieves accurate estimation of linear and nonlinear impairments over a broad range, representing a significant advancement in the field of optical performance monitoring (OPM).
Current WiFi-based respiration detection algorithms may experience performance degradation due to variations in user location, as the relationship between user location and patterns of respiration has not been adequately considered. To overcome this limitation, this paper proposes a spatially directional respiration detection approach, named Wi-locind. Wi-locind employs antenna arrays on commercial WiFi receivers to achieve directional enhancement of respiration signals. Combined with post-filtering techniques, Wi-locind is capable of extracting respiration patterns that are independent of changes in the user's location. Specifically, the Minimum Variance Distortionless Response algorithm is used to identify the arrival angle of the target user and directionally enhance the received signal in the corresponding direction. The Empirical Mode Decomposition algorithm is subsequently utilized to suppress the environmental noise and time domain artifacts caused by the enhancement method, enabling the extraction of the target's respiration pat-tern. Our results show that the proposed approach consistently achieves an average absolute error of less than 0.3 breaths per minute across all positions, significantly outperforming the baseline approaches.
This paper proposes a composite preamble structure based on the fusion of the original Ultra Wide Band (UWB) preamble structure and the Nonlinear Frequency Modulated (NLFM) pulse signal to improve the range-sensing ability of the UWB signal. We focus on improving the range-sensing performance of UWB signals with composite preamble structure. The short-time NLFM signal is utilized in the composite preamble structure to locate the target, so as to lower the target misjudgment rate. Simulation results show that the UWB signal with a composite preamble structure improves the target detection accuracy and reduces probability of target misjudgment, outperforming the original UWB signal.
Multimodal human activity recognition attracts wide attention in human-computer interaction. However, in the collected multimodal signals, not all modal signals contain useful feature information; some irrelevant and redundant information may negatively impact the model's performance, reducing the accuracy of activity recognition. This paper designs a self-attention mechanism-based multimodal fusion network for combining Wi-Fi signals and image streams based on video signals. The self-attention mechanism possesses the capability to capture spatio-temporal local features within multimodal signals. It dynamically learns the weights of different modalities, assigning higher weights to relatively important modalities. This process effectively fuses features extracted from individual modalities, resulting in a more comprehensive feature set. Through extensive experiments, we evaluate the performance of the proposed multimodal human activity recognition network from various perspectives. The experimental results indicate the effectiveness of the proposed method for data collected in different scenarios.
Deep joint source-channel coding (DeepJSCC) has shown promise in wireless transmission of text, speech, and images within the realm of semantic communication. However, wireless video transmission presents greater challenges due to the difficulty of extracting and compactly representing both spatial and temporal features, as well as its significant bandwidth and computational resource requirements. In response, we propose a novel video DeepJSCC (VDJSCC) approach to enable end-to-end video transmission over a wireless channel. Our approach involves the design of a multi-scale vision Transformer encoder and decoder to effectively capture spatial-temporal representations over long-term frames. Additionally, we propose a dynamic token selection module to mask less semantically important tokens from spatial or temporal dimensions, allowing for content-adaptive variable-length video coding by adjusting the token keep ratio. Experimental results demonstrate the effectiveness of our VDJSCC approach compared to digital schemes that use separate source and channel codes, as well as other DeepJSCC schemes, in terms of reconstruction quality and bandwidth reduction.
Speckle patterns generated by the intermodal interference of multimode fibers enable accurate broadband wavelength measurements. However, the measurement speed is limited by the frame rate of the camera that captures the patterns. We propose a compact and cost-effective ultrafast wavemeter based on multimode and multicore fibers, which employs spectral -spatial -temporal mapping. The speckle patterns generated by multimode fibers enable spectral-to-spatial mapping, which is then sampled by a multicore fiber into a pulse sequence to implement spatial-to-temporal mapping. A high-speed single-pixel photodetector is employed to capture the pulse sequence, which is analysed using a multilayer perceptron to estimate the wavelength. The feasibility of the proposed wavelength estimation method is experimentally verified, achieving a measurement rate of 100 MHz with a resolution of 2.7 pm in a 1 nm operation bandwidth.