Backscatter communications, an emerging low-cost and low-consumption green communication paradigm, has received significant attention in both industry and academia. In recent years, multiple-input-multiple-output (MIMO) technology has also been integrated into backscatter communications to meet the requirements of transmission performance in Internet of Things (IoT) scenarios. However, due to the hardware limitations of backscatter tags, one of the main challenges for multiantenna backscatter tags is to reduce the circuit complexity of space-time block codes (STBCs) while maintaining transmission performance. Although low-complexity STBCs for dual-antenna tags have been explored, to the best of our knowledge, research on high-performance, low-complexity STBCs for four-antenna tags remains in its infancy. In this article, we propose a 2x4 extended Alamouti code (EAC) with full-rate capability for MIMO backscatter communications. The results show that the required impedance of the proposed EAC on the backscatter tag is less than half of that of the conventional orthogonal STBC (OSTBC), which considerably reduces the complexity of the tag circuit. Additionally, we also derive the asymptotic closed-form expression of the symbol error rate (SER) of the proposed EAC to obtain the system insights. The derivation results show that the achievable diversity order of the proposed EAC in the 1x 4x N backscatter channel is 2xmin (2,N) , which indicates that the proposed EAC can achieve the same diversity performance as the conventional OSTBC by setting an appropriate number of receiver antennas. Finally, numerical simulations are performed to validate the accuracy of our analysis and the superiority of the proposed EAC in the SER.
The joint design of block-level unitary query (BUTQ) and orthogonal space-time block code (OSTBC) has been proposed to considerably improve system performance of MIMO backscatter communications over Rayleigh fading to satisfy the transmission performance requirements of various internet of things scenarios. However, the performance of BUTQ-OSTBC over Nakagami-m fading, a more general fading, is still unknown. In this paper, we derive the asymptotic closed-form expressions for the symbol error rate (SER) of BUTQ-OSTBC over Nakagami-m fading to obtain system insights. The derivation results showed that the achievable diversity order of BUTQ-OSTBC in M × L × N backscatter channels is L × min(Mmd, Nmu). Furthermore, we also investigate the maximum achievable diversity order of MIMO backscatter communications. Interestingly, we find that the maximum achievable diversity order matches exactly the diversity order achieved by BUTQ-OSTBC, i.e., the maximum achievable diversity order of MIMO backscatter communications is also L × min(Mmd, Nmu). Numerical simulations are also conducted to verify the accuracy of our derivation.
In the rapidly advancing realm of wireless communication, device-to-device (D2D) technology, an emerging approach for data exchange and connectivity, has been attracting increasing attention. Unmanned Aerial Vehicles (UAVs) can act as air relays or base stations, and integrate isolated D2D clusters into a cohesive network fabric in outdoor environments. However, in complex terrain, the communication signals are subject to irregular attenuation, and the signal propagation attenuation of different frequency bands in the same terrain is inconsistent. It is challenging to utilize UAVs to coverage D2D terrestrial users in complex terrain. In this paper, we propose the UAVs relaying for bridging the terrestrial D2D networks assisted by multi-frequency radio maps. From the real-world topographical data, we generate multi-frequency radio maps, which represent the distortion of different frequency band signals by rich information about land layouts. Next, we focus on the air-to-ground D2D network topology and formulate it into an optimization problem. Then, we decompose it into two subproblems. The first subproblem pertains to the design of the ground network structure. We employ the D2D frequency band radio map to assess the communication quality between user pairs, and propose a measure of D2D closeness centrality to select ‘cellular users’ that can communicate directly to a UAV. The second subproblem involves the UAVs’ deployment and the frequency selection. We present a multi-frequency radio map improved k-means method, which has lower algorithm complexity than the traversal method by reducing the utilization of the radio maps. Simulations validate the proposed scheme, demonstrating that: 1. Multi-frequency radio maps can provide efficient gains with real-world complex topography; 2. The proposed network structure and algorithm outperform other existing approaches.
The simultaneous quantum and classical communication (SQCC) system implements continuous variable quantum key distribution (CVQKD) and classical communication using the same basic communication facilities. However, the SQCC system has a very low tolerance for phase noise, which affects its transmission distance and secret key rate under the local local oscillator (LO) design. In order to reduce the phase noise in the SQCC system, this paper proposes a method to automatically compensate the signal phase based on the long-short-term memory network (LSTM) model. Firstly, the LSTM model is trained to predict the phase value of the reference pulse during operation. Then the predicted value of the LSTM model to compensate for the phase drift of the quantum signal, thus reducing the phase noise. The results demonstrate that the automatic phase compensation method based on LSTM exhibits excellent prediction performance and compensation accuracy, which can significantly improve the transmission distance and secret key rate of the SQCC system without requiring any additional quantum resources and extra experimental hardware.
Efficient 3D LiDAR point cloud compression (LPCC) and streaming are critical for edge server-assisted robotic systems, enabling real-time communication with compact data representations. A widely adopted approach represents LiDAR point clouds as range images, enabling the direct use of mature image and video compression codecs. However, because these codecs are designed with human visual perception in mind, they often compromise geometric details, which downgrades the performance of downstream robotic tasks such as mapping and object detection. Furthermore, rate-distortion optimization (RDO)-based rate control remains largely underexplored for range image compression (RIC) under dynamic bandwidth conditions. To address these limitations, we propose D-Compress, a new detail-preserving and fast RIC framework tailored for real-time streaming. D-Compress integrates both intra- and inter-frame prediction with an adaptive discrete wavelet transform approach for precise residual compression. Additionally, we introduce a new RDO-based rate control algorithm for RIC through new rate-distortion modeling. Extensive evaluations on various datasets demonstrate the superiority of D-Compress, which outperforms state-of-the-art (SOTA) compression methods in both geometric accuracy and downstream task performance, particularly at compression ratios exceeding 100x, while maintaining real-time execution on resource-constrained hardware. Moreover, evaluations under dynamic bandwidth conditions validate the robustness of its rate control mechanism.
This paper proposes an integrated sensing and communication (ISAC)-enabled grant-free uplink framework based on artificial-path delay modulation. A grant-free user equipment (g-UE) conveys uplink information by modulating the delay of a controllable artificial path derived from the scheduled downlink waveform. In contrast to conventional superposition-based schemes with successive interference cancellation, the proposed method enables uplink-downlink coexistence in the delay-sensing domain. By introducing a single weak artificial path confined within the cyclic prefix (CP), the g-UE allows the access point (AP) to decode uplink symbols from CSI perturbations while causing only limited degradation to the scheduled user equipment (s-UE) in the downlink. To support reliable finite-alphabet delay detection under unknown path gain and off-grid leakage, we develop a baseline delay calibration procedure and a normalized matched-filter detector. Results show that reflection power determines the reliability trade-off between the g-UE and the s-UE, whereas the delay step mainly controls the g-UE reliability-efficiency trade-off with little additional impact on the downlink s-UE. Even with an artificial path 15 dB weaker than the scheduled downlink signal, the g-UE achieves lower BER than the s-UE at an effective modulation order of 16-QAM. The proposed framework thus offers a low-complexity, SIC-free, and downlink-friendly solution for grant-free uplink in ISAC systems.
Orthogonal time frequency space (OTFS) modulation has demonstrated significant advantages in high-mobility scenarios in future 6G networks. However, existing channel estimation methods often overlook the structured sparsity and clustering characteristics inherent in realistic clustered delay line (CDL) channels, leading to degraded performance in practical systems. To address this issue, we propose a novel nonparametric Bayesian learning (NPBL) framework for OTFS channel estimation. Specifically, a stick-breaking process is introduced to automatically infer the number of multipath components and assign each path to its corresponding cluster. The channel coefficients within each cluster are modeled by a Gaussian mixture distribution to capture complex fading statistics. Furthermore, an effective pruning criterion is designed to eliminate spurious multipath components, thereby enhancing estimation accuracy and reducing computational complexity. Simulation results demonstrate that the proposed method achieves superior performance in terms of normalized mean squared error compared to existing methods.
To make cross-band channel prediction practical for AI-native RAN, algorithms must generalize across diverse environments and support real-time inference. Existing approaches achieve one but not both. To bridge this gap, we introduce GUIDE, a physics-guided deep unfolding framework that embeds wireless channel physics into differentiable layers. Without retraining in unseen environments, GUIDE achieves 2.75x beamforming gain than the deep learning-based baseline FIRE with only a slight increase in inference time, and 1.39x beamforming gain than the strongest model-based baseline R2F2 while running over 1610x faster.
Physical layer key generation (PLKG) has emerged as a promising solution for achieving highly secured and low-latency key distribution, offering information-theoretic security that is inherently resilient to quantum attacks. However, simultaneously ensuring a high data transmission rate and a high secret key generation rate under eavesdropping attacks remains a major challenge. In time-division duplex (TDD) systems with multiple antennas, we derive closed-form expressions for both rates by modeling the legitimate channel as a time-correlated autoregressive (AR) process. This formulation leads to a highly nonconvex and time-coupled optimization problem, rendering traditional optimization methods ineffective. To address this issue, we propose a multi-agent soft actor-critic (SAC) framework equipped with a long short-term memory (LSTM) adversary prediction module to cope with the partial observability of the eavesdropper's mode. Simulation results demonstrate that the proposed approach achieves superior performance compared with other benchmark algorithms, while effectively balancing the trade-off between secret key generation rate and data transmission rate. The results also confirm the robustness of the proposed framework against intelligent eavesdropping and partial observation uncertainty.
Transparent liquid manipulation in robotic pouring remains challenging for perception systems: specular/refraction effects and lighting variability degrade visual cues, undermining reliable level estimation. To address this challenge, we introduce RadarEye, a real-time mmWave radar signal processing pipeline for robust liquid level estimation and tracking during the whole pouring process. RadarEye integrates (i) a high-resolution range-angle beamforming module for liquid level sensing and (ii) a physics-informed mid-pour tracker that suppresses multipath to maintain lock on the liquid surface despite stream-induced clutter and source container reflections. The pipeline delivers sub-millisecond latency. In real-robot water-pouring experiments, RadarEye achieves a 0.35 cm median absolute height error at 0.62 ms per update, substantially outperforming vision and ultrasound baselines.
WiFi sensing faces a critical reliability challenge due to hardware-induced RF distortions, especially with modern, market-dominant WiFi cards supporting 802.11ac/ax protocols. These cards employ sensitive automatic gain control and separate RF chains, introducing complex and dynamic distortions that render existing compensation methods ineffective. In this paper, we introduce Domino, a new framework that transforms channel state information (CSI) into channel impulse response (CIR) and leverages it for precise distortion compensation. Domino is built on the key insight that hardware-induced distortions impact all signal paths uniformly, allowing the dominant static path to serve as a reliable reference for effective compensation through delay-domain processing. Real-world respiration monitoring experiments show that Domino achieves at least 2× higher mean accuracy over existing methods, maintaining robust performance with a median error below 0.24 bpm, even using a single antenna in both direct line-of-sight and obstructed scenarios.
WiFi sensing based on channel state information (CSI) collected from commodity WiFi devices has shown great potential across a wide range of applications, including vital sign monitoring and indoor localization. Existing WiFi sensing approaches typically estimate motion information directly from CSI. However, they often overlook the inherent advantages of channel impulse response (CIR), a delay-domain representation that enables more intuitive and principled motion sensing by naturally concentrating motion energy and separating multipath components. Motivated by this, we revisit WiFi sensing and introduce CIRSense, a new framework that enhances the performance and interpretability of WiFi sensing with CIR. CIRSense is built upon a new motion model that characterizes fractional delay effects, a fundamental challenge in CIR-based sensing. This theoretical model underpins technical advances for the three challenges in WiFi sensing: hardware distortion compensation, high-resolution distance estimation, and subcarrier aggregation for extended range sensing. CIRSense, operating with a 160 MHz channel bandwidth, demonstrates versatile sensing capabilities through its dual-mode design, achieving a mean error of approximately 0.25 bpm in respiration monitoring and 0.09 m in distance estimation. Comprehensive evaluations across residential spaces, far-range scenarios, and multi-target settings demonstrate CIRSense's superior performance over state-of-the-art CSI-based baselines. Notably, at a challenging sensing distance of 20 m, CIRSense achieves at least 3x higher average accuracy with more than 4.5x higher computational efficiency.
In current beamforming multi-objective optimization algorithms, parameter adjustment to satisfy the target pattern given the target pattern is essential. In order to avoid the influence of experience factors on the optimization target weight parameters, we proposed for the first time an intelligent optimization method that combines the optimization algorithm with the RL algorithm to assist in adjust the weight parameters. For different antenna array models, this method can also achieve excellent array beamforming. This article compares the results of PSO with auxiliary adjustment and the results of classic PSO without auxiliary adjustment on three array models and three groups of beam targets, proving the versatility and effectiveness of the algorithm for auxiliary adjustment of intelligent optimization algorithms.
This paper focuses on optimizing the long-term average age of information (AoI) in device-to-device (D2D) networks through age-aware link scheduling. The problem is naturally formulated as a Markov decision process (MDP). However, finding the optimal policy for the formulated MDP in its original form is challenging due to the intertwined AoI dynamics of all D2D links. To address this, we employ the Lyapunov optimization framework to develop a dynamic age-aware scheduling policy. Specifically, we explore two scenarios: known statistical channel state information (CSI) and known instantaneous CSI. For the statistical CSI case, we propose a message passing neural network (MPNN)-based policy for real-time scheduling. The MPNN is trained in an unsupervised manner with a loss function designed to minimize per-slot Lyapunov drift. For the instantaneous CSI case, we introduce a Gurobi-based policy, using the solver to minimize per-slot Lyapunov drift for scheduling decisions. To further reduce the computational complexity, we also propose a greedy heuristic policy that approximates drift minimization. Extensive simulation results show that our proposed age-aware scheduling policies have superior performance compared to the baselines, and can be applied to large-scale D2D networks.
Transparent objects are prevalent in everyday environments, but their distinct physical properties pose significant challenges for camera-guided robotic arms. Current research is mainly dependent on camera-only approaches, which often falter in suboptimal conditions, such as low-light environments. In response to this challenge, we present FuseGrasp, the first radar-camera fusion system tailored to enhance the transparent objects manipulation. FuseGrasp exploits the weak penetrating property of millimeter-wave (mmWave) signals, which causes transparent materials to appear opaque, and combines it with the precise motion control of a robotic arm to acquire high-quality mmWave radar images of transparent objects. The system employs a carefully designed deep neural network to fuse radar and camera imagery, thereby improving depth completion and elevating the success rate of object grasping. Nevertheless, training FuseGrasp effectively is non-trivial, due to limited radar image datasets for transparent objects. We address this issue utilizing large RGB-D dataset, and propose an effective two-stage training approach: we first pre-train FuseGrasp on a large public RGB-D dataset of transparent objects, then fine-tune it on a self-built small RGB-D-Radar dataset. Furthermore, as a byproduct, FuseGrasp can determine the composition of transparent objects, such as glass or plastic, leveraging the material identification capability of mmWave radar. This identification result facilitates the robotic arm in modulating its grip force appropriately. Extensive testing reveals that FuseGrasp significantly improves the accuracy of depth reconstruction and material identification for transparent objects. Moreover, real-world robotic trials have confirmed that FuseGrasp markedly enhances the handling of transparent items. A video demonstration of FuseGrasp is available at https://youtu.be/MWDqv0sRSok.
When acting as the dynamic aerial millimeter-wave (mmWave) base station, unmanned aerial vehicles (UAVs) can provide ubiquitous service to ground moving users (UEs). In this paper, we propose a low cost anti-blockage UAV 3D trajectory design algorithm that enables real-time blockage prediction and trajectory optimization without detailed building geometric information or additional equipment, with the aim of maximizing the throughput of UEs. Specifically, the UAV is conducted to collect UE locations and corresponding blockage status data and update the blockage prediction model periodically. A joint optimization algorithm is designed, integrating the soft actorcritic (SAC) method with the blockage prediction model. The SAC based algorithm alternates between blockage prediction and position optimization, effectively planning the UAV 3D path to improve UE throughput. This ensures that the UAV's position optimization aligns more closely with the actual blockage environment, thereby significantly enhancing the system's anti-blockage capabilities. The results of numerical simulation demonstrate that the proposed algorithm can considerably improve the accuracy of blockage prediction and throughput of UEs.
Robotic perception has been significantly advanced by integrating vision with additional sensing modalities such as acoustic and tactile sensors. However, existing methods largely emphasize external object properties, including appearance and geometry, while neglecting internal material attributes (e.g., composition) that are crucial for robust and reliable robotic manipulation. In this work, we propose augmenting robotic systems with radar sensing and introduce CRMaterial, a new camera-radar fusion framework powered by vision language models (VLMs) for accurate object material identification. Preliminary experiments demonstrate that our system improves material identification accuracy by 2.5x compared to a camera-only baseline.
Service robots often engage in liquid-related tasks such as pouring, which require accurate liquid identification. Vision-language models (VLM) exhibit strong performance in general object recognition, they struggle to reliably distinguish between visually similar liquids, limiting their effectiveness in such scenarios. As a promising solution, millimeter-wave (mmWave) radar combined with neural networks has been widely adopted for material recognition and liquid identification. However, training a radar-based liquid classifier demands a large volume of labeled radar-liquid data pairs and often faces challenges in generalizing to real-world environments. To address this challenge, we propose FuseLID, a VLM-based camera-radar late-fusion system designed to improve liquid identification performance. FuseLID leverages the weak penetration capability of mmWave signals through liquids and combines this with the robot's precise motion control to capture distinctive radar signatures of various liquids. Subsequently, a VLM-based late-fusion module is designed to combine camera and radar outputs for achieving enhanced liquid identification with limited radar-liquid data. Preliminary experiments show that FuseLID improves liquid identification accuracy from 51.6% to 95.7% when classifying several commonly consumed beverages, including visually similar ones.
WiFi sensing has attracted significant attention over the past decade for its potential to enable ubiquitous human activity monitoring. While prior approaches primarily rely on stationary transceivers, this paper is the first to investigate the feasibility of WiFi sensing on mobile robots. We introduce CornerSense, a new framework that empowers mobile robots to leverage their WiFi interfaces for detecting human proximity around corners, effectively addressing the non-line-of-sight challenges commonly encountered by traditional robotic sensors such as cameras and LiDAR. Our analysis demonstrates that the design of CornerSense fundamentally differs from, and is more challenging than, conventional WiFi sensing with stationary transceivers. In particular, the intertwined movement of the robot and nearby humans complicates the isolation of human-induced signal variations, which is critical for accurate proximity detection. To address this challenge, CornerSense develops a virtual path-augmented, two-stage dominant path extraction approach based on principal component analysis (PCA). In the first stage, through strategically introducing a virtual path, CornerSense extracts a reference path that incorporates only the robot's motion by applying PCA on the power of the virtual path-augmented channel state information (aug-CSI). During the second stage, the reference path is first subtracted from the aug-CSI. This subtraction allows for a second application of PCA to the power of the residual aug-CSI, thereby enabling the extraction of another distinct dominant path that is reflected off the human body before arriving at the robot. This path is referred to as the dominant human-reflected path. Finally, the reference path extracted in the first stage is employed to compensate for the hardware-induced phase offset in the dominant human-reflected path, yielding cleaned CSI ready for accurate and robust detection of human proximity. Real-world experimental evaluations conducted across nine different corners in three groups and under four distinct human walking patterns reveal that CornerSense achieves an average true positive rate (TPR) of 96% while maintaining a low average false positive rate (FPR) of 3%. In contrast, a baseline system that directly applies an algorithm intended for stationary transceiver-based proximity sensing only reaches a TPR of 84% and suffers from a markedly higher FPR of 46%, which is over an order of magnitude higher than that of CornerSense.
The range image has emerged as a dominant representation of 3D LiDAR data, enabling the direct application of well-established image and video compression techniques. However, existing compression methods, primarily optimized for human visual perception, often compromise the fidelity of physical distance information embedded in range images, which is critical for downstream robotic tasks. Additionally, rate-distortion optimization (RDO)-based rate control remains largely unexplored in range image-based LPCC. To address these limitations, we introduce D-Compress, a new framework for detail-preserving, fast, and precise compression of LiDAR range images tailored for real-time streaming. D-Compress focuses on preserving fine-grained range image details while achieving both high compression speed and geometric accuracy. Compared to state-of-the-art (SOTA) codecs, our approach delivers superior geometric precision, high compression ratios, and robust rate control.