Achieving zero-shot adversarial robustness without sacrificing generalization remains challenging for foundation models such as CLIP, especially under large adversarial perturbations. Through empirical analyses, we identify three critical yet overlooked issues: (1) Logit margins exhibit a stable offset between small and large adversarial perturbations, suggesting that explicitly adjusting margins could improve robustness against unseen large perturbations. (2) A significant negative correlation exists between logit margin and inter-class semantic similarity, indicating that semantic structures are insufficiently leveraged by existing methods. (3) Existing methods for adjusting text embeddings disrupt the intrinsic semantic consistency established by pre-trained models, undermining generalization capability. Motivated by these findings, we propose a novel Text-Image Mutual Awareness (TIMA) framework, including a Text-Aware Image (TAI) tuning module with an Adaptive Semantic-Aware Margin (ASAM) to explicitly calibrate logit margins, and an Image-Aware Text (IAT) tuning module with Semantic Consistent Minimum Hyperspherical Energy (SC-MHE) to preserve semantic consistency. Comprehensive experiments validate that TIMA significantly outperforms existing approaches by effectively addressing the identified limitations.
This work considers the scenario of using reconfigurable intelligent surface (RIS) to synthesize a massive MIMO downlink channel. This architecture has the potential of realizing massive MIMO with low hardware requirements at the base station (BS). We focus on multi-user downlink transmission, wherein the BS serves a group of primary users, while the RIS serves a group of secondary users and helps to improve the primary users’ service quality at the same time. The challenge lies in the need to address a large-scale discrete phase precoding problem, which is difficult to solve from an optimization viewpoint. We study an alternative approach using spatial Sigma-Delta (Σ∆) modulation, which shapes the quantization noise away from the users. The resultant signal model allows us to alleviate discrete optimization in designing the RIS coefficients. The problem turns out to be convex, and can be efficiently solved. Empirical results show that the proposed scheme works well, demonstrating that using RIS to synthesize massive MIMO is feasible.
The development of millimeter-wave and terahertz systems has pushed antenna arrays into the near-field region, where the conventional far-field (plane-wave) model becomes inaccurate. To implement such a large array with a large bandwidth, low-resolution digital-to-analog converters are unavoidable. Although both model mismatch and quantization have been studied individually, their combined impact remains poorly understood. This work investigates the loss in beamforming gain due to the quantization error and the channel model mismatch error in near-field scenarios. We derive analytical expressions for each type of loss and show that their relative dominance depends on system parameters such as user distance, angle, and quantization resolution. Our analysis reveals a crossover region, smaller than the Rayleigh distance, beyond which quantization error dominates, and far-field modeling becomes sufficiently accurate. Our results offer practical insights into when near-field modeling is beneficial and when it can be safely omitted in low-resolution systems.
This paper addresses the beam alignment problem for mono-static full-duplex backscatter communication systems. This is crucial for an access point (AP) to acquire data from a backscatter device (BD) without either information on the BD location or channel state information. Conventional beam alignment methods, including beam sweeping and channel estimation, require long preambles to achieve acceptable performance. To address this challenge, this work proposes an active sensing based method in which the AP actively interacts with the environment and uses historical observations collected thus far to adaptively adjust the beamforming vectors, thereby maximizing the signal-to-noise ratio (SINR) of the BD signal. Specifically, a long-short-term memory (LSTM) network is exploited to capture correlations over different observations, and deep neural networks are employed to map the hidden state to beamforming vectors. Simulation results demonstrate that the active sensing based method effectively optimizes the beamforming vectors that maximize SINR with only a few interactions with the environment, thus facilitating subsequent BD data transmission.
Initial beam alignment for analog arrays often relies on exhaustively sweeping beams over predefined codebooks when no prior information is available, leading to substantial pilot overhead before reliable data transmission and accurate localization can be achieved. This letter proposes a map-aided active beam-training framework for joint communication and localization in a single-user uplink system. During beam training, the BS exploits environmental geometry to adaptively design probing beams, thereby improving the communication-and-localization performance under a fixed pilot budget. The measurement policy is implemented by a long short-term memory (LSTM) network with a deep neural network (DNN) head that maps the history of observations to a beamforming vector, optimized under a tunable tradeoff between the localization Cramer-Rao bound (CRB) and the communication objective function. Simulation results show that the learned policy steers early beams to explore both direct and reflected paths and then concentrates energy on two paths simultaneously, achieving a higher rate and a lower CRB for the UE than baseline methods.
Dual-functional radar-communication (DFRC) signal design has received much attention lately. We consider the scenario of one-bit massive multi-input multi-output (MIMO) wherein one-bit DACs are employed for the sake of saving hardware costs. Specifically, a spatial Sigma-Delta $(\Sigma\Delta)$ modulation scheme is proposed for one-bit MIMO-DFRC waveform design. Unlike the existing approaches which require large-scale binary optimization, the proposed scheme performs $\Sigma\Delta$ modulation on a continuous-valued DFRC signal. The subsequent waveform design is formulated as a constrained least square problem, which can be efficiently solved. Moreover, we leverage quantization noise for radar probing purposes, rather than treating it as unwanted noise. Numerical results demonstrate that the proposed scheme performs well in both radar probing and downlink precoding.
Unlike existing methods that rely on source images as appearance references and use source speech to generate motion, this work proposes a novel approach that directly extracts information from the speech, addressing key challenges in speech-to-talking face. Specifically, we first employ a speech-to-face portrait generation stage, utilizing a speech-conditioned diffusion model combined with statistical facial prior and a sample-adaptive weighting module to achieve high-quality portrait generation. In the subsequent speech-driven talking face generation stage, we embed expressive dynamics such as lip movement, facial expressions, and eye movements into the latent space of the diffusion model and further optimize lip synchronization using a region-enhancement module. To generate high-resolution outputs, we integrate a pre-trained Transformer-based discrete codebook with an image rendering network, enhancing video frame details in an end-to-end manner. Experimental results demonstrate that our method outperforms existing approaches on the HDTF, VoxCeleb, and AVSpeech datasets. Notably, this is the first method capable of generating high-resolution, high-quality talking face videos exclusively from a single speech input.
Mobile users are prone to experience beam failure due to beam drifting in millimeter wave (mmWave) communications. Sensing can help alleviate beam drifting with timely beam changes and low overhead since it does not need user feedback. This work studies the problem of optimizing sensing-aided communication by dynamically managing beams allocated to mobile users. A multi-beam scheme is introduced, which allocates multiple beams to the users that need an update on the angle of departure (AoD) estimates and a single beam to the users that have satisfied AoD estimation precision. A deep reinforcement learning (DRL) assisted method is developed to optimize the beam allocation policy, relying only upon the sensing echoes. For comparison, a heuristic AoD-based method using approximated Cramér-Rao lower bound (CRLB) for allocation is also presented. Both methods require neither user feedback nor prior state evolution information. Results show that the DRL-assisted method achieves a considerable gain in throughput than the conventional beam sweeping method and the AoD-based method, and it is robust to different user speeds.
Extensive research on Reconfigurable Intelligent Surfaces (RIS) has primarily focused on optimizing reflective coefficients for passive beamforming in specific target directions. This optimization typically assumes prior knowledge of the target direction, which is unavailable before the target is detected. To enhance direction estimation, it is critical to develop array pattern synthesis techniques that yield a wider beam by maximizing the received power over the entire target area. Although this challenge has been addressed with active antennas, RIS systems pose a unique challenge due to their inherent phase constraints, which can be continuous or discrete. This work addresses this challenge through a novel array pattern synthesis method tailored for discrete phase constraints in RIS. We introduce a penalty method that pushes these constraints to the boundary of the convex hull. Then, the Minorization-Maximization (MM) method is utilized to reformulate the problem into a convex one. Our numerical results show that our algorithm can generate a wide beam pattern comparable to that achievable with per-power constraints, with both the amplitudes and phases being adjustable. We compare our method with a traditional beam sweeping technique, showing a) several orders of magnitude reduction of the MSE of Angle of Arrival (AOA) at low to medium Signal-to-Noise Ratio (SNR)s; and b) 8 dB SNR reduction to achieve a high probability of detection.
Deep learning draws heavily on the latest progress in semantic communications. The present paper aims to examine the security aspect of this cutting-edge technique from a novel shuffling perspective. Our goal is to improve upon the conventional secure coding scheme to strike a desirable tradeoff between transmission rate and leakage rate. To be more specific, for a wiretap channel, we seek to maximize the transmission rate while minimizing the semantic error probability under the given leakage rate constraint. Toward this end, we devise a novel semantic security communication system wherein the random shuffling pattern plays the role of the shared secret key. Intuitively, the permutation of feature sequences via shuffling would distort the semantic essence of the target data to a sufficient extent so that eavesdroppers cannot access it anymore. The proposed random shuffling method also exhibits its flexibility in working for the existing semantic communication system as a plugin. Simulations demonstrate the significant advantage of the proposed method over the benchmark in boosting secure transmission, especially when channels are prone to strong noise and unpredictable fading.
A major challenge in Fine-Grained Visual Classification (FGVC) is distinguishing various categories with high inter-class similarity by learning the feature that differentiates the details. Conventional cross-entropy trained Convolutional Neural Network (CNN) fails this challenge as they may suffer from producing inter-class invariant features in FGVC. In this work, we innovatively propose to regularize the training of CNN by enforcing the uniqueness of the features of each category from an information-theoretic perspective. To achieve this goal, we formulate a minimax loss based on a game-theoretic framework, where a Nash equilibrium is proved to be consistent with this regularization objective. Besides, to avoid getting a solution that produces redundant features, we present a Feature Redundancy Loss (FRL) based on the normalized inner product between each selected feature map pair to complement the proposed minimax loss. The proposed method is versatile, as it can be utilized as a regularizer for features in the mid-level or the penultimate layer, and can be combined with any architectures. Extensive experimental results on several influential benchmarks along with visualization show that our method obtains significant improvement over the baseline model without extra cost and achieves state-of-the-art results.
Fractional programming (FP) arises in various communications and signal processing problems because several key quantities in the field are fractionally structured, e.g., the Cramér-Rao bound, the Fisher information, and the signal-to-interference-plus-noise ratio (SINR). A recently proposed method called the quadratic transform has been applied to the FP problems extensively. The main contributions of the present paper are two-fold. First, we investigate how fast the quadratic transform converges. To the best of our knowledge, this is the first work that analyzes the convergence rate for the quadratic transform as well as its special case the weighted minimum mean square error (WMMSE) algorithm. Second, we accelerate the existing quadratic transform via a novel use of Nesterov's extrapolation scheme [1]. Specifically, by generalizing the minorization-maximization (MM) approach in [2], we establish a nontrivial connection between the quadratic transform and the gradient projection, thereby further incorporating the gradient extrapolation into the quadratic transform to make it converge more rapidly. Moreover, the paper showcases the practical use of the accelerated quadratic transform with two frontier wireless applications: integrated sensing and communications (ISAC) and massive multiple-input multiple-output (MIMO).
Extensive research on Reconfigurable Intelligent Surfaces (RIS) has primarily revolved around optimizing phase coefficients for passive beamforming in specific target directions. This optimization typically presumes prior knowledge of the target direction, a premise that is not valid during initial sensing phases where the target’s location is unknown. To enhance direction estimation, a broader beam is required, presenting the array pattern synthesis problem: the need to maximize received power over an entire target area. While this challenge has been addressed with active antennas, RIS systems pose a unique problem due to their inherent continuous or discrete phase constraints. Our study introduces a novel array pattern synthesis method tailored for discrete phase constraints in RIS. We introduce a penalty method that pushes these constraints to the boundary of the convex hull. Then the Minorization-Maximization (MM) method is utilized to reformulate the problem into a convex one. Our numerical results illustrate that our algorithm is capable of generating a wide beam pattern comparable to scenarios with per-power constraints where both the amplitudes and phases are adjustable, evidencing its proficiency in creating wide beams crucial for wireless communication and sensing in the absence of prior knowledge of the target’s location.
This study addresses the challenge of inaccurate gradients in computing the empirical Fisher Information Matrix during network pruning. We introduce SWAP, a formulation of Entropic Wasserstein regression (EWR) for network pruning, capitalizing on the geometric properties of the optimal transport problem. The “swap” of the commonly used linear regression with the EWR in optimization is analytically demonstrated to offer noise mitigation effects by incorporating neighborhood interpolation across data points with only marginal additional computational cost. The unique strength of SWAP is its intrinsic ability to balance noise reduction and covariance information preservation effectively. Extensive experiments performed on various networks and datasets show comparable performance of SWAP with state-of-the-art (SoTA) network pruning algorithms. Our proposed method outperforms the SoTA when the network size or the target sparsity is large, the gain is even larger with the existence of noisy gradients, possibly from noisy data, analog memory, or adversarial attacks. Notably, our proposed method achieves a gain of 6% improvement in accuracy and 8% improvement in testing loss for MobileNetV1 with less than one-fourth of the network parameters remaining.
This paper investigates the information theoretic limit of a reconfigurable intelligent surface (RIS) aided communication scenario in which the RIS and the transmitter either jointly or independently send information to the receiver. The RIS is an emerging technology that uses a large number of passive reflective elements with adjustable phases to intelligently reflect the transmit signal to the intended receiver. While most previous studies of the RIS focus on its ability to beamform and to boost the received signal-to-noise ratio (SNR), this paper shows that if the information data stream is also available at the RIS and can be modulated through the adjustable phases at the RIS, significant improvement in the {degree-of-freedom} (DoF) of the overall channel is possible. For example, for an RIS system in which the signals are reflected from a transmitter with $M$ antennas to a receiver with $K$ antennas through an RIS with $N$ reflective elements, assuming no direct path between the transmitter and the receiver, joint transmission of the transmitter and the RIS can achieve a DoF of $\min\left(M+\frac{N}{2}-\frac{1}{2},N,K\right)$ as compared to the DoF of $\min(M,K)$ for the conventional multiple-input multiple-output (MIMO) channel. This result is obtained by establishing a connection between the RIS system and the MIMO channel with phase noise and by using results for characterizing the information dimension under projection. The result is further extended to the case with a direct path between the transmitter and the receiver, and also to the multiple access scenario, in which the transmitter and the RIS send independent information. Finally, this paper proposes a symbol-level precoding approach for modulating data through the phases of the RIS, and provides numerical simulation results to verify the theoretical DoF results.
Transmitting data using the phases on reconfigurable intelligent surfaces (RIS) is a promising solution for future energy-efficient communication systems. Recent work showed that a virtual phased massive multiuser multiple-input-multiple-out (MIMO) transmitter can be formed using only one active antenna and a large passive RIS. In this paper, we are interested in using such a system to perform MIMO downlink precoding. In this context, we may not be able to apply conventional MIMO precoding schemes, such as the simple zero-forcing (ZF) scheme, and we typically need to design the phase signals by solving optimization problems with constant modulus constraints or with discrete phase constraints, which pose challenges in terms of incurring high computational costs. In this work, we propose an alternative approach based on Sigma-Delta (Σ∆) modulation, which is classically famous for its noise-shaping ability. Specifically, first-order Σ∆ modulation is applied in the spatial domain to handle phase quantization in generating constant envelope signals. Under some mild assumptions, the proposed phased Σ∆ modulator allows us to use the ZF scheme to synthesize the RIS reflection phases in a low complexity fashion. The proposed approach is empirically shown to achieve comparable bit error rate performance to the unquantized ZF scheme.
dArtificial intelligence (AI) is critical in evolving 5G and developing 6G networks, running on edge devices, and solving resource management challenges. The burgeoning number of edge devices draws attention to the potential of low-earth orbit (LEO) satellite networks with their onboard computing capabilities for edge inference. This paper explores LEO scenarios where multiple remote sensing edge AI inference tasks concurrently process data from a single source. However, due to there being parts with the same functions between different AI applications, traditional monolithic edge AI architecture must be deployed repeatedly and falls short in efficiently harnessing the heterogeneous resources of LEO satellite networks. To solve this problem, we utilize the microservice architecture to decouple a single AI application into several independent microservices to reuse these same functions. However, due to the high latency caused by multiple microservices' communication, we need to design a deployment strategy to fully utilize resources to reduce the service latency. We present a microservice deployment model to minimize the total service latency across all AI applications and meet resource constraints with the constraints of hardware, energy, and memory limitations. This latency optimization problem is rewritten as a Markov decision process (MDP) to effectively deal with the challenge posed by the time-varying transmission rate caused by satellite mobility. To increase the training data utilization, we employ a Proximal Policy Optimization (PPO) based reinforcement learning algorithm to meet the dynamic environment challenge. Finally, we obtain a sub-optimal solution with minimal accuracy loss and an acceptable solution time.
Speech-to-face generation is an intriguing area of research that focuses on generating realistic facial images based on a speaker's audio speech. However, state-of-the-art methods employing GAN-based architectures lack stability and cannot generate realistic face images. To fill this gap, we propose a novel speech-to-face generation framework, which leverages a Speech-Conditioned Latent Diffusion Model, called SCLDM. To the best of our knowledge, this is the first work to harness the exceptional modeling capabilities of diffusion models for speech-to-face generation. Preserving the shared identity information between speech and face is crucial in generating realistic results. Therefore, we employ contrastive pre-training for both the speech encoder and the face encoder. This pre-training strategy facilitates effective alignment between the attributes of speech, such as age and gender, and the corresponding facial characteristics in the face images. Furthermore, we tackle the challenge posed by excessive diversity in the synthesis process caused by the diffusion model. To overcome this challenge, we introduce the concept of residuals by integrating a statistical face prior to the diffusion process. This addition helps to eliminate the shared component across the faces and enhances the subtle variations captured by the speech condition. Extensive quantitative, qualitative, and user study experiments demonstrate that our method can produce more realistic face images while preserving the identity of the speaker better than state-of-the-art methods. Highlighting the notable enhancements, our method demonstrates significant gains in all metrics on the AVSpeech dataset and Voxceleb dataset, particularly noteworthy are the improvements of 32.17 and 32.72 on the cosine distance metric for the two datasets, respectively.
Fractional programming (FP) plays an important role in information science because of the Cramer-Rao bound,the Fisher information, and the signal-to-interference-plus-noise ratio (SINR). A state-of-the-art method called the quadratic transform has been extensively used to address the FP problems. This work aims to accelerate the quadratic transform-based iterative optimization via gradient projection and extrapolation. The main contributions of this work are three-fold. First, we relate the quadratic transform to the gradient projection, thereby eliminating the matrix inverse operation from the iterative optimization; our result generalizes the weighted sum-of-rates (WSR) maximization algorithm in [1] to a wide range of FP problems. Second, based on this connection to gradient projection, we incorporate Nesterov's extrapolation strategy [2] into the quadratic transform so as to accelerate the convergence of the iterative optimization. Third, from a minorization-maximization (MM) point of view, we examine the convergence rates of the conventional quadratic transform methods--which include the weighted minimum mean square error (WMMSE) algorithm as a special case--and the proposed accelerated ones. Moreover, we illustrate the practical use of the accelerated quadratic transform in two popular application cases of future wireless networks: (i) integrated sensing and communication (ISAC) and (ii) massive multiple-input multiple-output (MIMO).
Channel acquisition is one of the main challenges for the deployment of reconfigurable intelligent surface (RIS) aided communication systems. This is because an RIS has a large number of reflective elements, which are passive devices with no active transmitting/receiving abilities. In this paper, we study the channel estimation problem for the RIS aided multi-user millimeter-wave (mmWave) multi-input multi-output (MIMO) system. Specifically, we propose a novel channel estimation protocol for the above system to estimate the cascaded channels, which are the products of the channels from the base station (BS) to the RIS and from the RIS to the users. Further, since the cascaded channels are typically sparse, this allows us to formulate the channel estimation problem as a sparse recovery problem using compressive sensing (CS) techniques, thereby allowing the channels to be estimated with less training overhead. Moreover, the sparse channel matrices of the cascaded channels of all users have a common block sparsity structure due to the common channel between the BS and the RIS. To take advantage of the common sparsity pattern, we propose a two-step multi-user joint channel estimation procedure. In the first step, we make use of the common column-block sparsity and project the received signals onto the common column subspace. In the second step, we make use of the row-block sparsity of the projected signals and propose a multi-user joint sparse matrix recovery algorithm that takes into account the common channel between the BS and the RIS.