This letter investigates a fluid antenna (FA)-assisted integrated sensing and communication (ISAC) system, with joint antenna position optimization and waveform design. We consider enhancing the sum-rate maximization (SRM) and sensing performance with the aid of FAs. Although the introduction of FAs brings more degrees of freedom for performance optimization, its position optimization poses a non-convex programming problem and brings great computational challenges. This letter contributes to building an efficient design algorithm by the block successive upper bound minimization and majorization-minimization principles, with each step admitting closed-form update for the ISAC waveform design. In addition, the extrapolation technique is exploited further to speed up the empirical convergence of FA position design. Simulation results show that the proposed design can achieve state-of-the-art sum-rate performance with at least 60% computation cutoff compared to existing works with successive convex approximation (SCA) and particle swarm optimization (PSO) algorithms.
In this paper, we characterize the generalized degrees-of-freedom (GDoF) region of the two-user $(M,N_{1},N_{2})$ multiple-input multiple-output (MIMO) broadcast channel with delayed channel state information at the transmitter (CSIT), where there are one transmitter with $M$ antennas and two receivers with $N_{1}$ and $N_{2}$ antennas, respectively. Under delayed CSIT, different from the existing converse approaches in the multiple-input single-output (MISO) GDoF and MIMO degrees-of-freedom (DoF) models, we incorporate new components into traditional approaches for this MIMO GDoF converse. For the achievability, we generalize the existing MISO achievable scheme. Our result reveals how the channel strength and antenna configuration impact the GDoF region of the two-user MIMO broadcast channel with delayed CSIT. Furthermore, the extension of our converse to a GDoF outer region of the $K$ -user MIMO broadcast channel with delayed CSIT is also provided.
This letter investigates joint beamforming and power optimization in a vehicle-to-infrastructure (V2I) system for integrated passive sensing and communications techniques. Specifically, the base station (BS) with multiple antennas delivers downlink data to multiple vehicles, respectively, and a sensing receiver simultaneously utilizes the downlink data signal to sense the motion of one vehicle. The sensing receiver is deployed separately. It collects the line-of-sight (LoS) signals from the BS and the scattered signals via the target vehicle, such that the distance and velocity of the target vehicle can be detected in a passive manner. As a result, the joint downlink beamforming and power optimization at the BS would be able to maximize the weighted summation of downlink data rates, subject to constraints on the signal-to-interference-plus-noise ratios (SINRs) of the two signals for passive sensing. In order to solve the above non-convex optimization problem, we first derive the optimal receiving beams for the data and sensing receivers given arbitrary transmission beam design, and then propose a low-complexity iterative algorithm to find a sub-optimal design of transmission beams based on the successive convex approximation (SCA) method. Simulations demonstrate the good performance, fast convergence and useful design insights.
An active reconfigurable intelligent surface (RIS) has been shown to be able to enhance the sum-of-degrees-of-freedom (DoF) of a two-user multiple-input multiple-output (MIMO) interference channel (IC) with equal number of antennas at each transmitter and receiver. However, for any number of receive and transmit antennas, when and how an active RIS can help to improve the sum-DoF are still unclear. This paper studies the sum-DoF of an active RIS-assisted two-user MIMO IC with arbitrary antenna configurations. In particular, RIS beamforming, transmit zero-forcing, and interference decoding are integrated together to combat the interference problem. In order to maximize the achievable sum-DoF, an integer optimization problem is formulated to optimize the number of eliminating interference links by RIS beamforming. As a result, the derived achievable sum-DoF can be higher than the sum-DoF of two-user MIMO IC, leading to a RIS gain. Furthermore, a sufficient condition of the RIS gain is given as the relationship between the number of RIS elements and the antenna configuration.
This article investigates a distant proactive eavesdropping system in cooperative cognitive radio (CR) networks. Specifically, an amplify-and-forward (AF) full-duplex (FD) secondary transmitter assists to relay the received signal from suspicious users to legitimate monitor for wireless information surveillance. In return, the secondary transmitter is granted to share the spectrum belonging to the suspicious users for its own information transmission. To improve the eavesdropping, the transmitted secondary user’s (SU) signal can also be used as a jamming signal to moderate the data rate of the suspicious link. We consider two cases, i.e., nonnegligible processing delay (NNPD) and negligible processing delay (NPD) at the secondary transmitter. Our target is to maximize network energy efficiency (NEE) via jointly optimizing the AF relay matrix and precoding vector at the secondary transmitter, as well as the receiver combining vector at the monitor, subject to the maximum power constraint at the secondary transmitter and minimum data rate requirement of the SU. We also guarantee that the achievable data rate of the eavesdropping link should be no less than that of the suspicious link for efficient surveillance. Due to the nonconvexity of the formulated NEE maximization problem, we develop an efficient path-following algorithm and a robust alternating optimization (AO) method as solutions under perfect and imperfect channel state information (CSI) conditions, respectively. We also analyze the convergence and computational complexity of the proposed schemes. Numerical results are provided to validate the effectiveness of our proposed schemes.
Despite the widespread utilization of deep neural networks (DNNs) for speech emotion recognition (SER), they are severely restricted due to the paucity of labeled data for training. Recently, segment-based approaches for SER have been evolving, which train backbone networks on shorter segments instead of whole utterances, and thus naturally augments training examples without additional resources. However, one core challenge remains for segment-based approaches: most emotional corpora do not provide ground-truth labels at the segment level. To supervisely train a segment-based emotion model on such datasets, the most common way assigns each segment the corresponding utterance’s emotion label. However, this practice typically introduces noisy (incorrect) labels as emotional information is not uniformly distributed across the whole utterance. On the other hand, DNNs have been shown to easily over-fit a dataset when being trained with noisy labels. To this end, this work proposes a simple and effective iterative self-learning (ISL) framework, which comprises a procedure to progressively correct segment-level labels in an iterative learning manner. The ISL method produces dynamically-generated and soft emotion labels, leading to significant performance improvements. Experiments on three well-known emotional corpora demonstrate noticeable gains using the proposed method.
This paper investigates a distant proactive eavesdropping system in cooperative cognitive radio (CR) networks. Specifically, an amplify-and-forward (AF) full-duplex (FD) secondary transmitter assists to relay the received signal from suspicious users to legitimate monitor for wireless information surveillance. In return, the secondary transmitter is granted to share the spectrum belonging to the suspicious users for its own information transmission. To improve the eavesdropping, the transmitted secondary user's signal can also be used as a jamming signal to moderate the data rate of the suspicious link. We consider two cases, i.e., non-negligible processing delay (NNPD) and negligible processing delay (NPD) at secondary transmitter. Our target is to maximize network energy efficiency (NEE) via jointly optimizing the AF relay matrix and precoding vector at the secondary transmitter, as well as the receiver combining vector at monitor, subject to the maximum power constraint at the secondary transmitter and minimum data rate requirement of the secondary user. We also guarantee that the achievable data rate of the eavesdropping link should be no less than that of the suspicious link for efficient surveillance. Due to the non-convexity of the formulated NEE maximization problem, we develop an efficient path-following algorithm and a robust alternating optimization (AO) method as solutions under perfect and imperfect channel state information (CSI) conditions, respectively. We also analyze the convergence and computational complexity of the proposed schemes. Numerical results are provided to validate the effectiveness of our proposed schemes.
Edge nodes (ENs) in Internet of Things commonly serve as gateways to cache sensing data while providing accessing services for data consumers. This paper considers multiple ENs that cache sensing data under the coordination of the cloud. Particularly, each EN can fetch content generated by sensors within its coverage, which can be uploaded to the cloud via fronthaul and then be delivered to other ENs beyond the communication range. However, sensing data are usually transient with time whereas frequent cache updates could lead to considerable energy consumption at sensors and fronthaul traffic loads. Therefore, we adopt Age of Information to evaluate data freshness and investigate intelligent caching policies to preserve data freshness while reducing cache update costs. Specifically, we model the cache update problem as a cooperative multi-agent Markov decision process with the goal of minimizing the long-term average weighted cost. To efficiently handle the exponentially large number of actions, we devise a novel reinforcement learning approach, which is a discrete multi-agent variant of soft actor-critic (SAC). Furthermore, we generalize the proposed approach into a decentralized control, where each EN can make decisions based on local observations only. Simulation results demonstrate the superior performance of the proposed SAC-based caching schemes.
The recent emergence of orthogonal time frequency space (OTFS) modulation as a novel PHY-layer mechanism is more suitable in high-mobility wireless communication scenarios than traditional orthogonal frequency division multiplexing (OFDM). Although multiple studies have analyzed OTFS performance using theoretical and ideal baseband pulseshapes, a challenging and open problem is the development of effective receivers for practical OTFS systems that must rely on non-ideal pulseshapes for transmission. This work focuses on the design of practical receivers for OTFS. We consider a fractionally spaced sampling (FSS) receiver in which the sampling rate is an integer multiple of the symbol rate. For rectangular pulses used in OTFS transmission, we derive a general channel input-output relationship of OTFS in delay-Doppler domain without the common reliance on impractical assumptions such as ideal bi-orthogonal pulses and on-the-grid delay/Doppler shifts. We propose two equalization algorithms: iterative combining message passing (ICMP) and turbo message passing (TMP) for symbol detection by exploiting delay-Doppler channel sparsity and the channel diversity gain via FSS. We analyze the convergence performance of TMP receiver and propose simplified message passing (MP) receivers to further reduce complexity. Our FSS receivers demonstrate stronger performance than traditional receivers and robustness to the imperfect channel state information knowledge.
Introducing cooperative coded caching into small cell networks is a promising approach to reducing traffic loads. By encoding content via maximum distance separable (MDS) codes, coded fragments can be collectively cached at small-cell base stations (SBSs) to enhance caching efficiency. However, content popularity is usually time-varying and unknown in practice. As a result, cached content is anticipated to be intelligently updated by taking into account limited caching storage and interactive impacts among SBSs. In response to these challenges, we propose a multi-agent deep reinforcement learning (DRL) framework to intelligently update cached content in dynamic environments. With the goal of minimizing long-term expected fronthaul traffic loads, we first model dynamic coded caching as a cooperative multi-agent Markov decision process. Owing to the use of MDS coding, the resulting decision-making falls into a class of constrained reinforcement learning problems with continuous decision variables. To deal with this difficulty, we custom-build a novel DRL algorithm by embedding homotopy optimization into a deep deterministic policy gradient formalism. Next, to empower the caching framework with an effective trade-off between complexity and performance, we propose centralized, and partially and fully decentralized caching controls by applying the derived DRL approach. Simulation results demonstrate the superior performance of the proposed multi-agent framework.
We investigate a coded uplink non-orthogonal multiple access (NOMA) configuration in which groups of co-channel users are modulated in accordance with orthogonal time frequency space (OTFS). We take advantage of OTFS characteristics to achieve NOMA spectrum sharing in the delay-Doppler domain between stationary and mobile users. We develop an efficient iterative turbo receiver based on the principle of successive interference cancellation (SIC) to overcome the co-channel interference (CCI). We propose two turbo detector algorithms: orthogonal approximate message passing with linear minimum mean squared error (OAMP-LMMSE) and Gaussian approximate message passing with expectation propagation (GAMP-EP). The interactive OAMP-LMMSE detector and GAMP-EP detector are respectively assigned for the reception of the stationary and mobile users. We analyze the convergence performance of our proposed iterative SIC turbo receiver by utilizing a customized extrinsic information transfer (EXIT) chart and simplify the corresponding detector algorithms to further reduce receiver complexity. Our proposed iterative SIC turbo receiver demonstrates performance improvement over existing receivers and robustness against imperfect SIC process and channel state information uncertainty.
Human emotions are inherently ambiguous and impure. When designing systems to anticipate human emotions based on speech, the lack of emotional purity must be considered. However, most of the current methods for speech emotion classification rest on the consensus, e. g., one single hard label for an utterance. This labeling principle imposes challenges for system performance considering emotional impurity. In this paper, we recommend the use of emotional profiles (EPs), which provides a time series of segment-level soft labels to capture the subtle blends of emotional cues present across a specific speech utterance. We further propose the emotion profile refinery (EPR), an iterative procedure to update EPs. The EPR method produces soft, dynamically-generated, multiple probabilistic class labels during successive stages of refinement, which results in significant improvements in the model accuracy. Experiments on three well-known emotion corpora show noticeable gain using the proposed method.
Categorical speech emotion recognition is typically performed as a sequence-to-label problem, i. e., to determine the discrete emotion label of the input utterance as a whole. One of the main challenges in practice is that most of the existing emotion corpora do not give ground truth labels for each segment; instead, we only have labels for whole utterances. To extract segment-level emotional information from such weakly labeled emotion corpora, we propose using multiple instance learning (MIL) to learn segment embeddings in a weakly supervised manner. Also, for a sufficiently long utterance, not all of the segments contain relevant emotional information. In this regard, three attention-based neural network models are then applied to the learned segment embeddings to attend the most salient part of a speech utterance. Experiments on the CASIA corpus and the IEMOCAP database show better or highly competitive results than other state-of-the-art approaches.
In most Internet of Things (IoT) networks, edge nodes are commonly used as to relays to cache sensing data generated by IoT sensors as well as provide communication services for data consumers. However, a critical issue of IoT sensing is that data are usually transient, which necessitates temporal updates of caching content items while frequent cache updates could lead to considerable energy cost and challenge the lifetime of IoT sensors. To address this issue, we adopt the Age of Information (AoI) to quantify data freshness and propose an online cache update scheme to obtain an effective tradeoff between the average AoI and energy cost. Specifically, we first develop a characterization of transmission energy consumption at IoT sensors by incorporating a successful transmission condition. Then, we model cache updating as a Markov decision process to minimize average weighted cost with judicious definitions of state, action, and reward. Since user preference towards content items is usually unknown and often temporally evolving, we therefore develop a deep reinforcement learning (DRL) algorithm to enable intelligent cache updates. Through trial-and-error explorations, an effective caching policy can be learned without requiring exact knowledge of content popularity. Simulation results demonstrate the superiority of the proposed framework.
Explosive growth of mobile data demand may impose a heavy traffic burden on fronthaul links of cloud-based small cell networks (C-SCNs), which deteriorates users' quality of service (QoS) and requires substantial power consumption. This paper proposes an efficient maximum distance separable (MDS) coded caching framework for a cache-enabled C-SCNs, aiming at reducing long-term power consumption while satisfying users' QoS requirements in short-term transmissions. To achieve this goal, the cache resource in small-cell base stations (SBSs) needs to be reasonably updated by taking into account users' content preferences, SBS collaboration, and characteristics of wireless links. Specifically, without assuming any prior knowledge of content popularity, we formulate a mixed timescale problem to jointly optimize cache updating, multicast beamformers in fronthaul and edge links, and SBS clustering. Nevertheless, this problem is anti-causal because an optimal cache updating policy depends on future content requests and channel state information. To handle it, by properly leveraging historical observations, we propose a two-stage updating scheme by using Frobenius-Norm penalty and inexact block coordinate descent method. Furthermore, we derive a learning-based design, which can obtain effective trade-off between accuracy and computational complexity. Simulation results demonstrate the effectiveness of the proposed two-stage framework.
Unmanned aerial vehicles (UAVs) can be utilized as aerial base stations to provide communication service for remote mobile users due to their high mobility and flexible deployment. However, the line-of-sight (LoS) wireless links are vulnerable to be intercepted by the eavesdropper (Eve), which presents a major challenge for UAV-aided communications. In this paper, we propose a latency-minimized transmission scheme for satisfying legitimate users' (LUs') content requests securely against Eve. By leveraging physical-layer security (PLS) techniques, we formulate a transmission latency minimization problem by jointly optimizing the UAV trajectory and user association. The resulting problem is a mixed-integer nonlinear program (MINLP), which is known to be NP hard. Furthermore, the dimension of optimization variables is indeterminate, which again makes our problem very challenging. To efficiently address this, we utilize bisection to search for the minimum transmission delay and introduce a variational penalty method to address the associated subproblem via an inexact block coordinate descent approach. Moreover, we present a characterization for the optimal solution. Simulation results are provided to demonstrate the superior performance of the proposed design.
Human emotional speech is, by its very nature, a variant signal. This results in dynamics intrinsic to automatic emotion classification based on speech. In this work, we explore a spectral decomposition method stemming from fluid-dynamics, known as Dynamic Mode Decomposition (DMD), to computationally represent and analyze the global utterance-level dynamics of emotional speech. Specifically, segment-level emotion-specific representations are first learned through an Emotion Distillation process. This forms a multi-dimensional signal of emotion flow for each utterance, called Emotion Profiles (EPs). The DMD algorithm is then applied to the resultant EPs to capture the eigenfrequencies, and hence the fundamental transition dynamics of the emotion flow. Evaluation experiments using the proposed approach, which we call EigenEmo, show promising results. Moreover, due to the positive combination of their complementary properties, concatenating the utterance representations generated by EigenEmo with simple EPs averaging yields noticeable gains.
In this work, we consider a cognitive radio system, where the primary user (PU) owns the spectrum but has scarce energy while the secondary users (SUs) have adequate energy but lack of spectrum. Thus, a spectrum sharing and energy cooperation scheme is proposed, where the SUs help transfer energy to the PU in the first phase, and in return, the PU allows the SUs to access the spectrum in the second phase. This is particularly beneficial when the PU is energy-limited wireless sensor node or internet of things and the transmitters of SUs are base stations or access points with sufficient energy supply. Without loss of generality, we aim to maximize the minimum data rate among all SUs by jointly optimizing the time- splitting factor between the two phases, the transmission power at the primary transmitter (PT) and the precoding vectors for secondary transmitters (STs) under the minimum data rate requirement of the PU, and the power constraint at each ST. We also guarantee the energy causality constraint at the PT, i.e., the total consumed energy should be no larger than the total available energy. To solve this non-convex problem, we propose an efficient iterative algorithm by applying the successive convex approximation (SCA) and further show that the proposed algorithm is guaranteed to converge. Simulation results are finally presented to show the effectiveness of our proposed scheme.
This paper considers content delivery of the cache-enabled small cell networks (C-SCNs), where users with the same request form a multicast group and are served by a cluster of small-cell base stations (SBSs) under the coordination of the central processor. The performance of such a coordination is severely limited by the fronthaul link, which may be saturated and degrade quality of service (QoS). To improve user QoS, we propose a latency driven scheme by jointly optimizing fronthaul bandwidth allocation, multicast beamforming, and BS clustering. Accordingly, with min-max fairness among multicast groups, a latency minimization problem is formulated under the constraints of fronthaul bandwidth and transmission power. The resultant problem is a mixed-integer nonlinear program, which is NP-hard. To address such a complex problem, a quadratic penalty-based algorithm is proposed by using a reformulation of binary constraint. Meanwhile, we present the necessary condition for an optimal solution, which shows that fronthaul bandwidth allocation is inherently adaptive to cached contents and patterns of BS cooperation. Finally, simulation results demonstrate that the proposed scheme can effectively reduce latency under different caching strategies.
In this paper, we propose to combine the deep learning of feature representation with multiple instance learning (MIL) to recognize emotion from speech. The key idea of our approach is to first consciously classify the emotional state of each segment. Then the utterance-level classification is constructed as an aggregation of the segment-level decisions. For the segment-level classification, we attempt two different deep neural network (DNN) architectures called SegMLP and SegCNN, respectively. SegMLP is a multilayer perceptron (MLP) that extracts high-level feature representation from the manually designed perceptual features, and SegCNN is a convolutional neural network (CNN) that automatically learn emotion-specific features from the log Mel filterbanks. Extensive emotion recognition experiments are carried out on the CASIA corpus and the IEMOCAP database. We find that: (1) the aggregation of segment-level decisions provides richer information than the statistics over the low-level descriptors (LLDs) across the whole utterance; (2) automatic feature learning outperforms manual features. Our experimental results are also compared with those of state-of-the-art methods, further demonstrating the effectiveness of the proposed approach.