
We investigate energy- and spectral-efficiency trade-off by optimizing resource efficiency (RE) for single-cell massive multiple-input multiple-output (MIMO) downlink transmission with only statistical channel state information available at the transmitter (CSIT). We first show that beam domain transmission is favorable for RE optimization in massive MIMO downlink, by deriving a closed-form solution for the eigenvectors of the optimal transmit covariance matrices. With this conclusion, the RE optimization precoding design is reduced to a real-valued power allocation problem. Exploiting sequential optimization and random matrix theory, we further propose a low-complexity two-layer water-filling-structured power allocation algorithm. Numerical results demonstrate the effectiveness of our proposed statistical CSI aided RE optimization approach.
Pathological Hand Tremor (PHT) is one of the most prevalent symptoms of some neurological movement disorders such as Parkinson's Disease (PD) and Essential Tremor (ET). Characterization, estimation, and extraction of PHT is a crucial requirement for assistive and robotic rehabilitation technologies that aim to counteract or resist PHT as an input noise to the system. In general, research in the literature on the topic of PHT removal can be categorized into two major categories, namely, classic and data-driven methods. Classic techniques use hand-crafted features and statistical processing pipelines to model and then extract the tremor while data-driven approaches are trained based on a sizable dataset to allow a computational model (such as neural networks) learn how to estimate the PHT. Since the availability of large datasets, especially in PHT estimation field is a bottleneck, in this feasibility study, we investigate the possibility of combining different recording modalities of PHT to generate a neural network for this purpose. This work explores the potential of jointly using accelerometer data and gyroscope recordings to produce a larger dataset for training a relatively complex network, which can potentially be extended for a deeper generalization.
This paper considers the blind recognition of channel codes from noisy received signal. Specifically, based on the residual neural network (ResNet), the recurrent neural network (RNN), and the attention mechanism, three types of recognizers are proposed to recognize the type of a channel code through a classification process. Other code parameters (e.g., code rate and length) can be recognized by similar ways. Numerical experiments show that the proposed recognizers perform well even when the training set is a very small subset of all possible codewords.
In this paper, the issue of gridless compressed sensing (CS)-based direction-of-departure (DOD) and direction-of- arrival (DOA) estimation for bistatic multiple-input multiple- output (MIMO) radar is investigated with one snapshot, and an alternating direction method of multipliers (ADMM) for 2D parameter estimation using decoupled atomic norm minimization (DANM) is proposed. In the proposed algorithm, the decoupled atom set and atomic norm for DOD and DOA estimation are defined, then the atomic norm is transformed into a semi-definite programming (SDP) minimization problem. To decrease the computational complexity of solving SDP using interior point method based SDPT3 solver in CVX toolbox, the DANM with ADMM is deduced, which can greatly decrease the running time, especially in the presence of large scale arrays. Finally, the DOD and DOA are obtained via shift-invariance parameter estimation. It overcomes the grid-mismatch effect of grid-partition of the conventional CS methods, and outperforms the traditional subspace-based methods. Numerical simulations are presented to demonstrate the estimation performance of the proposed algorithm.
Due to the advances in dramatically increased computing power, machine learning technologies have been widely exploited to provide promising solutions in different application fields. However, the high performance of many of these machine learning technologies highly relies on the availability of sufficient annotated data. Considering the fact that data annotation is an extremely time-consuming process, this condition is not always practical. To address this challenge, in this paper we propose a novel computing method, called hybrid semi-supervision machine learning, that exploits the loose domain knowledge to enable the accurate results even in the presence of limited labeled data. Simulations results illustrate the effectiveness of our method.
We propose a method to estimate the 3D gaze of the observer onto the scene using a portable eye tracker with a monocular camera. We reconstruct the 3D scene using Structure from Motion (SfM) and use camera pose and 3D reconstruction information of the scene to localize the position and pose of the observer in the 3D reconstructed space. Along with the position, we can obtain the 2D gaze of the observer using an eye-tracker. Each person may have a different perspective of the same 3D object in the scene, observing it from different positions. We use this information to fuse these multiple perspectives in 3D space to get a better understanding of how differently each observer perceives the same scene, compared to others. In the entire system, we developed a convo-lutional neural network to detect and track eye movement and a camera re-localization method to localize the camera in 3D environment. Based on our novel eye tracking and camera re-localization methods, we can accurately localize the gaze in the 3D reconstructed environment.
We consider a scenario where signals that can be modeled as finite sums of linearly modulated signals that are non-orthogonal in both time and frequency are observed by a sensor. Under this model, when an antidiagonal slice of the trispectum of the sum of these signals is computed for multiple segments in time and stacked into a frontal symmetric block partitioned tensor (FSBPT) whose factor matrices characterize the received power spectrum and transmission activities of the signals. We propose a new alternating soft thresholding decomposition strategy that is tailored to FSBPTs and, unlike existing BPT decomposition algorithms, simultaneously estimates the factor partitions which is necessary since in our signal model the partitioning is not known a priori. We demonstrate the algorithm's performance via Monte Carlo simulations. We find that our proposed non-blind algorithm is able to estimate the tensor partitions and moderately accurately estimate the factor matrices under a range of factor matrix column collinearity values.
The operation of distribution networks is becoming increasingly volatile, due to fast variations of renewables and, hence, net-loading conditions. To perform a reliable state estimation under these conditions, this paper considers the case where measurements from meters, phasor measurement units, and distributed energy resources are collected and processed in real time to produce estimates of the state at a fast time scale. Streams of measurements collected in real time and at heterogenous rates render the underlying processing asynchronous, and poses severe strains on workhorse state estimation algorithms. In this work, a real-time state estimation algorithm is proposed, where data are processed on the fly. Starting from a regularized least-squares model, and leveraging appropriate linear models, the proposed scheme boils down to a linear dynamical system where the state is updated based on the previous estimate and on the measurement gathered from a few available sensors. The estimation error is shown to be always bounded under mild condition. Numerical simulations are provided to corroborate the analytical findings.
In this paper, we consider learning based regularization for compressive sensing reconstruction using focal plane array sensors. While many optimization algorithms employ proximal operators for regularization, they are often inadequate in fully capturing the characteristics of complex natural images. Recently, deep learning based approaches obtained promising results in different imaging problems, creating the possibility to use them for regularization in an optimization framework. Here, we utilize this approach in compressive sensing based spatial multiplexing cameras. This technique is motivated by the high cost of producing large focal plane arrays in infrared sensors. Reconstruction from undersampled measurements can be done using a spatial multiplexing camera which relies on multiple snapshots for super-resolving a scene. It acquires coded projections of a scene using a spatial light modulator and a low-resolution focal plane array. We first formulate the problem of finding a high resolution image from its undersampled measurements. Then, we develop a reconstruction method with learning based regularization to solve this problem using alternating direction method of multipliers framework. For this, we replace the proximal operator corresponding to the regularization function with a deep convolutional denoising network. We also enhance a previously proposed denoising network's training phase by introducing multiple noise realizations for each training patch, which results in better reconstruction performance. Numerical results for different imaging scenarios show successful recovery of high resolution images in terms of PSNR, SSIM and visual quality at significant noise levels.
In this paper, a new index modulation technique, referred to as the multi-mode generalized space-time index modulation (MM-GSTIM) is presented. MM-GSTIM is a high-rate index modulation scheme for multiantenna inter-symbol interference (MIMO-ISI) channels. The proposed MM-GSTIM is a single carrier scheme with zero padding (ZP). MM-GSTIM performs `mode-indexing' to efficiently increase the spectral efficiency. In a given MM-GSTIM frame, disjoint subsets of space-time slots are used to transmit modulation symbols from disjoint modulation alphabets (referred to as the modes). The choice of which space-time slots are used for transmission of a particular mode in an MM-GSTIM matrix conveys information bits. Thus, MM-GSTIM is a generalization of the well-known index modulation schemes such as spatial modulation (SM), generalized spatial modulation (GSM), and space-time index modulation. Here, an analytical characterization of the diversity order of the proposed MM-GSTIM scheme and the conditions required to maximize the transmission rate of multi-mode index modulation schemes are derived. Through simulations, it is shown that the proposed MM-GSTIM scheme can achieve better performance compared to other modulation schemes in MIMO-ISI channels.
Over the years, the number of consumers seeking personal financial advisory services has grown globally. However, recent studies indicate a worrying decline in consumers' trust and confidence in advisers and financial institutions, as well as low regulatory compliance rates. Inspiring consumer trust through increased vigilance of advice is not possible using current auditing practices as reviews are manual, time-consuming and complex. In this paper, we describe a generalised framework which leverages machine learning approaches to systematically characterise the risk status of financial advice documents prior to client delivery. We show how the framework presented provides a comprehensive, accurate and efficient compliance review of financial advice documents for financial advisers and compliance officers alike.
Unsupervised term discovery is the task of identifying and grouping reoccurring word-like patterns from the untranscribed audio data. It facilitates unsupervised acoustic model training in zero resource setting where no or minimal transcribed speech is available. In this paper, we investigate two-step bottom-up approaches for unsupervised discovery of word-like units. The first step discovers phone-like acoustic units from data and the second step combines the basic acoustic blocks to identify word-like units. We investigated Embedded Segmental K-means and Nested Hierarchical Pitman-Yor (PYR) model as bottom-up strategies. ESK-Means iteratively selects boundaries from an initial set to arrive at the word boundaries. The final performance critically depends on the quality of the initial boundaries. We used a segmentation method that discovers boundaries much closer to actual boundaries. PYR model has been used for word segmentation from space removed text data, and here we use it for word discovery from unsupervised acoustic units. The term discovery performance is evaluated on the Zero Resource 2017 challenge dataset, which consists of around 70 hours of unlabelled data. Our systems outperformed the baseline systems on all the languages without language-specific parameter tuning. We performed comprehensive experiments of the system parameters on the system performance.
Timely information updates are critical to time-sensitive applications in networked monitoring and control systems. In this paper, the problem of real-time status update is considered for a cognitive radio network (CRN), in which the secondary user (SU) can relay the status packets from the primary user (PU) to the destination. In the considered CRN, the SU has opportunities to access the spectrum owned by the PU to send its own status packets to the destination. The freshness of information is measured by the age of information (AoI) metric. The problem of minimizing the average AoI and energy consumption by developing new optimal status update and packet relaying schemes for the SU is addressed under an average AoI constraint for the PU. This problem is formulated as a constrained Markov decision process (CMDP). The monotonic and decomposable properties of the value function are characterized and then used to show that the optimal update and relaying policy is threshold-based with respect to the AoI of the SU. These structures reveal a tradeoff between the AoI of the SU and the energy consumption as well as between the AoI of the SU and the AoI of the PU. An asymptotically optimal algorithm is proposed. Numerical results are then used to show the effectiveness of the proposed policy.
Chest radiographs or X-ray images are a common diagnostic tool to identify different thoracic diseases and other abnormal cardiopulmonary conditions. The advancements of artificial intelligence paves the way to machine learning based computer-assisted systems that can support the radiologists in disease diagnosis and report generation from chest radio-graphs. In this work we report an implementation of a deep-learning based framework to interpret the disease signature from chest X-rays. The model was trained on a large dataset consisting of both frontal and lateral X-ray images of the chest with multiple thoracic disease labels. We report a mean area under ROC curve (AUC) of 0.86, with the AUC of individual diseases in the range of 0.76 to 0.93. We also generated disease-level colormaps to visually present the X-ray image region most indicative of the disease.
This paper proposes a rate control algorithm based on group of pictures (GOP) level quality dependency for high efficiency video coding (HEVC) low delay hierarchical coding structure. Firstly, this paper builds a GOP level quality dependency model. Secondly, a GOP level quality dependency model based GOP level rate-distortion optimization is introduced to make bit allocation more reasonable. Experimental results show that, compared with HM16.7 frame level rate control algorithm, proposed rate control algorithm shows 5.80% and 5.07% BD-rate saving on average under low delay B and low delay P configurations.
In this work, we tackle the problem of real-time compressive video reconstruction in spatial multiplexing cameras using spatially modulated and undersampled focal plane array data. In this setting, a spatial light modulator (SLM) modulates the scene in the image domain by blocking some of the pixels at a higher resolution level, using an SLM such as a digital micromirror device. Here, we first propose a practical warm-starting method to achieve real-time reconstruction of streaming video. We then extend it by adapting a multi-hypothesis technique that takes temporal consistency into account using residual reconstruction between frames. We implement the algorithms on a graphics processing unit and give an extensive analysis of the number of required iterations for different imaging settings. We conclude that the proposed methods achieve video reconstruction in real-time from highly undersampled compressed measurements and provide high quality frames in terms of PSNR and SSIM metrics.
Convolutional neural networks (CNNs) have been widely used in computer vision applications and achieved great success. However, large-scale CNN models usually consume a lot of computing and memory resources, which makes it difficult for them to be deployed on embedded devices. An efficient block-floating-point (BFP) arithmetic is proposed in this paper. compared with 32-bit floating-point arithmetic, the memory and off-chip bandwidth requirements during convolution are reduced by 50% and 72.37%, respectively. Due to the adoption of BFP arithmetic, the complex multiplication and addition operations of floating-point numbers can be replaced by the corresponding operations of fixed-point numbers, which is more efficient on hardware. A CNN model can be deployed on our accelerator with no more than 0.14% top-1 accuracy loss, and there is no need for retraining and fine-tuning. By employing a series of ping-pong memory access schemes, 2-dimensional propagate partial multiply-accumulate (PPMAC) processors, and an optimized memory system, we implemented a CNN accelerator on Xilinx VC709 evaluation board. The accelerator achieves a performance of 665.54 GOP/s and a power efficiency of 89.7 GOP/s/W under a 300 MHz working frequency, which outperforms previous FPGA based accelerators significantly.
Over the past decade, the increased adoption of Internet-connected technology has led to a substantial growth in the amount of videos produced and digested by people around the world. Subsequently, manipulated content within videos has become far more common and difficult to detect. Material of this nature poses a significant problem as it provides a way to falsely affect viewers' beliefs. In this work, we first develop a pipeline to generate manipulated video segments from preexisting videos and then develop a deep learning architecture to detect unique video manipulations on a frame-by-frame basis. Specifically, we develop a model which analyzes manipulated video segments represented as sequences of images via a joint Residual Network (ResNet) feature extractor and Long Short-Term Memory (LSTM) network to detect video frames that exhibit signs of manipulation and classify which type of manipulation was applied. We train our model on a modified subset of the UCF101 Action Recognition Dataset [1] which we alter with a set of four manipulation types: object insertion, compression, frame blackout, and blurring. We evaluate the classification accuracy of our model by creating manipulation "blocks", or sets of consecutive manipulated video frames, varying block length, distance between blocks, and manipulation distribution within blocks. Experimental results demonstrate that our model achieves high (>90%) classification accuracy in the presence of increased frame class transitions.
Beamforming with large-scale antenna arrays (LSAA) is one of the predominant operations in designing wireless communication systems. However, the implementation of a fully digital system significantly increases the number of required radio-frequency (RF) chains, which may be prohibitive. Thus, analog beamforming based on a phase-shifting network driven by a variable gain amplifier (VGA) is a potential alternative technology. In this paper, we cast the beamforming vector design problem as a beampattern matching problem, with an unknown power gain. This is formulated as a unit-modulus leastsquares (ULS) problem where the optimal gain of the VGA is also designed in addition to the beamforming vector. We also consider a scenario where the receivers have the additional processing capability to adjust the phases of the incoming signals to mitigate specular multipath components. We propose efficient majorization-minimization (MM) based algorithms with convergence guarantees to a stationary point for solving both variants of the proposed ULS problem. Numerical results verify the effectiveness of the proposed solution in comparison with the existing state-of-the-art techniques.
This work studies error exponent limits in hypothesis testing (HT) in a distributed scenario with partial communication constraints. We derive general conditions on the Type I error restriction under which the error exponent of the optimal Type II error has a closed-form characterization for the task of testing against independence. We show that the error exponent is preserved for a family of decreasing Type I error restrictions. Complementing this analysis, new expressions are derived to bound the optimal Type II error probability for a finite number of observations. These bounds shed light about the velocity at which error exponent limits are attained with the number of samples.