
Sensing and understanding all signals in an RF-secure military or civilian setup is important, this includes unintended RF emissions called emanations. Prior work in detecting emanations involves profiling, which is hardware (HW) specific and hence not a scalable approach. Our technique looks for a generic signature of harmonics in the frequency domain without knowledge of HW. It detects emanation and characterizes it by estimating the pitch frequency of harmonic. A signal model for emanations, HW, and channel effects is provided. A mathematical treatment showcases the removal of artifacts and extracts the harmonic structure. The performance is showcased on In-phase and Quadrature-phase (IQ) data collected using a Signal Hound software-defined radio (SDR) in a shielded room. Data is collected from 0.1-1.1 GHz and 100 ms duration. Emanations are detected for a laptop, and a desktop connected to a monitor. The unique contribution of this work is the detection of emanation without HW knowledge, with mathematical justification and demonstrated performance on wideband real data.
Signal/image reconstruction problems are typically posed as linear inverse problems with the objective of minimizing a least-squares (LS) loss function and a regularization term derived from a suitable prior. Several decades of work have gone into the design and efficient application of various penalty functions, whilst the design of the data-fidelity loss function has remained relatively unexplored. This paper proposes a distance-minimization approach to the design of the data-fidelity loss for solving linear inverse problems and discusses proximal algorithms for solving them. We discuss the implications of using the proposed formulation as opposed to the standard LS in terms of its effect on noise and convergence. We consider solving sparse signal recovery and image restoration tasks, in particular, image deblurring, and show that distance minimization techniques out-perform standard LS for the various models under consideration.
Generalized frequency division multiplexing (GFDM) offers more flexibility in choosing subsymbols, subcarriers, and pulse-shaping filters for efficient data transmission. Due to inherent non-orthogonality, GFDM produces intrinsic self-interference, namely inter-subsymbol interference (ISI) and inter-subcarrier interference (ICI). An Eigendecomposition-based GFDM (ED-GFDM) utilizes a matched filter (MF) receiver, this architecture can mitigate self-interference without any noise enhancement. Nevertheless, in wireless channels, the propagation of ED-GFDM is affected by channel characteristics like non-linearity and non-homogeneity. Usually, a generalized alpha - mu fading channel is modeled to capture the channel characteristics in terms of fading parameters alpha and mu. Specific combinations of alpha and mu result in various fading environments like Rayleigh, Exponential, Nakagami-m, etc. This paper aims to derive accurate closed-form expressions for Symbol Error Rate (SER) and ergodic channel capacity for ED-GFDM using M-PSK modulation with a MF receiver. Obtained simulation results are analyzed and validated with the expected theoretical results.
This work examines a cognitive millimeter-wave communication network with two operators, a primary and a secondary, each with different channel access priorities. The primary operator can transfer its licenced spectrum to a secondary opera-tor, while the secondary operator follows an optimal interference-threshold criteria which was pre-specified by primary operator in its secondary licenses. The study uses medium-access-probability and activity factor to quantify channel access opportunities and number of active secondary transmitter-receiver pairs. For better network performance, primary operators should promote adaptive directional sensing, allowing secondary users to switch on the narrower antenna beamwidth regimes based on their proximity to other primary devices.
Extensive use of voice assistants by children in their day-to-day life activities demand for better performance of Automatic Speech Recognition (ASR) for children's speech. The recent advancements in ASR perform better for adult speech. However, due to acoustic mismatch (in particular, higher $F_{0}$ and thus, poor spectral resolution, and less availability of children data), it remains a challenge to improve the performance of children's ASR. In this context, a shared task on ASR for children's speech 2021 was organized during INTERSPEECH 2021. In this paper, the voice conversion-based data augmentation technique using a cycle-consistent generative adversarial network (CycleGAN) is investigated for closed track in the challenge. Here, CycleGAN training is exploited for children-to-children voice conversion. Training of CycleGAN and ASR experiments are performed on the speech data provided during the challenge. In our experiments, CycleGAN-based data augmentation showed a relative improvement of 9% and 3% in WER compared to the baseline system on the Dev and Eval sets, respectively. This work may find its significance in ASR for low resource languages, where Librispeech or any other external dataset is not an option.
As per the existing literature, Orthogonal Time Frequency Space (OTFS) transceiver can be implemented using two different methods: i) the two step approach: includes inverse symplectic fast Fourier transform (ISFFT) and Heisenberg transform at the transmitter, along with SFFT and Wigner transform at the receiver, and ii) the direct approach: incorporates Inverse Zak (IZak) transformation at the transmitter and Zak transformation at the receiver. In this work, to expedite the implementation process of the OTFS transceiver in real-time wireless communication, we use field programmable gate array (FPGA) technology, leveraging the time acceleration benefits it provides. Additionally, while implementing the two aforementioned approaches, we employ the coordinate rotation digital computer (CORDIC) algorithm as it is flexible and requires minimum area. We compare the hardware performances among these two approaches in terms of resource utilization, timing, and power by implementing on the 7a200tiffg1156-1L FPGA board. We observe that, the direct approach exhibits significant improvements, with a 47.57% reduction in Look-Up Tables (LUTs) and a 17.63% reduction in Flip Flops (FF) compared to the two step approach in the OTFS transceiver design.
This paper conceives a new generalized simultane-ous orthogonal matching pursuit (GSOMP)-based algorithm to effectively estimate the sparse channel state information (CSI) in a multiuser (MU) THz hybrid MIMO system. The proposed framework also incorporates low-resolution analog-to-digital converters (ADCs) together with a sampled version of the transmit pulse shaping filter. The proposed techniques are based on a practical dual-wideband THz channel model that is developed initially, which considers both the spatial and frequency wide-band effects. The model also embraces the reflection, absorption and free-spaces losses that are a characteristic feature of the THz band. Subsequently, a novel MU hybrid transceiver design framework is advanced, based on the generalized alternating direction method of multipliers (MU-GADMM), which generates a new set of basis vectors toward robust approximation of the optimal precoder, in order to account for the beamsquint effect. Extensive simulations are conducted to evaluate the performance of the proposed CSI learning and beamforming techniques in a practical THz channel generated using the high-resolution transmission (HITRAN) database.
Photon-limited deblurring is a complex and demanding problem encountered in various applications where low-light conditions prevail. The scarcity of photons in such situations leads to the introduction of shot noise, resulting in a degradation of image quality. Solving this problem with Neural networks often involves constructing models empirically, making the behavior of the underlying architecture challenging to comprehend. A recent technique known as algorithm unrolling has enabled the connection of iterative algorithms with neural networks, where the Convolutional Neural Network (CNN) acts as a denoiser. This paper introduces a reduced parameter denoiser to enhance image quality and preserve finer details or avoid over-smoothing of the image during reconstruction. As a result, the unrolled model surpasses existing deblurring methods for improving image quality in low-light conditions. The proposed denoiser reduces the number of parameters by a factor of 3.84 and preserves the finer details while reconstructing. Our model improves computational efficiency and storage requirements compared to the state-of-the-art.
Among various speech disorders, dysarthria presents a unique challenge when it comes to end-to-end speech synthesis. Its severity and complexity further escalate if appropriate treatment is not taken. In this paper, we propose a speaker-adaptive dysarthric speech synthesis technique using the Tacotron2 model. We used this model to generate dysarthric speech utterances that already exist in the UASpeech database and for Out-Of-Vocabulary (OOV) words as well to expand the vocabulary size. By generating dysarthric speech from textual input, we achieved favorable Mean Opinion Score (MOS) ratings for both known and OOV words. This approach was successful in adapting to the inherent variability and intelligibility differences among speakers. To ensure the fidelity of our generated speech, we incorporated Dynamic Time Warping (DTW) to demonstrate the similarity between the original and generated waveforms. Additionally, we applied Waveform similarity-based Synchronized OverLap-Add (WSOLA) on the OOV words, ensuring a flawless evaluation process. This approach allows us to address the challenge of limited data availability by enhancing database and vocabulary size. This paper aims to generate synthetic dysarthric speech to expand the pool of available UASpeech database, facilitating advancements in automatic recognition of dysarthric speech. Audio samples for all generated words for 6 dysarthric speakers are available at https://github.com/samiulhaq4424/MOS
The outage probability analysis of orthogonal time frequency space (OTFS) in decode and forward (DF) and amplify and forward (AF) cooperative communication is investigated in this paper. The system is composed of a source node S, a relay node R and a destination node D, and it is considered that no direct communication link is feasible between the source and destination. The paper presents the closed form expression for the outage probability for both DF and AF OTFS cooperative communication with a single relay. The diversity order is found to be $K$ for both DF and AF, where $K$ is the number of multipaths. Therefore, OTFS AF and DF cooperative transmissions can achieve a diversity order equivalent to the number of resolvable paths within the channel. Monte-Carlo simulations are used to verify the theoretical results.
Code-mixing is the situation where words and phrases from two or more different languages are used inter-changeably within one sentence or utterance. With the increasing number of bilinguals and multilinguals in today's world, there is heavy evidences of code-mixing in many scenarios. However their presence affects the accuracy of current natural language processing and speech processing systems, since currently available speech-to-text and text-to-speech systems are not viable to translate speech which consists of two or more languages mixed together to a target language, as they assume the text and speech to all come from a single language. In the case of Dravidian languages, especially Kannada-English code-mixing, there are very few datasets available, and lesser so which consist of speech data. In order to get closer to rectifying this problem, we present a generative adversarial architecture to synthesize code-mixed speech in Kannada-English using monolingual utterances in Kannada. We are able to generate sentences with a good FAD (Frechet Audio Distance) score of around 14.490, using a very small training dataset.
Interstitial lung disease (ILD) is a collection of pulmonary adventitious conditions that induce scarring of the lung parenchyma, fibrosis, and inflammation. ILD encompasses over 200 chronic respiratory diseases that gradually damage the lung tissues and make it difficult to acquire adequate oxygen in the lungs. Therefore, it is essential to identify and diagnose diseases early to prevent their progression. ILDs are often characterized by abnormal respiratory sounds (RSs) such as crackles and squawks as a result of anatomical faults in the respiratory pathway produced by the disease. In this paper, for the first time, we propose a novel sinc convolution-based residual convolutional deep learning architecture, namely the ILDNet, for categorizing the ILD-affected RSs. The proposed framework comprises two major stages: (a) preprocessing of the input RS and (b) classification of the RSs using the proposed ILDNet. The proposed framework is extensively evaluated using the RSs from the publicly available BRACETS and KAUH datasets, and the experimental results show that our proposed ILDNet framework achieves an accuracy, sensitivity, and specificity of 81.25%, 78.85%, and 83.33%. These results also pave the way for future research on the potential use of RSs to identify reliable biomarkers for early-stage ILD identification.
This paper considers massive multiple-input multiple-output (MIMO) Internet of Things (IoT) networks where each sensor is equipped with a single antenna, while the fusion center (FC) is equipped with a massive array of antennas. Each sensor takes the linear noisy observation of the underlying unknown random quantity, pre-processes it, and then transmits its precoded observation to the FC over a fading wireless channel for efficient post-processing. Since these sensors are battery-operated tiny devices, energy efficiency (EE) becomes such a critical factor to optimize, while individual mean square error (MSE)-based quality of service (QoS) constraints ensures an efficient estimation of the random quantities at the FC. Quadratic transform theory is used to solve the resulting non-convex EE maximization problem, and the first-order Taylor series approximation is used to linearize the non-convex quantities. Our numerical results corroborate the analytical findings of this work.
We study a grouped bandit setting where each arm comprises multiple independent sub-arms referred to as attributes. Each attribute of each arm has an independent stochastic reward. We impose the constraint that for an arm to be deemed feasible, the mean reward of all its attributes should exceed a specified threshold. The goal is to find the arm with the highest mean reward averaged across attributes among the set of feasible arms in the fixed confidence setting. We propose a confidence interval-based policy to solve this problem and provide analytical guarantees for the policy. We compare the performance of the proposed policy with that of two suitably modified versions of action elimination via simulations.
The service-based architecture of 5G allows network operators to place softwarized network functions on commodity hardware with the help of software-defined networking (SDN) and network function virtualization (NFV) technologies. While several existing works focused on network function placement and service routing in a 5G network theoretically, more investi-gation is required to study the performance of the softwarized network functions when placed on commodity hardware. In this paper, we study the softwarized network security function placement in a 5G network and its impact on the network performance. Specifically, we focus on the implementation of pri-mary network security functions, such as intrusion detection and prevention systems (IDS/IPS) and network address translation (NAT), in a 5G network using open-source tools. We present a comprehensive method for the function placement, network configuration, and performance evaluation in the presence of synthetic network traffic. The extensive experiment results show that softwarized network security functions can be used to meet the quality-of-service requirements of specific applications in 5G. In contrast, dedicated hardware-based network functions may be required to support applications with stringent QoS requirements.
Hypertension or high blood pressure is a significant global health issue. Having high blood pressure is a big risk for conditions like coronary heart disease, including ischemic and hemorrhagic stroke. In general, the measurement of blood pressure is performed using a sphygmomanometer. However, this technique has several limitations in continuous and long-term monitoring due to bulky electronic devices with pneumatic systems (pump, valve, battery) to inflate and deflate the cuff. Cuffless blood pressure estimation has recently emerged as a good alternative to overcome these limitations. This paper proposes a machine learning-based approach using wavelet-based time-frequency features and adaptive boosting regression for cuffless blood pressure estimation from photoplethysmogram signals. The efficacy of the proposed approach is evaluated using various parameters concerning different state-of-the-art approaches. The proposed approach is found to perform better than various state-of-the-art methods. Furthermore, the proposed approach is implemented on the Xilinx PYNQ-Z2 board to validate the hardware compatibility.
In this work, we propose an Unequal Error Pro-tection technique suitable for scenarios that demand low power consumption and are equipped with limited computationally ca-pabilities. CosinePrism incorporates colour theory and techniques from lossy image compression methods towards optimizing the allocation of Forward Error Correction factor. CosinePrism can progressively deliver RGB and non-RGB images over half-duplex links, and includes a JPEG re-compression mode. The proposed technique achieves compression efficiency and image quality comparable to JPEG, and on an average offers DSSIM score improvement of 55.5% over Equal Error Protection method. An open-source reference encoder is provided for further research and development to the community.
Speech modification methods normalize children's speech towards adults' speech, enabling off-the-shelf generic automatic speech recognition (ASR) for this low-resource scenario. On the other hand, ASR models like Wav2Vec2 have shown remarkable robustness towards various speakers, thus streamlining their deployment. This paper examines the benefit of speech modification methods when using Wav2Vec2 models on children's speech. We experimented with prototypical speech modification methods and found that while models trained on large datasets exhibit similar performance across unmodified and modified children's speech, models trained on smaller datasets exhibit notably enhanced performance with modified speech. However, analyzing age effects on PF-Star and CMU Kids evaluation sets, we observe that all Wav2Vec2 variants still underperform for children under 10 years. In this scenario, speech modification methods and their combinations help improve performance for small and large Wav2Vec2 models but have plenty of room for improvement.
Rapid growth in demand for Internet of Things (IoT) applications has driven advancements in network tech-nologies. Mobile IoT applications, including tracking, intelli-gent wearables, and autonomous vehicles, necessitate high data throughput, privacy, low-latency data transfer, and prolonged network lifetime. In every IoT application, many sensor devices collect data from physical objects and send the acquired information to the access point (AP). A secure connection between IoT nodes and the AP is crucial for optimal data routing, transmission reliability, and network lifetime in IoT networks. Additionally, it is essential to detect and remove anomalous nodes to protect the network from unwanted intrusion. This work proposes a socially-aware anomalous node detection and data routing strategy, representing the mobile IoT network as a dynamic graph. A novel social relationship index (SRI)-factor is computed based on current and predicted future node encounters to identify and remove the imposture nodes. The network graph is then dynamically updated. Moreover, an optimization problem is formulated to maximize the network lifetime with respect to data transmission latency, optimum routing path, and energy budget. The proposed optimization problem is solved using the Q-learning framework. Q-learning-based routing algorithm adapts to the changes of the dynamic graph by continuously updating the Q-values on the basis of feedback. Finally, numerical results demonstrate a significant performance improvement compared to existing methods.
Voice OTP Authentication (VOA) is a two-way au-thentication framework to automatically authenticate the user by identifying the speaker and spoken One-Time Password (OTP). The work towards designing the VOA system is limited in the literature. This work proposed a novel framework, that uses state-of-the-art (SOTA) Emphasized Channel Attention, Propagation, and Aggregation Time Delayed Neural Network (ECAPA-TDNN) based x-vector as speaker representation and conformer-based digit representation to perform VOA. The performance of VOA using the pretrained publicly available speaker and digit representation extractor is 78.92%. The performance is improved to 93.21% after fine-tuning the frameworks with the in-domain training data. Further, to see the feasibility of the framework for deployment, the inference time and memory requirement are calculated by deploying the framework in three devices: (1) high-end server, (2) laptop CPU, and (3) low-end edge device (odroid board). The inference times in all three devices are 0.068s, 0.33s, and 1.74s, respectively.