
Gaps, dropouts and short clips of corrupted audio are a common problem and particularly annoying when they occur in speech. This paper uses machine learning to regenerate gaps of up to 320ms in an audio speech signal. Audio regeneration is translated into image regeneration by transforming audio into a Mel-spectrogram and using image in-painting to regenerate the gaps. The full Mel-spectrogram is then transferred back to audio using the Parallel-WaveGAN vocoder and integrated into the audio stream. Using a sample of 1300 spoken audio clips of between 1 and 10 seconds taken from the publicly-available LJSpeech dataset our results show regeneration of audio gaps in close to real time using GANs with a GPU equipped system. As expected, the smaller the gap in the audio, the better the quality of the filled gaps. On a gap of 240ms the average mean opinion score (MOS) for the best performing models was 3.737, on a scale of 1 (worst) to 5 (best) which is sufficient for a human to perceive as close to uninterrupted human speech.
Automated assessment of online spoken materials for use in listening training can be of benefit to language learners. ICALL systems with this facility can help learners sift through collections of online content to find materials that suit their proficiency level and learning goals. Our pilot study assesses the extent to which readability measures can contribute to the listenability assessment of online spoken materials for language learning. Extending these text-based measures to assess listenability is convenient, but their use must be reassessed against spoken materials available online. We compare assessments by four readability measures with expert-assigned assessments of difficulty on online spoken language learning materials. We also evaluate the robustness of these measures against errors arising from automatically generating transcripts and their capability to predict human-based judgments of listenability. Our findings show that these comprehensibility measures can be combined and used to discriminate between materials of different levels of difficulty. We are also able to illustrate a way to conduct a fully automated approach to listenability assessment of spoken language learning materials via text-based measures. Our study also highlights the need for an approach that looks beyond the text, particularly for assessing higher-level materials and spoken materials of different genres.
Current building codes require that new central heating appliances such as oil and gas fired boilers be accompanied by a room thermostat and programmer. The location of the thermostat can lead to problems in operating the central heating system when the temperature local to the thermostat is not consistent with the temperature in rooms where there is demand. Likewise, the interaction with other energy saving measures, such as thermostatic radiator valves (TRV) can produce control issues and lead to negative end user experiences. This paper considers an alternative approach using low cost temperature sensors, open source software and hardware, and interactive dashboard to assist end users in optimizing their energy use with a view to reducing demand and fuel costs.
The potential benefits of cognitive radio networks (CRNs) have urged the wireless community to incorporate them as a solution for potential scarcity and spectrum reutilization. An undersampled autocorrelation-based spectrum sensing scheme that provides orthogonal frequency division multiplexing (OFDM) signal detection in undersampled channels is proposed. The paper introduces a sampling factor that provides several undersampling levels. The resulting undersampled conventional autocorrelation-based (CAB) algorithm is analyzed. The statistical distribution of test statistic is evaluated by employing maximum likelihood estimation (MLE) to extract statistical parameters.
One of the most common tasks performed in VR is navigation through a virtual environment. Real walking, which allows the user to physically move through a physical tracking space as they navigate in a virtual environment, is one of the most natural navigation methods available for VR. However, while virtual worlds can take up vast amounts of space, the physical space the user can walk in is limited. Redirected walking seeks to alleviate this limitation by letting the user walk in a virtual environment that is larger than the physical tracking space available by decoupling the mapping of the user's real and virtual movements. In this paper, we describe an experiment system that evaluates peoples' ability to accurately turn at different levels of rotation gain. We further report on a commissioning experiment that tests the functioning of this research system. Specifically, we asked participants to rotate in two virtual environments, a minimal scene and a visually complex scene, with different levels of induced gain and repeat this task, as a control in the physical world, with their eyes open and then closed. It was found that participants rotated less at higher gain levels, especially in the complex scene.
The theory of partial coherence is a significant part of Fourier optics, which has been utilized in numerous areas and applications. Therefore, the simulation of partially coherent systems is important for the system analysis and design of optical signal processing and other applications. Therefore, it is useful to identify the differences and strengths of existing simulation methods. In this paper, we compared three partially coherent field simulation algorithms including the random screen, superposition, and coherent mode decomposition methods based on their simulation results. Finally, we identified the optimal usage scenarios for each algorithm.
Blood pressure (BP) monitoring via cuffless devices has gained significant attention in the last few years. Despite a plethora of works having been produced in this field based on traditional machine learning (ML) or deep learning (DL) models, very limited research has been carried out in terms of the external validation and reproducibility of said models to ensure that they are of clinical use. To the best of the authors’ knowledge, this is the first study to evaluate several of the currently most well cited ML/DL-based models for cuffless BP monitoring over multiple independent data sets. The results of this investigation in reproducibility are reported with particular recommendations provided regarding standardized data collection protocols, models and signals, data recording length, and open access data as potential steps to overcoming the challenge of reproducibility in ML/DL models in this field and the health domain in general.
Future energy-harvesting (EH) brain-machine interfaces (BMI) present fundamental design challenges to analog-to-digital converters (ADC) for neural signal digitisation and recording. Ultra-low power (ULP) consumption, maximised area density, bandwidth and dynamic range as well as amenability to ultra-low voltage (ULV) supply are all desirable performance specifications. This article presents a digitally intensive openloop voltage-controlled-oscillator (VCO)-based ADC operating at an ULV supply of only 0.2V in 28nm CMOS. We introduce a novel design framework harnessing the spatio-temporal dynamics of coupled oscillator ensembles (COE) for open-loop voltage-to-frequency (V -to-f) analog linearisation, as applied to high impedance input transconductor-driven VCOs. A machine learning (ML)-driven foreground calibration engine employing gradient descent automatically optimises the integrated digitally programmable ‘tuning knobs’ of the COE network to guide its self-adaptive linearisation. A Verilog-A calibration engine integrated with the transistor level COE was developed to perform full-system transient simulations (lasting several miliseconds) which reduced the dominant third-order harmonic distortion (HD3) from −35dBc to −70dBc for a near rail-to-rail differential input voltage swing of ±180mV, verifying the calibration process.
WebTransport is a new protocol which enables the opening of a persistent connection between a client and server for bi-directional, uni-directional and datagram data exchange. The WebTransport protocol is currently in a draft state within the Internet Engineering Task Force (IETF). WebTransport is designed to use the QUIC protocol as its transport mechanism, which provides packet ordering and delivery management. This article provides an overview of the WebTransport protocol, summarises prior work on benchmarking QUIC, and compares WebTransport and WebSockets. The findings indicate that WebTransport does not match the performance of WebSockets with the current language implementations.
This paper presents a technique to limit the noise multiplication of operational amplifier used in the bandgap core without adding any extra component. This is achieved by shifting the resistor used in the emitter of low current density bipolar to its base. Further this resistor is combined with the existing CTAT resistor in the current mode bandgap reference. Also a high gain self-bias opamp was presented to minimize the systematic offset of the opamp leading to an improved accuracy. The proposed circuit is designed in TSMC 28nm CMOS technology and post-layout simulation results were performed. The targeted output voltage is 250mV and TC of 18.4ppm/ ° C. The low frequency noise power spectral density (PSD) is 4.5 times lower compared to the convention architecture. Proposed architecture consumes 54.6µW power from 1V nominal supply and occupies an area of 3690 µm 2 silicon area including decoupling capacitors.
For the first time, proposed and experimentally demonstrated is a time compressed higher speed CAOS (Coded Access Optical Sensor) camera. Specifically, proposed are new Code Division Multiple Access (CDMA) CAOS mode Walsh-matrix based image irradiance encoding techniques that allow faster encoding times for One Dimensional (1-D) and Two Dimensional imaging (2-D). CAOS imaging and spectrometer optical designs are introduced for engaging these faster encoding methods, including a novel 2-D folded spectrum high resolution spectrometer. Experimentally demonstrated is a 16 times faster CAOS imager for 1-D imaging. Computer simulations verify a 16 times faster 2-D image recovery via the time compressed CAOS.
Fractional-N phase locked loops (PLLs) are widely used in electronic systems. The architecture of the divider controller has a significant influence on the performance of a fractional-N frequency synthesizer. The divider controller modulates the instantaneous division ratio of the multi-modulus divider in the feedback path. It works by performing quantization on its input signal. The process of quantization introduces noise. Although the quantization noise and its running sum can be made spur free, they are subjected to nonlinear distortions which results in spurious tones in the output signal. As spurious tones are undesired, different divider controllers have been promoted to reduce nonlinearity-induced spurs. Multi-stage noise shaping (MASH), Enhanced nonlinearity-induced noise performance (ENOP), Successive requantizers (SR), and Multi-stage noise shaping structure-successive requantizer (MASH-SR) are four divider controllers that have demonstrated spur immunity after polynomial distortion. This paper compares the simulated performances of these four divider controllers using MATLAB ® . Polynomial and piecewise-linear (PWL) nonlinearities are considered. We show that ENOP P3 exhibits the best performance among the five different divider controllers that are considered in this paper.
Epileptic seizures affect more than 50 million population worldwide. Automated methods for seizure detection from EEG can help detect seizures faster and reduce the diagnostic delay. While there are many deep learning solutions to seizure detection the transferability of the learnt representation across age groups has not been studied. In this paper, we first evaluate the performance of the state-of-the-art neonatal seizure detection model on a publicly available pediatric CHB-MIT EEG dataset. The obtained results are then contrasted with the performance of fine-tuned neonatal model on pediatric data. The developed patient-independent model achieved an average AUC score of 91.62% on the CHB-MIT dataset. This is the first study to assess whether a universal model is realizable for different age groups.
This study presents an analysis of the use of machine learning models in the identification and classification of ransomware encrypted files, differentiating them from standard encrypted or compressed files, and non-encrypted files (referred to as goodware). The study utilized a robust dataset of approximately 159,897 files, categorized into goodware, Chaos, Conti, and Xorist strains, and applied five machine learning models: Logistic Regression, Linear Discriminant Analysis, K-Nearest Neighbor, Naive Bayes, and Classification and Regression Trees to this dataset. The models were trained using an array of data points, including file headers and footers, entropy, Chi Squared, and file extensions. The analysis revealed high accuracy rates of between 97% and 100% in distinguishing ransomware encrypted files from other file types, demonstrating the importance of file extensions as a key determinant in this process. The study also draws attention to the increasing prevalence and complexity of ransomware strains, specifically those which do not alter file extensions, thereby posing additional challenges to identification and classification efforts. The research suggests further investigation and study into a wider array of ransomware strains and a more extensive range of file types. Special emphasis is recommended on strains that do not modify file extensions, as understanding these could significantly enhance the efficiency and effectiveness of machine learning models in ransomware detection.
This paper aims to intuitively explain and numerically verify the locking range of a sub-harmonic injection-locked oscillator (ILO) when applying different injection multiplying ratios (M). By considering an impulse sensitivity function (ISF), which is the phase response of an oscillator against current impulse perturbation obtained from a periodic small-signal analysis using a PXF engine together with an injection waveform, the relative phase of the locked oscillator and consequently its locking range can be accurately predicted without requiring any time-consuming transient simulations. Through this approach, the study of which waveform patterns would be suitable for a locking range enhancement of an ILO can easily be carried out. Based on the proposed method, we have illustrated and numerically verified that, when injecting a sinusoidal signal to drain of the main NMOS cross-coupled pair, the ILO can achieve a much wider locking range when M is odd (i.e. 3) compared to when M is even (i.e. 2).
Driver fatigue is a major factor in road accidents. To enhance road safety, this study proposes a novel deep learning model for detecting drivers' respiration rates using a thermal camera, an essential parameter for assessing drowsiness levels. Our approach predicts respiration rates directly without signal extraction from facial regions of interest, simplifying the detection process and potentially improving drowsiness detection systems. We evaluate and explored the model using capabilities on the new data acquired in a simulated driving environment which is divided in two subsets i.e. non-noisy and noisy datasets. Additionally, we introduce a unique data augmentation technique to reduce over-fitting in deep learning models utilizing temporal data. The implementation of this respiration detection model may contribute to driver drowsiness detection systems and enhance road safety.
The development of accurate accent-recognition technologies is important towards high-performance Automatic Speech Recognition (ASR) through the use of accent-specific models. Usually deep-learning models are employed to perform ASR and we intend to enhance such models by accurately capturing the accent-related idiosyncrasies of non-native English speakers. The typical features used for the deep-learning networks, in ASR, are Mel-Frequency Cepstral Coefficients (MFCCs) and Mel-Spectrogram of the audio samples. However, novel features based on the Hilbert-Huang Transform (HHT), an effective non-linear signal analysis technique can aid or even replace typical features used in Automatic Speech Recognition. In this paper, we seek to apply these features in the task of accent recognition for non-native English speakers.We propose a new novel input feature, the Hilbert Mel-Spectrogram, a logarithmic frequency representation of the signal derived using the HHT. The use of HHT creates an input feature free of the traditional limitations of Fourier analysis, namely limited frequency resolution, spectral energy leakage, and harmonic artifacts. A four-stage Convolution Neural Network (CNN) is used to model the features of accented speech samples spanning 5 native languages. Our results show that the models with the Hilbert-Mel-Spectrogram-based features outperform their Fourier-based counterparts for a wide range of datasets.
Image blurring prediction has a significant role in image restoration, forensics, and computer vision. Because of both the sharpened blurred image and the original image enlarge the image noise differently. To discern this difference, an autocorrelation measurement is suggested. This paper proposes a non-reference image blurring detection scheme based on the specificity of Moran statistics and UM (Unsharp Masking). Unsharp Masking is sensitive to image noise and extends noise in a sharpening process. The Moran's Z histogram has been successfully used to discern blurriness and sharpness in an image. The blurred image can be distinguished from images using an examination method that uses the Moran's Z score histogram on a UM pre-processed image. The Z histogram median values for blurred images are all higher than those from the original image, having a value of approximately 9. Based on these effects, image blur judgment factors are constructed. The features found in this work based on Moran statistics can be in the future fed into a support vector machine (SVM) classifier for classification.
Future wireless networks demand very high-speed and reliable connectivity to realize smart industrial and manufacturing systems. Furthermore, there has been a tremendous growth in the number of connected sensors and devices, requiring efficient utilization of radio resources. To achieve this, we propose a user grouping and scheduling framework for hybrid multiple access scheme that combines non-orthogonal multiple access (NOMA) with orthogonal frequency division multiple access (OFDMA). First, we analyse the impact of channel gain difference among a group of power domain NOMA users on the bit error rate (BER) performance and achieved throughput. Based on the analysis, we develop a throughput-aware user grouping scheme for multiple ultra reliable low latency communication (URLLC) users, each with its own demanded throughput, taking the channel gain differences into consideration. For low latency operation, we propose a scheduling method that aims to minimize the number of utilized time slots, for a given bandwidth. The proposed scheme is evaluated in terms of BER, achieved throughput, and fairness of the grouped users to indicate that it fulfils the demanded throughput of each user, while providing lower latency compared to OMA only scenario. It is shown that OFDMA-NOMA with the proposed scheme outperforms a previously proposed user pairing scheme for uplink NOMA.
This paper presents the performance of a battery-less near field communication (NFC) sensor for museum artifact monitoring. The radio frequency (RF) receiver sensitivity of the NFC sensor was measured using a commercial Tagformance measurement system. To respond to read commands, the NFC sensor requires a minimum magnetic field strength of 67.8 mA/m in free space, and 87.6 mA/m after integration within an archive box typically used for cultural heritage artifact storage. The integrated NFC loop antenna has a -3 dB bandwidth of 1.16 MHz in free space and 1.14 MHz bandwidth when integrated within the archive box. These measured RF sensitivity and -3 dB bandwidth values are in close agreement with the ISO/IEC 15693 standard. This work demonstrates a DC power consumption of 597 µW which is shown to be the lowest value compared to the state-of-the-art literature. In addition, the wireless communication range of approximately 5 cm was achieved which compares favourably with the maximum read ranges reported in literature. The developed NFC sensor has been deployed in European museums for wireless monitoring of valuable artifacts.