Tiny Signal-to-Interpretation (TinyS2I) has been recently introduced as an ultra low-footprint end-to-end spoken language understanding (SLU) model. This architecture is capable of running in ultra resource constrained environments like voice assistant devices, while at the same time reducing latency. In this work, we propose an extension to TinyS2I and train a multilingual system supporting several languages. Multilingual TinyS2I models show little to no degradation compared to their monolingual counterparts. Increasing the network size in width and depth improves the classification accuracy for mono- and multilingual setups, with the multilingual one improving beyond the monolingual accuracy. This enables users to interact with the device in the language of their choice and dynamically switch between languages without an explicit language setting or accuracy degradation.
To achieve robust far-field automatic speech recognition (ASR), existing techniques typically employ an acoustic front end (AFE) cascaded with a neural transducer (NT) ASR model. The AFE output, however, could be unreliable, as the beamforming output in AFE is steered to a wrong direction. A promising way to address this issue is to exploit the microphone signals before the beamforming stage and after the acoustic echo cancellation (post-AEC) in AFE. We argue that both, post-AEC and AFE outputs, are complementary and it is possible to leverage the redundancy between these signals to compensate for potential AFE processing errors. We present two fusion networks to explore this redundancy and aggregate these multi-channel (MC) signals: (1) Frequency-LSTM based, and (2) Convolutional Neural Network based fusion networks. We augment the MC fusion networks to a conformer transducer model and train it in an end-to-end fashion. Our experimental results on commercial virtual assistant tasks demonstrate that using the AFE output and two post-AEC signals with fusion networks offers up to 25.9% word error rate (WER) relative improvement over the model using the AFE output only, at the cost of <= 2% parameter increase.
We present dual-attention neural biasing, an architecture designed to boost Wake Words (WW) recognition and improve inference time latency on speech recognition tasks. This architecture enables a dynamic switch for its runtime compute paths by exploiting WW spotting to select which branch of its attention networks to execute for an input audio frame. With this approach, we effectively improve WW spotting accuracy while saving runtime compute cost as defined by floating point operations (FLOPs). Using an in-house de-identified dataset, we demonstrate that the proposed dual-attention network can reduce the compute cost by 90% for WW audio frames, with only 1% increase in the number of parameters. This architecture improves WW F1 score by 16% relative and improves generic rare word error rate by 3% relative compared to the baselines.
Neural contextual biasing for end-to-end neural ASR transducers has shown significant improvements in the recognition of named entities, such as contact names or device names. However, it comes with the cost of increased compute, as the biasing layers (which are usually based on cross-attention) add complexity to the neural transducers. In this paper, we propose gated contextual biasing models that can estimate at runtime when contextual biasing is needed and can toggle it on or off. That way, contextual biasing does not run on every audio frame, but only on the frames where it can be helpful for correct ASR recognition. We show that our gated contextual biasing models can maintain all the performance improvements of contextual biasing while offering significant compute-cost saving, as the contextual biasing needs to be executed for fewer than 15% of the audio frames.
The Internet of Things (IoT) has demonstrated promising growth, as it is crucial to numerous application domains in smart ecosystems, such as the smart city. Decentralization can help achieve growth, with previous research proposing the use of the decentralized blockchain technology in IoT ecosystems as, among others, it can offer cryptographically trustworthy interactions, nonrepudiable smart contracts, and interoperability between stakeholders. However, proposals are often theoretical or very specific to IoT ecosystem subareas, for instance, healthcare. In both cases, implementation may not be feasible in a larger ecosystem. This article investigates the performance, particularly throughput, aspect of feasibility. It proposes and demonstrates a platform that uses the blockchain technology as a building block in an IoT ecosystem. A permissioned blockchain network is built, performance is measured, and the measurements are interpreted in a smart city context, thus, providing valuable insights about real-world implementations and their degree of feasibility.
The concept of smart cities has gained popularity due to technological advances in areas such as the Internet of Things (IoT) and Big Data Analytics (BDA). Location-based services have emerged in such smart environments to improve people’s quality of life and generate statistics for mutual benefit. In this work, a stylized presence detection concept is proposed which uses Bluetooth Low Energy (BLE) beacons placed in locations of interest. Users can detect the BLE beacon identification number (ID) with personal devices such as cell phones and connected watches and transmit it along with a unique and randomly generated user ID. Blockchain technology is used for a storage back-end. Our proposal is by no means exhaustive and is intended to advance the discussion of location-based services that deal with big data.
On-device spoken language understanding (SLU) offers the potential for significant latency savings compared to cloud-based processing, as the audio stream does not need to be transmitted to a server. We present Tiny Signal-to-interpretation (TinyS2I), an end-to-end on-device SLU approach which is focused on heavily resource constrained devices. TinyS2I brings latency reduction without accuracy degradation, by exploiting use cases when the distribution of utterances that users speak to a device is largely heavy-tailed. The model is tailored to process on-device frequent utterances with support for dynamic contextual content, while deferring all other requests to the cloud. Compared to a powerful baseline, we demonstrate that TinyS2I achieves comparable performance, while offering latency gains due to local processing.
New healthcare record management (HRM) systems have been introduced as technology has evolved to provide more efficient care. Since medical data is usually sensitive and must be protected from unauthorized access, attention must be paid to data integrity, patient privacy, and storage. Blockchain technology has been proposed in the literature to integrate healthcare information systems through a decentralized and unified network. However, the literature on blockchain in healthcare is full of promises that may not be true under certain conditions. In our paper, we evaluate the veracity and sophistication of some of the claims made in the literature. We go beyond performing a literature review and shed light on the weak technical aspects claimed about blockchain. In addition, we benefit from our technical assessment and suggest some future research directions to improve healthcare systems that use blockchain and big data solutions.
We introduce Caching Networks (CachingNets), a speech recognition network architecture capable of delivering faster, more accurate decoding by leveraging common speech patterns. By explicitly incorporating select sentences unique to each user into the network's design, we show how to train the model as an extension of the popular sequence transducer architecture through a multitask learning procedure. We further propose and experiment with different phrase caching policies, which are effective for virtual voice-assistant (VA) applications, to complement the architecture. Our results demonstrate that by pivoting between different inference strategies on the fly, CachingNets can deliver significant performance improvements. Specifically, on an industrial-scale, VA ASR task, we observe up to 7.4% relative word error rate (WER) and 11% sentence error rate (SER) improvements with accompanied latency gains.
Blockchain technology is enabled by consensus algorithms to manage the relationships among several economic or business operators without human intervention. With the help of consensus algorithms, distributed systems can reliably reach agreement even if part of the system is faulty. Blockchain yields many benefits, among others, traceability, transparency, and security. We consider using the RAFT consensus algorithm to achieve robust and scalable decentralized applications, with focus on healthcare. We propose a stylized healthcare network, enabled by RAFT and built upon Hyperledger Fabric to showcase the use of RAFT in healthcare blockchain. However, RAFT is by no means limited to healthcare record systems, and can be applied to any other record system and value chain. Our paper offers several insights to those working in value chains and information management-related fields. In addition, we end our study with some future research avenues that may inspire managers and scholars to build or refine new decentralized systems in healthcare and other related fields.
With the rapid development of advanced biomedical sensors, the Internet of Things, and modern wireless communication technologies, smart healthcare systems provide feasible solutions to the problems of population aging and telemedicine services. However, physiological information involves personal privacy, and the data security concern during the transmission of information on public channels has become a critical issue that has restricted the wider acceptance of smart healthcare systems. In this article, a secure and trustworthy smart healthcare system interoperating with wireless body area networks based on multistage blockchain is proposed. The security scheme provides a completely secure and trustworthy environment covering the entire data flow from the front-end to the back-end. The evaluation results show that the encryption scheme based on piecewise linear chaotic map can resist common attack methods and reduce encryption time. The system has the advantages of low complexity for encryption scheme, and larger capacity and efficiency for the blockchain-based data transmission and storage system.
Compression and quantization is important to neural networks in general and Automatic Speech Recognition (ASR) systems in particular, especially when they operate in real-time on resource-constrained devices. By using fewer number of bits for the model weights, the model size becomes much smaller while inference time is reduced significantly, with the cost of degraded performance. Such degradation can be potentially addressed by the so-called quantization-aware training (QAT). Existing QATs mostly take into account the quantization in forward propagation, while ignoring the quantization loss in gradient calculation during back-propagation. In this work, we introduce a novel QAT scheme based on absolute-cosine regularization (ACosR), which enforces a prior, quantization-friendly distribution to the model weights. We apply this novel approach into ASR task assuming a recurrent neural network transducer (RNN-T) architecture. The results show that there is zero to little degradation between floating-point, 8-bit, and 6-bit ACosR models. Weight distributions further confirm that in-training weights are very close to quantization levels when ACosR is applied.
We present a systematic approach for audio mixing based on synchronized User-Generated audio Recordings (UGRs), e.g., audio recordings contributed by users attending the same public event. We discuss the challenges that relate to creating a mixture with such recordings, mainly due to the fact that each audio stream spans a different portion of the event of interest and comes with different signal level characteristics. We propose an approach to combine the available recordings based on a normalization step and a mixing step. The normalization step defines a fixed-with-time gain that is specific to each UGR. In the mixing step, a mechanism that reduces the master gain in accordance with the number of activated inputs at each time is employed. An approach called orthogonal mixing is presented, which is designed based on the assumption that the mixture components are mutually independent. The presented mixing process allows the combination of multiple short duration UGRs to produce a longer audio stream with potentially better quality than any one of its constituent parts.
Due to the increasing population of the elderly and patients with chronic diseases, more and more individuals are suffering from the limited service capabilities of the traditional medical systems. Benefiting from the rapid development of biomedical sensors, Internet of Things, and modern communication and network technologies, eHealthcare systems start appearing in the medical services, especially for the remote physical condition monitoring which improves the efficiency of the traditional medical systems. To provide a secure and low power healthcare solution, a blockchain-based eHealthcare system interoperating with wireless body area networks (WBAN) has been proposed, which utilizes the WBAN to network the devices of the patients and the blockchain technology as the data transmitting and storage method. The evaluation results show that the proposed system has the advantages of low hardware resources utilization, high-security protection level, and stable performance.
Securing the access in networks is a first-order concern that only gains importance with the advent of Internet of Things (IoT). In this paper, a security system is presented for password-free access over the secured link. It makes the connection faster than manual authentication and facilitates Machine-to-Machine (M2M) secure interactions, as required for IoT. The authentication procedure includes the exchange of certificate and challenge/response pairs, which are stored and computed in an external security coprocessor. The system enforces the authentication protocol, includes error detection, and handles multiple devices according to their Operating Systems (OS) through their connections/ disconnections. It also performs encryption, if necessary. It is applicable on application level for devices, including IoT based devices, sensors, Android, and iOS-based smartphones. The devices that have the correct certificate and can solve the challenge can connect to the network linked with the security system. The system security is hardened because the sensitive authentication elements such as keys, certificates, and challenge responses are invisible to users and are exchanged only using strong hashing algorithms that are irreversible. The proposed hardware security system can augment any supporting network, converting the entire insecure network into a secured one, as well as retrofit existing insecure Bluetooth devices for secure access. The system incurs low overhead in time and energy by performing security operations in an ASIC coprocessor, and can be shared to secure access to multiple devices, which reduces both energy and cost.
In recent years, the Wireless Body Area Network (WBAN) concept has attracted significant academic and industrial attention. WBAN specifies a network dedicated to collecting personal biomedical data from advanced sensors that are then used for health and lifestyle purposes. In 2012, the 802.15.6 WBAN standard was released by the Institute of Electrical and Electronics Engineers (IEEE), which regulates and specifies the configurations of WBAN. Compared to the prevailing wireless communication technologies such as Bluetooth and ZigBee, the WBAN standard has the advantages of ultra-low power consumption, high reliability, and high-security protection while transmitting sensitive personal data. Based on the standard specification, several implementations have been published. However, in terms of evaluation, different designs were implemented in proprietary evaluation environments, which may lead to unfair comparison. In this paper, a Software-Defined Radio (SDR) evaluation platform for WBAN systems is proposed to evaluate the RF channel specified in the IEEE 802.15.6 standard. A narrowband communication protocol demonstration with a security scheme in WBAN has been performed to successfully validate the design in the proposed evaluation platform.
Body Area Networks (BAN) have caused wide interest in academic and industrial areas in recent years due to the strong demands of health condition monitoring by people. Consequently, security becomes a critical consideration, which restricts the development of BAN while the data transferring in BAN is increasingly significant and their privacy is of the utmost importance. Therefore, an ASIC implementation of security scheme for BAN is proposed by the authors, based on the IEEE standard 802.15.6, which is synthesized using the SIMC 65nm CMOS technology.
In this paper, we consider the data-association problem for the localization of multiple sound sources in a wireless acoustic sensor network, where each node is a microphone array, using direction of arrival (DOA) estimates. The data-association problem arises because the central node that receives the multiple DOA estimates from the nodes cannot know to which source they belong. Hence, the DOAs from the different nodes that correspond to the same source must be found in order to perform accurate localization. We present a method to identify the correct association of DOAs to the sources and thus accurately estimate their locations. Our method results in high association and localization accuracy in realistic scenarios with missed detections, reverberation, noise, and moving sources and outperforms other recently proposed methods. It also incorporates a bitrate reduction scheme in order to keep the amount of information that needs to be transmitted in the network at low levels without affecting performance.
The development in the field of advanced biomedical sensors has resulted in a large volume of data being collected and transmitted wirelessly. The IEEE 802.15.6-2012 standard has specified a communication protocol for body area networks, however existing general-purpose communication protocols such as Bluetooth and Zigbee are more widely used due to multiple reasons. One of the critical issues is the lack of baseband processing hardware modules that implement the aforementioned standard. In this paper, the authors propose a baseband transceiver implementation in ASIC, which meets the 802.15.6-2012 standard requirements. Compared to other published designs, the proposed implementation exhibits better performance and low hardware cost, while also offering a complete standard implementation.
Body area network (BAN) consists of intercommunicating sensors, which are either wearable, or can be implanted into the human body. Being a network that features critical information, BAN requires encryption, however common encryption methods found in computer networks are not ideal due to complex algorithms and high power consumption. The authors propose a new encryption method based on the human body channel. This new encryption method has the advantages of low-power, dynamic updating, rapid operation, and easy implementation, which is suitable for BAN systems.