
Wireless Acoustic Sensor Networks (WASNs) have a wide range of audio processing applications. Due to the spatial diversity of the microphone and their relative position to the acoustic source, not all microphones are equally useful for subsequent audio signal processing tasks, nor do they all have the same wireless data transmission rates. Hence, a central task in WASNs is to balance a microphone’...
Informed speaker extraction aims to extract a target speech signal from a mixture of sources given prior knowledge about the desired speaker. Recent deep learning-based methods leverage a speaker discriminative model that maps a reference snippet uttered by the target speaker into a single embedding vector that encapsulates the characteristics of the target speaker. However, such modeling deliberately neglects the time-varying properties of the reference signal. In this work, we assume that a reference signal is available that is temporally correlated with the target signal. To take this correlation into account, we propose a time-varying source discriminative model that captures the temporal dynamics of the reference signal. We also show that existing methods and the proposed method can be generalized to non-speech sources as well. Experimental results demonstrate that the proposed method significantly improves the extraction performance when applied in an acoustic echo reduction scenario.
Recently proposed automatic pathological speech classification techniques use unsupervised auto-encoders to obtain a high-level abstract representation of speech. Since these representations are learned based on reconstructing the input, there is no guarantee that they are robust to pathology-unrelated cues such as speaker identity information. Further, these representations are not necessarily discriminative for pathology detection. In this paper, we exploit supervised auto-encoders to extract robust and discriminative speech representations for Parkinson’s disease classification. To reduce the influence of speaker variabilities unrelated to pathology, we propose to obtain speaker identity-invariant representations by adversarial training of an auto-encoder and a speaker identification task. To obtain a discriminative representation, we propose to jointly train an auto-encoder and a pathological speech classifier. Experimental results on a Spanish database show that the proposed supervised representation learning methods yield more robust and discriminative representations for automatically classifying Parkinson’s disease speech, outperforming the baseline unsupervised representation learning system.
This study is concerned with the relation between the information-theoretic notion of surprisal and articulatory gesture in Polish consonant-to-vowel transitions. It addresses the question of the influence of diphone predictability on spectral trajectories and articulatory gestures by relating the effect of surprisal with motor fluency. The study combines the computation of locus equations (LE) with kinematic data obtained from electromagnetic articulograph (EMA). The kinematic and acoustic data showed that a small coarticulation effect was present in the highand low-surprisal clusters. Regardless of some small discrepancies across the measures, a high degree of overlap of adjacent segments is reported for the mid-surprisal group in both domains. Two explanations of the observed effect are proposed. The first refers to low-surprisal coarticulation resistance and suggests the need to disambiguate predictable sequences. The second, observed in high surprisal clusters, refers to the prominence given to emphasize the unexpected concatenation.
In recent research, i-vectors have been shown to be significantly beneficial for speaker recognition and have been successfully applied in deep neural network (DNN) acoustic model (AM) training to improve the performance of automatic speech recognition (ASR). This paper describes our work in developing a bilingual i-vector extractor for training a German speech recognition system. A bilingual data...
In some cases, a speech communication link is necessary in underwater environments. While in terms of data communication advanced techniques are investigated and employed, mostly traditional (analog) forms of speech transmission are used for speech communication (at least in commercial products). In this contribution we present a mixed analog-digital approach, that tackles this problem by combinin...
In the human cortex, the auditory signal is segmented robustly into syllables using theta-oscillations, where the phase and instantaneous frequency of each theta-cycle corresponds to the position and duration of a syllable. Recently, in the superior temporal gyrus, ensamples of neurons sensitive to edge features have been detected, which spike at the maximal rise of the envelope of the auditory si...
Anglicisms pose a challenge in German speech recognition due to their irregular pronunciation compared to native German words. To solve this issue, we propose a comparative approach that uses both a German and an English grapheme-to-phoneme model to create Anglicism pronunciations. Comparing their confidence measures, we chose the best resulting pronunciations and added them to an Anglicism pronun...
This paper proposes to unify two deep-learning methods, Count- Net and Deep Clustering, designed for speaker count and separation respectively, in order to perform speaker count agnostic speech separation. Two approaches are compared, where the speaker count estimation and separation subnetworks are either trained separately or jointly. Training and evaluation are conducted on a tailored dataset W...
Acoustic feedback compensation in speech applications often suffers from an insufficient adaptation. As a result feedback whistling may be audible and the compensation gain and generally the speech quality is limited. In this work, we introduce the concept of an energy-decay operator and show how to use it for tap-selective step-size control of a multi-delay frequency-domain filter (MDF). Therefor...
Binaural recording using artificial heads or microphones on or near the ears of a person, e.g., on modern headphones and hearables, is a well-established technology for spatial audio capture. Our previously presented Binaural Cue Adaptation method improves the reproduction of binaural recordings by adapting the signals recorded for a fixed head orientation to a listener’s dynamic head movements. I...
There are several approaches for localising sound sources. This contribution utilises an eight-channel microphone array that is mounted on an artificial head and spatially samples the direct vicinity of both ears. This allows for a localisation approach that resembles the experience of a human listener while avoiding the limitations of an artificial head (e.g., with respect to front-back confusion...
Despite their small share in overall signal energy, plosives have been previously shown to be important for speech perception. We propose a simple, yet effective, model-based phase-aware speech enhancement approach specifically targeted at plosives. Starting from a model of the plosive burst as a unit impulse, we introduce three phase enhancement schemes: simple replacement of the noisy phase with...
State-of-the-art drone detection systems are generally combining different sensors, such as radar, acoustic, radio frequency (RF-) or optic sensors, each of which contributes individual capabilities. Acoustic sensors have a relatively low spatial range. On the other hand, they can be used in conditions of bad visibility and for scenarios without line-of-sight between sensor and drone. The developm...
A reproducible acoustic environment forms the basis for scientific evaluations of speech enhancement systems, as well as their influence on speech quality and intelligibility. Studies with people usually use binaural recordings played back through headphones. However, the effort required to evaluate such a speech enhancement system in this way can quickly become impractical, especially if the syst...
Speech communication under adverse conditions may be extremely stressful for the person located at the receiving (or near-end) side. There, background noise may originate from the given environment and cannot be reduced for the listener. Additionally, processed speech from the far-end might be degraded by the transmission network, due to e.g., transcoding or packet loss. The assessment of several ...
Spoken language skills are strong biomarkers for detecting cognitive decline. Studies like the Interdisciplinary Longitudinal Study of Adult Development and Aging (ILSE) are of particular interest to quantify the predictive power of biomarkers in terms of acoustic/linguistic features. ILSE consists of ca. 6500 hours of interviews and only 10% were manually transcribed. To extract linguistic featur...
Speech communication and dialog systems inside vehicles are usually optimized for a speaker in the driver’s seat who is permanently facing forward. However while driving, a speaker does repeatedly rotate his head left and right, which changes the properties of the signal picked up by the microphone. This work analyzes to what extent such a speaker head rotation impacts the signal in terms of frequ...
Recently a method has been proposed to blindly estimate the geometry of an array of distributed microphones using reverberant speech, which relies on estimating the coherence matrix of the reverberation using an iterative expectation conditional-maximization (ECM) approach. Instead of using a data-independent initial estimate of the coherence matrix and a matched beamformer to estimate the initial...
While Fourier phase has long been considered unimportant for speech enhancement in the short-time Fourier domain, phase-aware speech processing is receiving increasing attention in recent years. Among other advances, it has been shown that when using very short frames (< 2 ms), the phase spectrum carries enough information to allow an intelligible reconstruction from phase alone, whereas the ma...