
While speech recognition has become highly robust in the recent past, it is still a challenging task under very noisy or reverberant conditions. Augmenting speech recognition by lipreading from video input is hence a promising approach to make speech recognition more reliable. For this purpose, we consider slow feature analysis (SFA), an unsupervised machine learning method that finds temporally slowest varying features in sequential input data. It can automatically extract temporally slow features within a video sequence, such as lip movements, while at the same time removing quickly changing components such as noise. In this work, we apply SFA as an initial feature extraction step to the task of automatic lipreading. The performance is evaluated on small-vocabulary lipreading, both in the speaker-dependent and speaker-independent case, showing that the features are competitive to the often highly successful combination of a discrete cosine transform and a linear discriminant analysis, while also offering good interpretability.
Research of gender effects in the field of automatic speech emotion recognition (SER) has been subject to research in the past. Still, however, it is somewhat unclear to which degree speaker gender influences SER. Although we will prove it to be wrong for our SER model, usually it is assumed that a "gender-dependent" emotion recognizer performs better than a "gender-independent" one by using gender-dependent emotional speech training data. In this paper, we use a state-of-the-art emotion recognizer model to investigate the effects of speaker genders in training and in test. The gender-specific SER model performs well under matched-gender test conditions, as expected. Training with approximate half the amount of the training data from male and female speakers jointly, is about as good as separate training using gender-specific data and test in matched conditions with in total the same amount of training data. This however, only holds for an average over male and female speakers: Interestingly, we show that female voices emotions are recognized better on a mixed-gender SER model, than on a female SER model, indicating that female speakers express emotions in a wider variety than male speakers?
Supervectors represent speaker-specific Gaussian Mixture Models which are enrolled from a Universal Background Model (UBM) and approximate the unknown, underlying speech feature distributions. But as supervectors only consist of the stacked means of the Gaussian components, lowdimensional i-vectors which are derived from them do not completely capture the true feature distributions. In this work, the classical supervectors are extended with additional parameters before reducing their dimension to capture the feature distributions more accurately and complement the i-vectors more effectively. To extend a supervector, the mixture weights, the log-likelihood values of the UBM, a Bhattacharyya-distance based kernel and the Hellinger distance between each enrolled Gaussian component and the corresponding one of the UBM are used. In closed-set speaker identification experiments conducted on the NTIMIT corpus which consists of telephone quality speech, the extended supervectors provide significantly lower error rates than the standard supervectors, even after fusing them with i-vectors and the UBM.
The emerging field of wireless acoustic sensor networks (ASN) offers promising future applications, but at the same time entails several challenges for audio signal processing. One particular task is that of identifying the acoustical system between a source of interest and the receivers in the form of acoustical transfer functions (ATF). In ASN, ATF estimation is essentially a blind problem due to unknown sensor geometry and unavailable source signal and is further complicated by noisy environments and non-persistent excitation. In this paper, we therefore put the blind identification problem into a frame-based maximum-likelihood (ML) context, before we extend this data-driven method with a-priori information to a new maximum-a-posteriori (MAP) approach. The latter shall compensate for the non-persistent excitation in ASN. We further propose a measure for assessing the accuracy of blindly identified ATF and demonstrate in computer experiments that the MAP approach is superior to ML provided that accurate estimates of the source activity are available.
Stroke survivors often suffer from oro-facial impairments, affecting swallowing function and speech production. Measuring tongue pressure and position intraorally can help to improve therapy for both symptoms, but space inside the oral cavity is extremely limited and such devices can easily be prohibitively large and obstructive if too many sensors are needed. In this work, we present our efforts to sense the force of the tongue exerted against the hard palate and the tongue-palate distance, using only optical proximity sensors. To explore the feasibility and accuracy of this approach and to evaluate the selected sensor, we conducted a study with 10 subjects and measured the sensor's response to 10 discrete distances ranging from 0mm to 30mmbetween tongue and sensor, and to a continuously increasing tongue force against the sensor from 0.1N to 8N. For distance measurements, an existing in-situ calibration method was applied and verified that yielded errors of less than 2mm for the estimated distances in nearly every case. For force measurements, a Bayesian classification approach was adopted to map sensor data to two force regions (below and above a certain boundary value), where up to 84.1% (average: 71.7 %) of ADC values were classified correctly within-sample.
This paper presents a novel approach for post-filtering the output of a hearing aid (HA) beamformer (BF) using an external microphone (Emic) signal where the Emic is placed in front of the HA user. The proposed approach first estimates the target relative transfer function (RTF) between the Emic and the hearing aids. The target RTF is then used for target and noise signal estimation. The estimated target and noise powers are then used to enhance the output of the HA-BF via a Wiener-like post-filter. Objective measures and subjective testing are used to evaluate the performance of the proposed post-filter (PF) with respect to the state-of-the-art single-channel noise reduction and previously proposed Emic-based post-filter designs. These evaluations show that the RTF-based approach for post-filtering the binaural HA BF outperforms previous approaches for noise reduction using an Emic and ultimately supports the use of an Emic for signal enhancement in hearing aids.
Diagnosing dementia early is crucial in mitigating the consequences of the disease for patients, their care-givers and relatives. We present automatic screening for personal transition into dementia from speech using information from more than one point in time. Using conversational speech data from the ILSE corpus we screen subjects if they transition from a cognitively healthy state to a state of dementia. We use both acoustic and linguistic features from two pipelines of feature extraction: the manual pipeline uses manual transcriptions while the fully automatic pipeline uses transcriptions created by automatic speech recognition (ASR). Using these two different feature extraction pipelines we automatically screen for dementia transition, where the fully automatic pipeline performs the whole screening process fully automatically. Our results show the features extracted from automatic transcriptions outperform the features extracted from the manual transcriptions.
Improving the accuracy of dysarthric speech recognition is a challenging research field due to the high inter- and intra-speaker variability in disordered speech. In this work, we propose to use estimated articulatory-based representations to augment the conventional acoustic features for better modeling of the dysarthric speech variability in automatic speech recognition. To obtain the articulatory information, long short-time memory recurrent neural networks are employed to learn the acoustic-to-articulatory inverse mapping based on a simulated articulatory database. Experimental results show that the estimated articulatory features can provide consistent improvement for dysarthric speech recognition with more improvement observed for speakers with moderate and moderate-severe dysarthria.
To improve the sound quality of hearing devices, equalization algorithms can be used that aim at achieving acoustic transparency, i.e., listening with the device in the ear is perceptually similar to the open ear. The equalization filter needs to ensure that the superposition of the processed and equalized signal played by the device and the signal leaking through the device into the ear canal matches a processed version of the signal reaching the eardrum of the open ear. Since equalization using a single loudspeaker typically does not allow for perfect equalization, in this paper we propose to use a multi-loudspeaker equalization filter to achieve acoustic transparency in a custom multiloudspeaker hearing device. The equalization filter is computed by minimizing a regularized least-squares optimization problem. Experimental results using measured acoustic transfer functions show that the proposed multi-loudspeaker equalization filter is able to provide the desired signal at the eardrum for different gains and delays of the hearing device.
The multi-channel Wiener filter (MWF) is a commonly used speech enhancement technique for improving speech quality and intelligibility in reverberant and noisy environments. The MWF is typically implemented as a minimum variance distortionless response (MVDR) beamformer followed by a single-channel Wiener postfilter. Assuming that reverberation and ambient noise can be modeled as diffuse sound fields, estimates of the relative early transfer function (RETF) vector of the target speaker and of the diffuse power spectral density (PSD) are required to implement the MWF. RETF vector and diffuse PSD estimation methods are often decoupled, i.e., one of the quantities is estimated assuming that the other quantity is known. In this paper, we aim at jointly estimating the RETF vector and the diffuse PSD by minimizing the Frobenius norm of an error matrix based on the presumed signal model. To solve this minimization problem, we propose to use an alternating least-squares approach. Simulation results using artificial and real data show that the proposed method leads to a better performance than a state-of-the-art method based on covariance whitening.
Blind source separation (BSS) and blind dereverberation (BD) are known to improve a desired speech signal when it is degraded by concurrent sound sources and reverberant environments, respectively. However, the BSS performance suffers from strong reverberation and BD usually suffers when there are multiple sound sources active. Thus, it has been proposed to connect both methods in tandem, so that they mutually profit from their respective gains. These previously presented schemes, however, work in batch mode preventing their direct use in real-time applications. In this paper a RLS-based method for adaptively and jointly performing BD and BSS of speech mixtures is proposed. It runs in fully online mode and can adapt to abrupt changes of target speaker positions. Experimental results show improvements in signal-to-interference ratio by an average of about 2dB compared to a sequential use of BD and BSS, as well as improvements in direct-toreverberant ratio.
Parkinson’s Disease (PD) is a neurodegenerative disorder which gradually effects the neurological condition of the patient. In many cases the disease impairs the reliability of the articulatory system and the ability to pronounce vowels normally. One prominent way to measure the degree of the functioning of the articulatory system is the Vowel Space Area (VSA). However, the typical way to measure it, is to manually annotate sustained vowel recordings or phonetically annotated speech utterances of a speaker and then analyze the signals. However, it is often desirable to measure the VSA directly from unlabeled natural speech. Therefore an automatic model-based system is proposed in this paper to estimate the triangular Vowels Space Area (tVSA) and the underlying corner vowel formant frequencies directly from unlabeled natural speech. The proposed algorithm is able to estimate the tVSA automatically from the speech signals without the need of phonetical or vowel transcriptions. The i-Vectors are extracted from the signals as the speaker’s characteristic representation, from which the speaker’s corner vowel formant frequencies are estimated by regression classifiers. Two regression classifiers, namely Deep Neural Networks (DNN) and Support Vector Regression (SVR), are investigated in this work. The proposed configuration employs the SVR classifier, which is able to predict the corner vowel formant frequencies of the test speakers with R 2 up to 0.56719 and ρ up to 0.76485.
For historical reasons, today's telephony frequencies are mostly still restricted to a narrowband of 3.4 kHz. Meanwhile wideband speech coding, also called HD voice, with a bandwidth of 7 kHz, is increasingly on the move into the networks and terminals. In the transition phase quite often a wideband terminal is receiving a narrowband signal. In order to achieve speech quality as close as possible to HD voice, artificial bandwidth extension (ABWE) has been developed as add-on for wideband equipment. State-of-the-art ABWE algorithms estimate and reconstruct missing frequency components with the help of the source-filter model of speech production. In this paper, a new method for estimating the model filter is presented, which is based on geometrical interpolation of the well known acoustical lossless tube model of speech production. Objective evaluation and subjective listening tests prove similar or even better quality compared to traditional more complex approaches.
In conventional speech enhancement, statistical models for speech and noise are used to derive clean speech estimators. The parameters of the models are estimated blindly from the noisy observation using carefully designed algorithms. These algorithms generalize well to unseen acoustic conditions, but are unable to reduce highly non-stationary noise types. This shortcoming motivated the usage of machine-learning-based (ML-based) algorithms, in particular deep neural networks (DNNs). But if only limited training data are available, the noise reduction performance in unseen acoustic conditions suffers. In this paper, motivated by conventional speech enhancement, we propose to use the a priori and a posteriori signal-to-noise ratios (SNRs) for DNN-based speech enhancement systems. Instrumental measures show that the proposed features increase the robustness in unknown noise types even if only limited training data are available.
Many speech intelligibility models base predictions only on low-level physical properties of the signals and ignore the task syntax. However, empirical data indicates an influence of the complexity of speech tests on their outcome. The Framework for Auditory Discrimination Experiments, which simulates the speech recognition process with an automatic speech recognition system, was extended to predict speech reception thresholds (SRTs) for two extreme assumptions about the sentence structure: No knowledge and perfect knowledge. The outcomes of three speech tests were predicted and compared to empirical data: Digit triplet test, Matrix sentence test, and Goettingen (everyday) sentence test. The predicted SRTs were highest for the complex speech material of the Goettingen sentence test, with differences between the two extreme assumptions below 2 dB. The empirical outcome was found to lie in-between. The main effect of complexity is suspected to be due to the number of required concurrent hypothesis during recognition.
An objective evaluation of binaural noise reduction algorithms allows for directly comparing the performance of different algorithm realizations. Here, a binaural speech intelligibility model (BSIM), which mimics the effective binaural processing of a human listeners, is used to predict the performance of the binaural minimum-variance distortionless response beamformer with partial noise estimation (BMVDR-N), which aims at preserving the speech component in a reference microphone and a scaled version of the noise component. The BMVDR-N beamformer is evaluated with respect to a predicted change in SRT depending on the parameter eta, which controls a trade-off between noise reduction and binaural cue preservation of the noise component. The results show that BSIM benefits from the preserved binaural cues suggesting that the BMVDR-N beamformer can improve the spatial quality of a scene without affecting speech intelligibility.
In order to combat the degraded quality of sounds played in reverberant rooms, the methods of room impulse response equalization are used. New approaches utilize the properties of the human auditory system, such as temporal masking, for a better control of late echoes. There, a prefilter is used to modify the played signal, which renders the echoes inaudible for a given position. In case of spatial mismatch, the procedure fails and may even add additional distortions. In order to mitigate these effects, we propose to equalize subbands of the room impulse response independently. With heavy equalization of the low frequencies and only a little or none in the high frequencies, the mismatch and therefore the degradation can be reduced. The result is a bigger equalized volume where a potential human listener can move freely.
We present a spatio-temporally situated dialog that is implemented in a driver assistance system. The system supports the driver in turning left at a busy urban intersection by providing verbal information on the vehicles arriving from the right. This is a highly dynamic scenario, as the location of the vehicles significantly changes while the system is referring to them. Consequently, the time the system needs for producing the utterances and the time drivers need to comprehend them and prepare their action has to be taken into account in the dialog management. We introduce a dialog concept, which predicts the future traffic participants state and plans the utterances such that they align with the driver's expectations. We have implemented and evaluated this system in a prototype vehicle.
This paper presents the concept of fuzzy-membership value (FMV) aware delay-and-sum beamforming for source separation in reverberant environments using ad hoc distributed microphones. Our approach employs a previously proposed fuzzy clustering algorithm to assign microphones of ad hoc arrays to individual source-dominated clusters and to compute fuzzy-membership values for each microphone and cluster. For each source-dominated cluster we first estimate relative time-differences-of-arrival (TDOA) information from the observed microphone signals and then apply both the TDOA and the FMV information in the beamforming stage. We show that such weighted beamforming improves upon the unweighted case. In a second enhancement stage we then apply cluster-related spectral masks to the output of the beamformers. We validate the proposed approach in three realistically-simulated rooms of different sizes. The method is evaluated by informal listening tests as well as by instrumental quality and intelligibility measures.
The articulatory code is composed of three sub-codes describing the phones uttered during a cortical syllable. Each sub-code encodes clusters of 'constriction parameters' describing the constrictions of the vocal tract similarly to manner & place features. Each set of sub-codes describes a specific set of gestures defined as Opening, Vowel, and Closure gestures respectively, defining the set of OVC-gestures. They represent the smallest units used in speech production and speech perception. The timing of the OVC-gestures is steered by the articulatory rhythm composed of entrained ϴ - and nested ɤ-oscillations. In speech production, the articulatory rhythm determines the starting points of the OVC-gestures; in speech perception, the articulatory rhythm segments the auditory signal into OVC-gestures bottom up. To train models for speech production and speech perception, a reference speech corpus labeled into OVC-gestures with boundaries defined by the articulator rhythm is needed. To the author's knowledge, such a speech corpus does not exist. The paper presents an approach to produce such a corpus by enhancing an articulatory speech database. The main idea is to use the quasi-rhythmic opening and closing gesture of the jaw to retrieve the ϴ-oscillation together with the OVC-gestures.