Recordings in everyday life provide valuable insights for health-related applications, such as analyzing conversational behavior as an indicator of social interaction and well-being. However, these recordings require privacy preservation of both the speech content and the speaker’s identity of all persons involved. This article investigates privacy-preserving features feasible for power-constrained recording devices by combining smoothing and subsampling in the frequency and time domain with a low-cost speaker anonymization technique. A speech recognition and a speaker verification system are used to evaluate privacy protection, whereas a voice activity detection and a speaker diarization model are used to assess the utility for analyzing conversations. The evaluation results demonstrate that combining speaker anonymization with the aforementioned smoothing and subsampling protects speech privacy, albeit at the expense of utility performance. Overall, our privacy-preserving methods offer various trade-offs between privacy and utility, reflecting the requirements of different application scenarios.
Protecting speech privacy in real-life audio recordings is a growing concern. This contribution evaluates the effectiveness of three obfuscation techniques in protecting linguistic speech content, using digit recognition as a task-specific and practically motivated evaluation scenario. As a first baseline, a general-purpose speech recognition model and a digit-specific classifier were applied as informed attackers to recognise both single digits and concatenated digit sequences. Our experimental results demonstrate significant differences in recognition performance across digit modality, speech rate, and attack model. These findings emphasize the need for more comprehensive and application-oriented evaluation methods to ensure speech privacy.
Voice and speech disorders can be assessed through subjective perceptual measures or objective indices, such as the Acoustic Voice Quality Index (AVQI) or the Acoustic Breathiness Index (ABI). These objective measures reduce diagnostic variability by eliminating subjective evaluations. However, the impact of room acoustics on their robustness for therapy monitoring remains unclear. In this contribution, reverberation times, impulse responses, and background noise were recorded and analyzed in conjunction with speech samples using statistical models to evaluate their influence. Results indicate that room acoustics significantly affect AVQI and ABI, particularly for non-pathological voices and with distant microphones. For reliable therapeutic use, standardized measurement environments or robust analysis methods are essential. Optimized recording conditions with source-proximate microphones enhance accuracy, advancing objective voice quality assessments and evidence-based speech therapy.
Reverberation can severely degrade the quality of speech signals recorded using microphones in an enclosure. In acoustic sensor networks with spatially distributed microphones, a similar dereverberation performance may be achieved using only a subset of all available microphones. Using the popular convex relaxation method, in this paper we propose to perform microphone subset selection for the weighted prediction error (WPE) multi-channel dereverberation algorithm by introducing a group sparsity penalty on the prediction filter coefficients. The resulting problem is shown to be solved efficiently using the accelerated proximal gradient algorithm. Experimental evaluation using measured impulse responses shows that the performance of the proposed method is close to the optimal performance obtained by exhaustive search, both for frequency-dependent as well as frequency-independent microphone subset selection. Furthermore, the performance using only a few microphones for frequency-independent microphone subset selection is only marginally worse than using all available microphones.
Recordings in everyday life require privacy preservation of the speech content and speaker identity. This contribution explores the influence of noise and reverberation on the trade-off between privacy and utility for low-cost privacy-preserving methods feasible for edge computing. These methods compromise spectral and temporal smoothing, speaker anonymization using the McAdams coefficient, sampling with a very low sampling rate, and combinations. Privacy is assessed by automatic speech and speaker recognition, while our utility considers voice activity detection and speaker diarization. Overall, our evaluation shows that additional noise degrades the performance of all models more than reverberation. This degradation corresponds to enhanced speech privacy, while utility is less deteriorated for some methods.
Reverberation may severely degrade the quality of speech signals recorded using microphones in a room. For compact microphone arrays, the choice of the reference microphone for multi-microphone dereverberation typically does not have a large influence on the dereverberation performance. In contrast, when the microphones are spatially distributed, the choice of the reference microphone may significantly contribute to the dereverberation performance. In this paper, we propose to perform reference microphone selection for the weighted prediction error (WPE) dereverberation algorithm based on the normalized ℓ p -norm of the dereverberated output signal. Experimental results for different source positions in a reverberant laboratory show that the proposed method yields a better dereverberation performance than reference microphone selection based on the early-to-late reverberation ratio or signal power.
In recent years, the need for privacy preservation when manipulating or storing personal data, including speech , has become a major issue. In this paper, we present a system addressing the speaker-level anonymization problem. We propose and evaluate a two-stage anonymization pipeline exploiting a state-of-the-art anonymization model described in the Voice Privacy Challenge 2022 in combination with a zero-shot voice conversion architecture able to capture speaker characteristics from a few seconds of speech. We show this architecture can lead to strong privacy preservation while preserving pitch information. Finally, we propose a new compressed metric to evaluate anonymization systems in privacy scenarios with different constraints on privacy and utility.
Unlike model-based direction of arrival (DoA) estimation algorithms, supervised learning-based DoA estimation algorithms based on deep neural networks (DNNs) are usually trained for one specific microphone array geometry, resulting in poor performance when applied to a different array geometry. In this paper we illustrate the fundamental difference between supervised learning-based and model-based algorithms leading to this sensitivity. Aiming at designing a supervised learning-based DoA estimation algorithm that generalizes well to different array geometries, in this paper we propose a geometry-aware DoA estimation algorithm. The algorithm uses a fully connected DNN and takes mixed data as input features, namely the time lags maximizing the generalized cross-correlation with phase transform and the microphone coordinates, which are assumed to be known. Experimental results for a reverberant scenario demonstrate the flexibility of the proposed algorithm towards different array geometries and show that the proposed algorithm outperforms model-based algorithms such as steered response power with phase transform.
The analysis of conversations recorded in everyday life requires privacy protection. In this contribution, we explore a privacy-preserving feature extraction method based on input feature dimension reduction, spectral smoothing and the low-cost speaker anonymization technique based on McAdams coefficient. We assess the utility of the feature extraction methods with a voice activity detection and a speaker diarization system, while privacy protection is determined with a speech recognition and a speaker verification model. We show that the combination of McAdams coefficient and spectral smoothing maintains the utility while improving privacy.
In the last decades several multi-microphone speech dereverberation algorithms have been proposed, among which the weighted prediction error (WPE) algorithm. In the WPE algorithm, a prediction delay is required to reduce the correlation between the prediction signals and the direct component in the reference microphone signal. In compact arrays with closely-spaced microphones, the prediction delay is often chosen microphone-independent. In acoustic sensor networks with spatially distributed microphones, large time-differences-of-arrival (TDOAs) of the speech source between the reference microphone and other microphones typically occur. Hence, when using a microphone-independent prediction delay 1 the reference and prediction signals may still be significantly correlated, leading to distortion in the dereverberated output signal. In order to decorrelate the signals, in this paper we propose to apply TDOA compensation with respect to the reference microphone, resulting in microphone-dependent prediction delays for the WPE algorithm. We consider both optimal TDOA compensation using crossband filtering in the short-time Fourier transform domain as well as band-to-band and integer delay approximations. Simulation results for different reverberation times using oracle as well as estimated TDOAs clearly show the benefit of using microphone-dependent prediction delays.
Virtual and augmented realities are increasingly popular tools in many domains such as architecture, production, training and education, (psycho)therapy, gaming, and others. For a convincing rendering of sound in virtual and augmented environments, audio signals must be convolved in real-time with impulse responses that change from one moment in time to another. Key requirements for the implementation of such time-variant real-time convolution algorithms are short latencies, moderate computational cost and memory footprint, and no perceptible switching artifacts. In this engineering report, we introduce a partitioned convolution algorithm that is able to quickly switch between impulse responses without introducing perceptible artifacts, while maintaining a constant computational load and low memory usage. Implementations in several popular programming languages are freely available via GitHub.
Recently, exploring acoustic conditions of people in their everyday environments has drawn a lot of attention. One of the most important and disturbing sound sources is the test participant’s own voice. This contribution proposes an algorithm to determine the own-voice audio segments (OVS) for blocks of 125 ms and a method for measuring sound pressure levels (SPL) without violating privacy laws. The own voice detection (OVD) algorithm here developed is based on a machine learning algorithm and a set of acoustic features that do not allow for speech reconstruction. A manually labeled real-world recording of one full day showed reliable and robust detection results. Moreover, the OVD algorithm was applied to 13 near-ear recordings of hearing-impaired participants in an ecological momentary assessment (EMA) study. The analysis shows that the grand mean percentage of predicted OVS during one day was approx. 10% which corresponds well to other published data. These OVS had a small impact on the median SPL over all data. However, for short analysis intervals, significant differences up to 30 dB occurred in the measured SPL, depending on the proportion of OVS and the SPL of the background noise.
Aiming at estimating the direction of arrival (DOA) of a desired speaker in a multi-talker environment using a microphone array, in this paper we propose a signal-informed method exploiting the availability of an external microphone attached to the desired speaker. The proposed method applies a binary mask to the GCC-PHAT input features of a convolutional neural network, where the binary mask is computed based on the power distribution of the external microphone signal. Experimental results for a reverberant scenario with up to four interfering speakers demonstrate that the signal-informed masking improves the localization accuracy, without requiring any knowledge about the interfering speakers.
Ecological momentary assessment (EMA) was used in 24 adults with mild-to-moderate hearing loss who were seeking first hearing-aid (HA) fitting or HA renewal. At two stages in the aural rehabilitation process, just before HA fitting and after an average 3-month HA adjustment period, the participants used a smartphone-based EMA system for 3 to 4 days. A questionnaire app allowed for the description of the environmental context as well as assessments of various hearing-related dimensions and of well-being. In total, 2,042 surveys were collected. The main objectives of the analysis were threefold: First, describing the "auditory reality" of future and experienced HA users; second, examining the effects of HA fitting for individual participants, as well as for the subgroup of first-time HA-users; and third, reviewing whether the EMA data collected in the unaided condition predicted who ultimately decided for or against permanent HA use. The participants reported hearing-related disabilities across the full range of daily listening tasks, but communication events took the largest share. The effect of the HA intervention was small in experienced HA users. Generally, much larger changes and larger interindividual differences were observed in first-time compared with experienced HA users in all hearing-related dimensions. Changes were not correlated with hearing loss or with the duration of the HA adjustment period. EMA data collected in the unaided condition did not predict the cancelation of HA fitting. The study showed that EMA is feasible in a general population of HA candidates for establishing individual and multidimensional profiles of real-life hearing experiences.
Commercially available light-weight unmanned aerial vehicles (UAVs) present a challenge for public safety, e.g. espionage, transporting dangerous goods or devices. Therefore, countermeasures are necessary. Usually, detection of UAVs is a first step. Along many other modalities, acoustic detection seems promising. Recent publications show interesting results by using machine and deep learning methods. The acoustic detection of UAVs appears to be particularly difficult in adverse situations, such as in heavy wind noise or in the presence of construction noise. In this contribution, the typical feature set is extended to increase separation of background noise and the UAV signature noise. The decision algorithm utilized is support vector machine (SVM) classification. The classification is based on an extended training dataset labeled to support binary classification. The proposed method is evaluated in comparison to previously published algorithms, on the basis of a dataset recorded from different acoustic environments, including unknown UAV types. The results show an improvement over existing methods, especially in terms of false-positive detection rate. For a first step into real-time embedded systems a recursive feature elimination method is applied to reduce the model dimensionality. The results indicate only a slight decreases in detection performance.
Common methods to assess hearing deficits and the benefit of hearing devices include retrospective questionnaires and speech tests under controlled conditions. As typically applied, both approaches suffer from serious limitations regarding their ecological validity. An alternative approach rapidly gaining widespread use is ecological momentary assessment (EMA), which employs repeated assessments of individual everyday situations. Smartphones facilitate the implementation of questionnaires and rating schemes to be administered in the real life of study participants or customers, during or shortly after an experience. In addition, objective acoustical parameters extracted from head- or body-worn microphones and/or settings from the hearing aid's signal processing unit can be stored alongside the questionnaire data. The advantages of using EMA include participant-specific, context-sensitive information on activities, experienced challenges, and preferences. However, to preserve the privacy of all communication partners and bystanders, the law in many countries does not allow audio recordings, limiting the information about environmental acoustics to statistical data such as, for example, levels and averaged spectra. Other challenges for EMA are, for example, the unsupervised handling of the equipment, the trade-off between the accuracy of description and the number of similar listening situations when performing comparisons (e.g., with and without hearing aids), the trade-off between the duration of recording intervals and the amount of data collected and analyzed, the random or target-oriented reminder for subjective responses, as well as the willingness and ability of the participants to respond while doing specific tasks. This contribution reviews EMA in hearing research, its purpose, current applications, and possible future directions.
Near-end listening enhancement (NELE) algorithms aim to preprocess speech prior to playback via loudspeakers so as to maintain high speech intelligibility even when listening conditions are not optimal, e.g., due to noise or reverberation. Often NELE algorithms are designed for scenarios considering either only the detrimental effect of noise or only reverberation, but not both disturbances. In many typical applications scenarios, however, both factors are present. In this paper, we evaluate a new combination of a noise-dependent and a reverberationdependent algorithm implemented in a common framework. Specifically, we use instrumental measures as well as subjective ratings of listening effort for acoustic scenarios with different reverberation times and realistic signal-to-noise ratios. The results show that the noise-dependent algorithm also performs well in reverberation, and that the combination of both algorithms can yield slightly better performance than the individual algorithms alone. This benefit appears to depend strongly on the specific acoustic condition, indicating that further work is required to optimize the adaptive algorithm behavior.
In social interaction it is vitally important to be able to perceive the gaze direction or the head orientation of the people talking around us. As long as our dialogue partner is within our field of vision, visual cues are normally responsible for the head orientation estimation, here termed facing angle. The facing angle describes the angle of orientation for a directional source relative to the receiver position. The aim of this study was to measure the ability of participants to estimate the facing angle of a directional sound source based on acoustic cues for facing angles in 25 degrees steps at three source positions (frontal, lateral, rear left) in an anechoic and reverberant environment. The listeners performed generally poorly for a lateral or rear left sound source compared to a frontal sound source. They showed the best performance when the sound source pointed directly to the listener for all tested source positions. However, listeners perceived small facing angles as direct-facing. The facing angle estimation was more difficult in the anechoic room for facing angles where the loudspeaker was turned away from the listener. For a lateral sound source, the listeners were partly unable to distinguish between positive and negative facing angles.
This paper presents a novel algorithm to estimate the power spectral density (PSD) of stationary broadband noise disturbances in audio recordings. The proposed algorithm estimates the noise PSD as the mean value of an exponential distribution that corresponds to the truncated periodogram coefficients of the disturbed audio signal. An evaluation with a large number of speech and music test signals shows that a high PSD estimation accuracy can be obtained for a wide range of signal-to-noise ratios, allowing for unsupervised operation and thus constituting an important part of a fully automatic broadband noise restoration system for audio archives.