For automotive hands-free and speech recognition applications, distributed microphones are often mounted in the car where each of the speakers has a dedicated microphone close to his position. To provide additional control information for further speech enhancement, it is often advantageous to distinguish between the activity of the different passengers. In this contribution speaker activity is identified by the evaluation of the ratios between the powers of multiple speaker-dedicated microphone signals while further acoustic events are differentiated from single speaker's activity. An effective algorithm based on the exploitation of the expected range of power ratio values is presented that is able to detect local disturbances like scratch noise at the microphones. Furthermore, situations are identified where multiple passengers speak at the same time. Besides some examples, it can be shown that the proposed method for local disturbance detection is able to identify such noise within realistic driving conditions.
Supporting multiple active speakers in automotive hands-free or speech dialog applications is an interesting issue not least due to comfort reasons. Therefore, a multi-channel system for enhancement of speech signals captured by distributed distant microphones in a car environment is presented. Each of the potential speakers in the car has a dedicated directional microphone close to his position that captures the corresponding speech signal. The aim of the resulting overall system is twofold: On the one hand, a combination of an arbitrary pre-defined subset of speakers' signals can be performed, e.g., to create an output signal in a hands-free telephone conference call for a far-end communication partner. On the other hand, annoying cross-talk components from interfering sound sources occurring in multiple different mixed output signals are to be eliminated, motivated by the possibility of other hands-free applications being active in parallel. The system includes several signal processing stages. A dedicated signal processing block for interfering speaker cancellation attenuates the cross-talk components of undesired speech. Further signal enhancement comprises the reduction of residual cross-talk and background noise. Subsequently, a dynamic signal combination stage merges the processed single-microphone signals to obtain appropriate mixed signals at the system output that may be passed to applications such as telephony or a speech dialog system. Based on signal power ratios between the particular microphone signals, an appropriate speaker activity detection and therewith a robust control mechanism of the whole system is presented. The proposed system may be dynamically configured and has been evaluated for a car setup with four speakers sitting in the car cabin disturbed in various noise conditions.
In cars with integrated distributed microphone systems usually each speaker has a dedicated microphone. An often required broadband speaker activity detection can be performed by simply evaluating the power ratios among the microphones but transient interferers like indicator noise, outside crossing cars or speech from interfering speakers may be wrongly assigned to one speaker's activity. In this contribution a new method is presented that exploits patterns based on the characteristics of signal power inverted subbands at which compared to the closest microphone higher energy occurs in a distant one. By determination of a distance measure the currently observed pattern of these power inverted subbands is compared to online learned speaker position dependent reference patterns. During noise only periods as well as during transient interfering signals the patterns do not match the reference and false speaker activity detections can be reduced.
For instrumental quality assessment of speech enhancement systems it is common to process the signal components such as speech and noise independently using the filter coefficients obtained from the combined noisy signal. Thus, specific quality measures can be computed. A specified signal-to-noise ratio (SNR) can be set for the noisy signal by rescaling and adding the signal components. In this contribution a setup for a multi-channel instrumental quality assessment of distributed microphone systems is presented, where each of multiple active speakers has a dedicated microphone. For speech active periods of the related speaker the desired SNR is explicitly set in his dedicated microphone by rescaling the signal level, whereas the other channels are adjusted accordingly. The proposed setup allows to create realistic noisy input signals for a speech enhancement system and to preserve the acoustic characteristics of the multi-channel microphone arrangement. Based on this structure a noise reduction approach is evaluated. It performs a selective combination of the multi-channel signals and a frequency selective boosting of speech active bins of the related active speaker as an extension to a recursive Wiener filtering.
Distributed microphone systems in cars usually provide dedicated microphones for several speakers where each micro phone captures the desired speech signal at the best. The signal quality may differ strongly among the speaker channels depending on the microphone position, the microphone type, the kind of background noise, and the speaker himself. When combining these signals to a weighted mix annoying switching artifacts may result. In this contribution a new dynamic signal mixer is presented that uses spectral preprocessing to compensate both for different speech signal levels and for different background noise levels and colorations. Thus, artifacts are avoided and smooth transitions can be achieved for the various speech level and the background noise spectrum at speaker changes.
Hands-free systems in cars aim to capture the speech of different speakers at the best. Therefor distributed microphones can be aligned to each of these speakers and can be mounted in their vicinity. One application of this setup is to optimize speech for speech recognition. It is desirable to cancel the influence of possible interfering speakers in the microphone signals [1, 2]. For this task an adaptive filter is used to cancel the interfering signal from the target one. In contrast to the very similar and well known echo cancellation problem a noise component also occurs on the filter input signal. In this paper an optimal step size is proposed that considers this noise component for controlling the adaptation of a normalized least mean square (NLMS) algorithm [3] in the short-time frequency domain. A Signal-to-InterferenceRatio (SIR) based adaptation control similar to [1] is also investigated. The performance of both approaches is evaluated.