Proportionate normalized least mean squares (PNLMS) is an adaptive filter that has been shown to provide exceptionally fast convergence and tracking when the underlying system parameters are sparse. A good example of such a system is a network echo canceller. Principal components based PNLMS (PCP) extends this fast convergence property to certain nonsparse systems by applying PNLMS while using the principal components of the underlying system as basis vectors. An acoustic echo canceller is a possible example of this type of nonsparse system. Simulations of acoustic echo paths and cancellers indicate that PCP converges and tracks much faster than the classical normalized least mean squares (NLMS) and fast recursive least squares (FRLS) adaptive filters. However, when a basic parameter, like room temperature, changes, the underlying acoustic structure of the room changes as well and principal components of the room responses at one temperature are very different from those at another. This paper addresses this problem by using multiple sets of principle components as basis vectors and performing PNLMS in each basis set. Each set of principle components are derived from the room at a different temperature. The new algorithm, multiple principal components PNLMS (MPCP) is a generalization of PNLMS++. Simulations show the potential effectiveness of the approach.
Soft-phone technology for Internet protocol (IP) voice is growing in importance. However, soft phones exhibit poorer quality than public switched telephone network (PSTN) phones. A goal is to improve that quality, perhaps even to the point that the communication experience is better than with PSTN phones. This letter presents an analysis of soft-phone performance and describes acoustic echo cancellation and other technologies that improve soft-phone performance.
One may ask a legitimate question: why do we need multichannel sound for telecommunication? Let’s take the following example. When we are in a room with several people talking, laughing, or just communicating with each other, thanks to our binaural auditory system, we can concentrate on one particular talker (if several persons are talking at the same time), localize or identify a person who is talking, and somehow we are able to process a noisy or a reverberant speech signal in order to make it intelligible. On the other hand, with only one ear or, equivalently, if we record what happens in the room with one microphone and listen to this monophonic signal, it will likely make all of the above mentioned tasks more difficult. So, multichannel sound teleconferencing systems provide a realistic presence that mono-channel systems cannot offer.
While linear prediction has been successfully applied to many topics in signal processing, linear interpolation has received little attention. This chapter gives some results on linear interpolation and shows that many well-known variables or equations can be formulated in terms of linear interpolation. Also, the so-called principle of orthogonality is generalized. From this theory, we then give a generalized least mean square algorithm and a generalized affine projection algorithm.
Adaptive filters [60] play an important role in echo cancellation because we need to identify and track unknown and time-varying channels [24J . There are roughly two classes of adaptive algorithms. One class includes filters that are updated in the time domain, sample-by-sample in general, like the classical least mean square (LMS) [134] and recursive least-squares (RLS) [4], [66] algorithms. The other class contains filters that are updated in the frequency domain, block-by-block in general, using the fast Fourier transform (FFT) as an intermediary step. As a result of this block processing, the arithmetic complexity of the algorithms in the latter category is significantly reduced compared to time-domain adaptive algorithms. Use of the FFT is appropriate to the Toeplitz structure, which results from the time-shift properties of the filter input signal. Consequently, deriving a frequency-domain (FD) adaptive algorithm is just a matter of rewriting the time-domain error criterion in a way that Toeplitz and circulant matrices are explicitly shown.
Ideally, acoustic echo cancelers (AECs) remove undesired echoes that result from acoustic coupling between the loudspeaker and the microphone used in full-duplex hands-free telecommunication systems. Figure 6.1 shows a diagram of a single-channel AEC. The far-end speech signal x(n) goes through the echo path represented by a filter h(n) to produce the echo, y e(n), which is picked up by the microphone together with the near-end talker signal v(n) and ambient noise w(n). The composite microphone signal is denoted y(n).
When people with normal hearing converse in a room where many people are speaking simultaneously, their binaural hearing enables them to focus in on particular talkers according to the directions from which those talkers’ voices are coming. They can do this even when the signal-to-background-noise ratio is very low; noise, in this case being the voices of those ignored. This phenomenon of human audio perception is aptly called the cocktail party effect. In monophonic teleconferencing systems, this aid to audio communication is lost across the connection. A listener on one side hears all of the talkers of the far side coming from the same direction — the direction from a single local loudspeaker. So, when people on the far side talk simultaneously, it is impossible for the local listener to spatially separate their voices as he or she normally would. A stereo connection would solve this problem because with stereo the local listener (when located in the “sweet spot” — that area where the stereo effect is most clearly perceived) hears a reconstruction of the leftright positioning of the sound from the far side. Until very recently though, teleconferencing systems have been limited to monophonic connections because stereo acoustic echo cancellation was problematic [121]. However, with the advent of new techniques (see Chap. 5) such systems are now quite realizable.
A multichannel frequency-domain adaptive algorithm was presented in Chap. 8 (see also [15]). The multichannel frequency-domain algorithm has been shown to work very well in the two-channel acoustic echo cancellation application [40]. It has a fairly low computational complexity compared to the fast recursive least-squares algorithm (FRLS) [40]. Furthermore, it is an inherently stab le algorithm, well suite d for a fixed-point implementation. Our objective in this chapter is to provide a complete solution, based on the multichannel frequency-domain adaptive algorithm, which handles both echo cancellation and double-talk.
In this chapter, we present a family of fast-converging algorithms that are extensions of the proportionate normalized least mean square (PNLMS) algorithm introduced by Duttweiler [38], [54] This new family of algorithms is based on the affine projection algorithm/normalized least mean square (APA/NLMS) algorithm family [106], [100], [55], [125], [66]. What differentiates the new algorithms from the NLMS and APA algorithms is that they inherit the proportional step-size idea from the PNLMS algorithm, i.e., individually assigned step-sizes to each filter coefficient, where the step sizes are calculated from the previous estimate of the echo path. Because of these individual step sizes, the algorithms achieve a higher convergence rate by using the fact that the active part of a network echo path is usually much smaller (4–8 ms) than the possible echo path range (64–128 ms) that has to be covered. A natural extension of the basic PNLMS algorithm is a proportionate affine projection algorithm (PAPA) [51]. This algorithm (family) combines the fast converging APA with the proportional step size technique of PNLMS.
With rare exceptions, conversations take place in the presence of echoes. We hear echoes of our speech waves as they are reflected from the floor, walls, and other neighboring objects. If a reflected wave arrives a very short time after the direct sound, it is perceived not as an echo but as a spectral distortion, or reverberation. Most people prefer some amount of reverberation to a completely anechoic environment, and the desirable amount of reverberation depends on the application. (For example, much more reverberation is desirable in a concert hall than in an office.) The situation is very different, however, when the leading edge of the reflected wave arrives a few tens of milliseconds after the direct sound. In such a case, it is heard as a distinct echo. Such echoes are invariably annoying, and under extreme conditions can completely disrupt a conversation. It is such distinct echoes that we will be concerned with in this book.
In this paper a wideband stereophonic acoustic echo canceler is presented. The fundamental difficulty of stereophonic acoustic echo cancellation (SAEC) is described and an echo canceler based on a fast recursive least squares algorithm in a subband structure is proposed. This structure have been used in a real-time implementation, on which experiments have been performed. In the paper, simulation results of this implementation on real life recordings, with 8 kHz bandwidth, are studied. The results clearly verify that the theoretic fundamental problem of SAEC also applies in real-life situations. They also show that more sophisticated adaptive algorithms are needed in the lower frequency regions than in the higher regions.
Teleconferencing systems employ acoustic echo cancelers to reduce echoes that results from the coupling between loudspeaker and microphone. To enhance the sound realism, two-channel audio is necessary. However, stereophonic acoustic echo cancellation is more difficult to solve because of the necessity to uniquely identify two acoustic paths, which becomes problematic since the two excitation signals are highly correlated. In this paper a wideband stereophonic acoustic echo canceler is presented. The fundamental difficulty of stereophonic acoustic echo cancellation (SAEC) is described and an echo canceler based on a fast recursive least squares algorithm in a subband structure is proposed. The structure has been used in a real-time implementation, with which experiments have been performed. In the paper, simulation results of this implementation on real life recordings, with 8 kHz bandwidth, are studied. The results clearly verify that the theoretic fundamental problem of SAEC also applies in real-life situations. They also show that more sophisticated adaptive algorithms are needed in the lower frequency regions than in the higher regions.
Mark H. Jones合作论文数Department of Physics and Astronomy;Open University;Department of Physics and Astronomy, Open University1