This paper presents a novel algorithm to estimate the power spectral density (PSD) of stationary broadband noise disturbances in audio recordings. The proposed algorithm estimates the noise PSD as the mean value of an exponential distribution that corresponds to the truncated periodogram coefficients of the disturbed audio signal. An evaluation with a large number of speech and music test signals shows that a high PSD estimation accuracy can be obtained for a wide range of signal-to-noise ratios, allowing for unsupervised operation and thus constituting an important part of a fully automatic broadband noise restoration system for audio archives.
This article presents a new algorithm to classify whether each one-second long frame of an audio recording contains impulsive disturbances or not. The developed classification algorithm is based on supervised learning and appropriate prewhitening of the input signal. It is shown that existing impulse restoration algorithms suffer from degradation of the desired signal if the input SNR is high and if no manual parameter adjustment is possible, which makes automatic restoration of large amounts of diverse archive audio material infeasible. The proposed classification algorithm can be used as a supplement to an existing impulse restoration algorithm to alleviate this drawback. An evaluation with a large number of test signals shows that a high classification accuracy can be achieved, making fully automatic impulse restoration possible.
In this paper, we present a new approach for noise reduction. A binary time-frequency (T-F) masking threshold criterion is proposed and analyzed with respect to the average spectra of music and noise disturbances. Modified autoregressive (AR) detection and AR interpolation are then applied to the residual signal of the binary masking process. The proposed method is able to reduce supergaussian and impulsive noise while ensuring preservation of the desired signal, which is crucial for professional high-quality audio restoration, and it is also suitable for Gaussian noise to a certain extent. The approach is compared to a state-of-the-art restoration algorithm by means of the objective measures signal-to-noise ratio (SNR) improvement and perceptual quality, and by subjective listening tests. The objective results as well as the listening tests show that the proposed algorithm is especially suited for supergaussian, grainy-sounding noise types, e.g., optical soundtrack noise of celluloid movie footage, or rain noise.
This article examines the automatic detection of low frequency additive sinusoidal disturbances in audio signals, usually termed hum. We present a method to automatically determine whether an audio signal contains hum or not, and, if necessary; to determine its parameters - e.g., the fundamental frequency and the number of harmonics. The developed algorithm does not require a priori information, and we show its good detection capabilities by an evaluation with artificial signals and real recordings.
A new approach for noise reduction is presented. The method is capable of reducing noise of Gaussian, supergaussian and impulsive characteristics in degraded high-quality audio signals. The approach is based on classical autoregressive (AR) detection and interpolation, applied to the residual signal of a binary time-frequency (T-F) masking process. Analytic inspection allows for predicting the noise reduction level for white noise types and shows good accordance to simulation results. High reduction levels are achieved especially for supergaussian and impulsive disturbances having a higher sample kurtosis than Gaussian noise. The approach ensures high preservation of the underlying desired signal, satisfying the needs of high quality audio restoration. Furthermore, the approach is capable of reducing optical soundtrack noise of celluloid movie footage.
This Convention paper was selected based on a submitted abstract and 750-word precis that have been peer reviewed by at least two qualified anonymous reviewers. The complete manuscript was not peer reviewed. This convention paper has been reproduced from the author’s advance manuscript without editing, corrections, or consideration by the Review Board. The AES takes no responsibility for the contents. Additional papers may be obtained by sending request and remittance to Audio Engineering Society, 60 East 42 Street, New York, New York 10165-2520, USA; also see www.aes.org. All rights reserved. Reproduction of this paper, or any portion thereof, is not permitted without direct permission from the Journal of the Audio Engineering Society.
In this article, an evaluation of a recently published hum detection algorithm for audio signals is presented. To determine the performance of the method, large amounts of artificially generated and real-world audio data, containing a variety of music and speech recordings, are processed by the algorithm. By comparing the detection results with manually determined ground truth data, several error measures are computed: hit and false alarm rates, frequency deviation of the hum frequency estimation, offset of detected start and stop times and the accuracy of the SNR estimation.
In this paper, we investigate different types of spectral smoothing of the transfer function of single-channel noise reduction algorithms in terms of achieved audio quality. In order to determine the audio quality extensive listening tests have been conducted. Furthermore, we computed several existing objective quality measures based on technical measures or psychoacoustics. We examine whether the different forms of spectral smoothing of the weighting rule of the noise reduction algorithm are represented by the objective measures. We show that most of the known measures are insensitive to changes in the short-time spectra that are subtle in a technical way, but immediately noticeable by human listeners. The results of the listening test also indicate the optimal smoothing method for speech and audio enhancement.