This paper describes algorithms developed to apply the analysis-by-synthesis/overlap-add (ABS/OLA) sinusoidal modeling system to real-time speech and singing voice synthesis. As originally proposed, the ABS/OLA system is limited to unidirectional time-scaling, and relies on variable frame length to accomplish time-scale modification. For speech and voice synthesis applications, unidirectional time scaling makes effective looping to produce sustained vocal sounds difficult, and variable frame length makes real-time polyphonic synthesis problematic. This paper presents a reformulation of the basic ABS/OLA system to deal with these issues, which is termed fixed-rate ABS/OLA (ABS/OLA-FR)
Although sinusoidal models have been demonstrated to be capable of high-quality musical instrument synthesis, speech modification, and speech synthesis, little exploration of the application of these models to the synthesis of singing voice has been undertaken. We propose a system framework similar to that employed in concatenation-based text-to-speech synthesizers, and describe its extension to the synthesis of singing voice. The power and flexibility of the sinusoidal model used in the waveform synthesis portion of the system enables high-quality, computationally-efficient synthesis and the incorporation of musical qualities such as vibrato and spectral tilt variation. Modeling of segmental phonetic characteristics is achieved by employing a "unit selection" procedure that selects sinusoidally-modeled segments from an inventory of singing voice data collected from a human vocalist. The system, called LYRICOS, is capable of synthesizing very natural-sounding singing that maintains the characteristics and perceived identity of the analyzed vocalist.
This paper describes algorithms developed to implement the Analysis-by-Synthesis/Overlap-Add (ABS/OLA) sinusoidal modeling system in real-time polyphonic music synthesizer. As originally proposed, the ABS/OLA system is limited to unidirectional time-scaling, and relies on variable frame length to accomplish timescale modification. For MIDI-driven music synthesis applications, unidirectional time scaling makes effective looping to produce sustained notes difficult, and variable frame length makes real-time polyphonic synthesis problematic. This paper presents a reformulation of the ABS/OLA system to deal with these issues, which is termed Fixed-Rate ABS/OLA (ABS/OLA-FR).
Sinusoidal modeling has been successfully applied to a broad range of speech processing problems, and offers advantages over linear predictive modeling and the short-time Fourier transform for speech analysis/synthesis and modification. This paper presents a novel speech analysis/synthesis system based on the combination of an overlap-add sinusoidal model with an analysis-by-synthesis technique to determine the model parameters. It describes this analysis procedure in detail, and introduces an equivalent frequency-domain algorithm that takes advantage of the computational efficiency of the fast Fourier transform (FFT). In addition, a refined overlap-add sinusoidal model capable of shape-invariant speech modification is derived, and a pitch-scale modification algorithm is defined that preserves speech bandwidth and eliminates noise migration effects. Analysis-by-synthesis achieves very high synthetic speech quality by accurately estimating the component frequencies, eliminating sidelobe interference effects, and effectively dealing with nonstationary speech events. The refined overlap-add synthesis model correlates well with analysis-by-synthesis, and modifies speech without objectionable artifacts by explicitly controlling shape invariance and phase coherence. The proposed analysis-by-synthesis/overlap-add (ABS/OLA) system allows for both fixed and time-varying time-, frequency-, and pitch-scale modifications, and computational shortcuts using the FFT algorithm make its implementation feasible using currently available hardware.
In this paper, we propose a system for synthesizing the human singing voice and the musical subtleties that accompany it. The system, Lyricos, employs a concatenation-based text-to-speech method to synthesize arbitrary lyrics in a given language. Using information contained in a regular MIDI le, the system chooses units, represented as sinusoidal waveform model parameters, from an inventory of data collected from a professional singer, and concatenates these to form arbitrary lyrical phrases. Standard MIDI messages control parameters for the addition of vibrato, spectral tilt, and dynamic musical expression, resulting in a very natural-sounding singing voice. Presented at 103rd Meeting of the AES, Sept. 1997 { AES Preprint 4591 2
This paper presents a system for separating the cochannel speech of two talkers. The proposed harmonic enhancement and suppression (HES) system is based on a frame-by-frame speaker separation algorithm that exploits the pitch estimate of the stronger talker derived from the cochannel signal, The idea behind this: approach is to recover the stronger talker's speech by enhancing their harmonic frequencies and formants given a multiresolution pitch estimate, The weaker talker's speech is obtained from the residual signal created when the harmonics and formants of the stronger talker are suppressed, An automatic speaker assignment algorithm is used to place recovered frames from the target and interfering talkers in separate channels, Automatic speaker assignment performs reasonably well in most cochannel environments, including voiced-on-voiced, voiced-on-unvoiced, unvoiced-on-unvoiced, assignment after processing silence intervals, and single talker speech (no cochannel interference), The HES system has been tested at target-to-interferer ratios (TIR's) from -18 to 18 dB with widely available data bases, It has demonstrated improved performance in keyword spotting tests for TIR values of 6, 12, and 18 dB, and in human listening tests for TIR values of -6 and -18 dB.
This paper describes a real time MIDI music synthesis system using a low cost digital signal processor such as the TI TMS320C32. The system consists of a MIDI device, an IBM compatible personal computer, and a TMS320C32 development board where the core of the music synthesis engine resides. The MIDI device generates MIDI data which is decoded into music synthesis commands, the host computer handles communications between the MIDI device and the DSP where music samples are reproduced using sinusoidal modeling based music synthesis techniques. A detailed description of the interface between MIDI device, host computer and DSP is provided.
This paper describes our enhanced mixed excitation linear prediction (MELP) speech coder which is a candidate for the new U.S. Federal Standard at 2.4 kbits/s. The new coder is based on the MELP model, and it uses a number of enhancements as well as efficient quantization algorithms to improve performance while maintaining a low bit rate. In addition, the coder has been optimized for performance in acoustic background noise and in channel errors, as well as for efficient real-time implementation. Listening tests confirm that the enhanced 2.4 kbit/s MELP coder performs as well as the higher bit rate 4.8 kbit/s FS1016 CELP standard.
This paper presents a block-coded VFR strategy that removes the familiar "anchor frame" constraint in favor of a generalized delayed decision analysis. This approach is referred to as adaptive frame selection using dynamic programming (AFS/DP). The AFS/DP algorithm determines optimal breakpoint locations using a perceptually-based performance measure, accommodates fixed bit-rate operation, and operates with fixed delay and low complexity. The AFS/DP algorithm has been implemented in the context of the MELP vocoder at 2400 bits/sec, and has demonstrated comparable performance to block VFR methods, with lower coder delay.
This ]paper describes our enhanced Mixed Exatation Linear Pirediction (MELP) coder which is a candidate for the new U. S. Federal Standard at 2.4 kbits/s. The new coder is based on the MELP model, but it utilizes a number of enhancements as well as more efficient quantization algorithm!$ to improve performance while maintaining a low bit rate.
This paper presents an algorithm for single-sensor enhancment of speech corrupted by additive random noise, based on soft-decision and statistical signal processing concepts and incorporating fully automatic noise estimation/tracking algorithms. The soft-decision/variable attenuation (SDVA) algorithm uses a compressive noise reduction model within the framework of short-time Fourier processing. The SDVA algorithm is fast, effective and robust, and has been applied in realistic RF and telephone environments
This paper describes a system for the automatic separation of two-talker co-channel speech. This system is based on a frame-by-frame speaker separation algorithm that exploits a pitch estimate of the stronger talker derived from the co-channel signal. The concept underlying this approach is to recover the stronger talker's speech by enhancing harmonic frequencies and formants given a multi-resolution pitch estimate. The weaker talker's speech is obtained from the residual signal created when the harmonics and formants of the stronger talker are suppressed. A maximum likelihood speaker assignment algorithm is used to place the recovered frames from the target and interfering talkers in separate channels. The system has been tested at target-to-interferer ratios (TIRs) from -18 to 18 dB with human listening tests, and with machine-based tests employing a keyword spotting system on the Switchboard Corpus for target talkers at 6, 12, and 18 dB TIR
This paper reports on techniques used in the generation of a continuous speech, multi-speaker, cellular bandwidth database and describes its application to automatic speech recognition in the cellular environment. CTIMIT (cellular TIMIT) has been generated by transmitting the TIMIT speech database over the cellular network. The CTIMIT database can have widespread applicability in the design and development of speech processing and speech recognition products for the cellular market. It describes the preliminary collection of the CTIMIT database and reports on several studies designed to test the utility of the database in a phoneme recognition task. Two HMM-based phoneme recognizers were trained using utterances drawn from the TIMIT database and the CTIMIT database, respectively. Each recognizer was then tested using the test utterances from CTIMIT. Phoneme recognition accuracy for the TIMIT-trained recognizer dropped 58% from its baseline performance on TIMIT test utterances. By comparison, phoneme recognition accuracy of the CTIMIT-trained recognizer increased 82% compared to that of the TIMIT-trained recognizer.
Analysis-by-synthesis/overlap-add (ABS/OLA) sinusoidal modeling has been successfully demonstrated as an accurate, flexible, and computationally tractable representation for the purposes of speech modification and harmonic tone synthesis; however, the model formulation used to synthesize these signals does not take full advantage of the structure of quasi-harmonic music signals. This paper describes a generalized overlap-add sinusoidal model formulation that accounts for the time-frequency behavior of quasi-harmonic tones and which reduces to the previous formulation as a special case
An approach to coding the parameters of a harmonic sinusoidal model which incorporates error spectrum shaping in order to improve the subjective quality of the speech coded at low bit rates is presented. It is shown that the sinusoidal model formulation is very well suited to representing speech signals and provides knowledge of pertinent time-varying characteristics of speech, such as pitch and short-time spectral information, which is useful for speech coding. The model also has a particularly simple form which lends itself easily to analysis and includes an envelope signal which separately models syllabic volume changes, enhancing the performance of the model. The analysis procedure presented provides greater accuracy than other techniques, resulting in higher quality synthetic and coded speech. In addition, as more components are added, the synthetic speech signal is guaranteed to converge to the original speech signal. Using perceptual factors in coding the parameters of this model yields considerable improvement in the overall subjective performance of the coder