
This paper discusses the theoretical basis for representation of a speech signal by its short-time Fourier transform. The results of the theoretical studies were used to design a speech analysis-synthesis system which was simulated on a general-purpose laboratory digital computer system. The simulation uses the fast Fourier transform in the analysis stage and specially designed finite duration impulse response filters in the synthesis stage. The results of both the theoretical and computational studies lead to an understanding of the effect of several design parameters and elucidate the design tradeoffs necessary to achieve moderate information rate reductions.
High-speed algorithms to compute the discrete Hadamard and Walsh transforms of speech waveforms have been developed. Intelligible speech has been reconstructed from dominant Hadamard or Walsh coefficients on a medium sized computer in a non-real-time mode. Degradation of some phonemes was noted at low bit rates of reconstruction, but the reconstruction could be improved by varying the position of the sampling window. A digital processor, which allows real-time analysis of speech to be conducted on the system, is described.
A study has been made of linear applications of 4- layer p-n-p-n devices. Biased as transistor tetrodes, these elements are characterized as to dc and small-signal parameters. Tetrode audio automatic gain control AGC circuits are devised and their controlled gain characteristics are discussed.
A method for filter input-output data analysis is presented whereby its transfer function is computed under normally operating conditions. The method offers the advantages of simplicity of computation and accurate estimation in the present of noise.
The Padé approximant technique provides a quick design of recursive digital filters. An added advantage of the technique lies in that spectrum shaping requirements as well as linear phase constraints can be handled easily, even for higher order filters. This is important in supplying initial guesses of the filter parameters to iterative routines that would then seek a locally optimal design solution. These advantages are among those discussed in a partly tutorial presentation of the technique that relates to filter needs found in data transmission systems. In addition, the question of stability is treated and a new criterion is presented. The criterion provides sufficient conditions in establishing stability for a filter designed by using the Padé approximant technique.
A primary-recognition computer program has been written to provide segmentation, acoustical parameters, and phonetic features of continuous speech, together with classification of some vowel and consonantal segments. Based on fundamental frequency, level, and duration information provided by the primary recognition program from short-term spectra, a procedure to mark stressed and reduced vowels is proposed. Listener judgments of stress and vowel reduction can be correlated with the physical parameters, but talker differences are apparent. It is clear that feature extraction at the segment level and at the suprasegmental level are mutually interactive.
An integer matrix method for implementation of the bilinear transformation is discussed and extended to a more general case that is useful in the design of digital filters.
An equiripple error constraint was used to design linear approximations to . Several examples are presented including one having a peak error of less than one percent, which is significantly less than the error obtained by using other design criteria. The mean and standard deviation of the relative error are also tabulated as are earlier results obtained by other authors.
After more than two decades of research it is now possible to construct a high-performance reading system for the blind that will produce synthetic speech from printed text. The entire process can be carried out automatically by computer and associated special-purpose devices. As a first step toward the eventual deployment of a reading system, we have begun an evaluation study in collaboration with faculty and students at the University of Connecticut and with trainees at the Veterans Administration Eastern Blindness Rehabilitation Center. Questions to be answered concern the comprehensibility and educational uses of the output and the technical and economic resources required to make automated reading services accessible to progressively larger groups of blind people.
The ability of listeners to perform some speaker verification tasks has been measured experimentally and compared with the performance of an automatic system for speaker verification. A test presentation in the subjective experiments consists of a pair of utterances. One of these is drawn from the recordings of a group of speakers designated customers while the second utterance is either a distinct recording from the same customer or the recording of an impostor. Listeners must respond whether the utterances are from the same or different speakers. The impostor classes that have been considered are casual impostors making no attempt to mimic customers, trained professional mimics, and an identical twin of a customer. Listener performance is specified by the two types of error that can be committed.
The radiation characteristics of a planar array of concentric rings are examined. By employing optimization techniques, control of the beamwidth and sidelobe level is accomplished. Treating the energy in the sidelobes as a criterion for optimization, side-lobes of equal amplitude are obtained.
It is proved that the z transform of a sequence of data values cannot be exactly computed using a binary number representation for values of z on the unit circle, except z =±1, z = ±j. It is also proved that the discrete Fourier transform (DFT) of a sequence of data values cannot be evaluated with rational numbers.
An equiripple error constraint was used to design linear approximations to \sqrt{x^{2} + y^{2}} . Several examples are presented including one having a peak error of less than one percent, which is significantly less than the error obtained by using other design criteria. The mean and standard deviation of the relative error are also tabulated as are earlier results obtained by other authors.
The influence of coupling between flexural and extensional deformation and coupling between structure and acoustic volume on the dynamic response of piezoelectric ceramic transducer elements mounted on metal diaphragms is analyzed using three analytical methods: 1) classical boundary value techniques; 2) simple direct variational procedures; and 3) finite element methods. The analyses are able to predict the voltage output of the transducer, including resonant amplitudes and shapes, with reasonable accuracy and also to indicate critical front and back acoustic volume design parameters needed to control resonance. The finite element model includes a general formulation for axisymmetric layered shells of revolution (which degenerates to a circular plate), whose average normal displacement is coupled to the long wavelength motion of air in adjacent cavities (acoustic stiffness), ports (acoustic mass), and porous plugs (acoustic damping). The methods outlined here are also applicable to window-enclosure response to sonic boom excitation, skull-brain impact studies, and the study of respiratory mechanics.
We summarize work between 1969 and 1972 in a continuing project With two objectives: to produce acceptable synthetic speech directly from English text; and to demonstrate with speech synthesis a detailed model of human articulatory movements. Work in the four-year period has yielded moderately accurate rules for predicting the occurrence of pauses and lesser breaks in the sentence; rules for vowel duration in many conditions, not just primary stressed syllables immediately before a pause; rules for contextual variations of consonants; and rules for durational and other allophonic variations on consonants at word boundaries. Presently we are studying natural speech to quantify and add detail to these rules, and we are working to extend the vocal tract model to closer agreement with human articulation and vocal cord control.
Several new methods of realization of an arbitrary digital transfer function are proposed. The final realization is in the form of a two-input two-output digital filter configuration with one input variable constrained to be a multiple of one of the output variables. All realizations are canonic with respect to the delays but use different numbers of adders and multipliers. Examples illustrating the methods are included.
An experiment was performed in which the authors attempted to recognize a set of unknown sentences by visual examination of spectrograms and machine-aided lexical searching. Ninteen sentences representing data from five talkers were analyzed. An initial partial transcription in terms of phonetic features was performed. The transcription contained many errors and omissions: 10 percent of the segments were omitted, 17 percent were incorrectly transcribed, and an additional 40 percent were transcribed only partially in terms of phonetic features. The transcription was used by the experimenters to initiate computerized scans of a 200-word lexicon. A majority of the search responses did not contain the correct word. However, following extended interactions with the computer, a word-recognition rate of 96 percent was achieved by each investigator for the sentence material. Implications for automatic speech recognition are discussed. In particular, the differences between the phonetic characteristics of isolated words and of the same words when they appear in sentences are emphasized.
The electrotactile sound detector described here is designed to enable deaf persons to detect and localize sounds. Two microphones are worn bilaterally on the head, the sounds received are converted to electrical pulses, and the pulses are fed to two electrodes applied to the forehead. Differences in intensity of the pulses permit the wearer to localize the source of a sound. Additional information is furnished about the rhythmic patterning of sounds.
An important class of applications in digital signal processing involves the numerical solution of the convolution integral. These so-called numerical deconvolution problems are notoriously difficult to solve because of their inherent ill-conditioning. In this paper we present a characterization of this ill-conditioning based on a classical spectral decomposition of the discrete convolution. Factors prominently influencing the conditioning are identified and some explicit sensitivity measures are introduced.
Signal-to-noise level may be generalized to the form logSa/Nb, where S and N are respectively signal and noise values expressed in power units. As usually defined, the exponents a and b both assume values of +1. This paper presents some data analyses that were performed to study the optimum value of the ratio a/b as a predictor of speech transmission quality. In two separate experiments using simulated and normal telephone circuits, subjects rated the transmission quality in the presence of interference. In the first experiment, the interfering variables were crosstalk, speech, and noise. In the experiment, the variables were noise and loss of speech energy. A factor analysis of the rating data indicated that signal-to-noise level was the dominant factor. Both canonical correlation and multiple regression analyses were performed to find the optimum value of the ratio a/b as mentioned above. In all cases, it was found that a/b > 1 was a better predictor of transmission quality than a/b = 1.