The speaker verification is based on variations in formant frequencies at stationary fragments and transient processes of vowels, the spectral features of fricative sounds, and the duration of speech segments. The best features are chosen for each word from the fixed list of Russian numerals ranging from zero to nine. The password phrase is randomly generated by the system at each verification. The compensation for dynamic noise and the counteraction with respect to interference using the reproduction of the intercepted and recorded speech are provided by the repeated reproduction of several words. The total error probabilities for male and female voices are 0.006 and 0.025%, respectively, for 30 million tests, 429 speakers, and a maximum length of the password phrase of 10 words. Note that the probabilities of false identification and false rejection are almost equal.
An algorithm for estimating the vocal pulse positions and durations in an actual speech signal is described. Testing of the algorithm shows that it outperforms the best of the competitor algorithms in accuracy on the average by a factor of two. The algorithm is less sensitive to spectrum distortions in telephone channels, to various types of noise, and to instability in duration and amplitude of pulses produced by the voice source. The accuracy of the pulse position estimate is sufficient for a synchronous speech signal analysis, while the speed of signal processing makes the algorithm suitable for real-time operation.
Inverse problems with respect to parameters of the articulatory model are solved for all types of sounds: vowels, semi-vowels, nasals, stops and fricatives in various contexts. Acoustical parameters of the speech signal and trajectories of some reference points inside the vocal tract serve as input data. 3.7%, 3.8% and 2.6% average approximation error for the first three formants, 8.5% for the specific frequencies of fricative spectra, 2.8% for the coordinates of reference points for all kinds of phonemes are obtained when both – acoustic and articulatory data are used. 1.8%, 1.6%, and 1.1% error for the first three formant frequencies, and 6% for the coordinates of reference points are obtained when only acoustic data are used. Original and re-synthesized utterances are found to be very similar in appearance, according to subjective assessment.
The upper bound for errors of the least squares (LS) estimation is determined for arbitrary individual measurement errors with limited magnitudes. It is shown that the accuracy of the LS method may be substantially worse than an optimistic estimate of its value obtained under the assumption of independent errors of individual measurements, and worse than the estimated accuracy of the minimum data method, in which all excessive measurements are discarded. A fast algorithm is developed for calculating the upper bound of estimation errors for coefficients of an LS polynomial approximation.