Recent work has presented max-equivocation as a measure of the resistance of a cryptosystem to attacks when the attacker is aware of the encoder function and message distribution. Here we consider the vulnerability of a cryptosystem in the one-try attack scenario when the attacker has incomplete information about the encoder function and message distribution. We show that encoder functions alone yield information to the attacker, and combined with inferable information about the ciphertexts, information about the message distribution can be discovered. We show that the whole encoder function need not be fixed or shared a priori for an effective cryptosystem, and this can be exploited to increase the equivocation over an a priori shared encoder. Finally we present two algorithms that operate in these scenarios and achieve good equivocation results, ExPad that demonstrates the key concepts, and ShortPad that has less overhead than ExPad.
The deployment of systems for human-tomachine communication by voice requires overcoming a variety of obstacles that affect the speech-processing technologies. Problems encountered in the field might include variation in speaking style, acoustic noise, ambiguity of language, or confusion on the part of the speaker. The diversity of these practical problems encountered in the "real world" leads to the perceived gap between laboratory and "real-world" performance. To answer the question "What applications can speech technology support today?" the concept of the "degree of difficulty" of an application is introduced. The degree of difficulty depends not only on the demands placed on the speech recognition and speech synthesis technologies but also on the expectations of the user of the system. Experience has shown that deployment of effective speech communication systems requires an iterative process. This paper discusses general deployment principles, which are illustrated by several examples of human-machine communication systems. Speech-processing technology is now at the point at which people can engage in voice dialogues with machines, at least in limited ways. Simple voice communication with machines is now deployed in personal computers, in the automation of long-distance calls, and in voice dialing of mobile telephones. These systems have small vocabularies and strictly circumscribed task domains. In research laboratories there are advanced human-machine dialogue systems with vocabularies of thousands of words and intelligence to carry on a conversation on specific topics. Despite these successes, it is clear that the truly intelligent systems envisioned in science fiction are still far in the future, given the state of the art today. Human-machine dialogue systems can be represented as a four-step process, as shown in Fig. 1. This figure encompasses both the simple systems deployed today and the spoken language understanding we envision for the future. First, a speech recognizer transcribes sentences spoken by a person into written text (1, 2). Second, a language understanding module extracts the meaning from the text (3, 4). Third, a computer (consisting of a processor and a database) performs some action based on the meaning of what was said. Fourth, the person receives feedback from the computer in the form of a voice created by a speech synthesizer (5, 6). The boundaries between these stages of a dialogue system may not be distinct in practice. For instance, language-understanding modules may have to cope with errors in the text from the speech recognizer, and the speech recognizer may make use of grammar and semantic constraints from the language module in order to reduce recognition errors. In the 1993 "Colloquium on Human-Machine Communication by Voice," sponsored by the National Academy of Sciences (NAS), much of the discussion focused on practical difficulties in building and deploying systems for carrying on voice dialogues between humans and machines. Deployment of systems for human-to-machine communication by voice requires solutions to many types of problems that affect the speech-processing technologies. Problems encountered in the field might include variation in speaking style, noise, ambiguity of language, or confusion on the part of the speaker. There was a consensus at the colloquium that a gap exists between performance in the laboratory and accuracy in the field, because conditions in real applications are more difficult. However, there was little agreement about the cause of this gap in performance or what to do about it. A key point of discussion at the NAS colloquium concerned the factors that make a dialogue easy or difficult. Many such degrees of difficulty were mentioned in a qualitative way. To summarize the discussion in this paper it seems useful to introduce a more formal concept of the degree of difficulty of a human-machine dialogue and to list each dimension that contributes to the overall difficulty. The degree of difficulty is a useful concept, despite the fact that it is only a "fuzzy" (or qualitative) measure because of lack of precision in quantifying an overall degree of difficulty for an application. A second point of discussion during the NAS colloquium concerned the process of deployment of human-machine dialogue systems. In several cases such systems were built and then modified substantially as the designers gained experience in what the technology could support or in user-interface issues. This paper elaborates on this iterative deployment process and contrasts it with the deployment process of more mature technologies. DEGREE OF DIFFICULTY OF A VOICE DIALOGUE APPLICATION Whether a voice dialogue system is successful depends on the difficulty of each of the four steps in Fig. 1 for the particular application, as well as the technical capabilities of the computer system. There are several factors that can make each of these four steps difficult. Unfortunately, it is difficult to quantify precisely the difficulty of these factors. If the technology performs unsatisfactorily at any stage of processing of the speech dialogue, the entire dialogue will be unsatisfactory. We hope that technology will eventually improve to the point that there are no technical barriers whatsoever to complex speech-understanding systems. But until that time it is important to know what is easy, what is difficult but possible, and what is impossible, given today's technology. What are the factors that determine the degree of difficulty of a voice dialogue system? In practice, there are several factors for each step of the voice dialogue that may make the task difficult or easy. Because these factors are qualitatively independent, they can be viewed as independent variables in a multidimensional space. For a simple example of two dimensions of difficulty for speech recognition refer to Fig. 2. 10017 The publication costs of this article were defrayed in part by page charge payment. This article must therefore be hereby marked "advertisement" in accordance with 18 U.S.C. §1734 solely to indicate this fact. Proc. Natl. Acad. Sci. USA 92 (1995) FIG. 1. A human-machine dialogue system. There are many dimensions of difficulty for a speech recognition system, two of which are shown. Eight applications are rated according to difficulty of speaking mode (vertical axis) and vocabulary size (horizontal axis). "Voice dictation" refers to commercial 30,000-word voice typewriters; "V.R.C.P." stands for voice recognition call processing, as a way !to automate long-distance calling; "telephone number dialing" refers to connected digit recognition of phone numbers; "DARPA resource management" refers to the 991-word Naval Resource Management task with a constraining grammar; "A.T.I.S." stands for the DARPA (Defense Advanced Research Projects Agency) Air Travel Information System; and "natural spoken language" refers to conversational speech on any and every topic. Clearly, a task that is difficult in both the dimensions of vocabulary size and speaking style would be harder (and would have lower accuracy) than a small, isolated word recognizer, if all other factors are equal. The other factors are not equal, as discussed in the section "Dimensions of the Recognition Task." Note that telephone number dialing, despite its position on the two axes in this figure, is a difficult application because of dimensions not shown here, such as user tolerance of errors and grammar perplexity. Ideally, a potential human-machine dialogue could receive a numerical rating along each dimension of difficulty, and a cumulative degree of difficulty could be computed by summing the ratings along each separate dimension. Such a quantitative approach is overly simplistic. Nevertheless, it is a valuable exercise to evaluate potential applications qualitatively along each of the dimensions of difficulty. The problems for voice dialogue systems can be separated into those of speech recognition, language understanding, and speech synthesis, as in Fig. 1. (For the database access stage, a conventional computer is adequate for most voice dialogue tasks. The data-processing capabilities of today's machines pose no barriers to development of human-machine communication systems.) Let us examine the steps of speech recognition, language understanding, and speech synthesis in order
Science fiction has long been populated with conversational computers and robots. Now, speech synthesis and recognition have matured to where a wide range of real-world applications--from serving people with disabilities to boosting the nation's competitiveness--are within our grasp. Voice Communication Between Humans and Machines takes the first interdisciplinary look at what we know about voice processing, where our technologies stand, and what the future may hold for this fascinating field. The volume integrates theoretical, technical, and practical views from world-class experts at leading research centers around the world, reporting on the scientific bases behind human-machine voice communication, the state of the art in computerization, and progress in user friendliness. It offers an up-to-date treatment of technological progress in key areas: speech synthesis, speech recognition, and natural language understanding. The book also explores the emergence of the voice processing industry and specific opportunities in telecommunications and other businesses, in military and government operations, and in assistance for the disabled. It outlines, as well, practical issues and research questions that must be resolved if machines are to become fellow problem-solvers along with humans. Voice Communication Between Humans and Machines provides a comprehensive understanding of the field of voice processing for engineers, researchers, and business executives, as well as speech and hearing specialists, advocates for people with disabilities, faculty and students, and interested individuals.
The authors investigate the parallel computation of large-vocabulary speech recognition on a tree-structured parallel processor. Having seriously established results on parallel level-building with continuously variable hidden Markov models they extend this work to the case in which the spoken sentences are constrained by a finite state grammar. The two key ideas are: (1) a pipelined sorting function on the processor array that efficiently transmits a sorted list of the best scores from all processors to the host; and (2) a level-based pruning technique in which paths through the dynamic programming network are pruned only at the ends of words. These ideas are evaluated on the BT-100 processor, a binary-tree parallel processor. A performance model is presented that estimates execution time as a function of algorithm parameters. Real-time speech recognition has been achieved for a data-entry task with a 70-word vocabulary and average branching factor of 23, and for an airline reservation task with a vocabulary of 132 words. The performance model predicts real-time execution of the 991-word DARPA Resource Management Task on a 127-processor machine
The authors describe a parallel frame-synchronous level-building algorithm, utilizing HMM word-models, for connected-speech recognition on a tree-structured parallel computer. The algorithm is scalable in the sense that the source code and execution time remain essentially the same as vocabulary size increases, so long as the hardware is scaled proportionally. An illustrative sizing and timing analysis of a speaker-independent connected-digit recognizer on the ASPEN tree-machine is described. This algorithm executes in real-time and achieves 98.3% string accuracy on the Texas Instruments digit data base
Techniques for training hidden Markov model (HMM) parameters from a labeled training set of data are well established and include the forward-backward algorithm as well as the segmental K-means algorithm. These algorithms have been shown to be capable of estimating the parameters of an HMM based on mathematically well-founded techniques. In practice, however, difficulties are often encountered when estimating some of the HMM parameters. These difficulties are generally the result of having insufficient training data to give robust and reliable parameter estimates. Typically, the model parameters most affected by having insufficient training data are the spectral parameter variance estimates, and the estimates of parameters related to the modeling of state duration. Although techniques have been proposed for improving estimates of the variances due to the effects of insufficient training data, the results have not proven adequate in some cases. As such, improved training techniques (which give better recognition performance) have been devised for controlling the minimum variance estimate of any spectral parameter, and for thresholding and clipping state duration parameter estimates. These improved training methods have been tested on several databases with good success. In addition, advanced techniques for creating multiple HMMs from the training data (i.e., for speaker independent recognition) have been devised and have proven successful for modeling large databases of training material.