The performance of spoken language systems on utterances from the ATIS domain is evaluated by comparing system-produced responses with hand-crafted (and -verified) standard responses to the same utterances. The objective of SRI's annotation project is to provide SLS system developers with the range of correct responses to human utterances produced during experimental sessions with ATIS domain interactive systems. These correct responses are then used in system training and evaluation.
The objective of the CSR Data Collection effort is to collect and deliver a large corpus of continuous speech data to support ARPA research efforts in continuous speech recognition (CSR).
The objective of the CSR Corpus Development is to collect and deliver a large corpus of continuous speech data to support DARPA research efforts in continuous speech recognition (CSR). SRI's current goal is the completion of Phase 2, Part 1 of the planned CSR Corpus. This consists of 86,000 sentences from 275 speakers, including 8000 spontaneous sentences from 40 journalists.
The project goal is to collect and deliver a corpus of speech data that supports DARPA SLS system development. As of February 1991, SRI has set up a hardware and software environment for the collection of spoken interactions with a simulated Air Travel Information System (ATIS), established a data collection procedure, collected and distributed prototype data, and evaluated the prototype data with feedback from the SLS system developers. Having implemented revisions in the environment and procedures, SRI has begun collecting and distributing a corpus of data for ATIS SLS development.
SRI is developing a system that uses real time speech recognition to diag nose, evaluate and provide training in spoken English. The paper first describes the methods and results of a study of the feasibility of automati cally grading the performance of Japanese students when reading English aloud. Utterances recorded from Japanese speakers were independently rated by expert listeners. Speech grading software was developed from a speaker independent hidden-Markov-model speech recognition system. The auto matic grading procedure first aligned the speech with a model and then com pared the segments of the speech signal with models of those segments that have been developed from a database of speech from native speakers of English. The evaluation study showed that ratings of speech quality by experts are very reliable and automatic grades correlate well (r > 0.8) with those expert ratings. SRI is now extending this technology and integrating it in a spoken-language training system. This effort involves (1) porting SRI's DECIPHER speech recognition system to a microcomputer platform, and (2) extending the speech-evaluation software to more exactly diagnose a learner's pronuncia tion deficits and lead the learner through an appropriate regimen of exer cises.
SRI has developed a speaker-independent continuous speech, large vocabulary speech recognition system, DECIPHER, that provides state-of-the-art performance on the DARPA standard speaker-independent resource management training and testing materials. SRI's approach is to integrate speech and linguistic knowledge into the HMM framework. This paper describes performance improvements arising from detailed phonological modeling and from the incorporation of cross-word coarticulatory constraints.
The paper describes the methods and results of a study of the feasibility of automatically grading the performance of Japanese students when reading English aloud. SRI recorded 31 adult Japanese speakers: 22 men and 9 women. Each Japanese speaker read six sentences aloud. All 186 recorded utterances were presented in a random order for rating by three expert listeners who rated the utterances on two occasions. Speech-grading software was developed from an adaptive hidden-Markov-model (HMM) speech-recognition system. The grading procedure is a two-step process: First, the speech to be graded is aligned, then the segments of the speech signal that are located are compared with models of those segments that have been developed from a database of speech from native speakers of English. Important points in the results are: (1) ratings of speech quality by expert listeners are extremely reliable, and (2) automatic grades from the system correlate well (>0.8) with those ratings.
A database of continuous read speech has been designed and recorded within the DARPA strategic computing speech recognition program. The data is intended for use in designing and evaluating algorithms for speaker-independent, speaker-adaptive and speaker-dependent speech recognition. The data consists of read sentences appropriate to a naval resource management task built around existing interactive database and graphics programs. The 1000-word task vocabulary is intended to be logically complete and habitable. The database, which represents over 21000 recorded utterances from 160 talkers with a variety of dialects, includes a partition of sentences and talkers for training and for testing purposes.<>
Some speech recognition systems use alternative lexical forms to cover various pronunciations encountered when people speak. However, excess forms in the lexicon degrade recognition performance. Thus lexicons should maximize coverage of actual pronunciations that occur, while minimizing the number of alternative pronunciations. A set of phonological rules for English as been developed at SRI to cover a large body of transcriptions of sentences read by American speakers. This phonological rule set covers 98.5% of the segments in a separate set of transcriptions. Rule efficiency can be defined as the ratio of the number of times forms generated by a rule were spoken in a corpus of test data to the number of forms generated by the rule for the sentences in that test data. By this definition, palatalization of /t/ and /d/ before tautosyllabic /r/ is efficient, while flapping of /t/ that follows /l/ is not. There is a trade-off between the coverage of a rule set and the number of alternative pronunciations generated by the rule set. Rule efficiency can be used to rank rules in a rule set, to derive a family of rule sets that maximize coverage for a given number of alternative forms represented, or to select among alternative rule sets. [Work supported by DARPA.]
An understanding of the structure of pronunciation variation over a population of speakers and over time in the utterances of one speaker should be useful in designing speaker-independent speech recognizers. This paper reports a series of experiments designed to show different kinds of patterns observed in alternative forms of words in constant contexts (e.g., the presence or absence of frication in the “y” in “had your”). An analysis is presented of transcribed data from 630 speakers reading two sample sentences as well as data from four speakers reading the same two sentences 24 times each, separated by filler material, in three separate sessions. The analysis quantifies the relative usefulness of competing models of variation in information theoretic terms. The results indicate that (1) speakers can be clustered into low variation groups such that the variation within a group is significantly less than the population variation, and (2) individual speakers show greater consistency than comparable clustered subsets of the population. Finally, it is suggested how this structure may be used to guide rapid, automatic adaptation in speech recognition. [Work supported by NSE.]
This paper describes an alternative approach to lexical access in the CMU ANGEL speech recognition system. Using this approach, the asynchronous phonetic hypotheses generated by an acoustic-phonetics module are converted to a directed graph. This graph is compared to a pronunciation dictionary. Performance results for this approach and the original CMU approach are similar. An error analysis indicates promising directions for further work.
Recent analyses of the allophonic variants of phonemes in particular environments often seem to assume one or more of the following: (1) The proportions of variants encountered in a multispeaker sample represent an “irreducible” statistical component of phonology; (2) these proportions predict the likelihood of encountering these same allophones in new material; (3) the probability of encountering a particular allophone of some phoneme is independent of the observed allophones of other phonemes nearby. These related assumptions are questioned on theoretical and practical grounds, using transcribed data from 630 speakers reading two sample sentences. The frequencies of occurrence of all the allophones of certain phoneme tokens in the sample sentences were measured across all speakers; then the conditional co-occurrence of all the allophones of certain phoneme pairs in the sample sentences were analyzed. For example, first the occurrence of a flap in the words “suit in” was counted across all speakers; then the occurrence of a flap in “suit in” was counted for only those speakers who deleted /t/ in “don't ask,” and so on. Comparing these two kinds of analysis has implications for theories of variation. The application co-occurance analysis in speech recognition will be illustrated. [Work supported by DARPA.]
Assembling a speech data base that is both manageably small and sufficiently diverse can be a useful step in the development of speaker independent speech recognition systems. Yet there has been no data on what kind of speaker sample might be required to ensure a group whose speech includes certain phonetic or linguistic traits. The data gathered in this study suggests that some common and important dialect features will not be found even in a large number of speakers, if sampling is conducted at a single location. In order to compile a large pool of prospective speakers, 152 people were recorded for about one or two minutes speaking extemporaneously; the recordings were then rated by the three authors according to fifteen characteristics that form three classes: voice quality, manner of speaking, and dialect. Although a wide variety of voice characteristics and manners of speaking were evident among the 152 speakers, the dialect features covered a limited range. We discuss the possible causes of this distribution of characteristics in the sample and some of its implications for collecting adequate databases for speech recognition research.
Some speakers use different forms when training a speech recognizer than when speaking spontaneously to the device—this could be called “enrollment diglossia.” As a preliminary study of this phenomenon, we compared selected phonological and prosodic features of spontaneous speech to read and recited versions of the same sentences and paragraphs. Three subjects were interviewed and were later asked to read or memorize and recite, at various nominal rates, portions of the material that they had originally spoken spontaneously. We made detailed measurements of /t/ allophonics, speech rate, and the forms of certain words. For some speakers, there are considerable differences between spontaneous and prepared renditions; e.g., one speaker produced 45% vs 18% of /t/s as flaps, another speaker varies local speech rate and produces more nonsyntactic pauses in spontaneous paragraphs, and one speaker invokes “fast speech” forms at slower speech rates for spontaneous speech than for prepared speech. [Work supported by DARPA.]
It may be possible for a deaf person and a hearing person to converse over the telephone. The deaf user "speaks" by typing to a text-to-speech converter using a low-redundancy keyboard system. The hearing user speaks sentences one word at a time to a large-vocabulary, isolated-word-recognition system that displays a sentence lattice (a sequence of sets of likely matches for each word spoken). The deaf user then tries to find a sensible path through the sentence lattice. Successful implementation of such a system, under development at SRI International, requires adequate performance in text generation speed by deaf users, text-to-speech intelligibility, and word-at-a-time speaking by hearing users, as well as large-vocabulary speech recognition and disambiguation of sentence lattices by deaf users.
Douglas E. Appelt合作论文数Artificial Intelligence Center1