Rhythm patterns play an important role in the perception of second-language (L2) speech. This paper presents a novel approach to evaluating L2 speech rhythm using low-frequency spectral features inspired by the rhythmogram auditory model. In this paper we investigate several new feature sets for use in training rhythm-centric acoustic models. By capturing information over suprasegmental linguistic units appropriate for rhythmic analysis (including syllables and prosodic feet), these novel features can outperform traditional features in detecting rhythm errors on the ISLE corpus of learner English by 5-15% absolute.
This paper presents a novel student model intended to automate word-list-based reading assessments in a classroom setting, specifically for a student population that includes both native and nonnative speakers of English. As a Bayesian Network, the model is meant to conceive of student reading skills as a conscientious teacher would, incorporating cues based on expert knowledge of pronunciation variants and their cognitive or phonological sources, as well as prior knowledge of the student and the test itself. Alongside a hypothesized structure of conditional dependencies, we also propose an automatic method for refining the Bayes Net to eliminate unnecessary arcs. Reading assessment baselines that use strict pronunciation scoring alone (without other prior knowledge) achieve 0.7 correlation of their automatic scores with human assessments on the TBALL dataset. Our proposed structure significantly outperforms this baseline, and a simpler data-driven structure achieves 0.87 correlation through the use of novel features, surpassing the lower range of inter-annotator agreement. Scores estimated by this new model are also shown to exhibit the same biases along demographic lines as human listeners. Though used here for reading assessment, this model paradigm could be used in other pedagogical applications like foreign language instruction, or for inferring abstract cognitive states like categorical emotions.
Automatic literacy assessment is an area of research that has shown significant progress in recent years. Technology can be used to automatically administer reading tasks and analyze and interpret children's reading skills. It has the potential to transform the classroom dynamic by providing useful information to teachers in a repeatable, consistent, and affordable way. While most previous research has focused on automatically assessing children reading words and sentences, assessments of children's earlier foundational skills is needed. We address this problem in this research by automatically verifying preliterate children's pronunciations of English letter-names and the sounds each letter represents (“letter-sounds”). The children analyzed in this study were from a diverse bilingual background and were recorded in actual kindergarten to second grade classrooms. We first manually verified (accept/reject) the letter-name and letter-sound utterances, which serve as the ground-truth in this study. Next, we investigated four automatic verification methods that were based on automatic speech recognition techniques. We attained percent agreement with human evaluations of 90% and 85% for the letter-name and letter-sound tasks, respectively. Humans agree between themselves an average of 95% of the time for both tasks. We discuss the various confounding factors for this assessment task, such as background noise and the presence of disfluencies, that impact automatic verification performance.
The perception of rhythmic differences among languages relies on varieties in periodicity within prominence groups. But the consensus in phonetic research on rhythm is that existing measures don’t capture true rhythm by that definition instead, they merely measure short-term timing. This work proposes a new rhythm measure, the Generalized Variability Index (GVI), that examines durational contexts over arbitrarily long linguistic distances. To evaluate this new measure, we conducted a set of experiments in automatic language identification using large amounts of data from 11 languages in the Globalphone and TIMIT corpora. When added to baseline rhythm measures, these new GVI features offer absolute improvement in 11-way language classification accuracy by as much as 12%. Moreover, the addition of wider and wider durational context in the GVI continues to contribute information useful for automatic language ID, abating in usefulness only at a distance of about 10 syllables.
Motivated by a desire to assess the prosody of foreign language learners, this study demonstrates the benefit of highlevel syntactic information in automatically deciding where phrase breaks and pitch accents should go in text. The connection between syntax and prosody is well-established, and naturally lends itself to tree-based probabilistic models. With automatically-derived parse trees paired to tree transducer models, we found that categorical prosody tags for unseen text can be determined with significantly higher accuracy than they can with a baseline method that uses n-gram models of part-ofspeech tags. On the Boston University Radio News Corpus, the tree transducer outperformed the baseline by 14% overall for accents, and by 3% overall for breaks. These automatic results fell within this corpus’s range of inter-speaker agreement in assigning accents and breaks to text.
Parroting exercises in a foreign language are designed to make a student’s speech more native-like through imitation of specific native speech templates. In this paper we describe novel template-based methods for automatically estimating subjective scores for both intonation and rhythm in nonnative English. In terms of accuracy when automatically classifying a parroting speaker as a native or a learner, experimental results show that these new rhythm and intonation scores outperform similar baselines from nonnative speech assessment literature, and that they offer complementary discriminatory information when combined with automatic segment-level pronunciation scores, reaching a maximum classification accuracy of 89.8% on a corpus of parroting exercises. This suggests the general usefulness of these new scores in automatically assessing nonnative pronunciation in a computer-assisted pronunciation practice scenario.
Automatic literacy assessment technology can help children acquire reading skills by providing teachers valuable feedback in a repeatable, consistent manner. Recent research efforts have concentrated on detecting mispronunciations during word-reading and sentence-reading tasks. These token-level assessments are important since they highlight specific errors made by the child. However, there is also a need for more high-level automatic assessments that capture the overall performance of the children. These high-level assessments can be viewed as an interpretive extension to token-level assessments, and may be more perceptually relevant to teachers and helpful in tracking performance over time. In this paper, we model and predict the overall reading ability of young children reading a list of English words aloud. The data consist of audio recordings, collected in real kindergarten to second grade classrooms from children from native English- and Spanish-speaking households. This research is broken into two main parts. The first part is a user study, in which 11 human evaluators rated the children on their overall reading ability based on the audio recordings. The evaluators were volunteers from a diverse background, seven of whom were native speakers of American English and four that were fluent speakers of English as a secondary language. While none of the evaluators were trained reading experts or licensed teachers, a subset of them were linguists and researchers with experience in automatic literacy assessment. As part of this work, we analyzed the effect of the evaluator's background on inter-evaluator agreement. In the second part, we ran machine learning experiments to predict evaluators' scores using features automatically extracted from the audio. The features were human-inspired and correlated with cues human evaluators stated they used: pronunciation correctness, speaking rate, and fluency. We investigated various automated methods to verify the correctness of the word pronunciations and to detect disfluencies in the children's speech using held-out annotated data. Using linear regression techniques, we automatically predicted individual evaluators' high-level scores with a mean Pearson correlation coefficient of 0.828, and we predicted average evaluator's scores with correlation 0.946. Both these human-machine agreement statistics exceeded the mean inter-evaluator agreement statistics.
Articulatory Phonology’s link between cognitive speech planning and the physical realizations of vocal tract constrictions has implications for speech acoustic and duration modeling that should be useful in assigning subjective ratings of pronunciation quality to nonnative speech. In this work, we compare traditional phoneme models used in automatic speech recognition to similar models for articulatory gestural pattern vectors, each with associated duration models. What we find is that, on the CDT corpus, gestural models outperform the phonemelevel baseline in terms of correlation with listener ratings, and in combination phoneme and gestural models outperform either one alone. This also validates previous findings with a similar (but not gesture-based) pseudo-articulatory representation. Index Terms: pronunciation modeling, nonnative speech, articulatory phonology
Children need to master reading letter-names and letter-sounds before reading phrases and sentences. Pronunciation assessment of letter-names and letter-sounds read aloud is an important component of preliterate children's education, and automating this process can have several advantages. The goal of this work was to automatically verify letter-names spoken by kindergarteners and first graders in realistic classroom noise conditions. We applied the same techniques developed in our previous work on automatic letter-sound verification by comparing and optimizing different acoustic models, dictionaries, and decoding grammars. Our final system was unbiased with respect to the child's grade, age, and native language and achieved 93.1% agreement (0.813 kappa agreement) with human evaluators, who agreed among themselves 95.4% of the time (0.891 kappa).
To automate assessments of beginning readers, especially those still learning English, we have investigated the types of knowledge sources that teachers use and have tried to incorporate them into an automated system. We describe a set of speech recognition and verification experiments and compare teacher scores with automatic scores in order to decide when a novel pronunciation is best viewed as a reading error or as dialect variation. Since no one classroom teacher is expected to be familiar with as many dialect systems as might occur in an urban classroom, making progress in automated assessments in this area can improve the consistency and fairness of reading assessment. We found that automatic methods performed best when the acoustic models were trained on both native and non-native speech, and argue that this training condition is necessary for automatic reading assessment since a child's reading ability is not directly observable in one utterance. We also found assessment of emerging reading skills in young children to be an area ripe for more research!
Past studies have shown that a native Spanish speaker’s use of phrasal prominence is a good indicator of her level of English prosody acquisition. Because of the cross-linguistic differences in the organization of phrasal prominence and durational contrasts, we hypothesize that those speakers with English-like prominence in their L2 speech are also expected to have acquired English-like rhythm. Statistics from a corpus of native and nonnative English confirm that speakers with an Englishlike phrasal prominence are also the ones who use English-like rhythm. Additionally, two methods of automatic score generation based on vowel duration times demonstrate a correlation of at least 0.6 between these automatic scores and subjective scores for phrasal prominence. These findings suggest that simple vowel duration measures obtained from standard automatic speech recognition methods can be salient cues for estimating subjective scores of prosodic acquisition, and of pronunciation in general.
We present the first study of nonnative English speech using real-time MRI analysis. The purpose of this study is to investigate the articulatory nature of “phonological transfer”—a speaker’s systematic use of sounds from their native language (L1) when they are speaking a foreign language (L2). When a non-native speaker is prompted to produce a phoneme that does not exist in their L1, we hypothesize that their articulation of that phoneme will be colored by that of the “closest” phoneme in their L1’s set, possibly to the point of substitution. With data from three native German speakers and three reference native English speakers, we compare articulation of read phoneme targets well documented as “difficult” for German speakers of English (/w/ and /dh/) with their most common substitutions (/v/ and /d/, respectively). Tracking of vocal tract organs in the MRI images reveals that the acoustic variability in a foreign accent can indeed be ascribed to the subtle articulatory influence of these close substitutions. This suggests that studies in automatic pronunciation evaluation can benefit from the use of articulatory rather than phoneme-level acoustic models. [Work supported by NIH.]
Phonological transfer is the influence of a first language on phonological variations made when speaking a second language. With automatic pronunciation assessment applications in mind, this study intends to uncover evidence of phonological transfer in terms of articulation. Real-time MRI videos from three German speakers of English and three native English speakers are compared to uncover the influence of German consonants on close English consonants not found in German. Results show that nonnative speakers demonstrate the effects of L1 transfer through the absence of articulatory contrasts seen in native speakers, while still maintaining minimal articulatory contrasts that are necessary for automatic detection of pronunciation errors, encouraging the further use of articulatory models for speech error characterization and detection. Index Terms: real-time MRI, nonnative speech, articulation, phonological transfer
Automatic reading assessment software has the difficult task of trying to model human-based observations, which have both objective and subjective components. In this paper, we mimic the grading patterns of a "ground-truth" (average) evaluator in order to produce models that agree with many people's judgments. We examine one particular reading task, where children read a list of words aloud, and evaluators rate the children's overall reading ability on a scale from one to seven. We first extract various features correlated with the specific cues that evaluators said they used. We then compare various supervised learning methods that mapped the most relevant features to the ground-truth evaluator scores. Our final system predicted these scores with 0.91 correlation, higher than the average inter-evaluator agreement.
Evaluations of letter naming and letter sounding are commonly used to measure a young child's growing reading ability, since performance in them is well-correlated with future reading development. Assessing a child's oral reading skills requires teachers, as well as technologies that attempt to automate such assessment, to form an item-level accept/reject decision based on speech cues and prior knowledge of the child's literacy level and linguistic background. With data collected from 171 K-2 children, both learners and native speakers of American English, we designed and evaluated an automated letter naming assessment method using a simple word-loop HMM decoding for the word-level letter names. The automated accept/reject evaluation performance, 81.9%, approached the agreement of human raters, 83.2% (0.62 kappa). However, the task where children must produce the sound that the letter represents was more difficult: English orthography allows one-to many letter-to-sound mapping, teachers showed less agreement in their assessment (80.9%, 0.55 kappa), and the brief durations of some of the letter sounds made it difficult to distinguish them from each other and from background noises. Phone-level HMM based evaluation accuracy was 58.2%. Preprocessing the recordings into speech, silence, and noise improved these results, especially for plosive sounds. [Supported by NSF]
The use of speech technology in children’s reading assessment can help teachers to diagnose reading difficulties and plan appropriate interventions for a large number of students. We present a Bayesian Network model of student reading comprehension that can be used to estimate automatic scores for a child’s spoken answers to open-ended questions about a text. Through the use of features derived from language models capturing different degrees of comprehension, we found that on the TBALL dataset we could achieve 0.8 correlation with reference comprehension scores derived from teachers, exceeding the teachers’ own correlation with this same reference. This student model also proved to perform without bias due to a speaker’s native language, which was not the case for a comparable baseline method, nor for the teachers themselves.