Shadowing has become a well-known method to improve learners' overall proficiency. Our previous studies realized automatic scoring of shadowing speech using HMM phoneme posteriors, called GOP (Goodness of Pronunciation) and learners' TOEIC scores were predicted adequately. In this study, we enhance our studies from multiple angles: 1) a much larger amount of shadowing speech is collected, 2) manual scoring of these utterances is done by two native teachers, 3) DNN posteriors are introduced instead of HMM ones, 4) language -independent shadowing assessment based on posteriors -based DTW (Dynamic Time Warping) is examined. Experiments suggest that, compared to HMM, DNN can improve teacher -machine correlation largely by 0.37 and DTW based on DNN posteriors shows as high correlation as 0.74 even when posterior calculation is done using a different language from the target language of learning.
We demonstrate two systems, network-based and standalone, for collecting shadowing utterances and their automatic assessment. Since June 2016, these systems have been used in real English classes at several universities in Japan. The following features are highlighted: 1) To avoid pop noises from a speaker and reduce babble noises from other surrounding students, an ear-hook microphone is used with a USB audio device, 2) A network-based system and a standalone system were developed separately because network traffic not rarely causes technical errors when recording, 3) For beginners, easy-to-understand illustrations are prepared for them to get accustomed rapidly to shadowing practices, 4) DNN-based GOP calculation is run and its score is fed back to learners, and 5) To motivate them, the GOP score distribution over the learners are also fed back for them to compare their own scores with others’. In demonstration, each feature will be exhibited and explained in detail.
This study investigated relationships between teacher-selected stimulus sentences and machine-suggested ones in terms of the correlation between human ratings and GOP-based machine scores. In assessing shadowed speech consisting of 55 sentences recorded by 125 Japanese learners of English, it was examined which sentence combinations out of the 55 sentences could maximize the correlation between automatic scores and human ratings. A veteran teacher selected 10 sentences based on criteria such as sentence length, grammar, pronunciation, and prosodic features. The shadowed speech of the 10 sentences were manually rated by two native speakers of English focusing on pronunciation, prosody and lexical access (whether shadowing is done adequately or not after each word or phrase is identified). The same shadowed utterances were automatically assessed by the DNN-based GOP procedure. A significantly high correlation (r=0.738, p<.01) was found between manual ratings and automatic scores. Then groups of sentences were selected out of the original 55 sentences by greedy search (full search) so that correlation between their DNN-GOP scores and manual ratings of the selected 10 sentences could be maximized, and the top-ranked five combinations of groups of sentences were listed up. The number of shared sentences between the 10 teacher-selected sentences and machine-suggested ones in the top-ranked five combinations was only one, which was much smaller than expected, and thus the teacher’s strategies in selecting stimulus sentences turned out to be inappropriate in terms of maximizing correlation. Examining the sentence sets suggested by the machine revealed that some particular sentences frequently appeared in the sets and seemed to have something to do with maximizing correlation. Hence, the present study compared the 10 sentences selected by the teacher and his selection criteria with the sentence sets suggested by the machine and discussed what kind of sentence should be chosen to improve reliability in automatic assessment.
Shadowing is a task where the subject is required to repeat the presented speech as s/he hears it. Although shadowing is cognitively a challenging task, it is considered as an efficient way of language training since it includes processes of listening, speaking and comprehension simultaneously. Our previous study realized automatic assessment of shadowing speech using the average of Goodness of Pronunciation (GOP) scores. But the fact that shadowing often includes broken utterances makes this approach insufficient. This study attempts to improve automatic assessment and, at the same time, give corrective feedbacks to learners based on error detection. We first manually labeled shadowing speech of 10 female and 10 male speakers and defined ten typical error types including word omission, substitution etc.. Forced alignment with adjusted grammar and GOP scores are adopted to detect word omission errors and poorly pronounced words. In the experiments, GOP scores, Word Recognition Rate (WRR), silence ratio, forced alignment log-likelihood scores, word omission rate are used to predict the overall proficiency of the individual speakers. The mean correlation coefficient between automatic scores and the speaker's TOEIC scores is 0.81, improved by 13% relatively. The detection accuracy of word omission is 73%.