Word segmentation plays an important role in speech recognition as a text pre-processing step that helps decrease out-of-vocabulary items and lowers language model perplexity. Segmentation is applied mainly for agglutinative languages, but other morphologically rich languages, such as German, can also benefit from this technique. Using a relatively small, manually collected broadcast corpus of 134k tokens, the current study investigates how Finite-State Transducers (FSTs) can be applied to perform word segmentation in German. It is shown that FSTs incorporating word-formation rules can reach high segmentation performance with 0.97 precision and 0.93 recall rate. It is also shown that FSTs incorporating n-gram models of manually segmented data can reach even higher performance with accuracy and recall rates of 0.97. This result is remarkable considering the fact that the bottom-up approach performs on par with the expert system without requiring explicit knowledge about morphological categories or word formation rules.
This work is an addition to the relatively short line of research concerning second language prosody perception. Using a prominence marking experiment, the study demonstrates that Japanese learners of English can perceptually discriminate between different focus scopes. Perceptual score profiles imply that narrowly focused words are identified and discriminated relatively easily, while differentiation of different scopes of broad focus presents a greater challenge. An analysis of a range of acoustic cues indicates that perceptual scores correlate most strongly with F0-based features. While this result is in contradiction with previous research results, it is shown that the divergence is attributable to the particular acoustic characteristics of the stimulus.
This study presents a psycholinguistically motivated evaluation method for phoneme classifiers by using non-categorical perceptual data elicited in a Japanese sibilant matching 2AFC task. Probability values of a perceptual [s]-[integral] boundary, obtained from 42 speakers over a 7-step synthetic [s]-[integral] continuum, were compared to probability estimates of Gaussian mixture models (GMMs) of Japanese [s] and [integral]. The GMMs, trained on the Corpus of Spontaneous Japanese, differed in feature vectors (MFCC, PLP, acoustic features), covariance matrix types (full, tied, diagonal, spherical), and numbers of mixtures (1-20). Using ten-fold cross validation, it was found that GMMs trained on MFCC features had the best sibilant classification accuracies (87.4-90.4%), but their correlations with human perceptual data were non-conclusive (0.35-0.98). Acoustic feature-based GMMs with tied covariance matrices had near human-like synthetic stimuli perception (0.957-0.996), but their classification performance was poor (71.3-80.4%). Models trained on perceptual linear prediction (PLP) features were on par with the acoustic feature-based models in terms correlation to the perceptual experiment (0.884-0.995), while losing slightly on classification performance (86.1-88.9%) compared to MFCC models. Across the board correlation tests and mixture-effect models confirmed that GMMs with better sibilant classifying performance produced more human-like probability estimations on the synthetic sibilant continuum.
State-of-the-art pronunciation tutoring (CAPT) systems are based on ASR technology. Consequently, they can provide a distinguished learning feedback which is focused on phonetic features and the positions of articulation errors. In contrast with the relative success with segmental errors, the acquisition and assessment of second language (L2) prosody is still a challenging problem. Although prosodic parameters like f0 contour or duration measures are usually displayed, the consequential evaluation components are generally missing. Considering the strong variation in speech data, functional data analysis (FDA) is a useful concept which statistically analyses interrelations between principal components (e.g., given accentuation) and their contribution to superimposed forms (e.g., resultingf0 contour). This article describes baseline processing and preliminary results of a pilot study on the intonation-based proficiency classification of German by using FDA methods. The experimental part contains the FDA-based classification results compared to a perceptual classification by German natives. Index Terms: L2 prosody, proficiency, functional data analysis
This study provides a direct comparison of boundary and prominence perception strategies between Japanese EFL learners and native speakers of English using the Rapid Prosody Transcription (RPT) method. Although RPT experiments are available for both native English speakers [1], [2], [3] and Japanese EFL learners [4], a direct comparison of the available data is problematic as the stimuli sets used in the experiments are not identical. The present research addresses this issue by using identical stimuli sets across L1 and L2 listeners. The data for native English speakers was taken from RPT experiments carried out by Jennifer Cole with Yoonsook Mo and colleagues [1-3, 5-8]. The non-native data was collected by re-running a subset of RPT tasks reported in the work by Cole and Mo with 108 Japanese undergraduate students. Although native speakers perceived more boundaries and prominent words than L2 speakers, the results outlined surprisingly similar perceptual strategies A strong correlation was found between the responses of native speakers and Japanese learners of English in both boundary and prominence perception tasks. In boundary perception both groups relied heavily on silent pauses and vocal fillers. In prominence detection the responses correlated with vowel duration, maximum amplitude, and maximum pitch in this specific order for both language groups.