Large speech corpora (LSC) constitute an indispensable resource for conducting research in speech processing and for developing real-life speech applications. In 2004 the Spoken Dutch Corpus (CGN) became available, a corpus of standard Dutch as spoken by adult natives in the Netherlands and Flanders. Owing to budget constraints, CGN does not include speech of children, non-natives, elderly people and recordings of speech produced in human-machine interactions. Since such recordings would be extremely useful for conducting research and for developing HLT applications for these specific groups of speakers of Dutch, a new project, JASMIN-CGN, was started which aims at extending CGN in different ways: by collecting a corpus of contemporary Dutch as spoken by children of different age groups, non-natives with different mother tongues and elderly people in the Netherlands and Flanders and, in addition, by collecting speech material in a communication setting that was not envisaged in CGN: human-machine interaction. We expect that the knowledge gathered from these data can be generalized to developing appropriate systems also for other speaker groups (i.e. adult natives). One third of the data will be collected in Flanders and two thirds in the Netherlands.
We explore learning prepositional-phrase attachment in Dutch, to use it as a filter in prosodic phrasing. From a syntactic treebank of spoken Dutch we extract instances of the attachment of prepositional phrases to either a governing verb or noun. Using cross-validated parameter and feature selection, we train two learning algorithms, IB 1 and RIPPER, on making this distinction, based on unigram and bigram lexical features and a cooccurrence feature derived from WWW counts. We optimize the learning on noun attachment, since in a second stage we use the attachment decision for blocking the incorrect placement of phrase boundaries before prepositional phrases attached to the preceding noun. On noun attachment, IB 1 attains an F-score of 82; RIPPER an F-score of 78. When used as a filter for prosodic phrasing, using attachment decisions from IB 1 yields the best improvement on precision (by six points to 71) on phrase boundary placement.
This paper describes the results of an evaluation of PROS-3, a system that assigns prosodic structure to text on the basis of the output of a syntactic parser. In order to evaluate the performance of PROS-3 as such and in combination with a revised algorithm for prosodic phrasing, we compare it to the prosodic structure as assigned by human experts. Also, the results of a perception experiment are presented, which show that listeners have the same preference of prosodic realization as we would expect on the basis of the comparison of the prosodic structures as assigned by PROS-3 and by human experts.
In this paper we describe a study in which a comparison was made between prosodic structures as realized in a spoken version of a text and as assigned by annotators of this text on paper. The prosodic structures were assigned by experts. This study puts to test the strategy of annotating text on paper to obtain a HUMAN reference of the prosodic structure that would be assigned when reading text aloud. This strategy is less time consuming than the often used analysis of spoken versions to obtain the assigned prosodic structure. The results of the comparison described here show that speakers are fairly capable of predicting what prosodic structure they would assign when reading text aloud.