Adaptation of Language Resources and Tools for Closely Related Languages and Language Variants Proceedings of the Adaptation of Language Resources and Tools for Closely Related Languages and Language Variants Associated with the 9 Th International Conference on Recent Advances | AMiner
Adaptation of Language Resources and Tools for Closely Related Languages and Language Variants Proceedings of the Adaptation of Language Resources and Tools for Closely Related Languages and Language Variants Associated with the 9 Th International Conference on Recent Advances
In this paper we describe the construction of a parallel corpus between the standard and a non-standard language variety, specifically standard Austrian German and Viennese dialect. The resulting parallel corpus is used for statistical machine translation (SMT) from the standard to the non-standard variety. The main challenges to our task are data scarcity and the lack of an authoritative orthography. We started with the generation of a base corpus of manually transcribed and translated data from spoken text encoded in a specifically developed orthography. This data is used to train a first phrasebased SMT. To deal with out-of-vocabulary items we exploit the strong proximity between source and target variety with a backoff strategy that uses character-level models. To arrive at the necessary size for a corpus to be used for SMT, we employ a boot-strapping approach. Integrating additional available sources (comparable corpora, such as Wikipedia) necessitates to identify parallel sentences out of substantially differing parallel documents. As an additional task, the spelling of the texts has to be transformed into the above mentioned orthography of the target variety.