Towards better structured and less noisy Web data: Oscar with Register annotations.

Veronika Laippala,Anna Salmela,Samuel Rönnqvist,Alham Fikri Aji,Li-Hsin Chang, Asma Dhifallah,Larissa Goulart, Henna Kortelainen,Marc Pàmies,Deise Prina Dutra,Valtteri Skantsi,Lintang Sutawika,Sampo Pyysalo

International Conference on Computational Linguistics（2022）

引用 0|浏览54

暂无评分

摘要

Web-crawled datasets are known to be noisy, as they feature a wide range of language use covering both user-generated and professionally edited content as well as noise originating from the crawling process. This article presents one solution to reduce this noise by using automatic register (genre) identification -whether the texts are, e.g., forum discussions, lyrical or how-to pages. We apply the multilingual register identification model by Rönnqvist et al. (2021) and label the widely used Oscar dataset. Additionally, we evaluate the model against eight new languages, showing that the performance is comparable to previous findings on a restricted set of languages. Finally, we present and apply a machine learning method for further cleaning text files originating from Web crawls from remains of boilerplate and other elements not belonging to the main text of the Web page. The register labeled and cleaned dataset covers 351 million documents in 14 languages and is available at https://huggingface.co/datasets/TurkuNLP/register_oscar.

查看译文

AI 理解论文

溯源树

样例

生成溯源树，研究论文发展脉络

Chat Paper

正在生成论文摘要