NaijaSenti: A Nigerian Twitter Sentiment Corpus for Multilingual Sentiment Analysis.

Shamsuddeen Hassan Muhammad,David Ifeoluwa Adelani,Ibrahim Said Ahmad,Idris Abdulmumin,Bello Shehu Bello,Monojit Choudhury,Chris Chinenye Emezue,Anuoluwapo Aremu,Saheed Abdul,Pavel Brazdil

International Conference on Language Resources and Evaluation (LREC)（2022）

引用 38|浏览15

暂无评分

摘要

Sentiment analysis is one of the most widely studied applications in NLP, but most work focuses on languages with large amounts of data. We introduce the first large-scale human-annotated Twitter sentiment dataset for the four most widely spoken languages in Nigeria (Hausa, Igbo, Nigerian-Pidgin, and Yoruba) consisting of around 30,000 annotated tweets per language (except for Nigerian-Pidgin), including a significant fraction of code-mixed tweets. We propose text collection, filtering, processing, and labelling methods that enable us to create datasets for these low-resource languages. We evaluate a range of pre-trained models and transfer strategies on the dataset. We find that language-specific models and language-adaptive fine-tuning generally perform best. We release the datasets, trained models, sentiment lexicons, and code to incentivize research on sentiment analysis in under-represented languages.

查看译文

关键词

sentiment analysis, low-resource, twitter corpus, natural language processing

AI 理解论文

溯源树

样例

生成溯源树，研究论文发展脉络

Chat Paper

正在生成论文摘要