Which Nigerian-Pidgin does Generative AI speak?: Issues about Representativeness and Bias for Multilingual and Low Resource Languages
arxiv(2024)
摘要
Naija is the Nigerian-Pidgin spoken by approx. 120M speakers in Nigeria and
it is a mixed language (e.g., English, Portuguese and Indigenous languages).
Although it has mainly been a spoken language until recently, there are
currently two written genres (BBC and Wikipedia) in Naija. Through statistical
analyses and Machine Translation experiments, we prove that these two genres do
not represent each other (i.e., there are linguistic differences in word order
and vocabulary) and Generative AI operates only based on Naija written in the
BBC genre. In other words, Naija written in Wikipedia genre is not represented
in Generative AI.
更多查看译文
AI 理解论文
溯源树
样例
生成溯源树,研究论文发展脉络
Chat Paper
正在生成论文摘要