Combining the Power of Large Language Models with Finetuning Based on Strategically Collected Human Ratings: A Case Study about Age-of-acquisition Estimates of Spanish Words | AMiner
Combining the Power of Large Language Models with Finetuning Based on Strategically Collected Human Ratings: A Case Study about Age-of-acquisition Estimates of Spanish Words
This study examined the ability of a large language model, GPT-4o-mini, to predict age of acquisition (AoA) for Spanish words, as compared to human ratings. We found a strong correlation (rho=.75) between the model's AoA estimates and mean human ratings. This correlation was lower than the level of agreement observed between individual human raters (rho=.85), but we found that finetuning the model on a relatively small dataset of 2000 human AoA ratings has the potential to enhance the model's performance to a level comparable to human consensus. Our analyses further indicated that in large data analyses we risk confounding AoA with word familiarity if we use AoA estimates for words unknown to participants. In line with existing psycholinguistic theorizing, it is better to limit the study of AoA effects to words that are within the participants' vocabulary. We present a novel dataset of AoA estimates for 28,453 Spanish words likely known by adult speakers.
更多
查看译文
关键词
Large language models,AI,age of acquisition,word recognition,Spanish language