Arabic Semantic Textual Similarity (STS) research has been limited by the lack of large-scale, high-quality evaluation resources. We address this gap with the Comprehensive Arabic Semantic Similarity (CASS) dataset, comprising 3,048 manually annotated sentence pairs with fine-grained similarity scores (0–5) spanning six semantic categories and 42 subcategories. CASS is four times larger than existing Arabic STS datasets and provides structured taxonomic coverage supporting systematic model evaluation. We validate CASS by evaluating 21 large language models, including commercial systems, large open-source models, and Arabic-specific systems. Fine-tuning on CASS achieves state-of-the-art performance, with Fanar 9B reaching 0.93 Spearman correlation, exceeding GPT-4o’s zero-shot baseline (0.90). Cross-dataset experiments demonstrate CASS’s superior generalization, with models trained on CASS outperforming those trained on smaller benchmarks by 2–10
更多
查看译文
关键词
Large Language Models,Arabic Natural Language Processing,Semantic Textual Similarity,Benchmark Datasets,Model Evaluation,Cross-lingual Transfer Learning,Arabic Language Resources