Arabic diacritization is essential for ensuring accurate pronunciation, clarity, and disambiguation of texts. It is a vital task in Arabic natural language processing. Despite substantial progress in the field, existing models struggle to generalize across the diverse forms of Arabic and perform poorly in noisy, error-prone environments. These limitations may be tied to problems in training data and, more critically, to insufficient contextual understanding. To address these gaps, we present SukounBERT.v2, a BERT-based Arabic diacritization system that is built using a multi-phase approach. We refine the Arabic Diacritization (AD) dataset by correcting spelling mistakes, introducing a line-splitting mechanism, and by injecting various forms of noise into the dataset, such as spelling errors, transliterated non-Arabic words, and nonsense tokens. Furthermore, we develop a context-aware training dataset that incorporates explicit diacritic markings and the diacritic naming of classical grammar treatises. Our work also introduces the Sukoun Corpus, a large-scale, diverse dataset comprising over 5.2 million lines and 71 million tokens that were sourced from Classical Arabic texts, Modern Standard Arabic writings, dictionaries, poetry, and purpose-built contextual sentences. Complementing this is a token-level mapping dictionary that enables minimal diacritization without sacrificing accuracy. This is a previously unreported feature in Arabic diacritization research. Trained on this enriched dataset, SukounBERT.v2 delivers state-of-the-art performance with over 55% relative reduction in Diacritic Error Rate (DER) and Word Error Rate (WER) compared to leading models. These results underscore the impact of context-aware and noise-resilient modeling in advancing the field of Arabic text processing.
Toxic language detection in Arabic remains an underexplored yet critical task, complicated by the language’s dialectal diversity, morphological richness, and the prevalence of overlapping toxicity types. In this paper, AraTox is presented as a large-scale, hand-annotated Arabic dataset designed for multi-label and multi-dialect toxicity classification. The dataset comprises around 36,700 social media comments spanning Gulf, Levantine, Nile Basin, North African, Yemeni, and Modern Standard Arabic (MSA) varieties, each manually labeled by expert annotators across seven categories, including hatred, cussing, racial, appearance-based, and sexual insults. The categories were derived from an iterative refinement of overlapping harassment types. AraTox is benchmarked using a range of deep learning models and classical classifiers, highlighting the superior performance of a stacked meta-learning ensemble over standalone models. The meta-classifier achieves a macro-averaged F1-score of 96
The Orani dialect of Arabic is an under-resourced Algerian language variety which also suffers from a lack of systematic evaluation of morphological analyzers. This study attempts to fill a critical resource gap for dialectal Arabic NLP, linguistic research, and educational applications. It presents MADOran, a morphologically annotated corpus for the Orani dialect of Arabic (ORN), together with a systematic evaluation of morphological analyzers across conventional, deep, and transformer-based approaches. The dataset contains 30,919 words drawn from: written texts (41
This paper introduces a new morphologically annotated dataset for the Orani Arabic dialect (ORN), comprising 30,919 words gathered from diverse genres, including written language (41 %) which cover topics such as culture, history, politics, recipes, social issues, tourism, and traditions, as well as spoken language (59 %) spanning stories, songs, TV scenes, proverbs, relationships and daily conversation such as college life, family, and relationships. Each word is manually annotated with a fine-grained tagset that provides part-of-speech, root, pattern, and English and French translations of the glosses. The morphological annotation was performed using the Dialectal Word Annotation Tool for Arabic (DIWAN) and it followed guidelines established for The Dynamic Arabella Corpus (Arabella), with adaptations to fit the dialectal context. MADOran, the Morphologically Annotated Dataset of Oran, presents a valuable resource for natural language processing applications that are focused on Arabic dialect processing, including dialect identification, machine translation, text generation, and speech recognition. It enables the training of language models for capturing the unique morphological and syntactic characteristics of ORN. Moreover, the dataset supports any linguistic research where features of ORN and Modern Standard Arabic (MSA) are compared and whenever the features unique to each variety are discussed. It also aids in developing dialect-specific lexicographic resources and in facilitating Arabic language learning. The annotated dataset adheres to scientific data stewardship principles of findability, accessibility, interoperability, and reusability, which qualify MADOran for use across domains.
This paper aims to explore the meanings and functions of the pragmatic marker, "I think", in 19th century English literature. Through a contrastive corpus analysis, we will investigate its uses in character development, narrativity, and the interplay between language and culture, using a corpus of fiction as a tool. The translation of "I think" into Arabic will provide valuable insight into how meanings transfer between different language cultures. The analysis reveals that "I think" has nuanced meanings distinct from logical markers like "but", and that these meanings are inherently pragmatic, relating to structural or modal discourse functions rather than propositional content. These meanings can be conceptualized as part of a radial category or semantic network, shedding light on their role in 19th-century English literature.
This paper introduces the Morphological and Syntactical analysis for the Quran text. In this research we have constructed the MASAQ dataset, a comprehensive resource designed to address the scarcity of annotated Quranic Arabic corpora and facilitate the development of advanced Natural Language Processing (NLP) models. The Quran, being a cornerstone of classical Arabic, presents unique challenges for NLP due to its sacred nature and complex linguistic features. MASAQ provides a detailed syntactic and morphological annotation of the entire Quranic text that includes more than 131K morphological entries and 123K instances of syntactic functions, covering a wide range of grammatical roles and relationships. MASAQ's unique features include a comprehensive tagset of 72 syntactic roles, detailed morphological analysis, and context-specific annotations. This dataset is particularly valuable for tasks such as dependency parsing, grammar checking, machine translation, and text summarization. The potential applications of MASAQ are vast, ranging from pedagogical uses in teaching Arabic grammar to developing sophisticated NLP tools. By providing a high-quality, syntactically annotated dataset, MASAQ aims to advance the field of Arabic NLP, enabling more accurate and more efficient language processing tools. The dataset is made available under the Creative Commons Attribution 3.0 License, ensuring compliance with ethical guidelines and respecting the integrity of the Quranic text.
This paper introduces the Morphologically-Analyzed and Syntactically-Annotated Quran (MASAQ) dataset, a comprehensive resource designed to address the scarcity of annotated Quranic Arabic corpora and facilitate the development of advanced Natural Language Processing (NLP) models. The Quran, being a cornerstone of classical Arabic, presents unique challenges for NLP due to its sacred nature and complex linguistic features. MASAQ provides a detailed syntactic and morphological annotation of the entire Quranic text, utilizing a rigorously verified text from Tanzil.net. The dataset includes more than 131K morphological entries and 123K instances of syntactic functions, covering a wide range of grammatical roles and relationships. The annotation process involved a team of expert Arabic linguists who employed traditional i'rab methodologies to ensure high accuracy and consistency. The dataset is structured in multiple formats (tab-separated text file (tsv), SQLite3 database (.db), comma-separated file (csv), and JavaScript Object Notation (.JSON)) to cater to various research needs. MASAQ's unique features include a comprehensive tagset of 72 syntactic roles, detailed morphological analysis, and context-specific annotations. This dataset is particularly valuable for tasks such as dependency parsing, grammar checking, machine translation, and text summarization. The potential applications of MASAQ are vast, ranging from pedagogical uses in teaching Arabic grammar to developing sophisticated NLP tools. By providing a high-quality, syntactically annotated dataset, MASAQ aims to advance the field of Arabic NLP, enabling more accurate and more efficient language processing tools. The dataset is made available under the Creative Commons Attribution 3.0 License, which governs its use and distribution. It has been created in compliance with ethical guidelines and with respect for the integrity of the Quranic text.
This paper investigates the extent to which Arabic punctuation is rule-governed, with the aim of improving text comprehension, disambiguation, and machine translation. The study highlights the lack of systematic punctuation in Arabic written discourse, which may be attributed to difficulties in sentence boundary identification or inadequate differentiation between various conjunctions. The punctuation behavior of Arabic speakers is examined in relation to sentence boundary identification and the level of agreement among Arabic specialists is assessed. A quantitative analysis of paragraph and sentence lengths across genres, categories of writers, and in comparison to English is conducted using five corpora specifically compiled for this study. Additionally, a punctuation survey is carried out to evaluate specialists’ agreement on sentence boundary identification. The results indicate that writers of Arabic interpret punctuation rules differently and that Arabic punctuation practice is irregular. The study suggests that standardization of Arabic punctuation rules is necessary to facilitate comprehension and automatic text processing.
This paper explores the development, design, and reconstruction of a Historical Arabic Corpus (HAC), which covers more than 1600 years of uninterrupted language use. The study emphasizes the technical aspects followed to enhance the system and provide a usable concordancer, along with simple experiments conducted on the corpus and the concordancer. Arabic has a rich literary and cultural heritage spanning thousands of years. The inclusion of digital resources and the advancement in natural language processing (NLP) technology have made Arabic historical corpora increasingly crucial for researchers and learners worldwide. By integrating HAC and its tools into Arabic language learning, learners can delve deeper into vocabulary and culture and gain valuable insights that improve their language skills and understanding of Arabic. This combination of human guidance and NLP technology makes learning an engaging and enjoyable experience, offering a dynamic and authentic way to master the Arabic language.
Arabic, unlike many languages, suffers from punctuation inconsistency, posing a significant obstacle for Natural Language Processing (NLP). To address this, we present the Arabic Punctuation Dataset (APD), a large collection of annotated Modern Standard Arabic texts designed to train machine learning models in sentence boundary identification and punctuation prediction. APD leverages the "theme-rheme completion" principle, a grammatical feature closely linked to consistent punctuation placement. It consists of an annotated collection of Modern Standard Arabic (MSA) texts that encompass 312 million words in approximately 12 million sentences. It comprises three diverse components: Arabic Book Chapters (ABC): Manually annotated, non-fiction, book excerpts, constituting a gold-standard reference. Complete Book Translations (CBT): Parallel English-Arabic book translations with aligned sentence endings, ideal for machine translation training. Scrambled Sentences from the Arabic Component of the United Nations Parallel Corpus (SSAC-UNPC): Jumbled sentences for model training in automatic punctuation restoration. Beyond NLP, APD serves as a valuable resource for linguistics research, language learning, and real-time subtitling. Its authentic, grammar-based approach can enhance the readability and clarity of machine-generated text, opening doors for various applications such as automatic speech recognition, text summarization, and machine translation.
This study examines the employment of persuasive strategies in informational emails that market products and/or services, illustrating how these strategies influence target customers and persuade them to make purchases. A corpus of 850 emails, encompassing over a million words, was compiled and analyzed using a mixed-method approach that integrated both quantitative and qualitative measures. The emails were collected between 2020 and 2021. The categorization of persuasive strategies was directed by predefined operational definitions and criteria, informed by Aristotle's model of persuasion. The analysis identified 11 persuasive strategies utilized within the email corpus. Notably, the findings revealed that the offering appeal and the appeal to authority are the most commonly used strategies, whereas the contrasting appeal and romantic expressions are the least employed. These results underscore the importance of persuasive strategies in business communication, especially within informational emails. The insights derived from this study carry significant implications for businesses in crafting compelling marketing messages. Furthermore, the findings contribute to English for Business Purposes courses, particularly in English as a Foreign Language (EFL) contexts, by offering guidance on constructing persuasive business emails.
In order to accurately represent the meaning and pronunciation of Arabic words and sentences, the presence of diacritics plays a crucial role. Over the years, researchers have dedicated significant efforts to enhancing automated diacritization systems. This paper introduces a novel approach for Arabic diacritization utilizing Bidirectional Encoder representations from Transformers (BERT) models. To evaluate the effectiveness of the proposed approach, two publicly available datasets, namely the Arabic Diacritization (AD) dataset and the Tashkeela Processed (TP) dataset, were employed. The performance of the models was assessed using various error metrics, including Diacritic Error Rate (DER) and Word Error Rate (WER). The findings demonstrate the superior performance of BERT in the diacritization process, surpassing all models employed in other diacritization systems. On the AD dataset, the proposed system achieved state-of-the-art (SOTA) syntactic DER and WER of 1.14% and 3.34%, respectively. For morphological diacritization, the best results yielded a DER of 0.92% and a WER of 1.91%. These outcomes reflect a remarkable relative error reduction of over 30% compared to previous research. Additionally, on the TP dataset, the BERT models exhibited a substantial decrease in DER, reducing the benchmark from 4.0% to 1.11%. Furthermore, this study introduces a real-time diacritization system called SUKOUN, which offers diacritized text through a user-friendly website. A comparison with existing automatic diacritization tools, using six example texts, reveals the superior prediction accuracy and preservation of input format provided by SUKOUN.
Conceptualizations of jogging and strolling in English and Arabic may represent a key element in the successful performance of such actions. We aimed to demonstrate similarities and differences between the English verbs 'jogging' and 'strolling' and their Arabic counterparts 'yuharwil'and 'yatanazzah'. As our theoretical framework, we adopted the Natural Semantic Metalanguage Approach to meaning and relied on Arabic and English semantic primes to create a semantic template consisting of Lexicosyntactic Frame, Prototypical scenario, Manner, and Prototypical outcome to explicate these verbs. To facilitate our comparative analysis of the English verbs and their Arabic counterparts, we consulted four English monolingual dictionaries: Merriam-Webster, Macmillan, Oxford learner's Dictionary, and Cambridge Dictionary, and consulted three Arabic monolingual dictionaries: Mu'jam maqayis al-lughah by Ahmad Ibn Faris al-Qazwini, Lisan al-'Arab by ibn Manzur, and Mu'jam al-lughah al-'arabiyah al-mu'asirah by Ahmed Mukhtar Omar. Our analysis revealed that: A) The English and the Arabic verbs relied on similar conceptualization elements including the use of legs and feet, having a starting point, a destination, as well as alternation, and repetition. B) 'jogging' was conceptualized as nonurgent and slower in comparison to 'yuharwil'. C) Duration of contact with the ground in 'strolling' and 'yatanazzah' was found to be similar. The same was also true for 'jogging' and 'yuharwil' whose duration of contact was shorter than 'strolling' and 'yatanazzah'. D) As for purpose, 'jogging', and 'strolling' were found to be motivated by a desire to exercise or relax while 'yatanazzah' was found to be motivated by entrainment only, and 'yuharwil' did not state a purpose.
This study investigated the use and functions of metadiscourse markers in English as a foreign language (EFL) virtual classroom during the Covid-19 pandemic. The study examined which metadiscourse markers- interactive or interactional-were used more frequently and how they were employed in an EFL context. It explored two interactive metadiscourse resources (code glosses and evidentials) and two interactional metadiscourse resources (attitude and engagement markers). The study utilized a mixed -method approach, using Hyland's (2004) two -componential taxonomy, to analyze a corpus of 303,148 words from 35 online lectures (90 minutes each) delivered by three university instructors in the UAE. The Mann -Whitney U test was employed to determine any significant differences in the use of these resources and their subcategories. The results revealed that the three instructors used more interactional than interactive resources. The qualitative analysis showed that code glosses and evidentials were primarily used to manage the flow of information, provide elaboration on propositional content, and provide evidence to support arguments. They were also employed to achieve cohesion and logical coherence in online classrooms. In contrast, attitude and engagement markers were used to engage students and signal the instructors' attitudes toward their material and audience. The study concludes with pedagogical implications for EFL instructors, students, and syllabus designers to foster social justice and fairness in the online learning environment, ensuring all students feel valued and empowered in their educational journey.
ABSTRACTPresidential speeches are among the best ways to demonstrate heads of state’s competence in dealing publically with domestic as well as foreign issues, and thus contribute to their political survival. This study uses a new one-million-word parallel corpus of Arabic and English to investigate the political topics that King Abdullah II of Jordan discusses when using the Arabic language to address Jordanian or Arab audiences and when using the English language to address Western/international audiences. The corpus covers the period from 1999 to 2015. Using Wordsmith 7 and examining the most frequent 25 Arabic and English words, we found that King Abdullah tends to discuss particular issues on all occasions, locally, nationally, and globally. These include Arab and regional matters as well as the peace process in the Middle East. It is also found that local themes that are mainly associated with Jordan, its political system, and economic and social challenges were more frequent in the Arabic corpus when compared to its English counterpart. However, Jordan’s involvement in some issues in the Middle East, especially in Palestine and Iraq, was more important in the English texts. The study concludes that parallel corpora can be a rich resource for discourse studies.KEYWORDS: Arabic-Englishparallel corpusaudience adaptationfrequency analysiscorpus linguisticsKing Abdullah II of Jordanpresidential speeches Disclosure statementNo potential conflict of interest was reported by the author(s).Correction StatementThis article has been corrected with minor changes. These changes do not impact the academic content of the article.Additional informationNotes on contributorsAhmad S. HaiderAhmad S. Haider is an associate professor in the Department of English Language and Translation at the Applied Science Private University, Amman, Jordan. He is also a researcher at the MEU Research Unit, Middle East University, Amman, Jordan. He received his Ph.D. in Linguistics from the University of Canterbury/New Zealand. His main areas of interest include corpus linguistics, discourse analysis, pragmatics, and translation studies.Alia AhmadAlia Ahmad received her MA in Linguistics from the University of Jordan/Jordan. Her current research focuses on Computational Linguistics, Corpus Linguistics, (Critical) Discourse Analysis, and Pragmatics.Sane YagiSane Yagi is currently a Professor of Linguistics at the University of Sharjah and the University of Jordan. received his education in Jordan, USA, and New Zealand. The primary themes are corpus development, computational lexicography and lexicology, computational morphology, syntactic parsing, automatic punctuation, and machine learning. His research interests include computational linguistics, CMC, CALL, and TEFL.Bassam H. HammoProf. Bassam H. Hammo earned his Ph.D. in Computer Science from DePaul University/USA in 2002 and his M.Sc. in Computer Science from Northeastern Illinois University in 1993. His research interests include Arabic Natural Language Processing and Machine Learning. He teaches Database Management Systems, Human-Computer Interaction, Machine Learning, and Data Mining.
This paper compares Arabic and English speech rhythms to increase awareness of this neglected and often misunderstood topic in foreign language acquisition. Unlike previous studies, we adopt a phonological view of speech rhythm rather than an isochrony-based phonetic view. We detail the components of speech rhythm at the word and utterance levels in Arabic and English focusing on the rhythmical differences that would affect the learners’ rhythm of both languages negatively. Findings suggest that Modern Standard Arabic (MSA) and Jordanian-Ammani Arabic (JAA), unlike English, should be placed at the lower end of the rhythmic continuum. The study opens new directions for future research and concludes with pedagogical implications for learners of Arabic and English.
Word embeddings, which represent words as numerical vectors in a high-dimensional space, are contextualized by generating a unique vector representation for each sense of a word based on the surrounding words and sentence structure. They are typically generated using such deep learning models as BERT and trained on large amounts of text data and using self-supervised learning techniques. Resulting embeddings are highly effective at capturing the nuances of language, and have been shown to significantly improve the performance of numerous NLP tasks. Word embeddings represent textual records of human thinking, with all the mental relations that we utilize to produce the succession of sentences that make up texts and discourses. Consequently, the distributed representation of words within embeddings ought to capture the reasoning relations that hold texts together. This paper makes its contribution to the field by proposing a benchmark for the assessment of contextualized word embeddings that probes into their capability for true contextualization by inspecting how well they capture resemblance, contrariety, comparability, identity, relations in time and space, causation, analogy, and sense disambiguation. The proposed metrics adopt a triangulation approach, so they use (1) Hume's reasoning relations, (2) standard analogy, and (3) sense disambiguation. The benchmark has been evaluated against 22 Arabic contextualized embeddings and has proven to be capable of quantifying their differential performance in terms of these reasoning relations. Results of evaluation of the target embeddings revealed that they do take context into account and that they do reasonably well in sense disambiguation but have weakness in their identification of converseness, synonymy, complementarity, and analogy. Results also show that size of an embedding has diminishing returns because the highly frequent language patterns swamp low frequency patterns. Furthermore, the suggest that future research endeavors should not be concerned with the quantity of data as much as its quality, and that it should focus more on the representativeness of data, and on model architecture, design, and training.
Missing punctuation makes texts difficult to understand and confusing to read, be it in formal or casual writing. What punctuation marking does is that it defines sentence constituents and sentence boundaries, which is critical to such Natural Language Processing (NLP) downstream tasks as machine translation, automatic speech analysis and synthesis. Although there is a rising amount of accessible Arabic data on the Web, the majority of it either doesn’t contain any punctuation or it uses punctuation inappropriately. To the best of our knowledge, very few people in the Arabic NLP community have considered this issue. This paper aims to present a method for automatically punctuating Arabic text using a model based on a pre-trained transformer. The goal of the Arabic punctuation prediction is to identify and insert the appropriate punctuation marks in punctuation-stripped Arabic texts. Using a robust dataset with a large number of punctuated sentences, we trained and tested the system. The experimental results show that AraBERT v0.2 (Antoun et al. in Arabert: transformer-based model for Arabic language understanding, [1]) and AraELECTRA discriminator (Antoun et al. in AraELECTRA: pre-training text discriminators for Arabic language understanding, [2]) achieved the highest accuracies, 95.88
This paper reports on the findings of a study that aimed at investigating the conceptual metaphors used in the Arabic subtitling of 150 English TV series (1982-2017), adopting Conceptual Metaphor Theory (CMT) proposed by Lakoff and Johnson (1980) for data analysis. The data were examined by using WordSmith Tools (Scot 2012) which is compatible with Arabic data. The study revealed that the most frequently used source domains in the corpus were journey, building, war, illness, plants, and machine, respectively; whereas, the least frequently used source domains were body parts, game, water, supernatural creatures, fabrics, fire, and light, respectively. Besides, the most commonly used type of conceptual metaphor is structural metaphor. The study concluded that the vast majority of metaphorical expressions are lexicalized and conventional to make the subtitling easily accessible to the reader. The study recommends that future studies be conducted on the translation strategies adopted in subtitling English metaphors into Arabic.
AbstractModelling the distributional semantics of such a morphologically rich language as Arabic needs to take into account its introflexive, fusional, and inflectional nature attributes that make up its combinatorial sequences and substitutional paradigms. To evaluate such word distributional models, the benchmarks that have been used thus far in Arabic have mimicked those in English. This paper reports on a benchmark that we designed to reflect linguistic patterns in both Contemporary Arabic and Classical Arabic, the first being a cover term for written and spoken Modern Standard Arabic, while the second for pre-modern Arabic. The analogy items we included in this benchmark are chosen in a transparent manner such that they would capture the major features of nouns and verbs; derivational and inflectional morphology; high-, middle-, and low-frequency patterns and lexical items; and morphosemantic, morphosyntactic, and semantic dimensions of the language. All categories included in this benchmark are carefully selected to ensure proper representation of the language. The benchmark consists of 45 roots of the trilateral, all-consonantal, and semivowel-inclusive types; six morphosemantic patterns (’af‘ala; ifta‘ala; infa‘ala; istaf‘ala; tafa‘‘ala; and tafā‘ala); five derivations (the verbal noun, active participle, and the contrasts in Masculine-Feminine; Feminine-Singular-Plural; Masculine-Singular-Plural); and morphosyntactic transformations (perfect and imperfect verbs conjugated for all pronouns); and lexical semantics (synonyms, antonyms, and hyponyms of nouns, verbs, and adjectives), as well as capital cities and currencies. All categories include an equal proportion of high-, medium-, and low-frequency items. For the purpose of validating the proposed benchmark, we developed a set of embedding models from different textual sources. Then, we tested them intrinsically using the proposed benchmark and extrinsically using two natural language processing tasks: Arabic Named Entity Recognition and Text Classification. The evaluation leads to the conclusion that the proposed benchmark is truly reflective of this morphologically rich language and discriminatory of word embeddings.