
In this paper, we present a hybrid approach for Word Sense Disambiguation of Arabic Language (called WSD-AL), that combines unsupervised and knowledge-based methods. Some pre-processing steps are applied to texts containing the ambiguous words in the corpus (1500 texts extracted from the web), and the salient words that affect the meaning of these words are extracted. After that a Context Matching algorithm is used, it returns a semantic coherence score corresponding to the context of use that is semantically closest to the original sentence. The contexts of use are generated using the glosses of the ambiguous word and the corpus. The results found by the proposed system are satisfactory; we have achieved a precision of 79%.
In this study, our goal was to find out how the verbs 幫忙 bāngmáng and 幫助 bāngzhù differ in usage. In particular, we aimed to determine the characteristics of the co-occurrences for the two verbs to investigate whether the verb 幫忙 bāngmáng typically takes an event as an object with positive or neutral collocates and whether the verb 幫助 bāngzhù typically takes a person/an organization as an object with negative collocates. Data for the two synonymous verbs were collected from the Chinese GigaWord Corpus using Chinese Word Sketch. As we predicted, our corpus data showed that 幫忙 bāngmáng typically took an event as an object and 幫助 bāngzhù typically took a person/an organization as an object, and that the verb 幫助 bāngzhù was more often linked with lexical items that had negative meanings. Based on the results presented, this study has practical implications for second language acquisition, translation, and dictionary compiling.
A multi-level analysis of the polysemous Mandarin near-synonymous pair PÀO and JÌN is undertaken in this work. We provide a step-by-step account of meanings, which includes analyses of morphemes, argument structures, sense distributions and a later discussion of their possible extensions, as well as a representation of event modules of both verbs. In incorporating the Module-Attribute Representation of Verbal Semantics (MARVS) proposed by Huang and colleagues (2000), it was discovered that PÀO has a wider semantic extension than JÌN does. Although the event structure of 'soak objects in liquid' was found for both verbs, PÀO was found to be a composite event of a dual process-state, with the event-internal attribute 'toward saturation' [−saturation], while JÌN was found to be a simplex event with an inchoative state, reflecting the event-internal attribute 'saturated' [+saturation]. In addition, PÀO collocates with the goal of 'hot spring' but JÌN does not. Focusing on the core sense of 'soak objects in liquid', we also discussed the possible semantic extensions of PÀO and JÌN and discovered some metaphorical extensions, such as Part-Whole metonymy, CONTAINER metaphors, and a MASS-COUNT image schema.
In this paper, we present a simple mining technique named the Quran Mining Technique (QMT) in an attempt to automatically classify the Suras (i.e. chapters) of the Quran based on predefined set of 10 themes. QMT is composed mainly of two phases: a preprocessing phase and a classification phase. In the first phase, we manually label a set of representative words for ten predefined themes. In the second phase we use the QMT on a set of 14 Suras (the total number of Suras is 30) using a scoring function (SF) to identify their themes. The results of QMT are compared with the results obtained from expert scholars in the field of Quranic studies, which we used as a benchmark. The average accuracy of the QMT classifier shows a result close to 79%.
Romanization is used to phonetically translate names and technical terms from languages in non-Roman alphabets to languages in Roman alphabets. Because almost all dictionaries contain standard English forms for some Arabic names, this problem has been solved using machine transliteration. Several programs exist to deal with transliteration; they are based either on dictionary-based approach or on rule-based approach. In this study, a comparison between these two approaches is shown. Test data from the Yarmouk University library were used. Results show that while a rule-based Romanizer can romanize all names, a dictionary-based Romanizer romanizes (86%) of tested names. On the other hand, another kind of test was performed over the Romanization rules used by each Romanizer; the results show that the Romanization rules (in terms of accuracy and usability) used by the Dictionary-based Romanizer used in this study are better than the ones used by Rule-based Romanizer.