
In this paper, we present a hybrid approach for Word Sense Disambiguation of Arabic Language (called WSD-AL), that combines unsupervised and knowledge-based methods. Some pre-processing steps are applied to texts containing the ambiguous words in the corpus (1500 texts extracted from the web), and the salient words that affect the meaning of these words are extracted. After that a Context Matching algorithm is used, it returns a semantic coherence score corresponding to the context of use that is semantically closest to the original sentence. The contexts of use are generated using the glosses of the ambiguous word and the corpus. The results found by the proposed system are satisfactory; we have achieved a precision of 79%.
In this study, our goal was to find out how the verbs 幫忙 bāngmáng and 幫助 bāngzhù differ in usage. In particular, we aimed to determine the characteristics of the co-occurrences for the two verbs to investigate whether the verb 幫忙 bāngmáng typically takes an event as an object with positive or neutral collocates and whether the verb 幫助 bāngzhù typically takes a person/an organization as an object with negative collocates. Data for the two synonymous verbs were collected from the Chinese GigaWord Corpus using Chinese Word Sketch. As we predicted, our corpus data showed that 幫忙 bāngmáng typically took an event as an object and 幫助 bāngzhù typically took a person/an organization as an object, and that the verb 幫助 bāngzhù was more often linked with lexical items that had negative meanings. Based on the results presented, this study has practical implications for second language acquisition, translation, and dictionary compiling.
A multi-level analysis of the polysemous Mandarin near-synonymous pair PÀO and JÌN is undertaken in this work. We provide a step-by-step account of meanings, which includes analyses of morphemes, argument structures, sense distributions and a later discussion of their possible extensions, as well as a representation of event modules of both verbs. In incorporating the Module-Attribute Representation of Verbal Semantics (MARVS) proposed by Huang and colleagues (2000), it was discovered that PÀO has a wider semantic extension than JÌN does. Although the event structure of 'soak objects in liquid' was found for both verbs, PÀO was found to be a composite event of a dual process-state, with the event-internal attribute 'toward saturation' [−saturation], while JÌN was found to be a simplex event with an inchoative state, reflecting the event-internal attribute 'saturated' [+saturation]. In addition, PÀO collocates with the goal of 'hot spring' but JÌN does not. Focusing on the core sense of 'soak objects in liquid', we also discussed the possible semantic extensions of PÀO and JÌN and discovered some metaphorical extensions, such as Part-Whole metonymy, CONTAINER metaphors, and a MASS-COUNT image schema.
In this paper, we present a simple mining technique named the Quran Mining Technique (QMT) in an attempt to automatically classify the Suras (i.e. chapters) of the Quran based on predefined set of 10 themes. QMT is composed mainly of two phases: a preprocessing phase and a classification phase. In the first phase, we manually label a set of representative words for ten predefined themes. In the second phase we use the QMT on a set of 14 Suras (the total number of Suras is 30) using a scoring function (SF) to identify their themes. The results of QMT are compared with the results obtained from expert scholars in the field of Quranic studies, which we used as a benchmark. The average accuracy of the QMT classifier shows a result close to 79%.
Romanization is used to phonetically translate names and technical terms from languages in non-Roman alphabets to languages in Roman alphabets. Because almost all dictionaries contain standard English forms for some Arabic names, this problem has been solved using machine transliteration. Several programs exist to deal with transliteration; they are based either on dictionary-based approach or on rule-based approach. In this study, a comparison between these two approaches is shown. Test data from the Yarmouk University library were used. Results show that while a rule-based Romanizer can romanize all names, a dictionary-based Romanizer romanizes (86%) of tested names. On the other hand, another kind of test was performed over the Romanization rules used by each Romanizer; the results show that the Romanization rules (in terms of accuracy and usability) used by the Dictionary-based Romanizer used in this study are better than the ones used by Rule-based Romanizer.
This article focuses on the development of Natural Language Processing (NLP) tools for Computer Assisted Language Learning (CALL). First, we have developed some NLP tools: a labelled dictionary of Arabic (as complete as possible), a generator for morphological derivatives, a Conjugator and a morphological analyzer for Arabic. Second, we used these tools to create a number of educational applications for learning the Arabic language by using the proposed system SALA (an NLP-based authoring system, organized into three distinct layers: functions, scripts and activities).
Building large-scale semantic resource is one of the major tasks in Language Information Processing. We propose Feature Structure Theory, and apply this theory in building a large-scale Chinese semantic resource based on Penn Chinese Treebank corpus. The feature structure theory aims at addressing annotation problems from special sentence patterns, flexible word order, and serial noun phrase, etc., which are universal in Chinese. Annotation based on feature structure theory describes more semantic information than traditional approaches, and achieves higher annotating efficiency and higher accuracy.
Detecting and filtering e-mail alerts that are related to criminal or terrorist activities is of great interest for both security agencies and people. This paper evaluates and compares the performance of both the rule-based filter and Paul Graham statistical filter for detecting alerts in Arabic e-mail messages. To evaluate the two filters, a set of 1500 Arabic messages related to criminal activities were collected manually from some news websites such as Al-Jazeera Net and BBC Arabic news. The e-mails have been preprocessed, normalized, and then the relevant features were extracted from the collected e-mails by involving categorical proportional difference (CPD) and term frequency variance (TFV) as features weighting methods for the rule-based filter. To test the performance of the two filters, several experiments have been conducted and the result show that the Paul Graham statistical filter was more accurate. It was able to detect about 85% of the e-mail alerts used in the experiments. The rule-based filter has achieved 80% accuracy using the CPD method and 70% accuracy using the TFV method.
Module-Attribute Representation of Verbal Semantics (MARVS) is a theory of the representation of verbal semantics that is based on Mandarin Chinese data (Huang et al. 2000). In the MARVS theory, there are two different types of modules: Event Structure Modules and Role Modules. There are also two sets of attributes: Event-Internal Attributes and Role-Internal Attributes, which are linked to the Event Structure Module and the Role Module, respectively. In this study, we focus on four transitive verbs as chi1(eat), wan2(play), huan4(change) and shao1(burn) and explore their event structures by the MARVS theory.
Due to the lack of a rigorous methodology and explicit criteria to distinguish between classifiers (C) and measure words (M), previous inventories of Mandarin C's, or geti liangci 個體量詞, vary greatly. Based on the insight that an M in a Chinese [Num C/M N] phrase is semantically substantive, while a C is semantically redundant and thus does not block numeral quantification or adjectival modification to the noun, this paper further proposes that while C/M both function as a multiplicand mathematically, with Num as the multiplier, C's value is necessarily 1 and M is not, thus ~1. Cognitively, however, the semantically redundant C serves to profile an inherent semantic feature of N and thus selects a narrow class of N's. With these explicit distinctions between C and M, we then re-examine the inventory of C's put forth in 國語日報量詞典 Mandarin Daily News Dictionary of Measure Words and offer a much more reliable list of C's in Taiwan Mandarin.
In Buddhist Digital Archives, there are three core elements — lexicon, content and catalog that represent the knowledge of Buddhist Scriptures. However, the close relationship among these three core elements has not been explicitly and systematically highlighted. This paper aims to propose a framework for the integration of cross-language Buddhist Scriptures and traditional Buddhist taxonomic knowledge structure by applying current OntoLex (Ontological-Lexicon) techniques. In addition, an innovative attempt to import the concept of Ontology catalog in building cross-language Buddhist Tripitaka catalog in Chinese, Pali, Tibetan and Sanskrit is introduced. This paper starts with a portion of textual data from the CBETA, a comprehensive Chinese Buddhist Digital Archive. We believe that the ontological and lexical knowledge preserved in CBETA is the treasure to be discovered and explored. The mining of the large-scaled historical texts will not only enhance our understanding of Buddhist thought interpreted in varied temporal and geographical contexts, but also the interplay of human language and cognition in the diachronic multilingual contexts.
Out-of-vocabulary (OOV) terms, which do not exist in most dictionaries, usually cause failures in a cross language information retrieval (CLIR) system. Most existing approaches achieve a high performance when using web-mining to translate name entity type OOV terms. However, these methods gain a low performance when they are applied to medical OOV terms because they contain non-Chinese characters which are normally ignored by existing approaches, such as symbols, Roman alphabets and Arabic numbers. This paper presents a flexible rule-based approach towards the acquisition of medical OOV term translation. Our method uses a combination of a novel rule-based pattern extraction and brute force generation to identify the part of non-Chinese characters. To cope with the time-consuming task of ranking list and human extraction of OOV term translation, this paper presents a machine learning method to select correct translations automatically. In the method, twenty-one different features for each Chinese translation candidate are extracted, and the correct Chinese translations are selected by machine learning with our newly proposed statistics filter. By testing our method with 1,654 English ICD9 medical OOV terms, our proposed method (SF+F+W+B+P+S with the base machine learning algorithm SVM) outperforms the existing methods with a recall and precision value of 83.05% and 79.72%, respectively.
This paper explores the lexical semantic properties of five near-synonymous Chinese words expressing the emotion of SHAME. The concept of self-construal is vital in understanding emotions such as shame as it relies on the reflections of oneself. The interdependent self-construal is a view of the self through relationship with others and it is related to the characteristics of SHAME in Chinese context. The current study carried out an in-depth examination of how interdependent self-construal shapes the shame concept in Chinese, and how features concerning “self versus others” are encoded in Chinese shame words. The “self versus others” features that we look at include cause attribution (to self vs. others), probable relevant outcome (affect self vs. others), social relations (between self and others that cause shame), social norm (personal values vs. social norms), and presence (or absence) of audience. We examine whether and how these features play a crucial role in the Chinese SHAME concept and how they contribute to the differences between the shame words in Mandarin Chinese. The features can be described as a dimension that is implicit in the denotative meaning of these words.
Due to the rapid change of the real world, it is impossible for a Web map service to gather all recently appeared addresses in its spatial database, so as to provide an effective on-demand service. This paper proposes a novel method that takes advantage of Web mining to locate unregistered toponyms by utilizing address information available on the Internet. The design principle of this method consists of two key steps. First, multiple Web pages that contain the unregistered toponym are collected, and the recognized addresses, which already exist in the spatial database, are extracted from the Web pages. Second, the spatial relationship between each recognized address and the unregistered toponym is calculated to determine which recognized addresses have the highest weights, and accordingly locate the unregistered toponym on the map with high accuracy.
In this paper, we focus on the reliability of information encoded in a Web 2.0 community platform. Specifically, we aim to explore how linguistically encoded clues can contribute to the task. Given that evidentiality is the linguistic representation for the reliability of a statement, we propose to use evidential, the lexicalized evidentiality, to model text representation under the framework of machine learning based text classification. Based on the model, we conducted experiments to identify the best answers in Collaborative Question Answering (CQA) services. The experimental results confirm that, incorporating evidential for predicting text reliability is effective, since it shows the writer's self-judgment for information reliability. Moreover, our method can largely reduce the dimensionality of the vector space, and therefore provide an improvement in efficiency.
We introduce novel paraphrasing rules for Japanese light verb constructions (LVCs) that reduce the differences in the surface forms while retaining several of the crucial syntactic/semantic functions of these light verbs. An analysis of the linguistic properties of light verbs allows us to create paraphrasing patterns that map 151 different light verbs into 10 simple forms. Of these 10 forms, 7 convert complex noun-particle-verb structures into simple predicative forms. By constructing a list of 923 examples for ambiguous light verbs, we show that we can correctly distinguish real LVCs from those in which the light verbs were actually functioning as a main verb. The results of experiments indicate that our paraphrasing rules offer high accuracy. The experiments also reveal that our paraphrasing system works as a normalizer of complex predicates, which improves the recall rate of the predicate extraction task.
Analyzing the logical structure of a sentence is important for understanding natural language. In this paper, we present a task of Recognition of Requisite Part and Effectuation Part in Law Sentences, or RRE task for short, which is studied in research on Legal Engineering. The goal of this task is to recognize the structure of a law sentence. We investigate the RRE task regarding both the linguistic features and problem modeling aspects. We also propose solutions and present experimental results in a Japanese legal text domain. We got 88.58% with a supervised learning model and 88.84% with a semi-supervised learning model in the Fβ=1 score on the Japanese National Pension Law corpus.