Most question answering systems have been designed to answer short questions (straight answers such as dates, locations), but there are only a few pieces of research about complex questions. In this paper, we present a method for analyzing complex questions at the syntactic and the morphological levels with a pattern-based structure. These linguistic patterns allow us to annotate the question and its semantic features for extracting the focus and topic. We start with the implementation of the rules which identify and annotate the various medical named entities. Our Named Entity Recognizer (NER) is able to sift through references to people, places and organizations, diseases, viruses, as targets to extract the correct answer from the user. The NER is embedded in our question answering system. The task of our system is divided in four phases each of which plays a critical role in the general performance: question analysis, segmentation and passage retrieval, answer validation and, finally, answer extraction. We use the NooJ platform which represents a valuable linguistic development environment. The first evaluations show that the actual results are encouraging and could be deployed for further question types.
Language resources are a necessary component to language Development in NLP. They are useful for any empirical language study including linguistic analysis, language translation and language disambiguation. The linguistic development environment NooJ ( http://www.nooj4nlp.net/ ) allow formalizing complex linguistic phenomena such as compound words generation, processing as well as analysis. NooJ offers the possibility to use the dynamic library NoojEngine.dll or the command-line program: noojapply.exe. In this study, we will take advantage of the noojapply.exe program that is freely available in the Standard edition of NooJ. Noojapply.exe allows users to apply dictionaries and grammars automatically to texts from external environments. In this paper, we introduce a module for Arabic MWEs recognition that is based on rules grammar. MWEs module allows recognizing several types of morphosyntactic variations that can occur to a Multi Word Expression. Then, these linguistic resources are compiled to be used as parameters in the command-line noojapply.exe in order to be integrated within an Arabic language processing environment for linguistic disambiguation. Our work is divided into three sections. First, we deal with a literature review on disambiguation tasks in the Arabic language. Then, we give a detailed description of our Integrated NooJ environment for Arabic linguistic disambiguation and the associated grammars. Finally, a set of tests and experiments is carried out to measure the impact of multi- word expression recognition in Word disambiguation.
Nowadays, most question-answering systems have been designed to answer factoid or binary questions (looking for short and precise answers such as dates, locations), however little research has been carried out to study complex questions. In this paper, we present a method for analyzing medical opinion questions. The analysis of the question asked by the user by means of a pattern based analysis covered the syntactic as well as the morphological levels. These linguistic patterns allow us to annotate the question and the semantic features of the question by means of extracting the focus and topic of the question. We start with the implementation of the identifying rules and the annotation of the various medical named entities. Our named entity recognizer tool (NER) is able to find references to people, places and organizations, diseases, viruses, as targets to extract the correct answer from the user. The NER is embedded in our question answering system. The task of QA is divided into four phases: question analysis, segmentation, and passage retrieval & answer extraction. Each phase plays a crucial role in the overall performance. We use the NooJ platform which represents a valuable linguistic development environment. The first evaluations show that the actual results are encouraging and could be deployed for further question types.
Nowadays, the medical domain has a high volume of electronic documents. The exploitation of this large quantity of data makes the search of specific information complex and time consuming. This difficulty has prompted the development of new adapted research tools, as question-answering systems. Indeed, this type of system allows a user to ask a question in natural language and automatically identify a specific answer instead of a set of documents deemed pertinent, as is the case with search engines. For this purpose, we are developing a question answering system which is based on a linguistic approach. The use of the linguistic engine of NooJ in order to formalize the automatic recognition rules and then applying them to a dynamic corpus composed of arabic medical journalistic articles. In this paper, we present a method for analyzing medical Binary questions. The analysis of the question asked by the user by means the application of cascade of morpho-syntactic resources. The linguistic patterns (grammars) which allow us to annotate the question and the semantic features of the question of extracting the focus and topic of the question. We start with the implementation of the rules which identify and to annotate the various medical entities. The named entity recognizer (NER) is able to find references to people, places and organizations, diseases, viruses, as targets to extract the correct answer from the user. The NER is embedded in our question answering system in order to identify the answer and delimit the potential justification sequence the precision and recall show that the actual results are encouraging and could be integrated for more types of questions other than binary questions.
Document classification is a necessary task for most Natural Language Processing tools since it classifies documents content in a helpful and meaningful way. The main concern in this paper is to investigate the impact of using multi-words for text representation on the performances of text classification task. Two text classification strategies are proposed to observe the robustness of each of them. First, we will deal with the literature review of existing linguistic resources in Arabic language. Secondly, we will present a classification method that is based on domain candidate simple terms. These terms are automatically extracted from multiple specialized corpora depending on their appearance frequency. Then, we will present a detailed description of a classification method based on multi-word expressions dictionary. CompounDic, an Arabic multi-word expressions dictionary, will be used to automatically annotate multi-word expressions and variations in text. Finally, we carried out a series of experiments on classifying specialized text based on simple words and multi-word expressions for comparison purposes. Our experiments show that the use of multi-word expressions annotations enhances the text classification results.
This works deals with Arabic factoid Question Answering systems (QA). Commonly, the task of QA is divided into three phases: question analysis, answer pattern generation, and answer extraction. Each phase plays a crucial role in overall performance. In this paper, we focus on the two first phases: Question Analysis and Answer Pattern Generation. We used the NooJ platform which represents a valuable linguistic development environment. The first evaluations show that the actual results are encouraging and could be deployed for more types of questions other than factoid ones.
Due to their wide popularity and easy access to the published contents, social media such as Facebook and Twitter have attracted the interest of media to disseminate their information (information, news, events …). Nowadays, we are witnessing a much-accelerated rhythm of events shared on social media. These events are covering several topics like politics (e.g. presidential elections), epidemics (Zika virus), terrorism (DAECH attacks) or economy (Stock market). Given the importance and relevance of the shared information, several methods and tools are developed to detect and display information from social networks especially Twitter such as MABED1 [1], Twitter Monitor [2] and 3KeySEE [3].
Since the continuous proliferation of the journalistic content online and the changing political landscape in many Arabic countries, we started our current research in order to implement a media monitoring system about the opinion mining in political field. This system allows political actors, despite of the large volume of online data, to be constantly informed about opinions expressed on the web in order to properly monitor their actual standing, orient their communication strategy and prepare the election campaigns. The developed system is based on a linguistic approach using NooJ's linguistic engine to formalize the automatic recognition rules and apply them to a dynamic corpus composed of journalistic articles. The first implemented rules allow identifying and annotating the different political entities (political actors and organizations). Then these annotations are used in our system of media monitoring in order to identify the opinions associated with the extracted named entities. The system is mainly based on a set of local grammars developed for the identification of different structures of the political opinion phrases. These grammars are using the entries of the opinion lexicon that contain the different opinion words (verbs, adjectives, nouns) where each entry is associated with the corresponding semantic marker (polarity and intensity). Our developed system is able to identify and properly annotate the opinion holder, the opinion target and the polarity (positive or negative) of the phraseological expression (nominal or verbal) expressing the opinion. Our experiments showed that the adopted method of extraction is consistent with 0.83 F-measure.
CompounDic is an Arabic MWEs dictionary that lists many entries, divided into more than 20 domains. It lists only MWEs in their base form. With regard to syntactic and morphological flexibility, the lexicon covers 2 types of MWEs: Fixed MWEs (no variation allowed) and semi-fixed MWEs (variation in their structural pattern). Arabic presents distinctive features to deal with MWEs processing. A lot of possible derivations are possible (plural or dual forms, multiple irregular plurals). In addition, we need to process agglutination forms. In this paper, we will study the structural variability of semi-fixed multiword expressions in Arabic language in order to recognize the morphological and inflectional variations. We will adopt a recognition approach based on the use of a cascade of local grammars. The recognition system is based on NooJ’s local grammars as well as an Arabic MWEs dictionary covering more than 20 domains. The inflectional and derivational rules, which concern semi-fixed MWEs, use some specific morphological operators that will be described as well. Finally, we present new results showing the experimentation scores of morpho-lexical coverage enhancement.
NooJ is a linguistic development environment that allows formalizing complex linguistic phenomena such as compound words generation, processing as well as analysis. We will take advantage of NooJ's linguistic engine strength in order to create a new large coverage terminological compound word's dictionary for Modern Standard Arabic language. Classifying and annotating Arabic compound words would have a major impact on the disambiguation of applications working with Arabic texts. The diverse analyzers, based on morphological aspect, are not able to recognize multiword expressions. Morphological analyzers usually separate compound expressions into single terms. Therefore recognizing the entire compound words is essential to preserve the semantic of texts and to provide a crucial resource for a better analysis and understanding of Arabic language.Our work is composed of three sections. First, we will deal with a literature review on Arabic compound expression's categories which aims to dress a detailed topology. The structural variability of multiword expressions in Arabic language will be studied in order to measure the degree of morphological, lexical and grammatical flexibility of multiword expressions. Then, we will discuss the electronic thematic dictionary of compound Arabic expressions and give detailed description of our methodology and guidelines.
Since 2006 we have undertaken to describe the differences between 17th century English and contemporary English thanks to NLP software. Studying a corpus spanning the whole century (tales of English travellers in the Ottoman Empire in the 17th century, Mary Astell's essay A Serious Proposal to the Ladies and other literary texts) has enabled us to highlight various lexical, morphological or grammatical singularities. Thanks to the NooJ linguistic platform, we created dictionaries indexing the lexical variants and their transcription in CE. The latter is often the result of the validation of forms recognized dynamically by morphological graphs. We also built syntactical graphs aimed at transcribing certain archaic forms in contemporary English. Our previous research implied a succession of elementary steps alternating textual analysis and result validation. We managed to provide examples of transcriptions, but we have not created a global tool for automatic transcription. Therefore we need to focus on the results we have obtained so far, study the conditions for creating such a tool, and analyze possible difficulties. In this paper, we will be discussing the technical and linguistic aspects we have not yet covered in our previous work. We are using the results of previous research and proposing a transcription method for words or sequences identified as archaic.
This paper presents a cascade of morpho-syntactic tools to deal with Arabic natural language processing. It begins with the description of a large coverage formalization of the Arabic lexicon. The built electronic dictionary, named "El-DicAr", which stands for "Electronic Dictionary for Arabic", links inflectional, morphological, and syntactic-semantic information to the list of lemmas. Automated inflectional and derivational routines are applied to each lemma producing over 3 million inflected forms. El-DicAr represents the linguistic engine for the automatic analyzer, built through a lexical analysis module, and a cascade of morpho-syntactic tools including: a morphological analyzer, a spell-checker, a named entity recognition tool, an automatic annotator and tools for linguistic research and contextual exploration. The morphological analyzer identifies the component morphemes of the agglutinative forms using large coverage morphological grammars. The spell-checker corrects the most frequent typographical errors. The lexical analysis module handles the different vocalization statements in Arabic written texts. Finally, the named entity recognition tool is based on a combination of the morphological analysis results and a set of rules represented as local grammars.
In this paper, we propose an Arabic Question-Answering (Q-A) system called QASAL (Question-Answering system for Arabic Language). QASAL accepts as an input a natural language question written in Modern Standard Arabic (MSA) and generates as an output the most efficient and appropriate answer. The proposed system is composed of three modules: A question analysis module, a passage retrieval module and an answer extraction module. To process these three modules we use the NooJ Platform which represents a linguistic development environment.
In this paper, we describe the use of an incremental construction method of minimal, acyclic, deterministic FST. The approach consists in constructing a transducer in a single step by adding new strings one by one and minimizing the resultant automaton incrementally. Then, we present a new method to encode the morphological information associated with the dictionary entries. The new encoding unifies a large number of word forms' analyses, thus reducing the number of terminal states of the dictionary's FST, that triggers a more efficient minimization process. Finally, we present experimental results on the FST that represents the Arabic dictionary.
La langue arabe, bien que tres importante par son nombre de locuteurs, elle presente des phenomenes morpho-syntaxiques tres particuliers. Cette particularite est liee principalement a sa morphologie flexionnelle et agglutinante, a l’absence des voyelles dans les textes ecrits courants, et a la multiplicite de ses formes, et cela induit une forte ambiguite lexicale et syntaxique. Il s'ensuit des difficultes de traitement automatique qui sont considerables. Le choix d'un environnement linguistique fournissant des outils puissants et la possibilite d'ameliorer les performances selon nos besoins specifiques nous ont conduit a utiliser la plateforme linguistique NooJ. Nous commencons par une etude suivie d’une formalisation a large couverture du vocabulaire de l’arabe. Le lexique construit, nomme «El-DicAr», permet de rattacher l’ensemble des informations flexionnelles, morphologiques, syntactico-semantiques a la liste des lemmes. Les routines de flexion et derivation automatique a partir de cette liste produisent plus de 3 millions de formes flechies. Nous proposons un nouveau compilateur de machines a etats finis en vue de pouvoir stocker la liste generee de facon optimale par le biais d’un algorithme de minimisation sequentielle et d’une routine de compression dynamique des informations stockees. Ce dictionnaire joue le role de moteur linguistique pour l’analyseur morpho-syntaxique automatique que nous avons implante. Cet analyseur inclut un ensemble d’outils: un analyseur morphologique pour le decoupage des formes agglutinees en morphemes a l’aide de grammaires morphologiques a large couverture, un nouvel algorithme de parcours des transducteurs a etats finis afin de traiter les textes ecrits en arabe independamment de leurs etats de voyellation, un correcteur des erreurs typographiques les plus frequentes, un outil de reconnaissance des entites nommees fonde sur une combinaison des resultats de l’analyse morphologique et de regles decrites dans des grammaires locales presentees sous forme de reseaux augmentes de transitions (ATNs), ainsi qu’un annotateur automatique et des outils pour la recherche linguistique et l’exploration contextuelle. Dans le but de mettre notre travail a la disposition de la communaute scientifique, nous avons developpe un service de concordances en ligne «NooJ4Web: NooJ pour la Toile» permettant de fournir des resultats instantanes a differents types de requetes et d’afficher des rapports statistiques ainsi que les histogrammes correspondants. Les services ci-dessus cites sont offerts afin de recueillir les reactions des divers usagers en vue d’une amelioration des performances. Ce systeme est utilisable aussi bien pour traiter l’arabe, que le francais et l’anglais
Named entities (NE) occur frequently in Arabic texts, and their recognition is essential. Recognizing and categorizing NE requires both internal (morphological) and external (syntactic) evidences. This paper describes a system that combines a morphological parser and a syntactic parser, that are built with the NooJ linguistic development environment.
Cet article décrit un système de construction du lexique et d’analyse morphologique pour l’arabe standard. Ce système profite des apports des modèles à états finis au sein de l’environnement linguistique de développement NooJ pour traiter aussi bien les textes voyellés que les textes partiellement ou non voyellés. Il se base sur une analyse morphologique faisant appel à des règles grammaticales à large couverture.
From the beginning of the sixties, and starting with the first automatic analyzer proposed by David Cohen, one of the first theorists of NLP [1], research has continued with natural language processing and especially the automatic treatment of the Arabic language. In 1983, with a minimalist morphological analysis, based on the theory that any Arabic form is generated using root and pattern, researchers developed the first twolevel morphological analyzer for Arabic (Koskenniemi 1983); this work was included within the project ALPNET (Beesley and Buckwalter 1989) using finite-state technology allowing only the concatenation of morphemes in the morphotactics. Since 1996, the Xerox research centre has enhanced this system using an algorithm of automatic combination between roots and patterns generating stems; this research is based on the ALPNET’s dictionaries which were, considerably rebuilt using the Xerox finite-state technology (Beesley 2001). This technology is computationally very efficient for natural-language-processing; it’s used within the developmental environment NooJ (Silberztein 2006). The use of finite-state machines within NooJ was extremely attractive, they are used to generate and analyse several thousands of words per second. This linguistic platform will be described inside this paper as the tool used for vocabulary formalization and analysis of standard Arabic language.
This article describes the construction of a lexicon and a morphological description for standard Arabic. This system uses finite state technology to parse vowelled texts, as well as partially and not vowelled ones. It is based on large-coverage morphological grammars covering all grammatical rules.
Abdelmajid Ben Hamadou合作论文数Higher Institute of Computer Science and Multimedia, Sfax University1