This paper presents a set of linguistic resources that formalizes the morphological behavior of simple Rromani adjectives. We describe the formalization of the adjectives' morphology and the implementation with the NooJ linguistic platform of an electronic dictionary associated with a formal morpho-syntactic grammar. We can then apply this set of resources to a corpus to evaluate the resources and automatically annotate adjectival forms in Rromani texts. The final set of resources can then be used to identify each Rromani dialectal variant and can be used as a pedagogical tool to teach Rromani as a second language.
The article emphasizes the critical importance of language generation today, particularly focusing on three key aspects: Multitasking, Multilinguality, and Multimodality, which are pivotal for the Natural Language Generation community. It delves into the activities conducted within the Multi3Generation COST Action (CA18231) and discusses current trends and future perspectives in language generation.
The article emphasizes the critical importance of language generation today, particularly focusing on three key aspects: Multitasking, Multilinguality, and Multimodality, which are pivotal for the Natural Language Generation community. It delves into the activities conducted within the Multi3Generation COST Action (CA18231) and discusses current trends and future perspectives in language generation.
To describe the infinite set of sentences expressed in a Natural language, one needs to define the finite set of its atomic units, i.e., its vocabulary, and the rules that combine these atomic units to construct sentences, i.e., its grammar. However, separating the vocabulary from the grammar is not straightforward; one crucial problem is defining multiword units. Here, I present three reproducible criteria to characterize them and thus separate the vocabulary from the grammar in an operational way.
Nowadays, most Natural Language Processing software applications use empirical "black box" methods associated with training corpora to analyze texts written in natural languages. To analyze a sequence of text, they look for similar sequences in a corpus, select among them the most similar one according to some statistical measurement or some neural-network-based optimization state, and then bring forth its analysis as the new sequence analysis. Here, I first show that the limited size of the corpora used and their questionable quality explain why most NLP applications produce unreliable results. Next, I examine the principles which are at the basis of corpus-based methods and uncover their linguistic naiveté. I finally dispute the scientific validity of empirical approaches. I propose solutions to various problems that are based on the use of carefully handcrafted linguistic methods and resources.
Bakground The linguistic pursuit of describing natural languages stands as a commendable scientific endeavor, regardless of immediate software application prospects. It transcends mere documentation of possible sentences to establish connections between sentences derived from transformations. Methods Amid the dominance of Large Language Models (LLMs) in research and technology, which offer intriguing advancements in text generation, the approaches presented in this article confront challenges like opacity, limited human intervention, and adaptation difficulties inherent in LLMs. The alternative or complementary approaches highlighted here focus on the theoretical and methodological challenges of describing linguistic transformations and are firmly rooted in the field of linguistics, the science of language. We propose two solutions to address the problem of language transformations: (i) the procedural approach, which involves representing each transformation with a transducer, and (ii) the declarative method, which entails capturing all potential transformations in a single neutral grammar. Results These approaches simplify the generation of complex sentences from elementary ones and vice versa. Conclusion This work has benefited from research exchanges within the Multi3Generation COST Action (CA18231), and the resources produced can contribute to enhancing any language generation system.
The purpose of this article is to highlight the critical importance of language generation today. In particular, language generation is explored from the following three aspects: multi-modality, multilinguality, which play crucial role for NLG community. We present the activities conducted within the Multi3Generation COST Action (CA18231), as well as current trends and future perspectives for multitask, multilingual and multimodal language generation.
En nuestro CETEHIPL1 creado en 2018, asumimos que en la reflexión metalingüística que se despliega en la interacción entre el aprendiente y una herramienta informática, se genera un conocimiento lingüístico sobre una lengua determinada y esto tiene consecuencias en el aprendizaje. En este punto, la plataforma NooJ creada por Max Silberztein (2015) (2016) nos brinda el sustento ideal para poder establecer comparaciones desde el momento en que cada módulo lingüístico se organiza desde bases comunes para poder procesar en forma automática los datos de cada lengua, en este caso el español y el francés. Pretendemos superar la fragmentación de saberes que separa a Lengua materna de Lengua extranjera, sobre todo en la educación secundaria. Esta fragmentación se instala dentro de un supuesto que intentamos deconstruir: el aprendizaje de lenguas sobre todo hace foco en el contenido de lo que se enseña. Nosotros en cambio, entendemos que enseñar no es solamente apropiarse de un contenido, sino poder establecer relaciones. Enseñar lengua por tanto se redimensiona en el plurilingüismo, es decir, en la aceptación de la diversidad. De tal modo, la comparación entre lenguas como estrategia metodológica permite extraer hipótesis que se validan en máquina y generan un conocimiento determinado. Así en este trabajo nos proponemos presentar al adjetivo en cuanto a su morfología tomando dos lenguas: español y francés. El abordaje automático nos permite mostrar rasgos característicos de ambas lenguas.
Le probleme pose est de savoir pourquoi l'on peut dire aller ou etre au salon ou a la cuisine (entre autres) mais non *aller ou *etre a la chambre ou au couloir, par exemple. L'hypothese avancee est que la construction est inacceptable lorsque le nom n'est pas lie au predicat approprie « action » ou « activite » et que le verbe ne suppose pas un sujet actif : on ne dit pas Je vais a la chambre car chambre est lie au predicat approprie etat (dormir, etre couche, etre alite) et que, de ce fait, le sujet n'est pas interpretable comme une personne decidant de se rendre en ce lieu pour y exercer l'activiteparticuliere qui lui serait prototypiquement liee.
Dans la tradition linguistique slave, les formes perfectives et imperfectives des verbes sont traditionnellement inscrites séparément dans les dictionnaires. Cependant, il existe de forts liens morphologiques et sémantiques entre les deux formes verbales. Nous présentons une formalisation qui nous a permis de lier les deux formes. Nous avons construit un dictionnaire électronique qui contient plus de 13 000 entrées verbales associées à plus de 300 paradigmes morphologiques, qui peut être utilisé pour automatiquement lemmatiser les formes verbales dans les textes ukrainiens et relier les formes perfectives et imperfectives.
Aujourd’hui, la plupart des applications logicielles du Traitement Automatique des Langues (analyse du discours, extraction d’information, moteurs de recherche, etc.) analysent les textes comme étant des séquences de formes graphiques. Mais les utilisateurs de ces logiciels cherchent typiquement des unités de sens : concepts, entités, relations dans leurs textes. Il faut donc établir une relation entre les formes graphiques apparaissant dans les textes et les unités de sens qu’elles représentent. Cette mise en relation nécessite des ressources et des méthodes de traitement linguistiques, que je présente ici.
Les chercheurs en sciences humaines et sociales utilisent des logiciels d'analyse de texte pour détecter et analyser des informations « intéressantes » dans leurs corpus : classer des concepts ou des entités, mettre à jour des prédicats ou des relations, trouver des oppositions entre thèmes, des similarités ou des collocations de termes, etc. Nous montrons comment les dictionnaires DEM et LVF de Dubois et Dubois-Charlier peuvent être utilisés pour établir une relation entre les formes graphiques présentes dans les textes et les unités de sens qui intéressent les utilisateurs.
L'analyse distributionnelle des mots relevant du champ de l'habitation, comme appartement, maison, studio, montre qu'ils sont munis, en discours, de connotations différentes : ils reflètent la manière, parfois surprenante, dont la société se représente leur référent, mais cette représentation est susceptible de varier selon le discours. La valeur d'un lexème ne peut donc être saisie qu'au terme de l'analyse de corpus de genres divers.
In the original version of the book, The given name and family name of the first editor was tagged incorrectly. It has been corrected.
NooJ is a linguistic development environment that provides tools for linguists to construct linguistic resources that formalise a large gamut of linguistic phenomena: typography, orthography, lexicons for simple words, multiword units and discontinuous expressions, inflectional and derivational morphology, local, structural and transformational syntax, and semantics. For each resource that linguists create, NooJ provides parsers that can apply it to any corpus of texts in order to extract examples or counterexamples , to annotate matching sequences, to perform statistical analyses, etc. NooJ also contains generators that can produce the texts that these linguistic resources describe, as well as a rich toolbox that allows linguists to construct, maintain, test, debug, accumulate and reuse linguistic resources. For each elementary linguistic phenomenon to be described, NooJ proposes a set of computational formalisms, the power of which ranges from very efficient finite-state automata to very powerful Turing machines. This makes NooJ's approach different from most other computational linguistic tools that typically offer a unique formalism to their users. Silberztein's article " NooJ computational devices " compares the different tools NooJ offers with the theoretical grammars described by Chomsky-Schützenberger's hierarchy. Since it was released in 2002, NooJ has been enhanced with new features every year. Linguists, researchers in Social Sciences and more generally all professionals who analyse texts have contributed to its development and participated in the annual NooJ conference. Since 2011, the European project Meta-Net CESAR brought a new interest in NooJ as well as a new set of projects, both in linguistics and in computer science. The present volume contains 18 articles selected from the 32 papers presented at the International NooJ 2012 Conference which was held from June 14th to 16th at the Institut NAtional des Langues et Civilisations Orientales (INALCO) in Paris. These articles are organised in three parts: " Vocabulary and Morphology " contains five articles; " Syntax and Semantics " contains six articles; " NooJ Applications " contains six articles. In this volume, we decided to add a new part: eight short papers that present prototype NooJ modules developed by graduate students and could serve as bases for more ambitious projects. Formalising Natural Languages with Nooj ix The articles in the first part involve the construction of dictionaries for simple words, multiword units as well as discontinuous expressions as well as the development of morphological grammars: —Thierry Declerck and Karlheinz Mörth's article " Porting Persian Lexical Resources to NooJ " …
Abdelmajid Ben Hamadou合作论文数Higher Institute of Computer Science and Multimedia, Sfax University2