Lexical substitution task requires to substitute a target word by candidates in a given context. Candidates must keep meaning and grammatically of the sentence. The task, introduced in the SemEval 2007, has two objectives. The first objective is to find a list of substitutes for a target word. This list of substitutes can be obtained with lexical resources like WordNet or generated with a pre-trained language model. The second objective is to rank these substitutes using the context of the sentence. Most of the methods use vector space models or more recently embeddings to rank substitutes. Embedding methods use high contextualized representation. This representation can be over contextualized and in this way overlook good substitute candidates which are more similar on non-contextualized layers. SemDis 2014 introduced the lexical substitution task in French. We propose an application of the state-of-the-art method based on BERT in French and a novel method using contextualized and non-contextualized layers to increase the suggestion of words having a lower probability in a given context but that are more semantically similar. Experiments show our method increases the BERT based system on the OOT measure but decreases on the BEST measure in the SemDis 2014 benchmark.
L'Open Data fournit de nombreuses données publiques avec une couverture très large, mais aucune base n'a jamais été structurée à partir d'informations issues de l'actualité.À travers DataNews, notre objectif est d'aller chercher automatiquement des données afin d'offrir un moyen de les réutiliser.Pour ce faire, nous avons tout d'abord défini une typologie d'événements dans le contexte spécifique des décès dans des dépêches AFP.Puis, en se limitant aux catastrophes naturelles, nous avons regroupé ces dépêches par événement afin de pouvoir les identifier.La dernière étape a pour objectif de construire des patrons d'extraction afin de collecter les valeurs correspondant au nombre de morts, de même que le contexte associé à ces valeurs.Les résultats de nos évaluations nous ont confirmé le fort potentiel de notre méthode qui pourrait amener à l'élaboration de plusieurs applications.ABSTRACT.The Open Data allows the access to plentiful data, with a large coverage, but none of them offers a structured databased around news.Through DataNews, our goal is to seek for data automatically so as to provide means to reuse them.To do so, we first defined an event typology in the specific context of death in AFP wires.Then, by restraining ourselves to the natural disasters, we clustered these wires by events so as to identify them.The goal of the last step is to build extraction patterns so as to collect values corresponding to the death number, as well as the context associated to these values.The results of our evaluations reassured ourselves in the large potential of our method that could lead to several applications.MOTS-CLÉS.base de connaissances, extraction d'information, construction de patrons, détection d'événements.
Depuis plusieurs annees, Syllabs integre de nombreux composants au sein d’un agrefilter, utilisant des technologies d’extraction d’information developpees en interne et dans un contexte multilingue. Originellement concu pour agreger des contenus issus de la presse, SylNews peut etre utilise a des fins de veille, pour explorer des contenus, ou pour identifier d’une maniere plus globale les sujets chauds de l’ensemble ou d’une partie des contenus stockes.
Nous présentons la participation de Syllabs à la tâche de classification de tweets dans le domaine du transport lors de DEFT 2018. Pour cette première participation à une campagne DEFT, nous avons choisi de tester plusieurs algorithmes de classification état de l’art. Après une étape de prétraitement commune à l’ensemble des algorithmes, nous effectuons un apprentissage sur le seul contenu des tweets. Les résultats étant somme toute assez proches, nous effectuons un vote majoritaire sur les trois algorithmes ayant obtenus les meilleurs résultats.
Cet article présente une méthode permettant de collecter sur le web des informations complémentaires à une information prédéfinie, afin de remplir une base de connaissances. Notre méthode utilise des patrons lexico-syntaxiques, servant à la fois de requêtes de recherche et de patrons d’extraction permettant l’analyse de documents non structurés. Pour ce faire, il nous a fallu définir au préalable les critères pertinents issus des analyses dans l’objectif de faciliter la découverte de nouvelles valeurs.
This paper describes the participation of the Syllabs Team in the content analysis task of the CLEF MC2 Evaluation lab. In the current state of our work, we offer preliminary solutions to first detect the language of the microblogs used within the task, then extract the named entities that will be later used to recognize Wikipedia entities and finally, detect microblogs that deal with festivals.
This paper analyzes the content of the proceedings of the Language Resources and Evaluation Conference (LREC) over the past 17 years (1998–2014), with the goal of gaining a picture of the LREC community and the topics that are most relevant to the field. We follow the methodology used in similar studies, including the survey of the IEEE ICASSP conference proceedings from 1976 to 1990, the survey of the Association of Computational Linguistics conference proceedings over 50 years, and the survey of the proceedings of the conferences contained in the ISCA Archive over 25 years (1987–2012). We expand on results originally presented at LREC 2014, but include the proceedings of LREC 2014 itself in the study together with an analysis of various citation graphs. We show the evolution over time of the number of papers and authors, including their distribution by gender and affiliation, as well as collaborations and citation patterns among authors and papers, funding sources for reported research, and plagiarism and reuse in LREC papers; results for LREC are compared with similar results for major conferences in related fields. We also consider the evolution of research topics over time and identify the authors who introduced key terms. Finally, we propose and apply a measure of a researcher’s notability and provide the results for LREC authors. The study uses NLP methods that have been published in the corpus considered in the study. In addition to providing a revealing characterization of the LRE community, the study also demonstrates the need for establishing a system for unique identification of authors, papers and other sources to facilitate this type of analysis.
D7.4 reports on the evaluation of the different components integrated in the PANACEA third cycle of development as well as the final validation of the platform itself. All validation and evaluation experiments follow the evaluation criteria already described in D7.1. The main goal of WP7 tasks was to test the (technical) functionalities and capabilities of the middleware that allows the integration of the various resource-creation components into an interoperable distributed environment (WP3) and to evaluate the quality of the components developed in WP5 and WP6. The content of this deliverable is thus complementary to D8.2 and D8.3 that tackle advantages and usability in industrial scenarios. It has to be noted that the PANACEA third cycle of development addressed many components that are still under research. The main goal for this evaluation cycle thus is to assess the methods experimented with and their potentials for becoming actual production tools to be exploited outside research labs.
The PANACEA user’s workshop took place in Budapest on June 29, as a satellite event of the META-FORUM 2011. The choice for this collocated event was indeed strategic. META-FORUM is an event organised by another FP7 project (T4ME/META-NET) that shares a strong interest for the industrial world with PANACEA. This seemed very interesting in order to attract potential users’ interest and encourage them to join us for the workshop. Targeting two events was certainly a plus for some of the attendees.
This paper aims at analyzing the content of the LREC conferences contained in the ELRA Anthology over the past 15 years (1998-2013). It follows similar exercises that have been conducted, such as the survey on the IEEE ICASSP conference series from 1976 to 1990, which served in the launching of the ESCA Eurospeech conference, a survey of the Association of Computational Linguistics (ACL) over 50 years of existence, which was presented at the ACL conference in 2012, or a survey over the 25 years (1987-2012) of the conferences contained in the ISCA Archive, presented at Interspeech 2013. It contains first an analysis of the evolution of the number of papers and authors over time, including the study of their gender, nationality and affiliation, and of the collaboration among authors. It then studies the funding sources of the research investigations that are reported in the papers. It conducts an analysis of the evolution of the research topics within the community over time. It finally looks at reuse and plagiarism in the papers. The survey shows the present trends in the conference series and in the Language Resources and Evaluation scientific community. Conducting this survey also demonstrated the importance of a clear and unique identification of authors, papers and other sources to facilitate the analysis. This survey is preliminary, as many other aspects also deserve attention. But we hope it will help better understanding and forging our community in the global village.
This report elaborates on the exploitation of the PANACEA project assets. These assets have been clustered into a few items (a) the PANACEA Factory/Platform, (b) the web services integrated within the platform, (c) the associated workflows to manage the sequencing of web services (d) the tools developed during the project and last but not least (e) the data sets, i.e. Language Resources (LRs) produced mainly for Machine translation/localization, but not only, within the project, exploiting the platform, web services, and the workflows.
This deliverable defines how evaluation will be carried out at each integration cycle. As PANACEA aims at producing large scale resources, evaluation becomes a critical and challenging issue. Critical because it is important to assess the quality of the results that should be delivered to users. Challenging because we prospect rather new areas, and through a technical platform: some new methodologies will have to be explored or old ones to be adapted.
aim of this document is to list and explain the work developed in WP3 The platform and the resulting documentation.