The objective of this paper is to address the scarcity of labeled textual data and improve the performance of language models in classification tasks within a One Health context by using small domain-specific labeled corpora. To address these challenges, we propose a two-phase training pipeline for language models, in which the first phase involves post-training guided by selective masking (SM) strategies to adapt the model to a specific domain. For this purpose, we propose two novel masking strategies: SM-Lex-TFIDF, which masks domain lexicon terms with high TF-IDF (term frequency-inverse document frequency) values, and SM-NonLex-TFIDF, which masks non-domain lexicon terms with high TF-IDF values. The second phase focuses on fine-tuning the model for the target classification task using small amounts of labeled data. To demonstrate the effectiveness of our approach, we focus on two related application areas within the One Health context, i.e., (i) thematic content in integrated health, covering the biomedical, plant health, and syndromic surveillance domains, and (ii) epidemic misinformation, to achieve improved One Health monitoring. We conduct extensive evaluations to assess the performance of our approach using three language models: BERTBase, SciBERT, and BioBERT. Additionally, we compare our method with low-resource LLM-based approaches, including zero/few-shot classification. Experimental results demonstrate significant improvements in the performance of the language models across classification tasks in both targeted areas, even with limited labeled data. Our approach outperforms zero/few-shot classification using LLaMA-3.1-8B and Mistral-7B in four out of the five datasets evaluated. Furthermore, we provide a summary mapping each strategy to its most effective context.
Event-Based Surveillance (EBS) systems play an important role in the early detection of disease outbreaks by monitoring unstructured data sources. However, they face key challenges, including handling the overwhelming volume of collected articles, detecting false positives, and the lack of explainability in detected events. To address these limitations, we previously developed EpiDCA, an unsupervised model that integrates epidemiological and environmental data, as well as expert-defined parameters, into the classification process. Its initial application to avian influenza (AI) in Asia demonstrated very promising results, comparable to well-known supervised baseline methods. In this paper, we assess the robustness and genericity of EpiDCA by applying it to three different case studies: AI in France, African swine fever (ASF) in Europe, and West Nile virus disease (WND) in Europe. Results showed that EpiDCA effectively distinguishes relevant from irrelevant events, achieving weighted F-scores between 0.64-0.85 across all case studies. Sensitivity analysis demonstrated model robustness, with most parameters showing minimal influence on results. Notably, the incorporation of environmental data and finer spatial granularity significantly improved classification precision. Overall, EpiDCA remains robust and adaptable across diverse epidemiological contexts, further validating its effectiveness as a valuable tool for event-based surveillance, with improved interpretability and real-time classification.
Event Based Surveillance (EBS) monitors online sources such as broadcast, print, web news and generates early warning and response (EWAR) signals for use in disaster mitigation. These online sources provide a dynamic data source allowing for potential real-time EBS updates. However, in dealing with news articles, fragmented information exists in varied sources and redundant information are known to overburden EBS. In this study we propose a Large Language Model based approach that filters out redundancies while learning novel information from event centric online news corpora. We study this novelty task for events covering animal health, food security and climate change surveillance domains. Our approach focuses on features integrating spatio-temporal information (such as location and date of event) and thematic information (such as the name of disease, food insecurity triggers, climate change magnitude).We characterize novelty as presence of new and additional information (e.g., a newly mentioned disease name or additional location information) as distinguished from duplicate (e.g., an already seen disease name) and missing (expected but absent) information. To this regard, our approach proposes fine-grained classification of novelty in event surveillance and language modeling adoption with a multi-class classification objective to learn classifying of event information. Our LLM adoption strategy proposes question-based prompts whose extracted answers map to predefined feature types (e.g., location, date, name of disease) in order to enrich our classifier. In our empirical studies, we present comparative analysis with respect to language models and large language models for State-Of-The-Art performance in the event novelty classification task. Our findings demonstrates the ability of cross-domain novelty classification with our model EpidGPT (few-shot) achieving F1% scores of 82.3, 85.49 and 88.97 in animal health, food security and climate change domains while finetuned EpidGPT achieves F1% scores of 96.02, 86.0 and 88.45 on each respectively domains.
This document, based on feedback from UMR TETIS members and the scientific literature, provides a generic methodology for creating annotation guidelines and annotated textual datasets (corpora). It covers methodological aspects, as well as storage, sharing, and valorization of the data. It includes definitions and examples to clearly illustrate each step of the process, thus providing a comprehensive framework to support the creation and use of corpora in various research contexts.
The location of events in multilingual texts, particularly in Arabic, represents a challenge for epidemiological monitoring. Systems such as PADI-web rely on English translation to extract spatial entities, but the scarcity of annotated spatial entities in Arabic can hamper the reliability of translations and extraction. In this context, PADI-Location-AR-EN, which is a dataset of 328 spatial entities that were manually extracted from 96 Arabic-language news articles collected by the PADI-web epidemiological monitoring system, is presented in this paper. Each entity was manually translated into English, normalized using the GeoNames database, and then classified according to its type and spatial category. The dataset can be used to evaluate the translation quality of three machine translation systems (DeepL, Microsoft Azure and Reverso) as well as the performance of named entity recognition models on the translated texts.
Between 2019 and 2025, the European Commission funded 82 Research & Innovation projects in the agricultural sector across the globe with the main goal to support agrifood system transformation and sustainability transitions. To analyze the type of innovations developed and promoted by these projects, we used a text-mining approach on a corpus of various projects’ documents (including project descriptions, project reports, descriptions of project events, publications, and websites), produced at different stages of their implementation. Since these Research & Innovation projects addressed agricultural innovation challenges from different angles (i.e. different innovation purposes, strategies, stakeholders, processes and outcomes), we decided to create a comprehensive lexicon on innovation. This lexicon covers six entry points to analyze multi-dimensional innovation: innovation actors, triggers, processes, outputs, purposes and phases. This data paper presents the lexicon and the methodology used to create it. The lexicon is composed of 717 keywords in English, divided into 44 concepts and associated with the 6 entry points presented above. It has been built iteratively, through 3 workshops with experts in agricultural innovation. The lexicon is intended to be used as a semantic layer in knowledge management tools, for instance to identify the most cited keywords and the most frequent keyword associations according to the type of documents, the localization of projects and the phase of implementation. These results contribute to building typologies of Research & Innovation projects to better understand the contribution of large project portfolios funded by the European Commission to innovation and impact.
Online news sources are popular resources for learning about current health situations and developing event-based surveillance(EBS) systems. However, having access to diverse information originating from multiple sources can misinform stakeholders, eventually leading to false health risks. The existing literature contains several techniques for performing data quality evaluation to minimize the effects of misleading information. We mainly proposed three approaches to assess the quality of news sources. In our research, our primary focus was on ensuring data quality assessment at two levels: 1) News article level and 2) News source level. We explored data quality assessment at the news article level through two main approaches: 1) Data-driven score-based approach and 2) Metadata-based machine learning(ML) approach. The data-driven score-based approach aims to classify relevant and irrelevant news articles, adding an explainability aspect in the context of EBS. Similarly, the metadata approach is employed for classification, utilizing news article metadata features in ML models to highlight important metadata features. For source-level quality assessment, we identified exogenous metadata attributes such as source categorization and geographical coverage associated with news sources, extracting this information automatically. With the help of extracted source metadata, we conducted the classification of news sources. The obtained results hold significance in terms of prioritizing news sources within the context of EBS. Nevertheless, further investigation is required to enhance the methodology of this approach.
Source variables, or observable properties, used to describe agroecological experiments are often heterogeneous, non-standardized, and multilingual, making them challenging to understand, explain, and utilize in cropping system modeling and multicriteria evaluations of agroecological system performance. A potential solution is data annotation via a controlled vocabulary, known as candidate variables, from the Agroecological Global Information System (AEGIS). However, matching source and candidate variables via their textual descriptions remains a challenging task in agroecology. Domain-general language models, such as BERT, often struggle with domain-specific tasks due to their general-purpose training data. In the literature, these models are adapted to specialized domains through further pretraining, pretraining from scratch, and/or fine-tuning on downstream tasks. However, pretraining a domain-general model on a domain-specific corpus is resource-intensive, requiring substantial time, energy, and computational resources. To the best of our knowledge, no study has further pretrained a domain-general model on a small corpus (less than 100 MB) to adapt it to a domain-specific task and evaluated it on downstream tasks without fine-tuning. To address these shortcomings, this paper proposes further pretraining BERT and AgriBERT on a small agroecology-related corpus. This approach is designed to be both time- and resource-efficient while enhancing domain adaptation. We evaluate the pretrained models on the task of matching source and candidate variable descriptions without fine-tuning. Our results show that our further pretrained AgriBERT (+ Experts + Core) model outperforms all others by more than 8
Due to its highly contagious nature, Avian Influenza (AI) is considered an animal health emergency affecting commercial sector and wild bird populations. Several genome sequencing databases have been created to help researchers understand how AI viruses evolve, spread, and cause disease. However, for a global epidemic monitoring approach, they need to be combined to public health surveillance systems, the well-one being EMPRES-i from the World Organisation for Animal Health (WOAH) and the Food and Agriculture Organization of the United Nations (FAO). This paper presents a new AI dataset, in which EMPRES-i is enriched thanks to the genome sequence data of Avian Influenza cases affecting bird species from 2012 to 2021, publicly provided by the Bacterial and Viral Bioinformatics Resource Center (BV-BRC). This dataset is obtained by automatically linking sequence information in BV-BRC to the AI events in EMPRES-i, which results in "putatively" linked events between these two sources. The collected data is structured by nature, but it is preprocessed and normalized for the purpose of high-quality data linkage. Moreover, several data linkage strategies and missing information handling are introduced. To show the usefulness of our dataset, we quantitatively evaluate the proposed strategies in randomly sampled events and present in the end a diffusion network inference task.
To address the current crises (climatic, social, economic), the self-sufficiency – a set of practices that combine energy sobriety, self-production of food and energy, and self-construction - arouses an increasing interest. The CNRS STAY project (Savoirs Techniques pour l’Auto-suffisance sur YouTube) explores this topic by analyzing techniques shared on YouTube. We present Agro-STAY, a platform designed for the collection, processing, querying and visualization of data from YouTube videos and their comments. We propose a full methodology dedicated to processing YouTube videos, and apply Natural Language Processing (NLP) techniques and language models, which enable a fine-grained analysis of alternative agricultural practices described online. In addition, we provide a well-adapted graphical user interface to help the experts analyzing the extracted knowledge.
Land artificialization is a significant modern concern, as it is irreversible, diminishes agriculturally suitable land and causes environmental problems. Our project, Hérelles, aims to address this challenge by developing a framework for land artificialization management. In this framework, we associate urban planning rules in text form with clusters extracted from time series of satellite images. To achieve this, it is crucial to understand the planning rules with two key objectives: (1) to verify if the constraints derived from the rules are verifiable on satellite images and (2) to use these constraints to guide the labelling (or semantization) of clusters. The first step in this process involves the automatic extraction of rules from urban planning documents written in the French language. To solve this problem, we propose a method based on the multilabel classification of textual segments and their subsequent summarization. This method includes a special format for representing segments, in which each segment has a title and a subtitle. We then propose a cascade approach to address the hierarchy of class labels. Additionally, we develop several text augmentation techniques for French texts that can improve prediction results. Finally, we reformulate classified segments into concise text portions containing necessary elements for expert rule construction. We adapt an approach based on Abstract Meaning Representation (AMR) graphs to generate these portions in the French language and conduct a comparative analysis with ChatGPT. We experimentally demonstrate that the resulting framework correctly classifies each type of segment with more than 90
Named Entity Recognition (NER) in specialized domains poses major challenges due to the scarcity of annotated data and the limitations of existing models. One challenge lies in handling long documents, which often requires segmenting the text into smaller chunks that fit within the model's input window. Although several segmentation strategies exist, their impact on NER performance in new domains remains underexplored. Another challenge is domain adaptation: while openschema NER models can perform reasonably well under zero-shot settings, they often struggle in highly specialized contexts without additional supervision. To address these issues, we combine segmentation strategy selection with a semi-supervised finetuning pipeline based on pseudo-labeled annotations. First, we compare four segmentation strategies to identify which offers the best trade-off between precision and recall in zero-shot settings. Then, we fine-tune two open-schema NER models-GLiNER and NuNER-first on a manually annotated dataset, and subsequently on a large pseudo-labeled dataset built from model agreement. Both models are evaluated on a held-out test set and on the full manually annotated corpus. Experiments on French-language documents specifically focused on the underexplored domain of territorial food systems reveal that fine-tuning on pseudo-labeled data-obtained through cross-model agreement-yields better performance than relying solely on human-annotated data. The results also highlight the strengths and weaknesses of different segmentation strategies and confirm the importance of optimizing segmentation choices for NER in domain-specific low-resource settings. Code related to this work is available on GitHub11https://github.com/ibzodiaz/segmentation-strategies.
Understanding the environmental factors that facilitate the occurrence and spread of infectious diseases in animals is crucial for risk prediction. As part of the H2020 Monitoring Outbreaks for Disease Surveillance in a Data Science Context (MOOD) project, scoping literature reviews have been conducted for various diseases. However, pathogens continuously mutate and generate variants with different sensitivities to these factors, necessitating regular updates to these reviews. In this paper, we propose to evaluate the potential benefits of artificial intelligence (AI) for updating such scoping reviews. We thus compare different combinations of AI methods for solving this task. These methods utilize generative large language models (LLMs) and lighter language models to automatically identify risk factors in scientific articles.
Qualitative research, widely employed across various academic fields, explores phenomena using non numerical data, with a particular focus on understanding the meanings, experiences, and perspectives of participants. In contrast to other type of research, it seeks to answer how, where, what, when and why individuals behave or respond in certain ways toward specific issues or topics. Qualitative research involves collecting and analyzing textual data, with interviews playing a central role in gathering expert knowledge. An essential part of data analysis is coding, using specially developed code system hierarchy that helps to categorize and organize responses and facilitates the retrieval of insights. Manual data coding is labor-intensive, and to automate this process we developed the AgriCode tool based on machine learning and manually annotated data. To address data scarcity and improve the prediction quality of our offline classifiers, we perform data augmentation using Retrieval-Augmented Generation (RAG), a state-of-the-art method originally designed for online Q&A systems. Our tool automates the coding of interview responses within the Horizon Europe Agriloop project, which focuses on agricultural waste in the food industry. AgriCode predicts a subset of a predefined code system hierarchy, assisting a human coder by accelerating the process and identifying errors in manual coding. Although initially designed for the valorization of agricultural residues, AgriCode's methodology can be adapted for any qualitative research domain characterized by data scarcity and the need of automated textual analysis. To achieve this, responses from the first round of interviews must be manually annotated using dedicated code system hierarchy. They can then be used for fine-tuning the model, while the RAG method can be employed to address the lack of data for certain classes.
Amidst the overwhelming volume of health-related data available, diverse epidemiological surveillance strategies have been adopted to swiftly detect outbreak events. These strategies differ in terms of structure, type, and sources used. When combined, they offer a more comprehensive understanding of epidemiological events then when used alone. In this paper, we propose an unsupervised approach that allows epidemiological data to be combined with risk factors related to disease onset. We applied this method, named EpiDCA, to enhance the classification and early detection capabilities of Event-Based Surveillance (EBS) systems.EpiDCA is an adaptation of the Dendritic Cells Algorithm (DCA) inspired by the danger theory. The DCA has been applied in various studies and has shown promising results in real-time and binary classification problems. However, some stochastic elements in the algorithm, such as the random sampling and the migration threshold in the detection phase, have been criticized. To overcome these limitations, we integrated spatio-temporal information into the method. We then applied EpiDCA to an avian influenza case study and evaluated the results. These were very promising, and comparable to well-known, supervised baseline methods.
Source variables or observable properties used to describe agroecological experiments are heterogeneous, nonstandardized, and multilingual, which makes them challenging to understand, explain, and use in cropping system modeling and multicriteria evaluations of agroecological system performance. Data annotation via a controlled vocabulary, known as candidate variables from the agroecological global information system (AEGIS), offers a solution. Text similarity measures play crucial roles in tasks such as word-sense disambiguation, schema matching in databases, and data annotation. Commonly used measures include (1) string-based similarity, (2) corpus-based similarity, (3) knowledge-based similarity, and (4) hybrid-based similarity, which combine two or more of these measures. This work presents a hybrid approach called Matching Agroecological Experiment Variables (MAEVa), which combines well-known techniques (PLMs, multi-head attention, TF–IDF) tailored to the challenges of aligning source and candidate variables in agroecology. MAEVa integrates the following components: (1) Our key innovation, which consists of extending pretrained language models (PLMs) (i.e., BERT, SBERT, SimCSE) with an external multi-head attention layer for matching variable names; (2) An analysis of the relevance and impact of various data collection techniques (snippet extraction, scientific articles) and prompt-based data augmentation on TF–IDF for matching variable descriptions; (3) A linear combination of components (1) and (2); and (4) A voting-based method for selecting the final matching results. Experimental results demonstrate that extending PLMs with an external multi-head attention layer improves the matching of variable names. Furthermore, TF–IDF benefits consistently from the presence of an enriched corpus, regardless of the specific enrichment technique employed.
Language models now constitute essential tools for improving efficiency for many professional tasks such as writing, coding, or learning. For this reason, it is imperative to identify inherent biases. In the field of Natural Language Processing, five sources of bias are well-identified: data, annotation, representation, models, and research design. This study focuses on biases related to geographical knowledge. We explore the connection between geography and language models by highlighting their tendency to misrepresent spatial information, thus leading to distortions in the representation of geographical distances. This study introduces four indicators to assess these distortions, by comparing geographical and semantic distances. Experiments are conducted from these four indicators with eight widely used language models and their implementations are available on github ( https://github.com/tetis-nlp/geographical-biases-in-llms ). Results underscore the critical necessity of inspecting and rectifying spatial biases in language models to ensure accurate and equitable representations.
Pascal Poncelet合作论文数University Montpellier 2 - LIRMM51