Calaveritas are seen as a sort of poem or ode to the dead, since is a Mexican tradition linked to the Day of the death, in which through text you could see the personification of death and popular characters been mock and satirize, this topic is quite interesting because it have humor and the structure of the more serious poem. The classification of this textual genre could help locate text that contain humor with unconventional variants or structures and to preserver in a way this written scheme. To tackle this task, it was decided use a machine learning approach for the baseline, taking quite good results around 94% on the F1-score for the top methods of the baseline, in this case the main approach was to finetune Transformers like BETO or BERT-multilingual obtaining 98% and 97% on de F1-score and analyze the similarities and to observer the characteristic inherent to each class. The classes were quite separable since the calaveritas are more near related to humor than to the classical approach of poem, since the text of these classes contain words that are more easily identifiable. Given the observed degree of separability between classes, we sought to ensure that the classification was not primarily driven by topical information. To this end, we masked the most frequent words in each class as a preliminary control experiment, the produced results were broadly comparable to those obtained through fine-tuning our main models, suggesting that structural features may play a role in the classification process. In a way humour intervenes to create a poem-like structure with a humoristic content, A hybrid, perhaps. Accepted
Over the last few years, the SimpleText Track has created an active community of NLP and IR researchers collaborating to improve access to scientific text. Its benchmarks on scientific passage retrieval, scientific terminology detection and explanation, and scientific text simplification have become standard references. Following a similar track design from 2021 to 2024, we introduced substantial changes to the track's structure and tasks in 2025. We plan to continue this successful setup in 2026, plus add a new (pilot) task on research area classification of scientific papers. Hence, the CLEF 2026 SimpleText track will contain the following 3 tasks. Task 1 (Text Simplification): simplify scientific text. Task 2 (Controlled Creativity): identify and avoid hallucination. Task 3 (Research Area Classification): classification of scientific articles by research area.
Building on the success and insights from previous years, the CLEF 2025 SimpleText Track continues to advance the mission of making scientific information more accessible to a broader audience. In 2025, we introduced a new biomedical corpus, based on aligned Cochrane abstracts and plain language summaries, for the main scientific text simplification task. In addition, we devote particular attention to remaining issues of current generative models, focusing on indentifying and classifying overgeneration and other information distortion in the predictions, as well as promoting grounded text generation approaches. This paper presents an overview of the CLEF 2025 SimpleText Track. Task 1 focuses on Text Simplification, aiming to simplify complex scientific texts. Task 2 addresses Controlled Creativity, emphasizing the detection, classification, and avoidance of hallucinations in generated content. Task 3 revisits selected tasks from SimpleText 2024 by popular demand. We discuss the data and benchmarks provided for these tasks, along with preliminary insights and anticipated challenges.
The JOKER Track has created an active community of researchers in NLP and IR working together on the non-literal use of language in text – which is still challenging for both AI models and humans, as it requires understanding implicit cultural references and double meanings. Its benchmarks on humorous text analysis, retrieval, and translation have become standard references. We made significant changes to the track’s setup and tasks in 2024 and 2025, and propose continuing these to complete the test collections. The CLEF 2026 JOKER track will contain the following four tasks: Task 1 (Humour-aware Information Retrieval): Retrieve short humorous texts for a query, Task 2 (Pun Translation): translate puns from English to French and Spanish, Task 3 (Onomastic Wordplay Translation): translate onomastic wordplay from English to French, and Task 4 (Humour Generation): Guided Creativity.
Over the last few years, the SimpleText Track has created an active community of NLP and IR researchers collaborating to improve access to scientific text. Its benchmarks on scientific passage retrieval, scientific terminology detection and explanation, and scientific text simplification have become standard references. Following a similar track design from 2021 to 2024, we introduced substantial changes to the track’s structure and tasks in 2025. We plan to continue this successful setup in 2026, plus add a new (pilot) task on research area classification of scientific papers. Hence, the CLEF 2026 SimpleText track will contain the following 3 tasks. Task 1 (Text Simplification): Simplify Scientific Text. Task 2 (Controlled Creativity): Identify and Avoid Hallucination. Task 3 (Research Area Classification): Classification of scientific articles by research area.
The JOKER Track has created an active community of researchers in NLP and IR working together on the non-literal use of language in text which is still challenging for both AI models and humans, as it requires understanding implicit cultural references and double meanings. Its benchmarks on humorous text analysis, retrieval, and translation have become standard references. We made significant changes to the track's setup and tasks in 2024 and 2025, and propose continuing these to complete the test collections. The CLEF 2026 JOKER track will contain the following four tasks: Task 1 (Humour-aware Information Retrieval): retrieve short humorous texts for a query, Task 2 (Pun Translation): translate puns from English to French and Spanish, Task 3 (Onomastic Wordplay Translation): translate onomastic wordplay from English to French, and Task 4 (Humour Generation): guided creativity.
Humour poses a unique challenge for artificial intelligence, as it often relies on non-literal language, cultural references, and linguistic creativity. The JOKER Lab, now in its fourth year, aims to advance computational humour research through shared tasks on curated, multilingual datasets, with applications in education, computer-mediated communication and translation, and conversational AI. This paper provides an overview of the JOKER Lab held at CLEF 2025, detailing the setup and results of its three main tasks: (1) humour-aware information retrieval, which involves searching a document collection for humorous texts relevant to user queries in either English or Portuguese; (2) pun translation, focussed on humour-preserving translation of paronomastic jokes from English into French; and (3) onomastic wordplay translation, a task addressing the translation of name-based wordplay from English into French. The 2025 edition builds upon previous iterations by expanding datasets and emphasising nuanced, manual evaluation methods. The Task 1 results show a marked improvement this year, apparently due to participants' judicious combination of retrieval and filtering techniques. Tasks 2 and 3 remain challenging, not only in terms of system performance but also in terms of defining meaningful and reliable evaluation metrics.
This paper highlights the evolution and future directions of the SimpleText Track at CLEF, which, over the last few years, fostered an active NLP and IR research community focused on improving access to scientific text. Its benchmarks on scientific passage retrieval, scientific terminology detection and explanation, and scientific text simplification have become standard references. After using a similar setup of the track in 2021–2024, we propose substantial modifications to the track’s structure and tasks. The CLEF 2025 SimpleText track will contain the following three tasks. Task 1 on Text Simplification: Simplify Scientific Text. Task 2 on Controlled Creativity: Identify and Avoid Hallucination. Task 3 on SimpleText 2024 Revisited: Selected Tasks by Popular Request.
Errors in natural language generation, so-called hallucinations, remain a critical challenge, particularly in high-stakes domains such as healthcare or science communication. While several automatic metrics have been proposed to detect and quantify hallucinations, such as FactCC, QAGS, FEQA, and FactAcc, these metrics are often unavailable, difficult to reproduce, or incompatible with modern development workflows. We introduce MIRAGE, an open-source Python library designed to address these limitations. MIRAGE reimplements key hallucination evaluation metrics in a unified library built on the Hugging Face framework, offering modularity, reproducibility, and standardized inputs and outputs. By adhering to FAIR principles, MIRAGE promotes reproducibility, accelerates experimentation, and supports the development of future hallucination metrics. We validate MIRAGE by re-evaluating existing metrics on benchmark datasets, demonstrating comparable performance while significantly improving usability and transparency.
As digital competence becomes a core priority for lifelong learning in Europe, the need to teach AI literacy grows increasingly urgent. To address this, we present IILAP, a browser-based Interactive Information Literacy Assessment Platform, designed to support the teaching and assessment of critical reading of AI-generated content. IILAP includes a Teacher Tool for curating chatbot QA datasets with truth labels and sources, and a Student Interface that provides responses enriched with citations and trust indicators. The system logs interaction data-such as time on task, source clicks, verification attempts, and error detection-to help educators identify gaps in students' critical reading skills. After each session, automated Excel reports summarize these measures for easy assessment. Developed through an initial user study, IILAP enables classroom deployment of controlled chatbot interactions and provides structured analytics aligned with indicators of critical thinking. This demo showcases how the system bridges user behavior and educational evaluation to fostering AI literacy in education.
Pierre De Loor合作论文数 CNRS ;Lab-STICC ; ENIB6