
With the rising popularity of large language models (LLMs) for various applications ranging from every-day information seeking to high stakes applications such as complex medical question answering, hallucinated content poses a serious risk. Retrieval-Augmented Generation (RAG) aims to enhance the trustworthiness of LLMs by grounding their outputs in external documents, often using inline citations for verifiability. However if those citations are not faithful, i.e. if they do not accurately reflect the source of the information during the answer generation, user trust might be misguided by the sheer existence of such references to trusted sources. We argue that to understand citation faithfulness and to develop a reliable framework for the evaluation of citation faithfulness, a mechanistic approach that considers the model internals, rather than mere observations on the input/output of the model, is necessary. This paper offers the first mechanistic account of how a large language model decides whether to attach an inline citation while answering a factoid question. Through activation patching we identify an "attributional ensemble" of attention heads and MLP layers that are responsible for the citation generation process in Llama-3.1 -8B-Instruct. Our findings suggest that citation decisions rely heavily on shallow heuristics such as entity co-reference matching, raising concerns about the trustworthiness of such citations.
The JOKER Track has created an active community of researchers in NLP and IR working together on the non-literal use of language in text which is still challenging for both AI models and humans, as it requires understanding implicit cultural references and double meanings. Its benchmarks on humorous text analysis, retrieval, and translation have become standard references. We made significant changes to the track's setup and tasks in 2024 and 2025, and propose continuing these to complete the test collections. The CLEF 2026 JOKER track will contain the following four tasks: Task 1 (Humour-aware Information Retrieval): retrieve short humorous texts for a query, Task 2 (Pun Translation): translate puns from English to French and Spanish, Task 3 (Onomastic Wordplay Translation): translate onomastic wordplay from English to French, and Task 4 (Humour Generation): guided creativity.
What is an argument? Is an argument valid? Was a text manipulated to persuade? Since 2020, Touche fosters the development of support-technologies for decision-making and opinion-forming. To this end, the lab brings together researchers that develop systems to automatically answer questions like those above. At CLEF 2026 we do so in four tasks: (1) Fallacy Detection (new task), in which participants determine whether an argument follows a valid argument pattern; (2) Causality Extraction (new task), in which participants extract pro- and concausal claims from text; (3) Generalizability of Argument Identification in Context (new task), in which participants predict whether sentences would be annotated as an argument under different guidelines; and (4) Advertisement in Retrieval-Augmented Generation (2nd edition), in which participants detect and block advertisements in generated text. This paper details these tasks and summarizes the results of Touche 2025.
Over the last few years, the SimpleText Track has created an active community of NLP and IR researchers collaborating to improve access to scientific text. Its benchmarks on scientific passage retrieval, scientific terminology detection and explanation, and scientific text simplification have become standard references. Following a similar track design from 2021 to 2024, we introduced substantial changes to the track's structure and tasks in 2025. We plan to continue this successful setup in 2026, plus add a new (pilot) task on research area classification of scientific papers. Hence, the CLEF 2026 SimpleText track will contain the following 3 tasks. Task 1 (Text Simplification): simplify scientific text. Task 2 (Controlled Creativity): identify and avoid hallucination. Task 3 (Research Area Classification): classification of scientific articles by research area.
Learned sparse retrieval (LSR) models exhibit varying trade-offs between effectiveness and efficiency. But while standard tools exist for evaluating LSR effectiveness, there is none for evaluating efficiency. Also, datasets with high-quality relevance judgments are too large for repeated efficiency experiments, e.g., on different hardware configurations. To promote the evaluation of LSR models in terms of their effectiveness and efficiency, we introduce the lsr_benchmark, which measures retrieval efficiency at each step of an LSR pipeline (document embedding, indexing, query embedding, and retrieval) as well as its overall effectiveness. To ensure tractability and extensibility, we apply current corpus subsampling methods to eleven TREC tasks, precompute embeddings with eleven LSR models per task, and evaluate eight retrieval engines as baselines. For the benchmark’s hosted version, a modular API, along with tools for evaluating effectiveness and efficiency, facilitates the submission of new approaches. Our experiments show that the chosen embedding model significantly affects the efficiency of a retrieval engine and that LSR is more effective but less efficient than BM25—an efficiency gap that our benchmark now tracks as new LSR models are published.
This paper presents a principled and scalable framework for systematically generating complex Question Answering (QA) data. In the core of this framework is a graphlet-anchored generation process, where small subgraphs from a Knowledge Graph (KG) are used in a structured prompt to control the complexity and ensure the factual grounding of questions generated by Large Language Models. The first instantiation of this framework is BioGraphletQA, a new biomedical KGQA dataset of 119,856 QA pairs. Each entry is grounded in a graphlet of up to five nodes from the OREGANO KG, with most of the pairs being enriched with relevant document snippets from PubMed. We start by demonstrating the framework’s value and the dataset’s quality through evaluation by a domain expert on 106 QA pairs, confirming the high scientific validity and complexity of the generated data. Secondly, we establish its practical utility by showing that augmenting downstream benchmarks with our data improves accuracy on PubMedQA from 49.2 https://zenodo.org/records/17381119 ) and framework code ( https://github.com/ieeta-pt/BioGraphletQA ), are publicly available to facilitate use, reproducibility and extension.
Query logs are key resources for studying search engine interactions and improving retrieval effectiveness but are rarely publicly available. In the past, search providers only shared small subsets of their own logs to curb competition and to ensure privacy. The Archive Query Log (AQL) will become an open alternative: mining query logs from archived search engine result pages (SERPs). While the AQL-22 prototype demonstrated the feasibility of this approach, its limited scalability and maintainability hindered widespread adoption by the research community. We re-implement the crawling and parsing of the AQL on open infrastructure, using standard tools, a new framework for storing SERPs, and following FAIR data principles. The extended and continuously crawled AQL corpus currently contains 553 million SERPs from 775 search providers, mined from six web archives, where so far 223 million SERPs (44 TB; 40
Evaluating retrieval-augmented generation (RAG) systems is challenging in the legal domain due to the nuanced nature of reasoning and scarcity of domain-specific datasets. We investigate semantic similarity and legal correctness in Brazilian legal question answering. We develop a synthetic dataset of 3,012 evaluation instances from the Brazilian Civil Procedure Code spanning seven query types, validated through multi-agent critique. Our evaluation framework combines BERTScore, retrieval metrics, and domain-adapted LLM-as-Judge to assess semantic coherence and legal correctness. Analysis reveals that 53.3
Retrieval-Augmented Generation (RAG) enhances Large Language Models by grounding their responses in external knowledge, yet existing approaches struggle with complex reasoning tasks that require combining multiple sources of information. Although recent multi-retrieval methods integrate iterative retrieval and reasoning, they still lack mechanisms to decide when to stop retrieving and how to maintain reasoning coherence, often leading to error propagation and hallucinations. This research introduces a metacognitive multi-agent framework for RAG that models reasoning as a collaborative process guided by metacognitive control and shared memory systems. The framework enables dynamic coordination between retrieval and reasoning, allowing the system to monitor its progress, assess evidence sufficiency, and revise its reasoning when inconsistencies appear. By incorporating metacognitive regulation and explicit memory interaction, the proposed approach aims to improve reasoning reliability, factual grounding, and interpretability in multi-hop question answering.
AI is increasingly central to understanding and managing biodiversity and ecosystems. Since 2011, the LifeCLEF lab has provided large-scale benchmarks that stimulate progress in multimodal species recognition, ecological prediction, and knowledge extraction. The 2026 edition expands this scope with five complementary challenges spanning visual, acoustic, and textual data: (i) AnimalCLEF: discovery and re-identification of individual animals, (ii) BirdCLEF+: multi-taxonomic species recognition in complex soundscapes, (iii) FathomNetCLEF: detection of marine species in underwater imagery under positive-unlabeled constraints, (iv) PestCLEF: extraction of information on plant pests from heterogeneous textual sources, (v) PlantCLEF: multi-species plant identification in quadrat images. Together, these challenges address critical dimensions of biodiversity science and ecosystem management, while fostering collaboration between AI researchers, ecologists, and practitioners. This paper provides an overview of the LifeCLEF 2026 lab and its tasks, outlining their motivation, data, and evaluation methodology to guide participants and inform the wider research community.
The robustness of embedding-based retrieval systems is critical for reliable information access, yet these models remain vulnerable to distributional shifts and adversarial manipulations. My doctoral project focuses on addressing this challenge by investigating robustness along two complementary dimensions: generalizability, which examines performance consistency across diverse and evolving real-world scenarios, and stability, which assesses resilience against both unintentional perturbations and malicious attacks. Anchored in a two-phase program–(RQ1) Understanding and (RQ2) Enhancing Robustness–the project first performs a systematic empirical study to attribute failure modes in modern retrievers. Building on these insights, the second phase develops principled defense strategies, moving beyond standard data augmentation to explore novel training paradigms and task-specific robust adaptation. Ultimately, this work aims to establish methodological foundations for the next generation of trustworthy dense retrieval systems.
Personalized food recommendation can promote healthier, sustainable eating, but current systems often rely on sparse and unstructured data, limiting semantic expressiveness and diverse personalization. In this paper, we propose FoodNexus, a large-scale knowledge graph with nearly one billion triples designed to enrich food recommendation with structured, nutrition-aware, and user-contextual information. We built it via a multi-stage pipeline that combines and augments the largest public dataset of user–recipe interactions, HUMMUS, with extensive metadata from Open Food Facts by linking recipes to concrete food products, extracting user traits from their biographies and reviews, and mapping both data sources onto the same ontology. Experiments show that FoodNexus enables richer, nutrition-sensitive evaluation of recommendations. Code Resource: https://github.com/tail-unica/food-nexus .
Many components of information retrieval systems evolve over time. The LongEval Lab aims to provide a benchmark setting to the longitudinal evaluation of IR models. At its fourth edition, LongEval we focus on scholarly search and scholarly user models. We describe in this paper the tasks that are planned for the 2026 lab, the data necessary for each of the tasks, as well as the choice of evaluation activities.
Context attribution in retrieval-augmented generation (RAG) identifies and scores context sentences that support a model’s answer, thereby improving transparency and trustworthiness. However, existing attention-based methods often aggregate attention over all tokens in a sentence, which introduces noise and causes instability across varying model sizes. They also rely on manually tuned selection thresholds to determine supportive sentences, limiting scalability. This work addresses these issues with three key contributions. First, we systematically explore token selection strategies to enhance attention-based attribution scoring. Second, we introduce a think-twice mechanism that refines attention to capture previously overlooked sources. Third, we propose a greedy context ablation method that automatically determines the number of sources without relying on manual thresholds. Experiments on MESAQA, a multi-evidence dataset, show that our method achieves an improvement of up to 0.70 in F1-score and remains competitive at 0.68 even without any hyperparameter tuning, demonstrating its effectiveness and practicality for reliable context attribution.
Sequential recommender systems have achieved significant success in modeling temporal user behavior but remain limited in capturing rich user semantics beyond interaction patterns. Large Language Models (LLMs) present opportunities to enhance user understanding with their reasoning capabilities, yet existing integration approaches create prohibitive inference costs in real time. To address these limitations, we present a novel knowledge distillation method that utilizes textual user profile generated by pre-trained LLMs into sequential recommenders without requiring LLM inference at serving time. The resulting approach maintains the inference efficiency of traditional sequential models while requiring neither architectural modifications nor LLM fine-tuning.
The ELOQUENT lab for evaluation of generative language model quality and usefulness addresses high-level quality criteria for generative language models through a set of open-ended shared tasks implemented, where possible, to minimise human effort in assessment, and with an objective to study how much the languages that the foundation model has been trained on make a difference in its responses. In this third ELOQUENT edition, the three planned tasks investigate how human-like text generated by language models can be (the Voight-Kampff task), how reliably a language model handles varied but equivalent input across languages (the Robustness and Consistency task), and if a generative language model can be used productively to generate and score topical quizzes without diverging into general knowledge acquired in foundational training (the PISA task). All tasks are continued evolved versions of previous editions’ tasks.
Previous research on group recommender systems (GRSs) has shown that group dynamics strongly influence decision-making, yet collaborative filtering (CF)–based GRSs rarely account for social interactions, partly because suitable tools to capture and analyze live interaction traces are limited. This paper introduces a community resource for studying live groups engaging with a CF-based recommender system through a domain-independent graphical interface that records structured interaction signals (e.g., suggestions, views, and favorites) and integrates them into interaction-aware consensus strategies. A live user study with 72 participants organized into 18 groups illustrates the platform’s ability to capture and analyze user interactions, including interface engagement patterns and perceived social roles. Comparing two interaction-aware consensus strategies (mean vs. completeness), we observe differences in satisfaction distributions and interaction patterns; isolating the causal contribution of interaction signals would require an interaction-unaware baseline. Source code and dataset are available online at this link ( https://github.com/davidcontrerasaguilar/GREAT.git ).