The First Workshop on INformation access in Uncertainty ScEnarios (INFUSE) [Rieger et al., 2026], in conjunction with the Forty-eighth European Conference on Information Retrieval (ECIR) 2026, focused on the complexities people face when seeking information on issues that lack a clear, definitive answer. These situations, defined as "uncertainty scenarios", often involve scarce information, such as data voids, rapidly evolving knowledge, such as breaking news or rumours, or debated topics with multiple conflicting perspectives. A core problem is that modern information retrieval (IR) systems and large language models (LLMs) are primarily optimised to produce answers rather than communicate the limits of existing knowledge. Consequently, these systems may generate plausible-sounding responses even when reliable information is missing, thereby obscuring rather than exposing uncertainty. This dynamic introduces significant risks, including fostering a false sense of certainty, amplifying mis- and disinformation, and exposing users to viewpoint bias. To address these challenges, this full-day, in-person event brought together a cross-disciplinary group of researchers and practitioners. The primary objectives of the workshop included: • Highlighting emerging risks: Drawing attention to the potential harms caused by the rapid integration of LLM-enabled search services. • Building community: Connecting professionals and researchers from various backgrounds who are focused on information access under uncertainty. • Improving information systems: Contributing to the creation of systems that promote critical thinking, awareness of complex nuances, and an informed citizenry. • Fostering collaboration: Creating a dedicated space for attendees to discover shared interests and form impactful research partnerships. The INFUSE workshop was designed to be highly interactive, moving away from one-sided presentations. The program featured invited talks from interdisciplinary experts, participant presentations grouped by topic, "musical" roundtables for discussions, and a collaboration fair to help participants launch new research projects. The discussions and outcomes are compiled here as a collaborative report. Our aim is to share these insights with the broader IR community and help seed further dialogue on uncertainty in information access. Date: 2 April 2026. Website: https://sites.google.com/view/infuse-workshop/.
Large language models (LLMs) are increasingly embedded in information access systems, extending, and in cases such as ChatGPT, replacing the traditional “10 blue links” paradigm. These systems can enhance knowledge access by summarising large bodies of text or simplifying complex prose, yet they also amplify the risk of overreliance on generated output, potentially eroding users’ critical thinking and information-evaluation skills, an emerging phenomenon sometimes referred to as The ChatGPT Effect. Therefore, this research investigates how scaffolded and personalised cognitive support can help users engage more critically with content in LLM-based information access systems. Drawing on educational psychology and dual-process theories of reasoning, the project examines how nudging and boosting interventions can be designed to (1) adapt to users’ digital literacy and need for cognition, (2) promote meta-cognitive awareness and lateral reading strategies, and (3) preserve autonomy by gradually reducing system support as proficiency increases. The research employs mixed methods, combining user studies with adaptive interface prototyping to evaluate the effects of personalised scaffolds on users’ critical reasoning. The central aim is to introduce critical reasoning processes into LLM-based information access systems through the modelling and evaluation of adaptive scaffolding that initially guides users, analogous to educational instruction, and progressively withdraws assistance as these skills are internalised. By embedding cognitive and pedagogical principles into interactive IR systems, this research advances a new paradigm of human-centred information resilience, shifting the focus from algorithmic detection to the cultivation of critical reasoning skills.
Generative AI (GenAI) tools are transforming information seeking, but their fluent, authoritative responses risk overreliance and discourage independent verification and reasoning. Rather than replacing the cognitive work of users, GenAI systems should be designed to support and scaffold it. Therefore, this paper introduces an LLM-based conversational copilot designed to scaffold information evaluation rather than provide answers and foster digital literacy skills. In a pre-registered, randomised controlled trial (N=261) examining three interface conditions including a chat-based copilot, our mixed-methods analysis reveals that users engaged deeply with the copilot, demonstrating metacognitive reflection. However, the copilot did not significantly improve answer correctness or search engagement, largely due to a "time-on-chat vs. exploration" trade-off and users' bias toward positive information. Qualitative findings reveal tension between the copilot's Socratic approach and users' desire for efficiency. These results highlight both the promise and pitfalls of pedagogical copilots, and we outline design pathways to reconcile literacy goals with efficiency demands.
Many users struggle with effective online search and critical evaluation, especially in high-stakes domains like health, while often overestimating their digital literacy. Thus, in this demo, we present an interactive search companion that seamlessly integrates expert search strategies into existing search engine result pages. Providing context-aware tips on clarifying information needs, improving query formulation, encouraging result exploration, and mitigating biases, our companion aims to foster reflective search behaviour while minimising cognitive burden. A user study demonstrates the companion's successful encouragement of more active and exploratory search, leading users to submit 75
The DRAGUN Track at TREC 2025 targets the growing need for effective support tools that help users evaluate the trustworthiness of online news. We describe the UR_Trecking system submitted for both Task 1 (critical question generation) and Task 2 (retrieval-augmented trustworthiness reporting). Our approach combines LLM-based question generation with semantic filtering, diversity enforcement using clustering, and several query expansion strategies (including reasoning-based Chain-of-Thought expansion) to retrieve relevant evidence from the MS MARCO V2.1 segmented corpus. Retrieved documents are re-ranked using a monoT5 model and filtered using an LLM relevance judge together with a domain-level trustworthiness dataset. For Task 2, selected evidence is synthesized by an LLM into concise trustworthiness reports with citations. Results from the official evaluation indicate that Chain-of-Thought query expansion and re-ranking substantially improve both relevance and domain trust compared to baseline retrieval, while question-generation performance shows moderate quality with room for improvement. We conclude by outlining key challenges encountered and suggesting directions for enhancing robustness and trustworthiness assessment in future iterations of the system.
While it is often assumed that searching for information to evaluate misinformation will help identify false claims, recent work suggests that search behaviours can instead reinforce belief in misleading news, particularly when users generate queries using vocabulary from the source articles. Our research explores how different query generation strategies affect news verification and whether the way people search influences the accuracy of their information evaluation. A mixed-methods approach was used, consisting of three parts: (1) an analysis of existing data to understand how search behaviour influences trust in fake news, (2) a simulation of query generation strategies using a Large Language Model (LLM) to assess the impact of different query formulations on search result quality, and (3) a user study to examine how 'Boost' interventions in interface design can guide users to adopt more effective query strategies. The results show that search behaviour significantly affects trust in news, with successful searches involving multiple queries and yielding higher-quality results. Queries inspired by different parts of a news article produced search results of varying quality, and weak initial queries improved when reformulated using full SERP information. Although 'Boost' interventions had limited impact, the study suggests that interface design encouraging users to thoroughly review search results can enhance query formulation. This study highlights the importance of query strategies in evaluating news and proposes that interface design can play a key role in promoting more effective search practices, serving as one component of a broader set of interventions to combat misinformation.
We present GERestaurant, a novel dataset consisting of 3,078 German language restaurant reviews manually annotated for Aspect-Based Sentiment Analysis (ABSA). All reviews were collected from Tripadvisor, covering a diverse selection of restaurants, including regional and international cuisine with various culinary styles. The annotations encompass both implicit and explicit aspects, including all aspect terms, their corresponding aspect categories, and the sentiments expressed towards them. Furthermore, we provide baseline scores for the four ABSA tasks Aspect Category Detection, Aspect Category Sentiment Analysis, End-to-End ABSA and Target Aspect Sentiment Detection as a reference point for future advances. The dataset fills a gap in German language resources and facilitates exploration of ABSA in the restaurant domain.
We present a study in the context of computational social science that explores the topics debated in the context of the 2021 German Federal Election by using the topic modeling technique BERTopic. The corpus consists of German language tweets posted by political party accounts of the major German parties, as well as tweets by the general public mentioning the party accounts. We examined the textual content of the tweets but also included the text in images that were posted into the analysis by extracting the text using optical character recognition (OCR). Our results show that the most frequently discussed topics are party-oriented policies (including call-to-action content), climate policy and financial policy, with these topics being discussed in tweets by both, the political party accounts and tweets by accounts mentioning them. In addition, we observed that some topics were discussed consistently throughout the year, such as the COVID-19 pandemic, climate policy or digitization, while other topics, such as the return to power of the Taliban in Afghanistan or Israel were debated to a greater extent at limited time frames during the election year.
In this work, we investigate the efficacy of boost interventions rooted in information literacy principles to enhance user interaction with search results and promote knowledge acquisition on debated topics. We conducted a pre-registered online user study with a between-groups design involving 351 participants who each completed knowledge assessments before and after performing a search task. In total, 9 boost conditions were tested (consisting of 4 search tips x 2 presentations and one control without a boost). Our findings indicate that the tested boost interventions successfully steer users towards examining a greater number of search results, investing more time in their search, and achieving a more equitable distribution of arguments presented for each side of a debated topic. Nonetheless, in terms of overall knowledge gain, the interventions do not yield a significant difference when compared to the baseline. The results underscore that boosts could be useful in effectively restricting some of the biases involved when users perform searches on debated topics, and should therefore be tested in more naturalistic settings. However, additional support mechanisms are essential if the goal is to enhance overall knowledge acquisition.
The applicability of retrieval algorithms to real data relies heavily on the quality of the training data. Currently, the creation process of training and test collections for retrieval systems is often based on annotations produced by human assessors following a set of guidelines. Some concepts, however, are prone to subjectivity, which could restrict the utility of any algorithm developed with the resulting data in real world applications. One such concept is credibility, which is an important factor in user’s judgements on whether retrieved information helps to answer an information need. In this paper, we evaluate an existing set of assessment guidelines with respect to their ability to generate reliable credibility judgements across multiple raters. We identify reasons for disagreement and adapt the guidelines to create an actionable and traceable annotation scheme that i) leads to higher inter-annotator reliability, and ii) can inform about why a rater made a specific credibility judgement. We provide promising evidence about the robustness of the new guidelines and conclude that they could be a valuable resource for building future test collections for misinformation detection.
Featured snippets that attempt to satisfy users’ information needs directly on top of the first search engine results page (SERP) have been shown to strongly impact users’ post-search attitudes and beliefs. In the context of debated but scientifically answerable topics, recent research has demonstrated that users tend to trust featured snippets to such an extent that they may reverse their original beliefs based on what such a snippet suggests; even when erroneous information is featured. This paper examines the effect of featured snippets in more nuanced and complicated search scenarios concerning debated topics that have no ground truth and where diverse arguments in favor and against can legitimately be made. We report on a preregistered, online user study (N = 182) investigating how the stances and logics of evaluation (i.e., underlying reasons behind stances) expressed in featured snippets influence post-task attitudes and explanations of users without strong pre-search attitudes. We found that such users tend to not only change their attitudes on debated topics (e.g., school uniforms) following whatever stance a featured snippet expresses but also incorporate the featured snippet’s logic of evaluation into their argumentation. Our findings imply that the content displayed in featured snippets may have large-scale undesired consequences for individuals, businesses, and society, and urgently call for researchers and practitioners to examine this issue further.
Search engines often provide featured snippets, which are boxed and placed above other results with the aim of directly answering user queries. To learn about how users judge the credibility of such results and how they influence search outcomes, a controlled web-based user study (N = 96) was conducted. Using resources made available by scholars in the community, we study featured snippets in a medical context with participants being tasked with determining whether a named treatment is helpful for a specified medical condition both before and after viewing the search results. Experimental conditions varied the presence and credibility of featured snippets. Our findings indicate that participants tend to overestimate the credibility of information in featured snippets. Featured snippets are, moreover, shown to often change users’ opinion about a topic, especially if they are uncertain. Showing correct information inside featured snippets helped participants make more accurate decisions, whereas incorrect or contradicting information led to more harmful outcomes.