
Gender bias is a pervasive issue that continues to influence various aspects of society, including the outcomes of information retrieval (IR) systems. As these systems become increasingly integral to accessing and navigating the vast amounts of information available today, the need to understand and mitigate gender bias within them is paramount. This monograph provides a comprehensive examination of the origins, manifestations, and consequences of gender bias in IR systems, as well as the current methodologies employed to address these biases. Theoretical frameworks surrounding gender and its representation in artificial intelligence (AI) systems are explored, particularly focusing on how traditional gender binaries are perpetuated and reinforced through data and algorithmic processes. Metrics and methodologies used to identify and measure gender bias within IR systems are then analyzed, offering a detailed evaluation of existing approaches and their limitations. Subsequent sections address the sources of gender bias, including biased input queries, retrieval methods, and gold standard datasets. Various data-driven and method-level debiasing strategies are presented, including techniques for debiasing neural embeddings and algorithmic approaches aimed at reducing bias in IR system outputs. The monograph concludes with a discussion of the challenges and limitations faced by current debiasing efforts and provides insights into future research directions that could lead to more equitable and inclusive IR systems. This monograph serves as a valuable resource for researchers, practitioners, and students in the fields of information retrieval, artificial intelligence, and data science, providing the knowledge and tools needed to address gender bias and contribute to the development of fair and unbiased information systems.
Text classification stands as a cornerstone within the realm of Natural Language Processing (NLP), particularly when viewed through computer science and engineering. The past decade has seen deep learning revolutionize text classification, propelling advancements in text retrieval, categorization, information extraction, and summarization. The scholarly literature includes datasets, models, and evaluation criteria, with English being the predominant language of focus, despite studies involving Arabic, Chinese, Hindi, and others. The efficacy of text classification models relies heavily on their ability to capture intricate textual relationships and non-linear correlations, necessitating a comprehensive examination of the entire text classification pipeline. In the NLP domain, a plethora of text representation techniques and model architectures have emerged, with Large Language Models (LLMs) and Generative Pre-trained Transformers (GPTs) at the forefront. These models are adept at transforming extensive textual data into meaningful vector representations encapsulating semantic information. The multidisciplinary nature of text classification, encompassing data mining, linguistics, and information retrieval, highlights the importance of collaborative research to advance the field. This work integrates traditional and contemporary text mining methodologies, fostering a holistic understanding of text classification. This monograph provides an in-depth exploration of the text classification pipeline, with a particular emphasis on evaluating the impact of each component on the overall performance of text classification models. The pipeline includes state-of-the-art datasets, text preprocessing techniques, text representation methods, classification models, evaluation metrics, and future trends. Each section examines these stages, presenting technical innovations and recent findings. The work assesses various classification strategies, offering comparative analyses, examples and case studies. These contributions extend beyond a typical survey, providing a detailed and insightful exploration of the field.
Search systems are often designed to support simple look-up tasks, such as fact-finding and navigation tasks. However, people increasingly use search engines to complete tasks that require deeper learning. In recent years, the search as learning (SAL) research community has argued that search systems should also be designed to support information- seeking tasks that involve complex learning as an important outcome. This monograph aims to provide a comprehensive review of prior research in search as learning and related areas. Searching to learn can be characterized by specific learning objectives, strategies, and context. Therefore, we begin by reviewing research in education that has aimed at characterizing learning objectives, strategies, and context. Then, we review methods used in prior studies to measure learning during a search session. Here, we discuss two important recommendations for future work: (1) measuring learning retention and (2) measuring a learner's ability to transfer their new knowledge to a novel scenario. Following this, we discuss studies that have focused on understanding factors that influence learning during search and search behaviors that are predictive of learning. Next, we survey tools that have been developed to support learning during search. Searching for the purpose of learning is often a solitary activity. Research in self-regulated learning (SRL) aims to understand how people monitor and control their own learning. Therefore, we review existing models of SRL, methods to measure engagement with specific SRL processes, and tools to support effective SRL. We conclude by discussing potential areas for future research.
Mathematical information is essential for technical work, but its creation, interpretation, and search are challenging. To help address these challenges, researchers have developed multimodal search engines and mathematical question answering systems. This monograph begins with a simple framework characterizing the information tasks that people and systems perform as we work to answer math-related questions. The framework is used to organize and relate the other core topics of the monograph, including interactions between people and systems, representing math formulas in sources, and evaluation. We close by addressing some key questions and presenting directions for future work. This monograph is intended for students, instructors, and researchers interested in systems that help us find and use mathematical information.
Recent studies have shown that information retrieval systems may exhibit stereotypical gender biases in outcomes which may lead to discrimination against minority groups, such as different genders, and impact users' decisionmaking and judgements. In this tutorial, we inform the audience of studies that have systematically reported the presence of stereotypical gender biases in Information Retrieval (IR) systems and different pre-trained Natural Language Processing (NLP) models. We further classify existing work on gender biases in IR systems and NLP models as being related to (1) relevance judgement datasets, (2) structure of retrievalmethods, (3) representations learnt for queries and documents, (4) and pretrained embedding models. Based on the aforementioned categories, we present a host of methods from the literature that can be leveraged tomeasure, control, or mitigate the existence of stereotypical biases within IR systems and different NLP models that are used for down-stream tasks. Besides, we introduce available datasets and collections that are widely used for studying the existence of gender biases in IR systems and NLP models, the evaluation metrics that can be used for measuring the level of bias and utility of the models, and de-biasing methods that can be leveraged to mitigate gender biases within those models.
A contract is an economic tool used by a principal to incentivize one or more agents to exert effort on her behalf, by defining payments based on observable performance measures. A key challenge addressed by contracts known in economics as moral hazard is that, absent a properly set up contract, agents might engage in actions that are not in the principal's best interest. Another common feature of contracts is limited liability, which means that payments can go only from the principal who has the deep pocket to the agents. With classic applications of contract theory moving online, growing in scale, and becoming more data-driven, tools from contract theory become increasingly important for incentive- aware algorithm design. At the same time, algorithm design offers a whole new toolbox for reasoning about contracts, ranging from additional tools for studying the tradeoff between simple and optimal contracts, through a language for discussing the computational complexity of contracts in combinatorial settings, to a formalism for analyzing data-driven contracts. This survey aims to provide a computer science-friendly introduction to the basic concepts of contract theory. We give an overview of the emerging field of "algorithmic contract theory" and highlight work that showcases the potential for interaction between the two areas. We also discuss avenues for future research.
Search engines play a crucial role in organizing and delivering information to billions of users worldwide. However, these systems often reflect and amplify existing societal biases and stereotypes through their search results and rankings. This concern has prompted researchers to investigate methods for measuring and reducing algorithmic bias, with the goal of developing more equitable search systems. This monograph presents a comprehensive taxonomy of fairness in search systems and surveys the current research landscape. We systematically examine how bias manifests across key search components, including query interpretation and processing, document representation and indexing, result ranking algorithms, and system evaluation metrics. By critically analyzing the existing literature, we identify persistent challenges and promising research directions in the pursuit of fairer search systems. Our aim is to provide a foundation for future work in this rapidly evolving field while highlighting opportunities to create more inclusive and equitable information retrieval technologies.
With the emergence of various information access systems exhibiting increasing complexity, there is a critical need for sound and scalable means of automatic evaluation. To address this challenge, user simulation emerges as a promising solution. This half-day tutorial focuses on providing a thorough understanding of user simulation techniques designed specifically for evaluation purposes. We systematically review major research progress, covering both general frameworks for designing user simulators, and specific models and algorithms for simulating user interactions with search engines, recommender systems, and conversational assistants. We also highlight some important future research directions.
E-commerce (electronic commerce or EC) is the buying and selling of goods and services, or the transmitting of funds or data online. E-commerce platforms come in many kinds, with global players such as Amazon, Airbnb, Alibaba, eBay, JD.com and platforms targeting specific markets such as Bol.com and Booking.com. Information retrieval has a natural role to play in e-commerce, especially in connecting people to goods and services. Information discovery in e-commerce concerns different types of search (exploratory search vs. lookup tasks), recommender systems, and natural language processing in e-commerce portals. Recently, the explosive popularity of e-commerce sites has made research on information discovery in e-commerce more important and more popular. There is increased attention for e-commerce information discovery methods in the community as witnessed by an increase in publications and dedicated workshops in this space. Methods for information discovery in e-commerce largely focus on improving the performance of e-commerce search and recommender systems, on enriching and using knowledge graphs to support e-commerce, and on developing innovative question-answering and bot-based solutions that help to connect people to goods and services. Below we describe why we believe that the time is right for an introductory tutorial on information discovery in e-commerce, the objectives of the proposed tutorial, its relevance, as well as more practical details, such as the format, schedule and support materials.
The task of Question Answering (QA) has attracted significant research interest for long. Its relevance to language understanding and knowledge retrieval tasks, along with the simple setting makes the task of QA crucial for strong AI systems. Recent success on simple QA tasks has shifted the focus to more complex settings. Among these, Multi-Hop QA (MHQA) is one of the most researched tasks over the recent years. In broad terms, MHQA is the task of answering natural language questions that involve extracting and combining multiple pieces of information and doing multiple steps of reasoning. An example of a multi-hop question would be "The Argentine PGA Championship record holder has won how many tournaments worldwide?". Answering the question would need two pieces of information: "Who is the record holder for Argentine PGA Championship tournaments?" and "How many tournaments did [Answer of Sub Q1] win?". The ability to answer multi-hop questions and perform multi step reasoning can significantly improve the utility of NLP systems. Consequently, the field has seen a surge with high quality datasets, models and evaluation strategies. The notion of 'multiple hops' is somewhat abstract which results in a large variety of tasks that require multi-hop reasoning. This leads to different datasets and models that differ significantly from each other and makes the field challenging to generalize and survey. We aim to provide a general and formal definition of the MHQA task, and organize and summarize existing MHQA frameworks. We also outline some best practices for building MHQA datasets. This book provides a systematic and thorough introduction as well as the structuring of the existing attempts to this highly interesting, yet quite challenging task.
This monograph takes a step towards promoting the study of efficiency in the era of neural information retrieval by offering a comprehensive survey of the literature on efficiency and effectiveness in ranking, and to a limited extent, retrieval. This monograph was inspired by the parallels that exist between the challenges in neural network-based ranking solutions and their predecessors, decision forest-based learning to rank models, as well as the connections between the solutions the literature to date has to offer. We believe that by understanding the fundamentals underpinning these algorithmic and data structure solutions for containing the contentious relationship between efficiency and effectiveness, one can better identify future directions and more efficiently determine the merits of ideas. We also present what we believe to be important research directions in the forefront of efficiency and effectiveness in retrieval and ranking.
This monograph offers a survey of work to date to inform how interactions in information retrieval systems could afford inclusion of users who are neurodiverse. This existing work is positioned within a range of philosophies, frameworks and epistemologies which frame the importance of including neurodiverse users in all stages of research and development of Interactive Information Retrieval (IIR) systems. The monograph also offers examples and practical approaches to include neurodiverse users in IIR research, and explores the challenges ahead in the field.
The introduction of Quantum Theory (QT) provides a unified mathematical framework for Information Retrieval (IR). Compared with the classical IR framework, the quantum-inspired IR framework is based on user-centered modeling methods to model non-classical cognitive phenomena in human relevance judgment in the IR process. With the increase of data and computing resources, neural IR methods have been applied to the text matching and understanding task of IR. Neural networks have a strong learning ability of effective representation and generalization of matching patterns from raw data. However, these methods show some unavoidable defects, such as the inability to model user cognitive phenomena, large number of model parameters and the "black box" characteristics of network structure. These problems greatly limit the development of neural IR and related fields. Although the quantum-inspired retrieval framework can theoretically solve the above problems, it is faced with problems such as poor model efficiency and difficulty in integrating with neural network, which lead to a huge gap between QT and neural network modeling. This review gives a systematic introduction to quantum-inspired neural IR, including quantum-inspired neural language representation, matching and understanding. This is not only helpful to non-classical phenomena modeling in IR but also to break the theoretical bottleneck of neural networks and design more transparent neural IR models. We introduce the language representation method based on QT and the quantum-inspired text matching and decision making model under neural network, which shows its theoretical advantages in document ranking, relevance matching, multimodal IR, and can be integrated with neural networks to jointly promote the development of IR. The latest progress of quantum language understanding is introduced and further topics on QT and language modeling provide readers with more materials for thinking.
Conversational information seeking (CIS) is concerned with a sequence of interactions between one or more users and an information system. Interactions in CIS are primarily based on natural language dialogue, while they may include other types of interactions, such as click, touch, and body gestures. This monograph provides a thorough overview of CIS definitions, applications, interactions, interfaces, design, implementation, and evaluation. This monograph views CIS applications as including conversational search, conversational question answering, and conversational recommendation. Our aim is to provide an overview of past research related to CIS, introduce the current state-of-the-art in CIS, highlight the challenges still being faced in the community. and suggest future directions.
The core of information retrieval (IR) is to identify relevant information from large-scale resources and return it as a ranked list to respond to the user's information need. In recent years, the resurgence of deep learning has greatly advanced this field and leads to a hot topic named NeuIR (i.e., neural information retrieval), especially the paradigm of pre-training methods (PTMs). Owing to sophisticated pre-training objectives and huge model size, pre-trained models can learn universal language representations from massive textual data, which are beneficial to the ranking task of IR. Recently, a large number of works, which are dedicated to the application of PTMs in IR, have been introduced to promote the retrieval performance. Considering the rapid progress of this direction, this survey aims to provide a systematic review of pre-training methods in IR. To be specific, we present an overview of PTMs applied in different components of an IR system, including the retrieval component, the re-ranking component, and other components. In addition, we also introduce PTMs specifically designed for IR, and summarize available datasets as well as benchmark leaderboards. Moreover, we discuss some open challenges and highlight several promising directions, with the hope of inspiring and facilitating more works on these topics for future research.
With the rapid progress of deep neural models and the explosion of available data resources, dialogue systems that supports extensive topics and chit-chat conversations are emerging as a research hot-spot for many communities, e.g., information retrieval (IR), natural language processing (NLP), and machine learning (ML). Building a chit-chat system with retrieval techniques is an essential task and has achieved great success in the past few years. The advance of chit-chat systems, in turn, can support extensive IR tasks, e.g., conversational search and conversational recommendation. To facilitate the development of both retrieval-based chit-chat systems and IR tasks supported by these systems, we survey chit-chat systems from two perspectives: (1) techniques to build chit-chat systems, i.e., deep retrieval-based models, generative methods, and their ensembles, and (2) chit-chat components in completing IR tasks. In each aspect, we present cutting-edge neural methods and summarize the core challenges encountered and possible research directions.
Recommendation, information retrieval, and other information access systems pose unique challenges for investigating and applying the fairness and non-discrimination concepts that have been developed for studying other machine learning systems. While fair information access shares many commonalities with fair classification, the multistakeholder nature of information access applications, the rank-based problem setting, the centrality of personalization in many cases, and the role of user response complicate the problem of identifying precisely what types and operationalizations of fairness may be relevant, let alone measuring or promoting them. In this monograph, we present a taxonomy of the various dimensions of fair information access and survey the literature to date on this new and rapidly-growing topic. We preface this with brief introductions to information access and algorithmic fairness, to facilitate use of this work by scholars with experience in one (or neither) of these fields who wish to learn about their intersection. We conclude with several open problems in fair information access, along with some suggestions for how to approach research in this space.
Personalized recommender systems have become indispensable in today's online world. Most of today's recommendation algorithms are data-driven and based on behavioral data. While such systems can produce useful recommendations, they are often uninterpretable, black-box models, which do not incorporate the underlying cognitive reasons for user behavior in the algorithms' design. The aim of this survey is to present a thorough review of the state of the art of recommender systems that leverage psychological constructs and theories to model and predict user behavior and improve the recommendation process. We call such systems psychology-informed recommender systems. The survey identifies three categories of psychology-informed recommender systems: cognition-inspired, personality-aware, and affect-aware recommender systems. Moreover, for each category, we highlight domains, in which psychological theory plays a key role and is therefore considered in the recommendation process. As recommender systems are fundamental tools to support human decision making, we also discuss selected decision-psychological phenomena that impact the interaction between a user and a recommender. Besides, we discuss related work that investigates the evaluation of recommender systems from the user perspective and highlight user-centric evaluation frameworks. We discuss potential research tasks for future work at the end of this survey.
This monograph reviews research on the design and evaluation of search user interfaces that has been published within the past 10 years. Our primary goal is to integrate state-of-the-art research in the areas of information seeking behavior, information retrieval, and human-computer interaction on the topic of search interface. Specifically, this monograph (1) describes the history and background of the development of the search interface; (2) introduces information search behavior models that help conceptualize users’ information needs, and how people seek, select, and use information; (3) characterizes the major components of search interfaces that support different subprocesses based on Marchonini’s information seeking process model; (4) reviews the design of search interfaces for different user groups, especially that of vulnerable people, as well as personalized and adaptive search interfaces; (5) identifies evaluation methods of search interfaces and how they were implemented in research having different evaluation purposes. We also provide an outlook on the future trends of search interfaces including conversational search interfaces, search interfaces supporting serendipity and creativity, and searching in immersive and virtual reality environments.
Email has been an essential communication medium for many years. As a result, the information accumulated in our mailboxes has become valuable for all of our personal and professional activities. For years, researchers have developed interfaces, models, and algorithms to facilitate email search, discovery, and organization. This tutorial brings together these diverse research directions and provides both a historical background, as well as a high-level overview of the recent advances in the field. In particular, we lay out all of the components needed in the design of email search engines, including user interfaces, indexing, document and query understanding, retrieval, ranking, evaluation, and data privacy. The tutorial also goes beyond search, presenting recent work on intelligent task assistance in email and a number of interesting future directions.