Data-Driven Evaluation Metrics for Heterogeneous Search Engine Result Pages

Leif Azzopardi,Ryen W. White,Paul Thomas,Nick Craswell

CHIIR '20: Conference on Human Information Interaction and Retrieval Vancouver BC Canada March, 2020（2020）

引用 12|浏览128

暂无评分

摘要

Evaluation metrics for search typically assume items are homoge- neous. However, in the context of web search, this assumption does not hold. Modern search engine result pages (SERPs) are composed of a variety of item types (e.g., news, web, entity, etc.), and their influence on browsing behavior is largely unknown. In this paper, we perform a large-scale empirical analysis of pop- ular web search queries and investigate how different item types influence how people interact on SERPs. We then infer a user browsing model given people's interactions with SERP items - creating a data-driven metric based on item type. We show that the proposed metric leads to more accurate estimates of: (1) total gain, (2) total time spent, and (3) stopping depth - without requiring extensive parameter tuning or a priori relevance information. These results suggest that item heterogeneity should be accounted for when de- veloping metrics for SERPs. While many open questions remain concerning the applicability and generalizability of data-driven metrics, they do serve as a formal mechanism to link observed user behaviors directly to how performance is measured. From this approach, we can draw new insights regarding the relationship between behavior and performance - and design data-driven metrics based on real user behavior rather than using metrics reliant on some hypothesized model of user browsing behavior.

查看译文

AI 理解论文

溯源树

样例

生成溯源树，研究论文发展脉络

Chat Paper

正在生成论文摘要