Prediction of new outlinks for focused Web crawling

Thi Kim Nhung Dang,Doina Bucur,Berk Atil,Guillaume Pitel,Frank Ruis,Hamidreza Kadkhodaei,Nelly Litvak University of Twente, The Netherlands, Bogazici University, Turkey, Exensa, France, Eindhoven University of Technology

ArXiv(2021)

引用 0|浏览8
暂无评分
摘要
Discovering new hyperlinks enables Web crawlers to find new pages that have not yet been indexed. This is especially important for focused crawlers because they strive to provide a comprehensive analysis of specific parts of the Web, thus prioritizing discovery of new pages over discovery of changes in content. In the literature, changes in hyperlinks and content have been usually considered simultaneously. However, there is also evidence suggesting that these two types of changes are not necessarily related. Moreover, many studies about predicting changes assume that long history of a page is available, which is unattainable in practice. The aim of this work is to provide a methodology for detecting new hyperlinks effectively using a short history. To this end, we use a dataset of ten crawls at intervals of one week. Our study consists of three parts. First, we obtain insight in the data by analyzing empirical properties of the number of new outlinks. We observe that these properties are, on average, stable over time, but there is a large difference between emergence of hyperlinks towards pages within and outside the domain of a target page (internal and external outlinks, respectively). Next, we provide statistical models for three targets: the link change rate, the presence of new links, and the number of new links. These models include the features used earlier in the literature, as well as new features introduced in this work. We analyze correlation between the features, and investigate their informativeness. A notable finding is that, if the history of the target page is not available, then our new features, that represent the history of related pages, are most predictive for new hyperlinks in the target page. Finally, we propose ranking methods as guidelines for focused crawlers to efficiently discover new pages, and demonstrate that they achieve excellent performance with respect to the corresponding targets.
更多
查看译文
关键词
focused web crawling,new outlinks,prediction
AI 理解论文
溯源树
样例
生成溯源树,研究论文发展脉络
Chat Paper
正在生成论文摘要