Evidence based on web graph structure is reportedly used by the current generation of World-Wide Web (WWW) search engines to identify “high-quality”, “important” pages and to reject “spam” content. However, despite the apparent wide use of this evidence its application in web-based document retrieval is controversial. Confusion exists as to how to incorporate web evidence in document ranking, and whether such evidence is in fact useful. This thesis demonstrates how web evidence can be used to improve retrieval effectiveness for navigational search tasks. Fundamental questions investigated include: which forms of web evidence are useful, how web evidence should be combined with other document evidence, and what biases are present in web evidence. Through investigating these questions, this thesis presents a number of findings regarding how web evidence may be effectively used in a general-purpose web-based document ranking algorithm. The results of experimentation with well-known forms of web evidence on several small-to-medium collections of web data are surprising. Aggregate anchor-text measures perform well, but well-studied hyperlink recommendation algorithms are far less useful. Further gains in retrieval effectiveness are achieved for anchor-text measures by revising traditional full-text ranking methods to favour aggregate anchor-text documents containing large volumes of anchor-text. For home page finding tasks additional gains are achieved by including a simple URL depth measure which favours short URLs over long ones. The most effective combination of evidence treats document-level and web-based evidence as separate document components, and uses a linear combination to sum scores. It is submitted that the document-level evidence contains the author’s description of document contents, and that the web-based evidence gives the wider web community view of the document. Consequently if both measures agree, and the document is scored highly in both cases, this is a strong indication that the page is what it claims to be. A linear combination of the two types of evidence is found to be particularly effective, achieving the highest retrieval effectiveness of any query-dependent evidence on navigational and Topic Distillation tasks. However, care should be taken when using hyperlink-based evidence as a direct measure of document quality. Thesis experiments show the existence of bias towards the home pages of large, popular and technology-oriented companies. Further empirical evidence is presented to demonstrate how the authorship of web documents and sites directly affects the quantity and quality of available web evidence. These factors demonstrate the need for robust methods for mining and interpreting data from the web graph.
Hyperlink recommendation information is used by current search engines to provide a query-independent measure of document “importance”, and thus improve retrieval effectiveness [1]. However, as yet the research community has been unable to achieve the purported benefits of this evidence [5]. Effective evaluation of query-independent methods is hindered by the many ways in which this evidence can be combined with query-dependent evidence. In this paper we perform an initial study of how query-independent hyperlink recommendation evidence can be effectively combined with query-dependent baselines. We investigate whether hyperlink recommendation evidence is best incorporated as some kind of threshold, providing a minimum basis for query inclusion, or whether it should be included as a component in the document scoring function. Moreover, we investigate whether this evidence should play a large role in choosing candidates for retrieval, or simply be used to re-shuffle or “jitter” documents that already achieve high query dependent scores. PageRank is believed to be an important component of Google’s ranking algorithm and is used to provide a measure as to ‘whether other people on the web consider a page to be a high-quality site worth checking out’ [1]. A number of systematic biases present in query-independent link evidence on the Web have previously been reported in [4]. A bias towards homepages was observed, which is important in navigational search, however the use of URL length measures has previously been observed to provide superior gains [5]. PageRank has been observed to be highly correlated with indegree [5, 4], therefore we do not consider both measures in this work. In these experiments we eliminate the effect of homepage bias to investigate weaknesses introduced through the use of PageRank in navigational search. The task we examine is: Given the name of a company how well different methods retrieve that company’s homepage from the set of all company homepages. We compute three baselines which are then re-ranked by PageRank; content, anchor-text and a combination of content and anchor-text. We retrieved a set of candidate companies from three American stock exchanges; NYSE, NASDAQ and AMEX. For each company we sourced the homepage URL, stock name and homepage content (URLs sourced from http:
Okapi BM25 scoring of anchor text surrogate documents has been shown to facilitate effective ranking in navigational search tasks over web data. We hypothesize that even better ranking can be achieved in certain important cases, particularly when anchor scores must be fused with content scores, by avoiding length normalisation and by reducing the attentuation of scores associated with high tf . Preliminary results are presented.
Link information, especially anchor text, is known to be very useful for effective ranking of web pages, particularly in response to navigational queries. We investigated whether enterprise webs contain sufficient internal link information to adequately answer queries derived from the enterprise's site map or, alternatively, whether adding link evidence from the external Web can boost search effectiveness. Using 1266 navigational queries derived from Stanford University's A-Z site index, we found no difference between the quality of results returned by Stanford's Google appliance and those from an appropriately site-restricted search of the global Google service. Applying similar methodology to our own crawls of seven Australian organisations, we found that adding external link evidence made no significant difference to search effectiveness in five cases and a slight difference (in different directions) in the other two. We observed that external links to an organisation show very different patterns to internal links. Unlike enterprise web publishers, external web authors heavily favour directory default pages, particularly the organisation's home page and pages offering information or services likely to be useful on an ongoing basis. External links seldom reference the complex, parameterised URLs in common use in many organisations.
David Hawking合作论文数Australian National University9
John Blitzer合作论文数Google Research2
James Thom合作论文数RMIT University1