We describe the University of Amsterdam’s participation in the WebCLEF track at CLEF 2005. We submitted runs for both the mixed monolingual task and the multilingual task.
EuroGOV is a multilingual web corpus that was created to serve as the document collection for WebCLEF, the CLEF 2005 web retrieval task. EuroGOV is a collection of web pages crawled from the European Union portal, European Union member state governmental web sites, and Russian government web sites. The corpus contains over 3 million documents written in more than 20 different European languages. In this paper we provide a detailed description of the EuroGOV collection.
In most social networks, measuring similarity between users is crucial for providing new functionalities, understanding the dynamics of such networks, and growing them (e.g., people you may know recommendations depend on similarity, as does link prediction). In this paper, we study a large sample of Flickr user actions and compare tags across different explicit and implicit network relations. In particular, we compare tag similarities in explicit networks (based on contact, friend, and family links), and implicit networks (created by actions such as comments and selecting favorite photos). We perform an in-depth analysis of these five types of links specifically focusing on tagging, and compare different tag similarity metrics. Our motivation is that understanding the differences in such networks, as well as how different similarity metrics perform, can be useful in similarity-based recommendation applications (e.g., collaborative filtering), and in traditional social network analysis problems (e.g., link prediction). We specifically show that different types of relationships require different similarity metrics. Our findings could lead to the construction of better user models, among others.
This is a preliminary report on the University of Amster- dam's participation in the INEX 2005 Interactive Track. We participated in Task A, a common baseline system with the IEEE collection, as well as in Task B, in which the baseline system is compared to a home-grown XML element retrieval system, xmlfind.
This paper describes MediaFaces, a system that enables faceted exploration of media collections. The system processes semi-structured information sources to extract objects and facets, e.g. the relationships between two objects. Next, we rank the facets based on a statistical analysis of image search query logs, and the tagging behaviour of users annotating photos in Flickr. For a given object of interest, we can then retrieve the top-k most relevant facets and present them to the user. The system is currently deployed in production by Yahoo!'s image search engine 1 . We present the system architecture, its main components, and the application of the system as part of the image search experience.
The research described in this paper forms the backbone of a service that enables the faceted search experience of the Yahoo! search engine. We introduce an approach for a machine learned ranking of entity facets based on user click feedback and features extracted from three different ranking sources. The objective of the learned model is to predict the click-through rate on an entity facet. In an empirical evaluation we compare the performance of gradient boosted decision trees (GBDT) against a linear combination of features on two different click feedback models using the raw click-through rate (CTR), and click over expected clicks (COEC). The results show a significant improvement in retrieval performance, in terms of discounted cumulated gain, when ranking entity facets with GBDT trained on the COEC model. Most notably this is true when evaluated against the CTR test set.
In this paper we address the task of recommending additional tags to partially annotated media objects, in our case images. We propose an extendable framework that can recommend tags using a combination of different personalised and collective contexts. We combine information from four contexts: (1) all the photos in the system, (2) a user's own photos, (3) the photos of a user's social contacts, and (4) the photos posted in the groups of which a user is a member. Variants of methods (1) and (2) have been proposed in previous work, but the use of (3) and (4) is novel. For each of the contexts we use the same probabilistic model and Borda Count based aggregation approach to generate recommendations from different contexts into a unified ranking of recommended tags. We evaluate our system using a large set of real-world data from Flickr. We show that by using personalised contexts we can significantly improve tag recommendation compared to using collective knowledge alone. We also analyse our experimental results to explore the capabilities of our system with respect to a user's social behaviour.
This paper presents a mobile software application for the provision of mobile guidance, supporting functionalities, which are based on automatically extracted Collective Intelligence. Collective Intelligence is the intelligence which emerges from the collaboration, competition and coordination among individuals and can be extracted by the analysis of mass amount of user-contributed data currently available in Web 2.0 applications. More specifically, services including automatic Point of Interest (POI) detection, raking, search and aggregation with semi-structured sources (e.g. Wikipedia) are developed, which are based on lexical and statistical analysis of mass data coming from Wikipedia, Yahoo! Geoplanet, query logs and flickr tags. These services together with personalization functionalities are integrated in a travel mobile application, enabling their efficient usage exploiting on the same time user location information. Evaluation with real users depicts the application's potential for providing a higher degree of satisfaction compared to existing travel information management solutions and also directions for future enhancements.
On photo sharing websites like Flickr and Zooomr, users are offered the possibility to assign tags to their uploaded pictures. Using these tags to find interesting groups of semantically related pictures in the result set of a given query is a problem with obvious applications. We analyse this problem from a Minimum Description Length (MDL) perspective and develop an algorithm that finds the most interesting groups. The method is based on Krimp, which finds small sets of patterns that characterise the data using compression. These patterns are sets of tags, often assignedtogether to photos. The better a database compresses, the more structure it contains and thus the more homogeneous it is. Following this observation we devise a compression-based measure. Our experiments on Flickr data show that the most interesting and homogeneous groups are found. We show extensive examples and compare to clusterings on the Flickr website.
Christof Monz合作论文数Department of Computer Science
Queen Mary;University of London5