The goal of query-focused summarization is to extract a summary for a given query from the document collection. Although much work has been done for this problem, there are still many challenging issues: (1) The length of the summary is predefined by, for example, the number of word tokens or the number of sentences. (2) A query usually asks for information of several perspectives (topics); however existing methods cannot capture topical aspects with respect to the query. In this paper, we propose a novel approach by combining statistical topic model and affinity propagation. Specifically, the topic model, called qLDA, can simultaneously model documents and the query. Moreover, the affinity propagation can automatically discover key sentences from the document collection without predefining the length of the summary. Experimental results on DUC05 and DUC06 data sets show that our approach is effective and the summarization performance is better than baseline methods.
A folksonomy refers to a collection of user-defined tags with which users describe contents published on the Web. With the flourish of Web 2.0, folksonomies have become an important mean to develop the Semantic Web. Because tags in folksonomies are authored freely, there is a need to understand the structure and semantics of these tags in various applications. In this paper, we propose a learning approach to create an ontology that captures the hierarchical semantic structure of folksonomies. Our experimental results on two different genres of real world data sets show that our method can effectively learn the ontology structure from the folksonomies.
Previous chapter Next chapter Full AccessProceedings Proceedings of the 2009 SIAM International Conference on Data Mining (SDM)Multi-topic based Query-oriented SummarizationJie Tang, Limin Yao, and Dewei ChenJie Tang, Limin Yao, and Dewei Chenpp.1148 - 1159Chapter DOI:https://doi.org/10.1137/1.9781611972795.98PDFBibTexSections ToolsAdd to favoritesExport CitationTrack CitationsEmail SectionsAboutAbstract Query-oriented summarization aims at extracting an informative summary from a document collection for a given query. It is very useful to help users grasp the main information related to a query. Existing work can be mainly classified into two categories: supervised method and unsupervised method. The former requires training examples, which makes the method limited to predefined domains. While the latter usually utilizes clustering algorithms to find 'centered' sentences as the summary. However, the method does not consider the query information, thus the summarization is general about the document collection itself. Moreover, most of existing work assumes that documents related to the query only talks about one topic. Unfortunately, statistics show that a large portion of summarization tasks talk about multiple topics. In this paper, we try to break limitations of the existing methods and study a new setup of the problem of multi-topic based query-oriented summarization. We propose using a probabilistic approach to solve this problem. More specifically, we propose two strategies to incorporate the query information into a probabilistic model. Experimental results on two different genres of data show that our proposed approach can effectively extract a multi-topic summary from a document collection and the summarization performance is better than baseline methods. The approach is quite general and can be applied to many other mining tasks, for example product opinion analysis and question answering. Previous chapter Next chapter RelatedDetails Published:2009ISBN:978-0-89871-682-5eISBN:978-1-61197-279-5 https://doi.org/10.1137/1.9781611972795Book Series Name:ProceedingsBook Code:PR133Book Pages:1-1244
In this paper, we study a novel problem of staring people discovery from social networks, which is concerned with finding people who are not only authoritative but also sociable in the social network. We formalize this problem as an optimization programming problem. Taking the co-author network as a case study, we define three objective functions and propose two methods to combine these objective functions. A genetic algorithm based method is further presented to solve this problem. Experimental results show that the proposed solution can effectively find the staring people from social networks.
Web services are new paradigm of using the Web. One of emerging challenges is to discover Web services efficiently and precisely. Unfortunately, existing systems like UDDI often have the problems of low availability, low accuracy, and bad performance. In this paper, we propose an approach of semantic Web services discovery in P2P environment. Firstly, Web services are published and deployed in same Web server to guarantee the availability. Secondly, Web servers are organized into groups to form a structured P2P network. Thirdly, for improving the query performance of service discovering, we propose a 2-layers searching algorithm. Finally we developed a prototype system, and tested the system in real P2P test bed of Planet-lab. Experimental results show our approach's precision, recall, query time and scalability. By analysis of the performance, we observed that the system has good efficiency and scalability.
Web Services constitute a new computing model for Web application. The application based on SOA is a promising trend of distributed computing. The key and most difficult problem in SOA is how to automatically discover the services according to the end users’ query accurately and quickly. A new P2P and Semantic Web based service discovery mechanism is discussed in this paper in which the deployment and publication of a Web Service are bound together. When web service manager deploys his services, the embedded toolkit will create web service description files automatically and put them into the service metadata repository of the Peer. Profiting from using P2P network to exchange the metadata, the service providers are permitted to add and modify and delete services freely without conferring others. When a query is submitted by a service consumer, a two steps querying and two layers searching methods are used to improve the performance. Key words based matching and indexing in the Group layer are used during the first step, and semantic based Goal Capability matching in the Peer layer is used during the second step of querying. Two metrics, service growth time and service death time are also introduced in this paper to evaluate the discovery performance.
Semantic based Web services integration aims to add valuable semantic information in Web services Description and to implement the semantic based Web services discovery and integration. This paper puts forward a semantic based Web service integration framework in P2P-SEWSIP. SEWSIP can collect Web services in distributed web environment, lift the WSDL files into OWL-S by the use of XML schema in WSDL and domain ontology, and then deploy them in P2P. It also can provide service discovery, evaluation and composition based on semantic information in OWL-S.
Jie Tang (唐杰)合作论文数Department of Computer Science and Technology, Tsinghua University4