With the overwhelming volume of news stories created and stored electronically everyday, there is an increasing need for techniques to analyze and present news stories to the users in a more meaningful manner. Most previous research focus on organizing news set into flat collections (topics) of stories. However, a topic in news is more than a mere collection of stories: it is actually characterized by a definite structure of inter-related events. Unfortunately, it is very difficult to identify events and dependencies within a topic because stories about the same topic are usually very similar to each other irrespective of the events they belong to. This is because stories within a topic usually share some terms which are related to the topic other than a specific event. To deal with this problem, we propose two methods based on event key terms to identify events and discover event dependency accurately. For event identification, we first capture some tight term clusters as term committees of potential events, and then use them to find the core story sets of potential events. At last we assign all stories to an event. For event dependency discovery, we emphasize the terms closely related to a certain event. So similarity contributed by topic-popular terms can be decreased. The experimental results on two Linguistic Data Consortium (LDC) datasets show that both the proposed methods for event identification and event dependency discovery have significant improvement over previous methods. Categories and Subject Descriptors: H.3.3 [Information Systems]: Information Search and Retrieval; H. 4.2 [Information Systems Applications]: Types of Systems – decision support. General Terms: Algorithms, Experimentation
Recent years, the amount of semi-structured documents available electrically has increased dramatically. Semi-structured documents usually are difficult to reuse due to the lack of explicit metadata. To enable integration and retrieval over semi-structured documents, the essential aspects in the documents should be described by metadata explicitly. The metadata could be assigned to documents and present part of their information content using various IE techniques. This paper also provides flexible user interaction mechanism to achieve better performance over less training sample documents. In semantic view extraction, by using similarity based rule induction, we have been able to improve the rule learning procedure. Experimental results show that our approach can significantly outperform most of the existing wrapper methods. We make use of the semantics that resides in document logical structure to help find relations between semantic entities. After semantic annotations of the documents, TIPSI allows those to be indexed with respect to the extracted text entities. To answer the query, TIPSI applies semantic restrictions over the entities in the KB.
Sentiment analysis from large-scale networked data attracts increasing attention in recent years. Most previous works on sentiment prediction mainly focus on text or image data. However, voice is the most natural and direct way to express people's sentiments in real-time. With the rapid development of smart phone voice dialogue applications (e.g., Siri and Sogou Voice Assistant), the large-scale networked voice data can help us better quantitatively understand the sentimental world we live in. In this paper, we study the problem of sentiment prediction from large-scale networked voice data. In particular, we first investigate the data observations and underlying sentiment patterns in human-mobile voice communication. Then we propose a deep sparse neural network (DSNN) model to incorporate acoustic features, content information and geo-information to automatically predict sentiments. The effectiveness of the proposed model is verified by the experiments on a real dataset from Sogou Voice Assistant application.
Creating knowledge bases based on the crowd-sourced wikis, like Wikipedia, has attracted significant research interest in the field of intelligent Web. However, the derived taxonomies usually contain many mistakenly imported taxonomic relations due to the difference between the user-generated subsumption relations and the semantic taxonomic relations. Current approaches to solving the problem still suffer the following issues: (i) the heuristic-based methods strongly rely on specific language dependent rules. (ii) the corpus-based methods depend on a large-scale high-quality corpus, which is often unavailable. In this paper, we formulate the cross-lingual taxonomy derivation problem as the problem of cross-lingual taxonomic relation prediction. We investigate different linguistic heuristics and language independent features, and propose a cross-lingual knowledge validation based dynamic adaptive boosting model to iteratively reinforce the performance of taxonomic relation prediction. The proposed approach successfully overcome the above issues, and experiments show that our approach significantly outperforms the designed state-of-the-art comparison methods.
With the development of social networks, more and more users have a great need to search for people to follow (SPTF) to receive their tweets. According to our experiments, approximately 50% of social networks' lost users leave due to a lack of people to follow. In this paper, we define the problem of SPTF and propose an approach to give users tags and then deliver a ranked list of valuable accounts for them to follow. In the proposed approach, we first seek accounts related to keywords via expanding and predicting tags for users. Second, we propose two algorithms to rank relevant accounts: the first mines the forwarded relationship, and the second incorporates the following relationship into PageRank. Accordingly, we have built a search system that to date, has received more than 1.7 million queries from 0.2 million users. To evaluate the proposed approach, we created a crowd-sourcing organization and crawled 0.25 billion profiles, 15 billion messages and 20 billion links representing following relationships on Sina Microblog. The empirical study validates the effectiveness of our algorithms for expanding and predicting tags compared to the baseline. From query logs, we discover that hot queries include keywords related to academics, occupations and companies. Experiments on those queries show that PageRank-like algorithms perform best for occupation-related queries, forward-relationship-like algorithms work best for academic-related queries and domain-related headcount algorithms work best for company-related queries.
用户满意度是以用户为中心的搜索引擎性能评价的一个重要分支,区别于传统基于查询与文档相关性的评价方法,基于用户满意度的性能评价能够更加全面、客观地对搜索引擎性能进行评价。该文通过设计搜索实验平台,在尽量不影响用户正常搜索过程的前提下收集用户的搜索行为及其满意度评价,通过用户行为分析的方法挖掘用户群体行为特征与用户查询满意度之间的关联关系。相关结论对提高搜索引擎性能、改善用户查询体验具有一定的参考意义。
信息检索的效果很大程度上取决于用户能否输入恰当的查询来描述自身信息需求。很多查询通常简短而模糊,甚至包含噪音。查询推荐技术可以帮助用户提炼查询、准确描述信息需求。为了获得高质量的查询推荐,在大规模"查询-链接"二部图上采用随机漫步方法产生候选集合。利用摘要点击信息对候选列表进行重排序,使得体现用户意图的查询排在比较高的位置。最终采用基于学习的算法对推荐查询中可能存在的噪声进行过滤。基于真实用户行为数据的实验表明该方法取得了较好的效果。
Emotions are increasingly and controversially central to our public life. Compared to text or image data, voice is the most natural and direct way to express ones’ emotions in real-time. With the increasing adoption of smart phone voice dialogue applications (e.g., Siri and Sogou Voice Assistant), the large-scale networked voice data can help us better quantitatively understand the emotional world we live in. In this paper, we study the problem of inferring public emotions from large-scale networked voice data. In particular, we first investigate the primary emotions and the underlying emotion patterns in human-mobile voice communication. Then we propose a partially-labeled factor graph model (PFG) to incorporate both acoustic features (e.g., energy, f0, MFCC, LFPC) and correlation features (e.g., individual consistency, time associativity, environment similarity) to automatically infer emotions. We evaluate the proposed model on a real dataset from Sogou Voice Assistant application. The experimental results verify the effectiveness of the proposed model.
To address self-tagging concerns, some social networks' websites, such as LinkedIn and Sina Weibo, allow users to tag themselves as part of their profiles; however, due to privacy or other unknown reasons, most of the users take just a few tags. Self-tag sparsity refers to the problem of low recall obtained when searching for people on systems based on user profiles. In this paper, we use not only users' self-tags but also their friend relationships (which are often not hidden) to expand the tag list and measure the effectiveness of different types of friendship links and their self-tags. Experimental results show that friendship information (friendship links and profiles) can effectively improve the performance of tag expansion, especially for common users who have limited followers.
Click models are developed to interpret clicks by making assumptions on how users browse the search result page. Most existing click models implicitly assume that all users are homogeneous and act in the same way when browsing the search results. However, a number of researches have shown that users have diverse behavioral patterns, which is also observed in this paper by eye-tracking experiments and click-through log analysis. As a uniform click model for all users can hardly capture the diverse click behavior, in this paper we incorporate user preferences into both a variety of existing click models and a novel click model. The experimental results on a large-scale click-through data set show consistent and significant performance improvement of the click models with user preferences integrated.
In modern search engines, an increasing number of search result pages (SERPs) are federated from multiple specialized search engines (called verticals, such as Image or Video). As an effective approach to interpret users' click-through behavior as feedback information, most click models were designed to reduce the position bias and improve ranking performance of ordinary search results, which have homogeneous appearances. However, when vertical results are combined with ordinary ones, significant differences in presentation may lead to user behavior biases and thus failure of state-of-the-art click models. With the help of a popular commercial search engine in China, we collected a large scale log data set which contains behavior information on both vertical and ordinary results. We also performed eye-tracking analysis to study user's real-world examining behavior. According these analysis, we found that different result appearances may cause different behavior biases both for vertical results (local effect) and for the whole result lists (global effect). These biases include: examine bias for vertical results (especially those with multimedia components), trust bias for result lists with vertical results, and a higher probability of result revisitation for vertical results. Based on these findings, a novel click model considering these biases besides position bias was constructed to describe interaction with SERPs containing verticals. Experimental results show that the new Vertical-aware Click Model (VCM) is better at interpreting user click behavior on federated searches in terms of both log-likelihood and perplexity than existing models.
High space cost and ignoring synonyms in STC (Suffix Tree Clustering algorithm) are challenges for search results clustering. Aiming at these challenges, this paper proposes a WordNet-based suffix tree clustering algorithm (WNSTC). WNSTC can construct a suffix tree containing WordNet synsets. When constructing the suffix tree, WNSTC looks every feature word up in WordNet database. If the feature word is included in WordNet, its synsets will be added into corresponding node. The node in the suffix tree may be a set of words (strings) with similar meaning instead of a single word (string). Experiments executed on data sets show that WNSTC has better clustering quality and smaller suffix tree size than original STC algorithm.
Ad click-through rate (CTR) prediction is to estimate CTR with click log, which is influenced by the page information, the position, the user properties, the nature features of ad and some other factors. The right ads for the query and the order they are displayed greatly affects the revenue the company receives from these ads. Therefore, it is important to be able to estimating CTR precisely with click log in sponsored search advertising system. We present a useful CTR prediction model for ads of abundant history data. We also show that using our model improves the performance of an advertising system.
With rapid development of the Internet, much attention has been paid to the problem of children exposed to Internet pornography. Existing detection techniques, which mainly focus on pornography content analysis have obtained much success. However, they still meet challenges in practical Web environment due to the great computational costs and the difficulties in dealing with various pornography forms. We attempt to solve this problem from a new perspective with the wisdom of crowds in search engine click-through logs. Inspired by the idea that different pornography Web pages may be oriented by similar search keywords, a label propagation method on click-through bipartite graph is proposed which can locate pornography Web pages from a small set (a few hundreds) of manually labeled seed pages. Experiments performed on datasets collected from both English and Chinese search engines show that the proposed algorithm can identify different forms of Internet pornography both effectively and efficiently.
A new approach based on the modified particle swarm optimization (PSO) algorithm is presented for the protection capacity optimization problem. The modified PSO algorithm combines the elitist strategy and mutation operation of the genetic algorithm (GA) which avoids premature convergence of standard PSO algorithm. The efficiency of the new method is higher than the GA owning to its simpler structure, higher speed of solution and hunting efficiency. Simulated results indicate that the near global optimal solution is easily obtained with this new method which facilitates the practical engineering realization of the algorithm.
Search engine click-through data is a valuable source of implicit user feedback for relevance. However, not all user clicks are good indication of relevance. The clicks from search experts, who are more successful searching a query, tend to be more reliable in indicating document relevance than those of the non-experts. Therefore, knowing the expertise of search users is helpful to better understand their clicks. In this paper, we propose two probabilistic modelings of user expertise in the environment of web search. Inspired by the idea of evaluation metrics in classification, search users are treated as classifiers and result documents are viewed as the data samples to classify in our models. A click implies that the document is classified as relevant by the user. Therefore, the expertise of a user can be measured by how well he/she classifies the documents. We carry out experiments on a real-world click-through data of a Chinese search engine. The results show that modeling user expertise helps the click models with relevance inference, which also implies that our models are effective in identifying the user expertise.
Rapid data growth in web-based applications poses great challenges to the design and implementation of key-value storage systems (key-value stores). In this paper, we present the design of a highly efficient single machine key-value store called THUIR-DB, which features a highly-compacted index structure and a fast querying strategy. Like Google's LevelDB, THUIR-DB also doesn't focus on distributing, based on which we can build distributing key-value data store. Experimental results based on Google's N-gram dataset show great improvement with THUIR-DB in both time and memory efficiency compared with some widely adopted open source key-value stores such as LevelDB and Tokyo Cabinet. Experiment shows that THUIR-DB costs 1.06 bit index for every record and gains 1.2 million ops/sec throughputs based on the 0.7 billion-scale dataset. We have already applied THUIR-DB to our online system Weibo Xunren(xunren.thuir.org). Copyright © 2013 Binary Information Press.
Throughput and latency are the two important performance indicators for distributed file systems. Google has achieved a great success with Google File System (GFS) when operating big files, but the latency is too big when reading and writing small files. In this paper, we propose a small-file writing acceleration mechanism (SFWM) for distributed file system. SFWM optimizes access paths of writing to disks and uses an asynchronous method to accelerate the writing. Experiments in the distributed file system, SandBox show that the time to write small-file using SFWM is significantly reduced compared with the writing time in MooseFS and the throughput of writing using SFWM is improved.
Jie Tang (唐杰)合作论文数Department of Computer Science and Technology, Tsinghua University9