事件时序关系抽取是一项重要的自然语言理解任务,可以广泛应用于诸如知识图谱构建、问答系统等任务.已有事件时序关系抽取方法往往将该任务视为句子级事件对的分类问题,而基于有限的局部句子信息导致其抽取的事件时序关系的精度较低,且无法保证整体时序关系的全局一致性.针对此问题,提出一种融合上下文信息的篇章级事件时序关系抽取方法,使用基于双向长短期记忆(bidirectional long short-term memory,Bi-LSTM)的神经网络模型学习文章中事件对的时序关系表示,再利用自注意力机制融入上下文中其他事件对信息,从而得到更丰富的事件对时序关系表示用于时序关系分类通过 TB-Dense(timebank dense)和 M ATRES(multi-axis temporal relations for start-points)数据集的实验表明:此方法能够取得比当前主流的句子级方法更佳的抽取效果.
随着信息技术的飞速发展,互联网成为了舆情传播的主要载体.各种舆情事件不断涌现,并在网民的参与下广泛传播,由此可能引发强烈的社会反响.因此,如何实现网络舆情事件快速发现与个性化监测需求的精准推送,成为了当前舆情的重点关注内容.对于舆情场景下用户交互信息稀疏导致的兴趣难以刻画的问题,提出了一种基于层次知识的话题推荐模型.模型通过引入层次知识来扩充语义增加话题之间的潜在信息关联,分别对层次知识、话题和用户建模得到对应的嵌入向量表示,再结合多层感知机匹配模型预测用户点击率.实验结果表明,该模型在与多个基线算法的对比中,在F1(the balanced F score)和AUC(the area under curve)指标的平均值上分别提升了6.7%和4.9%.
利用事件报道描述内容高度相似的特点,提出了一种抽取式话题简短表示生成方法。把事件文档标题集中的标题作为处理对象,从不同的标题中抽取出保留原有语序的共性信息,并进一步融合这些共性信息,生成事件粒度的话题简短表示。在来自搜索引擎中的事件数据上,实验结果表明该方法能生成精练、准确、语义明确完整且可读性好的话题简短表示。
检测网页重要变化,判断页面核心内容是否发生变化,可有效降低数据采集中重复索引的数量,因此,文中提出基于视觉的网页重要变化检测方法,用于检测页面不同语义区域的变化,可将页面压缩表示为一个低维向量.从用户视觉的角度,理解页面不同区块语义重要度的差异.相比现有方法,文中方法独立于基于HTML类基础文档的分析方法,在新媒体,如移动互联网上,也有一定的适用性.实验也验证文中方法的有效性.
网络舆论对人们生活的影响程度与日俱增,通过结合多源数据进行事件发现可以更好地捕捉舆情事件,提高舆情系统的效果。针对在多源文本场景下如何将来自新闻、微博、微信等多通道的数据融合,文章根据事件的定义,提出了事件核心实体的概念,设计了事件核心实体识别方法,并且将事件核心实体应用到事件发现过程,提出了结合实体的事件发现方法 ESP(Entity Single-Pass)。该方法通过引入实体信息,丰富了多源文本中每篇文档的表达,从而提高了多源文本事件发现的效果。实验表明,在微博、新闻等数据上,我们的方法与K-means和SinglePass方法相比,在NMI与RI两项指标上分别提高了0.2和0.3,证明了ESP算法的有效性。
及时获取新增内容,是采集器的重要衡量指标。基于版块页-内容页架构设计的网络采集器通过定期重采入口的版块页,能够有效地快速识别新产生内容页面并进行扩展。然而获取内容的实时性与对网站访问的友好性存在一定的折中。传统的重采策略关注时效性,而忽略了对网站访问的友好性。该文提出了一种基于时间序列预测的改进重采策略兼顾时效性和友好性。实验表明,该方法可以在保证数据采集实时性的情况下,有效降低访问量,提升对网站访问的友好性。
Microblog, as a way of online communication, can generate large amounts of information in a very short period. Therefore, how to retrieve the latest relevant information becomes a hot research area. Different from traditional information retrieval (IR), the microblog retrieval emphasizes fresh contents of the information. In order to solve this problem, we extend the traditional IR methods by taking into account the posting time. We propose a time-sensitive retrieval model, which takes the time factor as a prior probability. In the retrieval model, we introduce the pseudo relevance feedback technology as a query expansion approach to improve retrieval performance. Furthermore, we introduce a strategy to filter the initial retrieval results, which takes post quality factors into account including entropy and link features. Experiments on Twitter corpus show that our algorithm is effective to improve the retrieval performance, and the retrieval results can meet the real time retrieval need well.
In order to solve problems which include the topic drift phenomenon and much higher level of noise in micro-blogs,an algorithm named the Streaming Dynamic Topic Model,which improves the dynamic topic model with MEntropy,was presented to track additional events on topics.The method of the dynamic topic model was first tried to update the topic in the whole tracking process,which enhanced the description power of the topic model by both positive and negative sides to overcome the topic drift problem.However,as a high level of neutral posts existed,MEntropy was defined and used to evaluate the importance of a microblog for tracking a topic,and was then extended to the dynamic topic model in order to make a better distinction between even micro-blogs and neutral ones.Topic tracking experiments on a collection of more than 170,000 users' 12 million microblogs show that our algorithm is more efficient and with lower noise compared with the traditional dynamic topic model.
: There are two search tasksin TREC2012 Microblog Track, namely: Real-time Adhoc and Real-time Filtering. The Tweets2011 corpus is used again and last year s results can be used as first officially labeled data for any participants to train their models. In this year s track, the former task has60 new queries given and the latter is first proposed. In the Real-time Adhoc task, we use indri retrieval toolkit to construct our retrieval system and propose a strategy of pseudo relevance feedback to expand original query, then we retrieved original tweets and their indri s scores as an important feature. Besides, we calculate lots of other features of these tweets, such abouturl, hash_tag, entropy, tfidf, bm25, language model and proximity. At last, we use two learning-to-rank methods, specifically RankSVM and ListNet, to combine all those featuresto sort them, returning the final ranked tweets to a specified query. In the Real-time Filtering task, we assuming this task is similar with the topic tracking in Twitter Stream, we build up two filtering models based on language model and Vector Space Model respectively. Each model is initialized by the start query and its relevant tweets. For each new coming tweet, the model will decide whether it is under the topic. If it is, we update the model to keep up with the development of the topic. The rest of this paper is organized as follows. In Section 2, we discuss the preprocessing of Tweets2011 corpus. In Section 3, the main method to rank the search results in Real-time Adhoc task is discussed. In Real-time Filtering task, we describe two filtering models in Section 4. Experiment resultresults areted in Section 5. And in the last section, we draw conclusions about our work.
当前我国正对足球赌球案件进行专项调查.针对网络赌博案情信息语义信息的不明确性和分析的复杂性,综合运用Web信息抽取技术、犯罪特征关系可视化分析技术和计算机取证技术,设计并实现了网络赌博案情分析系统.实验表明,该系统可以快速、有效地进行网络赌博案情信息的分析处理,更加直观地表现案情,为案件侦破提供重要线索.
In TREC 2011 Microblog Track, we explore the use of pseudo relevance feedback to expand original query terms in topics. Hyperlink is used to enhance the performance of the retrieval results. And we set a threshold of entropy to filter retrieval results. Microblog is a Realtime Adhoc Task, so we make use of average querytweettime that comes from pseudo relevance feedback to change retrieval score. We combine two models to improve retrieval results. The results show that our model is effective at realtime relevance retrieval.