We seek to improve the robustness and portability of temporal information extraction systems by incorporating data-driven techniques. We present two sets of experiments pointing us in this direction. The first shows that machine-learning-based recognition of temporal expressions not only achieves high accuracy on its own but can also improve rule-based normalization. The second makes use of a staged normalization architecture to experiment with machine learned classifiers for certain disambiguation sub-tasks within the normalization task.
We pursue two strategies for oine data collection for a temporal question answering system that uses both quantitative methods and fuzzy methods to reason about time and events. The first strategy extracts event descriptions from the structured year entries in the online encyclopedia Wikipedia, yielding clean quantitative temporal information about a range of events. The second strategy mines the web using patterns indicating temporal relations between events and times and between events. Web mining leverages the volume of data available on the web to find qualitative temporal relations between known events and new, related events and to build fuzzy time spans for events for which we lack crisp metric temporal information.
We present TweetMotif, an exploratory search applica- tion for Twitter. Unlike traditional approaches to in- formation retrieval, which present a simple list of mes- sages, TweetMotif groups messages by frequent signif- icant terms — a result set’s subtopics — which facili- tate navigation and drilldown through a faceted search interface. The topic extraction system is based on syn- tactic filtering, language modeling, near-duplicate de- tection, and set cover heuristics. We have used Tweet- Motif to deflate rumors, uncover scams, summarize sentiment, and track political protests in real-time. A demo of TweetMotif, plus its source code, is available at http://tweetmotif.com.
We describe our participation in the TREC 2004 Question Answering track. We provide a detailed account of the ideas underlying our approach to the QA task, especially to the so-called “other” questions. This year we made essential use of Wikipedia, the free online encyclopedia, both as a source of answers to factoid questions and as an importance model to help us identify material to be returned in response to “other” questions.
Franco Salvetti合作论文数Microsoft Live Search12
Ajay K. Gupta合作论文数Computer Science at Western Michigan University1