Automatic Detection Of News Articles Of Interest To Regional Communities

INTERNATIONAL JOURNAL OF COMPUTER SCIENCE AND NETWORK SECURITY(2012)

引用 27|浏览9
暂无评分
摘要
In this paper, we devise an approach for identifying and classifying contents of interest related to geographic communities from news articles streams. We first conduct a short study on related works, and then present our approach, which consists in 1) filtering out contents irrelevant to communities and 2) classifying the remaining relevant news articles. Using a confidence threshold, the filtering and classification tasks can be performed in one pass using the weights learned by the same algorithm. We use Bayesian text classification, and because of important empiric class imbalance in Web-crawled corpora, we test several approaches: Naive Bayes, Complementary Naive Bayes, use of {1,2,3}-Grams, and use of oversampling. We find out in our testing experiment on Japanese prefectures that 3-gram CNB with oversampling is the most effective approach in terms of precision, while retaining acceptable training time and testing time.
更多
查看译文
关键词
Web Intelligence, Natural Language Processing, Machine Learning, Semantic Web
AI 理解论文
溯源树
样例
生成溯源树,研究论文发展脉络
Chat Paper
正在生成论文摘要