2010 IEEE INTERNATIONAL SYMPOSIUM ON INFORMATION THEORY(2010)
Univ Calif San Diego
被引用12|浏览12
摘要
We consider the problem of classification, where the data of the classes are generated i.i.d. according to unknown probability distributions. The goal is to classify test data with minimum error probability, based on the training data available for the classes. The Likelihood Ratio Test (LRT) is the optimal decision rule when the distributions are known. Hence, a popular approach for classification is to estimate the likelihoods using well known probability estimators, e.g., the Laplace and Good-Turing estimators, and use them in a LRT. We are primarily interested in situations where the alphabet of the underlying distributions is large compared to the training data available, which is indeed the case in most practical applications. We motivate and propose LRT's based on pattern probability estimators that are known to achieve low redundancy for universal compression of large alphabet sources. While a complete proof for optimality of these decision rules is warranted, we demonstrate their performance and compare it with other well-known classifiers by various experiments on synthetic data and real data for text classification.
更多
查看译文
AI 解读
一键生成论文网页
Chat Paper
正在生成论文摘要
关键词
data compression,error statistics,pattern classification,statistical distributions,text analysis,Laplace estimators,good-Turing estimators,large alphabet source universal compression,likelihood ratio test,minimum error probability,optimal decision rule,pattern probability estimators,probability distributions,text classification