Textual documents representation is considered as a crucial phase in many text mining tasks. It intends generally to capture the semantic information of the whole document into a representation vector, which could be further utilized in such tasks. In this paper, we have proposed a deep semantic representation method which is based on two types of features. The first is derived from the biomedical widely used Structured Semantic Resource MeSH whereas the second is generated from a deep learning phase. At this end, we chose to combine the two models Word2Vec and Convolutional Neural Network which allow bringing out the semantic relationships existing in large and complex textual documents at high abstraction levels. To evaluate this method, we involved it into a document classification process. The results of experiments, carried out on a sub-set of the OHSUMED collection, show that our method performs well.
更多
查看译文
关键词
Deep semantic document representation,Convolution Neural Network,MeSH,Word2Vec,Textual Document Classification