Semi supervised classification of scientific and technical literature based on semi supervised hierarchical description of improved latent dirichlet allocation (LDA)

Cluster Computing(2018)

引用 0|浏览5
暂无评分
摘要
Chinese text classification problem was studied based on domain ontology graph (DOG) of semi-supervised conceptual clustering to solve the problem that English word disambiguation method cannot be applied to Chinese text classification. Structure model of domain ontology graph, text classification algorithm in HowNet dictionary and KLSeeker ontology and so on were used to realize accurate classification of Chinese text and display effectiveness of algorithm. Chinese text classification model in domain ontology graph based on conceptual clustering was developed from the angle of decreasing human participation in ontology construction as much as possible in the paper. Aimed at application domain of Chinese web text, the algorithm can generate DOG of knowledge conceptualization automatically. At the same time, document ontology graph (DocOG) was defined to represent contents of individual text document. DocOG extracting target realized text classification based on ontology by matching of single document ontology and domain ontology. Finally, example calculation analysis and actual data test set experiment were given in experimental stage. The result shows that proposed Chinese text classification method has higher classification accuracy and reflects effectiveness of design.
更多
查看译文
关键词
Scientific literature, LDA, Domain ontology graph, Word disambiguation, Semi-surprised, Conceptual clustering
AI 理解论文
溯源树
样例
生成溯源树,研究论文发展脉络
Chat Paper
正在生成论文摘要