Probabilistic topic models are a conventional standard for discovering hidden semantics from text documents. However, when applied on short texts, existing models work poorly and yield incoherent and repetitive topics because of the limited word co-occurrence information. Several attempts have been made in this regard, where some methods have been proposed to integrate supervision into the learning process, but these models are still unable to infer coherent topics that correlate with human judgment. To address this issue, we propose an embedded supervised topic model that encodes word embeddings and latent topic representations guided by class labels on the hypersphere. Unlike existing supervised topic models that focus on target prediction and neglect topic interpretability, our model integrates pre-trained word embeddings to infer interpretable topics. The resulting spherical representations can be used for prediction tasks, as the inferred topics are indicative of document labels. To further enhance the interpretability aspect of topics, we propose a second framework that integrates knowledge graph embeddings into our probabilistic model, which is crucial to further palliate the data sparsity problem. Experimental results on four benchmark datasets show that our proposed models outperform existing models on topic interpretability, while having competitive label prediction capability.
更多
查看译文
关键词
Topic models,Short text modeling,Spherical embedding,Knowledge graphs,Word embedding