2017 12th IEEE Conference on Industrial Electronics and Applications (ICIEA)(2017)
Wuhan Univ Sci & Technol
被引用2|浏览2
摘要
Due to the large number of XML data and information and its advantage of simplicity, semi-structure, extensibility and self-description, clustering XML document has become a hot issue in XML data mining. LSPX model is a kind of XML data structure representation model, which is simple in construction process and short in time. The incremental clustering algorithm based on this model and the similarity calculation gets high time efficiency and good clustering effect. But the weakness of sensitivity to the order of input appeared in traditional incremental clustering. In order to further improve the clustering efficiency of XML documents, in this paper, a new XML clustering algorithm (CO-LSPX) is proposed, which is based on the cluster core and LSPX. The experimental result shows that the proposed method can increase efficiency of clustering, reduce the time consumption greatly, as well as mute the sensitivity of input data order on the basis of ensuring the quality of clustering results.