The selection of initial focal point has great influence on the clustering results of traditional K-means algorithm,for it tends to get a local optimal solution when inappropriately assigned.In view of this issue,initial algorithm that can generate the initial cluster center was proposed,through introducing the density and nearest neighbor idea.These selected centers were used for K-means algorithm;a better text clustering algorithm called DN-K-means was put forward.The results of experiments indicate that the algorithm can lead to results with high and steady clustering quality.
Aim The traditional algorithm based on vector space model actually neglects the word order and structure in sentences,which will affect the accuracy of similarity computing.So this paper proposed a new textual case similarity algorithm.Methods The sentence,rather than the word,was used as the unit and the word order information was considered,sentence vector space model was proposed,which is the base of textual case similarity algorithm.Results The method is more consistent with the mode of human understanding and improves the accuracy of textual case similarity compatation.Conclusion The application in textual case classification proves that the method is feasible.