Chinese Document Clustering Based on Topic Concept Clustering
Fuding Xie
Abstract
Fuding Xie
Abstract
Nowadays,document clustering technology has been extensively used in text mining,information retrieval systems and etc.The conventional document clustering methods rely on the classical vector-space model using the key words as the feature.However,these methods ignore the semantic relations among the keywords,don′t really address the special problems of document clustering:high dimensionality of the data,and high computation complexity.To solve these problems,based on topic concept clustering,this paper proposes a method for Chinese document clustering.The topic concepts from a document are extracted by using HowNet and are clustered by Chameleon algorithm.The document classification has been performed in terms of the results of topic concept clustering.The document is presented by the concept vector to decrease the dependence relations among of documents.The computation complexity of the document clustering is reduced efficiently and effectively via this method.
OpenAlex reports 1 citations for this work. Citation counts describe recorded attention and do not establish research quality.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
Nowadays,document clustering technology has been extensively used in text mining,information retrieval systems and etc.The conventional document clustering methods rely on the classical vector-space model using the key words as the feature.However,these methods ignore the semantic relations among the keywords,don′t really address the special problems of document clustering:high dimensionality of the data,and high computation complexity.To solve these problems,based on topic concept clustering,this paper proposes a method for Chinese document clustering.The topic concepts from a document are extracted by using HowNet and are clustered by Chameleon algorithm.The document classification has been performed in terms of the results of topic concept clustering.The document is presented by the concept vector to decrease the dependence relations among of documents.The computation complexity of the document clustering is reduced efficiently and effectively via this method.
Key concepts: Cluster analysis, Document clustering, Computer science, Vector space model, Clustering high-dimensional data, Data mining, Correlation clustering, Information retrieval