2007Modern Electronics TechniqueRequires access

Chinese Document Clustering Based on Topic Concept Clustering

Fuding Xie

Open publisher page 1 citations

Abstract

Nowadays,document clustering technology has been extensively used in text mining,information retrieval systems and etc.The conventional document clustering methods rely on the classical vector-space model using the key words as the feature.However,these methods ignore the semantic relations among the keywords,don′t really address the special problems of document clustering:high dimensionality of the data,and high computation complexity.To solve these problems,based on topic concept clustering,this paper proposes a method for Chinese document clustering.The topic concepts from a document are extracted by using HowNet and are clustered by Chameleon algorithm.The document classification has been performed in terms of the results of topic concept clustering.The document is presented by the concept vector to decrease the dependence relations among of documents.The computation complexity of the document clustering is reduced efficiently and effectively via this method.

About this research paper

What this paper is about

Nowadays,document clustering technology has been extensively used in text mining,information retrieval systems and etc.The conventional document clustering methods rely on the classical vector-space model using the key words as the feature.However,these methods ignore the semantic relations among the keywords,don′t really address the special problems of document clustering:high dimensionality of the data,and high computation complexity.To solve these problems,based on topic concept clustering,this paper proposes a method for Chinese document clustering.The topic concepts from a document are extracted by using HowNet and are clustered by Chameleon algorithm.The document classification has been performed in terms of the results of topic concept clustering.The document is presented by the concept vector to decrease the dependence relations among of documents.The computation complexity of the document clustering is reduced efficiently and effectively via this method.

Why it matters

OpenAlex reports 1 citations for this work. Citation counts describe recorded attention and do not establish research quality.

Key contribution

A contribution statement is not available in the OpenAlex record.

Method / approach

Method details are not available in the OpenAlex metadata.

Main findings

Findings are not separately available in the OpenAlex metadata.

Limitations

Limitations are not available in the OpenAlex metadata.

Applications

Application details are not available in the OpenAlex metadata.

Available abstract

Nowadays,document clustering technology has been extensively used in text mining,information retrieval systems and etc.The conventional document clustering methods rely on the classical vector-space model using the key words as the feature.However,these methods ignore the semantic relations among the keywords,don′t really address the special problems of document clustering:high dimensionality of the data,and high computation complexity.To solve these problems,based on topic concept clustering,this paper proposes a method for Chinese document clustering.The topic concepts from a document are extracted by using HowNet and are clustered by Chameleon algorithm.The document classification has been performed in terms of the results of topic concept clustering.The document is presented by the concept vector to decrease the dependence relations among of documents.The computation complexity of the document clustering is reduced efficiently and effectively via this method.

Key concepts: Cluster analysis, Document clustering, Computer science, Vector space model, Clustering high-dimensional data, Data mining, Correlation clustering, Information retrieval

Related papers

Back to paper searchBrowse research topicsOriginal source
Chinese Document Clustering Based on Topic Concept Clustering — Research Paper | ScholarLens