Incremental hierarchical text clustering with privileged information
Ricardo Marcondes Marcacini, Solange Oliveira Rezende
Abstract
Ricardo Marcondes Marcacini, Solange Oliveira Rezende
Abstract
In many text clustering tasks, there is some valuable knowledge about the problem domain, in addition to the original textual data involved in the clustering process. Traditional text clustering methods are unable to incorporate such additional (privileged) information into data clustering. Recently, a new paradigm called LUPI - Learning Using Privileged Information - was proposed by Vapnik to incorporate privileged information in classification tasks. In this paper, we extend the LUPI paradigm to deal with text clustering tasks. In particular, we show that the LUPI paradigm is potentially promising for incremental hierarchical text clustering, being very useful for organizing large textual databases. In our method, the privileged information about the text documents is applied to refine an initial clustering model by means of consensus clustering. The initial model is used for incremental clustering of the remaining text documents. We carried out an experimental evaluation on two benchmark text collections and the results showed that our method significantly improves the clustering accuracy when compared to a traditional hierarchical clustering method.
OpenAlex reports 20 citations for this work. Citation counts describe recorded attention and do not establish research quality.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
In many text clustering tasks, there is some valuable knowledge about the problem domain, in addition to the original textual data involved in the clustering process. Traditional text clustering methods are unable to incorporate such additional (privileged) information into data clustering. Recently, a new paradigm called LUPI - Learning Using Privileged Information - was proposed by Vapnik to incorporate privileged information in classification tasks. In this paper, we extend the LUPI paradigm to deal with text clustering tasks. In particular, we show that the LUPI paradigm is potentially promising for incremental hierarchical text clustering, being very useful for organizing large textual databases. In our method, the privileged information about the text documents is applied to refine an initial clustering model by means of consensus clustering. The initial model is used for incremental clustering of the remaining text documents. We carried out an experimental evaluation on two benchmark text collections and the results showed that our method significantly improves the clustering accuracy when compared to a traditional hierarchical clustering method.
Key concepts: Cluster analysis, Computer science, Brown clustering, Document clustering, Conceptual clustering, Clustering high-dimensional data, Consensus clustering, Hierarchical clustering