A fuzzy threshold based unsupervised clustering algorithm for natural data exploration
Binu P. Thomas, G. Raju
Abstract
Binu P. Thomas, G. Raju
Abstract
Traditional clustering methods require the user to determine the number of clusters before we start any data exploration. In fuzzy clustering methods the performance efficiency of the algorithm depends mainly on the initial selection of number of clusters and cluster seeds. The real world data is almost never arranged in clear cut group and the initial selection of cluster count and centroids becomes a tedious task. In this paper we propose a new unsupervised clustering algorithm which works on the principles of fuzzy clustering. The new method we propose is using a modified form of popular fuzzy c-means algorithm for membership calculation. The algorithm begins with two initial cluster centers and forms many clusters based on a threshold value. It uses the fuzzy membership value of a cluster centre in another existing cluster to merge the clusters and finally converges to the optimum number of clusters. The algorithm is tested with the data for Gross National Happiness (GNH) program of Bhutan and found to be highly efficient in segmenting natural data sets.
OpenAlex reports 6 citations for this work. Citation counts describe recorded attention and do not establish research quality.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
Traditional clustering methods require the user to determine the number of clusters before we start any data exploration. In fuzzy clustering methods the performance efficiency of the algorithm depends mainly on the initial selection of number of clusters and cluster seeds. The real world data is almost never arranged in clear cut group and the initial selection of cluster count and centroids becomes a tedious task. In this paper we propose a new unsupervised clustering algorithm which works on the principles of fuzzy clustering. The new method we propose is using a modified form of popular fuzzy c-means algorithm for membership calculation. The algorithm begins with two initial cluster centers and forms many clusters based on a threshold value. It uses the fuzzy membership value of a cluster centre in another existing cluster to merge the clusters and finally converges to the optimum number of clusters. The algorithm is tested with the data for Gross National Happiness (GNH) program of Bhutan and found to be highly efficient in segmenting natural data sets.
Key concepts: Cluster analysis, Fuzzy clustering, Computer science, Single-linkage clustering, CURE data clustering algorithm, Data mining, Centroid, Canopy clustering algorithm