Optimal Clustering using Modified Fuzzy C-Means Clustering Algorithm
Shafeeq B M Ahamed
Abstract
Shafeeq B M Ahamed
Abstract
Fuzzy clustering has been widely studied and applied in a variety of applications and areas. In hard clustering, data is divided into distinct clusters, where each data element belongs to exactly one cluster. In fuzzy clustering, data elements can belong to more than one cluster, and associated with each element is a set of membership levels. These indicate the strength of the association between that data element and a particular cluster. Fuzzy clustering is a process of assigning these membership levels, and then using them to assign data elements to one or more clusters. The investigation is needed to reveal whether the optimal number of clusters can be found on the run based on the cluster quality measure. The silhouette coefficient is the one of measure used to measure the quality of clusters. In the practical scenario, it is very difficult to fix the number of clusters in advance. In this paper we propose an optimal clustering of data with modified Fuzzy C-Means algorithm. The proposed method works for both the cases i.e. for known number of clusters in advance as well as unknown number of clusters. The user has the flexibility either to fix the number of clusters or input the minimum number (K=2) of clusters required. In the former case it works same as Fuzzy C-means algorithm. In the latter case the algorithm computes the quality of clusters for each set of clusters. The process is repeated by incrementing the cluster counter by one in each iteration until it satisfies the validity of cluster quality. It is observed that the modified Fuzzy C-means algorithm produces quality clusters compared to the Fuzzy C-means clustering. It assigns the data point to their appropriate class or cluster more effectively.
OpenAlex reports 1 citations for this work. Citation counts describe recorded attention and do not establish research quality.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
Fuzzy clustering has been widely studied and applied in a variety of applications and areas. In hard clustering, data is divided into distinct clusters, where each data element belongs to exactly one cluster. In fuzzy clustering, data elements can belong to more than one cluster, and associated with each element is a set of membership levels. These indicate the strength of the association between that data element and a particular cluster. Fuzzy clustering is a process of assigning these membership levels, and then using them to assign data elements to one or more clusters. The investigation is needed to reveal whether the optimal number of clusters can be found on the run based on the cluster quality measure. The silhouette coefficient is the one of measure used to measure the quality of clusters. In the practical scenario, it is very difficult to fix the number of clusters in advance. In this paper we propose an optimal clustering of data with modified Fuzzy C-Means algorithm. The proposed method works for both the cases i.e. for known number of clusters in advance as well as unknown number of clusters. The user has the flexibility either to fix the number of clusters or input the minimum number (K=2) of clusters required. In the former case it works same as Fuzzy C-means algorithm. In the latter case the algorithm computes the quality of clusters for each set of clusters. The process is repeated by incrementing the cluster counter by one in each iteration until it satisfies the validity of cluster quality. It is observed that the modified Fuzzy C-means algorithm produces quality clusters compared to the Fuzzy C-means clustering. It assigns the data point to their appropriate class or cluster more effectively.
Key concepts: Fuzzy clustering, Cluster analysis, Correlation clustering, Complete-linkage clustering, Determining the number of clusters in a data set, k-medians clustering, Single-linkage clustering, FLAME clustering