Design of improved K-medoids algorithm for adaptive clustering number selection
Nan Wang, Wang Dawei, Lixia Wang, Qiang Gao, Hao Chen
Abstract
Nan Wang, Wang Dawei, Lixia Wang, Qiang Gao, Hao Chen
Abstract
As a classical clustering algorithm, K-medoids algorithm needs to manually input its clustering number when the program runs, so it is difficult to realize the adaptive calculation of clustering number. Therefore, an improved K-medoids algorithm considering distance and weight is proposed in this paper. The clustering algorithm uses dimension-weighted Euclidean distance to measure the distance between samples, and then obtains the density and weight of sample distance. Then, the point with the highest density in the sample was taken as the first cluster center, and all samples in the cluster were removed. The next cluster center was found according to the weight of the previous cluster center and the remaining sample points in the data set. Repeat the above process, when all the data sets are screened, multiple clustering centers will be automatically obtained. Simulation experiments on the UCI real and artificial simulated datasets show that the proposed algorithm has high accuracy and good stability.
OpenAlex reports 1 citations for this work. Citation counts describe recorded attention and do not establish research quality.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
As a classical clustering algorithm, K-medoids algorithm needs to manually input its clustering number when the program runs, so it is difficult to realize the adaptive calculation of clustering number. Therefore, an improved K-medoids algorithm considering distance and weight is proposed in this paper. The clustering algorithm uses dimension-weighted Euclidean distance to measure the distance between samples, and then obtains the density and weight of sample distance. Then, the point with the highest density in the sample was taken as the first cluster center, and all samples in the cluster were removed. The next cluster center was found according to the weight of the previous cluster center and the remaining sample points in the data set. Repeat the above process, when all the data sets are screened, multiple clustering centers will be automatically obtained. Simulation experiments on the UCI real and artificial simulated datasets show that the proposed algorithm has high accuracy and good stability.
Key concepts: k-medoids, Cluster analysis, Euclidean distance, k-medians clustering, Medoid, Computer science, Determining the number of clusters in a data set, Algorithm