An adaptive density clustering algorithm for massive data
Keyan Cao, Ibrahim Musa, Jiadi Liu, Yunting Zhang
Abstract
Keyan Cao, Ibrahim Musa, Jiadi Liu, Yunting Zhang
Abstract
In this paper, two clustering algorithms are proposed: DBSCAN Entropy-based (DBSCAN) and dynamic clustering algorithm (DBSCAN) to determine the optimal clustering results. The Optimal Number of Clusters ENDBSCAN (OP-ENDBSCAN). ENDBSCAN takes information entropy as the main consideration in clustering, and avoids the traditional DBSCAN algorithm needs to define two parameters of Eps (neighborhood radius) and Minpts (density threshold). At the same time, in order to solve the problem of huge amount of data, a data preprocessing method is proposed. The method divides the data into blocks and divides them into different computer nodes, so as to make full use of the data nodes. Computing power, and improve the efficiency and scalability of the clustering algorithm. OP-ENDBSCAN is an algorithm to determine the optimal number of clustering dynamically and to evaluate the quality of clustering. Based on the analysis of ENDBSCAN, it is found that this algorithm needs to determine the number of clustering by artificially. In order to avoid this problem, OP-ENDBSCAN The effect of anthropogenic parameters on the clustering results was improved and the quality of clustering was improved. Experiments show that both ENDBSCAN and OP-ENDBSCAN can show high efficiency under different data sets and show good clustering results.
OpenAlex reports 2 citations for this work. Citation counts describe recorded attention and do not establish research quality.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
In this paper, two clustering algorithms are proposed: DBSCAN Entropy-based (DBSCAN) and dynamic clustering algorithm (DBSCAN) to determine the optimal clustering results. The Optimal Number of Clusters ENDBSCAN (OP-ENDBSCAN). ENDBSCAN takes information entropy as the main consideration in clustering, and avoids the traditional DBSCAN algorithm needs to define two parameters of Eps (neighborhood radius) and Minpts (density threshold). At the same time, in order to solve the problem of huge amount of data, a data preprocessing method is proposed. The method divides the data into blocks and divides them into different computer nodes, so as to make full use of the data nodes. Computing power, and improve the efficiency and scalability of the clustering algorithm. OP-ENDBSCAN is an algorithm to determine the optimal number of clustering dynamically and to evaluate the quality of clustering. Based on the analysis of ENDBSCAN, it is found that this algorithm needs to determine the number of clustering by artificially. In order to avoid this problem, OP-ENDBSCAN The effect of anthropogenic parameters on the clustering results was improved and the quality of clustering was improved. Experiments show that both ENDBSCAN and OP-ENDBSCAN can show high efficiency under different data sets and show good clustering results.
Key concepts: DBSCAN, Cluster analysis, CURE data clustering algorithm, Computer science, Canopy clustering algorithm, Correlation clustering, Data stream clustering, Data mining