Study of High-Dimensional Data Analysis based on Clustering Algorithm
Ping Zong, Junyan Jiang, Jun Qin
Abstract
Ping Zong, Junyan Jiang, Jun Qin
Abstract
With the rapid development of big data, the scale, dimensions, diversity and sparsity of high-dimensional data restrict the effectiveness of traditional clustering algorithms. This paper mainly focuses on high-dimensional data clustering. Starting from the traditional K-means clustering algorithm and subspace clustering algorithm based on self-representation model, an improved algorithm is designed and implemented based on the existing clustering algorithm in this paper. The improved algorithm has better clustering quality by combining the "distance optimization method" and the "density method" to determine the initial clustering center. The feasibility and effectiveness of improved algorithm are verified through simulation experiments.
OpenAlex reports 5 citations for this work. Citation counts describe recorded attention and do not establish research quality.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
With the rapid development of big data, the scale, dimensions, diversity and sparsity of high-dimensional data restrict the effectiveness of traditional clustering algorithms. This paper mainly focuses on high-dimensional data clustering. Starting from the traditional K-means clustering algorithm and subspace clustering algorithm based on self-representation model, an improved algorithm is designed and implemented based on the existing clustering algorithm in this paper. The improved algorithm has better clustering quality by combining the "distance optimization method" and the "density method" to determine the initial clustering center. The feasibility and effectiveness of improved algorithm are verified through simulation experiments.
Key concepts: Cluster analysis, Canopy clustering algorithm, CURE data clustering algorithm, Correlation clustering, Computer science, Data stream clustering, Clustering high-dimensional data, Data mining