Data stream clustering algorithm based on dependent function
Hui DANG
Abstract
Hui DANG
Abstract
The traditional data stream clustering algorithms are mostly based on distance or density,so their clustering quality and processing efficiency are weak.To address the above problems,this paper proposed a data stream clustering algorithm based on dependent function.Firstly,the data points were modeled in the form of matter-element and dependent function was established to solve the problem.Secondly,the value of the dependent function was calculated.According to this value,the degree that data point belongs to a certain cluster was judged.Then,the proposed method was applied to online-offline framework of the data stream clustering.Finally,the proposed algorithm was tested by using the real data set KDD-CUP99 and randomly generated artificial data sets.The experimental results show that clustering purity of the proposed method is over 92%,and it can deal with about 6 300 records per second.Compared with the traditional algorithm,the processing efficiency of the algorithm is greatly improved.In the aspects of dimension and the number of cluster,the algorithm shows stronger scalability,and it is suitable for processing large dynamic data set.
A significance statement is not available in the OpenAlex record.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
The traditional data stream clustering algorithms are mostly based on distance or density,so their clustering quality and processing efficiency are weak.To address the above problems,this paper proposed a data stream clustering algorithm based on dependent function.Firstly,the data points were modeled in the form of matter-element and dependent function was established to solve the problem.Secondly,the value of the dependent function was calculated.According to this value,the degree that data point belongs to a certain cluster was judged.Then,the proposed method was applied to online-offline framework of the data stream clustering.Finally,the proposed algorithm was tested by using the real data set KDD-CUP99 and randomly generated artificial data sets.The experimental results show that clustering purity of the proposed method is over 92%,and it can deal with about 6 300 records per second.Compared with the traditional algorithm,the processing efficiency of the algorithm is greatly improved.In the aspects of dimension and the number of cluster,the algorithm shows stronger scalability,and it is suitable for processing large dynamic data set.
Key concepts: Cluster analysis, Data stream clustering, Computer science, CURE data clustering algorithm, Data mining, Correlation clustering, Canopy clustering algorithm, Algorithm