Data stream clustering based on feature selection
Lin Wang
Abstract
Lin Wang
Abstract
Clustering in the data stream,the redundant features will affect the quality of data clustering,removing redundant features to improve the clustering quality is very important.To solve this problem,it is proposed that a data stream clustering algorithm based on feature selection(DSCFC).It is one-pass clustering algorithms,these are applied that ranking feature,grading feature,detecting redundant features and removing the redundant features algorithm and so on.The experimental results indicated that DSCFC algorithm can detect hidden redundant features in data stream and remove redundant features;when there are redundant features in the data stream clustering,the algorithm is more efficient than CluStream,clustering quality is better.
A significance statement is not available in the OpenAlex record.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
Clustering in the data stream,the redundant features will affect the quality of data clustering,removing redundant features to improve the clustering quality is very important.To solve this problem,it is proposed that a data stream clustering algorithm based on feature selection(DSCFC).It is one-pass clustering algorithms,these are applied that ranking feature,grading feature,detecting redundant features and removing the redundant features algorithm and so on.The experimental results indicated that DSCFC algorithm can detect hidden redundant features in data stream and remove redundant features;when there are redundant features in the data stream clustering,the algorithm is more efficient than CluStream,clustering quality is better.
Key concepts: Data stream clustering, Cluster analysis, Computer science, CURE data clustering algorithm, Correlation clustering, Canopy clustering algorithm, Data mining, Fuzzy clustering