ADTHA: The improvement of clustering algorithm
Adtha Lawanna
Abstract
Adtha Lawanna
Abstract
Amounts of data increases perform data overload is a critical issue in IT organizations. Therefore, clustering techniques are proposed for reducing the size of data to keep the performance of the entire system by selecting groups of data that is depend on their similarity. According to this k-means and hierarchical clustering techniques have been designed. The well-known techniques are clustering large application-based upon randomized search and clustering and balanced iterative reducing and clustering using hierarchy. However, one of the remained problems is size of data still large. This can make the whole processes of data mining works slowly. Another is that the similarity value of the clusters is still not high enough for to be used, particularly, in the process of decision making. Therefore, this paper presents a model of clustering to provide higher efficiency of the process of k-mean and hierarchical clustering. The efficiency of the proposed model is better than traditional techniques by 0.5-1.3% approximately in term of giving higher similarity value. Besides, the running time of the model is also faster than the comparative studies about 2.7-7.4 times.
A significance statement is not available in the OpenAlex record.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
Amounts of data increases perform data overload is a critical issue in IT organizations. Therefore, clustering techniques are proposed for reducing the size of data to keep the performance of the entire system by selecting groups of data that is depend on their similarity. According to this k-means and hierarchical clustering techniques have been designed. The well-known techniques are clustering large application-based upon randomized search and clustering and balanced iterative reducing and clustering using hierarchy. However, one of the remained problems is size of data still large. This can make the whole processes of data mining works slowly. Another is that the similarity value of the clusters is still not high enough for to be used, particularly, in the process of decision making. Therefore, this paper presents a model of clustering to provide higher efficiency of the process of k-mean and hierarchical clustering. The efficiency of the proposed model is better than traditional techniques by 0.5-1.3% approximately in term of giving higher similarity value. Besides, the running time of the model is also faster than the comparative studies about 2.7-7.4 times.
Key concepts: Cluster analysis, Computer science, Data mining, CURE data clustering algorithm, Correlation clustering, Canopy clustering algorithm, Similarity (geometry), Data stream clustering