Clustering Techniques in Bioinformatics
Muhammad Ali Masood, M. N. A. Khan
Abstract
Open-access reader
Muhammad Ali Masood, M. N. A. Khan
Abstract
Open-access reader
Dealing with data means to group information into a set of categories either in order to learn new artifacts or understand new domains.For this purpose researchers have always looked for the hidden patterns in data that can be defined and compared with other known notions based on the similarity or dissimilarity of their attributes according to well-defined rules.Data mining, having the tools of data classification and data clustering, is one of the most powerful techniques to deal with data in such a manner that it can help researchers identify the required information.As a step forward to address this challenge, experts have utilized clustering techniques as a mean of exploring hidden structure and patterns in underlying data.Improved stability, robustness and accuracy of unsupervised data classification in many fields including pattern recognition, machine learning, information retrieval, image analysis and bioinformatics, clustering has proven itself as a reliable tool.To identify the clusters in datasets algorithm are utilized to partition data set into several groups based on the similarity within a group.There is no specific clustering algorithm, but various algorithms are utilized based on domain of data that constitutes a cluster and the level of efficiency required.Clustering techniques are categorized based upon different approaches.This paper is a survey of few clustering techniques out of many in data mining.For the purpose five of the most common clustering techniques out of many have been discussed.The clustering techniques which have been surveyed are: K-medoids, Kmeans, Fuzzy C-means, Density-Based Spatial Clustering of Applications with Noise (DBSCAN) and Self-Organizing Map (SOM) clustering.
OpenAlex reports 27 citations for this work. Citation counts describe recorded attention and do not establish research quality.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
Dealing with data means to group information into a set of categories either in order to learn new artifacts or understand new domains.For this purpose researchers have always looked for the hidden patterns in data that can be defined and compared with other known notions based on the similarity or dissimilarity of their attributes according to well-defined rules.Data mining, having the tools of data classification and data clustering, is one of the most powerful techniques to deal with data in such a manner that it can help researchers identify the required information.As a step forward to address this challenge, experts have utilized clustering techniques as a mean of exploring hidden structure and patterns in underlying data.Improved stability, robustness and accuracy of unsupervised data classification in many fields including pattern recognition, machine learning, information retrieval, image analysis and bioinformatics, clustering has proven itself as a reliable tool.To identify the clusters in datasets algorithm are utilized to partition data set into several groups based on the similarity within a group.There is no specific clustering algorithm, but various algorithms are utilized based on domain of data that constitutes a cluster and the level of efficiency required.Clustering techniques are categorized based upon different approaches.This paper is a survey of few clustering techniques out of many in data mining.For the purpose five of the most common clustering techniques out of many have been discussed.The clustering techniques which have been surveyed are: K-medoids, Kmeans, Fuzzy C-means, Density-Based Spatial Clustering of Applications with Noise (DBSCAN) and Self-Organizing Map (SOM) clustering.
Key concepts: Cluster analysis, Computer science, Consensus clustering, Data mining, Fuzzy clustering, Clustering high-dimensional data, Correlation clustering, CURE data clustering algorithm