QH-K Algorithm for News Text Topic Extraction
Fengxia, Wang Yongli, Huanhuan Yuan, Gong Xiaoze, Sun Shurong
Abstract
Fengxia, Wang Yongli, Huanhuan Yuan, Gong Xiaoze, Sun Shurong
Abstract
Text clustering method is the main measure for news topic extraction and trend tracking. For condensing advantages and disadvantages of the hierarchy Clustering algorithm and K-means algorithm on text Clustering, proposed a new news text Clustering optimization algorithm, called QH-K(Quick Hierarchical K-means Clustering)algorithm. First, Using word2vector model to train texts into word vector. Then, Using proposed hierarchical clustering algorithm to cluster text, get the initial number of clustering and the clustering center by proposed validity index ST. Finally, Using k-means algorithm to optimize the clustering results and improve the final clustering effect. The experiments show that the accuracy, recall rate and F value of QH-K clustering optimization algorithm are all improved to compared with the traditional algorithm. In addition, the running time of the algorithm is also reduced.
OpenAlex reports 2 citations for this work. Citation counts describe recorded attention and do not establish research quality.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
Text clustering method is the main measure for news topic extraction and trend tracking. For condensing advantages and disadvantages of the hierarchy Clustering algorithm and K-means algorithm on text Clustering, proposed a new news text Clustering optimization algorithm, called QH-K(Quick Hierarchical K-means Clustering)algorithm. First, Using word2vector model to train texts into word vector. Then, Using proposed hierarchical clustering algorithm to cluster text, get the initial number of clustering and the clustering center by proposed validity index ST. Finally, Using k-means algorithm to optimize the clustering results and improve the final clustering effect. The experiments show that the accuracy, recall rate and F value of QH-K clustering optimization algorithm are all improved to compared with the traditional algorithm. In addition, the running time of the algorithm is also reduced.
Key concepts: Cluster analysis, Canopy clustering algorithm, CURE data clustering algorithm, Correlation clustering, Computer science, Fuzzy clustering, Data stream clustering, Single-linkage clustering