Self-organized hierarchical k-means file clustering algorithm based on P2P sharing directories and its application
Kai Lei, Lu Han, Wenhan Chen, HuSheng Yuan, Tao Sun
Abstract
Kai Lei, Lu Han, Wenhan Chen, HuSheng Yuan, Tao Sun
Abstract
In order to improve the recall of search results and calculate file relevancy in P2P sharing systems, a file clustering algorithm using self-organized k-means method was proposed, which is based on the hierarchical structure of file sharing directories and file names' implication of classification. With uploaded file path information built into the indexes, a tree-structure like vector space model was designed. After analyzing the advantages and shortcomings of the traditional k-means method, we implemented a revised self-organized k-means algorithm. This algorithm can easily calculate the distances among file categories and file relevancies in a same category by adjusting two thresholds called as ¿combined factor¿ and ¿correlation factor¿. From the experiment and evaluation results, this model indicated that more target files can be found and improved recall rate to 83.54% and precision of the information retrieval to 85%.
OpenAlex reports 1 citations for this work. Citation counts describe recorded attention and do not establish research quality.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
In order to improve the recall of search results and calculate file relevancy in P2P sharing systems, a file clustering algorithm using self-organized k-means method was proposed, which is based on the hierarchical structure of file sharing directories and file names' implication of classification. With uploaded file path information built into the indexes, a tree-structure like vector space model was designed. After analyzing the advantages and shortcomings of the traditional k-means method, we implemented a revised self-organized k-means algorithm. This algorithm can easily calculate the distances among file categories and file relevancies in a same category by adjusting two thresholds called as ¿combined factor¿ and ¿correlation factor¿. From the experiment and evaluation results, this model indicated that more target files can be found and improved recall rate to 83.54% and precision of the information retrieval to 85%.
Key concepts: Computer science, Upload, Cluster analysis, Data mining, Torrent file, Vector space model, File sharing, Information retrieval