Parallel Web crawler system with increment update
Qingkui Chen
Abstract
Qingkui Chen
Abstract
This paper discussed the architecture of parallel Web crawler system.Incremental crawling method was used to the system to improve the efficiency of massive information updating.Meanwhile,considering the difference of crawler in the system and with the aim of fully usage of crawler in cluster system,Cosine vector parallel crawling model was introduced to solve this problem.After giving the definitions of crawling task vector and crawler vector,relevant parallel crawling algorithms were designed.The results confirm that the system is effective in distribution adaptability and runs well in maintaining the freshness of the Web repository.
A significance statement is not available in the OpenAlex record.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
This paper discussed the architecture of parallel Web crawler system.Incremental crawling method was used to the system to improve the efficiency of massive information updating.Meanwhile,considering the difference of crawler in the system and with the aim of fully usage of crawler in cluster system,Cosine vector parallel crawling model was introduced to solve this problem.After giving the definitions of crawling task vector and crawler vector,relevant parallel crawling algorithms were designed.The results confirm that the system is effective in distribution adaptability and runs well in maintaining the freshness of the Web repository.
Key concepts: Web crawler, Crawling, Focused crawler, Computer science, Adaptability, Task (project management), Information retrieval, Web page