2005Journal of Changchun Post and Telecommunication InstituteRequires access

Using Hyperlink Information to Improve Crawler's Searching Strategy

Wanli Zuo

Open publisher page 1 citations

Abstract

A crawler must face two problems when it searches pages in internet. One is that an internet search engine cannot contain entire Web pages due to huge volumes of data in internet. Because of the constraint of hardware resource, the other is that the Web pages stored in the internet search engine are limited. Crawling in Web space according to the strategy of the traditional breadth-first search, if a crawler respects the importance of every page equally, the quality of Web pages collected by the crawler is not high. The algorithm proposed in the present paper makes the best use of hyperlink information contained in the Web pages to the great extent, overcomes the limitations of the blind crawling strategy owned by the traditional breadth-first search. It is improved by using the hyperlink information contained in the Web pages on the base of the traditional breadth-first search. The experiment results show the Web pages crawled by using this algorithm those are relevant to a pre-defined set of topics are over 50%.

About this research paper

What this paper is about

A crawler must face two problems when it searches pages in internet. One is that an internet search engine cannot contain entire Web pages due to huge volumes of data in internet. Because of the constraint of hardware resource, the other is that the Web pages stored in the internet search engine are limited. Crawling in Web space according to the strategy of the traditional breadth-first search, if a crawler respects the importance of every page equally, the quality of Web pages collected by the crawler is not high. The algorithm proposed in the present paper makes the best use of hyperlink information contained in the Web pages to the great extent, overcomes the limitations of the blind crawling strategy owned by the traditional breadth-first search. It is improved by using the hyperlink information contained in the Web pages on the base of the traditional breadth-first search. The experiment results show the Web pages crawled by using this algorithm those are relevant to a pre-defined set of topics are over 50%.

Why it matters

OpenAlex reports 1 citations for this work. Citation counts describe recorded attention and do not establish research quality.

Key contribution

A contribution statement is not available in the OpenAlex record.

Method / approach

Method details are not available in the OpenAlex metadata.

Main findings

Findings are not separately available in the OpenAlex metadata.

Limitations

Limitations are not available in the OpenAlex metadata.

Applications

Application details are not available in the OpenAlex metadata.

Available abstract

A crawler must face two problems when it searches pages in internet. One is that an internet search engine cannot contain entire Web pages due to huge volumes of data in internet. Because of the constraint of hardware resource, the other is that the Web pages stored in the internet search engine are limited. Crawling in Web space according to the strategy of the traditional breadth-first search, if a crawler respects the importance of every page equally, the quality of Web pages collected by the crawler is not high. The algorithm proposed in the present paper makes the best use of hyperlink information contained in the Web pages to the great extent, overcomes the limitations of the blind crawling strategy owned by the traditional breadth-first search. It is improved by using the hyperlink information contained in the Web pages on the base of the traditional breadth-first search. The experiment results show the Web pages crawled by using this algorithm those are relevant to a pre-defined set of topics are over 50%.

Key concepts: Web crawler, Focused crawler, Computer science, Web page, World Wide Web, Hyperlink, Web search engine, Static web page

Related papers

Back to paper searchBrowse research topicsOriginal source
Using Hyperlink Information to Improve Crawler's Searching Strategy — Research Paper | ScholarLens