2020•Unpublished venueRequires access

A REVIEW OF WEB CRAWLERS FOR INFORMATION RETRIEVAL

B. M. Ajose-Ismail, Q. A. Osanyin

Open publisher page 0 citations

Abstract

Performance of any search engine relies heavily on its Web crawler. Web crawlers are the programs that get webpages from the web by following hyperlinks. These webpages are indexed by a search engine and can be retrieved by a user query. In the area of web crawling which is a subfield of Information Retrieval, we still lack an exhaustive study that covers all crawling techniques. This study follows the guidelines of systematic literature review and applies it to the field of Web crawling. Existing literature about the web crawler is classified into different key subareas. Each subarea is further divided according to the techniques being used. We have highlighted future areas of research. We call for an increased awareness in various fields of the web crawler.

About this research paper

What this paper is about

Performance of any search engine relies heavily on its Web crawler. Web crawlers are the programs that get webpages from the web by following hyperlinks. These webpages are indexed by a search engine and can be retrieved by a user query. In the area of web crawling which is a subfield of Information Retrieval, we still lack an exhaustive study that covers all crawling techniques. This study follows the guidelines of systematic literature review and applies it to the field of Web crawling. Existing literature about the web crawler is classified into different key subareas. Each subarea is further divided according to the techniques being used. We have highlighted future areas of research. We call for an increased awareness in various fields of the web crawler.

Why it matters

A significance statement is not available in the OpenAlex record.

Key contribution

A contribution statement is not available in the OpenAlex record.

Method / approach

Method details are not available in the OpenAlex metadata.

Main findings

Findings are not separately available in the OpenAlex metadata.

Limitations

Limitations are not available in the OpenAlex metadata.

Applications

Application details are not available in the OpenAlex metadata.

Available abstract

Performance of any search engine relies heavily on its Web crawler. Web crawlers are the programs that get webpages from the web by following hyperlinks. These webpages are indexed by a search engine and can be retrieved by a user query. In the area of web crawling which is a subfield of Information Retrieval, we still lack an exhaustive study that covers all crawling techniques. This study follows the guidelines of systematic literature review and applies it to the field of Web crawling. Existing literature about the web crawler is classified into different key subareas. Each subarea is further divided according to the techniques being used. We have highlighted future areas of research. We call for an increased awareness in various fields of the web crawler.

Key concepts: Web crawler, World Wide Web, Focused crawler, Computer science, Web page, Information retrieval, Web search engine, Crawling

Back to paper searchBrowse research topicsOriginal source
A REVIEW OF WEB CRAWLERS FOR INFORMATION RETRIEVAL — Research Paper | ScholarLens