2014International Journal of Knowledge and Web IntelligenceOpen access

Review of web crawlers

S.R. Sreeja, Sangita Chaudhari

Open full text 7 citations

Abstract

The web is a repository of large amount of data.Information available in the web is organised in the form of pages.Due to the presence of unlimited amount of information, searching and finding out appropriate information from the web is a task which needs expertise.Web crawlers are programmes that assist search engines by automating the task of visiting web pages and downloading their contents.They also help in ranking the downloaded web pages.Thus, the search engines can produce a list of web pages ordered by their relevance and can display this list as a result of the search.Crawling also helps to validate web pages, analyse them, notify about page-updation, visualise web pages and sometimes for collecting e-mail addresses for spam purposes.They can be of different types, each one using different strategies and techniques to crawl web pages.This paper presents a review of various types of web crawlers.

Open-access reader

About this research paper

What this paper is about

The web is a repository of large amount of data.Information available in the web is organised in the form of pages.Due to the presence of unlimited amount of information, searching and finding out appropriate information from the web is a task which needs expertise.Web crawlers are programmes that assist search engines by automating the task of visiting web pages and downloading their contents.They also help in ranking the downloaded web pages.Thus, the search engines can produce a list of web pages ordered by their relevance and can display this list as a result of the search.Crawling also helps to validate web pages, analyse them, notify about page-updation, visualise web pages and sometimes for collecting e-mail addresses for spam purposes.They can be of different types, each one using different strategies and techniques to crawl web pages.This paper presents a review of various types of web crawlers.

Why it matters

OpenAlex reports 7 citations for this work. Citation counts describe recorded attention and do not establish research quality.

Key contribution

A contribution statement is not available in the OpenAlex record.

Method / approach

Method details are not available in the OpenAlex metadata.

Main findings

Findings are not separately available in the OpenAlex metadata.

Limitations

Limitations are not available in the OpenAlex metadata.

Applications

Application details are not available in the OpenAlex metadata.

Available abstract

The web is a repository of large amount of data.Information available in the web is organised in the form of pages.Due to the presence of unlimited amount of information, searching and finding out appropriate information from the web is a task which needs expertise.Web crawlers are programmes that assist search engines by automating the task of visiting web pages and downloading their contents.They also help in ranking the downloaded web pages.Thus, the search engines can produce a list of web pages ordered by their relevance and can display this list as a result of the search.Crawling also helps to validate web pages, analyse them, notify about page-updation, visualise web pages and sometimes for collecting e-mail addresses for spam purposes.They can be of different types, each one using different strategies and techniques to crawl web pages.This paper presents a review of various types of web crawlers.

Key concepts: Computer science, World Wide Web, Web crawler, Information retrieval

Related papers

Back to paper searchBrowse research topicsOriginal source
Review of web crawlers — Research Paper | ScholarLens