2021Unpublished venueRequires access

A Survey on Crawlers used in developing Search Engine

Smita Deshmukh, Kantilal Vishwakarma

Open publisher page 7 citations

Abstract

A vast collection of information, which is developed over a period, using HTML (Hyper Text Markup Language) formatted documents that are interlinked to each other is called as World Wide Web (WWW). With the increasing size of World Wide Web (WWW), obtaining meaningful information from the web is becoming a tedious task. Search Engine is developed for extracting information from the web. It works as an interface between the web and the user. The three important components of a search engine are: Crawler, Indexer and Page Ranking. The ambiguity of data along with its vast availability on the web is increasing at a greater pace due to the tremendous increase in data on the World Wide Web (WWW). For a naïve user this implies a challenge while surfing on the web for retrieving relevant and required information based on the search. Crawler is a component of search engine responsible for traversing webpages and fetching relevant links from the web. This represents huge dependency of any search engine on the crawlers. So, a detailed study about all the available crawlers for understanding the drawbacks and the insights about the working methodology undertaken is necessary before proceeding to develop a smart crawler. For developing a smart crawler which is future scope, a comparative analysis of widely used crawlers like Focused Crawler, Inference-based Crawler, Incremental Crawler, Parallel Crawler and Distributed Crawler is done.

About this research paper

What this paper is about

A vast collection of information, which is developed over a period, using HTML (Hyper Text Markup Language) formatted documents that are interlinked to each other is called as World Wide Web (WWW). With the increasing size of World Wide Web (WWW), obtaining meaningful information from the web is becoming a tedious task. Search Engine is developed for extracting information from the web. It works as an interface between the web and the user. The three important components of a search engine are: Crawler, Indexer and Page Ranking. The ambiguity of data along with its vast availability on the web is increasing at a greater pace due to the tremendous increase in data on the World Wide Web (WWW). For a naïve user this implies a challenge while surfing on the web for retrieving relevant and required information based on the search. Crawler is a component of search engine responsible for traversing webpages and fetching relevant links from the web. This represents huge dependency of any search engine on the crawlers. So, a detailed study about all the available crawlers for understanding the drawbacks and the insights about the working methodology undertaken is necessary before proceeding to develop a smart crawler. For developing a smart crawler which is future scope, a comparative analysis of widely used crawlers like Focused Crawler, Inference-based Crawler, Incremental Crawler, Parallel Crawler and Distributed Crawler is done.

Why it matters

OpenAlex reports 7 citations for this work. Citation counts describe recorded attention and do not establish research quality.

Key contribution

A contribution statement is not available in the OpenAlex record.

Method / approach

Method details are not available in the OpenAlex metadata.

Main findings

Findings are not separately available in the OpenAlex metadata.

Limitations

Limitations are not available in the OpenAlex metadata.

Applications

Application details are not available in the OpenAlex metadata.

Available abstract

A vast collection of information, which is developed over a period, using HTML (Hyper Text Markup Language) formatted documents that are interlinked to each other is called as World Wide Web (WWW). With the increasing size of World Wide Web (WWW), obtaining meaningful information from the web is becoming a tedious task. Search Engine is developed for extracting information from the web. It works as an interface between the web and the user. The three important components of a search engine are: Crawler, Indexer and Page Ranking. The ambiguity of data along with its vast availability on the web is increasing at a greater pace due to the tremendous increase in data on the World Wide Web (WWW). For a naïve user this implies a challenge while surfing on the web for retrieving relevant and required information based on the search. Crawler is a component of search engine responsible for traversing webpages and fetching relevant links from the web. This represents huge dependency of any search engine on the crawlers. So, a detailed study about all the available crawlers for understanding the drawbacks and the insights about the working methodology undertaken is necessary before proceeding to develop a smart crawler. For developing a smart crawler which is future scope, a comparative analysis of widely used crawlers like Focused Crawler, Inference-based Crawler, Incremental Crawler, Parallel Crawler and Distributed Crawler is done.

Key concepts: Web crawler, Focused crawler, Computer science, World Wide Web, Web search engine, Information retrieval, Web page, Spamdexing

Related papers

Back to paper searchBrowse research topicsOriginal source
A Survey on Crawlers used in developing Search Engine — Research Paper | ScholarLens