2021•2021 International Conference on Smart Generation Computing, Communication and Networking (SMART GENCON)Requires access

An Approach to Design and Implement Parallel Web Crawler

Tithi Dhar, Sayan Mazumder, Susnigdha Dhar, Susovan Karak, Debraj Chatterjee

Open publisher page 5 citations

Abstract

Indefinite cyberspace crawling is an architecture for a parallel searching process occurred in various websites using internet. Web crawler is defined as a software or program which traverses the Web and downloads web documents in an automated manner. It’s depends on the type of questions asked by the user. The pertinence of Web Crawler in the field of web search of this architecture is to capably and successfully crawl the current set of publicly index able web pages to amplify the download rate on shrinking the parallelization. Now a days web grows is becoming extremely difficult to recover the substantial or significant portion of the Web using in a single method. There are many search engines like-Google, Bing, Yandex, Search Encrypt etcetera runs various threads in parallel to perform the above job, so that the download rate gets exploited. The requirement of web crawler is to gather and verify the data in order to store the huge amount of data in such a speed on the Web requires parallel computing. It is definite as a program or as a software that traverses the Website and downloads web documents in a systematic, robotic method. We are going to discuss about the relevance of the Web Crawler. The topic about web crawler is the option to do web search and a small examination on Web Crawler on different problematic fields or domains in web search as mention here.

About this research paper

What this paper is about

Indefinite cyberspace crawling is an architecture for a parallel searching process occurred in various websites using internet. Web crawler is defined as a software or program which traverses the Web and downloads web documents in an automated manner. It’s depends on the type of questions asked by the user. The pertinence of Web Crawler in the field of web search of this architecture is to capably and successfully crawl the current set of publicly index able web pages to amplify the download rate on shrinking the parallelization. Now a days web grows is becoming extremely difficult to recover the substantial or significant portion of the Web using in a single method. There are many search engines like-Google, Bing, Yandex, Search Encrypt etcetera runs various threads in parallel to perform the above job, so that the download rate gets exploited. The requirement of web crawler is to gather and verify the data in order to store the huge amount of data in such a speed on the Web requires parallel computing. It is definite as a program or as a software that traverses the Website and downloads web documents in a systematic, robotic method. We are going to discuss about the relevance of the Web Crawler. The topic about web crawler is the option to do web search and a small examination on Web Crawler on different problematic fields or domains in web search as mention here.

Why it matters

OpenAlex reports 5 citations for this work. Citation counts describe recorded attention and do not establish research quality.

Key contribution

A contribution statement is not available in the OpenAlex record.

Method / approach

Method details are not available in the OpenAlex metadata.

Main findings

Findings are not separately available in the OpenAlex metadata.

Limitations

Limitations are not available in the OpenAlex metadata.

Applications

Application details are not available in the OpenAlex metadata.

Available abstract

Indefinite cyberspace crawling is an architecture for a parallel searching process occurred in various websites using internet. Web crawler is defined as a software or program which traverses the Web and downloads web documents in an automated manner. It’s depends on the type of questions asked by the user. The pertinence of Web Crawler in the field of web search of this architecture is to capably and successfully crawl the current set of publicly index able web pages to amplify the download rate on shrinking the parallelization. Now a days web grows is becoming extremely difficult to recover the substantial or significant portion of the Web using in a single method. There are many search engines like-Google, Bing, Yandex, Search Encrypt etcetera runs various threads in parallel to perform the above job, so that the download rate gets exploited. The requirement of web crawler is to gather and verify the data in order to store the huge amount of data in such a speed on the Web requires parallel computing. It is definite as a program or as a software that traverses the Website and downloads web documents in a systematic, robotic method. We are going to discuss about the relevance of the Web Crawler. The topic about web crawler is the option to do web search and a small examination on Web Crawler on different problematic fields or domains in web search as mention here.

Key concepts: Web crawler, Computer science, Focused crawler, World Wide Web, Web page, Web search engine, Static web page, Web modeling

Related papers

Back to paper searchBrowse research topicsOriginal source
An Approach to Design and Implement Parallel Web Crawler — Research Paper | ScholarLens