An Approach to Design and Implement Parallel Web Crawler
Tithi Dhar, Sayan Mazumder, Susnigdha Dhar, Susovan Karak, Debraj Chatterjee
Abstract
Tithi Dhar, Sayan Mazumder, Susnigdha Dhar, Susovan Karak, Debraj Chatterjee
Abstract
Indefinite cyberspace crawling is an architecture for a parallel searching process occurred in various websites using internet. Web crawler is defined as a software or program which traverses the Web and downloads web documents in an automated manner. It’s depends on the type of questions asked by the user. The pertinence of Web Crawler in the field of web search of this architecture is to capably and successfully crawl the current set of publicly index able web pages to amplify the download rate on shrinking the parallelization. Now a days web grows is becoming extremely difficult to recover the substantial or significant portion of the Web using in a single method. There are many search engines like-Google, Bing, Yandex, Search Encrypt etcetera runs various threads in parallel to perform the above job, so that the download rate gets exploited. The requirement of web crawler is to gather and verify the data in order to store the huge amount of data in such a speed on the Web requires parallel computing. It is definite as a program or as a software that traverses the Website and downloads web documents in a systematic, robotic method. We are going to discuss about the relevance of the Web Crawler. The topic about web crawler is the option to do web search and a small examination on Web Crawler on different problematic fields or domains in web search as mention here.
OpenAlex reports 5 citations for this work. Citation counts describe recorded attention and do not establish research quality.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
Indefinite cyberspace crawling is an architecture for a parallel searching process occurred in various websites using internet. Web crawler is defined as a software or program which traverses the Web and downloads web documents in an automated manner. It’s depends on the type of questions asked by the user. The pertinence of Web Crawler in the field of web search of this architecture is to capably and successfully crawl the current set of publicly index able web pages to amplify the download rate on shrinking the parallelization. Now a days web grows is becoming extremely difficult to recover the substantial or significant portion of the Web using in a single method. There are many search engines like-Google, Bing, Yandex, Search Encrypt etcetera runs various threads in parallel to perform the above job, so that the download rate gets exploited. The requirement of web crawler is to gather and verify the data in order to store the huge amount of data in such a speed on the Web requires parallel computing. It is definite as a program or as a software that traverses the Website and downloads web documents in a systematic, robotic method. We are going to discuss about the relevance of the Web Crawler. The topic about web crawler is the option to do web search and a small examination on Web Crawler on different problematic fields or domains in web search as mention here.
Key concepts: Web crawler, Computer science, Focused crawler, World Wide Web, Web page, Web search engine, Static web page, Web modeling