Web Crawler for Searching Deep Web Sites
Tejaswini Arun Patil, Santosh V. Chobe
Abstract
Tejaswini Arun Patil, Santosh V. Chobe
Abstract
In World Wide Web deep web searching is a most important issue till date. Searching relevant information on a web require different techniques. Crawler is a technique which will help to find out relevant information on web. Nowaday, humans are searching data with the help of search engines such as Google and Yahoo but these search engines will not cover all information accurately. To avoid these problems, we propose crawler as a framework for deep web searching. To improve efficiency and to achieve wide coverage we design crawler which will overcome these problems. □ Our crawler first performs site based searching in this it will perform ranking of relevant sites. In second stage crawler achieves in-site searching using adaptive link ranking. These crawlers may contain duplicate URL's or contents. These URLs can be removed with the help of DUSTER technique. This framework focuses on improving the efficiency of web crawler by pre-query processing approach.
OpenAlex reports 7 citations for this work. Citation counts describe recorded attention and do not establish research quality.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
In World Wide Web deep web searching is a most important issue till date. Searching relevant information on a web require different techniques. Crawler is a technique which will help to find out relevant information on web. Nowaday, humans are searching data with the help of search engines such as Google and Yahoo but these search engines will not cover all information accurately. To avoid these problems, we propose crawler as a framework for deep web searching. To improve efficiency and to achieve wide coverage we design crawler which will overcome these problems. □ Our crawler first performs site based searching in this it will perform ranking of relevant sites. In second stage crawler achieves in-site searching using adaptive link ranking. These crawlers may contain duplicate URL's or contents. These URLs can be removed with the help of DUSTER technique. This framework focuses on improving the efficiency of web crawler by pre-query processing approach.
Key concepts: Web crawler, Computer science, Focused crawler, Information retrieval, Ranking (information retrieval), World Wide Web, Web search engine, Search engine