A Survey Paper on An Efficient Harvesting scheme for Deep Web Interfaces based on Two-Stage Crawler
Shinde Pavan Bhausaheb, Sonkar Shriniwas K
Abstract
Shinde Pavan Bhausaheb, Sonkar Shriniwas K
Abstract
The web pages available in the Internet are growing tremendously so that searching relevant information in the Internet is tedious job. A lot of this information is hidden behind query forms that interface to unexplored databases containing high quality structured data. General search engines cannot extract and index this hidden part of the Web, retrieving this hidden data is challenging task. So that large number of web data resources and the dynamic nature of deep web sites, achieving wide coverage and high efficiency is a challenging task. We propose a two-stage framework, namely SmartCrawler, for effective searching deep web interfaces. First stage of SmartCrawler performs site-based searching for pages with the help of web crawler, avoiding visiting a large number of sites. To produce more relavant results for a focused crawl, SmartCrawler ranks links to prioritize highly relevant pages for a given topic. Then in second stage, it achieves fast in-site searching by extracting most relevant links with an adaptive link-ranking.
A significance statement is not available in the OpenAlex record.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
The web pages available in the Internet are growing tremendously so that searching relevant information in the Internet is tedious job. A lot of this information is hidden behind query forms that interface to unexplored databases containing high quality structured data. General search engines cannot extract and index this hidden part of the Web, retrieving this hidden data is challenging task. So that large number of web data resources and the dynamic nature of deep web sites, achieving wide coverage and high efficiency is a challenging task. We propose a two-stage framework, namely SmartCrawler, for effective searching deep web interfaces. First stage of SmartCrawler performs site-based searching for pages with the help of web crawler, avoiding visiting a large number of sites. To produce more relavant results for a focused crawl, SmartCrawler ranks links to prioritize highly relevant pages for a given topic. Then in second stage, it achieves fast in-site searching by extracting most relevant links with an adaptive link-ranking.
Key concepts: Web crawler, Computer science, Information retrieval, Ranking (information retrieval), The Internet, Task (project management), Web page, Deep Web