A framework of deep Web crawler
Xiang Pei-su, Ke Tian, Huang Qinzhen
Abstract
Xiang Pei-su, Ke Tian, Huang Qinzhen
Abstract
As an ever-increasing amount of information on the web today is available through search interfaces, users have to key in a set of keywords in order to access the pages from certain web sites, which are often referred to as the hidden web or the deep web. Since there is no static links to the hidden web pages, search engines cannot discover and index such pages. However, according to recent studies, the content provided by many hidden web sites is often of very high quality and can be extremely valuable to many users. How to build an effective hidden web crawler that can autonomously discover and download pages from the hidden web is studied. A framework of deep web crawler is provided and we propose novel techniques to handle the actual mechanics of crawling the deep Web. Experiment shows that these policies are effective.
OpenAlex reports 26 citations for this work. Citation counts describe recorded attention and do not establish research quality.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
As an ever-increasing amount of information on the web today is available through search interfaces, users have to key in a set of keywords in order to access the pages from certain web sites, which are often referred to as the hidden web or the deep web. Since there is no static links to the hidden web pages, search engines cannot discover and index such pages. However, according to recent studies, the content provided by many hidden web sites is often of very high quality and can be extremely valuable to many users. How to build an effective hidden web crawler that can autonomously discover and download pages from the hidden web is studied. A framework of deep web crawler is provided and we propose novel techniques to handle the actual mechanics of crawling the deep Web. Experiment shows that these policies are effective.
Key concepts: Web crawler, Computer science, Web page, World Wide Web, Focused crawler, Deep Web, Static web page, Web search engine