2008Unpublished venueRequires access

A framework of deep Web crawler

Xiang Pei-su, Ke Tian, Huang Qinzhen

Open publisher page 26 citations

Abstract

As an ever-increasing amount of information on the web today is available through search interfaces, users have to key in a set of keywords in order to access the pages from certain web sites, which are often referred to as the hidden web or the deep web. Since there is no static links to the hidden web pages, search engines cannot discover and index such pages. However, according to recent studies, the content provided by many hidden web sites is often of very high quality and can be extremely valuable to many users. How to build an effective hidden web crawler that can autonomously discover and download pages from the hidden web is studied. A framework of deep web crawler is provided and we propose novel techniques to handle the actual mechanics of crawling the deep Web. Experiment shows that these policies are effective.

About this research paper

What this paper is about

As an ever-increasing amount of information on the web today is available through search interfaces, users have to key in a set of keywords in order to access the pages from certain web sites, which are often referred to as the hidden web or the deep web. Since there is no static links to the hidden web pages, search engines cannot discover and index such pages. However, according to recent studies, the content provided by many hidden web sites is often of very high quality and can be extremely valuable to many users. How to build an effective hidden web crawler that can autonomously discover and download pages from the hidden web is studied. A framework of deep web crawler is provided and we propose novel techniques to handle the actual mechanics of crawling the deep Web. Experiment shows that these policies are effective.

Why it matters

OpenAlex reports 26 citations for this work. Citation counts describe recorded attention and do not establish research quality.

Key contribution

A contribution statement is not available in the OpenAlex record.

Method / approach

Method details are not available in the OpenAlex metadata.

Main findings

Findings are not separately available in the OpenAlex metadata.

Limitations

Limitations are not available in the OpenAlex metadata.

Applications

Application details are not available in the OpenAlex metadata.

Available abstract

As an ever-increasing amount of information on the web today is available through search interfaces, users have to key in a set of keywords in order to access the pages from certain web sites, which are often referred to as the hidden web or the deep web. Since there is no static links to the hidden web pages, search engines cannot discover and index such pages. However, according to recent studies, the content provided by many hidden web sites is often of very high quality and can be extremely valuable to many users. How to build an effective hidden web crawler that can autonomously discover and download pages from the hidden web is studied. A framework of deep web crawler is provided and we propose novel techniques to handle the actual mechanics of crawling the deep Web. Experiment shows that these policies are effective.

Key concepts: Web crawler, Computer science, Web page, World Wide Web, Focused crawler, Deep Web, Static web page, Web search engine

Related papers

Back to paper searchBrowse research topicsOriginal source
A framework of deep Web crawler — Research Paper | ScholarLens