2009Computer and ModernizationRequires access

Design and Implementation of a Deep Web Crawler

Zhang Hua-xiang

Open publisher page 1 citations

Abstract

As the World Wide Web grows rapidly,more and more data become available in the Deep Web.The data can be obtained by submiting form in the Web pages and arise dynamicly from Deep Web database.Traditional Web crawler only can retrieve Surface Web page by following hyperlinks.Since there is no static links to the hidden Web pages,most search engines cannot discover and index such pages.However,compared to surface Web,the information provided by hidden Web sites is often of more high quality and can be more valuable to us.A method of designing deep Web crawler by use of HtmlUnit framework is proposed in this paper.The crawler which integrate several Web sites can analyze form and fill them automatically to retrieve relevant information from the database.The results of a number of experiments carried out with actual Deep Web sites demonstrate the accuracy of the method.

About this research paper

What this paper is about

As the World Wide Web grows rapidly,more and more data become available in the Deep Web.The data can be obtained by submiting form in the Web pages and arise dynamicly from Deep Web database.Traditional Web crawler only can retrieve Surface Web page by following hyperlinks.Since there is no static links to the hidden Web pages,most search engines cannot discover and index such pages.However,compared to surface Web,the information provided by hidden Web sites is often of more high quality and can be more valuable to us.A method of designing deep Web crawler by use of HtmlUnit framework is proposed in this paper.The crawler which integrate several Web sites can analyze form and fill them automatically to retrieve relevant information from the database.The results of a number of experiments carried out with actual Deep Web sites demonstrate the accuracy of the method.

Why it matters

OpenAlex reports 1 citations for this work. Citation counts describe recorded attention and do not establish research quality.

Key contribution

A contribution statement is not available in the OpenAlex record.

Method / approach

Method details are not available in the OpenAlex metadata.

Main findings

Findings are not separately available in the OpenAlex metadata.

Limitations

Limitations are not available in the OpenAlex metadata.

Applications

Application details are not available in the OpenAlex metadata.

Available abstract

As the World Wide Web grows rapidly,more and more data become available in the Deep Web.The data can be obtained by submiting form in the Web pages and arise dynamicly from Deep Web database.Traditional Web crawler only can retrieve Surface Web page by following hyperlinks.Since there is no static links to the hidden Web pages,most search engines cannot discover and index such pages.However,compared to surface Web,the information provided by hidden Web sites is often of more high quality and can be more valuable to us.A method of designing deep Web crawler by use of HtmlUnit framework is proposed in this paper.The crawler which integrate several Web sites can analyze form and fill them automatically to retrieve relevant information from the database.The results of a number of experiments carried out with actual Deep Web sites demonstrate the accuracy of the method.

Key concepts: Web crawler, Computer science, Web page, Focused crawler, World Wide Web, Static web page, Web search engine, Information retrieval

Related papers

Back to paper searchBrowse research topicsOriginal source
Design and Implementation of a Deep Web Crawler — Research Paper | ScholarLens