2006•Unpublished venueRequires access

A New Framework for Domain-Specific Hidden Web Crawling Based on Data Extraction Techniques

Ali Ibrahim El-Desouky, Hesham Ali, Sally M. Elghamrawy

Open publisher page 2 citations

Abstract

The World Wide Web continues to grow at an exponential rate which makes exploiting all useful information a standing challenge. Search engines like "Google" crawl and index a large amount of information, ignoring valuable data that represent 80% of the content on the Web, this portion of Web called Hidden Web (HW), they are "Hidden" in databases behind search interfaces. In this paper, a framework of a HW crawler is proposed to crawl and extract hidden Web pages. Two unique features of our framework are 1) the classification phase for grouping HW and Publicly Indexable Web (PIW) pages into distinct classes, so that making our crawler performs well in both the domain-specific and random mode of crawling and 2) the capability of dealing with single-attribute and multi-attribute databases. Three novel algorithms proposed in the framework, one for collecting Web pages, one for identifying relevant forms, and one for extracting labels. The effectiveness of proposed algorithms is evaluated through experiments using real Web sites. The preliminary results are very promising. For instance, one of these algorithms proves to be accurate (over 99% precision and 100 % recall).

About this research paper

What this paper is about

The World Wide Web continues to grow at an exponential rate which makes exploiting all useful information a standing challenge. Search engines like "Google" crawl and index a large amount of information, ignoring valuable data that represent 80% of the content on the Web, this portion of Web called Hidden Web (HW), they are "Hidden" in databases behind search interfaces. In this paper, a framework of a HW crawler is proposed to crawl and extract hidden Web pages. Two unique features of our framework are 1) the classification phase for grouping HW and Publicly Indexable Web (PIW) pages into distinct classes, so that making our crawler performs well in both the domain-specific and random mode of crawling and 2) the capability of dealing with single-attribute and multi-attribute databases. Three novel algorithms proposed in the framework, one for collecting Web pages, one for identifying relevant forms, and one for extracting labels. The effectiveness of proposed algorithms is evaluated through experiments using real Web sites. The preliminary results are very promising. For instance, one of these algorithms proves to be accurate (over 99% precision and 100 % recall).

Why it matters

OpenAlex reports 2 citations for this work. Citation counts describe recorded attention and do not establish research quality.

Key contribution

A contribution statement is not available in the OpenAlex record.

Method / approach

Method details are not available in the OpenAlex metadata.

Main findings

Findings are not separately available in the OpenAlex metadata.

Limitations

Limitations are not available in the OpenAlex metadata.

Applications

Application details are not available in the OpenAlex metadata.

Available abstract

The World Wide Web continues to grow at an exponential rate which makes exploiting all useful information a standing challenge. Search engines like "Google" crawl and index a large amount of information, ignoring valuable data that represent 80% of the content on the Web, this portion of Web called Hidden Web (HW), they are "Hidden" in databases behind search interfaces. In this paper, a framework of a HW crawler is proposed to crawl and extract hidden Web pages. Two unique features of our framework are 1) the classification phase for grouping HW and Publicly Indexable Web (PIW) pages into distinct classes, so that making our crawler performs well in both the domain-specific and random mode of crawling and 2) the capability of dealing with single-attribute and multi-attribute databases. Three novel algorithms proposed in the framework, one for collecting Web pages, one for identifying relevant forms, and one for extracting labels. The effectiveness of proposed algorithms is evaluated through experiments using real Web sites. The preliminary results are very promising. For instance, one of these algorithms proves to be accurate (over 99% precision and 100 % recall).

Key concepts: Web crawler, Crawling, Computer science, Web page, Web search engine, Information retrieval, Focused crawler, Domain (mathematical analysis)

Related papers

Back to paper searchBrowse research topicsOriginal source
A New Framework for Domain-Specific Hidden Web Crawling Based on Data Extraction Techniques — Research Paper | ScholarLens