A New Framework for Domain-Specific Hidden Web Crawling Based on Data Extraction Techniques
Ali Ibrahim El-Desouky, Hesham Ali, Sally M. Elghamrawy
Abstract
Ali Ibrahim El-Desouky, Hesham Ali, Sally M. Elghamrawy
Abstract
The World Wide Web continues to grow at an exponential rate which makes exploiting all useful information a standing challenge. Search engines like "Google" crawl and index a large amount of information, ignoring valuable data that represent 80% of the content on the Web, this portion of Web called Hidden Web (HW), they are "Hidden" in databases behind search interfaces. In this paper, a framework of a HW crawler is proposed to crawl and extract hidden Web pages. Two unique features of our framework are 1) the classification phase for grouping HW and Publicly Indexable Web (PIW) pages into distinct classes, so that making our crawler performs well in both the domain-specific and random mode of crawling and 2) the capability of dealing with single-attribute and multi-attribute databases. Three novel algorithms proposed in the framework, one for collecting Web pages, one for identifying relevant forms, and one for extracting labels. The effectiveness of proposed algorithms is evaluated through experiments using real Web sites. The preliminary results are very promising. For instance, one of these algorithms proves to be accurate (over 99% precision and 100 % recall).
OpenAlex reports 2 citations for this work. Citation counts describe recorded attention and do not establish research quality.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
The World Wide Web continues to grow at an exponential rate which makes exploiting all useful information a standing challenge. Search engines like "Google" crawl and index a large amount of information, ignoring valuable data that represent 80% of the content on the Web, this portion of Web called Hidden Web (HW), they are "Hidden" in databases behind search interfaces. In this paper, a framework of a HW crawler is proposed to crawl and extract hidden Web pages. Two unique features of our framework are 1) the classification phase for grouping HW and Publicly Indexable Web (PIW) pages into distinct classes, so that making our crawler performs well in both the domain-specific and random mode of crawling and 2) the capability of dealing with single-attribute and multi-attribute databases. Three novel algorithms proposed in the framework, one for collecting Web pages, one for identifying relevant forms, and one for extracting labels. The effectiveness of proposed algorithms is evaluated through experiments using real Web sites. The preliminary results are very promising. For instance, one of these algorithms proves to be accurate (over 99% precision and 100 % recall).
Key concepts: Web crawler, Crawling, Computer science, Web page, Web search engine, Information retrieval, Focused crawler, Domain (mathematical analysis)