2009Unpublished venueRequires access

Adaptive Web Information Extraction Based on DOM Tree

Yang Qin-yao

Open publisher page 1 citations

Abstract

Many Web information extraction methods are related to wrapper induction.It extracts the items by the rules learnt from the Web pages used for training.Although it can get the information accurately,it is hard to be maintained when the template of the Web site is changed,as it needs to learn the rules again.In our research,we put forward a new adaptive Web information extraction.It determines the block which contains all information about the merchandise by using the keywords of a certain topic,which is based on DOM tree structure.The experiments on a great amount of Web pages show that our method can not only extract the information efficiently,but also is irrelevant to the site structure,which can be widely used for many different Web information extractions.

About this research paper

What this paper is about

Many Web information extraction methods are related to wrapper induction.It extracts the items by the rules learnt from the Web pages used for training.Although it can get the information accurately,it is hard to be maintained when the template of the Web site is changed,as it needs to learn the rules again.In our research,we put forward a new adaptive Web information extraction.It determines the block which contains all information about the merchandise by using the keywords of a certain topic,which is based on DOM tree structure.The experiments on a great amount of Web pages show that our method can not only extract the information efficiently,but also is irrelevant to the site structure,which can be widely used for many different Web information extractions.

Why it matters

OpenAlex reports 1 citations for this work. Citation counts describe recorded attention and do not establish research quality.

Key contribution

A contribution statement is not available in the OpenAlex record.

Method / approach

Method details are not available in the OpenAlex metadata.

Main findings

Findings are not separately available in the OpenAlex metadata.

Limitations

Limitations are not available in the OpenAlex metadata.

Applications

Application details are not available in the OpenAlex metadata.

Available abstract

Many Web information extraction methods are related to wrapper induction.It extracts the items by the rules learnt from the Web pages used for training.Although it can get the information accurately,it is hard to be maintained when the template of the Web site is changed,as it needs to learn the rules again.In our research,we put forward a new adaptive Web information extraction.It determines the block which contains all information about the merchandise by using the keywords of a certain topic,which is based on DOM tree structure.The experiments on a great amount of Web pages show that our method can not only extract the information efficiently,but also is irrelevant to the site structure,which can be widely used for many different Web information extractions.

Key concepts: Computer science, Document Object Model, Information extraction, Web page, Information retrieval, Tree (set theory), World Wide Web, Web modeling

Related papers

Back to paper searchBrowse research topicsOriginal source
Adaptive Web Information Extraction Based on DOM Tree — Research Paper | ScholarLens