2012•World Automation CongressRequires access

Web information extraction based on IEBIDTech

Xiaoyan Ren, Yunxia Fu

Open publisher page 2 citations

Abstract

Based on the survey of contemporary web information extraction theory, this paper studies the frequently-discussed but insufficiently-solved problem: data extracting from web pages containing several structured records, and proposes a new approach called IEBIDTech(Information Extraction based on Improved Dom Tree) which is mainly composed of three steps. At step 1, the given page is initially segmented into several blocks according to html delimiters after the transformation of a DOM tree, and the redundant blocks are subsequently removed, which is then followed by the induction of extraction rules at step 2 and the extraction of structured data at step 3. Large numbers of experiments from diverse domains' web pages show that both recall and precision rates are greater than 90%. That is this approach is able to extract data more accurately.

About this research paper

What this paper is about

Based on the survey of contemporary web information extraction theory, this paper studies the frequently-discussed but insufficiently-solved problem: data extracting from web pages containing several structured records, and proposes a new approach called IEBIDTech(Information Extraction based on Improved Dom Tree) which is mainly composed of three steps. At step 1, the given page is initially segmented into several blocks according to html delimiters after the transformation of a DOM tree, and the redundant blocks are subsequently removed, which is then followed by the induction of extraction rules at step 2 and the extraction of structured data at step 3. Large numbers of experiments from diverse domains' web pages show that both recall and precision rates are greater than 90%. That is this approach is able to extract data more accurately.

Why it matters

OpenAlex reports 2 citations for this work. Citation counts describe recorded attention and do not establish research quality.

Key contribution

A contribution statement is not available in the OpenAlex record.

Method / approach

Method details are not available in the OpenAlex metadata.

Main findings

Findings are not separately available in the OpenAlex metadata.

Limitations

Limitations are not available in the OpenAlex metadata.

Applications

Application details are not available in the OpenAlex metadata.

Available abstract

Based on the survey of contemporary web information extraction theory, this paper studies the frequently-discussed but insufficiently-solved problem: data extracting from web pages containing several structured records, and proposes a new approach called IEBIDTech(Information Extraction based on Improved Dom Tree) which is mainly composed of three steps. At step 1, the given page is initially segmented into several blocks according to html delimiters after the transformation of a DOM tree, and the redundant blocks are subsequently removed, which is then followed by the induction of extraction rules at step 2 and the extraction of structured data at step 3. Large numbers of experiments from diverse domains' web pages show that both recall and precision rates are greater than 90%. That is this approach is able to extract data more accurately.

Key concepts: Computer science, Information extraction, Document Object Model, Web page, Data extraction, Information retrieval, Tree (set theory), Data mining

Related papers

Back to paper searchBrowse research topicsOriginal source
Web information extraction based on IEBIDTech — Research Paper | ScholarLens