2009Unpublished venueRequires access

Extraction of Information from Web Pages Based on Extended DOM Tree

Tian Wei

Open publisher page 3 citations

Abstract

A method of information extraction from Web pages was presented,and it is based on extended DOM tree.Web pages were firstly transformed to DOM tree,then the DOM tree was extended by adding semantic expression to node and influence degree was calculated for each node.According to influence degree of nodes,the DOM tree was pruned,and it can automatically extract the useful relevant content from Web pages.This approach is a universal me-thod,which does not require to pre-know the structure of the Web page.The results of the information extraction are used not only for browsing but also for further Web information process,such as internet data mining,topic-based search engine.

About this research paper

What this paper is about

A method of information extraction from Web pages was presented,and it is based on extended DOM tree.Web pages were firstly transformed to DOM tree,then the DOM tree was extended by adding semantic expression to node and influence degree was calculated for each node.According to influence degree of nodes,the DOM tree was pruned,and it can automatically extract the useful relevant content from Web pages.This approach is a universal me-thod,which does not require to pre-know the structure of the Web page.The results of the information extraction are used not only for browsing but also for further Web information process,such as internet data mining,topic-based search engine.

Why it matters

OpenAlex reports 3 citations for this work. Citation counts describe recorded attention and do not establish research quality.

Key contribution

A contribution statement is not available in the OpenAlex record.

Method / approach

Method details are not available in the OpenAlex metadata.

Main findings

Findings are not separately available in the OpenAlex metadata.

Limitations

Limitations are not available in the OpenAlex metadata.

Applications

Application details are not available in the OpenAlex metadata.

Available abstract

A method of information extraction from Web pages was presented,and it is based on extended DOM tree.Web pages were firstly transformed to DOM tree,then the DOM tree was extended by adding semantic expression to node and influence degree was calculated for each node.According to influence degree of nodes,the DOM tree was pruned,and it can automatically extract the useful relevant content from Web pages.This approach is a universal me-thod,which does not require to pre-know the structure of the Web page.The results of the information extraction are used not only for browsing but also for further Web information process,such as internet data mining,topic-based search engine.

Key concepts: Computer science, Document Object Model, Web page, Tree (set theory), Information retrieval, World Wide Web, Static web page, Node (physics)

Related papers

Back to paper searchBrowse research topicsOriginal source
Extraction of Information from Web Pages Based on Extended DOM Tree — Research Paper | ScholarLens