2012Unpublished venueRequires access

An approach for text extraction from web news page

Mingsheng Hu, Jia Zhijuan, Xiangyu Zhang

Open publisher page 5 citations

Abstract

With the rapid development of Internet and Web technology, Web page has become a main carrier of information publishing. In connection with the problems of current complex implementation, high error rate and low extraction speed of Web information extraction technology, this paper proposes a new method of Web extraction based on the characteristics of structure of Web page. This method is to use tree structure of DOM (Document Object Model) when analyzing web page, parsing the Web page into DOM tree to sort the scattered web pages, by the using of the characteristics of Chinese web pages similar in information structure and aggregated distribution to achieve simply with good versatility. At the same time, this method can reduce the complexity when dealing with the structure of web page and increase the speed of the Web information extraction. At present, the method has been applied to the news page automatic classification system, which is good to meet the system's requirements.

About this research paper

What this paper is about

With the rapid development of Internet and Web technology, Web page has become a main carrier of information publishing. In connection with the problems of current complex implementation, high error rate and low extraction speed of Web information extraction technology, this paper proposes a new method of Web extraction based on the characteristics of structure of Web page. This method is to use tree structure of DOM (Document Object Model) when analyzing web page, parsing the Web page into DOM tree to sort the scattered web pages, by the using of the characteristics of Chinese web pages similar in information structure and aggregated distribution to achieve simply with good versatility. At the same time, this method can reduce the complexity when dealing with the structure of web page and increase the speed of the Web information extraction. At present, the method has been applied to the news page automatic classification system, which is good to meet the system's requirements.

Why it matters

OpenAlex reports 5 citations for this work. Citation counts describe recorded attention and do not establish research quality.

Key contribution

A contribution statement is not available in the OpenAlex record.

Method / approach

Method details are not available in the OpenAlex metadata.

Main findings

Findings are not separately available in the OpenAlex metadata.

Limitations

Limitations are not available in the OpenAlex metadata.

Applications

Application details are not available in the OpenAlex metadata.

Available abstract

With the rapid development of Internet and Web technology, Web page has become a main carrier of information publishing. In connection with the problems of current complex implementation, high error rate and low extraction speed of Web information extraction technology, this paper proposes a new method of Web extraction based on the characteristics of structure of Web page. This method is to use tree structure of DOM (Document Object Model) when analyzing web page, parsing the Web page into DOM tree to sort the scattered web pages, by the using of the characteristics of Chinese web pages similar in information structure and aggregated distribution to achieve simply with good versatility. At the same time, this method can reduce the complexity when dealing with the structure of web page and increase the speed of the Web information extraction. At present, the method has been applied to the news page automatic classification system, which is good to meet the system's requirements.

Key concepts: Document Object Model, Web page, Computer science, Static web page, Backlink, Web modeling, World Wide Web, Web navigation

Related papers

Back to paper searchBrowse research topicsOriginal source
An approach for text extraction from web news page — Research Paper | ScholarLens