An approach for text extraction from web news page
Mingsheng Hu, Jia Zhijuan, Xiangyu Zhang
Abstract
Mingsheng Hu, Jia Zhijuan, Xiangyu Zhang
Abstract
With the rapid development of Internet and Web technology, Web page has become a main carrier of information publishing. In connection with the problems of current complex implementation, high error rate and low extraction speed of Web information extraction technology, this paper proposes a new method of Web extraction based on the characteristics of structure of Web page. This method is to use tree structure of DOM (Document Object Model) when analyzing web page, parsing the Web page into DOM tree to sort the scattered web pages, by the using of the characteristics of Chinese web pages similar in information structure and aggregated distribution to achieve simply with good versatility. At the same time, this method can reduce the complexity when dealing with the structure of web page and increase the speed of the Web information extraction. At present, the method has been applied to the news page automatic classification system, which is good to meet the system's requirements.
OpenAlex reports 5 citations for this work. Citation counts describe recorded attention and do not establish research quality.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
With the rapid development of Internet and Web technology, Web page has become a main carrier of information publishing. In connection with the problems of current complex implementation, high error rate and low extraction speed of Web information extraction technology, this paper proposes a new method of Web extraction based on the characteristics of structure of Web page. This method is to use tree structure of DOM (Document Object Model) when analyzing web page, parsing the Web page into DOM tree to sort the scattered web pages, by the using of the characteristics of Chinese web pages similar in information structure and aggregated distribution to achieve simply with good versatility. At the same time, this method can reduce the complexity when dealing with the structure of web page and increase the speed of the Web information extraction. At present, the method has been applied to the news page automatic classification system, which is good to meet the system's requirements.
Key concepts: Document Object Model, Web page, Computer science, Static web page, Backlink, Web modeling, World Wide Web, Web navigation