Information Extraction for the Web Sources Based on DOM and WebTemPlate
Tang Jianxiong
Abstract
Tang Jianxiong
Abstract
Information extraction studled by the Paper is based on D0M (Document object Model) and web template. According to the definition of DOM,the paper describes the structure of web Pages by constructing HTML Parsing tree. Before Information extraction,the noise information can be filtrated in web pages by inducting web template. Then,the paper uses the extraction rule based on relative path to extract information in web pages. At last,the paper presents the result of inducting web template3s and extracting web pages. From the result,it is evident that the way of inducting web templates and the way of extracting web pages are correct and effective.
OpenAlex reports 1 citations for this work. Citation counts describe recorded attention and do not establish research quality.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
Information extraction studled by the Paper is based on D0M (Document object Model) and web template. According to the definition of DOM,the paper describes the structure of web Pages by constructing HTML Parsing tree. Before Information extraction,the noise information can be filtrated in web pages by inducting web template. Then,the paper uses the extraction rule based on relative path to extract information in web pages. At last,the paper presents the result of inducting web template3s and extracting web pages. From the result,it is evident that the way of inducting web templates and the way of extracting web pages are correct and effective.
Key concepts: Computer science, Document Object Model, Web page, World Wide Web, Static web page, Web modeling, Information retrieval, Web standards