2009Computer Technology and DevelopmentRequires access

Extracting Web Data Using Tree Structure

Fu Tao

Open publisher page 0 citations

Abstract

Information retrieval is the technique that searches the useful information from abundant data.But the common way of web information retrieval is based on the analysis of HTML documents on the web.Provides a kind of method that transform HTML into XML first,and then get information.XML is a language standard,which is used to deseribe the format of data documents in data exehange.It makes the strueture,content and representation separated.The data can be identified uniquely by XML.And then,it helps people organize and seareh the data.The proposed approach can effectively extract desired data with high accuracies and with linear complexity.

About this research paper

What this paper is about

Information retrieval is the technique that searches the useful information from abundant data.But the common way of web information retrieval is based on the analysis of HTML documents on the web.Provides a kind of method that transform HTML into XML first,and then get information.XML is a language standard,which is used to deseribe the format of data documents in data exehange.It makes the strueture,content and representation separated.The data can be identified uniquely by XML.And then,it helps people organize and seareh the data.The proposed approach can effectively extract desired data with high accuracies and with linear complexity.

Why it matters

A significance statement is not available in the OpenAlex record.

Key contribution

A contribution statement is not available in the OpenAlex record.

Method / approach

Method details are not available in the OpenAlex metadata.

Main findings

Findings are not separately available in the OpenAlex metadata.

Limitations

Limitations are not available in the OpenAlex metadata.

Applications

Application details are not available in the OpenAlex metadata.

Available abstract

Information retrieval is the technique that searches the useful information from abundant data.But the common way of web information retrieval is based on the analysis of HTML documents on the web.Provides a kind of method that transform HTML into XML first,and then get information.XML is a language standard,which is used to deseribe the format of data documents in data exehange.It makes the strueture,content and representation separated.The data can be identified uniquely by XML.And then,it helps people organize and seareh the data.The proposed approach can effectively extract desired data with high accuracies and with linear complexity.

Key concepts: Computer science, Information retrieval, Document Structure Description, XML, XML validation, Efficient XML Interchange, HTML element, XML database

Related papers

Back to paper searchBrowse research topicsOriginal source
Extracting Web Data Using Tree Structure — Research Paper | ScholarLens