Extracting Web Data Using Tree Structure
Fu Tao
Abstract
Fu Tao
Abstract
Information retrieval is the technique that searches the useful information from abundant data.But the common way of web information retrieval is based on the analysis of HTML documents on the web.Provides a kind of method that transform HTML into XML first,and then get information.XML is a language standard,which is used to deseribe the format of data documents in data exehange.It makes the strueture,content and representation separated.The data can be identified uniquely by XML.And then,it helps people organize and seareh the data.The proposed approach can effectively extract desired data with high accuracies and with linear complexity.
A significance statement is not available in the OpenAlex record.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
Information retrieval is the technique that searches the useful information from abundant data.But the common way of web information retrieval is based on the analysis of HTML documents on the web.Provides a kind of method that transform HTML into XML first,and then get information.XML is a language standard,which is used to deseribe the format of data documents in data exehange.It makes the strueture,content and representation separated.The data can be identified uniquely by XML.And then,it helps people organize and seareh the data.The proposed approach can effectively extract desired data with high accuracies and with linear complexity.
Key concepts: Computer science, Information retrieval, Document Structure Description, XML, XML validation, Efficient XML Interchange, HTML element, XML database