Web Data Extraction Based on XBRL-GL Taxonomy
Hanyang Luo, Jinling Gao, Hanyang Luo
Abstract
Hanyang Luo, Jinling Gao, Hanyang Luo
Abstract
The Web has become one of the most important connections to various information resources. The most interesting challenge is how to extract important data from a large number of web pages and transform them to more structural, standard and semantic information, which can be queried and analyzed by using matured techniques in database, data warehouse and other fields. We design a wrapper generator by combining the data extraction technique with XBRL technology based on XBRL-GL taxonomy. The wrapper can transform HTML documents to XML forms according to the analysis of HTML document structure, and then use XPath to locate the data. In this way, we can extract the data accurately and store them in a standard form.
OpenAlex reports 3 citations for this work. Citation counts describe recorded attention and do not establish research quality.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
The Web has become one of the most important connections to various information resources. The most interesting challenge is how to extract important data from a large number of web pages and transform them to more structural, standard and semantic information, which can be queried and analyzed by using matured techniques in database, data warehouse and other fields. We design a wrapper generator by combining the data extraction technique with XBRL technology based on XBRL-GL taxonomy. The wrapper can transform HTML documents to XML forms according to the analysis of HTML document structure, and then use XPath to locate the data. In this way, we can extract the data accurately and store them in a standard form.
Key concepts: XBRL, XPath, Computer science, XML, Information retrieval, Data extraction, Database, Data warehouse