Auto-extraction methods of Web pagelet
Li Wei
Abstract
Li Wei
Abstract
Besides the needed data,there are lots of navigation information and advertisements in the Web pages.A DOM tree comparison algorithm was proposed.It compared several pages within a class,and recognized the main contents in pages.Experiment results show that it is feasible and effective.
A significance statement is not available in the OpenAlex record.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
Besides the needed data,there are lots of navigation information and advertisements in the Web pages.A DOM tree comparison algorithm was proposed.It compared several pages within a class,and recognized the main contents in pages.Experiment results show that it is feasible and effective.
Key concepts: Computer science, Web page, Document Object Model, Information retrieval, Class (philosophy), Tree (set theory), Data mining, World Wide Web