Extracting Template Properties using Agglomerative Clustering
R. Devika, T. Mohanraj
Abstract
R. Devika, T. Mohanraj
Abstract
World Wide Web is widely used to publish and access information on the Internet. Most of the web pages in the web sites are published using the common templates with contents. Templates are the readymade holders, which provide readers easy access to the contents guided by consistent structures. It provides common look and feel to the web pages. However, the accuracy and performance of the web applications are degraded due to the presence of irrelevant terms in the templates. Thus, template detection techniques have received a lot of attention recently to improve the performance of search engines, clustering, and classification of web documents. Hence, the proposed system presents a new clustering algorithm for grouping the web pages that are using similar templates. The web pages under same cluster have equal priority and they are homogeneous. Hence, all those pages will not be displayed. In order to prioritize any homogeneous web page, the properties of that particular website will be extracted and modified. By changing the properties, homogeneous web page can be converted to a heterogeneous web page.
A significance statement is not available in the OpenAlex record.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
World Wide Web is widely used to publish and access information on the Internet. Most of the web pages in the web sites are published using the common templates with contents. Templates are the readymade holders, which provide readers easy access to the contents guided by consistent structures. It provides common look and feel to the web pages. However, the accuracy and performance of the web applications are degraded due to the presence of irrelevant terms in the templates. Thus, template detection techniques have received a lot of attention recently to improve the performance of search engines, clustering, and classification of web documents. Hence, the proposed system presents a new clustering algorithm for grouping the web pages that are using similar templates. The web pages under same cluster have equal priority and they are homogeneous. Hence, all those pages will not be displayed. In order to prioritize any homogeneous web page, the properties of that particular website will be extracted and modified. By changing the properties, homogeneous web page can be converted to a heterogeneous web page.
Key concepts: Computer science, Web page, Template, World Wide Web, Static web page, Information retrieval, Cluster analysis, Web modeling