2022•2022 10th International Conference on Reliability, Infocom Technologies and Optimization (Trends and Future Directions) (ICRITO)Requires access

Information Extraction from Web Pages Using Hyperlinks

Suvarna Sharma, Amit Bhagat

Open publisher page 6 citations

Abstract

Web mining is a well-studied topic nowadays. Web mining is a data processing tool that is useful for gathering information from the internet. With the expansion of the Web, people now have access to vast amounts of data. Online structure mining deals with the analysis of web linkages and content. An essential component of web structure mining is the web crawler. A web crawler is an automated software that searches the internet to locate new web pages and to update the content of sites that have already been downloaded. With the aim of helping web crawlers extract fresh connections and data from the sites, we provide the fundamental method for web crawling utilizing a set of seed Urls in this work. We take into account a scenario that concentrates on both the difficulties of redundant Urls and metadata gathering. This study makes use of an algorithm that downloads and browses websites at random. The necessity to rapidly download pages, store them efficiently in a database, and traverse pages is just one of the issues with the design of a web crawler. The architecture of Web Crawler, several of its capabilities, and a basic crawling technique are all described in this article.

About this research paper

What this paper is about

Web mining is a well-studied topic nowadays. Web mining is a data processing tool that is useful for gathering information from the internet. With the expansion of the Web, people now have access to vast amounts of data. Online structure mining deals with the analysis of web linkages and content. An essential component of web structure mining is the web crawler. A web crawler is an automated software that searches the internet to locate new web pages and to update the content of sites that have already been downloaded. With the aim of helping web crawlers extract fresh connections and data from the sites, we provide the fundamental method for web crawling utilizing a set of seed Urls in this work. We take into account a scenario that concentrates on both the difficulties of redundant Urls and metadata gathering. This study makes use of an algorithm that downloads and browses websites at random. The necessity to rapidly download pages, store them efficiently in a database, and traverse pages is just one of the issues with the design of a web crawler. The architecture of Web Crawler, several of its capabilities, and a basic crawling technique are all described in this article.

Why it matters

OpenAlex reports 6 citations for this work. Citation counts describe recorded attention and do not establish research quality.

Key contribution

A contribution statement is not available in the OpenAlex record.

Method / approach

Method details are not available in the OpenAlex metadata.

Main findings

Findings are not separately available in the OpenAlex metadata.

Limitations

Limitations are not available in the OpenAlex metadata.

Applications

Application details are not available in the OpenAlex metadata.

Available abstract

Web mining is a well-studied topic nowadays. Web mining is a data processing tool that is useful for gathering information from the internet. With the expansion of the Web, people now have access to vast amounts of data. Online structure mining deals with the analysis of web linkages and content. An essential component of web structure mining is the web crawler. A web crawler is an automated software that searches the internet to locate new web pages and to update the content of sites that have already been downloaded. With the aim of helping web crawlers extract fresh connections and data from the sites, we provide the fundamental method for web crawling utilizing a set of seed Urls in this work. We take into account a scenario that concentrates on both the difficulties of redundant Urls and metadata gathering. This study makes use of an algorithm that downloads and browses websites at random. The necessity to rapidly download pages, store them efficiently in a database, and traverse pages is just one of the issues with the design of a web crawler. The architecture of Web Crawler, several of its capabilities, and a basic crawling technique are all described in this article.

Key concepts: Web crawler, World Wide Web, Computer science, Web page, Static web page, Focused crawler, Web modeling, Web mining

Related papers

Back to paper searchBrowse research topicsOriginal source
Information Extraction from Web Pages Using Hyperlinks — Research Paper | ScholarLens