Information Extraction from Web Pages Using Hyperlinks
Suvarna Sharma, Amit Bhagat
Abstract
Suvarna Sharma, Amit Bhagat
Abstract
Web mining is a well-studied topic nowadays. Web mining is a data processing tool that is useful for gathering information from the internet. With the expansion of the Web, people now have access to vast amounts of data. Online structure mining deals with the analysis of web linkages and content. An essential component of web structure mining is the web crawler. A web crawler is an automated software that searches the internet to locate new web pages and to update the content of sites that have already been downloaded. With the aim of helping web crawlers extract fresh connections and data from the sites, we provide the fundamental method for web crawling utilizing a set of seed Urls in this work. We take into account a scenario that concentrates on both the difficulties of redundant Urls and metadata gathering. This study makes use of an algorithm that downloads and browses websites at random. The necessity to rapidly download pages, store them efficiently in a database, and traverse pages is just one of the issues with the design of a web crawler. The architecture of Web Crawler, several of its capabilities, and a basic crawling technique are all described in this article.
OpenAlex reports 6 citations for this work. Citation counts describe recorded attention and do not establish research quality.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
Web mining is a well-studied topic nowadays. Web mining is a data processing tool that is useful for gathering information from the internet. With the expansion of the Web, people now have access to vast amounts of data. Online structure mining deals with the analysis of web linkages and content. An essential component of web structure mining is the web crawler. A web crawler is an automated software that searches the internet to locate new web pages and to update the content of sites that have already been downloaded. With the aim of helping web crawlers extract fresh connections and data from the sites, we provide the fundamental method for web crawling utilizing a set of seed Urls in this work. We take into account a scenario that concentrates on both the difficulties of redundant Urls and metadata gathering. This study makes use of an algorithm that downloads and browses websites at random. The necessity to rapidly download pages, store them efficiently in a database, and traverse pages is just one of the issues with the design of a web crawler. The architecture of Web Crawler, several of its capabilities, and a basic crawling technique are all described in this article.
Key concepts: Web crawler, World Wide Web, Computer science, Web page, Static web page, Focused crawler, Web modeling, Web mining