Web data mining1
Stefan Bosse, Lena Dahlhaus, Uwe Engel
Abstract
Stefan Bosse, Lena Dahlhaus, Uwe Engel
Abstract
Digital spaces change the way social interaction takes place. Communication now often takes place on social media, blogs, and websites. Access to textual data that reflect how people communicate, and about what subjects they communicate, is then gained only by computational means. The chapter gives a summary of how to gain this access using web data mining. In doing so, it outlines how modern dynamic web pages are built to facilitate understanding of why and how this method of data mining works. An example web page is used for tutorial reasons, using the R package and further software as tools for web scraping. The chapter is targeted at transferring orientation knowledge and know-how regarding how written content is automatically extractable from a web resource. This is accompanied by a discussion of related issues. In this regard, the chapter discusses the legal status of web scraping, threats to the validity of related inferences, and the challenge of transforming unstructured “found” data to a structure that enables text analytics.
A significance statement is not available in the OpenAlex record.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
Digital spaces change the way social interaction takes place. Communication now often takes place on social media, blogs, and websites. Access to textual data that reflect how people communicate, and about what subjects they communicate, is then gained only by computational means. The chapter gives a summary of how to gain this access using web data mining. In doing so, it outlines how modern dynamic web pages are built to facilitate understanding of why and how this method of data mining works. An example web page is used for tutorial reasons, using the R package and further software as tools for web scraping. The chapter is targeted at transferring orientation knowledge and know-how regarding how written content is automatically extractable from a web resource. This is accompanied by a discussion of related issues. In this regard, the chapter discusses the legal status of web scraping, threats to the validity of related inferences, and the challenge of transforming unstructured “found” data to a structure that enables text analytics.
Key concepts: Computer science, Web application, World Wide Web