To link or not to link: Ranking hyperlinks in Wikipedia using collective attention
Philip Thruesen, Jaroslav Čechák, Blandine Seznec, Roel Castaño, Nattiya Kanhabua
Abstract
Philip Thruesen, Jaroslav Čechák, Blandine Seznec, Roel Castaño, Nattiya Kanhabua
Abstract
Wikipedia is one of the fastest growing websites and a primary source of knowledge on the Internet. Being a wiki, its content is crowd-sourced by the users. This has many benefits and it is one of the main reasons it has grown to reach more than 5 million articles in its English version. Nevertheless, this also raises issues, like the overlinking of articles, which are difficult to deal with by editors. In this paper, we tackle overlinking in Wikipedia as a ranking problem. We apply Learning to Rank algorithms to evaluate the click frequency of links in an effort to distinguish the most useful links for users. To accomplish this, we develop a ground truth, which serves as baseline for our algorithm and compare hyperlink features to implement the most advantageous ones. The results show 86.2% accuracy with the top-6 most useful features and 87.7% accuracy with the complete feature set. Considering these results, we outline a solution to the overlinking problem. By removing the most inadequate links, we suggest that readability of Wikipedia articles could be improved while preserving most of its useful links.
OpenAlex reports 1 citations for this work. Citation counts describe recorded attention and do not establish research quality.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
Wikipedia is one of the fastest growing websites and a primary source of knowledge on the Internet. Being a wiki, its content is crowd-sourced by the users. This has many benefits and it is one of the main reasons it has grown to reach more than 5 million articles in its English version. Nevertheless, this also raises issues, like the overlinking of articles, which are difficult to deal with by editors. In this paper, we tackle overlinking in Wikipedia as a ranking problem. We apply Learning to Rank algorithms to evaluate the click frequency of links in an effort to distinguish the most useful links for users. To accomplish this, we develop a ground truth, which serves as baseline for our algorithm and compare hyperlink features to implement the most advantageous ones. The results show 86.2% accuracy with the top-6 most useful features and 87.7% accuracy with the complete feature set. Considering these results, we outline a solution to the overlinking problem. By removing the most inadequate links, we suggest that readability of Wikipedia articles could be improved while preserving most of its useful links.
Key concepts: Hyperlink, Computer science, Readability, Link analysis, Ranking (information retrieval), Link (geometry), Baseline (sea), Information retrieval