FPST: a new term weighting algorithm for long running and short lived events
Y. Jahnavi, Y. Radhika
Abstract
Y. Jahnavi, Y. Radhika
Abstract
Term weighting is a useful technique that extracts important features from textual documents, thereby providing a basis for different text mining approaches. While several term weighting algorithms based on their frequency and some other statistical measures have been proposed in the past, they are inaccurate in extracting hot terms from internet-based digitised news documents. To overcome that problem, this paper presents an innovative and effective term weighting algorithm by considering position, scattering and topicality along with frequency. Frequency considers the number of occurrences of a term; position focuses on where the term appears; scattering focuses on the distribution of a term in the entire document. Here topicality is calculated for both short lived events and long running events. Experimental evaluation shows that the proposed term weighting algorithm outperforms the existing term weighting algorithms.
OpenAlex reports 11 citations for this work. Citation counts describe recorded attention and do not establish research quality.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
Term weighting is a useful technique that extracts important features from textual documents, thereby providing a basis for different text mining approaches. While several term weighting algorithms based on their frequency and some other statistical measures have been proposed in the past, they are inaccurate in extracting hot terms from internet-based digitised news documents. To overcome that problem, this paper presents an innovative and effective term weighting algorithm by considering position, scattering and topicality along with frequency. Frequency considers the number of occurrences of a term; position focuses on where the term appears; scattering focuses on the distribution of a term in the entire document. Here topicality is calculated for both short lived events and long running events. Experimental evaluation shows that the proposed term weighting algorithm outperforms the existing term weighting algorithms.
Key concepts: Weighting, Term (time), Position (finance), Computer science, Algorithm, The Internet, Data mining, A-weighting