Research on Term Weighting Based on MapReduce
Trs Information
Abstract
Trs Information
Abstract
Term recognition is widely used in the ontology construction,dictionary construction and other fields.And term weighting is a key step in the term recognition.In this paper,several improvements have been made to TF-IDF algorithm,e.g.,the length of terms is considered in weighting,also with terms' correlations to documentation set.The candidate term weight is calculated in a distributed manner based on MapReduce on Hadoop.Experimental results show that the method proposed not only simplifies the steps of term weighting,but also improves the efficiency of the algorithm.
A significance statement is not available in the OpenAlex record.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
Term recognition is widely used in the ontology construction,dictionary construction and other fields.And term weighting is a key step in the term recognition.In this paper,several improvements have been made to TF-IDF algorithm,e.g.,the length of terms is considered in weighting,also with terms' correlations to documentation set.The candidate term weight is calculated in a distributed manner based on MapReduce on Hadoop.Experimental results show that the method proposed not only simplifies the steps of term weighting,but also improves the efficiency of the algorithm.
Key concepts: Weighting, Term (time), Computer science, Key (lock), Set (abstract data type), Data mining, Documentation, Ontology