Study on new term weighting method and new vector space model based on word space in spoken document retrieval
S. Takao, Jun Ogata, Yasuo Ariki
Abstract
S. Takao, Jun Ogata, Yasuo Ariki
Abstract
Recently, TV news programs are broadcast from all over the world owing to the broadcast digitization. In this situation, TV viewers want to select and watch the most interesting news. In order to satisfy this requirement, news database has to be constructed which has automatic topic segmentation and retrieval function. In this paper, focusing on topic retrieval function, we propose a new spoken document retrieval method. There are two types of problems in spoken document retrieval. The first problem is how to eliminate the occurrence of error words caused by speech recognition. The second problem is how to extract the important words from spoken documents. In order to solve these problems, term weighting methods play an important role, because it can decrease the dimension of vector space model and usually increase the retrieval performance. We propose, in this paper, a new term weighting method "mutual information incorporating TF-IDF". We compared it with conventional TF-IDF, mutual...
OpenAlex reports 5 citations for this work. Citation counts describe recorded attention and do not establish research quality.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
Recently, TV news programs are broadcast from all over the world owing to the broadcast digitization. In this situation, TV viewers want to select and watch the most interesting news. In order to satisfy this requirement, news database has to be constructed which has automatic topic segmentation and retrieval function. In this paper, focusing on topic retrieval function, we propose a new spoken document retrieval method. There are two types of problems in spoken document retrieval. The first problem is how to eliminate the occurrence of error words caused by speech recognition. The second problem is how to extract the important words from spoken documents. In order to solve these problems, term weighting methods play an important role, because it can decrease the dimension of vector space model and usually increase the retrieval performance. We propose, in this paper, a new term weighting method "mutual information incorporating TF-IDF". We compared it with conventional TF-IDF, mutual...
Key concepts: Vector space model, Computer science, Document retrieval, tf–idf, Information retrieval, Term Discrimination, Similarity (geometry), Weighting