2000Unpublished venueRequires access

Study on new term weighting method and new vector space model based on word space in spoken document retrieval

S. Takao, Jun Ogata, Yasuo Ariki

Open publisher page 5 citations

Abstract

Recently, TV news programs are broadcast from all over the world owing to the broadcast digitization. In this situation, TV viewers want to select and watch the most interesting news. In order to satisfy this requirement, news database has to be constructed which has automatic topic segmentation and retrieval function. In this paper, focusing on topic retrieval function, we propose a new spoken document retrieval method. There are two types of problems in spoken document retrieval. The first problem is how to eliminate the occurrence of error words caused by speech recognition. The second problem is how to extract the important words from spoken documents. In order to solve these problems, term weighting methods play an important role, because it can decrease the dimension of vector space model and usually increase the retrieval performance. We propose, in this paper, a new term weighting method "mutual information incorporating TF-IDF". We compared it with conventional TF-IDF, mutual...

About this research paper

What this paper is about

Recently, TV news programs are broadcast from all over the world owing to the broadcast digitization. In this situation, TV viewers want to select and watch the most interesting news. In order to satisfy this requirement, news database has to be constructed which has automatic topic segmentation and retrieval function. In this paper, focusing on topic retrieval function, we propose a new spoken document retrieval method. There are two types of problems in spoken document retrieval. The first problem is how to eliminate the occurrence of error words caused by speech recognition. The second problem is how to extract the important words from spoken documents. In order to solve these problems, term weighting methods play an important role, because it can decrease the dimension of vector space model and usually increase the retrieval performance. We propose, in this paper, a new term weighting method "mutual information incorporating TF-IDF". We compared it with conventional TF-IDF, mutual...

Why it matters

OpenAlex reports 5 citations for this work. Citation counts describe recorded attention and do not establish research quality.

Key contribution

A contribution statement is not available in the OpenAlex record.

Method / approach

Method details are not available in the OpenAlex metadata.

Main findings

Findings are not separately available in the OpenAlex metadata.

Limitations

Limitations are not available in the OpenAlex metadata.

Applications

Application details are not available in the OpenAlex metadata.

Available abstract

Recently, TV news programs are broadcast from all over the world owing to the broadcast digitization. In this situation, TV viewers want to select and watch the most interesting news. In order to satisfy this requirement, news database has to be constructed which has automatic topic segmentation and retrieval function. In this paper, focusing on topic retrieval function, we propose a new spoken document retrieval method. There are two types of problems in spoken document retrieval. The first problem is how to eliminate the occurrence of error words caused by speech recognition. The second problem is how to extract the important words from spoken documents. In order to solve these problems, term weighting methods play an important role, because it can decrease the dimension of vector space model and usually increase the retrieval performance. We propose, in this paper, a new term weighting method "mutual information incorporating TF-IDF". We compared it with conventional TF-IDF, mutual...

Key concepts: Vector space model, Computer science, Document retrieval, tf–idf, Information retrieval, Term Discrimination, Similarity (geometry), Weighting

Related papers

Back to paper searchBrowse research topicsOriginal source
Study on new term weighting method and new vector space model based on word space in spoken document retrieval — Research Paper | ScholarLens