2010Unpublished venueRequires access

A Term Cluster Query Expansion Model Based on Classification Information in Natural Language Information Retrieval

Jeon Wook Kang, Hyun-Kyu Kang, Myeong-Cheol Ko, Heung Seok Jeon, Junghyun Nam

Open publisher page 10 citations

Abstract

A natural language information retrieval system ranks related documents according to criteria based on user query keywords and document similarities. However, many efforts have been made to make more useful query keywords because users do not use many keywords in their natural language search query when retrieving information on the Web. Because a keyword does not provide much information, however, relevance feedback is generally used to complement the weakness of general retrieval methods. This paper proposes a term cluster query expansion model based on classification information of retrieved documents. This model generates classification information from the upper ranked n documents retrieved by retrieval system. On the basis of the extracted classification information, the term cluster (m) that represents each group is generated, and then the model allows user to select term cluster that corresponds to user information needs. The query keywords are expanded by using a relevance feedback algorithm based on the selected classification information. As a result of the experiments with test collection, the retrieval effectiveness was improved by 13.2% compared to the initial query when the Rocchio method was used.

About this research paper

What this paper is about

A natural language information retrieval system ranks related documents according to criteria based on user query keywords and document similarities. However, many efforts have been made to make more useful query keywords because users do not use many keywords in their natural language search query when retrieving information on the Web. Because a keyword does not provide much information, however, relevance feedback is generally used to complement the weakness of general retrieval methods. This paper proposes a term cluster query expansion model based on classification information of retrieved documents. This model generates classification information from the upper ranked n documents retrieved by retrieval system. On the basis of the extracted classification information, the term cluster (m) that represents each group is generated, and then the model allows user to select term cluster that corresponds to user information needs. The query keywords are expanded by using a relevance feedback algorithm based on the selected classification information. As a result of the experiments with test collection, the retrieval effectiveness was improved by 13.2% compared to the initial query when the Rocchio method was used.

Why it matters

OpenAlex reports 10 citations for this work. Citation counts describe recorded attention and do not establish research quality.

Key contribution

A contribution statement is not available in the OpenAlex record.

Method / approach

Method details are not available in the OpenAlex metadata.

Main findings

Findings are not separately available in the OpenAlex metadata.

Limitations

Limitations are not available in the OpenAlex metadata.

Applications

Application details are not available in the OpenAlex metadata.

Available abstract

A natural language information retrieval system ranks related documents according to criteria based on user query keywords and document similarities. However, many efforts have been made to make more useful query keywords because users do not use many keywords in their natural language search query when retrieving information on the Web. Because a keyword does not provide much information, however, relevance feedback is generally used to complement the weakness of general retrieval methods. This paper proposes a term cluster query expansion model based on classification information of retrieved documents. This model generates classification information from the upper ranked n documents retrieved by retrieval system. On the basis of the extracted classification information, the term cluster (m) that represents each group is generated, and then the model allows user to select term cluster that corresponds to user information needs. The query keywords are expanded by using a relevance feedback algorithm based on the selected classification information. As a result of the experiments with test collection, the retrieval effectiveness was improved by 13.2% compared to the initial query when the Rocchio method was used.

Key concepts: Computer science, Query expansion, Information retrieval, Web query classification, Relevance (law), Query language, Term (time), Relevance feedback

Related papers

Back to paper searchBrowse research topicsOriginal source
A Term Cluster Query Expansion Model Based on Classification Information in Natural Language Information Retrieval — Research Paper | ScholarLens