Relating the New Language Models of Information Retrieval to the Traditional Retrieval Models
Djoerd Hiemstra, Arjen P. de Vries
Abstract
Open-access reader
Djoerd Hiemstra, Arjen P. de Vries
Abstract
Open-access reader
During the last two years, exciting new approaches to information retrieval were introduced by a number of different research groups that use statistical language models for retrieval.This paper relates the retrieval algorithms suggested by these approaches to widely accepted retrieval algorithms developed within three traditional models of information retrieval: the Boolean model, the vector space model and the probabilistic model.The paper shows the existence of efficient retrieval algorithms that only use the matching terms in their computation.Under these conditions, the language models of information retrieval are surprisingly similar to both tf.idf term weighting as developed for the vector space model and relevance weighting as developed in the traditional probabilistic model.The paper suggests a new method for relevance weighting and a new method to rank documents giving Boolean queries.Experimental results on the TREC collection indicate that the language modelling approach outperforms the three traditional approaches.
OpenAlex reports 58 citations for this work. Citation counts describe recorded attention and do not establish research quality.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
During the last two years, exciting new approaches to information retrieval were introduced by a number of different research groups that use statistical language models for retrieval.This paper relates the retrieval algorithms suggested by these approaches to widely accepted retrieval algorithms developed within three traditional models of information retrieval: the Boolean model, the vector space model and the probabilistic model.The paper shows the existence of efficient retrieval algorithms that only use the matching terms in their computation.Under these conditions, the language models of information retrieval are surprisingly similar to both tf.idf term weighting as developed for the vector space model and relevance weighting as developed in the traditional probabilistic model.The paper suggests a new method for relevance weighting and a new method to rank documents giving Boolean queries.Experimental results on the TREC collection indicate that the language modelling approach outperforms the three traditional approaches.
Key concepts: Divergence-from-randomness model, Vector space model, Computer science, Standard Boolean model, Relevance (law), Weighting, Information retrieval, Language model