Theoretical analysis on pruning nearest neighbor candidates by locality sensitive hashing
T Mutohy, Masakazu Iwamura, K. Kise
Abstract
T Mutohy, Masakazu Iwamura, K. Kise
Abstract
Locality Sensitive Hashing (LSH) is one of the most popular methods of the approximate near neighbor search. In applications that require the nearest neighbors of queries in a short time, LSH is sometimes used in pruning of the candidates of nearest neighbors. While the pruning reduces the processing time greatly, it also reduces the chances of retrieving the exact nearest neighbors. However, the pruning of nearest neighbor candidates using LSH has not been considered theoretically. Thus in this paper, we investigate the pruning effect by deriving the formulae of retrieval accuracy and computational cost of distance calculation for uniformly distributed data. Furthermore, we make evaluations on the formulae by comparison between simulation results and the theoretical values.
A significance statement is not available in the OpenAlex record.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
Locality Sensitive Hashing (LSH) is one of the most popular methods of the approximate near neighbor search. In applications that require the nearest neighbors of queries in a short time, LSH is sometimes used in pruning of the candidates of nearest neighbors. While the pruning reduces the processing time greatly, it also reduces the chances of retrieving the exact nearest neighbors. However, the pruning of nearest neighbor candidates using LSH has not been considered theoretically. Thus in this paper, we investigate the pruning effect by deriving the formulae of retrieval accuracy and computational cost of distance calculation for uniformly distributed data. Furthermore, we make evaluations on the formulae by comparison between simulation results and the theoretical values.
Key concepts: Locality-sensitive hashing, Pruning, k-nearest neighbors algorithm, Nearest neighbor search, Computer science, Hash function, Locality, Pattern recognition (psychology)