What Is the Nearest Neighbor in High Dimensional Spaces
Alexander Hinneburg, Charų C. Aggarwal, Daniel A. Keim
Abstract
Open-access reader
Alexander Hinneburg, Charų C. Aggarwal, Daniel A. Keim
Abstract
Open-access reader
Nearest neighbor search in high dimensional spaces is an interesting and important problem which is relevant for a wide varietyofnovel database applications. As recentresults show, however, the problem is a very difficult one, not only with regards to the performance issue but also to the quality issue. In this paper, we discuss the quality issue and identify a new generalized notion of nearest neighbor search as the relevant problem in high dimensional space. In contrast to previous approaches, our new notion of nearest neighbor search does not treat all dimensions equally but uses a quality criterion to select relevant dimensions (projections) with respect to the given query. As an example for a useful quality criterion, we rate howwell the data is clustered around the query point within the selected projection. We then propose an efficient and effectivealgorithm to solve the generalized nearest neighbor problem. Our experiments based on a number of real and synthetic data sets sho...
OpenAlex reports 498 citations for this work. Citation counts describe recorded attention and do not establish research quality.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
Nearest neighbor search in high dimensional spaces is an interesting and important problem which is relevant for a wide varietyofnovel database applications. As recentresults show, however, the problem is a very difficult one, not only with regards to the performance issue but also to the quality issue. In this paper, we discuss the quality issue and identify a new generalized notion of nearest neighbor search as the relevant problem in high dimensional space. In contrast to previous approaches, our new notion of nearest neighbor search does not treat all dimensions equally but uses a quality criterion to select relevant dimensions (projections) with respect to the given query. As an example for a useful quality criterion, we rate howwell the data is clustered around the query point within the selected projection. We then propose an efficient and effectivealgorithm to solve the generalized nearest neighbor problem. Our experiments based on a number of real and synthetic data sets sho...
Key concepts: Nearest neighbor search, Best bin first, Fixed-radius near neighbors, k-nearest neighbors algorithm, Nearest neighbor graph, Cover tree, Large margin nearest neighbor, Computer science