2004Unpublished venueRequires access

High-dimensional OLAP: a minimal cubing approach

Xiaolei Li, Jiawei Han, Héctor González

Open publisher page 110 citations

Abstract

Data cube has been playing an essential role in fast OLAP (online analytical processing) in many multi-dimensional data warehouses. However, there exist data sets in applications like bioinformatics, statistics, and text pro-cessing that are characterized by high dimen-sionality, e.g., over 100 dimensions, and mod-erate size, e.g., around 106 tuples. No feasible data cube can be constructed with such data sets. In this paper we will address the problem of developing an e±cient algorithm to perform OLAP on such data sets. Experience tells us that although data analy-sis tasks may involve a high dimensional space, most OLAP operations are performed only on a small number of dimensions at a time. Based on this observation, we propose a novel method that computes a thin layer of the data cube together with associated value-list indices. This layer, while being manageable in size, will be capable of supporting °exi-ble and fast OLAP operations in the original high dimensional space. Through experiments we will show that the method has I/O costs that scale nicely with dimensionality. Further-more, the costs are comparable to that of ac-cessing an existing data cube when full mate-rialization is possible. 1

About this research paper

What this paper is about

Data cube has been playing an essential role in fast OLAP (online analytical processing) in many multi-dimensional data warehouses. However, there exist data sets in applications like bioinformatics, statistics, and text pro-cessing that are characterized by high dimen-sionality, e.g., over 100 dimensions, and mod-erate size, e.g., around 106 tuples. No feasible data cube can be constructed with such data sets. In this paper we will address the problem of developing an e±cient algorithm to perform OLAP on such data sets. Experience tells us that although data analy-sis tasks may involve a high dimensional space, most OLAP operations are performed only on a small number of dimensions at a time. Based on this observation, we propose a novel method that computes a thin layer of the data cube together with associated value-list indices. This layer, while being manageable in size, will be capable of supporting °exi-ble and fast OLAP operations in the original high dimensional space. Through experiments we will show that the method has I/O costs that scale nicely with dimensionality. Further-more, the costs are comparable to that of ac-cessing an existing data cube when full mate-rialization is possible. 1

Why it matters

OpenAlex reports 110 citations for this work. Citation counts describe recorded attention and do not establish research quality.

Key contribution

A contribution statement is not available in the OpenAlex record.

Method / approach

Method details are not available in the OpenAlex metadata.

Main findings

Findings are not separately available in the OpenAlex metadata.

Limitations

Limitations are not available in the OpenAlex metadata.

Applications

Application details are not available in the OpenAlex metadata.

Available abstract

Data cube has been playing an essential role in fast OLAP (online analytical processing) in many multi-dimensional data warehouses. However, there exist data sets in applications like bioinformatics, statistics, and text pro-cessing that are characterized by high dimen-sionality, e.g., over 100 dimensions, and mod-erate size, e.g., around 106 tuples. No feasible data cube can be constructed with such data sets. In this paper we will address the problem of developing an e±cient algorithm to perform OLAP on such data sets. Experience tells us that although data analy-sis tasks may involve a high dimensional space, most OLAP operations are performed only on a small number of dimensions at a time. Based on this observation, we propose a novel method that computes a thin layer of the data cube together with associated value-list indices. This layer, while being manageable in size, will be capable of supporting °exi-ble and fast OLAP operations in the original high dimensional space. Through experiments we will show that the method has I/O costs that scale nicely with dimensionality. Further-more, the costs are comparable to that of ac-cessing an existing data cube when full mate-rialization is possible. 1

Key concepts: Online analytical processing, Data cube, Computer science, Cube (algebra), Tuple, Curse of dimensionality, Data warehouse, Data mining

Related papers

Back to paper searchBrowse research topicsOriginal source
High-dimensional OLAP: a minimal cubing approach — Research Paper | ScholarLens