High-dimensional OLAP: a minimal cubing approach
Xiaolei Li, Jiawei Han, Héctor González
Abstract
Xiaolei Li, Jiawei Han, Héctor González
Abstract
Data cube has been playing an essential role in fast OLAP (online analytical processing) in many multi-dimensional data warehouses. However, there exist data sets in applications like bioinformatics, statistics, and text pro-cessing that are characterized by high dimen-sionality, e.g., over 100 dimensions, and mod-erate size, e.g., around 106 tuples. No feasible data cube can be constructed with such data sets. In this paper we will address the problem of developing an e±cient algorithm to perform OLAP on such data sets. Experience tells us that although data analy-sis tasks may involve a high dimensional space, most OLAP operations are performed only on a small number of dimensions at a time. Based on this observation, we propose a novel method that computes a thin layer of the data cube together with associated value-list indices. This layer, while being manageable in size, will be capable of supporting °exi-ble and fast OLAP operations in the original high dimensional space. Through experiments we will show that the method has I/O costs that scale nicely with dimensionality. Further-more, the costs are comparable to that of ac-cessing an existing data cube when full mate-rialization is possible. 1
OpenAlex reports 110 citations for this work. Citation counts describe recorded attention and do not establish research quality.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
Data cube has been playing an essential role in fast OLAP (online analytical processing) in many multi-dimensional data warehouses. However, there exist data sets in applications like bioinformatics, statistics, and text pro-cessing that are characterized by high dimen-sionality, e.g., over 100 dimensions, and mod-erate size, e.g., around 106 tuples. No feasible data cube can be constructed with such data sets. In this paper we will address the problem of developing an e±cient algorithm to perform OLAP on such data sets. Experience tells us that although data analy-sis tasks may involve a high dimensional space, most OLAP operations are performed only on a small number of dimensions at a time. Based on this observation, we propose a novel method that computes a thin layer of the data cube together with associated value-list indices. This layer, while being manageable in size, will be capable of supporting °exi-ble and fast OLAP operations in the original high dimensional space. Through experiments we will show that the method has I/O costs that scale nicely with dimensionality. Further-more, the costs are comparable to that of ac-cessing an existing data cube when full mate-rialization is possible. 1
Key concepts: Online analytical processing, Data cube, Computer science, Cube (algebra), Tuple, Curse of dimensionality, Data warehouse, Data mining