Efficient Big-Data Access: Taxonomy and a Comprehensive Survey
Anis Alazzawe, Amitangshu Pal, Krishna Kant
Abstract
Anis Alazzawe, Amitangshu Pal, Krishna Kant
Abstract
The emerging systems are not only generating huge amounts of data but also expect this data to be analyzed expeditiously to drive online decision-making and control. Thus, identifying the most relevant data and making it available close to the computation becomes a central challenge in driving the big data revolution. Storage systems play a crucial role in enabling efficient access to the stored data and intelligent storage management techniques are thus central to addressing the problem. Generally, as the data volume increases, the marginal utility of an “average” data item tends to decline, which requires greater effort in identifying the most valuable data items and making them available with minimal overhead and latency. Data driven mechanisms have a big role to play in solving this needle-in-the-haystack problem. In this paper we propose a taxonomy to provide a structure for understanding the common issues surrounding these techniques. We discuss these techniques and articulate many research challenges and opportunities.
OpenAlex reports 13 citations for this work. Citation counts describe recorded attention and do not establish research quality.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
The emerging systems are not only generating huge amounts of data but also expect this data to be analyzed expeditiously to drive online decision-making and control. Thus, identifying the most relevant data and making it available close to the computation becomes a central challenge in driving the big data revolution. Storage systems play a crucial role in enabling efficient access to the stored data and intelligent storage management techniques are thus central to addressing the problem. Generally, as the data volume increases, the marginal utility of an “average” data item tends to decline, which requires greater effort in identifying the most valuable data items and making them available with minimal overhead and latency. Data driven mechanisms have a big role to play in solving this needle-in-the-haystack problem. In this paper we propose a taxonomy to provide a structure for understanding the common issues surrounding these techniques. We discuss these techniques and articulate many research challenges and opportunities.
Key concepts: Computer science, Haystack, Big data, Data science, Data management, Data access, Taxonomy (biology), Data mining