Demystifying MapReduce
Christopher J. Garcia
Abstract
Open-access reader
Christopher J. Garcia
Abstract
Open-access reader
Recent innovations in Big Data have enabled major strides forward in our ability to glean important insights from massive amounts of data, and to use these insights to make better decisions. Underlying many of these innovations is a computational paradigm known as MapReduce, which enables computational processes to be scaled up to very large sizes and to take advantage of cloud computing. While very powerful, MapReduce also requires a nontrivial shift in algorithm design strategies. In this paper we provide an overview of MapReduce and types of problems it is suited for. We discuss general strategies for designing MapReduce-based algorithms and provide an illustration using social media analytics.
OpenAlex reports 5 citations for this work. Citation counts describe recorded attention and do not establish research quality.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
Recent innovations in Big Data have enabled major strides forward in our ability to glean important insights from massive amounts of data, and to use these insights to make better decisions. Underlying many of these innovations is a computational paradigm known as MapReduce, which enables computational processes to be scaled up to very large sizes and to take advantage of cloud computing. While very powerful, MapReduce also requires a nontrivial shift in algorithm design strategies. In this paper we provide an overview of MapReduce and types of problems it is suited for. We discuss general strategies for designing MapReduce-based algorithms and provide an illustration using social media analytics.
Key concepts: Computer science, Big data, Cloud computing, Data science, Analytics, Distributed computing, Data mining, Operating system