Relational Query Techniques for Distributed Data Stream:A Survey
Wang Chun
Abstract
Wang Chun
Abstract
The applications that require online processing continuous data stream are increasing.Data stream management systems which are used to deal with massive and variable data in real time have been produced.With the development of open processing platforms in the ear of big data,a number of distributed data stream processing systems have emerged for dealing with large scale and diverse data stream,such as S4,Storm,Spark Streaming,etc.However,we should construct relational query systems which have abstract query language on basis of the processing systems for improving the ease of use and processing capability of them,so as to build complete distributed data stream management systems.How to design and realize the high efficiency and easy-to-use query systems is a great challenge.In this survey,we first provide an overview of typical applications,data characteristics and achieve goals of distributed data stream query processing.Furthermore,we propose the framework of distributed data stream relational query systems.Based on the framework,we analyze the key techniques in several aspects:UDF query,query optimization,query-driven approaches,compiling techniques,operator management,scheduling management and parallel management.Then,there is the comparison of representative query systems including SPL,StreamingSQL,Squall and DBToaster.Finally,some new challenges are put forward,including optimization technique,execution strategy,real-time precise query and complex query analysis.
A significance statement is not available in the OpenAlex record.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
The applications that require online processing continuous data stream are increasing.Data stream management systems which are used to deal with massive and variable data in real time have been produced.With the development of open processing platforms in the ear of big data,a number of distributed data stream processing systems have emerged for dealing with large scale and diverse data stream,such as S4,Storm,Spark Streaming,etc.However,we should construct relational query systems which have abstract query language on basis of the processing systems for improving the ease of use and processing capability of them,so as to build complete distributed data stream management systems.How to design and realize the high efficiency and easy-to-use query systems is a great challenge.In this survey,we first provide an overview of typical applications,data characteristics and achieve goals of distributed data stream query processing.Furthermore,we propose the framework of distributed data stream relational query systems.Based on the framework,we analyze the key techniques in several aspects:UDF query,query optimization,query-driven approaches,compiling techniques,operator management,scheduling management and parallel management.Then,there is the comparison of representative query systems including SPL,StreamingSQL,Squall and DBToaster.Finally,some new challenges are put forward,including optimization technique,execution strategy,real-time precise query and complex query analysis.
Key concepts: Computer science, Query optimization, Sargable, Query language, Query expansion, Stream processing, Online aggregation, Web query classification