2016•Unpublished venueRequires access

Fast Processing SPARQL Queries on Large RDF Data

Guang Yang, Pingpeng Yuan, Hai Jin

Open publisher page 0 citations

Abstract

The RDF (Resource Description Framework) data model has been used in various domains, such as Web,government, biology etc. Now, the volume of RDF datasets is growing significantly. The explosion on the volume of RDF data raises serious challenges: how to answer SPARQL queries on large RDF data sets efficiently. Here, we present a large-scale RDF data system - TripleParallel, which implements blockbased parallel processing SPARQL queries on RDF data sets with billion triples. The system improves parallelism while strengthening the overlapping data and calculations and reduces the overall execution time of the query. TripleParallel also implements multiple parallel operations for parallel processing joins. Experimental studies with several RDF datasets, including the LUBM and the UniProt collection, demonstrate the performance gains of our approach, outperforming the previous fastest system by more than an order of magnitude.

About this research paper

What this paper is about

The RDF (Resource Description Framework) data model has been used in various domains, such as Web,government, biology etc. Now, the volume of RDF datasets is growing significantly. The explosion on the volume of RDF data raises serious challenges: how to answer SPARQL queries on large RDF data sets efficiently. Here, we present a large-scale RDF data system - TripleParallel, which implements blockbased parallel processing SPARQL queries on RDF data sets with billion triples. The system improves parallelism while strengthening the overlapping data and calculations and reduces the overall execution time of the query. TripleParallel also implements multiple parallel operations for parallel processing joins. Experimental studies with several RDF datasets, including the LUBM and the UniProt collection, demonstrate the performance gains of our approach, outperforming the previous fastest system by more than an order of magnitude.

Why it matters

A significance statement is not available in the OpenAlex record.

Key contribution

A contribution statement is not available in the OpenAlex record.

Method / approach

Method details are not available in the OpenAlex metadata.

Main findings

Findings are not separately available in the OpenAlex metadata.

Limitations

Limitations are not available in the OpenAlex metadata.

Applications

Application details are not available in the OpenAlex metadata.

Available abstract

The RDF (Resource Description Framework) data model has been used in various domains, such as Web,government, biology etc. Now, the volume of RDF datasets is growing significantly. The explosion on the volume of RDF data raises serious challenges: how to answer SPARQL queries on large RDF data sets efficiently. Here, we present a large-scale RDF data system - TripleParallel, which implements blockbased parallel processing SPARQL queries on RDF data sets with billion triples. The system improves parallelism while strengthening the overlapping data and calculations and reduces the overall execution time of the query. TripleParallel also implements multiple parallel operations for parallel processing joins. Experimental studies with several RDF datasets, including the LUBM and the UniProt collection, demonstrate the performance gains of our approach, outperforming the previous fastest system by more than an order of magnitude.

Key concepts: SPARQL, RDF, Computer science, RDF Schema, Joins, Linked data, RDF query language, RDF/XML

Related papers

Back to paper searchBrowse research topicsOriginal source
Fast Processing SPARQL Queries on Large RDF Data — Research Paper | ScholarLens