Evaluating Data Consistency with Matching Dependencies from Multiple Sources
Mi Ni Huang, Lingli Li, Ping Xuan
Abstract
Mi Ni Huang, Lingli Li, Ping Xuan
Abstract
With the rapid growth of data, data quality issues have attracted increasing attention in both industry and academia. Since data consistency is one of the critical issues in data quality, we study the problem of how to evaluate the consistency of target data from multiple relevant sources under matching dependencies (MDs). Since accessing data sources directly introduces a huge cost of data comparisons, so this paper aims to design an efficient approximate consistency evaluation method with linear-time complexity. Firstly, we build a signature for each data source to approximate the pattern sets in this source defined by the MDs. Secondly, we develop a signature-based evaluation method to compute the consistency of target data based on the signatures of all the data sources that are related to our target data. Experimental results on real datasets shows high performance on both accuracy and efficiency of our algorithm.
OpenAlex reports 1 citations for this work. Citation counts describe recorded attention and do not establish research quality.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
With the rapid growth of data, data quality issues have attracted increasing attention in both industry and academia. Since data consistency is one of the critical issues in data quality, we study the problem of how to evaluate the consistency of target data from multiple relevant sources under matching dependencies (MDs). Since accessing data sources directly introduces a huge cost of data comparisons, so this paper aims to design an efficient approximate consistency evaluation method with linear-time complexity. Firstly, we build a signature for each data source to approximate the pattern sets in this source defined by the MDs. Secondly, we develop a signature-based evaluation method to compute the consistency of target data based on the signatures of all the data sources that are related to our target data. Experimental results on real datasets shows high performance on both accuracy and efficiency of our algorithm.
Key concepts: Consistency (knowledge bases), Data consistency, Computer science, Data mining, Data quality, Matching (statistics), Signature (topology), Quality (philosophy)