2020Tehnicki vjesnik - Technical GazetteOpen access

Data Deduplication Technology for Cloud Storage

Qinlu He, Genqing Bian, Bilin Shao, Weiqi Zhang

Open full text 5 citations

Abstract

With the explosive growth of information data, the data storage system has stepped into the cloud storage era.Although the core of the cloud storage system is distributed file system in solving the problem of mass data storage, a large number of duplicate data exist in all storage system.File systems are designed to control how files are stored and retrieved.Fewer studies focus on the cloud file system deduplication technologies at the application level, especially for the Hadoop distributed file system.In this paper, we design a file deduplication framework on Hadoop distributed file system for cloud application developer.Proposed RFD-HDFS and FD-HDFS two data deduplication solutions process data deduplication online, which improves storage space utilisation and reduces the redundancy.In the end of the paper, we test the disk utilisation and the file upload performance on RFD-HDFS and FD-HDFS, and compare HDFS with the disk utilisation of two system frameworks.The results show that the two-system framework not only implements data deduplication function but also effectively reduces the disk utilisation of duplicate files.So, the proposed framework can indeed reduce the storage space by eliminating redundant HDFS file.

Open-access reader

About this research paper

What this paper is about

With the explosive growth of information data, the data storage system has stepped into the cloud storage era.Although the core of the cloud storage system is distributed file system in solving the problem of mass data storage, a large number of duplicate data exist in all storage system.File systems are designed to control how files are stored and retrieved.Fewer studies focus on the cloud file system deduplication technologies at the application level, especially for the Hadoop distributed file system.In this paper, we design a file deduplication framework on Hadoop distributed file system for cloud application developer.Proposed RFD-HDFS and FD-HDFS two data deduplication solutions process data deduplication online, which improves storage space utilisation and reduces the redundancy.In the end of the paper, we test the disk utilisation and the file upload performance on RFD-HDFS and FD-HDFS, and compare HDFS with the disk utilisation of two system frameworks.The results show that the two-system framework not only implements data deduplication function but also effectively reduces the disk utilisation of duplicate files.So, the proposed framework can indeed reduce the storage space by eliminating redundant HDFS file.

Why it matters

OpenAlex reports 5 citations for this work. Citation counts describe recorded attention and do not establish research quality.

Key contribution

A contribution statement is not available in the OpenAlex record.

Method / approach

Method details are not available in the OpenAlex metadata.

Main findings

Findings are not separately available in the OpenAlex metadata.

Limitations

Limitations are not available in the OpenAlex metadata.

Applications

Application details are not available in the OpenAlex metadata.

Available abstract

With the explosive growth of information data, the data storage system has stepped into the cloud storage era.Although the core of the cloud storage system is distributed file system in solving the problem of mass data storage, a large number of duplicate data exist in all storage system.File systems are designed to control how files are stored and retrieved.Fewer studies focus on the cloud file system deduplication technologies at the application level, especially for the Hadoop distributed file system.In this paper, we design a file deduplication framework on Hadoop distributed file system for cloud application developer.Proposed RFD-HDFS and FD-HDFS two data deduplication solutions process data deduplication online, which improves storage space utilisation and reduces the redundancy.In the end of the paper, we test the disk utilisation and the file upload performance on RFD-HDFS and FD-HDFS, and compare HDFS with the disk utilisation of two system frameworks.The results show that the two-system framework not only implements data deduplication function but also effectively reduces the disk utilisation of duplicate files.So, the proposed framework can indeed reduce the storage space by eliminating redundant HDFS file.

Key concepts: Data deduplication, Cloud computing, Cloud storage, Database, Computer science, Operating system

Related papers

Back to paper searchBrowse research topicsOriginal source
Data Deduplication Technology for Cloud Storage — Research Paper | ScholarLens