Combination of Replication and Scheduling in Data Grids
Nhan Nguyen Dang, Sang Boem Lim
Abstract
Nhan Nguyen Dang, Sang Boem Lim
Abstract
Data Grid environment is a geographically distributed that deal with date-intensive application in scientific and enterprise computing. Dealing with large amount of data makes the requirement for efficiency in data access more critical. The goal of replication is to shorten the data access not only for user accesses but enhancing the job execution performance. In this paper, we proposed a new approach to replication based on organizing the data in Data Grid based on its property. In this paper, we organized the data in to several data categories that it belongs to. And this information is used to help improving data replication placement strategy. We study our approach and evaluate it through simulation. The result shows that our algorithm has improved 30% over the current strategies.
OpenAlex reports 46 citations for this work. Citation counts describe recorded attention and do not establish research quality.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
Data Grid environment is a geographically distributed that deal with date-intensive application in scientific and enterprise computing. Dealing with large amount of data makes the requirement for efficiency in data access more critical. The goal of replication is to shorten the data access not only for user accesses but enhancing the job execution performance. In this paper, we proposed a new approach to replication based on organizing the data in Data Grid based on its property. In this paper, we organized the data in to several data categories that it belongs to. And this information is used to help improving data replication placement strategy. We study our approach and evaluate it through simulation. The result shows that our algorithm has improved 30% over the current strategies.
Key concepts: Data grid, Computer science, Replication (statistics), Distributed computing, Grid, Data access, Scheduling (production processes), Grid computing