Novel Data-Placement Scheme For Improving The Data Locality Of Hadoop In Heterogeneous Environments

CONCURRENCY AND COMPUTATION-PRACTICE & EXPERIENCE(2021)

引用 7|浏览14
暂无评分
摘要
To address the challenging needs of high-performance big data processing, parallel-distributed frameworks such as Hadoop are being utilized extensively. However, in heterogeneous environments, the performance of Hadoop clusters is below par. This is primarily because the blocks of the clusters are allocated equally to all nodes without regard to differences in the capability of individual nodes. This results in reduced data locality. Thus, a new data-placement scheme that enhances data locality is required for Hadoop in heterogeneous environments. This article proposes a new data placement scheme that preserves the same degree of data locality in heterogeneous environments as that of the standard Hadoop, with only a small amount of replicated data. In the proposed scheme, only those blocks with the highest probability of being accessed remotely are selected and replicated. The results of experiments conducted indicate that the proposed scheme incurs only a 20% disk space overhead and has virtually the same data locality ratio as the standard Hadoop, which has a replication factor of three and 200% disk space overhead.
更多
查看译文
关键词
data locality, data placement, Hadoop MapReduce, heterogeneous environment, replication
AI 理解论文
溯源树
样例
生成溯源树,研究论文发展脉络
Chat Paper
正在生成论文摘要