当前位置:主页 > 科技论文 > 软件论文 >

并行计算框架Spark的自动检查点策略

发布时间:2019-01-29 20:40
【摘要】:针对现有的Spark检查点机制需要编程人员根据经验选择检查点,具有一定的风险和随机性,可能导致恢复开销较大的问题,通过对RDD属性的分析,提出了自动检查点策略,包括权重生成(WG)算法和检查点自动选择(CAS)算法.首先,WG算法分析作业的DAG结构,获取RDD的血统长度和操作复杂度等属性,计算RDD权重;然后,CAS算法选择权重大的RDD作为检查点进行异步备份,来实现数据的快速恢复.结果表明:在使用CAS算法时,不同数据集执行时间和检查点容量大小都有所增加,其中Wiki-Talk由于其计算量较大,增幅明显;使用CAS算法设置检查点后,在单点失效恢复的情况下,数据集的恢复时间较短.因此,自动检查点策略在略微增加执行时间开销的基础上,能够有效地降低作业的恢复开销.
[Abstract]:In view of the fact that the existing Spark checkpoint mechanism needs the programmer to select the checkpoint according to the experience, it has certain risks and randomness, which may lead to the problem of large recovery overhead. Through the analysis of the RDD attribute, the automatic checkpoint strategy is put forward. Including weight generation (WG) algorithm and checkpoint automatic selection (CAS) algorithm. First, the WG algorithm analyzes the DAG structure of the job, acquires the properties of the RDD, such as the length of the lineage and the complexity of the operation, and calculates the RDD weight. Then, the CAS algorithm chooses RDD as a checkpoint to perform asynchronous backup to realize the fast data recovery. The results show that when CAS algorithm is used, the execution time and checkpoint capacity of different data sets are increased, among which, Wiki-Talk has a significant increase due to its large amount of calculation. After the checkpoint is set up by CAS algorithm, the recovery time of data set is shorter than that of single point failure recovery. Therefore, automatic checkpoint strategy can effectively reduce the cost of job recovery on the basis of a slight increase in execution time.
【作者单位】: 新疆大学信息科学与工程学院;新疆大学软件学院;
【基金】:国家自然科学基金资助项目(61462079,61262088,61562086,61363083,61562078) 新疆维吾尔自治区高校科研计划资助项目(XJEDU2016S106)
【分类号】:TP311.13

【相似文献】

相关期刊论文 前2条

1 何鑫星;谭理;;数字地形图质量自动检查方法探讨及软件开发[J];测绘;2013年04期

2 ;[J];;年期



本文编号:2417843

资料下载
论文发表

本文链接:https://www.wllwen.com/kejilunwen/ruanjiangongchenglunwen/2417843.html


Copyright(c)文论论文网All Rights Reserved | 网站地图 |

版权申明:资料由用户58dd3***提供,本站仅收录摘要或目录,作者需要删除请E-mail邮箱bigeng88@qq.com