CRIU - Checkpoint Restore in Userspace for Computational Simulations and Scientific Applications
26TH INTERNATIONAL CONFERENCE ON COMPUTING IN HIGH ENERGY AND NUCLEAR PHYSICS, CHEP 2023(2024)
摘要
Creating new materials, discovering new drugs, and simulating systems areessential processes for research and innovation and require substantialcomputational power. While many applications can be split into many smallerindependent tasks, some cannot and may take hours or weeks to run tocompletion. To better manage those longer-running jobs, it would be desirableto stop them at any arbitrary point in time and later continue theircomputation on another compute resource; this is usually referred to ascheckpointing. While some applications can manage checkpointingprogrammatically, it would be preferable if the batch scheduling system coulddo that independently. This paper evaluates the feasibility of using CRIU(Checkpoint Restore in Userspace), an open-source tool for the GNU/Linuxenvironments, emphasizing the OSG's OSPool HTCondor setup. CRIU allowscheckpointing the process state into a disk image and can deal with both openfiles and established network connections seamlessly. Furthermore, it cancheckpoint traditional Linux processes and containerized workloads. Thefunctionality seems adequate for many scenarios supported in the OSPool.However, some limitations prevent it from being usable in all circumstances.
更多查看译文
AI 理解论文
溯源树
样例
生成溯源树,研究论文发展脉络
Chat Paper
正在生成论文摘要