Skip to main content

WIP2010Published

Checkpointing Orchestration for Performance Improvement

Authors
H. Jin
Venue
40th Annual IEEE/IFIP International Conference on Dependable Systems and Networks (DSN 2010), Student Paper
Date
2010
Type
WIP

Abstract

Checkpointing is a mostly used mechanism for supporting fault tolerance of high performance computing (HPC), but notorious in its expensive disk access. Parallel file systems such as Lustre, GPFS, PVFS are widely deployed on super computers to provide fast I/O bandwidth for general data-intensive applications. However, the unique feature of checkpointing makes it impossible to benefit from the parallel file systems. In addition, the design of parallel file system introduces extra contention overhead for checkpointing and significantly degrades the performance. In this study, we propose checkpointing orchestration to mask the unnecessary overhead for a better performance. We extend Open MPI and PVFS to support the idea of checkpointing orchestration. The experimental results confirm the potential of the proposed checkpointing orchestration.

Knowledge Graph Connections