HydEE: Failure Containment without Event Logging for Large Scale Send-Deterministic MPI Applications - INRIA - Institut National de Recherche en Informatique et en Automatique Accéder directement au contenu
Communication Dans Un Congrès Année : 2012

HydEE: Failure Containment without Event Logging for Large Scale Send-Deterministic MPI Applications

Résumé

High performance computing will probably reach exascale in this decade. At this scale, mean time between failures is expected to be a few hours. Existing fault tolerant protocols for message passing applications will not be efficient anymore since they either require a global restart after a failure (checkpointing protocols) or result in huge memory occupation (message logging). Hybrid fault tolerant protocols overcome these limits by dividing applications processes into clusters and applying a different protocol within and between clusters. Combining coordinated checkpointing inside the clusters and message logging for the inter-cluster messages allows confining the consequences of a failure to a single cluster, while logging only a subset of the messages. However, in existing hybrid protocols, event logging is required for all application messages to ensure a correct execution after a failure. This can significantly impair failure free performance. In this paper, we propose HydEE, a hybrid rollback-recovery protocol for send-deterministic message passing applications, that provides failure containment without logging any event, and only a subset of the application messages. We prove that HydEE can handle multiple concurrent failures by relying on the send-deterministic execution model. Experimental evaluations of our implementation of HydEE in the MPICH2 library show that it introduces almost no overhead on failure free execution.
Fichier principal
Vignette du fichier
ipdps2012.pdf (351.97 Ko) Télécharger le fichier
Origine : Fichiers produits par l'(les) auteur(s)

Dates et versions

hal-01121941 , version 1 (02-03-2015)

Identifiants

Citer

Amina Guermouche, Thomas Ropars, Marc Snir, Franck Cappello. HydEE: Failure Containment without Event Logging for Large Scale Send-Deterministic MPI Applications. {26th IEEE International Parallel & Distributed Processing Symposium (IPDPS2012), 2012, Shanghai, China. ⟨10.1109/IPDPS.2012.111⟩. ⟨hal-01121941⟩
136 Consultations
180 Téléchargements

Altmetric

Partager

Gmail Facebook X LinkedIn More