Mostrar o rexistro simple do ítem

dc.contributor.authorLosada, Nuria
dc.contributor.authorBosilca, George
dc.contributor.authorBouteiller, Aurelien
dc.contributor.authorGonzález, Patricia
dc.contributor.authorMartín, María J.
dc.date.accessioned2021-03-24T14:59:20Z
dc.date.available2021-03-24T14:59:20Z
dc.date.issued2019-02
dc.identifier.citationNuria Losada, George Bosilca, Aurélien Bouteiller, Patricia González, María J. Martín, Local rollback for resilient MPI applications with application-level checkpointing and message logging, Future Generation Computer Systems, Volume 91, 2019, Pages 450-464, ISSN 0167-739X, https://doi.org/10.1016/j.future.2018.09.041.es_ES
dc.identifier.issn0167-739X
dc.identifier.issn1872-7115
dc.identifier.urihttp://hdl.handle.net/2183/27584
dc.description.abstract[Abstract] The resilience approach generally used in high-performance computing (HPC) relies on coordinated checkpoint/restart, a global rollback of all the processes that are running the application. However, in many instances, the failure has a more localized scope and its impact is usually restricted to a subset of the resources being used. Thus, a global rollback would result in unnecessary overhead and energy consumption, since all processes, including those unaffected by the failure, discard their state and roll back to the last checkpoint to repeat computations that were already done. The User Level Failure Mitigation (ULFM) interface – the last proposal for the inclusion of resilience features in the Message Passing Interface (MPI) standard – enables the deployment of more flexible recovery strategies, including localized recovery. This work proposes a local rollback approach that can be generally applied to Single Program, Multiple Data (SPMD) applications by combining ULFM, the ComPiler for Portable Checkpointing (CPPC) tool, and the Open MPI VProtocol system-level message logging component. Only failed processes are recovered from the last checkpoint, while consistency before further progress in the execution is achieved through a two-level message logging process. To further optimize this approach point-to-point communications are logged by the Open MPI VProtocol component, while collective communications are optimally logged at the application level—thereby decoupling the logging protocol from the particular collective implementation. This spatially coordinated protocol applied by CPPC reduces the log size, the log memory requirements and overall the resilience impact on the applications.es_ES
dc.description.sponsorshipThis research was supported by the Ministry of Economy and Competitiveness of Spain and FEDER funds of the EU (Projects TIN2016-75845-P and the predoctoral grants of Nuria Losada ref. BES-2014-068066 and ref. EEBB-I-17-12005); by EU under the COST Program Action IC1305 Network for Sustainable Ultrascale Computing (NESUS) and a HiPEAC Collaboration Grant and by the Galician Government (Xunta de Galicia) under the Consolidation Program of Competitive Research (ref. ED431C 2017/04). We gratefully thank Galicia Supercomputing Center for providing access to the FinisTerrae-II supercomputer. This material is also based upon work supported by the US National Science Foundation, Office of Advanced Cyberinfrastructure , under Grants No. #1664142 and #1339763es_ES
dc.description.sponsorshipXunta de Galicia; ED431C 2017/04es_ES
dc.description.sponsorshipUS National Science Foundation, Office of Advanced Cyberinfrastructure; 1664142es_ES
dc.description.sponsorshipUS National Science Foundation, Office of Advanced Cyberinfrastructure; 1339763es_ES
dc.language.isoenges_ES
dc.publisherElsevier BV * North-Hollandes_ES
dc.relationinfo:eu-repo/grantAgreement/MINECO/Plan Estatal de Investigación Científica y Técnica y de Innovación 2013-2016/TIN2016-75845-P/ES/NUEVOS DESAFIOS EN COMPUTACION DE ALTAS PRESTACIONES: DESDE ARQUITECTURAS HASTA APLICACIONES (II)
dc.relationinfo:eu-repo/grantAgreement/MINECO/Plan Estatal de Investigación Científica y Técnica y de Innovación 2013-2016/BES-2014-068066/ES/
dc.relationinfo:eu-repo/grantAgreement/AEI/Plan Estatal de Investigación Científica y Técnica y de Innovación 2013-2016/EEBB-I-17-12005/ES/
dc.relation.urihttps://doi.org/10.1016/j.future.2018.09.041es_ES
dc.rightsAtribución-NoComercial-SinDerivadas 4.0 Internacionales_ES
dc.rights.urihttp://creativecommons.org/licenses/by-nc-nd/4.0/*
dc.subjectMPIes_ES
dc.subjectResiliencees_ES
dc.subjectMessage logginges_ES
dc.subjectApplication-level checkpointinges_ES
dc.subjectLocal rollbackes_ES
dc.titleLocal Rollback for Resilient Mpi Applications With Application-Level Checkpointing and Message Logginges_ES
dc.typeinfo:eu-repo/semantics/articlees_ES
dc.rights.accessinfo:eu-repo/semantics/openAccesses_ES
UDC.journalTitleFuture Generation Computer Systemses_ES
UDC.volume91es_ES
UDC.startPage450es_ES
UDC.endPage464es_ES
dc.identifier.doi10.1016/j.future.2018.09.041


Ficheiros no ítem

Thumbnail
Thumbnail

Este ítem aparece na(s) seguinte(s) colección(s)

Mostrar o rexistro simple do ítem