Buscar

Mostrando ítems 1-3 de 3

Improving Scalability of Application-Level Checkpoint-Recovery by Reducing Checkpoint Sizes

Cores González, Iván; Rodríguez, Gabriel; Martín, María J.; González, Patricia; Osorio, Roberto (Springer Japan KK, 2013)

[Abstract] The execution times of large-scale parallel applications on nowadays multi/many-core systems are usually longer than the mean time between failures. Therefore, parallel applications must tolerate hardware failures ...

Failure Avoidance in MPI Applications Using an Application-Level Approach

Cores González, Iván; Rodríguez, Gabriel; González, Patricia; Martín, María J. (Oxford University Press, 2014)

[Abstract] Execution times of large-scale computational science and engineering parallel applications are usually longer than the mean-time-between-failures. For this reason, hardware failures must be tolerated by the ...

Compiler-Assisted Checkpointing of Parallel Codes: The Cetus and LLVM Experience

Rodríguez, Gabriel; Martín, María J.; González, Patricia; Touriño, Juan; Doallo, Ramón (Springer New York LLC, 2013)

[Abstract] With the evolution of high-performance computing, parallel applications have developed an increasing necessity for fault tolerance, most commonly provided by checkpoint and restart techniques. Checkpointing tools ...

Buscar

Filtros

Improving Scalability of Application-Level Checkpoint-Recovery by Reducing Checkpoint Sizes

Failure Avoidance in MPI Applications Using an Application-Level Approach

Compiler-Assisted Checkpointing of Parallel Codes: The Cetus and LLVM Experience