Checkpoint/restart approaches for a thread-based MPI runtime.

Saved in:
Bibliographic Details
Title: Checkpoint/restart approaches for a thread-based MPI runtime.
Authors: Adam, Julien1 (AUTHOR), Kermarquer, Maxime2 (AUTHOR), Besnard, Jean-Baptiste1 (AUTHOR) jbbesnard@paratools.fr, Bautista-Gomez, Leonardo3 (AUTHOR), Pérache, Marc2 (AUTHOR), Carribault, Patrick2 (AUTHOR), Jaeger, Julien2 (AUTHOR), Malony, Allen D.4 (AUTHOR), Shende, Sameer4 (AUTHOR)
Source: Parallel Computing. Jul2019, Vol. 85, p204-219. 16p.
Subjects: Software failures, Data replication, Parallel programming, Network performance, Run time systems (Computer science)
Abstract: • Transparent checkpoint restart can be applied to high-speed networks with collaboration from the MPI runtime (particularly network modularity). • Thread-based MPI runtimes can be checkpointed both transparently and at application-level without blocking difficulties when compared to their process-based counterpart. • We introduce an asynchronous checkpointing interface for transparent checkpointing. Fault-tolerance has always been an important topic when it comes to running massively parallel programs at scale. Statistically, hardware and software failures are expected to occur more often on systems gathering millions of computing units. Moreover, the larger jobs are, the more computing hours would be wasted by a crash. In this paper, we describe the work done in our MPI runtime to enable both transparent and application-level checkpointing mechanisms. Unlike the MPI 4.0 User-Level Failure Mitigation (ULFM) interface, our work targets solely Checkpoint/Restart and ignores other features such as resiliency. We show how existing checkpointing methods can be practically applied to a thread-based MPI implementation given sufficient runtime collaboration. The two main contributions are the preservation of high-speed network performance during transparent C/R and the over-subscription of checkpoint data replication thanks to a dedicated user-level scheduler support. These techniques are measured on MPI benchmarks such as IMB, Lulesh and Heatdis, and associated overhead and trade-offs are discussed. [ABSTRACT FROM AUTHOR]
Copyright of Parallel Computing is the property of Elsevier B.V. and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Engineering Source
FullText Text:
  Availability: 0
Header DbId: egs
DbLabel: Engineering Source
An: 136581236
AccessLevel: 6
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: Checkpoint/restart approaches for a thread-based MPI runtime.
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Adam%2C+Julien%22">Adam, Julien</searchLink><relatesTo>1</relatesTo> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Kermarquer%2C+Maxime%22">Kermarquer, Maxime</searchLink><relatesTo>2</relatesTo> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Besnard%2C+Jean-Baptiste%22">Besnard, Jean-Baptiste</searchLink><relatesTo>1</relatesTo> (AUTHOR)<i> jbbesnard@paratools.fr</i><br /><searchLink fieldCode="AR" term="%22Bautista-Gomez%2C+Leonardo%22">Bautista-Gomez, Leonardo</searchLink><relatesTo>3</relatesTo> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Pérache%2C+Marc%22">Pérache, Marc</searchLink><relatesTo>2</relatesTo> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Carribault%2C+Patrick%22">Carribault, Patrick</searchLink><relatesTo>2</relatesTo> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Jaeger%2C+Julien%22">Jaeger, Julien</searchLink><relatesTo>2</relatesTo> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Malony%2C+Allen+D%2E%22">Malony, Allen D.</searchLink><relatesTo>4</relatesTo> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Shende%2C+Sameer%22">Shende, Sameer</searchLink><relatesTo>4</relatesTo> (AUTHOR)
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="JN" term="%22Parallel+Computing%22">Parallel Computing</searchLink>. Jul2019, Vol. 85, p204-219. 16p.
– Name: Subject
  Label: Subjects
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Software+failures%22">Software failures</searchLink><br /><searchLink fieldCode="DE" term="%22Data+replication%22">Data replication</searchLink><br /><searchLink fieldCode="DE" term="%22Parallel+programming%22">Parallel programming</searchLink><br /><searchLink fieldCode="DE" term="%22Network+performance%22">Network performance</searchLink><br /><searchLink fieldCode="DE" term="%22Run+time+systems+%28Computer+science%29%22">Run time systems (Computer science)</searchLink>
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: • Transparent checkpoint restart can be applied to high-speed networks with collaboration from the MPI runtime (particularly network modularity). • Thread-based MPI runtimes can be checkpointed both transparently and at application-level without blocking difficulties when compared to their process-based counterpart. • We introduce an asynchronous checkpointing interface for transparent checkpointing. Fault-tolerance has always been an important topic when it comes to running massively parallel programs at scale. Statistically, hardware and software failures are expected to occur more often on systems gathering millions of computing units. Moreover, the larger jobs are, the more computing hours would be wasted by a crash. In this paper, we describe the work done in our MPI runtime to enable both transparent and application-level checkpointing mechanisms. Unlike the MPI 4.0 User-Level Failure Mitigation (ULFM) interface, our work targets solely Checkpoint/Restart and ignores other features such as resiliency. We show how existing checkpointing methods can be practically applied to a thread-based MPI implementation given sufficient runtime collaboration. The two main contributions are the preservation of high-speed network performance during transparent C/R and the over-subscription of checkpoint data replication thanks to a dedicated user-level scheduler support. These techniques are measured on MPI benchmarks such as IMB, Lulesh and Heatdis, and associated overhead and trade-offs are discussed. [ABSTRACT FROM AUTHOR]
– Name: AbstractSuppliedCopyright
  Label:
  Group: Ab
  Data: <i>Copyright of Parallel Computing is the property of Elsevier B.V. and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.)
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=egs&AN=136581236
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1016/j.parco.2019.02.006
    Languages:
      – Code: eng
        Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 16
        StartPage: 204
    Subjects:
      – SubjectFull: Software failures
        Type: general
      – SubjectFull: Data replication
        Type: general
      – SubjectFull: Parallel programming
        Type: general
      – SubjectFull: Network performance
        Type: general
      – SubjectFull: Run time systems (Computer science)
        Type: general
    Titles:
      – TitleFull: Checkpoint/restart approaches for a thread-based MPI runtime.
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Adam, Julien
      – PersonEntity:
          Name:
            NameFull: Kermarquer, Maxime
      – PersonEntity:
          Name:
            NameFull: Besnard, Jean-Baptiste
      – PersonEntity:
          Name:
            NameFull: Bautista-Gomez, Leonardo
      – PersonEntity:
          Name:
            NameFull: Pérache, Marc
      – PersonEntity:
          Name:
            NameFull: Carribault, Patrick
      – PersonEntity:
          Name:
            NameFull: Jaeger, Julien
      – PersonEntity:
          Name:
            NameFull: Malony, Allen D.
      – PersonEntity:
          Name:
            NameFull: Shende, Sameer
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 07
              Text: Jul2019
              Type: published
              Y: 2019
          Identifiers:
            – Type: issn-print
              Value: 01678191
          Numbering:
            – Type: volume
              Value: 85
          Titles:
            – TitleFull: Parallel Computing
              Type: main
ResultId 1