Checkpoint/restart approaches for a thread-based MPI runtime.
Saved in:
| Title: | Checkpoint/restart approaches for a thread-based MPI runtime. |
|---|---|
| Authors: | Adam, Julien1 (AUTHOR), Kermarquer, Maxime2 (AUTHOR), Besnard, Jean-Baptiste1 (AUTHOR) jbbesnard@paratools.fr, Bautista-Gomez, Leonardo3 (AUTHOR), Pérache, Marc2 (AUTHOR), Carribault, Patrick2 (AUTHOR), Jaeger, Julien2 (AUTHOR), Malony, Allen D.4 (AUTHOR), Shende, Sameer4 (AUTHOR) |
| Source: | Parallel Computing. Jul2019, Vol. 85, p204-219. 16p. |
| Subjects: | Software failures, Data replication, Parallel programming, Network performance, Run time systems (Computer science) |
| Abstract: | • Transparent checkpoint restart can be applied to high-speed networks with collaboration from the MPI runtime (particularly network modularity). • Thread-based MPI runtimes can be checkpointed both transparently and at application-level without blocking difficulties when compared to their process-based counterpart. • We introduce an asynchronous checkpointing interface for transparent checkpointing. Fault-tolerance has always been an important topic when it comes to running massively parallel programs at scale. Statistically, hardware and software failures are expected to occur more often on systems gathering millions of computing units. Moreover, the larger jobs are, the more computing hours would be wasted by a crash. In this paper, we describe the work done in our MPI runtime to enable both transparent and application-level checkpointing mechanisms. Unlike the MPI 4.0 User-Level Failure Mitigation (ULFM) interface, our work targets solely Checkpoint/Restart and ignores other features such as resiliency. We show how existing checkpointing methods can be practically applied to a thread-based MPI implementation given sufficient runtime collaboration. The two main contributions are the preservation of high-speed network performance during transparent C/R and the over-subscription of checkpoint data replication thanks to a dedicated user-level scheduler support. These techniques are measured on MPI benchmarks such as IMB, Lulesh and Heatdis, and associated overhead and trade-offs are discussed. [ABSTRACT FROM AUTHOR] |
| Copyright of Parallel Computing is the property of Elsevier B.V. and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.) | |
| Database: | Engineering Source |
| FullText | Text: Availability: 0 |
|---|---|
| Header | DbId: egs DbLabel: Engineering Source An: 136581236 AccessLevel: 6 PubType: Academic Journal PubTypeId: academicJournal PreciseRelevancyScore: 0 |
| IllustrationInfo | |
| Items | – Name: Title Label: Title Group: Ti Data: Checkpoint/restart approaches for a thread-based MPI runtime. – Name: Author Label: Authors Group: Au Data: <searchLink fieldCode="AR" term="%22Adam%2C+Julien%22">Adam, Julien</searchLink><relatesTo>1</relatesTo> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Kermarquer%2C+Maxime%22">Kermarquer, Maxime</searchLink><relatesTo>2</relatesTo> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Besnard%2C+Jean-Baptiste%22">Besnard, Jean-Baptiste</searchLink><relatesTo>1</relatesTo> (AUTHOR)<i> jbbesnard@paratools.fr</i><br /><searchLink fieldCode="AR" term="%22Bautista-Gomez%2C+Leonardo%22">Bautista-Gomez, Leonardo</searchLink><relatesTo>3</relatesTo> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Pérache%2C+Marc%22">Pérache, Marc</searchLink><relatesTo>2</relatesTo> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Carribault%2C+Patrick%22">Carribault, Patrick</searchLink><relatesTo>2</relatesTo> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Jaeger%2C+Julien%22">Jaeger, Julien</searchLink><relatesTo>2</relatesTo> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Malony%2C+Allen+D%2E%22">Malony, Allen D.</searchLink><relatesTo>4</relatesTo> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Shende%2C+Sameer%22">Shende, Sameer</searchLink><relatesTo>4</relatesTo> (AUTHOR) – Name: TitleSource Label: Source Group: Src Data: <searchLink fieldCode="JN" term="%22Parallel+Computing%22">Parallel Computing</searchLink>. Jul2019, Vol. 85, p204-219. 16p. – Name: Subject Label: Subjects Group: Su Data: <searchLink fieldCode="DE" term="%22Software+failures%22">Software failures</searchLink><br /><searchLink fieldCode="DE" term="%22Data+replication%22">Data replication</searchLink><br /><searchLink fieldCode="DE" term="%22Parallel+programming%22">Parallel programming</searchLink><br /><searchLink fieldCode="DE" term="%22Network+performance%22">Network performance</searchLink><br /><searchLink fieldCode="DE" term="%22Run+time+systems+%28Computer+science%29%22">Run time systems (Computer science)</searchLink> – Name: Abstract Label: Abstract Group: Ab Data: • Transparent checkpoint restart can be applied to high-speed networks with collaboration from the MPI runtime (particularly network modularity). • Thread-based MPI runtimes can be checkpointed both transparently and at application-level without blocking difficulties when compared to their process-based counterpart. • We introduce an asynchronous checkpointing interface for transparent checkpointing. Fault-tolerance has always been an important topic when it comes to running massively parallel programs at scale. Statistically, hardware and software failures are expected to occur more often on systems gathering millions of computing units. Moreover, the larger jobs are, the more computing hours would be wasted by a crash. In this paper, we describe the work done in our MPI runtime to enable both transparent and application-level checkpointing mechanisms. Unlike the MPI 4.0 User-Level Failure Mitigation (ULFM) interface, our work targets solely Checkpoint/Restart and ignores other features such as resiliency. We show how existing checkpointing methods can be practically applied to a thread-based MPI implementation given sufficient runtime collaboration. The two main contributions are the preservation of high-speed network performance during transparent C/R and the over-subscription of checkpoint data replication thanks to a dedicated user-level scheduler support. These techniques are measured on MPI benchmarks such as IMB, Lulesh and Heatdis, and associated overhead and trade-offs are discussed. [ABSTRACT FROM AUTHOR] – Name: AbstractSuppliedCopyright Label: Group: Ab Data: <i>Copyright of Parallel Computing is the property of Elsevier B.V. and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.) |
| PLink | https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=egs&AN=136581236 |
| RecordInfo | BibRecord: BibEntity: Identifiers: – Type: doi Value: 10.1016/j.parco.2019.02.006 Languages: – Code: eng Text: English PhysicalDescription: Pagination: PageCount: 16 StartPage: 204 Subjects: – SubjectFull: Software failures Type: general – SubjectFull: Data replication Type: general – SubjectFull: Parallel programming Type: general – SubjectFull: Network performance Type: general – SubjectFull: Run time systems (Computer science) Type: general Titles: – TitleFull: Checkpoint/restart approaches for a thread-based MPI runtime. Type: main BibRelationships: HasContributorRelationships: – PersonEntity: Name: NameFull: Adam, Julien – PersonEntity: Name: NameFull: Kermarquer, Maxime – PersonEntity: Name: NameFull: Besnard, Jean-Baptiste – PersonEntity: Name: NameFull: Bautista-Gomez, Leonardo – PersonEntity: Name: NameFull: Pérache, Marc – PersonEntity: Name: NameFull: Carribault, Patrick – PersonEntity: Name: NameFull: Jaeger, Julien – PersonEntity: Name: NameFull: Malony, Allen D. – PersonEntity: Name: NameFull: Shende, Sameer IsPartOfRelationships: – BibEntity: Dates: – D: 01 M: 07 Text: Jul2019 Type: published Y: 2019 Identifiers: – Type: issn-print Value: 01678191 Numbering: – Type: volume Value: 85 Titles: – TitleFull: Parallel Computing Type: main |
| ResultId | 1 |