Repair Tree: Fast Repair for Single Failure in Erasure-Coded Distributed Storage Systems.
Saved in:
| Title: | Repair Tree: Fast Repair for Single Failure in Erasure-Coded Distributed Storage Systems. |
|---|---|
| Authors: | Zhang, Huayu1, Li, Hui1, Li, Shuo-Yen Robert2 |
| Source: | IEEE Transactions on Parallel & Distributed Systems. Jun2017, Vol. 28 Issue 6, p1728-1739. 12p. |
| Subjects: | Reliability engineering software, Computer system failures, Distributed computing |
| Abstract: | In order to guarantee data reliability, erasure codes have been used in distributed storage systems. Nevertheless, this mechanism suffers from the repair problem that excess data are needed to repair a single failure, causing both high bandwidth consuming for the network and heavy computing load on the replacement node. To reduce repair traffic, researchers pointed out the tradeoff between storage and repair traffic and proposed regenerating codes by combining network coding. However, the combination only focuses on the storage terminal and the construction of the codes is quite complicated. Therefore, this paper further combines network coding with network structure and proposes a repair tree model based on general erasure codes to simplify the repair procedure. By decomposing repair computing and distributing it among the tree nodes, our model can mitigate the computing tension. The performance of repair tree is analyzed and evaluated by preliminary emulation. The result shows it can make about three times faster computing than conventional measure and the repair throughput is doubled if there are network bottlenecks. For proper topology, it can significantly reduce the repair traffic. We present algorithms to generate trees across the network topology. At last, we present the idea of extending repair tree to repair multiple failures. [ABSTRACT FROM AUTHOR] |
| Copyright of IEEE Transactions on Parallel & Distributed Systems is the property of IEEE and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.) | |
| Database: | Engineering Source |
| FullText | Text: Availability: 0 |
|---|---|
| Header | DbId: egs DbLabel: Engineering Source An: 123209380 AccessLevel: 6 PubType: Academic Journal PubTypeId: academicJournal PreciseRelevancyScore: 0 |
| IllustrationInfo | |
| Items | – Name: Title Label: Title Group: Ti Data: Repair Tree: Fast Repair for Single Failure in Erasure-Coded Distributed Storage Systems. – Name: Author Label: Authors Group: Au Data: <searchLink fieldCode="AR" term="%22Zhang%2C+Huayu%22">Zhang, Huayu</searchLink><relatesTo>1</relatesTo><br /><searchLink fieldCode="AR" term="%22Li%2C+Hui%22">Li, Hui</searchLink><relatesTo>1</relatesTo><br /><searchLink fieldCode="AR" term="%22Li%2C+Shuo-Yen+Robert%22">Li, Shuo-Yen Robert</searchLink><relatesTo>2</relatesTo> – Name: TitleSource Label: Source Group: Src Data: <searchLink fieldCode="JN" term="%22IEEE+Transactions+on+Parallel+%26+Distributed+Systems%22">IEEE Transactions on Parallel & Distributed Systems</searchLink>. Jun2017, Vol. 28 Issue 6, p1728-1739. 12p. – Name: Subject Label: Subjects Group: Su Data: <searchLink fieldCode="DE" term="%22Reliability+engineering+software%22">Reliability engineering software</searchLink><br /><searchLink fieldCode="DE" term="%22Computer+system+failures%22">Computer system failures</searchLink><br /><searchLink fieldCode="DE" term="%22Distributed+computing%22">Distributed computing</searchLink> – Name: Abstract Label: Abstract Group: Ab Data: In order to guarantee data reliability, erasure codes have been used in distributed storage systems. Nevertheless, this mechanism suffers from the repair problem that excess data are needed to repair a single failure, causing both high bandwidth consuming for the network and heavy computing load on the replacement node. To reduce repair traffic, researchers pointed out the tradeoff between storage and repair traffic and proposed regenerating codes by combining network coding. However, the combination only focuses on the storage terminal and the construction of the codes is quite complicated. Therefore, this paper further combines network coding with network structure and proposes a repair tree model based on general erasure codes to simplify the repair procedure. By decomposing repair computing and distributing it among the tree nodes, our model can mitigate the computing tension. The performance of repair tree is analyzed and evaluated by preliminary emulation. The result shows it can make about three times faster computing than conventional measure and the repair throughput is doubled if there are network bottlenecks. For proper topology, it can significantly reduce the repair traffic. We present algorithms to generate trees across the network topology. At last, we present the idea of extending repair tree to repair multiple failures. [ABSTRACT FROM AUTHOR] – Name: AbstractSuppliedCopyright Label: Group: Ab Data: <i>Copyright of IEEE Transactions on Parallel & Distributed Systems is the property of IEEE and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.) |
| PLink | https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=egs&AN=123209380 |
| RecordInfo | BibRecord: BibEntity: Identifiers: – Type: doi Value: 10.1109/TPDS.2016.2628024 Languages: – Code: eng Text: English PhysicalDescription: Pagination: PageCount: 12 StartPage: 1728 Subjects: – SubjectFull: Reliability engineering software Type: general – SubjectFull: Computer system failures Type: general – SubjectFull: Distributed computing Type: general Titles: – TitleFull: Repair Tree: Fast Repair for Single Failure in Erasure-Coded Distributed Storage Systems. Type: main BibRelationships: HasContributorRelationships: – PersonEntity: Name: NameFull: Zhang, Huayu – PersonEntity: Name: NameFull: Li, Hui – PersonEntity: Name: NameFull: Li, Shuo-Yen Robert IsPartOfRelationships: – BibEntity: Dates: – D: 01 M: 06 Text: Jun2017 Type: published Y: 2017 Identifiers: – Type: issn-print Value: 10459219 Numbering: – Type: volume Value: 28 – Type: issue Value: 6 Titles: – TitleFull: IEEE Transactions on Parallel & Distributed Systems Type: main |
| ResultId | 1 |