Repair Tree: Fast Repair for Single Failure in Erasure-Coded Distributed Storage Systems.

Saved in:
Bibliographic Details
Title: Repair Tree: Fast Repair for Single Failure in Erasure-Coded Distributed Storage Systems.
Authors: Zhang, Huayu1, Li, Hui1, Li, Shuo-Yen Robert2
Source: IEEE Transactions on Parallel & Distributed Systems. Jun2017, Vol. 28 Issue 6, p1728-1739. 12p.
Subjects: Reliability engineering software, Computer system failures, Distributed computing
Abstract: In order to guarantee data reliability, erasure codes have been used in distributed storage systems. Nevertheless, this mechanism suffers from the repair problem that excess data are needed to repair a single failure, causing both high bandwidth consuming for the network and heavy computing load on the replacement node. To reduce repair traffic, researchers pointed out the tradeoff between storage and repair traffic and proposed regenerating codes by combining network coding. However, the combination only focuses on the storage terminal and the construction of the codes is quite complicated. Therefore, this paper further combines network coding with network structure and proposes a repair tree model based on general erasure codes to simplify the repair procedure. By decomposing repair computing and distributing it among the tree nodes, our model can mitigate the computing tension. The performance of repair tree is analyzed and evaluated by preliminary emulation. The result shows it can make about three times faster computing than conventional measure and the repair throughput is doubled if there are network bottlenecks. For proper topology, it can significantly reduce the repair traffic. We present algorithms to generate trees across the network topology. At last, we present the idea of extending repair tree to repair multiple failures. [ABSTRACT FROM AUTHOR]
Copyright of IEEE Transactions on Parallel & Distributed Systems is the property of IEEE and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Engineering Source
FullText Text:
  Availability: 0
Header DbId: egs
DbLabel: Engineering Source
An: 123209380
AccessLevel: 6
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: Repair Tree: Fast Repair for Single Failure in Erasure-Coded Distributed Storage Systems.
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Zhang%2C+Huayu%22">Zhang, Huayu</searchLink><relatesTo>1</relatesTo><br /><searchLink fieldCode="AR" term="%22Li%2C+Hui%22">Li, Hui</searchLink><relatesTo>1</relatesTo><br /><searchLink fieldCode="AR" term="%22Li%2C+Shuo-Yen+Robert%22">Li, Shuo-Yen Robert</searchLink><relatesTo>2</relatesTo>
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="JN" term="%22IEEE+Transactions+on+Parallel+%26+Distributed+Systems%22">IEEE Transactions on Parallel & Distributed Systems</searchLink>. Jun2017, Vol. 28 Issue 6, p1728-1739. 12p.
– Name: Subject
  Label: Subjects
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Reliability+engineering+software%22">Reliability engineering software</searchLink><br /><searchLink fieldCode="DE" term="%22Computer+system+failures%22">Computer system failures</searchLink><br /><searchLink fieldCode="DE" term="%22Distributed+computing%22">Distributed computing</searchLink>
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: In order to guarantee data reliability, erasure codes have been used in distributed storage systems. Nevertheless, this mechanism suffers from the repair problem that excess data are needed to repair a single failure, causing both high bandwidth consuming for the network and heavy computing load on the replacement node. To reduce repair traffic, researchers pointed out the tradeoff between storage and repair traffic and proposed regenerating codes by combining network coding. However, the combination only focuses on the storage terminal and the construction of the codes is quite complicated. Therefore, this paper further combines network coding with network structure and proposes a repair tree model based on general erasure codes to simplify the repair procedure. By decomposing repair computing and distributing it among the tree nodes, our model can mitigate the computing tension. The performance of repair tree is analyzed and evaluated by preliminary emulation. The result shows it can make about three times faster computing than conventional measure and the repair throughput is doubled if there are network bottlenecks. For proper topology, it can significantly reduce the repair traffic. We present algorithms to generate trees across the network topology. At last, we present the idea of extending repair tree to repair multiple failures. [ABSTRACT FROM AUTHOR]
– Name: AbstractSuppliedCopyright
  Label:
  Group: Ab
  Data: <i>Copyright of IEEE Transactions on Parallel & Distributed Systems is the property of IEEE and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.)
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=egs&AN=123209380
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1109/TPDS.2016.2628024
    Languages:
      – Code: eng
        Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 12
        StartPage: 1728
    Subjects:
      – SubjectFull: Reliability engineering software
        Type: general
      – SubjectFull: Computer system failures
        Type: general
      – SubjectFull: Distributed computing
        Type: general
    Titles:
      – TitleFull: Repair Tree: Fast Repair for Single Failure in Erasure-Coded Distributed Storage Systems.
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Zhang, Huayu
      – PersonEntity:
          Name:
            NameFull: Li, Hui
      – PersonEntity:
          Name:
            NameFull: Li, Shuo-Yen Robert
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 06
              Text: Jun2017
              Type: published
              Y: 2017
          Identifiers:
            – Type: issn-print
              Value: 10459219
          Numbering:
            – Type: volume
              Value: 28
            – Type: issue
              Value: 6
          Titles:
            – TitleFull: IEEE Transactions on Parallel & Distributed Systems
              Type: main
ResultId 1