Generalisation of recursive doubling for AllReduce: Now with simulation.

Saved in:
Bibliographic Details
Title: Generalisation of recursive doubling for AllReduce: Now with simulation.
Authors: Ruefenacht, Martin1 m.ruefenacht@ed.ac.uk, Bull, Mark1 m.bull@ed.ac.uk, Booth, Stephen1 s.booth@ed.ac.uk
Source: Parallel Computing. Nov2017, Vol. 69, p24-44. 21p.
Subjects: Computer simulation, Cray computers, Data pipelining, Computer software execution, Recursive functions
Abstract: The performance of AllReduce is crucial at scale. The recursive doubling with pairwise exchange algorithm theoretically achieves O (log 2   N ) scaling for short messages with N peers, but is limited by improvements in network latency. A multi-way exchange can be implemented using message pipelining, which is easier to improve than latency. Using our method, recursive multiplying, we show reductions in execution time of between 8% and 40% of AllReduce on a Cray XC30 over recursive doubling. Using a custom simulator we further explore the dynamics of recursive multiplying. [ABSTRACT FROM AUTHOR]
Copyright of Parallel Computing is the property of Elsevier B.V. and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Engineering Source
Description
Abstract:The performance of AllReduce is crucial at scale. The recursive doubling with pairwise exchange algorithm theoretically achieves O (log 2   N ) scaling for short messages with N peers, but is limited by improvements in network latency. A multi-way exchange can be implemented using message pipelining, which is easier to improve than latency. Using our method, recursive multiplying, we show reductions in execution time of between 8% and 40% of AllReduce on a Cray XC30 over recursive doubling. Using a custom simulator we further explore the dynamics of recursive multiplying. [ABSTRACT FROM AUTHOR]
ISSN:01678191
DOI:10.1016/j.parco.2017.08.004