Distributed Newton Methods for Deep Neural Networks.
Saved in:
| Title: | Distributed Newton Methods for Deep Neural Networks. |
|---|---|
| Authors: | Wang, Chien-Chih, Tan, Kent Loong, Chen, Chun-Ting, Lin, Yu-Hsiang, Keerthi, S. Sathiya, Mahajan, Dhruv, Sundararajan, S., Lin, Chih-Jen |
| Source: | Neural Computation. Jun2018, Vol. 30 Issue 6, p1673-1724. 52p. 2 Diagrams, 5 Charts, 2 Graphs. |
| Subjects: | Deep learning, Mathematical optimization, Machine learning, Artificial neural networks, Newton-Raphson method |
| Abstract: | Deep learning involves a difficult nonconvex optimization problem with a large number of weights between any two adjacent layers of a deep structure. To handle large data sets or complicated networks, distributed training is needed, but the calculation of function, gradient, and Hessian is expensive. In particular, the communication and the synchronization cost may become a bottleneck. In this letter, we focus on situations where the model is distributedly stored and propose a novel distributed Newton method for training deep neural networks. By variable and feature-wise data partitions and some careful designs, we are able to explicitly use the Jacobian matrix for matrix-vector products in the Newton method. Some techniques are incorporated to reduce the running time as well as memory consumption. First, to reduce the communication cost, we propose a diagonalization method such that an approximate Newton direction can be obtained without communication between machines. Second, we consider subsampled Gauss-Newton matrices for reducing the running time as well as the communication cost. Third, to reduce the synchronization cost, we terminate the process of finding an approximate Newton direction even though some nodes have not finished their tasks. Details of some implementation issues in distributed environments are thoroughly investigated. Experiments demonstrate that the proposed method is effective for the distributed training of deep neural networks. Compared with stochastic gradient methods, it is more robust and may give better test accuracy. [ABSTRACT FROM AUTHOR] |
| Copyright of Neural Computation is the property of MIT Press and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.) | |
| Database: | Psychology and Behavioral Sciences Collection |
|
Full text is not displayed to guests.
Login for full access.
|
|
| FullText | Links: – Type: pdflink Text: Availability: 1 |
|---|---|
| Header | DbId: pbh DbLabel: Psychology and Behavioral Sciences Collection An: 129770349 AccessLevel: 6 PubType: Academic Journal PubTypeId: academicJournal PreciseRelevancyScore: 0 |
| IllustrationInfo | |
| Items | – Name: Title Label: Title Group: Ti Data: Distributed Newton Methods for Deep Neural Networks. – Name: Author Label: Authors Group: Au Data: <searchLink fieldCode="AR" term="%22Wang%2C+Chien-Chih%22">Wang, Chien-Chih</searchLink><br /><searchLink fieldCode="AR" term="%22Tan%2C+Kent+Loong%22">Tan, Kent Loong</searchLink><br /><searchLink fieldCode="AR" term="%22Chen%2C+Chun-Ting%22">Chen, Chun-Ting</searchLink><br /><searchLink fieldCode="AR" term="%22Lin%2C+Yu-Hsiang%22">Lin, Yu-Hsiang</searchLink><br /><searchLink fieldCode="AR" term="%22Keerthi%2C+S%2E+Sathiya%22">Keerthi, S. Sathiya</searchLink><br /><searchLink fieldCode="AR" term="%22Mahajan%2C+Dhruv%22">Mahajan, Dhruv</searchLink><br /><searchLink fieldCode="AR" term="%22Sundararajan%2C+S%2E%22">Sundararajan, S.</searchLink><br /><searchLink fieldCode="AR" term="%22Lin%2C+Chih-Jen%22">Lin, Chih-Jen</searchLink> – Name: TitleSource Label: Source Group: Src Data: <searchLink fieldCode="JN" term="%22Neural+Computation%22">Neural Computation</searchLink>. Jun2018, Vol. 30 Issue 6, p1673-1724. 52p. 2 Diagrams, 5 Charts, 2 Graphs. – Name: Subject Label: Subjects Group: Su Data: <searchLink fieldCode="DE" term="%22Deep+learning%22">Deep learning</searchLink><br /><searchLink fieldCode="DE" term="%22Mathematical+optimization%22">Mathematical optimization</searchLink><br /><searchLink fieldCode="DE" term="%22Machine+learning%22">Machine learning</searchLink><br /><searchLink fieldCode="DE" term="%22Artificial+neural+networks%22">Artificial neural networks</searchLink><br /><searchLink fieldCode="DE" term="%22Newton-Raphson+method%22">Newton-Raphson method</searchLink> – Name: Abstract Label: Abstract Group: Ab Data: Deep learning involves a difficult nonconvex optimization problem with a large number of weights between any two adjacent layers of a deep structure. To handle large data sets or complicated networks, distributed training is needed, but the calculation of function, gradient, and Hessian is expensive. In particular, the communication and the synchronization cost may become a bottleneck. In this letter, we focus on situations where the model is distributedly stored and propose a novel distributed Newton method for training deep neural networks. By variable and feature-wise data partitions and some careful designs, we are able to explicitly use the Jacobian matrix for matrix-vector products in the Newton method. Some techniques are incorporated to reduce the running time as well as memory consumption. First, to reduce the communication cost, we propose a diagonalization method such that an approximate Newton direction can be obtained without communication between machines. Second, we consider subsampled Gauss-Newton matrices for reducing the running time as well as the communication cost. Third, to reduce the synchronization cost, we terminate the process of finding an approximate Newton direction even though some nodes have not finished their tasks. Details of some implementation issues in distributed environments are thoroughly investigated. Experiments demonstrate that the proposed method is effective for the distributed training of deep neural networks. Compared with stochastic gradient methods, it is more robust and may give better test accuracy. [ABSTRACT FROM AUTHOR] – Name: AbstractSuppliedCopyright Label: Group: Ab Data: <i>Copyright of Neural Computation is the property of MIT Press and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.) |
| PLink | https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=pbh&AN=129770349 |
| RecordInfo | BibRecord: BibEntity: Identifiers: – Type: doi Value: 10.1162/neco_a_01088 Languages: – Code: eng Text: English PhysicalDescription: Pagination: PageCount: 52 StartPage: 1673 Subjects: – SubjectFull: Deep learning Type: general – SubjectFull: Mathematical optimization Type: general – SubjectFull: Machine learning Type: general – SubjectFull: Artificial neural networks Type: general – SubjectFull: Newton-Raphson method Type: general Titles: – TitleFull: Distributed Newton Methods for Deep Neural Networks. Type: main BibRelationships: HasContributorRelationships: – PersonEntity: Name: NameFull: Wang, Chien-Chih – PersonEntity: Name: NameFull: Tan, Kent Loong – PersonEntity: Name: NameFull: Chen, Chun-Ting – PersonEntity: Name: NameFull: Lin, Yu-Hsiang – PersonEntity: Name: NameFull: Keerthi, S. Sathiya – PersonEntity: Name: NameFull: Mahajan, Dhruv – PersonEntity: Name: NameFull: Sundararajan, S. – PersonEntity: Name: NameFull: Lin, Chih-Jen IsPartOfRelationships: – BibEntity: Dates: – D: 01 M: 06 Text: Jun2018 Type: published Y: 2018 Identifiers: – Type: issn-print Value: 08997667 Numbering: – Type: volume Value: 30 – Type: issue Value: 6 Titles: – TitleFull: Neural Computation Type: main |
| ResultId | 1 |