A Bregman Learning Framework for Sparse Neural Networks.

Saved in:
Bibliographic Details
Title: A Bregman Learning Framework for Sparse Neural Networks.
Authors: Bungert, Leon1 LEON.BUNGERT@HCM.UNI-BONN.DE, Roith, Tim2 TIM.ROITH@FAU.DE, Tenbrinck, Daniel2 DANIEL.TENBRINCK@FAU.DE, Burger, Martin2 MARTIN.BURGER@FAU.DE
Source: Journal of Machine Learning Research. 2022, Vol. 23, p1-43. 43p.
Subjects: Artificial neural networks, Stochastic convergence, Stochastic analysis
Abstract: We propose a learning framework based on stochastic Bregman iterations, also known as mirror descent, to train sparse neural networks with an inverse scale space approach. We derive a baseline algorithm called LinBreg, an accelerated version using momentum, and AdaBreg, which is a Bregmanized generalization of the Adam algorithm. In contrast to established methods for sparse training the proposed family of algorithms constitutes a regrowth strategy for neural networks that is solely optimization-based without additional heuristics. Our Bregman learning framework starts the training with very few initial parameters, successively adding only significant ones to obtain a sparse and expressive network. The proposed approach is extremely easy and effcient, yet supported by the rich mathematical theory of inverse scale space methods. We derive a statistically profound sparse parameter initialization strategy and provide a rigorous stochastic convergence analysis of the loss decay and additional convergence proofs in the convex regime. Using only 3:4% of the parameters of ResNet-18 we achieve 90:2% test accuracy on CIFAR-10, compared to 93:6% using the dense network. Our algorithm also unveils an autoencoder architecture for a denoising task. The proposed framework also has a huge potential for integrating sparse backpropagation and resource-friendly training. Code is available at https://github.com/TimRoith/BregmanLearning. [ABSTRACT FROM AUTHOR]
Copyright of Journal of Machine Learning Research is the property of Microtome Publishing and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Engineering Source
FullText Text:
  Availability: 0
Header DbId: egs
DbLabel: Engineering Source
An: 164775316
AccessLevel: 6
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: A Bregman Learning Framework for Sparse Neural Networks.
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Bungert%2C+Leon%22">Bungert, Leon</searchLink><relatesTo>1</relatesTo><i> LEON.BUNGERT@HCM.UNI-BONN.DE</i><br /><searchLink fieldCode="AR" term="%22Roith%2C+Tim%22">Roith, Tim</searchLink><relatesTo>2</relatesTo><i> TIM.ROITH@FAU.DE</i><br /><searchLink fieldCode="AR" term="%22Tenbrinck%2C+Daniel%22">Tenbrinck, Daniel</searchLink><relatesTo>2</relatesTo><i> DANIEL.TENBRINCK@FAU.DE</i><br /><searchLink fieldCode="AR" term="%22Burger%2C+Martin%22">Burger, Martin</searchLink><relatesTo>2</relatesTo><i> MARTIN.BURGER@FAU.DE</i>
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="JN" term="%22Journal+of+Machine+Learning+Research%22">Journal of Machine Learning Research</searchLink>. 2022, Vol. 23, p1-43. 43p.
– Name: Subject
  Label: Subjects
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Artificial+neural+networks%22">Artificial neural networks</searchLink><br /><searchLink fieldCode="DE" term="%22Stochastic+convergence%22">Stochastic convergence</searchLink><br /><searchLink fieldCode="DE" term="%22Stochastic+analysis%22">Stochastic analysis</searchLink>
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: We propose a learning framework based on stochastic Bregman iterations, also known as mirror descent, to train sparse neural networks with an inverse scale space approach. We derive a baseline algorithm called LinBreg, an accelerated version using momentum, and AdaBreg, which is a Bregmanized generalization of the Adam algorithm. In contrast to established methods for sparse training the proposed family of algorithms constitutes a regrowth strategy for neural networks that is solely optimization-based without additional heuristics. Our Bregman learning framework starts the training with very few initial parameters, successively adding only significant ones to obtain a sparse and expressive network. The proposed approach is extremely easy and effcient, yet supported by the rich mathematical theory of inverse scale space methods. We derive a statistically profound sparse parameter initialization strategy and provide a rigorous stochastic convergence analysis of the loss decay and additional convergence proofs in the convex regime. Using only 3:4% of the parameters of ResNet-18 we achieve 90:2% test accuracy on CIFAR-10, compared to 93:6% using the dense network. Our algorithm also unveils an autoencoder architecture for a denoising task. The proposed framework also has a huge potential for integrating sparse backpropagation and resource-friendly training. Code is available at https://github.com/TimRoith/BregmanLearning. [ABSTRACT FROM AUTHOR]
– Name: AbstractSuppliedCopyright
  Label:
  Group: Ab
  Data: <i>Copyright of Journal of Machine Learning Research is the property of Microtome Publishing and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.)
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=egs&AN=164775316
RecordInfo BibRecord:
  BibEntity:
    Languages:
      – Code: eng
        Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 43
        StartPage: 1
    Subjects:
      – SubjectFull: Artificial neural networks
        Type: general
      – SubjectFull: Stochastic convergence
        Type: general
      – SubjectFull: Stochastic analysis
        Type: general
    Titles:
      – TitleFull: A Bregman Learning Framework for Sparse Neural Networks.
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Bungert, Leon
      – PersonEntity:
          Name:
            NameFull: Roith, Tim
      – PersonEntity:
          Name:
            NameFull: Tenbrinck, Daniel
      – PersonEntity:
          Name:
            NameFull: Burger, Martin
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 01
              Text: 2022
              Type: published
              Y: 2022
          Identifiers:
            – Type: issn-print
              Value: 15324435
          Numbering:
            – Type: volume
              Value: 23
          Titles:
            – TitleFull: Journal of Machine Learning Research
              Type: main
ResultId 1