A Bregman Learning Framework for Sparse Neural Networks.

Saved in:
Bibliographic Details
Title: A Bregman Learning Framework for Sparse Neural Networks.
Authors: Bungert, Leon1 LEON.BUNGERT@HCM.UNI-BONN.DE, Roith, Tim2 TIM.ROITH@FAU.DE, Tenbrinck, Daniel2 DANIEL.TENBRINCK@FAU.DE, Burger, Martin2 MARTIN.BURGER@FAU.DE
Source: Journal of Machine Learning Research. 2022, Vol. 23, p1-43. 43p.
Subjects: Artificial neural networks, Stochastic convergence, Stochastic analysis
Abstract: We propose a learning framework based on stochastic Bregman iterations, also known as mirror descent, to train sparse neural networks with an inverse scale space approach. We derive a baseline algorithm called LinBreg, an accelerated version using momentum, and AdaBreg, which is a Bregmanized generalization of the Adam algorithm. In contrast to established methods for sparse training the proposed family of algorithms constitutes a regrowth strategy for neural networks that is solely optimization-based without additional heuristics. Our Bregman learning framework starts the training with very few initial parameters, successively adding only significant ones to obtain a sparse and expressive network. The proposed approach is extremely easy and effcient, yet supported by the rich mathematical theory of inverse scale space methods. We derive a statistically profound sparse parameter initialization strategy and provide a rigorous stochastic convergence analysis of the loss decay and additional convergence proofs in the convex regime. Using only 3:4% of the parameters of ResNet-18 we achieve 90:2% test accuracy on CIFAR-10, compared to 93:6% using the dense network. Our algorithm also unveils an autoencoder architecture for a denoising task. The proposed framework also has a huge potential for integrating sparse backpropagation and resource-friendly training. Code is available at https://github.com/TimRoith/BregmanLearning. [ABSTRACT FROM AUTHOR]
Copyright of Journal of Machine Learning Research is the property of Microtome Publishing and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Engineering Source
Be the first to leave a comment!
You must be logged in first