Multi-distribution noise quantisation: an extreme compression scheme for transformer according to parameter distribution.
Saved in:
| Title: | Multi-distribution noise quantisation: an extreme compression scheme for transformer according to parameter distribution. |
|---|---|
| Authors: | Yu, Zaiyang (AUTHOR), Li, Shuang (AUTHOR), Sun, Linjun (AUTHOR), Liu, Liang (AUTHOR), Haining, Wang (AUTHOR) |
| Source: | Connection Science. Dec2022, Vol. 34 Issue 1, p990-1004. 15p. |
| Subjects: | Deep learning, Natural language processing, Parameters (Statistics), Noise, Khat |
| Abstract: | With the development of deep learning, neural networks are widely used in various fields, and the improved model performance also introduces a considerable number of parameters and computations. Model quantisation is a technique that turns floating-point computing into low-specific-point computing, which can effectively reduce model computation strength, parameter size, and memory consumption but often bring a considerable loss of accuracy. This paper mainly addresses the problem where the distribution of parameters is too concentrated during quantisation aware training (QAT). In the QAT process, we use a piecewise function to statistics the parameter distributions and simulate the effect of quantisation noise in each round of training, based on the statistical results. Experimental results show that by quantising the Transformer network, we lose less precision and significantly reduce the storage cost of the model; compared with the full precision LSTM network, our model has higher accuracy under the condition of a similar storage cost. Meanwhile, compared with other quantisation methods on language modelling task, our approach is more accurate. We validated the effectiveness of our policy on the WikiText-103 and PENN Treebank datasets. The experiments show that our method extremely compresses the storage cost and maintains high model performance. [ABSTRACT FROM AUTHOR] |
| Copyright of Connection Science is the property of Taylor & Francis Ltd and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.) | |
| Database: | Psychology and Behavioral Sciences Collection |
|
Full text is not displayed to guests.
Login for full access.
|
|
| FullText | Links: – Type: pdflink Text: Availability: 1 |
|---|---|
| Header | DbId: pbh DbLabel: Psychology and Behavioral Sciences Collection An: 164286344 AccessLevel: 6 PubType: Academic Journal PubTypeId: academicJournal PreciseRelevancyScore: 0 |
| IllustrationInfo | |
| Items | – Name: Title Label: Title Group: Ti Data: Multi-distribution noise quantisation: an extreme compression scheme for transformer according to parameter distribution. – Name: Author Label: Authors Group: Au Data: <searchLink fieldCode="AR" term="%22Yu%2C+Zaiyang%22">Yu, Zaiyang</searchLink> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Li%2C+Shuang%22">Li, Shuang</searchLink> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Sun%2C+Linjun%22">Sun, Linjun</searchLink> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Liu%2C+Liang%22">Liu, Liang</searchLink> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Haining%2C+Wang%22">Haining, Wang</searchLink> (AUTHOR) – Name: TitleSource Label: Source Group: Src Data: <searchLink fieldCode="JN" term="%22Connection+Science%22">Connection Science</searchLink>. Dec2022, Vol. 34 Issue 1, p990-1004. 15p. – Name: Subject Label: Subjects Group: Su Data: <searchLink fieldCode="DE" term="%22Deep+learning%22">Deep learning</searchLink><br /><searchLink fieldCode="DE" term="%22Natural+language+processing%22">Natural language processing</searchLink><br /><searchLink fieldCode="DE" term="%22Parameters+%28Statistics%29%22">Parameters (Statistics)</searchLink><br /><searchLink fieldCode="DE" term="%22Noise%22">Noise</searchLink><br /><searchLink fieldCode="DE" term="%22Khat%22">Khat</searchLink> – Name: Abstract Label: Abstract Group: Ab Data: With the development of deep learning, neural networks are widely used in various fields, and the improved model performance also introduces a considerable number of parameters and computations. Model quantisation is a technique that turns floating-point computing into low-specific-point computing, which can effectively reduce model computation strength, parameter size, and memory consumption but often bring a considerable loss of accuracy. This paper mainly addresses the problem where the distribution of parameters is too concentrated during quantisation aware training (QAT). In the QAT process, we use a piecewise function to statistics the parameter distributions and simulate the effect of quantisation noise in each round of training, based on the statistical results. Experimental results show that by quantising the Transformer network, we lose less precision and significantly reduce the storage cost of the model; compared with the full precision LSTM network, our model has higher accuracy under the condition of a similar storage cost. Meanwhile, compared with other quantisation methods on language modelling task, our approach is more accurate. We validated the effectiveness of our policy on the WikiText-103 and PENN Treebank datasets. The experiments show that our method extremely compresses the storage cost and maintains high model performance. [ABSTRACT FROM AUTHOR] – Name: AbstractSuppliedCopyright Label: Group: Ab Data: <i>Copyright of Connection Science is the property of Taylor & Francis Ltd and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.) |
| PLink | https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=pbh&AN=164286344 |
| RecordInfo | BibRecord: BibEntity: Identifiers: – Type: doi Value: 10.1080/09540091.2021.2024510 Languages: – Code: eng Text: English PhysicalDescription: Pagination: PageCount: 15 StartPage: 990 Subjects: – SubjectFull: Deep learning Type: general – SubjectFull: Natural language processing Type: general – SubjectFull: Parameters (Statistics) Type: general – SubjectFull: Noise Type: general – SubjectFull: Khat Type: general Titles: – TitleFull: Multi-distribution noise quantisation: an extreme compression scheme for transformer according to parameter distribution. Type: main BibRelationships: HasContributorRelationships: – PersonEntity: Name: NameFull: Yu, Zaiyang – PersonEntity: Name: NameFull: Li, Shuang – PersonEntity: Name: NameFull: Sun, Linjun – PersonEntity: Name: NameFull: Liu, Liang – PersonEntity: Name: NameFull: Haining, Wang IsPartOfRelationships: – BibEntity: Dates: – D: 01 M: 12 Text: Dec2022 Type: published Y: 2022 Identifiers: – Type: issn-print Value: 09540091 Numbering: – Type: volume Value: 34 – Type: issue Value: 1 Titles: – TitleFull: Connection Science Type: main |
| ResultId | 1 |