Multi-distribution noise quantisation: an extreme compression scheme for transformer according to parameter distribution.

Saved in:
Bibliographic Details
Title: Multi-distribution noise quantisation: an extreme compression scheme for transformer according to parameter distribution.
Authors: Yu, Zaiyang (AUTHOR), Li, Shuang (AUTHOR), Sun, Linjun (AUTHOR), Liu, Liang (AUTHOR), Haining, Wang (AUTHOR)
Source: Connection Science. Dec2022, Vol. 34 Issue 1, p990-1004. 15p.
Subjects: Deep learning, Natural language processing, Parameters (Statistics), Noise, Khat
Abstract: With the development of deep learning, neural networks are widely used in various fields, and the improved model performance also introduces a considerable number of parameters and computations. Model quantisation is a technique that turns floating-point computing into low-specific-point computing, which can effectively reduce model computation strength, parameter size, and memory consumption but often bring a considerable loss of accuracy. This paper mainly addresses the problem where the distribution of parameters is too concentrated during quantisation aware training (QAT). In the QAT process, we use a piecewise function to statistics the parameter distributions and simulate the effect of quantisation noise in each round of training, based on the statistical results. Experimental results show that by quantising the Transformer network, we lose less precision and significantly reduce the storage cost of the model; compared with the full precision LSTM network, our model has higher accuracy under the condition of a similar storage cost. Meanwhile, compared with other quantisation methods on language modelling task, our approach is more accurate. We validated the effectiveness of our policy on the WikiText-103 and PENN Treebank datasets. The experiments show that our method extremely compresses the storage cost and maintains high model performance. [ABSTRACT FROM AUTHOR]
Copyright of Connection Science is the property of Taylor & Francis Ltd and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Psychology and Behavioral Sciences Collection
Full text is not displayed to guests.
FullText Links:
  – Type: pdflink
Text:
  Availability: 1
Header DbId: pbh
DbLabel: Psychology and Behavioral Sciences Collection
An: 164286344
AccessLevel: 6
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: Multi-distribution noise quantisation: an extreme compression scheme for transformer according to parameter distribution.
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Yu%2C+Zaiyang%22">Yu, Zaiyang</searchLink> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Li%2C+Shuang%22">Li, Shuang</searchLink> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Sun%2C+Linjun%22">Sun, Linjun</searchLink> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Liu%2C+Liang%22">Liu, Liang</searchLink> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Haining%2C+Wang%22">Haining, Wang</searchLink> (AUTHOR)
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="JN" term="%22Connection+Science%22">Connection Science</searchLink>. Dec2022, Vol. 34 Issue 1, p990-1004. 15p.
– Name: Subject
  Label: Subjects
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Deep+learning%22">Deep learning</searchLink><br /><searchLink fieldCode="DE" term="%22Natural+language+processing%22">Natural language processing</searchLink><br /><searchLink fieldCode="DE" term="%22Parameters+%28Statistics%29%22">Parameters (Statistics)</searchLink><br /><searchLink fieldCode="DE" term="%22Noise%22">Noise</searchLink><br /><searchLink fieldCode="DE" term="%22Khat%22">Khat</searchLink>
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: With the development of deep learning, neural networks are widely used in various fields, and the improved model performance also introduces a considerable number of parameters and computations. Model quantisation is a technique that turns floating-point computing into low-specific-point computing, which can effectively reduce model computation strength, parameter size, and memory consumption but often bring a considerable loss of accuracy. This paper mainly addresses the problem where the distribution of parameters is too concentrated during quantisation aware training (QAT). In the QAT process, we use a piecewise function to statistics the parameter distributions and simulate the effect of quantisation noise in each round of training, based on the statistical results. Experimental results show that by quantising the Transformer network, we lose less precision and significantly reduce the storage cost of the model; compared with the full precision LSTM network, our model has higher accuracy under the condition of a similar storage cost. Meanwhile, compared with other quantisation methods on language modelling task, our approach is more accurate. We validated the effectiveness of our policy on the WikiText-103 and PENN Treebank datasets. The experiments show that our method extremely compresses the storage cost and maintains high model performance. [ABSTRACT FROM AUTHOR]
– Name: AbstractSuppliedCopyright
  Label:
  Group: Ab
  Data: <i>Copyright of Connection Science is the property of Taylor & Francis Ltd and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.)
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=pbh&AN=164286344
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1080/09540091.2021.2024510
    Languages:
      – Code: eng
        Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 15
        StartPage: 990
    Subjects:
      – SubjectFull: Deep learning
        Type: general
      – SubjectFull: Natural language processing
        Type: general
      – SubjectFull: Parameters (Statistics)
        Type: general
      – SubjectFull: Noise
        Type: general
      – SubjectFull: Khat
        Type: general
    Titles:
      – TitleFull: Multi-distribution noise quantisation: an extreme compression scheme for transformer according to parameter distribution.
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Yu, Zaiyang
      – PersonEntity:
          Name:
            NameFull: Li, Shuang
      – PersonEntity:
          Name:
            NameFull: Sun, Linjun
      – PersonEntity:
          Name:
            NameFull: Liu, Liang
      – PersonEntity:
          Name:
            NameFull: Haining, Wang
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 12
              Text: Dec2022
              Type: published
              Y: 2022
          Identifiers:
            – Type: issn-print
              Value: 09540091
          Numbering:
            – Type: volume
              Value: 34
            – Type: issue
              Value: 1
          Titles:
            – TitleFull: Connection Science
              Type: main
ResultId 1