O⁴-DNN: A Hybrid DSP-LUT-Based Processing Unit With Operation Packing and Out-of-Order Execution for Efficient Realization of Convolutional Neural Networks on FPGA Devices.

Saved in:
Bibliographic Details
Title: O⁴-DNN: A Hybrid DSP-LUT-Based Processing Unit With Operation Packing and Out-of-Order Execution for Efficient Realization of Convolutional Neural Networks on FPGA Devices.
Authors: Haghi, Pouya1 (AUTHOR) pouya.haghi@ut.ac.ir, Kamal, Mehdi1 (AUTHOR) mehdikamal@ut.ac.ir, Afzali-Kusha, Ali1 (AUTHOR) afzali@ut.ac.ir, Pedram, Massoud2 (AUTHOR) pedram@usc.edu
Source: IEEE Transactions on Circuits & Systems. Part I: Regular Papers. Sep2020, Vol. 67 Issue 9, p3056-3069. 14p.
Subjects: Convolutional neural networks, Field programmable gate arrays, Artificial neural networks, Energy consumption
Abstract: In this paper, we propose O4-DNN, a high-performance FPGA-based architecture for convolutional neural network (CNN) accelerators relying on operation packing and out-of-order (OoO) execution for DSP blocks augmented with LUT-based glue logic. The high-level architecture is comprised of a systolic array of processing elements (PEs), supporting output stationary dataflow. In this architecture, the computational unit of each PE is realized by using a DSP block as well as a small number of LUTs. Given the limited number of DSP blocks in FPGAs, the combination (DSP block and some LUTs) provides more computational power obtainable through each DSP block. The proposed computational unit performs eight convolutional operations on five input operands where one of them is an 8-bit weight and the others are four 8-bit input feature (IF) maps. In addition, to improve the energy efficiency of the proposed computational unit, we present an approximate form of the unit suitable for neural network applications. To reduce the memory bandwidth as well as increase the utilization of the computational units, a data reusing technique based on the weight sharing is also presented. To improve the performance of the proposed computational unit further, an addressing approach for computing the partial sums out-of-order is proposed. The efficacy of the architecture is assessed using two FPGA devices executing four state-of-the-art neural networks. Experimental results show that this architecture leads to, on average (up to), $2.5\times $ ($3.44\times$) higher throughput compared to a baseline structure. In addition, on average (maximum of), 12% (40%) energy efficiency improvement is achievable by employing the O4-DNN compared to the baseline structure. [ABSTRACT FROM AUTHOR]
Copyright of IEEE Transactions on Circuits & Systems. Part I: Regular Papers is the property of IEEE and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Engineering Source
FullText Text:
  Availability: 0
Header DbId: egs
DbLabel: Engineering Source
An: 145399757
AccessLevel: 6
PubType: Periodical
PubTypeId: serialPeriodical
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: O⁴-DNN: A Hybrid DSP-LUT-Based Processing Unit With Operation Packing and Out-of-Order Execution for Efficient Realization of Convolutional Neural Networks on FPGA Devices.
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Haghi%2C+Pouya%22">Haghi, Pouya</searchLink><relatesTo>1</relatesTo> (AUTHOR)<i> pouya.haghi@ut.ac.ir</i><br /><searchLink fieldCode="AR" term="%22Kamal%2C+Mehdi%22">Kamal, Mehdi</searchLink><relatesTo>1</relatesTo> (AUTHOR)<i> mehdikamal@ut.ac.ir</i><br /><searchLink fieldCode="AR" term="%22Afzali-Kusha%2C+Ali%22">Afzali-Kusha, Ali</searchLink><relatesTo>1</relatesTo> (AUTHOR)<i> afzali@ut.ac.ir</i><br /><searchLink fieldCode="AR" term="%22Pedram%2C+Massoud%22">Pedram, Massoud</searchLink><relatesTo>2</relatesTo> (AUTHOR)<i> pedram@usc.edu</i>
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="JN" term="%22IEEE+Transactions+on+Circuits+%26+Systems%2E+Part+I%3A+Regular+Papers%22">IEEE Transactions on Circuits & Systems. Part I: Regular Papers</searchLink>. Sep2020, Vol. 67 Issue 9, p3056-3069. 14p.
– Name: Subject
  Label: Subjects
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Convolutional+neural+networks%22">Convolutional neural networks</searchLink><br /><searchLink fieldCode="DE" term="%22Field+programmable+gate+arrays%22">Field programmable gate arrays</searchLink><br /><searchLink fieldCode="DE" term="%22Artificial+neural+networks%22">Artificial neural networks</searchLink><br /><searchLink fieldCode="DE" term="%22Energy+consumption%22">Energy consumption</searchLink>
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: In this paper, we propose O4-DNN, a high-performance FPGA-based architecture for convolutional neural network (CNN) accelerators relying on operation packing and out-of-order (OoO) execution for DSP blocks augmented with LUT-based glue logic. The high-level architecture is comprised of a systolic array of processing elements (PEs), supporting output stationary dataflow. In this architecture, the computational unit of each PE is realized by using a DSP block as well as a small number of LUTs. Given the limited number of DSP blocks in FPGAs, the combination (DSP block and some LUTs) provides more computational power obtainable through each DSP block. The proposed computational unit performs eight convolutional operations on five input operands where one of them is an 8-bit weight and the others are four 8-bit input feature (IF) maps. In addition, to improve the energy efficiency of the proposed computational unit, we present an approximate form of the unit suitable for neural network applications. To reduce the memory bandwidth as well as increase the utilization of the computational units, a data reusing technique based on the weight sharing is also presented. To improve the performance of the proposed computational unit further, an addressing approach for computing the partial sums out-of-order is proposed. The efficacy of the architecture is assessed using two FPGA devices executing four state-of-the-art neural networks. Experimental results show that this architecture leads to, on average (up to), $2.5\times $ ($3.44\times$) higher throughput compared to a baseline structure. In addition, on average (maximum of), 12% (40%) energy efficiency improvement is achievable by employing the O4-DNN compared to the baseline structure. [ABSTRACT FROM AUTHOR]
– Name: AbstractSuppliedCopyright
  Label:
  Group: Ab
  Data: <i>Copyright of IEEE Transactions on Circuits & Systems. Part I: Regular Papers is the property of IEEE and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.)
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=egs&AN=145399757
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1109/TCSI.2020.2986350
    Languages:
      – Code: eng
        Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 14
        StartPage: 3056
    Subjects:
      – SubjectFull: Convolutional neural networks
        Type: general
      – SubjectFull: Field programmable gate arrays
        Type: general
      – SubjectFull: Artificial neural networks
        Type: general
      – SubjectFull: Energy consumption
        Type: general
    Titles:
      – TitleFull: O⁴-DNN: A Hybrid DSP-LUT-Based Processing Unit With Operation Packing and Out-of-Order Execution for Efficient Realization of Convolutional Neural Networks on FPGA Devices.
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Haghi, Pouya
      – PersonEntity:
          Name:
            NameFull: Kamal, Mehdi
      – PersonEntity:
          Name:
            NameFull: Afzali-Kusha, Ali
      – PersonEntity:
          Name:
            NameFull: Pedram, Massoud
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 09
              Text: Sep2020
              Type: published
              Y: 2020
          Identifiers:
            – Type: issn-print
              Value: 15498328
          Numbering:
            – Type: volume
              Value: 67
            – Type: issue
              Value: 9
          Titles:
            – TitleFull: IEEE Transactions on Circuits & Systems. Part I: Regular Papers
              Type: main
ResultId 1