Memory-accelerated parallel method for multidimensional fast fourier implementation on GPU.

Saved in:
Bibliographic Details
Title: Memory-accelerated parallel method for multidimensional fast fourier implementation on GPU.
Authors: Hu, Yichang1 (AUTHOR), Lu, Lu1 (AUTHOR) lul@scut.edu.cn, Li, Cuixu2 (AUTHOR)
Source: Journal of Supercomputing. Nov2022, Vol. 78 Issue 16, p18189-18208. 20p.
Subjects: Discrete Fourier transforms, Multidimensional databases, Fast Fourier transforms, Graphics processing units, Patternmaking
Abstract: Fast Fourier transform (FFT) is a well-known algorithm that calculates the discrete Fourier transform (DFT) of discrete data and is an essential tool in scientific and engineering computation. Due to the large amounts of data, parallelly executing FFT in graphics processing unit (GPU) can effectively optimize the performance. Following this approach, FFTW and some other FFT packages were designed, but the fixed computation pattern makes it hard to utilize the computing power of GPU. Additionally, the memory access pattern is not optimized to alleviate the bottleneck of data exchange. Motivated by these challenges, we propose an efficient GPU-accelerated multidimensional FFT library to achieve better performance in this paper. We present a detailed and clear implementation strategy and optimize FFT by having as few memory transfers as possible. The data will be reshuffled on the CPU, and the access mode is also optimized to coordinate with the GPU memory access pattern. Several optimizations are also demonstrated to enhance the performance of our approach for varying FFT sizes, and the evaluation shows that our approach consistently outperforms rocFFT with a speedup of about 25% to 250% on average in AMD Instinct MI100 GPU. [ABSTRACT FROM AUTHOR]
Copyright of Journal of Supercomputing is the property of Springer Nature and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Engineering Source
Full text is not displayed to guests.
FullText Links:
  – Type: pdflink
Text:
  Availability: 1
Header DbId: egs
DbLabel: Engineering Source
An: 159685729
AccessLevel: 6
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: Memory-accelerated parallel method for multidimensional fast fourier implementation on GPU.
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Hu%2C+Yichang%22">Hu, Yichang</searchLink><relatesTo>1</relatesTo> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Lu%2C+Lu%22">Lu, Lu</searchLink><relatesTo>1</relatesTo> (AUTHOR)<i> lul@scut.edu.cn</i><br /><searchLink fieldCode="AR" term="%22Li%2C+Cuixu%22">Li, Cuixu</searchLink><relatesTo>2</relatesTo> (AUTHOR)
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="JN" term="%22Journal+of+Supercomputing%22">Journal of Supercomputing</searchLink>. Nov2022, Vol. 78 Issue 16, p18189-18208. 20p.
– Name: Subject
  Label: Subjects
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Discrete+Fourier+transforms%22">Discrete Fourier transforms</searchLink><br /><searchLink fieldCode="DE" term="%22Multidimensional+databases%22">Multidimensional databases</searchLink><br /><searchLink fieldCode="DE" term="%22Fast+Fourier+transforms%22">Fast Fourier transforms</searchLink><br /><searchLink fieldCode="DE" term="%22Graphics+processing+units%22">Graphics processing units</searchLink><br /><searchLink fieldCode="DE" term="%22Patternmaking%22">Patternmaking</searchLink>
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: Fast Fourier transform (FFT) is a well-known algorithm that calculates the discrete Fourier transform (DFT) of discrete data and is an essential tool in scientific and engineering computation. Due to the large amounts of data, parallelly executing FFT in graphics processing unit (GPU) can effectively optimize the performance. Following this approach, FFTW and some other FFT packages were designed, but the fixed computation pattern makes it hard to utilize the computing power of GPU. Additionally, the memory access pattern is not optimized to alleviate the bottleneck of data exchange. Motivated by these challenges, we propose an efficient GPU-accelerated multidimensional FFT library to achieve better performance in this paper. We present a detailed and clear implementation strategy and optimize FFT by having as few memory transfers as possible. The data will be reshuffled on the CPU, and the access mode is also optimized to coordinate with the GPU memory access pattern. Several optimizations are also demonstrated to enhance the performance of our approach for varying FFT sizes, and the evaluation shows that our approach consistently outperforms rocFFT with a speedup of about 25% to 250% on average in AMD Instinct MI100 GPU. [ABSTRACT FROM AUTHOR]
– Name: AbstractSuppliedCopyright
  Label:
  Group: Ab
  Data: <i>Copyright of Journal of Supercomputing is the property of Springer Nature and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.)
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=egs&AN=159685729
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1007/s11227-022-04570-9
    Languages:
      – Code: eng
        Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 20
        StartPage: 18189
    Subjects:
      – SubjectFull: Discrete Fourier transforms
        Type: general
      – SubjectFull: Multidimensional databases
        Type: general
      – SubjectFull: Fast Fourier transforms
        Type: general
      – SubjectFull: Graphics processing units
        Type: general
      – SubjectFull: Patternmaking
        Type: general
    Titles:
      – TitleFull: Memory-accelerated parallel method for multidimensional fast fourier implementation on GPU.
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Hu, Yichang
      – PersonEntity:
          Name:
            NameFull: Lu, Lu
      – PersonEntity:
          Name:
            NameFull: Li, Cuixu
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 11
              Text: Nov2022
              Type: published
              Y: 2022
          Identifiers:
            – Type: issn-print
              Value: 09208542
          Numbering:
            – Type: volume
              Value: 78
            – Type: issue
              Value: 16
          Titles:
            – TitleFull: Journal of Supercomputing
              Type: main
ResultId 1