Memory-accelerated parallel method for multidimensional fast fourier implementation on GPU.
Saved in:
| Title: | Memory-accelerated parallel method for multidimensional fast fourier implementation on GPU. |
|---|---|
| Authors: | Hu, Yichang1 (AUTHOR), Lu, Lu1 (AUTHOR) lul@scut.edu.cn, Li, Cuixu2 (AUTHOR) |
| Source: | Journal of Supercomputing. Nov2022, Vol. 78 Issue 16, p18189-18208. 20p. |
| Subjects: | Discrete Fourier transforms, Multidimensional databases, Fast Fourier transforms, Graphics processing units, Patternmaking |
| Abstract: | Fast Fourier transform (FFT) is a well-known algorithm that calculates the discrete Fourier transform (DFT) of discrete data and is an essential tool in scientific and engineering computation. Due to the large amounts of data, parallelly executing FFT in graphics processing unit (GPU) can effectively optimize the performance. Following this approach, FFTW and some other FFT packages were designed, but the fixed computation pattern makes it hard to utilize the computing power of GPU. Additionally, the memory access pattern is not optimized to alleviate the bottleneck of data exchange. Motivated by these challenges, we propose an efficient GPU-accelerated multidimensional FFT library to achieve better performance in this paper. We present a detailed and clear implementation strategy and optimize FFT by having as few memory transfers as possible. The data will be reshuffled on the CPU, and the access mode is also optimized to coordinate with the GPU memory access pattern. Several optimizations are also demonstrated to enhance the performance of our approach for varying FFT sizes, and the evaluation shows that our approach consistently outperforms rocFFT with a speedup of about 25% to 250% on average in AMD Instinct MI100 GPU. [ABSTRACT FROM AUTHOR] |
| Copyright of Journal of Supercomputing is the property of Springer Nature and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.) | |
| Database: | Engineering Source |
|
Full text is not displayed to guests.
Login for full access.
|
|
| FullText | Links: – Type: pdflink Text: Availability: 1 |
|---|---|
| Header | DbId: egs DbLabel: Engineering Source An: 159685729 AccessLevel: 6 PubType: Academic Journal PubTypeId: academicJournal PreciseRelevancyScore: 0 |
| IllustrationInfo | |
| Items | – Name: Title Label: Title Group: Ti Data: Memory-accelerated parallel method for multidimensional fast fourier implementation on GPU. – Name: Author Label: Authors Group: Au Data: <searchLink fieldCode="AR" term="%22Hu%2C+Yichang%22">Hu, Yichang</searchLink><relatesTo>1</relatesTo> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Lu%2C+Lu%22">Lu, Lu</searchLink><relatesTo>1</relatesTo> (AUTHOR)<i> lul@scut.edu.cn</i><br /><searchLink fieldCode="AR" term="%22Li%2C+Cuixu%22">Li, Cuixu</searchLink><relatesTo>2</relatesTo> (AUTHOR) – Name: TitleSource Label: Source Group: Src Data: <searchLink fieldCode="JN" term="%22Journal+of+Supercomputing%22">Journal of Supercomputing</searchLink>. Nov2022, Vol. 78 Issue 16, p18189-18208. 20p. – Name: Subject Label: Subjects Group: Su Data: <searchLink fieldCode="DE" term="%22Discrete+Fourier+transforms%22">Discrete Fourier transforms</searchLink><br /><searchLink fieldCode="DE" term="%22Multidimensional+databases%22">Multidimensional databases</searchLink><br /><searchLink fieldCode="DE" term="%22Fast+Fourier+transforms%22">Fast Fourier transforms</searchLink><br /><searchLink fieldCode="DE" term="%22Graphics+processing+units%22">Graphics processing units</searchLink><br /><searchLink fieldCode="DE" term="%22Patternmaking%22">Patternmaking</searchLink> – Name: Abstract Label: Abstract Group: Ab Data: Fast Fourier transform (FFT) is a well-known algorithm that calculates the discrete Fourier transform (DFT) of discrete data and is an essential tool in scientific and engineering computation. Due to the large amounts of data, parallelly executing FFT in graphics processing unit (GPU) can effectively optimize the performance. Following this approach, FFTW and some other FFT packages were designed, but the fixed computation pattern makes it hard to utilize the computing power of GPU. Additionally, the memory access pattern is not optimized to alleviate the bottleneck of data exchange. Motivated by these challenges, we propose an efficient GPU-accelerated multidimensional FFT library to achieve better performance in this paper. We present a detailed and clear implementation strategy and optimize FFT by having as few memory transfers as possible. The data will be reshuffled on the CPU, and the access mode is also optimized to coordinate with the GPU memory access pattern. Several optimizations are also demonstrated to enhance the performance of our approach for varying FFT sizes, and the evaluation shows that our approach consistently outperforms rocFFT with a speedup of about 25% to 250% on average in AMD Instinct MI100 GPU. [ABSTRACT FROM AUTHOR] – Name: AbstractSuppliedCopyright Label: Group: Ab Data: <i>Copyright of Journal of Supercomputing is the property of Springer Nature and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.) |
| PLink | https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=egs&AN=159685729 |
| RecordInfo | BibRecord: BibEntity: Identifiers: – Type: doi Value: 10.1007/s11227-022-04570-9 Languages: – Code: eng Text: English PhysicalDescription: Pagination: PageCount: 20 StartPage: 18189 Subjects: – SubjectFull: Discrete Fourier transforms Type: general – SubjectFull: Multidimensional databases Type: general – SubjectFull: Fast Fourier transforms Type: general – SubjectFull: Graphics processing units Type: general – SubjectFull: Patternmaking Type: general Titles: – TitleFull: Memory-accelerated parallel method for multidimensional fast fourier implementation on GPU. Type: main BibRelationships: HasContributorRelationships: – PersonEntity: Name: NameFull: Hu, Yichang – PersonEntity: Name: NameFull: Lu, Lu – PersonEntity: Name: NameFull: Li, Cuixu IsPartOfRelationships: – BibEntity: Dates: – D: 01 M: 11 Text: Nov2022 Type: published Y: 2022 Identifiers: – Type: issn-print Value: 09208542 Numbering: – Type: volume Value: 78 – Type: issue Value: 16 Titles: – TitleFull: Journal of Supercomputing Type: main |
| ResultId | 1 |