A methodology for speeding up matrix vector multiplication for single/multi-core architectures.
Saved in:
| Title: | A methodology for speeding up matrix vector multiplication for single/multi-core architectures. |
|---|---|
| Authors: | Kelefouras, Vasilios1 kelefouras@ece.upatras.gr, Kritikakou, Angeliki2, Papadima, Elissavet1, Goutis, Costas1 |
| Source: | Journal of Supercomputing. Jul2015, Vol. 71 Issue 7, p2644-2667. 24p. |
| Subjects: | Complex multiplication, SIMD (Computer architecture), Multicore processors, Cache memory, Computer memory management |
| Abstract: | In this paper, a new methodology for computing the Dense Matrix Vector Multiplication, for both embedded (processors without SIMD unit) and general purpose processors (single and multi-core processors, with SIMD unit), is presented. This methodology achieves higher execution speed than ATLAS state-of-the-art library (speedup from 1.2 up to 1.45). This is achieved by fully exploiting the combination of the software (e.g., data reuse) and hardware parameters (e.g., data cache associativity) which are considered simultaneously as one problem and not separately, giving a smaller search space and high-quality solutions. The proposed methodology produces a different schedule for different values of the (i) number of the levels of data cache; (ii) data cache sizes; (iii) data cache associativities; (iv) data cache and main memory latencies; (v) data array layout of the matrix and (vi) number of cores. [ABSTRACT FROM AUTHOR] |
| Copyright of Journal of Supercomputing is the property of Springer Nature and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.) | |
| Database: | Engineering Source |
| FullText | Links: – Type: pdflink Text: Availability: 0 |
|---|---|
| Header | DbId: egs DbLabel: Engineering Source An: 103643988 AccessLevel: 6 PubType: Academic Journal PubTypeId: academicJournal PreciseRelevancyScore: 0 |
| IllustrationInfo | |
| Items | – Name: Title Label: Title Group: Ti Data: A methodology for speeding up matrix vector multiplication for single/multi-core architectures. – Name: Author Label: Authors Group: Au Data: <searchLink fieldCode="AR" term="%22Kelefouras%2C+Vasilios%22">Kelefouras, Vasilios</searchLink><relatesTo>1</relatesTo><i> kelefouras@ece.upatras.gr</i><br /><searchLink fieldCode="AR" term="%22Kritikakou%2C+Angeliki%22">Kritikakou, Angeliki</searchLink><relatesTo>2</relatesTo><br /><searchLink fieldCode="AR" term="%22Papadima%2C+Elissavet%22">Papadima, Elissavet</searchLink><relatesTo>1</relatesTo><br /><searchLink fieldCode="AR" term="%22Goutis%2C+Costas%22">Goutis, Costas</searchLink><relatesTo>1</relatesTo> – Name: TitleSource Label: Source Group: Src Data: <searchLink fieldCode="JN" term="%22Journal+of+Supercomputing%22">Journal of Supercomputing</searchLink>. Jul2015, Vol. 71 Issue 7, p2644-2667. 24p. – Name: Subject Label: Subjects Group: Su Data: <searchLink fieldCode="DE" term="%22Complex+multiplication%22">Complex multiplication</searchLink><br /><searchLink fieldCode="DE" term="%22SIMD+%28Computer+architecture%29%22">SIMD (Computer architecture)</searchLink><br /><searchLink fieldCode="DE" term="%22Multicore+processors%22">Multicore processors</searchLink><br /><searchLink fieldCode="DE" term="%22Cache+memory%22">Cache memory</searchLink><br /><searchLink fieldCode="DE" term="%22Computer+memory+management%22">Computer memory management</searchLink> – Name: Abstract Label: Abstract Group: Ab Data: In this paper, a new methodology for computing the Dense Matrix Vector Multiplication, for both embedded (processors without SIMD unit) and general purpose processors (single and multi-core processors, with SIMD unit), is presented. This methodology achieves higher execution speed than ATLAS state-of-the-art library (speedup from 1.2 up to 1.45). This is achieved by fully exploiting the combination of the software (e.g., data reuse) and hardware parameters (e.g., data cache associativity) which are considered simultaneously as one problem and not separately, giving a smaller search space and high-quality solutions. The proposed methodology produces a different schedule for different values of the (i) number of the levels of data cache; (ii) data cache sizes; (iii) data cache associativities; (iv) data cache and main memory latencies; (v) data array layout of the matrix and (vi) number of cores. [ABSTRACT FROM AUTHOR] – Name: AbstractSuppliedCopyright Label: Group: Ab Data: <i>Copyright of Journal of Supercomputing is the property of Springer Nature and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.) |
| PLink | https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=egs&AN=103643988 |
| RecordInfo | BibRecord: BibEntity: Identifiers: – Type: doi Value: 10.1007/s11227-015-1409-9 Languages: – Code: eng Text: English PhysicalDescription: Pagination: PageCount: 24 StartPage: 2644 Subjects: – SubjectFull: Complex multiplication Type: general – SubjectFull: SIMD (Computer architecture) Type: general – SubjectFull: Multicore processors Type: general – SubjectFull: Cache memory Type: general – SubjectFull: Computer memory management Type: general Titles: – TitleFull: A methodology for speeding up matrix vector multiplication for single/multi-core architectures. Type: main BibRelationships: HasContributorRelationships: – PersonEntity: Name: NameFull: Kelefouras, Vasilios – PersonEntity: Name: NameFull: Kritikakou, Angeliki – PersonEntity: Name: NameFull: Papadima, Elissavet – PersonEntity: Name: NameFull: Goutis, Costas IsPartOfRelationships: – BibEntity: Dates: – D: 01 M: 07 Text: Jul2015 Type: published Y: 2015 Identifiers: – Type: issn-print Value: 09208542 Numbering: – Type: volume Value: 71 – Type: issue Value: 7 Titles: – TitleFull: Journal of Supercomputing Type: main |
| ResultId | 1 |