Perfect Hashing Schemes for Mining Association Rules.
Saved in:
| Title: | Perfect Hashing Schemes for Mining Association Rules. |
|---|---|
| Authors: | Chin-Chen Chang1 ccc@cs.ccu.edu.tw, Chih-Yang Lin1 |
| Source: | Computer Journal. Mar2005, Vol. 48 Issue 2, p168-179. 12p. 1 Chart, 1 Graph. |
| Subjects: | Hashing, Data mining -- Social aspects, Databases, Database searching, Encoding, Search engines, Information resources management |
| Abstract: | Hashing schemes are widely used to improve the performance of data mining association rules, as in the DHP algorithm that utilizes the hash table in identifying the validity of candidate itemsets according to the number of the table's bucket accesses. However, since the hash table used in DHP is plagued by the collision problem, the process of generating large itemsets at each level requires two database scans, which leads to poor performance. In this paper we propose perfect hashing schemes to avoid collisions in the hash table. The main idea is to employ a refined encoding scheme, which transforms large itemsets into large 2-itemsets and thereby makes the application of perfect hashing feasible. Our experimental results demonstrate that the new method is also efficient (about three times faster than DHP), and scalable when the database size increases. We also propose another variant of the perfect hash scheme with reduced memory requirements. The properties and performances of several perfect hashing schemes are also investigated and compared. [ABSTRACT FROM PUBLISHER] |
| Copyright of Computer Journal is the property of Oxford University Press / USA and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.) | |
| Database: | Engineering Source |
| FullText | Links: – Type: pdflink Text: Availability: 0 |
|---|---|
| Header | DbId: egs DbLabel: Engineering Source An: 44442395 AccessLevel: 6 PubType: Academic Journal PubTypeId: academicJournal PreciseRelevancyScore: 0 |
| IllustrationInfo | |
| Items | – Name: Title Label: Title Group: Ti Data: Perfect Hashing Schemes for Mining Association Rules. – Name: Author Label: Authors Group: Au Data: <searchLink fieldCode="AR" term="%22Chin-Chen+Chang%22">Chin-Chen Chang</searchLink><relatesTo>1</relatesTo><i> ccc@cs.ccu.edu.tw</i><br /><searchLink fieldCode="AR" term="%22Chih-Yang+Lin%22">Chih-Yang Lin</searchLink><relatesTo>1</relatesTo> – Name: TitleSource Label: Source Group: Src Data: <searchLink fieldCode="JN" term="%22Computer+Journal%22">Computer Journal</searchLink>. Mar2005, Vol. 48 Issue 2, p168-179. 12p. 1 Chart, 1 Graph. – Name: Subject Label: Subjects Group: Su Data: <searchLink fieldCode="DE" term="%22Hashing%22">Hashing</searchLink><br /><searchLink fieldCode="DE" term="%22Data+mining+--+Social+aspects%22">Data mining -- Social aspects</searchLink><br /><searchLink fieldCode="DE" term="%22Databases%22">Databases</searchLink><br /><searchLink fieldCode="DE" term="%22Database+searching%22">Database searching</searchLink><br /><searchLink fieldCode="DE" term="%22Encoding%22">Encoding</searchLink><br /><searchLink fieldCode="DE" term="%22Search+engines%22">Search engines</searchLink><br /><searchLink fieldCode="DE" term="%22Information+resources+management%22">Information resources management</searchLink> – Name: Abstract Label: Abstract Group: Ab Data: Hashing schemes are widely used to improve the performance of data mining association rules, as in the DHP algorithm that utilizes the hash table in identifying the validity of candidate itemsets according to the number of the table's bucket accesses. However, since the hash table used in DHP is plagued by the collision problem, the process of generating large itemsets at each level requires two database scans, which leads to poor performance. In this paper we propose perfect hashing schemes to avoid collisions in the hash table. The main idea is to employ a refined encoding scheme, which transforms large itemsets into large 2-itemsets and thereby makes the application of perfect hashing feasible. Our experimental results demonstrate that the new method is also efficient (about three times faster than DHP), and scalable when the database size increases. We also propose another variant of the perfect hash scheme with reduced memory requirements. The properties and performances of several perfect hashing schemes are also investigated and compared. [ABSTRACT FROM PUBLISHER] – Name: AbstractSuppliedCopyright Label: Group: Ab Data: <i>Copyright of Computer Journal is the property of Oxford University Press / USA and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.) |
| PLink | https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=egs&AN=44442395 |
| RecordInfo | BibRecord: BibEntity: Identifiers: – Type: doi Value: 10.1093/comjnl/bxh074 Languages: – Code: eng Text: English PhysicalDescription: Pagination: PageCount: 12 StartPage: 168 Subjects: – SubjectFull: Hashing Type: general – SubjectFull: Data mining -- Social aspects Type: general – SubjectFull: Databases Type: general – SubjectFull: Database searching Type: general – SubjectFull: Encoding Type: general – SubjectFull: Search engines Type: general – SubjectFull: Information resources management Type: general Titles: – TitleFull: Perfect Hashing Schemes for Mining Association Rules. Type: main BibRelationships: HasContributorRelationships: – PersonEntity: Name: NameFull: Chin-Chen Chang – PersonEntity: Name: NameFull: Chih-Yang Lin IsPartOfRelationships: – BibEntity: Dates: – D: 01 M: 03 Text: Mar2005 Type: published Y: 2005 Identifiers: – Type: issn-print Value: 00104620 Numbering: – Type: volume Value: 48 – Type: issue Value: 2 Titles: – TitleFull: Computer Journal Type: main |
| ResultId | 1 |