Perfect Hashing Schemes for Mining Association Rules.

Saved in:
Bibliographic Details
Title: Perfect Hashing Schemes for Mining Association Rules.
Authors: Chin-Chen Chang1 ccc@cs.ccu.edu.tw, Chih-Yang Lin1
Source: Computer Journal. Mar2005, Vol. 48 Issue 2, p168-179. 12p. 1 Chart, 1 Graph.
Subjects: Hashing, Data mining -- Social aspects, Databases, Database searching, Encoding, Search engines, Information resources management
Abstract: Hashing schemes are widely used to improve the performance of data mining association rules, as in the DHP algorithm that utilizes the hash table in identifying the validity of candidate itemsets according to the number of the table's bucket accesses. However, since the hash table used in DHP is plagued by the collision problem, the process of generating large itemsets at each level requires two database scans, which leads to poor performance. In this paper we propose perfect hashing schemes to avoid collisions in the hash table. The main idea is to employ a refined encoding scheme, which transforms large itemsets into large 2-itemsets and thereby makes the application of perfect hashing feasible. Our experimental results demonstrate that the new method is also efficient (about three times faster than DHP), and scalable when the database size increases. We also propose another variant of the perfect hash scheme with reduced memory requirements. The properties and performances of several perfect hashing schemes are also investigated and compared. [ABSTRACT FROM PUBLISHER]
Copyright of Computer Journal is the property of Oxford University Press / USA and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Engineering Source
FullText Links:
  – Type: pdflink
Text:
  Availability: 0
Header DbId: egs
DbLabel: Engineering Source
An: 44442395
AccessLevel: 6
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: Perfect Hashing Schemes for Mining Association Rules.
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Chin-Chen+Chang%22">Chin-Chen Chang</searchLink><relatesTo>1</relatesTo><i> ccc@cs.ccu.edu.tw</i><br /><searchLink fieldCode="AR" term="%22Chih-Yang+Lin%22">Chih-Yang Lin</searchLink><relatesTo>1</relatesTo>
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="JN" term="%22Computer+Journal%22">Computer Journal</searchLink>. Mar2005, Vol. 48 Issue 2, p168-179. 12p. 1 Chart, 1 Graph.
– Name: Subject
  Label: Subjects
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Hashing%22">Hashing</searchLink><br /><searchLink fieldCode="DE" term="%22Data+mining+--+Social+aspects%22">Data mining -- Social aspects</searchLink><br /><searchLink fieldCode="DE" term="%22Databases%22">Databases</searchLink><br /><searchLink fieldCode="DE" term="%22Database+searching%22">Database searching</searchLink><br /><searchLink fieldCode="DE" term="%22Encoding%22">Encoding</searchLink><br /><searchLink fieldCode="DE" term="%22Search+engines%22">Search engines</searchLink><br /><searchLink fieldCode="DE" term="%22Information+resources+management%22">Information resources management</searchLink>
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: Hashing schemes are widely used to improve the performance of data mining association rules, as in the DHP algorithm that utilizes the hash table in identifying the validity of candidate itemsets according to the number of the table's bucket accesses. However, since the hash table used in DHP is plagued by the collision problem, the process of generating large itemsets at each level requires two database scans, which leads to poor performance. In this paper we propose perfect hashing schemes to avoid collisions in the hash table. The main idea is to employ a refined encoding scheme, which transforms large itemsets into large 2-itemsets and thereby makes the application of perfect hashing feasible. Our experimental results demonstrate that the new method is also efficient (about three times faster than DHP), and scalable when the database size increases. We also propose another variant of the perfect hash scheme with reduced memory requirements. The properties and performances of several perfect hashing schemes are also investigated and compared. [ABSTRACT FROM PUBLISHER]
– Name: AbstractSuppliedCopyright
  Label:
  Group: Ab
  Data: <i>Copyright of Computer Journal is the property of Oxford University Press / USA and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.)
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=egs&AN=44442395
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1093/comjnl/bxh074
    Languages:
      – Code: eng
        Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 12
        StartPage: 168
    Subjects:
      – SubjectFull: Hashing
        Type: general
      – SubjectFull: Data mining -- Social aspects
        Type: general
      – SubjectFull: Databases
        Type: general
      – SubjectFull: Database searching
        Type: general
      – SubjectFull: Encoding
        Type: general
      – SubjectFull: Search engines
        Type: general
      – SubjectFull: Information resources management
        Type: general
    Titles:
      – TitleFull: Perfect Hashing Schemes for Mining Association Rules.
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Chin-Chen Chang
      – PersonEntity:
          Name:
            NameFull: Chih-Yang Lin
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 03
              Text: Mar2005
              Type: published
              Y: 2005
          Identifiers:
            – Type: issn-print
              Value: 00104620
          Numbering:
            – Type: volume
              Value: 48
            – Type: issue
              Value: 2
          Titles:
            – TitleFull: Computer Journal
              Type: main
ResultId 1